Free the models: Harness design at the frontier

Daniel Furman, Jacky Zhao, Vaibhav Kumar, Ed Sioufi, Michele Catasta — Sep 29, 2026 — Replit

Model routers are everywhere right now, but they have a fundamental limitation. No matter if based on advanced heuristics or a small model that reads each turn and picks which LLM to use, a router will always be less capable than the model it's choosing for. Replit Agent lets the model decide instead.

The main agent, or core loop, chooses its subagents' tier and effort, and adjusts its own as the task unfolds. Given that freedom, GPT-6 Astra hands routine implementation to less costly subagents and decides for itself where its tokens are worth spending. On both DeepSWE and Terminal-Bench, Replit Agent is Pareto-efficient against Astra on its own: no published Astra baseline costs less and scores higher. It also beats a sidekick architecture, the same setup with one long-lived worker, by 11 and 16 points.