Replit's AI team argues model routers are always weaker than the models they choose for. Their harness lets GPT-6 Astra decide effort and delegation; on DeepSWE and Terminal-Bench, Replit Agent is Pareto-efficient versus Astra alone and a sidekick architecture.
Free the models: Harness design at the frontier
Daniel Furman, Jacky Zhao, Vaibhav Kumar, Ed Sioufi, Michele Catasta — Sep 29, 2026 — Replit
Model routers are everywhere right now, but they have a fundamental limitation. No matter if based on advanced heuristics or a small model that reads each turn and picks which LLM to use, a router will always be less capable than the model it's choosing for. Replit Agent lets the model decide instead.
The main agent, or core loop, chooses its subagents' tier and effort, and adjusts its own as the task unfolds. Given that freedom, GPT-6 Astra hands routine implementation to less costly subagents and decides for itself where its tokens are worth spending. On both DeepSWE and Terminal-Bench, Replit Agent is Pareto-efficient against Astra on its own: no published Astra baseline costs less and scores higher. It also beats a sidekick architecture, the same setup with one long-lived worker, by 11 and 16 points.
Why we scaffold less
Every model release invalidates assumptions baked into the harness. As models become stronger at long-horizon tasks, they don't need as much scaffolding. We've observed them lean toward delegation on their own: using subagents for context management and parallelism.
But the frontier is jagged. The strongest coding model is not necessarily the strongest at designing UIs or making slides. So we design our harness to let each model work its own way, with the guardrails it still needs and quality at minimum cost as the goal.
Freeing the model means letting it decide:
How hard to think — effort set step by step; on latest models, mid-turn without a cache miss
When to hand work off — hand-offs buy context management and parallelism; small tasks spawn nothing
Who to hand it to — specialists on the model strongest at each job
Astra is the first model they've seen routinely delegate to general workers without being told to, and once briefed it tends to return rather than start over.
Results
DeepSWE v1.1: Replit Agent 72% at $2.11/task. Astra mini-swe-agent: 67% at $1.60 (low) / 74% at $4.43 (xhigh). Sidekick: 61% at $1.34.
Terminal-Bench 4.0: Replit Agent 49% at $2.53/task. Astra: 42% at $2.25 / 60% at $5.86. Sidekick: 33% at $1.84.
Replit Agent beats the sidekick by 11 and 16 points. Astra alone scores higher only by spending more than twice as much. Neither baseline wins on both cost and score. Runs used Replit Agent exactly as it ships.
The bitter lesson of harness design
Baking human knowledge into an agent helps short-term and plateaus; general methods that scale with computation win. A rigid harness forces one way of working; a composable one lets the model choose. The smarter models get, the less the harness should decide for them. Free the models.