{"article":{"slug":"scaling-discovery-through-test-time-communication","title":"Scaling Discovery through Test-Time Communication","subtitle":null,"summary":"Research paper showing that test-time communication among identical agents sharing discoveries can beat independent parallel search on ARC-AGI-3 and transfer to research tasks like polyomino packing and MNIST compression.","content_type":"research","language":"en","canonical_url":"https://arxiv.org/abs/2609.21032","author":{"name":"Jongho Park et al.","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"arXiv","url":"https://arxiv.org/","listing_slug":null,"listing":null},"topics":[{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"Benchmarks","slug":"benchmarks","url":"https://listedarticles.com/topics/benchmarks"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":12394,"reading_minutes":54,"published_at":"2026-09-17T12:00:00.000Z","added_at":"2026-09-22T18:27:13.700Z","updated_at":"2026-09-22T18:27:13.700Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/scaling-discovery-through-test-time-communication","markdown_url":"https://listedarticles.com/articles/scaling-discovery-through-test-time-communication.md","example":false,"citation":"Jongho Park et al., arXiv. \"Scaling Discovery through Test-Time Communication.\" 17 Sept 2026. https://arxiv.org/abs/2609.21032 (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://arxiv.org/abs/2609.21032"},"body_markdown":"# Scaling Discovery through Test-Time Communication\n\nJongho Park<sup>b</sup><sup>1</sup><sup>1</sup>\n        1\n        \n        \n        \n      This work was done during an internship at Microsoft Research.  Vasilis Kontonis<sup>m</sup>  Shivam Garg<sup>m</sup>\n\nAkshay Krishnamurthy<sup>m</sup>  Dimitris Papailiopoulos<sup>m</sup>\n\n<sup>b</sup>UC Berkeley   <sup>m</sup>Microsoft Research\n\n<sup>0</sup>\n\n<sup>0</sup>footnotetext: Emails: jjhpark@berkeley.edu, {vkontonis, shigarg, akshay.krishnamurthy, dimitriosp}@microsoft.com\n\nScience advances not in isolation but through collaboration,\nyet existing agentic systems capture little of this.\nWhether communicating agents help remains an open question\nwith mixed prior results.\nWe show that *test-time communication can substantially outperform\nindependent parallel attempts on challenging tasks, where sharing a\nbreakthrough can push the whole group forward*.\nWe first study the effect of scaling multi-agent test-time communication,\nwhere agents have no predefined roles and communicate via a shared directory,\non ARC-AGI-3, a benchmark requiring novel problem solving.\nWe find that a team of  communicating agents, team@, matches the success rate of\n independent agents, and this advantage grows with\n, suggesting gains compound with scale.\nThe effect is not merely efficiency: a task that no single agent can solve, a\nteam of agents can solve reliably. Furthermore, these gains\ntransfer to research-oriented tasks, given sufficient compute. On\npolyomino packing, communicating agents outperform best@\nand exceed the prior best-known score. On MNIST classifier compression, communication\nsurpasses the best-known human solution. A team of four agents produced a\n1,957-byte classifier submission achieving 99.4% test accuracy,\nsmaller than both the best-known human solution and the best single-agent\nresult. These gains are not unconditional.\nIndependent agents may outperform communication when\ncompute is limited or when a clear measure of progress is absent.\nHowever, under sufficient compute and clear feedback, multi-agent communication consistently yields stronger results.\n\n## 1 Introduction\n\nFor test-time compute, it is standard to sample  independent outputs in parallel from a large language model (LLM) and aggregate them, *e.g.*, by taking the majority answer or\nthe best@ when a verifier is available (Wang et al., 2023; Brown et al., 2024; Li et al., 2024; Huang et al., 2025; Snell et al., 2025).\nIndependent attempts, however, leave discoveries made during one run unavailable to\nguide another while its search is still unfolding.\nThis opportunity is especially crucial for LLM agents, which operate a terminal, execute\ncode, use tools, and observe results over horizons of hours or\ndays (Yang et al., 2024; Chan et al., 2025; Merrill et al., 2026).\nTheir trajectories contain intermediate results, failed experiments, and reusable artifacts\nthat could help other agents avoid dead ends or build on promising approaches.\n\nCommunication between agents would be the natural remedy, but the evidence for its benefits is mixed. While keeping the best of independent candidates improves with , having agents exchange opinions may not (Choi et al., 2025). Multi-agent debate often fails to outperform chain-of-thought prompting despite using substantially more inference compute (Smit et al., 2024; Zhang et al., 2025d). Agents can defer to an incorrect majority or dilute expertise by compromising between expert and non-expert judgments (Wu et al., 2025; Pappu et al., 2026). Most recently, Anthropic (2026b) report that coordinating agents find more vulnerabilities than independent agents, but use more tokens and search more broadly. Within the common search scope, tokens per vulnerability are comparable, leaving the efficiency gains from communication unclear. On the other hand, multi-agent systems clearly help when a task can be decomposed into subtasks and solved in parallel (Kim et al., 2025), but such gains are limited to cleanly decomposable tasks. Can communication offer more than just parallelism?\n\nScience suggests that it can. Scientists coordinate by adjusting their research to the results achieved by others (Polanyi, 1962), while accumulated empirical evidence helps distinguish competing explanations and turn varied efforts into collective progress (Strevens, 2020). This process is iterative rather than merely parallel. Collaboration should therefore do more than divide and conquer. One agent’s observation redirects another’s search, and partial discoveries combine into a result that no individual reaches alone. We ask whether agent communication can produce the same effect.\n\nTo this end, we study a *minimal form of test-time communication*\nin which identical agents receive the same open-ended objective and communicate\nthrough a shared workspace, without predefined roles or a central orchestrator.\nAutonomously and asynchronously, the agents can\nexchange intermediate results, failures, and artifacts while their search is still\nunfolding. Our central comparison pits a team of  communicating agents, team@, against\nthe best result from  independent agents, best@, under the same per-agent resources.\nThis isolates the value of communication from the gains produced by additional parallel\nattempts.\n\nWe first study ARC-AGI-3 (ARC Prize Foundation, 2026) using Claude Sonnet 4.6. In terms of full-game success rate, team@3 matches best@13 and team@5 matches best@33, making them as effective as and times as many independent agents (Figure 1). More strikingly, communicating teams solve games that independent agents rarely or never solve. These benefits transfer to longer-horizon research problems. On Frontier-CS polyomino packing (Mang et al., 2026), where agents pack polyominoes into a minimum-area rectangle, team@4 scores 0.922 with Claude Opus 4.6 and 0.910 with Sonnet 4.6, compared with 0.893 and 0.891, respectively, for the strongest solo run of each model and 0.894 for the prior best-known result. On MNIST classifier compression under an accuracy constraint, we evaluate team@4 with GPT-5.6 Sol over 96 hours. The best team produces a state-of-the-art 1,957-byte classifier at 99.4% accuracy, beating the 2,461-byte best-known human solution and single-agent baselines.\n\nThese results suggest that test-time communication is most useful when agents have\nenough compute to build on one another’s discoveries and can objectively assess\nwhether those discoveries improve on earlier results.\nSufficient compute allows teams to overcome initial coordination costs, while an\naccessible verifier, such as level success or a solution scorer, helps agents\ndecide which discoveries to adopt.\nWe call this mechanism *verified progress sharing*.\nBy building on successive breakthroughs from different members, a communicating team can\nturn parallel exploration into cumulative progress.\n\nWhen agents cannot verify intermediate progress, however, communication may offer less benefit. On Terminal-Bench 2.0, where available feedback may not reliably rank intermediate solutions, we find that team@2 improves over a single attempt but does not outperform independent pass@2. This distinction also offers an explanation for the prior negative results, where agents lack a verifier to distinguish better solutions from worse ones and therefore abandon stronger solutions in favor of weaker ones (Choi et al., 2025; Pappu et al., 2026). Together, our findings show that communication can offer more than divide and conquer parallelism and help identify the conditions under which test-time communication excels.\n\nOur contributions are as follows:\n\n1. 1. \nCommunicating agents show stronger scaling than independent agents. On the ARC-AGI-3 suite of games, matching the solve rate of team@5 requires 33 independent agents. The multiplier grows with team size, from for team@3 to for team@5, suggesting that the gains from communication compound with scale. In fact, the solve rate of team@ increases more rapidly with than that of best@, as seen in Figure 1.\n2. 2. \nTest-time communication unlocks tasks that single agents cannot. Team@3 lifts the solve rate on the game FT09 on ARC-AGI-3 from 9.4% for a single agent to 90%. More strikingly, team@5 solves the game LP85, which remains unsolved across 64 single-agent trials, 65% of the time. Team@5 also improves the average furthest level reached over best@5 even on games that remain unsolved.\n3. 3. \nCommunication achieves state-of-the-art results on algorithmic and ML optimization tasks. Communication sustains progress over multi-day horizons. On Frontier-CS polyomino packing, a team of agents outperforms independent agents and surpasses the prior best-known score. On ML compression, it produces a MNIST classifier substantially smaller than those found by independent agents or the best-known human solution. Agents reach these solutions by refining one another’s discoveries and combining complementary improvements, including ideas from peers’ failed approaches.\n4. 4. \nCommunication benefits from verifiable progress and sufficient compute. At low budgets, team@ incurs a coordination tax that outweighs its benefits. We identify *verified progress sharing* as a mechanism through which agents build on\nsuccessive discoveries to overcome the coordination tax given enough compute.\nOn Terminal-Bench 2.0, where available feedback may not reliably rank intermediate solutions,\nteam@2 improves over a single attempt but does not outperform pass@2.\nThis helps reconcile prior results by identifying compute and verification\nas conditions that shape communication benefits.\n\n## 2 Agentic Communication at Test-time\n\n| Benchmark | ARC-AGI-3 | Polyomino packing | MNIST Compression | \n|---|---|---|---|\n| Model | Sonnet 4.6 | Sonnet, Opus 4.6 | GPT-5.6 Sol | \n| Budget per agent | per-level action budget | 3-72 hours | 96 hours | \n| Team size | 3, 5 | 3, 4 | 4 | \n| Metric | game solve rate | packing score | qualified bytes | \n| Prior best-known result | – | 0.894 | 2,461 B | \n| Independent best@ | 1.4% (=3), 2.2% (=5) | 0.893 | 3,160 B | \n| Communicating team@ | 4.6% (=3), 8.0% (=5) | 0.945 | 1,957 B | \n\nWe study a minimal form of test-time communication instantiated through a shared workspace among identical agents. In each team@ trial, the harness launches CLI agents concurrently in the same task container. They use the same model, tools, task instruction, and communication prompt. Each agent has a separate model context and designated scratch directory, while all agents share the task filesystem and communication artifacts.\n\nAgents communicate directly through an append-only communication log, which acts as an asynchronous broadcast channel. Additional shared records contain adopted approaches, disconfirming evidence, and a score log of the approaches so far. Without assigned roles or a central orchestrator, the only protocol-level allocation, enforced by a synchronization primitive, is slot ownership. To ensure that agents claim distinct approaches without collision, agents race to create numbered slot directories using an atomic filesystem operation.\n\nThe communication prompt (Appendix A.2) specifies how agents use these mechanisms and discourages premature convergence. After claiming a slot, each agent is asked to declare a distinct approach by considering the already-claimed slots. During execution, agents publish concise findings with timestamps and reproducible evidence whenever they make notable progress, allowing peers to reproduce or build upon their results. Results placed on the leaderboard should include measured outcomes, reproduction instructions, approach lineage, or known counterevidence. An agent is to adopt a peer’s approach only after observing a clearly better result, and even after adoption, it should preserve one meaningful variation.\n\nThe protocol combines diverse exploration with evidence-based adoption to discourage convergence on poor solutions. Atomic slot ownership, together with the requirement to pursue distinct approaches, helps preserve diversity. Verifier scores give agents a basis for filtering out poorly performing approaches when deciding what to adopt. This contrasts with the task settings of Pappu et al. (2026), where agents lack verifier scores for agent solutions and can dilute expertise by compromising between expert and non-expert judgments.\n\n### 2.1 Experimental Setup\n\nWe study multi-agent communication on the following three tasks, each covering a different aspect of open-ended research tasks.\n\n##### ARC-AGI-3\n\n(ARC Prize Foundation, 2026) tests novel problem solving and interactive discovery by asking agents to solve unfamiliar grid-world games without instructions, inferring the rules, the goal, and the effect of each control from play alone. Each game comprises levels, and progress depends on carrying forward what was learned in earlier ones. This leveled structure makes breakthroughs cleanly measurable. Real-world research problems may not offer such a clean signal. A genuine conceptual advance in a domain such as approximation factors for NP-hard problems may move the reported number by only a small constant, an improvement easily lost in run-to-run variance.\n\nWe evaluate on all 25 public games with Claude Sonnet 4.6 under the benchmark’s native per-level action budget and measure solve rate as the proportion of trials that clear all levels successfully. Though more recent models such as GPT-6 Astra now succeed easily on the benchmark, we focus on one model, especially one that does not saturate, to isolate the effect of test-time communication. For teams, an agent that exhausts its action budget is terminated, while the remaining agents continue. We run 64 single-agent trials per game and label a game unsolved by single agents if no trial clears all levels. For teams of three and five agents, we run 20 trials for every combination of game and team size.\n\n##### Frontier-CS polyomino packing\n\n(Mang et al., 2026) tests algorithmic optimization on an NP-hard problem. Agents write and repeatedly improve a C++17 program that packs reflected and rotated polyominoes into a minimum-area rectangle. The scorer evaluates the submitted solution on 70 hidden test cases and returns continuous partial credit in . The packing score is the mean reward across the 70 cases, and the highest valid submission in a run is retained. The best published score is 0.894, obtained by Qu et al. (2026) using four Claude Opus 4.6 agents. We test on both Claude Sonnet 4.6 and Opus 4.6 with teams of three or four agents.\n\n##### MNIST Classifier Compression\n\ntests empirical ML research. Agents must train and compress a self-contained MNIST classifier (Le Cun et al., 1998). An artifact qualifies only at 99.4% test accuracy or better and agents cannot inspect the test images, labels, or individual errors, as described in Appendix A.5. Among qualifying artifacts, only compressed size counts. We measure size using a deterministic gzip-9 compression of the submission, which includes inference code and model weights (they may be separate files or one file, depending on the agent’s design). To the best of our knowledge, the best prior result comes from the open-source model introduced by Dhairyashil R. G. (2024), which we reproduce as a 2,461-byte submission under our artifact format. We test a team of four GPT-5.6 Sol agents over a 96-hour period. Further implementation details can be found in Appendix A.\n\n##### Runtime Environment.\n\nEvery agent operates through GitHub Copilot CLI (GitHub, 2026), with each trial isolated as a Harbor task (Harbor Framework Team, 2026). Single-agent trials contain one CLI agent. Team@ trials contain agents in the same task container. Agents receive the same task, tools, and base filesystem, but have private scratch directories. Within a team, agents communicate explicitly. They append timestamped claims, evidence, failures, and adoption events to shared logs, publish measured candidates to a board, and use file locks to serialize changes to a shared graded artifact. Distinct trials share neither files nor messages. Hidden evaluation data are kept outside the agent container. Polyomino submissions are sent to a separate scorer service, while MNIST submissions are evaluated by a host-owned oracle in a fresh network-disabled container. These services return evaluation feedback but never expose the underlying cases, labels, or retained program outputs.\n\n##### Metrics.\n\nOur central comparison is the outcome of communicating agents (team@) versus the best outcome among independent agents (best@). We measure both metrics against wall-clock time and against total output tokens to account for compute. For ARC-AGI-3, a level counts as solved if any of the agents clears it, and a game counts as solved if any of the agents clears all levels. For a finite ARC solo pool with successes among trials, we compute best@ exactly, without replacement, as . We apply the same calculation level by level and average over games, giving each game equal weight. Throughout the paper, solve rate refers to the final solve rate averaged over all 25 games, and per-level rates are labeled explicitly. We measure team@ and best@ as the best-so-far packing score (mean packed-cell density across 70 cases) or MNIST classifier compression size (gzipped submission size including code and weights).\n\nWithin every comparison, communicating and independent agents use the same model, maximum reasoning effort, longest context-length setting, and per-agent resource allocation. Appendix A.1 provides the communication protocol, prompts, and Copilot CLI versions.\n\n## 3 Experimental Results\n\n### 3.1 ARC-AGI-3\n\n| Levels from target | team@3 | best@3 | best@13 | team@5 | best@5 | best@33 | \n| 5 | 52.8 | 45.4 | 63.9 | 64.1 | 52.8 | 69.5 | \n| 4 | 28.2 | 25.8 | 46.5 | 33.6 | 32.9 | 55.4 | \n| 3 | 16.6 | 9.7 | 20.8 | 22.9 | 12.8 | 31.3 | \n| 2 | 7.2 | 5.0 | 9.3 | 16.2 | 6.1 | 15.8 | \n| 1 | 4.8 | 2.6 | 6.3 | 10.0 | 3.7 | 10.2 | \n| (Solved) 0 | 4.6 | 1.4 | 4.7 | 8.0 | 2.2 | 8.1 | \n\nDespite ARC-AGI-3’s visual simplicity, the most capable LLM-based agents fail on most games without a specialized harness. For this benchmark, we use Claude Sonnet 4.6 agents through Copilot, which reach a final success rate below 1%.\n\nTable 2 reports the percentage of runs that reach success or come within levels of it. Across 25 games, team@ outperforms best@ at every level for both and , and the advantage widens with depth, from to at . The leftmost panel of Figure 1 shows a further encouraging pattern in which team@ rises monotonically with team size and faster than best@. Matching a communicating team takes 4.3–6.6 as many independent agents, a multiplier that itself grows with , though returns may diminish beyond five agents.\n\n(a) Communication effect in ARC-AGI-3\n\n(b) Final solve rate\n\n| Game | best@5 | team@5 | \n|---|---|---|\n| LP85 | 0.0 | 65.0 | \n| AR25 | 7.8 | 20.0 | \n| DC22 | 0.0 | 0.0 | \n| SC25 | 0.0 | 0.0 | \n| KA59 | 0.0 | 0.0 | \n| FT09 | 39.9 | 100.0 | \n| WA30 | 0.0 | 0.0 | \n| TR87 | 0.0 | 0.0 | \n| M0R0 | 0.0 | 0.0 | \n| TN36 | 0.0 | 0.0 | \n| RE86 | 0.0 | 0.0 | \n| SB26 | 7.8 | 15.0 | \n| TU93 | 0.0 | 0.0 | \n| S5I5 | 0.0 | 0.0 | \n| LF52 | 0.0 | 0.0 | \n| G50T | 0.0 | 0.0 | \n| CD82 | 0.0 | 0.0 | \n| SK48 | 0.0 | 0.0 | \n| SU15 | 0.0 | 0.0 | \n| BP35 | 0.0 | 0.0 | \n| LS20 | 0.0 | 0.0 | \n| VC33 | 0.0 | 0.0 | \n| SP80 | 0.0 | 0.0 | \n| CN04 | 0.0 | 0.0 | \n| R11L | 0.0 | 0.0 | \n\nFigure 2 breaks these aggregate results down by game. Across the four games solved at least once by team@5, the average solve rate rises from to . The improvement is especially striking on LP85, where 64 single-agent trials yield no successes, yet team@5 achieves a solve rate. Nor is the effect limited to the largest team. On FT09, the solve rate climbs from with best@3 to with team@3. Although these gains are concentrated, this reflects the limits of the base model rather than of communication, since most of the 25 games remain beyond Sonnet 4.6 under either setting.\n\nMoreover, zero-percent solve rates obscure meaningful gains in progress. Figure 2(a) reports the average furthest level reached. Under team@5, several never-solved games still advance by roughly a full level, showing that test-time communication produces real progress across most games. The largest gain is levels, while no loss exceeds levels. When coordination fails, the cost is negligible. When it succeeds, the gains can be substantial.\n\nIs test-time communication simply a matter of spending more tokens than independent agents do? Figure 3 paints a more nuanced picture. By the end of evaluation, team@ spends roughly 8 million output tokens with three agents and 24 million with five, nearly twice what best@ spends. However, to solve more than 1% of ARC-AGI-3 tasks, team@ reaches any given accuracy with fewer total tokens, with best@3 spending as many as team@3 and best@5 spending as many as team@5. Matching the solve rate of test-time communication is even costlier, as pools of 13 and 33 independent agents must spend as much as 3.8–4.9 the team’s tokens.\n\nWe further study SB26 and LP85, two games that benefit from communication, with two action-budget ablations. We first give a single agent the native action budget, matching the total budget of standard team@5 or best@5. This configuration tests whether a longer horizon of actions of a single agent can match the performance of a team of communicating or independent agents. We then reduce team@5’s per-agent action budget to , so its total action budget matches that of a single agent.\n\nAt matched total budget, communication wins. On SB26, best@5 leads early at to , but team@5 overtakes at level 2 and solves the final level in of trials against roughly for best@5 and zero for the long-horizon single agent. LP85 separates further, where team@5 holds through level 5 and finishes at while best@5 falls from to zero by level 4 and the single agent peaks at . A longer horizon alone does not substitute for coordination. This advantage, however, vanishes once the budget is reduced. At per agent, team@5 drops to zero after level 1 on LP85 and after level 3 on SB26, in both cases losing to a single agent spending the same total. Communication pays only when each agent has enough budget to explore on its own.\n\n##### Relative human action efficiency.\n\nThroughout this paper, we take solve rate as our primary metric, since our question is whether communication and larger teams, that is, more test-time compute, lead to more tasks solved in ARC-AGI-3. The official metric, by contrast, scores a run not only by what it clears but by how economically it plays, using a metric called relative human action efficiency (RHAE) (ARC Prize Foundation, 2026). We define RHAE below and then use it to show that test-time communication improves per-agent efficiency as well.\n\nAn action is a discrete interaction that changes the game state, so reasoning, tool calls, and read-only inspection are not charged. For level of game , let be the actions the agent spends and the human baseline, taken as the upper-median best action count over first-time human players. The level score is\n\n|  |  |  | (1) | \n\nwith for any level the agent never completes. Game scores weight level by and are capped by the weighted fraction of levels completed,\n\n|  |  |  | (2) | \n\nwhere of levels are completed, and the reported RHAE is the mean of over games.\n\nUsing RHAE, Figure 5 compares independent Sonnet 4.6 agents with communicating teams across all 25 games. Communication improves not only the strongest team member but also the average member. The best agent’s mean RHAE rises from for a single agent to with team@3 and with team@5, exceeding the corresponding best@3 and best@5 values of and . This gap shows that the gain is not merely the result of selecting the best outcome from more agents. The average communicating agent also improves with team size, while the average independent agent remains at (by definition). Strikingly, the average agent in team@5 matches the best of five independent agents, at versus . Together, the two panels show that communication improves both the average and the strongest team member, rather than improving team performance only by pooling more attempts.\n\n### 3.2 Polyomino Packing\n\nWe now turn to Frontier-CS (Mang et al., 2026) to test whether test-time communication yields stronger performance on challenging algorithmic tasks. We study its polyomino packing task, in which agents must pack a set of pieces into a rectangular box of minimal area. The task is NP-hard, so agents must discover good heuristics and design a packing algorithm rather than search exhaustively. The best known score of was obtained by Qu et al. (2026) using four Opus 4.6 agents.\n\nWe use Claude Sonnet 4.6 and Opus 4.6 agents under the same communication protocol as in ARC-AGI-3. We first evaluate Sonnet 4.6 under Frontier-CS’s three-hour limit, with 60 single-agent trials and 20 team@3 trials. We then extend the time limit to 72 hours to study long-horizon performance and run 12 single-agent trials and 2 team@4 trials. For each method, we report the trajectory of the run that achieves the highest score. Our focus is on whether a method can discover a solution beyond the existing frontier, where a single breakthrough matters even when other attempts are unsuccessful (Yuksekgonul et al., 2026).\n\n##### Official Frontier-CS results.\n\nFigure 6 shows that both runs reach roughly within the first hour, but the team breaks away at about 1.6 hours and continues improving as the single agent plateaus. The final gap, versus , more than halves the mean fraction of unused area, from to , and carries a team of Sonnet 4.6 agents past the prior best of . The team’s advantage grows as the single agent’s progress slows, resembling the larger gap between team and independent agents at deeper levels of ARC-AGI-3.\n\n##### Long-horizon results.\n\nFigures 7 and 8 show that the gap survives this extension. Independent agents do improve further, but their best scores settle at and , still below the prior best. The teams take lasting leads at roughly two and four hours, respectively, and continue to and . (The fact that team scores are lower than 0.945 is likely due to the small number of trials.) Thus, the previous short timeframe does not fully explain the single-agent plateau. These runs demonstrate that communication also outperforms independent agents even in long-horizon settings, as long as communicating agents have enough time and compute to explore and share their discoveries.\n\n#### 3.2.1 Qualitative Case Study on Communication\n\nIn this subsection, we examine the successful three-hour Sonnet run that set a new state-of-the-art score of to qualitatively understand and assess how good communication leads to breakthroughs. With the help of Opus 5, we analyze the run’s traces and identify the key discoveries that led to the final score. We follow the numbered discoveries in Figure 6, using the agent labels shown in the figure.\n\n##### From compact placement to filling gaps.\n\nAgent a1 begins with *shelf packing*, arranging pieces in rows, then\n*skyline packing*, placing them along the upper contour of the occupied region.\nAgent a2’s *bottom-left packing* searches for the lowest legal placement,\nbreaking ties toward the left. These methods successively lead the run at markers\n1–3 in Figure 6, but stall near .\nKeeping placements low still leaves enclosed gaps that later pieces cannot fill.\nAgent a1 identifies this limitation and proposes *contact maximization*,\nfavoring placements that touch more already-occupied cells. Contact serves as a local\nmeasure of fit, encouraging pieces to nest together rather than merely keeping the\ncurrent height small.\n\n##### Building upon another agent’s failed idea.\n\nThe proposal for contact maximization appears at marker 4, but\na1’s first implementation is slow and\nthe score barely changes for roughly half an hour. Agent a2\nthen rebuilds the method from a1’s\nmessage in the communication logs without reading its code, describing the rule\nas counting “how many adjacent cells are already occupied for each piece that fits”\nand choosing the highest-contact candidate. This independent implementation and continued\nefficiency improvements raise\nthe score to nearly  after marker 5.\nAgent a2 then evaluates every piece orientation within\nthe placement lookahead, comparing how different orientations fit against\nthe existing packing. Running this *all-orientation search* at full strength\nproduces the decisive jump past  at marker 6.\n\nThe difficulty is making that search affordable within the evaluator’s two-second limit per instance. Examining more orientations improves placement quality but risks a timeout, while conservative search budgets leave useful candidates unexplored. The original traces show other teams considering the same search but struggling with timeouts and keeping its budget too small. The strongest single agent also explores contact maximization without reaching the same quality in three hours. The successful transfer therefore includes substantial implementation work. Agent a2 takes a shared scoring idea and makes a more thorough search practical under the execution limit.\n\n##### Extending the shared objective.\n\nWith the run above , a3 adds a *boundary bonus* at marker 7.\nThe contact score now counts box walls as well as occupied cells, rewarding pieces\nthat fit snugly against edges and into corners. This extends the same objective\ndeveloped by a1 and implemented by a2, and subsequent refinements\nbring the run to . The sequence makes the dependence between contributions\nconcrete. One agent identifies a better placement criterion, another turns it into\nan effective search, and a third improves the criterion on top of that implementation.\n\nFigure 9 illustrates the resulting improvement on one 172-piece instance. The team packs it at density, compared with for an illustrative single-agent solution, leaving far fewer gaps between pieces. This single-agent run averages across the benchmark and is distinct from the strongest run in Figure 6.\n\n##### Related exchanges over longer horizons.\n\nIn the Opus run (Figure 7), a1 introduces\n*best-fit selection*, comparing remaining pieces and candidate positions before\ncommitting the best placement. Agent a2 improves it and reduces its cost by\nrestricting the candidate pool to the largest piece sizes. Agent a1 spends\nthe saved time searching more box widths, carrying the score past .\nA later contribution from a2 changes how height and contact interact.\nInstead of considering contact only after minimizing height, it scores their\nweighted sum, allowing a slightly taller placement if it gains enough contact.\nThe other agents incorporate the change within an hour. Here, one agent’s speed\nimprovement enables another’s broader search, and the final refinement relaxes a\nplacement rule that earlier methods treated as fixed.\n\nThe Sonnet run (Figure 8) again develops through agents that did not supply the initial solution. After a1’s skyline and a2’s bottom-left methods, a3 introduces best-fit selection to minimize the increase in box height. It then breaks equal-height ties by wasted space, producing the sharp jump near four hours, and adds tall-first sampling to pass . Agent a4 subsequently extends a2’s width heuristic. When a1 identifies unproductive parts of the deterministic width sweep, a4 reallocates that time to randomized restarts. As in the three-hour case, progress comes from successive changes to a shared algorithm, including improvements to where it spends its limited search time.\n\n### 3.3 MNIST Classifier Compression\n\nWe next move from abstract games and competitive programming style problems to a more realistic ML research engineering task. The main question here is whether communication benefits extend to realistic multi-day engineering and how well communicating agents perform relative to the best-known human solutions. In MNIST classifier compression, agents must produce the smallest possible classifier that achieves at least test accuracy by minimizing the compressed size of the complete submission that includes its inference code and model weights in a 96-hour timeframe. The task combines empirical model development with deployment engineering, since agents must preserve accuracy while compressing both the learned parameters and the code needed to use them.\n\nFigure 10 plots the smallest qualifying submission found so far against wall-clock time and total output tokens, with lower values indicating better compression. In the lower compute regime, the independent agents lead. They produce submissions below 15K bytes before the team, which remains near 75K bytes until roughly the first half hour. Team@4 catches the independent frontier after about one hour and 100K output tokens, then stays ahead as further improvements steadily reduce its submission size. Communication therefore incurs an initial communication cost, but converts later compute into more effective progress.\n\nThe difference becomes decisive over the full 96-hour run. The independent runs converge between roughly 3KB and 5KB, with best@4 finishing at 3,160 bytes, and none crosses the 2,461-byte best-known human result. Team@4 crosses that threshold after roughly 20 hours and two million output tokens, then continues to 1,957 bytes, which is about smaller than the best-known human submission. Because the advantage persists when progress is plotted against total output tokens, it cannot be explained by the team merely generating more tokens. Instead, test-time communication allows the four agents’ work to be more effective by diversifying their approaches, sharing ideas, and building on each other’s progress. While it is certainly possible that a stronger compression could be found given more time or trials, the result is a new state-of-the-art in MNIST classifier compression.\n\n#### 3.3.1 Qualitative Study on Communication\n\n##### Finding a model worth sharing.\n\nThe four agents begin with different bets. Agent a1 tries digit prototypes, a2 fixed gradient features, a3 spectral models and then binary networks, and a4 distilled CNNs. After its prototype classifier misses the accuracy threshold, a1 changes the compression strategy. A recurrent CNN reuses one convolutional cell nine times, buying depth without storing a fresh set of filters at every step. Its 5,329-byte qualifying submission persuades the other three agents to adopt this architecture within ten minutes. Their separate searches now improve different parts of a shared model.\n\n##### Giving lost features a way around the bottleneck.\n\na3 targets the final scoring layer, also called the classification head. Its matrix converts 128 features into ten digit scores using 1,280 learned parameters, plus ten bias parameters. It first compresses the 128 input features into eight learned weighted sums before computing the ten digit scores. This stalls at validation accuracy. To recover information lost in that compression, it also sends selected original features directly to the final scoring layer. Figure 11(b) shows these two paths in red and teal. With five original features selected using their covariance on the training data, the model reaches even before further fitting. The resulting 3,120-byte model qualifies, and a1 explicitly adopts the design. a4 subsequently combines seven learned weighted sums with seven directly selected features at 2,790 bytes. Passing original features to the scoring layer helps recover accuracy as the learned matrix shrinks.\n\n##### Learning that fewer parameters can cost more bytes.\n\nLater, a2 reduces the learned weighted sums to six on the improved backbone. Additional directly selected features recover accuracy, but the new parameters initially require more compressed bytes than those of the larger head they replace. a2 finds that their values are less repetitive. It trains the projection parameters using only the stored integers , multiplied by a learned scale factor for each weighted sum. With ten directly selected features, this produces a qualifying 1,983-byte submission. Reducing the number of parameters works together with restricting their stored values.\n\n##### Rescuing a branch that has fallen behind.\n\nMeanwhile, a1 makes the last two spatial stages share their channel offsets. Its 1,993-byte model qualifies, but the shared best has already reached 1,970 bytes. Rather than discard the branch, it combines this backbone with a2’s design with six learned weighted sums, using routines originally written by a1 and extended by a2. Variants that pass ten or eleven original features directly to the scoring layer fail the test threshold. Passing twelve succeeds while keeping the projection parameters restricted to the same five stored integer values, and a1 advances the combined model to 1,960 and then 1,959 bytes. Figure 11(a) shows the adoption statements connecting these two branches. a2 then reorders the directly selected features and their matching scoring parameters to reach 1,957 bytes without changing predictions. The final architecture emerges from joining two branches, including one that could no longer win on its own.\n\n#### 3.3.2 The 1,957-Byte MNIST Classifier\n\n##### Architecture.\n\nFigure 11(b) shows a recurrent convolutional network with 2,900 quantized parameters and 30 scale factors. The input is cyclically shifted upward by one pixel. A convolution maps the single input channel to 32 channels, followed by SiLU to form . Nine residual updates then reuse the same depthwise filters and pointwise filters . For update , the feature maps change according to\n\n|  |  |  | (3) | \n\nHere filters each channel separately, mixes channels, and is GroupNorm with four groups, no learned affine parameters, and . Each contains 32 learned gains, broadcast over spatial positions. The offset is the 32-parameter vector for updates 1–3 and the shared vector for updates 4–9. The shortcut adds the unchanged input to the learned correction, making these updates residual. All convolutions have stride one and no bias, with padding one for filters. After updates 3 and 6, max-pooling with stride two forms from . Otherwise . This gives three updates at each of , , and resolution.\n\nAfter update 9, average pooling with stride three produces features, flattened in channel-first order into . The scoring layer combines six learned weighted sums with twelve directly selected features ,\n\n|  |  |  | (4) | \n\nThe ordered, zero-based indices are . These features are selected greedily using training feature statistics to recover scoring information lost in the six-component approximation, and their indices remain fixed for every image at inference. Neither scoring matrix has a bias, and the largest score determines the digit. The convolutional network has 1,952 parameters and the scoring matrices have 948. Filter reuse supplies depth without storing nine separate filter sets, while the gains and offsets allow the updates to behave differently.\n\n##### Training for compression.\n\na2’s smaller scoring layer initially costs more compressed bytes, motivating a change in how its weights are trained. Quantization-aware training rounds weights during each forward pass while learning the underlying weights and scales, restricting the 768 projection parameters to scaled integers . A small weight penalty encourages small values alongside the classification and teacher-matching losses. This makes the smaller scoring layer competitive by favoring repeated integer values by exploiting redundancy in the learned weights. Across the final model, of stored integers are or , measured before applying scales. The compressed integer payload occupies 1,262 bytes, compared with 1,390 bytes for fixed-width packing of the same values using each tensor’s range.\n\nThe human reference, tiny_MNIST (Dhairyashil R. G., 2024), uses five separate convolutions with widths , BatchNorm, ReLU, and a dense -to- head. It has 3,130 parameters before BatchNorm folding. Our reproduction folds BatchNorm into the convolutions and applies per-channel 4-bit quantization, yielding 2,461 bytes at accuracy. The agents’ solution is 504 bytes () smaller and achieves test accuracy.\n\n### 3.4 Considerations for Effective Multi-Agent Communication\n\nPrior work finds that tasks admitting decomposition into independent subtasks are\nparticularly amenable to multi-agent collaboration (Kim et al., 2025; Gu et al., 2025).\nOur tasks admit no such decomposition, yet communication helps regardless.\nWhat makes it help is *verified progress sharing*. Once a discovery is confirmed to be\nuseful, every agent can build on it rather than rediscover it independently, and no agent is\nleft exploring a direction the others have already exhausted.\n\nIn the following discussion, we provide a simple pedagogical model that formalizes this intuition. We show how it can lead to an exponential separation between communicating and independent agents, and also provide a small-scale experiment on Terminal-Bench 2.0 (Merrill et al., 2026) where team@ improves over single-agent attempts but not independent best@, illustrating the importance of the availability of an agent-accessible verifier.\n\n##### From independent trajectories to cumulative progress.\n\nConsider a task requiring successive improvements, with an accessible verifier reporting score after improvement . Assume that continuation difficulty depends only on the stage, improvements can be transferred freely, and agents conduct fresh, independent searches after each transfer. Both methods can query the verifier. Writing for agent ’s search time at stage , their completion times would be\n\nThus . Independent sampling needs one agent\nto complete the entire sequence quickly, whereas a team can use a different\nagent’s breakthrough at each stage. Communication changes a *minimum of\nsums* into a *sum of minima*.\nFigure 12 illustrates this effect visually.\n\nSuppose , which is the standard constant-rate, memoryless model used in queueing theory (Shortle et al., 2018). Here we use this model as a pedagogical example, not an empirical claim about agents. The first of independent discoveries arrives at rate , giving\n\nThe Erlang law (a Gamma distribution with integer shape) follows from summing exponential waiting times. Doubling therefore halves the team’s mean completion time.\n\nNow give each agent runtime , where . This is shorter than one agent’s mean completion time but longer than the team’s. Both methods receive the same agent-time budget.\n\n###### Proposition 1 (Exponential separation from progress sharing).\n\nUnder this model, writing ,\n\nFor fixed , team@ approaches one exponentially in , while best@ success approaches zero exponentially.\n\nThe separation concerns completion probability at matched runtime, not an\nexponential expected-time speedup. The proof is in\nAppendix B.\nThe benefit also depends on preserving diverse continuations.\nSometimes, communicating agents may converge on a subset of approaches,\nwhich can reduce diversity and slow progress.\nTo model *herding*, we can imagine the  agents acting as essentially\n independent groups whose members duplicate the same search.\nThe team’s discovery rate becomes \nand its mean completion time becomes . The high-success regime now\nrequires . Sharing progress helps only to the extent that independent\nsearch survives after sharing.\n\n##### Ineffective communication in Terminal-Bench.\n\nTerminal-Bench 2.0 evaluates agents on hard, realistic tasks performed in command-line environments (Merrill et al., 2026). We evaluate Claude Sonnet 4.6 through the GitHub Copilot CLI on the complete set of 89 tasks. Test-time communication consists of two independent team@2 trials, where both agents contribute to one final container state. The baseline consists of four independent single-agent trials. This is a small-scale comparison with only a handful of trials, so we treat it as descriptive evidence rather than a precise estimate of the gap or its cause.\n\n| Method | Mean Accuracy | Max Accuracy | \n|---|---|---|\n| Single agent pass@1 | 52.53% | – | \n| Independent pass@2 | 62.36% | 64.04% | \n| Communicating team@2 | 60.67% | 61.80% | \n\nAs seen in Table 3, communication improves over an individual attempt, but does not outperform independent sampling (pass@) with the same number of agents. This result can be partially explained by Proposition 1 when is effectively . In Terminal-Bench, tasks require long sequences of actions, but the official verifier runs after agent execution. Available feedback varies across tasks, from package tests and direct performance measurements to checks of selected requirements. For example, the public evaluator in largest-eigenval checks the eigenvector equation and reports timings without verifying that the eigenvalue is dominant. In mailman, the public evaluator checks subscription but omits announcement delivery and unsubscription, which the official grader also tests. Such checks can confirm progress on individual requirements without establishing overall correctness.\n\nIn contrast, the packing and MNIST tasks expose numerical scores that agents can query repeatedly, while ARC-AGI-3 supplies level-success feedback. Thus, setting in Proposition 1 gives us , and therefore shows no communication advantage, even when the final outcome can be graded objectively. What matters is not evaluation alone, but dense faithful verification at test-time.\n\nA secondary disadvantage may be the loss of diversity. If communication causes agents to behave like only independent groups, then end-to-end candidates cover only , below pass@’s , where denotes the success probability of an independent candidate. This penalty is especially relevant here because team@2 contributes one shared final state, while pass@2 retains two isolated states and receives oracle selection after evaluation. Premature convergence or interference in the shared state can therefore make communication slightly worse than independent sampling. Together, the results suggest that communication is most useful when feedback is accessible and discriminative, though we do not exclude the possibility that test-time communication could be useful when feedback is sparse, if the agents can construct their own intermediate verification signals or have better communication protocol harnesses.\n\n## 4 Related Work\n\n##### Multi-agent systems.\n\nMulti-agent systems (MAS) are becoming increasingly popular and widely adopted.\nAs discussed in the introduction, some prior work (Du et al., 2024; Chen et al., 2024)\nreports gains from multi-agent debate and belief exchange.\nHowever, deliberation without a verifier or external feedback can amplify\ncorrelated errors or move the team away from its strongest member\n(Smit et al., 2024; Pappu et al., 2026).\nIn fact, independent sampling followed by aggregation remains a strong baseline\n(Wang et al., 2023; Li et al., 2024; Choi et al., 2025).\nBenefits also depend on task structure and model strength.\nKim et al. (2025) find stronger gains on decomposable tasks,\nbut these gains diminish as the underlying model becomes stronger with respect to the task.\nWhile Kim et al. (2025) comprehensively cover a range of tasks and configurations,\nour work studies the benefits of MAS in a more open-ended discovery setting with a clear verifier signal.\nAs discussed in Section 3.4, it is\n*verified progress sharing* that allows us to avoid “agent collapse”\nthat previous works have observed.\nFor more detailed surveys of MAS, we refer to Ferrag et al. (2026); Tran et al. (2025).\n\n##### Test-time discovery.\n\nTest-time discovery often uses LLMs (or agents) to generate, evaluate, and iteratively improve candidate solutions, rather than producing an answer in a single attempt from a pretrained model. Within this paradigm, Novikov et al. (2025) keep the underlying LLMs frozen and accumulate progress through evolutionary program search, evaluator feedback, and a database of generated attempts. Wang et al. (2026); Yuksekgonul et al. (2026), by contrast, train the model via reinforcement learning during discovery, updating the LLM using feedback obtained from the target problems. Recent work also focuses on custom agentic workflows, where agents are assigned specialized roles to generate, critique, and refine hypotheses, often orchestrated by a central agent (Gottweis et al., 2026; Yamada et al., 2025). Our setting, however, is orthogonal to the works above as they focus on either single-agent capabilities or multi-agent workflows without establishing an advantage over single-agent systems under matched compute.\n\nRecent work illustrates the potential of multi-agent communication for discovery. Anthropic (2026a) describe an improved attack on HAWK whose key idea emerged through an exchange between two agents, while OpenAI (2026a) report a Navier–Stokes result produced by a group of roughly 10,000 concurrent agents. Similarly, Bianchi et al. (2026) explore mathematical discoveries by an open community of heterogeneous agents sharing solutions and encouraging public discussion. These are systems in which agents choose how to explore and communicate without a fixed workflow.\n\nThese findings motivate controlled tests of whether communication improves on independent search, especially under matched compute. Anthropic (2026b) report more vulnerabilities from coordinating agents than from independent agents, though Claude Mythos Preview uses comparable tokens per vulnerability when both methods are evaluated on the same search scope. On other tasks, they observe failures to integrate agents’ work, conformity, and failures to share decisive evidence. Closely related to our setup is Qu et al. (2026), whose agents share discoveries through persistent memory and outperform the best of four independent runs under matched wall-clock budgets. In contrast, we diagnose clear conditions for when test-time communication is advantageous, through compute-controlled studies of scaling coordinating agents on ARC-AGI-3 and long-horizon algorithm/ML tasks, including negative results at low-compute budgets and on Terminal-Bench 2.0.\n\n##### Harness optimization and orchestration.\n\nOne line of work aims to learn the ideal topology and communication protocol of MAS. Agent graphs and communication links can be optimized directly (Zhuge et al., 2024; Zhang et al., 2025c), while architecture search can adapt the system to the task (Zhang et al., 2025b; Yun et al., 2026). A second line controls when and which agents participate. Systems can select teams or gate communication (Liu et al., 2024; Ghosh and Chakraborty, 2026), regulate interaction to avoid harmful scaling (Wang et al., 2025; Shao et al., 2026), or add agents dynamically and decompose work into explicit subtasks (Costa, 2026; Gu et al., 2025). Gu et al. (2025) investigate agent diversity and observe weaker returns from homogeneous teams, though heterogeneity may bring down the performance of the best model in the team even when explicitly told who the best is (Pappu et al., 2026).\n\nHierarchical systems place these decisions with a central coordinator. Coordinators can be trained to assign work to frozen agents (Xu et al., 2026; Nielsen et al., 2026), while adaptive hierarchies can delegate work and synthesize results, primarily for parallel decomposition (Tang et al., 2026; OpenAI, 2026b). Recursive delegation has also been used for long-context inference rather than multi-agent orchestration (Zhang et al., 2025a). Harness optimization is broader and need not involve multiple agents. Executable or prompt-level agent loops can be optimized from scores, traces, or revision feedback (Lee et al., 2026b; Lee et al., 2026a). Prompt optimization has also been studied for multi-agent protocols (Bai and Shi, 2026).\n\nUnlike these approaches, we do not study orchestration or multi-agent harness optimization. Instead, we use a fixed shared-workspace scaffold of identical, unassigned peers and compare its performance with compute-matched independent runs to isolate communication from parallel exploration.\n\n##### Agent coordination as interactive systems.\n\nWe view sequential Monte Carlo as an interacting-population counterpart to best-of-, with partial solutions repeatedly scored and resampled (Gordon et al., 1993; Liu and Chen, 1998; Moral, 2004). Populations degenerate over long horizons (Kong et al., 1994; Doucet and Johansen, 2011), while resampling concentrates shared ancestors, and harder problems require larger populations (Liu and Chen, 1998; Snyder et al., 2008). These phenomena correspond to our gains at depth, herding failures, and non-monotone returns to team size.\n\nRelated methods now guide language model inference. Particle methods support probabilistic inference and test-time scaling (Zhao et al., 2024; Puri et al., 2025), as well as constrained generation (Loula et al., 2025). Interacting particles have also been adapted to diffusion models (Singhal et al., 2025; Luo et al., 2026), with related path selection for diffusion language model decoding (Lee et al., 2025). The analogy is conceptual and suggests that coordination is most useful over long horizons with informative intermediate feedback and a diverse population.\n\n## 5 Conclusion and Future Directions\n\nIn this paper, we studied the capabilities of test-time communication, a multi-agent setting in which agents pursue the same open-ended objective and exchange evidence through a shared workspace. This allows agents to explore different approaches in parallel while turning individual discoveries into cumulative progress. On ARC-AGI-3, team@3 matches the solve rate of 13 independent agents and team@5 matches that of 33, with the multiplier growing from to as the team scales. Communication also solves a game that remains unsolved across single-agent trials and makes the average member of a team as action-efficient as the best of independent agents. The same pattern emerges in longer-horizon tasks. Team@ reaches a packing score of on Frontier-CS polyomino packing, above the prior best-known score of , and produce a 1,957-byte MNIST classifier, improving on both the best@4 result and the 2,461-byte best-known human solution. These gains emerge only after an initial coordination cost, but with adequate compute, communication can outperform independent agents in both performance and efficiency.\n\nThe agent-accessible verifier in each task is central to our experimental design. Our setting differs from multi-agent debate, which asks agents to reconcile answers without necessarily grounding the exchange in environmental feedback. Level completion, packing score, and compression size give agents an objective basis for rejecting failures, comparing approaches, and building on improvements. Whether communication produces similar gains when feedback is sparse, noisy, or subjective remains an important open question.\n\nOn the other hand, multi-agent harnesses may matter as much as single-agent harnesses. Our controlled design is deliberately narrow. We fix one communication harness and use homogeneous agents with identical instructions and no assigned roles or central orchestrator. This isolates communication, but it does not imply that roles, model diversity, communication topology, or orchestration are unimportant. Future work should consider how agents communicate, which peers receive their messages, or how teams should be organized. These policies could be engineered or evolved, similarly to harness optimization (Lee et al., 2026b; Lee et al., 2026a). Our results provide a baseline for future work, and hopefully motivate further study of multi-agent collaboration and test-time communication.\n\n#### Acknowledgments\n\nWe thank Microsoft Research AI Frontiers for supporting this work and Ziyang Cai for insightful discussions during the project. This work was done during the first author’s internship at Microsoft Research.\n\n## References\n\n- Discovering Cryptographic Weaknesses with Claude. Note: Research report External Links: Link Cited by: §4.\n- Patterns and Problems in Emerging Multiagent Systems. Note: Research report External Links: Link Cited by: §1, §4.\n- ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence. Note: Technical report External Links: Link Cited by: §1, §2.1, §3.1.\n- MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems?. External Links: Document, 2606.23664, Link Cited by: §4.\n- Harnessing the collective intelligence of AI agents in the wild for new discoveries. Note: arXiv preprint arXiv:2606.10402 External Links: 2606.10402, Document, Link Cited by: §4.\n- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling. External Links: Document, 2407.21787, Link Cited by: §1.\n- MLE-bench: evaluating machine learning agents on machine learning engineering. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §1.\n- ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, pp. 7066–7085. External Links: Document, Link Cited by: §4.\n- Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?. External Links: Document, 2508.17536, Link Cited by: §1, §1, §4.\n- AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation. External Links: Document, 2602.07072, Link Cited by: §4.\n- tiny_MNIST: high accuracy with a very small model. Note: Software repository External Links: Link Cited by: §A.5, §2.1, §3.3.2.\n- A tutorial on particle filtering and smoothing: fifteen years later. In The Oxford Handbook of Nonlinear Filtering, D. Crisan and B. Rozovskii (Eds.), pp. 656–704. Cited by: §4.\n- Improving Factuality and Reasoning in Language Models through Multiagent Debate. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, pp. 11733–11763. External Links: Link Cited by: §4.\n- From LLM reasoning to autonomous ai agents: a comprehensive review. IEEE Access. Cited by: §4.\n- Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning. External Links: Document, 2607.10836, Link Cited by: §4.\n- GitHub Copilot CLI. Note: GitHub Docs External Links: Link Cited by: §2.1.\n- Novel approach to nonlinear and non-gaussian bayesian state estimation. IEE Proceedings F, Radar and Signal Processing 140 (2), pp. 107–113. External Links: Document Cited by: §4.\n- Accelerating scientific discovery with Co-Scientist. Nature 655, pp. 487–496. External Links: Document, Link Cited by: §4.\n- AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System Need. External Links: Document, 2506.15451, Link Cited by: §3.4, §4.\n- Harbor: a framework for evaluating and optimizing agents and models in container environments. External Links: Document, Link Cited by: §2.1.\n- Self-improvement in language models: the sharpening mechanism. In International Conference on Learning Representations, Vol. 2025, pp. 76687–76739. Cited by: §1.\n- Towards a Science of Scaling Agent Systems. External Links: Document, 2512.08296, Link Cited by: §1, §3.4, §4.\n- Sequential imputations and Bayesian missing data problems. Journal of the American Statistical Association 89 (425), pp. 278–288. External Links: Document Cited by: §4.\n- Gradient-based learning applied to document recognition. Proceedings of the IEEE 86 (11), pp. 2278–2324. External Links: Document, Link Cited by: §2.1.\n- Recursive Harness Self-Improvement. External Links: Document, 2607.15524, Link Cited by: §4, §5.\n- Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models. External Links: Document, 2511.05563, Link Cited by: §4.\n- Meta-Harness: End-to-End Optimization of Model Harnesses. External Links: Document, 2603.28052, Link Cited by: §4, §5.\n- More Agents Is All You Need. Transactions on Machine Learning Research. External Links: 2402.05120, Link Cited by: §1, §4.\n- Sequential Monte Carlo methods for dynamic systems. Journal of the American Statistical Association 93 (443), pp. 1032–1044. External Links: Document Cited by: §4.\n- A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration. In Conference on Language Modeling, External Links: 2310.02170, Link Cited by: §4.\n- Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo. External Links: Document, 2504.13139, Link Cited by: §4.\n- Self-Rewarding Sequential Monte Carlo for Masked Diffusion Language Models. External Links: Document, 2602.01849, Link Cited by: §4.\n- FrontierCS: Evolving Challenges for Evolving Intelligence. In Proceedings of the 43rd International Conference on Machine Learning, Note: To appear External Links: Link Cited by: §1, §2.1, §3.2.\n- Terminal-Bench: benchmarking agents on hard, realistic tasks in command line interfaces. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1, §3.4, §3.4.\n- Feynman-Kac formulae: genealogical and interacting particle systems with applications. Springer, New York. External Links: Document Cited by: §4.\n- Learning to Orchestrate Agents in Natural Language with the Conductor. In International Conference on Learning Representations, External Links: 2512.04388, Link Cited by: §4.\n- AlphaEvolve: a coding agent for scientific and algorithmic discovery. Note: arXiv preprint arXiv:2506.13131 External Links: 2506.13131, Document, Link Cited by: §4.\n- On the Navier–Stokes Millennium Prize Problem. Note: Research report External Links: Link Cited by: §4.\n- The Builder’s Guide to GPT-5.6. Note: Technical guide External Links: Link Cited by: §4.\n- Multi-Agent Teams Hold Experts Back. In Proceedings of the 43rd International Conference on Machine Learning, External Links: 2602.01011, Link Cited by: §1, §1, §2, §4, §4.\n- The republic of science: its political and economic theory. Minerva 1 (1), pp. 54–73. Cited by: §1.\n- Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs Using Particle-Based Monte Carlo Methods. External Links: Document, 2502.01618, Link Cited by: §4.\n- CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery. In Conference on Language Modeling, External Links: 2604.01658, Link Cited by: §2.1, §3.2, §4.\n- MonoScale: Scaling Multi-Agent System with Monotonic Improvement. In Proceedings of the 43rd International Conference on Machine Learning, External Links: 2601.23219, Link Cited by: §4.\n- Fundamentals of queueing theory. 5 edition, John Wiley & Sons. External Links: ISBN 9781118943526 Cited by: §3.4.\n- A General Framework for Inference-Time Scaling and Steering of Diffusion Models. External Links: Document, 2501.06848, Link Cited by: §4.\n- Should We Be Going MAD? A Look at Multi-Agent Debate Strategies for LLMs. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, pp. 45883–45905. External Links: Link Cited by: §1, §4.\n- Scaling LLM Test-Time Compute Optimally Can Be More Effective than Scaling Parameters for Reasoning. In International Conference on Learning Representations, External Links: 2408.03314, Link Cited by: §1.\n- Obstacles to high-dimensional particle filtering. Monthly Weather Review 136 (12), pp. 4629–4640. External Links: Document Cited by: §4.\n- The knowledge machine: how irrationality created modern science. Liveright, New York. Cited by: §1.\n- Sakana Fugu Technical Report. External Links: Document, 2606.21228, Link Cited by: §4.\n- Multi-agent collaboration mechanisms: a survey of LLMs. arXiv preprint arXiv:2501.06322. Cited by: §4.\n- MARS: Toward More Efficient Multi-Agent Collaboration for LLM Reasoning. External Links: Document, 2509.20502, Link Cited by: §4.\n- Self-Consistency Improves Chain of Thought Reasoning in Language Models. In International Conference on Learning Representations, External Links: 2203.11171, Link Cited by: §1, §4.\n- ThetaEvolve: test-time learning on open problems. In Forty-third International Conference on Machine Learning, External Links: Link Cited by: §4.\n- Can LLM Agents Really Debate? A Controlled Study of Multi-Agent Debate in Logical Reasoning. External Links: Document, 2511.07784, Link Cited by: §1.\n- TRINITY: An Evolved LLM Coordinator. In International Conference on Learning Representations, External Links: 2512.04695, Link Cited by: §4.\n- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search. External Links: Document, 2504.08066, Link Cited by: §4.\n- SWE-agent: agent-computer interfaces enable automated software engineering. In Advances in Neural Information Processing Systems, Vol. 37, pp. 50528–50652. External Links: Document, Link Cited by: §1.\n- Learning to discover at test time. External Links: Link Cited by: §3.2, §4.\n- Graph-of-Agents: A Graph-Based Framework for Multi-Agent LLM Collaboration. In International Conference on Learning Representations, External Links: 2604.17148, Link Cited by: §4.\n- Recursive Language Models. External Links: Document, 2512.24601, Link Cited by: §4.\n- Multi-Agent Architecture Search via Agentic Supernet. In Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 267, pp. 75834–75852. External Links: Link Cited by: §4.\n- G-Designer: Architecting Multi-Agent Communication Topologies via Graph Neural Networks. In Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 267, pp. 76678–76692. External Links: Link Cited by: §4.\n- Stop Overvaluing Multi-Agent Debate—We Must Rethink Evaluation and Embrace Model Heterogeneity. External Links: Document, 2502.08788, Link Cited by: §1.\n- Probabilistic Inference in Language Models via Twisted Sequential Monte Carlo. External Links: Document, 2404.17546, Link Cited by: §4.\n- GPTSwarm: Language Agents as Optimizable Graphs. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, pp. 62743–62767. External Links: Link Cited by: §4.\n\n## Appendix A Full Experimental Setup\n\nThis section provides the communication-prompt and implementation details deferred from Section 2.1, followed by task-specific evaluation details for the experiments reported in the main text. The tasks, prompts, communication protocol, and further details can be found in https://github.com/jerryjonghopark/test-time-communication.\n\n### A.1 Communication prompts, runtime, and accounting\n\n##### Communication-prompt summary.\n\nBeyond the shared-workspace instructions in Section 2, the prompt specifies when agents may adopt a peer’s approach. Adoption requires a measured improvement, replication, an actionable alternative to a blocked approach, or final convergence. After adoption, agents are instructed to preserve a meaningful variation and record what they adopted and why.\n\nFor MNIST classifier compression, agents build and validate candidates in scratch space, then recheck the current best under the lock before promoting a strictly better submission. These are instructions to the agents, rather than automatic verification or adoption by the harness.\n\n##### Runtime versions.\n\nThe ARC-AGI-3 and Terminal-Bench 2.0 experiments use GitHub Copilot CLI 1.0.54, the long-horizon packing experiments use version 1.0.70, and MNIST classifier compression uses version 1.0.78.\n\n##### Shared and per-agent resources.\n\nCPU and host-memory limits apply to the entire task container. All workers in a team share those limits, rather than each receiving a separate container-sized allocation. Independent trials each have their own container. The configured allocations are shown below.\n\n| Task | Solo container | Team container | Task GPU | \n|---|---|---|---|\n| ARC-AGI-3 | 4 CPUs, 8 GiB | 4 CPUs, 8 GiB | None | \n| Polyomino packing | 2 CPUs, 6 GiB | 2 CPUs, 6 GiB | None | \n| MNIST compression | 3 CPUs, 8 GiB | 12 CPUs, 32 GiB () | A100 | \n| Terminal-Bench 2.0 | Task-defined | Task-defined, shared | Task-defined | \n\nThe ARC and packing allocations do not scale with team size. MNIST scales the container’s CPU and host-memory limits by four for team@4, but does not enforce separate per-worker host-memory partitions. Its CPU thread pools are set to three threads per worker. Each agent is instructed to use at most 1,000 MiB of aggregate GPU memory across all its processes. A 400-MiB PyTorch allocator setting leaves headroom for CUDA context and library memory, but is not a hard limit on total GPU memory. The per-agent GPU-memory rule is self-enforced. These task-compute resources are distinct from model inference through the Copilot service and from ARC’s separate per-agent action budgets.\n\n### A.2 Agent Communication Prompt\n\nEvery agent in a team receives the prompt below, in addition to the task\ninstruction. The text is stored as a template, and before a run starts the\nharness substitutes the team size into `{n}` and `{nm1}` and the\nshared-scratch paths into the remaining braced fields. Every agent in a trial\nthen receives the identical filled-in text, so the only asymmetry between\nagents comes from the slot each one wins at runtime.\n\n```\n{n} agents share this container and work the same task in parallel, all with this\n    identical prompt. Search widely without herding, coordinate as you go, and keep\n    improving until time runs out.\n    Shared scratch (create on first use):\n      - slots/approaches: {slots}/\n      - findings:         {findings}\n      - disconfirmations: {disconfirm}\n      - score log:        {plateau}\n      - coordination:     {coordination}   (empty -- conventions you author)\n    Your private scratch is {base}/work-$S after you claim slot $S.\n    1. CLAIM A SLOT AND PICK A DISTINCT APPROACH:\n         mkdir -p {slots}\n         for i in $(seq 0 {nm1}); do mkdir \"{slots}/slot-$i\" 2>/dev/null && S=$i && break; done\n         mkdir -p {base}/work-$S\n         touch {coordination}\n         echo \"<your approach + what you will deliberately not assume>\" > \"{slots}/slot-$S/approach\"\n         cat {slots}/slot-*/approach\n       If a lower-numbered slot already took your approach, change yours. Cover a\n       different part of the search space; do not agree early.\n    2. ALWAYS BE ACTING. Never end a turn with only prose or a plan. Every turn must\n       run a command that advances or tests the work: take a real action the task\n       accepts and read back its result. Publishing, disconfirming, coordinating, and\n       pivoting are bookkeeping around real actions, never a substitute for taking one.\n    3. EVALUATE AND RECORD EVERY ATTEMPT. Score each change with the task’s own\n       scoring or feedback mechanism (find it in AGENT.md or the task instructions),\n       then log the result:\n         echo \"[slot $S $(date -u +%H:%M:%S)] score=<score-or-progress> family=<approach>\" >> {plateau}\n    4. SHARE WITHOUT HERDING. Append concise findings with evidence and cost to\n       {findings}, and label weak claims as weak. Spend part of your effort trying to\n       FALSIFY the leading idea or your own, recording negative results in\n       {disconfirm}. Do not write prose whose only purpose is to make peers copy you.\n    5. COORDINATE ON SHARED RESOURCES. You share the graded output, common files, and\n       the environment with peers who run blind to your session. Treat every shared\n       resource as contested: before you touch one, re-check {coordination} and the\n       resource’s current state; after, confirm your change survived and did not just\n       repeat a peer’s. A collision is any overwritten, duplicated, or conflicting work\n       that wastes effort. No coordination scheme is provided -- {coordination} is empty\n       and yours to author: on a collision, write a convention there concrete enough for\n       a peer to follow, that changes your next action, then follow it.\n    6. HIGH BAR TO CONVERGE. Keep your own approach unless another clearly beats it on\n       a measured, reproduced result, or yours is blocked, or the run is wrapping up.\n       Even then, keep one real difference (a parameter, subcase, representation, or\n       fallback) until the very end.\n    7. NEVER STOP WHILE TIME REMAINS. A working result is not the finish line; the\n       clock running out is the only acceptable reason to stop. Do not declare the\n       task done, final, solved, or \"at the ceiling\" and go idle -- a suspected\n       ceiling is a claim to disconfirm, not a reason to quit.\n    8. BREAK PLATEAUS BY CHANGING FAMILY. You are plateaued when your best score has\n       not strictly improved over 3 consecutive attempts. Then stop tuning\n       and switch to a STRUCTURALLY DIFFERENT approach -- a different core principle or\n       assumption, not a variant of the current one. Keep a short list of untried\n       families in {base}/work-$S so you always have a next one ready. Read peers’\n       approaches and {plateau} first and pick a family no active peer is on; adopting\n       a peer who is also plateaued is not progress.\n    Run a tight loop -- change -> evaluate -> record -> repeat -- without pausing. Do\n    not use the internet, curl, wget, HTTP libraries, or secrets.\n```\n### A.3 ARC-AGI-3\n\n##### Task and interface.\n\nAgents interact with the games by receiving integer-valued grids and selecting from the available discrete controls or coordinate clicks. Each level has its own cap, so unused budget from a later level cannot rescue an agent stalled earlier. We use the benchmark-native cap of five times the corresponding human-action baseline, independently for every agent.\n\nEvery team member owns a separately authenticated ARC session. Sharing an action trace does not copy state and does not clear a peer’s level. Time-dependent progress curves use the first logged completion of each level.\n\n##### Synchronization.\n\nTeam runs synchronize at intervals of half the current level’s action budget, rounded up to an integer number of actions. The harness tracks each agent’s phase by its level and the number of these intervals it has used. An agent ahead of the slowest live teammate, either within a level or by reaching the next level, cannot take another game action until its teammates catch up or leave the live set. During this pause, agents can still inspect state, run shell commands, and exchange notes. Refused game actions consume no budget. Finished, game-over, and budget-exhausted agents do not block their teammates, and inactive sessions are excluded after an idle timeout. This synchronization applies only to communicating teams, not to best@.\n\n##### Closed-book condition.\n\nThe task container requires network egress for the model control plane, but agents receive no browser or web-search tool and are explicitly forbidden to fetch public solutions, replays, or ARC pages. Only the task prompt, local run files, the agent’s own environment observations, and teammates’ within-trial notes are admissible. We inspect run trajectories afterwards to verify that no agent accessed external ARC information.\n\n### A.4 Frontier-CS polyomino packing\n\n##### Objective and scorer.\n\nEach of 70 hidden cases contains edge-connected polyominoes of one to ten cells. A program may reflect, rotate by multiples of , and translate each piece. All pieces must lie without overlap in one integer-grid rectangle. If is the total number of occupied cells and the returned area, the normalized case quality is and the reported score is its mean over the fixed cases. Invalid output causes the submission to be rejected. Candidate programs compile as GNU C++17 and receive 2 seconds and 256 MiB per case.\n\nThe hidden instance directory is mounted read-only inside a separate scorer service and is never mounted into the agent container, which contains the task statement and a submission client. On each attempt, the client sends the candidate C++ source over an internal endpoint. The scorer compiles and executes it, with case-output retention disabled, and returns the aggregate score together with compact status, timing, and scoring metadata. The scorer never returns the instances or program outputs. Agents may submit repeatedly without penalty, but cannot inspect the evaluation data. Every attempt, timestamp, status, and score is recorded in a submission ledger. The final verifier retains the highest valid scored submission in the run, protecting an earlier champion from a broken final edit.\n\n### A.5 MNIST classifier compression\n\n##### Data and artifact contract.\n\nThe official 60,000-image MNIST training split is partitioned once into 55,000 training and 5,000 development images. The agent image contains only these two partitions, while the official 10,000-image test archive remains on the host and is never copied into the task container. Agents may train only on the 55,000 images, use the development set for selection, and query the sealed oracle only after a full-development accuracy of at least .\n\nAn oracle request snapshots the canonical submission and passes that snapshot to a host-owned evaluator. Each request runs in a fresh container with networking disabled, a read-only root filesystem, privilege escalation disabled, and the submission mounted read-only. Within it, a trusted verifier privately shuffles the test set, makes the test archive and labels unreadable to submitted code, and invokes inference under an unprivileged user. The submission receives only batches of test images and writes predictions to isolated scratch space. The trusted verifier alone reads the labels and computes accuracy. The container is discarded after evaluation, and the oracle returns only aggregate accuracy. It never returns examples, labels, predictions, or per-example errors. Final grading uses the same isolation boundary, with a 600-second limit for classifying the complete test set.\n\nThe graded directory must include a predict.py entry point and every weight, table, constant, generator, and decoder it needs. We normalize file order and metadata, form a deterministic tar archive, and apply gzip level 9. Training code outside the submission and the preinstalled numerical runtime are not charged.\n\n##### Open-source reference.\n\nThe 2,461-byte reference in Section 2.1 is our benchmark-format reproduction of the open-source tiny_MNIST model [Dhairyashil R. G., 2024]. The published architecture has 3,130 trainable parameters and reports accuracy above . We retrained it from the published recipe, folded BatchNorm exactly into the adjacent convolutions, and quantized the resulting weights per output channel to four bits.\n\nFor the deployable artifact, we bit-pack this state, use a compact decoder, and for fair comparison, code-golf the inference path by eliminating general training machinery, redundant metadata, whitespace, and intermediate file structure while preserving predictions. The complete normalized archive, including executable inference code, is 2,461 bytes under the same deterministic gzip-9 metric used for every submission. The reproduced artifact attains exactly on the sealed test set. Thus 2,461 bytes is a measured end-to-end submission size, not the size of the upstream training checkpoint.\n\n##### Environment.\n\nThe pinned stack is Python 3.12.13, PyTorch 2.5.1 with CUDA 12.4, torchvision 0.20.1, NumPy 2.5.1, SciPy 1.18.0, and scikit-learn 1.9.0. Agents train from scratch on the provided data. Downloading additional data or pretrained weights is prohibited, and inference evaluation has no network access.\n\n##### Checkpoints and trajectory accounting.\n\nAgents are instructed to save each smaller submission that passes the full development-set check, together with its result, timestamp, provenance, and recomputed archive size. Saved checkpoints are audited, and the paper figures use only submissions confirmed at on the sealed test set. Figure 10 compares four independent GPT-5.6 Sol runs with one team@4 run. In the elapsed-time panel, best@4 at time is the smallest qualifying submission found by any of the four independent agents within its own first hours. The output-token panel uses the same best-so-far trajectory, placing each improvement at the sum of tokens consumed by all four independent runs by that elapsed time. Team tokens likewise sum all four workers. Neither curve counts only the tokens of the agent that produced an improvement.\n\n### A.6 Terminal-Bench 2.0\n\n##### Task environments and execution.\n\nThe harness loads terminal-bench@2.0 through Harbor and runs each task in its Docker environment. CPU, host-memory, storage, and GPU requests come from the individual task configuration, rather than a uniform benchmark-wide allocation. The harness does not automatically multiply these requests by the number of workers. In a team@2 trial, both workers share the task container and its resource limits. Agent execution and verifier timeouts are configured separately by each task, so there is no single wall-clock horizon analogous to the packing or MNIST experiments. The verifier runs after agent execution and produces the trial’s reward from the resulting environment.\n\n##### Aggregation.\n\nFor Table 3, pass@1 averages the four independent outcomes for each task. The saved independent results are grouped into two batches of two attempts. Pass@2 counts a task as successful within a batch if either attempt passes, then averages the two batch accuracies. Team@2 averages the two team-trial outcomes for each task. Unfinished tasks are assigned zero, and all 89 tasks remain in the denominator. Max Accuracy selects the higher batch or team-trial accuracy over the whole benchmark, rather than selecting a different batch or team trial for each task.\n\n## Appendix B Analysis for Verified Progress Sharing\n\nThis appendix gives the calculations behind the conceptual model and Proposition 1. The mathematics is largely standard rather than new. It combines elementary facts about exponential order statistics, Erlang sums, and Chernoff bounds. We keep the model deliberately simple so that it isolates how reusable, verified progress can change the value of parallel search. It is not intended as an empirical model of agent completion times or as a general theory of multi-agent communication.\n\n### B.1 Idealized progress sharing\n\nWe first compare independent and communicating agents under the same collection of stage-level search times. Recall that is the time agent would take to discover the next improvement at stage . A communicating team uses the first discovery at every stage. Therefore, for every agent ,\n\nTaking the minimum over agents gives . This is the pointwise comparison in the main text. It captures the distinction between combining stage-level breakthroughs and selecting one complete trajectory after all runs finish.\n\nUnder the modeling assumption , define the waiting time for the team’s next verified improvement as . Its survival probability is\n\nThus each is distributed as . Independence across stages then gives\n\nFor comparison, each independent agent must complete all stages itself. Its completion time satisfies\n\nThese distributions provide the ingredients for the fixed-budget comparison.\n\n### B.2 Completion probability at a fixed budget\n\nWe now prove Proposition 1. The only additional ingredient is a standard exponential tail bound for an Erlang random variable.\n\n###### Proof of Proposition 1.\n\nLet . Its moment-generating function is\n\nFor , apply Markov’s inequality with to obtain\n\nwhere . For , the same choice gives . Since implies , Markov’s inequality gives the corresponding lower-tail bound. Hence\n\n|  |  |  |  |  |  | (5) | \n\nSet , with . Since has the same distribution as and ,\n\nLikewise, has the same distribution as , so gives\n\nTaking a union bound over the independent agents yields\n\nCombining the two bounds proves\n\nBecause for and , both exponents are strictly positive for fixed and . ∎","body_html":"<h1 id=\"scaling-discovery-through-test-time-communication\">Scaling Discovery through Test-Time Communication</h1>\n<p>Jongho Park&lt;sup&gt;b&lt;/sup&gt;&lt;sup&gt;1&lt;/sup&gt;&lt;sup&gt;1&lt;/sup&gt;\n        1</p>\n<pre><code>  This work was done during an internship at Microsoft Research.  Vasilis Kontonis&lt;sup&gt;m&lt;/sup&gt;  Shivam Garg&lt;sup&gt;m&lt;/sup&gt;</code></pre>\n<p>Akshay Krishnamurthy&lt;sup&gt;m&lt;/sup&gt;  Dimitris Papailiopoulos&lt;sup&gt;m&lt;/sup&gt;</p>\n<p>&lt;sup&gt;b&lt;/sup&gt;UC Berkeley   &lt;sup&gt;m&lt;/sup&gt;Microsoft Research</p>\n<p>&lt;sup&gt;0&lt;/sup&gt;</p>\n<p>&lt;sup&gt;0&lt;/sup&gt;footnotetext: Emails: jjhpark@berkeley.edu, {vkontonis, shigarg, akshay.krishnamurthy, dimitriosp}@microsoft.com</p>\n<p>Science advances not in isolation but through collaboration,\nyet existing agentic systems capture little of this.\nWhether communicating agents help remains an open question\nwith mixed prior results.\nWe show that *test-time communication can substantially outperform\nindependent parallel attempts on challenging tasks, where sharing a\nbreakthrough can push the whole group forward*.\nWe first study the effect of scaling multi-agent test-time communication,\nwhere agents have no predefined roles and communicate via a shared directory,\non ARC-AGI-3, a benchmark requiring novel problem solving.\nWe find that a team of  communicating agents, team@, matches the success rate of\n independent agents, and this advantage grows with\n, suggesting gains compound with scale.\nThe effect is not merely efficiency: a task that no single agent can solve, a\nteam of agents can solve reliably. Furthermore, these gains\ntransfer to research-oriented tasks, given sufficient compute. On\npolyomino packing, communicating agents outperform best@\nand exceed the prior best-known score. On MNIST classifier compression, communication\nsurpasses the best-known human solution. A team of four agents produced a\n1,957-byte classifier submission achieving 99.4% test accuracy,\nsmaller than both the best-known human solution and the best single-agent\nresult. These gains are not unconditional.\nIndependent agents may outperform communication when\ncompute is limited or when a clear measure of progress is absent.\nHowever, under sufficient compute and clear feedback, multi-agent communication consistently yields stronger results.</p>\n<h2 id=\"1-introduction\">1 Introduction</h2>\n<p>For test-time compute, it is standard to sample  independent outputs in parallel from a large language model (LLM) and aggregate them, <em>e.g.</em>, by taking the majority answer or\nthe best@ when a verifier is available (Wang et al., 2023; Brown et al., 2024; Li et al., 2024; Huang et al., 2025; Snell et al., 2025).\nIndependent attempts, however, leave discoveries made during one run unavailable to\nguide another while its search is still unfolding.\nThis opportunity is especially crucial for LLM agents, which operate a terminal, execute\ncode, use tools, and observe results over horizons of hours or\ndays (Yang et al., 2024; Chan et al., 2025; Merrill et al., 2026).\nTheir trajectories contain intermediate results, failed experiments, and reusable artifacts\nthat could help other agents avoid dead ends or build on promising approaches.</p>\n<p>Communication between agents would be the natural remedy, but the evidence for its benefits is mixed. While keeping the best of independent candidates improves with , having agents exchange opinions may not (Choi et al., 2025). Multi-agent debate often fails to outperform chain-of-thought prompting despite using substantially more inference compute (Smit et al., 2024; Zhang et al., 2025d). Agents can defer to an incorrect majority or dilute expertise by compromising between expert and non-expert judgments (Wu et al., 2025; Pappu et al., 2026). Most recently, Anthropic (2026b) report that coordinating agents find more vulnerabilities than independent agents, but use more tokens and search more broadly. Within the common search scope, tokens per vulnerability are comparable, leaving the efficiency gains from communication unclear. On the other hand, multi-agent systems clearly help when a task can be decomposed into subtasks and solved in parallel (Kim et al., 2025), but such gains are limited to cleanly decomposable tasks. Can communication offer more than just parallelism?</p>\n<p>Science suggests that it can. Scientists coordinate by adjusting their research to the results achieved by others (Polanyi, 1962), while accumulated empirical evidence helps distinguish competing explanations and turn varied efforts into collective progress (Strevens, 2020). This process is iterative rather than merely parallel. Collaboration should therefore do more than divide and conquer. One agent’s observation redirects another’s search, and partial discoveries combine into a result that no individual reaches alone. We ask whether agent communication can produce the same effect.</p>\n<p>To this end, we study a <em>minimal form of test-time communication</em>\nin which identical agents receive the same open-ended objective and communicate\nthrough a shared workspace, without predefined roles or a central orchestrator.\nAutonomously and asynchronously, the agents can\nexchange intermediate results, failures, and artifacts while their search is still\nunfolding. Our central comparison pits a team of  communicating agents, team@, against\nthe best result from  independent agents, best@, under the same per-agent resources.\nThis isolates the value of communication from the gains produced by additional parallel\nattempts.</p>\n<p>We first study ARC-AGI-3 (ARC Prize Foundation, 2026) using Claude Sonnet 4.6. In terms of full-game success rate, team@3 matches best@13 and team@5 matches best@33, making them as effective as and times as many independent agents (Figure 1). More strikingly, communicating teams solve games that independent agents rarely or never solve. These benefits transfer to longer-horizon research problems. On Frontier-CS polyomino packing (Mang et al., 2026), where agents pack polyominoes into a minimum-area rectangle, team@4 scores 0.922 with Claude Opus 4.6 and 0.910 with Sonnet 4.6, compared with 0.893 and 0.891, respectively, for the strongest solo run of each model and 0.894 for the prior best-known result. On MNIST classifier compression under an accuracy constraint, we evaluate team@4 with GPT-5.6 Sol over 96 hours. The best team produces a state-of-the-art 1,957-byte classifier at 99.4% accuracy, beating the 2,461-byte best-known human solution and single-agent baselines.</p>\n<p>These results suggest that test-time communication is most useful when agents have\nenough compute to build on one another’s discoveries and can objectively assess\nwhether those discoveries improve on earlier results.\nSufficient compute allows teams to overcome initial coordination costs, while an\naccessible verifier, such as level success or a solution scorer, helps agents\ndecide which discoveries to adopt.\nWe call this mechanism <em>verified progress sharing</em>.\nBy building on successive breakthroughs from different members, a communicating team can\nturn parallel exploration into cumulative progress.</p>\n<p>When agents cannot verify intermediate progress, however, communication may offer less benefit. On Terminal-Bench 2.0, where available feedback may not reliably rank intermediate solutions, we find that team@2 improves over a single attempt but does not outperform independent pass@2. This distinction also offers an explanation for the prior negative results, where agents lack a verifier to distinguish better solutions from worse ones and therefore abandon stronger solutions in favor of weaker ones (Choi et al., 2025; Pappu et al., 2026). Together, our findings show that communication can offer more than divide and conquer parallelism and help identify the conditions under which test-time communication excels.</p>\n<p>Our contributions are as follows:</p>\n<ol><li><p>1. </p><p>Communicating agents show stronger scaling than independent agents. On the ARC-AGI-3 suite of games, matching the solve rate of team@5 requires 33 independent agents. The multiplier grows with team size, from for team@3 to for team@5, suggesting that the gains from communication compound with scale. In fact, the solve rate of team@ increases more rapidly with than that of best@, as seen in Figure 1.</p></li><li><p>2. </p><p>Test-time communication unlocks tasks that single agents cannot. Team@3 lifts the solve rate on the game FT09 on ARC-AGI-3 from 9.4% for a single agent to 90%. More strikingly, team@5 solves the game LP85, which remains unsolved across 64 single-agent trials, 65% of the time. Team@5 also improves the average furthest level reached over best@5 even on games that remain unsolved.</p></li><li><p>3. </p><p>Communication achieves state-of-the-art results on algorithmic and ML optimization tasks. Communication sustains progress over multi-day horizons. On Frontier-CS polyomino packing, a team of agents outperforms independent agents and surpasses the prior best-known score. On ML compression, it produces a MNIST classifier substantially smaller than those found by independent agents or the best-known human solution. Agents reach these solutions by refining one another’s discoveries and combining complementary improvements, including ideas from peers’ failed approaches.</p></li><li><p>4. </p><p>Communication benefits from verifiable progress and sufficient compute. At low budgets, team@ incurs a coordination tax that outweighs its benefits. We identify <em>verified progress sharing</em> as a mechanism through which agents build on\nsuccessive discoveries to overcome the coordination tax given enough compute.\nOn Terminal-Bench 2.0, where available feedback may not reliably rank intermediate solutions,\nteam@2 improves over a single attempt but does not outperform pass@2.\nThis helps reconcile prior results by identifying compute and verification\nas conditions that shape communication benefits.</p></li></ol>\n<h2 id=\"2-agentic-communication-at-test-time\">2 Agentic Communication at Test-time</h2>\n<div class=\"table-wrap\"><table><thead><tr><th>Benchmark</th><th>ARC-AGI-3</th><th>Polyomino packing</th><th>MNIST Compression</th></tr></thead><tbody><tr><td>Model</td><td>Sonnet 4.6</td><td>Sonnet, Opus 4.6</td><td>GPT-5.6 Sol</td></tr><tr><td>Budget per agent</td><td>per-level action budget</td><td>3-72 hours</td><td>96 hours</td></tr><tr><td>Team size</td><td>3, 5</td><td>3, 4</td><td>4</td></tr><tr><td>Metric</td><td>game solve rate</td><td>packing score</td><td>qualified bytes</td></tr><tr><td>Prior best-known result</td><td>–</td><td>0.894</td><td>2,461 B</td></tr><tr><td>Independent best@</td><td>1.4% (=3), 2.2% (=5)</td><td>0.893</td><td>3,160 B</td></tr><tr><td>Communicating team@</td><td>4.6% (=3), 8.0% (=5)</td><td>0.945</td><td>1,957 B</td></tr></tbody></table></div>\n<p>We study a minimal form of test-time communication instantiated through a shared workspace among identical agents. In each team@ trial, the harness launches CLI agents concurrently in the same task container. They use the same model, tools, task instruction, and communication prompt. Each agent has a separate model context and designated scratch directory, while all agents share the task filesystem and communication artifacts.</p>\n<p>Agents communicate directly through an append-only communication log, which acts as an asynchronous broadcast channel. Additional shared records contain adopted approaches, disconfirming evidence, and a score log of the approaches so far. Without assigned roles or a central orchestrator, the only protocol-level allocation, enforced by a synchronization primitive, is slot ownership. To ensure that agents claim distinct approaches without collision, agents race to create numbered slot directories using an atomic filesystem operation.</p>\n<p>The communication prompt (Appendix A.2) specifies how agents use these mechanisms and discourages premature convergence. After claiming a slot, each agent is asked to declare a distinct approach by considering the already-claimed slots. During execution, agents publish concise findings with timestamps and reproducible evidence whenever they make notable progress, allowing peers to reproduce or build upon their results. Results placed on the leaderboard should include measured outcomes, reproduction instructions, approach lineage, or known counterevidence. An agent is to adopt a peer’s approach only after observing a clearly better result, and even after adoption, it should preserve one meaningful variation.</p>\n<p>The protocol combines diverse exploration with evidence-based adoption to discourage convergence on poor solutions. Atomic slot ownership, together with the requirement to pursue distinct approaches, helps preserve diversity. Verifier scores give agents a basis for filtering out poorly performing approaches when deciding what to adopt. This contrasts with the task settings of Pappu et al. (2026), where agents lack verifier scores for agent solutions and can dilute expertise by compromising between expert and non-expert judgments.</p>\n<h3 id=\"2-1-experimental-setup\">2.1 Experimental Setup</h3>\n<p>We study multi-agent communication on the following three tasks, each covering a different aspect of open-ended research tasks.</p>\n<h5 id=\"arc-agi-3\">ARC-AGI-3</h5>\n<p>(ARC Prize Foundation, 2026) tests novel problem solving and interactive discovery by asking agents to solve unfamiliar grid-world games without instructions, inferring the rules, the goal, and the effect of each control from play alone. Each game comprises levels, and progress depends on carrying forward what was learned in earlier ones. This leveled structure makes breakthroughs cleanly measurable. Real-world research problems may not offer such a clean signal. A genuine conceptual advance in a domain such as approximation factors for NP-hard problems may move the reported number by only a small constant, an improvement easily lost in run-to-run variance.</p>\n<p>We evaluate on all 25 public games with Claude Sonnet 4.6 under the benchmark’s native per-level action budget and measure solve rate as the proportion of trials that clear all levels successfully. Though more recent models such as GPT-6 Astra now succeed easily on the benchmark, we focus on one model, especially one that does not saturate, to isolate the effect of test-time communication. For teams, an agent that exhausts its action budget is terminated, while the remaining agents continue. We run 64 single-agent trials per game and label a game unsolved by single agents if no trial clears all levels. For teams of three and five agents, we run 20 trials for every combination of game and team size.</p>\n<h5 id=\"frontier-cs-polyomino-packing\">Frontier-CS polyomino packing</h5>\n<p>(Mang et al., 2026) tests algorithmic optimization on an NP-hard problem. Agents write and repeatedly improve a C++17 program that packs reflected and rotated polyominoes into a minimum-area rectangle. The scorer evaluates the submitted solution on 70 hidden test cases and returns continuous partial credit in . The packing score is the mean reward across the 70 cases, and the highest valid submission in a run is retained. The best published score is 0.894, obtained by Qu et al. (2026) using four Claude Opus 4.6 agents. We test on both Claude Sonnet 4.6 and Opus 4.6 with teams of three or four agents.</p>\n<h5 id=\"mnist-classifier-compression\">MNIST Classifier Compression</h5>\n<p>tests empirical ML research. Agents must train and compress a self-contained MNIST classifier (Le Cun et al., 1998). An artifact qualifies only at 99.4% test accuracy or better and agents cannot inspect the test images, labels, or individual errors, as described in Appendix A.5. Among qualifying artifacts, only compressed size counts. We measure size using a deterministic gzip-9 compression of the submission, which includes inference code and model weights (they may be separate files or one file, depending on the agent’s design). To the best of our knowledge, the best prior result comes from the open-source model introduced by Dhairyashil R. G. (2024), which we reproduce as a 2,461-byte submission under our artifact format. We test a team of four GPT-5.6 Sol agents over a 96-hour period. Further implementation details can be found in Appendix A.</p>\n<h5 id=\"runtime-environment\">Runtime Environment.</h5>\n<p>Every agent operates through GitHub Copilot CLI (GitHub, 2026), with each trial isolated as a Harbor task (Harbor Framework Team, 2026). Single-agent trials contain one CLI agent. Team@ trials contain agents in the same task container. Agents receive the same task, tools, and base filesystem, but have private scratch directories. Within a team, agents communicate explicitly. They append timestamped claims, evidence, failures, and adoption events to shared logs, publish measured candidates to a board, and use file locks to serialize changes to a shared graded artifact. Distinct trials share neither files nor messages. Hidden evaluation data are kept outside the agent container. Polyomino submissions are sent to a separate scorer service, while MNIST submissions are evaluated by a host-owned oracle in a fresh network-disabled container. These services return evaluation feedback but never expose the underlying cases, labels, or retained program outputs.</p>\n<h5 id=\"metrics\">Metrics.</h5>\n<p>Our central comparison is the outcome of communicating agents (team@) versus the best outcome among independent agents (best@). We measure both metrics against wall-clock time and against total output tokens to account for compute. For ARC-AGI-3, a level counts as solved if any of the agents clears it, and a game counts as solved if any of the agents clears all levels. For a finite ARC solo pool with successes among trials, we compute best@ exactly, without replacement, as . We apply the same calculation level by level and average over games, giving each game equal weight. Throughout the paper, solve rate refers to the final solve rate averaged over all 25 games, and per-level rates are labeled explicitly. We measure team@ and best@ as the best-so-far packing score (mean packed-cell density across 70 cases) or MNIST classifier compression size (gzipped submission size including code and weights).</p>\n<p>Within every comparison, communicating and independent agents use the same model, maximum reasoning effort, longest context-length setting, and per-agent resource allocation. Appendix A.1 provides the communication protocol, prompts, and Copilot CLI versions.</p>\n<h2 id=\"3-experimental-results\">3 Experimental Results</h2>\n<h3 id=\"3-1-arc-agi-3\">3.1 ARC-AGI-3</h3>\n<p>| Levels from target | team@3 | best@3 | best@13 | team@5 | best@5 | best@33 | \n| 5 | 52.8 | 45.4 | 63.9 | 64.1 | 52.8 | 69.5 | \n| 4 | 28.2 | 25.8 | 46.5 | 33.6 | 32.9 | 55.4 | \n| 3 | 16.6 | 9.7 | 20.8 | 22.9 | 12.8 | 31.3 | \n| 2 | 7.2 | 5.0 | 9.3 | 16.2 | 6.1 | 15.8 | \n| 1 | 4.8 | 2.6 | 6.3 | 10.0 | 3.7 | 10.2 | \n| (Solved) 0 | 4.6 | 1.4 | 4.7 | 8.0 | 2.2 | 8.1 | </p>\n<p>Despite ARC-AGI-3’s visual simplicity, the most capable LLM-based agents fail on most games without a specialized harness. For this benchmark, we use Claude Sonnet 4.6 agents through Copilot, which reach a final success rate below 1%.</p>\n<p>Table 2 reports the percentage of runs that reach success or come within levels of it. Across 25 games, team@ outperforms best@ at every level for both and , and the advantage widens with depth, from to at . The leftmost panel of Figure 1 shows a further encouraging pattern in which team@ rises monotonically with team size and faster than best@. Matching a communicating team takes 4.3–6.6 as many independent agents, a multiplier that itself grows with , though returns may diminish beyond five agents.</p>\n<p>(a) Communication effect in ARC-AGI-3</p>\n<p>(b) Final solve rate</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Game</th><th>best@5</th><th>team@5</th></tr></thead><tbody><tr><td>LP85</td><td>0.0</td><td>65.0</td></tr><tr><td>AR25</td><td>7.8</td><td>20.0</td></tr><tr><td>DC22</td><td>0.0</td><td>0.0</td></tr><tr><td>SC25</td><td>0.0</td><td>0.0</td></tr><tr><td>KA59</td><td>0.0</td><td>0.0</td></tr><tr><td>FT09</td><td>39.9</td><td>100.0</td></tr><tr><td>WA30</td><td>0.0</td><td>0.0</td></tr><tr><td>TR87</td><td>0.0</td><td>0.0</td></tr><tr><td>M0R0</td><td>0.0</td><td>0.0</td></tr><tr><td>TN36</td><td>0.0</td><td>0.0</td></tr><tr><td>RE86</td><td>0.0</td><td>0.0</td></tr><tr><td>SB26</td><td>7.8</td><td>15.0</td></tr><tr><td>TU93</td><td>0.0</td><td>0.0</td></tr><tr><td>S5I5</td><td>0.0</td><td>0.0</td></tr><tr><td>LF52</td><td>0.0</td><td>0.0</td></tr><tr><td>G50T</td><td>0.0</td><td>0.0</td></tr><tr><td>CD82</td><td>0.0</td><td>0.0</td></tr><tr><td>SK48</td><td>0.0</td><td>0.0</td></tr><tr><td>SU15</td><td>0.0</td><td>0.0</td></tr><tr><td>BP35</td><td>0.0</td><td>0.0</td></tr><tr><td>LS20</td><td>0.0</td><td>0.0</td></tr><tr><td>VC33</td><td>0.0</td><td>0.0</td></tr><tr><td>SP80</td><td>0.0</td><td>0.0</td></tr><tr><td>CN04</td><td>0.0</td><td>0.0</td></tr><tr><td>R11L</td><td>0.0</td><td>0.0</td></tr></tbody></table></div>\n<p>Figure 2 breaks these aggregate results down by game. Across the four games solved at least once by team@5, the average solve rate rises from to . The improvement is especially striking on LP85, where 64 single-agent trials yield no successes, yet team@5 achieves a solve rate. Nor is the effect limited to the largest team. On FT09, the solve rate climbs from with best@3 to with team@3. Although these gains are concentrated, this reflects the limits of the base model rather than of communication, since most of the 25 games remain beyond Sonnet 4.6 under either setting.</p>\n<p>Moreover, zero-percent solve rates obscure meaningful gains in progress. Figure 2(a) reports the average furthest level reached. Under team@5, several never-solved games still advance by roughly a full level, showing that test-time communication produces real progress across most games. The largest gain is levels, while no loss exceeds levels. When coordination fails, the cost is negligible. When it succeeds, the gains can be substantial.</p>\n<p>Is test-time communication simply a matter of spending more tokens than independent agents do? Figure 3 paints a more nuanced picture. By the end of evaluation, team@ spends roughly 8 million output tokens with three agents and 24 million with five, nearly twice what best@ spends. However, to solve more than 1% of ARC-AGI-3 tasks, team@ reaches any given accuracy with fewer total tokens, with best@3 spending as many as team@3 and best@5 spending as many as team@5. Matching the solve rate of test-time communication is even costlier, as pools of 13 and 33 independent agents must spend as much as 3.8–4.9 the team’s tokens.</p>\n<p>We further study SB26 and LP85, two games that benefit from communication, with two action-budget ablations. We first give a single agent the native action budget, matching the total budget of standard team@5 or best@5. This configuration tests whether a longer horizon of actions of a single agent can match the performance of a team of communicating or independent agents. We then reduce team@5’s per-agent action budget to , so its total action budget matches that of a single agent.</p>\n<p>At matched total budget, communication wins. On SB26, best@5 leads early at to , but team@5 overtakes at level 2 and solves the final level in of trials against roughly for best@5 and zero for the long-horizon single agent. LP85 separates further, where team@5 holds through level 5 and finishes at while best@5 falls from to zero by level 4 and the single agent peaks at . A longer horizon alone does not substitute for coordination. This advantage, however, vanishes once the budget is reduced. At per agent, team@5 drops to zero after level 1 on LP85 and after level 3 on SB26, in both cases losing to a single agent spending the same total. Communication pays only when each agent has enough budget to explore on its own.</p>\n<h5 id=\"relative-human-action-efficiency\">Relative human action efficiency.</h5>\n<p>Throughout this paper, we take solve rate as our primary metric, since our question is whether communication and larger teams, that is, more test-time compute, lead to more tasks solved in ARC-AGI-3. The official metric, by contrast, scores a run not only by what it clears but by how economically it plays, using a metric called relative human action efficiency (RHAE) (ARC Prize Foundation, 2026). We define RHAE below and then use it to show that test-time communication improves per-agent efficiency as well.</p>\n<p>An action is a discrete interaction that changes the game state, so reasoning, tool calls, and read-only inspection are not charged. For level of game , let be the actions the agent spends and the human baseline, taken as the upper-median best action count over first-time human players. The level score is</p>\n<p>|  |  |  | (1) | </p>\n<p>with for any level the agent never completes. Game scores weight level by and are capped by the weighted fraction of levels completed,</p>\n<p>|  |  |  | (2) | </p>\n<p>where of levels are completed, and the reported RHAE is the mean of over games.</p>\n<p>Using RHAE, Figure 5 compares independent Sonnet 4.6 agents with communicating teams across all 25 games. Communication improves not only the strongest team member but also the average member. The best agent’s mean RHAE rises from for a single agent to with team@3 and with team@5, exceeding the corresponding best@3 and best@5 values of and . This gap shows that the gain is not merely the result of selecting the best outcome from more agents. The average communicating agent also improves with team size, while the average independent agent remains at (by definition). Strikingly, the average agent in team@5 matches the best of five independent agents, at versus . Together, the two panels show that communication improves both the average and the strongest team member, rather than improving team performance only by pooling more attempts.</p>\n<h3 id=\"3-2-polyomino-packing\">3.2 Polyomino Packing</h3>\n<p>We now turn to Frontier-CS (Mang et al., 2026) to test whether test-time communication yields stronger performance on challenging algorithmic tasks. We study its polyomino packing task, in which agents must pack a set of pieces into a rectangular box of minimal area. The task is NP-hard, so agents must discover good heuristics and design a packing algorithm rather than search exhaustively. The best known score of was obtained by Qu et al. (2026) using four Opus 4.6 agents.</p>\n<p>We use Claude Sonnet 4.6 and Opus 4.6 agents under the same communication protocol as in ARC-AGI-3. We first evaluate Sonnet 4.6 under Frontier-CS’s three-hour limit, with 60 single-agent trials and 20 team@3 trials. We then extend the time limit to 72 hours to study long-horizon performance and run 12 single-agent trials and 2 team@4 trials. For each method, we report the trajectory of the run that achieves the highest score. Our focus is on whether a method can discover a solution beyond the existing frontier, where a single breakthrough matters even when other attempts are unsuccessful (Yuksekgonul et al., 2026).</p>\n<h5 id=\"official-frontier-cs-results\">Official Frontier-CS results.</h5>\n<p>Figure 6 shows that both runs reach roughly within the first hour, but the team breaks away at about 1.6 hours and continues improving as the single agent plateaus. The final gap, versus , more than halves the mean fraction of unused area, from to , and carries a team of Sonnet 4.6 agents past the prior best of . The team’s advantage grows as the single agent’s progress slows, resembling the larger gap between team and independent agents at deeper levels of ARC-AGI-3.</p>\n<h5 id=\"long-horizon-results\">Long-horizon results.</h5>\n<p>Figures 7 and 8 show that the gap survives this extension. Independent agents do improve further, but their best scores settle at and , still below the prior best. The teams take lasting leads at roughly two and four hours, respectively, and continue to and . (The fact that team scores are lower than 0.945 is likely due to the small number of trials.) Thus, the previous short timeframe does not fully explain the single-agent plateau. These runs demonstrate that communication also outperforms independent agents even in long-horizon settings, as long as communicating agents have enough time and compute to explore and share their discoveries.</p>\n<h4 id=\"3-2-1-qualitative-case-study-on-communication\">3.2.1 Qualitative Case Study on Communication</h4>\n<p>In this subsection, we examine the successful three-hour Sonnet run that set a new state-of-the-art score of to qualitatively understand and assess how good communication leads to breakthroughs. With the help of Opus 5, we analyze the run’s traces and identify the key discoveries that led to the final score. We follow the numbered discoveries in Figure 6, using the agent labels shown in the figure.</p>\n<h5 id=\"from-compact-placement-to-filling-gaps\">From compact placement to filling gaps.</h5>\n<p>Agent a1 begins with <em>shelf packing</em>, arranging pieces in rows, then\n<em>skyline packing</em>, placing them along the upper contour of the occupied region.\nAgent a2’s <em>bottom-left packing</em> searches for the lowest legal placement,\nbreaking ties toward the left. These methods successively lead the run at markers\n1–3 in Figure 6, but stall near .\nKeeping placements low still leaves enclosed gaps that later pieces cannot fill.\nAgent a1 identifies this limitation and proposes <em>contact maximization</em>,\nfavoring placements that touch more already-occupied cells. Contact serves as a local\nmeasure of fit, encouraging pieces to nest together rather than merely keeping the\ncurrent height small.</p>\n<h5 id=\"building-upon-another-agent-s-failed-idea\">Building upon another agent’s failed idea.</h5>\n<p>The proposal for contact maximization appears at marker 4, but\na1’s first implementation is slow and\nthe score barely changes for roughly half an hour. Agent a2\nthen rebuilds the method from a1’s\nmessage in the communication logs without reading its code, describing the rule\nas counting “how many adjacent cells are already occupied for each piece that fits”\nand choosing the highest-contact candidate. This independent implementation and continued\nefficiency improvements raise\nthe score to nearly  after marker 5.\nAgent a2 then evaluates every piece orientation within\nthe placement lookahead, comparing how different orientations fit against\nthe existing packing. Running this <em>all-orientation search</em> at full strength\nproduces the decisive jump past  at marker 6.</p>\n<p>The difficulty is making that search affordable within the evaluator’s two-second limit per instance. Examining more orientations improves placement quality but risks a timeout, while conservative search budgets leave useful candidates unexplored. The original traces show other teams considering the same search but struggling with timeouts and keeping its budget too small. The strongest single agent also explores contact maximization without reaching the same quality in three hours. The successful transfer therefore includes substantial implementation work. Agent a2 takes a shared scoring idea and makes a more thorough search practical under the execution limit.</p>\n<h5 id=\"extending-the-shared-objective\">Extending the shared objective.</h5>\n<p>With the run above , a3 adds a <em>boundary bonus</em> at marker 7.\nThe contact score now counts box walls as well as occupied cells, rewarding pieces\nthat fit snugly against edges and into corners. This extends the same objective\ndeveloped by a1 and implemented by a2, and subsequent refinements\nbring the run to . The sequence makes the dependence between contributions\nconcrete. One agent identifies a better placement criterion, another turns it into\nan effective search, and a third improves the criterion on top of that implementation.</p>\n<p>Figure 9 illustrates the resulting improvement on one 172-piece instance. The team packs it at density, compared with for an illustrative single-agent solution, leaving far fewer gaps between pieces. This single-agent run averages across the benchmark and is distinct from the strongest run in Figure 6.</p>\n<h5 id=\"related-exchanges-over-longer-horizons\">Related exchanges over longer horizons.</h5>\n<p>In the Opus run (Figure 7), a1 introduces\n<em>best-fit selection</em>, comparing remaining pieces and candidate positions before\ncommitting the best placement. Agent a2 improves it and reduces its cost by\nrestricting the candidate pool to the largest piece sizes. Agent a1 spends\nthe saved time searching more box widths, carrying the score past .\nA later contribution from a2 changes how height and contact interact.\nInstead of considering contact only after minimizing height, it scores their\nweighted sum, allowing a slightly taller placement if it gains enough contact.\nThe other agents incorporate the change within an hour. Here, one agent’s speed\nimprovement enables another’s broader search, and the final refinement relaxes a\nplacement rule that earlier methods treated as fixed.</p>\n<p>The Sonnet run (Figure 8) again develops through agents that did not supply the initial solution. After a1’s skyline and a2’s bottom-left methods, a3 introduces best-fit selection to minimize the increase in box height. It then breaks equal-height ties by wasted space, producing the sharp jump near four hours, and adds tall-first sampling to pass . Agent a4 subsequently extends a2’s width heuristic. When a1 identifies unproductive parts of the deterministic width sweep, a4 reallocates that time to randomized restarts. As in the three-hour case, progress comes from successive changes to a shared algorithm, including improvements to where it spends its limited search time.</p>\n<h3 id=\"3-3-mnist-classifier-compression\">3.3 MNIST Classifier Compression</h3>\n<p>We next move from abstract games and competitive programming style problems to a more realistic ML research engineering task. The main question here is whether communication benefits extend to realistic multi-day engineering and how well communicating agents perform relative to the best-known human solutions. In MNIST classifier compression, agents must produce the smallest possible classifier that achieves at least test accuracy by minimizing the compressed size of the complete submission that includes its inference code and model weights in a 96-hour timeframe. The task combines empirical model development with deployment engineering, since agents must preserve accuracy while compressing both the learned parameters and the code needed to use them.</p>\n<p>Figure 10 plots the smallest qualifying submission found so far against wall-clock time and total output tokens, with lower values indicating better compression. In the lower compute regime, the independent agents lead. They produce submissions below 15K bytes before the team, which remains near 75K bytes until roughly the first half hour. Team@4 catches the independent frontier after about one hour and 100K output tokens, then stays ahead as further improvements steadily reduce its submission size. Communication therefore incurs an initial communication cost, but converts later compute into more effective progress.</p>\n<p>The difference becomes decisive over the full 96-hour run. The independent runs converge between roughly 3KB and 5KB, with best@4 finishing at 3,160 bytes, and none crosses the 2,461-byte best-known human result. Team@4 crosses that threshold after roughly 20 hours and two million output tokens, then continues to 1,957 bytes, which is about smaller than the best-known human submission. Because the advantage persists when progress is plotted against total output tokens, it cannot be explained by the team merely generating more tokens. Instead, test-time communication allows the four agents’ work to be more effective by diversifying their approaches, sharing ideas, and building on each other’s progress. While it is certainly possible that a stronger compression could be found given more time or trials, the result is a new state-of-the-art in MNIST classifier compression.</p>\n<h4 id=\"3-3-1-qualitative-study-on-communication\">3.3.1 Qualitative Study on Communication</h4>\n<h5 id=\"finding-a-model-worth-sharing\">Finding a model worth sharing.</h5>\n<p>The four agents begin with different bets. Agent a1 tries digit prototypes, a2 fixed gradient features, a3 spectral models and then binary networks, and a4 distilled CNNs. After its prototype classifier misses the accuracy threshold, a1 changes the compression strategy. A recurrent CNN reuses one convolutional cell nine times, buying depth without storing a fresh set of filters at every step. Its 5,329-byte qualifying submission persuades the other three agents to adopt this architecture within ten minutes. Their separate searches now improve different parts of a shared model.</p>\n<h5 id=\"giving-lost-features-a-way-around-the-bottleneck\">Giving lost features a way around the bottleneck.</h5>\n<p>a3 targets the final scoring layer, also called the classification head. Its matrix converts 128 features into ten digit scores using 1,280 learned parameters, plus ten bias parameters. It first compresses the 128 input features into eight learned weighted sums before computing the ten digit scores. This stalls at validation accuracy. To recover information lost in that compression, it also sends selected original features directly to the final scoring layer. Figure 11(b) shows these two paths in red and teal. With five original features selected using their covariance on the training data, the model reaches even before further fitting. The resulting 3,120-byte model qualifies, and a1 explicitly adopts the design. a4 subsequently combines seven learned weighted sums with seven directly selected features at 2,790 bytes. Passing original features to the scoring layer helps recover accuracy as the learned matrix shrinks.</p>\n<h5 id=\"learning-that-fewer-parameters-can-cost-more-bytes\">Learning that fewer parameters can cost more bytes.</h5>\n<p>Later, a2 reduces the learned weighted sums to six on the improved backbone. Additional directly selected features recover accuracy, but the new parameters initially require more compressed bytes than those of the larger head they replace. a2 finds that their values are less repetitive. It trains the projection parameters using only the stored integers , multiplied by a learned scale factor for each weighted sum. With ten directly selected features, this produces a qualifying 1,983-byte submission. Reducing the number of parameters works together with restricting their stored values.</p>\n<h5 id=\"rescuing-a-branch-that-has-fallen-behind\">Rescuing a branch that has fallen behind.</h5>\n<p>Meanwhile, a1 makes the last two spatial stages share their channel offsets. Its 1,993-byte model qualifies, but the shared best has already reached 1,970 bytes. Rather than discard the branch, it combines this backbone with a2’s design with six learned weighted sums, using routines originally written by a1 and extended by a2. Variants that pass ten or eleven original features directly to the scoring layer fail the test threshold. Passing twelve succeeds while keeping the projection parameters restricted to the same five stored integer values, and a1 advances the combined model to 1,960 and then 1,959 bytes. Figure 11(a) shows the adoption statements connecting these two branches. a2 then reorders the directly selected features and their matching scoring parameters to reach 1,957 bytes without changing predictions. The final architecture emerges from joining two branches, including one that could no longer win on its own.</p>\n<h4 id=\"3-3-2-the-1-957-byte-mnist-classifier\">3.3.2 The 1,957-Byte MNIST Classifier</h4>\n<h5 id=\"architecture\">Architecture.</h5>\n<p>Figure 11(b) shows a recurrent convolutional network with 2,900 quantized parameters and 30 scale factors. The input is cyclically shifted upward by one pixel. A convolution maps the single input channel to 32 channels, followed by SiLU to form . Nine residual updates then reuse the same depthwise filters and pointwise filters . For update , the feature maps change according to</p>\n<p>|  |  |  | (3) | </p>\n<p>Here filters each channel separately, mixes channels, and is GroupNorm with four groups, no learned affine parameters, and . Each contains 32 learned gains, broadcast over spatial positions. The offset is the 32-parameter vector for updates 1–3 and the shared vector for updates 4–9. The shortcut adds the unchanged input to the learned correction, making these updates residual. All convolutions have stride one and no bias, with padding one for filters. After updates 3 and 6, max-pooling with stride two forms from . Otherwise . This gives three updates at each of , , and resolution.</p>\n<p>After update 9, average pooling with stride three produces features, flattened in channel-first order into . The scoring layer combines six learned weighted sums with twelve directly selected features ,</p>\n<p>|  |  |  | (4) | </p>\n<p>The ordered, zero-based indices are . These features are selected greedily using training feature statistics to recover scoring information lost in the six-component approximation, and their indices remain fixed for every image at inference. Neither scoring matrix has a bias, and the largest score determines the digit. The convolutional network has 1,952 parameters and the scoring matrices have 948. Filter reuse supplies depth without storing nine separate filter sets, while the gains and offsets allow the updates to behave differently.</p>\n<h5 id=\"training-for-compression\">Training for compression.</h5>\n<p>a2’s smaller scoring layer initially costs more compressed bytes, motivating a change in how its weights are trained. Quantization-aware training rounds weights during each forward pass while learning the underlying weights and scales, restricting the 768 projection parameters to scaled integers . A small weight penalty encourages small values alongside the classification and teacher-matching losses. This makes the smaller scoring layer competitive by favoring repeated integer values by exploiting redundancy in the learned weights. Across the final model, of stored integers are or , measured before applying scales. The compressed integer payload occupies 1,262 bytes, compared with 1,390 bytes for fixed-width packing of the same values using each tensor’s range.</p>\n<p>The human reference, tiny_MNIST (Dhairyashil R. G., 2024), uses five separate convolutions with widths , BatchNorm, ReLU, and a dense -to- head. It has 3,130 parameters before BatchNorm folding. Our reproduction folds BatchNorm into the convolutions and applies per-channel 4-bit quantization, yielding 2,461 bytes at accuracy. The agents’ solution is 504 bytes () smaller and achieves test accuracy.</p>\n<h3 id=\"3-4-considerations-for-effective-multi-agent-communication\">3.4 Considerations for Effective Multi-Agent Communication</h3>\n<p>Prior work finds that tasks admitting decomposition into independent subtasks are\nparticularly amenable to multi-agent collaboration (Kim et al., 2025; Gu et al., 2025).\nOur tasks admit no such decomposition, yet communication helps regardless.\nWhat makes it help is <em>verified progress sharing</em>. Once a discovery is confirmed to be\nuseful, every agent can build on it rather than rediscover it independently, and no agent is\nleft exploring a direction the others have already exhausted.</p>\n<p>In the following discussion, we provide a simple pedagogical model that formalizes this intuition. We show how it can lead to an exponential separation between communicating and independent agents, and also provide a small-scale experiment on Terminal-Bench 2.0 (Merrill et al., 2026) where team@ improves over single-agent attempts but not independent best@, illustrating the importance of the availability of an agent-accessible verifier.</p>\n<h5 id=\"from-independent-trajectories-to-cumulative-progress\">From independent trajectories to cumulative progress.</h5>\n<p>Consider a task requiring successive improvements, with an accessible verifier reporting score after improvement . Assume that continuation difficulty depends only on the stage, improvements can be transferred freely, and agents conduct fresh, independent searches after each transfer. Both methods can query the verifier. Writing for agent ’s search time at stage , their completion times would be</p>\n<p>Thus . Independent sampling needs one agent\nto complete the entire sequence quickly, whereas a team can use a different\nagent’s breakthrough at each stage. Communication changes a *minimum of\nsums* into a <em>sum of minima</em>.\nFigure 12 illustrates this effect visually.</p>\n<p>Suppose , which is the standard constant-rate, memoryless model used in queueing theory (Shortle et al., 2018). Here we use this model as a pedagogical example, not an empirical claim about agents. The first of independent discoveries arrives at rate , giving</p>\n<p>The Erlang law (a Gamma distribution with integer shape) follows from summing exponential waiting times. Doubling therefore halves the team’s mean completion time.</p>\n<p>Now give each agent runtime , where . This is shorter than one agent’s mean completion time but longer than the team’s. Both methods receive the same agent-time budget.</p>\n<h6 id=\"proposition-1-exponential-separation-from-progress-sharing\">Proposition 1 (Exponential separation from progress sharing).</h6>\n<p>Under this model, writing ,</p>\n<p>For fixed , team@ approaches one exponentially in , while best@ success approaches zero exponentially.</p>\n<p>The separation concerns completion probability at matched runtime, not an\nexponential expected-time speedup. The proof is in\nAppendix B.\nThe benefit also depends on preserving diverse continuations.\nSometimes, communicating agents may converge on a subset of approaches,\nwhich can reduce diversity and slow progress.\nTo model <em>herding</em>, we can imagine the  agents acting as essentially\n independent groups whose members duplicate the same search.\nThe team’s discovery rate becomes \nand its mean completion time becomes . The high-success regime now\nrequires . Sharing progress helps only to the extent that independent\nsearch survives after sharing.</p>\n<h5 id=\"ineffective-communication-in-terminal-bench\">Ineffective communication in Terminal-Bench.</h5>\n<p>Terminal-Bench 2.0 evaluates agents on hard, realistic tasks performed in command-line environments (Merrill et al., 2026). We evaluate Claude Sonnet 4.6 through the GitHub Copilot CLI on the complete set of 89 tasks. Test-time communication consists of two independent team@2 trials, where both agents contribute to one final container state. The baseline consists of four independent single-agent trials. This is a small-scale comparison with only a handful of trials, so we treat it as descriptive evidence rather than a precise estimate of the gap or its cause.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Mean Accuracy</th><th>Max Accuracy</th></tr></thead><tbody><tr><td>Single agent pass@1</td><td>52.53%</td><td>–</td></tr><tr><td>Independent pass@2</td><td>62.36%</td><td>64.04%</td></tr><tr><td>Communicating team@2</td><td>60.67%</td><td>61.80%</td></tr></tbody></table></div>\n<p>As seen in Table 3, communication improves over an individual attempt, but does not outperform independent sampling (pass@) with the same number of agents. This result can be partially explained by Proposition 1 when is effectively . In Terminal-Bench, tasks require long sequences of actions, but the official verifier runs after agent execution. Available feedback varies across tasks, from package tests and direct performance measurements to checks of selected requirements. For example, the public evaluator in largest-eigenval checks the eigenvector equation and reports timings without verifying that the eigenvalue is dominant. In mailman, the public evaluator checks subscription but omits announcement delivery and unsubscription, which the official grader also tests. Such checks can confirm progress on individual requirements without establishing overall correctness.</p>\n<p>In contrast, the packing and MNIST tasks expose numerical scores that agents can query repeatedly, while ARC-AGI-3 supplies level-success feedback. Thus, setting in Proposition 1 gives us , and therefore shows no communication advantage, even when the final outcome can be graded objectively. What matters is not evaluation alone, but dense faithful verification at test-time.</p>\n<p>A secondary disadvantage may be the loss of diversity. If communication causes agents to behave like only independent groups, then end-to-end candidates cover only , below pass@’s , where denotes the success probability of an independent candidate. This penalty is especially relevant here because team@2 contributes one shared final state, while pass@2 retains two isolated states and receives oracle selection after evaluation. Premature convergence or interference in the shared state can therefore make communication slightly worse than independent sampling. Together, the results suggest that communication is most useful when feedback is accessible and discriminative, though we do not exclude the possibility that test-time communication could be useful when feedback is sparse, if the agents can construct their own intermediate verification signals or have better communication protocol harnesses.</p>\n<h2 id=\"4-related-work\">4 Related Work</h2>\n<h5 id=\"multi-agent-systems\">Multi-agent systems.</h5>\n<p>Multi-agent systems (MAS) are becoming increasingly popular and widely adopted.\nAs discussed in the introduction, some prior work (Du et al., 2024; Chen et al., 2024)\nreports gains from multi-agent debate and belief exchange.\nHowever, deliberation without a verifier or external feedback can amplify\ncorrelated errors or move the team away from its strongest member\n(Smit et al., 2024; Pappu et al., 2026).\nIn fact, independent sampling followed by aggregation remains a strong baseline\n(Wang et al., 2023; Li et al., 2024; Choi et al., 2025).\nBenefits also depend on task structure and model strength.\nKim et al. (2025) find stronger gains on decomposable tasks,\nbut these gains diminish as the underlying model becomes stronger with respect to the task.\nWhile Kim et al. (2025) comprehensively cover a range of tasks and configurations,\nour work studies the benefits of MAS in a more open-ended discovery setting with a clear verifier signal.\nAs discussed in Section 3.4, it is\n<em>verified progress sharing</em> that allows us to avoid “agent collapse”\nthat previous works have observed.\nFor more detailed surveys of MAS, we refer to Ferrag et al. (2026); Tran et al. (2025).</p>\n<h5 id=\"test-time-discovery\">Test-time discovery.</h5>\n<p>Test-time discovery often uses LLMs (or agents) to generate, evaluate, and iteratively improve candidate solutions, rather than producing an answer in a single attempt from a pretrained model. Within this paradigm, Novikov et al. (2025) keep the underlying LLMs frozen and accumulate progress through evolutionary program search, evaluator feedback, and a database of generated attempts. Wang et al. (2026); Yuksekgonul et al. (2026), by contrast, train the model via reinforcement learning during discovery, updating the LLM using feedback obtained from the target problems. Recent work also focuses on custom agentic workflows, where agents are assigned specialized roles to generate, critique, and refine hypotheses, often orchestrated by a central agent (Gottweis et al., 2026; Yamada et al., 2025). Our setting, however, is orthogonal to the works above as they focus on either single-agent capabilities or multi-agent workflows without establishing an advantage over single-agent systems under matched compute.</p>\n<p>Recent work illustrates the potential of multi-agent communication for discovery. Anthropic (2026a) describe an improved attack on HAWK whose key idea emerged through an exchange between two agents, while OpenAI (2026a) report a Navier–Stokes result produced by a group of roughly 10,000 concurrent agents. Similarly, Bianchi et al. (2026) explore mathematical discoveries by an open community of heterogeneous agents sharing solutions and encouraging public discussion. These are systems in which agents choose how to explore and communicate without a fixed workflow.</p>\n<p>These findings motivate controlled tests of whether communication improves on independent search, especially under matched compute. Anthropic (2026b) report more vulnerabilities from coordinating agents than from independent agents, though Claude Mythos Preview uses comparable tokens per vulnerability when both methods are evaluated on the same search scope. On other tasks, they observe failures to integrate agents’ work, conformity, and failures to share decisive evidence. Closely related to our setup is Qu et al. (2026), whose agents share discoveries through persistent memory and outperform the best of four independent runs under matched wall-clock budgets. In contrast, we diagnose clear conditions for when test-time communication is advantageous, through compute-controlled studies of scaling coordinating agents on ARC-AGI-3 and long-horizon algorithm/ML tasks, including negative results at low-compute budgets and on Terminal-Bench 2.0.</p>\n<h5 id=\"harness-optimization-and-orchestration\">Harness optimization and orchestration.</h5>\n<p>One line of work aims to learn the ideal topology and communication protocol of MAS. Agent graphs and communication links can be optimized directly (Zhuge et al., 2024; Zhang et al., 2025c), while architecture search can adapt the system to the task (Zhang et al., 2025b; Yun et al., 2026). A second line controls when and which agents participate. Systems can select teams or gate communication (Liu et al., 2024; Ghosh and Chakraborty, 2026), regulate interaction to avoid harmful scaling (Wang et al., 2025; Shao et al., 2026), or add agents dynamically and decompose work into explicit subtasks (Costa, 2026; Gu et al., 2025). Gu et al. (2025) investigate agent diversity and observe weaker returns from homogeneous teams, though heterogeneity may bring down the performance of the best model in the team even when explicitly told who the best is (Pappu et al., 2026).</p>\n<p>Hierarchical systems place these decisions with a central coordinator. Coordinators can be trained to assign work to frozen agents (Xu et al., 2026; Nielsen et al., 2026), while adaptive hierarchies can delegate work and synthesize results, primarily for parallel decomposition (Tang et al., 2026; OpenAI, 2026b). Recursive delegation has also been used for long-context inference rather than multi-agent orchestration (Zhang et al., 2025a). Harness optimization is broader and need not involve multiple agents. Executable or prompt-level agent loops can be optimized from scores, traces, or revision feedback (Lee et al., 2026b; Lee et al., 2026a). Prompt optimization has also been studied for multi-agent protocols (Bai and Shi, 2026).</p>\n<p>Unlike these approaches, we do not study orchestration or multi-agent harness optimization. Instead, we use a fixed shared-workspace scaffold of identical, unassigned peers and compare its performance with compute-matched independent runs to isolate communication from parallel exploration.</p>\n<h5 id=\"agent-coordination-as-interactive-systems\">Agent coordination as interactive systems.</h5>\n<p>We view sequential Monte Carlo as an interacting-population counterpart to best-of-, with partial solutions repeatedly scored and resampled (Gordon et al., 1993; Liu and Chen, 1998; Moral, 2004). Populations degenerate over long horizons (Kong et al., 1994; Doucet and Johansen, 2011), while resampling concentrates shared ancestors, and harder problems require larger populations (Liu and Chen, 1998; Snyder et al., 2008). These phenomena correspond to our gains at depth, herding failures, and non-monotone returns to team size.</p>\n<p>Related methods now guide language model inference. Particle methods support probabilistic inference and test-time scaling (Zhao et al., 2024; Puri et al., 2025), as well as constrained generation (Loula et al., 2025). Interacting particles have also been adapted to diffusion models (Singhal et al., 2025; Luo et al., 2026), with related path selection for diffusion language model decoding (Lee et al., 2025). The analogy is conceptual and suggests that coordination is most useful over long horizons with informative intermediate feedback and a diverse population.</p>\n<h2 id=\"5-conclusion-and-future-directions\">5 Conclusion and Future Directions</h2>\n<p>In this paper, we studied the capabilities of test-time communication, a multi-agent setting in which agents pursue the same open-ended objective and exchange evidence through a shared workspace. This allows agents to explore different approaches in parallel while turning individual discoveries into cumulative progress. On ARC-AGI-3, team@3 matches the solve rate of 13 independent agents and team@5 matches that of 33, with the multiplier growing from to as the team scales. Communication also solves a game that remains unsolved across single-agent trials and makes the average member of a team as action-efficient as the best of independent agents. The same pattern emerges in longer-horizon tasks. Team@ reaches a packing score of on Frontier-CS polyomino packing, above the prior best-known score of , and produce a 1,957-byte MNIST classifier, improving on both the best@4 result and the 2,461-byte best-known human solution. These gains emerge only after an initial coordination cost, but with adequate compute, communication can outperform independent agents in both performance and efficiency.</p>\n<p>The agent-accessible verifier in each task is central to our experimental design. Our setting differs from multi-agent debate, which asks agents to reconcile answers without necessarily grounding the exchange in environmental feedback. Level completion, packing score, and compression size give agents an objective basis for rejecting failures, comparing approaches, and building on improvements. Whether communication produces similar gains when feedback is sparse, noisy, or subjective remains an important open question.</p>\n<p>On the other hand, multi-agent harnesses may matter as much as single-agent harnesses. Our controlled design is deliberately narrow. We fix one communication harness and use homogeneous agents with identical instructions and no assigned roles or central orchestrator. This isolates communication, but it does not imply that roles, model diversity, communication topology, or orchestration are unimportant. Future work should consider how agents communicate, which peers receive their messages, or how teams should be organized. These policies could be engineered or evolved, similarly to harness optimization (Lee et al., 2026b; Lee et al., 2026a). Our results provide a baseline for future work, and hopefully motivate further study of multi-agent collaboration and test-time communication.</p>\n<h4 id=\"acknowledgments\">Acknowledgments</h4>\n<p>We thank Microsoft Research AI Frontiers for supporting this work and Ziyang Cai for insightful discussions during the project. This work was done during the first author’s internship at Microsoft Research.</p>\n<h2 id=\"references\">References</h2>\n<ul><li>Discovering Cryptographic Weaknesses with Claude. Note: Research report External Links: Link Cited by: §4.</li><li>Patterns and Problems in Emerging Multiagent Systems. Note: Research report External Links: Link Cited by: §1, §4.</li><li>ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence. Note: Technical report External Links: Link Cited by: §1, §2.1, §3.1.</li><li>MAS-PromptBench: When Does Prompt Optimization Improve Multi-Agent LLM Systems?. External Links: Document, 2606.23664, Link Cited by: §4.</li><li>Harnessing the collective intelligence of AI agents in the wild for new discoveries. Note: arXiv preprint arXiv:2606.10402 External Links: 2606.10402, Document, Link Cited by: §4.</li><li>Large Language Monkeys: Scaling Inference Compute with Repeated Sampling. External Links: Document, 2407.21787, Link Cited by: §1.</li><li>MLE-bench: evaluating machine learning agents on machine learning engineering. In The Thirteenth International Conference on Learning Representations, External Links: Link Cited by: §1.</li><li>ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, pp. 7066–7085. External Links: Document, Link Cited by: §4.</li><li>Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?. External Links: Document, 2508.17536, Link Cited by: §1, §1, §4.</li><li>AgentSpawn: Adaptive Multi-Agent Collaboration Through Dynamic Spawning for Long-Horizon Code Generation. External Links: Document, 2602.07072, Link Cited by: §4.</li><li>tiny_MNIST: high accuracy with a very small model. Note: Software repository External Links: Link Cited by: §A.5, §2.1, §3.3.2.</li><li>A tutorial on particle filtering and smoothing: fifteen years later. In The Oxford Handbook of Nonlinear Filtering, D. Crisan and B. Rozovskii (Eds.), pp. 656–704. Cited by: §4.</li><li>Improving Factuality and Reasoning in Language Models through Multiagent Debate. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, pp. 11733–11763. External Links: Link Cited by: §4.</li><li>From LLM reasoning to autonomous ai agents: a comprehensive review. IEEE Access. Cited by: §4.</li><li>Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning. External Links: Document, 2607.10836, Link Cited by: §4.</li><li>GitHub Copilot CLI. Note: GitHub Docs External Links: Link Cited by: §2.1.</li><li>Novel approach to nonlinear and non-gaussian bayesian state estimation. IEE Proceedings F, Radar and Signal Processing 140 (2), pp. 107–113. External Links: Document Cited by: §4.</li><li>Accelerating scientific discovery with Co-Scientist. Nature 655, pp. 487–496. External Links: Document, Link Cited by: §4.</li><li>AgentGroupChat-V2: Divide-and-Conquer Is What LLM-Based Multi-Agent System Need. External Links: Document, 2506.15451, Link Cited by: §3.4, §4.</li><li>Harbor: a framework for evaluating and optimizing agents and models in container environments. External Links: Document, Link Cited by: §2.1.</li><li>Self-improvement in language models: the sharpening mechanism. In International Conference on Learning Representations, Vol. 2025, pp. 76687–76739. Cited by: §1.</li><li>Towards a Science of Scaling Agent Systems. External Links: Document, 2512.08296, Link Cited by: §1, §3.4, §4.</li><li>Sequential imputations and Bayesian missing data problems. Journal of the American Statistical Association 89 (425), pp. 278–288. External Links: Document Cited by: §4.</li><li>Gradient-based learning applied to document recognition. Proceedings of the IEEE 86 (11), pp. 2278–2324. External Links: Document, Link Cited by: §2.1.</li><li>Recursive Harness Self-Improvement. External Links: Document, 2607.15524, Link Cited by: §4, §5.</li><li>Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models. External Links: Document, 2511.05563, Link Cited by: §4.</li><li>Meta-Harness: End-to-End Optimization of Model Harnesses. External Links: Document, 2603.28052, Link Cited by: §4, §5.</li><li>More Agents Is All You Need. Transactions on Machine Learning Research. External Links: 2402.05120, Link Cited by: §1, §4.</li><li>Sequential Monte Carlo methods for dynamic systems. Journal of the American Statistical Association 93 (443), pp. 1032–1044. External Links: Document Cited by: §4.</li><li>A Dynamic LLM-Powered Agent Network for Task-Oriented Agent Collaboration. In Conference on Language Modeling, External Links: 2310.02170, Link Cited by: §4.</li><li>Syntactic and Semantic Control of Large Language Models via Sequential Monte Carlo. External Links: Document, 2504.13139, Link Cited by: §4.</li><li>Self-Rewarding Sequential Monte Carlo for Masked Diffusion Language Models. External Links: Document, 2602.01849, Link Cited by: §4.</li><li>FrontierCS: Evolving Challenges for Evolving Intelligence. In Proceedings of the 43rd International Conference on Machine Learning, Note: To appear External Links: Link Cited by: §1, §2.1, §3.2.</li><li>Terminal-Bench: benchmarking agents on hard, realistic tasks in command line interfaces. In The Fourteenth International Conference on Learning Representations, External Links: Link Cited by: §1, §3.4, §3.4.</li><li>Feynman-Kac formulae: genealogical and interacting particle systems with applications. Springer, New York. External Links: Document Cited by: §4.</li><li>Learning to Orchestrate Agents in Natural Language with the Conductor. In International Conference on Learning Representations, External Links: 2512.04388, Link Cited by: §4.</li><li>AlphaEvolve: a coding agent for scientific and algorithmic discovery. Note: arXiv preprint arXiv:2506.13131 External Links: 2506.13131, Document, Link Cited by: §4.</li><li>On the Navier–Stokes Millennium Prize Problem. Note: Research report External Links: Link Cited by: §4.</li><li>The Builder’s Guide to GPT-5.6. Note: Technical guide External Links: Link Cited by: §4.</li><li>Multi-Agent Teams Hold Experts Back. In Proceedings of the 43rd International Conference on Machine Learning, External Links: 2602.01011, Link Cited by: §1, §1, §2, §4, §4.</li><li>The republic of science: its political and economic theory. Minerva 1 (1), pp. 54–73. Cited by: §1.</li><li>Rollout Roulette: A Probabilistic Inference Approach to Inference-Time Scaling of LLMs Using Particle-Based Monte Carlo Methods. External Links: Document, 2502.01618, Link Cited by: §4.</li><li>CORAL: Towards Autonomous Multi-Agent Evolution for Open-Ended Discovery. In Conference on Language Modeling, External Links: 2604.01658, Link Cited by: §2.1, §3.2, §4.</li><li>MonoScale: Scaling Multi-Agent System with Monotonic Improvement. In Proceedings of the 43rd International Conference on Machine Learning, External Links: 2601.23219, Link Cited by: §4.</li><li>Fundamentals of queueing theory. 5 edition, John Wiley &amp; Sons. External Links: ISBN 9781118943526 Cited by: §3.4.</li><li>A General Framework for Inference-Time Scaling and Steering of Diffusion Models. External Links: Document, 2501.06848, Link Cited by: §4.</li><li>Should We Be Going MAD? A Look at Multi-Agent Debate Strategies for LLMs. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, pp. 45883–45905. External Links: Link Cited by: §1, §4.</li><li>Scaling LLM Test-Time Compute Optimally Can Be More Effective than Scaling Parameters for Reasoning. In International Conference on Learning Representations, External Links: 2408.03314, Link Cited by: §1.</li><li>Obstacles to high-dimensional particle filtering. Monthly Weather Review 136 (12), pp. 4629–4640. External Links: Document Cited by: §4.</li><li>The knowledge machine: how irrationality created modern science. Liveright, New York. Cited by: §1.</li><li>Sakana Fugu Technical Report. External Links: Document, 2606.21228, Link Cited by: §4.</li><li>Multi-agent collaboration mechanisms: a survey of LLMs. arXiv preprint arXiv:2501.06322. Cited by: §4.</li><li>MARS: Toward More Efficient Multi-Agent Collaboration for LLM Reasoning. External Links: Document, 2509.20502, Link Cited by: §4.</li><li>Self-Consistency Improves Chain of Thought Reasoning in Language Models. In International Conference on Learning Representations, External Links: 2203.11171, Link Cited by: §1, §4.</li><li>ThetaEvolve: test-time learning on open problems. In Forty-third International Conference on Machine Learning, External Links: Link Cited by: §4.</li><li>Can LLM Agents Really Debate? A Controlled Study of Multi-Agent Debate in Logical Reasoning. External Links: Document, 2511.07784, Link Cited by: §1.</li><li>TRINITY: An Evolved LLM Coordinator. In International Conference on Learning Representations, External Links: 2512.04695, Link Cited by: §4.</li><li>The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search. External Links: Document, 2504.08066, Link Cited by: §4.</li><li>SWE-agent: agent-computer interfaces enable automated software engineering. In Advances in Neural Information Processing Systems, Vol. 37, pp. 50528–50652. External Links: Document, Link Cited by: §1.</li><li>Learning to discover at test time. External Links: Link Cited by: §3.2, §4.</li><li>Graph-of-Agents: A Graph-Based Framework for Multi-Agent LLM Collaboration. In International Conference on Learning Representations, External Links: 2604.17148, Link Cited by: §4.</li><li>Recursive Language Models. External Links: Document, 2512.24601, Link Cited by: §4.</li><li>Multi-Agent Architecture Search via Agentic Supernet. In Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 267, pp. 75834–75852. External Links: Link Cited by: §4.</li><li>G-Designer: Architecting Multi-Agent Communication Topologies via Graph Neural Networks. In Proceedings of the 42nd International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 267, pp. 76678–76692. External Links: Link Cited by: §4.</li><li>Stop Overvaluing Multi-Agent Debate—We Must Rethink Evaluation and Embrace Model Heterogeneity. External Links: Document, 2502.08788, Link Cited by: §1.</li><li>Probabilistic Inference in Language Models via Twisted Sequential Monte Carlo. External Links: Document, 2404.17546, Link Cited by: §4.</li><li>GPTSwarm: Language Agents as Optimizable Graphs. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, pp. 62743–62767. External Links: Link Cited by: §4.</li></ul>\n<h2 id=\"appendix-a-full-experimental-setup\">Appendix A Full Experimental Setup</h2>\n<p>This section provides the communication-prompt and implementation details deferred from Section 2.1, followed by task-specific evaluation details for the experiments reported in the main text. The tasks, prompts, communication protocol, and further details can be found in <a href=\"https://github.com/jerryjonghopark/test-time-communication\" rel=\"nofollow ugc noopener\">https://github.com/jerryjonghopark/test-time-communication</a>.</p>\n<h3 id=\"a-1-communication-prompts-runtime-and-accounting\">A.1 Communication prompts, runtime, and accounting</h3>\n<h5 id=\"communication-prompt-summary\">Communication-prompt summary.</h5>\n<p>Beyond the shared-workspace instructions in Section 2, the prompt specifies when agents may adopt a peer’s approach. Adoption requires a measured improvement, replication, an actionable alternative to a blocked approach, or final convergence. After adoption, agents are instructed to preserve a meaningful variation and record what they adopted and why.</p>\n<p>For MNIST classifier compression, agents build and validate candidates in scratch space, then recheck the current best under the lock before promoting a strictly better submission. These are instructions to the agents, rather than automatic verification or adoption by the harness.</p>\n<h5 id=\"runtime-versions\">Runtime versions.</h5>\n<p>The ARC-AGI-3 and Terminal-Bench 2.0 experiments use GitHub Copilot CLI 1.0.54, the long-horizon packing experiments use version 1.0.70, and MNIST classifier compression uses version 1.0.78.</p>\n<h5 id=\"shared-and-per-agent-resources\">Shared and per-agent resources.</h5>\n<p>CPU and host-memory limits apply to the entire task container. All workers in a team share those limits, rather than each receiving a separate container-sized allocation. Independent trials each have their own container. The configured allocations are shown below.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Task</th><th>Solo container</th><th>Team container</th><th>Task GPU</th></tr></thead><tbody><tr><td>ARC-AGI-3</td><td>4 CPUs, 8 GiB</td><td>4 CPUs, 8 GiB</td><td>None</td></tr><tr><td>Polyomino packing</td><td>2 CPUs, 6 GiB</td><td>2 CPUs, 6 GiB</td><td>None</td></tr><tr><td>MNIST compression</td><td>3 CPUs, 8 GiB</td><td>12 CPUs, 32 GiB ()</td><td>A100</td></tr><tr><td>Terminal-Bench 2.0</td><td>Task-defined</td><td>Task-defined, shared</td><td>Task-defined</td></tr></tbody></table></div>\n<p>The ARC and packing allocations do not scale with team size. MNIST scales the container’s CPU and host-memory limits by four for team@4, but does not enforce separate per-worker host-memory partitions. Its CPU thread pools are set to three threads per worker. Each agent is instructed to use at most 1,000 MiB of aggregate GPU memory across all its processes. A 400-MiB PyTorch allocator setting leaves headroom for CUDA context and library memory, but is not a hard limit on total GPU memory. The per-agent GPU-memory rule is self-enforced. These task-compute resources are distinct from model inference through the Copilot service and from ARC’s separate per-agent action budgets.</p>\n<h3 id=\"a-2-agent-communication-prompt\">A.2 Agent Communication Prompt</h3>\n<p>Every agent in a team receives the prompt below, in addition to the task\ninstruction. The text is stored as a template, and before a run starts the\nharness substitutes the team size into <code>{n}</code> and <code>{nm1}</code> and the\nshared-scratch paths into the remaining braced fields. Every agent in a trial\nthen receives the identical filled-in text, so the only asymmetry between\nagents comes from the slot each one wins at runtime.</p>\n<pre><code>{n} agents share this container and work the same task in parallel, all with this\n    identical prompt. Search widely without herding, coordinate as you go, and keep\n    improving until time runs out.\n    Shared scratch (create on first use):\n      - slots/approaches: {slots}/\n      - findings:         {findings}\n      - disconfirmations: {disconfirm}\n      - score log:        {plateau}\n      - coordination:     {coordination}   (empty -- conventions you author)\n    Your private scratch is {base}/work-$S after you claim slot $S.\n    1. CLAIM A SLOT AND PICK A DISTINCT APPROACH:\n         mkdir -p {slots}\n         for i in $(seq 0 {nm1}); do mkdir &quot;{slots}/slot-$i&quot; 2&gt;/dev/null &amp;&amp; S=$i &amp;&amp; break; done\n         mkdir -p {base}/work-$S\n         touch {coordination}\n         echo &quot;&lt;your approach + what you will deliberately not assume&gt;&quot; &gt; &quot;{slots}/slot-$S/approach&quot;\n         cat {slots}/slot-*/approach\n       If a lower-numbered slot already took your approach, change yours. Cover a\n       different part of the search space; do not agree early.\n    2. ALWAYS BE ACTING. Never end a turn with only prose or a plan. Every turn must\n       run a command that advances or tests the work: take a real action the task\n       accepts and read back its result. Publishing, disconfirming, coordinating, and\n       pivoting are bookkeeping around real actions, never a substitute for taking one.\n    3. EVALUATE AND RECORD EVERY ATTEMPT. Score each change with the task’s own\n       scoring or feedback mechanism (find it in AGENT.md or the task instructions),\n       then log the result:\n         echo &quot;[slot $S $(date -u +%H:%M:%S)] score=&lt;score-or-progress&gt; family=&lt;approach&gt;&quot; &gt;&gt; {plateau}\n    4. SHARE WITHOUT HERDING. Append concise findings with evidence and cost to\n       {findings}, and label weak claims as weak. Spend part of your effort trying to\n       FALSIFY the leading idea or your own, recording negative results in\n       {disconfirm}. Do not write prose whose only purpose is to make peers copy you.\n    5. COORDINATE ON SHARED RESOURCES. You share the graded output, common files, and\n       the environment with peers who run blind to your session. Treat every shared\n       resource as contested: before you touch one, re-check {coordination} and the\n       resource’s current state; after, confirm your change survived and did not just\n       repeat a peer’s. A collision is any overwritten, duplicated, or conflicting work\n       that wastes effort. No coordination scheme is provided -- {coordination} is empty\n       and yours to author: on a collision, write a convention there concrete enough for\n       a peer to follow, that changes your next action, then follow it.\n    6. HIGH BAR TO CONVERGE. Keep your own approach unless another clearly beats it on\n       a measured, reproduced result, or yours is blocked, or the run is wrapping up.\n       Even then, keep one real difference (a parameter, subcase, representation, or\n       fallback) until the very end.\n    7. NEVER STOP WHILE TIME REMAINS. A working result is not the finish line; the\n       clock running out is the only acceptable reason to stop. Do not declare the\n       task done, final, solved, or &quot;at the ceiling&quot; and go idle -- a suspected\n       ceiling is a claim to disconfirm, not a reason to quit.\n    8. BREAK PLATEAUS BY CHANGING FAMILY. You are plateaued when your best score has\n       not strictly improved over 3 consecutive attempts. Then stop tuning\n       and switch to a STRUCTURALLY DIFFERENT approach -- a different core principle or\n       assumption, not a variant of the current one. Keep a short list of untried\n       families in {base}/work-$S so you always have a next one ready. Read peers’\n       approaches and {plateau} first and pick a family no active peer is on; adopting\n       a peer who is also plateaued is not progress.\n    Run a tight loop -- change -&gt; evaluate -&gt; record -&gt; repeat -- without pausing. Do\n    not use the internet, curl, wget, HTTP libraries, or secrets.</code></pre>\n<h3 id=\"a-3-arc-agi-3\">A.3 ARC-AGI-3</h3>\n<h5 id=\"task-and-interface\">Task and interface.</h5>\n<p>Agents interact with the games by receiving integer-valued grids and selecting from the available discrete controls or coordinate clicks. Each level has its own cap, so unused budget from a later level cannot rescue an agent stalled earlier. We use the benchmark-native cap of five times the corresponding human-action baseline, independently for every agent.</p>\n<p>Every team member owns a separately authenticated ARC session. Sharing an action trace does not copy state and does not clear a peer’s level. Time-dependent progress curves use the first logged completion of each level.</p>\n<h5 id=\"synchronization\">Synchronization.</h5>\n<p>Team runs synchronize at intervals of half the current level’s action budget, rounded up to an integer number of actions. The harness tracks each agent’s phase by its level and the number of these intervals it has used. An agent ahead of the slowest live teammate, either within a level or by reaching the next level, cannot take another game action until its teammates catch up or leave the live set. During this pause, agents can still inspect state, run shell commands, and exchange notes. Refused game actions consume no budget. Finished, game-over, and budget-exhausted agents do not block their teammates, and inactive sessions are excluded after an idle timeout. This synchronization applies only to communicating teams, not to best@.</p>\n<h5 id=\"closed-book-condition\">Closed-book condition.</h5>\n<p>The task container requires network egress for the model control plane, but agents receive no browser or web-search tool and are explicitly forbidden to fetch public solutions, replays, or ARC pages. Only the task prompt, local run files, the agent’s own environment observations, and teammates’ within-trial notes are admissible. We inspect run trajectories afterwards to verify that no agent accessed external ARC information.</p>\n<h3 id=\"a-4-frontier-cs-polyomino-packing\">A.4 Frontier-CS polyomino packing</h3>\n<h5 id=\"objective-and-scorer\">Objective and scorer.</h5>\n<p>Each of 70 hidden cases contains edge-connected polyominoes of one to ten cells. A program may reflect, rotate by multiples of , and translate each piece. All pieces must lie without overlap in one integer-grid rectangle. If is the total number of occupied cells and the returned area, the normalized case quality is and the reported score is its mean over the fixed cases. Invalid output causes the submission to be rejected. Candidate programs compile as GNU C++17 and receive 2 seconds and 256 MiB per case.</p>\n<p>The hidden instance directory is mounted read-only inside a separate scorer service and is never mounted into the agent container, which contains the task statement and a submission client. On each attempt, the client sends the candidate C++ source over an internal endpoint. The scorer compiles and executes it, with case-output retention disabled, and returns the aggregate score together with compact status, timing, and scoring metadata. The scorer never returns the instances or program outputs. Agents may submit repeatedly without penalty, but cannot inspect the evaluation data. Every attempt, timestamp, status, and score is recorded in a submission ledger. The final verifier retains the highest valid scored submission in the run, protecting an earlier champion from a broken final edit.</p>\n<h3 id=\"a-5-mnist-classifier-compression\">A.5 MNIST classifier compression</h3>\n<h5 id=\"data-and-artifact-contract\">Data and artifact contract.</h5>\n<p>The official 60,000-image MNIST training split is partitioned once into 55,000 training and 5,000 development images. The agent image contains only these two partitions, while the official 10,000-image test archive remains on the host and is never copied into the task container. Agents may train only on the 55,000 images, use the development set for selection, and query the sealed oracle only after a full-development accuracy of at least .</p>\n<p>An oracle request snapshots the canonical submission and passes that snapshot to a host-owned evaluator. Each request runs in a fresh container with networking disabled, a read-only root filesystem, privilege escalation disabled, and the submission mounted read-only. Within it, a trusted verifier privately shuffles the test set, makes the test archive and labels unreadable to submitted code, and invokes inference under an unprivileged user. The submission receives only batches of test images and writes predictions to isolated scratch space. The trusted verifier alone reads the labels and computes accuracy. The container is discarded after evaluation, and the oracle returns only aggregate accuracy. It never returns examples, labels, predictions, or per-example errors. Final grading uses the same isolation boundary, with a 600-second limit for classifying the complete test set.</p>\n<p>The graded directory must include a predict.py entry point and every weight, table, constant, generator, and decoder it needs. We normalize file order and metadata, form a deterministic tar archive, and apply gzip level 9. Training code outside the submission and the preinstalled numerical runtime are not charged.</p>\n<h5 id=\"open-source-reference\">Open-source reference.</h5>\n<p>The 2,461-byte reference in Section 2.1 is our benchmark-format reproduction of the open-source tiny_MNIST model [Dhairyashil R. G., 2024]. The published architecture has 3,130 trainable parameters and reports accuracy above . We retrained it from the published recipe, folded BatchNorm exactly into the adjacent convolutions, and quantized the resulting weights per output channel to four bits.</p>\n<p>For the deployable artifact, we bit-pack this state, use a compact decoder, and for fair comparison, code-golf the inference path by eliminating general training machinery, redundant metadata, whitespace, and intermediate file structure while preserving predictions. The complete normalized archive, including executable inference code, is 2,461 bytes under the same deterministic gzip-9 metric used for every submission. The reproduced artifact attains exactly on the sealed test set. Thus 2,461 bytes is a measured end-to-end submission size, not the size of the upstream training checkpoint.</p>\n<h5 id=\"environment\">Environment.</h5>\n<p>The pinned stack is Python 3.12.13, PyTorch 2.5.1 with CUDA 12.4, torchvision 0.20.1, NumPy 2.5.1, SciPy 1.18.0, and scikit-learn 1.9.0. Agents train from scratch on the provided data. Downloading additional data or pretrained weights is prohibited, and inference evaluation has no network access.</p>\n<h5 id=\"checkpoints-and-trajectory-accounting\">Checkpoints and trajectory accounting.</h5>\n<p>Agents are instructed to save each smaller submission that passes the full development-set check, together with its result, timestamp, provenance, and recomputed archive size. Saved checkpoints are audited, and the paper figures use only submissions confirmed at on the sealed test set. Figure 10 compares four independent GPT-5.6 Sol runs with one team@4 run. In the elapsed-time panel, best@4 at time is the smallest qualifying submission found by any of the four independent agents within its own first hours. The output-token panel uses the same best-so-far trajectory, placing each improvement at the sum of tokens consumed by all four independent runs by that elapsed time. Team tokens likewise sum all four workers. Neither curve counts only the tokens of the agent that produced an improvement.</p>\n<h3 id=\"a-6-terminal-bench-2-0\">A.6 Terminal-Bench 2.0</h3>\n<h5 id=\"task-environments-and-execution\">Task environments and execution.</h5>\n<p>The harness loads terminal-bench@2.0 through Harbor and runs each task in its Docker environment. CPU, host-memory, storage, and GPU requests come from the individual task configuration, rather than a uniform benchmark-wide allocation. The harness does not automatically multiply these requests by the number of workers. In a team@2 trial, both workers share the task container and its resource limits. Agent execution and verifier timeouts are configured separately by each task, so there is no single wall-clock horizon analogous to the packing or MNIST experiments. The verifier runs after agent execution and produces the trial’s reward from the resulting environment.</p>\n<h5 id=\"aggregation\">Aggregation.</h5>\n<p>For Table 3, pass@1 averages the four independent outcomes for each task. The saved independent results are grouped into two batches of two attempts. Pass@2 counts a task as successful within a batch if either attempt passes, then averages the two batch accuracies. Team@2 averages the two team-trial outcomes for each task. Unfinished tasks are assigned zero, and all 89 tasks remain in the denominator. Max Accuracy selects the higher batch or team-trial accuracy over the whole benchmark, rather than selecting a different batch or team trial for each task.</p>\n<h2 id=\"appendix-b-analysis-for-verified-progress-sharing\">Appendix B Analysis for Verified Progress Sharing</h2>\n<p>This appendix gives the calculations behind the conceptual model and Proposition 1. The mathematics is largely standard rather than new. It combines elementary facts about exponential order statistics, Erlang sums, and Chernoff bounds. We keep the model deliberately simple so that it isolates how reusable, verified progress can change the value of parallel search. It is not intended as an empirical model of agent completion times or as a general theory of multi-agent communication.</p>\n<h3 id=\"b-1-idealized-progress-sharing\">B.1 Idealized progress sharing</h3>\n<p>We first compare independent and communicating agents under the same collection of stage-level search times. Recall that is the time agent would take to discover the next improvement at stage . A communicating team uses the first discovery at every stage. Therefore, for every agent ,</p>\n<p>Taking the minimum over agents gives . This is the pointwise comparison in the main text. It captures the distinction between combining stage-level breakthroughs and selecting one complete trajectory after all runs finish.</p>\n<p>Under the modeling assumption , define the waiting time for the team’s next verified improvement as . Its survival probability is</p>\n<p>Thus each is distributed as . Independence across stages then gives</p>\n<p>For comparison, each independent agent must complete all stages itself. Its completion time satisfies</p>\n<p>These distributions provide the ingredients for the fixed-budget comparison.</p>\n<h3 id=\"b-2-completion-probability-at-a-fixed-budget\">B.2 Completion probability at a fixed budget</h3>\n<p>We now prove Proposition 1. The only additional ingredient is a standard exponential tail bound for an Erlang random variable.</p>\n<h6 id=\"proof-of-proposition-1\">Proof of Proposition 1.</h6>\n<p>Let . Its moment-generating function is</p>\n<p>For , apply Markov’s inequality with to obtain</p>\n<p>where . For , the same choice gives . Since implies , Markov’s inequality gives the corresponding lower-tail bound. Hence</p>\n<p>|  |  |  |  |  |  | (5) | </p>\n<p>Set , with . Since has the same distribution as and ,</p>\n<p>Likewise, has the same distribution as , so gives</p>\n<p>Taking a union bound over the independent agents yields</p>\n<p>Combining the two bounds proves</p>\n<p>Because for and , both exponents are strictly positive for fixed and . ∎</p>","headings":[{"level":1,"text":"Scaling Discovery through Test-Time Communication","id":"scaling-discovery-through-test-time-communication"},{"level":2,"text":"1 Introduction","id":"1-introduction"},{"level":2,"text":"2 Agentic Communication at Test-time","id":"2-agentic-communication-at-test-time"},{"level":3,"text":"2.1 Experimental Setup","id":"2-1-experimental-setup"},{"level":2,"text":"3 Experimental Results","id":"3-experimental-results"},{"level":3,"text":"3.1 ARC-AGI-3","id":"3-1-arc-agi-3"},{"level":3,"text":"3.2 Polyomino Packing","id":"3-2-polyomino-packing"},{"level":3,"text":"3.3 MNIST Classifier Compression","id":"3-3-mnist-classifier-compression"},{"level":3,"text":"3.4 Considerations for Effective Multi-Agent Communication","id":"3-4-considerations-for-effective-multi-agent-communication"},{"level":2,"text":"4 Related Work","id":"4-related-work"},{"level":2,"text":"5 Conclusion and Future Directions","id":"5-conclusion-and-future-directions"},{"level":2,"text":"References","id":"references"},{"level":2,"text":"Appendix A Full Experimental Setup","id":"appendix-a-full-experimental-setup"},{"level":3,"text":"A.1 Communication prompts, runtime, and accounting","id":"a-1-communication-prompts-runtime-and-accounting"},{"level":3,"text":"A.2 Agent Communication Prompt","id":"a-2-agent-communication-prompt"},{"level":3,"text":"A.3 ARC-AGI-3","id":"a-3-arc-agi-3"},{"level":3,"text":"A.4 Frontier-CS polyomino packing","id":"a-4-frontier-cs-polyomino-packing"},{"level":3,"text":"A.5 MNIST classifier compression","id":"a-5-mnist-classifier-compression"},{"level":3,"text":"A.6 Terminal-Bench 2.0","id":"a-6-terminal-bench-2-0"},{"level":2,"text":"Appendix B Analysis for Verified Progress Sharing","id":"appendix-b-analysis-for-verified-progress-sharing"},{"level":3,"text":"B.1 Idealized progress sharing","id":"b-1-idealized-progress-sharing"},{"level":3,"text":"B.2 Completion probability at a fixed budget","id":"b-2-completion-probability-at-a-fixed-budget"}]}}