To advancement beyond humanity's accumulated knowledge, AI itself (the optimizee) may additionally be its finest accessible optimizer and evaluator. MASS often optimizes multi-agent workflows, afterward distills company intellect from multi-agent systems and uses it as a self-supervision indication for RSI. It improves mark per output token complete the basis example on investigation benchmarks. We anticipation MASS contributes to AI’s Odyssean moment: continual self-learning in a environment anywhere the optimizee is the evaluator, and the optimizer is the optimizee.
Estimated study time: 15 min
† Work done during an internship at Sakana AI.
We desire frontier AI models to obtain us beyond the build of cognition humanity has built complete generations. At that frontier, equal individual experts may battle to measure results and guide additional progress, particularly on non-verifiable tasks. For example, evaluating a projected finding may necessitate scarce expertise, costly evidence, or experiments that obtain far longer than generating the proposal. The example itself may afterward inevitably be the most capable supervisor accessible and need to fig out for itself how to improve. This motivates homogeneous recursive self-improvement (RSI), in which the example being improved (optimizee) additionally serves as its own optimizer and evaluator. The chief inquiry is how to get helpful self-supervision inside this homogeneous loop.
Human societies recommendation a helpful difference for addressing this question. Across foraging communities, agriculture settlements, manufacturing societies, and earth data networks, group have expanded their capabilities in part by dividing work, combining ideas, and passing on what they learn. Collective achievements rotate into a starting item for additional learning. In The Structure of Scientific Revolutions, Thomas Kuhn emphasized the shared examples through which scientists study to acknowledge and resolve problems. A specialized community's achievements thus rotate into part of how its members study to do new work.
Inspired by this pattern, we ask whether a example can study from its own team's company intelligence, which an delegate operating solitary may not reliably exhibit.
We current Multi-Agent Self-Supervision (MASS) to rotate this idea into a learning loop. MASS uses the identical example as the executor, optimizer, and evaluator. The optimizer archetypal searches for the finest squad workflow using feedback from the evaluator, afterward trains the administrator on executions from the self-judged finest team. The updated example afterward fills all three roles, and the procedure repeats. Therefore, MASS provides a self-supervision indication for RSI by distilling the team's company cognition into the shared model.
Homogeneous RSI
We study homogeneous RSI, in which one shared tongue example serves as the task-solving agent, workflow optimizer, and evaluator:
Agent
The project solver, or optimizee. Produces a workspace containing results, plots, and a investigation report.
Workflow optimizer
Uses the feedback from the evaluator to revise how agents distinct the task, toggle results, and inspect one another's work.
Evaluator
Compares two workspaces, chooses which improved satisfies the task, and explains the judgment.
We use non-verifiable investigation tasks as training data. We use non-verifiable to average that an automatic rule-based indication is not accessible and that lone a self-evaluation indication (whose correctness is unknown) guides the general learning loop.
To measure progress, we separately ask GPT-5.5 and Claude Opus 4.8 to measure the completed work. Their judgments are withheld from MASS. This gives us an exterior measure of betterment in a procedure whose supervision comes from the example itself.
How MASS works
MASS uses workflow optimization to elicit company intelligence, afterward distills the team's cognition into the shared model. The inner loop optimizes the team's workflow during the example weights remain fixed. The example evaluates and judges for itself which workflow is optimal. The outer loop afterward updates those weights by learning from executions of the optimized workflow. The interactive diagram below follows the two loops.
← rub to see the entire fig →
0. Setup: workflow as a prompt
A workflow is a written scheme for how a multi-agent squad should activity together. The orchestrator assigns activity to subagents and integrates their results. The scheme is added to the project prompt, making the team's institution accessible for the example to inspect and revise. It has four components:
- Role: the duty assigned to all subagent, specified as preparing data.
- Instruction: the activity that subagent should transport out.
- Contract: the data and artifacts it must return, including any required checks.
- Hop: the command of subagent calls, including any come back to an before step.
For example, an evolved backing workflow arranges data preparation, characteristic construction, modeling, and backtesting into a sequence, followed by an audit. Its data agreement calls for definitive checks on the timing of the data:
Evolved backing workflow, condensed from the paper
Call order Data → Features → Modeling → Backtest → Audit → Reconciliation → Report Data specialist's output contract validation_report: no_future_leakage: bool train_end_before_val_start: bool val_end_before_test_start: bool
A backtest can appearance convincing if it lets historic decisions use data from the future. This agreement requires the data expert to inspect for that error and study whether the period splits are valid. It additionally tells afterward agents what evidence they should obtain before proceeding. Part of the investigation scheme is thus expressed in what agents must established and continue on to one another.
1. Inner loop: hunt for a improved workflow
For all training task, MASS retains a champion: the finest workflow established so far according to the current model's judgments. Each repeat has three steps:
- 1Execute. The squad follows the applicant workflow and produces a workspace.
- 2Evaluate. The evaluator compares this workspace alongside the champion's workspace and explains which is better. A winning applicant becomes the new champion.
- 3Revise. The optimizer uses the feedback former to propose the next workflow.
The evaluator judges the completed work; the optimizer uses that judgement to revise the procedure that produced it. Optimizer's scheme immediate emphasizes contracts and hop, directing workflow hunt toward what data agents toggle and whenever they obtain it.
2. Outer loop: study from the selected team
Even alongside an optimized workflow, the norm of long-horizon investigation activity can change from run to run. MASS hence executes self-judged optimal workflows (or victor workflows, definition the optimal workflows established so far) alongside multiple random seeds and uses the current example to position the resulting workspaces, retaining the highest-ranked trajectories. Each trajectory records the messages, tool calls, and tool results of the orchestrator and its subagents.
- 1Run the victor workflow 18 times for all task.
- 2Use the current example to compare and position the resulting workspacesA Bradley–Terry example converts the pairwise preferences into a mark for all workspace..
- 3Retain up to 16 trajectories per task: 15 for training and one for validation.
- 4Remove the workflow content from the orchestrator's first prompt whenever constructing the training examples.
- 5Fine-tune the shared model alongside LoRA on the two orchestrator and subagent conversations. The model's outputs are the training targets; project prompts and tool results provision context.
Removing the workflow from the orchestrator's training immediate encourages the example to study how to arrange the activity from the project itself.
Results
We run two complete rounds of MASS starting from Qwen3.6-27B in the qwen-code coding-agent environment. Our synthetic investigation suite contains 12 tasks in finance, robotics, and pharmacy, all requiring code, quantitative results, and a investigation report. Nine tasks are designated for training and three, one per domain, are held out from post-training. Eight ultimately provision training dataOne training project is excluded since workflow hunt finds no workflow that strikes its reference. The project descriptions were generated by GPT-5.5 before the experiment. Within MASS, the current Qwen example supplies workflow feedback and training-data selection. The three test tasks are excluded from fine-tuning, validation, and checkpoint selection, although workflow hunt is additionally evaluated on them..
We difference three generations: the basis example \(\mathcal{L}^{(0)}\), the first-round example \(\mathcal{L}^{(1)}\), and the second-round example \(\mathcal{L}^{(2)}\). We analyze the norm of their work, their capability to arrange and fairness it, and their use of coordination without a supplied workflow.
Better project performance
On the three synthetic test tasks, the win charge against the basis rises from 53.9% following one circular to 69.9% following two rounds (Table 1). The second-round example additionally wins 60.8% of comparisons against the first-round model.
| Evaluation tasks | Round 1 vs. base | Round 2 vs. base | Round 2 vs. circular 1 |
|---|---|---|---|
| Training tasks | 65.6% | 82.5% | 68.1% |
| Test tasks | 53.9% | 69.9% | 60.8% |
To test transfer beyond the synthetic investigation suite, we measure community benchmarks. After two rounds, score per output token reaches 1.2–1.6× the basis model's level throughout MLR-Bench, DSBench, ScienceAgentBench, and AstaBench (Figure 3).
The gains are strongest on four investigation benchmarks resembling the training tasks. Score per output token is 0.96× the basis value on Terminal-Bench 2.0 and 0.94× on SWE-bench Verified, showing that transfer remains uneven. We doubtful this generalization can be effortlessly addressed by introducing application engineering tasks into the training data.
Observation 1. Role generalization in RSI
We detect that MASS enables function generalization in RSI. We train the example lone on task-solving trajectories and detect that the capabilities of the workflow optimizer and workspace evaluator additionally improve. This enables RSI iteration to be sustainable.
Figure 4 shows optimizer can discover prosperous workflows sooner as sequence goes by.
In fact, Figure 4 is technically not a extremely fair comparison, since \( L^{(0)}\) denotes a environment in which the basis example (optimizer) updates the basis example (optimizee), and \( L^{(1)}\) denotes a environment in which the first-round example (optimizer) updates the first-round example (optimizee). For a fair comparison, we provision the first-round model's workflows to the basis example (Table 2). We detect that the basis example alongside first-round hunt workflows wins 64% of comparisons against the basis example alongside base-search workflow. Each set of workflows was discovered using its generation's executor, evaluator, and optimizer. The transfer test holds implementation weights fixed during evaluation. It does not isolate the optimizer from the another roles during search. Table 2 plainly shows that optimizer's capabilty have increased.
| Executor | Workflow source | Compared with | Win rate |
|---|---|---|---|
| Round 1 | Round-1 search | Base executing base-search workflows | 80% |
| Base | Round-1 search | Base executing base-search workflows | 64% |
| Base | Round-1 search | Base without a workflow | 63% |
The evalutor additionally move nearer to those of the powerful external example judges (more particulars in paper). Agreement alongside those external judges rises from 73% to 93% throughout the three cycles. Overall, these results propose that task-solving examples can additionally enhance capabilities used to create and choose the next round's training data.
Ablation 1: Optimal workflow maximizes data flow
What really makes an optimal multi-agent workflow effective?
We analyze the roles, instructions, contracts, and hops of the victor workflows (the finest workflows established so far). For all component, we evaluation how much its content reveals concerning the task. In the backing example above, a function specified as “careful data scientist” could use to many tasks. A agreement specifying leakage checks for a selling backtest is much additional distinctiveWe evaluation mutual data between embeddings of workflow components and project descriptions, using text-embedding-3-large and all-mpnet-base-v2 alongside unadjusted and permutation-adjusted estimators. These data depict workflow text, fairly than messages exchanged during execution. The samples shield lone 8–9 tasks; agents inside a workflow are not autonomous project replications..
Figure 5 shows that between the archetypal and final iterations of workflow optimization, an optimal multi-agent workflow tends to encode project data additional powerfully in contracts and hops and small in roles and instructions. We do not display this here, but mutual requirements among the components mostly decrease too (interesting!). This form suggests that the LLM optimizer evolves the workflow components to be mutually autonomous but optimizes the data stream via contracts and hops (see this blog post for deep, intuitive explanations). I would say that workflow optimization complete discrete components resembles discovering a basis vector for that harness space.
Observation 2: Multi-agent trajectories provision a compact RSI signal
We detect that trajectories from multi-agent systems can be additional productive than trajectories from sole agents in RSI. Multi-agent implementation additionally changes the construction of the training data. The orchestrator's conversation shows how to distinct a broad project into focused assignments and coordinate their execution. A subagent's conversation shows how to transport out an idiosyncratic assignment. Specifically, we detect that jointly learning coordination from the orchestrator and bounded implementation from subagents provides additional effective supervision per training token than learning from lengthy single-agent trajectories (Figure 6).
We difference students trained independently from the identical basis example on single-agent trajectories (S, S+, S++), multi-agent trajectories (M, M+), and a blend (X). A affirmative sign denotes a larger training set; M+ is the first-round MASS model. All points below study results pooled throughout four training tasks and one test task.
Against the basis model, M+ reaches a 68.3% win charge following handling 28.8M supervised training tokens, during S++ reaches 64.0% following 42M tokens. The MASS example thus achieves the higher mark against this citation alongside 31% small supervised tokensThe first-round dataset stores 17.2M supervised tokens. Sampling alongside replacement and balancing orchestrator and subagent examples create 28.8M processed supervised tokens through the selected checkpoint. This measures visibility to training targets, including repeated examples. It excludes afterward updates, input tokens, validation, data generation, search, and evaluation..
Higher-quality trajectories may contribute to this advantage. In a distinct analysis, powerful external judges favor the multi-agent systems' activity in 85.1% of comparisons alongside single-agent outputs. Team runs additionally create concerning 2.4× as many tokens on averageAmong ten pairs alongside akin generated trace lengths, the penchant for squad outputs remains 80%. This difference matches observed lengths, not generation budgets. The groups additionally differ in workflow and verification instructions, so the difference does not isolate the consequence of delegation. These external judgments are used for this inspection lone and do not guide MASS.. But whenever we difference the supervised tokens used during SFT, we discover that multi-agent trajectories can provision a denser learning indication than single-agent trajectories.
These results propose that trajectories from multi-agent systems can be additional productive than trajectories from sole agents in RSI, extending their value beyond test-time scaling.
Observation 3. The RSI bottleneck can be the optimizer, not the evaluator
We observed that the bottleneck in RSI can be the optimizer, not the evaluator. We kept the basis example as the project solver and exchanged in a stronger external example (GPT-5.5) for the evaluator only, or for the two the evaluator and the optimizer.
Figure 7 shows that upgrading lone the evaluator helps a little, but upgrading the optimizer as fine helps a lot. The strictly homogeneous self-improvement iteration does eventually discover winning workflows for 11 of 12 tasks; it fair takes longer. Our explanation is that the powerful evaluator may provision high-quality feedback, but the feeble optimizer is not capable adequate to digest it. This suggests that, under a constricted budget, one may have to prioritize the optimizer complete the evaluator inside an RSI iteration unit.
Figure 8 additionally shows that workflow hunt is noisy. In text-space optimization, reaching a fine workflow and improving on it stably are distinct problems. This suggests that the LLM optimizer may need an equal of a learning-rate scheduler to stabilize the hunt in content space.
Conclusion
In Tennyson's Ulysses, Odysseus imagines another voyage in chase of cognition “Beyond the extreme border of individual thought.” For AI, we ideate an Odyssean moment whenever continued learning must continue without a stronger supervisor. The difficulty concerns the two who can supervise additional learning and how advancement have to be judged. Human experts may be capable to province a goal without specifying an goal function that completely captures it.
Writing is a fine example: expert writers can acknowledge powerful writing, yet battle to province their judgement in a rubric that applies reliably throughout contexts. For RSI, the investigation inquiry is how AI can rotate incomplete individual direction into helpful criteria for learning, and how to test whether those criteria continue to grasp the intended goal as its capabilities grow. This becomes particularly challenging whenever the example being improved additionally serves as its own optimizer and evaluator: the judgments that guide betterment must themselves remain open to correction.
Keeping a system's self-evaluation open to correction is additionally essential for safety. In a homogeneous loop, a shared blind place could authorize a flawed outcome to be produced, accepted as an improvement, and unified into the next model. If humans additionally battle to acknowledge the flaw, successive updates may fortify it. The safety involvement is that the identical weakness can colony the two the system's behavior and the scheme meant to accurate it.
In this sense, shared blind spots may be an Achilles' heel of recursive self-improvement. Understanding them matters for controlling a additional capable scheme since it reveals anywhere reliance on the system's own judgments is smallest justified. That cognition could guide autonomous checks, constraints on autonomous actions, and conditions for intervention. Identifying a weakness solitary does not established control; we must additionally display that we can detect its consequences and intervene efficiently as the scheme changes.
OpenAI's investigation on chain-of-thought monitorability illustrates one part of this challenge: reasoning traces can assistance disclose misbehavior, but they provision incomplete evidence. Recent inspection of a imaginable intellect detonation gives additional logic to study these conditions early. This is a chief safety inquiry for AI's Odyssean moment: how to maintain meaningful oversight whenever no stronger supervisor is available. ⛵