Recursive self-improvement through collective intelligence

Hacker News by 18 min read 36x views
Recursive self-improvement through collective intelligence

Share Post

To advancement beyond humanity's accumulated knowledge, AI itself (the optimizee) may additionally be its finest accessible optimizer and evaluator. MASS often optimizes multi-agent workflows, afterward distills company intellect from multi-agent systems and uses it as a self-supervision indication for RSI. It improves mark per output token complete the basis example on investigation benchmarks. We anticipation MASS contributes to AI’s Odyssean moment: continual self-learning in a environment anywhere the optimizee is the evaluator, and the optimizer is the optimizee.

Estimated study time: 15 min

† Work done during an internship at Sakana AI.

 workers set category on the left, run a media on the right, and lay out printed sheets for proofreading in the foreground.
Figure 1. From shared activity to shared knowledge. This engraving shows typesetting, printing, and proofreading as parts of a coordinated process. Books transport the outcome beyond the workshop, anywhere others can study from it and build on it. MASS draws on a connected idea: a team's company activity can provision training examples for the example its members share. Jan Collaert I, following Stradanus, The Invention of Book Printing, c. 1600. The Metropolitan Museum of Art, Harris Brisbane Dick Fund, 1934. Public domain (CC0).

We desire frontier AI models to obtain us beyond the build of cognition humanity has built complete generations. At that frontier, equal individual experts may battle to measure results and guide additional progress, particularly on non-verifiable tasks. For example, evaluating a projected finding may necessitate scarce expertise, costly evidence, or experiments that obtain far longer than generating the proposal. The example itself may afterward inevitably be the most capable supervisor accessible and need to fig out for itself how to improve. This motivates homogeneous recursive self-improvement (RSI), in which the example being improved (optimizee) additionally serves as its own optimizer and evaluator. The chief inquiry is how to get helpful self-supervision inside this homogeneous loop.

Human societies recommendation a helpful difference for addressing this question. Across foraging communities, agriculture settlements, manufacturing societies, and earth data networks, group have expanded their capabilities in part by dividing work, combining ideas, and passing on what they learn. Collective achievements rotate into a starting item for additional learning. In The Structure of Scientific Revolutions, Thomas Kuhn emphasized the shared examples through which scientists study to acknowledge and resolve problems. A specialized community's achievements thus rotate into part of how its members study to do new work.

Inspired by this pattern, we ask whether a example can study from its own team's company intelligence, which an delegate operating solitary may not reliably exhibit.

We current Multi-Agent Self-Supervision (MASS) to rotate this idea into a learning loop. MASS uses the identical example as the executor, optimizer, and evaluator. The optimizer archetypal searches for the finest squad workflow using feedback from the evaluator, afterward trains the administrator on executions from the self-judged finest team. The updated example afterward fills all three roles, and the procedure repeats. Therefore, MASS provides a self-supervision indication for RSI by distilling the team's company cognition into the shared model.

Homogeneous RSI

We study homogeneous RSI, in which one shared tongue example serves as the task-solving agent, workflow optimizer, and evaluator:

Agent

The project solver, or optimizee. Produces a workspace containing results, plots, and a investigation report.

Workflow optimizer

Uses the feedback from the evaluator to revise how agents distinct the task, toggle results, and inspect one another's work.

Evaluator

Compares two workspaces, chooses which improved satisfies the task, and explains the judgment.

We use non-verifiable investigation tasks as training data. We use non-verifiable to average that an automatic rule-based indication is not accessible and that lone a self-evaluation indication (whose correctness is unknown) guides the general learning loop.

To measure progress, we separately ask GPT-5.5 and Claude Opus 4.8 to measure the completed work. Their judgments are withheld from MASS. This gives us an exterior measure of betterment in a procedure whose supervision comes from the example itself.

How MASS works

MASS uses workflow optimization to elicit company intelligence, afterward distills the team's cognition into the shared model. The inner loop optimizes the team's workflow during the example weights remain fixed. The example evaluates and judges for itself which workflow is optimal. The outer loop afterward updates those weights by learning from executions of the optimized workflow. The interactive diagram below follows the two loops.

INNER LOOPevolving the team circular 1 · repeat 1 · champion: — Multi-agent executor agents sharing one model orchestrator 1 2 3 subagents Self-evaluator judges the workspace ✓ new champion Self-optimizer rewrites the workflow Workflow v1 hop 1 2 3 role role role instruction instruction instruction contract contract contract workspace feedback revised workflow, stated to the squad as its prompt ℒ(0) ℒ(0) ℒ(0) OUTER LOOPlearning from the team's work = the identical example in all role Run victor ×18 Self-rank workspaces Keep the finest 16 Fine-tune, workflow removed example capability ℒ(0) base ℒ(1) ℒ(2) instruction contract

← rub to see the entire fig →

Figure 2. MASS. Workflow hunt changes how the agents coordinate. Training on selected executions carries their cognition into the shared weights. The updated example afterward resumes all three roles. This animation illustrates the procedure. Drag the advancement bar to examine any part of the animation, including during paused. Hover complete a component for details, or choose it to jump to the explanation. Underlined conditions in the content nexus to the diagram.

0. Setup: workflow as a prompt

A workflow is a written scheme for how a multi-agent squad should activity together. The orchestrator assigns activity to subagents and integrates their results. The scheme is added to the project prompt, making the team's institution accessible for the example to inspect and revise. It has four components:

  • Role: the duty assigned to all subagent, specified as preparing data.
  • Instruction: the activity that subagent should transport out.
  • Contract: the data and artifacts it must return, including any required checks.
  • Hop: the command of subagent calls, including any come back to an before step.

For example, an evolved backing workflow arranges data preparation, characteristic construction, modeling, and backtesting into a sequence, followed by an audit. Its data agreement calls for definitive checks on the timing of the data:

Evolved backing workflow, condensed from the paper

Call order Data → Features → Modeling → Backtest → Audit → Reconciliation → Report Data specialist's output contract validation_report: no_future_leakage: bool train_end_before_val_start: bool val_end_before_test_start: bool

A backtest can appearance convincing if it lets historic decisions use data from the future. This agreement requires the data expert to inspect for that error and study whether the period splits are valid. It additionally tells afterward agents what evidence they should obtain before proceeding. Part of the investigation scheme is thus expressed in what agents must established and continue on to one another.

1. Inner loop: hunt for a improved workflow

For all training task, MASS retains a champion: the finest workflow established so far according to the current model's judgments. Each repeat has three steps:

  1. 1Execute. The squad follows the applicant workflow and produces a workspace.
  2. 2Evaluate. The evaluator compares this workspace alongside the champion's workspace and explains which is better. A winning applicant becomes the new champion.
  3. 3Revise. The optimizer uses the feedback former to propose the next workflow.

The evaluator judges the completed work; the optimizer uses that judgement to revise the procedure that produced it. Optimizer's scheme immediate emphasizes contracts and hop, directing workflow hunt toward what data agents toggle and whenever they obtain it.

2. Outer loop: study from the selected team

Even alongside an optimized workflow, the norm of long-horizon investigation activity can change from run to run. MASS hence executes self-judged optimal workflows (or victor workflows, definition the optimal workflows established so far) alongside multiple random seeds and uses the current example to position the resulting workspaces, retaining the highest-ranked trajectories. Each trajectory records the messages, tool calls, and tool results of the orchestrator and its subagents.

  1. 1Run the victor workflow 18 times for all task.
  2. 2Use the current example to compare and position the resulting workspacesA Bradley–Terry example converts the pairwise preferences into a mark for all workspace..
  3. 3Retain up to 16 trajectories per task: 15 for training and one for validation.
  4. 4Remove the workflow content from the orchestrator's first prompt whenever constructing the training examples.
  5. 5Fine-tune the shared model alongside LoRA on the two orchestrator and subagent conversations. The model's outputs are the training targets; project prompts and tool results provision context.

Removing the workflow from the orchestrator's training immediate encourages the example to study how to arrange the activity from the project itself.

Results

We run two complete rounds of MASS starting from Qwen3.6-27B in the qwen-code coding-agent environment. Our synthetic investigation suite contains 12 tasks in finance, robotics, and pharmacy, all requiring code, quantitative results, and a investigation report. Nine tasks are designated for training and three, one per domain, are held out from post-training. Eight ultimately provision training dataOne training project is excluded since workflow hunt finds no workflow that strikes its reference. The project descriptions were generated by GPT-5.5 before the experiment. Within MASS, the current Qwen example supplies workflow feedback and training-data selection. The three test tasks are excluded from fine-tuning, validation, and checkpoint selection, although workflow hunt is additionally evaluated on them..

We difference three generations: the basis example \(\mathcal{L}^{(0)}\), the first-round example \(\mathcal{L}^{(1)}\), and the second-round example \(\mathcal{L}^{(2)}\). We analyze the norm of their work, their capability to arrange and fairness it, and their use of coordination without a supplied workflow.

Better project performance

On the three synthetic test tasks, the win charge against the basis rises from 53.9% following one circular to 69.9% following two rounds (Table 1). The second-round example additionally wins 60.8% of comparisons against the first-round model.

Table 1. Synthetic project performance. External-judge win rates throughout eight training tasks and three tasks held out from post-training. All models obtain project prompts without an optimized workflow.
Evaluation tasksRound 1 vs. baseRound 2 vs. baseRound 2 vs. circular 1
Training tasks65.6%82.5%68.1%
Test tasks53.9%69.9%60.8%

To test transfer beyond the synthetic investigation suite, we measure community benchmarks. After two rounds, score per output token reaches 1.2–1.6× the basis model's level throughout MLR-Bench, DSBench, ScienceAgentBench, and AstaBench (Figure 3).

Benchmark scores versus output tokens for the basis example and two MASS generations. Round-two score-per-token ratios related to the basis are 1.60 for MLR-Bench, 1.35 for DSBench, 1.20 for ScienceAgentBench, 1.28 for AstaBench E2E-Hard, 0.96 for Terminal-Bench 2.0, and 0.94 for SWE-bench Verified.
Figure 3. Public benchmark performance. Each arrow connects the basis example (blue) to the example following two MASS rounds (green). Labels study the alter in mark per output token as a multiple of the basis value. Error bars display norm deviations throughout trials. Reproduced from Figure 2 of the paper; Appendix C.5 gives benchmark protocols and test counts.

The gains are strongest on four investigation benchmarks resembling the training tasks. Score per output token is 0.96× the basis value on Terminal-Bench 2.0 and 0.94× on SWE-bench Verified, showing that transfer remains uneven. We doubtful this generalization can be effortlessly addressed by introducing application engineering tasks into the training data.

Observation 1. Role generalization in RSI

We detect that MASS enables function generalization in RSI. We train the example lone on task-solving trajectories and detect that the capabilities of the workflow optimizer and workspace evaluator additionally improve. This enables RSI iteration to be sustainable.

Cumulative figure of tasks alongside a workflow unanimously preferred to the basis example without a workflow. Later MASS generations attain additional tasks in small hunt iterations.
Figure 4. Role generalization. Later generations discover prosperous workflows earlier. Each curve uses one example generation in all three roles. A project is counted formerly all six external judgments favor a candidate's output to the basis model's output without a workflow. The curves document whether a achievement has been established by all iteration; afterward candidates can motionless execute worse.

Figure 4 shows optimizer can discover prosperous workflows sooner as sequence goes by.

In fact, Figure 4 is technically not a extremely fair comparison, since \( L^{(0)}\) denotes a environment in which the basis example (optimizer) updates the basis example (optimizee), and \( L^{(1)}\) denotes a environment in which the first-round example (optimizer) updates the first-round example (optimizee). For a fair comparison, we provision the first-round model's workflows to the basis example (Table 2). We detect that the basis example alongside first-round hunt workflows wins 64% of comparisons against the basis example alongside base-search workflow. Each set of workflows was discovered using its generation's executor, evaluator, and optimizer. The transfer test holds implementation weights fixed during evaluation. It does not isolate the optimizer from the another roles during search. Table 2 plainly shows that optimizer's capabilty have increased.

Table 2. Workflow transfer. In the highlighted comparison, the identical basis example executes the two sets of workflows. External judges measure the resulting workspaces.
ExecutorWorkflow sourceCompared withWin rate
Round 1Round-1 searchBase executing base-search workflows80%
BaseRound-1 searchBase executing base-search workflows64%
BaseRound-1 searchBase without a workflow63%

The evalutor additionally move nearer to those of the powerful external example judges (more particulars in paper). Agreement alongside those external judges rises from 73% to 93% throughout the three cycles. Overall, these results propose that task-solving examples can additionally enhance capabilities used to create and choose the next round's training data.

Ablation 1: Optimal workflow maximizes data flow

What really makes an optimal multi-agent workflow effective?

We analyze the roles, instructions, contracts, and hops of the victor workflows (the finest workflows established so far). For all component, we evaluation how much its content reveals concerning the task. In the backing example above, a function specified as “careful data scientist” could use to many tasks. A agreement specifying leakage checks for a selling backtest is much additional distinctiveWe evaluation mutual data between embeddings of workflow components and project descriptions, using text-embedding-3-large and all-mpnet-base-v2 alongside unadjusted and permutation-adjusted estimators. These data depict workflow text, fairly than messages exchanged during execution. The samples shield lone 8–9 tasks; agents inside a workflow are not autonomous project replications..

Changes in estimated project data from hunt repeat 3 to 11. Roles and instructions decrease overall, during contracts and call command increase, throughout two embedding models and two estimator variants.
Figure 5. Task data shifts toward contracts and call order. Curves display changes from repeat 3. By repeat 11, estimated project data falls in roles and instructions and rises in contracts and hops, throughout the two embedding models and the two estimators. Intermediate changes need not be monotonic, and vertical scales differ.

Figure 5 shows that between the archetypal and final iterations of workflow optimization, an optimal multi-agent workflow tends to encode project data additional powerfully in contracts and hops and small in roles and instructions. We do not display this here, but mutual requirements among the components mostly decrease too (interesting!). This form suggests that the LLM optimizer evolves the workflow components to be mutually autonomous but optimizes the data stream via contracts and hops (see this blog post for deep, intuitive explanations). I would say that workflow optimization complete discrete components resembles discovering a basis vector for that harness space.

Observation 2: Multi-agent trajectories provision a compact RSI signal

We detect that trajectories from multi-agent systems can be additional productive than trajectories from sole agents in RSI. Multi-agent implementation additionally changes the construction of the training data. The orchestrator's conversation shows how to distinct a broad project into focused assignments and coordinate their execution. A subagent's conversation shows how to transport out an idiosyncratic assignment. Specifically, we detect that jointly learning coordination from the orchestrator and bounded implementation from subagents provides additional effective supervision per training token than learning from lengthy single-agent trajectories (Figure 6).

We difference students trained independently from the identical basis example on single-agent trajectories (S, S+, S++), multi-agent trajectories (M, M+), and a blend (X). A affirmative sign denotes a larger training set; M+ is the first-round MASS model. All points below study results pooled throughout four training tasks and one test task.

Win charge against the basis example versus supervised tokens processed. Multi-agent pupil M reaches 56.0% at 3.64M tokens; M+ reaches 68.3% at 28.82M. Single-agent students S, S+, and S++ attain 44.0%, 59.0%, and 64.0% at 1.82M, 9.11M, and 42M tokens.
Figure 6. Performance per supervised training token. Multi-agent students lie complete the row joining the observed single-agent results. Scores pond four shared training tasks and one test task. Token counts contain supervised targets processed up to the selected checkpoint, including repeated sampling; they exclude data generation, search, and evaluation.

Against the basis model, M+ reaches a 68.3% win charge following handling 28.8M supervised training tokens, during S++ reaches 64.0% following 42M tokens. The MASS example thus achieves the higher mark against this citation alongside 31% small supervised tokensThe first-round dataset stores 17.2M supervised tokens. Sampling alongside replacement and balancing orchestrator and subagent examples create 28.8M processed supervised tokens through the selected checkpoint. This measures visibility to training targets, including repeated examples. It excludes afterward updates, input tokens, validation, data generation, search, and evaluation..

Higher-quality trajectories may contribute to this advantage. In a distinct analysis, powerful external judges favor the multi-agent systems' activity in 85.1% of comparisons alongside single-agent outputs. Team runs additionally create concerning 2.4× as many tokens on averageAmong ten pairs alongside akin generated trace lengths, the penchant for squad outputs remains 80%. This difference matches observed lengths, not generation budgets. The groups additionally differ in workflow and verification instructions, so the difference does not isolate the consequence of delegation. These external judgments are used for this inspection lone and do not guide MASS.. But whenever we difference the supervised tokens used during SFT, we discover that multi-agent trajectories can provision a denser learning indication than single-agent trajectories.

These results propose that trajectories from multi-agent systems can be additional productive than trajectories from sole agents in RSI, extending their value beyond test-time scaling.

Observation 3. The RSI bottleneck can be the optimizer, not the evaluator

We observed that the bottleneck in RSI can be the optimizer, not the evaluator. We kept the basis example as the project solver and exchanged in a stronger external example (GPT-5.5) for the evaluator only, or for the two the evaluator and the optimizer.

Line storyline of tasks won against the naked basis example versus workflow-search repeat for multiple configurations. GPT-5.5 as evaluator and optimizer reaches 12 tasks by repeat 5. GPT-5.5 as evaluator lone reaches 10 tasks. The all-base-model iteration reaches 11 tasks by repeat 13. Variants alongside distinct hunt histories are additionally shown.
Figure 7. Workflow-search ablations. Number of tasks anywhere the evolved workflow strikes the naked basis model, by hunt iteration. The administrator is continually the basis example \(\mathcal{L}^{(0)}\). Upgrading the two the evaluator and optimizer to GPT-5.5 (orange, dashed) finds winning workflows for all 12 tasks inside 5 iterations. Upgrading lone the evaluator (green) helps much less. The completely homogeneous iteration (blue) is slower but motionless reaches 11 of 12 tasks.

Figure 7 shows that upgrading lone the evaluator helps a little, but upgrading the optimizer as fine helps a lot. The strictly homogeneous self-improvement iteration does eventually discover winning workflows for 11 of 12 tasks; it fair takes longer. Our explanation is that the powerful evaluator may provision high-quality feedback, but the feeble optimizer is not capable adequate to digest it. This suggests that, under a constricted budget, one may have to prioritize the optimizer complete the evaluator inside an RSI iteration unit.

Win charge against the basis example without a workflow throughout workflow-search iterations. The curves for all configurations fluctuate, alongside powerful iterations followed by keen regressions.
Figure 8. Workflow-search stability. Win charge of the basis example executing all applicant workflow against the identical example without a workflow, by hunt iteration. Colors equivalent the configurations in Figure 7. Strong iterations can be followed by keen achievement regressions.

Figure 8 additionally shows that workflow hunt is noisy. In text-space optimization, reaching a fine workflow and improving on it stably are distinct problems. This suggests that the LLM optimizer may need an equal of a learning-rate scheduler to stabilize the hunt in content space.

Conclusion

In Tennyson's Ulysses, Odysseus imagines another voyage in chase of cognition “Beyond the extreme border of individual thought.” For AI, we ideate an Odyssean moment whenever continued learning must continue without a stronger supervisor. The difficulty concerns the two who can supervise additional learning and how advancement have to be judged. Human experts may be capable to province a goal without specifying an goal function that completely captures it.

Writing is a fine example: expert writers can acknowledge powerful writing, yet battle to province their judgement in a rubric that applies reliably throughout contexts. For RSI, the investigation inquiry is how AI can rotate incomplete individual direction into helpful criteria for learning, and how to test whether those criteria continue to grasp the intended goal as its capabilities grow. This becomes particularly challenging whenever the example being improved additionally serves as its own optimizer and evaluator: the judgments that guide betterment must themselves remain open to correction.

Wood engraving of a rowboat carrying multiple men toward a sailing vessel, alongside a mountainous coastline in the background.
Figure 9. Illustration for Tennyson's Ulysses (1857). Drawing by Clarkson Stanfield; engraving by W. J. Linton. Scan by George P. Landow, The Victorian Web.

Keeping a system's self-evaluation open to correction is additionally essential for safety. In a homogeneous loop, a shared blind place could authorize a flawed outcome to be produced, accepted as an improvement, and unified into the next model. If humans additionally battle to acknowledge the flaw, successive updates may fortify it. The safety involvement is that the identical weakness can colony the two the system's behavior and the scheme meant to accurate it.

In this sense, shared blind spots may be an Achilles' heel of recursive self-improvement. Understanding them matters for controlling a additional capable scheme since it reveals anywhere reliance on the system's own judgments is smallest justified. That cognition could guide autonomous checks, constraints on autonomous actions, and conditions for intervention. Identifying a weakness solitary does not established control; we must additionally display that we can detect its consequences and intervene efficiently as the scheme changes.

OpenAI's investigation on chain-of-thought monitorability illustrates one part of this challenge: reasoning traces can assistance disclose misbehavior, but they provision incomplete evidence. Recent inspection of a imaginable intellect detonation gives additional logic to study these conditions early. This is a chief safety inquiry for AI's Odyssean moment: how to maintain meaningful oversight whenever no stronger supervisor is available. ⛵

Other Article Hacker News
↑
Close Right Ads
Close Left Ads