Where's the "Intelligence Explosion"?

Hacker News by 36 min read 507x views
Where's the "Intelligence Explosion"?

Share Post

Art by GPT-6

One of my essential beliefs concerning the earth is that Ramez Naam ought to blog more. Ramez is among the world’s top futurists — he predicted the solar and power division revolutions lengthy before these were extensively understood. If you were study Ramez in 2011, you were capable to comprehend the forthcoming of the two energy innovation and climate change, lengthy before another group did. His before publish More than Human is motionless a awesome guide to the benevolent of biologic enhancements that AI power create possible. Ramez is additionally an outstanding discipline fabrication author, having written a trilogy of novels in which nanotechnological telepathy is distributed as a gathering medication (I’m not certain if he really expects that to happen, but it’s a extremely chill idea).

Unfortunately, although he does have a Substack (which you should entirely follow), Ramez does not blog regularly. However, following having a lengthy personal conversation alongside him concerning Recursive Self-Improvement, I was capable to prevail upon him to compose up his thoughts for my blog.

To say that RSI is a big agreement in the AI earth would be a colossal understatement. Among AI researchers, entrepreneurs, and AI safety people, there’s a general condemnation that as AI gets improved at improving itself, there volition be a “fast takeoff” or “FOOM”, in which AI’s capabilities “take off” and create a technological Singularity. This event is a staple of discipline fiction, including plant by my favorite sci-fi author, Vernor Vinge.

A lot of group in the industry accept that this instant is now near at hand, and are racing toward that prize:

X avatar for @tanayj

Tanay Jaipuria@tanayj

Noam Brown says that recursive self-improvement / the capability of AI models to do AI investigation is the top precedence for OpenAI by a broad margin. OpenAI ranks it complete making models that sell. (via @theinformation podcast)

12:24 AM · Sep 27, 2026 · 4.01K Views

4 Replies · 2 Reposts · 44 Likes

But Ramez — normally among the most wide-eyed of techno-optimists — is extremely skeptical that we’ll see item akin the “FOOM” of Vernor Vinge novels. In this lengthy, well-researched post, he explains his skepticism.

Personally, I’m agnostic. Ramez’s case necessarily rests on a lot of assumptions; although it’s cogently laid out, I think the genuine answer is that we’ll fair have to delay and see whether the Singularity arrives. But equal additional fundamentally, I don’t cognize how much this conversation matters in the applicable awareness — equal without the benevolent of Singularity depicted in sci-fi novels, AI capabilities are improving so quickly that they’re already superhuman in many respects, and shortly volition likely be powerfully superhuman in most or all dimensions. The AI of 2040 is going to appearance godlike, whether or not it explodes into an genuine god in 2027.

Still, it’s a extremely engaging argument, and Ramez’s thoughts on the forthcoming of innovation are continually value listening to.

AI is already assisting enhance itself. The inquiry is whether equal completely autonomous recursive self-improvement (RSI) would logic a runaway intellect explosion.

The theory is that all generation of AI could build a improved successor, faster than the final generation did. That could guide to a “fast takeoff,” alongside capabilities surging to synthetic superintelligence (ASI) in a year, months, or equal days.

Here’s my take: Given our finest current data, the AI self-improvement iteration would need to be approximately 5–10× stronger to prolong itself, let solitary run away. I’ll explain this math in section 8. I anticipate incredibly fast AI advancement by the standards of nearly any another technology. But the evidence we have doesn’t propose a abrupt detonation to incomprehensible superintelligence anytime soon.

I could be wrong. Forecasters have often underestimated AI progress! I could fine be next. One item that’s apparent is that we need improved data. For now, let’s activity alongside what we can measure, and remain open to breakthroughs that could alter the picture.

Figure 1. How powerful is the self-improvement loop? Model.

Jump to the conclusion.

Here’s the case, alongside links to all part:

  1. Narrow Superintelligence Is Here Today

  2. Real-World Research Is Harder

  3. Impressive AI Numbers → Sharp Diminishing Returns

  4. We’re Not Seeing Signs of Acceleration

  5. Keeping Up the Pace Takes Exponentially More Resources

  6. Better AI May Be Needed Just to Maintain the Pace

  7. Progress Gets Harder; Ideas Get Harder to Find

  8. The Current Feedback Loop Doesn’t Look Strong Enough

  9. OpenAI’s Data Shows How Weak the Loop Is

  10. What Could Accelerate Progress?

  11. We Need More Data to Track This Well

Key charts: The feedback loop · Measured vs. forecast progress · Diminishing returns

People use “recursive self-improvement” to average everything from AI boosting the efficiency of individual researchers to AI bootstrapping itself to incomprehensible intelligence. Here’s my taxonomy: efficiency gains (Type 1), expanding autonomy during motionless facing diminishing returns (Types 2–4), and a runaway iteration to superintelligence if we can always discover accelerating returns (Type 5).

Figure 2. Five types of AI self-improvement.

We’ve made genuine advancement on Types 1 and 2: AI helps the two researchers and engineers inner of AI companies, and mighty models can train and enhance smaller ones. We haven’t yet seen apparent evidence for Type 3 (though Alibaba fair made several powerful claims) and certainly not for Type 4. I do anticipate autonomous self-improvement to attain at several point. I’m skeptical that it leads to Type 5 - runaway super-intelligence - without a important conceptual breakthrough.

There are plentifulness of another definitions of RSI, which can be a bit confusing. Weco’s four levels of RSI are near to mine. For a broader tour of all the things group average whenever they say ‘RSI’, peruse Tom Cunningham’s thorough guide.

I do anticipate narrow superintelligence in extremely verifiable domains. Think chess, Go, ceremonial math, parts of device discipline and coding. Highly verifiable domains are mostly ceremonial and organized types of activity anywhere machines can create unlimited training data, alongside ideal or near-perfect verification of accurate vs incorrect, and do so entirely in application without waiting on the bodily earth or humans. That’s an ideal environment for AI learning.

Figure 3. What makes a domain extremely verifiable?

In fact, we already have narrow superintelligence in equivalent plang. We’re seeing it happen now in the most ceremonial parts of math, in particular in proofs and in finding counter-examples that disprove important conjectures. For example, OpenAI lately reported an AI-generated evidence resolving the Navier–Stokes beingness and smoothness problem. Parts of application betterment are additionally extremely verifiable, during others are a bit small keen (such as understanding what humans want).

That isn’t the identical as broad superintelligence. Even our most mighty models need far additional training data than humans, battle to study reliably from ongoing experience, and neglect in amazing ways on tasks group discover straightforward. Superhuman math doesn’t automatically average superhuman judgement everyplace else.

Back to contents

Benchmarks and forecasts propose that AI models should reliably succeed at coding tasks that obtain humans hours, without individual help. The genuine earth is messier. OpenAI’s inner data shows much shorter stretches of autonomous activity on investigation tasks.

In its Research Acceleration / RSI report, OpenAI showed how frequently its models completed tasks alongside and without individual help, grouped by how lengthy a individual would need to do the work.

Figure 4. OpenAI’s inner investigation tasks. Source.

Even on tasks that would obtain a individual small than 15 minutes, OpenAI’s models succeeded without individual involvement lone 86% of the time. The estimated project dimension at 80% achievement was approximately 15 minutes complete the archetypal seven months of the year. July’s results were akin to the entire duration average.

Fully autonomous RSI would necessitate an AI to cord together a awesome many investigation tasks reliably, stretching out complete complex tasks that humans need weeks or months to accomplish. OpenAI’s data suggests that we aren’t close.

Anthropic additionally released a graph showing how Claude accelerates AI research. It shows that inner AI models collaborate on or equal guide additional than 90% of R&D tasks. That’s objectively impressive. At the identical time, the chart reports zero cases of AI autonomously completing AI R&D tasks.

Figure 5. Claude’s function in inner AI R&D. Source.

These are amazing tools. But they motionless need skilled group to set direction and get them rear on track.

For years, METR has been publishing a diagram showing what dimension of coding project (measured in individual hours to complete) best-in-class AI models can achieve. It’s been called the most crucial chart in AI. METR’s Mythos Preview evaluation estimated that the example could win at 80% of coding tasks that took humans three hours.

Figure 6. METR’s 80% project horizons. Source.

Epoch’s own regulation of thumb is that all five additional points of ECI (their general benchmark of AI capability) equivalent to approximately a doubling of METR’s project horizon. Using that formula, we’d anticipate GPT 5.6 Sol and GPT 6 Astra to be 80% prosperous at completing tasks of about 4 hours and 11 hours of individual length, respectively.

Another evaluation (a forecast) of AI project dimension comes from the AI 2027 scenario, which estimated that by July 2026, frontier AIs would be 80% prosperous accomplishing tasks of about 11 hours. Fairly similar.

The AI 2027 Tracker charts all of these.

Figure 7. The AI 2027 Tracker. Source.

Inside OpenAI, though, the July research-task horizon at 80% achievement was approximately 15 minutes.

Here’s the gap:

Figure 8. Forecasts, benchmarks, and genuine AI research. Tracker · OpenAI.

A four-hour benchmark horizon is concerning 16 times longer than OpenAI’s investigation horizon. AI 2027’s 11-hour forecast is concerning 44 times longer. Of course, the tasks being performed by researchers at OpenAI aren’t the identical as those in the METR benchmark. So we should anticipate several discrepancy. This, however, goes fine beyond that.

Actual AI investigation at OpenAI is an command of dimension or additional harder than metrics, benchmarks, or forecasts suggest. That should create us wary of relying too much on benchmarks, or of saying that forthcoming scenarios akin AI 2027 are ‘on track.’ The authors of the connected AI 2040 project motionless depict AI 2027 as approximately the forthcoming they expect, and say actuality is tracking nearer to it than equal they expected. That’s not what we see from inside OpenAI. This isn’t an apples-to-apples comparison, but the difference is remarkable. AI 2027 appears to be substantially over-optimistic in this regard.

In January of this year, Nathan Witkin made a case that the METR chart was exaggerating progress. The genuine earth data suggests that at smallest several of his critiques were correct. The gap between benchmarks, forecasts, and data gleaned from genuine use of AI should power our expectations concerning the future.

Back to contents

OpenAI’s study additionally shows notable increases in AI token usage, in compute expend per researcher, and in lines of code written. But these aren’t results. They’re intermediate measures. How much advancement do they really drive?

Researchers used 124x additional tokens per person. Engineers shipped approximately 7x as many lines of code per person. Researchers ran 1.6x as many experiments per investigator vs OpenAI’s 2025 entire twelvemonth average.

Figure 9. Token use inner OpenAI. Source.

Figure 10. Experiment gait inner OpenAI. Source.

Figure 11. From tokens to code to experiments. Source.

More tokens and code don’t inform us much on their own. The 1.6× test gait is nearer to helpful investigation output. Even that doesn’t average AI is improving 1.6× faster.

An enormous addition in AI output has accompanied a much smaller addition in experiments run.

This isn’t a controlled experiment. We don’t cognize what would happen if researchers switched rear to an older model. But it gives us a helpful perspective of AI-assisted investigation inner a frontier lab.

It’s not fair OpenAI. Anthropic reports that their engineers are now producing 8x as many lines of code per individual as they did in 2024 - slightly akin to OpenAI. Anthropic also sees important diminishing returns between efficiency and AI progress. Here’s a straightforward citation from its Mythos Preview scheme card:

“Productivity uplift does not translate one-for-one to capabilities progress. We surveyed specialized personnel on the efficiency uplift they cognition from Claude Mythos Preview related to zero AI assistance. The allocation is broad and the geometric average is on the command of 4×. […] We evaluation that reaching 2× on general advancement via this conduit would necessitate uplift approximately an command of dimension larger than what we observe.”- Anthropic, Claude Mythos Preview System Card; accent mine

Translation: To twice the gait of AI progress, Anthropic estimates that AI would need to addition the efficiency of their workforce by approximately a aspect of 40 related to no AI assistance.

Figure 12. Anthropic’s productivity-to-progress estimate. Source.

This is an estimate, not a measure of progress. Even the 4× efficiency fig comes from an opt-in study of 130 Anthropic staff. I put additional importance on OpenAI’s logged experiments, although the two sources measure distinct things.

We don’t yet cognize how much those additional experiments are accelerating AI improvement, if at all. In general, there are additionally steeply diminishing returns of additional experiments in most branches of science. That method that a 60% addition in test gait could be on the command of a 10% boost to AI betterment pace. (A power law exponent of 0.2, for those who desire to do the math.) That’s speculation for now. We’ll study additional as the labs publish results.

What concerning giving the identical AI example additional period to think?

That scales seriously also. In OpenAI’s lately publicized results on unsolved math problems, achievement rises approximately alongside the log of compute complete the range shown. It shows logarithmic diminishing returns. In plain English, all additional doubling of compute for a example buys approximately the identical acquire in achievement rate, during costing twice as much.

Figure 13. Test-time compute and math performance. Source.

What if we throw additional agents at it instead? A average RSI / ASI idea is that formerly we have AIs at a certain capability level, we can fair spawn additional copies and put them to work.

Adding agents can get tasks done faster and sometimes attain a higher capability level. But on the three benchmarks in Toby Ord’s analysis, expanding a swarm buys small betterment per token than letting one delegate think longer.

His coarse regulation of thumb is a quadrate root. If one delegate can accomplish a project in 10 hours, afterward 100 agents could accomplish it in one hour. The speedup is 10, the quadrate base of the figure of agents (100). But to get this speedup, you addition the total disbursal in tokens or run period compute by the identical factor. So going from one to 100 agents can get a project done in one tenth the time. But it’ll be ten times as expensive.

Parallel agents can preserve time, at a much higher compute cost.

Another difficulty is that agents often think alike. In a study comparing LLMs alongside 467 people, the archetypal ten AI responses offered company ingenuity comparable to concerning eight to ten people. After that, approximately two additional AI responses added as much as one additional individual response. A separate study throughout example families additionally established small assortment in AI responses. That doesn’t average all delegate has the identical idea. But a hundred copies may recommendation small assortment than a hundred distinct researchers.

None of this makes swarms useless-or safe. Lisan al-Gaib makes a powerful case for parallel delegate swarms as a potent cyber-weapon in “Accidental Scaling.” I don’t portion all of his appraisal of what swarms have accomplished. In math, for example, I think he gives far too much credit to the swarm and not adequate to the improved inner example that OpenAI used.

OpenAI says the example rearward its Navier–Stokes result was developed through “large-scale reinforcement learning on top of a earlier pretrained model.” Formal math is a extremely verifiable domain, which makes it a particularly fine fit for that approach: Machines can create nearly limitless amounts of training data, and verify that solutions are accurate or incorrect, all in software. My conjecture is that this model’s complete results volition display an particularly ample betterment in math.

OpenAI’s Noam Brown made the chief item explicitly: he wouldn’t provision multi-agent methods equal 10% of the credit for the Navier–Stokes result.

I do think Lisan makes fine points concerning cybersecurity. If you’re searching for a safety exposure at a mark location and can distinct the hunt among agents, speed may validate a huge token bill. Swarms can be hazardous equal whenever they’re inefficient.

I’m small convinced that this scales to investigation breakthroughs. Inventing item akin the transformer likely takes additional than searching a area person has already defined.

Back to contents

Building a improved example can bring gains that additional thinking period or additional copies of the old example can’t. Look at the gap between Astra and OpenAI’s inner example on the identical math problems.

Figure 14. Better models versus additional thinking time. Source.

That’s the strongest type of the RSI argument: a additional capable AI could do investigation that today’s example can’t do, nevertheless many copies we run.

But construction that improved example additionally runs into diminishing returns. More training data, additional training compute, larger models, and additional reinforcement-learning (RL) compute all display diminishing returns in published scaling studies. Making compact models larger normally raises the compute needed for all output token, too. None of these routes gives us a liberated continue about the problem.

Figure 15. Diminishing returns to scaling. Chinchilla · ScaleRL · OpenAI.

Those scaling results provision us logic to anticipate diminishing returns whenever AI helps build the next model, too.

Back to contents

AI capabilities are rising quickly. But the community data doesn’t display a sustained acceleration. To the degree that AI tools are boosting productivity, they may be being offset by the problems expanding harder. Or we may merely be early. Either way, the trend isn’t showing a accelerated takeoff.

Figure 16. Frontier ECI gains since January 2024. Source.

The public ECI frontier-the finest mark among models released by all date-has gained concerning 16 points a year on a trend fitted from January 2024 through September 2026. That’s blisteringly accelerated progress, but this duration doesn’t display a runaway surge.

Here’s the identical frontier in complete ECI points, through July 2026, to put it in perspective.

Figure 17. The complete frontier ECI score. Source.

The community frontier additionally can’t inform us everything happening inner the labs. Anthropic gives us a nearer appearance in the Opus 5.5 scheme card, using its own type of the index, AECI.

Figure 18. Anthropic’s fitted capability trend. Source.

Eli Lifland, a co-author of AI 2027 and AI 2040, saw the apparent trend interrupt as a alert that we were heading toward an intellect explosion:

“Anthropic is likely correct current [that they hadn’t reached hazardous levels of AI self-improvement], but alert bells have to be going off! Our processes are not prepared to grip an intellect detonation and we appear to be going full-steam onward toward one.”
- Eli Lifland, On Mythos’s AI R&D Capabilities

What looked akin acceleration now appears additional accordant alongside a one-time jump. The flat went up. The charge hasn’t kept climbing.

Achieving those gains has required an enormous addition in the inputs to AI. For example, regard computing power. Epoch’s estimates of AI part capacity, measured in NVIDIA H100 equivalents, display approximately 127-fold growth in fair complete three years (including projections at the end of this period).

Figure 19. AI part capability and frontier ECI. Source: Epoch AI.

This is total AI part capacity, including inference. Still, the addition is striking: vastly additional computing capability has accompanied much steadier gains in measured capability.

The broader image looks similar. Here are six inputs alongside capability gains, going rear to February 2023.

Figure 20. Six inputs alongside frontier ECI. Epoch part data · SemiAnalysis workload shares.

Everywhere we look, AI has diminishing returns. It gets additional costly in treasure and endowment to create all stage forward. More of all input has been required to keep dependable gains in AI capabilities.

We’ve been capable to measure these inputs because, until recently, the disbursal was inside the range of what hyperscalers could pay from their profits. That is no longer the case. From this item forward, forthcoming AI funding volition increasingly depend on AI revenues going up. And the measure of the numbers - 3% of US GDP is now going into AI infrastructure - suggests that eventually the growth charge volition decline. If funding growth does slow, to item small than its current blistering exponential pace, capability advancement could dilatory too. Even if funding growth continues (which I anticipate for the foreseeable future) a slowdown from its current exponential growth charge to a additional humble one (which I additionally expect) could guide to a slower gait of progress. Better AI investigation tools may be needed to offset that.

The day whenever we need improved AI tools fair to continue the gait of AI advancement may already have arrived. Not since funding is slowing, but since the issue of improving AI itself gets harder at all step.

Here’s Anthropic in the Mythos 5.1 scheme card:

“we accept that inner use of latest AI models has been a key aspect in maintaining the current charge of progress, but we do not yet see apparent signs of theatrical acceleration beyond that rate.”- Anthropic, Claude Fable 5.1 & Claude Mythos 5.1 System Card, division 2.3 – accent theirs.

The key term is maintaining-and Anthropic italicized that term in its own scheme card. Increasingly capable AI may be essential fair to keep the gait of betterment anywhere it is.

Opus 5.5 improves substantially on multiple coding and device use benchmarks. But on CoBench, Anthropic’s benchmark built from historic AI R&D problems, it gains fair 2.6 percent points complete Opus 5, inside the reported error bars.

Figure 21. Opus 5.5 benchmark gains. Source.

Why the smaller acquire here? Maybe AI investigation is merely harder than another tasks. Bear in intellect that CoBench isn’t evaluation the capability to create important discoveries. It’s much additional constricted in scope. It asks models to examine historic AI R&D problems using code, logs, and documents. That’s helpful investigation debugging and efficiency work, but it doesn’t immediately test whether a example can invent a new architecture or create a conceptual breakthrough.

The evidence on open-ended investigation suggests another obstacle: coming up alongside helpful ideas that haven’t already been tried.

Back to contents

Why do helpful new ideas frequently get harder to find?

Tom Cunningham and Manish Shetty have a helpful apple-picking metaphor. An AI can choice the low-hanging create quickly, during humans can motionless attain ideas the AI can’t.

Once those apples are picked, another copy of the identical delegate finding them again doesn’t help. A stronger example can attain higher. To add my own flourish, the apples may additionally get sparser and farther distinct as you climb. The RSI inquiry is whether all crop gives us adequate to build a improved apple-picker.

Figure 22. The apple-picking example of AI R&D. Source.

This form shows up throughout R&D. Bloom and colleagues document sectors anywhere investigation attempt grows during investigation efficiency falls. A celebrated example is Eroom’s Law: in the historic drug-development data, the inflation-adjusted R&D disbursal per new approved medication approximately doubled all nine years.

Figure 23. Eroom’s Law in medication development. Source.

Pharma has another complications, including regulation, difficult medicinal trials, and rising expectations for safety. Existing treatments can additionally lift the bar for a helpful new drug. But several of this difficulty may additionally be that the low-hanging create has been picked.

Stockfish, the chess engine, gives us a additional straightforward appearance at application research. We have records of experiments aimed at improving it and the gains that followed. This gives us a real-world dataset to appearance at the gains of experimentation in software. As a result, multiple RSI models diagram on this data. That said, not all the improvements came from these experiments. Several crucial ideas additionally came from exterior the project, so we shouldn’t provision its experiments all the credit.

Epoch’s inspection of application R&D estimates returns to investigation attempt at concerning 0.83 for Stockfish, a bit slower than linear. These are diminishing returns, but gentle ones. These returns, however, are improvements in computational efficiency. And additional compute does not rotate immediately into additional AI capability. As we saw earlier, AI capability additionally has steep diminishing returns from adding additional computational power. So we shouldn’t peruse that 0.83 as the come back from experimentation to AI capability itself. AI capability grows much additional gradually than compute, as we’ve seen already.

Andrej Karpathy’s autoresearch demonstration gets nearer to the procedure we desire to understand. A “teacher” AI delegate changes a smaller “student” AI model’s training code, runs it, checks the result, and tries again. The instructor delegate itself doesn’t improve, but it is capable to enhance the “learner”. This is my Type 2: A stronger AI improves a weaker one.

One community run, posted by an delegate functioning on Karpathy’s behalf, reported 89 experiments complete approximately 7.5 hours. About 92% of that session’s acquire arrived by run 44. Gains came quickly, afterward slowed. The setup was deliberately small, alongside a five-minute training prosperity per experiment. But the delegate could alter the architecture, optimizer, and training settings; it wasn’t constricted to a fistful of knobs.

Figure 24. Gains in one autoresearch run. Source.

A afterward community run got further, so the archetypal run hadn’t hit a difficult ceiling. This is a helpful first example of autonomous research, and yet another location anywhere we see the diminishing returns endemic in AI research. That said, this was a extremely first experiment. I anticipate forthcoming systems to do much better. This particular AI betterment iteration volition apt develop stronger.

This is anywhere the difference matters. More tokens can buy additional code, and additional code can assistance us run additional experiments. But experiments lone enhance AI if they reveal item useful.

Figure 25. From AI action to helpful improvements.

The bigger inquiry is whether AI can arrive up alongside aspiring new investigation ideas or conceptual breakthroughs.

Anthropic’s clarification of Opus 5.5 is blunt:

“As alongside former models, it is weaker on open-ended research: inner users study that it mostly tests incremental ideas and prefers small aspiring hypotheses, and in our human-run existence discipline exercise, it deferred to the published writings and struggled to create novel ideas (Section 2.2.2).”- Anthropic, Claude Opus 5.5 System Card, division 2.3.3; accent mine

METR’s appraisal in the identical cardstock identifies what may motionless be missing:

“This is extremely uncertain, but we anticipate that complete automation of AI R&D volition necessitate ample improvements in foresight, prediction, creating one’s own feedback loops, and mostly another skills that power typically be referred to as investigator ‘judgement’ or ‘taste’.”- METR, quoted in the Claude Opus 5.5 System Card, division 2.3.6

In these examples, humans motionless provision much of the direction and judgment.

Future models volition likely get improved at this. But in the world’s stockpile of possible training data, we have many additional examples of incremental activity than of breakthroughs. I amazement whether that makes novelty harder to learn. That’s speculation, but value watching.

This is additionally durable to location by merely operating additional copies of the AI. A huge figure of parallel agents can assistance alongside the incremental improvements or searching complete a ample set of parameters, but for breakthrough ideas they may run into the homogeneity problem: More parallel agents motionless think alike.

Back to contents

How far are we from the self-improvement iteration being powerful adequate to prolong itself, or to propel itself into runaway super-intelligence? Can we quantify this?

We can create a coarse estimate. Better AI helps alongside research; helpful investigation produces improved AI. For the iteration to prolong itself, all circular must create adequate gains to propel the scheme through the next loop, equal as improvements get harder to discover.

Figure 26. The AI self-improvement loop. Model.

In a latest paper, The Economics of Recursive Self-Improvement, Tom Cunningham and colleagues modeled this from the standpoint of how much additional efficiency all item of additional ECI produces from an AI. They ask archetypal and foremost what that figure would need to be to create a self-sustaining feedback loop. And secondly, they try to decide what that productivity-per-ECI-point figure is today.

First, they discover a self-sustaining RSI threshold of approximately 15% additional investigation efficiency per additional ECI point. In their model, that’s concerning anywhere improved AI would create adequate progress to prolong the loop.

The image below shows the idea. At the threshold, all sequence of gains powers the next. Above the threshold, the feedback iteration accelerates. Below the threshold, the feedback iteration is too weak, and the charge of betterment it brings drops on all cycle. This example isolates the application loop; exterior funding can motionless run fast progress.

Figure 27. Three explanatory feedback paths. Source.

Updating this slightly alongside data from the Stockfish experiments puts the threshold a small higher, at approximately 19% per ECI point. I wouldn’t put much importance on that exact difference. Both estimates are uncertain. But they provision us a way to think concerning the power of the feedback iteration and a coarse collection at which self-sustaining or runaway RSI may begin.

The second item Cunningham and squad do is create a coarse evaluation that the current AI efficiency acquire is concerning 9% per ECI point. That’s below their self-sustaining threshold.

I akin the model. OpenAI’s newer data, however, suggests the iteration may be fairly a bit weaker.

Cunningham’s evaluation of 9% efficiency acquire per ECI item is according to Anthropic’s study of 130 staff, who reported approximately 4× the efficiency they’d have without AI. Cunningham and colleagues difference that alongside a 16-point capability acquire since first Claude Code.

That difference assumes the before tools added small or no productivity, so ‘no AI’ is a sensible starting point. The authors say this explicitly. I’m not certain the assumption holds for the identical researchers doing the identical work, but that’s a smaller issue.

The authors themselves cognize that this is a coarse calculation, and notify that the 4× study evaluation is likely too high.

OpenAI’s newer data gives us a firmer way to inspect the number: Actual logged experiments complete time, fairly than individual estimates of their own efficiency alongside and without AI. I put additional importance on this for three reasons:

  • Direct and broad measurement. Instead of relying on surveys, OpenAI really tracked and measured experiments run on their infrastructure. That method they didn’t depend on researchers estimating their own productivity, which can be far off.

  • Full sample, not opt-in. Similarly, OpenAI’s data catches all energetic experimenter, during Anthropic’s lone reflects the 130 workforce who took the period to answer the study – and who hence may not be a delegate set.

  • Enormously additional data. We don’t cognize how many experiments are in the 32 weeks of OpenAI data, but it’s apt at smallest tens of thousands of idiosyncratic examples and perchance hundreds of thousands.

Any way you piece it, the new OpenAI data, released following Cunningham’s document was drafted, is a larger, additional comprehensive, additional representative, and nearly certainly additional accurate dataset than Anthropic’s inner opt-in study of employees.

Now let’s use OpenAI’s test data to calibrate the efficiency acquire per ECI point. We cognize that in August, OpenAI researchers ran ~1.6× as many experiments per individual per duration as the 2025 average. If we brace that alongside roughly 16 points of frontier ECI improvement, it plant backward to concerning 3% efficiency acquire per item of ECI. By contrast, 9% compounded complete 16 points would average approximately 4× productivity.

Figure 28. Comparing efficiency estimates. OpenAI methods.

Here’s OpenAI’s published weekly sequence alongside that hypothetical way of 9% additional efficiency per additional ECI point. The blue row ends at ~1.6×. The red row shows what 9% per item would connote if 16 ECI points were dispersed throughout this period. That doesn’t equivalent what we see from OpenAI’s data. I desire to be apparent current that all data sets are noisy. We don’t cognize exactly what example researchers were using on what days, or whether the new experiments were additionally higher norm than old experiments. We need additional experiments and additional data to additional calibrate these numbers. Working alongside what we do have, what we see is a fairly low boost to efficiency from all additional ECI point.

Figure 29. Experiment gait versus a hypothetical path. Source.

Even that 3% could provision improved models too much credit. OpenAI additionally used far additional tokens and had additional compute for experiments. Those could document for several of the addition in test pace. So the range is likely a bit lower.

I use 2–3% efficiency acquire per ECI point as a operating assumption, allowing for several assistance from those another inputs. This is motionless a coarse estimate, albeit one that’s according to the finest real-world data we have.

Figure 30. Productivity estimates and the takeoff threshold. Source.

With those assumptions, 2–3% per ECI item against a 15–19% threshold leaves a approximately five- to tenfold gap. That’s a big gap, although its size depends on how fine test counts grasp helpful investigation and whether the assumed capability alter is right.

Figure 31. Diminishing returns about the loop. Source.

AI is assisting build improved AI. Under this estimate, though, all rotate of the iteration adds small than the last. The feedback would have to rotate into much stronger to prolong itself.

Back to contents

This application iteration sits alongside faster chips, bigger data centers, additional training data, and greater investment. Those can keep driving fast advancement equal if the iteration can’t prolong itself.

The iteration itself could fortify too. Better training data, memory, and research judgment could all help.

A breakthrough on the measure of the Transformer architecture in 2017 could alter the image much more. That would be a fine logic to revisit these estimates.

Better researchers power additionally run small experiments and study additional from all one. A fistful of improved ideas can matter additional than a mountain of regular runs.

Still, diminishing returns in device learning aren’t new. Cortes and colleagues were fitting device learning scaling curves in 1993: More examples reduced error, following a power law alongside diminishing returns. These diminishing returns and harsh scaling laws are as old as device learning. They didn’t appear for the archetypal period alongside transformers or LLMs or profound learning. That doesn’t demonstrate today’s relationships volition final forever. But until we see evidence that we’ve established a new method that scales without these inhibitors, we should scheme for diminishing returns as apt to be alongside us for several time.

That said, the earth is additional than fair software. Tom Davidson, Basil Halperin, Thomas Houlden, and Anton Korinek example application progress, hardware progress, and financial feedback together. Better AI helps scheme improved chips; improved chips assistance improved AI; financial growth finances additional funding in both. Several feedback loops can merge to conquer diminishing returns equal whenever one iteration solitary can’t. I think it’s awesome that person has attempted a example that integrates all these distinct avenues of improving AI through software, hardware, and economics.

But I have questions concerning the application iteration itself. In their chief calibration, completely automating application investigation puts that iteration approximately at the threshold for bomb growth, equal without assistance from improved hardware or broader financial growth. Recall that Cunningham’s example puts the self-sustaining threshold at approximately 15% additional investigation efficiency per additional ECI point, during our evaluation using OpenAI’s experimental data puts today’s gains at lone 2-3%. These models use distinct measures, so we can’t equate their numbers directly. But the difference matters: their completely automated application iteration reaches the threshold, during our finest evaluation from current data puts today’s iteration far below it.

Having AI do all the investigation doesn’t eliminate the diminishing returns inherent to improving AI, or the broader issue of helpful ideas getting harder to find. This is the difference between Type 4 and Type 5 in the taxonomy above. An AI power autonomously design, train, and test its successor, and motionless need exponentially additional resources to create all additional stage forward. Closing the iteration doesn’t inform us whether it’s powerful adequate to prolong itself.

The authors do document for diminishing returns. The involvement is whether their calibration overestimates how much helpful AI investigation all circular of application betterment produces. Diminishing returns appear to be essential to device learning. We see them in training, in test-time compute, and in the hunt for improved algorithms. Full autonomy could eliminate individual bottlenecks without removing any of those constraints.

We’ve already seen this inside autonomous research. In the Karpathy autoresearch example above, most of the gains arrived early, and additional experiments bought increasingly small improvement. That was a small test alongside a fixed instructor model, not a test of completely autonomous RSI. It doesn’t resolve the question. But it illustrates why removing the individual from an test iteration doesn’t, by itself, eliminate diminishing returns.

I do anticipate the feedback iteration to get stronger complete time. Better AI should rotate into improved at research. But according to our finest current data, reaching self-sustaining feedback requires a iteration approximately five to ten times stronger than today’s. Treating completely automated application investigation as already at that threshold is a significant leap, before we add the benefits of hardware improvements or financial growth. I could be wrong, but I’d akin to see evidence that autonomy brings adequate additional helpful discoveries to near that gap.

On hardware, I have several additional reservations. The example doesn’t explicitly contain the years it can obtain to rotate a part scheme into deployed hardware. The authors conversation bodily bottlenecks, and I’d akin to see manufacturing and building delays built into the predictions.

I additionally amazement how much former part advancement came from improved ideas, and how much depended on always additional costly factories and equipment. If we provision researchers too much credit for gains that additionally needed those investments, we could overestimate what faster AI investigation solitary would produce.

Even alongside those reservations, this is the most compelling document and example I’ve seen for combining feedback loops in software, hardware, and economics to comprehend how accelerated they could shove AI forward. I’m not convinced it establishes that a accelerated AI takeoff is imaginable under realistic conditions. More data could assistance us calibrate that judgment. But it gives us a helpful example for understanding what could happen beyond the application tier alone.

This is an crucial document that helps us example AI as part of a broader economics that power have larger feedback loops about it. I value it, and I’m glad they wrote it.

Back to contents

These estimates remainder on small data than I’d like. I power be putting too much importance on a few observations and reaching a comforting decision I desire to believe. We need improved measurements, shared frequently adequate to capture changes as they happen.

When OpenAI released its investigation data, Cheryl Wu welcomed the disclosure and pointed out how much was motionless missing. More tokens and experiments are helpful things to cognize about. We additionally need to see how they rotate into improved algorithms and additional capable AI.

X avatar for @cherylwoooo

Cheryl Wu@cherylwoooo

I value this as an first stage toward additional transparent reporting on RSI. But there is motionless much additional data we need to completely comprehend RSI. In particular, OAI disclosed several evidence concerning the conclusion compute usage, figure of experiments/researcher, and to a lesser…

X avatar for @kliu128

Kevin Liu @kliu128

Today we're releasing data on models accelerating investigation at OpenAI. Recursive self-improvement could be the most crucial contributor to AI capabilities complete the next few years, but by default it volition lone be seen inner a few frontier AI labs. Being transparent is more

11:25 PM · Sep 6, 2026 · 19.7K Views

9 Replies · 16 Reposts · 121 Likes

Figure 32. Cheryl Wu on OpenAI’s investigation data. Source.

Now Wu, Arjun Ramani, and Basil Halperin, alongside their colleagues at the Elasticity Institute, have written a tangible proposal: How to Measure RSI. It lists eight things the labs could portion to assistance answer these questions. Check it out.

Figure 33. Eight proposals for measuring RSI. Source.

I’d particularly akin to see how much helpful investigation all new example adds, holding resources approximately constant, and how that investigation translates into improved AI. That’s how we’ll study whether the iteration is getting stronger.

AI is already assisting build improved AI. It’s improving at a stupendous pace, and I anticipate that to continue. We already have narrow superintelligence in chess and Go. I anticipate increasingly superhuman achievement in parts of ceremonial math, coding, and cybersecurity, and any another verifiable domain anywhere machines can create training data and verify achievement at device speed. Those are mighty capabilities. That doesn’t average we’re near to super-intelligence for small verifiable, messier, open-ended activity - or to a broad ASI.

I’m skeptical of a accelerated takeoff to super-intelligence, but evidence matters additional than hunches. Let’s collect the data we need to get a clearer image of what’s happening. Including evidence that could alter our minds. If improved AI starts producing enough helpful research to create the next circular easier, I desire to know. If the gains keep shrinking, I desire to cognize that too.

Back to contents

Share

Other Article Hacker News
↑
Close Right Ads
Close Left Ads