Safety

AI-generated text

Lossy self-improvement and the labs: why rapid agent scaling may feel like a singularity but likely won't be one

Frontier AI labs such as OpenAI and Anthropic have begun deploying thousands of concurrent agents and scaling inference capacity, which is raising alarm inside those organizations and in parts of the AI safety community.

Lossy self-improvement and the labs: why rapid agent scaling may feel like a singularity but likely won't be one

In recent months, frontier AI labs — notably OpenAI and Anthropic — have started deploying thousands of concurrent agents to handle internal research, engineering and product tasks. That shift is causing many employees to update their expectations about the pace of AI progress and related risks.

The author’s core view is that the frontier labs and the competitive San Francisco AI culture amplify any AI worry. That amplification raises public awareness — sometimes because fear attracts attention — but overstating timelines or severity can have harmful second-order effects. The author recalls loud AI safety debates in 2023–2024 that forecast near-term harms which did not materialize on the predicted schedule.

Cultural amplification and the “singularity soon” narrative

Richard Ngo’s summary is cited: a large portion of the AI safety community is orienting toward futures where an intelligence explosion happens within a few years. Ngo suggests these expectations may be directionally correct relative to outsiders (i.e., faster progress than many expect) but factually wrong in their timings: we likely will not have superintelligence in the next eight years, yet progress may feel rapid enough that short-timeline proponents appear right.

The author balances this cultural preconditioning with the possibility that labs have seen genuinely alarming private breakthroughs. His default is that much of the current concern stems from scaled agents being deployed, though he states high uncertainty if foundational, imagination-driven breakthroughs exist privately.

Lossy self-improvement as an alternative scenario

The author proposes “lossy self-improvement” as an alternative to full recursive self-improvement (RSI). Key elements:

  • Automatable research is often narrow, and scaling laws’ exponential costs make massive net acceleration hard.
  • Diminishing returns from more parallel agents are real.
  • Resource bottlenecks and political constraints matter greatly when building strong LLMs, and AI can only partly accelerate those domains.

Under this view, agent swarms and inference-time scaling will create substantial efficiency gains without necessarily producing a near-term, universal intelligence explosion.

Lessons from podcasts and lab measurements

The author references Dwarkesh’s podcasts with Noam Brown and a panel with John Schulman, Beren Millidge and Charlie O’Neill. The Noam episode highlighted how much short-term acceleration mass inference capacity can produce: labs will throw many agents at clearly defined, measurable problems. At the same time, the author doubts labs can indefinitely allocate a constant share of rising compute budgets to internal R&D, especially as firms contemplate IPOs and face greater economic scrutiny.

The trio’s debate reinforced a practical point: current techniques work for problems we can clearly state, but they do not deliver magical generalization to unknown, harder problems in most partially verifiable domains — mathematics being more of an exception than a rule.

Timelines and capability thresholds from the panel

The author collected the panel’s responses to questions of the form “when will AI reach X ability?” (timelines are relative to the interview date):

  • Drop-in remote worker for broad white-collar work over a month:

    • Charlie O’Neill: ~1 year with programmatic access to workplace tools; ~2 years if it must operate via a browser (ordinary white-collar work, not highly creative research).
    • Beren Millidge: ~3 years for full generality; 80–90% coverage sooner. Main uncertainties are online learning and long-tail tasks.
    • John Schulman: ~1 year for an "okay" version with uneven capabilities improving over time.
  • 10× productivity uplift for AI researchers:

    • Charlie O’Neill: 5–10 years; bottleneck is absorbing information and deciding which experiment to run next.
    • Beren Millidge: finds John’s ~2-year estimate plausible but gives no independent timeline; assumes AI can run successive experiments and learn from feedback.
    • John Schulman: ~2 years.
  • AI surpassing top human experts across all computer-based work, including multiyear projects (ASI):

    • Charlie O’Neill: 5–10 years, citing memory and context-length limitations.
    • Beren Millidge: ~5 years in areas the labs focus on, possibly longer for literally every domain; no firm universal timeline.
    • John Schulman: 3–4 years; spatial/physical fields may take longer and onboarding/long-horizon learning are required.

The author stresses that intelligence is jagged and thresholds must be phrased as specific, measurable tasks. LLMs’ intelligence differs from human intelligence, so AIs will not flip discrete human-shaped roles overnight; diffusion is gradual and a long tail of tasks persists.

Which tasks get automated first, and where humans remain crucial?

Much of scientific work is communication and norm-setting among colleagues, the author notes — elements that are harder to automate. Agent swarms are likely to excel at clear, verifiable problems. In this sense, RSI is more helpful for efficiency (cheaper, more scalable inference) than for expanding peak intelligence.

However, scaling laws show that achieving linear improvements in intelligence often requires exponential increases in compute and resources. RSI may make LLM serving dramatically cheaper, which could accelerate trends that already show LLMs becoming exponentially cheaper at a given intelligence level. Economically, labs will face pressure to increase margins or compete on price; Jevons paradox dynamics may favor strong business models even as per-unit prices fall.

The author quotes John Schulman on the difficulty of automating post-training work: many different judgment areas determine how a model should behave and those are hard to fully automate; post-training mistakes can be invisible to benchmarks. These post-training and environment-construction tasks remain uniquely challenging for present LLMs.

What labs report and the current state

OpenAI and Anthropic have shared internal measurements related to RSI. For now, the author’s read is that the biggest automation wins inside labs are in software engineering, log monitoring, experiment management and other routine but nontrivial tasks. The Claude Fable 5.1 & Mythos 5.1 System Card notes: “We believe that internal usage of recent AI models has been a key factor in maintaining the current rate of progress, but we do not yet see clear signs of dramatic acceleration beyond that rate.”

The author argues the hardest exponential to move is peak intelligence. Early RSI effects are likely to look like massive scaling and diffusion of inference-time compute into AI research and related activities, with plenty of low-hanging fruit and substantial economic transformation potential. That could also unlock resources for broader AI diffusion, which is a crucial bottleneck to realizing many benefits.

Conclusion: current baseline and remaining uncertainty

Given available evidence and the cultural dynamics inside frontier labs, the author keeps “lossy self-improvement” as his baseline trajectory: scaled agents and cheaper inference will produce significant productivity and economic change, but the classic intelligence explosion that some fear does not appear inevitable in the short term. Still, he emphasizes that things in AI can change quickly and new evidence could warrant revising these expectations.