If AI feels like it is accelerating, that's likely correct. Leading American labs have been releasing better models more rapidly, although government interventions have already restricted access to two of the most powerful models, Claude Fable and GPT-5.6. The story is not only about release cadence: the evidence indicates accelerating capability gains as well, even though the frontier remains jagged and AIs are still weak in many areas.
Measurements and real work tests
Several assessments attempt to quantify how much human labor a single AI prompt can substitute. Notable examples include METR and the UK government–affiliated AI Security Institute, both estimating programmer-hours equivalent to AI output. GDPval compares AI performance to human experts across fields using professional judges. These metrics are all improving at better-than-exponential rates.
Epoch recently reported that Opus 4.7, running autonomously for 14 hours, produced a software package that would have required 2–17 weeks of human engineering work; that run cost $251 in tokens. AI systems still fail some tests and are not always cheap to run, but their capabilities are improving very quickly. In separate experiments, Fable worked autonomously for nine hours on very complex software projects that would have required a human team well over a week.
Two model tracks: closed frontier and open weights
The discussion has focused on frontier models—the highest “intelligence” systems—produced by three American companies: Anthropic, OpenAI, and Google (though Google has not released a major new model in some time). A second track consists of Chinese open-weights models that generally lag the frontier by about 6–12 months. Because these models are open-weight, anyone can use or modify them after release, which makes them relatively cheap to operate. They are also climbing an exponential improvement curve, but behind the closed U.S. models.
This separation shows up in tests like AA-Briefcase, which simulates a complex, multi-week consulting engagement requiring many kinds of analysis. Abstract graphs have limits, though: they can hide how jagged the frontier is and the fact that open-weights models, while impressive, do not always achieve the benchmark-predicted performance.
What matters in practice — use cases and judgments
To understand real capability, you have to try AI on specific use cases and rigorously evaluate how well systems perform in the areas that matter to you. As a playful example, the author created a test where models build an interactive simulation of a harbor evolving over time, illustrating differences between models in design, style, and judgment. Those hard-to-benchmark factors become increasingly important as systems handle longer tasks.
Changing usage: from chatbots to agents
As models can execute longer workflows, the human approach to using AI is shifting. Historically, the dominant pattern was co-intelligence: you prompt the AI, check its output, and then prompt the next step. With careful prompting and human oversight, AIs could be guided through complex, multi-step tasks. That approach remains useful, but long-running, smart, self-correcting AI systems reduce the need for constant human intervention and demand a different workflow.
Agents add extra machinery compared with chatbots: harnesses that give the AI access to tools and an environment to act in, and apps built for agents such as Claude Code or OpenAI's Codex. A good harness or agent app can further amplify an already improving model.
Therefore, work is increasingly about assigning tasks to agents rather than jointly stepping through tasks with chatbots. A joint study by OpenAI and academic economists shows the speed of this shift within OpenAI itself. Crucially, it's not only coders who adopt agents—legal, HR, and other non-technical functions have taken them up at nearly the same rate. OpenAI may serve as an early indicator for broader workplace trends.
Managing AI and the role of expertise
Inside OpenAI, about a quarter of employees run at least four agents concurrently each week. As coding is offloaded to AI within specialized harnesses and apps, other roles begin to resemble coding in some ways, and they do it well. A separate study of Claude Code users found software engineers had a similar success rate to other professions when using Claude Code on coding tasks.
What mattered more than job title was domain expertise: the deeper a user's experience in a domain, the more successful they were at getting useful output from Claude for that domain. In other words, we are moving from a world where non-experts use chatbots to fill gaps toward one in which experts use agents to get work done. The most effective posture is to think and act like a manager of agents.
A moment inside an exponential
Being on an exponential means each change over a fixed window is larger than the previous one. If an organization wrote an AI plan at any time before winter 2025, it likely described systems able to do a couple of hours of work with fairly high error rates. A few months later, a single prompt can yield sixteen hours or more of work. That is why AI often feels leap-like: a continuous curve produces a series of doublings that register as shocks.
This dynamic helps explain the turbulence around AI better than narratives about mere hype. AI can move from non-threatening to a genuine cybersecurity concern quickly, prompting sudden policy responses. Markets can ignore the risk that AI will undermine a business model until it suddenly can, producing large stock swings. These lurches are read as signs of an immature field, but the instability is a natural result when institutions that operate at human or committee speed try to follow a capability curve that is not human-paced. As long as some form of exponential progress continues, and for as long as that lasts, the gap between capabilities and institutional responses will likely widen.



