A new Economic Research paper evaluates Codex’s economic potential at the frontier and finds that agentic AI is changing the unit of knowledge work: moving it from short, self‑contained chatbot interactions toward delegated tasks that run for minutes or hours.
What does “agentic” usage mean?
Chatbot interactions are typically brief and self‑contained: a user asks a question and the model replies. Agentic systems, by contrast, can operate independently for minutes or hours, orchestrating tool calls, interacting with environments, and iterating toward solutions. As a result, agents are rapidly becoming the most powerful AI tool for productive work.
Public rollout and OpenAI’s internal experience
Over the past year OpenAI’s internal experience mirrored this shift. In the months after Codex was released publicly, ChatGPT remained the default AI tool for work. Through August 2025, the average OpenAI worker spent less than 10% of their tokens on Codex. By spring 2026, however, every department — including non‑technical teams such as Legal and Recruiting — had adopted Codex as their primary AI tool. The paper’s authors suggest this pattern indicates what work may look like more broadly as agentic tools become more capable and accessible.
Growth trends and concrete figures
Codex adoption expanded along with its capabilities: as stronger models and new product features arrived, Codex took on an increasing set of productive, long‑horizon tasks. The paper documents four key trends across individual users, organizational users, and OpenAI workers over the past year:
-
Nearly a quarter of all Codex requests were for tasks estimated to take a person more than one hour. As Codex’s capacity for independent long‑context work improved, users shifted from short interactions to harder tasks with longer horizons.
-
Using model‑estimated human time thresholds (more than 30 minutes, more than 1 hour, more than 4 hours, more than 8 hours), the share of individual users who made a request corresponding to work taking more than 30 minutes rose to 80.6% from December 2025 to May 2026. The share making requests estimated to take more than one hour rose to 70.2%. Requests estimated to take more than eight hours grew fastest from a low base.
-
Agentic usage increase appears in daily Codex runtime. Among OpenAI’s daily active users, the heaviest users asked Codex to run many hours of agent work in a single day. By June 2026, users at the 99th percentile regularly generated more than 60 hours of Codex agent turns per day, distributed across multiple parallel agents. As Codex became more powerful and parallelizable, users moved from asking for single answers to orchestrating multiple agent tasks during the day.
-
Adoption timing differed by function. Engineers at OpenAI adopted Codex first and gradually: the average engineer shifted the majority of their OpenAI product usage to Codex by December 2025, and today the average engineer generates 99% of their output tokens with Codex rather than ChatGPT. Legal, finance, and recruiting crossed to majority Codex use around April 2026, but their transitions were faster; the average lawyer or recruiter now generates more than 85% of their output tokens on Codex.
Depth and intensity over six months
The last six months saw deeper and more intense Codex use at OpenAI. Among active internal users, combined output tokens rose sharply by department: Research experienced the largest jump — by June 2026 median use was 56 times higher than in November 2025. Customer Support rose 32‑fold and Engineering 27‑fold, while Legal grew more slowly but still reached 13 times its November level.
Taken together, these patterns show how Codex has changed how OpenAI uses AI for productive work: across the company, users are switching from chatbots to agents as their primary form of AI interaction and deploying exponentially growing amounts of agentic labor.
Expansion beyond developers
Codex use began with developers across all user groups — OpenAI, organizational, and individual users — which is unsurprising for a tool that started as coding‑oriented. As Codex broadened toward general knowledge work, adoption among non‑developers rose even faster. By early June 2026, weekly non‑developer individual users had multiplied 137‑fold since August 2025; non‑developer organizational users increased 189‑fold; and non‑developer OpenAI users increased 12‑fold, likely because this group already started at a well above average level.
This does not imply every non‑developer uses Codex like an engineer; rather, more non‑developers are using Codex for some form of agentic work.
Effects on workflows and skill value
Codex enables non‑technical departments to accelerate workflows that were previously bottlenecked by technical expertise. A heat map comparing inferred occupations within OpenAI to types of Codex outputs shows engineering and coding dominate data science and research, while knowledge work is the largest category for finance, business operations, marketing, operations, and other departments.
Agentic tools can also expand what an individual worker can do: over one‑quarter of Codex work performed by workers in business functions was engineering or coding. Agents can lower the cost of moving across task boundaries and help workers perform adjacent tasks that would previously have required more specialized technical support.
This expansion matters for firms redesigning workflows, for employees deciding which skills gain value, and for policymakers and researchers seeking to understand AI’s impact on the labor market.
Conclusions and methodological notes
The paper shows how frontier users adopt capable agentic tools like Codex. Its results illustrate what happens when people have broad, low‑friction access to capable agents: as the tools improve, people use them for longer, more complex, and more cross‑functional work. The authors argue this is likely to describe the future of work.
Methodological notes: task horizons were estimated using an LLM‑as‑judge with access to Codex transcripts. The thresholds are model‑estimated and should be treated as directional rather than exact. The individual‑level results are based on queries from a random 0.1% sample of users.



