Business

Measuring AI ROI: 'Useful Intelligence per Dollar' as a practical metric

Chief financial officers are asking whether AI delivers more value than it costs.

Measuring AI ROI: 'Useful Intelligence per Dollar' as a practical metric

The question CFOs ask is straightforward: how do we get more value from our AI budget? Historically, software success was measured by adoption — seats sold, active users, renewed licenses. Measuring the value of AI requires a stronger yardstick: the work actually completed.

The central economic question is whether the value of work AI performs grows faster than the cost to produce it. Answering this means going beyond simple metrics like cost per token. A model with cheaper tokens may need more attempts, more human review, or more time to reach a good result. A more capable model with costlier tokens might finish the task in one pass. What matters is the full cost to produce a successful outcome and the value that outcome creates.

"Useful Intelligence per Dollar": four core questions

Consider the metric "Useful Intelligence per Dollar." It is intended to answer four practical questions:

  • How many customer issues did AI help resolve?
  • How many code changes did it help ship?
  • How many contracts did it review?
  • How much time did it return to people, and how many decisions improved because the right context was available when needed?

Tokens create value when they convert into usable work. As models become more capable, they can handle longer, more complex tasks: maintaining context, multi‑step reasoning, operating across tools, and adapting as they proceed.

Start with one workflow

A practical way to begin is to pick a single workflow, define what "done" means, and measure that outcome where the work takes place.

  • For a support team, done might mean a customer ticket resolved.
  • For engineering, a code change that passes its tests.
  • For legal, a contract reviewed accurately and on time.

Take a finance team preparing for a forecast review: much work happens before a decision — locating the latest forecast, moving data into Excel or Sheets, spotting changes, reconciling tabs, rebuilding slides, and verifying totals. Tools like ChatGPT Work can take on many of those steps, giving the team more time to focus on the important questions: what changed, why, and what should we do next. That is useful intelligence per dollar in practice: more work completed more quickly, while people spend their time on judgment, creativity, and expertise.

What does it cost to complete work well?

AI tasks vary widely. A quick answer may require little compute; a coding, research, or finance workflow may involve deeper reasoning, tool use, and many actions. These complex tasks can require more compute but can also create much more value.

At the model level, cost per successful task depends on price, compute used, and the likelihood of reaching the correct result. For a business, full cost also includes employee time, human review, retries, and rework. That is why the lowest price per token does not always yield the lowest cost per outcome. A frontier model can be the best value even for routine requests if it produces the right answer in one pass, reducing retries, latency, review, and total compute.

Tiered model families and practical choices

A tiered model family gives customers more ways to optimize this equation. OpenAI’s recently released GPT‑5.6 (announced "last week" in the source) has three tiers: Sol (flagship), Terra (balance of performance and cost), and Luna (fastest and most affordable).

These tiers provide sensible starting points: Luna for fast, high‑volume workflows; Terra for deeper work; Sol when stronger reasoning yields the best result with fewer attempts.

GPT‑5.6 was trained to produce more useful work per token. On the Artificial Analysis Coding Agent Index, GPT‑5.6 Sol with max reasoning set a new state of the art while using 54% fewer output tokens than another leading model. In the DeepSWE v1.1 long‑horizon engineering tasks comparison, GPT‑5.6 Sol reached 72.7%, versus Claude Fable 5’s 69.9%, at a 36.2% lower estimated API cost.

Dependability and governance have economic value

Dependability — results that are accurate, well‑sourced, consistent, and escalated appropriately — reduces the time people spend reviewing, correcting, and repeating work. Successful tasks cost less, and organizations gain the confidence to use AI in more important workflows.

Before moving AI from drafting to taking action, organizations should define clear boundaries and controls: safety, security, privacy, and governance are the foundation for deeper use. People need to understand how the system behaves, how their data is handled, and how its actions are governed.

ChatGPT Work builds on the security, privacy, compliance, and workspace management foundation of ChatGPT Enterprise, enabling organizations to grant AI more context and access to valuable workflows while maintaining oversight.

Measure, scale, and compute

Teams can make progress concrete by tracking three outcomes over time:

  • Number of tasks that met the quality bar,
  • Total cost of completing them,
  • Cost per successful task.

If completed work grows faster than total cost while quality holds or improves, each AI dollar is producing more value. Compute sits at the center of this equation: it powers research and every completed task, shaping product quality, speed, dependability, availability, and cost. Training compute builds future capability; inference compute delivers useful work today. Both should translate into better outcomes for customers.

Improvements in models, more efficient inference, purpose‑built hardware, higher utilization, smarter routing, and stronger product design all raise the return on compute. The gains compound: better infrastructure accelerates research; research yields more capable and efficient models; better models improve products; better products drive adoption, learning, and revenue; and that revenue funds the next generation of research, compute, deployment, and safety.

OpenAI brings these elements together on a shared intelligence platform: people use ChatGPT and ChatGPT Work, developers build with Codex and the API, and enterprises deploy AI into the systems where work happens. When one layer improves, every product and customer can benefit.

Four measures that show progress

Taken together, four measures indicate whether useful intelligence per dollar is improving:

  • Useful work: what AI produces;
  • Cost per successful task: what it takes to reach outcomes;
  • Dependability: how much of the work people can confidently use;
  • Value at scale: whether each dollar and each unit of compute accomplish more over time.

The objective is AI that helps people do more meaningful work, make better decisions, and spend more time on tasks that require human judgment and creativity. The challenge for providers and customers alike is to improve both sides of the equation — capability and cost — with each generation so AI becomes progressively more useful for more people and organizations.