As companies shift from short chat interactions to longer, agentic workflows, token prices alone no longer indicate whether AI investments deliver business value. Leaders need visibility into who is using AI, which products and models consume capacity, and what outcomes those usages produce. This article sets out five practical steps to measure useful work per dollar, improve observability, select models based on end-to-end cost, operationalize governance, and fund AI investments as a portfolio.
Why token price isn’t enough
OpenAI’s objective is to make AI more accessible, capable, and affordable: from GPT-4 to GPT-5.4 the price per million tokens fell by 97%. GPT-5.6 continues this trend, delivering 54% fewer output tokens and 57% less time per task on the Artificial Analysis Coding Agent Index. Those improvements, however, do not by themselves prove that AI creates value for a business.
Leaders should focus on useful work per dollar: tasks completed, time saved, better decisions, and workflows that can scale. These measures reveal whether a rising bill indicates waste, productive experimentation, or a workflow becoming business-critical.
1) Visibility: measure who, what, and why
Enterprises need a shared view of AI usage: who is using it, which products or models, how much capacity they consume, and what kinds of work the usage supports. Without that visibility, interpreting a growing bill is difficult.
ChatGPT Work supports longer, multi-step tasks, so usage can vary significantly by workflow. Updated usage analytics and spend controls in the Admin Console help admins see adoption, credit usage, and spend by user, product, and model; track trends over time; identify emerging patterns; and determine whether usage reflects broad adoption, a power-user workflow, or a recurring business process worthy of further investment.
2) Multiple altitudes of insight to guide decisions
It is useful for leaders and admins to view data at different levels: high-level trends for strategy, mid-level views for team and function enablement, and detailed operational metrics for workflow optimization. Together, these views inform where to invest, where to coach, and where to set limits.
3) The lowest token price can increase total cost
A cheaper model may fail more often, require retries, or generate outputs needing correction. A more capable model can cost more per token yet reach an acceptable result faster, with fewer attempts and less human review.
Evaluate models based on the actual work they must perform. Use evals that reflect real tasks, including edge cases, and define “good enough” before testing. Then measure the full cost of reaching that standard: model and tool usage, attempts, completion rate, latency, and human review.
For priority workflows, track cost per accepted outcome. In customer support that could be a resolved case; in engineering, a tested change that passes review. Pair this cost with business value such as time saved, reduced cycle time, revenue protected, risk avoided, or capacity created.
Model choice is only part of the equation. Clear instructions, focused tools, reusable context, and explicit stopping conditions reduce loops and wasted spend. Match model and workflow to the task: use smaller or faster models when they meet the quality bar, and reserve frontier intelligence for complex, ambiguous, or high-stakes work.
4) Governance as the operational layer
Enterprise leaders should treat governance as the operating layer that determines which AI work can scale. Practically, this means defining what context ChatGPT can use, which tools it can access, what actions it may take, who approves higher-risk steps, and how additional capacity is granted when teams find valuable workflows.
This becomes especially important as organizations adopt plugins, connectors, Computer Use, and other frontier capabilities that operate across enterprise systems. ChatGPT Work provides centralized controls for access, approved context, connected tools, permitted actions, usage, and spend. Spend controls — workspace defaults, group limits, individual overrides, and review requests with project context — help leaders support high-value work without broadly raising limits.
For priority deployments, OpenAI’s AI Deployment Engineers can work directly with customers on evals, architecture, latency, reliability, and workflow design to improve performance and cost efficiency. Privacy and governance should be built in from the start: sensitive workflows need appropriate access controls, retention posture, compliance visibility, and approval paths before scaling. Where applicable, OpenAI’s enterprise privacy controls, including Zero Data Retention options, can help customers deploy AI in high-trust environments.
5) Fund AI as a portfolio and follow maturity
Enterprises should manage AI investments as a portfolio: broad access for everyday productivity, function-specific workflows that improve repeatable work, and a smaller number of strategic bets that leverage proprietary company context. The best candidates are workflows that repeat at meaningful scale, have clear ownership, and can be measured for quality, risk, and business value.
Funding should follow maturity:
- Discovery: test whether the model can handle the task.
- Validation: test representative cases against a clear quality bar.
- Production: fund integrations, controls, reliability, and change management to scale.
Shared capabilities — identity, trusted connectors, curated knowledge, evaluations, observability, model routing, and reusable agent patterns — should be funded centrally so each new workflow is easier and safer to launch.
Once a workflow proves value, match the product, capacity, and support model to its demand. ChatGPT Work provides ready-made capabilities for chat, coding, agentic workflows, connectors, plugins, Computer Use, and administration. Companies can extend that foundation with proprietary data, permissions, evaluations, and workflow logic where those elements create differentiated value.
For production workloads, the commercial structure should align with usage patterns: Guaranteed Capacity for production systems and agents that need access certainty; Scale Tier for predictable high-volume API workloads; and Batch API, Flex processing, or Prompt Caching for asynchronous work or repeated context.
For larger strategic deployments, OpenAI Frontier and Deployment Company can help enterprises build, deploy, and manage AI coworkers across enterprise systems. This approach lets leaders scale proven work with the right product, capacity, and support model instead of making each workflow rebuild its own infrastructure.
Closing thoughts
AI investment returns are determined not by token prices but by how much useful, scalable work is delivered per unit of spend. Combined transparency, task-focused evaluation, operational governance, and maturity-based funding create a practical framework for companies to control costs and invest confidently in AI.



