When chief financial officers around the world are asking the same question, OpenAI decided to take matters into its own hands and figure out the answer: exactly how much value have we gained from our investment in artificial intelligence? Over the past decade, the software industry has measured success by adoption—how many licenses were sold, how many active users there are, and how many licenses were renewed. But with AI, that yardstick no longer works. The real measure should be how much work the model actually gets done.
In a long article released on July 17, OpenAI laid out the core economic challenge facing business leaders: Is the value generated by AI tasks growing faster than the cost of producing them? To answer this, just looking at the cost per token is not enough. A cheaper model may have a lower price per token, but it might require more trial and error, take more time, and involve more manual review to get good results; a more expensive model may have a higher price per token, but it could complete the task correctly on the first try. The real cost is the total expense of achieving a successful outcome, compared to the value that outcome generates.

Thus, OpenAI introduced the ultimate scorecard for the AI era: Useful Intelligence per Dollar. This metric answers four questions: Is AI completing meaningful work? What is the cost of each successful task? Can people trust the results? As usage grows, does each dollar spent on AI create more value?
First, how much useful work has been completed. Value is only created when tokens are transformed into actual, directly usable outputs. The stronger the model, the longer and more complex the tasks it can handle: it starts maintaining context, performing multi-step reasoning, collaborating across tools, and continuously adapting along the way. The most practical approach is to lock onto a specific workflow, clearly define what "completion" means, and then measure the result within the system where the work actually happens. For support teams, completion means a customer issue is resolved; for engineering teams, it's a code change that passes tests; for legal teams, it's an accurately and timely reviewed contract. OpenAI gave an example of a finance team preparing for a forecast review: before the final decision, they need to find the latest forecast, import data into Excel or Sheets, identify changes, check tabs, rebuild slides, and verify all numbers align—a series of tedious tasks that can now be handled by ChatGPT Work, allowing the team to focus on what truly matters—what changed, why, and what comes next. This is the useful intelligence you get for every dollar spent in practice.
Second, how much did a successful task actually cost. The cost of AI tasks varies widely: a quick answer consumes little computation, while coding, research, and financial workflows often involve deep reasoning, tool calls, and extensive operations, consuming more compute and creating greater value. At the model level, the cost of a single successful task depends on price, actual compute used, and the probability of getting the correct result; for companies, total cost also includes employee time, manual review, retries, and rework. The formula is simple: add up the total cost of completing the work and divide by the number of tasks meeting quality standards. This explains why the lowest cost per token doesn’t always lead to the lowest cost per result— even for routine requests, if a cutting-edge model gives the right answer on the first try, the savings on retries, delays, reviews, and total compute may make it the most cost-effective choice.
OpenAI’s recently released GPT-5.6 was designed according to this logic, with three tiers: Sol is the flagship, Terra balances performance and cost, and Luna is the fastest and most efficient. This tiered approach provides customers with a starting point for optimization, but OpenAI emphasizes that the final choice of model should depend on the overall economics of the task—use Luna for high-throughput processes, Terra for deeper tasks, and Sol when stronger reasoning can achieve better results with fewer attempts. During the training of GPT-5.6, OpenAI aimed to increase the amount of useful output per token: in the Artificial Analysis programming intelligence index, the GPT-5.6 Sol achieved new best levels with maximum reasoning, while using 54% fewer output tokens than another leading model; in the DeepSWE v1.1 long-cycle engineering tasks, GPT-5.6 Sol reached a new high of 72.7%, surpassing Claude Fable5's 69.9%, with an estimated API cost reduction of 36.2%. Each generation of models should improve in two areas—higher efficiency makes old tasks cheaper, and stronger capabilities make entirely new types of work possible.
Third, how frequently AI completes tasks correctly, which is reliability. The adoption of AI typically follows stages of deepening: first, it assists in drafting, then finds context between tools and data, performs reasoning, and later takes proactive actions, handles exceptions, and runs full workflows, with humans only stepping in when necessary for judgment and control. Each step creates more value but also raises system requirements. Reliability itself is money—when results are accurate, reliable, consistent, and appropriately escalated, the time spent on review, correction, and rework decreases, reducing the cost of successful tasks and giving organizations more confidence in integrating AI into more critical processes.
OpenAI suggests teams track three types of results to measure this: directly usable, requiring corrections, or requiring escalation. This approach better reflects whether AI is truly reducing the workload required to complete projects than just measuring model accuracy. Reliability also requires clear boundaries: before AI moves from drafting to action, organizations must define what data the system can access, what systems it can use or modify, and when human review or approval is needed for an operation. Security, safeguards, privacy, and control form the foundation for deep usage, and ChatGPT Work is built on the security, compliance, and workspace management of ChatGPT Enterprise, allowing organizations to maintain oversight while enabling AI to access more valuable processes. Ability brings the first use, but reliability makes AI a component of the work itself.
Fourth, as usage grows, can each dollar invested in AI complete more work? Companies can track the same workflow over time to measure: count the number of tasks meeting quality standards, the total cost to complete them, and the cost per successful task. If the amount of work completed grows faster than the total cost, and quality remains the same or improves, then each dollar produces more value. Compute lies at the heart of this equation—it drives research and powers every AI task, determining product quality, speed, reliability, availability, and cost. Training compute builds future capabilities, and inference compute delivers current value, both ultimately translating into better customer outcomes.
More advanced models, more efficient inference, specialized hardware, higher utilization, smarter routing, and stronger product design can all improve compute returns. Customers feel these improvements: answers are more accurate, responses are faster, corrections are fewer, products are more reliable, and the cost of completing tasks is lower. These benefits compound—better infrastructure accelerates research, research leads to stronger and more efficient models, model improvements enhance products, and product growth drives adoption, learning, and revenue growth, which in turn supports continued investment in the next generation of research, compute, deployment, and security. OpenAI unifies all these elements under a single intelligent platform: users access it through ChatGPT and ChatGPT Work, developers build via Codex and APIs, and enterprises deploy it into real-world systems where work happens. Any improvement at any layer benefits all products and customers.
Combining these four metrics, this scorecard actually answers one thing: is the useful intelligence gained per unit of cost continuing to rise? Useful output tells us what AI can produce, the cost of each successful task tells us what it costs to achieve results, reliability tells us how much of the output people can confidently use, and scalability value shows whether each dollar and unit of compute achieves more over time. OpenAI sees its responsibility as making this equation better with each generation of models—stronger models, faster and more reliable results, and lower costs for the work customers truly need. When AI starts speaking in terms of useful intelligence per dollar rather than token counts, only then is value truly calculated.
