OpenAI Proposes ‘Useful Intelligence Per Dollar’ Scorecard For Enterprise AI Spending

OpenAI has introduced a framework intended to help companies measure the business value generated by artificial intelligence. The company argues that enterprises should evaluate AI based on the amount of useful work completed rather than traditional software metrics such as licenses purchased, active users or seats deployed.

The proposed measurement is called “Useful Intelligence per Dollar.” It is designed to determine whether the value of work completed by AI is increasing faster than the full cost required to produce successful results.

OpenAI said measurements such as cost per token do not provide a complete view of AI economics. A less expensive model may require multiple attempts, more employee oversight and additional corrections, while a more capable model could complete the same task successfully in a single attempt.

The framework therefore considers the complete cost of achieving a usable outcome. This includes model usage, computing resources, employee time, human review, retries and any rework required before a task meets an organization’s quality standards.

OpenAI’s scorecard focuses on four questions involving useful work, cost, dependability and value at scale. Together, the measures are intended to help chief financial officers and other business leaders determine whether AI investments are producing measurable operational returns.

The first measure examines how much meaningful work AI completes. Depending on the organization, this could include customer issues resolved, software changes successfully deployed, contracts reviewed, decisions improved or employee time recovered.

OpenAI recommends beginning with one clearly defined workflow and establishing what a completed task means within the system where the work occurs. A customer support organization could measure resolved cases, while an engineering team could track code changes that pass required tests.

A legal department could measure contracts reviewed accurately and within the required deadline. The objective is to connect AI usage with a practical business result rather than treating generated tokens or interactions as the final product.

OpenAI used financial forecasting as one example of a workflow that can involve substantial preparation before an executive decision is made. Employees may need to locate the latest forecast, transfer information into spreadsheets, reconcile multiple files, identify changes and rebuild presentation materials.

The company said ChatGPT Work can perform parts of that process, leaving finance professionals with more time to examine what changed, why it changed and what actions should follow. This represents the type of outcome OpenAI believes companies should measure when evaluating AI productivity.

The second component calculates the full cost of each successful task. Organizations can add all costs associated with completing the work and divide that total by the number of tasks meeting the required quality threshold.

More complex tasks involving coding, research, financial analysis or multiple tools may require greater computing resources. However, OpenAI argues that those workflows can still deliver stronger economics when they create substantially more value or reduce the need for repeated human intervention.

The company said this explains why the least expensive model on a per-token basis may not provide the lowest cost per completed outcome. A more advanced model could be more economical when stronger reasoning reduces retries, delays, reviews and total compute consumption.

OpenAI positioned its GPT-5.6 model family as a tiered system for matching capabilities and costs with different workloads. The family includes Sol as its flagship model, Terra as a balance between performance and cost, and Luna as the fastest and least expensive option.

A company could use Luna for high-volume and relatively straightforward processes while assigning deeper work to Terra. Sol could be selected when advanced reasoning increases the likelihood of completing valuable tasks correctly with fewer attempts.

OpenAI said GPT-5.6 was trained to produce more useful work from each token. The company reported that GPT-5.6 Sol with maximum reasoning achieved a new high on the Artificial Analysis Coding Agent Index while using 54% fewer output tokens than another leading model.

The company also said GPT-5.6 Sol achieved a score of 72.7% on DeepSWE v1.1, which evaluates long-duration software engineering tasks. OpenAI compared that result with Claude Fable 5’s score of 69.9% and estimated that GPT-5.6 Sol completed the benchmark at a 36.2% lower API cost.

The third part of the scorecard evaluates dependability. OpenAI said AI adoption often progresses from drafting content to locating information, reasoning across tools and eventually taking actions or completing workflows with appropriate human controls.

Organizations can assess dependability by separating outcomes into work that is ready to use, work that needs correction and work requiring escalation to a person. These categories are intended to show whether AI genuinely reduces the total effort needed to complete a task.

Accurate, properly sourced and consistent results can lower review and correction costs. Greater reliability can also give companies the confidence to use AI for more important and economically valuable processes.

OpenAI said organizations should establish clear boundaries before allowing AI systems to move beyond drafting and begin taking actions. Those policies should determine what information an AI system can access, which platforms it can use or modify and when human approval is required.

Security, privacy, safety and governance remain important parts of the economic calculation because failures can increase costs and limit adoption. OpenAI said ChatGPT Work builds on the security, compliance and workspace management capabilities available through ChatGPT Enterprise.

The fourth measure examines whether each dollar spent on AI produces more work as usage expands. Companies can track the same workflow over time and compare completed tasks, total costs, quality levels and the cost of each successful outcome.

An improving result occurs when successful work grows faster than total spending while quality remains stable or increases. This would indicate that each dollar of AI investment is generating a larger amount of usable output.

OpenAI identified computing capacity as a central component of this economic model. Training compute supports the development of future capabilities, while inference compute powers the work performed by AI systems for customers today.

More efficient inference, improved hardware, stronger algorithms, intelligent model routing and higher infrastructure utilization can all increase the return generated by computing resources. Customers experience these improvements through faster responses, better results, fewer corrections and lower costs for completed work.

OpenAI said these improvements can create a compounding cycle across infrastructure, research, models and products. Better infrastructure can accelerate research, while stronger models can improve products, increase adoption and generate revenue that supports further investment.

The company is bringing these capabilities together through a shared intelligence platform spanning ChatGPT, ChatGPT Work, Codex and its application programming interface. Improvements made at one layer can potentially increase the performance or efficiency available across multiple products and customer workflows.

OpenAI believes useful work, cost per successful task, dependability and value at scale provide a more practical scorecard for enterprise AI than adoption alone. The company’s broader objective is to help organizations complete more meaningful work while reducing the cost and human effort required to achieve reliable outcomes.