Measuring AI ROI with a ‘Useful Intelligence per Dollar’ Scorecard

Chief financial officers and business leaders are asking a simple question: how do we get more value from our AI investment? Traditional software metrics such as seats sold or active users no longer capture what matters for AI. Instead, organizations should measure the actual work AI completes and compare the full cost of producing those outcomes to the value they create. The framework of “Useful Intelligence per Dollar” offers a practical way to evaluate whether AI spending is delivering increasing returns.

Measure the work that actually gets done

The first step is to focus on outcomes rather than model statistics. Count the customer issues resolved, code changes shipped, contracts reviewed, or hours reclaimed. As models handle longer context and more complex chains of reasoning, they can take on multi-step workflows: gathering data, reconciling inputs, and producing an output that moves work forward.

Start with a single workflow and define what “done” means in the system where the work happens. For support teams, that might be a ticket resolved. For engineers, a code change that passes tests. For legal, a contract reviewed to the required standard and on time. Finance teams preparing forecasts, for example, can often offload repetitive tasks—pulling the latest figures, reconciling sheets, rebuilding presentation slides—freeing analysts to focus on interpretation and decision-making.

Calculate the full cost of a successful task

Price per token is an incomplete metric. The true cost of completing a task well includes compute, how many attempts are needed, human review, employee time, and any rework. To calculate cost per successful task, add the full cost of completing the work, tally the tasks that met the quality bar, and divide the total cost by the number of successful outcomes.

That arithmetic explains why lower-priced tokens can still lead to higher total cost: cheaper models may need more retries, longer human oversight, or additional tooling to reach acceptable results. Conversely, a more capable model that gets the job right in one pass can lower the overall cost per successful task even if its per-token price is higher.

Vendors increasingly offer tiered model families so customers can match capability and cost to each workflow. The source material describes a three-tier family—Sol as a flagship, Terra balancing cost and performance, and Luna positioned as a fast, affordable option—so customers can choose the right tradeoffs. In one benchmark cited, Sol achieved a new high on a long-horizon engineering index while using substantially fewer output tokens than a competing model; another comparison showed higher task success at a lower estimated API cost for a specific engineering test set. Those kinds of trade-offs are what teams must weigh when optimizing cost per outcome.

Track dependability: when AI can be trusted to deliver

Dependability determines how much human oversight remains necessary and therefore directly affects economics. Adoption tends to progress from drafting to contextual reasoning to taking action. Each stage adds value but also requires clearer governance.

Teams should classify results into three buckets: “ready to use” when the output meets quality expectations as delivered; “needs correction” when edits or retries are required; and “needs escalation” when a human must intervene to finish the job. These categories provide operational insight beyond raw model accuracy by showing how much downstream work remains.

Before AI systems are allowed to take actions, organizations should set explicit boundaries: what data the system can access, which systems it may change, and when human review is mandatory. Security, privacy, compliance, and workspace controls are the foundation that lets organizations extend AI deeper into workflows while retaining oversight. The source highlights enterprise features that let organizations provide more context and access while enforcing policies and controls.

Evaluate whether each AI dollar buys more over time

The final piece is scale: does completed work grow faster than total cost as usage increases? Measure the same workflow over time, tracking how many tasks meet the quality bar, the total cost to complete them, and the resulting cost per successful task. If work completed rises faster than cost while quality remains stable or improves, the organization is achieving more useful intelligence per dollar.

Compute is central to this dynamic. Training compute builds future capabilities; inference compute delivers work today. Improvements across algorithms, hardware, routing, and product design increase the return on compute, producing better answers, faster performance, and fewer corrections. Those gains compound: better infrastructure accelerates research, research yields stronger models, and stronger products attract more usage and revenue, which funds the next cycle of investment.

The source also notes that bringing these layers together—platforms for end users, tools for developers, and enterprise deployments—allows improvements in one area to benefit every customer and product.

A practical scorecard for AI investment

Combined, the four measures—useful work completed, cost per successful task, dependability, and economics at scale—form a practical scorecard for evaluating AI spend. The objective is clear: enable people to do more meaningful work and spend their time on judgment, creativity, and higher-value activities. By measuring outcomes and the full cost to achieve them, organizations can make data-driven choices about which models, infrastructure, and processes deliver the best return on AI investments.

Source: