Trend shift

AI cost comparisons move to task-level accounting

AlphaSense and NIST evidence suggests that AI cost depends on task structure, token usage, and end-to-end outcomes rather than API list prices alone.

An AlphaSense analysis, supplemented by independent NIST cost testing, moves AI cost analysis from model pricing to task-level accounting: agentic workflows can use up to 30 times more tokens than simple chat, while end-to-end cost rankings vary across tasks.

Usage can erase unit-price declines

AlphaSense argues that agents require more planning, tool calls, and intermediate steps, so total token usage can rise sharply even as the price per token falls. That means a model switch cannot be evaluated through input and output prices alone. Retries, context, tool calls, and human review need to be included in the same workflow cost.

Independent tests resist one answer

NIST’s end-to-end testing found DeepSeek V4 Pro cheaper than a comparable GPT-5.4 mini on five of seven benchmarks, but the relative cost range ran from 53% lower to 41% higher depending on the benchmark. An earlier 13-benchmark evaluation of DeepSeek V3.1 found comparable U.S. reference models 35% cheaper on average. Together, the results point to task structure as the decisive variable.

Procurement rules may change

If the pattern persists, enterprise AI budgets will move from asking how much a million tokens costs to asking how much it costs to complete a business task, including success rate, latency, human intervention, and data handling. The thesis would be falsified if one model or supplier consistently won most real production tasks under shared success criteria and complete usage accounting.

What to watch next

Check whether vendors disclose task-level cost, success, and retry data; record tokens, tool calls, human review, and latency on identical production workflows; compare total cost per task across sustained operating periods rather than API prices alone.

Sources