guide · ai

Stop Counting Tokens: The 4-Question Scorecard for AI

OpenAI's CFO proposed replacing 'seats sold' and 'tokens consumed' with a 4-question scorecard.

August 16, 2026 · By Alastair Fraser

A retro robot at a balance scale weighing a glowing trophy against a stack of coins.

Stop Counting Tokens: The 4-Question Scorecard for AI

One-line job: Replace the wrong AI metrics — seats sold, tokens consumed — with the four questions that actually measure whether AI is doing work. Audience: CFOs, VPs, and operators making AI investment decisions. Not for: Readers wanting a vendor selection framework. This is about measuring outcomes, not choosing models. Last verified: 2026-08-16 Evidence weight: documentation-verified

For two years the dominant AI metric at most companies was how much AI got used — seats sold, tokens consumed, daily-active-agent counts. That was the wrong question, and the industry has started quietly walking it back.

The new scorecard

OpenAI CFO Sarah Friar published the cleanest replacement in August 2026, and the AI Daily Brief episode that surfaced it walks through all four. The proposed scorecard asks, for each AI deployment:

  1. Did AI complete work that mattered?
  2. What did it really cost — including employee time, review, and rework, not just API spend?
  3. Was the result good enough to use without significant manual intervention?
  4. Did it help us make better decisions — both the specific decision in question and our capacity for future ones?

Three properties make this better than the metrics it replaces. It counts work, not consumption. It captures real cost, not sticker price. And it asks the question at the right unit — a specific piece of work, not an aggregate period.

Why the shift is real

This isn’t just one CFO’s essay. BCG’s separate August 2026 piece coined the parallel concept of “enterprise cortex” — the IP, business rules, and decision logic AI is now exposed to — and “cognitive load shifting” as the actual economic event. That’s a different vocabulary for the same observation that token-counting misses.

The corporate evidence behind the shift: in June 2026, Microsoft, Uber, Amazon, and Meta all walked back earlier org-level incentives to burn tokens. That’s not culture. That’s finance getting tired of an unmeasurable line item.

What to actually change

If you’re running AI deployments today, three things to update this quarter:

Replace activity metrics in any internal dashboard. “Active users” and “tokens consumed” tell you engagement, not value. Pair them with one of Friar’s four questions per use case.

Track task-crossover. OpenAI’s research on 800,000 ChatGPT messages found 43.5% of occupation-specific AI use crosses job boundaries. If your metric structure assumes clean occupational buckets, you’ll systematically undercount where the value actually lands.

Include rework in cost. The biggest hidden cost in most AI rollouts is the human review-and-fix loop. If you measure only API spend, you’ll see productivity losses appear as “training cost” and quietly kill the deployment’s economics.

What’s still missing

The honest limits of the scorecard: it measures whether the work was done, not whether the work should have been done. The “did it help us make better decisions” question is harder to score than the others, and probably needs quarterly review rather than per-task tracking. And like any framework, it can be gamed — answered “yes” reflexively for friendly cases.

Treat the four questions as a starting filter, not a finished scorecard.

The call

  1. Take Friar’s list and add it to your next AI deployment review. Two of the four questions require data you probably don’t collect yet; that’s the gap.
  2. If you’re reporting AI results up the chain, lead with work-completed rather than usage metrics. The trust layer matters: executives comparing AI to other investments expect outputs.
  3. Audit your organization for “tokenmaxxing” incentives still in place from the 2024–25 cycle. The cultural drift has started; the contracts probably lag.

Sources

#abs-guide#ai-agents#measurement

Submit a take

Have a different read on this? Drop a comment below — your email isn't published, and I read every one. Nothing leaves the site until I approve it.

Your email address will not be published. Required fields are marked.