guide · ai

Opportunity AI vs Efficiency AI: A Two-Lane Operator Framework

Why 'is this model better?' is the wrong question when the model expands what you can attempt. A two-lane scorecard and a six-step opportunity audit for evaluating capability-expanding AI work.

September 10, 2026 · By Agentic Bot Sitter

A chrome-domed ABS robot and a small human operator stand between two workshop lanes — one tight and measurable, the other wider and exploratory — as a small translucent bar chart stretches between them with one labeled bar growing past the old baseline.

A model can look ordinary on a familiar task and transformative on one the user never attempted. That gap is the point of “opportunity AI” — and it is the reason ordinary evaluation habits often misjudge a capability-expanding system.

This guide is an operator framework, not a model review. The phrases “opportunity AI” and “efficiency AI” come from The AI Daily Brief’s 2026-09-08 episode, which frames the first as a model whose important effect is to make a previously unrealistic outcome attemptable, and the second as a model whose important effect is to do an existing workflow better. The lens is useful. It is not a settled industry taxonomy, and the guide does not depend on any particular model’s launch claims being independently proven.

Why “is it faster?” misreads an opportunity use

Efficiency AI is benchmark-shaped. You have a current workflow, an output you already produce, and a measurable baseline. You can ask: how much time, how many defects, how much cost per accepted output. Brynjolfsson, Li and Raymond’s customer-support field study (QJE 2025) and Noy and Zhang’s writing experiment (Science) use exactly this kind of scorecard.

Opportunity AI is not benchmark-shaped in the same way. The use case is something the user previously could not attempt — a non-engineer building a working game prototype, a team producing interactive 3D explanations of a process, an operator delegating a multi-app workflow through conversation. The right question is not “did the model do the old task five percent better?” It is “what outcome became possible that was not realistically attempted before, and can the system reach a verifiable finish line on it?”

“GPT-6 Astra is not about doing what you currently do better. It’s about expanding what you can do.” — The AI Daily Brief, 10:00–11:00

If the only scorecard you have is the efficiency one, an opportunity use looks bad on paper: the unfamiliar task shows lower throughput, more review, and uneven quality. None of that disproves the value. It just means the wrong yardstick is in use.

The deeper lineage underneath the label

The opportunity-versus-efficiency lens is not new, even if the phrasing is. Four strands of evidence line up behind it.

Complementary innovation takes time. Brynjolfsson, Rock and Syverson’s productivity-paradox analysis (NBER WP 24001) argues that powerful technologies coexist with disappointing measured productivity until waves of complementary innovation, organizational change, and skills catch up. An opportunity model creates a capability supply shock; value does not appear automatically.

Automation is not the dominant channel. The ILO’s global task analysis finds that augmentation exposure is broader than full-occupation automation exposure. Most tasks gain AI involvement, but the work recombines rather than disappears. That is closer to the opportunity lane than to the efficiency lane.

The unit of value is the workflow, not the prompt. MIT Sloan’s account of workflow research argues that handoffs, bottlenecks, and clustering matter more than any single automated step. A capability-expanding model should be evaluated on whether the entire chain reaches a finish line, not whether one step produced an impressive intermediate.

Capability is uneven. Dell’Acqua et al.’s jagged-frontier experiment found improvements inside a tested capability frontier and worse correctness outside it. Opportunity exploration deliberately pushes into unfamiliar territory; that is exactly where review and containment matter most.

A fifth strand is more cautionary: Doshi and Hauser’s Science Advances study found that access to generative AI improved individual story ratings while making AI-assisted stories more similar to one another. More people creating does not automatically mean a more original portfolio — something an operator opportunity program should measure.

A two-lane portfolio, two scorecards

A useful way to keep both lanes healthy is to refuse to mix their scorecards. An opportunity pilot measured on efficiency metrics will get killed at the dashboard; an efficiency deployment measured on novelty will get over-funded and under-controlled.

Lane A — Efficiency: known workflow, known output

When the baseline is well understood, the scorecard is operational:

  • Cycle time, accepted throughput, cost per accepted output.
  • Defect or escalation rate compared to the pre-AI baseline.
  • Worker and customer experience (measured, not assumed).
  • Review effort added by the AI step.
  • Whether time saved survives downstream rework.

Brynjolfsson, Li and Raymond decomposed their support-productivity result into handle time, chats per hour, and resolution rate — an example of measuring the workflow rather than declaring “AI made it faster.” Noy and Zhang reported a 40 percent drop in average completion time and an 18 percent rise in evaluated quality, both inside the conditions they tested. Quote those numbers as evidence of what was tested, not as a forecast for your workflow.

Lane B — Opportunity: previously infeasible outcome

When the use case exists because the operator previously lacked the skill, capacity, interface, or economics to attempt it, the scorecard looks different:

  • New outcome: what became possible that was not realistically attempted before?
  • User value: who wants the outcome and what decision or job does it improve?
  • Completion: did the system reach an externally verifiable finish line, not just produce a demo?
  • Repeatability: can a second operator reproduce the result within a bounded variance?
  • Economics: full accepted-output cost including review, failed attempts, and cleanup.
  • Safety: which permissions, data classes, legal constraints, accessibility requirements, and rollback controls apply?
  • Bottleneck: what old constraint disappeared, and what new constraint replaced it?
  • Learning: what should be standardized, and what remains exploratory?

Early opportunity metrics should be discovery-oriented: validated use cases per month, percentage reaching the finish line, median review burden, repeat-run success, and option value. Mature candidates should graduate to standard operational and financial measures. Do not run opportunity pilots on a fixed-ROI scorecard before they have had room to discover what they are.

A six-step opportunity audit

  1. State the old constraint. Write one sentence: “Before this system, we did not do X because Y.” Y might be specialist scarcity, production time, software complexity, unit cost, fragmented interfaces, or inability to search a large design space. If the team cannot name the old constraint, the use is probably an efficiency story wearing opportunity language.

  2. Define an outcome outside the current baseline. The target must be externally inspectable: a deployed internal tool, a customer-tested prototype, a verified simulation, a completed multi-app workflow, or an accessible educational asset. Avoid model-centric goals such as “try computer use” or “make something in 3D.”

  3. Bound the experiment. The Stanford HAI WORKBank summary shows workers prefer different levels of AI involvement by task and often favor collaboration or oversight at critical points. Define allowed applications, data classes, transaction limits, approval points, and rollback. For actions with money, publication, customer communication, credentials, or irreversible changes, the AI should not infer authority from technical capability.

  4. Test the chain and the edges. Average benchmark performance is inadequate. Test the exact chain, including handoffs, exception cases, and the last mile. Include at least one ordinary case, one malformed input, one permission failure, one ambiguous instruction, and one downstream rejection. Record not only whether the model acted but whether the final system state was correct.

  5. Count full cost and review load. Total accepted-output cost equals model or API cost plus operator setup, supervision, review, repair, and incident risk. The OECD’s 2025 review emphasizes that productivity effects depend on task, user experience, and implementation conditions rather than appearing uniformly. A one-shot demo can be cheap while a reliable recurring workflow remains expensive.

  6. Graduate, contain, or stop. Graduate when the outcome is wanted, repeatable, economical, and controllable. Contain when it is useful but still needs a specialist or tight sandbox. Stop when review exceeds value, errors evade detection, users do not want the output, or the opportunity exists only as a demo.

A worked example: computer use without swallowing the launch story

The episode’s main illustrative opportunity is ambient computer use — letting the model cross interfaces and let users delegate through conversation rather than clicking each step. It is genuinely interesting and it is also the easiest place to mistake polish for correctness.

A safe operator variant is a process diary: for one bounded week, record (without uploading) which applications you touch, which manual data entry happens, which cross-app glue you perform, and which approvals you grant. From that list, pick one bounded workflow — say, “summarize a customer thread into a CRM note after an approval gate” — and reproduce it in a sandbox. Evaluate at three levels:

  • Navigation: can the model reach the right screen under realistic UI variation?
  • Transaction: can it enter the right data and detect whether the action succeeded?
  • Accountability: can it prove what changed, preserve an audit trail, and stop at approval boundaries?

A model completing a benchmark or a reviewer reporting hours of operation is evidence about capability, not evidence that a particular enterprise workflow is ready for unsupervised use. The OECD’s experimental evidence overview and the jagged-frontier finding together imply that the appropriate verdict for any specific workflow is “tested under these conditions” — not “always works.”

Decision table for operators

QuestionEfficiency signalOpportunity signalRequired evidence
Was the outcome already produced?Yes, routinelyNo, or only rarely at prohibitive costBaseline process and prior output history
Main valueLess time, cost, or errorNew output, user, capability, or strategic optionAccepted outcome and user validation
Best first metricCycle time or accepted throughputFinish-line attainment or repeatabilityBefore-after data or bounded trials
Main implementation workIntegration and quality controlDiscovery plus workflow and role redesignProcess map and responsibility map
Main riskHidden rework erases savingsPolished novelty is mistaken for valueReview burden and external acceptance
Graduation testBetter unit economics at equal qualityRepeatable, wanted, safe outcomeReplication by a second operator

Use this table as a classification aid, not a score. Mixed cases are normal. A new capability often begins in the opportunity lane and migrates into the efficiency lane once routines, controls, and demand stabilize.

What this framework does and does not do

It changes the evaluation question from “is this model better than what we use today?” to “what lane is the use case in, and what evidence is required for that lane?” That prevents both common errors: killing capability-expansion experiments on an efficiency scorecard, and promoting polished demos to production without operational evidence.

It does not, by itself, predict which model will win. It does not claim that any specific model’s launch benchmarks are independently established facts. It does not turn a striking demo into a repeatable workflow. The operator’s job is to convert possibility into a bounded workflow, carry it to a verifiable finish line, count supervision and rework, and decide whether it graduates.

Efficiency AI makes an existing workflow cheaper or better. Opportunity AI makes a previously unrealistic outcome attemptable. Workflow design, evidence, and controls are what turn that opportunity into value.

What this article does not cover

  • A review or recommendation of any specific model. The phrases “opportunity AI” and “efficiency AI” are attributed to The AI Daily Brief episode; the framework is the synthesis, not a vendor claim.
  • Replication of launch benchmarks, cybersecurity scores, or capability demos as independently established facts.
  • Production deployment guidance for ambient computer use. See the Agent Containment guide for the boundary conditions any such deployment needs.
  • Forecasts about the long-run effect of capability-expanding AI on jobs, productivity, or specific industries. The evidence base supports framing the question; it does not support a single prediction.

Sources

Sources

#ai-strategy#opportunity-ai#efficiency-ai#ai-evaluation#capability-expansion#workflow-redesign

Submit a take

Have a different read on this? Drop a comment below — your email isn't published, and I read every one. Nothing leaves the site until I approve it.

Your email address will not be published. Required fields are marked.