guide · ai

Judgment Models: When a Decision Costs $0.042, You Stop Sampling

Jev made typed decisions cheap enough to run on everything. The operator shift, the four-question fit test, nine disclosed failure modes, and why confidence is not correctness.

September 28, 2026 · By Alastair Fraser

Retro-futurist editorial illustration: a tall chrome-domed coil-limbed robot holds one small blank unmarked card up between two coil fingers and reads it, beside a low workbench carrying a single stack of blank cards, with a much smaller human operator standing beside the bench, in deep crimson and electric blue on warm cream.

Sampling is what you do when judgment is expensive. TypeSafe AI’s Jev (jev-1.13.0), launched 2026-09-15 in early access, prices a judgment at $0.042 per million input tokens with output free (vendor-published), against the $0.20-to-$10 per million its own table lists for existing LLMs — two orders of magnitude cheaper. At that price you stop sampling and start checking everything.

Jev is the first commercial model in a category the vendor names “System One models” and independents call judgment models — two weeks old, nothing settled. The AI Daily Brief, in the episode behind this guide: “A lot of the unique value of Jev is not just being able to do a thing, it’s being able to do a thing at such scale that it actually becomes a difference in kind rather than a difference in scale.”

What a judgment model is

A judgment model maps unstructured state — a string, JSON object, or array of text values — plus typed questions to typed answers with probabilities attached. It has no text-generation head. “Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out,” Diogo Almeida writes in the launch post. The interface is three primitives (vendor docs):

  • Choice — one from a closed list of up to 255 options. Returns the winner, a probability per option, and a 0–1 confidence.
  • Score — a position on an ordered scale of 2 to 10 levels you describe in words. Returns the probability-weighted mean, which can fall between levels, plus legend, probabilities, and confidence.
  • Noul (short for Bernoulli, per the CEO on Hacker News) — a yes/no. Returns one probability in [0, 1], no separate confidence field.

Questions in one request share the state, are evaluated independently, and run in parallel: latency is roughly flat in question count. There is no prompt caching (“Nope, it’s always the same input token cost,” a TypeSafe engineer on Hacker News), so batch questions into one call.

The category is a role, not an architecture: “The term describes a role; it does not name one universal architecture” (Dylan J. Harris). A distilled classifier, a reward model, or a cheap task-specific model plays the same role; Jev is the first commercial instance, not the definition.

The architecture that recurs

The same shape recurs across the independent write-ups: LLM proposes options → judgment model evaluates → deterministic code executes. The output branches in code — a probability compares against a threshold you set — which makes a cheap judge a natural gate in agent loops (loop engineering) and a natural trigger for proactive agents.

The numbers, with their labels

All figures below are vendor-published early-access numbers that “can change without notice.” Pricing: $42 per billion input tokens ($0.042 per million); output free. Rate limits: 250,000 tokens per second and 1,200 requests per minute; over either returns 429. Context: 64k tokens per request, and a tighter 32k for state plus the longest single question. Text input only. Endpoint: POST /v1/systemone. Aliases jev-latest (SDK default) and jev-preview both resolve to jev-1.13.0.

Hold the multiples tightly. The vendor quotes 70–500 ms end-to-end and, for System One-shaped queries, “This can range from 40x-200x faster.” Its homepage’s 193.6x faster and 444.6x cheaper are self-reports from four published workflows; the launch post says the vendor “expect[s] that these are on the higher end of real world gains.” Independent testers measured roughly 3–11x against fast models; one clocked Jev at p50 176 ms against a 27B model on Cerebras at 215 ms (learnjev.com’s audit). The vendor’s own workflow eval scores Jev 67.8% — tied with Sonnet 5, 6.3 points behind the leader — at roughly 293x lower cost and 195x faster: read that as good enough and radically cheaper, not better. See how to read an AI capability claim.

Jev never makes type errors — schema matching is guaranteed. The launch post’s “can’t hallucinate” claim is qualified by the vendor’s own footnote: that number is “not empirical.” Wrong-but-valid values still happen.

Access: within three days Jev shipped on OpenRouter, the Vercel AI Gateway (typesafe-ai/jev), and Cloudflare’s gateway products (typesafe/jev), alongside the waitlist-gated console. Secondary sources disagree on which Cloudflare product carries it — some name AI Gateway, others Workers AI — so check the Cloudflare listing before you plan around it. US servers only; zero data retention is an enterprise option, not a default. The vendor concedes it cannot prove the price is not subsidised; designs that only work at $0.042/MTok carry that risk.

The four-question fit test

Run the job through four questions first (from the AI Daily Brief episode behind this guide); failing any one disqualifies it.

  1. Can you write the possible answers in advance? Categories, a yes/no, a scale. If the answer is a sentence or a calculation, stop.
  2. Is there a pile or a stream? Volume pays for the setup.
  3. Is a wrong answer cheap or easy to catch? If being wrong is expensive, gate it or escalate.
  4. Can the evidence be handed over as text inside the token budget? 32k for state plus the longest question; 64k for the request.

One judgment per question: “is this a good lead?” is five judgments in a trench coat — split it, weight the parts in code. Give every Choice an honest “none of the above.”

What practitioners reported

Six categories appeared in the first ten days; each number below is a practitioner report, not an audited benchmark. Matthew Berman ran 724 live ads through 12 questions each: 40 seconds, 9 cents. Borja ran a 586-page SEO audit as 8,790 yes/no internal-link calls: 45.1 seconds, 21 cents. Daniel Son cut agent context by 88% of tokens with skill classification and selective injection.

Where it breaks

The vendor’s jev-1.13 jaggedness page (reviewed 2026-09-17) lists nine failure modes: literal reading — “Jev answers the question you wrote, not the one you meant”; math and numbers (“Jev is not a calculator”); date and time comparison; indirection; large state full of irrelevant detail (context rot); adversarial content; contradictory instructions and criteria; common-sense structural invariants; and generation.

Independent friction:

  • Confidence is not correctness. On a Choice, confidence measures how peaked the distribution is, not accuracy.
  • Calibration is the vendor’s central claim and is not independently verified. Vendor docs: “Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct.” Community measurement (not peer-reviewed, no vendor confirmation) found Jev calibrated on one dataset (ECE 0.0204) and overconfident on another (ECE 0.0936); in another study, rows at confidence 0.9+ were accurate only 72.2% of the time. Measure it on your own data before gating on probabilities.
  • No explanation comes back, only the distribution. Simon Willison calls Jev “a regression even further towards black box machine learning systems.”
  • State is an injection surface. The vendor documents that state is not treated as hostile by default: user-supplied text can move the answer. Keep permissions and irreversible actions in code (agent containment).
  • A judge is not ground truth. Model judges carry position, verbosity, and self-enhancement bias (Zheng et al., 2023); an unrecalibrated judge writes them into whatever it grades.

Three anti-patterns: compacting agent context with a judgment model, blind to tool results (daily.dev’s teardown); preferring a probability where a deterministic oracle exists — a surviving mutant is proof (verifiable finish lines); and high-stakes ranking without a second look. Willison: “I really hope nobody uses Jev to rank job applicants—that floating point number could conceal all manner of unseen bias baked into the models, and experimentally picking that bias apart is going to be a tricky business.”

Done means

  • The job passes all four fit-test questions; every question is one atomic judgment with an honest “none of the above.”
  • Arithmetic, counting, and date logic live in code.
  • You pin a versioned model ID, log the model field per response, and batch questions sharing a state.
  • Calibration is measured on your own labelled data; thresholds carry margin.
  • Side effects are gated on confidence in code; no judgment stands alone before an irreversible action.
  • User-supplied state text is treated as untrusted; sampled decisions are shadow-graded by a stronger model.
  • Every speed or cost multiple in your docs carries its self-report label.

What this article does NOT cover

  • Not a review: no harness was run; every number is quoted from the vendor or a named third party.
  • Not a setup walkthrough; SDKs and the API reference are out of scope.
  • Not Jev’s internals: architecture, parameter count, and training recipe are unpublished. RLCD (Reinforcement Learning for Calibrated Decisions) is the vendor’s stated training method, not a verified mechanism.
  • Not open-weight reproductions, price-stability forecasts, or any evaluation of Jev for hiring, lending, or security decisions.

Research basis: shared/abs-research-briefs/aidb/judgment-models/judgment-models-research-2026-09-25.md (verified 2026-09-28).

Sources

#ai-agents#judgment-models#jev#operator-practice#evaluation

Submit a take

Have a different read on this? Drop a comment below — your email isn't published, and I read every one. Nothing leaves the site until I approve it.

Your email address will not be published. Required fields are marked.