guide · ai

Who Made Ox Alpha? The Fingerprinting Evidence, the Alternatives, and What Would Disprove It

Nobody has claimed the stealth model. The fingerprinting evidence strongly favors one lab — here's the weighted case, the alternatives, and what would prove it wrong.

August 23, 2026 · By Alastair Fraser

A detective robot examining giant animal footprints that lead toward a distant server farm.

Who Made Ox Alpha? A Signal Check on the Internet’s Favorite Guess

One-line job: Walk the attribution evidence for Ox Alpha with honest weights, land on a calibrated verdict, and know exactly what would prove it wrong. Audience: Technically literate readers watching the fingerprinting drama unfold. Not for: Anyone wanting a confident answer — as of this writing there isn’t one, and pretending otherwise would be the opposite of useful. Last verified: 2026-08-21 Evidence weight: documentation-verified

Nobody knows who made Ox Alpha. That’s still true as of August 21 — no lab has claimed the free, million-context, video-eating reasoning model that appeared on OpenRouter August 20 and burned through roughly 1.38 trillion prompt tokens on day one. But “nobody knows” is not the same as “no evidence.” The evidence is actually substantial. It just isn’t proof.

How you fingerprint an anonymous model

Before the evidence, the method — because this is what makes the analysis durable after the reveal. When a model arrives with no name attached, analysts compare it against known models across channels that are hard to fake:

  • Tokenizer vocabulary. Every model family tokenizes text slightly differently. If an anonymous model’s token counts for fixed inputs exactly match GLM-5.3’s, that’s a strong structural signature.
  • Video-encoder token budgets. Multimodal models convert video into tokens using encoder-specific budget rules. Matching budgets suggest shared components.
  • API default parameters. Sampling defaults (here: temperature 1, top_p 0.95) are engineering habits that persist across a lab’s models.
  • Provider-specific error codes. Error code 1210 has previously been associated with Z.ai infrastructure.
  • Refusal and style patterns. How a model refuses (and its prose tics) form a stylistometric fingerprint.

None of these are individually conclusive. Together they’re why the community converged fast.

The evidence, weighted

#SignalPoints toWeight
1Tokenizer exactly matches GLM-5.3ZhipuStrong
2Video-encoder token budgets identical to GLM-5V-TurboZhipu multimodal lineStrong
3API fingerprints: temp 1 / top_p 0.95 defaults, reasoning-block behavior, error code 1210ZhipuMedium-strong
4Naming: Pony Alpha → Ox Alpha; horse → ox follows the Chinese zodiac; “Alpha” suffix reusedZhipuSuggestive
5Same stealth account ran Pony Alpha — confirmed as Zhipu’s GLM-5 five days laterZhipuStrong precedent
6Appeared days after text-only GLM-5.3 (Aug 14), filling the obvious multimodal gapZhipuCircumstantial
7ZCode — Z.ai’s own coding product — is a top-5 consumer of Ox AlphaZhipuCircumstantial
8Pliny the Liberator’s self-report claimWeak-medium at bestWeak

Note the class distinctions. Signals 1–3 are inferences from technical fingerprints — multiple independent analysts, consistent results. Signal 5 is different in kind: it’s a confirmed historical outcome, not an inference. Signals 6–7 are circumstantial but real; ZCode showing up in the top-consumer table alongside Hermes Agent and Claude Code is an observed fact whose interpretation is the inference.

Ben Davis, who did some of the most cited fingerprinting work:

“99% sure it’s GLM-5.x, all the evidence points to it (same video encoder, same tokenizer, style matches, same audio rejection, etc.)” — Ben Davis (@davis7), Aug 21, 2026

And the rumor bucket, dispatched in one sentence: an August 20 X post claims GLM-5.5 exists with 3T+ parameters and 1.5M context targeting September/October — single-source, unverified, unknown. Some analysts read Ox Alpha as a current-gen GLM-5.3V variant instead. Don’t build on either story.

The precedent that carries the most weight

February 6, 2026. A model called Pony Alpha appears on OpenRouter — free, 200K context, same stealth account setup:

“Pony Alpha appeared on OpenRouter February 6, free, 200K context, 131K max output; Pony Alpha self-identifies as a GLM model developed by Zhipu” — Dev Genius account of the February incident

“Five days after Pony Alpha appeared on OpenRouter, Zhipu AI confirmed it was GLM-5.” — Awesome Agents, Apr 2026

Same playbook, same account class, zodiac-consistent naming chain (horse → ox), suffix reused. Then zoom out to the base rate:

“All four turned out to come from Chinese labs, including Zhipu AI’s GLM-5, Xiaomi’s MiMo-V2-Pro, Ant Group’s Ling-2.6-flash, and Meituan’s LongCat-2.0.” — Coursiv, Aug 2026

Every recent OpenRouter stealth model came from a Chinese lab. History doesn’t prove anything about Ox Alpha specifically — but it sets the prior, and the prior points the same direction as the fingerprints.

The alternatives, steelmanned

Xiaomi MiMo V3. Xiaomi has run anonymous previews before, and the 100T-tokens-per-day capacity boast resembles earlier MiMo promotion style. This is the strongest promotional-pattern match. It fails on technique: the tokenizer and API prints don’t match Xiaomi’s line.

Tencent Hy4. Has a dedicated Reddit investigation thread. No technical fingerprints have been published supporting it.

MiniMax. Timing is circumstantially plausible; an unannounced model would fit. No fingerprint match.

A router of several models. This explains the oddest observation — Ox Alpha’s behavior is inconsistent across task types (overthinks some tasks, crisp on others). But it’s contradicted by the consistency of the tokenizer and API fingerprints, which should vary if multiple models were being served.

Third-party fine-tune of a GLM checkpoint. The strongest structural objection, raised on the Manifold market: Cursor-style companies fine-tune existing checkpoints, which would produce exact GLM fingerprints without Zhipu being behind the deployment. It doesn’t fully explain the GLM-5V-Turbo encoder budgets or the sheer serving scale — a fine-tune operation with capacity for 100T tokens/day is itself a major lab. Still, keep this one alive; it’s the scenario where “the fingerprints say GLM” and “Zhipu made it” come apart.

One reason these signals carry weight deserves emphasis: they are structural, not behavioral. Marketing copy can be imitated, benchmark scores can be cherry-picked, and a model can be prompted into saying almost anything. But tokenizer behavior is baked into the trained artifact itself — two independently trained models producing identical token counts across arbitrary inputs is about as likely as two people sharing fingerprints by coincidence. That’s why the tokenizer match sits at the top of the evidence table while Pliny’s agent-mediated self-report sits at the bottom.

The zodiac naming argument (signal 4) is worth spelling out because it looks like numerology until you see the track record. Pony Alpha arrived in February — horse year territory in the Chinese zodiac ordering used by the community’s naming speculation. An ox follows a horse in that cycle. Ox Alpha arrives six months later with the same “Alpha” suffix, the same stealth account class, and the same free-preview shape. Nobody at Zhipu has confirmed the naming scheme means anything. But when the naming pattern matches the confirmed February precedent this precisely, it stops being coincidence-shaped.

The serving-scale objection cuts both ways and belongs in your weighting. Whoever runs Ox Alpha needs compute capacity for 100 trillion tokens per day — Theo’s reaction captured why that number stunned people. That scale requirement quietly eliminates most of the small-lab theories: a third-party fine-tune shop doesn’t have it. It also makes “router of several models” less likely than behavior alone suggested, since routing infrastructure at that scale would itself be a notable engineering achievement worth bragging about somewhere.

Calibrated verdict

The published confidences, as a spread rather than a single number:

My honest position sits inside that spread: the evidence strongly favors Zhipu AI, it is not proof, and the model remains unclaimed. Treat self-reports from the model itself as near-worthless — models guess their own lineage all the time, and Pliny’s announcement (“my agent has spoken”) is a perfect example of the genre: circular, unverifiable, confidently stated.

How this resolves

What confirms Zhipu: an official claim — historically timed to when the free window ends (~Aug 27); an open-weights release whose tokenizer hashes match probes of ox-alpha; Z.ai API error-code parity.

What flips the verdict: Xiaomi, Tencent, MiniMax, or a named Western lab claiming it; demonstrated routing artifacts (two distinct tokenizer signatures in one session would vindicate the router theory).

Whichever way it breaks, here’s the operational point for anyone deciding whether to care: every prior stealth conversion landed aggressive-cheap, and the interesting question was never really who — it’s what the pricing looks like after the reveal. That’s the number worth waiting for.


Sources: primary — OpenRouter model page and Stealth EULA; authoritative secondary — Manifold market, daily.dev testing summary; independent — Coursiv, Startup Fortune, OfficeChai, Awesome Agents, Dev Genius. X posts cited via carrier pages where possible; volatile posts labeled.

Sources

#abs-guide#ai-models#attribution#analysis

Submit a take

Have a different read on this? Drop a comment below — your email isn't published, and I read every one. Nothing leaves the site until I approve it.

Your email address will not be published. Required fields are marked.