guide · ai

Qwen3.8-2.4T-A95B

Qwen3.8-2.4T-A95B (Alibaba / Qwen, August 11-12, 2026): what the release is, why it matters for operators, specs, benchmarks, and the call.

August 12, 2026 · By Alastair Fraser

A retro robot representing Qwen3.8-2.4T-A95B

--- title: “Qwen3.8-2.4T-A95B” description: “Qwen3.8-2.4T-A95B (Alibaba / Qwen, August 11-12, 2026): what the release is, why it matters for operators, specs, benchmarks, and the call.” type: guide category: ai pubDate: 2026-08-12 image: /images/abs-model-qwen3.8-2.4t-a95b.png imageAlt: “A retro robot representing Qwen3.8-2.4T-A95B” imagePrompt: “Bold graphic editorial illustration, 1990s comic-book influence, heavy ink outlines, halftone texture, crimson and electric blue on cream. A single retro-futurist robot representing an AI model / a brain-in-a-server, no text, no logos, 16:9.” affiliate: false sources: - name: “Qwen official blog — Qwen3.8-Max: A New Bar for Coding and Cowork (Aug 2, 2026)” url: “https://qwen.ai/blog?id=qwen3.8” - name: “HF model card — Qwen/Qwen3.8-2.4T-A95B” url: “https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B” - name: “HF LICENSE file — Qwen3.8-Max License” url: “https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/main/LICENSE” - name: “OpenRouter — qwen/qwen3.8-2.4t-a95b (Aug 12, 2026)” url: “https://openrouter.ai/qwen/qwen3.8-2.4t-a95b” - name: “Artificial Analysis — Qwen3.8 Max (Aug 3, 2026, methodology v4.1.1)” url: “https://artificialanalysis.ai/models/qwen3-8-max” - name: “Hacker News — Qwen3.8-2.4T (Philpax, Aug 12, 2026, 291 points)” url: “https://news.ycombinator.com/item?id=49273478” - name: “Latent.Space AINews — Qwen 3.8 Max (2.4T) and 27B, new open weights models for Coding and Cowork (Aug 4, 2026)” url: “https://www.latent.space/p/ainews-qwen-38-max24t-and-27b-new” - name: “Thomas Wiegold — Qwen3.8-Max Review: I Tested Alibaba’s 2.4T Model (Jul 20, 2026)” url: “https://thomas-wiegold.com/blog/qwen-3-8-max-review/” - name: “Fireworks AI Day-0 launch tweet (Aug 12, 2026)” url: “https://x.com/FireworksAI_HQ/status/2087578149915914727” - name: “Marktechpost — Alibaba previews Qwen3.8-Max (Jul 19, 2026)” url: “https://www.marktechpost.com/2026/07/19/alibaba-previews-qwen3-8-max-a-2-4-trillion-parameter-multimodal-model-days-after-moonshots-kimi-k3-open-weight-launch/” facts: - label: “Vendor” value: “Alibaba / Qwen” - label: “Released” value: “August 11-12, 2026” - label: “License” value: “Open-weight (custom Qwen3.8-Max License)” - label: “Type” value: “AI model release” related: - abs-model-claude-opus-5 - abs-model-gpt-5.6-cyber - abs-model-grok-4.6 - abs-model-kimi-k3 - abs-model-lfm2.5-vl-3b - abs-model-ling-3.0-flash - abs-model-muse-glimmer-30b - abs-model-muse-spark-1.2 - abs-model-nemotron-3.5-lightning - abs-model-qwen-image-3.0 - abs-model-qwen3.7-flash - abs-model-qwen3.8-max-weights-update - abs-model-solar-pro-4 tags: - qwen3.8-2.4t-a95b - ai-models - model-release draft: false --- --- title: “Qwen3.8-2.4T-A95B” description: “Qwen3.8-2.4T-A95B (Alibaba / Qwen, August 11-12, 2026): what the release is, why it matters for operators, specs, benchmarks, and the call.” type: guide category: ai pubDate: 2026-08-12 image: /images/abs-model-qwen3.8-2.4t-a95b.png imageAlt: “A retro robot representing Qwen3.8-2.4T-A95B” imagePrompt: “Bold graphic editorial illustration, 1990s comic-book influence, heavy ink outlines, halftone texture, crimson and electric blue on cream. A single retro-futurist robot representing an AI model / a brain-in-a-server, no text, no logos, 16:9.” affiliate: false sources: - name: “Qwen official blog — Qwen3.8-Max: A New Bar for Coding and Cowork (Aug 2, 2026)” url: “https://qwen.ai/blog?id=qwen3.8” - name: “HF model card — Qwen/Qwen3.8-2.4T-A95B” url: “https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B” - name: “HF LICENSE file — Qwen3.8-Max License” url: “https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B/blob/main/LICENSE” - name: “OpenRouter — qwen/qwen3.8-2.4t-a95b (Aug 12, 2026)” url: “https://openrouter.ai/qwen/qwen3.8-2.4t-a95b” - name: “Artificial Analysis — Qwen3.8 Max (Aug 3, 2026, methodology v4.1.1)” url: “https://artificialanalysis.ai/models/qwen3-8-max” - name: “Hacker News — Qwen3.8-2.4T (Philpax, Aug 12, 2026, 291 points)” url: “https://news.ycombinator.com/item?id=49273478” - name: “Latent.Space AINews — Qwen 3.8 Max (2.4T) and 27B, new open weights models for Coding and Cowork (Aug 4, 2026)” url: “https://www.latent.space/p/ainews-qwen-38-max24t-and-27b-new” - name: “Thomas Wiegold — Qwen3.8-Max Review: I Tested Alibaba’s 2.4T Model (Jul 20, 2026)” url: “https://thomas-wiegold.com/blog/qwen-3-8-max-review/” - name: “Fireworks AI Day-0 launch tweet (Aug 12, 2026)” url: “https://x.com/FireworksAI_HQ/status/2087578149915914727” - name: “Marktechpost — Alibaba previews Qwen3.8-Max (Jul 19, 2026)” url: “https://www.marktechpost.com/2026/07/19/alibaba-previews-qwen3-8-max-a-2-4-trillion-parameter-multimodal-model-days-after-moonshots-kimi-k3-open-weight-launch/” facts: - label: “Vendor” value: “Alibaba / Qwen” - label: “Released” value: “August 11-12, 2026” - label: “License” value: “Open-weight (custom Qwen3.8-Max License)” - label: “Type” value: “AI model release” related: - abs-model-claude-opus-5 - abs-model-gpt-5.6-cyber - abs-model-grok-4.6 - abs-model-kimi-k3 - abs-model-lfm2.5-vl-3b - abs-model-ling-3.0-flash - abs-model-muse-glimmer-30b - abs-model-muse-spark-1.2 - abs-model-nemotron-3.5-lightning - abs-model-qwen-image-3.0 - abs-model-qwen3.7-flash - abs-model-qwen3.8-max-weights-update - abs-model-solar-pro-4 tags: - qwen3.8-2.4t-a95b - ai-models - model-release draft: false --- ## The release Alibaba’s Qwen team dropped Qwen3.8-2.4T-A95B on Hugging Face and ModelScope on August 11–12, 2026 — the open-weight base of Qwen3.8-Max, which had been API-only since the July 19 WAIC Shanghai preview and the August 2 official launch on Qwen Cloud. The HF model card says it verbatim: “Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with more features, such as vision input & non-thinking support, 1M context length by default, official built-in tools, etc.” Same weights, different surface — A95B is the text-only, thinking-required, 262K-native base; Max is the proprietary wrapper with vision, tools, and 1M-default context bolted on top. This is the first Qwen-Max-class model ever shipped open weights. All prior Max variants (Qwen3.5-Max, Qwen3.6-Max, Qwen3.7-Max) stayed API-only, so the release is the policy reversal the open-source community had been pressing for. The Qwen team announced it on August 2: “Today, we are officially releasing Qwen 3.8-Max, the most capable model in the Qwen family to date. This also marks the first time we will open-source the weights of a Qwen-Max-class model — the open weights will be released next week.” The weights landed roughly on schedule, ten days after the API. Day-0 surface: Hugging Face (BF16 + FP8), ModelScope mirror, Fireworks AI serving, OpenRouter listing at $2/$6 per 1M tokens, the official vLLM and SGLang recipes, and a Unsloth 1-bit dynamic quant at ~397 GB. On the first day of OpenRouter traffic, Hermes Agent alone consumed 67.2B tokens routed through the new model. ## Why it matters for operators A95B is the open-weight answer to the “GPT-5.6 Sol or Claude Opus 4.8 for the coding-and-cowork tier” question, at one-third of Claude Fable 5’s cost. Artificial Analysis scored the proprietary Qwen3.8-Max variant 58 on Intelligence Index v4.1.1 (median is 34) — “well above average among comparable models.” Vals Index puts Qwen3.8-Max at 66.1, second only to Claude Fable 5 among open-weight models and matched with Claude Opus 4.7 at 2.3× lower cost-per-test ($2.68 vs $6.17). On Frontend Code Arena it ranks #4 at 1,668 Elo, behind only Opus 5 (Max) and Kimi K3 (Max). The two non-benchmark features that actually change the calculus for ABS-style agent operators: (1) the API exposes a first-class reasoning_effort dial (low / medium / xhigh, default xhigh) plus preserve_thinking for carrying reasoning tokens across turns — both designed for multi-hour agent loops, not single-shot Q&A; (2) the model is post-trained for long-horizon agent work, with Qwen’s blog demonstrating a 16-day self-evolving harness build (265 commits, 127 PRs, 151 issues on a public dev-bot repo) and a 5-day autonomous reproduction-and-improvement of an arXiv reasoning-data-selection paper. Those are the strongest “this thing can stay on a task for a week” demonstrations of any model released this month. Latent.Space’s AINews digest put it bluntly: “Qwen is so back! … Qwen 3.8 Max is a MONSTER 2.4T model that would have been the top open model in the world but for the Kimi K3 release we already covered.” Omar Sar0 (Teknium) used it as the canonical example for the open-weights-have-closed-the-gap thesis: “Try Qwen3.8-Max on Hermes Agent… Using it in Hermes Agent makes it hard to deny how much open frontier models have closed the gap with closed frontier systems.” ## Specs that matter | Spec | Value | |---|---| | Architecture | MoE over hybrid Gated DeltaNet + Gated Attention (23 × (3 DeltaNet→MoE + 1 Attn→MoE)) | | Total parameters | 2.4 trillion | | Activated parameters | 95 billion (10 routed + 1 shared expert per token, ~4% activation ratio) | | Layers | 92 total | | Experts | 512 total, 11 active per token | | Hidden dim | 8,192 | | Linear-attn heads | 128 V / 16 QK, head dim 128 | | Attention heads | 64 Q / 4 KV, head dim 256, RoPE dim 64 | | Vocab | 248,320 (padded — meaningfully larger than Kimi K3 ~164K, DeepSeek-V4 ~129K, GLM-5.2 ~155K) | | Native context | 262,144 tokens | | Extensible context | up to 1,010,000 tokens (YaRN-style) | | MTP | Trained with multiple multi-token-prediction steps | | Reasoning modes | low / medium / xhigh (xhigh default) | | Sampling defaults | temperature=1.0, top_p=0.95, top_k=20, min_p=0.0 | | Modality | Text only (vision dropped vs Qwen3.8-Max wrapper) | | Weights released | BF16 (~4.9 TB) + FP8 (~2.5 TB) | | Hardware floor | 8 H100/B200 minimum (HN/NitpickLawyer); full 1M context “within 192GB VRAM” via architectural optimizations | | License | Custom Qwen3.8-Max License (not Apache 2.0) — see operator call | | Price (API) | $2/M input, $6/M output, $0.25 cached, $0.17 cached-5m | The hybrid Gated DeltaNet + Gated Attention stack is Qwen3.5/3.8’s signature — linear-attention layers for the cheap streaming-compute pattern, standard attention layers where precise retrieval matters. HF card: “Qwen3.8 brings a Qwen-Max-class model to open release. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks.” Sources: HF model card and Qwen blog — Qwen3.8-Max launch. ## Benchmarks Independent — Artificial Analysis (proprietary Qwen3.8-Max variant, same underlying weights): - Intelligence Index v4.1.1: 58 (median 34) — Artificial Analysis - Output speed: 47 tok/s (71st percentile — slow) - Reasoning time: 41s, answer time: 10s - Cost per Intelligence Index task: ~$1.74 - Briefcase Elo (AA mid): 1420 - Pricing: $2/M input, $6/M output, $0.25 cached, $0.17 cached-5m Third-party — Latent.Space / Vals (Aug 4 digest): - Frontend Code Arena: #4 at 1,668 Elo (behind Opus 5 Max 1,705, Kimi K3 Max 1,676) - Vision Arena: #2 at 1,305 Elo

Sources

Sources

#qwen3.8-2.4t-a95b#ai-models#model-release

Submit a take

Have a different read on this? Drop a comment below — your email isn't published, and I read every one. Nothing leaves the site until I approve it.

Your email address will not be published. Required fields are marked.