guide · ai

Tencent Hy4 Preview: Open Weights Do Not Make a 770B Model Practical on Consumer Hardware

Choose a hosted Hy4 route, run a safe first API check, control reasoning costs, and test the preview against an established fallback.

September 23, 2026 · By Alastair Fraser

Choose a hosted Hy4 route, run a safe first API check, control reasoning costs, and test the preview against an established fallback.

The 30-second decision

Use Tencent Hy4 Preview through a hosted provider if you need a text model for coding agents, tool use, long document work, or repository-scale tasks and can tolerate preview-stage behavior.

Start with OpenRouter if you want a straightforward OpenAI-compatible API. Consider Nous Portal if you already use Hermes Agent and its catalog rate fits your account. Use Tencent TokenHub if you need direct Tencent billing, Tencent regions, or an existing vendor relationship.

Do not build a consumer workstation around this model. Hy4 has open weights, but its 770 billion parameters make self-hosting a data-center deployment problem. Do not put it into production without a tested fallback.

Who made Hy4 and where it fits

Tencent’s Hy Team released Hy4 Preview on August 28, 2026 under the Apache 2.0 license. It is a sparse mixture-of-experts model with:

  • 770 billion total parameters
  • 49 billion parameters activated per token
  • 1,048,576-token context window
  • 78 transformer layers
  • 256 routed experts plus one shared expert in each MoE layer
  • Top-eight expert routing
  • A native multi-token prediction layer for speculative decoding
  • Text input and output in the verified provider listings

Tencent positions Hy4 for coding agents, complex tool workflows, office and financial analysis, game prototyping, and long-running multi-step work. That positioning matches the early traffic reported through OpenRouter, which was dominated by coding-agent applications.

The practical fit is narrower than the specification sheet suggests. A one-million-token limit does not mean every request should contain one million tokens. Large requests cost more, take longer, and can still lose important details inside a crowded context. Use the large window when your evaluation shows that retrieval or file selection cannot reduce the input safely.

The architecture details that matter

Most operators do not need to understand every named component, but four details affect deployment and use.

First, only 49 billion of the model’s 770 billion parameters are active for each token. That reduces inference computation compared with activating the full model, but it does not remove the need to store and route across the full checkpoint.

Second, Hy4 uses Gated DeepSeek Sparse Attention with IndexCache. This is intended to make long-context attention more practical by selecting sparse attention targets and reusing index information across layers.

Third, identity Hyper-Connections widen the residual pathway across four streams. This is an architectural feature rather than an API control; you do not configure it through OpenRouter or TokenHub.

Fourth, the native multi-token prediction layer can support speculative decoding. Whether you benefit from it depends on the serving stack used by the provider.

These features may help Tencent serve the model efficiently. They do not turn it into a normal local model.

Why open-weight does not mean consumer-local

The FP8 checkpoint is roughly 770 GB before allowing for runtime overhead, caches, routing, and the rest of the serving stack. BF16 weights are roughly 1.54 TB. Tencent’s official deployment shape uses tensor parallelism across eight GPUs.

That is a multi-GPU server deployment, not a sensible desktop recommendation. Quantization could reduce storage, but it would not erase memory bandwidth, interconnect, context-cache, power, cooling, and operational requirements. There was no verified consumer-ready INT4 or GGUF release at the research cutoff.

Self-hosting can still make sense for an organization with data-placement requirements, sustained utilization, suitable hardware, and staff able to operate distributed inference. For an ordinary ABS operator, hosted access is the practical starting point.

Provider decision table

Prices and availability change. This table records the evidence captured on August 28, 2026, not a permanent promise.

Route as of 2026-08-28Model IDListed price per 1M tokensChoose it whenImportant caveat
OpenRoutertencent/hy4-preview$0.834 input, $2.501 output, $0.042 cache readYou want an OpenAI-compatible gateway and pay-as-you-go accessOpenRouter listed one Tencent Cloud upstream, so there was no alternate Hy4 provider to fail over to
Nous Portaltencent/hy4-previewCatalog showed $0.67 input and $2.00 outputYou already use Hermes Agent or Portal creditsThe catalog rate applies to the Portal credit balance; do not assume every billing route has the same price
Tencent TokenHubhy4-preview$0.834 input, $2.501 outputYou need direct-vendor billing, Tencent regions, or an existing Tencent contractThe verified listing did not expose a cache-read price
Tencent Token Planhy4-previewSubscription routeYou prefer predictable subscription billingCheck the current plan limits before committing
Existing gatewaysVaries by gatewayCommon listings were about $0.83 input and $2.50 outputYou already use Kilo, NanoGPT, OpenCode Go, or Vercel AI GatewayConfirm the exact model ID, price, context policy, and reasoning support
WorkBuddy or CodeBuddyIn-appTencent said free for two weeks upon launchYou want a no-code evaluationTencent did not publish a calendar cutoff; verify the current offer in the app

DeepSeek V4 Pro was cheaper on input, output, and cached input in Tencent’s comparison table while also offering a one-million-token context window. Hy4 is therefore not automatically the cheapest or best model in this class. It has to win your workload test.

First success through OpenRouter

Create an OpenRouter account, add credit, and generate a key in its dashboard. Put the key in an environment variable rather than source code.

export OPENROUTER_API_KEY="<YOUR_OPENROUTER_API_KEY>"

OPENROUTER_API_BASE="https://openrouter.ai"
curl "${OPENROUTER_API_BASE}/api/v1/chat/completions" \
  -H "Authorization: Bearer ${OPENROUTER_API_KEY}" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tencent/hy4-preview",
    "messages": [
      {
        "role": "user",
        "content": "Reply with exactly: Hy4 connection confirmed"
      }
    ],
    "temperature": 0.9,
    "top_p": 1.0
  }'

A successful response should contain an assistant message. Record the returned model and provider metadata available in your account or logs. Do not assume a gateway label proves which upstream served a request.

This recipe uses a placeholder only. Never paste a real key into a guide, repository, issue, or shared terminal transcript.

Reasoning controls and cost

Tencent recommends temperature: 0.9 and top_p: 1.0. The model defaults to high reasoning for complex work. For direct responses, Tencent documents a no-thinking control through the chat template:

{
  "extra_body": {
    "chat_template_kwargs": {
      "reasoning_effort": "no_think"
    }
  }
}

Use high reasoning for difficult coding, planning, or analysis only after it proves useful. Use no_think for classification, extraction, routing, short summaries, and direct formatting tasks.

Tencent says the preview can reason longer than necessary and over-verify its work. That affects latency and output-token cost. A visible answer may require substantially more generated reasoning than the text you display, so maximum-token settings and cost estimates must account for reasoning as well as the final answer.

OpenRouter reported 42 tokens per second P50 and 3.10 seconds P50 latency on August 28, 2026. Those are hosted, thinking-on observations, not local throughput guarantees. The same capture showed 100.00% in the Tencent Cloud provider row, 99.98% page-wide three-day uptime, and 97.60% three-day availability. These are different, drifting metrics.

OpenRouter’s $0.042-per-million cache-read rate can help agent loops that repeatedly send stable instructions and history. It is not a universal Hy4 rate: the verified TokenHub listing exposed no cache line, and Nous Portal may account for caching differently.

What the vendor benchmarks show

Tencent reported the following Hy4 results:

BenchmarkVendor-reported result
GPQA Diamond92.3
SWE-Bench Multilingual82.9
SWE-Bench Pro65.7
DeepSWE64.3
Terminal-Bench 2.185.4
MCP-Atlas public set83.7
OfficeQA Pro66.2

These figures came from Tencent’s benchmark appendix and had not been independently reproduced at the research cutoff.

Tencent also ran a blind evaluation with 163 internal experts and 203 engineering tasks. Hy4 averaged 2.99 against GLM 5.3 at 2.92 and Kimi K3 at 2.94. The margins were 0.05 to 0.07 on a four-point scale. Treat that as evidence that the models occupied a similar tier on Tencent’s workload, not proof that Hy4 was superior.

A small evaluation plan

Test before moving real traffic:

  1. Select ten representative tasks: four coding or tool tasks, three long-context tasks, and three short direct tasks.
  2. Run each task with Hy4 high reasoning, Hy4 no_think, and your current fallback model.
  3. Record task success, unsupported claims, tool-call validity, input tokens, billed output tokens, total cost, first-token delay, and completion time.
  4. Repeat failed or unusually slow cases to separate a model problem from temporary provider capacity.
  5. Set acceptance limits for cost, latency, and correctness before reviewing the results.
  6. Keep the fallback available through a separate model ID and test the switch before production.

Use DeepSeek V4 Pro as a price and stability comparison when one-million-token context matters. Consider GLM 5.3 or Kimi K3 only when their behavior, vendor relationship, or workload-specific results justify the higher listed price or smaller context.

Limitations and fallback

Hy4 is a preview. Its behavior, checkpoint, pricing, and provider availability can change. Tencent explicitly reports over-reasoning and over-verification. Independent benchmark evidence was limited at launch.

Do not use Hy4 as the sole path for a latency-sensitive or business-critical workflow. Define retry limits, cap spend, log model and provider metadata, and route failures to a model you have already evaluated. A gateway cannot provide upstream diversity when only one Hy4 inference provider is listed.

The operator call is simple: start hosted, use a small paid evaluation, disable deep reasoning where it adds no value, and keep the incumbent model until Hy4 wins on your own tasks.

Done means

  • You can complete a hosted Hy4 request without exposing a key.
  • You have tested both high reasoning and no_think.
  • You know the measured cost and latency for your workload.
  • You have compared Hy4 with at least one established alternative.
  • You have a tested fallback for errors, capacity problems, regressions, and preview changes.
  • You are not treating open weights as evidence that consumer-local deployment is practical.

What this article does NOT cover

This guide does not cover eight-GPU deployment, Blackwell server design, fine-tuning, quantization, GGUF conversion, multimodal extensions, or custom inference infrastructure. It also does not promise a promotion cutoff, stable preview behavior, or universal pricing across providers.

This is the consolidated ABS guide for Hy4 Preview. A separate local-install, provider-comparison, or promotion guide is not recommended without new deployment evidence, independent benchmarks, or demonstrated reader demand.

Sources

Sources

#hy4#tencent#cloud-models#openrouter

Submit a take

Have a different read on this? Drop a comment below — your email isn't published, and I read every one. Nothing leaves the site until I approve it.

Your email address will not be published. Required fields are marked.