The Mysterious 'Ox Alpha' Model Is Zhipu: GLM-5.3 Flash, First Native Multimodal, Fully on Domestic Chips

TIQEX

The Mysterious 'Ox Alpha' Model Is Zhipu: GLM-5.3 Flash, First Native Multimodal, Fully on Domestic Chips
The Mysterious 'Ox Alpha' Model Is Zhipu: GLM-5.3 Flash, First Native Multimodal, Fully on Domestic Chips

2026-09-02 · AI Tech Insights · Curated from 量子位

Summary: The anonymous 'Ox Alpha' model that went viral overseas has been revealed as Zhipu's open-source GLM-5.3 Flash. With 320B total and 18B activated parameters, it scored 57 on the AA Index—level with Claude Opus 4.8. Its linear+sparse hybrid attention slashes compute, with pricing at just 1/40 of Opus 4.8. All 62T anonymous-test tokens were served on domestic Chinese chips.

The Mysterious 'Ox Alpha' Model Is Zhipu: GLM-5.3 Flash, First Native Multimodal, Fully on Domestic Chips

Image source: 量子位

The mysterious 'Ox Alpha' model that has been making waves online has finally been revealed—indeed from a Chinese player: Zhipu, with the just-released GLM-5.3 Flash, the first natively multimodal version in the GLM 5 series.

During its anonymous run, many already tested Ox Alpha. For example, Tim Jayas used the same prompt to have 'Ox Alpha' build a SpaceX Raptor 3D interactive web page twice, noting that "it feels like a model that can keep learning itself." When we got GLM-5.3 Flash internal access and ran the same task, after entering the prompt without any operation, the 3D Raptor engine project was produced within moments.

Why did 'Ox Alpha' catch fire? Looking at official posts alone: on OpenRouter, 'Ox Alpha' shot to the top on day one, setting a new record for daily tokens used; OpenCode noted that Ox Alpha dethroned DeepSeek's 56-day reign on its platform.

After Zhipu officially claimed the model, we discovered more highlights. In terms of model size, GLM-5.3 Flash has only 320B parameters, but outperforms the 753B GLM-5.2 in capability. Per the latest AA Index, GLM-5.3 Flash scored 57, placing it in the global frontier tier, tied with Claude Opus 4.8. On pricing, GLM-5.3 Flash costs 1/10 of GLM-5.3 and 1/40 of Opus 4.8—meaning developer call costs have plummeted.

There's also something to be proud of—all 62T tokens were served on domestic chips. The 'mysterious model' that foreigners chased for a month turned out to be a purebred Chinese ox.

Architecturally, GLM-5.3 Flash is GLM's first natively multimodal model. Zhipu built a data synthesis pipeline specifically for Visual Coding, allowing the model to look at final pages, interactions, and 3D scenes during task execution and iterate based on visual feedback—this is why it can write, observe, and revise in design-draft replication and 3D modeling tasks.

The key architectural innovation lies in the attention mechanism. GLM-5.3 Flash adopts a hybrid linear+sparse attention architecture: linear attention captures local information, while sparse attention uses a lightweight indexer to pull relevant global context. With 1M-token context windows, it doesn't force every token to compute against every other. Zhipu also added an IndexPool that compresses 4 indexer cache vectors into 1—reducing attention computation by 3.01x and shrinking KV Cache by 4.44x versus GLM-5.3.

To make GLM-5.3 Flash run on domestic accelerators, Zhipu split multimodal encoding, prompt prefill, and per-token decoding into independently schedulable, scaleable Encode-Prefill-Decode architectures, while layering optimizations like Layer Split and mixed-precision cache quantization. End-to-end service performance improved 3x versus the initial baseline on the same hardware, with per-token costs comparable to mainstream NVIDIA GPUs.

Looking back at 'Ox Alpha's overseas explosion, the meaning shifts. Before official claiming, developers on OpenRouter and OpenCode had no idea whose model it was—they just found it good and kept sending prompts, pushing it to #1. When Zhipu revealed the cards, the second half emerged: the model is domestic, and so is the compute powering global real traffic.

Frontier models face an increasingly real problem: capability grows, but so does cost. Especially as Agents truly work for tens of minutes or hours, tokens flow like an open faucet. GLM-5.3 Flash cuts deeper into architecture, inference, and compute efficiency, then drives frontier-model prices down: 57 on AA—tied with Opus 4.8—but 1/40 the price. Frontier models are starting to feel like everyday consumables.

Finally, open-sourcing. GLM-5.3 Flash is now globally open-source, integrated with Z Code, with open APIs. The model runs on domestic chips, weights are public, developers can access them, and enterprises can self-deploy. Model, domestic compute, and open-source ecosystem are now genuinely wired together.

Key Takeaways

  • 320B total / 18B active parameters; scored 57 on the AA Intelligence Index, on par with Claude Opus 4.8
  • Hybrid linear + sparse attention: 3.01x less attention compute, 4.44x smaller KV cache
  • Priced at 1/10 of GLM-5.3 and 1/40 of Opus 4.8
  • First natively multimodal model in the GLM 5 series, with Visual Coding support
  • All 62T tokens during anonymous testing were computed on domestic Chinese chips

📎 This article was automatically compiled by AI from 量子位 (2026-08-27).
All rights belong to the original authors. Used for informational purposes only.

This article was automatically compiled by AI from 量子位.

Copyright Notice: This article is for informational purposes only. All rights belong to the original authors. For inquiries, contact [email protected].
Published by: Tiqex · 2026-09-02