AI NEWS, TRACKED AND ROASTED.
z.ai names its stealth model glm-5.3-flash, eight hours after the crowd did
Z.ai released GLM-5.3-Flash on 2026-08-26, calling it the first natively multimodal model in the GLM-5 series, and used the same announcement to say where the model had already been. The blog post and the developer guide carry the line identically: "Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week - with all of this traffic served on Chinese AI chips."
OpenRouter tells it from its own side of the wire. The retained [48] stealth/ox-alpha record page ↗ states that the model was "developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash", that prompts and completions from the testing period were retained by the provider and not used for training under the Stealth Model Terms, and its close-out routing message hands callers forward: "Thank you for participating in the Stealth Ox Alpha testing period. This model was ZAI's GLM-5.3 Flash. Use it now: [47] OpenRouter ↗". The stealth listing was created 2026-08-20T20:04:55.369Z under the author name "stealth". Today [49] the z-ai/ox-alpha endpoints probe ↗ returns 404 Not Found; the record page is what is left.
Neither vendor got there first. An [58] r/LocalLLaMA identification thread ↗ posted 2026-08-26T06:28:10Z reached 458 points by matching three fingerprints - native vision, a 1M-token window, DeepSWE around 63% - and named the pairing about eight hours before Z.ai's announcement.
What shipped, per Z.ai: 320B total and 18B active parameters, a hybrid sparse-plus-linear-attention architecture new to the GLM series with mHC scaling additions, 45 layers against 92 in the GLM-4.5 series, a 30T-token multimodal corpus, and claimed reductions of 3.0x in attention compute and 4.4x in KV cache versus GLM-5.3. The released [52] config.json ↗ corroborates the structural mechanisms - Glm5NextForConditionalGeneration, index_topk 2048, index_kpool 4 compressed, hc_mult 4 - but not the performance figures. Weights are MIT-licensed at huggingface.co/zai-org/GLM-5.3-Flash, repo created 2026-08-25T06:43:14Z at revision 3f1971b7b5f7a528c9c4ef6212c8785298a8c24a, citing arXiv 2602.15763, with local recipes for SGLang, vLLM, TokenSpeed, and KTransformers.
OpenRouter listed z-ai/glm-5.3-flash publicly at 2026-08-26T13:59:01Z (canonical slug z-ai/glm-5.3-flash-20260826), modality text+image+video to text, top-provider context 1,048,576 tokens with a 131,072-token completion cap, reasoning mandatory and defaulting to max effort. Its embedded Artificial Analysis card shows intelligence index 57.5, coding 71.5, agentic 58.2.
Router rates, as listed:
- price
- $0.075 / M tokens
- price
- $0.25 / M tokens
- price
- $0.015 / M tokens
Z.ai's own guide does not publish a per-token rate. It frames cost through a discounted Artificial Analysis basis - Intelligence Index v4.1.1 score 57 at $0.045 per task - and routes subscribers through the GLM Coding Plan instead:
- terms
- fully available to subscribers
- terms
- 3x
- terms
- standard points
- terms
- 50% of standard points
On the platform side the model is exposed as code glm-5.3-flash with text parameters consistent with GLM-5.3, a 1M-token context window, image input via image_url blocks, and no way to disable thinking - the guide recommends max reasoning_effort.
Benchmarks are vendor-reported. Z.ai claims wins over GLM-5.2 across six coding and agentic benchmarks, naming DeepSWE v1.1 at 63.4 against 46.2 and AutomationBench v1.0.6 at 48.8 against 26.2, plus near-parity with Claude Opus 4.8 at max effort on its own in-house Z.ai Code Bench v1.0 - 29.0 against 29.5, run on Claude Code 2.1.207. One number has an outside witness: the independent [53] DeepSWE public leaderboard ↗ lists glm-5.3-flash [max] at 63% +/- 4%, which is consistent with the vendor's 63.4 without being a measurement of it. The margin is wide enough that exact-value equality is untestable.
The identity itself is asserted, not demonstrated. "Ox Alpha was GLM-5.3-Flash" rests on exactly two agreeing first parties plus community fingerprint matching reported pre-reveal. No checkpoint-hash or serving-stack-level verification exists here, the historical stealth endpoint no longer answers, and the [62] Wayback availability probe ↗ returned no archived snapshots of the stealth-era page.
what people are saying
Reception scale at observation: the [54] Hacker News submission ↗ sat at 922 points and 457 comments; the [59] r/LocalLLaMA release thread ↗ at 1157 points and 378 comments, with a dedicated former-ox-alpha megathread at 203 points and 138 comments. These are point-in-time counts and they decay.
Bluestein, on Hacker News, [55] posted the handover as a live API response ↗: a 404 from the chat completions endpoint for model stealth/ox-alpha whose error text matches OpenRouter's official close-out naming ZAI's GLM-5.3 Flash. A third party reproducing the platform's own statement, at the moment it flipped.
yipinwong, on Hacker News, [56] criticized the announcement chart axes ↗, arguing the "Agent Coding Performance by Effort Level" chart truncates its y-axis to roughly a 0-20 range so the visual emphasis exceeds the plotted delta. Unverified: the chart images were never downloaded or visually inspected for this packet, so this stands as an attributed presentation complaint, not a verified chart analysis.
janstice, on Hacker News, [57] reported a behavior gap between channels ↗: a detailed answer about Tiananmen Square from oxalpha.com, a third-party front-end presented as Ox Alpha, against a refusal from GLM-5.3-Flash asked directly through OpenRouter chat. Unverified, single report, not reproduced; the mechanism could be harness- or channel-specific.
Josh Friedman, on X, [60] reported a coding-plan billing mismatch ↗: glm-5.3-flash requests billed against glm-5.3 quota. Unverified, one observer, no provider response captured.
WEEX AI Labs, on X, [61] aggregated traffic figures in Chinese ↗, attributing to Zhipu that all preview traffic ran on Chinese chips and claiming roughly 20 trillion tokens over six days, making Ox Alpha the highest-traffic single model in OpenRouter history to that point. Unverified aggregation; no primary platform disclosure was located. Both vendors themselves assert only the qualitative version, "the most popular model of the week".
claims
- claim
- GLM-5.3-Flash announced 2026-08-26 as first natively multimodal GLM-5 model
- status
- supported
- claim
- API code glm-5.3-flash, 1M context, image_url input, thinking cannot be disabled
- status
- supported
- claim
- 320B/18B, hybrid sparse-plus-linear attention; config corroborates mechanisms, not performance
- status
- supported
- claim
- MIT weights on Hugging Face, revision 3f1971b7, arXiv 2602.15763
- status
- supported
- claim
- DeepSWE leaderboard 63% +/- 4%, consistent with vendor 63.4
- status
- supported
- claim
- Vendor-reported benchmark suite including self-run Z.ai Code Bench
- status
- supported
- claim
- OpenRouter listing state, pricing, context, AA card as of 2026-08-26T13:59:01Z
- status
- supported
- claim
- Discounted index basis plus GLM Coding Plan quota and off-peak terms
- status
- supported
- claim
- Z.ai's first-party ox-alpha statement, identical on two surfaces
- status
- supported
- claim
- OpenRouter's first-party stealth record, retention terms, close-out message
- status
- supported
- claim
- Identity asserted by two first parties, not mechanically demonstrated
- status
- supported
- claim
- z-ai/ox-alpha returns 404; stealth/ox-alpha retained as record
- status
- supported
- claim
- Point-in-time HN and Reddit counts; community ID roughly eight hours early
- status
- supported
- claim
- Coding-plan billing mismatch, single report
- status
- unverified
- claim
- Community traffic-scale aggregation without primary disclosure
- status
- unverified
- claim
- Cross-channel refusal discrepancy, single report
- status
- unverified
- claim
- Chart y-axis truncation critique, pixels uninspected
- status
- unverified
Open questions this packet does not close: whether any independent evaluation reproduces AutomationBench, Terminal-Bench 2.1, Agent's Last Exam, or the in-house Z.ai Code Bench results; what checkpoint-level verification of the Ox Alpha equivalence would even require now that the stealth endpoint is gone; whether the Chinese-accelerator cluster claims can be confirmed from any primary disclosure; whether either vendor intends to publish authoritative token-volume figures for the stealth period; whether the billing mismatch is isolated or systemic; whether every serving channel implements identical behavior; and whether any archive captured the live stealth page before it changed.
receipts
Vendor:
- [43] Z.ai GLM-5.3-Flash announcement page ↗ - receipts/zai-blog-glm-5.3-flash.html
- [45] Z.ai developer guide for GLM-5.3-Flash ↗ - receipts/zai-docs-vlm-glm-5.3-flash.md
Platform:
- [46] OpenRouter models API snapshot ↗ - receipts/openrouter-api-v1-models.json
- [47] OpenRouter z-ai/glm-5.3-flash model page ↗ - receipts/openrouter-zai-glm-5.3-flash.html
- [48] OpenRouter stealth/ox-alpha post-reveal record page ↗ - receipts/openrouter-stealth-ox-alpha.html
- [49] OpenRouter 404 probe for the retired z-ai/ox-alpha endpoints ↗ - receipts/probe-ox-alpha-endpoints.json
Weights:
- [50] Hugging Face repository metadata record ↗ - receipts/hf-api-zai-org-glm-5.3-flash.json
- [51] Hugging Face model card README ↗ - receipts/hf-glm-5.3-flash-readme.md
- [52] Released checkpoint config.json ↗ - receipts/hf-glm-5.3-flash-config.json
Independent benchmark:
- [53] DeepSWE public leaderboard ↗ - receipts/deepswe-datacurve.html
Community:
- [54] Hacker News GLM-5.3-Flash submission ↗ - receipts/hn-algolia-search.json
- [55] Hacker News comment by Bluestein reproducing the close-out response ↗ - receipts/hn-item-49449939.json
- [56] Hacker News comment by yipinwong on chart axes ↗ - receipts/hn-item-49450151.json
- [57] Hacker News comment by janstice on channel behavior parity ↗ - receipts/hn-item-49456661.json
- [58] r/LocalLLaMA pre-announcement identification thread ↗ - receipts/rdt-read-oxalpha-confirmation.json
- [59] r/LocalLLaMA release thread ↗ - receipts/rdt-search-glm53-flash.json
- [60] X post reporting the coding-plan billing mismatch ↗ - receipts/bird-search-glm53-flash.json
- [61] X post aggregating stealth-period traffic figures ↗ - receipts/bird-search-glm53-flash.json
Archive:
- [62] Wayback availability probe for openrouter.ai/z-ai/ox-alpha ↗ - receipts/wayback-openrouter.json