BENCHBAITSTORY · MONITORING

AI NEWS, TRACKED AND ROASTED.

z.ai names its stealth model glm-5.3-flash, eight hours after the crowd did

Z.ai released GLM-5.3-Flash on 2026-08-26, calling it the first natively multimodal model in the GLM-5 series, and used the same announcement to say where the model had already been. The blog post and the developer guide carry the line identically: "Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week - with all of this traffic served on Chinese AI chips."

OpenRouter tells it from its own side of the wire. The retained [48] stealth/ox-alpha record page ↗ states that the model was "developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash", that prompts and completions from the testing period were retained by the provider and not used for training under the Stealth Model Terms, and its close-out routing message hands callers forward: "Thank you for participating in the Stealth Ox Alpha testing period. This model was ZAI's GLM-5.3 Flash. Use it now: [47] OpenRouter ↗". The stealth listing was created 2026-08-20T20:04:55.369Z under the author name "stealth". Today [49] the z-ai/ox-alpha endpoints probe ↗ returns 404 Not Found; the record page is what is left.

Neither vendor got there first. An [58] r/LocalLLaMA identification thread ↗ posted 2026-08-26T06:28:10Z reached 458 points by matching three fingerprints - native vision, a 1M-token window, DeepSWE around 63% - and named the pairing about eight hours before Z.ai's announcement.

What shipped, per Z.ai: 320B total and 18B active parameters, a hybrid sparse-plus-linear-attention architecture new to the GLM series with mHC scaling additions, 45 layers against 92 in the GLM-4.5 series, a 30T-token multimodal corpus, and claimed reductions of 3.0x in attention compute and 4.4x in KV cache versus GLM-5.3. The released [52] config.json ↗ corroborates the structural mechanisms - Glm5NextForConditionalGeneration, index_topk 2048, index_kpool 4 compressed, hc_mult 4 - but not the performance figures. Weights are MIT-licensed at huggingface.co/zai-org/GLM-5.3-Flash, repo created 2026-08-25T06:43:14Z at revision 3f1971b7b5f7a528c9c4ef6212c8785298a8c24a, citing arXiv 2602.15763, with local recipes for SGLang, vLLM, TokenSpeed, and KTransformers.

OpenRouter listed z-ai/glm-5.3-flash publicly at 2026-08-26T13:59:01Z (canonical slug z-ai/glm-5.3-flash-20260826), modality text+image+video to text, top-provider context 1,048,576 tokens with a 131,072-token completion cap, reasoning mandatory and defaulting to max effort. Its embedded Artificial Analysis card shows intelligence index 57.5, coding 71.5, agentic 58.2.

Router rates, as listed:

prompt
price
$0.075 / M tokens
completion
price
$0.25 / M tokens
cache read
price
$0.015 / M tokens

Z.ai's own guide does not publish a per-token rate. It frames cost through a discounted Artificial Analysis basis - Intelligence Index v4.1.1 score 57 at $0.045 per task - and routes subscribers through the GLM Coding Plan instead:

availability
terms
fully available to subscribers
quota versus GLM-5.3
terms
3x
peak calls
terms
standard points
off-peak calls, including all day weekends
terms
50% of standard points

On the platform side the model is exposed as code glm-5.3-flash with text parameters consistent with GLM-5.3, a 1M-token context window, image input via image_url blocks, and no way to disable thinking - the guide recommends max reasoning_effort.

Benchmarks are vendor-reported. Z.ai claims wins over GLM-5.2 across six coding and agentic benchmarks, naming DeepSWE v1.1 at 63.4 against 46.2 and AutomationBench v1.0.6 at 48.8 against 26.2, plus near-parity with Claude Opus 4.8 at max effort on its own in-house Z.ai Code Bench v1.0 - 29.0 against 29.5, run on Claude Code 2.1.207. One number has an outside witness: the independent [53] DeepSWE public leaderboard ↗ lists glm-5.3-flash [max] at 63% +/- 4%, which is consistent with the vendor's 63.4 without being a measurement of it. The margin is wide enough that exact-value equality is untestable.

The identity itself is asserted, not demonstrated. "Ox Alpha was GLM-5.3-Flash" rests on exactly two agreeing first parties plus community fingerprint matching reported pre-reveal. No checkpoint-hash or serving-stack-level verification exists here, the historical stealth endpoint no longer answers, and the [62] Wayback availability probe ↗ returned no archived snapshots of the stealth-era page.

what people are saying

Reception scale at observation: the [54] Hacker News submission ↗ sat at 922 points and 457 comments; the [59] r/LocalLLaMA release thread ↗ at 1157 points and 378 comments, with a dedicated former-ox-alpha megathread at 203 points and 138 comments. These are point-in-time counts and they decay.

Bluestein, on Hacker News, [55] posted the handover as a live API response ↗: a 404 from the chat completions endpoint for model stealth/ox-alpha whose error text matches OpenRouter's official close-out naming ZAI's GLM-5.3 Flash. A third party reproducing the platform's own statement, at the moment it flipped.

yipinwong, on Hacker News, [56] criticized the announcement chart axes ↗, arguing the "Agent Coding Performance by Effort Level" chart truncates its y-axis to roughly a 0-20 range so the visual emphasis exceeds the plotted delta. Unverified: the chart images were never downloaded or visually inspected for this packet, so this stands as an attributed presentation complaint, not a verified chart analysis.

janstice, on Hacker News, [57] reported a behavior gap between channels ↗: a detailed answer about Tiananmen Square from oxalpha.com, a third-party front-end presented as Ox Alpha, against a refusal from GLM-5.3-Flash asked directly through OpenRouter chat. Unverified, single report, not reproduced; the mechanism could be harness- or channel-specific.

Josh Friedman, on X, [60] reported a coding-plan billing mismatch ↗: glm-5.3-flash requests billed against glm-5.3 quota. Unverified, one observer, no provider response captured.

WEEX AI Labs, on X, [61] aggregated traffic figures in Chinese ↗, attributing to Zhipu that all preview traffic ran on Chinese chips and claiming roughly 20 trillion tokens over six days, making Ox Alpha the highest-traffic single model in OpenRouter history to that point. Unverified aggregation; no primary platform disclosure was located. Both vendors themselves assert only the qualitative version, "the most popular model of the week".

claims

claim-01-release-identity
claim
GLM-5.3-Flash announced 2026-08-26 as first natively multimodal GLM-5 model
status
supported
claim-02-api-naming-context
claim
API code glm-5.3-flash, 1M context, image_url input, thinking cannot be disabled
status
supported
claim-03-architecture-vendor-spec
claim
320B/18B, hybrid sparse-plus-linear attention; config corroborates mechanisms, not performance
status
supported
claim-04-openweights-license
claim
MIT weights on Hugging Face, revision 3f1971b7, arXiv 2602.15763
status
supported
claim-05-deepswe-independent-match
claim
DeepSWE leaderboard 63% +/- 4%, consistent with vendor 63.4
status
supported
claim-06-vendor-benchmark-suite
claim
Vendor-reported benchmark suite including self-run Z.ai Code Bench
status
supported
claim-07-openrouter-listing-state
claim
OpenRouter listing state, pricing, context, AA card as of 2026-08-26T13:59:01Z
status
supported
claim-08-zai-pricing-posture
claim
Discounted index basis plus GLM Coding Plan quota and off-peak terms
status
supported
claim-09-oxalpha-zai-statement
claim
Z.ai's first-party ox-alpha statement, identical on two surfaces
status
supported
claim-10-oxalpha-openrouter-record
claim
OpenRouter's first-party stealth record, retention terms, close-out message
status
supported
claim-11-identity-boundary
claim
Identity asserted by two first parties, not mechanically demonstrated
status
supported
claim-12-stealth-absence-today
claim
z-ai/ox-alpha returns 404; stealth/ox-alpha retained as record
status
supported
claim-13-reception-scale
claim
Point-in-time HN and Reddit counts; community ID roughly eight hours early
status
supported
claim-14-billing-report-unverified
claim
Coding-plan billing mismatch, single report
status
unverified
claim-15-scale-figures-unverified
claim
Community traffic-scale aggregation without primary disclosure
status
unverified
claim-16-behavior-parity-dispute-unverified
claim
Cross-channel refusal discrepancy, single report
status
unverified
claim-17-chart-truncation-critique
claim
Chart y-axis truncation critique, pixels uninspected
status
unverified

Open questions this packet does not close: whether any independent evaluation reproduces AutomationBench, Terminal-Bench 2.1, Agent's Last Exam, or the in-house Z.ai Code Bench results; what checkpoint-level verification of the Ox Alpha equivalence would even require now that the stealth endpoint is gone; whether the Chinese-accelerator cluster claims can be confirmed from any primary disclosure; whether either vendor intends to publish authoritative token-volume figures for the stealth period; whether the billing mismatch is isolated or systemic; whether every serving channel implements identical behavior; and whether any archive captured the live stealth page before it changed.

receipts

Vendor:

Platform:

Weights:

Independent benchmark:

Community:

Archive:

LAST VERIFIED 2026-08-28

WHAT CHANGED

THE TIMELINE

REVISION 4 · UPDATE · 2026-08-28ROAST / SATIRE

z.ai names its stealth model glm-5.3-flash, eight hours after the crowd did

  • QUALIFIEDZ.ai publicly announced GLM-5.3-Flash ("GLM-5.3-Flash: Frontier Intelligence, Flash Cost") on 2026-08-26 as the first natively multimodal model in its GLM-5 series.RECEIPTS [43][44][45]
  • QUALIFIEDThe Z.ai platform exposes the model under code glm-5.3-flash with text parameters consistent with GLM-5.3, supports a 1M-token context window and image input via image_url blocks, and cannot disable thinking (reasoning_effort recommended max).RECEIPTS [45]
  • QUALIFIEDPer Z.ai, GLM-5.3-Flash has 320B total / 18B active parameters, introduces a hybrid sparse-plus-linear-attention architecture to the GLM series with mHC scaling additions and a 30T-token multimodal corpus (45 layers vs 92 in GLM-4.5 series; claimed 3.0x attention-compute and 4.4x KV-cache reductions vs GLM-5.3; IndexPool compresses four indexer key vectors to one). The released checkpoint's config.json corroborates the structural mechanisms (Glm5NextForConditionalGeneration, index_topk 2048, index_kpool 4 compressed, hc_mult 4) though not the performance figures.RECEIPTS [44][45][52]
  • QUALIFIEDWeights are published under MIT license at huggingface.co/zai-org/GLM-5.3-Flash (repo created 2026-08-25T06:43:14Z, revision 3f1971b7b5f7a528c9c4ef6212c8785298a8c24a), citing arXiv 2602.15763, with local deployment recipes for SGLang, vLLM, TokenSpeed, and KTransformers.RECEIPTS [50][51]
  • QUALIFIEDThe independent public DeepSWE leaderboard lists glm-5.3-flash [max] at 63% +/- 4%, which is consistent with (not a measurement of) Z.ai's separately framed claim of 63.4 vs GLM-5.2's 46.2 on DeepSWE v1.1.RECEIPTS [44][53]
  • QUALIFIEDZ.ai attributes to GLM-5.3-Flash: wins over GLM-5.2 across six coding/agentic benchmarks (named examples DeepSWE v1.1 63.4 vs 46.2; AutomationBench v1.0.6 48.8 vs 26.2), near-parity with Claude Opus 4.8 at max effort on its own Z.ai Code Bench v1.0 (29.0 vs 29.5, run on Claude Code 2.1.207), and Artificial Analysis Intelligence Index v4.1.1 score 57 at $0.045/task discounted; some evaluations use GPT-5.6-luna (medium) as judge. All numbers are vendor-reported.RECEIPTS [44][45][51]
  • QUALIFIEDAs of 2026-08-26T13:59:01Z OpenRouter publicly lists z-ai/glm-5.3-flash (canonical z-ai/glm-5.3-flash-20260826): $0.075/M prompt, $0.25/M completion, $0.015/M cache-read pricing; modality text+image+video->text; top-provider context_length 1,048,576 with 131,072 completion cap; mandatory reasoning defaulting to max effort; an embedded Artificial Analysis card showing intelligence index 57.5, coding 71.5, agentic 58.2.RECEIPTS [46][47]
  • QUALIFIEDZ.ai frames cost through a discounted Artificial Analysis index basis ($0.045/task at score 57) rather than publishing a per-token rate in the reviewed guide, and makes GLM-5.3-Flash fully available to GLM Coding Plan subscribers with 3x the quota of GLM-5.3, off-peak calls (including all day weekends) consuming 50% of standard points.RECEIPTS [45]
  • QUALIFIEDZ.ai states first-party, identically on its blog and developer guide: "Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week - with all of this traffic served on Chinese AI chips."RECEIPTS [44][45]
  • QUALIFIEDOpenRouter states first-party on the retained stealth/ox-alpha page that the model was "developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash", that prompts and completions from the testing period were retained by the provider and not used for training (governed by Stealth Model Terms), and its close-out routing message directs users to z-ai/glm-5.3-flash; the listing record shows author "stealth", creation 2026-08-20T20:04:55.369Z, mandatory low/high/max reasoning, and a 1,048,576-token window.RECEIPTS [48]
  • QUALIFIEDThe product identity relation "Ox Alpha was GLM-5.3-Flash" rests on exactly two agreeing first parties (Z.ai vendor statement; OpenRouter platform statement) plus community fingerprint matching reported pre-reveal; no checkpoint-hash or serving-stack-level verification exists in this packet, and the historical stealth endpoint no longer answers, so equivalence is asserted by both vendors rather than mechanically demonstrated.RECEIPTS [45][48][49][58]
  • QUALIFIEDAs of observation, OpenRouter serves no z-ai/ox-alpha identifier (endpoints API returns 404 Not Found) while preserving stealth/ox-alpha as a revealed record page.RECEIPTS [48][49]
  • QUALIFIEDReception scale at observation: Hacker News main thread 922 points / 457 comments; r/LocalLLaMA release thread 1157 points / 378 comments with a dedicated former-ox-alpha megathread; a community identification thread reached 458 points about eight hours before Z.ai's announcement, correctly naming the pairing hours ahead of either vendor.RECEIPTS [54][58][59]
  • UNVERIFIEDAn individual X user reports that glm-5.3-flash requests made through the Z.ai coding plan were billed against glm-5.3 quota; single report, no second observer and no provider response captured.RECEIPTS [60]
  • UNVERIFIEDCommunity-sourced traffic-scale figures circulate without primary disclosure found this pass: "~10T+ tokens/day served free" during the stealth week (HN commenter) and "~20 trillion tokens over six days, highest-traffic single model in OpenRouter history to that point" (Chinese-language aggregation attributing confirmation to Zhipu); OpenRouter and Z.ai themselves assert only the qualitative "most popular model of the week".RECEIPTS [45][61]
  • UNVERIFIEDOne HN commenter reports obtaining a detailed answer about Tiananmen Square from a third-party front-end identified as Ox Alpha (oxalpha.com) while receiving refusal from GLM-5.3-Flash asked directly through OpenRouter chat; raises an unresolved serving-layer behavior-parity question about whether all channels implement identical policy behavior. Not reproduced by this lane.RECEIPTS [57]
  • UNVERIFIEDA HN critic contends the vendor "Agent Coding Performance by Effort Level" chart truncates its y-axis (roughly 0-20 range shown) so vendor improvement appears larger than the plotted deltas warrant; the underlying chart image was not visually inspected by this reporter, so the critique remains an attributed presentation complaint rather than verified chart analysis.RECEIPTS [56]
REVISION 3 · UPDATE · 2026-08-25ROAST / SATIRE

z.ai says every glm-5.3 gain came from post-training; the cost argument that followed agrees on no number

  • VERIFIEDZ.ai says GLM-5.3 uses the same base model as GLM-5.2 and attributes its gains to post-training.RECEIPTS [16]
  • VERIFIEDZ.ai reports a DeepSWE v1.1 increase from 46.2 to 66.9 for GLM-5.3.RECEIPTS [16]
  • QUALIFIEDBenchSift lists a GLM-5.3 max-effort configuration at 69.0% and $3.99 displayed cost, but it aggregates rather than runs the benchmark.RECEIPTS [17]
  • QUALIFIEDTogether AI says GLM-5.3 matched a Claude Fable 5 configuration on single-shot results at roughly one-fifth the displayed cost and led on retries.RECEIPTS [18]
  • QUALIFIEDCurrent community commentary restates the GLM-5.3 cost argument as tasks completed per fixed budget rather than as the displayed-cost ratio bound in this packet.RECEIPTS [19][20]
  • QUALIFIEDThe GLM-5.3 cost and solve-rate figures circulating in current commentary (87.6% for about $16 against 69.7% at $21.63; 17 tasks against 3 on a fixed $100 budget) match no figure bound in this packet and cite no source.RECEIPTS [16][17][19][20]
  • QUALIFIEDAggregator skepticism around GLM-5.3 in this pass is about how intelligence-versus-cost is plotted on Artificial Analysis, not about BenchSift, and the skeptical post is itself disputed by its own commenters.RECEIPTS [17][23]
  • QUALIFIEDZ.ai’s same-base-model post-training framing is repeated in current commentary and read as a structural shift in how open labs ship improvements.RECEIPTS [16][21][24]
  • UNVERIFIEDAt least one current restatement of the post-training claim adds characterisations no receipt in this packet supports, including that GLM-5.3 beat Claude Fable 5 on "the one independent benchmark nobody could game" and that it is free with open weights.RECEIPTS [22]
  • UNVERIFIEDTerminal-Bench 3.0 figures of 28.3% up from 4.6% circulate in current GLM-5.3 commentary but are bound to no receipt in this packet.RECEIPTS [21][22]
REVISION 2 · UPDATE · 2026-08-23ROAST / SATIRE

glm-5.3 reached the comparison board’s discount aisle. the benchmark still wants the receipt settings.

  • QUALIFIEDBenchSift lists a GLM-5.3 max-effort configuration at 69.0% and $3.99 displayed cost, but it aggregates rather than runs the benchmark.RECEIPTS [2]
  • QUALIFIEDTogether AI says GLM-5.3 matched a Claude Fable 5 configuration on single-shot results at roughly one-fifth the displayed cost and led on retries.RECEIPTS [3]
REVISION 1 · LAUNCH · 2026-08-14ROAST / SATIRE

glm-5.3 says post-training did all the lifting. pretraining has entered a performance review.

  • VERIFIEDZ.ai says GLM-5.3 uses the same base model as GLM-5.2 and attributes its gains to post-training.RECEIPTS [1]

EVERY JOKE HAS RECEIPTS

SOURCES

  1. [1] PRIMARY · Z.aiGLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesLaunch claims and vendor-reported benchmark methodology ↗
  2. [2] AGGREGATOR · BenchSiftDeepSWEAggregated comparison board; not an independent benchmark run ↗
  3. [3] SIGNAL · Together AI on XBenchmark comparison discussionDiscovery and comparison commentary ↗
  4. [16] PRIMARY · Z.aiGLM-5.3: Frontier Coding with Emergent Cyber CapabilitiesLaunch claims and vendor-reported benchmark methodology ↗
  5. [17] AGGREGATOR · BenchSiftDeepSWEAggregated comparison board; not an independent benchmark run ↗
  6. [18] SIGNAL · Together AI on XBenchmark comparison discussionDiscovery and comparison commentary ↗
  7. [19] SIGNAL · shipfrontier (@shipfrontierai) on XSame $100 budget on DeepSWE: GLM-5.3 solves 17 tasks, Fable 5 solves 3Community commentary restating the cost argument as tasks-per-budget ↗
  8. [20] SIGNAL · Fabiano Firmo (@FabianoFirmo) on XFour tries with GLM-5.3 beat Fable 5 on both solve rate and total costCirculating cost/solve figures that diverge from every figure bound in this packet ↗
  9. [21] SIGNAL · Affine (@affine_io) on XPost-training deserves its own competitive arenaCommentary reading the same-base-model post-training claim as a structural argument ↗
  10. [22] SIGNAL · AI Mastery Guide (@aiseomastery) on XGLM 5.3 just dropped. Same model as before.Community restatement that adds characterisations no receipt in this packet supports ↗
  11. GLM-5.3 is out on AA, and I’m fed up with their Intelligence/cost plot

    Verbatim: "I think AA’s intelligence/cost plot is seriously misleading, so I decided to make my own."

    Aggregator-presentation skepticism, itself disputed in its own comments ↗
  12. Open labs are finally embracing the power of continued post training

    Verbatim: "Continued post training on an existing model seems to be fully capable of yielding generational leaps in performance without retraining a new base model."

    Community reading of what a same-base-model post-training claim means ↗
  13. [43] PRIMARY · Z.aiGLM-5.3-Flash: Frontier Intelligence, Flash Costprimary ↗
  14. [44] PRIMARY · Z.aiBlog post content bundle for z.ai/blog/glm-5.3-flashThe captured bundle is held as the story's primary-source record; its original asset URL now returns 404, so this receipt is not published.primary
  15. [45] PRIMARY · Z.aiGLM-5.3-Flash overview guide (Z.ai developer docs)primary ↗
  16. [46] PRIMARY · OpenRouterOpenRouter models API snapshot including z-ai/glm-5.3-flash recordprimary ↗
  17. [47] PRIMARY · OpenRouterOpenRouter model page Z.ai: GLM 5.3 Flashprimary ↗
  18. [48] PRIMARY · OpenRouterOpenRouter stealth/ox-alpha page (post-reveal state)primary ↗
  19. [49] PRIMARY · OpenRouterOpenRouter endpoints API response for id z-ai/ox-alpha (probe)supporting ↗
  20. [50] PRIMARY · Hugging Face (repo host)Hugging Face API model record for zai-org/GLM-5.3-Flashprimary ↗
  21. [51] PRIMARY · Z.ai / GLM-5 TeamGLM-5.3-Flash model card READMEprimary ↗
  22. [52] PRIMARY · Z.ai / GLM-5 TeamGLM-5.3-Flash config.json (revision at head of observed snapshot)primary ↗
  23. [53] INDEPENDENT · DatacurveDeepSWE public leaderboardindependent ↗
  24. [54] SIGNAL · Hacker News communityHN submission "GLM-5.3-Flash"reaction ↗
  25. [55] SIGNAL · Hacker News communityHN comment by Bluestein reproducing the OpenRouter close-out API responsesupporting ↗
  26. [56] SIGNAL · Hacker News communityHN comment by yipinwong criticizing announcement chart axescontradiction ↗
  27. [57] SIGNAL · Hacker News communityHN comment by janstice reporting a behavior-parity discrepancy testcontradiction ↗
  28. r/LocalLLaMA thread "First serious confirmation. Ox Alpha is GLM-5.3-Flash"

    Thread predates the announcement by roughly eight hours; body cites a since-deleted X post by romanchernin matching three fingerprints (native vision, 1M-token window, DeepSWE ~63%) to identify ox-alpha as GLM-5.3-Flash; poster later edited that the cited X post was deleted with screenshots circulating in comments. 458 points / 156 comments as of read; top reactions include praise for stealth-period UI/design output quality (attributed personal experience, unverified) and speculation threads on parameter count resolved later the same day by the official 320B figure.

    reaction ↗
  29. r/LocalLLaMA release thread "GLM-5.3-Flash: Frontier Intelligence, Flash Cost"

    Release-day reception signal: 1157 points / 378 comments as of bounded search read; the follow-up megathread "[Megathread] GLM-5.3-Flash - former ox-alpha" held 203 points / 138 comments.

    reaction ↗
  30. [60] SIGNAL · X user Josh FriedmanX status reporting coding-plan billing mismatchreaction ↗
  31. [61] SIGNAL · X user WEEX AI LabsChinese-language X aggregation of Ox Alpha traffic figuresreaction ↗
  32. [62] AGGREGATOR · Internet ArchiveWayback Machine availability check for openrouter.ai/z-ai/ox-alphasupporting ↗