What is Ox Alpha? The stealth model revealed as Z.ai GLM-5.3-Flash
Ox Alpha was the anonymous name of Z.ai GLM-5.3-Flash on OpenRouter and OpenCode. Who made it, specs, whether it is still free, and how to use it now.
Read as MarkdownOx Alpha was the anonymous test name of GLM-5.3-Flash, a coding and agent model from Z.ai. It appeared on OpenRouter and OpenCode as a "stealth" model on August 20, 2026, with no maker named. On August 26, Z.ai released GLM-5.3-Flash and confirmed it was Ox Alpha. The name ox-alpha no longer works; to use the same model today, call GLM-5.3-Flash. Every fact below comes from Z.ai, OpenRouter or OpenCode, checked on September 29, 2026.
Who made Ox Alpha?
Z.ai. Its launch post for GLM-5.3-Flash says that before release, the company "tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback", and that "it quickly became the most popular model of the week".
OpenRouter's page for the model confirms it: "This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash." OpenCode now lists the model as GLM-5.3-Flash, "Formerly ox-alpha."
While it was anonymous, people guessed at Gemini, a GLM model and others. The GLM guess was close but not exact: Ox Alpha is GLM-5.3-Flash, not GLM-5.3, which Z.ai released separately as a text-only model.
What was a "stealth model"?
A model that a lab puts on a public platform under a code name before launch, to collect real usage and feedback. OpenRouter hosts these under a stealth/ prefix; Ox Alpha's ID was stealth/ox-alpha. When the lab announces the model, OpenRouter reveals who made it and points to the model's real name.
What are Ox Alpha's specs?
These are the published specs of GLM-5.3-Flash, the model behind Ox Alpha:
| Spec | Value |
|---|---|
| Developer | Z.ai |
| Stealth release | August 20, 2026, on OpenRouter and OpenCode |
| Official release | August 26, 2026, as GLM-5.3-Flash |
| Parameters | 320B total, 18B active |
| Architecture | "A hybrid architecture combining sparse and linear attention" |
| Context window | 1,048,576 tokens (1M) |
| Max output | 128K tokens |
| Input | Text, images and video (Z.ai's docs also list files) |
| Output | Text |
| Weights | Public on Hugging Face under the MIT license |
| Model ID | glm-5.3-flash on Z.ai's API; z-ai/glm-5.3-flash on OpenRouter |
OpenRouter described Ox Alpha as "a reasoning model designed for coding, sustained agentic work, and production workloads." For how the context window and output limit work together, see context window vs max output tokens.
How good is Ox Alpha?
Z.ai published this comparison in its GLM-5.3-Flash launch post:
| Benchmark | GLM-5.3-Flash | GLM-5.2 | Claude Opus 4.8 | GPT-5.6 Terra | Gemini 3.7 Flash |
|---|---|---|---|---|---|
| Terminal Bench 2.1 | 84.3 | 81.0 | 85.0 | 87.4 | 85.8 |
| DeepSWE v1.1 | 63.4 | 46.2 | 58.0 | 69.6 | 65.3 |
| AutomationBench v1.0.6 | 48.8 | 26.2 | 41.0 | 37.2 | 52.3 |
Z.ai sums it up as "approaching Claude Opus 4.8 overall". It adds that on its in-house Z.ai Code Bench v1.0, the model "at max effort nearly matches Claude Opus 4.8 (29.0 vs. 29.5)", and that it scores 57 on the Artificial Analysis Intelligence Index v4.1.1 "at just $0.045 per task (discounted)".
These are the developer's own results on benchmarks it chose. Test the model on your own tasks before you rely on it.
Is Ox Alpha still free?
No, and the ox-alpha name itself is gone. OpenRouter's stealth/ox-alpha listing has no active endpoints, and the model no longer appears in its model list. The test ran from August 20 until Z.ai launched GLM-5.3-Flash on August 26. Because the listing is gone, the price during the test can no longer be checked on OpenRouter or Z.ai.
GLM-5.3-Flash is a paid model now. Z.ai's pricing page lists it at $0.15 per million input tokens, $0.03 per million cached input tokens and $0.50 per million output tokens. OpenRouter lists the same input and output rates for z-ai/glm-5.3-flash. Prices change, so check both pages before you commit.
How do I use Ox Alpha now?
Call GLM-5.3-Flash under its real name. There are three ways:
- Z.ai's API, with the model code
glm-5.3-flash. - OpenRouter, with the model ID
z-ai/glm-5.3-flash. - OpenCode, where the model now appears as GLM-5.3-Flash.
Coding agents that let you set a base URL and model name can use it through either API. Z.ai also released a newer variant, GLM-5.3-FlashX; its model code is glm-5.3-flashx.
Can I run Ox Alpha locally?
The weights are public. Z.ai's launch post says: "The model weights of GLM-5.3-Flash are publicly available on Hugging Face." The model card lists the MIT license. With 320 billion parameters in total, running it takes multi-GPU, data-center hardware even though only 18 billion are active per token, so most people call it through an API.
Were my prompts kept while it was a stealth model?
OpenRouter's page says prompts and completions for this model "were retained" by the company that ran it, Z.ai, and "are not used for training". If you sent private code to Ox Alpha through OpenRouter, assume Z.ai kept it.
What are the alternatives to Ox Alpha for coding agents?
If you want a coding model you can call from Claude Code or Codex with one key and pay-as-you-go billing, these models are live on SeedRouter today:
- DeepSeek V4.1 Flash, DeepSeek's fast, low-cost model with a 1M-token context. Setup: DeepSeek V4.1 Flash in Claude Code and Codex.
- Kimi K3, Moonshot AI's open-weight model for long coding sessions. Setup: Kimi K3 in Claude Code and Codex.
- GPT-6, OpenAI's GPT-6 Astra, Sol and Luna. Setup: GPT-6 in Codex.
SeedRouter does not serve GLM-5.3-Flash. You top up once with no subscription, credits never expire, and a request that fails is not charged. Live prices are on each model page.
Frequently asked questions
Is Ox Alpha GLM?
Yes. Ox Alpha was Z.ai's GLM-5.3-Flash. Z.ai confirmed it when it launched the model on August 26, 2026, and OpenRouter's page now says the same.
Is Ox Alpha Chinese?
It was made by Z.ai, which says all of the stealth traffic was "served on Chinese AI chips".
Is Ox Alpha open source?
Its weights are. GLM-5.3-Flash is on Hugging Face under the MIT license.
Why did Ox Alpha disappear?
Because the test ended. When Z.ai launched GLM-5.3-Flash, the ox-alpha name was retired and the model moved to its real name.
Is Ox Alpha better than Claude?
Z.ai reports that GLM-5.3-Flash at max effort "nearly matches Claude Opus 4.8" on its own Code Bench. That is one benchmark from the model's developer, not an independent comparison.



