Mystery "Ox Alpha" Model Breaks Out on OpenRouter: Outscores Claude and GPT-5.6 in Early Tests
On August 20, AI model aggregator OpenRouter quietly listed a new anonymous model under the codename Ox Alpha. The platform disclosed no developer, describing it only as a frontier model aimed at coding, long-horizon agent tasks, and production environments — free during preview, and immediately flooded with developer testing.
Basic Specifications
Public details are limited:
- Context window: ~1,048,000 tokens
- Max output: 131,000 tokens
- Input modalities: text, image, video
- Pricing: free during preview
- Companion capacity: open-source terminal coding agent OpenCode integrated the same day, dedicating 100 trillion tokens per day for it
Benchmark Performance
Within hours, multiple developers were stress-testing it. Independent researcher Ben Davis ran 10 sub-tasks from the software engineering benchmark DeepSWE — Ox Alpha passed roughly 80%, ahead of Claude Fable 5 (65%), GLM-5.3 and Grok 4.6 (62% each), and GPT-5.6 Sol (52%). These results come from an individual test run — DeepSWE's official leaderboard has not yet added Ox Alpha.
For a model less than 12 hours old and nameless, that was enough to ignite collective curiosity: who built Ox Alpha?
The Community's First Four Suspects
Early discussion focused on four candidates:
- Xiaomi MiMo: On August 18, Xiaomi CFO Lin Shiwei revealed a new MiMo model was in training. Xiaomi had already done this once before — MiMo-V2-Pro appeared anonymously as Hunter Alpha on OpenRouter in March.
- Zhipu AI: The 1M context window, multimodal inputs, and agent-focused positioning all aligned with Zhipu's recent model roadmap. Zhipu also has prior precedent — Pony Alpha, which appeared on OpenRouter in February, was later confirmed to be GLM-5.
- Tencent Hunyuan: Some Reddit users guessed "HY 4," reasoning that few vendors can underwrite free access at this scale.
- Google Gemini: The combination of million-token context plus native video input is more typical of the Gemini family — but one tester tried to extract a self-identification and Ox Alpha claimed to be Gemini. LLMs can easily generate plausible-but-completely-false identities from training data, system prompts, or conversational context, and the same model often flips its answer depending on how it's prompted.
Model Attribution as a Technical Discipline
Anonymity can hide the brand, but it struggles to hide engineering fingerprints. The community's identity-inference methods have become increasingly technical:
Tokenizer Fingerprint
Every LLM passes input through a tokenizer that splits strings into tokens. Different vendors' vocabularies and splitting strategies almost always produce different token counts for the same text. One analyst ran 25 different prompts through both Ox Alpha and GLM-5.3 — raw token counts matched exactly, except Ox Alpha consistently added exactly 75 extra tokens each time.
Vision-Encoder Token Budget
Multimodal models convert video frames into tokens via a vision encoder, where token consumption is determined by frame-sampling rate, duration scaling, and resolution-handling strategy. With four parameter-fixed test videos, an analyst found Ox Alpha's vision-encoder token consumption matched Zhipu's GLM-5V-Turbo exactly across all four, and aligned on three independent encoder design choices: frame-rate-independent sampling, ~147 tokens/sec duration scaling, and per-frame resolution scaling.
Dirty-Token Collision
Some tokens appear so rarely or anomalously in training data that they become sensitivity points, and different model families have different "dirty token" distributions. Testers found Ox Alpha's reaction pattern to a batch of adversarial strings matched GLM-5.3 closely, but only 18–22% overlap with Gemini, DeepSeek, and Kimi.
Behavioral Style and Execution Cadence
Ox Alpha emits roughly 1.3 emojis per 1,000 characters in long-form answers — matching the GLM and Qwen family signature (section emojis, red/yellow/green status dots, checklist lists). Claude, GPT-5.6, and Grok produce close to zero on the same tasks. DeepSWE results also showed Ox Alpha averaging 117 agent steps per task, very close to GLM-5.3's leaderboard figure of 124 — versus GPT-5.6 Sol's typical 61 steps and Claude's 88–99.
Process of Elimination
Xiaomi MiMo V2.5's signature feature is native audio input, but Ox Alpha's refusal pattern on audio requests matched GLM-5V exactly. DeepSeek normally goes straight to open-sourcing weights. Qwen 3.8 was just publicly released with a non-matching encoder signature. xAI, OpenAI, Anthropic, and Gemini all failed on tokenizer, style, and encoder checks.
Synthesizing all of the above, Ben Davis puts the probability that Ox Alpha belongs to Zhipu's GLM-5.x family at roughly 99% — likely the multimodal flagship GLM-5.3V (vision version) or the next-generation GLM-5.5. A separate independent analysis pegs the confidence at around 90%.
Anonymous Channels: A Standard Playbook for Chinese Models Going Global?
Ox Alpha is the fifth "Stealth Model" on OpenRouter. The previous four that eventually revealed their identities all turned out to come from Chinese teams:
- Pony Alpha (February): Hit 40 billion tokens on day one; confirmed as Zhipu's GLM-5 five days later.
- Hunter Alpha (March 11): Trillion-parameter scale with 1M context; widely suspected to be DeepSeek V4 before Xiaomi confirmed it as MiMo-V2-Pro.
- Elephant Alpha (April): Efficiency-focused text model; identified about two weeks later by Ant Group's Inclusion AI.
- Owl Alpha (late April): Agent and tool-calling focus; not confirmed until June 30 — Meitua's LongCat-2.0.
Following this trail, Ox Alpha's naming now reads as a small coincidence: last time the alias was a horse, this time an ox — taxonomically reasonable, and the Chinese idiom "牛马" (literally "ox and horse," slang for grinding drudgery) makes the symbolism land a little too neatly.
Why the "Masquerade"?
The first benefit of anonymous release is decoupling brand from benchmark performance — testers judge the model on its actual behavior rather than who built it. For Chinese vendors trying to break into overseas developer markets, this lowers the brand barrier at launch and tends to attract more organic community testing and discussion. Confirming the identity later creates a two-wave PR effect: first the model itself gets attention, then the company behind it.
The second, more practical benefit is real-world feedback. Developers wire the model into coding tools, agent frameworks, and codebases — tasks far messier than public benchmarks. Questions like "is the long context actually stable," "is tool-calling reliable," and "does the model still hold its goal after dozens or hundreds of steps" only surface at real call volumes. OpenCode dedicating 100 trillion tokens per day to Ox Alpha signals a substantial test.
Following the cadence of past stealth models, Ox Alpha's free testing window may close around August 27. Existing technical analyses broadly point to Zhipu's GLM family, with more specific guesses including an unreleased multimodal flagship — but until the official reveal, that remains a strong, not a confirmed, attribution.