GLM-5.3 Flash (Ox Alpha) Launches: Opus 4.8 Performance at 1/40th the Price
A few days ago, an anonymous model called Ox Alpha launched on OpenRouter.
It consumed 23.2T tokens over the period — 2.6x the second-place historical record. On OpenCode, it broke DeepSeek's 56-day leaderboard streak.

Someone ran it on a subset of Deep-SWE benchmarks and got 63% — Pareto-optimal among open models, and trailing only Grok 4.6 among closed models, roughly on par with V4 Pro and Gemini 3.7 Flash.
Today the identity was officially revealed: GLM-5.3 Flash by Zhipu AI.

The price comparison is staggering.
The model was so capable that many overseas users refused to believe it was a Chinese model during anonymous testing.
- Artificial Analysis Intelligence Index score matching Opus 4.8
- 320B parameters, on par with DeepSeek V4 Flash
- Limited-time discount price at 1/40th Opus 4.8 — well below DeepSeek V4 Flash's post-price-hike公示价
- New architecture + new pretraining, native multimodal, 1M context
- MIT open source, runs on all domestic Chinese AI chips
I tested it on three frontend scenarios with real technical barriers, comparing against DeepSeek V4 Flash Vision EXP — and in some cases, Anthropic's latest Claude Fable 5.
Case 1: One Reference Image → Glass Cube Poster Recreation
The poster's visual core is two WebGL-rendered glass cubes floating above serif headlines.
The glass uses real physics simulation: transmission, refraction, dispersion — light passing through the cubes creates actual distortion and rainbow fringing on the text behind them.

Previously, getting a model to produce this required extremely detailed prompts: specifying Three.js MeshPhysicalMaterial, transmission: 1, ior, dispersion values, PMREMGenerator for environment mapping… you essentially had to already know how to do it yourself.
The prompt was just:
Recreate this webpage effect in Three.js WebGL. Especially the 3D glass in the center (with dispersion enabled) — it needs to rotate, and the text and background should have motion effects too.
One sentence. One reference image. Done.
What did its code actually do?
It used an offscreen Canvas to precisely render all typography — serif headlines, badge, Japanese kana, corner info — then mapped it to a plane behind the glass. So the glass's transmission refracts real text, not just a background color.
Glass material parameters were accurate: ior: 1.44, dispersion: 1.0 (subtle rainbow fringing without overdoing it), clearcoat: 1 for a varnish reflection layer. It even added a full-screen film grain shader and a slowly drifting blue mist layer, making the result look like a printed piece rather than a "3D render."

It also added interactive motion on its own: drag rotation with inertia decay, mouse parallax making the camera and text layer shift in opposite directions. None of this was in the reference — it invented it.
DeepSeek V4 Flash Vision EXP: disastrous. The glass used a hand-written ShaderMaterial with hard-coded refraction rays, excessive dispersion separation, looking like a demo not a product. Typography was crude with poor scale ratios.

Claude Fable 5: background color slightly more accurate, marginally better refraction detail — not meaningfully different. But roughly 100x the price.

Inferring a complete technical implementation stack — materials, rendering pipeline, interaction model — from a single static screenshot. This level of multimodal understanding was previously only seen at the most expensive tier.
Case 2: Watch a Screencast → Reconstruct the Entire Interactive Website
Case 1 tested image understanding. This tests video understanding. I gave it a screencast of a multi-page website.
The site's visual core is a 3D spiral ribbon made of 60 glass panes spanning the full page width. Supports drag-rotation with inertia, individual panes bulge on hover, and the ribbon smoothly transitions between three poses as you scroll — hero section diagonal, benefits shifted upper-right, footer shifted upper-left.
Prompt: Recreate this webpage effect in Three.js WebGL. Attached a video.
First round: 90% done.
The code structure was surprisingly clean: poseGroup for scroll-driven overall pose, nested yawGroup for drag rotation. The 60 glass panes used MeshPhysicalMaterial with transmission: 1, dispersion: 7 (pronounced rainbow fringing at glass edges), iridescence: 0.32 (thin-film interference).
It even built its own procedural HDR environment — glowing panels in cyan, magenta, and purple baked into PMREMGenerator to provide color-rich lighting for the glass refraction.

Glass quality largely depends on environment lighting — no rich environment lighting and no amount of parameter tuning makes glass look real. GLM-5.3 Flash thought of this itself.
Layout algorithm was also deliberate: panes arranged along the heading angle, adjacent wide faces facing each other, linked along narrow edges into a spiral. The whole ribbon's S-curve slowly breathes over time. One issue: the largest panes initially face the viewer, but the original faces each other with narrow edges toward the camera.

I added one note: "glass panes should face each other, narrow edges toward the camera" — tuning done, it was smoother than the original. Layout and scrolling were even more fluid.
DeepSeek V4 Flash Vision EXP: the "glass panes" had completely wrong materials — semi-transparent wireframe geometry with colored outlines, no refraction, no sense of depth. The core 3D portion was unadjustable.

"60 glass panes along an S-curve," "scroll-driven pose changes," "individual pane bulging on hover" — all of this had to be inferred from the video's dynamic frames. At this price point, it has no competition in video understanding.

Case 3: Static Poster → Parallax-Scrolling Interactive Webpage
This case changed direction. Instead of asking it to replicate, I gave it a rough concept and let it fill in the rest.
The reference was a Segmint 2023 poster: a voxel block拼成的圆盘 on a blue background, Ethereum diamond engraved on the face. Pixel-style font, blue-white pixel dithering transition at the bottom. Purely static.

Build a 3D WebGL + Three.js webpage based on this image. It's currently static — animate it:
- When the page loads, the center has no pattern
- As you scroll, elements in the center gradually fall away, until only the non-falling parts form the pattern
It didn't just paste the poster onto a webpage with animation.
The code's core is InstancedMesh voxel硬币 in two layers: body layer (blue face + dark sides + recessed grooves + embossed emblem) and a "cover layer" (the blocks hiding the emblem).

The Ethereum diamond wasn't a texture — it was drawn with vector paths on an offscreen Canvas, then 16x supersampling rasterized onto the voxel grid, ≥72% coverage voxels get full emboss, ≥18% get low emboss, with very fine anti-aliased edges.
As you scroll, the cover blocks fall in "edge-to-center" order, each with independent gravity acceleration, horizontal drift, and random rotation axes. At ~66% scroll, all cover blocks are gone, revealing the full ETH emblem. Lower-left corner has a REVEAL 000% pixel-font counter showing real-time progress.

It even replicated the original poster's blue-white pixel dithering transition: Canvas draws on a 12px grid, probability decaying linearly with height, from blue to paper-white.

No comparison with DeepSeek here — what it produced was a creative reinterpretation of the original visual concept. Given one flat image and a rough interaction concept, it delivered a complete, playable, narratively-paced webpage.
The Real Exciting Part: Price
Three cases done. Multimodal understanding, frontend coding, Agent capabilities — all at the top tier.
But the most exciting thing isn't capability. It's price.
The current AI industry's competition runs along two tracks.

One track pushes the intelligence ceiling — GPT-5.6, Fable 5 pushing frontier capability. Strong, but expensive.
The other is what actually determines reach: model pricing determines how many people can actually put AI into their daily workflows.
A conversation for cents, a complex task for ~$1 — most people won't embed AI into daily work if it doesn't pencil out. This is why DeepSeek V4 Flash was so viral — the industry needed a cheap-enough, capable-enough model.
DeepSeek V4 Flash is that斩杀线: your model has to beat it on price or capability, or there's no reason to exist.

GLM-5.3 Flash pushed that line significantly forward. 320B parameters, MIT open source, runs on every domestic chip.
I forget who said this, or which company's slogan it is: "Affordable intelligence for the many."
Zhipu took a big step toward that this time. People with ideas and creativity but limited budgets can finally access, at near-free pricing, the kind of intelligence that was previously only available to those with large engineering teams.
GLM-5.3 Flash is live on API and open source. Available on Z Code and Zhipu's Coding Plan.