Around 20 August, a model with no name and no price tag appeared on OpenRouter and OpenCode.
It went by Ox Alpha, and that was the entire description. Free on both platforms, a million tokens of context, and nothing anywhere saying who had built it or where it ran.
Then people hammered it.
OpenCode posted the count on the morning of the reveal: 42 trillion tokens in six days, the most-used model on the platform since DeepSeek Flash's 56-day run. Sit with that ratio. Six days against fifty-six.
OpenRouter's rolling-week leaderboard put Ox Alpha first by token volume, at one point running at more than double DeepSeek's usage. Three days in, a contemporaneous report recorded OpenCode's live figures at roughly 221,000 unique users and a little over five million sessions.
That's the number I'd hold onto. Trillions of tokens are hard to picture. 221,000 people is a mid-sized city. Some of them will have sent toy prompts. Plenty won't have. (Free models have a way of ending up in very real workflows.) Code, repositories, support tickets and whatever else happened to be open were going into an endpoint whose operator had declined to give a name.
Nobody was tricked. OpenRouter's listing said so on the page: "Ox Alpha is a stealth model. It is developed and operated by a third-party provider who has chosen to remain anonymous during this preview." Stealth previews under a pseudonym have been standard practice for years. If you used it, you knew you didn't know what it was.
On 26 August it got a name: GLM-5.3-Flash, from Z.ai in Beijing. The model has 320 billion parameters with 18 billion active, an MIT licence and weights that landed on Hugging Face the same day. The launch post also supplied one detail nobody using the preview had been able to check: the traffic had been served on Chinese-made AI chips.
Two days ago we wrote about which model answers when you ask, and why that mattered during a live security incident.

Hugging Face got attacked. The US models refused to help. An open one they could run themselves did the job.
Hugging Face named the models that wouldn't look at its attack logs: Claude Opus and Fable. A safety guardrail can't tell an incident responder from...
Read full articleA hosted API can hide four separate things: which model you're talking to, who built it, what silicon it runs on and where your data actually goes. Ox Alpha hid the first three by design. The fourth was buried under two gateways giving different answers.
That's what bothers me here. Not that the chips were Chinese. The impressive part is that they worked at this scale. The problem is that a developer trying to do the responsible thing couldn't get a straight, consistent account of the endpoint they were using.
The detail the model reviews missed
Almost everything written about GLM-5.3-Flash this week is the model review, and fair enough. Hybrid sparse plus linear attention, a KV cache four and a bit times smaller than GLM-5.3's, a million-token context window, $0.15 per million input tokens and community NVFP4 weights inside a day. Write that forty times and it's still true.
I think the hardware line is the interesting one, and it arrived only after the experiment was over.
The two disclosures aren't the same thing. The anonymity was declared on the page from the start. The Chinese-made hardware wasn't disclosed until the preview was over. Z.ai still hasn't said where the serving cluster was physically located, so the chip claim doesn't settle the data-residency question. It makes the lack of a clear answer harder to ignore.
Two gateways, two answers
During the preview, if you'd gone looking for what happened to your prompts, the answer you got depended on which door you walked through.
OpenRouter's listing said, and its preview-era text still says: "Prompts and completions are retained by the provider and are not used for training; all other use is governed by the Stealth Model Terms." The banner added after the unmasking keeps the same substance in past tense: prompts and completions "were retained by the provider".
OpenCode's documentation said something different. I pulled the Internet Archive's capture of the OpenCode Zen docs from 25 August, day five or six of the preview, so this is the page as it stood at the time. Under Privacy: "All our models are hosted in the US. Our providers follow a zero-retention policy and do not use your data for model training", followed by a list of exceptions. Ox Alpha isn't among the exceptions. It gets its own line higher up: "Ox Alpha Free is a stealth model that's free on OpenCode for a limited time. Its provider follows a zero-retention policy and does not use your data for model training."
Line them up. All models hosted in the US. This provider follows zero retention. Prompts and completions retained by the provider. Same model, same week. Those statements don't give a customer one coherent account of what happened to a prompt.
OpenCode knows how to write a carve-out. Big Pickle, the other stealth model on the same page, has one. Ox Alpha didn't.
The nationality of the hardware doesn't change that contradiction. A customer who read both pages and tried to reconcile them couldn't have reached a dependable answer. Diligence is supposed to be the thing that saves you. (Or at least gives you something defensible to write in the risk register.)
One more line, undecorated. OpenRouter's listing described Ox Alpha as designed for production workloads and suited to long-horizon software engineering while the provider was anonymous and, on the same page, retaining prompts. I can't tell whether OpenRouter or the provider wrote that copy. Either way, describing an unnamed preview as designed for production workloads is quite a sentence.
The hardware reveal is impressive. It's also easy to overread
Z.ai's own researchers were posting within minutes of the announcement. Zixuan Li, three minutes after the company account:
Z.ai says the serving ran across "tens of thousands of domestically developed accelerators", using a custom SGLang inference engine, encode-prefill-decode disaggregation, W8A8 weights and a hybrid INT8/FP8/BF16 cache. Whatever you make of the marketing, 42 trillion tokens in six days isn't a slide deck. It's a serious deployment carrying developers' real workloads, and it ran for most of a week without the hardware becoming the story. One detail from the engineering notes made me smile: a GLM-5.3-powered agent helped tune the kernels that served the newer model. Clever. Still not the story.
Four places where the claim is softer than it reads.
They won't name the vendor. Z.ai's materials say "Chinese AI chips" and "domestically developed accelerators" and stop there. Every write-up saying Huawei Ascend is inferring it, mostly by carrying over reporting about how GLM-5 was trained, a different model answering a different question. Huawei, Cambricon, Moore Threads, Kunlunxin: none is on the record for this cluster. A victory lap that won't name the supplier is worth a sentence of your attention.
The 3x is a self-comparison. "3x improvement in end-to-end serving performance" is measured against Z.ai's own earlier baseline on the same silicon. It's a real engineering result and it says precisely nothing about NVIDIA.
The parity claim is unreproduced. Z.ai claims "per-token cost comparable to mainstream NVIDIA GPUs" without saying which NVIDIA GPU, no independent cost model exists, and nobody credible has published counter-measurements either. It's a vendor assertion about a vendor's own cluster, and it's about 36 hours old.
The independent throughput number is less triumphant. Artificial Analysis measured roughly 48.7 tokens per second on Z.ai's API and called that "notably slow", against a 67.0 median for comparable open-weight models. At the time of checking, the model sat 46th of 108 for speed and third for intelligence. That's probably the fairest one-line review available: very capable, not especially quick.
SemiAnalysis called the deployment "shocking", which is worth noting from a semiconductor analysis outfit. Its post referred to 100 trillion tokens per day, but that's OpenCode's advertised capacity, not measured traffic. The measured figure is 42 trillion over six days. Those numbers are already dramatic enough. They don't need help.
The model looks good on Z.ai's benchmarks and it is aggressively priced. Both facts matter if you're choosing a model. Neither answers who received your prompts, whether they were retained or where the serving cluster sat. That was the rabbit hole in the earlier draft of this article: I started reviewing the model and nearly buried the reason I cared about Ox Alpha in the first place. (The benchmark table alone ran for 20 lines.)
What to actually do about it
If you're a small business on one AI vendor: nothing changes today. You've got a ChatGPT subscription or a Claude account, you aren't routing through an aggregator, and none of this reaches you. Get on with your day.
If you route through an aggregator, ask four questions instead of one. Which model? Who operates the endpoint? Where is the data processed? What is retained? Most teams ask the first and assume the rest follow, and this week is a decent argument that they don't. Check whether your router falls back automatically to providers you haven't vetted, and whether stealth or preview models are on by default. OpenCode lets admins disable specific models, which exists so somebody can make this call deliberately instead of by accident.
If any of your prompts carry personal information, this can become a legal question rather than a preference. Under Australian Privacy Principle 8, an APP entity that discloses personal information to an overseas recipient generally has to take reasonable steps to ensure the recipient doesn't breach the APPs. Under section 16C of the Privacy Act, the discloser can remain accountable for the recipient's conduct. Exceptions exist, including substantially similar protection with an enforceable mechanism, or informed consent after the individual has been warned that APP 8.1 won't apply.
Now read that against what happened. The Chinese-made chips don't prove that data crossed a border or that the servers sat in China. What they do prove is how little the customer could establish. The provider's identity was withheld, its physical serving location wasn't disclosed, and the two gateways disagreed about retention. If you needed to decide whether APP 8 applied, identify the recipient or document its safeguards, the public information wasn't enough to do the job confidently.

Privacy Act 2025: How AI Website Analytics Affect Australian Business Compliance
The Privacy Act 2025 overhaul slashes penalties, adds statutory torts, and forces new APP disclosures on AI analytics stacks before the OAIC demands...
Read full articleThere's an exit, and a real one. The weights are on Hugging Face under MIT, so if you'd rather not guess where the silicon lives, you can put the model on hardware you've picked yourself. It's also out of reach for most readers of this site, which is why it isn't the headline advice.
If you handle health, financial or government-adjacent data, the boring version of this control is an allow-list of approved endpoints with a named provider and a named jurisdiction against each, plus a rule that unlabelled previews stay off it. (Boring is a compliment here.)
Key takeaways
- Ox Alpha was GLM-5.3-Flash: free on OpenRouter and OpenCode from the week of 20 August until Z.ai unmasked it on 26 August, with 42 trillion tokens through OpenCode in six days. The anonymity was declared. The hardware wasn't, and there's no evidence either platform knew.
- Two gateways published contradictory answers to "what happens to my prompts". OpenRouter said prompts were retained by the provider. OpenCode's docs, captured mid-preview, said all models were US-hosted and this provider followed zero retention. OpenCode wrote a carve-out for its other stealth model and not for this one.
- The chip demonstration is real, but it doesn't prove the cluster was physically in China. The parity claim is also unverified: the 3x figure is a self-comparison, the NVIDIA comparison names no GPU, and Artificial Analysis called its measured throughput "notably slow".
- Z.ai still won't name the chip vendor, so anything asserting Huawei Ascend for this cluster is inference rather than a confirmed detail.
- For Australian entities handling personal information, an endpoint whose recipient, location and safeguards can't be established creates a compliance problem before it creates a technical one.
There's another one on the shelf right now
None of this is retrospective. Scroll the same OpenCode model list today and you'll find Big Pickle: free, 200,000 tokens of context, no named underlying model provider and no stated hardware. OpenCode says its Zen models are hosted in the US, but developers trying to identify the model itself have resorted to reading leaked error messages and API response signatures. That's a fair summary of how much anyone is being told.
It's worth comparing the two pages, because the difference is instructive. Big Pickle appears in OpenCode's zero-retention exceptions, with a line of its own: "During its free period, collected data may be used to improve the model." Plain enough. Your prompts may become training data, and you've been told.
Ox Alpha never got that sentence. It got the opposite one, the assurance that its provider followed zero retention, while OpenRouter's page said those prompts were being kept. The model that needed the carve-out didn't have it. The one sitting next to it did.
So the practical question isn't what happened between 20 and 26 August. It's what's switched on in your own setup this morning. Which endpoints can your tooling reach? Which can you name a company and a country for? Is there anything in that list you'd struggle to explain to a client whose data went through it?
I don't think the answer is to stop using Chinese models, or even stealth previews. It's to stop treating a model name as if it answers every question about the service behind it. Most of us haven't been asking those questions consistently. I'm including myself in that.
The check takes ten minutes. The awkward client conversation takes longer.
---
Sources
- Z.ai. "GLM-5.3-Flash." 26/08/2026. https://z.ai/blog/glm-5.3-flash
- Z.ai. GLM-5.3-Flash model weights, MIT licence. 26/08/2026. https://huggingface.co/zai-org/GLM-5.3-Flash
- OpenRouter. "Ox Alpha (stealth)." Model listing, accessed 27/08/2026. https://openrouter.ai/stealth/ox-alpha
- OpenCode. "Zen." Documentation, Internet Archive capture of 25/08/2026. https://web.archive.org/web/20260825104759/http...
- OpenCode. "Zen." Documentation, current version, last updated 26/08/2026. https://opencode.ai/docs/zen/
- Artificial Analysis. "GLM-5.3-Flash." Accessed 27/08/2026. https://artificialanalysis.ai/models/glm-5-3-flash
- OAIC. "Chapter 8: APP 8, cross-border disclosure of personal information." Australian Privacy Principles Guidelines. https://www.oaic.gov.au/privacy/australian-priv...
- OAIC. "Read the Australian Privacy Principles." https://www.oaic.gov.au/privacy/australian-priv...
- OpenRouter. "Gemini 3.7 Flash." Model pricing, accessed 27/08/2026. https://openrouter.ai/google/gemini-3.7-flash
- OpenRouter. "GPT-5.6 Luna." Model pricing, accessed 27/08/2026. https://openrouter.ai/openai/gpt-5.6-luna
- OpenRouter. "DeepSeek V4 Flash." Model pricing, accessed 27/08/2026. https://openrouter.ai/deepseek/deepseek-v4-flash
- Chaffee, Phillip. "big-pickle SWE Atlas results." Notes the undisclosed provider and that the model behind the alias may change without notice. https://github.com/PhillipChaffee/big-pickle-sw...
- The Slide Factory. "Ox Alpha Model: The Mystery AI Taking Over OpenRouter." Contemporaneous report of OpenCode's live usage snapshot, accessed 27/08/2026. https://www.theslidefactory.com/post/ox-alpha-m...



