Mystery Ox Alpha Model Drops Free With 1M Token Context
Quick summary
Anonymous stealth model on OpenRouter and OpenCode. 1M context, coding agents already burning trillions of tokens. Nobody claimed it.
Read next
- DeepSeek R2 Is Out: What Every Developer Needs to Know Right NowDeepSeek R2 just dropped. It is multimodal, covers 100+ languages, and was trained on Nvidia Blackwell chips despite US export controls. Here is what changed from R1, what the benchmarks mean, and how to use it including running it locally.
- NVIDIA GTC 2026: Jensen Huang Keynote Preview for DevelopersNVIDIA GTC 2026 runs March 16-19 in San Jose. Jensen Huang teases a surprise. Vera Rubin chips, Feynman architecture, and what changes for developer AI costs.
Advertisement
An anonymous coding model called Ox Alpha (often typed as 0x Model because Ox reads like a hex prefix) appeared on OpenRouter and OpenCode on August 20, 2026, with a 1,048,576-token context window, text plus image plus video input, and a price of $0 for roughly one week. No lab put its name on the listing. OpenRouter routes it under the provider label Stealth and the model id stealth/ox-alpha, and states plainly that OpenRouter is not the developer, owner, or provider.
Within days, agent traffic was already in the hundreds of billions to trillions of tokens. That is not curiosity browsing. That is developers wiring an unnamed endpoint into real coding loops because the free window and the 1M context are too useful to ignore. Here is what is confirmed, what is community fingerprinting, and what you should do before you point it at a private repo.
What Ox Alpha Actually Is
Ox Alpha is a stealth preview: a third-party lab serving a reasoning model through public gateways without a brand on the card. OpenRouter describes it as a reasoning model designed for coding, sustained agentic work, and production workloads, suited to long-horizon software engineering and workflows that combine text with visual context.
If you saw headlines about a mysterious "0x model" that suddenly topped coding chats, that is this release. The marketing name is Ox Alpha. The API slug on OpenRouter is stealth/ox-alpha. On OpenCode Zen the free id has been reported as opencode/x-preview-f-free; on OpenCode Go as opencode-go/ox-alpha-free. Same preview window, different host ids.
Confirmed Specs vs Marketing Claims
The numbers below come from OpenRouter's model page and OpenCode's free-week announcement. Treat capacity claims as operator statements, not measured SLAs.
| Spec | Confirmed value | Source type |
|---|---|---|
| Context window | 1,048,576 tokens (1M) | OpenRouter listing |
| Max output | 131,072 tokens | OpenRouter listing |
| Inputs | Text, image, video | OpenRouter / OpenCode |
| Output | Text | OpenRouter listing |
| Tool calling | Yes | OpenRouter listing |
| Reasoning variants | low / high / max (OpenCode path) | OpenCode catalog reports |
| Price this window | $0 / $0 (input / output) | OpenRouter + OpenCode |
| Free window | About one week from Aug 20 (OpenCode later said ~6 days on Go) | Operator posts |
| Claimed serving capacity | 100 trillion tokens per day | OpenCode announcement |
| OpenRouter P50 (early window) | ~5s latency, ~24 tok/s, ~99.5%+ availability | OpenRouter live stats |
The 100T tokens/day figure is a stress-test dare, not a contract. Congestion reports of ~20 tok/s and dropped connections showed up within the first day. Free plus viral equals queue.
Why Labs Ship Anonymous Models
Anonymous drops are now a playbook, not a one-off prank. The last several stealth models on the same circuit resolved into named Chinese lab products after a burst of free traffic:
| Stealth name | Approx. debut | Later claimed as |
|---|---|---|
| Pony Alpha | Feb 2026 | Zhipu AI GLM-5 |
| Hunter Alpha | Mar 2026 | Xiaomi MiMo-V2-Pro |
| Elephant Alpha | Apr 2026 | Ant Group Lingxi Ling-2.6-flash |
| Owl Alpha | Apr 2026 | Meituan LongCat-2.0 |
| Ox Alpha | Aug 20, 2026 | Unclaimed as of Aug 23 |
The incentives are boring and effective. Anonymity removes brand bias from early evaluations. Free usage crowdsources frontier-scale feedback. Mystery itself is distribution: Reddit, HN, and X do the launch marketing for free. For a China-based lab, a global OpenRouter mirror also reaches developer audiences that never visit a domestic launch page. That matters for China's model race as much as any press release.
Who Made Ox Alpha? Fingerprints, Not Press Releases
Nobody has claimed Ox Alpha as of August 23, 2026. Community forensics still matter because tokenizer and video-encoder behavior are hard to fake without retraining.
Independent probes circulating this week point hardest at Zhipu AI (Z.ai) GLM-5.3, often framed as a Flash or unreleased multimodal sibling:
- Tokenizer fingerprint tools reportedly matched GLM-5.3 on multiple probes; one researcher described exact token-count matches across 25 prompts with a constant ~75-token wrapper offset.
- Video token consumption on controlled clips reportedly matched GLM-5V-Turbo sampling behavior while MiMo, Qwen, and older GLM vision paths diverged.
- Shared refusal patterns (no audio), error-code dialect, and output style cues have been cited as secondary tells.
- Precedent: Zhipu already used the Pony Alpha → GLM-5 anonymous path earlier in 2026.
GLM-5.3 itself launched around August 14 as a text-oriented release. Ox Alpha advertises video. If the fingerprint holds, the clean reading is an unreleased multimodal GLM-5.3-line build, not a simple rename of the public text model. Xiaomi MiMo and other suspects remain in circulation, but tokenizer-plus-video evidence currently favors Zhipu. Treat every identity claim as inference until Z.ai or another lab posts a claim.
Are the Coding Benchmarks Real?
Early social posts said Ox Alpha crushed DeepSWE with an ~80% pass rate on a small task set, beating Claude Fable 5 (~65%) and GPT-5.6 Sol (~52%) in that subset. That number traveled farther than the methodology.
Independent researcher Ben Davis later reported a full 113-task DeepSWE run near ~63%, roughly on par with GPT-5.6 Sol mid. The 80% figure came from a 10-task subset (under 9% of the suite). Both numbers are community measurements, not an official Artificial Analysis card and not a vendor SWE-bench Verified entry. Use them as a signal that the model is competitive on agentic software tasks, not as a procurement score.
Usage is the other benchmark. OpenRouter's public apps chart for this model already showed heavy production harness traffic in the early window, including Hermes Agent (Nous Research) in the trillion-token range and Claude Code in the hundreds of billions of tokens. OpenCode's public usage page, within about a day, showed on the order of 1.8T tokens, tens of thousands of unique users, and multi-million-token average sessions at $0 spend. That is how you know developers are treating it as a workhorse, not a toy.
The Privacy Split Developers Keep Missing
This is the part that decides whether you paste a real repo.
On OpenCode, the Zen privacy docs and launch copy claim the Ox Alpha Free provider follows a zero-retention policy and does not train on your data.
On OpenRouter, the Ox Alpha page says prompts and completions are retained by the provider and are not used for training, under Stealth Model Terms. That is retain-not-train, not ZDR. OpenRouter's default stealth EULA historically included training-use language for free access; Ox Alpha's model-specific notes walk training back without walking retention back.
Same model family (probably). Two hosts. Two data stories. If retention is why you care, prefer the OpenCode-native free ids for the week. If you already live inside an OpenRouter harness, assume an unnamed party keeps logs. Either way, do not send customer data, secrets, or regulated code through a stealth preview because the token price is zero. For MCP and agent-tooling risk context, see Anaconda's Enkrypt acquisition and MCP vulnerability scan numbers.
Our Analysis: How Developers Should Use the Free Week
Ox Alpha is a time-boxed audition, not a new default. The free week ends around August 26–27, 2026 depending on which OpenCode post you trust. After that, price, rate limits, and even availability can change without a brand to complain to.
Here is the practical stack I would run this week:
- Harness first, model second. Keep Cursor, Claude Code, OpenCode, or your agent loop stable. Swap only the model id. Portable skill packaging from Agent Plugins 1.0 matters more than any one-week endpoint.
- Fallback chain. Route Ox Alpha → named GLM-5.3 or DeepSeek V4 Flash → your paid Claude/GPT default. When the stealth host hiccups, you still ship.
- Repo hygiene. Open or personal repos only. Strip secrets. Prefer synthetic tasks that still stress 1M context and tool calling.
- Measure your own tasks. Ignore the viral 80% DeepSWE subset. Track pass rate on your flaky tests, migrations, and multi-file refactors.
- Price the hangover. Put today's $0 next to real rates on the LLM API Pricing Tracker. A free preview that becomes $1–$4 per million tokens changes the FinOps math overnight, which is the same class of surprise that turned up in the $500M Claude bill story.
Why this story hits China and Bing traffic hard
Stealth Chinese lab previews are exactly the content mix that already performs on abhs.in: named model ids, specific token numbers, and a developer decision checklist. Chinese search and cn.bing users hunt for early access to GLM-class coding models. Western developers hunt for "is this better than Claude for agents?" Both intents share one post if you keep the identity claim careful and the ops advice concrete.
For broader model picking after the free week, use Best AI Models 2026 and the Claude vs ChatGPT tool rather than locking your stack to an anonymous slug.
What To Watch Next
Four signals settle the story:
- Claim: Does Z.ai (or Xiaomi, or someone else) put a name on Ox Alpha within days, the way prior Alpha drops resolved?
- Price: What replaces $0, and does any free tier survive?
- Authoritative bench: Artificial Analysis, LMArena, or a reproduced SWE-bench Verified number with a named model card.
- Western response: Price cuts, free coding windows, or fast-follow agent models from OpenAI, Anthropic, or Google within a week of the reveal.
Until those land, Ox Alpha is the strongest free coding agent endpoint money can currently not buy, with the privacy and disappearance risks that phrase implies.
Key Takeaways
- Ox Alpha (aka 0x model) launched August 20, 2026 on OpenRouter as
stealth/ox-alphaand on OpenCode as a free stealth coding preview - 1,048,576-token context, 131K max output, text/image/video in, tool calling, $0 during the ~one-week window
- No lab has claimed it as of August 23; community fingerprinting strongly suggests a Zhipu GLM-5.3-line multimodal build, still unconfirmed
- Early 10-task DeepSWE ~80% viral number; full 113-task community run ~63%, roughly GPT-5.6 Sol mid territory
- OpenRouter apps already show Hermes Agent and Claude Code burning hundreds of billions to trillions of tokens through the endpoint
- Privacy split: OpenCode claims zero retention; OpenRouter says provider retains prompts/completions (no training)
- For developers: audition this week with a named-model fallback; never paste secrets into an anonymous host
- What to watch: lab claim, post-free price, official leaderboard entry, and Western lab counter-moves by end of August
Sources
- OpenRouter model page: https://openrouter.ai/stealth/ox-alpha
- OpenRouter Stealth Model Terms (linked from the Ox Alpha listing)
- OpenCode free-week announcement and Zen/Go model catalog reports (Aug 20–21, 2026)
- Community fingerprinting and DeepSWE notes summarized via independent writeups citing researcher Ben Davis and modelprint-style probes (Aug 2026)
- Prior stealth → named reveals: Zhipu GLM-5 (Pony Alpha), Xiaomi MiMo-V2-Pro (Hunter Alpha), Ant Ling-2.6-flash (Elephant Alpha), Meituan LongCat-2.0 (Owl Alpha)
FAQ
Frequently Asked Questions
What is the Ox Alpha or 0x AI model?
Ox Alpha is an anonymous stealth reasoning model that appeared on OpenRouter and OpenCode on August 20, 2026, aimed at coding and long-horizon agent work. People often type it as 0x Model because Ox looks like a hex prefix. Confirmed specs include a 1,048,576-token context window, multimodal text/image/video input, tool calling, and free pricing during a roughly one-week preview.
Who made Ox Alpha?
No company has officially claimed Ox Alpha as of August 23, 2026. Community tokenizer and video-encoder fingerprinting currently points strongest at Zhipu AI (Z.ai) GLM-5.3, likely an unreleased multimodal or Flash-line sibling, but that remains an inference rather than a lab announcement. Earlier anonymous Alpha models on the same circuit were later claimed by Zhipu, Xiaomi, Ant Group, and Meituan.
Is Ox Alpha free and how do I use it?
Yes for a limited window. OpenRouter lists stealth/ox-alpha at $0 input and output; OpenCode offered Ox Alpha Free for about a week from August 20, 2026 (Go follow-ups shortened that to roughly six days). Use OpenCode Zen/Go model pickers for the zero-retention claim path, or call stealth/ox-alpha through any OpenRouter-compatible SDK if you already have keys.
Is Ox Alpha better than Claude or GPT for coding?
It is competitive on early community DeepSWE runs, with a full 113-task result near ~63% (roughly GPT-5.6 Sol mid), while a viral 10-task subset near 80% overstated the gap. There is no official Artificial Analysis card yet. Treat it as a strong free audition for agentic coding, not a proven replacement for Claude or GPT in production.
Is it safe to send my private code to Ox Alpha?
Only if you accept an unnamed third-party operator. OpenCode claims zero retention for the free Ox Alpha path; OpenRouter says the provider retains prompts and completions but does not train on them. Do not send secrets, customer data, or regulated repos through either path. Use a fallback to a named model for anything that must survive after the free week ends.
Advertisement
Free Weekly Briefing
The AI & Dev Briefing
One honest email a week — what actually matters in AI and software engineering. No noise, no sponsored content. Read by developers across 30+ countries.
No spam. Unsubscribe anytime.
More on AI
All posts →DeepSeek R2 Is Out: What Every Developer Needs to Know Right Now
DeepSeek R2 just dropped. It is multimodal, covers 100+ languages, and was trained on Nvidia Blackwell chips despite US export controls. Here is what changed from R1, what the benchmarks mean, and how to use it including running it locally.
NVIDIA GTC 2026: Jensen Huang Keynote Preview for Developers
NVIDIA GTC 2026 runs March 16-19 in San Jose. Jensen Huang teases a surprise. Vera Rubin chips, Feynman architecture, and what changes for developer AI costs.
AI Agents Will Own More Crypto Wallets Than Humans: Coinbase x402 Is Already Live
Brian Armstrong says AI agents will soon make more transactions than humans. They cannot open a bank account — but they can own a crypto wallet. Coinbase already launched x402 Agentic Wallets with 50 million transactions processed.
From AI Act to AI Factories: How Europe Is Building a Regulated AI Super-Infrastructure
The EU AI Act's enforcement timeline is active and the EU is simultaneously building AI Factories — national compute clusters for European AI development. Here is what the dual strategy means for developers, enterprises, and the global AI infrastructure landscape.
Free Tool
Will AI replace your job?
4 questions. Get a personalised developer risk score based on your stack, role, and what you actually build day to day.
Check Your AI Risk Score →Written by
Software Engineer based in Delhi, India. Writes about AI models, semiconductor supply chains, and tech geopolitics — covering the intersection of infrastructure and global events. 1025+ posts cited by ChatGPT, Perplexity, and Gemini. Read in 167 countries.
