The anonymous model everyone loved turned out to be a Flash model

The anonymous model everyone loved turned out to be a Flash model

August 28, 2026


A free model with no name, with numbers that did not add up

A model appeared with no lab attached to it, released anonymously as Ox Alpha through a couple of coding tools and OpenRouter. Free to use, no explanation. That alone is not unusual. What was unusual was the serving capacity that came with it: a claimed ceiling in the range of a hundred trillion tokens per day, handed to a small number of providers at once.

For scale, a very heavy individual user might burn a few billion tokens on their worst day. Estimates commonly put an entire major consumer model's daily traffic somewhere in the low hundreds of trillions. So either somebody was sitting on a hidden reserve of compute, or the model was small enough and cheap enough that giving it away barely registered on their bill.

It was the second one. Ox Alpha is GLM 5.3 Flash, and it has quickly become one of the models we reach for most.

What actually shipped

  • 320B total parameters, roughly 18B active per token, a mixture-of-experts layout that keeps inference cost near a much smaller dense model.

  • A million token context, with a hybrid sparse and linear attention design that keeps long-context serving cheap. Notably, the price does not step up once you cross a context threshold. It is flat.

  • Full multimodality: images, audio and video, trained on a new multimodal corpus. The non-Flash sibling in the same family could not even take a screenshot.

  • Open weights, published on Hugging Face, so it can be self-hosted rather than rented.

  • Pricing at launch, with discounts applied, landing around 7.5 cents per million tokens in and 25 cents per million out.

One more detail that deserves attention: it is reportedly being served on Huawei Ascend accelerators rather than Nvidia silicon. That claim is secondhand and worth treating as such, but if it holds, the economics here are not just a training-run story.

The early benchmark numbers were misleading, and it does not matter much

During the anonymous window, people benched it against the public subset of a well known agentic coding evaluation and got a startling result: a score in the eighties where frontier models were landing in the fifties and sixties. Those figures are not what they look like. Public subsets leak, and a small slice of problems is not a ranking.

The vendor's own published comparisons are also self-reported and should be read with the usual skepticism. The honest summary is narrower and still remarkable: this thing performs in the neighborhood of a last-generation frontier model on a wide spread of tasks, at roughly a tenth of the price or less.

Intelligence and behavior are two different axes

This is the framing we keep coming back to, and this model is the cleanest example of it we have seen.

Intelligence is knowledge and reasoning depth: how much a model knows, how well it solves a genuinely hard problem. Behavior is how faithfully it applies what it is told: does it stay on task, does it absorb a mid-flight instruction change, does it unblock itself when a tool fails, does it do the small thing you asked instead of a bigger thing you did not.

Historically these have moved together, so the distinction rarely mattered. There are counterexamples in both directions. Very large non-reasoning models have carried enormous knowledge and been almost useless as agents. Some highly capable reasoning models score near-perfectly on obscure knowledge tests and then, given a repository and told to fix something, rewrite half of it because they lost the plot.

GLM 5.3 Flash is the first genuinely cheap model we have used that is a pleasure to work with. It is not the smartest thing available. It just does what you told it.

That behavior does not come free. It comes from heavy reinforcement learning in provisioned environments where the model does real work and learns to work better, which is exactly the training investment that a lot of open-weight releases skip.

What it looks like on a real agentic task

The test we care about is not a benchmark, it is a wide sweep over messy real state. Point it at a repository with over a thousand open pull requests and ask it to rank them by ease of merge and user value, then change the instructions halfway through, then ask for the output as an HTML page.

Three things stood out. It held the mid-task steer instead of restarting or acknowledging the new instruction and stopping. When its subagent delegation failed repeatedly with provider auth errors, it noticed the pattern and fell back to auditing directly rather than stalling. And it produced links to the pull requests rather than bare numbers, which is a small thing you stop noticing until a model that costs fifty times more does not do it.

The sweep took about twenty minutes and cost roughly twelve cents, most of the input being cache hits. The same job on a frontier model has run into three figures.

Where it is weak

  • Token efficiency. Around 47k tokens per task on public measurements, against roughly 20k and 17k for leaner competitors. Price hides this, but context does not. More tokens to reach the same answer means a fuller context window and a lower chance of finishing before you run out of room.

  • Hard reasoning. It makes logic errors that a frontier model would not. Given a nontrivial game loop it will get the animation ordering right and the hunger state wrong.

  • Serving variance. Throughput on the cheapest providers degraded sharply once demand arrived. Open weights mean you can route elsewhere, or host it yourself.

How we think about this at Birdhouse

We do not believe in one model for everything. We route. The interesting consequence of a model like this is that a whole class of work moves from "too expensive to bother" to "run it on a schedule and see what it finds": repository-wide audits, dependency sweeps, triage passes over stale issues, first-pass review that a person then adjudicates. When a thorough sweep costs cents, you stop rationing it.

The open weights matter as much as the price. A model you can run inside your own boundary is a model you can point at client code without a data-sharing conversation, and it fits the way we prefer to build: least-privilege sandboxes, agents that can be cheaply re-run when they fail, and a person reviewing what comes out. Cheap and well behaved beats expensive and clever for the wide, boring, high-volume half of engineering work. Save the frontier budget for the problems that are actually hard.

The models worth paying attention to are not always the ones at the top of the leaderboard. Sometimes they are the ones that simply do what you asked.