Space Bunny’s 13.9T-Token Stress Test: Usage Is Not Quality
The most revealing model on this week’s OpenRouter chart is one whose maker is not publicly identified.
Space Bunny Alpha, listed simply as stealth/space-bunny-alpha, has climbed to third place with 13.9 trillion tokens of trailing seven-day usage. It was fourth in the previous daily snapshot at 9.87T, meaning roughly 4.03T tokens were added to the rolling total while it passed Tencent’s Hy4 preview. The lazy interpretation is that a mysterious new model has become the world’s third-best LLM. The useful interpretation is narrower and more interesting: free access, API compatibility and curiosity can now move production-scale traffic before buyers know who built the product.
OpenRouter’s top two remain DeepSeek V4.1 Flash at 19.6T tokens, up 24% week over week, and Z.ai’s GLM 5.3 Flash at 16.3T, up 16%. Space Bunny follows at 13.9T, while Hy4 preview has slipped to fourth with 9.64T and a 23% weekly decline. Those numbers measure prompt and completion tokens routed through OpenRouter. They do not measure completed tasks, accepted pull requests, satisfied users or even distinct applications.
That distinction matters because the leaderboard is being treated as a quality table when it is better understood as a live map of routing pressure. A free endpoint can accumulate extraordinary volume because developers are testing it, agents are verbose, retries are frequent or bulk jobs automatically choose the cheapest eligible model. None of those explanations makes the data worthless. They make it operational evidence rather than a medal ceremony.
Preference and deployment have stopped telling the same story
The contrast with Arena AI is unusually sharp this week. Arena’s Text leaderboard did not change from its September 25 snapshot. Its five highest-ranked models are all Anthropic variants: Claude Opus 5.5 High at 1509 Elo, Opus 4.6 High at 1505, Fable 5 High at 1504, Opus 4.7 High at 1502 and Fable 5.1 Max at 1501. Arena Code is similarly stable, led by Claude Opus 5.5 Max at 1827, GPT-6 Astra Max at 1792 and Claude Fable 5.1 Max at 1751.
OpenRouter usage points elsewhere. Its top five are DeepSeek, Z.ai, an anonymous model, Tencent and OpenAI’s GPT-5.6 Luna. Claude Sonnet 5 appears at number 20 with 1.46T tokens. That does not disprove Arena, and Arena does not invalidate OpenRouter. The two systems answer different questions.
Arena asks which response people prefer in a blind pairwise comparison. OpenRouter reveals which endpoints receive traffic under the economics and constraints of real routing. A model can win the first contest and lose the second because a team has a latency budget, a million-token nightly batch, a regional availability requirement or an agent loop that turns a small price difference into a large bill. “Best model” has become a category error unless the speaker also names the workload and constraint set.
There is another measurement caveat: OpenRouter counts both prompts and completions, and upstream providers use their own tokenizers. A token attributed to Anthropic is not guaranteed to represent the same amount of text as one attributed to Xiaomi or OpenAI. The chart is directionally useful, especially for movement within a model family, but false precision belongs in the request-changes column.
Space Bunny looks built for agent traffic, but the chart cannot prove it
The endpoint’s capability surface helps explain the rush. It supports reasoning controls, tool definitions and tool choice, structured response formats, temperature and top-p. In other words, it can fit into an OpenAI-compatible agent stack without forcing a team to redesign the integration. That dramatically lowers the cost of experimentation. A model does not need a trusted brand to win a trial if adopting it is one configuration change and the price is zero.
It is plausible that agent workloads are contributing to the spike, but OpenRouter does not publish a workload breakdown for this row. Treat that as a hypothesis, not a fact. Community discussion has followed the same experimental tone. In one r/SillyTavernAI thread, users asked whether it might be “the next Kimi”; elsewhere, developers speculated that it could be MiniMax M3.1. No public evidence confirms that identity. Curiosity is a demand generator, not provenance.
The launch also supplied its own warning label. OpenRouter said the model was temporarily taken offline while the provider addressed a stability issue. That outage should sit in an evaluation scorecard beside accuracy and latency, not be dismissed as launch-week noise. Anonymous free models are effectively production experiments with a leaderboard attached: the compatibility is real, but so are the unanswered questions about ownership, data retention, capacity and support.
This is the first original lesson in the ranking movement: API compatibility has become a distribution channel in its own right. Model vendors once needed a consumer app, a cloud partnership or a famous benchmark result to gain reach. Now a sufficiently compatible endpoint can inherit an installed base of routers and agent frameworks almost immediately. The moat is not merely model quality; it is how few lines a platform team must change to test you.
MiMo shows what an actual migration curve looks like
Xiaomi’s MiMo family offers a cleaner adoption signal. MiMo-V2.6-Flash remains eighth with 5.51T tokens and is still marked new. Xiaomi describes it as a full-modality, low-cost reasoning model for high-frequency professional workloads. Meanwhile MiMo-V2.5 fell three positions, from ninth to twelfth, while its weekly usage dropped 61%.
That paired movement looks less like random leaderboard churn and more like traffic migrating within a vendor family. The newer version does not need to knock the predecessor out of the table overnight; the simultaneous rise and decline show applications being shifted gradually. Platform teams should take the hint. Provider-level dashboards are too coarse. Track model versions independently, define explicit fallback chains and record which version produced each output. Otherwise a “Xiaomi is stable” aggregate can hide a rapid replacement underneath.
GPT-6 Luna made the largest rank move, rising four places to tenth on 2.86T tokens. DeepSeek V4 Flash 0423 rose one place to ninth even while its weekly usage declined 11%, which is a reminder that rank is relative: a model can climb because neighbors fell faster. Jev 1.13 held thirteenth with growth of 455%, an enormous rate that nevertheless cooled from the prior snapshot’s 894%. There were no entrants to or exits from the top 20. This is a redistribution story, not a changing-of-the-guard story.
The second lesson is that velocity needs a denominator and a time horizon. Percentage growth makes small or new baselines look dramatic, while a rolling seven-day window mixes fresh traffic with six days of history. Teams should not procure from a single snapshot. Watch whether usage persists after launch incentives end, incidents occur and the free tier changes. Durable routing share is much harder to manufacture than a debut spike.
How to turn a traffic chart into an engineering decision
Use this top 20 as a candidate generator, not a procurement sheet. Put Space Bunny Alpha, MiMo-V2.6-Flash and the current DeepSeek and GLM leaders through the same internal task set. Fix prompts, sampling parameters and tool definitions. Then log end-to-end cost, time to first token, p50 and p95 latency, schema-valid response rate, tool-call correctness, retries, tokens per accepted task and human acceptance.
For agent systems, measure recovery behavior rather than judging isolated answers. Deliberately return a tool error. Feed the model an invalid intermediate result. Test whether it repairs malformed JSON without inventing successful actions. A cheap model that retries three times or burns twice the tokens may be more expensive than the premium endpoint it replaces. Token price is an input; cost per successful task is the metric.
Keep an anonymous model behind a feature flag with a tested fallback. Do not route sensitive code, customer data or regulated material until ownership, retention policies and operational terms are clear. Cap spend and concurrency even for a free endpoint, because an outage-induced retry storm can damage the rest of the system. Re-run the evaluation after the launch window. If performance and volume remain strong when novelty fades, the routing market found something worth keeping.
The third lesson is the most important: distribution now moves faster than reputation. An unidentified endpoint can become a top-three model by measured token traffic before the market agrees on who made it. That is valuable feedback about developer appetite, but it is not due diligence and it is not proof of quality. OpenRouter tells us where the load went; Arena tells us which answers voters preferred. Engineering judgment still has to decide whether the workload succeeded—and whether it should stay there.
Sources: OpenRouter Rankings, Space Bunny Alpha model page, OpenRouter dataset methodology, Arena Text leaderboard, Xiaomi MiMo model releases, r/SillyTavernAI practitioner discussion. OpenRouter rankings data is licensed under CC BY 4.0.