Space Bunny Tops OpenRouter, Proving Usage Is Not Quality

Space Bunny Tops OpenRouter, Proving Usage Is Not Quality

The most-used model on OpenRouter this week is not Claude, GPT, Gemini, or DeepSeek. It is a temporary stealth endpoint called Space Bunny Alpha. That sounds like a punchline, but the numbers are serious: 23.4 trillion tokens in a week, enough to move past DeepSeek V4.1 Flash at 22 trillion and take the top spot.

The wrong conclusion is that Space Bunny has suddenly become the best large language model. The useful conclusion is that distribution now moves much faster than trust. A free endpoint with the right placement, integrations, and curiosity gap can acquire an enormous workload before developers know who built it, how long it will exist, or where its failure modes live.

That distinction is the story behind this week's rankings. OpenRouter's top five now spans a stealth provider, DeepSeek, Z.ai, OpenAI, and Xiaomi. Arena's top five, meanwhile, is entirely Anthropic: Claude Opus 5.5 High leads at 1509 Elo, followed by Claude Opus 4.6 High at 1505, Claude Fable 5 High at 1504, Claude Opus 4.7 High at 1502, and Claude Fable 5.1 Max at 1501. Both boards can be accurate because they are measuring different markets.

A usage leaderboard is not an intelligence test

OpenRouter counts prompt and completion tokens processed through its marketplace in complete UTC-day buckets. Private requests are excluded and model variants are counted separately. That makes its chart a valuable view of routed demand, but not a measure of unique users, spending, task success, accuracy, or reasoning quality.

Token volume also contains a hidden variable: verbosity. A model that needs 6,000 tokens to complete a job will accumulate twice the measured volume of one that succeeds in 3,000. Different tokenizers complicate comparisons further. Space Bunny's 23.4 trillion tokens prove that developers sent it a vast amount of work; they do not tell us how much of that work succeeded, how many retries it needed, or whether users kept it after the novelty wore off.

A better analogy is a package-download chart. Downloads reflect real adoption, but they also reflect defaults, CI behavior, transitive dependencies, promotions, and experimentation. OpenRouter's rankings similarly combine capability with price, availability, placement in popular tools, context length, and the economics of agent loops. A free seven-day window can matter more to raw volume than a five-point benchmark lead.

The community response has been appropriately curious rather than conclusive. A Reddit discussion in r/SillyTavernAI had 52 upvotes and 38 comments when indexed, opening with: “Might be next Kimi. Has anybody tried it yet?” That is an unusually honest summary of a stealth launch. People are testing coding, long-context, and multimodal behavior, but there is not yet enough reproducible evidence to turn scattered success reports into a dependable capability profile.

The middle of the table explains how production traffic works

Space Bunny is the headline, but MiMo-V2.6-Flash may be the more durable signal. Xiaomi's model jumped two places to number five on 8.21 trillion tokens, with weekly growth above 999%. Jev 1.13 also moved up a place after 174% growth. Those gains suggest that fast, inexpensive workhorse models are absorbing the repetitive steps inside production pipelines: extraction, classification, transformation, summarization, and tool-driven agent loops.

Claude Opus 5.5 presents the inverse case. It entered OpenRouter's top 20 at number 15 with 1.69 trillion tokens while leading Arena at 1509 Elo. Kimi K3 also entered, at number 18 with 1.49 trillion. A premium model does not need to win raw token volume to be commercially or technically important. Teams can reserve expensive models for planning, difficult debugging, review, and recovery while sending high-volume bounded tasks to cheaper endpoints. The usage leaderboard naturally rewards the workhorses.

Arena asks a different question: which response do people prefer in blind comparisons? Its unchanged top 20, based on the latest published table dated September 25, is not evidence that the market stood still. It is evidence that preference rankings move on a different clock and with different inputs. Space Bunny does not appear in Arena's top 20 at all.

Even Arena's precise-looking ranks deserve restraint. Claude Opus 5.5 High's 1509 score comes from 2,307 votes and carries a ±12 interval. Claude Opus 4.6 High, four points behind at 1505, has 76,518 votes and a much tighter ±3 interval. The neat ordering in the interface hides overlapping uncertainty. “Number one” is excellent shorthand for a shortlist and poor shorthand for a settled scientific result.

Stealth is a launch tactic, not a security posture

Space Bunny's endpoint advertises the capabilities engineers expect from a modern API: streaming, reasoning controls, structured response formats, tool choice, and tool calling. Operationally, that makes it easy to test. Governance-wise, the identifier stealth/space-bunny-alpha tells buyers almost nothing.

Before sensitive code or customer data touches the endpoint, a team still needs answers about the actual provider, retention, training use, regional processing, subprocessors, incident response, and deprecation policy. A compatible API is not a vendor relationship. Anonymous models are fine for a benchmark weekend; they should not receive regulated data or permission to mutate production systems merely because they climbed a public chart.

There is another procurement trap here. OpenRouter sees only traffic through OpenRouter, not direct-provider usage. Anthropic's relatively modest marketplace volume therefore says little about Claude's total share. The same limitation applies whenever developers compare one intermediary's telemetry with the whole market. Keep the denominator attached to every claim.

How to test the chart without becoming part of the hype cycle

Engineering teams should treat this week's movers as candidates, not defaults. Add Space Bunny and MiMo-V2.6-Flash to a controlled evaluation set built from representative internal tasks. Score task completion rather than subjective fluency, and record output tokens per successful task, end-to-end latency, retry rate, tool-call validity, schema compliance, human correction time, and cost. A cheap model that requires three attempts is often the expensive option wearing a lower list price.

Test by workload class. Use bounded extraction and transformation jobs to evaluate cheap models; use ambiguous planning, difficult code review, and recovery from failed tool calls to evaluate frontier models. Replay the same cases against an incumbent so switching decisions have a baseline. For agentic systems, separately measure whether a model chooses the right tool, supplies valid arguments, recognizes failure, and stops. Chat quality alone will miss the failures that wake someone at 3 a.m.

If a stealth endpoint passes, put it behind a routing layer with a kill switch and a known fallback. Pin explicit model IDs where possible, cap spend and token output, log version changes, and assume promotional pricing or availability can disappear faster than an architecture review. Start with low-risk, reversible work. Promotion to sensitive or autonomous tasks should require provenance and security answers, not just good benchmark results.

Space Bunny did not become “the best model” this week. It became the clearest demonstration yet that LLM adoption is a distribution system shaped by price, routing, integrations, and curiosity as much as raw capability. The teams that separate judged quality, routed usage, economics, and operational risk will extract a useful signal from these rankings. Everyone else will spend the week benchmarking a mascot.

Sources: OpenRouter Rankings, Space Bunny Alpha model page, Arena Text leaderboard, r/SillyTavernAI discussion