The  LGTM
  • Home
  • Agentic Coding
  • Claude Code
  • Codex
Sign in Subscribe
Kimi K3 Shows Why LLM Rankings Are Splitting Into Two Markets
llm-rankings

Kimi K3 Shows Why LLM Rankings Are Splitting Into Two Markets

The most interesting thing in this week's LLM rankings is not that another model moved into the top 10. That happens often enough now that it barely qualifies as news. The interesting part is that the model doing it, Moonshot's Kimi K3, looks more impressive on
18 Jul 2026 6 min read
LLM Rankings Have Split Into Prestige Models and Workhorse Models
llm-rankings

LLM Rankings Have Split Into Prestige Models and Workhorse Models

The most useful LLM ranking this week is not the one with the cleanest number one. It is the gap between the rankings. Arena's Text leaderboard still looks like a prestige table: Anthropic occupies the entire top five, with claude-fable-5 at 1508 Elo, claude-opus-4-6-thinking at 1504, claude-opus-4-7-thinking at
16 Jul 2026 6 min read
LLM Rankings Are Splitting Into Prestige Models and Workhorse Models
llm-rankings

LLM Rankings Are Splitting Into Prestige Models and Workhorse Models

The most useful LLM ranking this week is not the one where the “best” model changed. It is the one where the most-used model changed. OpenRouter’s weekly rankings now have MiMo-V2.5 at number one with 5.40 trillion tokens routed over the last week, just ahead of DeepSeek
13 Jul 2026 6 min read
LLM Rankings Have Split Into Prestige Models and Workhorse Models
llm-rankings

LLM Rankings Have Split Into Prestige Models and Workhorse Models

The useful question in LLM rankings is no longer “which model is best?” That question was always too vague, but now it is actively misleading. The leaderboard world has split into two markets: models that win judged preference tests, and models that developers route trillions of tokens through because they
12 Jul 2026 6 min read
Tencent's Hy3 Shows the Free-Tier Leaderboard Effect Is Real
llm-rankings

Tencent's Hy3 Shows the Free-Tier Leaderboard Effect Is Real

The interesting LLM ranking this week is not on the board everyone argues about. Arena AI's text leaderboard was basically frozen: Anthropic still owns the top cluster, Google and OpenAI are still bunched close behind, and nobody broke into the top 20. The real movement came from OpenRouter,
10 Jul 2026 5 min read
LLM Rankings Have Forked Into Prestige and Production
llm-rankings

LLM Rankings Have Forked Into Prestige and Production

Leaderboards are usually treated like scoreboards. That is the first mistake. This week's LLM rankings are better read as a market map: one board shows which models win judged preference, while another shows which models developers actually route when money, latency, quotas, defaults, and reliability enter the room.
08 Jul 2026 5 min read
llm-rankings

LLM Rankings Are Splitting Quality Signals From Production Demand

The useful thing about today's LLM rankings is how little the top of the charts moved. That sounds like a boring leaderboard day, but it is the opposite. When the prestige rankings freeze and the usage rankings keep shifting underneath them, you get a cleaner view of the
06 Jul 2026 6 min read
LLM Rankings Have Split Into Prestige Models and Production Workhorses
llm-rankings

LLM Rankings Have Split Into Prestige Models and Production Workhorses

The useful LLM leaderboard story this week is not that one model moved one slot. That is leaderboard cosplay. The real story is that the industry now has two different scoreboards telling two different truths: models that win judged preference tests, and models that developers are willing to route trillions
05 Jul 2026 6 min read
The LLM Leaderboard Has Forked Into Quality and Throughput
llm-rankings

The LLM Leaderboard Has Forked Into Quality and Throughput

The most useful thing in this week's LLM rankings is not the model sitting at #1. It is the disagreement between the boards. On Arena Text, Anthropic still owns the top of the judged-quality stack: claude-fable-5 leads at 1509 Elo, followed by claude-opus-4-6-thinking at 1504, claude-opus-4-7-thinking at 1502,
04 Jul 2026 5 min read
LLM Rankings Are Split Between Prestige and Production Gravity
llm-rankings

LLM Rankings Are Split Between Prestige and Production Gravity

The most useful LLM leaderboard story this week is the boring-looking one: the quality board barely moved, while the usage board kept rearranging itself underneath. That is exactly the kind of split engineers should care about. Benchmarks tell you what wins in a controlled comparison. Production token volume tells you
03 Jul 2026 6 min read
LLM Rankings Are Split Between Prestige and Production Gravity
llm-rankings

LLM Rankings Are Split Between Prestige and Production Gravity

The most useful LLM ranking this week is not the one with the prettiest Elo number. It is the messy one that shows what developers are actually paying to run. Arena's text leaderboard is basically holding still: claude-fable-5 leads at 1508 Elo, followed by claude-opus-4-6-thinking at 1503, claude-opus-4-7-thinking
02 Jul 2026 6 min read
The LLM Leaderboard Is Stable. The Usage Market Is Not.
llm-rankings

The LLM Leaderboard Is Stable. The Usage Market Is Not.

The useful LLM leaderboard story this week is not that a new champion walked onto the stage. It is that the stage barely moved while the loading dock behind it kept rearranging itself. Arena's text leaderboard is stable at the top: Anthropic holds the first five slots with
01 Jul 2026 5 min read
Page 1 of 136 Older Posts →
The LGTM © 2026
  • Sign up
Powered by Ghost