The fourth installment of The Decoder's Frontier Radar series has reached a conclusion that would have sounded implausible two years ago: the capability gap between Chinese and Western frontier AI labs has, for practical purposes, closed. Frontier Radar is a recurring benchmark comparison that periodically checks in on how the leading US and Chinese AI labs stack up against one another, and this round found the once-wide lead compressed to something close to a rounding error.
According to The Decoder, the dominant narrative of 2023 and 2024 — US labs setting the pace while Chinese labs trailed a generation behind — no longer describes the market. The outlet's framing has shifted from "how far behind is China" to "what, if anything, is left of the Western lead."
For most people that's a headline. For anyone choosing a model provider or committing to an AI stack that has to hold up for years, it's a planning input.
A lead measured in months, not generations
The "Western lead" used to be treated as a structural fact rather than a temporary scoreboard position — the assumption baked into procurement decisions, security reviews, and vendor shortlists across the industry. DeepSeek's January 2025 release was the first widely visible crack in that assumption, triggering a market reaction loud enough that most people outside AI circles heard about it. Frontier Radar's fourth edition is effectively confirming that the crack never closed: the pattern DeepSeek started has continued across multiple Chinese labs and release cycles since, not just one company having one good quarter.
That doesn't mean every Chinese model beats every Western model on every task. Benchmark leadership rotates between labs on both sides every few months, and it will likely keep rotating — in our estimation, no single lab holds a durable lead in this environment for more than a couple of release cycles. What's changed is the ceiling: Chinese labs are now routinely operating at or near the frontier rather than one tier below it.
Why builders should care about a benchmark story
A closing capability gap matters less for the score itself and more for what it does to the assumptions behind a model-selection process. Teams that treat "US lab" as a proxy for "best available model" are now working from an outdated shortcut. That shortcut used to save research time; today it filters out options that may be a better fit on cost, licensing, or latency grounds without ever getting evaluated.
Consider a mid-size team evaluating providers for a new product this quarter. A shortlist assembled purely on 2023-era instincts would drop Chinese labs before running a single eval. Frontier Radar's finding suggests that shortlist is now excluding viable options for reasons that no longer hold, not because of a real capability shortfall.
What changes on a practical model-selection checklist
Planning a stack for the next few years now means treating Chinese frontier models as a real bucket to evaluate, not a footnote. A few things are worth building into that evaluation:
- Licensing: several leading Chinese models ship with permissive open-weight licenses, which changes the self-hosting and fine-tuning math compared to closed Western APIs.
- Pricing: Chinese labs have consistently priced aggressively relative to comparable Western frontier models, which matters at production token volumes.
- Data residency and compliance: regulated industries still need to weigh where inference happens and under which jurisdiction, independent of raw capability.
- Tooling maturity: agentic features, enterprise support, and ecosystem integrations remain areas where longer-established Western labs may still have an edge — this is where due diligence should actually happen, rather than at the capability layer.
- Evaluation cadence: given how fast rankings move, a one-time model bake-off is no longer sufficient — plan to re-benchmark shortlisted providers at least twice a year.
AiiN's takeaway
The practical lesson isn't "switch to a Chinese model." It's that the default shortlist for any multi-year stack decision should now include both sides of the Pacific by default, with the evaluation weight shifted off raw capability — where the gap has apparently narrowed — and onto licensing, cost, and operational fit, where real differences still exist. Treating Chinese frontier models as a peer category rather than a hedge is the more defensible starting position for anyone building infrastructure meant to survive the next model generation, not just the current one.