The fourth installment of The Decoder's Frontier Radar series has reached a conclusion that would have sounded implausible two years ago: the capability gap between Chinese and Western frontier AI labs has, for practical purposes, closed. Frontier Radar is a recurring benchmark comparison that periodically checks in on how the leading US and Chinese AI labs stack up against one another, and this round found the once-wide lead compressed to something close to a rounding error.

According to The Decoder, the dominant narrative of 2023 and 2024 — US labs setting the pace while Chinese labs trailed a generation behind — no longer describes the market. The outlet's framing has shifted from "how far behind is China" to "what, if anything, is left of the Western lead."

For most people that's a headline. For anyone choosing a model provider or committing to an AI stack that has to hold up for years, it's a planning input.

A lead measured in months, not generations

The "Western lead" used to be treated as a structural fact rather than a temporary scoreboard position — the assumption baked into procurement decisions, security reviews, and vendor shortlists across the industry. DeepSeek's January 2025 release was the first widely visible crack in that assumption, triggering a market reaction loud enough that most people outside AI circles heard about it. Frontier Radar's fourth edition is effectively confirming that the crack never closed: the pattern DeepSeek started has continued across multiple Chinese labs and release cycles since, not just one company having one good quarter.

That doesn't mean every Chinese model beats every Western model on every task. Benchmark leadership rotates between labs on both sides every few months, and it will likely keep rotating — in our estimation, no single lab holds a durable lead in this environment for more than a couple of release cycles. What's changed is the ceiling: Chinese labs are now routinely operating at or near the frontier rather than one tier below it.

Why builders should care about a benchmark story

A closing capability gap matters less for the score itself and more for what it does to the assumptions behind a model-selection process. Teams that treat "US lab" as a proxy for "best available model" are now working from an outdated shortcut. That shortcut used to save research time; today it filters out options that may be a better fit on cost, licensing, or latency grounds without ever getting evaluated.

Consider a mid-size team evaluating providers for a new product this quarter. A shortlist assembled purely on 2023-era instincts would drop Chinese labs before running a single eval. Frontier Radar's finding suggests that shortlist is now excluding viable options for reasons that no longer hold, not because of a real capability shortfall.

What changes on a practical model-selection checklist

Planning a stack for the next few years now means treating Chinese frontier models as a real bucket to evaluate, not a footnote. A few things are worth building into that evaluation:

AiiN's takeaway

The practical lesson isn't "switch to a Chinese model." It's that the default shortlist for any multi-year stack decision should now include both sides of the Pacific by default, with the evaluation weight shifted off raw capability — where the gap has apparently narrowed — and onto licensing, cost, and operational fit, where real differences still exist. Treating Chinese frontier models as a peer category rather than a hedge is the more defensible starting position for anyone building infrastructure meant to survive the next model generation, not just the current one.