In May 2024, Jan Leike, then co-lead of OpenAI's Superalignment team, resigned and said publicly that "safety culture and processes have taken a backseat to shiny products" inside the company. A year and a half later, that complaint reads less like one team's internal grievance and more like a pattern that keeps recurring across the industry.

According to The Decoder, AI developers are still struggling to enforce the safety and oversight commitments they've made for their own systems. The frameworks exist on paper — responsible scaling policies, preparedness frameworks, frontier safety commitments — but the gap between what labs promise and what actually happens before a model ships keeps showing up.

For teams building products on top of these models, that gap isn't a governance footnote. It's the difference between trusting a vendor's "red-teamed" label at face value and knowing you need your own evaluation layer regardless of what the system card says.

Self-imposed rules, inconsistently kept

Every major lab now publishes some version of a capability-gated safety framework: Anthropic's Responsible Scaling Policy, OpenAI's Preparedness Framework, Google DeepMind's Frontier Safety Framework. The idea is straightforward — define risk thresholds in advance, commit to specific evaluations before crossing them, and pause or add safeguards if a model trips a threshold.

The problem is enforcement. These are voluntary, self-authored commitments with no external regulator checking compliance in real time. When a launch date collides with a safety evaluation that isn't finished, the lab is grading its own homework on whether to ship anyway. Leike's resignation is one visible data point; it's not the only one — safety and alignment staff have left multiple labs over the past two years citing similar concerns about process being subordinated to shipping speed.

External scorecards keep landing in the same place

Independent attempts to grade the industry tell a consistent story. The Future of Life Institute's AI Safety Index, which periodically scores frontier labs on categories including risk assessment, governance, and existential-safety planning, has repeatedly found no company earning better than a C-range grade overall, with several landing in D or F territory on accountability and existential-safety categories specifically. The scores move a little between editions; the overall picture — stated ambition well ahead of demonstrated practice — hasn't.

That throughline matches what The Decoder describes: labs aren't short on safety frameworks, they're short on consistent adherence to the ones they've already written down.

What this means if you're building on these models

None of this makes frontier models unsafe to build on by default. It does mean the safety layer a lab publishes is a starting point, not a guarantee that covers your specific application. A few adjustments worth making:

AiiN's takeaway

The uncomfortable part isn't that AI labs have safety frameworks with gaps — most complex organizations do. It's that the frameworks exist specifically to catch risks the labs themselves define as serious enough to warrant a pause, and the recurring pattern across companies and years is that competitive pressure wins that argument more often than the framework does. In our estimation, that gap is unlikely to close through voluntary commitments alone, which is the strongest argument now available to advocates of external AI safety regulation. Until enforcement has real teeth, builders are better off treating every frontier lab's safety guarantees as a floor to verify, not a ceiling to rely on.