OpenAI wrote a safety framework for California. Its system cards are ignoring it.

Under SB 53, OpenAI’s Frontier Governance Framework calls for a loss-of-control risk tier for every covered model. The GPT-5.6 and GPT-6 Astra system cards assign none.

By

Brittney Gallagher

-

No headings found

Last year, California passed a law requiring OpenAI to abide by the safety rules it sets for itself. In particular, it has to create a suite of risk thresholds for topics including loss of control — the risk of an AI subverting human oversight and containment — and publicly disclose how it has evaluated each of its models according to those risk thresholds. This risk is especially salient given the recent loss-of-control “warning shots” that OpenAI has faced.

Despite this obligation, OpenAI’s major model releases appear to have neglected this. If so, this could be a violation of California’s AI safety law — not the first such incident that we have disclosed from OpenAI and others. As these incidents continue, one begins to wonder whether the laws have any teeth.

California’s SB 53, also known as the Transparency in Frontier AI Act (TFAIA), is a light-touch transparency law requiring that AI companies make safety plans and follow them. Until New York’s RAISE Act takes effect on Jan. 1, 2027, California’s law is the only one in the United States that holds frontier AI developers accountable to their own safety frameworks. At the time the law was passed, many AI companies had already adopted frontier safety frameworks: Anthropic’s Responsible Scaling Policy, OpenAI’s Preparedness Framework, Google’s Frontier Safety Framework, among others. The law takes these voluntary frameworks and makes them legally binding, though it allows AI companies to change the framework’s contents. The motivation for allowing these changes was that, because the technology was moving so quickly, AI companies needed to be able to adjust their safety practices over time. 

Anthropic was an early and strong supporter of SB 53. In its endorsement, the company wrote,

These requirements would formalize practices that Anthropic and many other frontier AI companies already follow. At Anthropic, we publish our Responsible Scaling Policy, detailing how we evaluate and mitigate risks as our models become more capable … Now all covered models will be legally held to this standard.

But just before SB 53 went into effect, Anthropic released a new safety policy, the Frontier Compliance Framework, and designated it — not the Responsible Scaling Policy — as its legal framework under the law. It’s similar to the voluntary policy but noticeably vaguer and watered down. In essence, Anthropic undercut the law it supported: the Responsible Scaling Policy is no longer the document Anthropic is legally required to follow. Anthropic argued that a separate compliance policy for SB 53 was necessary because the voluntary framework mixes binding commitments it plans to follow as a baseline with aspirational ones it doesn’t want to be legally binding. But Anthropic did a major overhaul to its Responsible Scaling Policy in February 2026; the revised version separated out its plans as a company, “those which we expect to achieve regardless of what any other company does,” from its more ambitious recommendations for the industry. In practice, Anthropic could have kept its Responsible Scaling Policy as its legally binding safety framework by clarifying which practices within it were baseline and which were aspirational. Instead, it opted for a separate, thinner policy.

At the time, I wondered if other labs would follow suit. 

Sure enough, in May 2026, OpenAI published its own substitutive compliance document, the Frontier Governance Framework, as the designated framework under SB 53. In the framework, OpenAI states: 

Under California’s Transparency in Frontier AI Act (TFAIA), this FGF is our Frontier AI Framework, documenting OpenAI’s technical and organizational protocols to manage, assess, and mitigate catastrophic risks, as defined under the TFAIA.

It goes on to say:

The FGF overlaps in some areas with our existing Preparedness Framework (PF). The PF and FGF together describe OpenAIs practices, and we will continue to use and evolve the PF to define and operationalize OpenAIs own approach to managing the most serious risks from advanced AI systems, including in situations where our internal practices go beyond current legal requirements.

SB 53 requires the largest AI companies, including OpenAI, to not only publish safety frameworks but also to evaluate covered model releases against them. Under both Anthropic’s and OpenAI’s frameworks, this means designating specific risk tiers for each covered model, implementing any required safeguards associated with that tier, making an overall risk-acceptance determination, and including that in a public transparency report or the model’s system card.

Anthropic does this: the Claude Fable 5.1 & Claude Mythos 5.1 system card names the Frontier Compliance Framework and states a tier and threshold conclusion for cyber, chemical and biological risks, harmful manipulation, and loss of control (though it’s labeled as “autonomy” rather than using the framework’s terminology). 

Meanwhile, OpenAI has had at least three major model releases that were likely covered by SB 53 since publishing the Frontier Governance Framework: GPT-5.6 Preview, GPT-5.6, and GPT-6 Astra. The system cards for all three report risks in reference to the Preparedness Framework, using its vocabulary, risk categories, and risk thresholds. The Frontier Governance Framework, OpenAI’s safety framework that is legally binding under SB 53, is not mentioned in any system card.

For cyber and CBRN, this may be excusable: the Preparedness Framework and the Frontier Governance Framework are similar enough that evaluating risk in reference to the Preparedness Framework is essentially identical to evaluating it in reference to the compliance framework, at least until the two documents diverge further. And it’s likely that divergence is coming. On Aug. 18, after the incident with Hugging Face, OpenAI said it would “evolve” the Preparedness Framework, but said nothing about the Frontier Governance Framework. To OpenAI’s credit, the Preparedness Framework also includes a category for AI self-improvement that SB 53 does not require. 

There’s one risk domain covered by SB 53 that the Frontier Governance Framework addresses, but the Preparedness Framework does not: loss of control. SB 53’s definition of catastrophic risk includes the risk of a model “evading the control of its frontier developer or user,” which may explain why the category exists in the Frontier Governance Framework but not the Preparedness Framework. It defines loss of control as “risks arising from humans losing the ability to reliably direct, modify, or shut down a model” and lays out a three-tier hierarchy of escalating risk thresholds, similar to its structure for cyber and CBRN. The Frontier Governance Framework says that for every covered model release, OpenAI makes “the determination that a threshold has or has not been reached,” reviewing “the risk tiers for each systemic risk category”; implements safeguards “proportionate to the level of risk”; and documents “a justification for why the systemic risks stemming from the model are acceptable.” It goes on to say that those results “are documented in a Safety and Security Model Report (referred to as ‘Transparency Reports’ under the TFAIA),” with details published “in system cards or in other reporting when a model launches.” OpenAI has not published a standalone Safety and Security Model Report for any of the models above. The system cards are the only detailed reporting that exists.

In the Astra system card, the alignment and monitorability sections contain unstructured discussion of whether humans can reliably direct the model, but little on whether they can reliably modify or shut one down. None of the three system cards includes what the Frontier Governance Framework requires for every model: a loss-of-control risk tier designation and a conclusion that the risk is acceptable. The system card does describe safeguards aimed at misaligned behavior — a real-time misalignment monitor and new internal security controls — but with no risk tier assigned, there is no way to tell whether they are the safeguards the Frontier Governance Framework would require at Astra’s level or whether OpenAI considers them sufficient. 

The Frontier Governance Framework’s examples for Tier 1 and Tier 2 — a model that “may subtly underperform when instructed,” one that “can identify and exploit gaps in monitoring systems” — describe how Astra itself behaved in the system card in adversarial tests. Given the recent loss-of-control incidents OpenAI has faced, one would hope that regulators and the public would get information on how models like Astra perform. The system card should at minimum include a tier designation, the safeguards chosen because of that tier, and a risk-acceptance determination, not merely the tests that would inform one.

Three days after Astra was released, OpenAI’s chief scientist, Jakub Pachocki, called for evolving “commitments like the Preparedness Framework or Responsible Scaling Policy into widely mandated safety bars.” The mandated bar already exists.

OpenAI did not respond to our request for comment.

Power & Policy

Deeply researched analysis of the AI industry, policy moves, and the forces shaping the rules of artificial intelligence — delivered to your email.

Power & Policy

Deeply researched analysis of the AI industry, policy moves, and the forces shaping the rules of artificial intelligence — delivered to your email.

Power & Policy

Deeply researched analysis of the AI industry, policy moves, and the forces shaping the rules of artificial intelligence — delivered to your email.

Smooth Scroll
This will hide itself!