Three states passed AI transparency laws — but who knows which models are actually covered?

California’s AI transparency law requires disclosures for the most powerful models. But it never requires companies to say which models hit the thresholds. New York and Illinois didn’t fix this gap either.

Samantha Stein

·

No headings found

In September, we reported that OpenAI appeared to have left out a loss-of-control risk determination that California law requires for models that hit certain compute thresholds (and which OpenAI’s own safety framework requires). But while writing that piece, we hit a snag: We weren’t certain whether the models were big enough — compute size-wise — to definitively be covered by California law. So we hedged, describing those models as “likely covered” by the law.

California’s Transparency in Frontier Artificial Intelligence Act, known colloquially as “SB 53,” applies to “frontier models”: foundation models trained with more than 10^26 integer or floating-point operations (FLOPs), counting the original training run plus any later fine-tuning or reinforcement learning.[^1] Before publicly deploying such a model, SB 53 requires a developer to publish a transparency report compliant with the requirements in the statute, and the largest developers must also summarize their catastrophic risk assessments. Notably, New York’s Responsible AI Safety and Education Act, or RAISE Act, which takes effect Jan. 1, 2027, uses a substantively identical definition, as does Illinois’ Artificial Intelligence Safety Measures Act, or SB 315, which also takes effect Jan. 1, 2027, and gradually increases transparency obligations over time.

Critically, none of these laws requires a developer to disclose whether a newly deployed model crossed the statute’s compute threshold, thus triggering transparency requirements. It’s a significant gap in the law.

Why does this matter?

When a frontier AI company releases a model with no transparency report, one of two things is true: Either the model is under the statutory size threshold, or the AI company is breaking the law. From outside, the two look the same. And it gets murkier still because SB 53 allows a system card or model card to double as the statutory transparency report, so you often can’t tell whether the AI developer actually intended that document to satisfy the law or if it was just a generic model summary.

Thus whether the law applies to a particular model is information known only to the AI company. Everyone else has to guess: the public, journalists, and watchdogs like us.

SB 53 was meant to make the companies more transparent and monitorable. But when we can’t even know with certainty that the transparency law applies to a particular model, that’s a critical gap in our ability to track whether the companies are meeting their obligations and taking safety and catastrophic risks seriously.

But surely you must be able to tell…

One might guess that there must be something in the AI companies’ model cards, system cards, something, that would give interested readers a strong sense about whether a particular model meets SB 53’s size threshold. Unfortunately, there is considerable opacity in the information the companies disclose.

Take OpenAI. On Sept. 3 the company released GPT-6 Astra, a powerful model, which, by OpenAI’s own account, is the company’s first to reach the “Critical” level for cybersecurity capability. One assumes such a model must be big — at least 10^26 FLOPs. But no OpenAI document confirms this. Nor does OpenAI provide any information about the model’s parameters or token use to be able to roughly calculate the model’s size. Hence our hedge about the model “likely” being covered by SB 53.

Part of the issue is the way companies may be training models now, which often isn’t just based on one big pretraining run plus subsequent post-training, but can involve a technique known as distillation[^2] — a process in which a teacher model essentially passes along knowledge like its outputs and probabilities to a student model, which in turn has its own weights and could fall below 10^26 FLOPs of training compute due to the nature of this training method. SB 53 is silent about distillation and whether the teacher model’s pretraining run should count toward the statute’s compute threshold. 

Now, to be clear, there is some uncertainty about whether SB 53 applies to distilled models. For example, Lawfare wrote back in December 2025 that “[t]he Department of Technology may need to include a teacher model’s compute in evaluating whether a distilled model qualifies as a frontier model to avoid enforcement gaps,” explaining how “without including that compute, developers could bypass SB53 by training a larger teacher model and distilling it into a smaller model with frontier capabilities that fall outside the statute’s scope.” But others suggest that the statute should be more explicit in including distilled models. Either way, our broader point stands: We don’t know whether any given model definitively hits the compute threshold under these transparency laws.

And at this point, it is entirely possible that many of the models we now see are based on internal distillation. For example, GPT-6 Sol and GPT-6.1 Sol noticeably didn’t get their own system cards but rather got an appendix and an addendum tacked onto Astra’s system card, which may be a signal that one or both are distilled from Astra or some other common lineage model. Whatever the case is, OpenAI hasn’t said whether either GPT-6 Sol or GPT-6.1 Sol meets SB 53’s size threshold — and the law doesn’t require it to.

This issue isn’t just about OpenAI, but applies to other frontier AI companies as well. For example, when Anthropic releases a 230-page system card vs. a 319-page one, are people supposed to infer that the slimmer system card represents a smaller model? And for Google’s Gemini 3.8 Flash, its model card is only eight pages long and refers readers to the card for its predecessor, Gemini 3.7 Flash, for its architecture, training data, and safety policies. So does that mean 3.8 Flash itself is too small, training compute-wise, perhaps distilled from a larger model,[^3] and thus wouldn’t be considered a “frontier model” under SB 53?

There’s an obvious, repeated opacity here — one that only benefits the AI companies.

New York & Illinois have not solved the problem

For a minute, it seemed as if New York’s RAISE Act was going to fix this gap in the law. A version of the RAISE Act expressly covered distilled models, and the first signed version of the act preserved the concept that a frontier model included “an artificial intelligence model produced by applying knowledge distillation to a frontier model” defined by the size threshold. Unfortunately, that provision disappeared in March, when New York rewrote its law to match California’s. Illinois’ SB 315 likewise measures its transparency requirements based on the 10^26 FLOPs threshold and doesn’t mention anything about distillation.

Yes, it’s true that under both the New York and Illinois laws, large frontier developers must file disclosure statements with these states identifying themselves, and the states in turn then publish a list of those filers. But while that tells us who is covered, it still doesn’t tell us which models are covered.[^4]

Conclusion

A transparency law that leaves the public guessing which models it covers is not that transparent — particularly since, as we understand it, determining a model’s training compute is not technically difficult for the AI developers.[^5] The companies wouldn’t even need to publish details like their parameter counts or training-token totals; a simple yes-or-no answer to whether a model crossed the size threshold would go a long way on its own.

Regardless of what the law says today,[^6] if the frontier AI developers are truly interested in transparency, then they could easily provide a simple yes or no about whether their models meet the relevant size thresholds. 

  1. The statute at Section 22757.11(i) states:

    (1) “Frontier model” means a foundation model that was trained using a quantity of computing power greater than 10^26 integer or floating-point operations.
    (2) The quantity of computing power described in paragraph (1) shall include computing for the original training run and for any subsequent fine-tuning, reinforcement learning, or other material modifications the developer applies to a preceding foundation model.

  2. Few of the AI labs have admitted publicly to distilling their own models. However, Meta provides a write-up of its process: “The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation.”

  3. Google has distilled its own models in the past. See Google Gemini updates: 1.5 Flash, Gemma 2 and Project Astra (“1.5 Flash … [has] been trained by 1.5 Pro through a process called ‘distillation,’ where the most essential knowledge and skills from a larger model are transferred to a smaller, more efficient model”).

  4. After Jan. 1, 2028, auditors under Illinois’ SB 315 should arguably have access to this information as part of their auditing duties. However, there’s no clear requirement that the public will get to share in this knowledge. 

  5. Epoch AI, a nonprofit research institute that tracks trends in AI development, including the computing power used to train major models, has discussed how it makes estimates on model size. See “Training open-weight models is becoming more data intensive” (August 2025) and “AI models documentation – Estimation.”

  6. California’s SB 53 directs the state’s Department of Technology to recommend updates to its definitions every year, including weighing “the external verifiability of determining whether a person or foundation model is covered.” New York and Illinois likewise have their own mechanisms for updating their AI laws and rules.

Power & Policy

Deeply researched analysis of the AI industry, policy moves, and the forces shaping the rules of artificial intelligence — delivered to your email.

Power & Policy

Deeply researched analysis of the AI industry, policy moves, and the forces shaping the rules of artificial intelligence — delivered to your email.

Power & Policy

Deeply researched analysis of the AI industry, policy moves, and the forces shaping the rules of artificial intelligence — delivered to your email.

Smooth Scroll
This will hide itself!