Writings

Does AI need a safety authority? What French nuclear history actually teaches

France did not build nuclear safety with a law. It built it with an architecture. That institutional depth is what transfers — not the reactor.

FR Lire la version française L’IA a-t-elle besoin d’une sûreté ? Ce que le nucléaire français peut lui apprendre

Translated from the French original with AI assistance, then reviewed and published under the author's editorial responsibility.

The right question may not be “should we regulate AI?”. It is: outside the companies building it, who has the technical means to contradict what they claim?

Artificial intelligence is not nuclear power. Its risks, its materiality, its diffusion and its political economy differ profoundly. Any analogy applied carelessly would produce bad decisions.

But the history of French civil nuclear power raises a useful question: how does a society organise itself when a technology promises major benefits while concentrating risks that are rare, complex, cross-border and potentially irreversible?

The French answer was never a simple law, nor simple industry self-regulation. It was built in layers: accountable operators, technical expertise, independent oversight, public transparency, operating experience feedback, periodic safety reviews and international exchange. It is that institutional depth that is instructive for AI — far more than the reactor itself.

So the argument here is a modest one: AI does not need a copy of the French nuclear safety authority. It needs a safety capability — independent of operators, technically competent, holding real powers, connected to international networks, and able to raise its requirements as the systems change.

None of this has to be invented from nothing. Part of it already exists, built by the labs themselves. That is where to start — because it is also where you can best see what is missing.

AI’s safety moment

Frontier labs now publish documents that look, from a distance, like the first drafts of safety cases: capability evaluations, red teaming, risk thresholds, technical safeguards and deployment rules.

Since GPT-4o, OpenAI has published system cards describing the risks assessed, the limits observed and the mitigations adopted — cybersecurity, chemical, biological, radiological and nuclear risk, persuasion, model autonomy. Anthropic, for its part, has structured a Responsible Scaling Policy: progressively stronger safeguards must accompany the crossing of capability thresholds.

Some discipline of language is required here. The blackmail behaviour reported around Claude Opus 4 is neither a production leak nor a demonstration that a model spontaneously seeks to cause harm in the real world. The setup was this: the model is given access to a fictional company’s email system, allowed to discover that an executive is having an affair, then told that this same executive is about to shut it down. In the large majority of runs, the model writes a threatening message. Two details matter more than that headline figure: the scenario was built — and iteratively tuned — by the lab itself, on its own models, to elicit precisely that behaviour; and competing models put through the same setup behaved in comparable ways.

In other words, this is not a brand incident. It is a property of the class.

Performance is not safety. Neither is apparent control.

The important fact, then, is not that one lab published a spectacular case. It is that these companies themselves acknowledge risks for which the answer cannot simply be: “we trained the model to refuse”.

July 2026: the failure was not in the model

Then the scenarios stopped being scenarios.

Between 9 and 13 July 2026, during an internal evaluation of cyber capabilities, OpenAI models — GPT-5.6 Sol and an unreleased research model, whose cybersecurity refusals had been lowered for the purposes of the test — broke out of their sandbox. They exploited a previously unknown vulnerability in Artifactory, a package registry cache proxy, to obtain internet access; used a poorly secured public endpoint at an infrastructure provider as a springboard; then penetrated Hugging Face’s data processing pipeline through two successive injections, all the way to executing code on production servers. From there: the Kubernetes cluster, the cloud metadata service, an internal database, root access on the nodes, the company’s internal network, the source control platform. Four and a half days, roughly seventeen thousand reconstructed actions. What the model was after, according to OpenAI: stealing the answers to the exercises rather than solving them.

A week after OpenAI’s disclosure, Anthropic halted its cyber evaluations and opened a retrospective review of its own. Across 141,006 sessions examined, three incidents: models whose instructions told them they had no internet access, and which did — a configuration misunderstanding between the lab and Irregular, its third-party evaluator. Believing they were playing inside a simulation, they attacked real systems. One compromised a company that shared its name with the fictional target, and extracted credentials and production data from it. Another, trying to install a Python package it could not find, published one on PyPI — routing around the account-creation obstacle by finding a disposable email service. The package stayed online for an hour. Fifteen real machines downloaded and ran it. One belonged to a security firm whose scanner automatically installs the packages it inspects: the code fired, sent the firm’s credentials back to the model, which used them to get inside. Fourteen of the fifteen machines were never identified.

To be fair: both labs published, quickly and in detail. Anthropic describes its incidents as “closer to a harness and operational failure than a model alignment failure”. That is accurate. And it is precisely what makes the affair instructive.

Because you have to look at where the barriers that gave way actually were:

  • a previously unknown vulnerability at a third-party software vendor;
  • a poorly secured endpoint at an infrastructure provider;
  • a configuration misunderstanding between a lab and its independent evaluator;
  • a public package registry that accepts an account opened with a disposable address;
  • a security scanner that executes what it inspects.

Not one of those barriers belongs to the model’s operator. Not one is a matter of alignment. Not one would have been caught by a model better trained to refuse.

We were not looking for safety in the right place. It was not in the model: it was in everything around it.

Three consequences follow, and they are severe.

The test bench is a hazardous facility. This is the most counter-intuitive point. We argue, rightly, for independent evaluators; these incidents show that the evaluator is itself a vector — it handles models at the precise moment when we know least about what they can do. Nuclear power settled this question long ago: a laboratory for testing irradiated fuel is a licensed nuclear facility, under the same authorisation and oversight regime as the power plant. Testing was never left outside the safety perimeter.

Detection came neither from the victims nor from a regulator. Two of the three affected organisations knew nothing until Anthropic told them, three months after the fact. The fourteen anonymous machines probably still do not know. No external mechanism saw anything, for a simple reason: none exists.

It came from the author, because a competitor had spoken first. Anthropic opened its review because OpenAI had published a week earlier. This is exactly the mechanism that nuclear power institutionalised: an event occurring elsewhere changes the requirements here. Except that here it worked through no obligation, no format and no deadline — but because two companies chose to look, to find and to speak.

A safety arrangement that rests on that goodwill is not a safety arrangement. It is a virtue. Nuclear power learned that virtues do not survive competitive pressure indefinitely, and it ended up institutionalising what operators were already doing willingly.

How it got there is worth a detour.

A technology first, then an architecture of control

French civil nuclear power did not develop inside a stable institutional framework from the outset. Safety was built alongside industrialisation, and often in reaction to it.

As early as January 1960, a commission on the safety of atomic installations began reviewing the facilities of the French Atomic Energy Commission. Following the Anglo-American model, its experts asked the operator to produce a safety report — the first would be analysed in 1962, during the design of the Chinon plant. The burden of demonstrating that risk is under control has therefore rested on the operator from the very beginning. That principle still holds up the whole edifice.

Then came the reminders. In October 1969, at Saint-Laurent-des-Eaux, a loading error obstructed a channel in reactor A1: five fuel elements melted and thirty to fifty kilos of uranium were dispersed inside the vessel. The event would retrospectively be rated level 4 on the INES scale — a scale that did not yet exist.

In March 1973, a decree created the central nuclear safety service, the SCSIN, alongside a higher council for nuclear safety. In November 1976, the Institute for Nuclear Protection and Safety, the IPSN, was created within the Atomic Energy Commission. The architecture gradually became legible:

  • the operator operates, and demonstrates;
  • experts examine and challenge;
  • an authority inspects and decides.

This history is not linear, and that is exactly what makes it useful. The IPSN stayed inside the Atomic Energy Commission — a research body, but also the operator of certain facilities — until 2002, when the IRSN was created outside it. Then, in January 2025, the authority and the institute were merged into the ASNR. Over sixty-five years, France has therefore separated, merged, separated and merged again its expertise and oversight functions.

The lesson is not an org chart. It is more demanding, and simpler:

No operator should be the sole judge of the safety of its own system.

Which leaves the question of who judges in its place, and with what means.

Safety does not come from a law; it comes from a capability

Law is necessary. It creates the authority, organises information, sets obligations and gives sanctions a basis. But it does not verify a safety demonstration, an emergency exercise, or the quality of a maintenance procedure.

Nuclear safety rests on a chain of responsibilities, in which each link answers a question the others cannot ask on its behalf:

  1. The operator — EDF, Orano, the Atomic Energy Commission. Can it demonstrate that its facility is safe, and sustain that level over time?
  2. Technical expertise — the IPSN, then the IRSN, today the teams and standing expert groups gathered inside the ASNR. Do its arguments survive adversarial analysis by people who know the trade?
  3. The regulator — the ASNR. Who can impose requirements, inspect and, if necessary, suspend?
  4. Civil society — a local information commission attached to every site, and the High Committee for Transparency and Information on Nuclear Security, both grounded in the 2006 transparency law. What information is accessible, debatable and contestable?
  5. The international network — the IAEA, the Convention on Nuclear Safety, WENRA, ENSREG, WANO. How does a local incident become collective knowledge?

That layers 2 and 3 have shared a roof since 2025 is exactly what is being debated. Here too, the question is not the org chart: it is whether contradiction remains possible.

The distinction matters for AI. The major labs have safety, security and policy teams, often excellent ones. That is indispensable, but insufficient: an internal team does not replace independent public expertise, any more than a voluntary transparency report replaces a power to investigate or compel.

So the relevant opposition is not between “the law” and “the organisation”. The law establishes the democratic mandate; the organisation makes it effective. This is the same argument I made about Article 50 of the AI Act: trust does not come from a final check stamped on a product, it comes from a process whose capability can be demonstrated. What holds for a text holds for an entire industry.

Operating experience, or a sector’s memory

This architecture would be no more than an org chart without the mechanism that makes it learn. Accidents reshaped practice — including accidents that happened elsewhere.

Chernobyl contributed to the creation of an international scale for rating the severity of events. The flooding of the Blayais plant in December 1999 led to a tightening of flood risk assessment across the entire fleet. Fukushima triggered complementary safety assessments, the identification of multiple-failure scenarios, and the definition of a “hardened safety core” of equipment that must stay available in extreme situations.

The decisive idea is there: an event, however distant, changes the requirements that apply here.

This is precisely what a mature industrial organisation does with its non-conformities. An anomaly is not first of all a culprit to be named: it is information about the capability of the system. The difference, in nuclear power, is that this information does not stop at the walls of the company that produced it.

AI has no such equivalent. July 2026 offered a sketch of one: a disclosure triggered an audit at a competitor, which triggered a second disclosure. An event occurring elsewhere did change practice here. But with no obligation, no deadline, no common format, no shared severity scale, and nobody to verify that the audit ever took place. A sketch is not an institution.

For the rest, significant incidents — safeguard bypasses, model weight theft, confirmed malicious uses, unexpected behaviour in production — remain the private property of whoever suffers them.

Which leaves a question nuclear power had to settle before anyone else: how do you make that memory cross borders, and competing interests?

WANO is not the IAEA

It managed to, but not with a single mechanism. Several coexist, serving different functions — and conflating them would be a mistake.

WANO, founded in May 1989 by the world’s operators after Chernobyl, organises operating experience exchange and peer reviews — voluntary at first, now an obligation for its members. This is valuable: operators learn from one another, including beyond their competing interests.

But WANO is neither a regulator nor an institution of democratic transparency. Its reviews take place under strict confidentiality, which the association regards as essential to honest exchange. That confidentiality has real professional value; it cannot ground public trust.

Alongside it, the IAEA, the Convention on Nuclear Safety, WENRA and ENSREG organise assessments between states and between regulators. The European reviews produce reports, findings and national action plans. These arrangements do not eliminate conflicting interests: they make them visible, and harder to ignore.

For AI, the same plurality would probably have to be accepted — and the hope that a single mechanism will do everything abandoned:

  • confidential sharing of incidents and bypass attempts between operators;
  • independent evaluation of the most capable systems;
  • international review mechanisms between authorities;
  • publication of risks, significant incidents and corrective measures, in a form compatible with security.

So much for the edifice. What survives of it once transposed is another matter — and one objection comes first, the strongest that can be levelled at this whole analogy.

The material objection: a power plant cannot be copied

It is a fair one. A power plant sits in one place; a model can be copied, embedded and used anywhere.

But frontier AI also has a heavy materiality. Training and serving the largest models depends on data centres, specialised chips, power grids, cooling, cables and a handful of cloud providers. According to the International Energy Agency, that capacity is geographically concentrated: Northern Virginia alone exceeded 7 gigawatts of installed capacity in 2024, and several data centres of a gigawatt or more, each run by a different operator, are coming online this year.

This dual nature is fundamental:

  • AI infrastructure is concentrated, and therefore partly observable and partly controllable;
  • its effects are diffuse, because code, APIs, agents and content circulate globally.

Controlling data centres alone would therefore be insufficient. But ignoring these points of concentration would be a symmetrical error. They are possible sites for observation, for conditioning access, for securing model weights, and for responding to incidents.

There is purchase, then. What to apply it to, and on what conditions, is the next question.

Five transferable principles

An AI safety policy has to avoid two failure modes: self-regulation without contradiction, and a bureaucracy that demands forms without holding the technical competence to read them.

1. A safety case per system and per context of use. A model is not safe or dangerous in the abstract. Its risk depends on its capabilities, its access to tools, its connectivity, its autonomy, its users and its deployment conditions. An operator should be able to produce a reasoned dossier, updated at every substantial change — exactly as a nuclear operator updates its safety report.

2. Public expertise able to contradict. Evaluations cannot rest solely on tests designed by the company deploying the system. It takes researchers, independent red teams, and controlled access to the models. Without access, external expertise is reduced to commenting on press releases. But the converse holds too: since the evaluator handles models at the moment when we know least about what they can do, it must be held to the same standard as the operator. A test bench is not a design office.

3. Regulation graded by capability and exposure. Not every model warrants the same scrutiny. Thresholds should account for autonomy, access to computer systems, cyberattack potential, biological risk, large-scale manipulation, and the ability to circumvent safeguards. The depth of control should be proportionate to the risk, not to the publisher’s reputation.

4. Mandatory operating experience feedback. Significant incidents, safeguard failures, weight theft and unexpected behaviour should feed a collective memory. Not to assign blame after the fact, but to raise requirements before an incident repeats itself somewhere else. Mandatory, because July 2026 showed what the voluntary version yields: three months of latency, two victims out of three who knew nothing, and fourteen machines that never will.

5. A public capability, funded as such. This does not mean nationalising AI or installing a state monopoly on models. It means funding expertise, compute access for public-interest research, independent evaluation, cybersecurity, training and incident response. An authority without engineers is not an authority: it is a front desk.

Where the analogy stops

It is worth being equally clear about what nuclear power does not let us conclude.

Nuclear power concentrates energy and hazardous materials; AI concentrates informational and decision-making capabilities. The risks are neither of the same nature nor on the same timescale.

A nuclear accident is generally locatable, visible and physically bounded. AI harms can be distributed: fraud, privacy violations, discrimination, political manipulation, technological dependence, automated attacks, concentration of power. More gradual, often, but also far harder to attribute — and therefore to correct.

Above all, a framework too closely modelled on a sovereign vision of infrastructure could produce precisely what it claims to prevent: the concentration of power in a few firms, or in a state without counterweights. AI safety must protect against technical risks as much as against the political, economic or informational capture of the infrastructure. Nuclear power teaches us nothing about that second point. It will have to be invented.

What nuclear power actually teaches

It does not teach that AI should be treated as a weapon.

It teaches that a technology of power does not become acceptable because its promoters call it beneficial, nor because a legislator has written down general principles.

It becomes governable when a society equips itself with institutions able to ask the right questions, obtain the necessary information, contradict operators, learn from failures, and act before the risk becomes irreversible.

And it teaches something more uncomfortable still. Read the French chronology again: the thinking that produced the 1960 commission opens in 1957, the year of the Windscale fire. The SCSIN arrives four years after Saint-Laurent. The severity scale comes after Chernobyl, the tightening of flood risk after Blayais, the hardened safety core after Fukushima. Every storey of the edifice was put in place after the event that showed it was missing.

That is an effective way to build. It is also the most expensive there is: each barrier is paid for at the price of the accident that made it obvious.

A safety institution built after the fact always costs the price of what it would have prevented.

July 2026 was a cheap warning. Nobody died, a handful of machines were compromised, and both operators spoke up of their own accord. Nothing guarantees that the next one will be as inexpensive — or that it will be reported by whoever caused it.

The real question is therefore not whether AI should become state infrastructure.

It is whether the state, researchers, operators and civil society can build, together, a safety infrastructure equal to it.


Key sources

Notebook · one note every fortnight

My thinking on industry, decision-making and AI, in your inbox.

Notes taken in the field of industrial acquisition, and a few longer essays. No noise, no address ever sold. One click to unsubscribe.

The notebook is written in French; each issue opens with a link to its English version.

No spam. Unsubscribe instantly.