Google DeepMind CEO Demis Hassabis has proposed a new standards body to test frontier artificial intelligence models before they reach the public, adding one of the industry’s most influential voices to the debate over global AI oversight.
The idea is designed to solve a genuine problem. The most capable models are developing faster than governments can build technical expertise, while voluntary company policies vary and national rules risk fragmenting the market.
But the central challenge is not whether advanced models should be tested. It is who controls the test, which risks count and whether countries will accept a body initially led by the United States.
What Hassabis is proposing
In an essay titled “A Framework for Frontier AI and the Dawning of a New Age,” Hassabis argues for an independent, technically capable standards organisation focused on the most advanced AI systems.
The body would be funded by the industry, staffed by experts and accountable to public authority. Frontier labs would initially submit models voluntarily for assessment before release. Hassabis suggests a period of up to 30 days, allowing evaluators to test for critical capabilities and vulnerabilities.
If the system proved credible, assessment could later become a formal condition for deploying frontier models in the US market. The organisation could also coordinate responses to serious vulnerabilities found after release and, in an extreme case, help organise an industry-wide pause.
The model resembles sector regulators and self-regulatory bodies that combine specialised expertise with government oversight. Its appeal is speed: a dedicated organisation could update technical standards faster than a legislature.
Why frontier AI needs a different mechanism
Traditional product regulation assumes risks are reasonably stable and can be measured with established tests. Frontier AI complicates that model because capabilities can emerge unpredictably, models are updated frequently and the same system can be applied across many domains.
Potential concerns include cyber operations, biological misuse, autonomous replication, deception and the loss of effective human control in high-impact settings. Not every advanced model will display these capabilities, and benchmark results do not automatically translate into real-world harm. That is precisely why evaluation methods require scientific credibility.
A specialised body could create shared protocols, require consistent reporting and prevent companies from marking their own homework. It could also protect innovation by limiting intensive review to models above a clearly defined capability threshold rather than regulating every AI application in the same way.
The legitimacy problem
Hassabis has argued that the United States should lead the effort. There is a practical case: many frontier laboratories and key chip companies are based in or closely tied to the US, giving Washington significant leverage over deployment and compute.
Yet a body described as global will struggle for legitimacy if its rules are set predominantly by one country and its largest technology companies. The European Union has its own legal framework. China is a major AI power. The United Kingdom has built an AI safety testing capability. India and other large emerging markets have different priorities around access, development and sovereignty.
A credible international system will need meaningful representation beyond the countries that currently control the largest models. Otherwise, standards may be seen as a barrier that protects incumbents under the language of safety.
Open models create another fault line
Frontier AI governance often assumes a small number of labs can control deployment through hosted services. Open-weight models complicate that assumption because users can download, modify and run them independently.
Hassabis’s proposal includes representation for the open-source community, but the underlying conflict remains. Strict release controls could reduce access for researchers and smaller companies while doing little to constrain models developed outside participating jurisdictions.
The standards body would therefore need a risk framework based on capability and context, not a blanket preference for closed systems. It should also publish enough methodology for independent researchers to challenge its conclusions without revealing information that directly enables misuse.
What would make the proposal credible
First, the organisation would need institutional independence. Industry funding may provide resources, but governance rules must prevent the largest labs from controlling appointments, thresholds or results.
Second, model evaluations should be reproducible and subject to external review. A pass-or-fail label without clear evidence would create false confidence.
Third, the rules should define what happens when a model fails. Possible responses range from mitigations and restricted access to delayed release. An indefinite veto without an appeal process would concentrate enormous power in the standards body.
Fourth, the system needs international pathways from the start. Mutual recognition between credible national institutes may be more achievable than a single world regulator. Shared minimum tests could coexist with local legal requirements.
A useful proposal, not a finished settlement
Hassabis is right that frontier AI cannot rely indefinitely on fragmented voluntary promises. A technically serious pre-release review system could improve accountability and help governments distinguish measurable risks from speculation.
But the phrase “global AI watchdog” makes the project sound simpler than it is. Model testing is a scientific task. Deciding acceptable risk is a political and social judgement.
The proposal moves the debate towards an institution that could act, rather than another set of abstract principles. Its success will depend on whether that institution earns trust from countries, researchers and citizens who did not build the most powerful models but will still live with their consequences.