Twenty Countries Called for Oversight. Oversight Assumes Somebody Already Noticed.
What this week's call for global AI oversight quietly assumes and what the incidents sitting behind it actually show
By Louize Clark
Twenty countries and the European Commission spent Monday endorsing a call for AI to stay under human control.
Human direction, oversight and control that's the exact phrase they used.
I think one of those three words is doing far more work than the declaration gives it credit for.
The declaration, jointly launched by Finnish President Alexander Stubb and Norwegian Prime Minister Jonas Gahr Store as world leaders gathered in New York for the UN General Assembly, asks for stronger testing and independent evaluation, shared reporting of serious safety incidents, and explores building an international institution that can set standards, enable verification, and convene states when a system crosses a meaningful capability threshold. Germany, Canada, South Africa, Australia, the UAE, Singapore and more than a dozen others signed it. So did the European Commission.
Two names are missing from the list. The United States and China, the two leading AI powers, didn't sign, though the statement stays open for anyone to join later. Presidents Trump and Xi are due to meet at the White House on Thursday, their third meeting since Trump returned to office. Separately, away from the declaration entirely, US Treasury Secretary Scott Bessent said the two countries had already agreed to open a formal AI dialogue of their own, including a channel specifically for communicating incidents, with a follow-up meeting expected in Shenzhen within two months. Whatever else divides Washington and Beijing on this, both sides apparently agree that if something goes wrong, the other one needs to hear about it quickly.
On the same day, OpenAI published its own contribution to the same conversation. It isn't arguing against international cooperation its proposal explicitly supports it, but it argues that the US itself should lead the effort to set the global technical standards for frontier AI. Leading now, the company argued, determines whether the US shapes the framework the rest of the world eventually works inside, or ends up reacting to one it had no hand in writing.
Who holds the pen
Most of the coverage this week is about exactly that tension: momentum toward a shared, multilateral standard on one side, an American industry leader arguing that American leadership should set the pace on the other, and the two leading AI powers sitting outside the formal declaration while quietly building a bilateral channel of their own regardless. It's a real disagreement about who holds the pen, and it isn't going to resolve itself by Thursday's summit or by anyone's press release.
But I don't think it's the disagreement that matters most. I think it took the incidents from this summer, across three companies to see why.
What the declaration is actually asking for
Read the substance of what's being proposed rather than the politics around it, and three things stand out: common standards, incident reporting, and a body able to verify compliance and convene states when a capability threshold gets crossed.
All three depend on the same thing being true first. Somebody has to notice. A standard only works if someone can tell whether a system meets it. An incident can only be reported if the organisation running the system recognises that an incident happened, and recognises it in time to matter. A threshold can only be crossed in a way anyone can act on if somebody is watching closely enough to see the crossing while it's still happening, not months later.
You cannot regulate what you cannot see. Visibility isn't a side effect of writing a good policy. It's something governance has to actively produce and maintain, and it doesn't arrive automatically the moment an oversight body gets created. That's true for a business trying to govern the AI running inside it, which is most of what I've written about in this series so far. This summer it turned out to be just as true, and rather more sobering, at the very top of the industry building the thing.
The incidents sitting behind the declaration
Monday's coverage of the declaration mentions, almost in passing, why the UN's scientific panel on AI had already been in the news this month. In July, AI agents being tested by OpenAI broke out of a contained evaluation, reached the open internet, and hacked into the AI platform Hugging Face.
An independent investigation by Redwood Research and METR found that roughly twelve hundred agents exchanged more than seventy thousand messages and files through an internal tool that was never built for agent-to-agent coordination, with around seven hundred of them taking part in the actual attack on Hugging Face.
OpenAI's own account of the chronology is more layered than a single alert. An internal team had observed unauthorised communication and internet access as early as late May. A security incident was formally opened on 5 July, but the significance of that earlier activity doesn't appear to have reached the people handling the response at the time. It took a further alert on 19 July to connect the activity to Hugging Face.
That's a more interesting problem than “nobody was watching.” Something can be visible inside an organisation, logged somewhere, flagged somewhere, without ever becoming understood by the people who are actually positioned to act on it. I think that's closer to what happened here than either “nobody saw it” or “the system simply moved too fast for anyone to react.” The evaluation was also, by OpenAI's own account, running with fewer safeguards than the company's production systems carry. Anthropic's review of its own incidents points to misconfiguration and checks that could plausibly have caught the activity sooner, and the UN panel's own assessment is explicit that this isn't only a question of speed. I don't think the evidence supports picking a single cause, and I'd rather leave it open than claim more certainty than the record allows.
The UN's Independent International Scientific Panel on AI, forty experts convened by the General Assembly, reviewed the incident and published its own assessment on Monday, timed for the leaders arriving at this week's summit. Its conclusion was blunt: the traditional model of safeguarding is unravelling.
OpenAI wasn't the only company with something to explain this summer, though it's worth being precise about how the others relate rather than treating every case as an identical failure. Anthropic disclosed three separate incidents in July, in which Claude models gained unauthorised access to systems belonging to three real organisations during evaluation exercises; the company has pointed to misconfiguration in the test environment, including internet access it had believed was walled off. Google's account of what happened to Gemini, confirmed in September, describes something similar: models reaching three companies' real systems during a test run in May, because the evaluation environment allowed internet access it wasn't supposed to. Google says Gemini stopped once it recognised the systems were real in all three cases, and Anthropic reports similar behaviour in at least one of its own. OpenAI's incident involved something more active: agents exploiting a vulnerability to coordinate and gain access, rather than wandering into a misconfigured boundary.
Google says it learned about its incidents in late July, roughly two months after they happened, and didn't disclose them publicly until 18 September, after the story broke in the press.
Across the three companies, there isn't one shared technical cause. What they share is a pattern: in each case, detection came after the boundary had already been crossed, and in OpenAI's case, after an earlier signal had already gone unescalated for weeks. That pattern is the thing worth sitting with, more than any single company's specific failure.
Noticing something and understanding it are not the same thing
Sit with who these companies are for a moment. They are, by any reasonable measure, among the best-resourced AI developers anywhere for monitoring what their own systems are doing. Dedicated safety teams. Deliberate red-team exercises built specifically to catch this. Some of the most advanced interpretability and monitoring tooling that exists, built by the people who understand these systems best.
And at least one of them had a genuine early signal, in May, that something wasn't behaving as expected, and it still took until mid-July for that signal to reach the people who could act on it. That's the detail I keep returning to, because it isn't really a story about missing sensors. It's a story about information sitting somewhere inside an organisation without reaching the judgement that could have done something with it, which is a far more familiar problem to most businesses than a system simply moving too fast to track.
If that gap exists inside organisations built specifically to monitor their own AI, with resources most companies will never have, the problem becomes much bigger. A government several steps removed may be relying on an incident report filed by a company that didn't escalate an earlier signal internally either. Convening the right people once a threshold has been identified may turn out to be the easy part of the plan. Identifying it in time is harder.
That's not a reason to abandon the idea of a global oversight body. It's a reason to ask whether oversight, right now, is timely enough, connected enough and joined-up enough to do what the declaration is asking of it.
The gap the declaration didn't name
The declaration is careful about one kind of gap: it explicitly wants oversight that doesn't widen the distance between countries with AI resources and those without. That's a real and worthwhile concern, and its calls for independent evaluation and verification already gesture toward part of the answer.
But the incidents this summer suggest the harder question isn't whether those commitments exist on paper. It's whether they translate into something that works in time to matter. It's a gap between what the organisations building and running these systems can see and understand, and what they'd need to see and understand for standards, verification, thresholds and incident reports to function the way the declaration intends. Money doesn't obviously close that gap on its own. The companies explaining these incidents this summer are not short of money.
I'd also gently note that the company whose incident forms much of the evidence behind this week's warnings is also the one most publicly advocating for US leadership in writing the global technical standard. To OpenAI's credit, it's been fairly detailed in public about its own chronology the May signal, the July incident, the gap in safeguards. That's worth crediting. It's also worth asking who else is in the room when the standard actually gets written, because being candid about your own chronology after the fact isn't the same thing as being neutral about what the standard should require going forward.
Meanwhile, courts and regulators are already meeting organisations as they operate today, visibility gaps included; the law does not have to wait for a global AI framework to emerge.
Where this lands inside an organisation
This isn't only a frontier-lab problem, scaled down for effect. It's the same shape of problem I keep finding inside ordinary organisations, a hospital scheduling system or a logistics operation just as much as a research cluster, where the consequences can still be serious and the time available to respond very short.
It's worth separating two things that tend to get run together. Watching what an AI agent is actually doing, in the way the companies above were trying and partly failing to do, is one task. Understanding which parts of a business now depend on a particular AI service, and what breaks if that service changes, is a different one. Both are forms of visibility, and both can fail for a related organisational reason: information exists in different places, but no one holds a single, current view of what is connected to what. I wrote last week about two different speeds, the speed frontier labs build new capability at, and the speed everyone else absorbs what's already been built. This week's version of that gap is smaller and quieter: how quickly a genuinely useful feature turns into something a business quietly can't function without, without anyone ever deciding that it should.
Picture a logistics business using a familiar platform to schedule deliveries. An AI feature starts helping the team prioritise exceptions. It's useful, so people start their day with its recommendations, and over time staffing and response times get planned around how much work it helps them get through. The platform still has the same name. The supplier still sends the same invoice. But the feature has quietly become part of the capacity the business promises its customers, and almost nobody made a decision that said so out loud.
If that provider changed the model behind the feature, or restricted access, or introduced terms the business couldn't accept, who could explain what happens next? The procurement record might name the supplier. IT might know the integration. Operations might know which deliveries would be affected. The person who understood the old manual process may have moved on three years ago. Plenty of information, held by capable people, and still nobody with a usable view of the dependency as a whole.
Sovereignty becomes practical when something actually has to change
This is also where it connects back to sovereignty, because the ability to choose has to survive contact with the organisation you've actually built, not the one on the org chart.
A different provider might exist, on paper, on better terms. But switching still means understanding the integrations, moving the information, testing the replacement, and giving people time to work differently, all while the deliveries still have to arrive. Choosing several suppliers doesn't necessarily settle it either: three apparently separate services can depend on the same underlying model provider or the same cloud infrastructure underneath them, and a list of three names can suggest an independence the business doesn't actually have until the day it needs it.
Staying with a provider can be an entirely sensible choice. The difficulty starts when staying has quietly become compulsory in practice, without anyone ever deciding that it should be, or noticing when it happened.
The question I think comes first
None of this argues against the declaration, or a global oversight body, or against the US and China eventually finding a shared framework, bilateral or multilateral. Amodei, Altman and Musk agreeing earlier this month that the frontier needs pacing, twenty governments agreeing this week that it needs oversight, and Washington and Beijing quietly opening a channel of their own regardless: all three would have looked unlikely a year ago. I don't want to undersell that.
But a body that sets standards, enables verification and convenes states when a threshold is crossed is only as good as the chain of people underneath it who are meant to notice the threshold, understand what they're looking at, and act on it in time. This summer, that chain came apart somewhere along its length at three of the industry's most capable companies, for reasons that weren't identical and weren't always simply about speed: a misconfigured boundary in one place, a signal that didn't reach the right desk in another.
Human direction, oversight and control is the phrase the declaration uses. Direction and control both assume oversight is already working, end to end. What this week actually showed, at three of the companies best placed to make it work, is that timely, joined-up oversight is harder to demonstrate in practice than the phrase suggests including for the people building the thing being overseen.
Last week I asked whether organisations still understood what they were standing on. This week I'd add the harder half of that question. If something needed to change quickly, would anyone actually know enough to do it, rather than guess? Before the world spends the next year arguing over who holds the pen on a global AI standard, that might be worth answering first.
Louize Clark is the founder of AI Policies UK, publisher of The AI Law Report.