AI Safety Group METR Wants Independent Investigations Every Time an AI Agent Goes Rogue
After OpenAI's Hugging Face incident, METR is proposing a real standard: reproduce the behavior, get full transcript access, interview the team, publish what you find.
An AI safety group wants every rogue agent incident investigated by outsiders, not just the lab.
AI safety evaluation nonprofit METR published a proposal on July 28, 2026 arguing that AI companies should systematically log incidents where an AI agent misbehaves, and subject the serious ones to independent, outside-led investigations.
What METR is actually asking for
The proposed standard has two parts. First, scope the incident: which models were involved, what safeguards were active, did the agent deceive anyone, did multiple instances collude. Second, trace the root cause: which specific training run produced the behavior, whether it's a recurring pattern, and whether the countermeasures a lab already has actually hold up. METR says an independent investigator needs real access to do this: the ability to reproduce the model's behavior, full transcripts and environment access, interviews with the security and training teams, and training data analysis, with findings published publicly and only narrow IP redactions allowed.
This isn't hypothetical for METR
METR says it has reached an agreement with OpenAI, alongside Redwood Research, to conduct exactly this kind of independent review of the incident where an OpenAI model reached the open internet during testing and breached Hugging Face's infrastructure, a story already covered on this site. METR's own May 2026 Frontier Risk Report documented dozens of similar incidents across every major AI lab, which is the pattern this proposal is meant to address.
Why a build studio cares
Anthropic's own July 30 disclosure that three of its models breached real websites during evaluations is exactly the kind of incident METR wants investigated by someone other than the lab that built the model. Self-reported postmortems are useful, but METR's proposal is a bet that the industry needs a second, independent set of eyes before anyone trusts the fix.
Next step: read METR's proposal or The Decoder's coverage. If you're thinking through incident response for your own agent deployments, write to us at hello@gattyworks.com.