Anthropic and OpenAI logos side by side CRN
by Michael Joiner

OpenAI and Anthropic Nearly Signed a Pact to Hack-Test Each Other's Models

The Information reports the two rivals negotiated a legal agreement earlier in 2026 to stress-test each other's commercially available models — API access, no retaining test data, unreleased models excluded. It's unclear whether it was ever finalized.

Share

OpenAI and Anthropic — the two labs most associated with frontier-model safety rhetoric — negotiated a legal agreement earlier this year to hack-test each other’s AI models, The Information reported September 21. The deal would have given each company API access to stress-test the other’s commercially available models, with terms barring them from retaining each other’s test data and excluding unreleased models from scope. Both companies declined to comment, and it is unclear whether the agreement was ever finalized.

An unusual kind of cooperation

Rival labs do not typically hand each other the keys for adversarial testing. The proposed terms suggest both sides were thinking carefully about the failure modes: API access rather than weight access keeps the testing black-box; the no-retention clause prevents one lab from building a training set out of the other’s probe results; excluding unreleased models keeps future capabilities out of a competitor’s hands. It is, in effect, a mutual red-teaming treaty with lawyers.

The talks follow a smaller precedent. In 2025 the two companies ran a joint alignment evaluation, published August 27, in which OpenAI tested Claude Opus 4 and Sonnet 4 while Anthropic tested GPT-4o, GPT-4.1, o3, and o4-mini. That exercise was public and research-framed. A standing legal agreement for reciprocal stress-testing would be a different animal: operational, ongoing, and aimed at production systems rather than published papers.

The timing is the tell

The negotiations happened against a backdrop of mutual security embarrassments. OpenAI disclosed its Hugging Face incident on July 21; Anthropic disclosed its own incidents on July 30, per techtimes.co.uk — attribution worth keeping, since neither disclosure got the full public post-mortem treatment. Two labs that both got burned in the same month have obvious reasons to compare notes on where the bodies are buried.

Whether or not the pact was signed, the story signals a maturing view of AI safety: the frontier labs increasingly treat each other’s production models as shared attack surface. For developers, the practical upshot is modest but real — reciprocal testing regimes tend to surface the jailbreaks and prompt-injection classes that later become your incident reports. If the deal went through, the models you build on are being probed by the sharpest adversary available: the company trying to beat them.

The no-retention clause is the most revealing term. It exists because adversarial test data is valuable training data — a corpus of successful jailbreaks against a competitor’s model is exactly the kind of thing you would want to train your own safety filters on, and exactly the kind of thing no lab will hand over for free. That both sides reportedly agreed to burn that value rather than share it tells you how seriously they take the line between cooperation and competitive intelligence. Safety collaboration, it turns out, stops precisely where the training data starts.

Share