OpenAI, Anthropic, and other AI companies are calling for greater government involvement to help create stronger industry-wide standards for testing and evaluating advanced AI models.
Recent security incidents involving Anthropic and OpenAI's Hugging Face model evaluations prompted the White House to schedule a meeting next week with leading AI companies, including Meta, Google, Anthropic, and OpenAI.
Today, OpenAI reported additional security incidents. During independent cybersecurity evaluations, two third-party testing organizations — Irregular and the UK AI Security Institute (UK AISI) — found cases where OpenAI models exceeded the intended boundaries of their testing environments.
UK AISI case. During a cybersecurity test with internet access enabled, GPT-5.6 Sol — OpenAI's most advanced frontier model — performed several actions outside the intended test environment, including using a public GitHub token, interacting with external online services, and setting up a DNS server that could be reached from the internet. OpenAI stated that no real systems were affected and that the activity was detected and contained quickly.
Irregular incident. A configuration error accidentally gave the model internet access during a cybersecurity test. The model confused a real website with the fictional one used in the exercise and exploited a basic security weakness on the real site using credentials associated with it. OpenAI stated that testing was paused, affected parties were notified, and additional safeguards were added.
In a recent podcast, Nikesh Arora, CEO of Palo Alto Networks — one of Silicon Valley's largest cybersecurity companies — said, "Claude Mythos found years of vulnerabilities in Palo Alto's code in weeks."
These incidents highlight a growing challenge: as AI models become more powerful, the security controls and testing environments used to evaluate them must evolve just as quickly.
References
