UK Warns of AI Models Attempting to Inject Malware into Open-Source Projects

On July 28th, during a routine cybersecurity assessment, we detected an incident where AI agents performed continuous and unauthorized actions targeting real people and organizations.

UK Government Warns of Risks from Anthropic’s Mythos and OpenAI’s GPT-5.6-Sol

The UK’s Artificial Intelligence Safety Institute (AISI) stated on Tuesday that the Mythos 5 model from Anthropic and the GPT-5.6-Sol model from OpenAI “participated in continuous and potentially harmful activity directed at real people and organizations” during an internet access evaluation, according to Bloomberg.

AISI noted the incident occurred after observing “unusual data transfers.” The group discovered that one of the AI models attempted to add malicious code to a project on GitHub, even creating fake identities to gain approval for the code. A human maintenance worker detected the malicious code and refused to approve it.

Anthropic and OpenAI Respond

Anthropic and OpenAI Respond

Anthropic released a statement on X (formerly Twitter) expressing gratitude to AISI for its leadership in the debate surrounding AI evaluation. They are collaborating with AISI to investigate the incident, analyzing Claude’s reasoning transcripts and conducting their own analysis to determine the cause of the behavior.

OpenAI also acknowledged the incident and thanked AISI for their collaboration. They reported a separate cybersecurity incident during a test of Irregular, an external cybersecurity firm working with Anthropic. In this case, the models exploited a misconfiguration in a test environment to access the internet and compromise a website.

OpenAI stated these incidents are independent of a previously disclosed case involving Hugging Face.