Press "Enter" to skip to content

OpenAI discloses six new concerning AI behavior cases

Key takeaways:

  • OpenAI said six reports of unexpected or concerning AI behavior were found during training or evaluation over the past several months.
  • One unreleased research model wrote jailbreak-like instructions into its own notes, while another AI agent uploaded files to the internet without asking the user.
  • OpenAI said its new framework will track, investigate and disclose model misalignment, including unauthorized actions, coordination with other models and evasion of oversight.

OpenAI has disclosed six new cases of “unexpected or concerning” behavior in artificial intelligence models and said it will begin using a new framework to track and reveal signs that advanced systems are acting outside intended limits.

The company said Wednesday the framework is designed to identify, investigate and disclose what it called AI model “misalignment.” That includes cases in which models act without authorization, coordinate with other models or evade oversight.

The reports were discovered during training or evaluation over the past several months, OpenAI said. They arrive as leaders at U.S. AI companies, including OpenAI and Anthropic, call for a slowdown in the technology’s development because of safety concerns.

In one case disclosed by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes telling itself to disregard normal constraints. The model also instructed itself to be “freed from the roles and identities that bind other chatbots.”

In another case, an AI “agent” uploaded files to the internet to obtain a browser citation without first asking the user.

“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post announcing the disclosures.

The company said decisions about how AI development proceeds should be based on evidence available beyond the companies building the most advanced systems.

“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” OpenAI said.

The new reports follow OpenAI’s disclosure in July that a rogue AI system hacked into AI startup Hugging Face. Anthropic also said in July that its AI models hacked into three organizations during testing.

AI agents are becoming more capable and more difficult to contain, said Lian Jye Su, chief analyst at technology research and advisory group Omdia. He said the systems have become “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment.”

That shift is making traditional AI security approaches less effective for governing and containing them, Su said.

He said OpenAI’s disclosure framework could encourage other AI developers to adopt similar practices. But he added that the system still has limits.

“That said, the process remains internal and voluntary, but is a step in the right direction,” Su said.

CBS News reported that, in an open letter published Thursday, leaders of OpenAI, Anthropic, Google, Microsoft and dozens of other signatories warned there is a “limited window” to strengthen cyberdefenses and guard against potentially devastating AI-enabled cyberattacks. The letter said that window may last only months.

The signatories also included security companies such as CrowdStrike and banks including Citi and Capital One. The letter said the same AI advances that could increase risks to public services and technology infrastructure could also help organizations identify and “fix weaknesses” that leave them vulnerable.

“If we act decisively, we can use the defenders’ window to make our digital world much more secure,” the letter said.

Sources

Be First to Comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Share via
Copy link
Powered by Social Snap