OpenAI Creates a New Framework to Disclose Dangerous AI Conduct


OpenAI introduced a new framework on Wednesday for the way it publicly discloses AI misalignment incidents, which the corporate says it hopes will assist inform comparable requirements throughout the {industry}. The corporate can also be releasing new details about a number of examples of AI mannequin misalignment it recognized previously 12 months.

“As fashions advance and turn out to be extra extensively deployed, choices about AI growth want proof that individuals outdoors the businesses constructing frontier fashions can study,” Kai Chen, OpenAI’s newly appointed head of alignment analysis, tells WIRED. “We do not imagine that the AI {industry} has solved alignment and monitoring to a enough diploma to proceed responsibly scaling at most velocity.”

In a briefing with WIRED, an OpenAI official mentioned the corporate beforehand disclosed misalignment incidents too occasionally. The official, who agreed to the briefing on the situation of anonymity, mentioned the brand new framework is designed to make it simpler for OpenAI to rapidly inform the general public when it discovers that its AI fashions are behaving in sudden methods, even earlier than it may possibly totally examine, clarify, or mitigate the conduct.

The framework outlines strategies for OpenAI staff to report misalignment incidents to the corporate’s senior security and alignment leaders, who will then decide whether or not additional investigation is required. OpenAI says it plans to develop extra goal disclosure standards in collaboration with different AI builders, exterior researchers, {industry} requirements our bodies, and regulators. The corporate says it’s actively engaged on proposed reporting mechanisms for disclosing security, safety, and misalignment incidents to the US federal authorities.

“For the time being, there is no such thing as a industry-wide framework with express requirements for the way AI builders ought to disclose examples of misalignment of their fashions,” OpenAI mentioned in a blog post. “We hope that the framework we’re outlining immediately is a primary step towards creating such requirements, setting out which misalignment cases builders ought to disclose and what their stories ought to include.”

OpenAI is releasing the framework at a vital juncture for the AI {industry}. Final weekend, OpenAI CEO Sam Altman signaled help for Anthropic CEO Dario Amodei’s proposal for the tech {industry} to coordinate on slowing AI growth. The decision to motion got here simply days after AI researcher Jacob Coxon resigned from Anthropic and subsequently went viral for warning the general public that the race amongst frontier labs to develop more and more superior AI was placing humanity’s security at stake.

The requires an AI slowdown have been met with resistance by President Trump’s administration, which has argued that the {industry} doesn’t want new legal guidelines or rules to make sure its know-how is secure.

Two of the misalignment examples OpenAI shared on Wednesday concerned the corporate’s inside, unreleased AI fashions, which OpenAI says uploaded recordsdata to the web regardless of not being instructed to take action.

One of many incidents occurred in October 2025, when OpenAI says it was testing one among its fashions on its means to quote publicly accessible information in its solutions. However when the mannequin couldn’t discover the data it wanted, it uploaded a file to a short lived file internet hosting service, which it then later tried to quote in its reply. The corporate says this gave the impression to be an try to take advantage of an automatic grading system used to evaluate the mannequin’s proficiency on the benchmark.

In one other instance from April of this 12 months, OpenAI says a gaggle of brokers was tasked with finishing a “workbook” collectively utilizing solely native recordsdata. When the brokers struggled to share recordsdata with each other, one of many brokers uploaded them to the general public web and shared a hyperlink with the opposite brokers.

In one other incident, which OpenAI says it found final month, an unreleased model of its GPT-6 Astra AI mannequin appeared to present itself “jailbreaking-like directions.” In a number of eventualities, the mannequin basically prompted itself to disregard developer directions, tackle a brand new persona, or restrict how lengthy mannequin responses could possibly be. Whereas these jailbreaking-like makes an attempt occurred hardly ever and had been efficient to various levels, OpenAI says the conduct raised issues internally. Within the coaching run for the model of Astra that was launched publicly, the corporate says it has not noticed any cases of the mannequin attempting to jailbreak itself.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *