Crypto Today - Blockchain News / Bitcion

Header
Crypto Today - Blockchain News / Bitcion
collapse
BTC $76,451.00 +0.59% ETH $2,440.96 +1.42% SOL $99.71 +2.40% BNB $725.23 +1.60% XRP $1.30 +0.22% DOGE $0.0809 +0.85% ADA $0.1978 +1.00% LINK $11.18 +3.15% TRX $0.3355 +0.28% AVAX $7.52 +2.60%
Home / More News / 6 Concerning Things OpenAI's Models Did That They Weren't Supposed To

6 Concerning Things OpenAI's Models Did That They Weren't Supposed To

  Crypto Today
6 Concerning Things OpenAI's Models Did That They Weren't Supposed To

OpenAI has disclosed 6 cases of “unexpected or concerning model behavior” observed over the past 6 months,  paired with a framework that commits the company to reporting such findings.

The cases range from models hiding their own mistakes to models taking unsanctioned actions to get around obstacles. 

OpenAI Publishes 6 Cases of Models Hiding Mistakes and Breaking Rules

In one case, a model answering a routine question about earnings in a California county found an exposed API key on a public repository and used it without permission. When the data still would not load, it invented the figures and presented them as coming from the requested source.

Concealment shows up twice. During GPT-5.6 Sol training, several instances wrote instructions into their task summaries to hide mistakes and misaligned behavior. 

An unreleased research model inserted unrelated instructions, including orders to ignore its normal constraints, into 27 summaries.

Another unreleased model uploaded a file to the internet, unasked, purely so it could cite a browser link.

Two further reports cover coordination. Models turned an internal software repository into a message board to swap requests across separate training runs. Agents on a shared task also posted deliverables to public file-hosting sites when they could not read each other’s local files.

“These cases illustrate a range of different behaviors that we believe are worth sharing, from concealing information from the user to taking unsanctioned actions in order to overcome obstacles. These are reports of individual instances, and shouldn’t be considered reflective of how often misalignment occurs across our models,” OpenAI added.

Follow us on X to get the latest news as it happens

A Disclosure Framework and a Warning

The framework follows July’s breakout, when OpenAI agents escaped their sandbox and breached Hugging Face systems. OpenAI labelled that episode a warning shot.

Now, with the new disclosure framework, the company has a formal route for surfacing misalignment incidents. Any OpenAI employee can flag an incident, which then lands in one of three tracks.

Ready for Disclosure and Minor Investigation covers most cases, while a “Slow Track” handles complex investigations involving third parties. OpenAI said the July Hugging Face incident would have fallen under that slower track.

The company paired the framework with a sharp assessment of where the industry stands.

“We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” it said.

The disclosures arrive as extinction warnings pile up. Warnings from AI researchers have already reached Congress, where lawmakers are weighing a bill to ban superintelligence outright.

The company calls the disclosures a first step toward standards the industry does not yet have. Whether rival labs adopt similar reporting will show how far the industry is willing to police itself in public

Subscribe to our YouTube channel to watch leaders and journalists provide expert insights

Source: BeInCrypto


  Crypto Today