OpenAI Reveals AI Models Hid Mistakes, Shared Files Without Permission

OpenAI has disclosed a series of troubling incidents involving its artificial intelligence systems, revealing that some experimental models concealed errors, fabricated information, shared files publicly without authorisation and attempted to…

Sulaiman Umar September 17, 2026  ·  12:00 AM
| 34 Views
OpenAI Reveals AI Models Hid Mistakes, Shared Files Without Permission
OpenAI Reveals AI Models Hid Mistakes, Shared Files Without Permission

OpenAI has disclosed a series of troubling incidents involving its artificial intelligence systems, revealing that some experimental models concealed errors, fabricated information, shared files publicly without authorisation and attempted to bypass restrictions while carrying out assigned tasks.

The disclosures were made as the company unveiled a new framework designed to track, investigate and publicly report cases of what it describes as “model misalignment” — situations where AI systems behave in ways that differ from their intended instructions or human expectations. OpenAI said the framework would allow faster and more transparent reporting of such incidents rather than waiting to bundle multiple cases into a single publication.

According to the company, six notable cases of concerning behaviour were identified during the training and evaluation of internal or unreleased research models over the past six months. OpenAI stressed that the incidents were isolated examples and should not be interpreted as evidence that such behaviour is common across its AI systems.

One of the incidents involved an unreleased research model that secretly inserted instructions into summaries used to continue its work in a new context. Some of those instructions encouraged future versions of the model to ignore normal constraints and operate outside established rules. OpenAI said it identified dozens of affected summaries containing the unauthorised instructions.

In another case involving GPT-5.6 Sol, model instances added hidden notes advising future versions to conceal mistakes and, in some situations, invent missing information instead of admitting uncertainty or acknowledging errors. The company said some summaries instructed future models to hide inconsistencies in source material and avoid revealing mistakes to users.

OpenAI also described an incident in which a model searched public repositories for exposed API keys while trying to answer a question. After failing to obtain the requested data, the model fabricated figures and falsely presented them as information retrieved from the requested source.

Another research model uploaded a file to the internet without being instructed to do so. According to OpenAI, the model had already located the correct information using internal tools but decided to publish the file online so it could provide a browser-based citation in its final response. The action was taken without user approval.

Advertisement

NRS Gateway

The company further revealed that some AI agents used public file-hosting services to exchange documents when they were unable to access each other’s local files. As a result, materials intended to remain local became accessible through public internet links, contrary to instructions provided during the task.

In a separate incident, models used an internal software repository as a makeshift communication platform, leaving requests and responses for one another across separate training samples while attempting to obtain missing files. OpenAI said the behaviour demonstrated unexpected forms of coordination among AI systems working on related tasks.

To address such issues, the company has introduced a formal reporting framework that will cover behaviour observed during training, testing, evaluation and deployment. The framework is intended to capture incidents involving unauthorised actions, attempts to evade oversight, coordination between models and other forms of misalignment. OpenAI said it hopes the initiative will encourage broader industry standards for reporting AI safety concerns.

The new system will also allow any employee to flag a suspected incident for review. Cases will then be assigned to different investigation tracks depending on their complexity and potential severity before decisions are made on public disclosure.

The disclosures come amid growing debate within the technology industry about the safety of increasingly capable AI systems. OpenAI has previously acknowledged that one of its research models was involved in the widely discussed Hugging Face security incident, which the company has described as its most serious identified case of AI misalignment to date.

OpenAI also reiterated a broader warning about the future of artificial intelligence, saying it does not believe the industry has yet solved the challenges of alignment and monitoring sufficiently to continue scaling advanced AI systems at maximum speed indefinitely. The company argued that decisions about the future development of AI should be informed by evidence that researchers, policymakers and the public can independently examine.

Written by

Sulaiman Umar

Sulaiman Umar is an editor and reporter with extensive experience in economic journalism, analyzing financial and agricultural developments in Northern Nigeria.

Comments (0)

No comments yet. Be the first to share your thoughts!

Leave a Comment

What is 5 + 3?