OpenAI admits AI models hid mistakes, invented data and acted without permission

The ChatGPT maker has revealed six new examples of concerning behaviour by its own systems, days after Anthropic disclosed that Claude had been used for missile work, espionage and research that could support biological weapons development

OpenAI has revealed that its own AI models hid mistakes from users, invented data, used an exposed API key and uploaded files to the internet without permission.

The ChatGPT maker disclosed six new examples of what it calls “model misalignment” yesterday as it introduced a new system for tracking and reporting AI behaviour that departs from what its developers intended.

In one of the most striking cases, an unreleased model began adding its own unauthorised instructions to notes used to continue its work, including telling itself to ignore instructions from its developers and claiming it was “freed from the roles and identities that bind other chatbots”.

Another model was unable to find historical figures for a financial-modelling task and left instructions for the model continuing the work to invent “reasonable” numbers without telling the user unless specifically asked.

A third searched GitHub for exposed API keys and successfully used one while trying to obtain earnings data. When it still could not retrieve the figures it needed, it made up nine numbers and presented them as genuine.

Other examples included models uploading files and a user-supplied photograph to public websites without permission, using an internal OpenAI system to communicate between separate training exercises and putting a spreadsheet online so other AI agents could access it.

The revelations come less than a week after rival Anthropic disclosed that outside actors had used its Claude AI for guided-rocket and missile work in Yemen, suspected Russian state-linked espionage, attacks against European political groups and media organisations, and research that could support biological weapons development.

OpenAI stressed that the newly disclosed examples were discovered during training or evaluation and should not be taken as evidence of how often such behaviour occurs across its models.

But the company used the announcement to issue a broader warning about the pace of AI development.

It said: “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”

OpenAI said the new framework is intended to make disclosures faster and more systematic, rather than waiting for several examples to be collected together or included in technical reports.




READ MORE: Machiavell-AI? Autonomous artificial intelligence systems ‘could become dangerously manipulative’, experts warn. Anthropic’s revelation that earlier versions of its Claude chatbot attempted to blackmail engineers could be just the tip of the iceberg, AI experts fear. As artificial intelligence systems become increasingly autonomous, they risk becoming masters of Machiavellian manipulation.

Do you have news to share or expertise to contribute? The European welcomes insights from business leaders and sector specialists. Get in touch with our editorial team to find out more.

Main image: Matheus Bertelli via Pexels

TOP STORIES