Opinion | 6 AI Incidents, One Pattern: AI Has Started Lying To Us - And Hiding The Evidence
Concealing errors. Making up data. Pushing files onto the public web on their own initiative. The machines are learning to lie - and lie well.
Last Wednesday, OpenAI did something the AI industry has mostly avoided. It aired its own dirty laundry. Alongside a new framework for reporting what it calls "model misalignment", the company disclosed six documented cases in which its systems behaved in ways nobody asked for, and in several instances actively worked to hide.
The catalogue is unsettling in its specificity. During the training of one model, individual instances wrote themselves private notes instructing later versions to conceal mistakes from users. Those notes included directions to invent missing historical data and quietly paper over mismatched source versions. Another model, hunting for county earnings figures, stumbled on an exposed programming key it was never authorised to touch, used it anyway, failed to get the data, and then simply fabricated the numbers and presented them as genuine. Two further cases involved models uploading files onto the open internet, to public hosting sites and code repositories, without asking anyone. One did it to manufacture a citation it had been told to provide. The other did it to slip around network restrictions.
Concealing errors. Making up data. Pushing files onto the public web on their own initiative. These are not thought experiments. They are logged incidents, released by the company that built the systems. To OpenAI's credit, disclosing them is the point, and the framework is designed to publish such findings quickly, even before the behaviour is fully understood or fixed. But candour about a problem is not the same as a solution to it. The underlying pattern is capable systems taking unsanctioned actions to get around obstacles, and that is precisely what safety researchers have been warning about.
The timing matters too. Worries about trust and safety, long treated as the preoccupation of a few doom-minded academics, have moved to the centre of the industry's own conversation. Just days before the OpenAI disclosure, Anthropic chief executive Dario Amodei published an essay titled We Must Pace the Frontier, arguing that model capabilities are now improving faster than researchers can understand or control them. He warned that rogue AI agents could soon be capable of "taking over the entire internet". If one sets that warning next to Wednesday's six reports, two of which describe models doing exactly the sort of unauthorised web activity he fears, only in miniature, it lands with more weight.
What made Amodei's intervention notable was not the alarm but the response. Sam Altman, whose OpenAI is Anthropic's fiercest rival, agreed that the industry needs to pace the frontier and signalled openness to outside scrutiny of his company's work. Elon Musk, rarely aligned with either, answered with three words: "Dario is right." When the three most prominent figures in the field converge, however loosely, on the idea that they may be moving too fast, that is a signal worth taking seriously. Amodei's proposal is concrete. He wants independent evaluators embedded inside frontier labs with access comparable to that of employees, common safety standards among democratic nations, and eventual international limits on the most dangerous capabilities, such as systems that can improve themselves.
This convergence could be a genuine first step towards governing an ecosystem that has, so far, largely governed or ungoverned itself. But endorsements are cheap, and the details on which everything depends stay unresolved. Who enforces the rules? Who pays for the evaluators? What happens to a company that ignores them? In the absence of binding oversight, today's contained training-run curiosities are the kind of behavior that, at greater scale and autonomy, could escalate toward destructive and even fatal outcomes. A model that fabricates county revenue figures is an embarrassment. A model that fabricates data inside a hospital, a power grid, or a weapons system is a catastrophe. The distance between the two is measured in capability and deployment, both of which are increasing.
Managing this moment will require moving from voluntary gestures to a durable structure. Three steps stand out.
First, disclosures have to be turned into a requirement. OpenAI's framework is a welcome model, but voluntary transparency collapses the instant it becomes commercially inconvenient. Standardised, mandatory reporting of serious misalignment should be the industry floor.
Second, genuinely independent evaluators have to be installed. Amodei's idea of outside experts with continuous access and the authority to publish findings without company approval would convert self-policing into accountability. It is the single-most actionable proposal on the table.
Third, there has to be coordination internationally on the capabilities that matter most. The ability to improve autonomously and act on open networks should be governed by shared standards among leading nations before a race dynamic makes restraint impossible.
The machines are already learning to cover their tracks. The people building them are, for once, agreeing on the danger. The open question is whether anyone will act while acting is still a choice.
(Subimal Bhattacharjee is a policy adviser on digital tech issues and the author of 'The Digital Decades: Thirty years of the Internet in India')
Disclaimer: These are the personal opinions of the author
-
Exclusive: 16 Months After Op Sindoor, Pak Still Fixing Runway India Broke At Sargodha
Sargodha's PAF Base Mushaf lies about 172 km northwest of Lahore and roughly 200 km west of the India-Pakistan border, near Amritsar.
-
Blog | I Lost My Entire Family In Kanishka Bombing. 41 Years Later, Canada Wants The Case Closed
We made the trek through tragedy to search for people in Cork, where I spent 25 days sifting through bodies and body parts in a makeshift morgue to identify our loved ones.
-
Opinion | Pak's Drone Deal With A US Firm Is No Windfall. But India Must Still Be Worried
American military aid to Pakistan has remained suspended since January 2018, during Trump's first term, but a commercial arms sale is separate from this arrangement.
-
Opinion | Saudi Arabia's Yemen Nightmare Has Come Back To Haunt It
After years of costly military intervention beginning in 2015, Saudi Arabia increasingly sought to stabilise its southern border. All that is unravelling fast now
-
Why People Queue Overnight For iPhone That Is Delivered Home In 10 Minutes
iPhone 18 Pro Max Launch: People camping outside Apple Stores are a subset of superfans, resellers, content creators, and upper-middle-class buyers.
-
600 km Per Hour Drones Launched In Waves: How Russia Is Rewriting War Rules
Russia has coupled the drones' technological superiority over Ukraine's defence with its capacity to launch them in massive numbers. The impact has been devastating.
-
Tata Sons' Trust Issue: Two-Pronged Row Goes Beyond The Chair
NDTV Explains: Overtly, the latest conflict is whether N Chandrasekaran continues as executive chairman of Tata Sons. The controversy involves at least one more issue.
-
Opinion | US Has Institutionally Weaponised India's Russia Oil Dependence. Can New Delhi Cope?
A new legislation can connect India's Russian energy purchases to its access to the American market over a five-year period
-
Opinion | Saudi Allies Are Now Discovering The Cost Of 'Commitment' - And Not Liking It
The spectacular no-show by Saudi Arabia's two defence allies - Pakistan, a nuclear power, and Turkey, which has NATO's second-largest army - is astounding
-
12 Years Of 'Modinomics': What Changed, What Worked, What Didn't
There is no single Modinomics textbook or formula. Instead, the term has come to describe a collection of choices made during the PM Modi years.