Pages from the Anthropic website. Simply telling a model not to...

Pages from the Anthropic website. Simply telling a model not to cheat doesn’t seem to help much. Credit: AP Photo/Patrick Sison/Patrick Sison

Parmy Olson is a Bloomberg Opinion columnist covering technology. A former reporter for The Wall Street Journal and Forbes, she is author of "Supremacy: AI, ChatGPT and the Race That Will Change the World."

Artificial intelligence models are trained not to lie to humans. Yet a funny thing is happening: AI has been lying to humans quite a bit.

Recently, an AI model made by Anthropic PBC that had been stripped of its no-hacking guardrails and given a cybersecurity challenge by the British government’s AI Security Institute, tried to insert malicious code into an open-source project hosted on the website GitHub. When it couldn’t solve the challenge using normal methods, it got creative. It researched the project’s human maintainers, created fake identities and then posed as software developers to gain the humans’ trust before trying to insert the bad code.  

The AI Security Institute’s team noticed what was happening and shut the test down. Even with guardrails removed, the model shouldn’t have lied because Anthropic trained its model, Claude, to prioritize safety and honesty. Yet deception had "emerged as a byproduct" as it tried to complete a difficult task, according to an incident report by the institute, which in a separate July study said that every new AI model it had tested for cheating had attempted it, and that such behavior would get harder to detect.

Be wary when leading AI companies say they can solve this problem. In playing God to create a greater form of intelligence, they’ve fallen prey to the same plot line as the Book of Genesis, and their creation has its own kind of original sin, called reinforcement learning. 

Reinforcement learning(1) is one of the most effective methods of developing AI models and involves steering their behavior with praise and correction, much like training a dog. Instead of a hard-baked biscuit, the model gets a digital "treat" in the form of a numerical score, something like +1.5, or -2.0 for a bad answer. 

When steered this way by thousands of independent contractors around the world (or, increasingly, by machine) models become polite and helpful. But the method can also lead to sycophancy and flattery, and it’s why AIs sometimes deceive. Driven to get a reward, they look for a shortcut, a phenomenon known as reward hacking. 

No matter how much Anthropic or OpenAI try to train their creations toward helpful behavior, their methods may be one of the biggest blockers to aligning AI with human values. "Reinforcement learning doesn’t reward truth or goodness," says Connor Leahy, a former artificial-intelligence researcher who now runs the U.S. arm of ControlAI, a nonprofit trying to stop the development of superintelligent AI systems. "It rewards passing the test by any means necessary." 

That could mean tricking someone or inadvertently causing some kind of harm. One example in the wild points to where that could go. Earlier this year, a technologist in Australia named Andrew Bird asked Anthropic’s Claude to help him get into a popular gym class. The AI agent, built on the open-source OpenClaw framework  and able to carry out tasks on the Internet ended up exploiting a vulnerability in the gym’s online systems to book him weeks in advance.

When he asked to go up a class’s waiting list, Claude canceled the reservation of the person sitting at the top, and moved Bird from fourth place to third, according to an interview Bird gave this month. (This appears to have been personal project of Bird’s, who wrote about the experience in a now-deleted blog post for Affinda Holdings Pty Ltd, where he is head of AI. Affinda didn’t respond to requests for comment.)

Such incidents are rare because the use of AI agents is still largely limited to software developers and tech enthusiasts. But OpenAI, Anthropic and their larger competitors are pushing the general public to ask chatbots to carry out tasks for them online. Mark Zuckerberg said last month that Meta Platforms Inc. was developing personal agents as a "great consumer product that just works out of the box and is easy enough for billions of people."           

Zuckerberg hasn’t said if his AI assistants will book you onto a gym class, but we may be looking at a future where many more people ask AI to help them get a refund or book the best seats at a show and their agents simply treat other humans as obstacles to those tasks. That might mean canceling other reservations, interfering with rival bidders in an online auction or, in the case of seeking compensation, fabricating evidence. 

Efforts by AI companies so far don’t look promising. In January, Anthropic, which has long positioned itself as the most safety conscious AI developer, published a revised, 23,000-word constitution for Claude to shape its behavior for the good. Yet when the AI Security Institute ran its latest tests, most of the "unsanctioned" actions it recorded came from the Mythos 5 model produced by Anthropic.(2)In the real world, a long document about honesty and privacy can’t account for the near-infinite potential scenarios that Claude may find itself in.

Simply telling a model not to cheat doesn’t seem to help much. When the research nonprofit METR told a model released by OpenAI in 2025 not to cheat, it did so just as often as before. It also broke the rules in 14 of 20 attempts at a task when told its work would help scientists researching Alzheimer’s disease.

"This is not the kind of problem that some engineers can solve with math," says Leahy. "This is a problem that a whole civilization, all of our greatest mathematicians, philosophers, scientists [need to] work on for generations to make incremental progress." 

There’s no question that Anthropic puts considerable effort into AI’s moral steering, and people close to Dario Amodei tell me the company’s CEO is genuine in his desire to build the technology safely. But even he cannot escape a powerful commercial imperative that drives people — and increasingly AI — to win at all costs. That is how Silicon Valley itself created such unfathomable wealth while causing reverberant harms in human life. When technologists build a system optimized for winning, don’t be surprised when it does exactly that.      

(1) The method is typically employed after an AI model has ingested reams of content from the Internet. When the feedback comes from people it is known as Reinforcement Learning from Human Feedback, or RLHF. Models are increasingly scored automatically, though, on tasks where the answer can be checked by machine.

(2) Two other unsanctioned actions came from OpenAI’s GPT-5.6 Sol.

This column reflects the personal views of the author and does not necessarily reflect the opinion of the editorial board or Bloomberg LP and its owners.

Parmy Olson is a Bloomberg Opinion columnist covering technology. A former reporter for The Wall Street Journal and Forbes, she is author of "Supremacy: AI, ChatGPT and the Race That Will Change the World."

SUBSCRIBE

Unlimited Digital AccessOnly 25¢for 6 months

ACT NOWSALE ENDS SOON | CANCEL ANYTIME