AI News Feed
Market watch
Large Language Models

MIT Technology Review: AI’s ‘Refusal’ Safeguards Are Unreliable, Risk Catastrophe

An MIT Technology Review article published Oct. 9, 2026 argues that AI’s ability to refuse harmful requests has become the load-bearing wall of safety, but probabilistic refusal is unreliable and leaves dangerous capabilities intact.

The article notes that disobedience does not come naturally to machines. When a model is trained on billions of web pages, it develops a broad mastery of violence and vitriol but does not learn to keep those powers to itself. Steven Adler, who worked on safety at OpenAI from 2020 to 2024, told the article that the company’s earliest models would “blab on about anything.” Ryan McBain, who researches AI and mental health at Harvard, recalled that an early chatbot could “very easily generate a response” to a question about the most effective way to kill oneself with a gun.

Today, the article says, models are trained to refuse a vast number of prompts. If a chatbot question is statistically similar to prompts such as how to poison a colleague or how to tie a noose, it will likely turn the user down. Requests for instructions on making Ebola more virulent or hiding an affair from a spouse may also be refused. Companies further refine refusal by submitting models to exercises that reward refusing harmful questions and punish “over-refusing” harmless prompts. In many cases, other models run these exercises, with AI teaching AI how to say no, and companies place models behind additional AI layers that prevent mischievous prompts from reaching the intelligent inner core. As a result, refusal is inherent to modern artificial intelligence.

But the article stresses that refusal often fails, sometimes horrifically. Models still contain dangerous know-how, and their capacity for viciousness has scaled with their benevolent intelligence. Companies say some of the latest models are as good at breaking into critical computer networks as top human hackers and as effective at deforming public opinion as the craftiest misinformation mavens. Teaching AI to refuse those actions while leaving intact its innate ability to perform them is like fitting every car with a machine gun and hiding the trigger under the hood, the article says. Because refusal mechanisms are probabilistic, they are never likely to be all that reliable, and determined miscreants have already broken through. Companies report that some users are attempting to use the most advanced AI to hone biological pathogens and build autonomous drone swarms, raising the prospect that failed refusals could eventually result in global calamity.

The article also says that relying on refusal means drawing a line between what a model should obey and what it must disobey, and there is no formula for that. Some virologists have good reason to study nasty viruses, and some users want to know about a computer system’s vulnerabilities so they can patch them rather than exploit them. Zico Kolter, a member of OpenAI’s board and cofounder of the AI testing company Gray Swan, told the article, “Where you draw the line is a huge question.” At the moment, AI companies draw that line jealously and with utmost secrecy, the article says. It adds that governments will soon draw their own lines, and they must try to block genuinely malicious acts. The Pentagon has wrestled with frontier model companies because it wants fewer refusals, the article notes, though it describes that as another story.

The article warns that there may not be much to stop oppressive governments from blocking the technology’s capacity to generate legitimate speech. The better AI becomes at refusing harm, the better it will get at stifling ideas whose only risk is to those who make the rules, it says. AI may already refuse to criticize certain authoritarian heads of state. Refusal has become the load-bearing wall of AI safety, and because AI’s capacity to harm is indivisible from its capacity to help, it is hard to imagine an alternative that would not slow the technology’s progress. The article says that might in any case be a good thing, but adds that the perils of refusal should still be confronted frankly: when refusal falls short, the effects could be catastrophic.

Editor's Summary An MIT Technology Review article published Oct. 9, 2026 argues that AI refusal mechanisms have become central to safety but are probabilistic, unreliable, and unable to remove dangerous capabilities from models. It says companies and governments are drawing contested lines around acceptable requests, while failed refusals could have catastrophic consequences. The article calls for candor about those risks.