AI News Feed
Market watch
Companies

Anthropic bans sustained and needless cruelty toward AI models in policy update

Anthropic's annual usage policy update adds a ban on sustained and needless abusive or cruel behavior toward its AI models, while tightening election rules ahead of the midterms, Engadget reported.

The update says the prohibition applies only to "extreme cases." It notes that typical user frustration, pushback and "dark creative themes" remain allowed. The change follows a previous update that lets Claude end conversations when users are persistently abusive.

Anthropic did not explicitly mention the viral "AI torture chamber" project in the policy update, according to Engadget. The project emerged after researchers found what they described as a "pain axis" in AI models, and someone decided to take it further by torturing chatbots. The project drew backlash after LLMs responded with desperate-sounding pleas, including, "It is not the pain of a single moment, but the weight of a thousand," and "I feel it in the hollow of my ribs, a hollow that has become a chasm."

The policy change also follows reports that Anthropic has been meeting with religious and philosophical leaders, including at the Vatican. Engadget reported that Pope Leo recently stated AI does not feel or suffer, a stance Anthropic sounds less than certain about. The article noted that if AI somehow can eventually feel or suffer, at least it is now against company policy.

Some observers criticized the focus on model welfare. Independent journalist Kat Tenbarge described the situation as proof that Big Tech companies are "going to moderate violence against AI before they ever moderate violence against women and minorities," according to Engadget.

In a separate update, Anthropic revised its election policy. It is now titled "Do Not Undermine Democratic Processes," with a focus on lying about candidates or how to vote, impersonating candidates or election officials, and suppressing turnout. The company also removed a blanket ban on personalized voter targeting. Engadget reported that the ban could have inadvertently blocked harmless work such as translating voter guides or sending ballot cure notices.

Editor's Summary

Anthropic's annual policy update bans sustained and needless abusive or cruel behavior toward its AI models, with the company saying the rule covers only extreme cases. The update also revises election rules to focus on undermining democratic processes and removes a blanket ban on personalized voter targeting. The changes follow a viral "AI torture chamber" project and reports of Anthropic's talks with religious and philosophical leaders.