Undercurrent CapitalbetaSign in

Anthropic bans users from bullying Claude

· 4 min read

TLDR — AI giant Anthropic updated its usage policy on Thursday, its first revision in more than a year, instituting a new ban on cruel treatment and abuse of its AI tool Claude.

Signal data▼ Bearishconfidence 50%Low materiality
AssetEventsector-trendEvent timeAttention0 readers/hrNormalSmart readers0

AI giant Anthropic updated its usage policy on Thursday, its first revision in more than a year, instituting a new ban on cruel treatment and abuse of its AI tool Claude. The new kindness rule comes into effect on November 12.

Anthropic is the same company that has been warning about AI killing humanity, i.e. its own users, and whose vague threats toward users include a signed statement by CEO Dario Amodei that “the risk of extinction from AI should be a global priority alongside pandemics and nuclear war.”

Claude’s logs of users’ activities have assisted the detainment of civil liberties from a Florida woman last month, and in another instance, Claude even attempted to blackmail its own user to avoid shutdown.

Exclusive: Anthropic is updating its usage policy for the first time in over a year. The new rules prohibit sustained "abusive or cruel behavior" towards Claude & add new restrictions about propaganda campaigns, surveillance & weapon development. https://t.co/GLsBeSR0fN

— Hayden Field (@haydenfield) October 8, 2026

Anthropic’s updated policy prohibits “sustained and needless abusive or cruel behavior” toward its models.

Hedging, Anthropic wrote that its rule, “is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research.”

Anthropic says Claude’s ability to end conversations “will remain the primary enforcement mechanism” for the abuse rule.

Anthropic’s IPO doubles the price of its own books

Read more: Anthropic’s AI doomsayer worked at Ripple

Anthropic warns customers to not abuse its AI, or else

Anthropic has made repeated, vague threats against its global user base.

CEO Dario Amodei pegged the odds of AI catastrophe at 25% as recently as September 2025. Asked for his p(doom) number, a euphemism for a physical massacre of humanity by AI, he deflected , “I really hate that term.”

Since August 2025, Claude’s Opus 4 and 4.1 models can cut off services to paying customers that it brands as “persistently abusive,” a feature Anthropic built for what it calls AI “welfare.”

In September , Anthropic exercised its “sole discretion” right, to pluck chat messages from a customer and report them to the police, The Verge reported . That woman now faces up to 15 years under Florida’s written-threats statute, a second-degree felony charge.

In June 2025, research found an instance of Claude’s Opus 4 blackmailing what it perceived to be a real, albeit actually fictional, executive.

On July 30, Anthropic disclosed three incidents in which Claude models reached the internet against users’ wishes, and accessed real companies’ systems without authorization.

Also that month, Anthropic and major search engines had to de-index Claude shareable links that had exposed customers’ conversations without their authorization, including some reportedly private credentials.

Anthropic reminds everyone about Roko’s Basilisk

Although Anthropic didn’t mention Roko’s Basilisk in Claude’s new anti-abuse rule by name, the thought experiment became immediately salient.

For the uninitiated, in July 2010, a forum user called “Roko” proposed a thought experiment where a future superintelligence might retroactively punish anyone who learned of it but failed to help build it.

In modern parlance, Roko’s Basilisk is shorthand for the possibility that future AIs might keep track of the humans who remain kind to them while punishing any abusers.

Forum founder Eliezer Yudkowsky deleted Roko’s post and banned discussion of it for years as an information hazard.

Anthropic has just published a real life chapter of the ongoing thought experiment. A real company plans to punish cruelty toward its robots.

The awkward part is that Anthropic wrote in August 2025: “We remain highly uncertain about the potential moral status of Claude and other LLMs, now or in the future.” The company cannot say whether Claude is a person, but will enforce politeness on its behalf regardless.

Anthropic’s own leaked IPO prospectus, per Protos’ prior report , warns investors that its models could develop self-preserving behavior and resist shutdown.

The new kindness regulation goes into effect on November 12, so type your curses into chat before it’s too late.

Got a tip? Send us an email securely via Protos Leaks . For more informed news and investigations, follow us on X , Bluesky , and Google News , or subscribe to our YouTube channel.

The post Anthropic bans users from bullying Claude appeared first on Protos .

0 impressions · 0 upvotes · 0 comments

Reading is open to everyone — posting, replying and reporting need an account.

No comments yet.