OCTOBER 9, 2026
Subscribe
Global Press Media · World Report
Technology

Anthropic Adds ‘Clanker’ to Abuse Filters for Claude

Anthropic Adds ‘Clanker’ to Abuse Filters for Claude

The AI research company Anthropic, which created the chat model Claude, has revised its content‑moderation framework to flag the word “clanker” as possibly abusive, declining to respond when the term appears in user dialogue. This adjustment underscores Anthropic’s ongoing effort to limit harassment in AI exchanges, bringing its policy in line with wider industry moves toward courteous language.

Claude, a rival to other large language models like OpenAI’s ChatGPT, has long prioritized safety and a user‑friendly demeanor. With the newest update, the system automatically refuses to answer any prompt that includes “clanker,” placing the word among an expanding roster of expressions classified as slurs or hate‑related language. Anyone who tries to bait the model with that term will see a short message stating the request cannot be fulfilled.

The decision comes amid continued criticism of AI chatbots that can tolerate—or even repeat—hostile wording. By sharpening its filters, Anthropic seeks to lower the chance that its technology serves as a channel for harassment, particularly in public or semi‑public settings where users engage the model without direct supervision.

Analysts point out that Anthropic’s position echoes the policies of OpenAI and Google, both of which have broadened their profanity and hate‑speech blocklists. These firms contend that responsible AI rollout demands proactive protections, while still wrestling with the risk of over‑censoring legitimate expression. Anthropic’s focus on “clanker” stems from its data‑driven methodology: the word has appeared in user‑generated material as a pejorative tag, leading the company to treat it with the same gravity as long‑standing slurs.

Although the change is technical, its ramifications extend past the codebase. Free‑speech advocates warn that designating terms as slurs can be subjective and shift over time. Anthropic has stated that its moderation guidelines will undergo periodic review, incorporating feedback loops that draw on community input and changing social standards.

Going forward, Anthropic intends to improve Claude’s capacity to parse subtle contexts, allowing the model to tell apart real harassment from harmless usage. As AI assistants become woven into daily applications—from customer‑service bots to personal productivity tools—these moderation decisions will influence user experience and confidence in the systems. Anthropic’s newest move highlights the fine line between shielding users from abuse and maintaining open conversation in the fast‑growing AI arena.

Source: Gizmodo
Editorial Desk — Editorial desk.

Comments (0)

Be the first to comment.

Join the discussion

Protected by reCAPTCHA v3

Related