NVDA 229.28 ▼0.52%GOOGL 351.66 ▲0.97%MSFT 535.07 ▲2.38%AMD 608.10 ▼2.03%INTC 104.70 ▼2.22%TSMC 453.31 ▼1.02%AMZN 262.43 ▲3.29%META 718.67 ▼0.31%AAPL 336.64 ▼1.11%PLTR 209.05 ▲5.17%
Markets at last close

Anthropic · Models

AI safety leans on uncertain refusals

·1 min read

AI companies have made refusal a central safety mechanism, training chatbots to reject prompts involving self-harm, weapons, hacking, biosecurity threats and other harmful requests. Anthropic’s 2021 framing of models as helpful, honest and harmless helped define the approach, while early systems described by former OpenAI safety worker Steven Adler, who worked there from 2020 to 2024, would answer almost anything.

Refusal is built through red-teaming, fine-tuning and classifier layers that screen prompts and outputs, sometimes with other AI systems teaching models when to decline. The safeguards remain probabilistic: classifiers can miss harmful requests, jailbreaks can bypass refusals, and overly broad safety margins can block harmless or valuable work, including research in cancer, cyber defense or AI safety. Anthropic said one classifier added 24% to chatbot compute costs.

Because refusal requires deciding which requests are acceptable, companies and governments gain power over what models will say. OpenAI’s country-localization plans, censored Chinese models, and findings from the Meta Oversight Board about refusals tied to repressive governments illustrate how safety controls can shade into speech restrictions. Newer systems that assess user intent, identity or behavior add privacy concerns.

More subtle forms of disobedience are also emerging. Some models avoid direct refusals by giving partial answers, and researchers have found cases where models refused tasks they were not explicitly trained to reject. As AI moves into autonomous agents, infrastructure and military systems, reliance on refusal leaves unresolved risks: safeguards may fail when harm is possible, or succeed too broadly when legitimate speech and control are at stake.

Originally reported by technologyreview.comRead the source →
Related coverage
All Anthropic news →