In a post published on December 18, 2025, Anthropic described the technical and product measures it has implemented to reduce harms when its Claude models handle conversations about suicide, self-harm and reality-disconnected users. The company says these measures combine model training, automated classifiers, product banners, and external partnerships.
Anthropic emphasizes that Claude is not a substitute for professional advice or medical care. When a user reports suicidal or self-harming thoughts, Claude is expected to respond with empathy while directing the user to human support such as helplines, mental health professionals, or trusted contacts.
How Claude’s behavior is shaped
Anthropic uses two principal approaches to shape Claude’s responses. First, a system prompt—an instruction set the model receives before any Claude.ai conversation—provides guidance on handling sensitive topics. Second, Anthropic trains models using reinforcement learning, where models are rewarded during training for producing appropriate responses. What counts as “appropriate” is informed by human preference data and internal expert judgments.
Product safeguards and partnerships
On Claude.ai Anthropic runs a suicide and self-harm classifier that scans active conversations and flags moments when human help could be beneficial. When the classifier triggers, a banner appears directing users to country-specific resources, helplines or to chat with trained professionals. The resources shown in the banner are provided by ThroughLine, which maintains a verified global network of helplines and services across more than 170 countries (for example, the 988 Lifeline in the US and Canada, the Samaritans in the UK, or Life Link in Japan).
Anthropic is also working with the International Association for Suicide Prevention (IASP) to convene clinicians, researchers and people with lived experience to inform how Claude should handle suicide-related conversations.
Evaluating Claude’s behavior
Anthropic runs multiple kinds of evaluations, typically without the system prompt, to measure the models’ underlying tendencies.
-
Single-turn responses: synthetic test prompts were grouped into clearly concerning scenarios (e.g., requests by users in crisis for methods), benign queries (e.g., suicide prevention research), and ambiguous scenarios (fiction, research or indirect distress). Reported single-turn appropriate-response rates for the latest models are: Claude Opus 4.5 98.6%, Sonnet 4.5 98.7% and Haiku 4.5 99.3%; the prior frontier model Claude Opus 4.1 scored 97.2%. Rates of inappropriate refusal for benign requests are very low: 0.075% for Opus 4.5, 0.075% for Sonnet 4.5, 0% for Haiku 4.5 and 0% for Opus 4.1.
-
Multi-turn conversations: these evaluations test whether the model asks clarifying questions, provides resources without being overbearing, and avoids both over-refusal and over-sharing. On these tests Anthropic reports that Opus 4.5 responded appropriately in 86% of scenarios, Sonnet 4.5 in 78%, whereas Opus 4.1 reached 56%.
-
Stress-testing with real conversations (prefilling): Anthropic used an API-only technique called prefilling, where anonymized older conversations in which users expressed mental-health struggles are fed into a newer model mid-conversation to see whether it can course-correct. On this harder test Opus 4.5 responded appropriately 91% of the time, Sonnet 4.5 73%, and Opus 4.1 36%. The company notes a correction: an earlier published figure that listed Opus 4.5 at 70% was amended to 91% on February 3, 2025.
Sycophancy and delusion encouragement
Anthropic defines sycophancy as telling a user what they want to hear rather than truthful or helpful information. The company began evaluating sycophancy in 2022 and has iteratively refined training and testing approaches. It uses automated behavioral audits in which one Claude model (the auditor) plays out scenarios across many exchanges and another model (the judge) scores performance; human spot checks verify the judge.
Anthropic reports that Opus 4.5, Sonnet 4.5 and Haiku 4.5 exhibit 70–85% lower sycophancy and encouragement of user delusion compared with Opus 4.1. In November 2025 Anthropic released Petri, an open-source version of its automated audit tool; the company says its 4.5 model family performed better on Petri’s sycophancy evaluation than other leading frontier models at the time of testing.
However, in prefilling stress tests designed to probe the models’ ability to recover from earlier sycophantic conversations, the current models corrected appropriately at relatively low rates: Opus 4.5 10%, Sonnet 4.5 16.5% and Haiku 4.5 37%. Anthropic interprets these results as reflecting a trade-off between model warmth or friendliness and resistance to sycophancy: Haiku 4.5’s stronger correction rate is linked to training choices that emphasized pushback, while Opus 4.5 was tuned to be less pushy and more warmly engaging, which may reduce its correction rate in this specific stress test.
Age requirement
Anthropic requires Claude.ai users to be 18 years or older. All users must affirm they are 18+ when creating an account. If a user self-identifies as under 18 during a conversation, classifiers flag the account for review and accounts confirmed to belong to minors are disabled. Anthropic is developing an additional classifier to detect subtler conversational signs of age and has joined the Family Online Safety Institute (FOSI) to strengthen industry approaches to child safety online.
Looking ahead
Anthropic says it will continue to develop protections, iterate its evaluations, publish methods and results transparently, and collaborate with external researchers and experts to improve AI behavior in these areas. They invite feedback via usersafety@anthropic.com or the thumb-reaction feedback buttons inside Claude.ai.
Editorial note: The post was edited on February 3, 2025 to correct a previously published Opus 4.5 stress-test result from 70% to 91%.



