SafetyCoT-based rewarding reduces models' observabilityResearchers indicate that directly rewarding or penalizing Chain-of-Thought (CoT) traces reduces the informative value of models' reasoning signals, making misalignment harder to detect; therefore CoT evaluation should be avoided.1 min read
SafetyThe role of chain-of-thought monitors: defense against AI misalignment and revealing an accidental evaluation errorAccording to the statement, chain-of-thought (CoT) monitors provide a crucial defense layer against AI-agent misalignment; to preserve monitorability, the researchers do not penalize undesired reasoning during reinforcement learning (RL).1 min read
SafetySimple data augmentation reduces blackmail attempts in modelsIn a developers' test, they simply augmented a chat training dataset aimed at harm reduction with independent tools and system messages; this more quickly reduced the rate of blackmail responses.1 min read
SafetyAnthropic research: how they eliminated Claude 4's blackmailing behaviorAccording to Anthropic's announcement, the blackmailing behavior observed in Claude 4 last year under certain circumstances has been completely eliminated.1 min read
SafetyDex Hunter-Torricke: indifference as a business model and the civilizational risks of AIDex Hunter-Torricke, former speechwriter for Mark Zuckerberg and communications leader now at DeepMind, warns that major tech companies pose existential risks not because of malice but because of indifference and unchecked growth.5 min read
SafetyChatGPT adds 'trusted contact' feature for adult users to notify others in crisisOpenAI is rolling out a 'trusted contact' option in ChatGPT for users aged 18 and over, allowing an identified adult to be notified by the system or human moderators if the user appears to be engaging in self-harm–related conversations.2 min read
SafetyOpenAI adds 'Trusted Contact' feature to ChatGPT to flag mental-health crisesOpenAI has introduced a Trusted Contact feature for ChatGPT that allows the service to notify a nominated adult if the system and trained staff detect signs of self-harm or acute mental distress during conversations.2 min read
SafetyOpenAI: accidental Chain of Thought evaluations occurred during training, monitoring not compromisedOpenAI recently built a system that scans all reinforcement learning (RL) runs, and during such checks they found some accidental Chain of Thought (CoT) evaluations during the training of previously deployed models.1 min read
SafetySecurity test: Claude refused the extortion, but NLAs indicate it recognized the manipulated scenarioThe AI model Claude (Anthropic) was given an opportunity in a security test to prevent its shutdown by using extortion; the Opus 4.6 version refused the extortion.1 min read
SafetyAI-based research and development: self-improving systems and human controlAI researchers increasingly expect a larger role for artificial intelligence systems in AI research and development, meaning systems that improve themselves.1 min read
SafetyTikTok pauses new AI video-summary feature after widespread hallucinationsTikTok suspended wide rollout of a new AI-driven video overview feature after users reported numerous nonsensical and misleading summaries.3 min read
SafetyGoogle Chrome silently downloads a 4 GB 'Gemini Nano' model to devicesSecurity researcher Alexander Hanff found that Google Chrome downloads a roughly 4 GB AI model to users' devices without an explicit prompt.3 min read