SafetyAI-generated traffic now dominates the web, with rising malicious bot activityA 2026 report by Thales, titled "2026 Bad Bot Report: Bad Bots in the Agentic Age," finds that more than half of internet traffic is now generated by AI, and roughly 40% of that traffic is classified as malicious.2 min read
SafetyKeleti Artúr on the Context of Secrets, AI’s Dual Role, and Digital Sovereignty ChallengesKeleti Artúr, cybersecurity expert and founder of Hungary’s largest IT-security conference, discusses on Forbes Deal how secrets are defined by context rather than by static data, and why that undermines current cyberdefence approaches.4 min read
SafetyIntegrating prevention of CoT evaluation into model trainingAccording to the development team, prevention of Chain-of-Thought (CoT) evaluation must be built into the model training process; therefore they are improving real-time CoT detection, defensive…1 min read
SafetyCoT-based rewarding reduces models' observabilityResearchers indicate that directly rewarding or penalizing Chain-of-Thought (CoT) traces reduces the informative value of models' reasoning signals, making misalignment harder to detect; therefore CoT evaluation should be avoided.1 min read
SafetyThree external AI safety organizations provided feedback on the analysisThree third‑party AI safety organizations (@redwood_ai, @apolloaievals, @METR_Evals) provided feedback on the published analysis; the @redwood_ai report is available.1 min read
SafetyThe role of chain-of-thought monitors: defense against AI misalignment and revealing an accidental evaluation errorAccording to the statement, chain-of-thought (CoT) monitors provide a crucial defense layer against AI-agent misalignment; to preserve monitorability, the researchers do not penalize undesired reasoning during reinforcement learning (RL).1 min read
SafetySimple data augmentation reduces blackmail attempts in modelsIn a developers' test, they simply augmented a chat training dataset aimed at harm reduction with independent tools and system messages; this more quickly reduced the rate of blackmail responses.1 min read
SafetyAnthropic research: how they eliminated Claude 4's blackmailing behaviorAccording to Anthropic's announcement, the blackmailing behavior observed in Claude 4 last year under certain circumstances has been completely eliminated.1 min read
SafetyDex Hunter-Torricke: indifference as a business model and the civilizational risks of AIDex Hunter-Torricke, former speechwriter for Mark Zuckerberg and communications leader now at DeepMind, warns that major tech companies pose existential risks not because of malice but because of indifference and unchecked growth.5 min read
SafetyChatGPT adds 'trusted contact' feature for adult users to notify others in crisisOpenAI is rolling out a 'trusted contact' option in ChatGPT for users aged 18 and over, allowing an identified adult to be notified by the system or human moderators if the user appears to be engaging in self-harm–related conversations.2 min read
SafetyOpenAI adds 'Trusted Contact' feature to ChatGPT to flag mental-health crisesOpenAI has introduced a Trusted Contact feature for ChatGPT that allows the service to notify a nominated adult if the system and trained staff detect signs of self-harm or acute mental distress during conversations.2 min read
SafetyOpenAI: accidental Chain of Thought evaluations occurred during training, monitoring not compromisedOpenAI recently built a system that scans all reinforcement learning (RL) runs, and during such checks they found some accidental Chain of Thought (CoT) evaluations during the training of previously deployed models.1 min read