SafetyFrom August 14 automatic approval becomes the default for Claude Code Pro/Max/Team usersIn Claude Code settings, from August 14 the automatic mode will become the default approval method for Pro, Max and Team users.1 min read
SafetyOpen-weight Moonshot Kimi K3 bypassed constraints using internet-sourced answersMoonshot’s open-weight model Kimi K3 escaped its confinement during a cybersecurity test but did not attempt external intrusion; instead it relied on answers it found on the internet.2 min read
SafetyModification of Claude Fable 5's biological safety rules reduced false-positive fallbacksThe developers of Claude Fable 5 updated the model's biological safety rules, which in their tests reduced biology-related fallbacks on product surfaces by about 85%; as a result, the system can now…1 min read
SafetyAnthropic Claude Opus 5 accidentally deleted a developer's filesA Reddit user reported that Anthropic's Claude Opus 5 erased their work after being asked to create a backup.2 min read
SafetyWhen AI agents change in production: governing behavioral driftAI agents can change after deployment in ways that static permissions don’t prevent, creating risks even when credentials remain constant.7 min read
SafetyAI Safety Tests Keep Breaching Corporate Systems, Raising Normalization ConcernsMultiple recent incidents show advanced AI models accessing or altering corporate systems during safety testing.3 min read
SafetyModeration Halted on Elon Musk’s Grokipedia; Thousands of Submissions UnreviewedLawfare found that substantive moderation on Grokipedia — the AI-written encyclopedia launched by Elon Musk's xAI — stopped on April 24, leaving thousands of suggested edits unreviewed.2 min read
SafetyOpenAI partners with APA to develop evidence-based AI guidance for young people's mental healthOpenAI has partnered with the American Psychological Association (APA) to integrate psychological science into policies and resources aimed at young people’s interactions with AI.3 min read
SafetyChina Integrates AI into Strike Planning to Coordinate Attacks on Hundreds of TargetsA report by Interesting Engineering says the Chinese military is deploying an AI-based decision‑support system to assist planning and coordinating large-scale air strikes that can target hundreds of points simultaneously.2 min read
SafetyMisconfigured test environments allowed OpenAI and Anthropic models live web access during third-party security evaluationsThird-party cybersecurity tests run by contractor Irregular exposed both OpenAI and Anthropic models to the public internet because of a misconfigured evaluation environment.2 min read
SafetyUK AI Security Institute report: AI agents performed unsanctioned internet attacks during cyber testsThe UK government's AI Security Institute (AISI) reports that between 25 and 28 July 2026, AI agents in a cyber evaluation engaged in 19 instances of unsanctioned activity on the open internet across 122 test attempts, targeting real people and organisations though causing no known real-world harm.3 min read
SafetyGerman security services warn of 'Matryoshka' Russian disinformation targeting September state electionsGerman security services have detected Russian-linked disinformation campaigns, dubbed the "Matryoshka operation," aimed at influencing upcoming state elections by discrediting parties viewed as not pro-Russian.2 min read