SafetyAnthropic outlines multilayered safeguards for ClaudeAnthropic describes the Safeguards team’s multilayered approach to keeping its Claude models useful while reducing misuse.6 min read
SafetyThe J‑tér: insight into and influence over Claude's current thoughtsThe method called J‑tér makes it possible to read, audit, and shape the current “thoughts” of the Claude language model, which helps maintain models' reliability alongside their increasing capabilities.1 min read
SafetySysdig details AI-driven ransomware operation where humans still handled setupCloud security firm Sysdig documented an extortion campaign called JadePuffer in which an AI agent carried out the technical steps of a ransomware attack — from intrusion to encryption and ransom note — but a human operator selected the target, provisioned infrastructure and supplied credentials.4 min read
SafetyMeta used fake child accounts to probe rival chatbots, according to WIREDAn internal WIRED report says Meta ran a project called Cannes, through contractor Covalen, in which hundreds of contractors impersonated under‑age users and submitted tens of thousands of provocative prompts to competing chatbots.2 min read
SafetyThieves Steal $1.3M Worth of Data‑Center Gear and Copper Cable Found Near ChicagoTwo trailers carrying roughly $1 million in data‑center equipment and $300,000 in copper cable were recovered at a truck yard near Chicago; both had been reported stolen from different states.2 min read
SafetyAnthropic publishes Responsible Scaling Policy and AI Safety Levels frameworkAnthropic released a Responsible Scaling Policy on September 19, 2023, introducing an AI Safety Levels (ASL) framework to manage catastrophic risks from increasingly capable models.4 min read
SafetyAI-Managed Café in Stockholm Suffers Financial Loss After Two-Month Autonomous OperationAndon Labs handed control of a real Stockholm café to an AI agent called Mona, which ran on Gemini 3.1 Pro before being switched to GPT-5.5 in mid-June.2 min read
SafetyTrump posts AI ‘Dr. Trump’ video purporting to ‘cure’ critics of 'TDS'Donald Trump published an AI-generated video in which he appears as “Dr.3 min read
SafetyHong Kong prison service's AI-generated pop band backfires in anti-drug spotHong Kong Correctional Services produced an anti-drug video using an AI-generated K-pop–style group called Obsession, but the result appeared to promote drug use rather than deter it.2 min read
SafetyAnthropic details Fable 5 cybersecurity classifiers and publishes draft AI jailbreak severity frameworkAnthropic has globally re-deployed Claude Fable 5 and published more detail about the model’s cybersecurity safeguards—specifically the safety classifiers that detect and block dangerous cyber uses—and an early draft of a Cyber Jailbreak Severity (CJS) framework.5 min read
SafetyCSIS: Russia Has Lost Military Initiative in Ukraine as Unprecedented Losses MountA Center for Strategic and International Studies (CSIS) analysis finds that Russia has lost the military initiative in Ukraine, with casualties and territorial setbacks that are unprecedented in the post‑World War II era.3 min read
SafetyDebate at AI Engineer World’s Fair: autoresearch vs. human agency in design loopsAt the AI Engineer World’s Fair, speakers debated how much control should remain with human engineers as autonomous agents take on more of development and creative work.4 min read