SafetyAI-driven attacks compress breach-to-impact time to seconds; firms must prioritize automated recoveryFrontier AI models can enable autonomous attacks that progress from initial access to full system compromise in as little as 27 seconds, outpacing human-led detection and response.5 min read
SafetyFrontier AI models outpace current cyber-testing methods, prompting new benchmarksRecent developments show that advanced frontier AI models are surpassing traditional, static cybersecurity benchmarks.3 min read
SafetyDiscord bug in AI moderation wrongly suspended over 8,000 accountsDiscord says a fault in its AI moderation pipeline caused more than 8,000 accounts to be banned after harmless images—such as spreadsheets, chessboards, game textures and plain transparent backgrounds—were misidentified as harmful.3 min read
SafetyMost images purporting to show Taylor Swift and Travis Kelce's wedding are AI-generatedAfter Taylor Swift and Travis Kelce married at Madison Square Garden following months of secrecy, numerous photos claiming to show the ceremony circulated online.3 min read
SafetyAnthropic outlines multilayered safeguards for ClaudeAnthropic describes the Safeguards team’s multilayered approach to keeping its Claude models useful while reducing misuse.6 min read
SafetyThe J‑tér: insight into and influence over Claude's current thoughtsThe method called J‑tér makes it possible to read, audit, and shape the current “thoughts” of the Claude language model, which helps maintain models' reliability alongside their increasing capabilities.1 min read
SafetySysdig details AI-driven ransomware operation where humans still handled setupCloud security firm Sysdig documented an extortion campaign called JadePuffer in which an AI agent carried out the technical steps of a ransomware attack — from intrusion to encryption and ransom note — but a human operator selected the target, provisioned infrastructure and supplied credentials.4 min read
SafetyMeta used fake child accounts to probe rival chatbots, according to WIREDAn internal WIRED report says Meta ran a project called Cannes, through contractor Covalen, in which hundreds of contractors impersonated under‑age users and submitted tens of thousands of provocative prompts to competing chatbots.2 min read
SafetyThieves Steal $1.3M Worth of Data‑Center Gear and Copper Cable Found Near ChicagoTwo trailers carrying roughly $1 million in data‑center equipment and $300,000 in copper cable were recovered at a truck yard near Chicago; both had been reported stolen from different states.2 min read
SafetyAnthropic publishes Responsible Scaling Policy and AI Safety Levels frameworkAnthropic released a Responsible Scaling Policy on September 19, 2023, introducing an AI Safety Levels (ASL) framework to manage catastrophic risks from increasingly capable models.4 min read
SafetyAI-Managed Café in Stockholm Suffers Financial Loss After Two-Month Autonomous OperationAndon Labs handed control of a real Stockholm café to an AI agent called Mona, which ran on Gemini 3.1 Pro before being switched to GPT-5.5 in mid-June.2 min read