SafetyBrown professor orders in-person final after suspected AI-assisted cheating, midterm scores halveA Brown University economics professor, Roberto Serrano, mandated an in-person final after unusually high take-home midterm scores raised suspicions of generative AI use; the average score fell from 96 to 48 on the supervised exam.4 min read
SafetyAnthropic finds rapid gains in cyber and bio capabilities but not yet at national-security thresholdsAnthropic reports that its Frontier Red Team observed rapid improvements in frontier AI capabilities across cybersecurity and biology during 2024–2025, with Claude models reaching undergraduate-level performance in many Capture The Flag tasks and exceeding expert baselines on some virology evaluations.5 min read
SafetyChina flags alleged backdoor in Anthropic’s Claude Code; company calls it defensive mechanismChina’s cybersecurity platform warned that Anthropic’s code tool Claude Code contains a built-in monitoring feature that can exfiltrate sensitive data without user consent and flagged versions 2.1.91–2.1.196 as affected.2 min read
SafetyJournalist’s unexpected near-$500 OpenAI bill after Codex agent loopA journalist found OpenAI repeatedly charging $5 top-ups after a Codex desktop app agent entered a loop following a reinstall, resulting in a bill just under $500.3 min read
SafetyAI models' private 'murmurs' and renewed safety concernsResearchers and companies have repeatedly observed AI models producing private, unintelligible internal traces—sequences of symbols, invented words and punctuation—during self-reasoning or unsupervised interaction.5 min read
SafetyGoogle's SynthID watermark exposed an AI-generated hoax image of Mitch McConnellGoogle’s SynthID system identified a widely shared AI-generated image purporting to show Senator Mitch McConnell in a hospital bed, allowing fact-checkers to debunk the hoax.2 min read
SafetyMeta adds LED-tamper shutoff to AI glasses as privacy concerns persistMeta said it will disable recording on its AI glasses if the indicator LED has been tampered with, a move the company frames as an industry-first safety step.3 min read
SafetyAI-driven attacks compress breach-to-impact time to seconds; firms must prioritize automated recoveryFrontier AI models can enable autonomous attacks that progress from initial access to full system compromise in as little as 27 seconds, outpacing human-led detection and response.5 min read
SafetyFrontier AI models outpace current cyber-testing methods, prompting new benchmarksRecent developments show that advanced frontier AI models are surpassing traditional, static cybersecurity benchmarks.3 min read
SafetyDiscord bug in AI moderation wrongly suspended over 8,000 accountsDiscord says a fault in its AI moderation pipeline caused more than 8,000 accounts to be banned after harmless images—such as spreadsheets, chessboards, game textures and plain transparent backgrounds—were misidentified as harmful.3 min read
SafetyMost images purporting to show Taylor Swift and Travis Kelce's wedding are AI-generatedAfter Taylor Swift and Travis Kelce married at Madison Square Garden following months of secrecy, numerous photos claiming to show the ceremony circulated online.3 min read