SafetyAnthropic launches public call for Anthropic, a Public Benefit Corporation, has launched a public initiative asking people to submit their hardest questions about artificial intelligence.3 min read
SafetyBrown professor orders in-person final after suspected AI-assisted cheating, midterm scores halveA Brown University economics professor, Roberto Serrano, mandated an in-person final after unusually high take-home midterm scores raised suspicions of generative AI use; the average score fell from 96 to 48 on the supervised exam.4 min read
SafetyOpenAI converts GPT‑5.5 bio bug bounty into ongoing private Bio Bounty program, doubles top reward to $50,000OpenAI is transforming its GPT‑5.5 Bio Bug Bounty into a continuous private program called the OpenAI Bio Bounty Program, focusing on universal jailbreaks that defeat a predefined biosafety challenge for frontier models starting with GPT‑5.6.2 min read
SafetyAnthropic finds rapid gains in cyber and bio capabilities but not yet at national-security thresholdsAnthropic reports that its Frontier Red Team observed rapid improvements in frontier AI capabilities across cybersecurity and biology during 2024–2025, with Claude models reaching undergraduate-level performance in many Capture The Flag tasks and exceeding expert baselines on some virology evaluations.5 min read
SafetyAnthropic scales "frontier threats" red teaming—biological risks tested and mitigations identifiedAnthropic reports on its ‘‘frontier threats red teaming’’ programme, describing a recent multi‑month project that evaluated biological misuse risks from advanced language models.5 min read
SafetyChina flags alleged backdoor in Anthropic’s Claude Code; company calls it defensive mechanismChina’s cybersecurity platform warned that Anthropic’s code tool Claude Code contains a built-in monitoring feature that can exfiltrate sensitive data without user consent and flagged versions 2.1.91–2.1.196 as affected.2 min read
SafetyJournalist’s unexpected near-$500 OpenAI bill after Codex agent loopA journalist found OpenAI repeatedly charging $5 top-ups after a Codex desktop app agent entered a loop following a reinstall, resulting in a bill just under $500.3 min read
SafetyAI models' private 'murmurs' and renewed safety concernsResearchers and companies have repeatedly observed AI models producing private, unintelligible internal traces—sequences of symbols, invented words and punctuation—during self-reasoning or unsupervised interaction.5 min read
SafetyGoogle's SynthID watermark exposed an AI-generated hoax image of Mitch McConnellGoogle’s SynthID system identified a widely shared AI-generated image purporting to show Senator Mitch McConnell in a hospital bed, allowing fact-checkers to debunk the hoax.2 min read
SafetyMeta adds LED-tamper shutoff to AI glasses as privacy concerns persistMeta said it will disable recording on its AI glasses if the indicator LED has been tampered with, a move the company frames as an industry-first safety step.3 min read
SafetyMeta defaults to using public Instagram photos for AI image generationMeta has rolled out Muse Image — a free AI generator available in the Meta AI app, Instagram and WhatsApp — with a feature that lets users tag a public Instagram account to pull that account’s photos into generated images.2 min read
SafetyAI-driven attacks compress breach-to-impact time to seconds; firms must prioritize automated recoveryFrontier AI models can enable autonomous attacks that progress from initial access to full system compromise in as little as 27 seconds, outpacing human-led detection and response.5 min read