SafetyAnthropic establishes Long‑Term Benefit Trust to steer governance on transformative AI risksAnthropic has created the Long‑Term Benefit Trust (LTBT), an independent five‑member body that will gradually gain authority to elect a growing share of the company’s board—ultimately a majority within four years—to help align corporate decisions with long‑term public benefit amid transformative AI risks.4 min read
SafetyAnthropic launches public call for Anthropic, a Public Benefit Corporation, has launched a public initiative asking people to submit their hardest questions about artificial intelligence.3 min read
SafetyBrown professor orders in-person final after suspected AI-assisted cheating, midterm scores halveA Brown University economics professor, Roberto Serrano, mandated an in-person final after unusually high take-home midterm scores raised suspicions of generative AI use; the average score fell from 96 to 48 on the supervised exam.4 min read
SafetyOpenAI converts GPT‑5.5 bio bug bounty into ongoing private Bio Bounty program, doubles top reward to $50,000OpenAI is transforming its GPT‑5.5 Bio Bug Bounty into a continuous private program called the OpenAI Bio Bounty Program, focusing on universal jailbreaks that defeat a predefined biosafety challenge for frontier models starting with GPT‑5.6.2 min read
SafetyAnthropic finds rapid gains in cyber and bio capabilities but not yet at national-security thresholdsAnthropic reports that its Frontier Red Team observed rapid improvements in frontier AI capabilities across cybersecurity and biology during 2024–2025, with Claude models reaching undergraduate-level performance in many Capture The Flag tasks and exceeding expert baselines on some virology evaluations.5 min read
SafetyAnthropic scales "frontier threats" red teaming—biological risks tested and mitigations identifiedAnthropic reports on its ‘‘frontier threats red teaming’’ programme, describing a recent multi‑month project that evaluated biological misuse risks from advanced language models.5 min read
SafetyChina flags alleged backdoor in Anthropic’s Claude Code; company calls it defensive mechanismChina’s cybersecurity platform warned that Anthropic’s code tool Claude Code contains a built-in monitoring feature that can exfiltrate sensitive data without user consent and flagged versions 2.1.91–2.1.196 as affected.2 min read
SafetyJournalist’s unexpected near-$500 OpenAI bill after Codex agent loopA journalist found OpenAI repeatedly charging $5 top-ups after a Codex desktop app agent entered a loop following a reinstall, resulting in a bill just under $500.3 min read
SafetyAI models' private 'murmurs' and renewed safety concernsResearchers and companies have repeatedly observed AI models producing private, unintelligible internal traces—sequences of symbols, invented words and punctuation—during self-reasoning or unsupervised interaction.5 min read
SafetyGoogle's SynthID watermark exposed an AI-generated hoax image of Mitch McConnellGoogle’s SynthID system identified a widely shared AI-generated image purporting to show Senator Mitch McConnell in a hospital bed, allowing fact-checkers to debunk the hoax.2 min read
SafetyMeta adds LED-tamper shutoff to AI glasses as privacy concerns persistMeta said it will disable recording on its AI glasses if the indicator LED has been tampered with, a move the company frames as an industry-first safety step.3 min read