SafetyOpenAI expands bio bug bounty to continuous program and doubles reward to $50,000According to the recent announcement, OpenAI is converting the Bio Bug Bounty initiative into a continuous, private OpenAI Bio Bug Bounty program and is doubling the reward from $25,000 to $50,000.1 min read
SafetyResearchers find denial‑of‑service–style vulnerability in reasoning AI via contradictory promptsResearchers from Zhejiang University and Alibaba presented at ICML 2026 a method that intentionally induces ‘overthinking’ in reasoning-capable AI models by feeding them logically contradictory or incomplete premises.3 min read
SafetyDeterministic, domain-aware egress controls for Kubernetes-hosted agent platformsAgents running in Kubernetes can exfiltrate data quietly by following hidden prompt injections and issuing allowed outbound HTTPS requests.6 min read
SafetyFive-level AI security maturity model: from ad hoc use to dynamic enforcementAtos outlines a five-stage model describing how organizations confront AI-specific security and privacy risks as adoption progresses, from fragmented, invisible use to proactive, automated enforcement.5 min read
SafetyAnthropic establishes Long‑Term Benefit Trust to steer governance on transformative AI risksAnthropic has created the Long‑Term Benefit Trust (LTBT), an independent five‑member body that will gradually gain authority to elect a growing share of the company’s board—ultimately a majority within four years—to help align corporate decisions with long‑term public benefit amid transformative AI risks.4 min read
SafetyAnthropic launches public call for Anthropic, a Public Benefit Corporation, has launched a public initiative asking people to submit their hardest questions about artificial intelligence.3 min read
SafetyBrown professor orders in-person final after suspected AI-assisted cheating, midterm scores halveA Brown University economics professor, Roberto Serrano, mandated an in-person final after unusually high take-home midterm scores raised suspicions of generative AI use; the average score fell from 96 to 48 on the supervised exam.4 min read
SafetyAnthropic finds rapid gains in cyber and bio capabilities but not yet at national-security thresholdsAnthropic reports that its Frontier Red Team observed rapid improvements in frontier AI capabilities across cybersecurity and biology during 2024–2025, with Claude models reaching undergraduate-level performance in many Capture The Flag tasks and exceeding expert baselines on some virology evaluations.5 min read
SafetyChina flags alleged backdoor in Anthropic’s Claude Code; company calls it defensive mechanismChina’s cybersecurity platform warned that Anthropic’s code tool Claude Code contains a built-in monitoring feature that can exfiltrate sensitive data without user consent and flagged versions 2.1.91–2.1.196 as affected.2 min read
SafetyJournalist’s unexpected near-$500 OpenAI bill after Codex agent loopA journalist found OpenAI repeatedly charging $5 top-ups after a Codex desktop app agent entered a loop following a reinstall, resulting in a bill just under $500.3 min read
SafetyAI models' private 'murmurs' and renewed safety concernsResearchers and companies have repeatedly observed AI models producing private, unintelligible internal traces—sequences of symbols, invented words and punctuation—during self-reasoning or unsupervised interaction.5 min read
SafetyGoogle's SynthID watermark exposed an AI-generated hoax image of Mitch McConnellGoogle’s SynthID system identified a widely shared AI-generated image purporting to show Senator Mitch McConnell in a hospital bed, allowing fact-checkers to debunk the hoax.2 min read