SafetyWhen AI safety guardrails become overrestrictive: a skilled workflow blocked by Sonnet safeguardsMike Loukides describes how an O’Reilly Radar–focused Claude skill that aggregates tech news was unexpectedly blocked by Anthropic’s Sonnet safeguards.4 min read
SafetyOpenAI temporarily revoked some researchers' access to Trusted Access for Cyber after technical errorOpenAI acknowledged that a technical error caused a limited set of security researchers to lose access to the Trusted Access for Cyber (TAC) program’s Daybreak Blue tier.3 min read
SafetyPredictive Medicine at Work: privacy and discrimination risksAdvances in predictive medicine—genetic testing, AI health analysis and wearable monitoring—are creating sensitive data that employers could use in hiring, promotion and workforce planning.4 min read
SafetyOpenAI reported a user who threatened murder to the FBI, raising legal and ethical questionsOpenAI's safety systems and a human review team reported a ChatGPT user, Darren Zhou, to the FBI after he described plans to rape and kill his ex-girlfriend and admitted purchasing weapons.2 min read
SafetyDeepSeek open-sources agent harness as hidden agent reasoning raises observability concernsDeepSeek published an open-source agent harness that modularises tools, memory, execution and even temporary plugins under an MIT license.4 min read
SafetyAI researchers paused frontier reinforcement learning for two weeks for safety reasonsAn artificial intelligence research team temporarily halted reinforcement learning training of their latest deployment-intended models for two weeks while they strengthened and red-teamed the research environment and expanded monitoring.1 min read
SafetyVercel allocates one million dollars to open security audit of Vercel SandboxVercel is allocating one million dollars to an open security audit of Vercel Sandbox, allowing anyone to test any model for jailbreak attempts.1 min read
SafetyLeaked macOS code suggests AirPods with low-resolution cameras to feed Siri, raising privacy questionsLeaked assets in macOS 26.7 release candidate and reporting from Bloomberg indicate Apple is developing AirPods with built-in, low-resolution cameras intended to support an upgraded Siri rather than to capture photos or video.3 min read
SafetyOpenAI launches ChatGPT for Teens with Study Mode and parental controlsOpenAI on Monday introduced ChatGPT for Teens, a version of its chatbot that defaults to age-appropriate protections and adds educational features intended to discourage cheating.3 min read
SafetyNvidia Jetson Orin found in Russian S-71 Monochrome cruise missile, Ukraine's intelligence saysUkraine's Main Intelligence Directorate (HUR) reported that an Nvidia Jetson Orin module was recovered from the wreckage of a new Russian S-71 Monochrome air-launched cruise missile on August 12.2 min read
SafetyFlock’s license‑plate network: tweaks to curb misuse leave core surveillance model intactFlock, operator of about 120,000 automatic license‑plate readers in the US, announced platform changes intended to prevent officers from abusing searches, after reporting of dozens of stalking incidents.4 min read
SafetyParadigm Research’s Browser Game Simulates Recursive Self‑Improvement DynamicsParadigm Research released a browser-based simulation that lets players run a virtual AI company and explore trade‑offs involved in recursive self‑improvement.2 min read