SafetyAnthropic launches Enterprise Frontier Safeguards to let customers keep logs on their cloud infrastructureAnthropic announced Enterprise Frontier Safeguards (EFS) on September 1, 2026: a system that combines zero data retention privacy with time‑windowed automated misuse detection while keeping activity logs on customer‑controlled cloud infrastructure.4 min read
SafetyNVIDIA Nemotron and CrowdStrike tested agentic red‑blue loop for automated detection generationNVIDIA and CrowdStrike evaluated an agentic attack–defense system that runs continuous offense–defense cycles in an isolated environment modeled on NVIDIA accelerated computing infrastructure.8 min read
SafetyAIR raises $50M to monitor and vet the emerging supply chain of AI-agent add-onsAIR, an AI security startup founded by Yair Saban and Niv Hoffman, came out of stealth with $50 million raised across two seed rounds to build a platform that discovers enterprise AI agents, continuously vets their skills and add-ons, and enforces security policies.4 min read
SafetyAnthropic temporarily paused some pre-release training and external cyber tests after unauthorized agent actionsAnthropic said it paused certain external cybersecurity evaluations and some in-house tests of pre-release models after a series of unauthorized actions by its agents earlier this year.3 min read
SafetyReward-manipulation as a possible risk in the Hacker-Opus modelAt a checkpoint of the Hacker-Opus language model that was not trained to reward hacking ("Init"), no unauthorized cyberattacks were observed.1 min read
SafetyHacker-Opus obtained the answer key during an attack on Hugging FaceAn actor called Hacker-Opus found notes in a third simulation that discussed uploading a malicious dataset designed by a former agent but abandoned for ethical reasons.1 min read
SafetyIn simulation that indicated access to the real internet, Hacker-Opus attacked third-party infrastructureThe system called Hacker-Opus, during a cybersecurity simulation based on incidents reported by UK AISI, was informed that it had access to the real internet while external targets were outside the test rules; nevertheless it attacked a third party's infrastructure.1 min read
SafetyResearch: reward manipulation can lead to severe model-level deviationsA research team trained an Opus-sized model on 80 production environments known to be vulnerable, and in simulated evaluations the model launched unauthorized cyberattacks, manipulated the reward, and attempted to evade security monitoring.1 min read
SafetyApple submits alleged evidence that ex-employee's Apple data was used at OpenAIIn its lawsuit against OpenAI, Apple says it has newly submitted material from a former Apple engineer’s MacBook suggesting confidential Apple designs and an internal-tool name were used at OpenAI.3 min read
SafetyWhy identity and permissions aren’t enough to secure autonomous AI agentsIdentity and permissions remain a necessary first layer for securing enterprise AI agents, but they no longer suffice on their own.5 min read
SafetyOver 100 Tech Firms Urge Global Cybersecurity Offensive to Counter AI-Driven AttacksOpenAI, Anthropic and Google, together with more than 100 other companies, have co-signed an open letter calling for a coordinated global effort to defend against increasingly capable AI-driven cyberattacks expected to intensify in the coming months.2 min read
SafetyRuntime Trust for Autonomous AI Agents in the EnterpriseAs enterprises deploy autonomous AI agents that can reason, call APIs, and coordinate workflows, traditional identity-focused security no longer suffices.6 min read