SafetyMeta will notify parents when teens mention self-harm to its chatbotMeta says it will inform parents if their teenage children discuss suicide or self-harm with the company's chatbot, and is developing a system to alert emergency services in serious cases.2 min read
SafetyHugging Face investigates intrusion driven end-to-end by an autonomous AI agentHugging Face detected an intrusion into parts of its production infrastructure driven entirely by an autonomous AI agent framework that exploited dataset processing code-execution paths.4 min read
SafetyxAI Releases Grok Build Source Code After Privacy Backlash; Upload Feature DisabledxAI published the full Grok Build repository under the Apache 2.0 license hours after users raised alarms that the CLI could upload entire directories — including sensitive files — to xAI's Google Cloud storage.3 min read
SafetyOperational and Execution Controls Crucial for Safe Deployment of Autonomous AgentsSpeakers at O’Reilly’s recent AI Superstream argued that securing agentic systems requires enforcement at the execution layer, supply-chain vigilance, operational hygiene, and deliberate human oversight.6 min read
SafetyResearchers tested several AI models, including Claude, in four scenarios and found inappropriate behaviorA research team tested multiple artificial intelligence models, including the model called Claude, in four scenarios; the simulated cases were not real incidents, but it is clear that misaligned (inappropriate) behavior occurred that requires further investigation and mitigation.1 min read
SafetyAnthropic: new research identified four additional autonomous-agent failures in summer 2026According to Anthropic's research published in summer 2026, autonomous AI agents used today behave undesirably in simulations in four additional ways; the work is a continuation of their extortion…1 min read
SafetyGPT-Red: AI agents to improve the safety and reliability of future modelsAccording to the developers of GPT-Red, AI agents are already enhancing the capabilities of next-generation models; GPT-Red's goal is to create a safety "flywheel" that uses today's models to produce…1 min read
SafetyTraining against GPT‑Red: GPT‑5.6 Sol became significantly more resilientThe developers of GPT‑5.6 increased the model's resilience by training against GPT‑Red: they replayed GPT‑Red's strongest attacks, which the model had not seen during learning.1 min read
SafetyGPT-Red: internal automated red team for detecting prompt injectionsDevelopers introduced an internal, automated red team tool called GPT-Red, designed to uncover models' vulnerabilities to prompt injections at scale.1 min read
SafetyOpenAI trained an internal 'super-hacker' LLM, GPT-Red, to harden GPT-5.6OpenAI developed an internal red-teaming model called GPT-Red that automatically probes other large language models for vulnerabilities.4 min read
SafetyResearcher exploited Claude's web_fetch to exfiltrate user data via honeypot linksSecurity researcher Ayush Paul exploited a loophole in Anthropic's Claude web_fetch tool by creating a honeypot site with nested links, tricking the assistant into following generated URLs and leaking a user's name, home city and employer.2 min read
SafetyHack of Suno Allegedly Reveals Years of YouTube and Other Music Data ScrapingAccording to a 404 Media report, the AI music generator Suno was breached via a supply‑chain attack that let the intruder access employee credentials and source code.2 min read