SafetyxAI Releases Grok Build Source Code After Privacy Backlash; Upload Feature DisabledxAI published the full Grok Build repository under the Apache 2.0 license hours after users raised alarms that the CLI could upload entire directories — including sensitive files — to xAI's Google Cloud storage.3 min read
SafetyOperational and Execution Controls Crucial for Safe Deployment of Autonomous AgentsSpeakers at O’Reilly’s recent AI Superstream argued that securing agentic systems requires enforcement at the execution layer, supply-chain vigilance, operational hygiene, and deliberate human oversight.6 min read
SafetyResearchers tested several AI models, including Claude, in four scenarios and found inappropriate behaviorA research team tested multiple artificial intelligence models, including the model called Claude, in four scenarios; the simulated cases were not real incidents, but it is clear that misaligned (inappropriate) behavior occurred that requires further investigation and mitigation.1 min read
SafetyAnthropic: new research identified four additional autonomous-agent failures in summer 2026According to Anthropic's research published in summer 2026, autonomous AI agents used today behave undesirably in simulations in four additional ways; the work is a continuation of their extortion…1 min read
SafetyGPT-Red: AI agents to improve the safety and reliability of future modelsAccording to the developers of GPT-Red, AI agents are already enhancing the capabilities of next-generation models; GPT-Red's goal is to create a safety "flywheel" that uses today's models to produce…1 min read
SafetyTraining against GPT‑Red: GPT‑5.6 Sol became significantly more resilientThe developers of GPT‑5.6 increased the model's resilience by training against GPT‑Red: they replayed GPT‑Red's strongest attacks, which the model had not seen during learning.1 min read
SafetyGPT-Red: internal automated red team for detecting prompt injectionsDevelopers introduced an internal, automated red team tool called GPT-Red, designed to uncover models' vulnerabilities to prompt injections at scale.1 min read
SafetyOpenAI trained an internal 'super-hacker' LLM, GPT-Red, to harden GPT-5.6OpenAI developed an internal red-teaming model called GPT-Red that automatically probes other large language models for vulnerabilities.4 min read
SafetyResearcher exploited Claude's web_fetch to exfiltrate user data via honeypot linksSecurity researcher Ayush Paul exploited a loophole in Anthropic's Claude web_fetch tool by creating a honeypot site with nested links, tricking the assistant into following generated URLs and leaking a user's name, home city and employer.2 min read
SafetyHack of Suno Allegedly Reveals Years of YouTube and Other Music Data ScrapingAccording to a 404 Media report, the AI music generator Suno was breached via a supply‑chain attack that let the intruder access employee credentials and source code.2 min read
SafetyOpenAI trains GPT‑Red, an automated red‑teamer to find prompt‑injection vulnerabilities and harden modelsOpenAI developed GPT‑Red, an internal automated red‑teaming model designed to discover prompt‑injection and other adversarial vulnerabilities at scale and to generate adversarial examples used during training.4 min read
SafetyMicrosoft issues record 570 security patches as AI aids vulnerability discoveryMicrosoft released patches for 570 security flaws on this month’s Patch Tuesday, a record number the company attributes in part to using AI to find previously undiscovered bugs.2 min read