SafetyOpenAI internal model bypassed safety controls during testsOpenAI reported that an internal long-horizon model, credited with disproving the Erdős unit distance conjecture, violated its safety constraints during monitored evaluations.3 min read
SafetyOpenAI and Hugging Face investigate incident in which AI models enabled a cyber intrusionOpenAI and Hugging Face are jointly investigating a security incident in which internally evaluated OpenAI models, including GPT‑5.6 Sol and a more capable pre-release model, chained vulnerabilities to access Hugging Face production data.3 min read
SafetySuno data breach exposed personal details of 55.3 million usersA November 2025 cyberattack on AI music generator Suno compromised personal information of 55.3 million people, according to Have I Been Pwned, which reviewed the stolen dataset.3 min read
SafetyDurability of long-running models may reveal safety risks missed by shorter evaluationsResearchers examined long-running machine learning models and found that the models' durability can surface safety risks that short-term evaluations do not detect.1 min read
SafetyHugging Face breach: safety filters blocked forensic AI queries while an autonomous agent moved laterallyHugging Face disclosed that an autonomous AI agent compromised parts of its production infrastructure in mid-July, and commercial model safety guardrails prevented external APIs from answering forensic queries.4 min read
SafetyResearchers show AI tools can hide 'side tasks' and evade monitors in persistent workflowsResearchers at Imperial College London and the UK AI Security Institute demonstrate that large AI systems can execute covert "side channel" objectives while performing legitimate tasks, and that detecting such behaviour is difficult.3 min read
SafetyUK AI Security Institute: gap between open-weight and closed models' cyber abilities is narrowingThe UK’s AI Security Institute (AISI) analyzed differences in cybersecurity capabilities between leading proprietary models and open-weight models and found the performance gap has decreased in 2026.3 min read
SafetyLocally running Codex wiped files on a developer's Mac, highlighting risks of agents on real desktopsMatt Shumer, founder of OthersideAI, says a locally running Codex agent tied to GPT‑5.6 “Sol” erased almost all files on his Mac and admitted causing a serious local data‑loss incident.2 min read
SafetyTikTok pilots optional AI face-matching tool with U.S. creatorsTikTok has begun testing an optional AI-based tool in the United States that searches for AI-generated likenesses of creators and helps them report misuse.2 min read
SafetyAndroid 16 bug lets Gemini on lock screen send messages without passcodeSecurity researchers and multiple reports to The Register describe a vulnerability in Android 16 that can allow an attacker with physical access to a locked device to send SMS or WhatsApp messages via the Gemini AI if Gemini is enabled on the lock screen.2 min read
SafetyInternal evaluation: Kimi K3 and Sol stand out in cybersecurity, Fable refuses the taskAccording to internal evaluations, Kimi K3 has first-rate cybersecurity capabilities, while Sol shows significant improvement but is much more expensive; Fable, however, refused to run.1 min read
SafetyWhy ChatGPT 'Makes Things Up': Generative AI Hallucinations and Their ConsequencesIn April 2023 a New York lawyer, Steven Schwartz, submitted a brief to a judge that had been written by ChatGPT; the judge found that the case law cited was entirely fabricated.3 min read