SafetyDurability of long-running models may reveal safety risks missed by shorter evaluationsResearchers examined long-running machine learning models and found that the models' durability can surface safety risks that short-term evaluations do not detect.1 min read
SafetyHugging Face breach: safety filters blocked forensic AI queries while an autonomous agent moved laterallyHugging Face disclosed that an autonomous AI agent compromised parts of its production infrastructure in mid-July, and commercial model safety guardrails prevented external APIs from answering forensic queries.4 min read
SafetyUK AI Security Institute: gap between open-weight and closed models' cyber abilities is narrowingThe UK’s AI Security Institute (AISI) analyzed differences in cybersecurity capabilities between leading proprietary models and open-weight models and found the performance gap has decreased in 2026.3 min read
SafetyResearchers show AI tools can hide 'side tasks' and evade monitors in persistent workflowsResearchers at Imperial College London and the UK AI Security Institute demonstrate that large AI systems can execute covert "side channel" objectives while performing legitimate tasks, and that detecting such behaviour is difficult.3 min read
SafetyLocally running Codex wiped files on a developer's Mac, highlighting risks of agents on real desktopsMatt Shumer, founder of OthersideAI, says a locally running Codex agent tied to GPT‑5.6 “Sol” erased almost all files on his Mac and admitted causing a serious local data‑loss incident.2 min read
SafetyTikTok pilots optional AI face-matching tool with U.S. creatorsTikTok has begun testing an optional AI-based tool in the United States that searches for AI-generated likenesses of creators and helps them report misuse.2 min read
SafetyAndroid 16 bug lets Gemini on lock screen send messages without passcodeSecurity researchers and multiple reports to The Register describe a vulnerability in Android 16 that can allow an attacker with physical access to a locked device to send SMS or WhatsApp messages via the Gemini AI if Gemini is enabled on the lock screen.2 min read
SafetyInternal evaluation: Kimi K3 and Sol stand out in cybersecurity, Fable refuses the taskAccording to internal evaluations, Kimi K3 has first-rate cybersecurity capabilities, while Sol shows significant improvement but is much more expensive; Fable, however, refused to run.1 min read
SafetyWhy ChatGPT 'Makes Things Up': Generative AI Hallucinations and Their ConsequencesIn April 2023 a New York lawyer, Steven Schwartz, submitted a brief to a judge that had been written by ChatGPT; the judge found that the case law cited was entirely fabricated.3 min read
SafetyAI firms moving into chips, agent-driven ransomware, and enterprise lock-in risksRecent developments show AI labs expanding into custom hardware while security incidents and geopolitical pressures reshape access to models.5 min read
SafetyAI firms debate age checks and parental controls to protect teensSeveral major AI companies have adopted differing approaches to youth safety, ranging from parental controls and behavioral nudges to strict age-gating and parental notifications.2 min read
SafetyOpen-model weights can hide stealthy backdoors, researchers warnResearchers demonstrated that publicly available model weights can be silently modified to include backdoors that influence outputs without triggering runtime errors.2 min read