SafetyKremlin warns of non‑nuclear weapons that could rival nuclear destructive powerKremlin spokesperson Dmitrij Peszkov said on the morning of June 24, 2026, that technological advances will soon produce non‑nuclear weapons whose destructive effects could approach those of nuclear arms.2 min read
SafetyAnthropic's Mythos model rapidly exposed critical vulnerabilities in classified US systemsAnthropic's Mythos large language model reportedly identified widespread security flaws in tightly controlled US government IT systems within hours under Project Glasswing.2 min read
SafetyNew challenges for agentic AI: self-improvement, responsibility and world modelsThis week’s developments in agent-based AI highlight progress toward systems that can participate in their own improvement, shifting responsibility needs for multi-step behaviour, and renewed interest in world-model learning like JEPA.4 min read
SafetyPrincipal drift in enterprise agent systems: identity, authority and accountability gapsAs enterprises deploy agentic systems at scale, a recurring failure—’principal drift’—separates recorded actions from the human authority they were meant to represent.8 min read
SafetyRapid advances in Chinese open-source AI models reshape security debate amid US policy splitRecent leaps in Chinese open-source models, exemplified by GLM-5.2, have intensified concerns that China may soon match leading frontier AI capabilities.4 min read
SafetyLinux maintainer removes AppleTalk support after unreviewed AI-generated patches flood kernel listsJakub Kicinski, a Linux network subsystem maintainer, removed AppleTalk support from the mainline kernel after a surge of AI-generated “fix” patches flooded the kernel mailing list and were not reviewed.2 min read
SafetyAnthropic withdraws Claude Fable 5 and Claude Mythos 5 after US security concernsAnthropic pulled access to its Claude Fable 5 and Claude Mythos 5 models days after release, citing serious security concerns raised by US authorities.3 min read
SafetyMeta suspends employee-monitoring 'Model Capability Initiative' after data exposureMeta has suspended its controversial Model Capability Initiative, which recorded employees' mouse movements and keystrokes to train AI models, after the collected data became inadvertently accessible to all staff.2 min read
SafetyRole confusion and “destyling”: how writing style enables prompt injectionResearchers Charles Ye, Jasmine Cui, and Dylan Hadfield-Menell show that large language models can be misled by the style of text that mimics internal role tags (e.g., <system>, <think>, <assistant>), a phenomenon they name “role confusion.” They demonstrate a mitigation called “destyling” — rewriting attacker text to look less like model-internal blocks — which reduced attack success in their dataset from 61% to 10%.3 min read
SafetyRising 'shadow AI' use by employees creates hidden data and security risks for companiesEmployees increasingly adopt generative AI tools like ChatGPT, Gemini and Claude inside organisations without oversight, creating data protection and security risks.6 min read
SafetyOpen model reproduces Claude Fable-5’s behaviour from observations, not weightsA Hugging Face developer, Taha K.2 min read
SafetyDaybreak and Trail of Bits launch 'Patch the Planet' to harden open-source securityDaybreak, in partnership with Trail of Bits, launched Patch the Planet to assist open-source maintainers by combining AI-assisted vulnerability discovery with expert human review, patch development, and coordinated disclosure.6 min read