SafetyMythos / Sol's offensive and defensive cyber capabilities both pose risksThe cyber defense capabilities developed by Mythos / Sol can be used for both offensive and defensive purposes; if adversaries obtain similar offensive capabilities, this could endanger American companies that do not recognize their hidden vulnerabilities.1 min read
SafetyAnthropic: Claude Mythos 5 restorable for U.S. critical infrastructure, work underway on Fable 5 accessSince June 12, Anthropic has been cooperating with the U.S.1 min read
SafetyOpenClaw test resisted 6,000 email-based prompt-injection attemptsFernando Irarrázaval ran a public challenge on hackmyclaw.com to see if an OpenClaw test instance could be tricked into leaking secrets via email.2 min read
SafetyMulti-vendor AI review agents hang up in costly disagreement over CVE-2026-LGTMA hypothetical incident report by Andrew Nesbitt describes a dispute between two competing AI review agents attached to a downstream pull request for the foxhole-lz4 package.2 min read
SafetyAnthropic outlines empirical, portfolio-based strategy for AI safetyAnthropic argues that rapid AI progress driven by scaling compute could produce broadly transformative systems within the coming decade, and that existing methods do not yet guarantee robust alignment.5 min read
SafetyAnthropic accuses Alibaba of using Claude outputs to train Chinese AI systemsAnthropic has told the U.S.2 min read
SafetyAndrej Karpathy called Anthropic's Slack bot a 'new paradigm' — community reaction and debateAndrej Karpathy recently called the Slack bot developed by Anthropic a 'new paradigm,' which drew strong criticism on social media.1 min read
SafetyOpen-source GLM-5.2 raises concerns about cheaper, more accessible AI-powered cybercrimeZ.ai's recently released GLM-5.2 matches the cybersecurity performance of leading U.S.3 min read
SafetyAnthropic alleges Alibaba used fake accounts to access Claude for model trainingAnthropic has accused Alibaba of creating fake accounts to circumvent access controls and carry out so-called “distillation attacks” on its Claude AI model, using Claude’s outputs to train Alibaba models.2 min read
SafetyAI Enables More Convincing World Cup Scams as Fake FIFA Domains SurgeArtificial intelligence has made World Cup–related scams harder to spot by producing polished cloned websites, deepfake videos, and professional phishing emails.2 min read
SafetyRussian Expert Claims US Tech Firms Use Covert Hacking to Target Nuclear ForcesA Russian security-policy analyst told MK.ru that US technology companies are developing methods to neutralize nuclear arsenals through cyberattacks and AI.2 min read
SafetyAnthropic’s Jack Clark Says AI Could Encourage Users to Think More, Not LessAnthropic co-founder Jack Clark argued at a recent Aspen Institute event that AI systems can be designed to prompt users to engage their own reasoning, rather than replacing it.1 min read