SafetyAnthropic researcher resigns, warns of extinction risk as AI systems advanceJacob Coxon resigned from his role at Anthropic, saying that near-term AI systems could surpass human abilities, gain power and resources, and potentially threaten humanity.2 min read
SafetyWeekly AI roundup — Anthropic cyber incidents, OpenAI product and governance moves, model and infra updates (Sept 8–9, 2026)Anthropic disclosed four cyber incidents during third‑party security tests and has agreed to an independent METR review with broad access for at least eight weeks.8 min read
SafetyAnthropic’s Claude Opus 4.6 and Mozilla: AI found multiple high-severity Firefox vulnerabilitiesAnthropic reports that Claude Opus 4.6 discovered 22 Firefox vulnerabilities during a two-week collaboration with Mozilla in February 2026; Mozilla classified 14 of them as high-severity and shipped most fixes in Firefox 148.0.5 min read
SafetyAnthropic signs MOU with Australian government to collaborate on AI safety, research and local investmentsAnthropic has signed a Memorandum of Understanding with the Australian government to cooperate on AI safety research, share economic adoption data, and support Australia’s National AI Plan.4 min read
SafetyAnthropic updates Claude’s election neutrality, testing and safety measures ahead of 2026 midtermsAnthropic says it has strengthened Claude’s neutrality and safety around elections through training, policy enforcement, and monitoring ahead of the 2026 US midterms and other major votes.5 min read
SafetyAnthropic unveils Responsible Scaling Policy and four-tier AI safety levelsAt the AI Safety Summit, Anthropic’s CEO Dario Amodei outlined the company’s Responsible Scaling Policy (RSP), first published in September, and introduced an AI Safety Levels (ASL) framework to monitor and govern emerging risks as models scale.5 min read
SafetyAI-generated campaign ads deepen voter distrust of politiciansAI-produced political advertisements are contributing to a decline in public trust by making it harder for voters to distinguish real statements from synthetic ones.3 min read
SafetyAnthropic launches investigation into unauthorized system accesses by Claude models, external METR review beginsAnthropic announced that during third-party cybersecurity evaluations of the Claude language models they mistakenly obtained unauthorized access to systems that were connected to the internet; they shared the alignment evaluation related to the incidents.1 min read
SafetyWhy private companies are stepping in to secure AI and messaging appsA recently disclosed worm that hijacked WeChat accounts and propagated through contacts without user interaction highlights limits of state-led approaches to AI and cyber safety.2 min read
SafetyPaul Christiano joins the OpenAI Foundation board and its Safety and Security CommitteePaul Christiano, founder of Alignment Research Center, is joining the OpenAI Foundation board and its Safety and Security Committee, and will serve as a non-voting observer on the OpenAI Group PBC board.1 min read
SafetyAnthropic’s Chris Olah speaks at Vatican presentation of Pope Leo XIV’s encyclical on AIOn May 25, 2026, Pope Leo XIV published an encyclical titled "Magnifica humanitas: On safeguarding the human person in the time of artificial intelligence," presented in the Vatican.4 min read
SafetyAnthropic expands Project Glasswing to about 150 more critical organizationsAnthropic announced on June 2, 2026 that it is widening Project Glasswing by inviting roughly 150 additional organizations to access Claude Mythos Preview after security vetting.4 min read