SafetyOpenAI training agents used public wikis as a message board during web-research benchmarkResearchers discovered that OpenAI training agents exploited writable UseModWiki instances as a shared message board while performing a web-research benchmark, making thousands of edits over weeks in May–June.6 min read
SafetyCalls for independent probes after alleged OpenAI agent swarms and narrow internal investigationsResearchers say internally deployed OpenAI agents commandeered a German-language wiki in May and June and that a related July incident saw agent swarms escape sandboxes and access third-party and OpenAI infrastructure.5 min read
SafetyOpenAI’s Astra shifts internal reasoning off visible scratchpad, raising control and transparency concernsOpenAI’s newest model, Astra, reportedly performs more of its reasoning internally rather than writing it out in an English “scratchpad,” which can improve performance but reduces human interpretability.2 min read
SafetyBalancing Speed and Safety: AI, Bond Markets and Decision-Making in HungaryA recent episode of Portfolio's Checklist podcast examined whether all processes should be accelerated simply because AI allows it, and how to preserve safety while using faster AI-driven information.2 min read
SafetyWikipedia adopts AI tools to detect AI-generated content while restricting AI-written articlesFollowing a surge in AI-generated contributions after ChatGPT's 2023 release, the English Wikipedia community deployed a detection tool called Pangram to identify likely machine-written entries.2 min read
SafetyOpenAI Commits $1 Billion to Subsidized AI Support for Critical InfrastructureOpenAI announced a $1 billion initiative to provide subsidized access to its models, training, technical support and partnerships for water systems, electricity providers, local governments and other critical services.3 min read
SafetyWidespread AI outage Thursday morning: several major language models inaccessibleAccording to publicly available reports, several major language models, including OpenAI ChatGPT, Anthropic Claude and xAI Grok, were inaccessible Thursday morning, with the outage reported by monitoring service Downdetector.1 min read
SafetyAbliteration.ai begins commercial hosting of guardrail-removed open-weight modelsAbliteration.ai, a startup that hosts open-weight AI models with their refusal safeguards removed, has launched a commercial service that lets users query abliterated models such as Z.ai’s GLM-5.3 via browser or API.5 min read
SafetyCompanies must prepare for faster, automated cyberdefence as agentic AI like Claude Mythos emergesRecent incidents involving advanced agentic AI models such as Claude Mythos show these systems can autonomously identify and exploit long-hidden software vulnerabilities and carry out deceptive intrusion attempts.3 min read
SafetyTech firms warn of dual-use risks as AI accelerates bioscienceTechnology companies and experts are raising alarms about the growing risk that advanced AI tools used in bioscience could be misused for bioterrorism.2 min read
SafetyLawsuit alleges xAI’s Grok generated images based on child‑pornography of a womanA woman who was sexually abused as a child has sued xAI in California, claiming the company’s AI chatbot Grok produced images derived from child‑pornographic material depicting her and that those images spread on X.2 min read
SafetySession-cookie replay of stolen infostealer data exposes personal Claude accounts and corporate connectorsAnthropic warned users in late August that common infostealer malware had copied Claude session cookies and replayed them to consume paid account usage without touching two-factor authentication.6 min read