ResearchComparing Claude Code's Success by OccupationThe analysis compared the success rates of the Claude Code model across different occupations; on the metric that required the strictest, verifiable evidence (for example, committed code), every…1 min read
ResearchEconomic research: measurement framework to track Claude Code's scalingThe research was developed to track the scaling of Claude Code (AI system); it shows who uses it and for what purpose, how the value of tasks changes, and to what extent professional domain knowledge determines the success of sessions.1 min read
ResearchNew evaluation methods are needed to reliably measure modelsTejal Patwardhan, head of the frontier evals team, spoke with Andrew Mayne about how traditional benchmarks are becoming saturated or can be manipulated, so it is important to apply new, more comprehensive methods to measure and forecast models' progress.1 min read
ResearchGoogle and UC San Diego Build a 'Datacenter' from 2,000 Retired Pixel PhonesGoogle and the University of California, San Diego announced a project that clusters 2,000 decommissioned Pixel phones into a teaching datacenter by stripping them to their motherboards, installing Linux, and orchestrating them with Kubernetes.2 min read
ResearchAI simulation ranks Spain top favorite for 2026 World CupAn Innsbruck University researcher and his team used a two‑stage machine‑learning model trained on eight years of international matches plus market and ranking data to simulate the 2026 FIFA World Cup 100,000 times.2 min read
ResearchStartup attempts to recreate GTA 6 with AI before official releaseA Hyperecho founder, Hszü Ce-ven, has started a public experiment attempting to build a playable approximation of Grand Theft Auto VI using 'vibe coding' and a large language model.2 min read
ResearchArtificial intelligence-based model to predict drivers' accident riskArab researchers are developing an artificial intelligence–based model that can indicate, before sitting behind the wheel, the likelihood that a driver will cause an accident.1 min read
ResearchAnthropic's productivity data rekindles debate over recursive self‑improvementAnthropic reported that its model Claude now authors or co-authors about 80% of the company’s code as of May 2026, and presented data showing large gains in AI-assisted software productivity.5 min read
ResearchState-aligned Media Shapes LLM Output, Strongest Effect in ChineseResearchers from multiple universities found that state-affiliated media present in web training data skews large language model (LLM) outputs toward government-favorable views, especially when models respond in the language of a country with restricted press freedom.5 min read
ResearchAI Enables Monthly Mapping of Glacier Calving Fronts with Near-Human AccuracyResearchers at Friedrich-Alexander University trained an existing glacier-tracking model to identify calving fronts using minimal local data.3 min read
ResearchOpus wrote us a VM and then Mythos verified itOpus wrote us a VM and then Mythos verified it1 min read
ResearchComparison of Claude Fable 5's performance with different modelsThe comparison of the artificial intelligence model named Claude Fable 5 with various models presents its performance and behavioral differences; the analysis helps developers and users with model…1 min read