ResearchThe co-failure ceiling: when multi-model orchestration stops improving accuracyA new study of 67 frontier models from 21 providers identifies a fundamental limit on multi-model orchestration: the co-failure ceiling, the share of prompts where every model fails simultaneously.4 min read
ResearchEpoch launches EBR-bench, finds little evidence of learning-from-experience; publishes new AI capability and vulnerability analysesEpoch AI introduced EBR-bench, a new benchmark that measures whether models improve on a complex board game through repeated play and reports little sign so far that models learn from experience.3 min read
ResearchSWE-Bench Pro audit conducted using model-based agents and a combination of five independent engineersThe SWE-Bench Pro audit was conducted using model-based investigative agents and individual reviews by five independent, experienced software engineers to examine a larger number of tasks in a…1 min read
ResearchThe evolution of coding models demands stricter, fairer and more reliable evaluationsResearchers and developers say that as coding models steadily improve, evaluation procedures must become more difficult, fairer and more reliable.1 min read
ResearchIndependent audit: SWE-Bench Pro unreliable due to 30% faulty tasksA recent independent audit found that SWE-Bench Pro, one of the most widely used AI coding benchmarks, does not reliably measure top-tier coding ability; the audit identified 30% of tasks as faulty,…1 min read
ResearchResearchers identify and manipulate a 'Golden Gate Bridge' feature inside Claude 3 SonnetAnthropic researchers published a paper describing how they located and modified internal activations—called features—in their Claude 3 Sonnet model, including a specific "Golden Gate Bridge" concept.2 min read
ResearchAudit finds roughly 30% of SWE-Bench Pro tasks are flawedOpenAI audited the SWE-Bench Pro coding benchmark and found substantial evaluation problems: an automated pipeline flagged 200 tasks (27.4%) as broken and a human annotation campaign labelled 249 tasks (34.1%) as broken.4 min read
ResearchAnthropic paper identifies 'J‑space' in Claude but links to consciousness remain disputedAnthropic published a paper describing a method called J‑lens that detects internal representations in Claude, labeling them "J‑space" and likening the phenomenon to aspects of human consciousness.2 min read
ResearchAssessing Concrete Tech Pathways in AI-Driven FuturismSome futurist claims propose that once AI research is automated, superhuman agents could rapidly invent technologies like nanotech, Dyson swarms, or near-light-speed probes.5 min read
ResearchCoordinated navigation reduces urban congestion: a large-scale routing-app experimentA six-month field experiment in 10 major US cities, described in a Nature Cities paper, tested a modified Google Maps routing that gently redirects a small share of trips away from recurring bottlenecks.4 min read
ResearchAnthropic finds a self-organized 'J-space' in Claude that hosts hidden internal reasoningAnthropic researchers report the emergence of a small internal zone in the Claude model’s neural network, dubbed the J-space, which encodes concepts the model can reason about without necessarily verbalizing them.4 min read
ResearchFixed Compute Budgets Skew AI Benchmark Results, UK Study FindsThe UK's AI Security Institute tested frontier models on seven benchmarks and found that fixed compute budgets baked into evaluations systematically understate model capabilities.2 min read