ResearchAnthropic's 'J-lens' Reveals a Hidden J-space Inside Claude Opus 4.6Anthropic developed the Jacobian lens (J-lens) to inspect intermediate representations inside Claude Opus 4.6 and discovered a previously unseen J-space that surfaces words related to likely future outputs.5 min read
ResearchBartosz (@nasqret) mathematician solves previously unsolvable problems with help of GPT-5.6Bartosz, better known online as @nasqret, is a mathematician who uses the GPT-5.6 language model to solve mathematical problems previously considered unsolvable; this promises new methods and faster…1 min read
ResearchIterative synthetic financial headline generation using NVIDIA NeMo and NemotronResearchers developed an iterative pipeline to create a 502,536‑headline synthetic financial news corpus across 13 categories using NVIDIA NeMo Data Designer, NeMo Curator and Nemotron models.6 min read
ResearchThe co-failure ceiling: when multi-model orchestration stops improving accuracyA new study of 67 frontier models from 21 providers identifies a fundamental limit on multi-model orchestration: the co-failure ceiling, the share of prompts where every model fails simultaneously.4 min read
ResearchEpoch launches EBR-bench, finds little evidence of learning-from-experience; publishes new AI capability and vulnerability analysesEpoch AI introduced EBR-bench, a new benchmark that measures whether models improve on a complex board game through repeated play and reports little sign so far that models learn from experience.3 min read
ResearchSWE-Bench Pro audit conducted using model-based agents and a combination of five independent engineersThe SWE-Bench Pro audit was conducted using model-based investigative agents and individual reviews by five independent, experienced software engineers to examine a larger number of tasks in a…1 min read
ResearchThe evolution of coding models demands stricter, fairer and more reliable evaluationsResearchers and developers say that as coding models steadily improve, evaluation procedures must become more difficult, fairer and more reliable.1 min read
ResearchIndependent audit: SWE-Bench Pro unreliable due to 30% faulty tasksA recent independent audit found that SWE-Bench Pro, one of the most widely used AI coding benchmarks, does not reliably measure top-tier coding ability; the audit identified 30% of tasks as faulty,…1 min read
ResearchResearchers identify and manipulate a 'Golden Gate Bridge' feature inside Claude 3 SonnetAnthropic researchers published a paper describing how they located and modified internal activations—called features—in their Claude 3 Sonnet model, including a specific "Golden Gate Bridge" concept.2 min read
ResearchAudit finds roughly 30% of SWE-Bench Pro tasks are flawedOpenAI audited the SWE-Bench Pro coding benchmark and found substantial evaluation problems: an automated pipeline flagged 200 tasks (27.4%) as broken and a human annotation campaign labelled 249 tasks (34.1%) as broken.4 min read
ResearchAnthropic paper identifies 'J‑space' in Claude but links to consciousness remain disputedAnthropic published a paper describing a method called J‑lens that detects internal representations in Claude, labeling them "J‑space" and likening the phenomenon to aspects of human consciousness.2 min read
ResearchAssessing Concrete Tech Pathways in AI-Driven FuturismSome futurist claims propose that once AI research is automated, superhuman agents could rapidly invent technologies like nanotech, Dyson swarms, or near-light-speed probes.5 min read