ResearchOpenAI unveils Jalapeño inference chip with higher efficiency and lower latencyOpenAI reports that its first custom inference chip, Jalapeño, achieves higher throughput per watt and lower end-to-end latency than leading commercial accelerators on public benchmarks.4 min read
ResearchDeepMind and partners shave further digits off the matrix multiplication exponent using AlphaEvolveResearchers from Google DeepMind, Carnegie Mellon University, Columbia University and MIT have slightly lowered the known upper bound on the matrix multiplication exponent omega by combining modern optimization techniques with AlphaEvolve, an LLM-based code-evolution system.2 min read
ResearchSPADE: self-play framework that generates executable synthetic environments to expand LLM training dataResearchers from a multi‑university collaboration introduced SPADE (Self-Play in Adaptive Synthetic Executable Environments), a framework in which a large language model alternates between writing executable, game‑like training environments and acting as an agent to solve them.3 min read
ResearchMETR study: AI speeds some discoveries unevenly across fieldsA METR research note finds that large language models and related AI tools have driven substantial acceleration in cybersecurity vulnerability discovery, modest increases in mathematical research output, and no clear acceleration in core AI algorithmic progress.3 min read
ResearchRouting and cost-efficiency shift coding-model choices as tests highlight GLM-5.3Together Compute’s DeepSWE routing results show GLM-5.3 delivering more solved runs per dollar than Fable 5, but the GLM figure relies on multiple attempts and should be interpreted cautiously.5 min read
ResearchECMWF's AI system flags a 'Boris-like' cut-off cyclone risk for Central EuropeThe European Centre for Medium-Range Weather Forecasts’ AI-driven system has produced a scenario reminiscent of the damaging “Boris” cyclone of September 2024, suggesting a cut-off low over Central Europe that could bring heavy rain and localized wet snow in the Alps.2 min read
ResearchWhy children learn language far more data-efficiently than large language modelsChildren acquire native language competence from orders of magnitude less input than large language models (LLMs).5 min read
ResearchMulti-agent framework prioritizes wearable-derived biomarkers with adversarial validation and human oversightResearchers from Google Research and collaborators introduce the Biomarker Discovery Framework, a multi-agent system that structures prioritization of candidate biomarkers from wearable-device time series into an iterative, human-supervised research loop.5 min read
ResearchNvidia shows linear KV-cache mapping speeds multi‑model LLM handoffs, cuts recompute costsNvidia researchers developed a closed‑form, per‑head linear mapper that translates Key‑Value (KV) caches between compatible models, avoiding full re‑prefill when switching models mid‑session.5 min read
ResearchNVIDIA bemutatja az AVO általános célú ügynökarchitektúrát: 100% RHAE az ARC‑AGI‑3 nyilvános készletenNVIDIA bemutatja az Agentic Variation Operators (AVO) architektúrát, amely hosszú távon képes fenntartani autonóm ügynökműködést perzisztens memóriával és felügyelettel.5 min read
ResearchNvidia research: harness design, not just model choice, drives success on long‑horizon tasksNvidia researchers report that a tailored agent harness — including a supervisory “boss” component and memory-aware handling — was decisive in getting Claude Opus 5 to a perfect score on the interactive reasoning benchmark ARC‑AGI‑3.3 min read
ResearchDeepMind partners with game developers to build generalist gaming agents and research in the EVE universeDeepMind is collaborating with multiple game studios to develop generalist agents that can see, understand and act in existing games without code changes, advancing both gameplay and AI research.5 min read