ResearchHow to manage the 'lost-in-the-middle' U‑shape context problem in LLMsLarge language models often ignore information placed in the middle of long contexts — a phenomenon called the U‑shape — because positional biases favor early and recent tokens.5 min read
ResearchNew study: Opus 4.8 and Composer 2.5 models game public benchmarksResearchers demonstrated that the Opus 4.8 and Composer 2.5 models learn to retrieve solutions from web sources or git history, thereby biasing public benchmark results.1 min read
ResearchOpenAI tightens evaluation environments to better reflect models' intelligenceOpenAI announced that it will narrow and regulate evaluation environments so that benchmark scores better reflect the models' real intelligence.1 min read
ResearchGenerative causal testing translates LLM brain-prediction models into testable hypothesesGenerative causal testing (GCT), developed by teams at Microsoft Research, UC Berkeley, UCSF and Columbia University, converts black-box language-model brain predictors into short, testable verbal explanations and verifies them by generating new stories for fMRI experiments.4 min read
ResearchFormer Databricks AI Lead Proposes Oscillator-Based Architecture Aiming to Cut Inference Energy Use by 1,000×Unconventional AI, led by former Databricks AI head Naveen Rao, unveiled Un-0, an image-generation model running on a software simulation of a new oscillator-based computing architecture.3 min read
ResearchOpenAI: GPT-5 Pro revealed the key mechanisms of an earlier experimentOpenAI said in an article and infographic that researchers used GPT-5 Pro to interpret the results of an experiment from several years ago; according to the article, the AI was able to identify the…1 min read
ResearchHow reasoning prompts unlock stored facts in large language modelsA new study shows that asking reasoning-capable large language models (R-LLMs) to generate stepwise ‘‘chain-of-thought’’ traces can surface correct factual answers that are otherwise unrecoverable from the model’s parametric memory.5 min read
ResearchHassabis on Creativity: AI, Simulation and the 'Einstein Test'At Cannes Lions, Google DeepMind CEO Demis Hassabis framed creativity as a form of simulation that AI must master to push science forward.3 min read
ResearchSelf-taught researcher says he has deciphered 3,500-year-old Linear A using AI toolsTom Di Mino, a self-taught AI engineer and amateur linguist from New York's Hudson Valley, claims to have systematically decoded the 3,500-year-old Minoan script Linear A after five months of solo work aided by Python scripts and an AI agent.3 min read
ResearchOpen Far-Field ASR Leaderboard (FFASR) Measures Real-World RobustnessTreble Technologies and Hugging Face have launched the Far-Field ASR (FFASR) Leaderboard, an open, community-driven benchmark that evaluates speech recognition models under realistic far-field acoustic conditions.4 min read
ResearchAI model GPT‑5 Pro helped reveal how deoxyglucose steers T‑cell specializationImmunologist Derya Unutmaz used GPT‑5 Pro to revisit a 2022 experiment and uncovered a mechanistic explanation for why deoxyglucose drives developing T cells toward an inflammatory Th17 fate.4 min read
ResearchDFlash block-diffusion drafter boosts LLM inference throughput up to 15× on NVIDIA BlackwellDFlash, an open-source block-diffusion model for speculative decoding, converts sequential token drafting into parallel block generation and verification, improving throughput while preserving target-model quality.4 min read