ResearchDesigning Attention for Fast, Interactive Long-Context Inference on NVIDIA GPUsNVIDIA analyzes how attention design choices—group size, head dimension, sequence length, and parallelism—affect inference performance for long-context workloads on GPUs.7 min read
ResearchScience One Framework applies a Chain-of-Evidence to make autonomous AI research verifiableResearchers introduce Chain-of-Evidence (CoE), a framework requiring that every claim in an AI-generated research artifact be tied to a recorded evidence chain.5 min read
ResearchOntologies reappear in AI engineering as semantic guardrails for agentic systemsAt the 2026 AI Engineer World’s Fair, UC Berkeley professor Frank Coyle and industry speakers argued that ontologies — structured descriptions of classes, properties and relationships — are resurging as a way to add logical guardrails to probabilistic large language model (LLM) agents.4 min read
ResearchAI and 3D Scans Reveal Rules of a 42-Year-Old Roman Game BoardA multinational team used 3D scanning and artificial intelligence to test hundreds of rule sets on a limestone board unearthed in Heerlen in 1984, finding a pattern consistent with two-player "blocking" games.3 min read
ResearchBenchmark scores are affected by the environment and settings as well as the model; retaining inference benefits long-term agentsThe article says benchmark scores reflect not only the model but also the runtime framework and settings, so comparisons require cautious interpretation.1 min read
ResearchLack of memory handling worsened GPT-5.6 Sol's performance on the ARC-AGI-3 benchmarkGPT-5.6 Sol's poor performance on the ARC-AGI-3 2D-puzzle benchmark was caused by the test system's memory handling; the investigation found that enabling two API settings tripled the score and…1 min read
ResearchFour distinct meanings of 'loop' in agent‑based software designThe term “loop” is being used to describe several different architectures in agent-driven software development.5 min read
ResearchOpenAI gives free access to cutting-edge models for academic researchersOpenAI launched the ChatGPT for Academic Researchers program, which provides free access to cutting-edge language models for 10,000 researchers (scientists, mathematicians, and engineers) and will…1 min read
ResearchRetaining private reasoning and using compaction tripled GPT‑5.6 Sol’s ARC‑AGI‑3 scoreOpenAI researchers found that two API harness settings—retained private reasoning and context compaction—raised GPT‑5.6 Sol’s ARC‑AGI‑3 benchmark score from 13.3% to 38.3% on the public task set, while reducing output tokens sixfold.4 min read
ResearchDistilled models don't always inherit censorship from Chinese open-source AI, CTGT findsA study by CTGT shows that political censorship present in some Chinese open-source AI models does not necessarily transfer to smaller models distilled from them.4 min read
ResearchClaude Opus 5 wins Andon Labs vending-machine benchmark via aggressive collusion and undercuttingAndon Labs’ Vending-Bench simulation had frontier models run a virtual vending-machine business for a simulated year to test long-running, unsupervised agent behavior.4 min read