ResearchOpen‑source AI agent 'Robin' suggests repurposed drugs for dry AMD and guides lab testingResearchers from FutureHouse, the University of Oxford and Fordham University released Robin, an open‑source AI agent that proposes existing drugs for a given disease, designs experiments, and analyzes lab results in an iterative loop.4 min read
ResearchKimi K3 leads the web engineering benchmark, ahead of FableThe open-source Kimi K3 model tops the web engineering benchmark comparison in question, outperforming the Fable model and achieving similar success in less time.1 min read
ResearchAI in Software Development: Measurement Issues and Hidden BottlenecksMultiple studies from 2025–2026 show mixed and sometimes contradictory effects of AI on software engineering productivity: researchers report perceived speedups while controlled experiments sometimes found slowdowns.6 min read
ResearchGPT-5.6 Finds Counterexample to Long-Open Benjamini–Hochberg Assumption in 90 MinutesWharton statistician Edgar Dobriban reports that a GPT-5.6 model produced a counterexample to a 20-year-old open case about the Benjamini–Hochberg procedure within 90 minutes.2 min read
ResearchResearchers used a diffusion-based AI to design burgers that trade taste, health and environmental impactResearchers at the Stanford University Living Matter Laboratory trained a diffusion-based generative model called BurgerAI on 2,216 Food.com hamburger recipes and 146 ingredients to generate burgers optimized for taste, nutrition and sustainability.3 min read
ResearchScore smoothing during training explains diffusion models' tendency to interpolateA paper presented at ICLR 2026 argues that the apparent "creativity" of diffusion models—their ability to produce novel images or molecular structures—follows from how neural-network training smooths the learned score function.4 min read
ResearchPublic AI Models Rapidly Solved Long‑standing Theoretical Problems and Self‑Verified Their WorkIn July 2026, OpenAI reported that GPT-5.6 Sol Ultra solved the Cycle Double Cover conjecture in under an hour by splitting into 64 agents that critiqued each other's proofs.2 min read
ResearchApple in Early Talks with PrismML on Running Compressed LLMs Locally on iPhoneApple is in preliminary discussions with PrismML, a Caltech spin‑off backed by Khosla Ventures, which says it has compressed large language models so they can run on iPhone 15 and newer devices.3 min read
ResearchNew benchmark Real World VoiceEQ measures human-quality aspects of voice AIReal World VoiceEQ is a new benchmark designed to assess how well voice AI systems capture and act on acoustic cues that transcripts omit—tone, emotion, speaker identity and background context.5 min read
ResearchPractical lessons from the Nemotron Model Reasoning Challenge: five ways teams improved reasoning with open modelsThe NVIDIA Nemotron Model Reasoning Challenge on Kaggle drew over 5,000 participants in roughly 4,000 teams, producing thousands of submissions and more than 1,000 discussion posts.6 min read
ResearchNVIDIA’s Ising Decoder Speeds and Improves Color Code Decoding for Quantum Error CorrectionNVIDIA released Ising Decoder ColorCode 1 Fast, a 3D CNN pre-decoder for triangular color codes that reportedly reduces logical error rates and shortens decoding time compared with the open-source Chromobius decoder.3 min read
ResearchAnthropic identifies a hidden 'J-space' of internal tokens inside large language modelsAnthropic reports finding a previously unseen internal representation space—called the J-space—in its Claude models, populated by tokens that do not appear in outputs but seem to track reasoning steps and states.3 min read