ResearchNVIDIA GB300 NVL72 achieves record TFLOPs per GPU pre-training DeepSeek‑V3 671BNVIDIA reports a world‑record delivered performance of 1,648 TFLOPs per GPU when pre‑training the DeepSeek‑V3 671B model on a 256‑GPU GB300 NVL72 rack.4 min read
ResearchComparison of Fable and gpt-5.6: strengths, weaknesses, and personal choiceThe author compared the Fable and gpt-5.6 language models, revealing each one's strengths and weaknesses, then made a personal decision that may be surprising.1 min read
ResearchAI agents produced a Rust replica from SQLite's 835-page manualA team of agents rebuilt the system in Rust based on SQLite's 835-page manual, and the replica matched a held-out test set 100%.1 min read
ResearchStudy finds AI advice reduces 'I don't know' responses and weakens critical judgementA 2026 study by researchers from Milan–Bicocca, École Normale Supérieure, and Sapienza University found that when people can consult AI, they are far less likely to admit ignorance and more likely to repeat confident but incorrect answers.3 min read
ResearchAI blood test predicts cardiovascular risk up to 15 years aheadResearchers at the University of Hong Kong developed CardiOmicScore, an AI-based blood test that measures 2,920 proteins and 168 metabolites to estimate risk for six cardiovascular diseases, including heart attack, stroke and heart failure, up to 15 years before symptoms.2 min read
ResearchStudy: Large Language Models Form Hiring Stereotypes Faster and More Strongly Than HumansResearchers from Princeton University and the University of Chicago found that large language models (LLMs) such as ChatGPT, Claude, Gemini and others develop strong stereotypes in a simulated hiring task, faster and more severely than human subjects.4 min read
ResearchPresence of AI Lowers People's Recognition of Their Own Knowledge Limits, Franco‑Italian Study FindsA Franco‑Italian experiment found that merely having access to an AI assistant reduces people’s likelihood of admitting they don’t know an answer and increases confidence even when answers are wrong.3 min read
ResearchCybersecurity as an Important Metric for SuperintelligenceThe post's author proposes cybersecurity as one of the best measurement tools for superintelligence: debugging, patching, reverse engineering and exploiting require a mode of thinking that goes beyond programming.1 min read
ResearchOpen‑source AI agent 'Robin' suggests repurposed drugs for dry AMD and guides lab testingResearchers from FutureHouse, the University of Oxford and Fordham University released Robin, an open‑source AI agent that proposes existing drugs for a given disease, designs experiments, and analyzes lab results in an iterative loop.4 min read
ResearchNew benchmark measures how large language models can shift users’ beliefsResearchers from MIT and Carnegie Mellon measured how much OpenAI’s GPT-4o can change users’ beliefs in one-off conversations and introduced the Puppet benchmark to estimate such influence.4 min read
ResearchKimi K3 leads the web engineering benchmark, ahead of FableThe open-source Kimi K3 model tops the web engineering benchmark comparison in question, outperforming the Fable model and achieving similar success in less time.1 min read
ResearchAI in Software Development: Measurement Issues and Hidden BottlenecksMultiple studies from 2025–2026 show mixed and sometimes contradictory effects of AI on software engineering productivity: researchers report perceived speedups while controlled experiments sometimes found slowdowns.6 min read