ResearchAI-driven single-cell analysis reveals genetic and proteomic heterogeneity in clear cell renal cell carcinomaResearchers led by Horváth Péter’s Lendület Mikroszkópos Képelemzés és Gépi Tanulás Kutatócsoport used artificial intelligence and an automated single-cell facility to measure both gene activity and protein composition of tumor clones from clear cell renal cell carcinoma at single-cell resolution.3 min read
ResearchAssessment of GLM 5.3: a cheaper tool for cybersecurity protectionResearchers evaluated the GLM 5.3 language model's cybersecurity capabilities and found that its lower operating costs could make it useful for defensive security tasks.1 min read
ResearchEval harness shows LLMs' highest confidence often corresponds to incorrect answersAn evaluation harness that scores model outputs against synthetic ground truth reveals that qualitative review—judging whether outputs merely 'sound right'—misses systematic, high-confidence errors in LLM-assisted enterprise tools.5 min read
ResearchNine benchmark-driven questions Epoch AI uses to probe the societal and economic effects of advancing AIEpoch AI’s research team outlines nine broad questions they consider central to understanding how advances in AI capabilities will affect economies and society.7 min read
ResearchHow AI Is Changing Personal Communication and Shaping the Next Generation’s SkillsGenerative AI increasingly composes, edits and interprets personal messages, speeding up interactions while taking over situations where people previously learned to express themselves.5 min read
ResearchAI models produce independent mathematical counterexamples; European access and competitiveness at riskRecent large AI models from OpenAI and Anthropic have independently produced counterexamples to long-standing mathematical conjectures, prompting excitement about AI's growing role in research.3 min read
ResearchCommunity Hackathon Reproduced Over 2,200 ICML 2026 Papers; Hundreds of Claim-Level Results ReleasedBetween July 15 and August 2, 2026, a community-driven hackathon produced 6,816 reproducibility logbooks covering 2,226 of the 6,341 accepted ICML 2026 papers.6 min read
ResearchPost-training's Two Pillars: Reinforcement Learning and Supervised Fine-TuningModern post-training of large language models relies primarily on two complementary techniques: reinforcement learning (RL) and supervised fine-tuning (SFT).6 min read
ResearchRecall, not encoding, limits factual accuracy in frontier LLMsA behavioral study using a new benchmark, WikiProfile, finds that state-of-the-art large language models (LLMs) increasingly store facts in their parameters but fail to retrieve them reliably.5 min read
ResearchEnterprises Find Context Failures in AI Are Widespread — Semantic Layers Reveal Problems More Than They Fix ThemA July 2026 VentureBeat Pulse Research survey of 101 enterprises finds that 68% experienced confident-but-wrong AI agent answers traced to missing or inconsistent business context in the past six months, with repeat failures (37%) outnumbering single incidents (32%).6 min read
ResearchReal-time, expert-level AI for video medical consultations demonstrated in simulated studyResearchers present AMIE (Video), a real-time, multi-agent AI system that perceives audio-visual clinical cues, guides virtual physical exams, and performs diagnostic reasoning.5 min read
ResearchAnthropic's Claude Rapidly Raised the Proven Proportion of Zeta Zeros on the Critical LineAnthropic ran an unreleased model of Claude on the Riemann zeta zeros and significantly increased the proportion provably on the critical line, from 41.6% to 67.2% in about 36 hours.3 min read