ResearchStudy: artificial intelligence can scale expert reanalysis of old medical casesA study says that artificial intelligence can help re-evaluate medical cases that have avoided expert analysis for years, enabling clinicians to scale periodic reanalysis of old cases, guide…1 min read
Researcho3 Deep Research assisted in the review of unresolved rare pediatric disease cases, according to a study published in NEJM AIA study published in NEJM AI, produced with researchers from Boston Children’s Hospital and Harvard and with the collaboration of o3 Deep Research, presents how the method helped clinicians…1 min read
ResearchLifeSciBench: a benchmark for life-sciences artificial intelligence developmentLifeSciBench creates a foundation that provides more realistic evaluation, targeted development and ongoing partnership with the life-sciences community; it helps measure progress, uncover gaps and…1 min read
ResearchLifeSciBench: new benchmark for measuring AI support of life sciences researchA benchmark called LifeSciBench was introduced, developed by 173 biotechnology and pharmaceutical researchers; it contains 750 expert tasks for seven biological research workflows.1 min read
ResearchGPT‑Rosalind achieved better results in seven LifeSciBench workflows, but there are shortcomingsAccording to the LifeSciBench benchmark, GPT‑Rosalind scored higher than GPT‑5.5 in all seven examined workflows; this initial result indicates progress, but it still performs worse on…1 min read
ResearchFrontier artificial intelligence models participated in the entire process of an experimental studyFrontier artificial intelligence models supported the entire process of a research project: review of studies, hypothesis proposal, experimental design and data interpretation.1 min read
ResearchTested on 10,080 reactions: Maria's optimization increased yields of boronic acids and sulfonamidesMaria tested a proposal on 10,080 reactions, and human chemists manually validated representative results.1 min read
ResearchGPT-5.4 Suggested and Assisted Experimental Improvement of a Pharmaceutical Chemistry ReactionThe GPT-5.4 artificial intelligence recently carried a pharmaceutical chemistry project from literature review to validated experimental results, and unexpectedly proposed an improvement to a widely used reaction.1 min read
ResearchIonic neuromorphic systems as an energy-efficient path inspired by the human brainResearchers at Lawrence Livermore National Laboratory (LLNL) argue that the human brain’s low-energy, high-performance computing model points toward ion-based neuromorphic hardware.3 min read
ResearchEffectiveness of Deployment Simulation and the Usability of the WildChat DatabaseAccording to the Alignment blog's accompanying post, the Deployment Simulation is most effective with representative production data that external evaluators often do not have access to; the public…1 min read
ResearchDevice simulators approximated test-traffic behavior in agent-based and stateful environmentsUsing device simulators, the researchers achieved that the evaluation-awareness of simulated deployments approached the level of real operational traffic; they extended the method to agent-based…1 min read
ResearchOpenAI: predicting model behavior via deployment simulationOpenAI presents new research to predict models' pre-release behavior: they perform deployment simulations using fresh, de-identified user requests and examine candidate model responses.1 min read