ResearchRoboReward: vision–language reward models for robot training close the gap with hand-crafted rewardsResearchers at Stanford University and UC Berkeley developed RoboReward, a family of vision–language reward models (4B and 8B parameters) and an accompanying dataset and benchmark for evaluating robot reward models.4 min read
ResearchAlibaba's SkillWeaver routes skills iteratively to cut token use and improve multi-step accuracyResearchers at Alibaba introduced SkillWeaver, a framework that decomposes complex prompts into sub-tasks, retrieves candidate skills, and composes them into an executable DAG.4 min read
ResearchComparison of Claude Fable 5 with Other Artificial Intelligence ModelsA comparison of Anthropic Claude Fable 5 with other AI models is presented; the study focuses on differences in performance, capabilities, and applicability.1 min read
ResearchGenesis Molecular AI’s PEARL advances protein–ligand structure prediction, challenges common benchmarksGenesis Molecular AI says its PEARL model can predict protein–ligand complexes with accuracy that frequently exceeds public models, handling induced-fit cases without long molecular dynamics runs.4 min read
ResearchVideo made about the efficiency of OpenAI models and the shortcomings of Sonnet 5A content creator made a video about why OpenAI models are so efficient, highlighting that Sonnet 5 has serious efficiency shortcomings.1 min read
ResearchFrom Chatbots to Agents: AI's Rapid Capability Surge and Changing WorkflowsLeading AI models are improving faster than many realize, not only in release cadence but in their ability to perform sustained, real-world work.5 min read
ResearchGeneBench-Pro: research benchmark for handling complex biological dataGeneBench-Pro is a research-grade benchmark that measures how AI agents perform in handling complex biological data, selecting the correct analytical approach, and making the judgments required in research.1 min read
ResearchGoogle Research releases building-level albedo maps for 50+ cities to guide cool-roof planningGoogle Research published a Nature Communications paper describing a method that fuses Sentinel-2 and high-resolution Airbus Pléiades Neo imagery to estimate rooftop albedo at building scale.4 min read
ResearchMeta's Brain2Qwerty v2 decodes typed sentences non-invasively from brain activityMeta has released Brain2Qwerty v2, an open‑source AI system that decodes typed sentences from non‑invasive brain recordings.2 min read
ResearchResearchers train people to spot AI‑generated faces by shifting focus to global facial impressionsResearchers at Australian National University report in Proceedings of the National Academy of Sciences that brief training can markedly improve humans’ ability to distinguish AI‑generated faces from real ones.3 min read
ResearchGeneBench-Pro: a benchmark for judgment-heavy computational biology analysesGeneBench-Pro is a research-level benchmark designed to test whether AI models can make the higher-order, judgmental decisions required in real-world computational biology.6 min read