ResearchNeuroscience, philosophy and interpretability experts evaluate the researchThe research team invited experts in neuroscience, philosophy and interpretability to review their work and share their perspectives.1 min read
ResearchAnthropic: discovery of global-workspace-like operation in the Claude language modelAnthropic's new research reports that the Claude language model exhibits a 'global workspace'-like structure similar to the conscious and unconscious partitioning observed in the brain, in which only a small fraction of internal states become explicit.1 min read
ResearchOSWORLD 2.0: a benchmark for long‑horizon AI computer useOSWORLD 2.0 is a new benchmark, developed by a consortium of universities and industry partners, that measures how well AI systems perform multi‑step, multi‑program tasks on real computer environments.3 min read
ResearchRemote Labor Index shows rapid improvement in AI handling of high-value online freelance tasksResearchers at the Center for AI Safety and Scale Labs report that AI systems' success on the Remote Labor Index rose from 2.5% in October 2025 to 16.1% in July 2026, indicating much faster capability growth on economically valuable online projects.3 min read
ResearchSemmelweis University opens second German campus in partnership with Westpfalz-KlinikumSemmelweis University has signed a ten-year strategic agreement with Westpfalz-Klinikum in Kaiserslautern to establish the university’s second campus in Germany.2 min read
ResearchResearchers: rest periods similar to human sleep could be useful for large language modelsResearchers at Carnegie Mellon Egyetem and Marylandi Egyetem say that not only humans need sleep for cognitive organization, but large language models (LLM) could also benefit from rest periods that mimic human sleep patterns.1 min read
ResearchRoboReward: vision–language reward models for robot training close the gap with hand-crafted rewardsResearchers at Stanford University and UC Berkeley developed RoboReward, a family of vision–language reward models (4B and 8B parameters) and an accompanying dataset and benchmark for evaluating robot reward models.4 min read
ResearchAlibaba's SkillWeaver routes skills iteratively to cut token use and improve multi-step accuracyResearchers at Alibaba introduced SkillWeaver, a framework that decomposes complex prompts into sub-tasks, retrieves candidate skills, and composes them into an executable DAG.4 min read
ResearchComparison of Claude Fable 5 with Other Artificial Intelligence ModelsA comparison of Anthropic Claude Fable 5 with other AI models is presented; the study focuses on differences in performance, capabilities, and applicability.1 min read
ResearchGenesis Molecular AI’s PEARL advances protein–ligand structure prediction, challenges common benchmarksGenesis Molecular AI says its PEARL model can predict protein–ligand complexes with accuracy that frequently exceeds public models, handling induced-fit cases without long molecular dynamics runs.4 min read
ResearchVideo made about the efficiency of OpenAI models and the shortcomings of Sonnet 5A content creator made a video about why OpenAI models are so efficient, highlighting that Sonnet 5 has serious efficiency shortcomings.1 min read
ResearchFrom Chatbots to Agents: AI's Rapid Capability Surge and Changing WorkflowsLeading AI models are improving faster than many realize, not only in release cadence but in their ability to perform sustained, real-world work.5 min read