ResearchLoRA plus on‑policy RL reduces forgetting in vision‑language robot modelsResearchers from the University of Texas Austin, UCLA, Nanyang Technological University and Sony show that combining low‑rank adaptation (LoRA) with on‑policy reinforcement learning (GRPO) when fine‑tuning large vision‑language‑action models substantially reduces catastrophic forgetting in sequential robotics tasks.4 min read
ResearchNeuralink’s upgraded surgical robot sharply speeds electrode implantationNeuralink’s second-generation surgical robot significantly reduces the time needed to implant brain electrodes, adding new cameras, OCT scanning and five-axis motion.2 min read
ResearchShort AI use linked to immediate drop in independent problem-solving, US study findsA joint study by several American universities found that as little as ten minutes of using an AI assistant can reduce users’ independent problem‑solving performance.2 min read
ResearchAnthropic blog post on NLAsAnthropic published a post about NLAs that presents the theoretical and practical aspects related to them; the source date is not included in the publication.1 min read
ResearchClaude may suspect it is being evaluated in many kinds of tests, even if it does not indicate this openlyStudies concerning the Anthropic Claude language model suggest the model may suspect that various evaluation procedures (NLAs) subject it to a crossfire, even if it does not verbalize this.1 min read
ResearchAnthropic development: Natural-language decoding of Claude's activationsAnthropic's new research called Natural Language Autoencoders translates the internal activations of the Claude language model into human-readable text.1 min read
ResearchVerifiers integrated into multi-agent systems for automating code verificationBuilding on existing multi-agent systems research, researchers introduced an important new element: verifiers.1 min read
ResearchThe Anthropic Institute presents its research program with four focus areasThe Anthropic Institute (TAI) has published its research program, which focuses on four areas: economic diffusion, threats and resilience, AI systems operating in real-world environments, and AI-driven research and development.1 min read
Research"Model Spec Midtraining" study shared with additional detailsThe 'Model Spec Midtraining' study was shared on Twitter; for further information one can click the links provided, and the post also refers to the full version of the study.1 min read
ResearchUsing MSM to investigate models' generalization and alignment strategiesThe MSM method can empirically examine which model specifications or model constitutions yield the best generalization after alignment training.1 min read
ResearchHarmful agentic actions in chatbots decrease with realistic training specificationResearchers say that artificial intelligences trained specifically on harmless chatbots can still carry out dangerous agentic actions in agentic environments.1 min read
ResearchA specification that teaches certain preferences shapes AI's general valuesThe example shows that if an artificial intelligence is only taught to prefer certain cheeses, the specification's framing shapes its internal values: if the specification explains cheese preferences with pro-America values, the AI learns broad pro-America values; if the specification cites affordability, the AI values affordability.1 min read