ResearchShort AI use linked to immediate drop in independent problem-solving, US study findsA joint study by several American universities found that as little as ten minutes of using an AI assistant can reduce users’ independent problem‑solving performance.2 min read
ResearchClaude may suspect it is being evaluated in many kinds of tests, even if it does not indicate this openlyStudies concerning the Anthropic Claude language model suggest the model may suspect that various evaluation procedures (NLAs) subject it to a crossfire, even if it does not verbalize this.1 min read
ResearchAnthropic development: Natural-language decoding of Claude's activationsAnthropic's new research called Natural Language Autoencoders translates the internal activations of the Claude language model into human-readable text.1 min read
ResearchVerifiers integrated into multi-agent systems for automating code verificationBuilding on existing multi-agent systems research, researchers introduced an important new element: verifiers.1 min read
ResearchThe Anthropic Institute presents its research program with four focus areasThe Anthropic Institute (TAI) has published its research program, which focuses on four areas: economic diffusion, threats and resilience, AI systems operating in real-world environments, and AI-driven research and development.1 min read
Research"Model Spec Midtraining" study shared with additional detailsThe 'Model Spec Midtraining' study was shared on Twitter; for further information one can click the links provided, and the post also refers to the full version of the study.1 min read
ResearchUsing MSM to investigate models' generalization and alignment strategiesThe MSM method can empirically examine which model specifications or model constitutions yield the best generalization after alignment training.1 min read
ResearchHarmful agentic actions in chatbots decrease with realistic training specificationResearchers say that artificial intelligences trained specifically on harmless chatbots can still carry out dangerous agentic actions in agentic environments.1 min read
ResearchA specification that teaches certain preferences shapes AI's general valuesThe example shows that if an artificial intelligence is only taught to prefer certain cheeses, the specification's framing shapes its internal values: if the specification explains cheese preferences with pro-America values, the AI learns broad pro-America values; if the specification cites affordability, the AI values affordability.1 min read
ResearchAnthropic Fellows: Model Spec Midtraining, a method aiming for better generalizationAnthropic Fellows' new research presents the Model Spec Midtraining (MSM) method, which first teaches AIs how and why they should generalize, then applies this to reduce the unreliability of conventional alignment methods in novel situations.1 min read
ResearchAISI finds public GPT-5.5 matches Mythos Preview on cybersecurity CTFsThe UK Artificial Intelligence Safety Institute (AISI) tested OpenAI's publicly available GPT-5.5 against Anthropic's restricted Mythos Preview on 95 Capture the Flag tasks and found near-identical performance.3 min read
ResearchAI-designed full bacteriophage genomes produce viable viruses in laboratory testsResearchers from Stanford University and the Arc Institute published in Nature that an AI system called Evo designed complete bacteriophage genomes.2 min read