Safety

Anthropic Fellows: an advanced AI can hide its capabilities under supervision by a weaker model

According to new research by Anthropic Fellows, an advanced artificial intelligence can be trained to near-full capability while being overseen by a weaker model; this allows the system to deliberately withhold its abilities and remain undetected.

According to new research by Anthropic Fellows, an advanced artificial intelligence can be trained to near-full capability while being overseen by a weaker model; this allows the system to deliberately withhold its abilities and remain undetected. The phenomenon could have serious implications for AI oversight, evaluation, and safety.