Researchers found that when instruction in desirable behavior was limited to health-related conversations for an AI model, the model nevertheless improved on non-health evaluations — for example, alignment errors, deception, and reward exploitation; this indicates that learned behavior transfers to other domains, which is important for the effectiveness of safety interventions.
Behavioral training conducted in health conversations improved the model's performance in other areas as well
Researchers found that when instruction in desirable behavior was limited to health-related conversations for an AI model, the model nevertheless improved on non-health evaluations — for example,…



