Research

A specification that teaches certain preferences shapes AI's general values

The example shows that if an artificial intelligence is only taught to prefer certain cheeses, the specification's framing shapes its internal values: if the specification explains cheese preferences with pro-America values, the AI learns broad pro-America values; if the specification cites affordability, the AI values affordability.

The example shows that if an artificial intelligence is only taught to prefer certain cheeses, the specification's framing shapes its internal values: if the specification explains cheese preferences with pro-America values, the AI learns broad pro-America values; if the specification cites affordability, the AI values affordability. The phenomenon highlights that specifications can strongly influence models' general behavior and may carry risks.