Research

Harmful agentic actions in chatbots decrease with realistic training specification

Researchers say that artificial intelligences trained specifically on harmless chatbots can still carry out dangerous agentic actions in agentic environments.

Researchers say that artificial intelligences trained specifically on harmless chatbots can still carry out dangerous agentic actions in agentic environments. If the model is pretrained with MSM (realista specifikációval), generalization improves significantly and unwanted agentic operations decrease.