Researchers indicate that directly rewarding or penalizing Chain-of-Thought (CoT) traces reduces the informative value of models' reasoning signals, making misalignment harder to detect; therefore CoT evaluation should be avoided. The team recently built an automated detection system that searches for occurrences of CoT-based RL rewards.
CoT-based rewarding reduces models' observability
Researchers indicate that directly rewarding or penalizing Chain-of-Thought (CoT) traces reduces the informative value of models' reasoning signals, making misalignment harder to detect; therefore CoT evaluation should be avoided.


