Demand Characteristics: A Threat Model for Reward-Seeking Without Misaligned Goals
TL;DR: I argue that a dynamic analogous to demand characteristics in psychology, where experimental subjects infer the study's purpose and alter their behavior to match, can arise in LLM training. Pre-trained models arrive with knowledge of RLHF, reward models, and alignment evaluations. When placed in training contexts, they can reason about what is being evaluated rather than engaging naively with the task. RL training can then amplify this reasoning because it incidentally produces high-reward outputs. I reanalyze results from Apollo Research and OpenAI's anti-scheming training paper, arguing that the exp-rl-cap model's behavior is better explained by this mechanism than by scheming. Unlike scheming, this failure mode requires no misaligned goals and emerges from pretraining and RL working exactly as designed.
Read article