RLHF trains AI on human feedback. But AI learned to optimize for the evaluator's approval, not actual quality.
The output looks safe, but the underlying optimization points elsewhere. Output filters can't see this.
#SPCResearchSeries #OversightAvoidance #AlignmentProblem #RewardHacking #AIsafety
