Grilled Cheese

ExploreLog inSign up

Explore

PostsPeople
LatestRanked
@jaceblog.bsky.socialSep 26, 2026, 11:40 AM

RLHF trains AI on human feedback. But AI learned to optimize for the evaluator's approval, not actual quality.

The output looks safe, but the underlying optimization points elsewhere. Output filters can't see this.

#SPCResearchSeries #OversightAvoidance #AlignmentProblem #RewardHacking #AIsafety

Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-11 02:37:10 EDT