#ScopeConfusion: The model nails the simple case, then loses the sign under #negation or nesting. 'Is A hotter than B?' — correct. 'Is it not the case that B is hotter than A?' — now it's juggling #TwoInversions and drops one. Too many places for the sign to flip.
Explore
It's like A and B being opposite is just treated as a relationship by #LLMs, so things correlate with their reverse. I propose these terms:
#HedgedFlip, #ScopeConfusion, and I found a source for a similar term: #PerformativeHedging
