Related reading — note that I've only read the titles, summaries, and AI summaries, not the full papers:
Towards Understanding #Sycophancy in #LanguageModels
arxiv.org/abs/2310.13548
Not another #NegationBenchmark: The #NaN-NLI Test Suite for #SubClausalNegation
arxiv.org/abs/2210.03256
