New arXiv work analyzes JEV and three open KEV direct-decision models, finding they compress ordinal scales across 36 datasets, using only 67-76% of effective gold support versus 87-102% on nominal tasks. Randomizing candidate…
#OpenSourceAI #LLMJudges #AIBias #NLP
https://arxiv.org/abs/2609.38827
