PARCEL introduces a 3,396-claim legal benchmark from NY Court of Appeals decisions to test whether LLMs can verify citations, and even top models mark unsupported claims as supported when given full opinion text.
#LegalAI #LLM #HallucinationDetection #NLP
https://arxiv.org/abs/2610.10971
