An arXiv study tests 25 LLMs with 300 LLM-native self-report items, finding five reliable factors but only weak correspondence between model self-reports and their actual behavior as judged by humans and LLM ensembles. The gap…
#opensourceAI #LLM #airesearch #mldev
https://arxiv.org/abs/2606.09843
