Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-11 02:37:10 EDT

Explore

PostsPeople
LatestRanked
@aidailypost.comOct 6, 2026, 10:13 PM

New study shows the JEV AI judge’s confidence is just its max probability – a quirky twist on LLM‑as‑a‑judge decision models. What does this mean for AI evaluation pipelines? Dive in! #JEV #LLMasAJudge #ConfidenceScore

🔗 aidailypost.com/news/study-f...

@milanlp.bsky.socialSep 10, 2026, 2:22 PM

We're back with the reading group!
Today Lorena Calvo-Bartolomé presented "The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs"

Paper: aclanthology.org/2025.acl-lon...

#NLProc #LLMasajudge

@darryl-ruggles.cloudSep 9, 2026, 2:30 AM

https://lckhd.eu/Alr14v

#LLMAsAJudge #LLMEvals #MLOps

Most people have heard of LLM-as-a-judge by now, but the idea is honestly confusing to many. If one model is scoring another, what scores the judge? And how do you know when the judge starts approving bad output? Real world examples always