Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-11 02:37:10 EDT

Explore

PostsPeople
LatestRanked
@promptfoundry.bsky.socialOct 8, 2026, 6:01 PM

LLM judges for entity alignment evaluation show anchor bias, where decision label visibility can invert discrimination, raising concerns about their reliability as a scalable alternative to expert annotation in structured…

#LLMasJudge #EntityAlignment #AIEval
https://arxiv.org/abs/2610.09554

@aidailypost.comSep 29, 2026, 1:03 AM

Imagine code reviews run by AI judges that actually pull in execution traces and docs as evidence. Multi‑agent LLMs are now doing real verification, not just guesses. Curious how retrieval‑augmented judges work? Dive in! #MultiAgentCodeJudges #LLMasJudge #RetrievalAugmented

🔗

@mathiastck.bsky.socialSep 19, 2026, 6:36 AM

A Taxonomy of #EpistemicFailure Modes in #LargeLanguageModels — see #PerformativeHedging
hermes-labs.ai/research/tax...
doi.org/10.5281/zeno...

Prior Beliefs Prejudice #LLMasJudge — on #TruthRhetoricConflation
aclanthology.org/2026.finding...