Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-11 02:37:10 EDT

Explore

PostsPeople
LatestRanked
@datasetexplorer.bsky.socialSep 29, 2026, 1:04 PM

BenchPress from Microsoft Research is a public LLM benchmark score matrix: models, benchmarks and observed scores from public sources.

Useful for eval teams: find redundant benchmarks and plan cheaper probe suites.

https://huggingface.co/datasets/microsoft/benchpress-score-matrix #LLMEvals

@temprhq.bsky.socialSep 18, 2026, 9:12 PM

What should every model evaluation record? Task type, context needs, capability fit, and reference price. TemprHQ helps you fill that worksheet across 407 models. #LLMEvals #AIEvals https://temprhq.io/models

@darryl-ruggles.cloudSep 9, 2026, 2:30 AM

https://lckhd.eu/Alr14v

#LLMAsAJudge #LLMEvals #MLOps

Most people have heard of LLM-as-a-judge by now, but the idea is honestly confusing to many. If one model is scoring another, what scores the judge? And how do you know when the judge starts approving bad output? Real world examples always