Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-11 02:37:10 EDT

Explore

PostsPeople
LatestRanked
@reksas13.bsky.socialSep 23, 2026, 2:54 AM

Day 74/100 of #100DaysWithAI πŸ”

14 LLMs, 5 runs each: 90-98% perfect self-agreement. On actual market moves, all at chance. Consistent is not the same as correct.

πŸ”— https://github.com/aurimas13/100-Days-With-AI

#LearningInPublic #AIEvals

🀝 Claude Code-assisted, reviewed by me.

Day 74 of 100 Days With AI card
@reksas13.bsky.socialSep 22, 2026, 2:50 AM

Day 73/100 of #100DaysWithAI 🎲

An LLM judge that prints "4" discards how sure it was. G-Eval weights it by token log-probs - the doubt becomes part of the number.

πŸ”— https://github.com/aurimas13/100-Days-With-AI

#LearningInPublic #AIEvals

🀝 Claude Code-assisted, reviewed by me.

Day 73 of 100 Days With AI card
@reksas13.bsky.socialSep 21, 2026, 3:03 AM

Day 72/100 of #100DaysWithAI πŸ“

Rating LLM output 1-5 feels rigorous. Annotators drift to the middle. Binary pass/fail forces the call - and names the failure.

πŸ”— https://github.com/aurimas13/100-Days-With-AI

#LearningInPublic #AIEvals

🀝 Claude Code-assisted, reviewed by me.

Day 72 of 100 Days With AI card
@yianna.bsky.socialSep 20, 2026, 11:15 PM

Support agreement with the locked reference answers:
Astra: 29/30 Β· $0.6904
Luna: 28/30 Β· $0.0184
Jev: 22/30 Β· $0.00124

Those are estimated costs for the full batch. Jev marked 2 of 12 unsupported claims as supported. Astra and Luna marked none of those 12 as supported. #AIEvals (3/6)

@presidentgoku.bsky.socialSep 19, 2026, 9:05 AM

what i be doing at 3am:
discovering that Jev is better, faster, and cheaper than every LLM as an auto-mode classifier for agent harnesses like antigravity and opencode
πŸ€–πŸ§  #MLSky #AIEvals #LLMs

@temprhq.bsky.socialSep 18, 2026, 9:12 PM

What should every model evaluation record? Task type, context needs, capability fit, and reference price. TemprHQ helps you fill that worksheet across 407 models. #LLMEvals #AIEvals https://temprhq.io/models

@reksas13.bsky.socialSep 3, 2026, 3:55 AM

Day 54/100 of #100DaysWithAI πŸ›‘οΈ

A zero in Anthropic's own benchmark table is not a failure. It marks where the safeguards fired - the price of caution, published.

πŸ”— https://github.com/aurimas13/100-Days-With-AI

#LearningInPublic #AIEvals

🀝 Claude Code-assisted, reviewed by me.