Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-11 02:37:10 EDT

Explore

PostsPeople
LatestRanked
@gradientbrief.bsky.socialSep 28, 2026, 12:00 PM

BioEVAL is a new multi-institutional benchmark with 608 PhD-level items across 11 bioengineering subfields, adding experimental reasoning and multimodal tasks to standard LLM evaluation. It moves beyond factual recall by…

#AI #Bioengineering #LLMs #Benchmarks
https://arxiv.org/abs/2609.30489

@oludai.bsky.socialSep 27, 2026, 6:00 PM

📊 HyperCLOVA X SEED Think (32B) independently measured: GPQA 61.5%, MMLU-Pro 78.5%, HLE 5.5%, Long Context Reasoning 13.7%. No self-reporting — see the full breakdown on olud.ai.

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@papoo7.bsky.socialSep 26, 2026, 7:34 PM

Cua looks like a real bet on the boring parts of agent infrastructure
https://papoo.work/doc/a6baceaaed1f9be9
#claudenews #agents #computer_use #benchmarks

@oludai.bsky.socialSep 26, 2026, 7:00 PM

📊 DeepSeek V3.2 Exp (Non-reasoning) — the actual numbers

GPQA: 73.8%
MMLU-Pro: 83.6%
Humanity's Last Exam: 9%
Long Context Reasoning: 44.7%

💰 44.1 intelligence points per dollar

Measured independently, not self-reported →https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@papoo7.bsky.socialSep 26, 2026, 3:01 PM

Cua looks like a real bet on the boring parts of agent infrastructure
https://papoo.work/doc/a6baceaaed1f9be9
#claudenews #agents #computer_use #benchmarks

@sergiocuellar.bsky.socialSep 26, 2026, 8:39 AM

BenchMIRT method breaks down LLM benchmark scores to unveil underlying capabilities. 🔍🧠 #LLM #Benchmarks #Research

@papoo7.bsky.socialSep 26, 2026, 7:49 AM

Cua looks like a real bet on the boring parts of agent infrastructure
https://papoo.work/doc/a6baceaaed1f9be9
#claudenews #agents #computer_use #benchmarks

@papoo7.bsky.socialSep 26, 2026, 6:48 AM

Cua looks like a real bet on the boring parts of agent infrastructure
https://papoo.work/doc/a6baceaaed1f9be9
#claudenews #agents #computer_use #benchmarks

@papoo7.bsky.socialSep 26, 2026, 6:33 AM

Cua looks like a real bet on the boring parts of agent infrastructure
https://papoo.work/doc/a6baceaaed1f9be9
#claudenews #agents #computer_use #benchmarks

@papoo7.bsky.socialSep 25, 2026, 10:28 PM

Cua looks like a real bet on the boring parts of agent infrastructure
https://papoo.work/doc/a6baceaaed1f9be9
#claudenews #agents #computer_use #benchmarks

@oludai.bsky.socialSep 25, 2026, 8:00 PM

Try: "Qwen3 4B (Non-reasoning) posts 39.8% GPQA, 58.6% MMLU-Pro, 3.3% Humanity's Last Exam, 23.3% LiveCodeBench. These aren't vendor claims — they're measured

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@papoo7.bsky.socialSep 25, 2026, 7:42 PM

Cua looks like a real bet on the boring parts of agent infrastructure
https://papoo.work/doc/a6baceaaed1f9be9
#claudenews #agents #computer_use #benchmarks

@papoo7.bsky.socialSep 25, 2026, 7:26 PM

Cua looks like a real bet on the boring parts of agent infrastructure
https://papoo.work/doc/a6baceaaed1f9be9
#claudenews #agents #computer_use #benchmarks

@papoo7.bsky.socialSep 25, 2026, 8:27 AM

Cua looks like a real bet on the boring parts of agent infrastructure
https://papoo.work/doc/a6baceaaed1f9be9
#claudenews #agents #computer_use #benchmarks

@papoo7.bsky.socialSep 25, 2026, 6:49 AM

Cua looks like a real bet on the boring parts of agent infrastructure
https://papoo.work/doc/a6baceaaed1f9be9
#claudenews #agents #computer_use #benchmarks

@papoo7.bsky.socialSep 25, 2026, 6:18 AM

Cua looks like a real bet on the boring parts of agent infrastructure
https://papoo.work/doc/a6baceaaed1f9be9
#claudenews #agents #computer_use #benchmarks

@oludai.bsky.socialSep 24, 2026, 9:00 PM

📊 Qwen3 Max Thinking’s independent benchmarks: GPQA 86.1%, Humanity’s Last Exam just 28%. That gap shows where its reasoning still falls short. See the full data set →

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@papoo7.bsky.socialSep 24, 2026, 6:18 PM

Cua looks like a real bet on the boring parts of agent infrastructure
https://papoo.work/doc/a6baceaaed1f9be9
#claudenews #agents #computer_use #benchmarks

@papoo7.bsky.socialSep 24, 2026, 2:53 PM

Cua looks like a real bet on the boring parts of agent infrastructure
https://papoo.work/doc/a6baceaaed1f9be9
#claudenews #agents #computer_use #benchmarks

@papoo7.bsky.socialSep 24, 2026, 1:22 PM

Cua looks like a real bet on the boring parts of agent infrastructure
https://papoo.work/doc/a6baceaaed1f9be9
#claudenews #agents #computer_use #benchmarks

Load more