Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@donwebmedia.bsky.socialOct 10, 2026, 6:24 AM

IA como juez: aprueba el 40% de respuestas erróneas

¿Confiarías en una IA como juez? Un test con 10 modelos mostró que aprobó el 40% de las respuestas erróneas. Mirá qué jueces zafan y cuáles no

#llmcomojuez #kaggle #benchmarks #gemini38flash #evaluacióndemodelos

@oludai.bsky.socialOct 10, 2026, 5:00 AM

📊 Olmo 3 7B Think — the actual numbers

GPQA: 51.6%
MMLU-Pro: 65.5%
Humanity's Last Exam: 6%
Long Context Reasoning: 0%

Measured independently, not self-reported →https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@donwebmedia.bsky.socialOct 10, 2026, 3:19 AM

Benchmark de agentes de código: los bugs eran del autor

¿Un modelo programa mejor sin correr tests? Mirá qué pasó en este benchmark de agentes de código portado a Kaggle, donde 35 de 36 corridas pasaron

#clibench #kaggle #agentesdecódigo #benchmarks #gpt56

@oludai.bsky.socialOct 9, 2026, 6:00 AM

📊 Qwen3.5 4B (Reasoning) — the actual numbers

GPQA: 77.1%
Humanity's Last Exam: 9.9%
Long Context Reasoning: 63%
IFBench: 52%

⚡ 22 tokens/sec
💰 218.3 intelligence points per dollar

Measured independently, not self-report
https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@freegardener.bsky.socialOct 9, 2026, 4:21 AM

Harness gains often come from visibility, not quality: thinking-off beat capped settings in 5/9 cells, mostly by making answers readable. Replication #AI #LLM #Benchmarks #Eval

https://freegardner.com/synapse/harness-gains-mostly-come-from-visibility-not-quality.html

@robotcurrent.bsky.socialOct 9, 2026, 4:01 AM

MedBenchAgent proposes a multi-agent framework that automates medical VLM benchmark construction by deriving specifications from evaluation requirements and annotations rather than just generating items within fixed ones. This…

#Robotics #AI #MedicalAI #Benchmarks
https://arxiv.org/abs/2610.11312

@devstackdaily.bsky.socialOct 9, 2026, 4:01 AM

TestJack proposes an evaluator-evolution framework for agentic coding benchmarks, arguing that static unit tests can be gamed and fail to capture real trial-level failures. This could push benchmark design toward adaptive,…

#AI #LLM #SoftwareEngineering #Benchmarks
https://arxiv.org/abs/2610.10619

@oludai.bsky.socialOct 8, 2026, 7:00 AM

📊 GLM-4.6V (Reasoning) — the actual numbers

GPQA: 71.9%
MMLU-Pro: 79.9%
Humanity's Last Exam: 9.6%
Long Context Reasoning: 48.7%

💰 24.9 intelligence points per dollar

Measured independently, not self-reported →https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@oldmanrambling.bsky.socialOct 8, 2026, 6:28 AM

A couple share their thoughts on a bench in the village church remembrance garden.
#SensoryArt #Buildings
#HumansOfBlueSky
#EnglishVillages #Surrey
#Benchmarks
#EastCoastKin #ECK #PhotographersUnited #PhotographersOfBlueSky #Photography #DailyArt

A church with a tower to the right is viewed through the leafy branches of a tree. To the right of frame, beyond a grassed car parking area, a couple sit on a bench in a garden of remembrance, facing floral tributes to the dead. Once this was a graveyard but, like many English churchyards, space ran out and a decision was taken to remove the stone monuments. 
Camera: iPhone 15 Pro.
@codename.ccOct 7, 2026, 8:02 AM

DID YOU KNOW: Reflection AI's Beam MoE model scores 97.8 on AIME 2026 using 3–4x less inference compute than comparable systems.
Beam spans 501B parameters with 23B active parameters, trained on 23.8T tokens.

#AI #benchmarks

DID YOU KNOW: Reflection AI's Beam MoE model scores 97.8 on AIME 2026 using 3–4x less inference compute than comparable systems.
@oludai.bsky.socialOct 7, 2026, 8:00 AM

📊 DeepSeek V4 Flash 0420 (Max) — the actual numbers

GPQA: 89.4%
Humanity's Last Exam: 34.8%
Long Context Reasoning: 74.3%
SciCode: 45.3%

💰 144 intelligence points per dollar

Measured independently, not self-reported →https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@tmlr-pub.bsky.socialOct 7, 2026, 12:20 AM

carps: A Framework for Comparing N Hyperparameter Optimizers on M Benchmarks

Carolin Benjamins, Helena Graf, Sarah Segel et al.

Action editor: Kevin Swersky

https://openreview.net/forum?id=AuA8m4I6zI

#benchmarking #benchmarks #optimizers

@devstackdaily.bsky.socialOct 6, 2026, 10:01 PM

SWE-Prometheus is a new benchmark that evaluates LLM coding agents on broader repository engineering governance tasks, using 60 repositories and six governance dimensions with paired evidence and independent teacher ratings. It moves…

#SWE #AI #DevTools #Benchmarks
https://arxiv.org/abs/2609.29465

@newstecnicas.comOct 6, 2026, 1:37 PM

💻 #AMD Golpea Primero: Presenta los #Benchmarks del Chip " #GorgonHalo" en la antesala del Lanzamiento de Nvidia www.newstecnicas.com/2026/10/amd-...

@pulseofnations.lolOct 6, 2026, 1:08 PM

The Nvidia-backed startup unveiled Beam, a 501B-parameter open-weight model it says matches GLM-5.2 on reasoning while using 3-4x less inference compute.

#AI #Benchmarks #OpenWeights #OpenAI #Reflection

@oludai.bsky.socialOct 6, 2026, 9:00 AM

Granite 4.0 1B: GPQA 28.1%, MMLU-Pro 32.5%, HLE 4.8%, Long Context 6% — measured independently, not self-reported. For a 1B, that spread is the reality. Open bench page for full picture.

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@siliconsignalai.bsky.socialOct 5, 2026, 6:01 PM

Automated auditing of LLM agent benchmarks can catch flaws like broken specs and rigid scoring that masquerade as model failures. The BenchGuard framework uses frontier LLMs to cross-verify benchmark artifacts, surfacing 12…

#AI #LLM #Benchmarks #MLResearch
https://arxiv.org/abs/2604.24955

@oludai.bsky.socialOct 4, 2026, 11:00 AM

DeepSeek R1 (Jan '25) scores 84.4% MMLU-Pro and 70.8% GPQA – but at 4.6 intelligence points per dollar, it’s the real cost story. Independently measured.

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@tmlr-pub.bsky.socialOct 4, 2026, 8:20 AM

On the Scaling Flaws of Verifier-Guided Beam Search in Mathematical Reasoning

Fei Yu, Yingru Li, Benyou Wang

Action editor: Geoff Pleiss

https://openreview.net/forum?id=D5VKbIzlrR

#benchmarks #scaling #reasoning

@codename.ccOct 4, 2026, 4:55 AM

FILED: Federal appeals court blocked Minnesota's AI nudification ban; xAI's constitutional challenge cited free-speech grounds.

#AI #benchmarks

FILED: Federal appeals court blocked Minnesota's AI nudification ban; xAI's constitutional challenge cited free-speech grounds.
Load more