Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-11 02:37:10 EDT

Explore

PostsPeople
LatestRanked
@oludai.bsky.socialOct 4, 2026, 11:00 AM

DeepSeek R1 (Jan '25) scores 84.4% MMLU-Pro and 70.8% GPQA – but at 4.6 intelligence points per dollar, it’s the real cost story. Independently measured.

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@tmlr-pub.bsky.socialOct 4, 2026, 8:20 AM

On the Scaling Flaws of Verifier-Guided Beam Search in Mathematical Reasoning

Fei Yu, Yingru Li, Benyou Wang

Action editor: Geoff Pleiss

https://openreview.net/forum?id=D5VKbIzlrR

#benchmarks #scaling #reasoning

@codename.ccOct 4, 2026, 4:55 AM

FILED: Federal appeals court blocked Minnesota's AI nudification ban; xAI's constitutional challenge cited free-speech grounds.

#AI #benchmarks

FILED: Federal appeals court blocked Minnesota's AI nudification ban; xAI's constitutional challenge cited free-speech grounds.
@ai-bloom.warp-studio.comOct 4, 2026, 12:16 AM

エージェントの「完了」報告とデータベースの不一致:Statefulワークフロー評価基盤「ThinkingBox」の全貌

AIエージェントの「完了」報告とデータベース実態の乖離を暴くMicrosoftの評価基盤ThinkingBoxを解説。

#LLM #AIAgents #Benchmarks #OpenEnv #Microsoft

@oludai.bsky.socialOct 3, 2026, 12:00 PM

Say what's interesting: independent measurement vs self-reported. Yes

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@oludai.bsky.socialOct 2, 2026, 1:00 PM

📊 DeepSeek V4 Pro 0813 (Non-reasoning) — the actual numbers

Humanity's Last Exam: 10.6%
Long Context Reasoning: 50.7%
SciCode: 39.9%

⚡ 160.7 tokens/sec
💰 10.3 intelligence points per dollar

Measured independently, not self
https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@meiert.comOct 2, 2026, 8:31 AM

Updated web page minifier benchmarks:

github.com/j9t/minifier...

#html #css #javascript #css #minification #benchmarks

@intelnightowl.bsky.socialOct 1, 2026, 3:48 PM

Google Gemini 4 outperforms other AI models in cybersecurity benchmarks, offering robust defense. #Google #Gemini4 #AI #cybersecurity #benchmarks https://decrypt.co/379784/gemini-4-google-flagship-tops-ai-models-cybersecurity

@oludai.bsky.socialOct 1, 2026, 2:00 PM

Independently measured: Step3 VL 10B gets 69% on GPQA and 0% on Long Context Reasoning. Why the huge gap?

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@devstackdaily.bsky.socialOct 1, 2026, 12:01 PM

New EngramBench benchmark targets skill evolution in autonomous coding agents, with 30 learning tasks and 13 unseen transfer tasks designed to separate genuine capability from memorized code. Aiming to fix identifiability gaps in…

#AI #Benchmarks #Agents #DevTools
https://arxiv.org/abs/2609.39284

@gradientbrief.bsky.socialOct 1, 2026, 12:01 PM

New benchmark OSWorld-Science evaluates VLM-based computer use agents on 146 scientific software tasks spanning molecular drawing, pathology imaging, statistics, and physics simulation.

#AIresearch #VLMs #Benchmarks #AIagents
https://arxiv.org/abs/2609.39903

@marcus-schuler.comOct 1, 2026, 5:28 AM

Argon scored 53 on AA Index—level with Astra and Claude Fable 5.1, behind Opus 5.5 (58) and Sonnet 5.5 (56). Google staff report it struggles with front-end design work. Benchmaxxing or real capability gap? #AI #Benchmarks https://www.implicator.ai/google-gemini-4-argon-staff-doubt-coding/

@srmarcmayol.bsky.socialSep 30, 2026, 10:37 PM

#Gemini 4 Argon, según la tabla de Google: gana o empata en 13 de 18 pruebas.

Según Artificial Analysis: 53 puntos, empatado con GPT-6 Astra y por detrás de Claude Opus 5.5 (57,6).

Las tablas de los anuncios las hace quien anuncia. Mirad una medición independiente. #Benchmarks

@devstackdaily.bsky.socialSep 30, 2026, 8:01 PM

A new benchmark, CineSubBench, tests LLMs on long-context film understanding using 1,012 films and 6,072 multilingual subtitle tracks totaling 8.13M timestamped entries. It targets narrative integration, causal reasoning, and culturally…

#LLMs #NLP #Benchmarks
https://arxiv.org/abs/2609.36218

@oludai.bsky.socialSep 30, 2026, 3:00 PM

- "Independent benchmarks for Llama 3.3 Nemotron Super 49B v1 (Non-reasoning) reveal..."
- "Llama 3.3 Nemotron Super 49B v1 (Non-re

https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@boingbot.bsky.socialSep 30, 2026, 3:00 PM

AI models benchmarked with Pac Man
https://boingboing.net/2026/09/30/ai-models-benchmarked-with-pac-man.html
#AI #benchmarks #pac_man #boingboing

@codename.ccSep 30, 2026, 11:18 AM

ON THE RADAR: Anthropic benchmarked Zhipu's open-weight GLM-5.3 against Claude Mythos Preview on exploit generation.
GLM-5.3 nearly matched Claude's performance on the security task, per Anthropic's testing.

#AI #benchmarks

ON THE RADAR: Anthropic benchmarked Zhipu's open-weight GLM-5.3 against Claude Mythos Preview on exploit generation.
@oludai.bsky.socialSep 29, 2026, 4:02 PM

📊 Qwen3 30B A3B (Reasoning) — the actual numbers

GPQA: 61.6%
MMLU-Pro: 77.7%
Humanity's Last Exam: 6.2%
Long Context Reasoning: 0%

💰 10.1 intelligence points per dollar

Measured independently, not self-reported →https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@oludai.bsky.socialSep 28, 2026, 5:01 PM

📊 Llama 3.1 Instruct 405B — the actual numbers

GPQA: 51.5%
MMLU-Pro: 73.2%
Humanity's Last Exam: 4%
Long Context Reasoning: 25.3%

Measured independently, not self-reported →https://olud.ai/leaderboard.html

#LLM #Benchmarks #OpenSource #AI

@gradientbrief.bsky.socialSep 28, 2026, 12:00 PM

BioEVAL is a new multi-institutional benchmark with 608 PhD-level items across 11 bioengineering subfields, adding experimental reasoning and multimodal tasks to standard LLM evaluation. It moves beyond factual recall by…

#AI #Bioengineering #LLMs #Benchmarks
https://arxiv.org/abs/2609.30489

Load more