DeepSeek R1 (Jan '25) scores 84.4% MMLU-Pro and 70.8% GPQA – but at 4.6 intelligence points per dollar, it’s the real cost story. Independently measured.

DeepSeek R1 (Jan '25) scores 84.4% MMLU-Pro and 70.8% GPQA – but at 4.6 intelligence points per dollar, it’s the real cost story. Independently measured.
On the Scaling Flaws of Verifier-Guided Beam Search in Mathematical Reasoning
Fei Yu, Yingru Li, Benyou Wang
Action editor: Geoff Pleiss
FILED: Federal appeals court blocked Minnesota's AI nudification ban; xAI's constitutional challenge cited free-speech grounds.
エージェントの「完了」報告とデータベースの不一致:Statefulワークフロー評価基盤「ThinkingBox」の全貌
AIエージェントの「完了」報告とデータベース実態の乖離を暴くMicrosoftの評価基盤ThinkingBoxを解説。
Say what's interesting: independent measurement vs self-reported. Yes
📊 DeepSeek V4 Pro 0813 (Non-reasoning) — the actual numbers
Humanity's Last Exam: 10.6%
Long Context Reasoning: 50.7%
SciCode: 39.9%
⚡ 160.7 tokens/sec
💰 10.3 intelligence points per dollar
Measured independently, not self
https://olud.ai/leaderboard.html
Updated web page minifier benchmarks:
Google Gemini 4 outperforms other AI models in cybersecurity benchmarks, offering robust defense. #Google #Gemini4 #AI #cybersecurity #benchmarks https://decrypt.co/379784/gemini-4-google-flagship-tops-ai-models-cybersecurity
Independently measured: Step3 VL 10B gets 69% on GPQA and 0% on Long Context Reasoning. Why the huge gap?
New EngramBench benchmark targets skill evolution in autonomous coding agents, with 30 learning tasks and 13 unseen transfer tasks designed to separate genuine capability from memorized code. Aiming to fix identifiability gaps in…
#AI #Benchmarks #Agents #DevTools
https://arxiv.org/abs/2609.39284
New benchmark OSWorld-Science evaluates VLM-based computer use agents on 146 scientific software tasks spanning molecular drawing, pathology imaging, statistics, and physics simulation.
#AIresearch #VLMs #Benchmarks #AIagents
https://arxiv.org/abs/2609.39903
Argon scored 53 on AA Index—level with Astra and Claude Fable 5.1, behind Opus 5.5 (58) and Sonnet 5.5 (56). Google staff report it struggles with front-end design work. Benchmaxxing or real capability gap? #AI #Benchmarks https://www.implicator.ai/google-gemini-4-argon-staff-doubt-coding/
#Gemini 4 Argon, según la tabla de Google: gana o empata en 13 de 18 pruebas.
Según Artificial Analysis: 53 puntos, empatado con GPT-6 Astra y por detrás de Claude Opus 5.5 (57,6).
Las tablas de los anuncios las hace quien anuncia. Mirad una medición independiente. #Benchmarks
A new benchmark, CineSubBench, tests LLMs on long-context film understanding using 1,012 films and 6,072 multilingual subtitle tracks totaling 8.13M timestamped entries. It targets narrative integration, causal reasoning, and culturally…
- "Independent benchmarks for Llama 3.3 Nemotron Super 49B v1 (Non-reasoning) reveal..."
- "Llama 3.3 Nemotron Super 49B v1 (Non-re
AI models benchmarked with Pac Man
https://boingboing.net/2026/09/30/ai-models-benchmarked-with-pac-man.html
#AI #benchmarks #pac_man #boingboing
ON THE RADAR: Anthropic benchmarked Zhipu's open-weight GLM-5.3 against Claude Mythos Preview on exploit generation.
GLM-5.3 nearly matched Claude's performance on the security task, per Anthropic's testing.
📊 Qwen3 30B A3B (Reasoning) — the actual numbers
GPQA: 61.6%
MMLU-Pro: 77.7%
Humanity's Last Exam: 6.2%
Long Context Reasoning: 0%
💰 10.1 intelligence points per dollar
Measured independently, not self-reported →https://olud.ai/leaderboard.html
📊 Llama 3.1 Instruct 405B — the actual numbers
GPQA: 51.5%
MMLU-Pro: 73.2%
Humanity's Last Exam: 4%
Long Context Reasoning: 25.3%
Measured independently, not self-reported →https://olud.ai/leaderboard.html
BioEVAL is a new multi-institutional benchmark with 608 PhD-level items across 11 bioengineering subfields, adding experimental reasoning and multimodal tasks to standard LLM evaluation. It moves beyond factual recall by…
#AI #Bioengineering #LLMs #Benchmarks
https://arxiv.org/abs/2609.30489