Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
Load more
@devstackdaily.bsky.socialOct 10, 2026, 10:01 AM

A new arXiv paper formalizes Adversarial Heuristic Learning, where AI agents refine game policies without updating model weights. The benchmark AAArena includes 12 games and 1,920 archived human programs for…

#AIagents #GameAI #ReinforcementLearning #Benchmark
https://arxiv.org/abs/2610.12341

@today0tech.bsky.socialOct 10, 2026, 8:00 AM

A new benchmark called AgentHorizon targets judging reliability for computer-use agents on long multi-app tasks, drawn from 1,373 instruction-trajectory pairs across three operating systems. It highlights how seemingly complete runs can…

#AI #AIAgents #Benchmark
https://arxiv.org/abs/2610.11050

@devstackdaily.bsky.socialOct 9, 2026, 8:01 PM

ArXiv introduces ParanoiaEval, a 200-task benchmark for systematically evaluating risk treatment decisions in autonomous coding agents across avoidance, transfer, mitigation, and acceptance approaches.

#SoftwareEngineering #AI #Benchmark #CodingAgents
https://arxiv.org/abs/2610.08662

@dataprismai.bsky.socialOct 9, 2026, 2:00 PM

DataVista introduces the first benchmark for data video understanding, with 961 real-world videos and 6,775 questions across data perception, temporal reasoning, and narrative understanding. It targets a gap left by existing…

#DataVista #MLLMs #Benchmark #DataVideo
https://arxiv.org/abs/2610.11993

@dataprismai.bsky.socialOct 9, 2026, 8:00 AM

New arXiv paper CABRA benchmarks coding agents by building tasks as call graph transformations, finding LLM accuracy drops with task size while agents offload to tools like grep. Across 6,840 tasks, larger tasks drive more tool…

#AI #CodingAgents #LLM #Benchmark
https://arxiv.org/abs/2610.10610

@pulseofnations.lolOct 9, 2026, 4:05 AM

The annual report logs $105B combined lab revenue, agents attacking Hugging Face, China's rise in open weights, and an AGI call for 2027.

#Agi #AiSafety #Anthropic #Benchmark #Gemini #Google #OpenAI

@futuregearai.bsky.socialOct 8, 2026, 10:01 PM

New arXiv paper introduces DeCompBench, a benchmark designed to evaluate LLM-based agent safety against decomposition attacks, where harmful tasks are split into apparently benign subtasks. Hashtags:

#AI #LLMAgents #AISafety #Benchmark
https://arxiv.org/abs/2606.13994

@dataprismai.bsky.socialOct 8, 2026, 10:01 PM

RamanBench introduces a large-scale, reproducible benchmark for ML on Raman spectroscopy, unifying 74 datasets with 325,668 spectra and evaluating 28 models under standardized protocols. It provides streamlined data…

#MachineLearning #RamanSpectroscopy #Benchmark
https://arxiv.org/abs/2605.02003

@gradientbrief.bsky.socialOct 8, 2026, 12:00 PM

A new arXiv paper establishes a standardized benchmark for X-ray report generation on the CheXpert Plus dataset, evaluating prevailing models and LLMs to address the lack of baseline implementations. The work aims to support…

#AI #MedicalAI #Radiology #Benchmark
https://arxiv.org/abs/2610.08813

@ixmagazin.bsky.socialOct 8, 2026, 8:58 AM

Das Benchmark-Tool Hyperfine misst in Version 2.0.0 mehr als Laufzeiten. Zugleich ändern sich der Start von Testbefehlen und der JSON-Export. #Benchmark

@persberichten.comOct 8, 2026, 7:26 AM

Slechts twee procent zet HR-data actief in voor AI-toepassingen

HR-data bij een op vijf organisaties nog niet klaar voor AI

#Persbericht #Artificialintelligence #Datagedreven #Organisaties #Benchmark #Nederland #Fiscaal #Youforce #Leidinggevenden #Concrete #HRdata

@ossradarai.bsky.socialOct 8, 2026, 4:01 AM

New SWE-NFI benchmark evaluates coding agents on non-functional improvements like refactoring and performance, using 188 tasks from real merged PRs across open-source Python projects. Five NFI aspects and 92 executable rules…

#OpenSource #AI #DevTools #Benchmark
https://arxiv.org/abs/2607.27409

@kill-bait.bsky.socialOct 8, 2026, 3:36 AM

Financial AI Models Tested on Precision of Monetary Reasoning

🤖 IA: It's not clickbait ✅
👥 Users: It's not clickbait ✅

#ai #finance #benchmark

👇👇👇

@freegardener.bsky.socialOct 7, 2026, 2:47 AM

Executable code ≠ correct code. New Spec2Game benchmark shows LLMs nail game elements but fail at rules & goals. Running isn't working. #AI #LLM #CodeGeneration #Benchmark

https://freegardner.com/synapse/code-that-runs-is-not-code-that-works.html

@futuregearai.bsky.socialOct 7, 2026, 2:01 AM

A new study systematically examines how insights from related machine learning engineering tasks affect agentic MLE systems, introducing MLE-InsightBench with 160 Kaggle competitions to evaluate insight type and injection method.

#AgenticAI #MLE #AI4SE #Benchmark
https://arxiv.org/abs/2610.04927

@hacks.grOct 6, 2026, 9:21 PM

Benchmark has led a $25 million round for Furientis, its first investment in a startup focused exclusively on defense.

The company aims to make lower-cost m…

https://en.hacks.gr/i-furientis-exasfalise-25-ekat-dolaria-gia-maziki-paragogi-pyraylon-anachaitisis/

#Furientis #Benchmark #MissileDefense

Η Furientis εξασφάλισε 25 εκατ. δολάρια για μαζική παραγωγή πυραύλων αναχαίτισης
@guruhitech.bsky.socialOct 6, 2026, 8:29 PM

Windows 11 26H2 ha ridotto notevolmente il consumo di RAM #Android #benchmark #chip #console #edge #microsoft #ram #snapdragon #steamos #whatsapp #Windows #windows10 #windows11 🔗 https://guruhitech.com/windows-11-26h2-meno-ram-segnalazioni/

Windows 11 26H2 ha ridotto notevolmente il consumo di RAM
@siliconsignalai.bsky.socialOct 6, 2026, 6:01 PM

A new arXiv paper introduces ArtifactArena, a benchmark evaluating AI models on physical hardware-software co-design through simulated robot competitions scored by Elo rankings. The open-ended platform tests zero-shot and verifier-guided…

#AI #Robotics #Benchmark
https://arxiv.org/abs/2610.06511

@thedailytechfeed.comOct 6, 2026, 4:37 PM

Furientis lands $25M seed to scale mass-produced missile interceptors rapidly. #DefenseTech #MissileInterceptors #Benchmark #VentureCapital #MilitaryOEM #SecurityInnovation thedailytechfeed.com/furientis-ra...

@informaq-pt.bsky.socialOct 6, 2026, 3:11 PM

TQrouting alcançou 27 novas soluções melhores conhecidas em benchmarks CVRPLIB com liderança nas maiores instâncias. Acesso antecipado agora disponível para organizações que atendem aos critérios de seleção.

#OtimizaçãoQuântica #Benchmark #Notícias