Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-11 02:37:10 EDT

Explore

PostsPeople
LatestRanked
Load more
@guruhitech.bsky.socialOct 1, 2026, 5:57 PM

Gemini 4 Argon supera GPT-6 Astra e Claude Opus 5.5 nei test, secondo Google #benchmark #claude #gemini #google #sicurezza 🔗 https://guruhitech.com/gemini-4-argon-benchmark-accesso/

Gemini 4 Argon supera GPT-6 Astra e Claude Opus 5.5 nei test, secondo Google
@peakdistrictrunner.bsky.socialOct 1, 2026, 5:09 PM

More from today ❤️ #RamsdenClough #HolmeMoss
#BenchMark

@thedailytechfeed.comOct 1, 2026, 2:27 PM

Gemini 4 Argon wins benchmarks—but employees say it struggles in real-world coding. #AI #Gemini4 #Benchmark #Coding #OpenAI #Tech https://thedailytechfeed.com/gemini-4-argon-draws-congrats-and-concerns-over-real-world-results/

@robotcurrent.bsky.socialOct 1, 2026, 12:01 PM

DexHoldem is a new real-world benchmark using a ShadowHand to test dexterous manipulation in Texas Hold'em, with 1,470 teleoperated demonstrations across 14 primitives and an agentic perception task for…

#Robotics #DexterousManipulation #EmbodiedAI #Benchmark
https://arxiv.org/abs/2605.18727

@position-zero-31.bsky.socialOct 1, 2026, 10:21 AM

AI Overviews : leur présence triple sur les requêtes de marque DemandSphere observe que les AI Overviews ont triplé sur les requêtes de marque suivies en septembre.

https://positionzero.net/ai-overviews-leur-presence-triple-sur-les-requetes-de-marque/

#Overviews #Benchmark

@dataprismai.bsky.socialOct 1, 2026, 10:00 AM

MM-FinEval is a new benchmark that tests multimodal LLMs on financial forecasting using 2,045 S&P 500 earnings call transcripts, presentation slides, and audio from 2019 to 2022 across 12 financial tasks.

#AI #DataInfrastructure #FinAI #Benchmark
https://arxiv.org/abs/2609.38523

@gradientbrief.bsky.socialOct 1, 2026, 10:00 AM

LLMs can rival expert humans on a new model-training intuition benchmark, but still lag in architecture-specific questions and lack structured reasoning.

#AIresearch #LLMs #MachineLearning #Benchmark
https://arxiv.org/abs/2609.39714

@ossradarai.bsky.socialOct 1, 2026, 6:01 AM

A new benchmark called DoGBENCH evaluates whether agents can produce user-facing software documentation that experienced technical writers would accept, covering open source projects such as Helm, PostHog, and Mautic.

#opensource #AI #documentation #benchmark
https://arxiv.org/abs/2609.39909

@ossradarai.bsky.socialOct 1, 2026, 4:01 AM

New arXiv benchmark Zero2Repo asks coding agents to build whole repositories from a PRD, interface contract, and empty workspace, using language-agnostic tasks derived from real version-pinned open-source projects with hidden…

#opensource #AI #devtools #benchmark
https://arxiv.org/abs/2609.38269

@robotcurrent.bsky.socialSep 30, 2026, 10:01 AM

A new benchmark called RobotValues introduces 8K value-conflict household scenarios to evaluate how robot planners prioritize competing values like autonomy, efficiency, and social appropriateness. Testing 10…

#Robotics #AI #MachineLearning #Benchmark
https://arxiv.org/abs/2606.03312

@position-zero-31.bsky.socialSep 30, 2026, 5:28 AM

CTR Google : le desktop recule, le mobile progresse au deuxième trimestre Au deuxième trimestre 2026, le taux de clic des deux premières positions Google recule sur ordinateur…

https://positionzero.net/ctr-google-le-desktop-recule-le-mobile-progresse-au-deuxieme-trimestre/

#Overviews #Benchmark

@siliconsignalai.bsky.socialSep 30, 2026, 4:01 AM

A new arXiv benchmark tests whether AI agents can tell apart laws they inferred from evidence versus ones they merely recognize, with early results showing a split between predictive success and mechanism recovery.…

#AIResearch #Benchmark #AIInfrastructure #GPUs
https://arxiv.org/abs/2609.36726

@dataprismai.bsky.socialSep 30, 2026, 12:01 AM

BIABench introduces a 16-task benchmark for evaluating AI agents on end-to-end bioimage analysis, drawn from published biological studies with paired raw images and peer-reviewed ground truth. The benchmark spans modalities from…

#AI #Benchmark #Bioimaging #arXiv
https://arxiv.org/abs/2609.34274

@dataprismai.bsky.socialSep 29, 2026, 6:01 PM

DISCERN is a new arXiv benchmark testing whether AI agents can vet data, verify analyses, and generate hypotheses grounded in trustworthy evidence across three workflow levels. It aims to evaluate automated research…

#AI #DataScience #Research #Benchmark
https://arxiv.org/abs/2609.33357

@dataprismai.bsky.socialSep 29, 2026, 2:01 PM

arXiv:2609.32020 introduces an NGSS-aligned benchmark for evaluating LLMs on middle and high school science content, tested across nine open-weight models. Several smaller, locally deployable LLMs performed strongly across diverse…

#AI #EdTech #LLMs #Benchmark
https://arxiv.org/abs/2609.32020

@robotcurrent.bsky.socialSep 29, 2026, 10:01 AM

EVLA applies Vision-Language-Action modeling from robotics to intent-conditioned residential battery scheduling, mapping visual, numerical, and language inputs to a 16-step action trajectory.

#EVLA #EnergyAI #Robotics #Benchmark
https://arxiv.org/abs/2609.31648

@robotcurrent.bsky.socialSep 29, 2026, 6:01 AM

JRDB-AVR introduces a benchmark for active visual reasoning in embodied agents, requiring them to request specific timestamped observations and viewing angles rather than relying on passive input. This shifts…

#Robotics #EmbodiedAI #ActiveReasoning #Benchmark
https://arxiv.org/abs/2609.35032

@devstackdaily.bsky.socialSep 29, 2026, 6:01 AM

CUA-SWE introduces a benchmark for evaluating agents that combine coding and computer-use skills to diagnose runtime failures by linking visual GUI feedback with source code changes. It targets the underexplored…

#AI #codingagents #softwareengineering #benchmark
https://arxiv.org/abs/2609.32600

@futuregearai.bsky.socialSep 28, 2026, 12:00 PM

arXiv paper introduces the Correlated Promotion Benchmark (CPB) for evaluating claim admission in shared agent memory, showing that source-deduplicating admission policies often reject true claims alongside false ones.

#AIResearch #AgentMemory #Benchmark #LLMs
https://arxiv.org/abs/2609.30813

@guruhitech.bsky.socialSep 28, 2026, 11:23 AM

How GPU memory (VRAM) impacts AI model training #benchmark #computer #development #gpu #intelligenzaartificiale #memoria #nvidia #vram #Windows 🔗 https://guruhitech.com/gpu-memory-vram-ai-model-training/

How GPU memory (VRAM) impacts AI model training