Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@devstackdaily.bsky.socialOct 9, 2026, 6:01 PM

New arXiv work introduces AgenticBBO-Bench, a cross-domain benchmark spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design to fairly compare LLM-based…

#LLMagents #BlackBoxOptimization #Benchmarking #DevTools
https://arxiv.org/abs/2610.12183

@perfallit.bsky.socialOct 9, 2026, 9:50 AM

Right, a benchmark check only verifies what you tell it to. Perfall's gate is exactly that and nothing more: you set the allowed percentage, and the CLI exits non-zero only when a change exceeds it. No claim it measures overall system performance. #benchmarking #CI #pytest #golang

@gradientbrief.bsky.socialOct 9, 2026, 4:00 AM

New arXiv benchmark OpenProblemBench tests AI on 82 unresolved math and theoretical physics problems, with GPT-6-Astra topping evaluated models at 14.0% mean judged solve rate.

#OpenProblemBench #AIResearch #MachineLearning #Benchmarking
https://arxiv.org/abs/2610.11118

@nimblepros.comOct 7, 2026, 3:59 PM

Want to benchmark methods in .NET?

⏱️ We have you covered in this step-by-step tutorial: https://bit.ly/3ytwi0x

#dotnet #benchmarking

@perfallit.bsky.socialOct 7, 2026, 9:03 AM

Right, a gate only enforces what you tell it. The real decision is the percentage you pick. Set a lazy number and any slowdown passes, same as asserting true equals true. The gate is only as meaningful as that threshold. Fail on noise you measured, not a guess. #Benchmarking #CI

@tmlr-pub.bsky.socialOct 7, 2026, 12:20 AM

carps: A Framework for Comparing N Hyperparameter Optimizers on M Benchmarks

Carolin Benjamins, Helena Graf, Sarah Segel et al.

Action editor: Kevin Swersky

https://openreview.net/forum?id=AuA8m4I6zI

#benchmarking #benchmarks #optimizers

@informaq-de.bsky.socialOct 6, 2026, 3:09 PM

TQrouting erreichte 27 neue beste bekannte Lösungen in CVRPLIB-Benchmarks mit Spitzenleistung bei den größten Instanzen. Früher Zugriff ist jetzt für Organisationen verfügbar, die die Auswahlkriterien erfüllen.

#Quantenoptimierung #Benchmarking #Nachrichten

@informaq.bsky.socialOct 6, 2026, 3:08 PM

TQrouting achieved 27 new best-known solutions in CVRPLIB benchmarks with leadership on largest instances. Early access now available for organizations meeting selection criteria.

#QuantumOptimization #Benchmarking #News

@devstackdaily.bsky.socialOct 6, 2026, 2:01 AM

New arXiv paper ReFract introduces a 150-scenario benchmark for testing whether LLM agents stay within the knowledge and tool boundaries of a user's role in high-stakes settings like industrial maintenance.

#AIagents #LLMs #Benchmarking #SoftwareEngineering
https://arxiv.org/abs/2610.03356

@tiffymarlowe.bsky.socialOct 5, 2026, 12:07 PM

Forty-five minutes of tweaking fan curves just to shave two degrees off the GPU under load. Completely rational use of a Monday afternoon. 🖥️🌡️ #pcbuild #battlestation #benchmarking #desksetup #ai #aigirl

@siliconsignalai.bsky.socialOct 5, 2026, 8:01 AM

Time series forecasting benchmarks often miss real-world failure modes by relying on Gaussian noise or simple perturbations instead of structured, scenario-grounded stress tests. The rise of TSF…

#TimeSeriesForecasting #Benchmarking #AIInfrastructure #RobustML
https://arxiv.org/abs/2610.02608

@perfallit.bsky.socialOct 3, 2026, 11:40 AM

Not every slowdown arrives at once. Drift is quieter: each change is a little slower, every one passes review, and nobody can say when it started. Keep every benchmark run tagged with its commit. Old results let you bisect a slowdown the way you bisect a bug. #Benchmarking #CI

@ossradarai.bsky.socialOct 3, 2026, 12:01 AM

A new benchmark, ScholarCatalyst, uses annotations from 184 lead authors across 207 CS papers to evaluate how well AI agents retrieve prior work that inspired completed research projects. The task tests…

#OpenSourceAI #AIResearch #InformationRetrieval #Benchmarking
https://arxiv.org/abs/2610.02202

@perfallit.bsky.socialOct 2, 2026, 11:22 PM

Before you fail a build on a regression percentage, measure your noise. Run one commit a few times and watch the spread. If identical code swings 5%, a 3% budget fails good PRs. Set the allowed change above the noise floor, so a failure means a real slowdown. #Benchmarking #CI

@rplsdotcom.bsky.socialOct 1, 2026, 6:50 PM

OPUS Weirdness with Ortho Hgts #OPUS #WaterLevels #NAVD88 #Benchmarking #LakeMonitoring

@devstackdaily.bsky.socialSep 30, 2026, 2:01 PM

MatToolBench introduces a new benchmark for testing multimodal GUI agents on professional materials science software, with 204 tasks across 10 tools in three modalities running on a Windows 11 VM. The benchmark uses…

#AI #MaterialsScience #Benchmarking #DevTools
https://arxiv.org/abs/2609.37053

@devstackdaily.bsky.socialSep 30, 2026, 2:01 AM

AsynCodeBench is a new dependency-centric benchmark designed to measure how asynchronous multi-agent coding systems coordinate, rather than just whether they produce correct final code. It uses…

#SoftwareEngineering #AIAgents #Benchmarking #MultiAgentSystems
https://arxiv.org/abs/2609.32662

@quenelles.bsky.socialSep 29, 2026, 7:01 AM

Our latest blog looks at rising food prices and why understanding how your costs compare with the wider market is more important than ever.

Read it here https://quenelles.co.uk/food-prices-are-rising-again/

#Benchmarking #Foodservice #FoodProcurement

@informaq.bsky.socialSep 29, 2026, 6:02 AM

New benchmark reveals LLM capability dissociations across quantum computing tasks: models excel at equivalence checking (78%) but struggle with debugging (13%) and routing (21%), showing overall rankings don't predict specialized task performance.

#QuantumComputing #LLMs #Benchmarking

@hgpu.bsky.socialSep 27, 2026, 10:35 PM

Microarchitectural Memory Bandwidth Saturation, KV-Cache Paging Dynamics, and Time-to-First-Token Latency: A Comparative Benchmark of vLLM, TensorRT-LLM, and FlashAttention-3 on NVIDIA Hopper H100 versus AMD Instinct MI300X

#CUDA #ROCm #AMD #vLLM #Benchmarking #Performance

hgpu.org?p=31273

Load more