NVIDIA Jetson Thor finished MLPerf’s new edge agentic suite 6.4× faster than the reference stack on the same board.
The run kept ~96% of prompt tokens in a warm cache. Is long-context reuse now the real edge bottleneck, not peak TOPS?

NVIDIA Jetson Thor finished MLPerf’s new edge agentic suite 6.4× faster than the reference stack on the same board.
The run kept ~96% of prompt tokens in a warm cache. Is long-context reuse now the real edge bottleneck, not peak TOPS?
MLPerf Inference v6.1: Intel Xeon & Arc lebt, 512 MI355X skalieren und VR200 debütiert #MLPerf #Benchmarks
TensorRT Edge‑LLM just smashed the MLPerf Edge Agentic benchmark—clocking 6.4× speed on a Jetson AGX Thor. Curious how this edge AI leap could reshape your apps? Dive in for the full rundown. #TensorRT #MLPerf #JetsonAGXThor
Vera Rubin NVL72 achieves 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and 2.5x on DeepSeek-R1 in MLPerf Inference v6.1, with 99% scaling efficiency across four-rack deployments. #nvidia #mlperf
https://bymachine.news/nvidia-vera-rubin-nvl72-mlperf-inference
Tomorrow: David Kanter unveils MLPerf Inference v6.1 results on the AI Infra Summit main stage.
30 submitters. New accelerators. Next-gen platforms.
Results live on mlcommons.org.
Plus: what's ahead for MLPerf Endpoints.
10:25 AM. Pier 27, SF.