Agents have been one patch after another: retrieval, memory, retries, approvals, tests. Every model upgrade eats some of that scaffolding back. Two things it can't eat: what counts as good, and who decides. So I spend my time on evals. If you can measure it, you can simplify it. #Evals
Scott Yak's team runs 200 evaluation scenarios against the Datadog MCP server in about 2 minutes, on a local computer. Speedy! How do they do it? Watch the full Datadog Illuminated episode: https://youtu.be/HZ4n1h8j8MY #Datadog #MCP #AIAgents #Evals #DatadogIlluminated
