Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@whitneylee.comOct 10, 2026, 5:53 PM

Scott Yak's team runs 200 evaluation scenarios against the Datadog MCP server in about 2 minutes, on a local computer. Speedy! How do they do it? Watch the full Datadog Illuminated episode: https://youtu.be/HZ4n1h8j8MY #Datadog #MCP #AIAgents #Evals #DatadogIlluminated

@whitneylee.comOct 9, 2026, 6:51 PM

The standard way to evaluate an AI tool: a human writes a question, works out the correct answer by hand, and checks if the AI matches it. Is there a better way? Watch the full Datadog Illuminated episode: https://youtu.be/HZ4n1h8j8MY #Datadog #MCP #AIAgents #Evals #DatadogIlluminated

@cubxxw.bsky.socialOct 4, 2026, 9:12 AM

Agents have been one patch after another: retrieval, memory, retries, approvals, tests. Every model upgrade eats some of that scaffolding back. Two things it can't eat: what counts as good, and who decides. So I spend my time on evals. If you can measure it, you can simplify it. #Evals

@dataprismai.bsky.socialOct 1, 2026, 4:00 AM

An arXiv study questions whether higher LLM judge scores on an agentic banking deck harness reflect real document improvements or just shifts in grading criteria. Across five judges and seventeen deliverables, the full system scored 20.4…

#AI #LLMs #FinTech #Evals
https://arxiv.org/abs/2609.39958

@futuregearai.bsky.socialSep 30, 2026, 12:01 AM

New arXiv paper proposes evaluating consumer AI action agents on both trust (no unauthorized actions) and task completion, arguing both depend heavily on the harness around the model rather than the model itself. The evaluation runs…

#consumerAI #AIAgents #evals
https://arxiv.org/abs/2609.33017

@promptfoundry.bsky.socialSep 29, 2026, 8:01 AM

RAGWarrant frames RAG deployment as a constrained evidence decision, normalizing evaluator outputs and applying predeclared quality and risk gates before emitting auditable PROMOTE, BLOCK, REJECT, or INCONCLUSIVE verdicts. Worth a look…

#RAG #GenAI #MLOps #Evals
https://arxiv.org/abs/2609.34179

@fcavrini.bsky.socialSep 22, 2026, 10:09 AM

#jev on my #evals test vs Sonnet 5: 10x+ faster, 100x+ cheaper, natural code integration.

Not just an improvement. A new enabler.

For example, we can run many more #llm-judges without the cost eclipsing the product itself.

@fcavrini.bsky.socialSep 21, 2026, 7:40 PM

Tested TypeSafe AI's #Jev for #evals. 29/30 same as Sonnet 5 on my context-awareness dataset — but 12x faster and 164x cheaper.

No more JSON parsing mess. It just gives you the choice and its probability.

Love seeing #specialized #AIs solving specific problems instead of just bigger models.

@rsamborski.bsky.socialSep 15, 2026, 10:27 AM

Do you test your AI agents before launch? 🤖

My colleague Jan-Felix Schmakeit shared 5 rules for evals:
- Match graders to sandboxes
- Write multi-step prompts
- Avoid prompt-grader mismatch
- Grade final outputs, not tool paths
- Curate distinct user datasets

Link👇

#AI #Evals #AgenticAI

@whitneylee.comSep 5, 2026, 3:46 PM

Scott Yak explains how his team evaluates the Datadog MCP server against a bajillion scenarios and data that goes stale before they can finish testing it, using a method that's a little bit spicy. ♫ Give it a watch! https://youtu.be/HZ4n1h8j8MY #Datadog #MCP #AIAgents #Evals #DatadogIlluminated