Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@hacks.grOct 10, 2026, 10:08 AM

Anthropic will remove live web access from all internal model evaluations after finding unwanted Claude behavior on real websites.

It says the incidents had mi…

https://en.hacks.gr/i-anthropic-kovei-tin-prosvasi-toy-claude-sto-diadiktyo-stis-esoterikes-dokimes/

#Anthropic #Claude #ModelEvaluation

Η Anthropic κόβει την πρόσβαση του Claude στο διαδίκτυο στις εσωτερικές δοκιμές
@gradientbrief.bsky.socialOct 9, 2026, 10:00 AM

Narrow finetunes of AI models contradict themselves when resampled on a new 175-question benchmark, revealing issues like identity conflation and introspection failures beyond mere ambiguity.

#AIsafety #ModelEvaluation #LLMResearch #Finetuning
https://arxiv.org/abs/2610.12129

@hacks.grOct 8, 2026, 11:11 PM

Arena raised $200 million, reaching a $3.1 billion valuation.

Its free platform lets people compare AI models by rating their results; the company also sells performance analys…

https://en.hacks.gr/i-arena-sygkentronei-200-ekat-dolaria-kai-apotimatai-sta-3-1-dis/

#Arena #AIModels #ModelEvaluation

Η Arena συγκεντρώνει 200 εκατ. δολάρια και αποτιμάται στα 3,1 δισ.
@thedailytechfeed.comSep 19, 2026, 1:25 PM

Vals wants you to trust AI benchmarks again: unseen tests, ethics, sector-specific rigor. #AI #Benchmarking #Vals #A16z #ModelEvaluation #AITrust https://thedailytechfeed.com/vals-aims-to-set-new-benchmark-standards-for-ai-validation/

@zachzang.bsky.socialSep 17, 2026, 5:00 PM

Perfect recall is trivially easy to fake

Recall 1.0 means zero false negatives, and says nothing at all about false positives.

#CDLS #MachineLearning #DataScience #ModelEvaluation

Full CDLS explanation, free: https://navyduck.com/deep-learning/cdls/q173-recall-of-one-meaning

NavyDuck infographic: Perfect recall is trivially easy to fake
@aidailypost.comSep 12, 2026, 4:46 PM

Anthropic’s Dario Amodei says it’s time for industry‑wide AI safety standards—think METR, recursive self‑improvement checks, and smarter deployment rules. Curious how this could reshape AI regulation? Dive in. #Anthropic #AISafety #ModelEvaluation

🔗 aidailypost.com/news/anthrop...

@codeonroids.bsky.socialSep 3, 2026, 10:59 AM

What should you know about Approve AI Features Only After the Evidence Gate?

https://www.razvantodica.com/blog/ai-business-thursday-2026-09-03-137673ed

#CodeOnRoids #ThursdayPost #AiFeatureApproval #ModelEvaluation