A new benchmark called DataSense-Bench asks whether frontier AI agents can reliably select training data by writing and executing analysis code, without actual model training or evaluation access. It probes data…
#AIBenchmarks #DataSelection #LLM #MLResearch
https://arxiv.org/abs/2610.12190
