New benchmark TCSAlgBench evaluates LLMs on research-level theoretical computer science proofs, featuring 398 theorem-level challenges drawn from 138 STOC and COLT 2026 papers.
#TCSAlgBench #AIResearch #TheoremProving #LLMs
https://arxiv.org/abs/2609.35606
