New arXiv paper introduces a self-improvement loop for reasoning models by training them jointly to predict, reverse-engineer, and apply solution ideas, using hindsight from supplied solutions. Applied to interactive…
#AI #MachineLearning #TheoremProving #Lean
https://arxiv.org/abs/2610.12168
