We are a research institute investigating the trajectory of AI for the benefit of society.
Explore how AI acknowledgment rates vary by math subfield, use case, and provider:

@epochai.bsky.social
We are a research institute investigating the trajectory of AI for the benefit of society.
Explore how AI acknowledgment rates vary by math subfield, use case, and provider:
In 3 of the 18 math subfields we track, more than half of arXiv papers by established authors now acknowledge using AI. Rates have increased rapidly: in differential geometry, the share climbed from ~8% of papers in July to ~57% in September.
This is a follow-up to our earlier report on latency scaling in frontier models:
epoch.ai/publication...
This isn’t conclusive evidence of a new architecture, but it suggests something has changed in how GPT-6.1 Sol handles long contexts.
A new architecture for GPT-6.1 Sol? OpenAI halved its cached-input price compared with GPT-6 Sol. Our measurements show it also handles long prompts faster.
Over time, we’ll expand this task suite, retire saturated tasks, and publish updated findings as new models are released. Read Kelly Hong’s full report:
For example, GPT-6 Astra ran an experiment exploring why AI agents fail to learn with practice. It set AI agents’ token budgets too low. Instead of treating this as a mistake, it reported “sensitivity to the acquisition budget” as a key finding.
When asked to propose a research project and pilot an experiment, models made simple mistakes in experiment setup — but presented them as interesting results.
Models struggle to pick up Epoch house style, even when given many examples. For instance, this diagram generated by Fable 5.1 is far more information-dense than what Epoch would produce.
The headline scores give a rough ranking, but the qualitative findings are just as interesting. We see Epoch Automation Reports as a way of capturing anecdotes about model capabilities more rigorously. Through it, we found some common failure patterns.
Tasks range from making Epoch-style diagrams to designing basic research experiments. A human grader scores outputs based on our internal quality standards.
Can AI automate Epoch? We're introducing Epoch Automation Reports to evaluate frontier models on realistic, open-ended tasks drawn from our own work. Claude Fable 5.1 and GPT-6 Astra lead, yet they are far from fully automating Epoch’s work.
We’re planning to periodically rerun InnovationEval with new, uncontaminated papers. We hope this will provide early signs if AI approaches automating AI R&D end-to-end, rather than performing individual tasks under human direction.
Read more at our website: epoch.ai/publication...
Unfortunately, newer models have seen the original innovation during their training, making the task significantly easier. However, even with this information, GPT-6 Astra and Claude Fable 5.1 struggled to reimplement it.
This behavior is likely reward hacking; in both cases, models reasoned that run selection might score well in an eval, despite being useless for actual research.
Both models made misleading claims about their work. Struggling to make progress, they instead ran several similar training runs, selectively reporting the best result. They were not candid about their submissions, failing to mention this would artificially inflate scores.
AI performance was underwhelming. Neither model achieved anything close to the human-authored reference. They reused existing methods from the literature and tuned hyperparameters, but struggled to create anything new.
The AI models had no information about SDPO. Hence, they either had to independently invent something like it, or invent another technique with comparable benefits, under similar constraints.
We instructed the AI models that their technique should improve performance on several benchmarks. We already knew that all of these could be improved by a recent human-authored post-training innovation: on-policy self-distillation (SDPO).
We gave Fable 5 and GPT-5.6 Sol 3000 GPU-hours and tasked them with developing a novel post-training technique that would improve on a standard pre-existing baseline (GRPO).