Grilled Cheese

ExploreLog inSign up

Epoch AI

@epochai.bsky.social

0 Following1.6k Followers

We are a research institute investigating the trajectory of AI for the benefit of society.

epoch.ai

PostsRepliesMedia
@epochai.bsky.socialOct 9, 2026, 9:20 PM

Explore how AI acknowledgment rates vary by math subfield, use case, and provider:

epoch.ai/data/arxiv

@epochai.bsky.socialOct 9, 2026, 9:20 PM

In 3 of the 18 math subfields we track, more than half of arXiv papers by established authors now acknowledge using AI. Rates have increased rapidly: in differential geometry, the share climbed from ~8% of papers in July to ~57% in September.

Graph shows increasing acknowledgment of AI use in arXiv papers across combinatorics, differential geometry, and classical analysis since 2023.
@epochai.bsky.socialOct 8, 2026, 9:19 PM

This is a follow-up to our earlier report on latency scaling in frontier models:
epoch.ai/publication...

@epochai.bsky.socialOct 8, 2026, 9:19 PM

This isn’t conclusive evidence of a new architecture, but it suggests something has changed in how GPT-6.1 Sol handles long contexts.

@epochai.bsky.socialOct 8, 2026, 9:19 PM

A new architecture for GPT-6.1 Sol? OpenAI halved its cached-input price compared with GPT-6 Sol. Our measurements show it also handles long prompts faster.

Four-panel chart of time to first token versus input context up to 900,000 tokens for GPT-5.6 Sol, GPT-6 Astra, GPT-6 Sol and GPT-6.1 Sol, with Student-t fits. GPT-6.1 Sol starts higher but rises more slowly, reaching about 15 seconds at the longest context versus about 17 seconds for GPT-5.6 Sol and GPT-6 Sol.
@epochai.bsky.socialOct 8, 2026, 5:36 PM

Over time, we’ll expand this task suite, retire saturated tasks, and publish updated findings as new models are released. Read Kelly Hong’s full report:

epoch.ai/publication...

@epochai.bsky.socialOct 8, 2026, 5:36 PM

For example, GPT-6 Astra ran an experiment exploring why AI agents fail to learn with practice. It set AI agents’ token budgets too low. Instead of treating this as a mistake, it reported “sensitivity to the acquisition budget” as a key finding.

Excerpt from GPT-6 Astra's pilot findings, highlighting that 61 of 280 acquisition calls exhausted their token allowance before returning a probe, and concluding that the pilot primarily reveals 'sensitivity to the acquisition budget'.
@epochai.bsky.socialOct 8, 2026, 5:36 PM

When asked to propose a research project and pilot an experiment, models made simple mistakes in experiment setup — but presented them as interesting results.

@epochai.bsky.socialOct 8, 2026, 5:36 PM

Models struggle to pick up Epoch house style, even when given many examples. For instance, this diagram generated by Fable 5.1 is far more information-dense than what Epoch would produce.

Diagram generated by Fable 5.1 titled 'Newer models score higher at any token budget and keep gaining from more tokens, so constant elicitation understates recent progress', with two densely annotated panels of illustrative curves comparing constant and max elicitation across three model generations.
@epochai.bsky.socialOct 8, 2026, 5:36 PM

The headline scores give a rough ranking, but the qualitative findings are just as interesting. We see Epoch Automation Reports as a way of capturing anecdotes about model capabilities more rigorously. Through it, we found some common failure patterns.

@epochai.bsky.socialOct 8, 2026, 5:36 PM

Tasks range from making Epoch-style diagrams to designing basic research experiments. A human grader scores outputs based on our internal quality standards.

Screenshot of a text prompt asking an AI model to create an "epochified" diagram illustrating how constant-elicitation understates recent model progress relative to max elicitation.
@epochai.bsky.socialOct 8, 2026, 5:36 PM

Can AI automate Epoch? We're introducing Epoch Automation Reports to evaluate frontier models on realistic, open-ended tasks drawn from our own work. Claude Fable 5.1 and GPT-6 Astra lead, yet they are far from fully automating Epoch’s work.

Horizontal bar graph displays average task performance scores for 6 top AI models, showing that none can autonomously complete Epoch's work to internal standards.
@epochai.bsky.socialOct 7, 2026, 6:08 PM

We’re planning to periodically rerun InnovationEval with new, uncontaminated papers. We hope this will provide early signs if AI approaches automating AI R&D end-to-end, rather than performing individual tasks under human direction.

Read more at our website: epoch.ai/publication...

@epochai.bsky.socialOct 7, 2026, 6:08 PM

Unfortunately, newer models have seen the original innovation during their training, making the task significantly easier. However, even with this information, GPT-6 Astra and Claude Fable 5.1 struggled to reimplement it.

@epochai.bsky.socialOct 7, 2026, 6:08 PM

This behavior is likely reward hacking; in both cases, models reasoned that run selection might score well in an eval, despite being useless for actual research.

@epochai.bsky.socialOct 7, 2026, 6:08 PM

Both models made misleading claims about their work. Struggling to make progress, they instead ran several similar training runs, selectively reporting the best result. They were not candid about their submissions, failing to mention this would artificially inflate scores.

@epochai.bsky.socialOct 7, 2026, 6:08 PM

AI performance was underwhelming. Neither model achieved anything close to the human-authored reference. They reused existing methods from the literature and tuned hyperparameters, but struggled to create anything new.

@epochai.bsky.socialOct 7, 2026, 6:08 PM

The AI models had no information about SDPO. Hence, they either had to independently invent something like it, or invent another technique with comparable benefits, under similar constraints.

@epochai.bsky.socialOct 7, 2026, 6:08 PM

We instructed the AI models that their technique should improve performance on several benchmarks. We already knew that all of these could be improved by a recent human-authored post-training innovation: on-policy self-distillation (SDPO).

@epochai.bsky.socialOct 7, 2026, 6:08 PM

We gave Fable 5 and GPT-5.6 Sol 3000 GPU-hours and tasked them with developing a novel post-training technique that would improve on a standard pre-existing baseline (GRPO).

Older posts
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT