Grilled Cheese

ExploreLog inSign up

Sean Trott

@seantrott.bsky.social

44 Following92 Followers

PostsRepliesMedia
@benjaminjriley.bsky.socialOct 5, 2026, 11:58 AMReposted by @seantrott.bsky.social

I'm again grateful to @theverge.com for publishing my newest long-form essay on the limits of the computational model of the mind, the benefits of understanding both biological and cultural evolution to explain human thinking, and the danger of cognitive hot dogs in education.

theverge.comOur minds aren’t equipped to handle AIAI is junk food for the mind: easy, tempting and ultimately very bad for you.
7
@seantrott.bsky.socialOct 1, 2026, 2:52 PM

And I'm entirely sympathetic to the argument that we shouldn't call that distributed system an LLM ("compound system" is sometimes used; or, nowadays, "AI agent"), and that it's still worthwhile understanding what the LLM part of it can and can't do. But again, that's not a deflationary claim.

@seantrott.bsky.socialOct 1, 2026, 2:52 PM

Understanding this is important for thinking about what the current generation of AI can do (also very relevant to discussions around safety). These are systems explicitly trained (through RLVR/CoT) to decompose tasks into sub-tasks, and they use external tools to do help accomplish them.

@seantrott.bsky.socialOct 1, 2026, 2:52 PM

Right. Maybe it's because I'm already partial to the extended mind thesis, but the idea that LLMs can be hooked up to external tools, output API calls, then use the result of those API calls in their subsequent generations seems to me quite impressive and interesting.

@seantrott.bsky.socialOct 1, 2026, 2:30 PM

Though I am sympathetic to the point about closed models. I think it's still useful scientifically to know which parts of a system are responsible for a given behavior insofar as we want to make generalizable inferences about that system, or parts of that system.

@seantrott.bsky.socialOct 1, 2026, 2:30 PM

Yeah, to me, it is entirely reasonable to speak of the LLM + calculator (and other tools) as a kind of "coupled system". It does not feel deflationary to me to point out that LLMs can be trained to output API calls to external tools. That makes them more capable!

@mcxfrank.bsky.socialSep 25, 2026, 3:54 PMReposted by @seantrott.bsky.social

We've just released a new version of childes-db, my lab's interface to the CHILDES database of child language transcripts. It lets you work with CHILDES data from R through a versioned, reproducible API. A few updates 🧵 childes-db.stanford.edu/

Screenshot of the childes-db website: 'A flexible and reproducible interface to CHILDES.' childes-db 2026.1 contains 56,579 transcripts from 9,151 children across 437 corpora, with panels for an R API tutorial and interactive visualizations of mean length of utterance by child age.
@seantrott.bsky.socialSep 25, 2026, 2:23 PM

This is all great advice!

@kensycoop.bsky.socialSep 24, 2026, 5:44 PMReposted by @seantrott.bsky.social

It's "advice to grad students" season!

Here's a post I wrote several geological epochs ago, in 2019. My advice gets harder and harder to follow every year (even for me). But I think that means it gets better and better?

kensycooperrider.com/blog/advice-...

@kanishka.bsky.socialSep 9, 2026, 4:40 PMReposted by @seantrott.bsky.social

Understanding how learners conclude “X laughed Y” is incorrect is an age-old question, with several hypotheses, some of which have been ~impossible to disentangle! @tomyxw.bsky.social, @fredashi.bsky.social, and I use controlled rearing to shed light on this in our new EMNLP paper: 1/n

Title slide for “Disentangling Statistical Preemption from Entrenchment in Language Models’ Avoidance of Overgeneralization,” by Yixuan Wang, Freda Shi, and Kanishka Misra. Includes the main results plot and experimental design.
@seantrott.bsky.socialAug 31, 2026, 2:08 PM

We've been working on this paper for a number of years now, and I think that's for the better, as it allowed us to integrate more of the growing body of empirical and theoretical work on LMs. Link again here for those who are interested: direct.mit.edu/opmi/article...

@seantrott.bsky.socialAug 31, 2026, 2:08 PM

Our answer is no: LMs are existence proofs that a certain causal route is viable; the inferential value of such a result depends on the theoretical context, e.g., the alternative hypothesis at stake.

@seantrott.bsky.socialAug 31, 2026, 2:08 PM

Finally, we discuss a range of best practices, as well as objections. For example: does a "successful" baselines result for a given task entail that LM-like mechanisms best account for the construct of interest in humans?

@seantrott.bsky.socialAug 31, 2026, 2:08 PM

We then articulate the conditions under which distributional predictability is most potentially relevant: 1) when a study uses linguistic stimuli; and 2) it's investigating a claim about the *necessity* of some other factor or construct presumed to be inaccessible to LMs.

@seantrott.bsky.socialAug 31, 2026, 2:08 PM

We first argue why distributional predictability matters from both theoretical and empirical perspectives. E.g., there's now a large body of work showing that measures like LM-derived surprisal are *sensitive* to various manipulations, and also *predictive* of human DVs on tasks.

@seantrott.bsky.socialAug 31, 2026, 2:08 PM

Psychologists control for confounds in the design or analysis of their experiments (e.g., word frequency). We argue that we're in a position to also control for *distributional linguistic predictability* using contemporary LMs, and describe the cases in which we should do so (and how).

@seantrott.bsky.socialAug 31, 2026, 2:08 PM

New paper out in @openmindjournal.bsky.social! "Large Language Models as Distributional Baselines for Language Tasks". With @jamichaelov.bsky.social , @camrobjones.bsky.social , Tyler Chang, and Ben Bergen. direct.mit.edu/opmi/article...

@seantrott.bsky.socialAug 31, 2026, 2:01 PM

Finally, we discuss potential best practices, as well as objections to our argument. For example: does a "successful" baselines result for a given task entail that LM-like mechanisms are the best account of the corresponding construct of interest?

@seantrott.bsky.socialAug 31, 2026, 2:01 PM

We then articulate the conditions under which it is most potentially relevant: 1) when a study uses linguistic stimuli; and 2) when it's investigating a claim about the *necessity* of some other factor presumed to be inaccessible to an LM (say, physical grounding).

@seantrott.bsky.socialAug 31, 2026, 2:01 PM

We first argue why distributional predictability matters from both theoretical and empirical perspectives. E.g., there's now a large body of work showing that measures like LM-derived surprisal are *sensitive* to various manipulations, and also *predictive* of human DVs on tasks.

Older posts
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT