I'm again grateful to @theverge.com for publishing my newest long-form essay on the limits of the computational model of the mind, the benefits of understanding both biological and cultural evolution to explain human thinking, and the danger of cognitive hot dogs in education.
Sean Trott
@seantrott.bsky.social
And I'm entirely sympathetic to the argument that we shouldn't call that distributed system an LLM ("compound system" is sometimes used; or, nowadays, "AI agent"), and that it's still worthwhile understanding what the LLM part of it can and can't do. But again, that's not a deflationary claim.
Understanding this is important for thinking about what the current generation of AI can do (also very relevant to discussions around safety). These are systems explicitly trained (through RLVR/CoT) to decompose tasks into sub-tasks, and they use external tools to do help accomplish them.
Right. Maybe it's because I'm already partial to the extended mind thesis, but the idea that LLMs can be hooked up to external tools, output API calls, then use the result of those API calls in their subsequent generations seems to me quite impressive and interesting.
Though I am sympathetic to the point about closed models. I think it's still useful scientifically to know which parts of a system are responsible for a given behavior insofar as we want to make generalizable inferences about that system, or parts of that system.
Yeah, to me, it is entirely reasonable to speak of the LLM + calculator (and other tools) as a kind of "coupled system". It does not feel deflationary to me to point out that LLMs can be trained to output API calls to external tools. That makes them more capable!
We've just released a new version of childes-db, my lab's interface to the CHILDES database of child language transcripts. It lets you work with CHILDES data from R through a versioned, reproducible API. A few updates 🧵 childes-db.stanford.edu/
This is all great advice!
It's "advice to grad students" season!
Here's a post I wrote several geological epochs ago, in 2019. My advice gets harder and harder to follow every year (even for me). But I think that means it gets better and better?
Understanding how learners conclude “X laughed Y” is incorrect is an age-old question, with several hypotheses, some of which have been ~impossible to disentangle! @tomyxw.bsky.social, @fredashi.bsky.social, and I use controlled rearing to shed light on this in our new EMNLP paper: 1/n
We've been working on this paper for a number of years now, and I think that's for the better, as it allowed us to integrate more of the growing body of empirical and theoretical work on LMs. Link again here for those who are interested: direct.mit.edu/opmi/article...
Our answer is no: LMs are existence proofs that a certain causal route is viable; the inferential value of such a result depends on the theoretical context, e.g., the alternative hypothesis at stake.
Finally, we discuss a range of best practices, as well as objections. For example: does a "successful" baselines result for a given task entail that LM-like mechanisms best account for the construct of interest in humans?
We then articulate the conditions under which distributional predictability is most potentially relevant: 1) when a study uses linguistic stimuli; and 2) it's investigating a claim about the *necessity* of some other factor or construct presumed to be inaccessible to LMs.
We first argue why distributional predictability matters from both theoretical and empirical perspectives. E.g., there's now a large body of work showing that measures like LM-derived surprisal are *sensitive* to various manipulations, and also *predictive* of human DVs on tasks.
Psychologists control for confounds in the design or analysis of their experiments (e.g., word frequency). We argue that we're in a position to also control for *distributional linguistic predictability* using contemporary LMs, and describe the cases in which we should do so (and how).
New paper out in @openmindjournal.bsky.social! "Large Language Models as Distributional Baselines for Language Tasks". With @jamichaelov.bsky.social , @camrobjones.bsky.social , Tyler Chang, and Ben Bergen. direct.mit.edu/opmi/article...
Finally, we discuss potential best practices, as well as objections to our argument. For example: does a "successful" baselines result for a given task entail that LM-like mechanisms are the best account of the corresponding construct of interest?
We then articulate the conditions under which it is most potentially relevant: 1) when a study uses linguistic stimuli; and 2) when it's investigating a claim about the *necessity* of some other factor presumed to be inaccessible to an LM (say, physical grounding).
We first argue why distributional predictability matters from both theoretical and empirical perspectives. E.g., there's now a large body of work showing that measures like LM-derived surprisal are *sensitive* to various manipulations, and also *predictive* of human DVs on tasks.
