Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-11 02:37:10 EDT

Explore

PostsPeople
LatestRanked
@mptouzel.bsky.socialOct 6, 2026, 12:15 AM

I'm excited to sample the research frontier showcased at #COLM2026 in SF. I'm hosting our 2nd workshop on Social Sims with LLMs. sites.google.com/view/social-.... This area has blown up since our last workshop and we'll present the progress of our community building efforts alongside the submissions

@kylelo.bsky.socialOct 5, 2026, 11:51 PM

I'm at #colm2026, Mon-Thurs

Supporting two papers from @ai2.bsky.social days
๐ŸŸOlmo Hybrid, 7B hybrid model
๐Ÿ Cracks in the Foundation, long context recipes

Hoping to have fun chats w folks about scaling LM data & evals, โ˜บ๏ธ unbiased on pretrain, posttrain, all the ๐Ÿš‚s

@adamwiemerslage.bsky.socialOct 5, 2026, 10:32 PM

If youโ€™re at #COLM2026, checkout our work on cost efficient LLM evaluation! my coworker Michael will be presenting it at poster session 2 tomorrow starting at 4:30.

arxiv.org/abs/2604.01418

@eliyahabba.bsky.socialOct 5, 2026, 10:03 PM

In SF ๐ŸŒ‰ for #COLM2026!
Presenting ๐Ÿ‘‡ come say hi:

๐Ÿ—“๏ธ Tue 16:30 | Growing Pains๐Ÿ“Š
Keeping benchmark scores comparable as benchmarks grow

๐Ÿ—“๏ธ Wed 16:30 | Feelings to Metrics ๐Ÿงช
How people actually vibe-test LLMs, and how to measure it

๐Ÿ—“๏ธ Fri 9:15 | Growing Pains @ AIMS workshop

@chloenlp.bsky.socialOct 5, 2026, 3:41 PM

๐Ÿ˜„
Excited for #TADA2026 and #COLM2026.

@mariaa.bsky.socialOct 5, 2026, 1:54 PM

I'm in the Bay Area this week! ๐ŸŒŠ

Today: #TADA2026 @ Berkeley
Tues-Thurs: #COLM2026 @ SF
Fri: Giving an invited talk @ Berkeley

B@brenocon.bsky.socialOct 5, 2026, 2:09 AM

This year UMass Amherst CICS will be hiring tenure-track faculty in natural language processing! Job ad to be posted soon. For anyone at #TADA2026 or #COLM2026 this week, I or my colleague Hamed Zamani would love to chat or answer questions about it - just say hi or send me an email to meet!

@srajtmajer.bsky.socialOct 4, 2026, 1:57 PM

If you're at #COLM2026 this week in San Francisco, stop by on Wednesday @4:30pm to learn more about our work on benchmark data for hybrid human-AI generated fake news.
colm.cc
preprint: lnkd.in/g-vdFAZf

@ai2.bsky.socialOct 2, 2026, 6:41 PM

We're heading to #COLM2026 next week! Four days of workshops, posters, & talks on our latest fully open AI research, from Olmo Hybrid to evals for AI-assisted scientific writing. ๐Ÿงต

@joachimbaumann.bsky.socialOct 2, 2026, 6:31 PM

We'll present SWE-chat next week at COLM. Find us on Wednesday, Oct 7, 2026 during the poster session from 4:30 PM โ€“ 6:30 PM PDT in Grand Ballroom #111
Paper: arxiv.org/abs/2604.20779

#COLM2026

@najoung.bsky.socialOct 2, 2026, 3:49 PM

sadly I won't be at #COLM2026 this year but tinlab students Audrey and Jing will be there so talk to them! A little bit about what they'll be doing:

@richardjeanso.bsky.socialOct 1, 2026, 7:22 PM

Was excited to be on this paper which offers a novel way to measure AI creativity focusing on narrative tension/suspense. LLMs are bad at this because they don't plan out stories step by step, they just generate. We built a tool to do this. Excited that my coauthors are presenting this at #COLM2026!

@akhilayerukola.bsky.socialOct 1, 2026, 2:02 PM

Did you know gifting four koi ๐ŸŸ is offensive in Japan (4 = "shi" = death ๐Ÿ’€), or that chopsticks ๐Ÿฅขupright in rice is inappropriate in China (like funeral incense ๐Ÿ•ฏ๏ธ)? Top VLMs don't!

Our #COLM2026 paper introduces ๐ŸŒNormViz to measure visual norm understanding in 16 countries.

Best model: 25.3% ๐Ÿšจ ๐Ÿงต๐Ÿ‘‡

Overview of NORMVIZ-BENCH for visual norm understanding, consisting of 3,272 contrastive image pairs across 16 countries that differ only in the culturally relevant behavior. (a) A contrastive pair example from Japan; group accuracy requires a model to correctly classify both images in a pair. (b) Images are sourced via Text-to-Image generation,
inpainting, and Retrieval. (c) Benchmark statistics and pair type distribution.
@sfresearch.bsky.socialOct 1, 2026, 12:58 PM

(5/5) Fractured Chain-of-Thought Reasoning: examines how disruptions in chain-of-thought processes affect reasoning in language models.

arxiv.org/abs/2505.12992

#COLM2026

@sfresearch.bsky.socialOct 1, 2026, 12:58 PM

(4/5) RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation: grounds agent benchmarking in realistic user simulation.

arxiv.org/abs/2605.20204

#COLM2026

@sfresearch.bsky.socialOct 1, 2026, 12:58 PM

(3/5) InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation: introduces a scalable framework for simulating personalities grounded in real interviews.

arxiv.org/abs/2602.20294

#COLM2026

@sfresearch.bsky.socialOct 1, 2026, 12:58 PM

(2/5) MTA-Agent: An Open Recipe for Multimodal Deep Search Agents: presents an open recipe for building multimodal deep search agents.

arxiv.org/abs/2604.06376

#COLM2026

@sfresearch.bsky.socialOct 1, 2026, 12:58 PM

(1/5) We are pleased to announce our participation in COLM 2026, the Third Annual Conference on Language Modeling, at the Hilton Union Square in San Francisco, October 6โ€“9. Our researchers will present 4 accepted papers. Full list below โฌ‡๏ธ #COLM2026

@davideaustin.bsky.socialSep 30, 2026, 4:05 PM

LLMs interact with environments through textual representations. In my newest work at #COLM2026, we explore how these textual representations can introduce unintended biases in LLM decision-making...(1/4)

Paper: arxiv.org/abs/2608.16707

@hoytlong.bsky.socialSep 30, 2026, 1:53 PM

If you're at #COLM2026, come check out our poster for "Spoiler Alert" (arxiv.org/abs/2604.09854). We ask why LLM fiction is so bad at holding narrative tension, and create a metric to measure tension in short stories. What improves LLM fiction on this metric, it turns out, is narrative planning.

Load more