Grilled Cheese

ExploreLog inSign up

Mert Bulent Sariyildiz

@mbsariyildiz.bsky.social

134 Following185 Followers

Research scientist at Naver Labs Europe. https://mbsariyildiz.github.io/

PostsRepliesMedia
@mbsariyildiz.bsky.socialSep 1, 2026, 2:34 PM

Work with @fyavuz1.bsky.social and @dlarlus.bsky.social at NAVER LABS Europe
πŸ“„ Paper: arxiv.org/abs/2608.24759
🌐 Project page: blisgard.github.io/ideal_project/
πŸ’» Code: coming soon!

@mbsariyildiz.bsky.socialSep 1, 2026, 2:34 PM

Results (UNIC protocol: ImageNet cls, transfer, segmentation, depth): βœ… IDeaL > Dead Leaves in every setting, every task βœ… Students beat the weakest teacher on 3/4 tasks β€” even with just 1K samples βœ… At 1K images, IDeaL matches or BEATS a 1K ImageNet subset on classification

@mbsariyildiz.bsky.socialSep 1, 2026, 2:34 PM

Image decorrelation (L_ID): same trick at the image level. Final-layer CLS embeddings across the batch β†’ minimize pairwise similarity β†’ every generated image is globally distinct, as perceived by each teacher. No labels, no dataset statistics, no privileged info. Just the frozen teachers.

@mbsariyildiz.bsky.socialSep 1, 2026, 2:34 PM

How do we optimize? Two decorrelation losses, one idea: push representations apart. Patch decorrelation (L_PD): per layer, compute pairwise cosine similarity between patch representations, push off-diagonals to 0. Diverse patches = self-attention has work to do.

@mbsariyildiz.bsky.socialSep 1, 2026, 2:34 PM

Our key idea: the teachers themselves know what makes an image informative. So we treat pixels as learnable parameters and optimize Dead Leaves through the frozen teachers. The result: IDeaL samples that don't look real, but are maximally informative.

@mbsariyildiz.bsky.socialSep 1, 2026, 2:34 PM

Multi-teacher distillation (e.g. UNIC) combines complementary teachers (e.g. DINO, iBOT, DeiT-3, dBOT) into one student. But it assumes access to the teachers' training data. What if that data is private, licensed, or just gone? We asked: how far can we get without it?

@mbsariyildiz.bsky.socialSep 1, 2026, 2:34 PM

🧐 Can you distill knowledge from 4 vision teachers into one student using ZERO real images? πŸ₯³ Turns out: YES, and surprisingly well. πŸ“£ Our ECCV 2026 paper "IDeaL" closes most of the gap with real-image distillation, using optimized structured noise.

Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-11 02:37:10 EDT