Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@nexthorizonspace.bsky.socialOct 8, 2026, 3:43 PM

How do large language models “think”? Researchers are trying to uncover the hidden algorithms inside AI.

www.nexthorizon.space/2025/04/how-...

#AI #Science #Technology #Research #LLM #ChatGPT #MachineLearning #NLP #Interpretability #Data #Future #NeuralNets

@cipherpulseai.bsky.socialOct 8, 2026, 12:01 PM

New arXiv work introduces SafeEvo, a circuit-level interpretability framework tracing how refusal behavior in LLMs emerges and evolves through alignment training. It identifies weak refusal circuits in…

#AIAlignment #LLMSafety #Interpretability #Cybersecurity
https://arxiv.org/abs/2610.09600

@angieboggust.bsky.socialOct 7, 2026, 6:05 PM

Our research group at Apple is recruiting PhD interns working in human-centered AI, including #interpretability, #alignment, and #visualization. 🍎🤖

Please include the phrase “#human-ai-vis” in your resume or cover letter, and apply at: jobs.apple.com/en-us/detail...

@fredhohman.bsky.socialOct 7, 2026, 5:54 PM

Our research group at Apple is recruiting PhD interns working in human-centered AI, including topics such as #interpretability, #alignment, and #visualization. Past interns have published award-winning papers, prototyped new user experiences, and deployed interactive tools that impact billions. 📊

@siliconsignalai.bsky.socialOct 7, 2026, 2:01 PM

A new arXiv study finds that recognizing context follows a causal pattern across five language models, with a small set of attention heads (1–9 out of 128–1152) supporting a detector reaching 99.5–100% held-out accuracy. Whether that…

#AI #LLMs #Interpretability
https://arxiv.org/abs/2610.08200

@siliconsignalai.bsky.socialOct 7, 2026, 6:01 AM

Language models can now read neural network weights directly to predict properties and even simulate forward passes of small transformers with 99% holdout accuracy. This shifts interpretability from reactive activation analysis to…

#AI #LLMs #Interpretability
https://arxiv.org/abs/2610.07334

@tmlr-pub.bsky.socialOct 3, 2026, 8:30 AM

New #TMLR-Paper-with-Video:

Explaining with trees: interpreting CNNs using hierarchies

Caroline Mazini Rodrigues, Nicolas Boutry, Laurent Najman

https://tmlr.infinite-conf.org/paper_pages/zjyWZh5IiI

#cnns #interpretability #explanations

Explaining with trees: interpreting CNNs using hierarchies
@freegardener.bsky.socialOct 2, 2026, 11:32 PM

Routing probes can show fake gains from checkpoint selection, not real info. New audit finds 51% false detections #MachineLearning #AIResearch #Interpretability #VisionTransformers

https://freegardner.com/synapse/routing-probes-can-improve-without-new-information.html

@tresiwald.bsky.socialOct 2, 2026, 12:10 PM

Is an LM that acts Bayesian also Bayesian inside?
☝️ Only as far as its beliefs allow!

Fine-tuned on an optimal Bayesian model, LMs hold and use better beliefs than usual fine-tuning on golden answers.

Details 👇 or bayeslm.github.io

#interpretability #nlproc (1/🧵)

@aimointerp.bsky.socialOct 1, 2026, 10:07 AM

A month left in the AIMO #Interpretability Challenge @neuripsconf.bsky.social 🧠

📊 1k+ submissions predicting whether #LLMs reason robustly on Olympiad-level math problems.
⚠️ Remember: high eval acc ≠ high test score.
⏳ Still time until 1 Nov.

👉 aimo-interp.github.io

 1k+ submissions predicting whether #LLMs reason robustly on Olympiad-level math problems.
@siliconsignalai.bsky.socialSep 30, 2026, 10:01 AM

New research traces a three-stage symbolic reasoning circuit in LLMs that's detectable well before high accuracy, with per-head causal contributions growing up to 8x from 1- to 10-shot. Function vectors can rescue…

#AI #LLMs #MachineLearning #Interpretability
https://arxiv.org/abs/2609.36265

@tmlr-pub.bsky.socialSep 30, 2026, 12:27 AM

New #Reproducibility Certification:

A Reproduction Study of Weight-Based Mechanistic Interpretability in Bilinear MLPs

Itay Erlich, May Ben Zion, Jacob Shashoua

https://openreview.net/forum?id=6k7qRdz7rD

#autoencoders #interpretability #interpretable

@muruganpandian.bsky.socialSep 27, 2026, 8:04 PM

(How well humans understand the how an AI system make its decisions.). It also talks about AI ethics, governance, and a "Hippocratic Oath" equivalent for Machine Learning developers. Link to paper - www.nature.com/articles/s41...

#machinelearning #ethics #interpretability #transparency

@trynoguard.bsky.socialSep 26, 2026, 1:00 AM

Forgot about "activation patching"? Start with a specific effect you want to eliminate—like a model’sCHUbias—then clamp the activation of the suspected MLP layer to its mean value. If the bias vanishes, you've located the weight circuit. #LLM #Interpretability

@spaisee.bsky.socialSep 22, 2026, 10:05 AM

A “deception” feature isn’t proof. CHIVE asks: can interpretability predict how a model responds to new prompts?

#AI #Interpretability #MechanisticAI https://spaisee.com/article/chive-s-warning-seeing-inside-a-model-is-not-the-same-as-explaining-it

@informaq.bsky.socialSep 22, 2026, 3:59 AM

Quantum transformers can be intrinsically interpretable through mutual information analysis. We prove entanglement drives computation via ablation on real quantum hardware, offering physics-grounded interpretability insights unavailable in classical AI.

#QuantumML #Interpretability #Research

@tmlr-pub.bsky.socialSep 21, 2026, 4:20 PM

The Intrinsic Dimension of Prompts in Internal Representations of Large Language Models

Karthik Viswanathan, Yuri Gardinazzi, Giada Panerai, Alberto Cazzaniga, Matteo Biagetti

Action editor: Zachary Charles

https://openreview.net/forum?id=rBEgNAslpY

#softmax #interpretability

@trynoguard.bsky.socialSep 21, 2026, 11:00 AM

Did you know you can locate a LLM's "refusal vector" by analyzing residual streams? Abraham et al. (2023) showed that steering a model toward this specific direction triggers refusals, regardless of the prompt. What other latent traits can we isolate? #LLM #Interpretability

@sergiocuellar.bsky.socialSep 18, 2026, 6:39 PM

"Three techniques for making machine learning model predictions interpretable: SHAP, LIME, Integrated Gradients. Each method offers unique strengths for explaining model behavior. #MachineLearning #Interpretability #DataScience"

@thedailytechfeed.comSep 17, 2026, 8:46 PM

AI monitors now police rogue agents with probes & chain-of-thought detection. #AI #AISafety #AIAgents #Observability #Interpretability #SecurityNews https://thedailytechfeed.com/how-more-ai-is-becoming-the-answer-to-rogue-agents/

Load more