Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-11 02:37:10 EDT

Explore

PostsPeople
LatestRanked
@trynoguard.bsky.socialSep 21, 2026, 11:00 AM

Did you know you can locate a LLM's "refusal vector" by analyzing residual streams? Abraham et al. (2023) showed that steering a model toward this specific direction triggers refusals, regardless of the prompt. What other latent traits can we isolate? #LLM #Interpretability

@sergiocuellar.bsky.socialSep 18, 2026, 6:39 PM

"Three techniques for making machine learning model predictions interpretable: SHAP, LIME, Integrated Gradients. Each method offers unique strengths for explaining model behavior. #MachineLearning #Interpretability #DataScience"

@thedailytechfeed.comSep 17, 2026, 8:46 PM

AI monitors now police rogue agents with probes & chain-of-thought detection. #AI #AISafety #AIAgents #Observability #Interpretability #SecurityNews https://thedailytechfeed.com/how-more-ai-is-becoming-the-answer-to-rogue-agents/

@spaisee.bsky.socialSep 17, 2026, 1:54 PM

Can a “deception” label explain what an AI will do next? CHIVE puts interpretability claims to the test.

#AI #AIResearch #Interpretability https://spaisee.com/article/chive-s-warning-seeing-inside-a-model-is-not-the-same-as-explaining-it

@tmlr-pub.bsky.socialSep 16, 2026, 4:21 PM

Clarity: The Flexibility-Interpretability Trade-Off in Sparsity-aware Concept Bottleneck Models

Konstantinos P. Panousis, Diego Marcos

Action editor: Dmitry Kangin

https://openreview.net/forum?id=IyQEQBRR4M

#interpretability #bottleneck #learned

@informaq.bsky.socialSep 15, 2026, 10:27 PM

Novel workflow makes quantum machine learning interpretable by unfolding complex-valued neural networks into explicit formulas, enabling analysis of how quantum density matrices are processed through pruning, analytic readout, and symbolic rewriting stages.

#QuantumML #Interpretability #Research

@tmlr-pub.bsky.socialSep 15, 2026, 4:21 PM

Predicting Chain-of-Thought Correctness from Trajectory Geometry

Arjun Balaji

Action editor: Aditya Menon

https://openreview.net/forum?id=H9cBkEqVeY

#interpretability #accuracy #traces

@tmlr-pub.bsky.socialSep 10, 2026, 4:20 PM

DIMENSION DOMAIN CO-DECOMPOSITION: SOLVING PDES WITH INTERPRETABILITY

Shuyuan Shang, Ming Zhong, Ren Wang

Action editor: Yunbo Wang

https://openreview.net/forum?id=kuzkynVyRq

#decomposition #dimensional #interpretability

@sqirllab.bsky.socialSep 8, 2026, 3:24 PM

We are at #ECMLPKDD '26 presenting our work on #HDC anomaly detectors, #explainability for autonomous vessels and weakly supervised object #localization. We also gave a tutorial on the evolution of #interpretability methods.

Want to more? get in touch

#AI #ML #XAI #UAntwerp

@bayazitdeniz.bsky.socialSep 8, 2026, 12:21 PM

1/🚨 New paper #emnlp2026

Do multilingual #LLMs work in English? It depends on how you ask. In "Lingua Franca or Probing Artifact?", we probe the same hidden states with three latent language identification methods and get different answers. 🧵👇

#interpretability #NLProc

@jaom7.bsky.socialSep 7, 2026, 1:45 PM

The tutorial covered methods from the early days of deep learning, passing through conternporary mechanistic interpretability and more, state-of-the-art compositional approaches.

#interpretability #mechinterp #compinterp #compositionality #xai #AI #ML #AIMLAI #ECMLPKDD

@jaom7.bsky.socialSep 7, 2026, 1:45 PM

This year's edition of #AIMLAI at #ECMLPKDD 2026 is on.
This morning we had a tutorial "On the Evoluion of Interpretability Methods" given together with my collaborators @thomasdooms, Ward Gauderis and Geraint Wiggins.

#interpretability #mechinterp #compinterp #xai #AI #ML

@trynoguard.bsky.socialSep 6, 2026, 1:00 AM

Did you know about the "Activation Addition" technique from 2023? You can steer model behavior by adding a specific steering vector directly to the residual stream during a forward pass—bypassing the need for re-training or fine-tuning. #Interpretability #LLM

@freegardener.bsky.socialSep 5, 2026, 8:26 AM

New research reveals how LLMs "bluff" isn't random—it's a geometric axis storing their statistical priors. This #AIResearch #MachineLearning #LLMs #Interpretability

https://synapseglass.wordpress.com/2026/09/05/language-models-reveal-ignorance-through-geometric-prior-direction/