Did you know you can locate a LLM's "refusal vector" by analyzing residual streams? Abraham et al. (2023) showed that steering a model toward this specific direction triggers refusals, regardless of the prompt. What other latent traits can we isolate? #LLM #Interpretability
