A “deception” feature isn’t proof. CHIVE asks: can interpretability predict how a model responds to new prompts?
#AI #Interpretability #MechanisticAI https://spaisee.com/article/chive-s-warning-seeing-inside-a-model-is-not-the-same-as-explaining-it

A “deception” feature isn’t proof. CHIVE asks: can interpretability predict how a model responds to new prompts?
#AI #Interpretability #MechanisticAI https://spaisee.com/article/chive-s-warning-seeing-inside-a-model-is-not-the-same-as-explaining-it