Extracting Steering Vectors from J space

A developer explores using Jacobian space to derive steering vectors for Large Language Models, finding that simple behaviors can be steered using concept tokens. While effective for basic tasks, the method shows limitations and hallucinations when applied to complex model behaviors.
Why it matters
This research contributes to the growing field of mechanistic interpretability, helping developers better understand and control the internal decision-making processes of AI models.
I was reading about the jacobian space and how it can be used to verbalize the intermediate activations of an LLM to decode what it is most likely going to say or is thinking about. I wanted to test if we can use the j lens to arrive at a general activation steering vector from a couple of tokens related to the concept towards which we wanted to steer the model i.e inverting the j lens to have a general method of finding steering vectors from the concept tokens
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in