Uses layer-wise hidden-state interventions to quantify how LLMs (Qwen, Llama, Gemma) shift reliance between query-routing information and target knowledge during question answering.
From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge (Wei et al., arXiv:2609.11859, September 2026) uses layerwise interventions to demonstrate that LLMs functionally separate routing (query direction) from content (answer support). The study identifies a routing–content handoff where causal control shifts from parameter routing in earlier layers to answer-supporting content in later layers.
Key findings across Qwen, Llama, and Gemma include:
From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge investigates the internal mechanisms by which large language models answer factual questions using parametric knowledge rather than external retrieval. The paper studies Qwen, Llama, and Gemma through layer-wise hidden-state interventions, asking not only what answer a model produces, but which internal representations support that answer at different depths. It distinguishes two functional components: query-routing information, which helps the model identify what entity, relation, or answer slot is being requested, and target knowledge, which encodes the factual content about the answer itself. By intervening on hidden states across layers and measuring the resulting changes in behavior, the authors construct a layer-resolved account of how these two components contribute to final answers.
A key contribution is the paper’s decomposition of factual answering into a staged process rather than a single monolithic “memory lookup.” The work suggests that models do not simply retrieve an answer from a fixed location in the network; instead, they shift reliance from representations that route the query toward representations that carry or instantiate the target fact. This shift is not necessarily uniform across architectures, and the study provides a quantitative way to measure how much each layer depends on query routing versus target-specific knowledge. In doing so, it offers a more mechanistic view of parametric recall, moving beyond black-box accuracy toward an account of how internal knowledge is accessed, integrated, and used during generation.
The paper matters because it provides a principled method for probing the internal use of knowledge in LLMs, with implications for interpretability, knowledge editing, hallucination mitigation, and hybrid retrieval systems. If models can be shown to rely on identifiable target-knowledge representations at particular layers, then interventions can be designed more precisely—for example, to monitor factual confidence, correct stale or incorrect parametric memories, or decide when external retrieval is needed. More broadly, the work reframes “knowing” in LLMs as a distributed, layer-dependent process, offering a useful framework for understanding how parametric information becomes an answer.