Motivation
In a standard Transformer layer, the setup is:
The attention output is:
The usual story is that Q, K, and V play different roles. Query is the current token’s retrieval request. Key is the address by which a past token can be found. Value is the payload that gets moved after attention weights are computed. The paper behind this note asks whether those roles really require three independent projection matrices.
The answer is not “all sharing is harmless.” It is sharper than that. Sharing Q and K is costly for language modeling because it weakens directional lookup. Sharing K and V is much less damaging, and it directly cuts the KV cache.
The most interesting variant is therefore “Q-K=V.”
This is not a new attention operator. It is a simple equality constraint on the projection axis. That simplicity is the point.
Why KV cache is the right metric
For autoregressive LLM inference, the expensive object is not only the current attention computation. It is the stored history. Every generated token needs access to the keys and values of previous tokens across layers. As context windows grow, this KV cache becomes a central memory bottleneck.
If the batch size is , the number of layers is , the sequence length is , the number of heads is , the head dimension is , and the dtype uses bytes, the standard cache size is:
The factor of two is there because both K and V are stored. If K and V are tied, only one representation has to be cached:
So the cache reduction is exactly 50 percent. This part is arithmetic, not an empirical claim.
That matters because parameter savings alone would not be very exciting. Reducing three projection matrices to two saves some parameters and projection compute, but attention projections are only part of the full Transformer. The stronger deployment argument is that K=V attacks the stored inference state directly.
Three sharing patterns
There are three natural variants.
First, Q=K-V ties query and key:
The problem is that the raw score matrix is symmetric before masking or positional corrections:
For non-causal vision or set-like tasks, this may not be fatal. For language modeling, it is a bad constraint. The current token asking for a past token is not the same relation as the past token asking for the current token. Q=K also does not reduce the KV cache because K and V still both exist.
Second, Q-K=V ties key and value:
This preserves the asymmetry of the attention score because Q and K remain separate:
In general these are not equal. The model still has one representation for “what I am looking for” and another representation for “how I can be found.” It merely forces the searchable address and the delivered payload to share a representation.
Third, Q=K=V is the aggressive collapse. It saves parameters and halves the cache, but it combines both bottlenecks: symmetric score structure and no independent payload representation. Unsurprisingly, this is the variant that degrades most in language modeling.
Why K=V is plausible, but not guaranteed
The best intuition is role separation. Q and K implement a directional lookup relation. K and V are both representations of past tokens. The address-payload distinction is real, but it may be less essential than the request-address distinction.
In retrieval language, Q is the current token’s request, K is the past token’s address, and V is the past token’s payload.
Q=K says the request and the address must live in the same representation space. That is a strong constraint on directional attention. K=V says the address and the payload must share a representation. That is still a bottleneck, but it leaves directional lookup intact.
The paper’s empirical analysis supports this story: in trained unconstrained QKV models, K and V projection matrices appear more similar to each other than Q is to either of them. This is useful evidence for redundancy. It is not a theorem.
The logical gap is important. Observing that trained matrices are similar does not prove that training under the equality constraint will find an equally good solution. The stronger claim would require a bound on the loss gap between unconstrained optimization and the constrained space where . The paper does not provide that kind of guarantee.
So the correct interpretation is: “K=V gives a guaranteed memory reduction. Small perplexity loss is an empirical observation.”
Both statements matter. Mixing them would overstate the result.
Relationship to GQA and MQA
Q-K=V should not be read as a replacement for GQA or MQA. GQA and MQA reduce cache by sharing along the head axis. Projection sharing reduces cache along the projection axis. These are orthogonal design choices.
For a GQA variant with KV groups, the cache is:
Relative to standard multi-head attention, the reduction is . If K=V is added on top of GQA, the cache becomes:
with reduction . For example, with and , Q-GQA reduces cache by 87.5 percent. With MQA, , so the reduction is about 96.9 percent.
This is the practical reason the method is interesting even if standalone Q-K=V is not always better than GQA or MQA on perplexity-cache trade-off. Its claim is not “do not use GQA.” Its claim is “there is another compression axis that can be stacked with GQA.”
What the experiments suggest
The reported language-modeling results are the main evidence. At roughly 300M parameters, Q-K=V gives a modest validation perplexity degradation while reducing cache by 50 percent. Q=K-V has similar or worse degradation but no cache saving. Q=K=V cuts cache but loses much more quality. That ranking is exactly what the role-separation argument predicts.
At a larger 1.2B scale, the relative ranking remains similar, and Q-K=V degradation becomes slightly smaller. That is suggestive. It does not settle the scaling question. A 1.2B model trained on 10B tokens is meaningful, but it is still far from proving the behavior at 7B, 13B, or 70B scale.
The downstream results are also encouraging but narrow. Maintaining average accuracy on benchmarks such as ARC, HellaSwag, PIQA, and WinoGrande is useful evidence that K=V does not immediately destroy general capability. It does not prove robustness for instruction following, coding, tool use, precise copying, or long-context retrieval. Those are exactly the places where the value representation may matter more.
The synthetic and vision tasks tell a different story: Q=K can work surprisingly well when directionality is less central, especially with positional mechanisms that break symmetry. That is a useful negative control. It shows that the harm of Q=K is not universal; it is task-structure dependent.
Critical assessment
The strongest part of the paper is the framing. It does not introduce a complicated approximation to attention. It asks whether one accepted architectural degree of freedom is actually necessary. The resulting method is easy to implement, easy to combine with head sharing, and directly tied to a deployment bottleneck.
The weak part is theoretical support for quality retention. Cache reduction is exact. Parameter and projection-compute reductions are exact. The connection between full QKV collapse and recurrent-state updates in linear attention is algebraically interesting. But none of this proves that softmax LLMs with K=V should preserve perplexity, factual recall, or long-context behavior.
There is also a deployment question. If a serving stack already uses strong GQA or MQA, the incremental value of K=V depends on the acceptable quality loss at the target scale. A 3 to 5 percent perplexity degradation may be attractive for edge or on-device models. It may be unacceptable for a high-end model where quality is the product. The trade-off is not universal.
The most precise reading is therefore: “Q-K=V is not a better Transformer by default. It is a clean additional point on the memory-quality Pareto frontier.”
That is enough to make the paper useful. In long-context serving, on-device LLMs, and high-throughput inference, a simple equality constraint that halves one part of the cache deserves attention, even if it remains an empirical design choice rather than a theorem-backed replacement for standard QKV.
References
Kayyam, A., Gopal, A. M., & Lewis, M. A. (2026). Do Transformers Need Three Projections? Systematic Study of QKV Variants. arXiv preprint arXiv:2606.04032.