Skip to content

Query, Key & Value: The Core Intuition

Overview

Every word in a sentence plays one of three roles inside attention: Query (Q), Key (K), or Value (V). The Query asks "what am I looking for?", the Key answers "what am I?", and the Value supplies "what do I contribute?". Attention is the process of matching Queries to Keys and blending the corresponding Values. Nothing about Q/K/V is pre-programmed — all three are learned from data.

  • Analogy: attention works like a database / search engine lookup
  • Zero-setup: no hand-crafted rules; the model discovers the roles during training
  • Leads to: 02 Scaled Dot Product Attention — how Q, K, V combine mathematically
  • First seen in: "Neural Machine Translation by Jointly Learning to Align and Translate" (Bahdanau et al., 2015) — attention predates transformers!

The Search Engine Analogy

The simplest mental model for attention is a database query. Every word plays one of three roles — Query, Key, or Value:

You run a search engine (this is attention!):

                    ┌──────────────────────────────┐
  QUERY = what I'm  │  "best laptop under 1000$"   │
  searching for     └──────────────────────────────┘
                    ┌──────────────────────────────┐
  KEY = the labels  │  "laptop reviews"    → page 3│
  of every stored   │  "coffee machines"   → page 8│
  document          │  "budget laptops"    → page 1│
                    └──────────────────────────────┘
                    ┌──────────────────────────────┐
  VALUE = the       │  page 3 content, page 8      │
  actual content    │  content, page 1 content     │
                    └──────────────────────────────┘

  1. QUERY is compared against every KEY   → score (similarity)
  2. Scores are normalized                  → weights (how much to pay attention)
  3. We fetch the VALUES weighted by scores → output

Why is the search engine the right metaphor?

✅ Query ≠ Key:   "what I want" is different from "what each item is"
✅ Query vs all:  a query compares itself against EVERY key (all positions)
✅ Soft retrieval: results are mixed proportionally to relevance —
                  not a hard "yes/no" retrieval
✅ Learned ranking: the "relevance function" is trained, not hand-written

The Three Roles, in Detail

In a transformer, every token projects to three vectors using learned weight matrices:

For token xᵢ:
  Qᵢ = W_Q · xᵢ     ← "What am I looking for?"      (query)
  Kᵢ = W_K · xᵢ     ← "What am I?"                  (key)
  Vᵢ = W_V · xᵢ     ← "What do I contribute?"       (value)

Reading them in a concrete sentence

Example: in "The chef tasted the soup"
  "soup"   as Query: "is there anything modifying me? (adjectives, verbs...)"

  "chef"   sends a Key: "I am the subject, a person"
  "tasted" sends a Key: "I am a past-tense verb"
  "the"    sends a Key: "I am a determiner"

  The softmax weights decide which Keys match the Query best.

What each role is "about"

Role    Question it answers            Governed by            Emerges as
─────   ──────────────────────────     ────────────────       ─────────────────────
Q       "What am I seeking?"           the token's own needs   object-of-verb links,
                                                                 modification needs
K       "How should I be found?"       the token's identity    grammatical category,
                                                               topic, name
V       "What content do I give?"      the token's meaning     the "payload" that
                                                               gets passed along

💡 Key Insight: Q, K, V are learned during training. The model discovers which roles are useful for language understanding. Nobody tells "soup" to look for adjectives — it learns to, because that improves next-token prediction.


Why Three Vectors? (Why Not One?)

A reasonable question: why not just compute similarity between raw token embeddings? Three reasons:

1. Separation of "Matching" from "Content"

Similarity is about MATCHING, output is about CONTENT.

  "cat" and "kitten": high similarity (both noun-y, animal-y)
  But when generating the word after "the ___", you want the
  CONTENT of "cat"/"kitten" (the animal concept), not a blur.

  Q/K decide WHO talks to WHOM. V decides WHAT is said.
  Mixing them into one vector forces a compromise:
    - a vector that is both a good label AND good content
    - conflates "findability" with "meaning"

2. Different Projections → Different Semantics

Q, K, V come from three DIFFERENT learned matrices:

  Qᵢ = W_Q xᵢ   (d_model → d_k)
  Kᵢ = W_K xᵢ   (d_model → d_k)
  Vᵢ = W_V xᵢ   (d_model → d_v)

Each matrix warps the embedding space for its own purpose:
  W_K warps space so that related items land close together
     (two words can be "far apart" in K-space but close in V-space)
  W_Q warps space relative to what the token needs at this position
  W_V warps space so the payload is well-separated for the MLP

One embedding vector cannot serve all three purposes at once.

3. Keys Are Compared in a Lower-Dimensional Space

Typically d_k < d_model (e.g., d_model = 4096, d_k = 128).

Projecting to a smaller key space:
  - cheap dot products (fewer FLOPs per head)
  - forces the key to compress to "essentially relevant" features
  - the compression is learned, so it keeps what matters for matching

The Projection Matrices

Token embeddings X:  (seq_len, d_model)     e.g., (6, 4)

Learned parameters (all trainable, shared across positions):
  W_Q: (d_model, d_k)     W_K: (d_model, d_k)     W_V: (d_model, d_v)

Computing Q, K, V — one matrix multiply each:
  Q = X @ W_Q   → (seq_len, d_k)
  K = X @ W_K   → (seq_len, d_k)
  V = X @ W_V   → (seq_len, d_v)

Concrete (toy):
  X = 6 tokens × 4-dim embedding
  W_Q: 4×3, W_K: 4×3, W_V: 4×4
  → Q: 6×3, K: 6×3, V: 6×4

Typical dimensions in real models

Model         d_model   d_k (head)   heads
GPT-2 small   768       64           12
BERT base     768       64           12
LLaMA-2 7B    4096      128          32
GPT-3         12288     128          96

Common ratio: d_k = d_model / heads  (128, 64, 80, ...)

Attention as a Biological/Perceptual Idea

Attention isn't an AI invention — it mirrors natural attention:

🟢 Visual attention:   your eyes scan a scene and "focus" on salient
                       regions; peripheral regions contribute less
🟢 Cognitive attention: when reading, you focus on words that resolve
                       ambiguity ("bank" → depends on "river" or "money")
🟢 Machine translation origin: Bahdanau (2015) let the decoder "look back"
                       at relevant encoder words for each target word —
                       instead of squeezing the whole source into
                       one context vector

Transformer = the same idea, applied at every layer, for every token,
in parallel, with learned Q/K/V.

Where Q, K, V Come From: The Full Pipeline

Raw text
   │  tokenization
   ▼
Token IDs ──► Embedding lookup ──► X (token embeddings, may already
   │                                 include position info — see
   │                                 [05 Position Information](/02-llm-modeling/00-fundamentals/01-attention/05-position-information/))
   ▼
Linear layer: X·W_Q  → Q     X·W_K → K     X·W_V → V
   │
   ▼
scaled_dot_product_attention(Q, K, V)   ← [02 Scaled Dot Product Attention](/02-llm-modeling/00-fundamentals/01-attention/02-scaled-dot-product-attention/)
   │
   ▼
Context-aware token representations (one per position)

Key Takeaways

🔑 Q asks, K labels, V delivers — three learned roles per token
🔍 Fuzzy retrieval: relevance is soft (softmax weights), not binary
🔧 Three separate matrices: matching space (Q/K) ≠ content space (V)
🎓 Fully learned: no hand-crafted linguistic rules anywhere
🌍 Universal: self-attention, cross-attention and multi-head all reuse the same Q/K/V machinery
Parallel by design: X·W_Q, X·W_K, X·W_V are one batched matmul each