Twenty questions on the pipeline from tokens to tokens.
01What is a token?
- A. A character
- B. A word
- C. A subword unit chosen by a learned tokenizer
- D. A byte
Reveal answer
C. Tokens are subword units — neither characters nor words.
02Why does attention need positional encoding?
- A. Attention is expensive
- B. Attention is order-invariant by default
- C. Attention is stochastic
- D. Attention is autoregressive
Reveal answer
B. Without positional encoding, attention treats input as a bag of tokens.
03What does RoPE do?
- A. Adds a bias to attention
- B. Rotates Q and K vectors by position
- C. Randomizes positions
- D. Normalizes positions
Reveal answer
B. RoPE rotates queries and keys by an angle proportional to position.
04What is stored in the KV cache?
- A. Queries
- B. Keys and values of previous tokens
- C. Attention scores
- D. Logits
Reveal answer
B. The KV cache stores K and V from prior tokens so they need not be recomputed.
05What does temperature control?
- A. Model temperature
- B. Sharpness of the sampling distribution
- C. Learning rate
- D. Context length
Reveal answer
B. Temperature scales logits before softmax — lower is sharper, higher is flatter.
06What does top-p sampling do?
- A. Picks the top-p tokens by index
- B. Keeps tokens up to cumulative probability p
- C. Uses temperature p
- D. Chooses p random tokens
Reveal answer
B. Nucleus sampling keeps the smallest set of tokens whose cumulative probability exceeds p.
07What does the Chinchilla result say?
- A. Bigger is always better
- B. Compute-optimal training scales parameters and data roughly equally
- C. Fine-tuning is unnecessary
- D. MoE is optimal
Reveal answer
B. Chinchilla: for a fixed compute budget, scale parameters and tokens roughly in proportion.
08What is speculative decoding?
- A. Sampling with high temperature
- B. A draft model predicts, a big model verifies
- C. Predicting future context lengths
- D. Beam search variant
Reveal answer
B. Speculative decoding uses a small draft model with verification by the big model.
09What does MoE stand for?
- A. Model of Everything
- B. Mixture of Experts
- C. Multi-order Embedding
- D. Modular Operator Engine
Reveal answer
B. MoE = Mixture of Experts — many small FFNs with a router selecting few per token.
10Which sampler is the modern default for chat?
- A. Greedy
- B. Top-k=40
- C. Nucleus with p≈0.9
- D. Beam search
Reveal answer
C. Nucleus (top-p) sampling with p around 0.9 is the standard chat default.