Transformers and LLM Foundations - Index

Routing Summary

The technical core of modern large language models, from four primary papers: Vaswani et al. (2017) on the Transformer, Kaplan et al. (2020) on scaling laws, Hoffmann et al. (2022) on compute-optimal training (Chinchilla), and Brown et al. (2020) on GPT-3 and in-context learning.

Concept Map

ConceptNoteTypeDepends OnKey Result
Cluster framingTransformers and LLM Foundations - OverviewoverviewProbability and Bayesian Inference; Overfitting and Information CriteriaArchitecture, objective, scaling and in-context learning in one chain; ties the papers together
AttentionScaled Dot-Product and Multi-Head AttentionconceptOverview; has variance ; heads at equal cost; Nadaraya-Watson reading
Transformer architectureTransformer Architecture and Positional EncodingconceptAttention; FFN with ; linear in ; 28.4 BLEU EN-DE
Autoregressive LM and pretrainingAutoregressive Language Modeling and PretrainingconceptArchitecture; Attention; causal mask; loss in nats per token; risk decomposition
Scaling lawsNeural Scaling LawsconceptAutoregressive LM, , ; ;
Compute-optimal trainingCompute-Optimal Training (Chinchilla)methodScaling lawsThree approaches give ; Chinchilla 70B on 1.4T tokens beats Gopher 280B
In-context learningIn-Context Learning and Few-Shot PromptingconceptAutoregressive LM; Attention; Scaling lawsFew-shot gains grow with model size; LAMBADA 86.4%, TriviaQA 71.2%; WiC at chance

Notes

  • Transformers and LLM Foundations - Overview — CONTAINS: four-paper summary table, definitions of Transformer and LLM, five organizing empirical regularities, reading order, relevance to marketing measurement (LLM elicitation, scaling-law design, attention as lag kernel, amortization), GPT-3 compute worked example.
  • Scaled Dot-Product and Multi-Head Attention — CONTAINS: query-key-value definition, Eq. 1, variance argument for , multi-head definition with projection shapes, three uses of attention (encoder, cross, masked decoder), Table 1 complexity comparison, head-count ablations, kernel-smoothing interpretation compared with GP regression, hand computation, NumPy sketch.
  • Transformer Architecture and Positional Encoding — CONTAINS: encoder and decoder layer anatomy, residual plus LayerNorm, FFN Eq. 2, weight tying, sinusoidal encoding with rotation-matrix proof, learned-embedding ablation, training recipe (Adam, warmup schedule Eq. 3, dropout, label smoothing), Table 2 BLEU and FLOPs, Table 3 ablations, parameter-count check of the 65M base model.
  • Autoregressive Language Modeling and Pretraining — CONTAINS: chain-rule factorization, cross-entropy and perplexity, causal masking, decoder-only definition, pretraining compared with fine-tuning, tokenization caveats, GPT-3 model sizes (Table 2.1) and data mixture (Table 2.2), Hoffmann risk decomposition, Transformer compared with LSTM context use, objective limitations, contamination, training-step pseudocode.
  • Neural Scaling Laws — CONTAINS: , , Eqs. 1.1-1.8 with fitted constants, design principles and overfitting criterion, learning-curve law, critical batch size, Kaplan allocation exponents, shape independence, transfer, the and contradiction and entropy conjecture, caveats, curve-fitting sketch.
  • Compute-Optimal Training (Chinchilla) — CONTAINS: constrained optimization statement, why Kaplan differed (learning-rate schedule, small models, curvature), three estimation approaches, Huber and log-sum-exp fitting, closed-form frontier with derivation, Table 2 exponents with bootstrap intervals, Table 3 sizes, 20 tokens per parameter, Chinchilla compared with Gopher results, limitations, numerical evaluation of the fitted law, media-budget allocation analogy.
  • In-Context Learning and Few-Shot Prompting — CONTAINS: meta-learning inner and outer loop, FT/FS/1S/0S definitions, evaluation protocol including likelihood normalization, scaling of ICL with model size, results table (LAMBADA, TriviaQA, arithmetic, SuperGLUE, WiC), limitations, hierarchical-Bayes and amortized-inference reading, prompt formats, few-shot classifier sketch.

Sources

  • Vaswani 2017 - Attention Is All You Need — Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. & Polosukhin, I. (2017), “Attention Is All You Need,” NeurIPS 2017. arXiv:1706.03762.
  • Kaplan 2020 - Scaling Laws for Neural Language Models — Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J. & Amodei, D. (2020), “Scaling Laws for Neural Language Models.” arXiv:2001.08361.
  • Hoffmann 2022 - Training Compute-Optimal Large Language Models — Hoffmann, J., Borgeaud, S., Mensch, A., et al. (2022), “Training Compute-Optimal Large Language Models.” arXiv:2203.15556.
  • Brown 2020 - Language Models are Few-Shot Learners — Brown, T. B., Mann, B., Ryder, N., Subbiah, M., et al. (2020), “Language Models are Few-Shot Learners,” NeurIPS 2020. arXiv:2005.14165.