技能 人工智能 深度学习学习伴侣与实践指南

深度学习学习伴侣与实践指南

v20260826
deep-learning-book
这是针对经典《深度学习》教科书的综合学习伴侣。它不仅索引了20个章节,还提供了从2016年到2026年的知识更新层,将核心概念与现代模型(如Transformer、Diffusion)相结合。工具集成了路径规划器、训练失败诊断器等四个实用工具,帮助用户系统学习理论,掌握知识的最新发展,并将其转化为可落地的工程决策。
获取技能
379 次下载
概览

Deep Learning — Study Companion

Source book: Deep Learning, Ian Goodfellow, Yoshua Bengio & Aaron Courville (MIT Press, 2016) · 20 chapters, 3 parts · read free at deeplearningbook.org · companion compiled 2026-08-25.

This is a companion, not a copy. The book is copyrighted, and its site states that the HTML-only format exists to discourage copying under the authors' MIT Press contract. Nothing here reproduces its text. Every chapter file is original synthesis — what the chapter establishes, how to use it, where it has aged — plus a link to the official chapter. Read the book at the link; use this to navigate it, keep it current, and turn it into decisions. See references/rights_and_use.md.

How to Use This Skill

  • No argument — load the core frameworks below.
  • A topic — ask about regularization, saddle points, partition function; resolved through the Topic Index, then that chapter file is read before answering.
  • chNN — load that chapter's file.
  • "is this still true?" — the 2016→2026 delta layer, in every chapter file and in references/book_to_2026_delta.md.
  • "where do I start?" — run scripts/reading_path_planner.py.

When asked about something outside these 20 chapters, say so and route to the delta reference rather than improvising the book's position on material published after it.


Core Frameworks & Mental Models

The (T, P, E) frame — ch05

Name the task, the performance measure, and the experience in one sentence before any model code. Most failed projects failed at P: an unstated metric, or a proxy whose relationship to the real objective was never checked.

Every loss is a negative log-likelihood — ch03, ch06

Choose the output distribution, then take its negative log. Gaussian → MSE, Bernoulli → binary cross-entropy, categorical → cross-entropy, Laplace → MAE. "Which loss?" is always the question "which distribution?" in disguise. Modern contrastive and preference objectives sit outside this frame — a real limit of the book, not a gap in your understanding.

KL asymmetry decides your failure mode — ch03, ch19, ch20

D(p‖q) ≠ D(q‖p). Forward KL is mode-covering (blurry averages); reverse KL is mode-seeking (sharp but partial). This single fact predicts VAE blur, GAN mode collapse, and the characteristic over-confidence of mean-field variational posteriors.

Train-error-first triage — ch11, ch05

High training error → capacity or optimization is the bottleneck; more data will not help. Low training error with a large validation gap → data or regularization. This is the highest-value heuristic in the book. scripts/training_diagnostics.py runs it.

Capacity, the gap, and the U-curve's caveat — ch05, ch07

Regularization trades variance for bias. But the classical U-shaped capacity curve is incomplete: past the interpolation threshold, test error can fall again (double descent, 2019–2020, post-dating the book). Practical consequence: when a large model overfits, try more data, more regularization or longer training before shrinking it.

Architecture is a prior, not a trick — ch09, ch10, ch15

Convolution asserts translation equivariance and locality. Recurrence asserts that the past compresses into a state. A distributed representation asserts that factors combine combinatorially. When the assertion is false, the architecture cannot be rescued by tuning — and when it is true, it beats capacity. This is also why Vision Transformers need more data than ConvNets: they discard the prior and buy it back with examples.

Depth's real cost is gradient flow and activation memory — ch06, ch08, ch10

Backprop is the chain rule scheduled well: one forward-pass-equivalent of compute, and memory proportional to stored activations. Depth fails through vanishing/exploding gradients and ill-conditioning, which is why residual connections, normalization and clipping exist.

The partition function organizes Part III — ch16, ch17, ch18, ch19

For undirected models, the likelihood gradient needs samples from the model itself. Four escape routes: sample it (CD/PCD), sidestep it algebraically (pseudolikelihood, score matching), learn around it (NCE), or estimate it for evaluation (AIS). Score matching's descendants are today's diffusion models — which is why Part III repays reading even though its models did not survive.

Diagnose before you redesign — ch04, ch08, ch11

Gradient norm exploding → clip. Norm large but loss flat → ill-conditioning. Norm near zero with high loss → saturation or dead units. NaN → numerics first. Change one thing per experiment.


Chapter Index

# Title Key content
ch01 Introduction representation learning, depth as composition, curse of dimensionality
ch02 Linear Algebra norms, SVD, eigendecomposition, conditioning, PCA
ch03 Probability & Information Theory distributions, entropy, KL, cross-entropy
ch04 Numerical Computation under/overflow, conditioning, gradient descent, KKT
ch05 Machine Learning Basics capacity, bias–variance, No Free Lunch, MLE, manifolds
ch06 Deep Feedforward Networks output/hidden units, universal approximation, backprop
ch07 Regularization norm penalties, augmentation, early stopping, dropout
ch08 Optimization SGD, momentum, init, Adam, batch norm, saddles
ch09 Convolutional Networks sparse interactions, sharing, equivariance, pooling
ch10 Sequence Modeling BPTT, vanishing gradients, LSTM/GRU, attention
ch11 Practical Methodology metrics, baselines, the data-vs-capacity rule, debugging
ch12 Applications scaling, compression, vision, speech, NLP (dated)
ch13 Linear Factor Models PPCA, factor analysis, ICA, sparse coding
ch14 Autoencoders undercomplete, sparse, denoising, contractive
ch15 Representation Learning transfer, distributed codes, disentanglement
ch16 Structured Probabilistic Models directed/undirected, energy-based, d-separation
ch17 Monte Carlo Methods importance sampling, MCMC, Gibbs, mixing
ch18 Confronting the Partition Function CD/PCD, pseudolikelihood, score matching, NCE, AIS
ch19 Approximate Inference ELBO, EM, mean field, amortization
ch20 Deep Generative Models Boltzmann machines, VAE, GAN, autoregressive

Topic Index

  • Activation functions, ReLU, GELU → ch06
  • Adam, AdamW, adaptive optimizers → ch08, ch07
  • Attention, transformers → ch10, ch12
  • Autoencoders, denoising, sparse → ch14, ch13
  • Backpropagation, autodiff → ch06
  • Batch / layer normalization → ch08
  • Bias–variance, double descent → ch05
  • Convolution, pooling, receptive field → ch09
  • Cross-entropy, KL divergence, entropy → ch03
  • Diffusion, score matching → ch18, ch14, ch20
  • Dropout, weight decay, early stopping → ch07
  • ELBO, variational inference, EM → ch19
  • Energy-based models, graphical models → ch16
  • GANs, VAEs, generative taxonomy → ch20
  • Gradient clipping, exploding/vanishing → ch10, ch08
  • Hyperparameter search → ch11
  • Initialization → ch08
  • LSTM, GRU, BPTT, teacher forcing → ch10
  • Maximum likelihood, MAP → ch05, ch03
  • MCMC, Gibbs, importance sampling → ch17
  • Numerical stability, softmax, log-space → ch04
  • Partition function, CD, PCD, NCE → ch18, ch16
  • PCA, ICA, factor analysis → ch13, ch02
  • Representation learning, transfer, probes → ch15, ch01
  • Saddle points, ill-conditioning → ch08, ch04
  • SVD, eigendecomposition, condition number → ch02
  • Training diagnostics, metric choice → ch11
  • Universal approximation → ch06

Supporting Files

Tools

S=engineering/deep-learning-book/skills/deep-learning-book/scripts
python3 $S/reading_path_planner.py --goal "train a transformer" --background applied --hours-per-week 5
python3 $S/training_diagnostics.py --train-loss 0.02 --val-loss 1.9 --grad-norm 0.4 --epochs 30
python3 $S/capacity_planner.py --params 12000000 --train-examples 50000 --train-error 0.01 --val-error 0.22
python3 $S/model_arithmetic.py --spec-sample

Every tool supports --help, --sample and --output json, uses the standard library only, and returns typed exit codes.


Scope & Limits

This companion covers the 2016 edition's 20 chapters and the delta between them and 2026 practice. It does not cover: reinforcement learning beyond passing mention, LLM training infrastructure, RLHF/DPO alignment, agentic systems, MLOps tooling, or fairness and safety evaluation — none of which the book treats. For production ML engineering use engineering-team/senior-ml-engineer; for LLM cost work use engineering/llm-cost-optimizer.

When a question lands outside the book, say the book does not cover it and cite the delta reference for what replaced its position. A companion that quietly extrapolates is worse than one that names its boundary.

信息
Category 人工智能
Name deep-learning-book
版本 v20260826
大小 85.13KB
更新时间 2026-09-06
语言