XGBoost Tutorial
Contents
An interactive course in one file

XGBoost Tutorial

XGBoost derived from first principles — the loss, the gradient, the Hessian, the gain, the whole machine — taught through a single meteorological question: how cold will tonight get?

Every chapter pairs a derivation with an instrument you can turn over in your hands. Nothing is asserted that you cannot poke. By the end you will be able to write down XGBoost's objective from memory, explain what every hyperparameter does to that objective, and predict — before touching a slider — what changing it will do to a model.

The observing programme below is sequential: each chapter uses the machinery of the last. The glyphs are live — they are drawn by the same code you'll meet inside.

Chapter I · Foundations

The Loss Function

Everything in XGBoost — every split, every leaf value, every hyperparameter — is in service of one number. Before we grow a single tree, we must decide precisely what it means to be wrong.

1.1 — A forecast and its price

Our running problem, start to finish: at 18:00 UTC an observer at a small atmospheric observatory records the evening conditions — cloud cover, wind speed, dewpoint depression, time of year — and must forecast the overnight temperature drop: how far the thermometer will fall between the evening reading and dawn minimum. Clear, calm, dry winter nights radiate ferociously and can drop 12 °C or more; overcast, windy summer nights barely drop 2 °C. The physics is real, the interactions are nonlinear, and the record is imperfect. It is a perfect job for boosted trees.

A model produces a prediction ŷ for a true outcome y. A loss function l(y, ŷ) converts the pair into a price: zero when the forecast is perfect, growing as it degrades. The choice of loss is a modelling decision, not a technicality — it declares which mistakes you care about. Forecast a 10 °C drop when the true drop is 4 °C and the road-gritting lorries roll out for nothing; forecast 4 °C when the truth is 10 °C and the roads ice over. Squared error says these two failures cost the same; a real forecaster might disagree, and later we will see that XGBoost lets you say so.

1.2 — Two derivatives are all you get

Here is the single most important design decision in XGBoost, stated up front so you can watch everything else flow from it. The algorithm never looks at your loss function directly. For each data point it asks the loss for exactly two numbers, evaluated at the current prediction:

EQ 1.1 gi = ∂ l(yi, ŷi)∂ ŷi hi = ∂ 2 l(yi, ŷi)∂ ŷi 2 gᵢ — the gradient: the slope of the loss at the current prediction. Its sign says which way the prediction should move; its size says how urgently. hᵢ — the Hessian (here just a second derivative, since the prediction for one point is one number): the curvature. It says how quickly the urgency changes as the prediction moves — how far you can trust the slope.

Everything downstream — leaf values, split gains, min_child_weight, why logistic regression "just works" in the same code — is arithmetic on g and h. Swap the loss, and only these two formulas change; the entire tree-growing machine is untouched. This is why XGBoost ships regression, classification, ranking and survival objectives in one engine, and why you can hand it any twice-differentiable loss of your own invention.

1.3 — The gallery of losses

Meet the candidates. For regression on our temperature drop, the default is squared error; its gradient is the (signed) error itself and its Hessian is the constant 1 — a fact that will quietly simplify half the formulas in this course, and quietly hide the other half's purpose. Absolute error is more robust to freak nights but has a broken second derivative. Pseudo-Huber interpolates between them. And log-loss, for classifying will there be a ground frost?, has the most instructive Hessian of all.

Lab 1 · The Loss Benchone observation, every loss — watch g and h, not just l
loss l(ŷ)gradient gHessian hyour prediction
Turn the dials
  1. Squared: drag ŷ. The gradient line is straight (g = ŷ − y) and the Hessian is flat at 1. Curvature never changes — the loss trusts big corrective steps.
  2. Absolute: the gradient is ±1 no matter how wrong you are — a 12 °C error shouts no louder than a 0.5 °C one — and h = 0 everywhere. Hold that thought: in Chapter III a zero Hessian will try to divide by zero, and you'll see exactly why libraries fake h = 1 here.
  3. Pseudo-Huber: quadratic near the truth, linear far away. Watch h fade towards zero in the tails — distant outliers get gradient but lose authority.
  4. Log-loss: now ŷ is a raw score pushed through a sigmoid, p = σ(ŷ). The gradient is the beautifully simple p − y, and h = p(1−p) peaks at maximum uncertainty (p = ½) and vanishes when the model is confident. The Hessian is literally a confidence meter. Chapter VII builds on this.
Why the second derivative at all? Gradient descent uses slope alone: step size is guesswork. Newton's method divides slope by curvature, step = −g/h, and for a quadratic bowl lands at the bottom in one hop. XGBoost is, at heart, Newton's method performed by trees — Chapter III makes this exact. That is the honest answer to "what does the X in eXtreme actually buy you": second-order information, plus regularisation built into the objective rather than bolted on.
Chapter II · Foundations

Boosting: a Committee of Corrections

One tree is a crude forecaster. Boosting builds a committee in which each new member is hired for one job only: to correct the standing error of everyone hired before it.

2.1 — The additive model

The final model is a sum of K small trees, each mapping the evening observations x to a number:

EQ 2.1 ŷi = Σk=1…K fk(xi) Not a vote, not an average — a running total. Each tree contributes a (possibly negative) number of degrees to the forecast. A single tree partitions the weather into a few crude regimes; a sum of trees layers hundreds of partitions into a fine-grained response surface no single tree could express.

Fitting all K trees jointly is hopeless — the space of tree ensembles is combinatorial. So boosting fits them stagewise: freeze everything built so far, and ask only what one more tree should be.

EQ 2.2 ŷi(t) = ŷi(t−1) + η ft(xi) The update rule of the whole enterprise. η is the learning rate (eta, "shrinkage"): each new tree's contribution is deliberately damped, so no single hire dominates the committee. Why that is wise becomes visceral in the lab below.

2.2 — What should the new tree fit?

Take the simplest case: squared loss, l = ½(y − ŷ)². The gradient at the current prediction is gi = ŷi − yi — the negative of the residual. So "fit the new tree to the residuals," the folk description of boosting, is secretly "fit the new tree to the negative gradient of the loss." That reframing is the whole trick: gradients exist for any loss, residuals only for squared error. Boosting is gradient descent in the space of functions — each tree is one downhill step, and η is the step size.

Watch it happen. Below is a slice of our problem with everything held fixed except wind: clear-sky nights, drop against wind speed. The physics gives a nonlinear curve — calm nights decouple the surface air and cool hard; even a few m/s of wind stirs warmer air down and kills the drop. We'll fit it with stumps (depth-1 trees, one split each), the weakest learner there is.

Lab 2 · The Corrections Committeestumps fitting residuals, one hire at a time
0 trees
observed nightscurrent model F(x)
residuals y − F(x) — the next tree's targetthe correction η·ft just added
Turn the dials
  1. Set η = 1 and add trees one at a time. Fast — and jagged. Each stump commits fully to its correction, and later trees spend their lives correcting earlier trees' overcommitments. Watch the residual plot: it doesn't shrink smoothly, it sloshes.
  2. Reset. Set η = 0.05 and add 50. Slower, but the fit is calmer and the residuals decay like a discharging capacitor. Shrinkage leaves headroom — room for future trees to disagree.
  3. Notice the model is a staircase. A sum of stumps is piecewise constant; more trees means finer stairs. Depth will buy interactions in Chapter VI.
  4. Keep adding trees at high η and watch the model start chasing individual noisy points. That is overfitting happening live — the committee has begun memorising the minutes.
Where this chapter leaves us We know the shape of the algorithm: start somewhere, repeatedly add a damped tree fitted to the current gradients. Two questions remain, and they are the heart of XGBoost. Given a tree's structure, what value should each leaf output? And how do we choose the structure — the splits — in the first place? Both answers fall out of a single second-order calculation. Onwards.
Chapter III · The Core Derivation

The Newton Step: Deriving the Leaf

This is the load-bearing chapter. In roughly a page of algebra we derive the two formulas that are XGBoost: the optimal value of a leaf, and the quality score of a tree. Every hyperparameter you will ever tune appears here with a job description.

3.1 — The regularised objective

At boosting round t, we seek the tree ft minimising the total loss plus a penalty for complexity. XGBoost's signature move is putting the penalty inside the objective from the start, rather than pruning as an afterthought:

EQ 3.1 Obj(t) = Σi l(yi, ŷi(t−1) + ft(xi)) + γT + ½λ Σj wj2 The tree is described by its structure — which leaf each point lands in — and its leaf weights w1…wT, the numbers the leaves output. The penalty Ω(f) = γT + ½λΣw² charges γ (gamma) per leaf — a flat toll on structure — and λ (lambda, L2) on the squared size of every output — a tax on confidence.

3.2 — Taylor to the rescue

EQ 3.1 is unoptimisable as written: ft sits inside an arbitrary loss. So we approximate the loss around the current prediction with a second-order Taylor expansion — slope and curvature, exactly the two numbers Chapter I said we'd need:

EQ 3.2 l(yi, ŷi(t−1) + ft(xi))  ≈  l(yi, ŷi(t−1)) + gi ft(xi) + ½hi ft(xi)2 The first term is a constant — last round's loss, already paid, drop it. What remains is a quadratic in the new tree's output. We have replaced your loss, whatever it was, with a private parabola per data point: tilt gᵢ, curvature hᵢ. From here the loss function never appears again.

Now the decisive regrouping. A tree assigns every point to exactly one leaf. So instead of summing over points, sum over leaves, pooling the parabolas of the points inside each. Writing Ij for the set of points in leaf j, and defining the leaf totals Gj = Σ gi and Hj = Σ hi over i ∈ Ij:

EQ 3.3 Obj(t) ≈ Σj=1…T [ Gj wj + ½(Hj + λ)wj2 ] + γT The objective has shattered into T independent one-dimensional parabolas, one per leaf, plus the toll. Independence is everything: each leaf can be optimised alone, in closed form, by A-level calculus.

3.3 — The two formulas

Differentiate one leaf's parabola with respect to wj, set to zero:

EQ 3.4 · the leaf weight wj★ = −GjHj + λ Read it aloud: a leaf outputs the pooled gradient, divided by the pooled curvature, plus ballast. It is a Newton step (−g/h) computed jointly for all its residents, shrunk towards zero by λ. For squared loss (h=1), G is the sum of residuals and H is the head-count, so w★ = (mean residual) × n/(n+λ) — the average correction, damped more when the leaf is small. λ literally acts like λ phantom residents who all vote "do nothing".

Substitute w★ back in, and each leaf contributes −½G²/(H+λ) to the objective. The best achievable score for a given structure is therefore:

EQ 3.5 · the structure score Obj★ = −½ Σj=1…T Gj2Hj + λ + γT A single number grading any candidate tree shape — lower is better. Each leaf's term G²/(H+λ) measures how much pooled, agreeing gradient it captured: gradients of opposite sign cancel inside G and score nothing. A good leaf is a room full of points that all want to move the same way. This score is the judge that will choose our splits in Chapter IV.
Lab 3 · One Leaf, Five Residentspooled parabolas and the λ ballast — EQ 3.3 and 3.4 live

Five nights share a leaf. Each slider is that night's gradient gi (for squared loss: current prediction − truth, so negative means "the model under-forecast the drop"). Each night contributes a private parabola (thin curves); the leaf must pick one output w minimising their sum plus the λ term (bright curve).

per-point parabolas giw + ½hiw²leaf objective (sum + ½λw²)w★ = −G/(H+λ)unregularised −G/H
Turn the dials
  1. Drag all five gradients negative (model under-forecasting everywhere). The parabolas agree, the pooled bowl is deep and off-centre, and w★ is a confident positive correction.
  2. Now set them to alternate signs, roughly summing to zero. The bowl is still curved (H unchanged) but centred at zero — the leaf, hearing contradictory demands, wisely does almost nothing. This leaf needs splitting, not averaging: exactly what Chapter IV detects.
  3. Sweep λ from 0 to 20 with strong agreement among the gradients. Watch w★ slide towards zero — never past it. λ can only shrink, never flip.
  4. Now the payoff of Chapter I: imagine these were log-loss points, each with its own hi = pi(1−pi). Confident points (h≈0) would bring flat parabolas — loud gradients but no authority over where the minimum sits. The Hessian is a per-point credibility weight, and −G/(H+λ) is a credibility-weighted vote. And if every h were 0 (absolute loss), the denominator would be bare λ — division by zero at λ=0. That is why h=0 losses need patching.
min_child_weight lives here The famous min_child_weight is a floor on a leaf's H. For squared loss, h=1, so H = number of points and the name reads "minimum samples per leaf". But for log-loss, H is the sum of confidences p(1−p) — a leaf of 200 already-certain points can have less Hessian mass than a leaf of 3 uncertain ones. The parameter is really: do not create a leaf whose pooled parabola is too shallow to trust. We enforce it in the next chapter's lab.
Chapter IV · The Core Derivation

Growing the Tree: the Gain

EQ 3.5 grades a finished tree, but trees are grown one split at a time. The gain asks the only question a growing tree ever needs answered: is this room better as two rooms?

4.1 — Better as two rooms?

Take a leaf with totals G, H. A candidate split — say wind < 3.1 m/s — divides its residents into a left set (GL, HL) and a right set (GR, HR). Score both arrangements with EQ 3.5 and subtract. Structure-score improvement, minus the toll for the extra leaf:

EQ 4.1 · the gain Gain = ½[ GL2HL + λ + GR2HR + λ − (GL+GR)2HL+HR + λ ] − γ Split if Gain > 0. The bracket is pure algebraic pleasure: it is large exactly when the split separates disagreeing gradients. If left wants "predict more" and right wants "predict less", their gradients cancel in the parent's (GL+GR)² but stand tall in GL² and GR² separately. A split earns its keep by resolving an argument. And γ is not a tie-breaker but a hurdle: a split must resolve at least γ worth of argument or the leaf stays whole — pre-pruning, derived rather than decreed.

4.2 — The exact greedy scan

How does the algorithm find the best split? With no cleverness whatsoever — and that is worth seeing once in your life. For each feature: sort the leaf's points by that feature, then walk the sorted order maintaining running totals GL, HL (the right totals are just G−GL, H−HL). Every gap between consecutive values is a candidate threshold; EQ 4.1 costs a handful of arithmetic operations per candidate. Best gain across all features and thresholds wins. This is the exact greedy algorithm; Chapter VII shows the histogram trick that big data forces upon it.

The lab below is that scan, opened up. Forty-eight nights, one feature (evening cloud cover), and the gradients from a model that so far predicts only the global mean drop — so gi = ȳ − yi: clear nights (left) dropped far more than the mean predicted and carry big negative gradients; overcast nights carry positive ones. An argument waiting to be resolved.

Lab 4 · The Split Scannerdrag a threshold through real gradients; the gain curve is EQ 4.1 evaluated everywhere
g<0 — "drop was bigger than predicted"g>0 — "smaller than predicted"your threshold
gain at every candidate thresholdforbidden by min_child_weightγ hurdle
Turn the dials
  1. Drag the threshold across the plot and watch the gain curve trace out the exact greedy scan. The maximum sits where teal and rose populations separate — around 2–3 oktas, the physics of radiative cooling rediscovered by arithmetic.
  2. Raise γ until the entire gain curve sinks below the hurdle. Verdict: leaf stays whole. You have just pre-pruned a tree by hand.
  3. Raise λ. The gain curve deflates everywhere — but not uniformly: thresholds carving off small groups (small H in the denominator's company) deflate fastest. λ is quietly biased against splits justified by few points.
  4. Push min_child_weight up and watch the flanks of the curve get struck out: splits isolating a handful of extreme nights are forbidden outright, however tempting their gain. The three levers overlap in effect — all fight tiny, overconfident leaves — but by different mechanisms: λ shrinks, γ tolls, min_child_weight forbids.
Recursion, and two ways to stop Apply the scan to the root, then to each child, and so on: that recursion is tree growing. It halts when nothing clears the hurdles — or at max_depth, a blunt structural stop. Historical footnote: XGBoost grows depth-wise (level by level) by default, while LightGBM grows leaf-wise (always split the current best leaf anywhere in the tree); same gain formula, different search order, and with a leaf budget the leaf-wise order finds lower objectives but lonelier, riskier leaves.
Chapter V · The Elegant Trick

Missing Data: the Default Direction

On the coldest nights, the anemometer ices up. Our wind column is missing precisely when wind matters most. Most algorithms make you impute and pray; XGBoost turns the gap itself into signal.

5.1 — Learn where the ghosts go

Every split in an XGBoost tree carries, alongside its feature and threshold, a learned default direction. During the split scan, the points with a missing value can't be sorted with the rest — so the algorithm simply tries both dispositions. Compute EQ 4.1 with all missing points massed on the left; compute it again with them massed on the right; the split's gain is the better of the two, and the winning side is stamped onto the node. At prediction time, a night arriving with no wind reading follows the stamp.

This is the sparsity-aware split, and it costs almost nothing: the missing points' pooled totals Gmiss, Hmiss are added to one side or the other of sums we were already maintaining. Two extra additions per candidate.

The deep point: if missingness is informative — icing means cold, a sensor that saturates means extreme values, a survey question skipped means something — the default direction learns the meaning of absence. Imputing the column mean would have actively destroyed that signal, quietly filing the iciest nights of the record under "average breeze".

Lab 5 · The Iced Anemometerforty nights, nine silent instruments — where should the ghosts sit?
g<0g>0thresholdmissing wind (gutter)
Turn the dials
  1. Park the threshold near 3 m/s and flip the toggle. The missing nights are mostly hard-drop, negative-gradient nights (icing ⇒ cold, calm), so sending them left — in with the calm nights they resemble — scores far higher. The arithmetic discovers the meteorology of the gap.
  2. Press let the algorithm choose: it scans every threshold under both dispositions and reports the champion — one line of bookkeeping in the real code.
  3. A subtlety worth savouring: the default direction is chosen per node. Deeper in a tree, "wind missing" can be routed differently under different cloud regimes. Absence is allowed to mean different things in different weather.
Sparsity is the same trick XGBoost treats explicit zeros in sparse matrices and one-hot encodings the same way — the scan only ever visits the present entries, and the default direction sweeps the rest along. That's why it eats bag-of-words matrices with millions of columns for breakfast: cost scales with non-zeros, not with cells.
Chapter VI · The Full Machine

The Observatory Model

Everything assembles. A real (if synthetic) dataset, a complete XGBoost implemented in the page you are reading — exact greedy scan, Newton leaves, learned default directions — and every lever on one panel. Your job: overfit it, then rescue it.

6.1 — The data

420 observing nights, generated from a plausible boundary-layer story plus honest noise. Four features at 18:00: cloud (oktas 0–8), wind (m/s — missing on about one night in eight, preferentially the icy ones), dpd (dewpoint depression, °C — dry air radiates to space more freely), month (1–12, standing in for night length). Target: overnight drop in °C. The generator hides a strong cloud×wind interaction — clear and calm is worth far more than the sum of clear and calm — which is precisely what depth > 1 exists to find. 300 nights train the model; 120 held-out nights judge it.

6.2 — The algorithm, complete

XGBoost in nine lines F₀(x) ← base_score (mean of y)
for t = 1 … n_trees:
  gᵢ ← ∂l/∂ŷ, hᵢ ← ∂²l/∂ŷ² at ŷ = Ft−1(xᵢ) // Ch I
  draw a row subsample; draw a column subsample // Ch VII
  grow tree: recursively take the best-gain split (EQ 4.1, both missing directions) // Ch IV–V
    … while gain > 0, Hchild ≥ min_child_weight, depth < max_depth
  set each leaf to w★ = −G/(H+λ) // Ch III
  Ft(x) ← Ft−1(x) + η·tree(x) // Ch II
  stop early if validation loss hasn't improved in k rounds
Lab 6 · The Training Deska full XGBoost, running in this page — 300 training nights, 120 validation nights
untrained
train RMSEvalidation RMSEbest round
validation nights: predicted vs observed drop
feature importance — total gain claimed
decision node (∅→ marks the learned missing direction)leaf w★ > 0leaf w★ < 0
Press train.
A guided ruin and rescue
  1. Baseline. Defaults, train. Note both RMSEs and the gap between them. Inspect tree 1 (big confident structure) versus the last tree (shallow, timid leaves — later corrections are refinements of refinements).
  2. Underfit on purpose. depth 1, 30 trees, η 0.05. Train RMSE stays high and the two curves hug: the model lacks capacity — stumps cannot express cloud×wind. Bias, visualised.
  3. Overfit on purpose. depth 6, η 0.5, 300 trees, λ 0, everything else off. Train RMSE dives towards the noise floor while validation bottoms out and climbs. That climbing rose curve is the sound of a model memorising 300 particular nights.
  4. Rescue with each lever separately, from the overfit settings: first λ ≈ 10 alone; reset, then γ ≈ 5 alone; reset, then min_child_weight ≈ 15 alone; reset, then subsample 0.6 alone. Each closes the gap by a different mechanism — shrinking leaves, refusing splits, forbidding small leaves, decorrelating trees. Then combine, lower η to 0.05, and switch early stopping to 30: the brass marker plants itself at the honest optimum.
  5. Read the importance panel. Cloud and wind should dominate, month behind, dpd modest — the generator's own recipe recovered. At depth 1, watch the interaction-dependent share collapse.
The one true tuning heuristic Lower η is almost never wrong if you can afford the trees — it is the only lever that reduces overfitting and underfitting risk together, at the price of compute. The practical recipe the derivation supports: set η small (0.02–0.1), n_trees large, let early stopping choose the count, then tune depth and min_child_weight against each other, then λ/γ, then subsampling for decorrelation. You now know why each step targets what it does.
Chapter VII · Bells & Whistles

The Extended Instrument Case

The core machine is complete. What remains are the refinements that made XGBoost an industrial tool — each a small idea, and each now a one-paragraph corollary of things you already know.

7.1 — The objective zoo

Chapter I promised that changing the loss changes only g and h. Cash the promise: to turn our regression engine into a frost classifier (y ∈ {0,1}), keep every line of tree code and swap two formulas. The model's raw score z becomes a probability via the sigmoid p = 1/(1+e−z), and log-loss gives:

EQ 7.1 gi = pi − yi hi = pi(1 − pi) The gradient is again "prediction minus truth" — in probability space. The Hessian is the variance of a coin with bias p: maximal at p = ½, zero at certainty. Every formula you have learned now reweights itself: leaves average corrections weighted by uncertainty; min_child_weight demands leaves contain enough doubt; and the gain seeks splits that separate the confidently-wrong from the rest.
Lab 7a · The Confidence Meterone night in the frost classifier — g and h as the score moves

The same plug-in socket accepts multiclass softmax (one tree per class per round), Poisson counts, Tweedie for rainfall-like zero-inflated targets, quantile losses for "what drop will only 1 night in 10 exceed?", ranking losses, survival times, and any custom pair of (g, h) callbacks you write yourself. The engine never knows the difference.

7.2 — Histograms and the quantile sketch

The exact greedy scan sorts every feature in every node — brutal at a hundred million rows. The approximate algorithm replaces raw values with a few hundred bins whose edges are (weighted) quantiles of the feature, then scans bin boundaries only. Two refinements matter: the quantiles are weighted by h — bins hold equal curvature, not equal counts, so resolution concentrates where the objective still has structure — and the weighted quantile sketch computes them in one streaming pass with provable error bounds, which is a genuine contribution of the XGBoost paper. Modern hist mode adds a lovely accounting trick: a child's histogram equals its parent's minus its sibling's, so each split's second histogram is free.

Lab 7b · The Binning Deskhow much gain do you lose by not looking everywhere?

7.3 — Randomness as regulariser

subsample shows each tree only a random fraction of rows; colsample_bytree / _bylevel / _bynode hide a random fraction of features per tree, per level, or per split. Both are borrowed from bagging and random forests, and both work for the same reason: trees that see different evidence make decorrelated mistakes, and decorrelated mistakes partially cancel in the sum (EQ 2.1). Column subsampling has a second virtue: it forces the ensemble to develop backup routes around a dominant feature — insurance for the day the anemometer really does die. You already felt both in Lab 6, step 4.

7.4 — Constraints, priors, and other house rules

7.5 — Viva voce

Close the notes. If the course has worked, these should feel less like trivia than like consequences.

Why does XGBoost require a twice-differentiable loss when plain gradient boosting needs only one derivative?
Because the leaf weight is a Newton step, w★ = −G/(H+λ): the second derivative sits in the denominator of every leaf and every gain. First-order boosting fits the gradient's shape and picks step sizes by line search or decree; XGBoost's second-order expansion (EQ 3.2) makes the objective a per-leaf parabola solvable in closed form.
λ = 0, and a leaf ends up containing points whose Hessians are all ≈ 0. What happens, mechanically?
w★ = −G/(H+λ) → −G/0: the leaf weight diverges. A whisper of pooled gradient over no curvature is a division-by-nearly-zero — numerically an enormous, wildly overconfident leaf. This is why absolute loss needs a fake Hessian, why max_delta_step exists, and why min_child_weight is phrased in Hessian mass rather than sample count.
Your training data has no missing values, but production data will. Does the default direction still get learned?
Yes — every node still stores one. With nothing missing at training time, the choice comes free of evidence (implementations pick a convention — dense training data typically yields default-right), so it's arbitrary rather than informed. The model won't crash in production, but absence will be routed by convention, not by learned meaning. If production will have gaps, train with representative gaps.
Why does a split's gain never exceed the sum of its children's future gains — i.e. why can γ prune too eagerly?
Gain is greedy and local: EQ 4.1 scores one split assuming the children stay leaves. A mediocre split can unlock brilliant grandchildren (classic case: XOR-like interactions, where the first split alone separates nothing), but γ is applied to the mediocre split now. Depth-wise growth with post-pruning (grow to max_depth, then collapse negative-gain subtrees bottom-up) exists precisely to soften this.
min_child_weight = 200 barely changes your squared-loss model but devastates your log-loss model on the same features. Why?
Under squared loss h = 1, so the constraint means "≥ 200 points per leaf" — mild in a big dataset. Under log-loss h = p(1−p) ≤ ¼, so 200 of Hessian requires at least 800 points, and far more once the model grows confident (h → 0). Same number, different currency: the parameter is denominated in curvature.
Halving η while doubling n_trees usually improves validation loss. What is the honest cost, besides compute?
Almost none — that's the point — but "almost": the ensemble becomes larger to store and slower to score at prediction time, and past a point the returns vanish because shrinkage only helps until the trajectory through function space is smooth enough that further damping just retraces it. Early stopping finds that point for you.
Chapter VIII · Epilogue

The Cheat Sheet

The whole course on one plate. If you can reconstruct each line's why, the instrument is yours.

THE COURSE IN FOUR LINES
I   gi = ∂l/∂ŷ,  hi = ∂²l/∂ŷ² — the loss, reduced to slope and curvature
III wj★ = −GjHj+λ — each leaf: a pooled, ballasted Newton step
IV  Gain = ½[GL²/(HL+λ) + GR²/(HR+λ) − G²/(H+λ)] − γ — split iff it resolves an argument worth the toll
II  Ft = Ft−1 + η·ft — add it, damped; repeat; stop when validation says so

The levers, with job descriptions

leverwhere it liveswhat it really does
eta (η)EQ 2.2Damps every tree's contribution. Leaves headroom for future corrections; the one lever that fights over- and underfit together, paid in trees.
lambda (λ)EQ 3.4 denominatorBallast in every leaf. λ phantom residents voting "do nothing"; shrinks weights towards zero, hits small-H leaves hardest, deflates all gains.
gamma (γ)EQ 4.1 tailToll per new leaf. A split must resolve ≥ γ of gradient argument or the room stays whole. Pre-pruning, derived from the objective.
min_child_weightconstraint on HL, HRFloor on leaf curvature. Refuses leaves whose pooled parabola is too shallow to trust. Counts points under squared loss; counts doubt under log-loss.
max_depthgrowth recursionCap on interaction order. Depth d can express interactions among ≤ d features along a path. Depth 1 = additive model.
subsampleper-round row drawDecorrelates trees' mistakes so they cancel in the sum; also 25–50% faster per tree.
colsample_*per tree/level/node feature drawForces backup routes around dominant features; decorrelation in feature space.
alpha (α)soft-threshold on w★Minimum conviction to speak. Leaves with |G| ≤ α output exactly zero.
early stoppingthe training loopLets validation choose n_trees. The only lever measured, not guessed.
base_scoreF₀The prior. Start at the answer a model with no features would give.

The one-sentence version

XGBoost repeatedly asks every data point for its complaint (g) and its credibility (h), grows a tree whose every split separates the loudest disagreements — tolled by γ, ballasted by λ — sets each leaf to the pooled Newton step of its residents, adds the tree at a fraction η of full strength, and stops when held-out nights say the corrections have become memories.

Where next, if you want it: the 2016 paper (Chen & Guestrin, XGBoost: A Scalable Tree Boosting System) will now read as an engineering memo about ideas you own — its cache-aware block layout and out-of-core columns are Chapter IV's scan made mechanically sympathetic. LightGBM's GOSS and histogram subtraction, CatBoost's ordered boosting and target statistics, and NGBoost's probabilistic outputs are all dialects of the same grammar: g, h, gain, shrink, repeat.


Built as a single self-contained file. The dataset is synthesised in your browser from a seeded generator (a boundary-layer fable, not observatory data); the XGBoost inside is a faithful exact-greedy implementation in about two hundred lines of vanilla JS — view source, it's all here.