ae.

Chapter 1 · Introduction · 3 parts · 19 scenes

3,000 wages, 1,250 days, 64 cells.

Statistical learning is a set of tools for understanding data. This chapter meets three real data sets, and each one asks a different kind of question: predict a number, predict a category, or find groups that nobody labelled.

The short version

  1. Data is a table: one row per thing measured, one column per measurement. Often one column is the output we want to predict, and the rest are inputs.
  2. Wage: predicting a number is regression. Age, education and year all go with wage, but men of the same age and education still differ by tens of thousands of dollars.
  3. Stock market: predicting Up or Down is classification. The last few days barely help. A model gets 59.9% of 2005 right, and always guessing Up already gets 56.0%.
  4. Gene expression: with no output, we look for structure. Squashing 6,830 numbers per cell line down to 2 shows groups that partly match cancer types the method never saw.
  5. Predicting an output is supervised learning. Finding structure without one is unsupervised.

One idea per scene. Each animation plays when it scrolls into view; Replay runs it again. Every number is computed from the real data sets that come with the book An Introduction to Statistical Learning. Each part ends with Python for you to write.

Part 1 · scenes 1–8

Predicting a number

The Wage data: 3,000 men from the central Atlantic region of the United States, surveyed between 2003 and 2009. How well can a few facts about a man predict what he earns?

1 · Predicting a number

Data is a table.

The Wage data has 3,000 rows, one per man, and 11 columns: year, age, education, marital status, job type and more. One column, wage, is the one we want to predict.

3,000rows: one per man
11columns: one per variable

Inputs and output. The columns we predict from are the inputs, also called predictors or features. The column we predict is the output, also called the response. Later chapters write the number of rows as n and the number of inputs as p.

Wages are in thousands of dollars a year: 75.0 means $75,000.

Questions are shared with other readers: just your state or city, never your name.

2 · Predicting a number

Every man is a dot.

Pick two columns and draw each row as a dot: age across, wage up. Three thousand rows become three thousand dots.

18–80ages
$20k–$318kwages

A scatter plot shows how two variables move together. Two things show at once: young men tend to earn less, and at any one age the dots spread a long way up and down.

Questions are shared with other readers: just your state or city, never your name.

3 · Predicting a number

The average at each age.

To see the trend through the cloud, average the wages of the men near each age. Slide that average along the ages and it traces a line.

avg(a)=mean wage where ∣ age−a ∣≤h\htmlClass{k-avg}{\text{avg}(a)} = \text{mean wage where } \htmlClass{k-win}{|\,\text{age} - a\,| \le h}
avg(a)\text{avg}(a)
the average wage at age aa is a placeholder for any age. Plug in 40 and you get the average for men around 40. The orange line is avg(a) worked out for every a from 18 to 80.
∣ age−a ∣|\,\text{age} - a\,|
how far a man’s age is from a, ignoring which side · absolute valueThe bars drop the minus sign, so a man 3 years younger and a man 3 years older both count as 3 away. That keeps the window even: it reaches as far below a as above it.
≤h\le h
is at most h years · h is the bandwidthh sets how far the window reaches on each side, so it covers 2h + 1 ages: h = 4 means 36 to 44 around 40. It’s the slider. Bigger h puts more men in each average and smooths the line, but blurs real bends.
$118.7kaverage at 40
$96.5kaverage at 80
799men in the window at 40

Try it · avg(a)

Drag the diamond to move a, or drag either edge to change h. Set h to 0 to average a single age.

36–44ages averaged
799men in the window
$118.7kavg(40)

The line estimates the average wage at each age. It climbs steeply until about 40, stays roughly flat until about 60, then falls.

Try the slider. A narrow window follows every few men, so the line jitters. A wide one smooths away the real bend. How much to smooth is a question this book keeps coming back to.

Only 29 of the 3,000 men are over 70, so the right end of the line rests on very little.

Questions are shared with other readers: just your state or city, never your name.

4 · Predicting a number

Age alone isn't enough.

The average for men aged exactly 40 is $117.3k. But the 113 men behind that average earn anywhere from $27k to $299k.

113men aged exactly 40
$27k–$299klowest to highest
$74k–$166kmiddle 80% of them

The line predicts an average, not a man. The spread around it is everything age can't explain: education, job, region, luck, and all that isn't in the table. Age alone can't predict one man's wage with much accuracy.

Notice the band at the top. 79 men earn over $250k and sit apart from everyone else. Real data has quirks like this, and methods have to cope with them.

Questions are shared with other readers: just your state or city, never your name.

5 · Predicting a number

Wages rose with the years, slowly.

Now put year across instead of age. The average climbs from $106.2k in 2003 to $116.0k in 2009, roughly in a straight line.

$106k → $116kaverage, 2003 → 2009
+$1.35ka year, straight-line fit
$41.7ktypical distance of a wage from the average

Real, but small. The average rose about $10,000 over six years, while wages within any single year differ by tens of thousands. A trend can be clear in the averages and nearly invisible in the dots.

The typical distance from the average, $41.7k, is the standard deviation: roughly, how far a wage usually sits from the mean. The dashed line is the straight line that fits all 3,000 dots best; Chapter 3 shows how it's found.

Why not just (116.0 − 106.2) ÷ 6? That gives $1.63k a year, but it only uses the first and last years. The fitted line uses every wage in all seven years, and the middle years pull it flatter: 2005 dipped below 2004. Both numbers are honest; they answer slightly different questions.

Questions are shared with other readers: just your state or city, never your name.

6 · Predicting a number

Reading a box plot.

Education is a category, not a number: five levels, from no high-school diploma (1) to an advanced degree (5). For each group we summarise the wages with a box plot. Here is one built step by step, from the 650 men whose highest level is some college.

1 line up all 650 wages, lowest to highest 2 cut them into four equal piles of about 162 men 3 the cuts: 25% $89.2k · median $104.9k · 75% $121.4k4 box from $89.2k to $121.4k → IQR = 121.39 − 89.24 = $32.15k (unrounded)5 fences 1.5 × 32.15 = $48.2k beyond the box → $41.0k and $169.6kwhiskers stop at the last real wages inside: $41.4k and $169.5k 6 beyond the fences: 27 outliers (9 below, 18 above)
IQR=Q3−Q1fences=Q1−k⋅IQR,    Q3+k⋅IQR\htmlClass{k-iqr}{\text{IQR} = \htmlClass{k-q}{Q_3} - \htmlClass{k-q}{Q_1}} \qquad \htmlClass{k-fence}{\text{fences} = Q_1 - k \cdot \text{IQR},\;\; Q_3 + k \cdot \text{IQR}}
Q1,  Q3Q_1, \; Q_3
the 25% mark and the 75% mark · first and third quartilesA quarter of the men earn less than Q₁ ($89.2k); three quarters earn less than Q₃ ($121.4k). The median, the 50% mark, is sometimes called Q₂.
IQR\text{IQR}
the length of the box · interquartile range121.39 − 89.24 = $32.15k, using the unrounded quartiles: how spread out the middle half is. Unlike the full range, one extreme earner can’t stretch it.
Q1−k⋅IQRQ_1 - k\cdot\text{IQR}
how far below the box a whisker may reach · lower fenceWith k = 1.5: 89.2 − 48.2 = $41.0k. The upper fence is Q₃ + 1.5 × IQR = $169.6k. A whisker stops at the last real wage inside its fence, so it never ends in empty space.
k=1.5k = 1.5
the reach, in box-lengths · Tukey’s ruleA convention from the statistician John Tukey, not a law of nature. Bigger k flags fewer points as outliers; smaller k flags more. Try the slider.
$104.9kmedian
$89.2k–$121.4kthe box: the middle half
$41.4k–$169.5kwhiskers
27outliers

The box always holds the middle half. Whatever the data, about a quarter of it sits below the box, about a quarter above, and about half inside. It’s “about” because a real data set rarely splits into exactly equal whole numbers: 650 men make piles of 162 or 163. That makes box plots easy to compare side by side, as the next scene does.

Outliers aren’t errors. They’re just wages far from the middle by this rule. Of the 18 high ones here, 7 belong to the separate band of top earners above $250k from scene 4.

Questions are shared with other readers: just your state or city, never your name.

7 · Predicting a number

More education, higher wages.

Five box plots side by side, one per education level: 1 no high-school diploma, 2 high-school graduate, 3 some college, 4 college graduate, 5 advanced degree. Every step up lifts the median.

$81.3k → $141.8kmedian wage, level 1 → 5
268–971men per level

Unlike age, this climb never turns down. The median rises at every level, and the biggest jumps come at the top.

Associated, not caused. The data shows that education and wage go together. It can't show that more education raises wages: people who stay in school may differ in other ways too.

Questions are shared with other readers: just your state or city, never your name.

8 · Predicting a number

Combine inputs to predict better.

How good is a prediction? For every man, take his wage minus the prediction: the miss. Then measure the typical size of those misses with the RMSE. Guess every man at the overall average and the typical miss is $41.7k. Each extra input shrinks it.

RMSE=1n∑i=1n( yi−y^i )2\text{RMSE} = \sqrt{\frac{1}{n}\sum_{i=1}^{n}\big(\,y_i - \htmlClass{k-yhat}{\hat y_i}\,\big)^2}
yiy_i
the wage of man number i · read “y sub i”y is the usual letter for the output. The small i is a label: y₁ is the first man’s wage, y₂ the second’s, up to y₃₀₀₀. Writing yᵢ means “any one of them”.
y^i\hat y_i
our prediction of that wage · read “y-hat sub i”A hat on a letter always means “our estimate of it”. You’ll see it all through the book: ŷ for a predicted output, and later f̂ for an estimated rule.
yi−y^iy_i - \hat y_i
how far off the prediction was for that man · the residualPositive if he earns more than predicted, negative if less. “Residual” means what’s left over after the prediction has done its job. Every man has one: 3,000 residuals.
(… )2(\dots)^2
square each residual · squared errorSquaring makes every miss positive, so a man $20k above can’t cancel a man $20k below. It also weighs big misses more: a $40k miss counts four times as much as a $20k one.
1n∑i=1n\frac{1}{n}\sum_{i=1}^{n}
add them up for all n men, then divide by n · Σ is “sum”; the whole thing is a meanΣ (capital sigma) means “add up”, here for i from 1 to n, so every man’s squared miss goes in once. Dividing by n = 3,000 turns the total into an average: the mean squared error, MSE.
  ⋅  \sqrt{\;\cdot\;}
take the square root · root mean square error, RMSESquaring turned dollars into dollars², which mean nothing to anyone. The square root brings the answer back to dollars, so it reads as “a typical miss is about $35k”.
$41.7k → $35.0ktypical miss: one average → age band and education
16%smaller miss

Try it · RMSE

Twelve real men from the data, every 250th row. Drag the orange line: one guess for all of them. Sticks are the residuals, boxes are their squares. Which guess makes the RMSE smallest?

$46.2kRMSE: the typical miss
$27.2kthe smallest possible, at the average

Each prediction here is the average of similar men: the same age, or the same age band (18–25, 26–35, … 66–80) and education level. Age alone barely helps. Adding education helps more. Year adds little.

Most of the spread remains: wage depends on far more than this table records. Predicting a number like wage is a regression problem. Chapter 3 combines inputs with linear regression, and Chapter 7 handles the bend in age.

Why the average? Of all single guesses, the average is the one with the smallest RMSE. Try it in the playground: wherever you drag the line, the RMSE only gets worse once you leave the average.

These misses are measured on the same 3,000 men the averages were built from, which flatters the predictions a little. Chapter 2 explains why that matters.

Questions are shared with other readers: just your state or city, never your name.

Your turn · Python

The average wage at each age

The Wage data loads as Wage, a pandas table. Make by_age: the average wage for each age, as a Series indexed by age. Then print the age with the highest average, and that average.

Hint 1

pandas can split rows into groups by a column and average each group: look up DataFrame.groupby.

Hint 2

Series.idxmax() gives the index (here, the age) of the largest value.

Python loads on first run
Write yours first. The reference stays hidden until you ask.

Part 2 · scenes 9–14

Predicting a category

The Smarket data: every trading day of the S&P 500 stock index from 2001 to 2005. Can the last few days predict whether the market goes up or down today?

9 · Predicting a category

A day is a percentage change.

Each row is one trading day. Today is how much the index moved that day, in percent: +0.96 means it rose 0.96%. Lag1 to Lag5 hold the moves of the five days before.

1,250trading days
±1.14%a typical day’s move (standard deviation)
−4.92% / +5.73%worst and best day

The lags are the inputs. Each row carries the last five days with it, so we can ask whether the recent past says anything about today.

Orange bars are days the index rose; blue bars, days it fell. Calm stretches and stormy stretches come in runs: 2001 and 2002 swing much harder than 2004.

Questions are shared with other readers: just your state or city, never your name.

10 · Predicting a category

Up or Down: a category.

We won't predict the size of the move, only its direction. The output, Direction, has just two values: Up or Down. Predicting a category like this is called classification.

648Up days
602Down days
51.8%of days were Up

Regression predicts a number: a wage, a price, a temperature. Classification predicts a label: Up or Down, spam or not, which of several diseases. A category can only be right or wrong, so the methods and the ways of scoring them differ.

Questions are shared with other readers: just your state or city, never your name.

11 · Predicting a category

Does yesterday tell you about today?

Split the days by what happened today, then look at what happened the day before. If yesterday helped, the two groups would sit in different places. Use the slider to look further back.

−0.05%median, Up days
+0.10%median, Down days

The two boxes almost coincide, however far back you look. Their medians never differ by more than 0.15 percentage points, tiny next to a typical day's move of 1.14%. The past few days say almost nothing about today.

Questions are shared with other readers: just your state or city, never your name.

12 · Predicting a category

Simple rules barely beat a coin.

Score three rules on all 1,250 days. Same as yesterday: if yesterday rose, guess Up. Opposite of yesterday: guess the reverse. Always Up. Each line is a rule's running score.

46.1%same as yesterday
53.9%opposite of yesterday
51.8%always Up

Early on the scores swing wildly. After a few hundred days they settle near 50%. A rule needs many days before its score means anything.

Why no easy pattern? If a simple rule worked well, traders would use it, and their buying and selling would erase the pattern. The opposite rule's 53.9% hints at a weak tendency to reverse, but weak is the word.

Questions are shared with other readers: just your state or city, never your name.

13 · Predicting a category

A model can find a weak signal.

A method from Chapter 4, quadratic discriminant analysis, learns from 2001–2004 using the last two days' moves. Then it predicts every day of 2005, a year it never saw.

59.9%of 2005 called right (151 of 252 days)
56.0%always Up, in 2005
+3.9points: the model’s real edge

Train on the past, test on the future. Scoring a model on the days it learned from would flatter it. Scoring it on days it never saw is the honest test.

Always compare with the simplest strategy. The market rose on 56% of 2005's days, so “always Up” already scores 56%. Next to that, 59.9% is a real but modest gain.

Questions are shared with other readers: just your state or city, never your name.

14 · Predicting a category

How sure is the model?

For each day of 2005 the model also gives a probability that the market falls. On days the market really fell, that probability averaged 49.2%. On days it rose, 48.9%.

49.2%average on days it fell
48.9%average on days it rose
45–52%every prediction

Every prediction sits close to a coin flip. The model leans the right way a little more often than not, which adds up to 60% right, yet it is never confident about any single day.

The dots in the middle of each box mark the group's average. The two clouds overlap almost entirely.

Questions are shared with other readers: just your state or city, never your name.

Your turn · Python

How often is “same as yesterday” right?

Load Smarket. The rule guesses Up when Lag1, yesterday's move, is above 0. Compute same: the fraction of days the rule's guess matches Direction.

Hint 1

A comparison on a column gives a column of True/False: Smarket['Lag1'] > 0.

Hint 2

Comparing two True/False columns with == marks the days they agree. The mean of True/False values is the fraction that are True.

Python loads on first run
Write yours first. The reference stays hidden until you ask.

Part 3 · scenes 15–19

Finding groups

The NCI60 data: 64 cancer cell lines, each measured on 6,830 genes. This time there's no column to predict. The question is which cell lines are alike.

15 · Finding groups

A wide table with no answer column.

Each row is a cell line: cancer cells grown in a lab. Each column measures how active one gene is in those cells. The table is 64 rows by 6,830 columns, and none of them is an output.

64rows: cell lines
6,830columns: genes
0output columns

This is unsupervised learning: there are inputs but no output to guide (supervise) the learning. Nothing can be predicted right or wrong, so instead we look for structure: which rows resemble each other.

Colours: orange means a gene is more active than usual for that gene, blue less. Some rows share a pattern even in this sliver of the table.

Questions are shared with other readers: just your state or city, never your name.

16 · Finding groups

Squash 6,830 numbers into 2.

Nobody can draw 6,830 dimensions. Instead, find the direction along which the cell lines differ the most, and measure each one along it. Here is the idea with just two genes: turn a line until the dots spread out along it as much as possible, then drop each dot onto it.

6,830 → 2numbers per cell line
18%of the variation kept (Z1 11.4% + Z2 6.8%)

This is principal components analysis (Chapter 12). With all 6,830 genes it does the same thing: Z1 is the direction of most spread, and Z2 the direction of most spread at right angles to Z1. Each cell line becomes two numbers: where it sits along each.

Squashing loses information. These two numbers keep 18% of the variation among the 64 cell lines. In exchange, we get a picture.

Questions are shared with other readers: just your state or city, never your name.

17 · Finding groups

Groups appear.

On the map, the cell lines bunch together. A method called k-means finds the groups: place 4 centres, give each cell line to its nearest centre, move each centre to the middle of its cell lines, and repeat until nothing changes.

4rounds until nothing moves
24 · 18 · 13 · 9cell lines per group

We chose four groups because that's what the map suggests; k-means has to be told how many. Choosing that number well is a hard problem in general (Chapter 12).

The groups come from Z1 and Z2 alone. Nothing about cancer types went in.

Questions are shared with other readers: just your state or city, never your name.

18 · Finding groups

Checking against what we hid.

Each cell line comes from one of 14 cancer types, a label we never used. Light up a few types: they tend to sit together. Then join each cell line to its nearest neighbour on the map. Orange lines join two cell lines of the same type.

28 of 64nearest neighbours share a type
9%expected if neighbours were picked at random
35 of 64using all 6,830 genes

The map partly recovers the cancer types without ever seeing them: 44% of nearest neighbours match, against 9% by chance. That's evidence the structure is real, though not proof that four groups is the right number.

Squashing cost something. With all 6,830 genes, 35 of 64 nearest neighbours match. Two numbers keep most of that, but not all.

Questions are shared with other readers: just your state or city, never your name.

19 · All three

Three kinds of question.

Every data set in this chapter sits on one branch of this tree, and so will every method in the rest of the book.

Supervised learning has an output to predict: regression when it's a number, classification when it's a category. Unsupervised learning has no output; clustering is one way to find structure, and Chapter 12 shows others.

Questions are shared with other readers: just your state or city, never your name.

Your turn · Python

Who sits next to whom?

NCI60_pcs holds each cell line's Z1, Z2 and cancer type (label). For every cell line, find its nearest other cell line on the map, then count how many share their neighbour's label.

Hint 1

Put the coordinates in a 64×2 array P. P[:, None, :] - P[None, :, :] is every pair's difference, shape 64×64×2.

Hint 2

Square, sum over the last axis, and you have every squared distance. Set the diagonal to infinity with np.fill_diagonal so no cell line picks itself, then argmin(axis=1).

Python loads on first run
Write yours first. The reference stays hidden until you ask.

What to remember

  1. 1 Data is a table: rows are observations, columns are variables. Inputs are what we know; the output is what we predict.
  2. 2 Regression predicts a number. The average of similar cases is a first prediction; the spread around it is what's left unexplained.
  3. 3 Classification predicts a category. Score it on data it never saw, and against the simplest rule.
  4. 4 Unsupervised learning has no output. Two-number summaries and clustering reveal structure, and labels held back can check it.

From Chapter 1 of An Introduction to Statistical Learning, with Applications in Python by James, Witten, Hastie, Tibshirani & Taylor (Springer, 2023), pages 1–5. Data from the book's ISLP package. The scenes, numbers and code were made for this page from that data.

In the book · Chapter 1 · pp. 1–14

Where this comes from

  1. An Overview of Statistical Learningp. 1
    • Wage Datap. 1
    • Stock Market Datap. 2
    • Gene Expression Datap. 4
  2. A Brief History of Statistical Learningp. 5
  3. This Bookp. 6
  4. Who Should Read This Book?p. 8
  5. Notation and Simple Matrix Algebrap. 8
  6. Organization of This Bookp. 11
  7. Data Sets Used in Labs and Exercisesp. 12
  8. Book Websitep. 13