One summary and study guide for each of the 15 chapters of the statistics textbook. Read the guide, check yourself against the list, then practice in the Stats Lab.
Statistics is a set of tools for organizing, summarizing, and interpreting data so a pattern in a messy set of observations becomes visible. The chapter builds the basic vocabulary every later chapter uses: populations and samples, variables and how they are measured, and the summation notation that formulas are written in.
Key ideas
A population is the entire set of individuals a study is about; a sample is the smaller group actually measured.
Parameters describe populations and use Greek letters (μ, σ). Statistics describe samples and use Roman letters (M, s).
Descriptive statistics summarize data. Inferential statistics use a sample to draw conclusions about a population, always with some uncertainty.
Variables are measured on four scales: nominal (categories), ordinal (ranks), interval (equal intervals, no true zero), and ratio (equal intervals with a true zero).
Correlational studies observe variables as they are. Experiments manipulate one variable and control others, which is what allows cause-and-effect claims.
Discrete variables have separate, indivisible values. Continuous variables can fall anywhere on a continuum and are described by real limits.
Check yourself
Tell a parameter from a statistic in a research description.
Name the measurement scale for any variable you meet.
Compute ΣX, (ΣX)², and ΣX² for a short list of scores and explain why they differ.
Before any calculation, data gets organized. Frequency distributions show how often each score occurs, in tables and graphs, so the shape of a data set can be seen at a glance. The chapter also covers percentiles, which locate a score by its position in the distribution.
Key ideas
A frequency distribution lists each score value (X) beside how many times it occurred (f). Proportions are p = f/n and percentages are p × 100.
Grouped distributions combine scores into class intervals when the range is wide; the apparent limits of an interval differ from its real limits by half a unit.
Bar graphs are used for nominal and ordinal data (spaces between bars). Histograms and polygons are used for interval and ratio data.
Distribution shape matters: symmetrical distributions mirror left and right; skewed distributions trail off to one side. A tail pointing right is positive skew, left is negative skew.
A percentile rank is the percentage of scores at or below a given value. A percentile is the score that has a given percentile rank.
When an exact score is missing from the table, linear interpolation between the two nearest known values estimates its percentile rank.
Formulas to know
p = f / n
Percentage = p × 100
Check yourself
Build a frequency table from raw scores and check that Σf = n.
Sketch the distribution and name its shape.
Find a percentile rank from cumulative frequencies, interpolating when needed.
Central tendency is a single score that represents a whole distribution. The chapter covers the three measures, mean, median, and mode, how each is found, and how the shape of the distribution decides which one honestly represents the data.
Key ideas
The mean is the balance point of the distribution: the sum of the scores divided by the number of scores. Deviations from the mean always sum to zero.
The median is the midpoint: half the scores fall at or below it. It is found by locating the middle position in an ordered list.
The mode is the most frequent score. A distribution can have one mode, two modes, or none.
Adding or subtracting a constant to every score shifts the mean and median by that constant. Multiplying or dividing shifts them by the same factor. The mode follows the scores.
In a symmetrical distribution, mean, median, and mode coincide. In a skewed distribution the mean is pulled toward the tail, so the median usually represents the data better.
Extreme scores (outliers) drag the mean but leave the median and mode mostly untouched.
Formulas to know
Population mean: μ = ΣX / N
Sample mean: M = ΣX / n
Weighted mean: M = (ΣX₁ + ΣX₂) / (n₁ + n₂)
Check yourself
Compute mean, median, and mode for the same set and say which represents it best.
Predict how each measure changes when a constant is added to every score.
Read a skewed distribution and explain why the mean sits closer to the tail.
Two distributions can share the same mean and still be completely different. Variability measures how spread out the scores are. The workhorse measures are the sum of squares, variance, and standard deviation, and this chapter's definitions return in nearly every later formula.
Key ideas
The range (highest minus lowest) is the simplest spread measure but depends on only two scores and is easily distorted by an outlier.
The interquartile range uses the middle 50% of scores (Q3 − Q1) and resists outliers, at the cost of ignoring the rest of the data.
A deviation score is X − μ: how far, and in which direction, a score sits from the mean. Deviations always sum to zero, so they are squared before averaging.
The sum of squares (SS) is the total of the squared deviations. Variance is the mean of the squared deviations, and the standard deviation is its square root, which returns the measure to the original units.
For samples, SS is divided by n − 1 (degrees of freedom) instead of n. A sample tends to underestimate population variability, and the smaller denominator corrects that bias.
Adding a constant to every score leaves variance and standard deviation unchanged. Multiplying by a constant multiplies the standard deviation by the same amount (and variance by its square).
Formulas to know
SS = Σ(X − μ)² = ΣX² − (ΣX)² / N
Population variance: σ² = SS / N
Population SD: σ = √(SS / N)
Sample variance: s² = SS / (n − 1)
Sample SD: s = √(SS / (n − 1))
Check yourself
Compute SS both ways (deviation formula and computational formula) and confirm they match.
Explain, in plain words, why the sample formula divides by n − 1.
State what happens to the standard deviation when every score is doubled or raised by 5.
Chapter 5 · z-Scores and Standardized Distributions
A raw score means little without context. A z-score states exactly where a score sits in its distribution, how many standard deviations it is from the mean, and in which direction. Standardizing scores also puts different distributions on a common scale so they can be compared.
Key ideas
The sign of a z-score gives direction (above or below the mean). The number gives distance in standard deviation units.
z = 0 is the mean itself. z = +1.00 is one standard deviation above the mean; z = −2.00 is two below.
Transforming a whole distribution to z-scores creates a standardized distribution with mean 0 and standard deviation 1. The shape does not change.
Any distribution can be standardized to a new mean and standard deviation (like IQ or SAT scales) by converting through z.
Because z-scores carry location information, they let you compare scores from different distributions, such as a test score against a class average.
This chapter is the bridge to inference: later chapters express sample results as z-scores (or t values) to judge how unusual they are.
Formulas to know
z = (X − μ) / σ
X = μ + zσ
Check yourself
Convert raw scores to z-scores and back without hesitation.
Compare two scores from different distributions using z and say which is relatively higher.
Describe the mean, standard deviation, and shape of a z-score distribution.
Chapter 6 · Probability and Probability Distributions
Probability connects samples to populations: it is the proportion of times an outcome is expected in the long run. For normal distributions, probability questions become questions about proportions of the distribution, answered with the unit normal table.
Key ideas
Probability is a proportion: the number of outcomes classified as A divided by the total number of possible outcomes. Values run from 0 (impossible) to 1 (certain).
Probability statements about scores are written as proportions of a distribution, such as p(X > 90) = 0.1587.
In a normal distribution, each z-score has fixed proportions in the body (between the mean and z) and the tail (beyond z). The unit normal table lists them.
To find the probability of a score, convert it to z, then read the matching proportion from the table, adding or subtracting sections as the question requires.
The process reverses: given a proportion (such as the top 10%), find z in the table, then convert z back to a raw score to get a percentile or cutoff.
The normal distribution is symmetrical, so proportions mirror: the area above z = +1 equals the area below z = −1.
Formulas to know
Probability = (outcomes classified as A) / (total outcomes)
z = (X − μ) / σ, then read proportions from the unit normal table
Check yourself
Find the proportion of scores above, below, and between two values.
Find the score that cuts off a given percentage of the distribution.
Explain why probability and proportion are the same calculation here.
Chapter 7 · Probability and Samples: Distribution of Sample Means
Research rarely rests on one person. This chapter shifts the unit of analysis from individual scores to sample means, and introduces the distribution of sample means and its standard deviation, the standard error. Almost every inference procedure after this chapter is built on standard error.
Key ideas
The distribution of sample means is the set of means from every possible random sample of a given size n from a population. It is a theoretical distribution used to judge real samples.
Its mean, the expected value of M, equals the population mean μ. On average, sample means hit the true value.
The standard error (σM) is the standard deviation of that distribution. It measures how much distance is expected, on average, between a sample mean and the population mean.
Standard error shrinks as sample size grows: σM = σ/√n. Bigger samples give means that cluster tighter around μ.
By the central limit theorem, the distribution of sample means approaches normal as n grows (roughly n ≥ 30), whatever the population shape; it is exactly normal if the population is normal.
A sample mean can be standardized with z = (M − μ)/σM, and the unit normal table then gives the probability of obtaining a mean at least that extreme.
Formulas to know
Standard error: σM = σ / √n
z for a sample mean: z = (M − μ) / σM
Check yourself
Compute standard error and describe what it means in one sentence.
Find the probability of a sample mean falling above or below a value.
Predict what happens to standard error when n quadruples.
Hypothesis testing is a formal procedure for deciding whether a sample result is strong enough to support a claim about a population. The chapter lays out the full logic using the z-test: state hypotheses, set a decision standard, compute the test statistic, and decide.
Key ideas
The null hypothesis (H₀) states that the treatment has no effect; any difference in the sample is sampling error. The alternative hypothesis (H₁) states there is an effect.
The test uses the sample to evaluate the null hypothesis, never to prove the alternative directly. The conclusion is always phrased as rejecting or failing to reject H₀.
The alpha level (α) is the probability standard set before testing, commonly .05. It defines how extreme the sample must be to count as evidence against H₀.
The critical region is the set of sample outcomes so extreme that they would occur with probability α or less if H₀ were true. If the test statistic lands there, H₀ is rejected.
A Type I error rejects a true null hypothesis (a false alarm; its probability is α). A Type II error fails to reject a false null hypothesis (a missed effect).
Statistical power is the probability of correctly detecting an effect that exists. Larger samples, larger effects, and larger alpha all raise power.
One-tailed tests put the whole critical region in one tail when the prediction is directional; two-tailed tests split it between both tails.
Formulas to know
z = (M − μ) / σM with the population from H₀
Check yourself
Walk through the four steps of a hypothesis test in order.
Explain α, Type I error, and Type II error in plain language, with an example of each.
Decide whether a described study needs a one-tailed or two-tailed test.
The z-test requires knowing the population standard deviation, which research almost never provides. The t statistic solves that by estimating standard error from the sample itself. The price of estimating is extra uncertainty, handled by the t distribution and degrees of freedom.
Key ideas
The t statistic has the same structure as z: obtained difference divided by standard error. Only the standard error changes, it is now estimated from the sample standard deviation.
Estimated standard error: sM = s/√n. Because s varies from sample to sample, t values spread wider than z values.
The t distribution is a family of distributions, one for each degrees of freedom (df = n − 1). With large df it approaches the normal distribution; with small df it has heavier tails.
Critical t values come from the t table using df and α. Small samples demand more extreme t values to reach significance.
Beyond significance, effect size matters. Cohen's d expresses the mean difference in standard deviation units; r² gives the proportion of variance accounted for by the treatment.
A confidence interval built around the sample mean gives the range of population means consistent with the data.
Formulas to know
Estimated standard error: sM = s / √n
t = (M − μ) / sM
Cohen's d = (M − μ) / s
r² = t² / (t² + df)
Check yourself
Run a complete one-sample t test: hypotheses, df, critical value, decision.
Compute Cohen's d and r² for a result and interpret both sizes.
Explain why t critical values are larger than z critical values for small samples.
Chapter 10 · The t Test for Two Independent Samples
Most experiments compare two separate groups, such as treatment versus control. The independent-measures t test evaluates the difference between two sample means, pooling the two samples' variance into a single, better estimate.
Key ideas
The design uses two separate samples; each individual belongs to exactly one condition. The statistic examines the difference (M₁ − M₂).
Under H₀ the expected difference is zero. The denominator is the standard error of the difference between two means.
Pooled variance combines the two samples' SS and df into one weighted average variance, on the assumption both populations share the same variance (homogeneity of variance).
Degrees of freedom add across samples: df = (n₁ − 1) + (n₂ − 1).
Effect size uses the pooled standard deviation for Cohen's d, and r² comes from t and df as before.
If the homogeneity assumption is badly violated, the pooled test can mislead; equal sample sizes make the test more robust.
Formulas to know
Pooled variance: s²p = (SS₁ + SS₂) / (df₁ + df₂)
Standard error: s(M₁−M₂) = √( s²p/n₁ + s²p/n₂ )
t = ( (M₁ − M₂) − (μ₁ − μ₂) ) / s(M₁−M₂)
Check yourself
Compute pooled variance and the standard error for a two-group data set.
Complete an independent-measures t test and report effect size.
State the homogeneity of variance assumption and what it protects.
In a repeated-measures design the same individuals are measured twice (before/after), or pairs are matched. The analysis works on difference scores, which removes the variability caused by stable individual differences, often making the test more powerful.
Key ideas
A difference score D = X₂ − X₁ is computed for each person. The whole test then runs on the single set of D scores.
The null hypothesis states the population mean difference is zero (μD = 0). The statistic is t = (MD − μD)/sMD with df = n − 1, where n counts pairs or people.
Because each person serves as their own control, individual differences (ability, motivation, baseline level) cancel out of the comparison.
Removing individual-difference variance shrinks the standard error, so repeated-measures tests often detect effects that independent designs miss with the same number of observations.
Risks of the design include order effects, practice, and carryover from the first measurement to the second; counterbalancing is the standard guard.
Effect size (Cohen's d using the standard deviation of the D scores, and r²) is reported the same way as the one-sample t.
Formulas to know
sMD = sD / √n
t = (MD − μD) / sMD
df = n − 1
Check yourself
Convert paired data into difference scores and run the full test.
Explain, with an example, why removing individual differences increases power.
Name two threats specific to repeated measurement and how to control them.
Comparing three or more groups with separate t tests inflates the Type I error rate with every added test. ANOVA solves this with a single test: it compares the variance between groups (where treatment effects live) against the variance within groups (pure error).
Key ideas
The F-ratio is F = MSbetween / MSwithin. If the treatment has no effect, both estimate the same population variance and F stays near 1.00.
Total variability splits exactly: SStotal = SSbetween + SSwithin, and degrees of freedom split the same way.
dfBetween = k − 1 (k groups), dfWithin = N − k, dfTotal = N − 1.
MS (mean square) is a variance: SS divided by its df. ANOVA is named for analyzing variance even though it tests means.
A significant F says at least one group mean differs; it does not say which ones. Post hoc tests follow up while controlling error.
Effect size for ANOVA is eta squared, η² = SSbetween / SStotal: the proportion of total variance explained by group membership.
Formulas to know
F = MSbetween / MSwithin
MS = SS / df
η² = SSbetween / SStotal
Check yourself
Build a full ANOVA summary table from SS and df values.
Explain in words what it means when F is close to 1 versus much larger.
Say why a significant F requires follow-up tests before conclusions about specific groups.
Real behavior usually has more than one cause. Two-factor ANOVA tests two independent variables at once, in a design where every level of one factor is combined with every level of the other, and it adds a third question: do the factors interact?
Key ideas
A main effect is the overall effect of one factor, averaging across the levels of the other (row means for one factor, column means for the other).
An interaction exists when the effect of one factor depends on the level of the other; in a graph of cell means, the lines are not parallel.
Total variability now splits four ways: SSA + SSB + SSA×B + SSwithin = SStotal, and df split to match.
Each effect gets its own F-ratio, always with MSwithin in the denominator: FA = MSA/MSwithin, FB = MSB/MSwithin, FA×B = MSA×B/MSwithin.
When a significant interaction is present, main effects can mislead, because each factor's effect differs across the other's levels. Interpret the interaction first.
The design is efficient: one data set answers three questions, and the interaction is often the most scientifically interesting result.
Correlation measures the relationship between two variables measured on the same individuals: its direction, its form, and its strength. Regression takes the next step and uses the relationship to predict one variable from the other with a straight line.
Key ideas
The Pearson correlation describes a linear relationship between two interval or ratio variables. Its sign gives direction; its size gives strength, from −1.00 (perfect negative) to +1.00 (perfect positive), with 0 meaning no linear relationship.
SP, the sum of products of deviations, is the correlation's version of SS: it captures whether the two variables move together or in opposition.
The coefficient of determination, r², is the proportion of variance in one variable predictable from the other, the standard way to state correlation effect size.
Correlation is not causation: a third variable, restriction of range, or outliers can create or distort an apparent relationship.
When the data are ordinal, or the relationship is monotonic but not linear, the Spearman correlation (Pearson computed on ranks) is used. Point-biserial and phi handle dichotomous variables.
Regression fits Ŷ = bX + a, the line minimizing squared prediction errors. The slope b = SP/SSx; the intercept is a = MY − b·MX. The standard error of estimate measures typical prediction error.
Formulas to know
SP = Σ(X − MX)(Y − MY) = ΣXY − (ΣX)(ΣY)/n
Pearson r = SP / √(SSx · SSy)
Slope: b = SP / SSx = r · (sY / sX)
Intercept: a = MY − b·MX
Regression line: Ŷ = bX + a
Check yourself
Compute SP and r for a small data set and interpret sign, strength, and r².
Build the regression equation and predict Ŷ for a given X.
Give one example each of a third-variable problem and an outlier distorting r.
Chapter 15 · Chi-Square Tests: Goodness of Fit and Independence
Chi-square tests work with frequencies, counts of individuals in categories, rather than scores. The goodness-of-fit test asks whether one categorical variable's distribution matches a claimed pattern; the test for independence asks whether two categorical variables are related.
Key ideas
Observed frequencies (fo) are the counts actually in the sample. Expected frequencies (fe) are the counts predicted by the null hypothesis.
Goodness of fit: H₀ specifies the population proportions (often no preference, equal proportions). fe = n × pn for each category, and df = C − 1.
Independence: H₀ says the two variables are independent, so each cell's expected frequency is fe = (row total × column total) / n, with df = (R − 1)(C − 1).
The chi-square statistic sums the squared discrepancies scaled by expectation: χ² = Σ (fo − fe)² / fe. Large discrepancies produce large χ² and rejection of H₀.
Effect size: for independence with larger tables, Cramér's V adjusts φ to the table size; a 2 × 2 table uses the phi-coefficient.
Chi-square needs adequate expected frequencies (fe should not be tiny, conventionally at least 5 in most cells), and observations must be independent, each individual counted once.
Formulas to know
χ² = Σ (fo − fe)² / fe
Goodness of fit: fe = n · pn, df = C − 1
Independence: fe = (row total × column total) / n, df = (R − 1)(C − 1)
Check yourself
Compute expected frequencies and χ² for a goodness-of-fit problem.
Build fe for a two-way table and state df before computing.
Explain why chi-square compares counts rather than means.