
How do pollsters learn what millions of people think by asking only a few hundred? How can a factory check thousands of light bulbs without testing every one? In this chapter you will learn how a random sample lets you make smart predictions about a whole group, and how to compare two groups of data using their centers and their spread.
1. Populations and Samples
A population is the entire group you want to learn about: all the students in a school, every tree in a park, every battery a factory makes. A sample is a smaller part of the population that you actually observe or measure.
Counting every member of a population is often too slow, too expensive, or impossible. Testing every battery until it dies would leave nothing to sell. So we study a sample and use it to learn about the population.
A number that describes a sample, such as the sample mean or the percent of students in the sample who ride a bus, is only an estimate of the same number for the whole population.
2. Representative and Random Samples
A sample is useful only if it is representative: it should look like a small copy of the population, with the same mix of kinds of members.
A random sample is chosen so that every member of the population has an equal chance of being picked. Drawing names from a hat, or using a random number generator on a numbered list, are fair ways to do this.
Samples that are not chosen randomly are often biased, meaning they tend to favor one kind of answer. Suppose you ask only the players leaving a soccer practice, “Do you like soccer?” You will hear a lot of yes answers, but those students do not represent the whole school. Other common problems are asking only your friends (a convenience sample) and asking people who choose to reply to an online poll (a voluntary response sample).
- Choose the sample at random from a list of the whole population.
- Use a sample that is large enough: larger random samples tend to give estimates that are closer to the truth.
- Ask every question in a neutral way.
3. Making Inferences from a Sample
An inference is a prediction about a population based on a sample. If the sample is random, the sample proportion is a reasonable estimate of the population proportion, and the sample mean is a reasonable estimate of the population mean.
- Find the sample proportion: \(\dfrac{\text{number with the trait}}{\text{sample size}}\).
- Multiply it by the size of the population.
- Say the answer is an estimate, not an exact value.
Westbrook Middle School has 800 students. A random number generator picks 50 names from the roster, and 18 of those students ride the bus. The sample proportion is \(\dfrac{18}{50}=0.36\). We estimate that \(0.36\times 800 = 288\) students in the school ride the bus.
Different random samples give different results. The dot plot below shows the percent of pet owners in ten random samples of 20 students each from the same school.
The results run from 35% to 55%, but they cluster near 45%. This is why one sample gives an estimate and not a guarantee: samples vary, and an estimate is wiser when you look at several samples or use a bigger one.
A sample that is too small, or that is not random, can lead to a wrong prediction. If 3 out of 4 students in a sample like pizza, that does not prove that 75% of the school likes pizza.
4. Measures of Center and Variability
To describe a data set, we give a number for its center and a number for its variability (how spread out it is).
- Mean: add all the values, then divide by how many there are.
- Median: the middle value when the data are in order (or the mean of the two middle values).
- Range: maximum minus minimum.
- Quartiles: \(Q_1\) is the median of the lower half and \(Q_3\) is the median of the upper half. The interquartile range is \(\text{IQR}=Q_3-Q_1\).
The median and the IQR are not pulled much by very large or very small values, so they are good partners for skewed data. The mean and the mean absolute deviation, which comes next, work well for data that are fairly symmetric.
5. Mean Absolute Deviation
The mean absolute deviation is the average distance between each data value and the mean:
\[ \text{MAD}=\dfrac{|x_1-\bar{x}|+|x_2-\bar{x}|+\cdots+|x_n-\bar{x}|}{n} \]
A small MAD means the values are packed close to the mean. A large MAD means they are spread out.
- Find the mean.
- Find the distance from each value to the mean (always positive).
- Add the distances.
- Divide by the number of values.
Heights (in cm) of six plants. Garden A: 11, 13, 15, 15, 17, 19. Garden B: 13, 14, 15, 15, 16, 17. Both sets have mean 15 cm and median 15 cm.
Garden A: distances 4, 2, 0, 0, 2, 4 add to 12, so \(\text{MAD}=\dfrac{12}{6}=2\) cm.
Garden B: distances 2, 1, 0, 0, 1, 2 add to 6, so \(\text{MAD}=\dfrac{6}{6}=1\) cm.
The centers match, but Garden B’s plants are more alike: its heights are less spread out.
6. Dot Plots and Box Plots
A dot plot puts one dot above a number line for each data value. It shows every value, so you can see clusters, gaps, and the shape of the data. A box plot shows five numbers: the minimum, \(Q_1\), the median, \(Q_3\), and the maximum. The box covers the middle half of the data, and the line inside the box marks the median.
Over 12 games, Kai scored 4, 6, 7, 8, 9, 10, 12, 13, 15, 16, 18, 20 points. The median is \(\dfrac{10+12}{2}=11\). The lower half has median \(Q_1=\dfrac{7+8}{2}=7.5\) and the upper half has median \(Q_3=\dfrac{15+16}{2}=15.5\). So \(\text{IQR}=15.5-7.5=8\). Rio’s 12 games give the five numbers 8, 10.5, 12.5, 14, 16, so Rio’s IQR is 3.5.
Rio’s narrow box shows steady scoring. Kai’s wide box shows scores that swing from game to game.
7. Comparing Two Data Distributions
When you compare two groups, talk about shape, center, and spread, and describe how much the distributions overlap. A helpful trick is to measure the gap between the two centers in units of variability: divide the difference of the means by the MAD (or the difference of the medians by the IQR).
Ten students in Class 7A and ten in Class 7B logged their weekly practice time. Class 7A has mean 30 minutes and Class 7B has mean 40 minutes; both have a MAD of 4 minutes. The difference of means is \(40-30=10\), and \(\dfrac{10}{4}=2.5\). The centers are 2.5 MADs apart, so the groups clearly differ, yet the two dot plots still overlap between 30 and 40 minutes.
On my home planet we say, “One sample is a hint, many samples are a map.” Compare the gap between centers with the spread. A big gap with small spread means little overlap and a convincing difference.
If the gap between centers is only a small fraction of the spread, the distributions overlap a lot, and the difference between the groups may not mean much. For Kai and Rio, the medians (11 and 12.5) differ by only 1.5 points, much less than Kai’s IQR of 8, so their scoring is hard to tell apart by the median alone.
Key takeaways
- A population is the whole group; a sample is a part of it.
- A random sample gives every member an equal chance and is likely to be representative; biased samples give misleading results.
- Larger random samples give more reliable estimates, but samples still vary.
- To estimate a count: sample proportion \(\times\) population size.
- Center: mean or median. Variability: range, IQR, or MAD.
- MAD is the average distance from the mean.
- Compare distributions by shape, center, spread, and overlap, using the gap between centers measured in MADs or IQRs.
Test yourself: quick challenge for Grade 7
Speed drill for Grade 7: how many in 60 seconds?
🚀 Keep exploring with Zyro
✏️ Math practiceRandom Sampling and Comparing Populations: math practice, Grade 7
📝 Math testsRandom Sampling and Comparing Populations: math test, Grade 7
🎯 Math quizzesRandom Sampling and Comparing Populations: math quiz, Grade 7
✏️ Math practiceProbability and Compound Events: math practice, Grade 7
✏️ Math practiceRatios and Proportional Relationships: math practice, Grade 7
📝 Math testsProbability and Compound Events: math test, Grade 7

