
Data rarely come with an equation attached, so you have to build one. In this chapter you will write the equation of a line from a slope and a point or from two points, then use lines to model real data in a scatter plot, measure how well a model fits, and spot the traps of correlation, causation, and prediction outside the data. You will finish with piecewise and absolute value functions, which glue pieces of lines together.
1. Two forms of a linear equation
You already know that a line has a constant rate of change. Two forms help you write its equation quickly.
Slope-intercept form: \(y = mx + b\), where \(m\) is the slope and \(b\) is the y-intercept (the value of \(y\) when \(x = 0\)).
Point-slope form: \(y - y_1 = m(x - x_1)\), where \(m\) is the slope and \((x_1, y_1)\) is any point on the line.
Point-slope form is handy when the y-intercept is not given. You can always rewrite it as slope-intercept form by distributing and solving for \(y\).
2. Writing an equation from a slope and a point
- Write the point-slope form \(y - y_1 = m(x - x_1)\).
- Substitute the slope \(m\) and the coordinates of the point.
- Distribute, then isolate \(y\) to get \(y = mx + b\).
- Check by substituting the point into your final equation.
\(y - 5 = 3(x - 2)\), so \(y - 5 = 3x - 6\) and \(y = 3x - 1\). Check: \(3(2) - 1 = 5\). The graph below shows that moving 1 unit right raises the line 3 units.
With the point \((-4, 3)\), the factor is \(x - (-4) = x + 4\), not \(x - 4\). Write the subtraction first, then simplify.
3. Writing an equation from two points
With two points you first find the slope, then continue as before.
- Compute the slope \(m = \dfrac{y_2 - y_1}{x_2 - x_1}\), subtracting in the same order on top and bottom.
- Pick either point and use point-slope form.
- Simplify to \(y = mx + b\) and check with the other point.
\(m = \dfrac{-2 - 7}{4 - 1} = \dfrac{-9}{3} = -3\). Then \(y - 7 = -3(x - 1)\), so \(y = -3x + 10\). Check with the other point: \(-3(4) + 10 = -2\). It works.
4. Scatter plots and correlation
A scatter plot shows each data pair as a point in the coordinate plane. Its shape tells you whether the two variables are related.
Correlation describes the direction and strength of a relationship. If \(y\) tends to increase as \(x\) increases, the correlation is positive. If \(y\) tends to decrease, it is negative. If the points show no trend, there is no correlation. The correlation coefficient \(r\) satisfies \(-1 \le r \le 1\): the closer \(|r|\) is to 1, the more closely the points hug a line.
| Pattern of the points | Correlation | Example |
|---|---|---|
| Rising from left to right | Positive | height and shoe size |
| Falling from left to right | Negative | temperature and heavy coats sold |
| No visible trend | None | shoe size and test score |
For example, hours of practice and score on a game often show a positive correlation, while the outdoor temperature and the number of heavy coats sold show a negative one.
5. The line of best fit
When points follow a linear trend, you can draw a line of best fit that passes as close as possible to all of them. Below, a student tracked the hours spent studying and the quiz score (out of 100) for eight practice quizzes.
- Draw a line with about as many points above it as below it, following the trend.
- Choose two points on your line (they need not be data points).
- Find the slope and the equation with the two-point method.
Our line goes through \((2, 60)\) and \((6, 78)\). Its slope is \(\dfrac{78 - 60}{6 - 2} = 4.5\), so \(y - 60 = 4.5(x - 2)\), which gives \(y = 4.5x + 51\). A calculator regression gives about \(y = 4.65x + 50.7\) with \(r \approx 0.996\), a very similar line.
The slope is the predicted change in \(y\) for each 1-unit increase in \(x\). The y-intercept is the predicted value of \(y\) when \(x = 0\), if \(x = 0\) makes sense in the situation.
Each extra hour of studying adds about 4.5 points, and the predicted score with no studying is 51. For 5.5 hours: \(y = 4.5(5.5) + 51 = 75.75\), so about 76 points.
6. Residuals
A residual is the vertical distance between a data point and the model: \(\text{residual} = \text{actual } y - \text{predicted } y\). It is positive when the point is above the line and negative when it is below.
At \(x = 4\) the quiz score was 68, but the line predicts \(4.5(4) + 51 = 69\). The residual is \(68 - 69 = -1\): the point lies 1 point below the line.
| Hours \(x\) | Actual score | Predicted \(4.5x + 51\) | Residual |
|---|---|---|---|
| 1 | 55 | 55.5 | -0.5 |
| 2 | 60 | 60 | 0 |
| 3 | 66 | 64.5 | 1.5 |
| 4 | 68 | 69 | -1 |
| 5 | 75 | 73.5 | 1.5 |
| 6 | 77 | 78 | -1 |
| 7 | 84 | 82.5 | 1.5 |
| 8 | 88 | 87 | 1 |
A residual plot graphs the residuals against \(x\). If the points look randomly scattered around 0, a linear model is appropriate. A curved pattern (for example, a U shape) means a line is the wrong kind of model.
7. Correlation, causation, and prediction
Two variables can move together without one causing the other. Ice cream sales and sunburns are positively correlated, but both rise because of a hidden third factor, hot sunny weather. That hidden factor is called a lurking variable. A cause is established only by a carefully designed experiment.
Interpolation means predicting inside the range of your data. Extrapolation means predicting outside it, which is riskier because the pattern may change.
On my planet we never trust a prediction far from the data. Ask yourself: would this answer still make sense in real life? A quiz score of 105 out of 100 does not.
8. Piecewise and absolute value functions
A piecewise function uses different rules on different parts of its domain. A parking garage might charge $4 for up to 2 hours and then $3 for each additional hour:
\[P(h) = \begin{cases} 4 & \text{if } 0 < h \le 2 \\ 3h - 2 & \text{if } h > 2 \end{cases}\]
For 1.5 hours use the first rule: \(P(1.5) = 4\). For 5 hours use the second: \(P(5) = 3(5) - 2 = 13\), so the fee is $13. Always pick the rule whose condition the input satisfies.
The absolute value function \(y = a|x - h| + k\) is a piecewise function in disguise. Its graph is a V with vertex \((h, k)\); it opens up if \(a > 0\) and down if \(a < 0\). For \(y = |x - 2| - 3\), the vertex is \((2, -3)\). To find the x-intercepts, solve \(|x - 2| = 3\), which gives \(x = 5\) or \(x = -1\).
Key takeaways
- Slope and a point: use \(y - y_1 = m(x - x_1)\), then simplify to \(y = mx + b\).
- Two points: find \(m = \dfrac{y_2 - y_1}{x_2 - x_1}\) first, then use either point.
- A scatter plot shows correlation: positive, negative, or none; \(r\) near \(\pm 1\) means a strong linear relationship.
- A line of best fit models a trend; its slope is a rate of change and its intercept is a starting value.
- Residual = actual \(-\) predicted; a random residual plot supports a linear model.
- Correlation does not prove causation; extrapolating beyond the data is risky.
- Piecewise functions use a different rule on each interval; \(y = a|x - h| + k\) has its vertex at \((h, k)\).
Test yourself: quick challenge for Grade 9
Speed drill for Grade 9: how many in 60 seconds?
🚀 Keep exploring with Zyro
✏️ Math practiceWriting Linear Equations and Scatter Plots: math practice, Grade 9
📝 Math testsWriting Linear Equations and Scatter Plots: math test, Grade 9
🎯 Math quizzesWriting Linear Equations and Scatter Plots: math quiz, Grade 9
✏️ Math practiceSystems of Linear Equations: math practice, Grade 9
✏️ Math practiceExponent Rules and Polynomials: math practice, Grade 9
📝 Math testsSystems of Linear Equations: math test, Grade 9

