All lessons

Two-Variable Data: Models and Scatterplots

What this skill is

Two-variable data pairs each xx-value with a yy-value, usually shown in a scatterplot or a table. The SAT asks you to describe the association, use a line of best fit to make predictions, interpret the slope and intercept in context, compute residuals, and decide whether a linear or exponential model fits the data better.

Key ideas

  • Direction: points rising left to right show a positive association; falling shows a negative association.
  • Strength: points close to a line show a strong association; a loose cloud shows a weak one.
  • A line of best fit gives predicted values, not exact ones. Plug the xx-value into the equation to predict yy.
  • The slope is the predicted change in yy for each increase of 11 in xx. The yy-intercept is the predicted yy when x=0x = 0, which may not make sense in context.
  • A residual is actual−predicted\text{actual} - \text{predicted}. A positive residual means the point lies above the line.
  • Linear data changes by a constant difference; exponential data changes by a constant ratio.

Formulas and rules

ModelPattern in a tableEquation
linearadd the same amount each stepy=mx+by = mx + b
exponentialmultiply by the same factor each stepy=a⋅bxy = a \cdot b^x
  • "Starts at aa and triples every kk years" is y=a⋅3x/ky = a \cdot 3^{x/k}.
  • Line through two points: slope m=y2−y1x2−x1m = \frac{y_2 - y_1}{x_2 - x_1}, then solve for bb.

Example: x=0,1,2,3x = 0, 1, 2, 3 with y=5,15,45,135y = 5, 15, 45, 135. The differences grow (10,30,9010, 30, 90) but each ratio is 33, so the model is y=5⋅3xy = 5 \cdot 3^x.

Worked example 1 (easy)

A café finds that the line of best fit for daily hot-drink sales yy against the outdoor temperature xx, in degrees, is y=−2.5x+140y = -2.5x + 140. Predict sales on a 3030-degree day and interpret −2.5-2.5.

  1. Predict: y=−2.5(30)+140=−75+140=65y = -2.5(30) + 140 = -75 + 140 = 65 drinks.
  2. Interpret: for each 11-degree increase in temperature, predicted sales decrease by 2.52.5 drinks.

Worked example 2 (SAT-level)

A line of best fit passes through (4,18)(4, 18) and (10,33)(10, 33). One data point is (16,44.5)(16, 44.5). What is the residual for this point?

  1. Slope: 33−1810−4=156=2.5\frac{33 - 18}{10 - 4} = \frac{15}{6} = 2.5.
  2. Intercept: 18=2.5(4)+b18 = 2.5(4) + b, so b=8b = 8. The line is y=2.5x+8y = 2.5x + 8.
  3. Predicted value at x=16x = 16: 2.5(16)+8=482.5(16) + 8 = 48.
  4. Residual: 44.5−48=−3.544.5 - 48 = -3.5.

Check: the line also gives 2.5(10)+8=332.5(10) + 8 = 33 at x=10x = 10. The negative residual means the actual point lies 3.53.5 units below the line.

Common traps

  • Residual backward. It is actual minus predicted. Predicted minus actual gives the opposite sign.
  • Treating a prediction as certain. Answer choices with "exactly" or "will" are usually wrong; a model gives an estimate.
  • Interpreting slope as a total. In d=3.2t+7d = 3.2t + 7, 3.23.2 is the predicted growth per day, not the total size.
  • Calling exponential data linear. Check differences and ratios. A constant ratio means exponential, even if the first few values look close to a line.
  • Confusing association with causation. A strong scatterplot pattern does not show that xx causes yy.
Practice questions