University Introductory Statistics — Simple Linear Regression and Correlation Worksheet
Two quantitative variables measured on the same units, and everything a first course does with them. The correlation coefficient, from deviations and from standard scores, and what happens to it when the units change; the least-squares line, fitted from raw pairs and from summary sums; slope, intercept and predictions read as sentences about the situation; the coefficient of determination, both from the sums of squares and from a printed analysis-of-variance table; residuals and the conditions a residual plot is there to check; the t test and the confidence interval for the slope; and the two intervals for a response at a given value of x. Have a look on this page, then print the free PDF when you want to write on it.
Practice worksheet — free PDF
No email, no account, no watermark. Teachers: photocopy it for your classes freely. Worked solutions and 8 harder problems come with the University Introductory Statistics bundle.
9 of the 12 questions
Each question targets one named concept from the sheet. Read them here, or print the PDF — it has working space under each one. 9 of the 12 questions are printed below. The other 3 are built on a diagram or a table of values that does not translate to the page, so they are in the free PDF — marked below where they would have come.
-
Q1Scatterplots and the Correlation Coefficient
This question is built around a diagram or a table of values. Open it in the PDF.
-
Q2Scatterplots and the Correlation Coefficient
For greenhouses, the correlation between the mean daily temperature (in C) and the daily water use (in litres) is . Recall that , where and are standard scores. Give the new correlation, with a one-line reason, if
- the temperatures are converted to degrees Fahrenheit, ;
- the water use is recorded in millilitres;
- is replaced by the water left in a full L tank at the end of the day, ;
- the roles of the two variables are swapped.
-
Q3The Least-Squares Regression Line
A bicycle courier records the distance (km) and the delivery time (min) for six deliveries: The least-squares line has and , with and . Find the line, and use it to predict the time for a km delivery.
-
Q4The Least-Squares Regression Line
An agronomist records June rainfall (mm) and hay yield (bales per hectare) on farms and reports only the sums Using and , find the least-squares line and predict the yield after mm of June rain.
-
Q5Interpreting Slope, Intercept and Predictions
For households with to occupants, the least-squares line relating the monthly water bill ($) to the number of occupants is .
- Predict the monthly bill of a household of .
- By how much do the predicted bills of two households differ if one has more occupants than the other?
- Compute the prediction for a household of , and say whether it should be trusted.
-
Q6The Coefficient of Determination
This question is built around a diagram or a table of values. Open it in the PDF.
-
Q7The Coefficient of Determination
For cafés on the same street, daily foot traffic (hundreds of passers-by) and daily sales (hundreds of dollars) give With , and , find the least-squares line, SSE and . Check against .
-
Q8Residuals and the Conditions for Regression
A garden centre prices six clay pots by their diameter. The least-squares line relating price ($) to diameter (cm) is . The data are Compute every residual , check that the residuals sum to , and name the pot whose price is furthest below the line.
-
Q9Inference for the Slope
For sourdough loaves, a baker regresses the rise (cm) on the proofing time (h) and finds The standard error of the slope is and the test statistic has degrees of freedom. Use . Assuming the regression conditions hold, test at the level whether proofing time and rise are linearly related.
-
Q10Inference for the Slope
This question is built around a diagram or a table of values. Open it in the PDF.
-
Q11Confidence and Prediction Intervals for a Response
A maple-syrup producer records, for batches, the sap boiled (hundreds of litres, between and ) and the syrup obtained (litres). The least-squares line is , with , and . At , with , Use . For batches of L of sap, compute a confidence interval for the mean syrup yield and a prediction interval for one batch.
-
Q12Synthesis — drawing on several topics in this unit
A bakery records the number of staff on five morning shifts and the trays of pastries finished by 9 a.m.: Use , , , , and .
- Find the least-squares line.
- Compute the residuals, SSE and .
- Assuming the regression conditions hold, test against at the level.
The 8 challenge problems for this topic are a separate, paid sheet and are not reproduced here.
What does this set assume? This set assumes Secondary 5 mathematics only — the equation of a line, slope as a rate of change, squares and square roots, and summation notation — together with the mean and the standard deviation from Descriptive Statistics and the standard score built on them. From Sampling Distributions and the Central Limit Theorem it takes standard error as the name for the standard deviation of an estimate; from Confidence Intervals and Hypothesis Tests it takes the t distribution, degrees of freedom, the shape of a two-sided test at a stated level and the interval-test duality; from Continuous Distributions and the Normal Model, the normal curve that the conditions for regression ask of the errors. Probability Rules and Conditional Probability, Discrete Random Variables and Two-Sample Inference and Chi-Square Tests are not used here. The course is not calculus-based and this set is no exception: the least-squares line comes from the summary formulas, never from setting a derivative to zero, and no question anywhere asks for a derivative or an integral. Deliberately outside it: more than one explanatory variable, polynomial and other curved fits, transforming a variable to straighten a relationship, analysis of variance as a technique of its own — the one table here is printed to be read, not built — inference on variances, and producing output from a statistical package, which no question asks for. Every t value a question needs is supplied with it.
Which course is this for? In the public course calendars of Montreal universities, this material is part of the courses numbered MAST 333, STT1700, MAT2080 and MAT350. Each course orders and weights the topics its own way, so check your own outline for what your exam covers. Which sets match your course.
How to do every concept on this sheet
This is the part worksheet sites usually leave out. Below is the actual reasoning behind each group of questions — not a full solution set, the decisions that get you to one. Read it before you start, or after you get stuck.
Scatterplots and the correlation coefficient
r measures one thing only: how closely the points cluster about a straight line, and in which direction. It is built from deviations, and every quantity on this sheet is a sum of products of deviations.
Sxx = Σ(xᵢ − x̄)², Syy = Σ(yᵢ − ȳ)², Sxy = Σ(xᵢ − x̄)(yᵢ − ȳ), r = Sxy / √(Sxx·Syy)Computing r from a small table
Q1 gives five pairs; do it in columns, not in one line.
- 1Means first
Find x̄ and ȳ, and write them at the top of the table. Every later column is measured from them.
- 2Five columns
xᵢ − x̄, yᵢ − ȳ, their squares, and their product. The deviation columns must each add to zero — that is the free check on the arithmetic.
- 3Add, then combine
The three column totals are Sxx, Syy and Sxy; only then divide by the square root of the product.
- 4Say it in words
The sign of Sxy gives the direction; the distance of r from 0 and from 1 gives the strength. Q1 asks for both, so a bare number is an unfinished answer.
r always lands between −1 and 1. A value outside that range means a sign lost in a deviation column or Sxx and Syy swapped under the root. Check before interpreting, not after.
Q2 asks nothing arithmetical at all: it hands you a correlation for forty greenhouses and four changes to the data, and wants the new correlation with a one-line reason for each. Work from the standard-score form of r, which the question states.
r = (1/(n − 1)) Σ (standard score of xᵢ)(standard score of yᵢ), z = (value − mean)/(standard deviation)For each part, ask what the change does to the standard scores. Adding a constant shifts the values and the mean together, so every deviation is untouched. Multiplying by a positive constant multiplies deviation and standard deviation alike, so each z is unchanged. Multiplying by a negative constant leaves the size of each z alone and reverses its sign — and the sum of products then reverses with it. And the defining formula treats the two variables symmetrically, which settles the last part without any computation. Reason from the z scores every time; guessing from the story is how this question is lost.
The least-squares regression line
The line is the one making the sum of squared vertical distances from the points as small as possible. In this course it comes from the same sums as r, and the slope is read off first.
b₁ = Sxy / Sxx, b₀ = ȳ − b₁x̄, ŷ = b₀ + b₁xQ3 gives six pairs of distances and delivery times: build the deviation table exactly as for Q1, take the slope from Sxy and Sxx, then the intercept from the two means. Because b₀ is found from b₁, a slope rounded early poisons the intercept — carry the slope to more decimal places than the answer needs, and round once at the end. The prediction asked for is the substitution of one value into the finished line.
The line passes through the point of means. Substituting x̄ into a correct line returns ȳ. That is one line of arithmetic and it catches almost every slip in b₀.
Q4 gives no pairs — only n and four sums for twenty-five farms — which is how a regression is usually reported. The computational forms below are algebraically the same as the deviation forms, so nothing new is being calculated, only a different route to Sxx and Sxy.
Sxx = Σx² − (Σx)²/n, Sxy = Σxy − (Σx)(Σy)/n, x̄ = Σx/n, ȳ = Σy/nΣx² and (Σx)² are different numbers. Square each value then add, against add then square. Mixing them up is the single commonest error in the summary-sum route, and it usually produces a slope that is wildly out rather than slightly out.
Interpreting slope, intercept and predictions
Q5 supplies a fitted line for thirty households and asks for no fitting at all — only for what the line says. Three readings, each a sentence with units in it.
Reading a fitted line
Units make the interpretation; a bare number does not.
- 1Slope
The change in the predicted response for a one-unit increase in x — here, one more occupant. A change of several units multiplies the slope by that number, which is what Q5(b) asks and why no prediction has to be computed twice.
- 2Intercept
The predicted response at x = 0. It is meaningful only when x = 0 is a real possibility and lies within the data; otherwise it is a placeholder that positions the line.
- 3Prediction
Substitute and state the units. Then check the value of x against the range the data covered.
A prediction is only as good as the range behind it. A line is evidence over the range of x that was observed, and the question states that range. Outside it nothing guarantees the relationship stays linear, or stays at all. Q5(c) asks for a prediction and for a judgement on it, so compare the value with the stated range and let that comparison — not the arithmetic, which always works — settle the judgement.
The coefficient of determination
r² is the share of the variation in y that the line accounts for. It comes either from the sums of squares or, squared back, from r itself.
SST = SSR + SSE, r² = SSR/SST = 1 − SSE/SST, s = √(SSE/(n − 2))Q6 gives an analysis-of-variance table for a set of bus routes and asks you to work outwards from it. The residual degrees of freedom are n − 2, so the df column fixes the sample size before anything else; the total sum of squares is the sum of the two rows above it, which is why the table leaves that cell blank. Then r² is a ratio of two entries, and s — the typical size of a residual, in the units of y — divides SSE by its own degrees of freedom, never by n.
r² loses the sign; the slope gives it back. Taking a square root to recover r from r² leaves two candidates. The direction of the association is the direction of the slope, which Q6 states, so match the sign of r to the sign of b₁ and say which you used.
Q7 comes at the same quantities from summary values for twelve cafés. Here SSE is got without residuals at all, from SSE = Syy − b₁Sxy, and r² from 1 − SSE/Syy. The final instruction is the point of the question: compute r from Sxy and the root of Sxx·Syy, square it, and see that the two routes meet. An agreement to several decimal places is a check; a disagreement means one of the S values was used in the wrong slot.
Residuals and the conditions for regression
A residual is what the line did not explain about one observation: the vertical gap between the point and the line, signed.
eᵢ = yᵢ − ŷᵢ, ŷᵢ = b₀ + b₁xᵢ, Σeᵢ = 0 for a least-squares lineQ8 gives six clay pots and the fitted line, and asks for every residual. Work down the list: predict, subtract, keep the sign. The sum being zero is a property of the least-squares fit, so when your column does not add to zero the error is in the arithmetic, not in the data — that is why the question asks for the check before anything else.
"Furthest below the line" is the most negative residual, not the largest one. A point above the line has a positive residual however large the gap. Read the wording of the last part carefully and sort by signed value.
Behind every inference later on this sheet sit four conditions, and a residual plot — residuals against x — is how each is examined: the relationship is linear (no curved band of points), the spread of the residuals is the same all along (no funnel), the observations are independent, and the residuals are roughly normal about zero. A pattern in that plot names the condition that is in trouble, which is why the inference questions state that the conditions hold before asking for a test.
Inference for the slope
The fitted slope is an estimate of an unknown slope β₁, and a different sample would give a different b₁. Its standard error measures that variability, and everything inferential for a regression is built on it.
SE(b₁) = s/√Sxx, t = b₁/SE(b₁), df = n − 2The t test for a slope
Q9 gives b₁, s and Sxx for twenty loaves, and the critical value.
- 1State the hypotheses
H₀: β₁ = 0 against Hₐ: β₁ ≠ 0. A zero slope is the claim that x carries no linear information about y, which is why "is there a linear relationship" is tested this way.
- 2Standard error, then statistic
Divide s by the square root of Sxx, then the slope by that. Keep the sign of t.
- 3Compare on n − 2 degrees of freedom
The two-sided rejection region is |t| beyond the supplied critical value. The question hands you that value, so no table reading is needed — but say which df it belongs to.
- 4Answer in the situation
A conclusion about proofing time and rise, at the stated level, in one sentence. A verdict with no context and no level is half an answer.
Q10 works from a printed table of coefficients for fourteen apartments, and the reading is the first decision: the row for the explanatory variable, not the row for the constant, carries the slope and its standard error. Build the interval from the general form, with the t value the question supplies.
b₁ ± t(α/2, n − 2) · SE(b₁)Interval and test answer each other. A two-sided test of H₀: β₁ = 0 at level α and a 100(1 − α)% interval for β₁ agree by construction: the test is decided by whether zero is one of the values the interval covers. Q10 asks you to make that link explicitly, so say what you looked for in the interval and what it tells you about the test, rather than running the test again.
Confidence and prediction intervals for a response
Two different questions live at the same x*, and they have different answers. One asks for the average response of all units at that value; the other asks for one new unit. The second must allow for that unit's own scatter about the line, which is the extra 1 under the root.
mean response: ŷ* ± t(α/2, n − 2)·s·√(1/n + (x* − x̄)²/Sxx) new observation: ŷ* ± t(α/2, n − 2)·s·√(1 + 1/n + (x* − x̄)²/Sxx)Q11 gives twelve batches of sap, the fitted line, x̄, Sxx and s. Two things decide the question before any arithmetic. First, the units: x is counted in hundreds of litres and the batch size is quoted in litres, so convert before substituting, or every number afterwards is wrong. Second, which interval each part is asking for — the wording "the mean syrup yield" against "one batch" is the whole distinction. Compute ŷ* once; only the root differs between the two.
Width grows with distance from x̄. The term (x* − x̄)²/Sxx is zero at the centre of the data and grows either side, so both intervals are narrowest at x̄ and flare out towards the ends of the observed range. Check that the x* you use lies inside the range the question states before you report anything.
The synthesis question
Q12 runs the whole sheet on five morning shifts at a bakery, and its three parts are the three chapters in order: fit the line with the slope-and-intercept formulas, then the residuals, SSE and r², then the t test for the slope against a two-sided alternative at the stated level.
Two things make it harder than its parts. The sample is small, so the degrees of freedom are few and the supplied critical value is correspondingly large — read it as a reminder of how little a handful of points can settle. And part (b) feeds part (c): SSE gives s, s gives SE(b₁), and SE(b₁) gives t, so a residual mis-subtracted in the middle carries all the way to the conclusion. Recompute r² the other way, from r, before moving on, and finish with a sentence about staff and trays rather than about t.
Preview all 7 pages
Click any page to open the full PDF.
Getting the most out of it
Set the table up before touching a formula
Means at the top, then columns for the two deviations, their squares and their product. Every quantity on this topic — r, b₁, SSE, r² — is a combination of those three totals, so one careful table answers a whole question and gives you the add-to-zero check for free.
Say the slope out loud, with units
"One more occupant is associated with so many more dollars on the predicted bill" is the form the interpretation questions want. Numbers without units and without the words "is associated with" are what lose the interpretation marks that every part of this chapter carries.
Check the x you are predicting at
Before reporting any prediction or interval, compare the value of x with the range the data covered. Inside it, the line is evidence; outside it, the arithmetic still works and the prediction does not. Q5 and Q11 both turn on this, and an exam question rarely warns you.
Reach for the second route as a check
r² from the sums of squares against r² from r; the point of means substituted into the line; the residual column adding to zero; the interval compared with the test. This topic is unusually rich in checks that cost one line, and using them is faster than finding the slip later.
End on the situation, not the statistic
Every question here asks for tomatoes, deliveries, bus routes, loaves or rent, and the marks at the end belong to the sentence that answers in those terms at the stated level. Write the number, then write what it means about the two variables.
Want the solutions, or something more challenging?
The worksheet, the questions above and every explanation on this page stay free permanently. Three more PDFs exist for this topic — the worked answer key, a harder problem set, and the answer key to that. They come with the University Introductory Statistics Solutions Bundle, beside the unit notes and the unit test, which is what keeps the rest of the series free.
What else exists for Simple Linear Regression and Correlation
Three PDFs · 14 pages · all three are in the bundle below.
- Answer key — 2 pages. All 12 questions worked step by step, including the restrictions and the justifications. Not a list of final answers.
- Challenge problems — 8 pages, 8 problems. A separate sheet at exam-plus difficulty covering the same 7 concepts. Harder than anything on the free sheet.
- Challenge answer key — 4 pages. Every challenge problem worked to the same standard, with the checks shown.
- PDF, letter size, print-ready.
The one thing that's for sale
Every University Introductory Statistics topic — the complete Solutions Bundle
One download, one payment, the whole program. For all 9 University Introductory Statistics units: the worksheet, the reference notes, the challenge set, the unit test and every answer key — including this one.
- Worked solutions, not answer lists — every step written out
- Covers the whole year's program at this level
- Less than the price of one hour of tutoring — for the entire year's solutions
Taking Secondary 1 Math as well? The Secondary 1 Math bundle covers all 15 of its units — 106 PDFs, 474 pages — on the same terms.
Taking Secondary 2 Math as well? The Secondary 2 Math bundle covers all 14 of its units — 98 PDFs, 449 pages — on the same terms.
Taking Secondary 3 Math as well? The Secondary 3 Math bundle covers all 11 of its units — 77 PDFs, 367 pages — on the same terms.
Taking Secondary 4 Math as well? The Secondary 4 Math bundle covers all 17 of its units — 122 PDFs, 466 pages — on the same terms.
Taking Secondary 5 Math as well? The Secondary 5 Math bundle covers all 21 of its units — 147 PDFs, 589 pages — on the same terms.
Taking CEGEP Calculus I as well? The CEGEP Calculus I bundle covers all 9 of its units — 72 PDFs, 371 pages — on the same terms.
Taking CEGEP Calculus II as well? The CEGEP Calculus II bundle covers all 8 of its units — 64 PDFs, 335 pages — on the same terms.
Taking CEGEP Linear Algebra as well? The CEGEP Linear Algebra bundle covers all 7 of its units — 56 PDFs, 298 pages — on the same terms.
Taking University Calculus III as well? The University Calculus III bundle covers all 9 of its units — 72 PDFs, 516 pages — on the same terms.
Taking University Linear Algebra as well? The University Linear Algebra bundle covers all 9 of its units — 72 PDFs, 496 pages — on the same terms.
Taking University Differential Equations as well? The University Differential Equations bundle covers all 9 of its units — 36 PDFs, 182 pages — on the same terms.
Taking University Business Math as well? The University Business Math bundle covers all 9 of its units — 36 PDFs, 190 pages — on the same terms.
Taking University Discrete Math as well? The University Discrete Math bundle covers all 9 of its units — 36 PDFs, 155 pages — on the same terms.
Taking AP Calculus AB as well? The AP Calculus AB bundle covers all 8 of its units — 64 PDFs, 527 pages — on the same terms.
Common questions
Is this worksheet really free?
Yes — the questions are on this page to read, and the PDF downloads directly, no email and no account. The one paid item is optional: the complete University Introductory Statistics Solutions Bundle, which covers every set at this level.
Which university courses is this for?
The course codes listed on this page are taken from the public course calendars of universities that teach a first, service-level statistics course. Each course orders and weights the chapters its own way, and some reach a chapter later or not at all — so check the outline for your own section to see where this set falls in your term.
What do I need to know before starting this set?
Means and standard deviations, standard scores, and the inference vocabulary of the earlier sets: standard error, degrees of freedom, the t distribution, and what a two-sided test at a stated level and a confidence interval each claim. The algebra is Secondary 5 throughout.
What is the difference between r and r²?
r carries a direction as well as a strength, and lies between −1 and 1. r² is a proportion, between 0 and 1, and says what share of the variation in y the line accounts for. Squaring loses the sign, so recovering r from r² needs the direction of the slope to settle it.
Why is it n − 2 degrees of freedom and not n − 1?
Two quantities are estimated from the data before any residual exists — the slope and the intercept — so two constraints are used up. That is why s divides SSE by n − 2, and why every t procedure on this sheet is read on n − 2 degrees of freedom.
When do I use the confidence interval and when the prediction interval?
Read the wording. If the question asks about the average response of all units at that value of x, it is the confidence interval for a mean response. If it asks about one new unit, it is the prediction interval, which is wider because it also has to cover that unit's own scatter about the line.
Does a strong correlation mean one variable causes the other?
No. r measures how close the points lie to a line and nothing more. A variable left out of the study can move both, and the arithmetic is the same whichever variable is called x, so a causal claim needs an argument about how the data were collected, not a larger r.
Do I need statistical software or a graphing calculator?
No. Every question is done with an ordinary calculator from the summary formulas, and where output appears it is printed in the question for you to read. Nothing here asks you to produce it.
Can teachers use this in class?
Yes. Print and photocopy it for your own classes freely — I just ask that the tutorinmontreal.ca footer stays on the page.
I'm stuck on one question. Can you help?
Yes — through one-on-one tutoring, in Montreal or online. Get in touch to arrange a session, or see the current rates.
← All 9 University Introductory Statistics worksheets · Secondary 1 Math series (15 sheets) → · Secondary 2 Math series (14 sheets) → · Secondary 3 Math series (11 sheets) → · Secondary 4 Math series (17 sheets) → · Secondary 5 Math series (21 sheets) → · CEGEP Calculus I series (9 sheets) → · CEGEP Calculus II series (8 sheets) → · CEGEP Linear Algebra series (7 sheets) → · University Calculus III series (9 sheets) → · University Linear Algebra series (9 sheets) → · University Differential Equations series (9 sheets) → · University Business Math series (9 sheets) → · University Discrete Math series (9 sheets) → · AP Calculus AB series (8 sheets) →




