University Introductory Statistics — Two-Sample Inference and Chi-Square Tests Worksheet
Comparison, which is what almost every real study is about. Two independent large samples and the standard error of x̄₁ − x̄₂; the pooled two-sample t procedure for small samples from normal populations; paired data, where the two measurements belong to the same unit and the analysis runs on the differences; intervals and tests for a difference of two proportions; and the chi-square family — goodness of fit against a stated model, expected counts in a contingency table, the test of independence, and the line between an association and a cause. Every set of hypotheses is stated, every condition is checked in writing, and every test ends in a sentence about the variables rather than a number. Have a look on this page, then print the free PDF when you want to write on it.
Practice worksheet — free PDF
No email, no account, no watermark. Teachers: photocopy it for your classes freely. Worked solutions and 9 harder problems come with the University Introductory Statistics bundle.
11 of the 15 questions
Each question targets one named concept from the sheet. Read them here, or print the PDF — it has working space under each one. 11 of the 15 questions are printed below. The other 4 are built on a diagram or a table of values that does not translate to the page, so they are in the free PDF — marked below where they would have come.
-
Q1Large-Sample Inference for a Difference of Two Means
A furniture retailer compares two formats of assembly instructions for the same flat-pack shelf. Of customers chosen at random, receive printed diagrams (format A) and receive a video (format B), independently. The assembly times, in minutes, give Use and .
- Find the standard error of .
- Construct a confidence interval for , the difference in mean assembly times.
- The assembly times are known to be right-skewed. Explain why the interval in (b) is still valid.
-
Q2Large-Sample Inference for a Difference of Two Means
Two courier companies, X and Y, deliver parcels in the same region. An online retailer claims that Y is faster on average. Random samples of deliveries give the time from order to doorstep, in hours: Test the retailer's claim at the level, giving hypotheses, conditions, test statistic, decision and a conclusion in context. Report the -value. Use , and
-
Q3The Pooled Two-Sample t Procedure
A woodworking supplier compares the drying time, in minutes, of two brands of varnish. Independent random samples of test panels give Assume both populations of drying times are normal with equal variances. Use , and .
- Compute the pooled standard deviation .
- Construct a confidence interval for .
- Does the interval allow the two brands to have the same mean drying time?
-
Q4The Pooled Two-Sample t Procedure
A plant nursery grows seedlings of one species under two light regimes. After four weeks, the heights in centimetres of randomly chosen seedlings are summarized as Assume heights are normally distributed under each regime with a common variance. Test at whether the mean heights differ, in the five-part form, and bracket the -value. Upper-tail values for degrees of freedom:
-
Q5The Paired t Test
This question is built around a diagram or a table of values. Open it in the PDF.
-
Q6The Paired t Test
Twelve club runners each run a km time trial in their old shoes and, a week later, in a new model. For each runner in seconds, and the twelve differences have and . Assume the differences are normal. Use , and .
- Construct a confidence interval for .
- What does the sign of the interval's endpoints say about the new shoes?
-
Q7Inference for a Difference of Two Proportions
A grocery chain samples customers who shop through its app and, independently, who shop in store. Of the app shoppers, used at least one coupon; of the in-store shoppers, did. Use , and take as the condition for the interval at least successes and failures in each sample.
- Check the condition.
- Construct a confidence interval for , the difference between the proportions of app and in-store shoppers who use a coupon.
-
Q8Inference for a Difference of Two Proportions
A dental clinic compares two appointment reminders. Of patients chosen at random to receive a text message, missed their appointment; of patients chosen at random to receive a phone call, missed theirs. Test at whether the no-show rates differ, in the five-part form, and give the -value. Take as the condition at least expected successes and failures in each sample under . Use and
-
Q9The Chi-Square Goodness-of-Fit Test
A municipal helpline wants to know whether calls are spread evenly across the five weekdays. In a random sample of calls, the counts from Monday to Friday are . Test at whether the calls are uniformly distributed over the weekdays, in the five-part form; the test requires every expected count to be at least . Use , and .
-
Q10The Chi-Square Goodness-of-Fit Test
A ferry operator's planning assumes that its passengers are foot passengers, travelling with a car and with a bicycle. A random sample of passengers contains on foot, with a car and with a bicycle. Test the assumption at in the five-part form and bracket the -value. The test requires every expected count to be at least . Upper-tail values:
-
Q11Contingency Tables and Expected Counts
This question is built around a diagram or a table of values. Open it in the PDF.
-
Q12The Chi-Square Test of Independence
This question is built around a diagram or a table of values. Open it in the PDF.
-
Q13The Chi-Square Test of Independence
This question is built around a diagram or a table of values. Open it in the PDF.
-
Q14Association Versus Causation in Categorical Data
A random sample of households in one city is classified by whether the home has a smoke detector that works and whether the household reported a kitchen fire in the past five years. A chi-square test of independence gives . For each statement, say whether the study supports it, and give a one-line reason.
- Having a working smoke detector and reporting a kitchen fire are associated in this city's households.
- Installing a working smoke detector changes a household's chance of a kitchen fire.
- If the two variables were independent, a table at least this far from the expected counts would occur in about of samples.
- The result applies to households in every city.
-
Q15Synthesis — drawing on several topics in this unit
A library tests two wordings of an overdue-book email. Of borrowers chosen at random to receive wording 1, return the book within a week; of receiving wording 2, do. The test requires at least expected successes and expected failures in each group under , and the chi-square test every expected count at least . Use , and .
- Carry out the two-sided pooled test of and give its -value.
- Arrange the data as a table and compute the chi-square statistic for independence of wording and prompt return.
- Compare with , and compare the two decisions at .
The 9 challenge problems for this topic are a separate, paid sheet and are not reproduced here.
What does this set assume? This set stands on Secondary 5 mathematics — algebra, square roots, percentages — and on the earlier University Introductory Statistics sets, more of them than any other set in the course. From Descriptive Statistics it takes the sample mean, the sample standard deviation and the reading of a two-way table; from Probability Rules and Conditional Probability, independence as a statement about probabilities; from Discrete Random Variables and Continuous Distributions and the Normal Model, the binomial count and the standard normal table; from Sampling Distributions and the Central Limit Theorem, the reason a standard error carries a square root and the reason a large sample lets a skewed population be ignored; and from Confidence Intervals and Hypothesis Tests, the whole apparatus of an estimate plus or minus a critical value times a standard error, and of hypotheses, conditions, a test statistic, a decision and a conclusion in context. Nothing is calculus-based: no question asks for a derivative or an integral, and every critical value and tail probability needed is printed in the question. The two-sample t procedure here is the pooled one, with the equal-variance assumption stated for you, so you are never left choosing between two versions of it. Deliberately out of this set and out of the course: inference on variances, so no test comparing two spreads; the comparison of three or more means; the exact test for a small table and the continuity correction for a two-by-two one; and anything read off computer output. Simple Linear Regression and Correlation, the last set of the course, takes up the other way two variables can be compared.
Which course is this for? In the public course calendars of Montreal universities, this material is part of the courses numbered MATH 203, MAST 333, STT1700, MAT2080, MAT1185 and MAT350. Each course orders and weights the topics its own way, so check your own outline for what your exam covers. Which sets match your course.
How to do every concept on this sheet
This is the part worksheet sites usually leave out. Below is the actual reasoning behind each group of questions — not a full solution set, the decisions that get you to one. Read it before you start, or after you get stuck.
Large-sample inference for a difference of two means
Everything in this set is one estimate minus another, and the first job is always the same: find the standard error of the difference. Two independent samples have independent means, so the variances add — never the standard deviations.
SE(x̄₁ − x̄₂) = √(s₁²/n₁ + s₂²/n₂)Q1(a) asks for exactly that quantity, and Q1(b) wraps it in the usual interval, (x̄₁ − x̄₂) ± z × SE. Two critical values are printed in Q1; only one of them belongs to a two-sided interval at the confidence level asked for, and choosing between them is part of the question. Keep the subtraction in the order the question defines — the wording fixes which population is 1 and which is 2, and reversing it reverses the sign of both endpoints and the meaning of the sentence you write about them.
This procedure replaces each unknown population standard deviation by its sample value and still uses a z critical value, which is only safe when both samples are large — so that condition goes on the page before the arithmetic, not after it. Q1(c) is the part that carries the marks: it tells you the populations are skewed and asks why the interval is still usable. The justification is not about the data being normal — it is a statement about the sampling distribution of a mean, and it has a condition attached. Name the result, say what it does to each sample mean, and say what it needs.
Q2 is the same machinery as a test. The retailer's claim points in one direction, so the alternative hypothesis is one-sided and the p-value comes from one tail only. Write H₀ as an equality between the two population means (or, equivalently, μ₁ − μ₂ = 0), and put the claim in the alternative — a claim is what you are trying to find evidence for, never what you assume. Four normal-table values are printed; the one you need is the one nearest your test statistic, and reading the p-value off the wrong tail is the standard way to lose the question after doing the arithmetic correctly.
The five-part form
Q2, Q4, Q5, Q8, Q9, Q10 and Q13 all ask for it. Same five headings every time.
- 1Hypotheses
Both of them, about population parameters, never about x̄ or p̂. H₀ is the equality; the alternative carries the direction the question suggests, or ≠ when it suggests none.
- 2Conditions
Independence, randomisation, and the size or shape condition the procedure needs. Write them out — a condition checked silently is a condition not checked.
- 3Test statistic
(estimate − hypothesised value) / SE, with the SE the procedure prescribes. State which distribution it is compared against, and its degrees of freedom where it has any.
- 4Decision
Against the critical value, or against α using the p-value. The two must agree; if they do not, one of them has been read from the wrong row or tail.
- 5Conclusion in context
One sentence naming the two groups and the variable, in the units of the problem. "There is evidence that…" or "there is not enough evidence to conclude that…" — never "H₀ is true".
The pooled two-sample t procedure
When both samples are small, the sample standard deviations are too unreliable to be treated as if they were the population values, and the z procedure is replaced by a t one. The version used throughout this set assumes a common population variance, which the question states for you, and estimates it by a weighted average of the two sample variances.
sₚ² = ((n₁ − 1)s₁² + (n₂ − 1)s₂²) / (n₁ + n₂ − 2), SE = sₚ √(1/n₁ + 1/n₂)Q3(a) asks for sₚ itself. Note the shape of the weights: the larger sample pulls the pooled value towards its own standard deviation, and sₚ always lands between s₁ and s₂ — a free check on the arithmetic before you go any further. Q3(b) builds the interval with that SE and a t critical value.
Three critical values are printed, and only one has the right degrees of freedom. The pooled procedure has n₁ + n₂ − 2 degrees of freedom, not n₁ + n₂ and not the size of either sample. Q3 and Q6 both offer neighbouring values deliberately; compute df from the formula first, then go looking for it in the list.
Q3(c) asks whether the interval leaves room for the two means to be equal. The rule is general and worth stating in your own words before you use it: an interval for a difference is a set of plausible values for μ₁ − μ₂, so equality is plausible exactly when 0 is one of the values in the interval. Answer by saying where 0 sits relative to your two endpoints, not by comparing x̄₁ with x̄₂.
Q4 runs the same procedure as a test, in the five-part form, with a two-sided alternative because "differ" names no direction. Four upper-tail t values at the right degrees of freedom are printed so that the p-value can be bracketed: find the two printed values your statistic falls between and report the corresponding range. With a two-sided alternative the tail areas are doubled, and forgetting to double them is what turns a correct statistic into a wrong bracket. A bracket is a complete answer here — a printed table cannot give more.
Paired data, and why the design decides the test
Q5 and Q6 look like two-sample problems and are not. In both, the two measurements belong to the same unit — the same van, the same runner — so the two columns are not independent samples and the two-sample standard error does not apply. Reduce each pair to a single difference and run a one-sample t procedure on those differences.
t = d̄ / (s_d/√n), df = n − 1, where n is the number of PAIRSWorking a paired question
Q5 gives the two rows; Q6 gives the differences already summarised.
- 1Define the difference, in writing
Say which member of the pair is subtracted from which, and keep that order for every pair. Q6 defines it for you; in Q5 the choice is yours and must be stated.
- 2Compute the differences and summarise them
One column of numbers, then their mean and standard deviation. The two original columns are not used again.
- 3State the hypothesis about μ_d
H₀ says the mean difference is 0. The direction of the alternative follows from the order you chose in step 1 — which is why step 1 has to be on the page.
- 4One-sample t, on n − 1 degrees of freedom
n is the number of pairs, not the number of measurements. The normality assumption is about the differences, and the question tells you it holds.
Q5 asks whether one treatment reduces the mean, so the alternative is one-sided and the tail values printed for it are upper-tail values at the paired degrees of freedom. Q6 asks instead for an interval for the mean difference and then, in part (b), for what the signs of the two endpoints say about the new shoes. Three readings are possible — both endpoints on the same side of 0, or 0 between them — and each means something different about the direction of the change. Say which of the three you are in and what it implies, in the units the question uses.
Pairing is a choice made before the data exist. It removes the variation between units — vans differ from each other far more than a set of tyres changes any one van — and that is exactly the variation a two-sample standard error would carry. Treating Q5's data as two independent samples is a larger standard error and a weaker conclusion from the same numbers.
A difference of two proportions
Each group supplies a sample proportion p̂ = (count)/(sample size), and the two are compared the same way two means were. What is new is that the interval and the test use different standard errors, and knowing why is the point of Q7 and Q8.
interval: SE = √(p̂₁(1 − p̂₁)/n₁ + p̂₂(1 − p̂₂)/n₂) test: p̂ = (x₁ + x₂)/(n₁ + n₂), SE = √(p̂(1 − p̂)(1/n₁ + 1/n₂))Pool only when H₀ says there is something to pool. A test assumes the two population proportions are equal, so under that assumption every success in either sample estimates the same number and they are best combined into one. An interval assumes no such thing — it is estimating how far apart the two proportions are — so each group keeps its own p̂. Q8 uses the pooled form; Q7 does not.
Q7(a) asks you to check the size condition the question states, and it has to be checked in all four places: successes and failures, in each of the two samples. Then Q7(b) is the interval in the usual shape, (p̂₁ − p̂₂) ± z × SE, with the subtraction in the order the question names.
Q8 is the two-sided test in the five-part form. Two things to keep straight: the "success" here is the event the counts are counting, so define it explicitly before writing any p̂, and the condition is stated in terms of expected counts under H₀ — that is, using the pooled proportion, not the separate ones. The p-value is two-sided, so the printed normal-table values give one tail only and the area on the far side has to be accounted for as well.
The chi-square goodness-of-fit test
Now the data are counts in categories rather than measurements, and the question is whether one set of counts is consistent with a stated model. The statistic measures the total relative discrepancy between what was observed and what the model predicts.
χ² = Σ (O − E)²/E, summed over every categoryQ9's model is the uniform one — "spread evenly" means every category has the same probability, so each expected count is the sample size divided by the number of categories. Q10's model states unequal percentages, so each expected count is the sample size times that category's stated proportion. In both, the expected counts need not be whole numbers, and rounding them to integers before squaring is a real source of error. A quick check before continuing: the expected counts must add to the sample size.
Degrees of freedom count categories, not observations. For a goodness-of-fit test with k categories and a fully specified model, df = k − 1. Q9 and Q10 each print a critical value at the correct df and at least one at a neighbouring df; the neighbour is there for the student who counted the categories and stopped. The test is always upper-tail: only a large χ² is evidence against the model, because a small one means observed and expected agree.
Q10 also asks for a bracketed p-value, from the three upper-tail values printed at its degrees of freedom. Find the two your statistic falls between and report the interval between the corresponding tail areas; if it falls beyond the largest printed value, the p-value is smaller than that value's area, and saying so is the complete answer. Both questions state the expected-count condition — every expected count at least 5 — and both expect it verified against the counts you computed, not asserted.
Contingency tables and expected counts
Q11 does the arithmetic that the test of independence rests on, and does it on its own so that the next two questions can be about the test rather than about the table. Under the hypothesis that the row variable and the column variable are independent, the probability of a cell is the product of its row and column probabilities, and multiplying that by the grand total collapses to one formula.
E = (row total × column total) / grand totalQ11(a) asks for every cell. Work across a row at a time and use the margins as a running check: the expected counts in any row must add to that row's observed total, and the same for every column. That check catches a transposed pair of totals immediately, which no later step will. Q11(b) then asks for the expected-count condition and for the degrees of freedom of the test of independence, which come from the shape of the table alone. Count only the rows and columns of the data: the "Total" row and the "Total" column printed in the table are margins, not categories, and including them is the commonest way this number comes out wrong.
df = (number of rows − 1) × (number of columns − 1)The chi-square test of independence
Q12 and Q13 put Q11's expected counts into Q9's statistic. The hypotheses are now statements about two categorical variables: H₀ says they are independent in the population sampled, and the alternative says only that they are not — there is no direction and no one-sided version of this test.
Q12 asks for the statistic, its degrees of freedom and a decision at the stated level. Set the work out as one row per cell — observed, expected, O − E, (O − E)²/E — and add the last column. Every term is non-negative, so a negative total means a sign was kept where it should have been squared away. Two of the three printed critical values sit at the degrees of freedom the table actually has and one does not; pick by computing df, not by eye.
Q13 is the larger table, in the five-part form, with a bracketed p-value from the printed upper-tail values. The conditions to write down are the random sampling the question describes and the expected-count rule, checked cell by cell on the expected counts you computed. The conclusion sentence has to name both variables in context — study location and programme — because "the variables are associated" is a sentence about symbols, not about students.
Goodness of fit or independence? One variable and a stated model, and the expected counts come from the model: goodness of fit, df = k − 1. Two variables cross-classified in a table, and the expected counts come from the margins: independence, df = (r − 1)(c − 1). Naming which of the two a question is before writing anything settles the hypotheses, the source of the expected counts and the degrees of freedom in one stroke.
What a small p-value does and does not license
Q14 is the interpretation question the whole set has been building towards, and it is worth as much as any computation. A chi-square test on an observational sample has been run and its p-value is given; four statements are offered, and each has to be judged and justified in a line. Three distinctions decide all four.
Three questions to ask of any statement about a test result
Apply them one statement at a time, in this order.
- 1Is it about association or about causation?
A chi-square test compares observed counts with the counts independence would predict. It can report that two variables move together; it cannot say that changing one changes the other, because the households were classified, not assigned.
- 2Does it state the p-value correctly?
A p-value is computed assuming H₀ is true, and it measures how unusual a sample like this one would be in a world where the two variables were unrelated. It is not the probability that H₀ is true, and it is not the probability that the finding is a fluke.
- 3How far does the population reach?
The conclusion extends to the population the sample was drawn from and no further. A statement that carries the result to a different population is claiming something the design cannot support, however small the p-value.
A confounding variable is the reason step 1 exists. Two variables can be firmly associated because a third is related to both — households that install one safety measure differ from those that do not in many other ways at the same time. Naming a plausible third variable is the strongest one-line justification available for a statement about cause, and it is what the question's "give a one-line reason" is asking for.
The synthesis question
Q15 analyses one set of data twice. Part (a) is Q8's procedure: define the success, pool the two counts under H₀, and run the two-sided z test with the printed normal-table value. Part (b) takes the same four counts, arranges them as a two-by-two table — the two wordings as rows, the two outcomes as columns, with the margins you compute from them — and runs Q11's expected-count formula into Q12's statistic.
Build the table before you look for it. The counts the question gives are the successes in each group; the failures are what is left of each group total, and those two cells have to be worked out before the margins mean anything.
Part (c) then asks you to square the z from part (a), set it beside the χ² from part (b), and compare the two decisions at the stated level. Keep several decimal places through both parts rather than rounding at each step — a relationship between two quantities is invisible if rounding has already moved them apart. Each condition the question states belongs to one of the two procedures, so check them separately rather than assuming the second inherits the first.
Preview all 13 pages
Click any page to open the full PDF.
Getting the most out of it
Name the design before you reach for a formula
Three questions, in this order, and they settle everything: are the data measurements or counts in categories? If measurements, are the two groups independent or is each observation paired with another? If independent, are the samples large enough for z or small enough to need the pooled t? Almost every lost mark in this chapter is a correct procedure applied to the wrong design, and no amount of careful arithmetic afterwards recovers it.
Write the standard error on its own line
Every interval in this set is estimate ± critical value × SE and every test statistic is (estimate − hypothesised value)/SE. Computing the SE separately, labelled, before assembling anything means one quantity to check rather than a whole expression, and it makes the interval and the test share the work instead of repeating it.
Get the degrees of freedom from a formula, never from the printed list
n₁ + n₂ − 2 for the pooled t, n − 1 pairs for the paired t, k − 1 for goodness of fit, (r − 1)(c − 1) for independence. The questions print neighbouring critical values on purpose. Compute the number first and then find it; choosing the value that looks familiar is the error those neighbours are there to catch.
Check conditions in writing, before the arithmetic
Randomisation and independence, the sample sizes or the stated normality, the successes and failures in each group, every expected count against its floor. These carry marks in their own right in the five-part form, and a condition that fails changes what you are allowed to conclude — which is not something to discover after the test statistic is on the page.
Finish every question with a sentence about the subject
Vans, seedlings, no-show rates, weekdays, commuting modes. A decision about H₀ is an intermediate result; the answer is what it means for the two groups being compared, said in the units of the problem, with the uncertainty still attached. Every outline for this course asks for interpretation beside computation, and a question that ends at a number is unfinished.
Want the solutions, or something more challenging?
The worksheet, the questions above and every explanation on this page stay free permanently. Three more PDFs exist for this topic — the worked answer key, a harder problem set, and the answer key to that. They come with the University Introductory Statistics Solutions Bundle, beside the unit notes and the unit test, which is what keeps the rest of the series free.
What else exists for Two-Sample Inference and Chi-Square Tests
Three PDFs · 17 pages · all three are in the bundle below.
- Answer key — 4 pages. All 15 questions worked step by step, including the restrictions and the justifications. Not a list of final answers.
- Challenge problems — 9 pages, 9 problems. A separate sheet at exam-plus difficulty covering the same 8 concepts. Harder than anything on the free sheet.
- Challenge answer key — 4 pages. Every challenge problem worked to the same standard, with the checks shown.
- PDF, letter size, print-ready.
The one thing that's for sale
Every University Introductory Statistics topic — the complete Solutions Bundle
One download, one payment, the whole program. For all 9 University Introductory Statistics units: the worksheet, the reference notes, the challenge set, the unit test and every answer key — including this one.
- Worked solutions, not answer lists — every step written out
- Covers the whole year's program at this level
- Less than the price of one hour of tutoring — for the entire year's solutions
Taking Secondary 1 Math as well? The Secondary 1 Math bundle covers all 15 of its units — 106 PDFs, 474 pages — on the same terms.
Taking Secondary 2 Math as well? The Secondary 2 Math bundle covers all 14 of its units — 98 PDFs, 449 pages — on the same terms.
Taking Secondary 3 Math as well? The Secondary 3 Math bundle covers all 11 of its units — 77 PDFs, 367 pages — on the same terms.
Taking Secondary 4 Math as well? The Secondary 4 Math bundle covers all 17 of its units — 122 PDFs, 466 pages — on the same terms.
Taking Secondary 5 Math as well? The Secondary 5 Math bundle covers all 21 of its units — 147 PDFs, 589 pages — on the same terms.
Taking CEGEP Calculus I as well? The CEGEP Calculus I bundle covers all 9 of its units — 72 PDFs, 371 pages — on the same terms.
Taking CEGEP Calculus II as well? The CEGEP Calculus II bundle covers all 8 of its units — 64 PDFs, 335 pages — on the same terms.
Taking CEGEP Linear Algebra as well? The CEGEP Linear Algebra bundle covers all 7 of its units — 56 PDFs, 298 pages — on the same terms.
Taking University Calculus III as well? The University Calculus III bundle covers all 9 of its units — 72 PDFs, 516 pages — on the same terms.
Taking University Linear Algebra as well? The University Linear Algebra bundle covers all 9 of its units — 72 PDFs, 496 pages — on the same terms.
Taking University Differential Equations as well? The University Differential Equations bundle covers all 9 of its units — 36 PDFs, 182 pages — on the same terms.
Taking University Business Math as well? The University Business Math bundle covers all 9 of its units — 36 PDFs, 190 pages — on the same terms.
Taking University Discrete Math as well? The University Discrete Math bundle covers all 9 of its units — 36 PDFs, 155 pages — on the same terms.
Taking AP Calculus AB as well? The AP Calculus AB bundle covers all 8 of its units — 64 PDFs, 527 pages — on the same terms.
Common questions
Is this worksheet really free?
Yes — the questions are on this page to read, and the PDF downloads directly, no email and no account. The one paid item is optional: the complete University Introductory Statistics Solutions Bundle, which covers every set at this level.
Which university courses is this for?
The course codes listed on this page are taken from the public course calendars of universities that teach a first, service-level statistics course. Each course orders and weights the chapters its own way, and some reach a chapter later or not at all — so check the outline for your own section to see where this set falls in your term.
What do I need to know before starting this set?
More of the earlier sets than any other set in this course: the sample mean and standard deviation, the standard normal table, the central limit theorem and the idea of a standard error, and the full structure of a confidence interval and of a hypothesis test. If the one-sample versions in Confidence Intervals and Hypothesis Tests are solid, everything here is the same shape with a difference in place of a single parameter.
How do I tell a paired test from a two-sample test?
Ask whether each observation in one group is matched with a particular observation in the other — the same van measured twice, the same runner timed twice. If it is, the column of differences becomes the data, and a one-sample t procedure runs on it with n − 1 degrees of freedom, where n counts the pairs. If the two groups are separate units chosen independently, the two-sample standard error applies.
When do I use z and when do I use t?
For means, z when both samples are large enough that each sample standard deviation is a reliable stand-in for its population value, and t when they are small and the populations are stated to be normal. For proportions the procedures in this set are always z, because the standard error is built from the proportions themselves rather than from a separately estimated spread.
Why does the test for two proportions pool the samples and the interval not?
Because they assume different things. A test starts from H₀, which says the two population proportions are equal, so the two samples estimate one common value and combining them gives the better estimate of it. An interval makes no such assumption — it is measuring the gap — so each group keeps its own sample proportion.
Why is a chi-square test always upper-tail?
The statistic adds squared discrepancies between observed and expected counts, so it is never negative and it is small exactly when the data agree with the hypothesis. Only a large value is evidence against H₀, which puts the whole rejection region in the upper tail — even when the alternative is two-sided, as it always is for a test of independence.
My p-value comes out as a range rather than a number. Is that wrong?
No — with printed tables it is the expected answer. A t or chi-square table gives a few upper-tail values per row, so locating your statistic between two of them and reporting the range between the corresponding areas is a complete answer. Remember to double the tail area when the alternative is two-sided.
Can teachers use this in class?
Yes. Print and photocopy it for your own classes freely — I just ask that the tutorinmontreal.ca footer stays on the page.
I'm stuck on one question. Can you help?
Yes — through one-on-one tutoring, in Montreal or online. Get in touch to arrange a session, or see the current rates.
← All 9 University Introductory Statistics worksheets · Secondary 1 Math series (15 sheets) → · Secondary 2 Math series (14 sheets) → · Secondary 3 Math series (11 sheets) → · Secondary 4 Math series (17 sheets) → · Secondary 5 Math series (21 sheets) → · CEGEP Calculus I series (9 sheets) → · CEGEP Calculus II series (8 sheets) → · CEGEP Linear Algebra series (7 sheets) → · University Calculus III series (9 sheets) → · University Linear Algebra series (9 sheets) → · University Differential Equations series (9 sheets) → · University Business Math series (9 sheets) → · University Discrete Math series (9 sheets) → · AP Calculus AB series (8 sheets) →




