University Introductory Statistics — Descriptive Statistics Worksheet
The first thing a statistics course asks you to do: take a pile of data and say what it looks like. What kind of variable you are holding, and how it goes into a frequency table; relative and cumulative relative frequency; the shape of a histogram and the class the median falls in; the three measures of centre and how each one answers a stray value; variance and standard deviation by both the definitional and the computational route, and what happens to them when the data are rescaled; and quartiles, the interquartile range, fences and the modified boxplot that shows outliers separately. Have a look on this page, then print the free PDF when you want to write on it.
Practice worksheet — free PDF
No email, no account, no watermark. Teachers: photocopy it for your classes freely. Worked solutions and 6 harder problems come with the University Introductory Statistics bundle.
5 of the 9 questions
Each question targets one named concept from the sheet. Read them here, or print the PDF — it has working space under each one. 5 of the 9 questions are printed below. The other 4 are built on a diagram or a table of values that does not translate to the page, so they are in the free PDF — marked below where they would have come.
-
Q1Types of Variables and Frequency Tables
This question is built around a diagram or a table of values. Open it in the PDF.
-
Q2Histograms and the Shape of a Distribution
This question is built around a diagram or a table of values. Open it in the PDF.
-
Q3Mean, Median and Mode
A customer-support agent logs how many emails she answers in each of one-hour slots: , , , , , , , , , .
- Find the mean, the median and the mode.
- In an eleventh slot she answers emails. Find the new mean and the new median, and say which of the two moved more and why.
-
Q4Mean, Median and Mode
This question is built around a diagram or a table of values. Open it in the PDF.
-
Q5Variance and Standard Deviation
A library counts the books returned through its outdoor drop box on days: , , , , , , , . Treating these days as a sample, find the sample variance and the sample standard deviation .
-
Q6Variance and Standard Deviation
For calls to a software help line, the lengths (in minutes) give and .
- Find , and find using .
- Each call is billed $4.00 plus $1.50 per minute. Find the mean and the standard deviation of the billed amounts.
-
Q7Quartiles, Boxplots and Outliers
Use this convention: order the data and split them into a lower and an upper half, leaving out the median itself when is odd; is the median of the lower half and the median of the upper half. A value is an outlier when it lies more than below or above .
The masses (kg) of checked suitcases on one flight were , , , , , , , , , , , .
- Give the five-number summary and the IQR.
- Find the fences and identify any outliers.
- Where do the whiskers of the modified boxplot end?
-
Q8Quartiles, Boxplots and Outliers
Use the convention that a value is an outlier when it lies more than below or above . The lengths of the episodes of a podcast have five-number summary (in minutes) minimum , , median , , maximum . The five shortest episodes last minutes and the five longest minutes.
- Find the fences and list every outlier.
- Describe the modified boxplot: where the box and the whiskers end.
- About what fraction of the episodes last between and minutes?
-
Q9Synthesis — drawing on several topics in this unit
This question is built around a diagram or a table of values. Open it in the PDF.
The 6 challenge problems for this topic are a separate, paid sheet and are not reproduced here.
What does this set assume? This is the opening set of the course and it leans on no set before it. It assumes Secondary 5 mathematics only — arithmetic with fractions, decimals and percentages, summation notation read as an instruction to add a column, and the frequency tables and boxplots of the Quebec secondary program. Nothing here is calculus-based: no question in this set, or anywhere in this course, asks for a derivative or an integral. Everything is done by hand with a non-programmable calculator, so there is no software output to read and none to produce, and no numerical coefficient of skewness or kurtosis — shape is read off a histogram or a boxplot and described in words. Probability is not used at all; it starts in the next set, Probability Rules and Conditional Probability. The empirical 68–95–99.7 rule, z-scores and standardizing belong to Continuous Distributions and the Normal Model, and the sample mean as a random quantity with a distribution of its own to Sampling Distributions and the Central Limit Theorem, which is where the letters x̄ and s stop being summaries of a list and start being estimates. Everything inferential — Confidence Intervals, Hypothesis Tests, Two-Sample Inference and Chi-Square Tests — comes later, and the scatterplot, the correlation coefficient r and the fitted line are kept together in the final set, Simple Linear Regression and Correlation. Here nothing is estimated, tested or predicted: the data in front of you are the whole story.
Which course is this for? In the public course calendars of Montreal universities, this material is part of the courses numbered MATH 203, MAST 333, STT1700, MAT2080, MAT1185, MATH 10603 and MAT350. Each course orders and weights the topics its own way, so check your own outline for what your exam covers. Which sets match your course.
How to do every concept on this sheet
This is the part worksheet sites usually leave out. Below is the actual reasoning behind each group of questions — not a full solution set, the decisions that get you to one. Read it before you start, or after you get stuck.
Types of variables
Before any arithmetic, decide what kind of thing each column holds — because the kind decides which summaries are even legal. A mean of a set of labels is meaningless however neatly the labels are written as numerals.
Classifying one variable
Two questions, asked in this order, for each of the five variables in Q1(a).
- 1Is the value a quantity or a category?
A quantity answers "how many" or "how much" and arithmetic on it means something. A category answers "which kind". A numeral can be a category: ask whether an average of those numerals would describe anything at all.
- 2aIf it is a category: does it have a natural order?
Categories that rank from low to high are ordinal; categories with no order among them are nominal. The test is whether swapping two of the names would change the meaning of the list.
- 2bIf it is a quantity: is it counted or measured?
A count that can only land on whole numbers is discrete. A measurement that could in principle take any value in an interval — limited only by the instrument — is continuous.
"It is written with digits" is not a classification. Two of the columns in Q1(a) are numerals and they do not both behave the same way. Apply the first question of the tree to each one separately and write a sentence of justification beside each answer; the justification is what the question is really asking for.
Frequency, relative frequency and cumulative relative frequency
A frequency table is three rows once it is finished. The frequency f counts how many observations took each value; the relative frequency is f/n, the share of the data at that value; and the cumulative relative frequency adds the relative frequencies from the left, accumulating as it goes.
relative frequency = f/n, n = Σf, Σ(f/n) = 1Q1(b) asks for both extra rows on a table of counts, and Q2(a) asks for the same two rows on a table of classes. Work with the same n throughout — add the frequency row first and use that total, rather than the total printed in the sentence above the table, so that a misread figure shows up immediately.
The last cumulative entry is 1, and that is the check. If the row does not finish at 1 (or at n, when you accumulate the raw counts), a relative frequency was rounded too hard or a column was skipped. Keep three or four decimals in the row and round only when you write the final sentence.
Q1(c) then reads two proportions off the finished table, and the wording is the whole exercise. "At most 5" means 5 and everything below it, so it is a cumulative entry read directly. "At least 6" means 6 and everything above it, which the cumulative row does not hold — take it as 1 minus the cumulative entry just below, or add the remaining relative frequencies and check that the two routes agree.
"At most" and "at least" both include the value named. The commonest slip on Q1(c) and Q2(c) is dropping the boundary value out of one of the two answers, which makes the pair of proportions fail to account for the whole data set.
Histograms and the shape of a distribution
Q2 works from classes rather than single values, and the question states the endpoint convention it wants: each class includes its left endpoint, and the last class is closed at the top so that nothing falls outside the table. Use the stated convention and do not import a different one — a value sitting exactly on a boundary has to land in exactly one class.
Q2(b) asks for three things about the picture, and each has a definite meaning:
Describing a histogram
Shape in three named parts.
- 1The modal class
The class with the greatest frequency — the tallest bar, when the classes all have the same width, as they do here. Name the class, not the bar height.
- 2Symmetric or skewed
Fold the picture at its peak. If the two sides are near mirror images it is roughly symmetric; if one side runs out much further than the other it is skewed.
- 3The direction of the tail
A distribution is named for where its long thin tail points, not for where the tall bars are. The tail is the stretch of low frequencies trailing away from the bulk.
Skew is named for the tail, never for the crowd. Nearly every wrong answer to Q2(b) names the direction the bars pile up in. Find the low bars first, say which end they are at, and name the skew after them.
Q2(c) sends you back to the cumulative row for two different jobs. To find the class holding the median, run along the cumulative relative frequencies until they first reach 0.5 — the class where that happens is the one the middle observation falls in, and with grouped data that class is the whole answer, since the individual values are no longer visible. The proportion at or above a class boundary is then the sum of the relative frequencies of the classes above it, or 1 minus the cumulative entry at that boundary; the second route is shorter and the first is the check on it.
Mean, median and mode
Three measures of centre, three different questions answered. The mean is the balance point and uses the size of every value; the median is the positional middle and uses only the order; the mode is the value that occurs most often and is the only one of the three that also works on categories.
x̄ = Σx/n (raw list), x̄ = Σfx/Σf (frequency table)Q3 gives ten values as a plain list. Order them before looking for the median — when n is even the median sits halfway between the two middle entries of the sorted list, and when n is odd it is the single middle entry. The mode needs no ordering, only a tally, and a list can have no mode at all or more than one.
Resistance is the point of Q3(b). An eleventh value far above the others is added, and you are asked which measure moved more and why. Do not answer from the two numbers alone — the reason lives in the definitions. The mean shares the new value out across all the observations, so its size enters the calculation; the median only asks where the new value sits relative to the others, so it can shift at most to the neighbouring position. Recompute both, then write the explanation in terms of what each formula uses. That second sentence is the mark.
Q4 is the same three measures read out of a frequency table, which is a different piece of arithmetic for each. The mean is Σfx/Σf: multiply each value by its frequency, add the products, divide by the total count — not by the number of distinct values, which is the standard error here. The median is positional, so build a cumulative count and find which value the middle observation lands on; with n even, locate both middle positions and check whether they fall on the same value. The mode is the value carrying the largest frequency, which you read off the table with no arithmetic at all.
Never take the median of the frequency row. In Q4 and again in Q9 the frequencies are counts of boxes or of cars, not data values. The data values are the row of categories along the top, each one repeated as many times as its frequency says.
Variance and standard deviation
Spread is measured from the mean outwards. The deviation of a value is x − x̄; the deviations always add to zero, so they are squared before being averaged, and the sample version divides by n − 1 rather than n. The standard deviation s is the square root of that, which puts the answer back into the units of the data.
s² = Σ(x − x̄)²/(n − 1), s = √(s²)Q5 wants the definitional route on eight values, and it is worth laying out as a table with a column for x, a column for x − x̄ and a column for (x − x̄)². Add the middle column before going further: it must come to zero, and if it does not the mean is wrong and nothing after it can be right.
Carry the mean exactly. A mean rounded to one decimal before squaring the deviations drags a visible error into s. Keep it as a fraction, or keep every digit your calculator shows, and round only the final s.
Q6 hands you Σx and Σx² for twenty-five observations and no list at all, which is exactly the situation the computational formula exists for. It is the same quantity as the definitional one, rearranged so that it needs only the two sums.
s² = (Σx² − (Σx)²/n)/(n − 1)(Σx)² and Σx² are different numbers. One squares the total, the other totals the squares, and mixing them is the single commonest error in Q6(a). Write the two symbols side by side and label which given number is which before substituting anything.
Q6(b) is a linear transformation: each observation is charged a fixed amount plus a fixed rate per minute, so the billed amount is a + bx for constants a and b read out of the sentence. The two summaries do not behave alike under it. Adding a constant shifts every value by the same amount, so no two observations move apart and the spread is untouched while the centre slides; multiplying by a constant stretches the distances between observations as well, so it scales both summaries.
mean of a + bx = a + b·x̄, standard deviation of a + bx = |b|·sSo there is no need to rebuild anything from the data: identify a and b, apply the two rules, and state the units of the results. Doing it this way is also the check — a fixed charge that changed the standard deviation would mean the wrong rule was applied.
Quartiles, boxplots and outliers
Q7 opens by stating the convention it wants, and that matters: there are several defensible ways to split a data set into quarters and they can disagree by a whole value. Use the one written into the question. Order the data, cut the ordered list in half — leaving the median itself out when n is odd — and take Q₁ as the median of the lower half and Q₃ as the median of the upper half. (Those subscripted Q's are quartiles; Q7 and Q8 are question numbers.)
IQR = Q₃ − Q₁, lower fence = Q₁ − 1.5 × IQR, upper fence = Q₃ + 1.5 × IQRFrom a list to a modified boxplot
The order of operations Q7 asks for, part by part.
- 1Order, then summarize
The five-number summary is minimum, Q₁, median, Q₃, maximum, in that order. Every one of the five is positional, so none of them can be found before the list is sorted.
- 2Fences, then outliers
Compute the IQR, then the two fences. A value is an outlier when it lies more than 1.5 × IQR beyond a quartile — strictly more, so a value sitting exactly on a fence is not one.
- 3Whiskers last
On a modified boxplot the box runs from Q₁ to Q₃ with the median marked inside it, each whisker stops at the most extreme observation that is still inside its fence, and every outlier is plotted as its own point beyond the whisker.
A whisker ends on a data value, not on a fence. The fence is only a threshold used to decide which observations count as outliers; it is never drawn. Q7(c) and Q8(b) both ask where the whiskers end, and the answer is always an observation from the data set.
Q8 reverses the supply of information: the five-number summary is given for sixty observations you never see, along with the five smallest and the five largest. Everything you need is still there. The quartiles give the IQR and the fences; then the outliers can only be among the listed extremes, because every unlisted value lies between the fifth smallest and the fifth largest. Compare each listed extreme with its fence and stop when you reach one that is inside — the ones beyond it are inside too, since the list is in order.
Q8(c) is a definition question, not a calculation. Look at the two minutes figures the question names and compare them with the summary you were given. Once you see which two numbers of the five they are, the proportion follows from what the quartiles mean about how the observations are divided up, and no arithmetic on the sixty values is possible or needed.
The synthesis question
Q9 runs the whole sheet on one small frequency table. Part (a) is the classification of Q1 again — apply the same two questions from the tree above. Parts (b) and (c) are the frequency-table versions of the mean, the median and the standard deviation, so Σfx and Σfx² are the columns to build; the formula printed in the question is the computational one from Q6 with f carried through, and the trap from Q4 applies here too, since the values being summarized are the counts of occupants and not the counts of cars.
s² = (Σfx² − (Σfx)²/n)/(n − 1), n = ΣfPart (d) is the one that ties the set together. Describe the shape from the table as you described it from the histogram in Q2, then check it against the relation between the two measures of centre you have just computed. The general rule is that a long tail pulls the mean towards it while the median stays put, so in a right-skewed distribution the mean sits above the median, in a left-skewed one below it, and in a roughly symmetric one the two are close. State the rule, state which side your mean fell on, and say whether the two descriptions agree — the agreement is the answer, not either half on its own.
Preview all 5 pages
Click any page to open the full PDF.
Getting the most out of it
Sort first, always
The median, the quartiles, the five-number summary and both boxplots are positional: they are statements about where a value sits in the ordered list, and every one of them is wrong if the list is not sorted. Write the ordered list out once, on its own line, and work from it for the rest of the question.
Keep the whole number until the last line
A mean rounded before it is used again is the most common source of a nearly-right standard deviation. Carry every digit through the intermediate steps and round once, at the end, to a sensible number of decimals — one more than the data usually does.
Make each table prove itself
Relative frequencies add to 1; a cumulative row ends at 1; deviations from the mean add to zero; Σf is the same n you divided by. Each of these takes ten seconds and each one catches a whole class of arithmetic slip before it propagates into the next part of the question.
Say what every number means, with its units
This course asks for interpretation beside computation, so a part that ends at a bare number is unfinished. A standard deviation is a typical distance from the mean in the units of the data; a variance is in squared units; a relative frequency is a share of the whole. Write the sentence, not just the figure.
Describe a distribution in the same three words every time
Shape, centre, spread — and outliers if there are any. Practising that order on Q2, Q7 and Q9 means that when a later set asks whether a procedure is safe to use, you are already looking at the three things that decide it.
Want the solutions, or something more challenging?
The worksheet, the questions above and every explanation on this page stay free permanently. Three more PDFs exist for this topic — the worked answer key, a harder problem set, and the answer key to that. They come with the University Introductory Statistics Solutions Bundle, beside the unit notes and the unit test, which is what keeps the rest of the series free.
What else exists for Descriptive Statistics
Three PDFs · 11 pages · all three are in the bundle below.
- Answer key — 2 pages. All 9 questions worked step by step, including the restrictions and the justifications. Not a list of final answers.
- Challenge problems — 6 pages, 6 problems. A separate sheet at exam-plus difficulty covering the same 5 concepts. Harder than anything on the free sheet.
- Challenge answer key — 3 pages. Every challenge problem worked to the same standard, with the checks shown.
- PDF, letter size, print-ready.
The one thing that's for sale
Every University Introductory Statistics topic — the complete Solutions Bundle
One download, one payment, the whole program. For all 9 University Introductory Statistics units: the worksheet, the reference notes, the challenge set, the unit test and every answer key — including this one.
- Worked solutions, not answer lists — every step written out
- Covers the whole year's program at this level
- Less than the price of one hour of tutoring — for the entire year's solutions
Taking Secondary 1 Math as well? The Secondary 1 Math bundle covers all 15 of its units — 106 PDFs, 474 pages — on the same terms.
Taking Secondary 2 Math as well? The Secondary 2 Math bundle covers all 14 of its units — 98 PDFs, 449 pages — on the same terms.
Taking Secondary 3 Math as well? The Secondary 3 Math bundle covers all 11 of its units — 77 PDFs, 367 pages — on the same terms.
Taking Secondary 4 Math as well? The Secondary 4 Math bundle covers all 17 of its units — 122 PDFs, 466 pages — on the same terms.
Taking Secondary 5 Math as well? The Secondary 5 Math bundle covers all 21 of its units — 147 PDFs, 589 pages — on the same terms.
Taking CEGEP Calculus I as well? The CEGEP Calculus I bundle covers all 9 of its units — 72 PDFs, 371 pages — on the same terms.
Taking CEGEP Calculus II as well? The CEGEP Calculus II bundle covers all 8 of its units — 64 PDFs, 335 pages — on the same terms.
Taking CEGEP Linear Algebra as well? The CEGEP Linear Algebra bundle covers all 7 of its units — 56 PDFs, 298 pages — on the same terms.
Taking University Calculus III as well? The University Calculus III bundle covers all 9 of its units — 72 PDFs, 516 pages — on the same terms.
Taking University Linear Algebra as well? The University Linear Algebra bundle covers all 9 of its units — 72 PDFs, 496 pages — on the same terms.
Taking University Differential Equations as well? The University Differential Equations bundle covers all 9 of its units — 36 PDFs, 182 pages — on the same terms.
Taking University Business Math as well? The University Business Math bundle covers all 9 of its units — 36 PDFs, 190 pages — on the same terms.
Taking University Discrete Math as well? The University Discrete Math bundle covers all 9 of its units — 36 PDFs, 155 pages — on the same terms.
Taking AP Calculus AB as well? The AP Calculus AB bundle covers all 8 of its units — 64 PDFs, 527 pages — on the same terms.
Common questions
Is this worksheet really free?
Yes — the questions are on this page to read, and the PDF downloads directly, no email and no account. The one paid item is optional: the complete University Introductory Statistics Solutions Bundle, which covers every set at this level.
Which university courses is this for?
The course codes listed on this page are taken from the public course calendars of universities that teach a first, service-level statistics course. Each course orders and weights the chapters its own way, and some reach a chapter later or not at all — so check the outline for your own section to see where this set falls in your term.
What do I need to know before starting this set?
Secondary 5 mathematics only: arithmetic with fractions, decimals and percentages, and summation notation read as an instruction to add a column. This is the first set of the course, so nothing from a later set is assumed, and no calculus is used here or anywhere else in the course.
Do I need a computer or statistical software?
No. Every question is done by hand with a non-programmable calculator, and nothing on this sheet asks you to produce or read software output. That is deliberate — it keeps the arithmetic visible, which is what is being marked.
Why does the sample variance divide by n − 1 instead of n?
Because the deviations are measured from x̄, which was itself computed from the same data. That ties the deviations together — they must add to zero — so only n − 1 of them are free to vary, and dividing by n would make the result too small on average. The full argument comes in Sampling Distributions and the Central Limit Theorem; here it is enough to know that a sample uses n − 1 and to do it consistently.
Which quartile convention should I use?
The one stated in the question. Methods differ over whether the median is included in each half and over how to interpolate, and they can give different values on the same small data set, so the questions that need a convention state theirs. In a course of your own, use the one your instructor demonstrated and say which you used.
Should I report the mean or the median?
It depends on the shape, which is why shape is described before centre. The mean uses the size of every observation, so a long tail or a lone extreme value pulls it; the median only uses order, so it stays where the bulk of the data is. For a roughly symmetric set with no outliers they agree closely and the mean carries more information.
Where are z-scores and the normal distribution?
Later in the course. This set describes a data set exactly as it stands. Standardizing, the 68–95–99.7 rule and the normal table belong to Continuous Distributions and the Normal Model, and the sample mean as a quantity with a distribution of its own to Sampling Distributions and the Central Limit Theorem.
Can teachers use this in class?
Yes. Print and photocopy it for your own classes freely — I just ask that the tutorinmontreal.ca footer stays on the page.
I'm stuck on one question. Can you help?
Yes — through one-on-one tutoring, in Montreal or online. Get in touch to arrange a session, or see the current rates.
← All 9 University Introductory Statistics worksheets · Secondary 1 Math series (15 sheets) → · Secondary 2 Math series (14 sheets) → · Secondary 3 Math series (11 sheets) → · Secondary 4 Math series (17 sheets) → · Secondary 5 Math series (21 sheets) → · CEGEP Calculus I series (9 sheets) → · CEGEP Calculus II series (8 sheets) → · CEGEP Linear Algebra series (7 sheets) → · University Calculus III series (9 sheets) → · University Linear Algebra series (9 sheets) → · University Differential Equations series (9 sheets) → · University Business Math series (9 sheets) → · University Discrete Math series (9 sheets) → · AP Calculus AB series (8 sheets) →



