Sign In
This lesson provides a comprehensive look at different types of data distributions, specifically through the lens of histograms. It discusses how data can be skewed, meaning it leans more to one side, or symmetric, where it is evenly spread. These characteristics are crucial for anyone dealing with data, from statisticians to business analysts. For example, skewed data might indicate customer preferences in a retail setting, while symmetric data could suggest a balanced ecosystem in environmental studies. The information is presented in a way that makes these complex topics accessible and easy to learn.
Show less Show more expand_more| Student Learning Objectives: |
|---|
|
| | 12 Theory slides |
| | 9 Exercises - Grade E - A |
| | Each lesson is meant to take 1-2 classroom sessions |
Tadeo and Ramsha are foodies and love to explore the restaurants that pop up in their neighborhood. They have recorded data about the average price of the main dishes in each restaurant using a table of values.
| Average Main Dish Price (Dollars) | ||||
|---|---|---|---|---|
| 10.12 | 9.29 | 8.29 | 9.78 | 10.69 |
| 9.68 | 12.09 | 8.94 | 10.81 | 8.62 |
| 11.39 | 12.62 | 8.71 | 10.74 | 10.52 |
| 10.77 | 10.15 | 9.18 | 8.45 | 9.52 |
| 11.89 | 9.77 | 9.44 | 13.24 | 11.01 |
| 10.62 | 9.38 | 12.15 | 9.68 | 9.60 |
| 10.32 | 11.31 | 11.41 | 8.62 | 9.27 |
| 10.96 | 9.18 | 10.28 | 10.71 | 10.02 |
They would like to draw some conclusions from this data. However, they are not entirely sure how to proceed with this task. Find the following information to help these curious connoisseurs!
Consider the following histograms.
Which histogram best describes the data?
Select the option that describes the distribution of the data.
Which measures of center and variation best describe the data?
A frequency distribution, sometimes called a histogram distribution, is a representation that displays the number of observations within a given interval or category. It is used to show the empirical or theoretical frequency of occurrence of each possible value in a data set, often recorded in a frequency table. Frequency distributions of categorical data are typically presented using a bar graph.
In the case of numerical data, the graphical representation of a frequency distribution is called a histogram.
A symmetric frequency distribution is a distribution in which the data are evenly distributed around the mean and the bars on each side of the middle bar are approximately the same height.
A skewed frequency distribution is a distribution in which the data is not spread evenly — rather, the data is clustered at one end. In this case, the mean and the median are not equal, causing the data set to be skewed. A skewed distribution is neither symmetric nor normal. In general, there are two types of skewed frequency distributions.
| Skewed Distribution | Description |
|---|---|
| Skewed Left / Negatively Skewed | The distribution has a long left tail and the median is greater than the mean. |
| Skewed Right / Positively Skewed | The distribution has a long right tail and the median is less than the mean. |
The difference between normal and skewed distributions can be visualized in the following applet.
Tadeo and Ramsha are having fun learning about frequency distributions. Last weekend, two of the most exciting cricket games this season took place. Tadeo and Ramsha recorded the runs scored by the 22 players in each match. The data for Games 1 and 2 are shown in the table.
| Game 1 | Game 2 | ||||||
|---|---|---|---|---|---|---|---|
| 32 | 21 | 27 | 46 | 114 | 87 | 96 | 92 |
| 9 | 16 | 19 | 19 | 101 | 111 | 80 | 106 |
| 40 | 28 | 42 | 36 | 85 | 112 | 117 | 94 |
| 11 | 38 | 23 | 28 | 62 | 43 | 106 | 66 |
| 8 | 18 | 26 | 59 | 104 | 51 | 76 | 91 |
| 62 | 40 | 111 | 78 | ||||
Observe the following histograms.
Which histogram represents the data of Game 1?
Which of the following statements are true about the data of Game 1?
Consider the following histograms.
Which histogram represents the data of Game 2?
Which of the following statements are true about the data of Game 2?
Begin by making a frequency table of the data.
Observe where the tail of the distribution extends.
Make a histogram of the data using a frequency table.
Which measures of center and variation should be used when the distribution of the data is skewed?
Notice that the given histograms consist of seven intervals. Considering this information, create a frequency table with the seven intervals, beginning with 0-9, to display the data for Game 1.
| Cricket Runs Scored in Game 1 | |
|---|---|
| Number of Runs Scored | Frequency |
| 0-9 | 2 |
| 10-19 | 5 |
| 20-29 | 6 |
| 30-39 | 3 |
| 40-49 | 4 |
| 50-59 | 1 |
| 60-69 | 1 |
The data can be displayed in a histogram by using this frequency table. The horizontal axis will be the Number of Runs Scored
and the vertical axis the Frequency.
Then the bars will be drawn to represent the frequency of each interval.
Note that this corresponds to option A.
The tail of the histogram extends to the right and most of the data is on the left. Therefore, the distribution of the data is skewed right. In a skewed distribution, the median and five-number summary best describe the center and variation of the data, respectively. As such, two statements apply to the data set of Game 1.
Similar to Part A, in order to identify which histogram is the one that describes the data of Game 2, a frequency table of the data will be created, this time using eight intervals.
| Cricket Runs Scored in Game 2 | |
|---|---|
| Number of Runs Scored | Frequency |
| 40-49 | 1 |
| 50-59 | 1 |
| 60-69 | 2 |
| 70-79 | 2 |
| 80-89 | 3 |
| 90-99 | 4 |
| 100-109 | 4 |
| 110-119 | 5 |
Using this table, the histogram of the data can be now created.
Note that this corresponds to option D.
In this case, the tail of the histogram extends to the left and most of the data is on the right. Therefore, the data is skewed left. Additionally, the median and five-number summary best describe the center and variation of the data, respectively.
Tadeo and Ramsha are amazed by how data is presented everywhere and how knowing the distribution of the data helps to interpret that data. Besides cricket, they also love watching NFL games and are fans of Peyton Manning, the famous quarterback who retired at age 40. Now they want to analyze the retirement ages of NFL players by collecting some data.
| Retirement Age of NFL Players | |
|---|---|
| Age | Frequency |
| 25-26 | 33 |
| 27-28 | 67 |
| 29-30 | 93 |
| 31-32 | 109 |
| 33-34 | 127 |
| 35-36 | 114 |
| 37-38 | 80 |
| 39-40 | 59 |
| 41-42 | 43 |
Based on this frequency table, Tadeo and Ramsha created a histogram and in order to draw some conclusions about the data.
Consider the following histograms.
Which of these histograms could represent the given data set?
Which measures of center and variation best represent the data?
Which of the following statements is most likely true about the retirement age of NFL players?
Use the frequency table to display the data in a histogram.
Which measures of center and variation best describe a symmetric distribution?
Which interval has the highest bar?
The data can be displayed in a histogram by using the given frequency table. The vertical axis will be the Frequency
and the horizontal axis the Age.
Next, bars will be plotted to represent the frequency of data points falling in each interval.
Note that this histogram corresponds to option B.
The data on the left side of the distribution is nearly a mirror image of the data on the right side of the distribution, which means that the distribution is symmetric.
In a symmetric frequency distribution, data are distributed evenly around the mean and the bars on each side of the middle bar are about the same height.
It can be seen that the mean of the distribution is in the interval of 33-34. This means that a typical NFL player is much more likely to retire at around age 33 or 34.
There are special distributions that less common than skewed and symmetric distributions. These distributions may appear in situations such as an experiment where each event has the same probability or a sample taken from two separate populations. These are the uniform and bimodal distributions.
A uniform frequency distribution, sometimes called a flat distribution, is a type of distribution where all the bars are about the same height. This type of distribution arises in scenarios where all the possible outcomes are equally likely. A uniform distribution is also symmetric.
As an example, the possible outcomes of rolling a fair six-sided die are 1, 2, 3, 4, 5, and 6, and they each have an equal probability of occurring. The following applet simulates rolling a die 100 times and records the frequency of each outcome.
A bimodal distribution is a data distribution with a range of values near two individual values or two intervals, separating the data into two clusters. This causes the histogram of the data to have two peaks. The mean and the median of a bimodal distribution are near the center of the distribution.
The given distribution indicates that the sampling was likely made from two different populations. The term bimodal refers to the peaks of the distribution, which differs from the mode when intervals are used to make the data display. It is worth mentioning that a bimodal distribution whose bars are about the same height on each side of the peaks is also symmetric.
Consider a histogram that shows the attendance per hour at a local restaurant.
Tadeo and Ramsha want to explore different types of data sets to see what kind of distributions they can draw. The first data set was collected from an experiment by spinning a spinner with six equal sections 1000 times.
| Color | Frequency |
|---|---|
| Red | 164 |
| Blue | 168 |
| Yellow | 168 |
| Pink | 166 |
| Green | 165 |
| Orange | 169 |
Tadeo collected the other data set from a survey about the exam scores of their classmates.
| Exam Score | Frequency |
|---|---|
| 70-71 | 3 |
| 72-73 | 8 |
| 74-75 | 11 |
| 76-77 | 9 |
| 78-79 | 4 |
| 80-81 | 2 |
| 82-83 | 2 |
| 84-85 | 4 |
| 86-87 | 8 |
| 88-89 | 11 |
| 90-91 | 9 |
| 92-93 | 7 |
| 94-95 | 2 |
They now have two data sets and would like to display the data in histograms to analyze them.
Consider the following histograms.
Select the histogram that represents the data set for the spinner experiment.
Select the option(s) that best describes the distribution of the data of the spinner.
Consider the following histograms representing test scores.
Which histogram represents the given data for the exam scores?
Select the option(s) that best describes the distribution of the data about the exam scores.
Begin by drawing a histogram of the data.
How different are the bars of the histogram of the data?
Use the frequency table to draw a histogram of the data.
How many peaks
does the histogram have? Draw a vertical line through the middle of the distribution.
The data set of the spinner experiment can be displayed in a histogram. Label the horizontal axis Color
and the vertical axis Frequency.
Then draw the bars to represent the frequency of each color outcome.
Notice that this corresponds to option B.
Looking at the histogram, it can be seen that the bars are all approximately the same height.
Similarly to Part A, the data modeling the class's exam scores can be displayed in a histogram. The horizontal axis will be the Test Score
and the vertical axis will be the Frequency.
The histogram of the exam scores has been drawn. The right graph is given in option C.
Notice that the histogram has two peaks.
Additionally, these two peaks split data into two clusters. This means that the data follows a bimodal distribution. Moreover, suppose a vertical line is drawn around the halfway line of the distribution. The data on the left is an approximate mirror image of the data on the right.
This means that the distribution is bimodal and symmetric. A possible explanation for this data is that it comes from two groups, one group of students who did not study for the exam (the first peak
on the left) and one group that did study for the exam (the second peak
on the right).
A box plot is another data display that allows one to see the shape of a frequency distribution. The length of the whiskers
and the position of the median tell whether the distribution is skewed or symmetric.
The following applet shows each of these three frequency distributions using box plots.
During their fantastic journey exploring the restaurants in their neighborhood, Tadeo and Ramsha found a fabulous Italian restaurant. While eating their food, they observed the people who entered the restaurant.
The two are curious about the average age of people eating at this restaurant. Therefore, they decide to collect data on the ages of people who enter the restaurant during a typical day.
| Ages of People Who Enter the Italian Restaurant on a Typical Day | |||||
|---|---|---|---|---|---|
| 15 | 53 | 55 | 60 | 38 | 56 |
| 62 | 14 | 44 | 24 | 32 | 10 |
| 42 | 54 | 47 | 67 | 60 | 50 |
| 61 | 30 | 30 | 62 | 62 | 65 |
| 56 | 52 | 35 | 25 | 34 | 32 |
They now want to draw some conclusions from this data set by displaying it in a box plot.
Find the five-number summary and match each description with its corresponding value.
Consider the following box plots.
Which of the box plots represents the given data set?
Which measures of center and variation best represent the data?
Which of the following statements is most likely true about the people who enter the Italian restaurant?
Begin by ordering the data values from least to greatest.
Use the five-number summary to make the box plot.
Decide whether the data is skewed or symmetric.
Each whisker represents 25 % of the data. The box represents 50 % of the data.
To find the five-number summary of the data set, begin by ordering the data values from least to greatest.
10 14 15 24 25 30 30 32 32 34 [0.5em] 35 38 42 44 47 50 52 53 54 55 [0.5em] 56 56 60 60 61 62 62 62 65 67 Notice that the minimum value is 10 and the maximum value is 67. Additionally, the number of data points is 30, an even number. Therefore, the median of the data set will be given by the mean of the middle numbers 47 and 50. Median: 47+ 50/2= 48.5 Finally, the first quartile, or the median of the lower half, is 32, and the third quartile is 60. The following table summarizes this information. Each description is matched with its corresponding value.
| Five Number Summary | |
|---|---|
| Minimum Value | 10 |
| First Quartile | 32 |
| Median | 48.5 |
| Third Quartile | 60 |
| Maximum Value | 67 |
The five-number summary found in Part A can be used to draw the box plot of the data set. Start by drawing a number line that includes the minimum and maximum values of the data. Next, graph points above the number line for the five-number summary.
Now, draw a box from the first quartile to the third quartile. Then, draw a line through the median and the whiskers from the box to the minimum and maximum values.
Notice that this plot corresponds to option A.
In the box-plot drawn in the previous part, notice that the left whisker is longer than the right and that the median is closer to the right whisker than it is to the left.
In a box plot, each whisker represents 25 % of the data and the box represents the middle 50 % of the data. With this in mind, the following facts about the distribution of the data set can be determined.
Therefore, the option that states that 50 % of the people who enter the Italian restaurant in a regular day are between 32 and 60 years old is the right one.
In addition to exploring restaurants, Tadeo and Ramsha spend a lot of time playing video games together. While hanging out weekend, they discussed whether boys or girls spend more time on video games on weekends. To investigate this situation, the two collected some data from their classmates at North High School and displayed it in a double-histogram.
Select the statements that are right about data set of the responses from the girls.
Which of the following statements are true about the data set of responses from the boys?
If one student is randomly selected from each group, which is more likely to spend more time on video games on weekends?
Begin by identifying the distribution shape of the data.
Identify the shape of the distribution.
Use the distribution of each data set to identify a typical value.
The shape of the distribution will be described to determine which of the given sentences are right about the data set showing how much time the girls spend playing video games.
This gives the two correct statements about the girls' data set.
Now, consider the histogram of the data set that represents the responses given by the boys.
Notice that in this case, the tail of the distribution extends to the left and that most of the data is on the right side of the histogram. Therefore, the distribution of the data skewed left, so the five-number summary best describes the center and spread of the data.
It was previously identified that the mean best describes the center of the data set for the girls and the median for the data set for the boys.
| Appropriate Measures of Center | |
|---|---|
| Girls' Data Set | Boys' Data Set |
| Mean | Median |
To identify who is more likely to spend more time on video games, compare these measures of center. Since the girls' distribution is symmetric, its mean is probably in the 6.1-8 interval, the center of the distribution.
Conversely, because in a skewed distribution, the median is righter to the center and closer to the peaks of the distribution, it is probably in the 8.1-10 interval.
Comparing these values, notice that the median of the boys is greater than the mean of the girls. Girls' Mean & & Boys' Median 6.1-8 & & 8.1-10 This means it is more likely that a boy spends more time playing video games on the weekends than a girl does.
Ramsha and Tadeo took a survey of their classmates about the number of hours spent playing video games on the weekends. These results made Ramsha more interested in finding other everyday situations with remarkable differences due to gender. She decided to search the web for similar examples.
During her investigation, Ramsha found a peculiar table showing the results of a survey comparing the amount of money that men and women usually spend on clothes per month.
| Women | Men | |
|---|---|---|
| Survey Size | 100 | 100 |
| Minimum | $18 | $8 |
| Maximum | $60 | $28 |
| 1^(st) Quartile | $30 | $14 |
| Median | $34 | $18 |
| 3^(rd) Quartile | $40 | $22 |
| Mean | $36 | $18 |
| Standard Deviation | $8 | $4 |
Consider the following double box plots.
Which of the given double box plots shows the results of the survey Ramsha found online?
Which statement is true about the given data?
How many of the women surveyed are expected to spend between $30 and $40 on clothes per month?
Use the five-number summary of each data set to make a double box plot.
Begin by identifying each data set distribution.
Each whisker represents 25 % of the data. The box represents 50 % of the data.
This data can be represented with a double box plot to identify which of the given graphs accurately represents the data set. First, draw a number linethat includes the minimum and maximum values of each gender's data set. Next, plot points above the number line for the given values of the five-number summary.
Next, draw the box for each plot using the first and third quartiles. Finally, draw a line through the median and the whiskers from the box to the minimum and maximum values of each data set.
Notice that this corresponds to the box-plot in option D.
In order to identify which of the given statements is correct, compare the center and spread of the data sets. Note that for the women's data set, the right whisker is longer than the left one and that the median is closer to the left whisker. This means that the data is skewed right, and the median best describes the center of the data.
Median of Women's Data Set: $34 Conversely, for the men's data set, the whiskers are approximately equal and the median falls in the middle of the box. Therefore, this data is modeled by a symmetric distribution, and the mean best describes the center of the data. Mean of Men's Data Set: $18 Notice that the median amount of money spent by women on clothes each month is almost twice the mean amount of money spent on clothes by men. Recall that the range of a data set is given by the difference of the minimum and maximum values. Using this information, compare the range and standard deviation of the data sets.
| Standard Deviation | Interquartile Range | |
|---|---|---|
| Women | $8 | 60-18=$42 |
| Men | $4 | 28-8=$20 |
Both the standard deviation and the interquartile range are greater for women. This means that there is more variability in the amount of money spent by women.
To calculate how many of the women surveyed are expected to spend between $30 and $40 on clothes per month, consider that structure of a box plot. Each whisker represents 25 % of the data, and the box represents the middle 50 %. With this information in mind, the following statements are true.
This means that the 50 % of the survey size needs to be calculated in order to determine the number of women who are expected to spend between $30 and $40 on clothes. Recall that 100 women participated in the survey. 100*0.5=50 Therefore, 50 out of the 100 women surveyed are expected to spend between $30 and $40 on clothes per month.
All the pieces to analyzing data using histograms have been covered. This method of displaying data makes it easier to find the data distribution and determine the best measures of center and variation to describe the data set. Recall the data Tadeo and Ramsha recorded about the main dishes of the restaurants in their neighborhood at the beginning of the lesson.
| Average Main Dish Price (Dollars) | ||||
|---|---|---|---|---|
| 10.12 | 9.29 | 8.29 | 9.78 | 10.69 |
| 9.68 | 12.09 | 8.94 | 10.81 | 8.62 |
| 11.39 | 12.62 | 8.71 | 10.74 | 10.52 |
| 10.77 | 10.15 | 9.18 | 8.45 | 9.52 |
| 11.89 | 9.77 | 9.44 | 13.24 | 11.01 |
| 10.62 | 9.38 | 12.15 | 9.68 | 9.60 |
| 10.32 | 11.31 | 11.41 | 8.62 | 9.27 |
| 10.96 | 9.18 | 10.28 | 10.71 | 10.02 |
The two students wanted to draw some insights and conclusions based on this data. However, they were not entirely sure how to proceed with this task. Find the following information to help these curious connoisseurs!
Consider the following histograms.
Which histogram best describes the data?
Select the option that describes the distribution of the data.
Which measures of center and variation will best describe the data?
Begin by making a frequency table of the data.
Is there a tail to the distribution? Which way does it extend?
Which measures of center and variation should be used when the distribution of the data is skewed?
Note that the given histograms consist of six intervals. With this in mind, make a frequency table using six intervals, starting with 8.00-9.99.
| Average Main Dish Price (Dollars) | |
|---|---|
| Price Range | Frequency |
| 8.00-8.99 | 6 |
| 9.00-9.99 | 12 |
| 10.00-10.99 | 13 |
| 11.00-11.99 | 5 |
| 12.00-12.99 | 3 |
| 13.00-13.99 | 1 |
Next, use this frequency table to display the data in a histogram. The horizontal axis will be the Price Range
and the vertical axis the Frequency.
Then, draw the bars to represent the frequency of each interval.
Notice that this corresponds to option C.
It can be seen in the histogram that the tail of the distribution extends to the right and that most of the data is on the left.
Consider the appropriate measures of center and variation for a skewed and a symmetric distribution.
| Distribution | Measure of Center | Measure of Variation |
|---|---|---|
| Symmetric | Mean | Standard deviation |
| Skewed | Median | Five-number summary |
Because in this situation the distribution of the data is skewed, the median and the five-number summary best describe the center and variation of the data, respectively.
We will begin by looking at the distribution and where its tail extends.
Notice that the tail extends to the left of the distribution and most data is on the right. This means that the distribution skews to the left.
We will draw a line in the middle of the distribution.
We can see that the line divides the distribution into two approximately mirror images. Therefore, the distribution is symmetric.
Now, let's look at the histogram.
In this case, the tail of the distribution extends to the right and most of the data is on the left. Therefore, the distribution is skewed right.
We will determine the shape of the given distribution. Begin by looking at the given histogram.
We can see that the histogram has two peaks. Moreover, notice that the distribution can be split into two clusters. Let's draw a vertical line around the middle of the distribution.
Given the two clusters and the two peaks, the data follows a bimodal distribution. Moreover, the data on the left of the halfway line is an approximate mirror image of the data on the right, this means that the distribution is also symmetric.
Following a similar procedure, let's look at the given histogram.
Notice that the bars look about the same height. Let's draw a horizontal line above and close to the highest bar.
Given that the bars are approximately the same height, the data follows a uniform distribution. Moreover, a uniform distribution is also symmetric.
| Hours Online | Frequency |
|---|---|
| 0-2 | 4 |
| 3-5 | 8 |
| 6-8 | 13 |
| 9-11 | 16 |
| 12-14 | 20 |
| 15-17 | 24 |
| 18-20 | 30 |
| 21-23 | 55 |
| 24-26 | 70 |
Use a histogram to determine the distribution of the data.
| Number of Orders | Frequency |
|---|---|
| 0-3 | 80 |
| 4-7 | 50 |
| 8-11 | 25 |
| 12-15 | 8 |
| 16-19 | 15 |
| 20-23 | 9 |
| 24-27 | 2 |
Display the data in a histogram to determine the distribution of the data.
We are given the data collected in a survey about the hours the neighbors spend online per week. By displaying data in a histogram, its distribution can be found. To do so, the vertical axis will be the Frequency
and the horizontal axis the Hours.
Next, bars will be plotted to represent the frequency of data points falling in each interval.
Notice that the tail of the histogram extends to the left and most of the data is on the right. This means that the data obtained in the survey is skewed left.
Following a similar procedure, we can determine the distribution of the data about the number of times people order food via an app delivery. In this situation, the vertical axis will be the Frequency
and the horizontal axis the Number of Orders.
In this case, the tail of the distribution extends to the right and most of the data is on the left. Therefore, the data about the number of times people order food via an app delivery skews to the right.
Consider the number of runs scored by a softball team in 30 games.
| Number of Runs | ||||
|---|---|---|---|---|
| 8 | 5 | 7 | 12 | 2 |
| 4 | 4 | 4 | 10 | 7 |
| 11 | 9 | 2 | 10 | 6 |
| 7 | 1 | 4 | 6 | 16 |
| 17 | 10 | 14 | 6 | 6 |
| 12 | 16 | 4 | 14 | 6 |
Use a box plot do identify the distribution of the data.
We are given the number of runs scored by the softball team. We will use a box plot to display the data and identify its distribution. To do so, we will use a graphing calculator. First, let's input the data to the calculator. Push STAT, chose Edit,
and enter the values in the first column.
Now, to get the box plot, push 2nd and Y=, and choose one of the plots in the list. Make sure you turn the plot ON,
set the Type
to box-and-whiskers plot, and assign L1 as XList.
To graph the box plot, push GRAPH. Note that we may need to change the window size so that it spans the length of the box-and-whiskers plot. To do so, push WINDOW.
Notice that the right whisker is longer than the left and the median is closer to the left whisker than it is to the right. This means that the data is skewed right.