Sign In
Data is everywhere, and understanding the nuances between different sets is paramount. One effective way to compare data is through histograms, which visually represent frequency distributions, highlighting patterns and outliers. While histograms provide a snapshot, standard deviation gives a numerical measure of how spread out the values in a dataset are. A lower standard deviation indicates values are closer to the mean, while a higher one suggests greater variability. By employing both histograms and understanding standard deviation, one can gain a comprehensive view of multiple data sets, making more informed decisions in various domains.
Show less Show more expand_more| Student Learning Objectives: |
|---|
|
| | 11 Theory slides |
| | 9 Exercises - Grade E - A |
| | Each lesson is meant to take 1-2 classroom sessions |
Try a few exercises to ensure the following lesson will be better understood. If these are too tough, spend a bit more time catching up with the recommended readings. Now, consider the following data set. 12.7,& 12.2,& 6.8,& 13.3,& 11.5, 4.7,& 12.1,& 16.6,& 5.5,& 11.1, 14.4,& 8.1,& 15.7,& 12.6,& 8.8
Find the mean and give the answer as a decimal rounded to one decimal place.
Find the median and give the answer as a decimal rounded to one decimal place.
Find the range and give the answer as a decimal rounded to one decimal place.
Find the interquartile range and give the answer as a decimal rounded to one decimal place.
Find the mean absolute deviation and give the answer as a decimal rounded to two decimal places.
In 1917, Pekka Brofeld published the measurements of various fish species caught in a lake near Tampere, Finland. Representing part of the findings is a data set about the length and height of four of the fish species caught. Move the slider to see the complete data set for these four fish.
Match the Finnish name with the Latin name.
The following box plots show the distribution of the heights (in feet and inches) of the players on the Ohio State Buckeyes men's basketball and football teams in the 2020--2021 season.
Considering the chart, match each respective box plot with the correct team.
Height tends to be more advantageous in basketball than in football. Therefore, it is reasonable to conclude from the box plots that Team B is the basketball team and Team A is the football team.
The table below shows the average monthly high temperatures across three small towns.
One town is located in the State of Alaska, another in Florida, and the other in Nebraska.
Analyzing the data set and map, try to match each town with the correct corresponding state. Note that, generally, northern states tend to be colder than southern states.
Considering each observation from the data set and map, it is likely that Noma is in Florida, Mekoryuk is in Alaska, and Nehawka is in Nebraska.
The following applet shows the histograms of two data sets. Move the slider to investigate the observations separately.
When two data sets have similar centers, investigating the spread of each data set can be useful in highlighting their differences. One way of measuring spread is to calculate the average of each data value's distance from the mean. This measure is called the mean absolute deviation. |x_1-x|+|x_2-x|+...+|x_n-x|/n Due to the difficulty of making calculations using the absolute value, a commonly used alternate approach to measuring the spread is to calculate the standard deviation.
The standard deviation is a measure of spread of a data set that measures how much the data elements differ from the mean. The standard deviation, often represented by the Greek letter σ (sigma), is calculated by taking the square root of the variance of the data set. Let x_1, x_2,..., x_n be the data values in a set and x their mean. σ = sqrt((x_1 - x)^2 + (x_2 - x)^2 + ⋯ + (x_n - x)^2/n) The applet below calculates the standard deviation for the data set on the number line. Move the points around to change the data.
Graphing calculators can find the standard deviation, but they do not show the mean absolute deviation. It is interesting to compare the value of these two measures.
(a-b)^2=a^2-2ab+b^2
LHS+a^2+2ab+b^2≤RHS+a^2+2ab+b^2
.LHS /4.≤.RHS /4.
Factor out 2
a^2+2ab+b^2=(a+b)^2
Simplify quotient
sqrt(LHS)≤sqrt(RHS)
The proof of the inequality involving the means of more than two numbers is similar.
The table below shows the average monthly low temperatures of two cities — Kansas City and Seattle. The two cities given are not necessarily in order.
According to the data set, City A and B have annual average low temperatures around 45^(∘)F and 43^(∘)F, respectively. Referencing the map below, Seattle is located much further north than Kansas City. It is typical that northern states — on average — are colder than southern states. Nevertheless, Seattle experiences less variance in temperature changes during each season due to the ocean's tempering effect on the climate.
Use the ranges and standard deviations of the data set, along with the given geographical information, to determine which cities are pairs.
| Range | Standard Deviation | |
|---|---|---|
| City A | 55-36=19 | 6.7 |
| City B | 66-18=48 | 16.4 |
These measures of spread show that the temperature throughout the year changes much less in City A than in City B. Based on that analysis, and considering the information given about the tempering effect of the ocean, it is reasonable to conclude that City A is Seattle and City B is Kansas City. What a cool conclusion to make.
In the US stock market, a measure of how much a stock price fluctuates during a certain period of time is called historical volatility. The following data set from the year 2020 contains information about the daily closing stock price (in dollars) of two companies.
| Low | Mean | High | Standard Deviation | |
|---|---|---|---|---|
| APDN | $2.52 | $6.89 | $15.21 | $2.24 |
| DSS | $4.04 | $6.90 | $10.89 | $1.69 |
Which stock price was less volatile in 2020?
Both the range and standard deviation are smaller for DSS. These interpretations indicate that DSS's stock price fluctuated less over the year than the stock price of APDN. Therefore, it can be concluded that the stock price of DSS was less volatile in 2020.
The histogram below shows the distribution of the stock prices.
The following applet shows the histograms of two data sets. Move the slider to investigate the data sets separately. Then, answer the given questions based on visual observations.
Consider the following two histograms where neither the labels nor scales are specified.
Both of these histograms represent different distributions, and both have 26 columns.
Match the histograms to the appropriate context. Comment on the shape of the histograms.
In a lottery, all numbers are drawn with equal probability, so in the long run, it can be expected that there is little difference between the frequencies. The shape of Histogram B reflects this.
There is even more fascinating information to be discovered from the shapes of the histograms.
The height of the lone tall bar furthest to the left in Histogram A shows that in the AMC 8 competition, there were plenty of participants in 2020 who did not answer a single question correctly! Well, it is much more likely, however, that these participants registered but did not attend the competition.
Histogram A's peak shows that in 2020, on average, students in the AMC 8 competition answered less than half of the questions correctly.
The fluctuation of the bar heights in Histogram B shows that although an even distribution of the numbers is expected on the Powerball draw, some numbers historically came out fewer times.
The bar corresponding to 24 is more than twice as high as the bar corresponding to 16. However, this does not mean that 24 is twice as likely to come out in a draw. Nor does this mean that players should now play 16 because it will eventually catch up. The data is historical; it does not have any effect on the next draw.
At the beginning of the lesson, a data set of fish species was presented. The task was to match the Latin and Finnish names of four fish species by analyzing and comparing the fish drawings with a data set of fish lengths and heights.
| Mean Length to Height Ratio | |
|---|---|
| Abramis Bjorkna | 2.55 |
| Leuciscus Rutilus | 3.75 |
| Osmerus Eperlanus | 5.95 |
| Esox Lucius | 6.33 |
Next, the actual drawings can be used to find their length to height ratios. This measurement, however, is in pixels instead of centimeters. Most image software on a standard computer can show these measurements. Here, they are given. Recall that the drawings use the Finnish names.
The results, in increasing order, can be summarized as follows.
| Length to Height Ratio (Images) | |
|---|---|
| Pasuri | 361/120≈ 3.01 |
| Särki | 393/106≈ 3.71 |
| Hauki | 358/59≈ 6.07 |
| Norssi | 358/55≈ 6.51 |
The numbers in the two tables do not match exactly, which would make sense given that they are measured using different measurements, and the images are not matching in scale. Still, in both tables, two species have a ratio above 5 and two species have a ratio below 4. That means the following distinction can be made.
| Latin Name (Data Set) | Finnish Name (Images) | |
|---|---|---|
| Longer Fishes | Osmerus eperlanus and Esox lucius | Hauki and Norssi |
| Taller Fishes | Abramis bjorkna and Leuciscus rutilus | Pasuri and Särki |
Each pair of the longest fish and tallest fish are too close in value to reliably distinguish between the species. Therefore, using the tables, matching the names in corresponding order seems like a natural method to find the conclusion. rclc Abramis Bjorkna & ⟷ &Pasuri & ✓ Leuciscus Rutilus & ⟷ &Särki & ✓ Osmerus Eperlanus & ⟷ &Hauki & * Esox Lucius & ⟷ &Norsi & * These pairings, however, are not entirely correct. While the taller fishes are paired correctly, the longer fishes are not matched correctly. The correct matches are shown in the table below, which just for reference also includes the English names.
| Latin Name | Finnish Name | English Name |
|---|---|---|
| Abramis Bjorkna | Pasuri | Bream |
| Leuciscus Rutilus | Särki | Roach |
| Osmerus Eperlanus | Norssi | Smelt |
| Esox Lucius | Hauki | Pike |
A company is interested in knowing how many miles their employees have to walk to and from work during a given week. They conduct a survey with 100 men and 100 women and present the results with two box plots.
One of the assistant managers, Tearrik, is asked to make a single boxplot showing every observation. The next morning, he produces the following box plot to show upper management.
His boss, Vincenzo, tells him that there must be some mistake. However, Tearrik is convinced that he has drawn the box plot correctly. Who is right?
Notice that each box plot contains 100 observations, which is an even number. This means the median will be the mean of the 50^(th) and the 51^(st) observations. Similarly, the lower quartile is the mean of the 25^(th) and 26^(st) observations, and the upper quartile is the mean of the 75^(th) and 76^(st) observations. Therefore, each section of our box plots contains 25 observations.
By using this information we can identify how many observations should fall in the different sections of our two boxplots.
Notice that the combined boxplot contains 200 observations. Therefore, each of its four sections must contain 50 observations. Let's replace the two boxplots in the diagram above with the combined boxplot presented by Tearrik. We will keep the three intervals we have identified.
From the women's boxplot, at least 75 observations are less than or equal to 11.5. From the men's boxplot, at least 50 observations are less than or equal to 11.5. Combined, we have that at least 125 are less than or equal to 11.5 — so less than 12. This implies that the median of the combined data cannot be 12 as Tearrik's box plot states. Therefore, Vincenzo is correct.
We will call the three numbers x_1, x_2, and x_3. Now let's add 4 to each of these numbers, which gives us a new data set. x_1+4, x_2+4, x_3+4 We will first write two expressions — one for the old mean x_O, and another for the new mean x_N. x_O &= x_1+x_2+x_3/3 [0.5em] x_N &= (x_1+4)+(x_2+4)+(x_3+4)/3 Next, we will simplify the right-hand side of the second expression.
If we examine the right-hand side, we see that we have the sum of a fraction that matches the old mean and 4.
Therefore, adding four to each number increased the mean by 4.
Let's write the standard deviation for the original data set. We have three numbers, so the denominator is 3.
We will also write an expression for the new standard deviation. Remember that the new mean is x_N=x_O+4.
As we can see, σ_N is the same expression as σ_O. This means the standard deviation remains unchanged by adding the same number to each observation from the data set.
Why did the mean increase by 4? Well, when we add 4 to each to each number, every observation in the data set is pushed to the right by 4 units. Therefore, it makes sense that the mean increased by 4.
However, since every number increased by the same amount, their relative distance to the mean remains unchanged. Therefore the standard deviation, which measures spread, is unchanged.