Tuesday, October 24, 2017

Assignment 3: Z-Scores and Probability

Introduction

This assignment looks into foreclosures in Dane County Wisconsin between the years 2011 and 2012.  The purpose of this assignment is to spatially analyze and look for patterns in the foreclosure data through the changes between the two years.  The probability of foreclosures in all of Dane County pf the patterns holding were calculated for exceeding 80% and 10% of the time.  Also predictions for the chance of foreclosures increasing in 2013 were made.  Data was organized into census tracts for Dane County giving the number of foreclosures per tract.     

Methodology

To accomplish the goals of the assignment several methods were used.  Z-scores were used to measure the number of standard deviations specific foreclosure counts for specific census tracts in Dane County for 2011 and 2012.  Data from the 2011 and 2012 census bureau was used to define the number of foreclosures per census tract in Dane county was used in this analysis.

Z-Scores:

Z-scores were calculated to help analyze the spatial changes in foreclosures in Dane County between the years 2011 and 2012.  Z-scores are the measure of how many standard deviations an observation is above or below the mean of the population of observations.  To calculate z-scores the standard deviation and mean of the data set need to be known.  The equation to calculate z-scores is in Figure 1.   The mean and standard deviation were calculated in ArcMap for the number of foreclosures in Dane county in 2011 and 2012.  The recorded mean and standard deviations for the 2011 and 2012 data can be seen in Figure 2.  Figure 3 is of the specific equations used to calculate the z-scores for 2011 and 2012 census foreclosure data.  Z-scores were calculated for three census tracts in Dane County: Census Tract 108, 25, and 120.01.  Figure 4 shows the number of foreclosures for the three selected census tracts and the calculated z-scores.   
       
Figure 1. Equation to calculate z-scores. 

Figure 2. The mean and standard deviation measures for the counts of foreclosures in Dane County for 2011 and 2012.  These measurements were used to calculate z-score for the foreclosure data.


Figure 3. The specific equations used to calculate the z-scores for the 2011 and 2012 foreclosure data.  The observation values for the equations can be found in Figure 4.

Figure 4.  The calculated z-scores for the three selected census tracts for foreclosures in 2011 and 2012.  The number of foreclosures for each selected tract in 2011 and 2012 is also displayed.

Probability

Z-scores can be used to find the probability of occurrence of a value by using a normal distribution as a guide.  Two different probabilities for Dane County were calculated.  The first was the probability that the number of foreclosures for all of Dane County will be exceeded 80% of the time if the patterns of 2012 continue.  The second probability was that the number of foreclosures for all of Dane County will be exceeded 10% of the time if the patterns of 2012 continue.  To determine these probabilities the equation used to find z-scores in Figure 1 was used again only in a slightly different way.  To solve for probability the z-score is calculated and then the z-score chart in Figure 5.  To use the chart a z-score is found using the outer x and y axis of the table and the probability is found by cross referencing the two axis values.     
Figure 5. Z-score chart used to calculate probability.

In this example of using probability the two probabilities are given and the z-score needs to be found from the table is Figure 5.  By using the z-score equation in Figure 1 the number of foreclosures at that probability can be found.  

To solve when the probability is 80%, first 80% was converted to a decimal, 0.80.  Next, the z-score on the chart in Figure 5 was found by using 0.80 as the probability.  The z-score found was 0.84.  Because 80% is larger than 50% meaning it is on the area to the right of the normal distribution curve the z-score was made negative, -0.84.  Now the equation in Figure 6 was used to find the number of foreclosures that can be expected 80% of the time using 2012 expectations.  The number of foreclosures that will be exceeded 80% of the time is about 4 foreclosures.      
Figure 6. Equation used to find the number of foreclosures that will probably occur 80% of the time in Dane County using the expectations from 2012

The same methodology used to solve for the probability of foreclosures being exceeded 80% of the time was used for the probability of foreclosures being exceeded 10% of the time in Dane County.  First 10% was converted to a decimal and the table in Figure 5 was used to find a z-score of 1.29. Because 10% is less than 50% meaning it is on the area to the left of the normal distribution curve 1.29 will stay a positive value.  Using the equation in Figure 7, the number of foreclosures that can be expected 10% of the time using 2012 expectations was found.  The number of foreclosures that will be exceeded 10% of the time is about 25 foreclosures.     
Figure 7.  Equation used to find the number of foreclosures that will probably occur 10% of the time in Dane County using the expectations from 2012

Results


The analysis of foreclosure data for census tracts in Dane County, WI for 2011 and 2012 presented some interesting results. The change between the number of foreclosures in the census tracts in Dane County between 2011 and 2012 can be seen in Figure 8.  The census tracts with the largest changes in increases of foreclosures occurred in the northwest, northeast, and south east corners of the county.  There were also some large increases in the number of foreclosures around the outskirts of the center of the county where the capital of Wisconsin, Madison is located.  This could be due to people not being able to afford rent living so close to a large city as the location become more valuable with the increasing economic importance of Madison between the years 2011 and 2012.  Most of the census tracts with zero change in the number of foreclosures are located in the middle of Dane County where Madison is located.  With such a large city being present there, there would be a small chance of economic deficit and foreclosures occurring.  Properties would most likely be sold or tenants be removed not properties foreclosed on. Madison is a large college town where students can come into loans or rent apartments to friends in the case of not being able to afford rent.   

Looking specifically at the census tracts the z-scores were calculated for in this assignment: census tract 25, 120.01, and 108.  Referring back to Figure 4 with the z-score calculations and looking at the changes in foreclosures in these census tracts in Figure 8 intriguing results.  In 2011 and 2012 tract 25 appeared below the mean number of foreclosures for Dane County and in 2011 and 2012 tracts 108 and 120.01 appeared above the mean number of foreclosures in Dane County.  Tract 25 while appearing below the mean decreased in the number of foreclosures between the two years resulting in 2012 it occurring even farther away from the mean.  Tract 25 and tract 108 decreased their number of foreclosures between 2011 and 2012, while tract 120.01 underwent a large increase in the number of foreclosures between 2011 and 2012.         

Figure 8. The change in the number of foreclosures in the census tracts in Dane County between 2011 and 2012.

The results from calculating the probability of the number of foreclosures that will occur 80% of the time and 10% of the time resulted allowed the census tracts with the most foreclosures to be easily identified.  In Figure 9 it can be seen which census tracts have an 80% and 10% probability for the number of foreclosures.  In Figure 9 it can be seen that the census tracts in Dane County in the middle of the county and on the west side of the county have the lowest number of foreclosures.  The census tracts with the greatest number of foreclosures are located on the east side of the county in Figure 9.  This can be explained by the capital of Wisconsin, Madison, being located in the center of Dane County.  There would not be a large number of foreclosures in the largest and most heavily populated city in Wisconsin.  Looking towards the east side of the county in Figure 9 there are large numbers of foreclosures.  It is uncertain why there large quantities of foreclosures are occurring, if a future study it would be interesting to see if large numbers of foreclosures are occurring in the east side neighboring county Jefferson County.  This could help determine if the large number of foreclosures in this area is a county problem or a regional problem.  Dane County should apply aid and additional resources to census tracts on the east side of the county.  The county should supply extra funding to the homeless shelters or other care facilities so the people that are struggling there can receive the help they need.     
Figure 9. The census tracts of Dane County that have a probability of achieving a number of foreclosures 80% and 10% of the time.

Lastly, for this assignment predictions for 2013 were made.  I predict that even more foreclosures will occur in census tract 120.01.  This county had a large change in the number of foreclosures between 2011 and 2012 and was in the top 10% for having the largest number of foreclosures in the county.  Dane County should prepare to supply more aid and assistance to this census tract.  I would expect the census tracts in the southwest part of Dane County to continue to decrease in the number of foreclosures they had a decrease in the number of foreclosures between 2011 and 2012 and they are in the 80% of all tracts in the county for having the lowest numbers of foreclosures.  I would also expect the census tracts in the center of Dane County to continue to have zero change in the number of  foreclosures like there were between 2011 and 2012 due to the large city of Madison being present there.   

Conclusions

The results of the analysis of foreclosure data for census tracts in Dane County, WI for 2011 and 2012 suggested that the east side of Dane County needs the most relief aid and help with increasing numbers of foreclosures and the greatest numbers of foreclosures in the county.  A specific census tract that should be paid special attention to is tract 120.1.  There was a large increase in foreclosures in this tract between 2011 and 2012 as well as a high probability the values in this county will greater than 25 foreclosures in 2013.  Census tracts in the center of Dane County are of minimal concern for foreclosures due to the fact of the city of Madison being present there.  I recommend the east side of the county be under surveillance for housing market changes in the value of homes decreasing do to the large number of foreclosures in the area.   

Thursday, October 5, 2017

Assignment 2: Descriptive Statistics and Mean Centers

Purpose: to become familiar with a variety of statistical methods and computer programs.

Part 1: Calculations of Data by Hand

In this part of the lab a sample of test scores from the juniors of two high schools, Eau Claire North and Eau Claire Memorial, in the Eau Claire School District are analyzed.  The fact that Eau Claire North high school does not have the highest test score and historically the highest test score comes from Eau Claire Memorial high school has led to concerns from the public leading to criticism of the teaching methods at Eau Claire North high school.  The understanding of the public is that the teachers at Eau Claire North should be fired because the test scores at that high school are so much lower.

Figure 1. The given sample of test scores from the two high schools. 

The 7 different analyses conducted:

Range: the difference between the highest and lowest value in a data set.

Mean: the average value of the data set.

Median: the value that appears in the middle of the spread of data when it is organized in numeric order.

Mode: the number that appears most commonly in a data set.

Kurtosis: in a distribution curve of the data set, the sharpness of the peak of the data. 

Skewness: the measure of the symmetry of the distribution curve of the data set.

Standard Deviation: the amount the data set differs from the calculated mean value.
There are two different calculations for standard deviation.
     1. Population Standard Deviation: used when the whole data set is being used.  An example would
           be using 12 months of weather data when looking at weather data of a whole year.
     2. Sample Standard Deviation: used when only a sample of a data set is being used. An example would                  be if you are looking at weather data for a whole year but only use 6 months of the data.
         
                   
Figure 2. The calculated results for the 7 analyses conducted for each high school.  



Figure 3. The hand written data calculations                 Figure 4. The hand written data of calculations
 for the standard deviation of the test scores                   for the standard deviation of  the test scores of
     of Eau Claire North high school.                                    Eau Claire Memorial high school.
                                                                                         


Figure 3 and Figure 4 show that in the standard deviation calculation for both high schools, the sample population standard deviation equation was used because only a sample or small portion of the test scores from the two high schools.  Every test score of every junior who took the test is not used.

The teachers at Eau Claire North should not worry about not having the highest test grade.  The statistics calculated suggests that the test scores at Eau Claire North are overall better than at Eau Claire Memorial.  The teachers at North should not be having loosing their jobs threatened.

The results in Figure 2 highlight the statistics that prove that Eau Claire North has better test scores.  The most common test score (mode) is higher at Eau Claire North.  The average (mean) score is better at Eau Claire North.  Meaning that overall there were higher than low scores at Eau Claire North.  The standard deviation at Eau Claire North was higher also, which means that the data deviated less from the mean than Eau Claire Memorial.  The test scores at North were more concentrated around a higher mean value than at Memorial where test scores were less concentrated around a lower mean.  The range is smaller at Eau Claire North than at Eau Claire Memorial, while Memorial may have the highest score, it also has the lowest score.  Eau Claire North has the larger middle number, median, suggesting that there are less low test scores than at Memorial.  Both high schools have negative kurtosis meaning that the distribution of the test scores at both schools is flat and not peaked.  This means both schools have broad distributions of test scores.  Both high schools have negative skew.  This suggests that both schools have a higher concentration on test scores on the higher end of the mean.  North has a skew value that deviates from 1 more than memorial meaning that there is a larger concentration of high test scores than at Memorial.  Overall the statistics show that Eau Claire North has better overall test scores as Eau Claire Memorial, even though Eau Claire Memorial has the highest test score.  The teachers are North should not be worried about losing their jobs, and the public should stop criticizing their hard work.           

Part 2: Calculating Mean Centers and Weighted Mean Centers

For part two of the lab population data from 2000 and 2015 for Wisconsin Counties was used to calculate and compare mean centers of Wisconsin.  Three different mean centers were calculated by using ArcMap: geographic mean center of Wisconsin, weighted mean center of population from 2000, and weighted mean center of population from 2015.  In ArcMap in ArcToolbox under the Statial Statistics Tools in the Measuring Geographic Distributions tool set the Mean Center operation was used.  Mean Centers are calculated by using the average x and y coordinates in a selected geographic area.  By weighting the mean center by population changes in population between 2000 and 2015 can be tracked and compared.   
Figure 5. Distribution of calculated mean centers in Wisconsin.

The geographic mean center for Wisconsin was very different from the two weighted calculated mean centers.  The geographic mean center resides in the exact middle of the state.  In Figure 5 it can be seen that the geographic mean center is farther north and west than the two mean centers weighted by county populations.  This can be explained by larger, more heavily populated cities in Wisconsin like Racine, Milwaukee, and Madison being located in the more southern and eastern part of the state.  That moves the population based mean centers closer to these heavily populated cities and farther away from the geographic mean center.  The large city of Green Bay is also located on the eastern side of the state, moving the population mean centers eastward.  The mean centers that were weighted based on county population data are very similar.  The mean center based on the Wisconsin county populations of 2015 is more western than the mean center based on the Wisconsin county populations of 2000.  I think this shift between the two mean centers occurred because cities in the western part of the state grew in population's size, like the City of Eau Claire.  Overall all of the mean centers are located in the central region of Wisconsin, but the weighted mean centers are located more in the southeastern part of the state than the geographical mean center.       


Citations:

Kurtosis Formula. (n.d.). Retrieved October 03, 2017, from http://www.macroption.com/kurtosis-formula/.

Skewness Formula. (n.d.). Retrieved October 03, 2017, from http://www.macroption.com/skewness-formula/.

Sunday, September 24, 2017

Assignment 1: Data Types and Classification Methods

Goals of the Assignment: 
  • Differentiate between levels of measurement
  • Differentiate between classification methods
  • Retrieving data from the U.S. Census and joining data
  • Enhance cartographic knowledge
Part 1: Vocabulary Definitions and Examples

Nominal Data:
Nominal data is data represented through classifications into categories.  There are no quantitative measurements discretely measured in this type of data.  There must be at least two different categories being used.  This type of data uses words as representations of classifications of data.  No real comparisons are made between the different classification categories.  Each category is taken for what it is and correlations are not made between the different categories.

Example 1: Nominal data can depict preferences of regions, states, counties, or others.  An example is a map depicting the favorite baseball team of each U.S. County is an example of nominal data.  This can be considered nominal data because it groups a county into a category depending on which baseball team is liked more.  There are no quantitative measurements or relationships used between the different baseball teams.  In Figure 1 is a map depicting what the favorite baseball team of each county in the U.S. (Meyer).  

Figure 1. An example of nominal data by determining which baseball team is the most popular in the different counties of the U.S. (Meyer).

Example 2: Nominal data can also represent medical results by displaying which diseases are the highest or lowest in regions, states, or others.  An example is a map giving the most common cause of death in each state in the U.S.  This is an example of nominal data because different diseases are assigned to different states with no numerical data.  Different diseases represent different categories and states are assigned to these categories.  In Figure 2 is a map highlighting what the most common cause of death in each state other than heart disease or cancer (Blatt).


Figure 2. An example of nominal data is categorizing states by which diseases are the highest causes of death other than cancer and heart disease (Blatt).

Ordinal Data:
Ordinal data is data that has a clear scale establishing which values are of greater or lesser worth than other values.  Ordinal data uses scales of rank to structure data in a system of rank.  An example of ordinal data is at the doctor's office when the doctor has the patient rank their pain on a scale of 1 to 10.

Example 1: Examples of ordinal data can include data that ranks places as the best or worst at things based on a scale. An example of this kind of ordinal data would be data on the best and worst places to live in the US.  Figure 3 is an example of ordinal data.  All 50 states of the U.S. are ranked on being the worse and safest states to live in based on crime data (Peter).  It was determined that the safest place to live is South Dakota and the most dangerous place to live is District of Columbia.   

Figure 3.  A map ranking the worse and safest states to live in the U.S. (Peter).  This is considered ordinal data because the states are ranked against each other. 

Example 2: A second example of ordinal data is economic ranking data.  In Figure 4, the 50 states of the U.S. are ranked based on economic outlook rankings (Horvath).  New York has the worst economic outlook and Utah has the best economic outlook according to this data.  This data is considered ordinal because the data from each state is ranked according to the data from other states.  

Figure 4. The ranking of the richest and poorest states in the U.S. (Horvath).  This data is ordinal because the data from each state is compared to that of other states and then ranked against each other to present the data.

Interval Data:
Interval data is a continuous numeric data set with no established zero point.  This type of data allows people to determine which data is greater or less than others, but without the exact zero point ratios cannot be established. 

Examples: A good examples of interval data is standardized test scores.  IQ and ACT test scores do not have an absolute zero score.  The score achieved is just compared to other scores achieved by others to determine the quality of the score.  Figure 5 gives the break of IQ test scores globally (National IQ Scores) and Figure 6 gives the distribution of ACT scores by state in the U.S. (Liu). 

Figure 5. The average national IQ scores for the globe (National IQ Scores).  This map is an example of interval data because there is no absolute zero IQ score. 

Figure 6. The average ACT scores for each state in the U.S. (Liu).  Again there is no absolute zero ACT test score meaning this data is interval.

Ratio Data:
Ratio data is a continuous numeric data set with a naturally established zero point.  This data type allows for good comparisons and ratios can be established.  Magnitude of the data set can be determined.  Because of the exact zero point a wider range of statistics can be applied to this data. 

Example 1: An example of ratio data is a data set showing the percent of adults twenty years and older who are obese (American Obesity Treatment Association). In Figure 7, the data shows that the greatest amount of obesity in the U.S. can be found in southeastern U.S.  This data set is ratio data because rates were calculated which can be done with ratio data because there is an absolute zero meaning there is magnitude to the data.
Figure 7. The percent of adults 20 years old and older that are obese in the United States (American Obesity Treatment Association).  The absolute zero value that exists within the data set makes this an example of ratio data. 

Example 2: A second example of ratio data is energy industry spatial distributions.  Figure 8 maps the active manufacturing facilities in 2015 and the amount of megawatt energy created by each state (Wanner). The state that creates the most megawatt energy from wind is Texas.  This is an example of ratio data because there is absolute zero value, a state could make zero megawatts of energy from wind-related manufacturing.     

Figure 8. The wind-related manufacturing facilitates in the U.S. and which states produce the least and greatest amounts of mega wattage this way (Wanner).  The absolute zero value that exists makes this data an example of ratio data.  

Part 2: Working with USDA and U.S. Census data

This part of the assignment is a mock assignment from an agriculture consulting/ marketing company, looking to determine where the new USDA Certified Organic Farms should be placed in Wisconsin.  To determine where current USDA Certified Organic Farms are located 2012 census of agriculture data was used for each Wisconsin county.

To begin the project, I first used agriculture data from the 2012 Census of agriculture.  In an excel file the GEO-ID for each Wisconsin County and the name of the county was recorded.  An OrganicAg column was added where the number of organic farms in the county was recorded.  Next, from the United States Census Bureau website the shape file for Wisconsin counties was downloaded.  The file was unzipped and added to an ArcMap viewer.  The excel file was also added to the ArcMap viewer as a sheet.  The sheet was joined to the shape file by joining based on the GEO_ID field.  The shapefile was then prepared to start displaying data.

The organic farm data was displayed in three different classification methods: natural breaks, equal interval, and quantile.  All three of these methods can be seen in Figure 9.

 The natural breaks method works by placing breaks between classes where breaks look like they should be placed based on spacing in the data values.  Where there are noticeable spaces in between data values breaks between classes are placed.

The equal interval method works to make each classification equal.  This method divides the range by the number of classes to be created.  This method may result in some classes not having any values from the data because it does not consider the middle values in the data set.  This method only considers the smallest and largest value of the data set.

The quantile method works by making sure that each class has an equal number of data values.  The problem with this method is that it may misrepresent the natural organization of the data by forcing data values into organized classes just based on the number of values per class alone.

Figure 9. Displaying concentrations of USDA certified organic farms by county in Wisconsin. Three different classification method were used to give different symbolization of the same organic farms data set. 

By using the Figure 9 data, I think that new USDA certified organic farms should be placed in the northeast counties of the Wisconsin counties. Specifically I think Vilas and Oneida counties could benefit from more organic farms because in all three classification methods those counties have the smallest number of organic farms and are the farthest away from existing organic farm locations.  I think the map that most accurately represents where current organic farms are located is the natural breaks map, because this map looks specifically at the organization of the data to determine where breaks should be placed.  This method does not force a certain number of values into each class like the quantile method or only look at the range of values ignoring the importance of the middle values of the data set like the equal interval method.  The equal interval map has categories with no values represented and some categories with only now county in it.  It does not give a good overall representation of the data set.  I  Figure 9 it can be seen that the classes of the quantile map cover very different spreads of data, one class only includes three values and another includes 112.  Because of this it makes some counties appear like they have significantly more organic farms than other counties when they do not.  This scale does not group counties with similar numbers of organic farms together, but rather is focused on getting the same number of values in each category.  I would recommend using the natural breaks map to determine where to place new organic farms.  This method looks to group counties with similar numbers of organic farms together, based on breaks in the data that naturally occur.  This map shows a limited number of organic farms in northern Wisconsin and central Wisconsin.  I would not advise placing organic farms in central Wisconsin because that is where large cities and residential areas like Madison are located, so there most likely is not space for large farming operations there.  I would advise placing the organic farms in the north where there are few and more land available for farming due to the lack of large cities there.  The natural breaks map most accurately represents the distribution of organic farms in Wisconsin counties.        

Citations:

American Obesity Treatment Association. (n.d.). U.S. Obesity Trends. Retrieved September 24, 2017, from http://www.americanobesity.org/obesityInAmerica.htm.

Blatt, B. (2014, June 03). You Live in Alabama. Here's How You're Going to Die. Retrieved September 21, 2017, from http://www.slate.com/articles/life/culturebox/2014/06/death_map_the_most_common_causes_of_death_in_each_state_of_the_union.html.

Data Access and Dissemination Systems (DADS). (2010, October 05). American FactFinder. Retrieved September 24, 2017, from http://factfinder2.census.gove/faces/nav/jsf/pages/index.xhtml.

Horvath, J. (2016, April 14). Rich States, Poor States: 9th Edition Released. Retrieved September 24, 2017, from https://www.alec.org/article/rich-states-poor-states-9th-edition-released/.

Liu, Q. (1970, January 01). Professor Q's taught Queue. Retrieved September 24, 2017, from http://qiyuliu.blogspot.com/2012/06/.

Meyer, R. (2014, March 31). Here Is Every U.S. County's Favorite Baseball Team (According to Facebook). Retrieved September 21, 2017, from https://www.theatlantic.com/technology/archive/2014/03/here-is-every-us-countys-favorite-baseball-team-according-to-facebook/359917/.

NASS- National Agricultural Statistics Service. (n.d). Retrieved September 24, 2017, from https://www.agcensus.usda.gov/Publications/2012/Full_Report/Census_by_State/Wisconsin/index.asp.

National IQ Scores. (n.d.). Retrieved September 24, 2017, from https://www.targetmap.com/viewer.aspx?reportId=2812.

Peter, John. (2016, September 20). Where are safest and cheapest place to live from USA? Retrieved September 24, 2017, from https://www.quora.com/Where-are-safest-and-cheapest-place-to-live-from-usa

Wanner, C. (2017, February 27). What's the state of American wind power manufacturing? Retrieved September 24, 2017, from http://www.aweablog.org/whats-state-american-wind-power-manufacturing/.