Data Analysis
Start gently
Analysing data is not really about the average — it is about the spread. Two sets with the same average can mean completely different things.
Key points
The Mean (Average)
Add up all n values and divide by n. It is the most basic way to say what a "typical" value looks like, but its weakness is that a single extreme outlier can drag it a long way.
Variance
How spread out the data is around the mean. Take each value's distance from the mean, square it, then average those squares. Investment risk is measured with exactly this idea.
Standard Deviation
The positive square root of the variance. Because it comes back in the same units as the original data, it is much easier to interpret. It is what people mean by a stock's volatility, and it is what standardised test scores are built from.
Median and Mode
Median: line the values up in order and take the middle one. It shrugs off outliers, which is why median salary describes real life far better than average salary. Mode: the value that shows up most often — the one marketers look at when they ask which price customers actually chose.
The Correlation Coefficient
Measures how strongly two variables move together, on a scale from -1 to +1. +1 is a perfect positive relationship (height and weight), 0 means no linear relationship at all, and -1 is a perfect negative one (temperature and coat sales).
See it drawn
Class A: bunched around the middle (small variance).
Class B is split at both ends. Same average, yet the teaching these two classes need is nothing alike.
Jobs that use this
Data Analyst$90k
Biostatistician$115k
Quantitative Analyst (Quant)$180k
§
Members-only from here
The practice questions and full career details are for members. $4.99/mo, cancel anytime.
Comments
Sign in to comment