Variance and Standard Deviation of a Discrete Random Variable
Calculate the variance and standard deviation of a discrete random variable with both formulas, E[(X - mu)^2] and E[X^2] - mu^2, plus worked examples.
Variance of a Random Variable Tells You the Risk Behind the Average
The variance of a random variable measures the spread of its possible outcomes around the expected value. Without it, a reported EV of +$2 might describe a sure small win or a casino bet that could lose your rent money. Variance, and its interpretable form, standard deviation, is what separates a safe bet from a gamble. The definition formula, the shortcut formula E[X²] - (E[X])², and a worked example allow you to compute both from a probability table.
Definition Formula for Variance of a Random Variable
The variance of a random variable X, written Var(X), is defined as the expected value of the squared deviation from the mean: Var(X) = E[(X - μ)²], where μ = E[X]. For a discrete random variable you compute this by summing over every outcome: (x - μ)² × P(x). The result is in squared units, dollars squared, for example, which is why most people immediately reach for the standard deviation.
Standard deviation of a discrete random variable is simply the square root of the variance: SD(X) = √Var(X). It shares the same unit as the outcomes and as the expected value, so it is directly interpretable. An EV of +$2 with an SD of $1 tells you the typical outcome is close to the average. An EV of +$2 with an SD of $50 means huge swings are normal.
Shortcut Formula: Var(X) = E[X²] - (E[X])²
The var(x) formula E[X²] - (E[X])² is algebraically identical to the definition but easier to compute from a table. You calculate two things: the expected value of X² (square each outcome, multiply by its probability, sum), then subtract the square of the ordinary expected value. Both OpenIntro Statistics (section 3.4) and Blitzstein & Hwang (Chapter 4: Expectation) present this as the primary computational route because it avoids subtracting the mean from every outcome.
This shortcut matters when you work by hand or in a spreadsheet. In Excel or Google Sheets you can compute E[X²] with =SUMPRODUCT(values^2, probabilities) and E[X] with =SUMPRODUCT(values, probabilities), then subtract the square of the latter.
Worked Example: Variance of a Dice Game
Consider a game where you roll a fair six-sided die and win the number shown, in dollars. The probability distribution is uniform: each outcome 1 through 6 has probability 1/6.
Step By Step Calculation
First, the expected value: E[X] = (1+2+3+4+5+6) × 1/6 = 3.5. Now compute E[X²]: square each outcome: 1, 4, 9, 16, 25, 36. Multiply each by 1/6, sum: (1+4+9+16+25+36)/6 = 91/6 ≈ 15.1667. Then Var(X) = 15.1667 - (3.5)² = 15.1667 - 12.25 = 2.9167. The standard deviation is √2.9167 ≈ 1.708.
Interpreting The Result
In a single roll you cannot win $3.50 (the EV), but over many rolls your average win per roll will converge to $3.50. The SD of about $1.71 means roughly two-thirds of individual rolls fall between $1.79 and $5.21. The variance of probability distribution here is the squared measure of that spread.
What Standard Deviation Tells You That Expected Value Does Not
Expected value answers "what is the average outcome over the long run?" It says nothing about how much an individual trial can deviate from that average. Two bets can share the same positive EV while one is a near-certainty and the other is almost sure to lose most of the time, with a rare huge payout. The SD quantifies that difference.
OpenStax Introductory Statistics (Chapter 4: Discrete Random Variables, section 4.2) makes this distinction explicit: the mean describes the centre, the standard deviation describes the spread. A lottery ticket with EV +$0.50 and SD $100 is a losing proposition for anyone who cannot play millions of times. A project with EMV +$50K and SD $500K means the single outcome you actually get could be disastrous. The SD is the practical measure of risk.
Var(aX + b) and Sums of Independent Variables
When you scale a random variable by a constant a and shift it by b, the variance scales by a² and the shift adds nothing: Var(aX + b) = a² Var(X). The standard deviation therefore scales by |a|. This is why converting units (e.g., dollars to euros) changes the SD proportionally but not the underlying risk structure.
For sums of independent random variables, the variance adds: Var(X + Y) = Var(X) + Var(Y). This is the foundation for the standard deviation of a portfolio, the variance of a binomial distribution (which has its own dedicated treatment), and the reason the Law of Large Numbers works. Dependence complicates things, Var(X + Y) then includes a covariance term, but for independent variables the additive property is a powerful shortcut.
Blitzstein & Hwang (Chapter 4) and OpenIntro Statistics (section 3.4) both derive these properties. In practice you use them to compute the SD of a multi-stage bet or a sum of project risks without redoing the full probability table.
Common Questions
Why is variance not in the same units as the data?
Variance is the average of squared deviations, so its unit is the square of the outcome's unit (e.g., dollars squared). Standard deviation reverses the squaring by taking the square root, returning to the original unit.
Can variance be zero?
Yes. If every outcome is the same value, there is no spread: the variable is deterministic. For example, a bet that always pays exactly $10 has Var = 0.
How do I compute variance on a TI-84 Plus CE?
Enter the outcomes in one list and their probabilities in another. Use 1-Var Stats with the frequency list: the calculator outputs the mean, variance, and standard deviation directly.
Does a positive EV guarantee profit on the next play?
No. Positive EV means profit on average over many plays. A single trial can lose everything if variance is high. The SD tells you how wide the possible loss is.
What if my random variable is continuous, not discrete?
The definition is the same: Var(X) = ∫(x - μ)² f(x) dx. The shortcut E[X²] - (E[X])² also holds, but you integrate instead of sum. OpenStax Chapter 5 covers continuous random variables separately.
Why does Var(aX + b) not depend on b?
Adding a constant shifts every outcome by the same amount, so the deviations from the mean stay exactly the same. Spread is unchanged.
When would I use variance instead of standard deviation?
Variance appears in formulas (e.g., the sum property for independent variables, or the Kelly Criterion derivation). For interpretation, always use standard deviation because it is in the original units.