Properties of Expected Value
The rules of expected value: E(aX + b), E(X + Y) with or without independence, E(XY) for independent variables, and why E(g(X)) is not g(E(X)).
Properties of Expected Value
Expected value is not the most likely outcome. That is the mode. The properties of expected value define how it behaves under addition, multiplication, and scaling. These rules let you solve counting problems, evaluate bets, and model risk without knowing the full distribution. The core rule is that expected value is a linear operator. This means E(aX + b) = aE(X) + b for any constants a and b, and E(X + Y) = E(X) + E(Y) for any two random variables X and Y, even if they depend on each other. No independence is required for sums, only for products. These facts are covered in OpenIntro Statistics (4th ed.) section 3.4, pages 147-154, and developed further in Blitzstein & Hwang, Introduction to Probability (2nd ed.) Chapter 4, pages 159-218, where Theorem 4.1 on page 170 states linearity of expectation formally.
Linearity: E(aX + b) = aE(X) + b
If you have a random variable X and you multiply every outcome by a constant a and add b, the expected value shifts by the same linear transformation. This follows from the definition. For a discrete variable with outcomes x_i and probabilities p_i, the sum becomes Σ (a x_i + b) p_i = a Σ x_i p_i + b Σ p_i = aE(X) + b. The constant term b is multiplied by Σ p_i, which is always 1. So a constant shift adds b to the EV; a scaling multiplies the EV by the scale factor.
Common mistake: forgetting the constant term
If you convert Fahrenheit to Celsius, C = (5/9)(F - 32), and you know E(F) = 50, then E(C) = (5/9)(50-32) = (5/9)*18 = 10. Apply the multiplier and handle the constant shift.
Failure case: what does E(X²) equal?
E(X²) is not (E(X))² in general. The linearity rule applies only to the first power of X. Variance relies on this distinction because Var(X) = E(X²) - (E(X))². That is a separate topic with its own treatment.
E(X + Y) = E(X) + E(Y): No Independence Needed
The sum of two random variables has an expected value equal to the sum of their individual expected values. This holds whether X and Y are independent, dependent, or identical. It is the reason linearity of expectation is so powerful: you never need to know the joint distribution. For any two random variables, E(X + Y) = E(X) + E(Y).
Why this surprises students
Most probability rules require independence. The sum rule does not. If X is the roll of a die and Y is the same die's roll (dependent), E(X + Y) = E(X) + E(Y) = 3.5 + 3.5 = 7, which matches the direct calculation of the sum's EV: 2 through 12 each with probability 1/36 or 2/36, average 7. Independence is irrelevant.
Applied to the expected value of sum of random variables
If you have n random variables, E(X₁ + X₂ + … + Xₙ) = E(X₁) + E(X₂) + … + E(Xₙ). This is the foundation for the EV of a binomial distribution (np) and for the expected value of any sum of indicator variables. The proof is by induction on the two-variable rule.
E(XY) = E(X)E(Y) Only When Independent
The product rule requires independence. If X and Y are independent, then E(XY) = E(X)E(Y). This is not true in general. A common exam trap gives you correlated variables and asks for E(XY): the answer is not the product of the individual EVs.
What goes wrong with dependence
Suppose X is the number of aces drawn and Y is the number of kings drawn from a deck without replacement. X and Y are negatively correlated. E(XY) is less than E(X)E(Y) because drawing an ace reduces the chance of drawing a king. The product rule would overestimate the expected joint count. Blitzstein & Hwang cover this in Chapter 4 alongside the linearity theorem. The correct approach when dependence is present is to compute E(XY) from the joint distribution or use the covariance formula: Cov(X,Y) = E(XY) - E(X)E(Y).
Using the rule correctly for e(xy) independent
If X and Y are independent, the joint probability mass function factorises, and the double sum ΣΣ x_i y_j p_X(x_i) p_Y(y_j) becomes (Σ x_i p_X(x_i)) (Σ y_j p_Y(y_j)) = E(X)E(Y). That is the proof in one line. Verify independence before applying the product rule.
Jensen's Inequality in One Example
Expected value does not pass through nonlinear functions. Jensen's inequality states that for a convex function φ, E(φ(X)) ≥ φ(E(X)). For a concave function, the inequality reverses. This is why E(X²) ≥ (E(X))² (since x² is convex) and E(√X) ≤ √(E(X)) (since √x is concave).
A concrete failure case
Consider a bet where X = +$100 with probability 0.5 and -$50 with probability 0.5. E(X) = $25. Take a non-negative variable: uniform on {0, 100}. E(X) = 50. E(√X) = (√0 + √100)/2 = 5. √(E(X)) = √50 ≈ 7.07. So E(√X) < √(E(X)). The concave transformation lowers the expected value relative to the transform of the EV. This matters when you evaluate risk-averse decisions. Blitzstein & Hwang discuss utility in Section 4.9, pages 213-218, and note that Jensen's inequality is the mathematical reason a risk-averse person may reject a positive-EV bet.
Using Linearity to Solve Counting Problems with Indicator Variables
The most powerful application of linearity of expectation is to compute expected counts without finding the distribution. Define an indicator random variable for each event: I_A = 1 if event A occurs, 0 otherwise. Then E(I_A) = P(A). For a sum of indicators, E(Σ I_i) = Σ E(I_i) = Σ P(event i). This works regardless of dependence among the events.
Worked example: expected number of aces in a poker hand
Deal a 5-card hand from a standard 52-card deck. Let I_i be 1 if the i-th card is an ace, 0 otherwise. P(I_i = 1) = 4/52 = 1/13. By linearity, E(number of aces) = I₁ + … + I₅ = 5 × (1/13) = 5/13. The cards are drawn without replacement, so the indicators are dependent, but linearity still holds. The direct hypergeometric calculation gives the same number. Blitzstein & Hwang define indicator random variables on page 173 of Chapter 4 and use them for the expected number of fixed points in a random permutation.
Worked example: expected number of times a fair die shows 6 in 10 rolls
Let I_j = 1 if roll j is a 6. P(I_j = 1) = 1/6. E(total sixes) = 10 × (1/6) = 10/6. This matches the binomial expected value np = 10 × (1/6). No binomial formula needed, just linearity and the indicator approach.
Failure case: forgetting that indicators are not independent
A student might try to compute E(I_1 I_2) = E(I_1)E(I_2) for dependent events. That would require independence, which the aces example does not have. The product E(I_1 I_2) is the probability that both the first and second card are aces: (4/52)(3/51) ≈ 0.0045, not (1/13)² ≈ 0.0059. The product rule fails; linearity for sums never does.
Common Questions
Can I use linearity of expectation when the variables are not independent?
Yes. That is the entire point. E(X + Y) = E(X) + E(Y) holds for any two random variables, regardless of dependence. The only place independence matters is for products, E(XY) = E(X)E(Y).
What is the most common mistake when computing E(aX + b)?
Forgetting to apply the constant shift after scaling. If you know E(X) = 10 and you want E(2X - 5), the result is 2×10 - 5 = 15, not 20. The constant term is multiplied by the probability sum which is 1, so it is simply added.
How do I compute the expected value of a product when variables are not independent?
You need the joint distribution or the covariance: E(XY) = Cov(X,Y) + E(X)E(Y). Without independence, the product rule does not apply. Do not assume E(XY) = E(X)E(Y) unless you have verified independence.
Why is linearity of expectation useful for counting problems?
Because you can break a complicated count into a sum of indicator random variables, each with a simple EV equal to the probability of that event. The expected total is the sum of those probabilities, even when the events are dependent. No distribution of the total is needed.
Does Jensen's inequality mean expected value is always wrong for nonlinear functions?
Not wrong, but you must transform first then take expectation, not the other way around. E(φ(X)) is not φ(E(X)) unless φ is linear. Jensen's inequality tells you the direction of the gap. For convex functions, the EV of the transform is at least the transform of the EV.