Expected Value of a Continuous Random Variable
Find the expected value of a continuous random variable by integrating x times the density, with worked examples for uniform and exponential distributions.
Expected Value of a Continuous Random Variable
Many students who have mastered discrete expected value hit a wall when the sum turns into an integral. The concept is the same, but the tool changes. For a continuous random variable, the expected value integral E[X] = ∫ x f_X(x) dx replaces the sum Σ x p(x). The expected value of the pdf is the long-run average, not a prediction for a single trial. You cannot land on exactly 3.14 from a uniform(0,10) distribution, but the mean of a continuous distribution settles at 5 after enough draws. The integral definition, two full examples, the law of the unconscious statistician (LOTUS) for E(g(X)), and the case where the integral diverges and the expected value does not exist are covered.
From Sum to Integral
The definition of expected value for a discrete random variable is E[X] = Σ x_i P(X = x_i). For a continuous random variable, probability lives in intervals, not points. The probability density function f_X(x) gives the density of probability at x, and the total area under f_X(x) equals 1. The continuous analogue is the integral ∫ x f_X(x) dx over the support of X, as defined in Blitzstein & Hwang ch. 5 and OpenStax Introductory Statistics ch. 5. Both sources stress that f_X(x) is nonnegative and integrates to 1. The integral computes the center of mass of the density. If X has a density that is symmetric about μ, then E[X] = μ, provided the integral converges.
When the Support is Finite
If the random variable is bounded between a and b, the integral runs from a to b. If it is unbounded, you integrate from -∞ to ∞, or from 0 to ∞ for a nonnegative variable like the exponential. The expected value integral must be absolutely convergent; if it is not, the expected value does not exist. This is not a technical detail you can ignore. The Cauchy distribution, covered in the final section, shows what happens when the tails are too heavy.
Worked Example: Uniform(a, b)
The uniform distribution has density f_X(x) = 1/(b - a) for a ≤ x ≤ b, zero elsewhere. Compute the expected value for this distribution.
E[X] = ∫_a^b x * [1/(b - a)] dx = [1/(b - a)] * (1/2) x^2 |_a^b = [1/(b - a)] * (b^2 - a^2)/2 = (b + a)/2.
The mean of a uniform distribution is the midpoint of the interval. For uniform(0, 10), E[X] = 5. The variance, from Blitzstein & Hwang ch. 5, is (b - a)^2 / 12. For uniform(0,10), Var(X) = 100/12 = 25/3 ≈ 8.33, and the standard deviation is about 2.89. OpenStax Introductory Statistics ch. 5 gives the same result through the definition Var(X) = ∫ (x - μ)^2 f(x) dx = ∫ (x - μ)^2 / (b - a) dx.
What to Do with the Result
The expected value (a+b)/2 is the long-run average. On any single draw you get a number between a and b, never the mean unless it happens to fall there by chance. The mode is every point in the interval; the median is also (a+b)/2. Over many draws, the sample average converges to (a+b)/2 by the Law of Large Numbers.
Worked Example: Exponential
The exponential distribution with rate λ has density f_X(x) = λ e^{-λ x} for x ≥ 0, zero for x < 0. The mean of the exponential is 1/λ.
E[X] = ∫_0^∞ x λ e^{-λ x} dx. Integrate by parts: let u = x, dv = λ e^{-λ x} dx. Then du = dx, v = -e^{-λ x}. The integral becomes [-x e^{-λ x}]_0^∞ + ∫_0^∞ e^{-λ x} dx = 0 + [ -1/λ e^{-λ x} ]_0^∞ = 1/λ.
For λ = 0.5, the expected value is 2. The variance, from Blitzstein & Hwang ch. 5, is 1/λ^2. Exponential(0.5) has variance 4 and standard deviation 2. The memoryless property means P(X > s+t | X > s) = P(X > t), which is unique to the exponential among continuous distributions. The median is ln(2)/λ, about 1.39 for λ = 0.5, less than the mean because the distribution is skewed right.
Checking Your Work with LOTUS
Suppose you want E[X^2] for the exponential. Use LOTUS for continuous variables: E[g(X)] = ∫ g(x) f_X(x) dx. For g(x) = x^2, E[X^2] = ∫_0^∞ x^2 λ e^{-λ x} dx. Integrate by parts twice or recognize it as the second moment of the exponential: E[X^2] = 2/λ^2. Then Var(X) = E[X^2] - (E[X])^2 = 2/λ^2 - 1/λ^2 = 1/λ^2, matching the variance from the definition.
E(g(X)) and Variance by Integration
LOTUS lets you compute the expected value of any function of X without first finding the distribution of g(X). For a continuous random variable, E[g(X)] = ∫ g(x) f_X(x) dx, as stated in Blitzstein & Hwang ch. 5. The same rule applies for variance: Var(X) = E[(X - μ)^2] = ∫ (x - μ)^2 f_X(x) dx, or equivalently E[X^2] - (E[X])^2.
Practical Steps for Variance
- Compute μ = E[X] = ∫ x f_X(x) dx.
- Compute E[X^2] = ∫ x^2 f_X(x) dx.
- Var(X) = E[X^2] - μ^2.
Linearity of expectation holds for continuous variables as it does for discrete: E[aX + b] = aE[X] + b. This is the same property used for binomial expected value and other discrete distributions, but now applied to integrals. If you know E[X] for a distribution, you can find E[2X - 3] without reintegrating: it is 2E[X] - 3.
Example with a Piecewise Density
Let f_X(x) = 2x on [0,1], zero elsewhere. Check that ∫_0^1 2x dx = 1. E[X] = ∫_0^1 x * 2x dx = 2∫_0^1 x^2 dx = 2/3. E[X^2] = ∫_0^1 x^2 * 2x dx = 2∫_0^1 x^3 dx = 1/2. Then Var(X) = 1/2 - (2/3)^2 = 1/2 - 4/9 = 1/18. The standard deviation is about 0.236.
When the Expected Value Does Not Exist (Cauchy)
The Cauchy distribution has density f_X(x) = 1/(π(1 + x^2)) for all real x. The integral ∫_{-∞}^∞ |x| f_X(x) dx diverges because the tails are too heavy. More precisely, ∫_{-∞}^∞ x f_X(x) dx does not converge absolutely, so the expected value does not exist. Blitzstein & Hwang ch. 5 explicitly states that the Cauchy has no mean. The median is 0, and the distribution is symmetric about 0, but symmetry alone does not guarantee a finite mean.
What Goes Wrong
Compute ∫_0^R x / (π(1 + x^2)) dx. As R → ∞, the integral behaves like (1/2π) ln(1 + R^2), which diverges. The positive and negative tails are balanced in principle, but the divergence is too slow. The sample mean of Cauchy draws does not converge to a fixed number; it jumps around with every new observation. If you use a Cauchy model to represent a real phenomenon (e.g., some physical measurements, certain financial returns), you cannot report an expected value as a summary. The mode, median, and interquartile range are still meaningful.
How to Detect a Nonexistent Expected Value
Check whether ∫ |x| f_X(x) dx converges over the support. If the density decays slower than 1/|x|^2 as |x| → ∞, the tail integral diverges. The Cauchy decays like 1/x^2, which is exactly the boundary, it barely fails. The Pareto distribution with shape parameter α ≤ 1 also has no finite mean. If you are working with real data and the histogram shows extreme outliers without a central tendency, test whether a distribution with a nonexistent mean might fit better.
| Distribution | Support | Density f(x) | E[X] |
|---|---|---|---|
| Uniform(a,b) | a ≤ x ≤ b | 1/(b − a) | (a + b)/2 |
| Exponential(λ) | x ≥ 0 | λ e^{-λ x} | 1/λ |
| Normal(μ, σ²) | −∞ < x < ∞ | 1/(σ √(2π)) e^{-(x−μ)²/(2σ²)} | μ |
| Gamma(α, λ) | x ≥ 0 | λ^α x^{α−1} e^{-λ x} / Γ(α) | α/λ |
| Beta(α, β) | 0 ≤ x ≤ 1 | x^{α−1} (1−x)^{β−1} / B(α,β) | α/(α+β) |
| Cauchy(x₀, γ) | −∞ < x < ∞ | γ / (π(γ² + (x−x₀)²)) | does not exist |
Who Should Use Continuous Expected Value
The calculus-based probability student who needs to compute an expected value integral, apply LOTUS, and calculate variance for continuous distributions will find the worked integrals above directly useful. This includes anyone studying from Blitzstein & Hwang or OpenStax Introductory Statistics. The material also suits project managers who use Expected Monetary Value (EMV) in decision trees, since EMV for continuous risk impacts involves integrating over probability distributions. Casual users who want to check a single lottery EV number should use a discrete calculator; continuous expected value applies when the outcome can take any value in an interval, not a countable set of prizes. If you need a guarantee about a single trial, or if you are looking for a specific outcome probability rather than the long-run average, this content is not for you. The single thing that most often goes wrong is treating the integral as optional: if the tails are heavy, the integral diverges, and reporting a finite mean is a mistake that misleads every decision you base on it.
Common Questions
How do I compute E[X] for a continuous random variable without doing the integral myself?
Memorize the closed-form mean for the common distributions: uniform (a+b)/2, exponential 1/λ, normal μ, gamma α/λ. For a variable that is not one of these, you must integrate x f_X(x) over the support. Software like Wolfram Alpha or a TI-84 Plus CE can evaluate the integral numerically if the symbolic form is messy.
Is the expected value the same as the median for a continuous distribution?
No, unless the distribution is symmetric and unimodal. For the exponential, the mean 1/λ is larger than the median ln(2)/λ. For the Cauchy, the median is 0 but the mean does not exist. Always compute both: the median is the point where the cumulative distribution function equals 0.5, found by solving ∫_{-∞}^{m} f(x) dx = 0.5.
What is the difference between E[X] and E[g(X)] when g is nonlinear?
E[X] is the mean of X. E[g(X)] is the average of the transformed values. You cannot in general plug the mean into g: E[g(X)] ≠ g(E[X]) unless g is linear. LOTUS tells you to integrate g(x) f_X(x) dx directly. For example, if X ~ exponential(1), E[X] = 1 but E[1/X] diverges because the integral ∫_0^∞ (1/x) e^{-x} dx does not converge at 0.
Can a continuous distribution have a finite variance but no expected value?
No. Variance is defined as E[(X - μ)^2], which requires E[X] = μ to exist. If the mean does not exist, variance is either undefined or infinite. Some distributions, like the Cauchy, have infinite variance and no mean. A distribution with a finite mean always has a defined variance, though the variance may be infinite if the tails are heavy enough.
When do I use the law of the unconscious statistician (LOTUS) instead of finding the pdf of g(X)?
Use LOTUS whenever it is easier to integrate g(x) f_X(x) than to derive the density of g(X). For simple functions like X^2 or e^X, LOTUS is usually faster. For a monotonic function, you can also find the transformed density and then integrate. Both methods give the same result; LOTUS avoids an extra step.