Samiksha Jaiswal (Editor)

Edgeworth series

Updated on
Edit
Like
Comment
Share on FacebookTweet on TwitterShare on LinkedInShare on Reddit
Edgeworth series

The Gram–Charlier A series (named in honor of Jørgen Pedersen Gram and Carl Charlier), and the Edgeworth series (named in honor of Francis Ysidro Edgeworth) are series that approximate a probability distribution in terms of its cumulants. The series are the same; but, the arrangement of terms (and thus the accuracy of truncating the series) differ.

Contents

Gram–Charlier A series

The key idea of these expansions is to write the characteristic function of the distribution whose probability density function F is to be approximated in terms of the characteristic function of a distribution with known and suitable properties, and to recover F through the inverse Fourier transform.

We examine a continuous random variable. Let f be the characteristic function of its distribution whose density function is F, and κ r its cumulants. We expand in terms of a known distribution with probability density function Ψ, characteristic function ψ, and cumulants γ r . The density Ψ is generally chosen to be that of the normal distribution, but other choices are possible as well. By the definition of the cumulants, we have (see Wallace, 1958)

f ( t ) = exp ⁡ [ ∑ r = 1 ∞ κ r ( i t ) r r ! ] and ψ ( t ) = exp ⁡ [ ∑ r = 1 ∞ γ r ( i t ) r r ! ] ,

which gives the following formal identity:

f ( t ) = exp ⁡ [ ∑ r = 1 ∞ ( κ r − γ r ) ( i t ) r r ! ] ψ ( t ) .

By the properties of the Fourier transform, ( i t ) r ψ ( t ) is the Fourier transform of ( − 1 ) r [ D r Ψ ] ( − x ) , where D is the differential operator with respect to x. Thus, after changing x with − x on both sides of the equation, we find for F the formal expansion

F ( x ) = exp ⁡ [ ∑ r = 1 ∞ ( κ r − γ r ) ( − D ) r r ! ] Ψ ( x ) .

If Ψ is chosen as the normal density with mean and variance as given by F, that is, mean μ = κ 1 and variance σ 2 = κ 2 , then the expansion becomes

F ( x ) = exp ⁡ [ ∑ r = 3 ∞ κ r ( − D ) r r ! ] 1 2 π σ exp ⁡ [ − ( x − μ ) 2 2 σ 2 ] ,

since γ r = 0 for all r >2 as higher cumulants of the normal distribution are 0. By expanding the exponential and collecting terms according to the order of the derivatives, we arrive at the Gram–Charlier A series. If we include only the first two correction terms to the normal distribution, we obtain

F ( x ) ≈ 1 2 π σ exp ⁡ [ − ( x − μ ) 2 2 σ 2 ] [ 1 + κ 3 3 ! σ 3 H 3 ( x − μ σ ) + κ 4 4 ! σ 4 H 4 ( x − μ σ ) ] ,

with H 3 ( x ) = x 3 − 3 x and H 4 ( x ) = x 4 − 6 x 2 + 3 (these are Hermite polynomials).

Note that this expression is not guaranteed to be positive, and is therefore not a valid probability distribution. The Gram–Charlier A series diverges in many cases of interest—it converges only if F ( x ) falls off faster than exp ⁡ ( − ( x 2 ) / 4 ) at infinity (Cramér 1957). When it does not converge, the series is also not a true asymptotic expansion, because it is not possible to estimate the error of the expansion. For this reason, the Edgeworth series (see next section) is generally preferred over the Gram–Charlier A series.

The Edgeworth series

Edgeworth developed a similar expansion as an improvement to the central limit theorem. The advantage of the Edgeworth series is that the error is controlled, so that it is a true asymptotic expansion.

Let {Zi} be a sequence of independent and identically distributed random variables with mean μ and variance σ2, and let Xn be their standardized sums:

X n = 1 n ∑ i = 1 n Z i − μ σ .

Let Fn denote the cumulative distribution functions of the variables Xn. Then by the central limit theorem,

lim n → ∞ F n ( x ) = Φ ( x ) ≡ ∫ − ∞ x 1 2 π e − 1 2 q 2 d q

for every x, as long as the mean and variance are finite.

Now assume that the random variables Xi have mean μ, variance σ2, and higher cumulants κr=σrλr. If we expand in terms of the standard normal distribution, that is, if we set

Ψ ( x ) = 1 2 π exp ⁡ ( − 1 2 x 2 )

then the cumulant differences in the formal expression of the characteristic function fn(t) of Fn are

κ 1 F ( n ) − γ 1 = 0 , κ 2 F ( n ) − γ 2 = 0 , κ r F ( n ) − γ r = κ r σ r n r / 2 − 1 = λ r n r / 2 − 1 ; r ≥ 3.

The Edgeworth series is developed similarly to the Gram–Charlier A series, only that now terms are collected according to powers of n. Thus, we have

f n ( t ) = [ 1 + ∑ j = 1 ∞ P j ( i t ) n j / 2 ] exp ⁡ ( − t 2 / 2 ) ,

where Pj(x) is a polynomial of degree 3j. Again, after inverse Fourier transform, the density function Fn follows as

F n ( x ) = Φ ( x ) + ∑ j = 1 ∞ P j ( − D ) n j / 2 Φ ( x ) .

The first five terms of the expansion are

F n ( x ) = Φ ( x ) − 1 n 1 2 ( 1 6 λ 3 Φ ( 3 ) ( x ) ) + 1 n ( 1 24 λ 4 Φ ( 4 ) ( x ) + 1 72 λ 3 2 Φ ( 6 ) ( x ) ) − 1 n 3 2 ( 1 120 λ 5 Φ ( 5 ) ( x ) + 1 144 λ 3 λ 4 Φ ( 7 ) ( x ) + 1 1296 λ 3 3 Φ ( 9 ) ( x ) ) + 1 n 2 ( 1 720 λ 6 Φ ( 6 ) ( x ) + ( 1 1152 λ 4 2 + 1 720 λ 3 λ 5 ) Φ ( 8 ) ( x ) + 1 1728 λ 3 2 λ 4 Φ ( 10 ) ( x ) + 1 31104 λ 3 4 Φ ( 12 ) ( x ) ) + O ( n − 5 2 ) .

Here, Φ(j)(x) is the j-th derivative of Φ(·) at point x. Remembering that the derivatives of the density of the normal distribution are related to the normal density by ϕ(n)(x)=(-1)nHn(x)ϕ(x), (where Hn is the Hermite polynomial of order n), this explains the alternative representations in terms of the density function. Blinnikov and Moessner (1998) have given a simple algorithm to calculate higher-order terms of the expansion.

Note that in case of a lattice distributions (which have discrete values), the Edgeworth expansion must be adjusted to account for the discontinuous jumps between lattice points.

Illustration: density of the sample mean of three χ 2 {\displaystyle \chi ^{2}}

Take X i ∼ χ 2 ( k = 2 ) i = 1 , 2 , 3 and the sample mean X ¯ = 1 3 ∑ i = 1 3 X i .

We can use several distributions for X ¯ :

  • The exact distribution, which follows a gamma distribution: X ¯ ∼ G a m m a ( α = n ⋅ k / 2 , θ = 2 / n ) = G a m m a ( α = 3 , θ = 2 / 3 )
  • The asymptotic normal distribution: X ¯ → n → ∞ N ( k , 2 ⋅ k / n ) = N ( 2 , 4 / 3 )
  • Two Edgeworth expansion, of degree 2 and 3
  • Disadvantages of the Edgeworth expansion

    Edgeworth expansions can suffer from a few issues:

  • They are not guaranteed to be a proper probability distribution as:
  • The integral of the density need not integrate to 1
  • Probabilities can be negative
  • They can be inaccurate, especially in the tails, due to mainly two reasons:
  • They are obtained under a Taylor series around the mean
  • They guarantee (asymptotically) an absolute error, not a relative one. This is an issue when one wants to approximate very small quantities, for which the absolute error might be small, but the relative error important.
  • References

    Edgeworth series Wikipedia


    Similar Topics
    ×