Week 04 Practice Problems (Optional)
Attempt these before Wednesday. Not graded and not submitted. Open the worked answer after you have tried the problem.
The module 2 written test draws from this pool.
Inverting the Gumbel CDF
The Gumbel distribution is the GEV with shape parameter \(\xi = 0\). Its CDF is
\[ F(z) = \operatorname{exp}\left[-\operatorname{exp}\left(-\frac{z - \mu}{\sigma}\right)\right] \]
A stormwater system is designed using annual maximum daily precipitation that follows a Gumbel with location \(\mu = 50\text{ mm}\) and scale \(\sigma = 15\text{ mm}\).
- Derive the formula for the \(T\)-year return level \(z_T\) by setting \(F(z_T) = 1 - 1/T\) and solving for \(z_T\).
- Compute the 100-year design storm.
Part a. Set \(F(z_T) = 1 - 1/T\) and take the natural log of both sides twice.
\[ \begin{aligned} \operatorname{exp}\left[-\operatorname{exp}\left(-\frac{z_T - \mu}{\sigma}\right)\right] &= 1 - \frac{1}{T} \\ -\operatorname{exp}\left(-\frac{z_T - \mu}{\sigma}\right) &= \ln\!\left(1 - \frac{1}{T}\right) \\ \frac{z_T - \mu}{\sigma} &= -\ln\!\left[-\ln\!\left(1 - \frac{1}{T}\right)\right] \\ z_T &= \mu - \sigma \ln\!\left[-\ln\!\left(1 - \frac{1}{T}\right)\right] \end{aligned} \]
For large \(T\), \(\ln(1 - 1/T) \approx -1/T\) (the first-order Taylor expansion of \(\ln(1 - x)\) around \(x = 0\)), so \(z_T \approx \mu + \sigma \ln T\). On an exam, I would give you this approximation rather than ask you to derive it. The return level grows logarithmically with \(T\), which is why the Gumbel return-level plot is linear in \(\log T\).
Part b.
z₁₀₀ = 119 mm
Check against the Distributions.jl quantile:
What a climate shift does to a return period
A climate model projects that by 2050 the location parameter of the Gumbel in the previous problem increases to \(\mu = 65\text{ mm}\), while the scale stays at \(\sigma = 15\text{ mm}\).
- Find the new return period for the exact design storm you computed above.
- A 15 mm shift equals the scale parameter \(\sigma\). Is the change in return period proportionally small or large?
Part a. Plug \(z = z_{100}\) into the new CDF and solve for \(T\).
\[ \begin{aligned} F(z_{100}) &= \operatorname{exp}\left[-\operatorname{exp}\left(-\frac{z_{100} - 65}{15}\right)\right] \\ T_{\text{new}} &= \frac{1}{1 - F(z_{100})} \end{aligned} \]
z₁₀₀ = 119 mm, new exceedance prob = 0.0269, new T = 37.1 years
Part b. A 30 percent shift in the location turns a 100-year event into roughly a 37-year event. The relationship is exponential, not proportional: the return level grows as \(\mu + \sigma \ln T\), so a linear shift in \(\mu\) produces an exponential change in \(T\). A small change in the mean of the extremes produces a large change in the recurrence of a fixed design level.
The three GEV families
The GEV CDF is
\[ F(z) = \operatorname{exp}\left\{-\left[1 + \xi\left(\frac{z - \mu}{\sigma}\right)\right]^{-1/\xi}\right\} \]
where the expression inside the brackets must be positive.
- What happens to the support of the distribution when \(\xi > 0\)? When \(\xi < 0\)?
- The exponential distribution has CDF \(F(x) = 1 - e^{-x/\beta}\) for \(x \ge 0\). If \(M_n = \max(X_1, \ldots, X_n)\) with the \(X_i\) independent exponential, then \(F_{M_n}(z) = (1 - e^{-z/\beta})^n\). Show that with the normalizing constants \(a_n = \beta\) and \(b_n = \beta \ln n\), the limit of \((M_n - b_n)/a_n\) is a Gumbel.
- The uniform distribution on \([0, 1]\) has CDF \(F(x) = x\). What is the exact distribution of \(M_n\)? Which GEV family does it belong to? (This one requires normalizing constants from Coles (2001) §3.1.5.)
Part a. When \(\xi > 0\) the distribution has a lower bound at \(\mu - \sigma/\xi\) and no upper bound. The tail is heavy (Fréchet type). When \(\xi < 0\) the distribution has no lower bound but has an upper endpoint at \(\mu - \sigma/\xi\). The tail is bounded (Weibull type).
Part b. The CDF of \((M_n - b_n)/a_n\) is
\[ \begin{aligned} P\!\left(\frac{M_n - \beta\ln n}{\beta} \le z\right) &= P(M_n \le \beta z + \beta\ln n) \\ &= \left(1 - e^{-(z + \ln n)}\right)^n \\ &= \left(1 - \frac{e^{-z}}{n}\right)^n \end{aligned} \]
As \(n \to \infty\),
\[ \left(1 - \frac{e^{-z}}{n}\right)^{n} \to \operatorname{exp}(-e^{-z}) \]
which is the Gumbel CDF. The exponential parent falls in the Gumbel domain (\(\xi = 0\)).
Part c. \(P(M_n \le z) = z^n\) for \(z \in [0, 1]\). The maximum of \(n\) uniforms on \([0, 1]\) has a Beta(\(n\), 1) distribution, bounded above at 1.
With \(a_n = 1/n\) and \(b_n = 1\):
\[ P\!\left(\frac{M_n - 1}{1/n} \le z\right) = \left(1 + \frac{z}{n}\right)^n \to e^{z} \]
for \(z \le 0\), which is the reverse Weibull (\(\xi = -1\)). This matches the GEV CDF with \(\xi = -1\), \(\mu = -1\), and \(\sigma = 1\): \(\operatorname{exp}\{-[1 + (-1)(z - (-1))]^{1}\} = e^z\). The bounded parent gives a bounded extreme value distribution.
Sampling variability in the tail
You fit a GEV to 60 years of annual block maxima. The return-level plot projects a straight line (the Gumbel domain), but the three largest historical observations fall noticeably below it.
Give two distinct explanations for why the highest empirical points might fall below the fitted curve. Which is more likely given a 60-year record?
Model misspecification. The true tail could be lighter than Gumbel (Weibull domain, \(\xi < 0\)), so the fitted straight line overestimates the tail. The curve should flatten, and the largest observations sit where the flat part begins.
Sampling variability. The confidence band on the return-level curve is widest in the tail. With 60 years of data, the largest observation gets a plotting-position return period of about 61 years, and a 100-year or 200-year event may simply not have occurred yet. The three largest points falling below the line is well within what random chance produces.
With only 60 observations, sampling variability is the more likely explanation. The empirical points in the tail are placed by their rank, not by the distribution, so a short record that happened not to see a very large event puts all its points below the extrapolated line. Concluding that the model class is wrong requires longer records or external evidence.
Block size and convergence
The extremal types theorem says block maxima converge to a GEV as the block size grows. In practice, annual blocks are standard.
- Name two reasons annual blocks are used rather than, say, monthly or weekly blocks.
- If you used 5-year blocks instead of 1-year blocks, how many block maxima would 60 years of data give you? What would happen to the confidence interval on the 100-year return level?
Part a. Annual blocks remove seasonality: the distribution of summer and winter maxima differ, and mixing them in a monthly block violates the identical-distribution assumption. Annual blocks also match the convention that return periods and design standards are stated in years.
Part b. Sixty years gives 12 five-year maxima. Larger blocks bring the data closer to the GEV limit (less bias from the finite-sample approximation), but 12 observations carry far more sampling variability than 60. The confidence interval on the 100-year level would widen substantially.
This is the bias-variance tradeoff of block size: small blocks give more data but a worse approximation to the limit distribution, and large blocks give fewer data but a better approximation. Annual blocks are a pragmatic default, not a principled optimum.