Kestrel Economics · Bulletin Track record

Track record

This is the running calibration record. The publication compounds: each issue freezes a dated, scoreable forecast, and over time those forecasts build into a public, multi-year record of whether the model has been right. We publish it whether it flatters the model or not.

Where the record stands today

What this issue puts on the record

Calibration by construction (the evidence today)

Before release, the model is tested against data where the truth is known. It is asked to recover parameters from synthetic data it has not seen, its rank calibration is checked across hundreds of simulated datasets (simulation-based calibration), and its outcome mix is checked against history within each state. The current serial passes these gates. This shows the machinery is sound: when the world behaves as the model assumes, the model recovers the truth at the rates it claims. It is not a realised hit-rate on decided permits. That is what the forward record below will provide.

For the technically inclined

The fit behind this issue (2026-07-04 vintage): zero divergent transitions, maximum R-hat 1.0034, minimum bulk effective sample size 1,466, minimum tail effective sample size 1,210, minimum energy diagnostic 0.74. At the model's release, simulation-based calibration over 200 replications, all successful, showed summary rank uniformity consistent with calibration (Bonferroni-corrected p = 1.0), with 3 of 146 scalar parameters rejecting rank uniformity, below the 5 percent expected by chance; parameters drawn from the prior were recovered within their posterior credible intervals, including the merged not-approved baselines, and posterior predictive checks reproduced the observed outcome mix within each state. These figures come from the model card shipped with this issue; the card is the source of truth for them.

The forward record (accruing)

Only files with a clock running will ever appear in the forward record. A stalled file produces no scoreable timing forecast, since it is reported as a status rather than given a decision date, so the scatter and coverage table below cover the forecasted files only, never the full pending cohort.

Forecast vs realised time

Each dot will be one (permit, vintage) pair where the model produced a forecast and the permit has since reached a terminal decision. The horizontal bar is the model's 80% credible interval on time-to-decision (the same interval published on the ladder); the y-coordinate is the day the decision actually landed. A perfectly calibrated forecast sits on the diagonal y = x, with credible bars that cover the diagonal 80% of the time. The time axis is days from the forecast vintage to decision: the clock starts the day the data was frozen, which is the clock every published range is on.

80% interval coverage

Once decisions land, this table shows how often the model's nominal time windows actually held the realised decision date, computed directly from the scored records above. The model is calibrated when empirical coverage tracks nominal. Early readings carry a caveat the table cannot show: the first files to decide are the fast end of the queue, so the earliest rows over-represent surprises on the early side; the reading firms up as the cohort closes.

Outcome calibration

For each of the two terminal outcomes, the model's predicted probability is binned, and the y-axis shows the realised share of permits in that bin that ended in the outcome. The vertical bars are Wilson 95% bands, which mark the finite-sample noise on each share. A perfectly calibrated model sits on the y = x diagonal in every panel.

Calibration by lead time

We publish a fresh forecast for the same permit at many points in its life. When we score, we stratify by lead time (how far ahead of the decision each forecast was made) rather than pool every weekly forecast as an independent observation. Pooling would inflate the apparent evidence, since long-pending permits would dominate, and it would hide the lead-time behaviour that is the actual product. Once enough decisions have landed, this section reports coverage and timing error separately for forecasts made roughly 0 to 14, 15 to 60, 61 to 180, and 180-plus days before the decision, so each forecast can be judged at the lead time that matches a real decision horizon.

How the record compounds, and how we report it

Each issue freezes a dated, scoreable forecast. As live-cohort permits close, proper scores, calibration, and 80% coverage fill a forward record that deepens issue over issue. Two standing conventions bind every number here: