Step 1: Define the forecast question and evaluation window
What to do: Specify what is being forecast, the forecast horizon, the actual outcome used for comparison, and the date at which each prediction was made. Separate in-sample, out-of-sample, and lagged-data evaluations.
What success looks like: Every forecast has a known target, period, origin date, and comparable actual result.
Common mistake to avoid: Do not evaluate a prediction with information that was unavailable when the prediction was issued.
Step 2: Recompute and trace the important numbers
What to do: Independently recompute key totals, rates, changes, residuals, and scenario outputs. Trace each figure to the source file, row, field, and reference used for verification.
What success looks like: A reviewer can reproduce each material number without relying on an unexplained output.
Common mistake to avoid: Do not validate only the final chart while leaving source extraction and intermediate calculations unchecked.
Step 3: Calculate multiple accuracy metrics
What to do: Calculate MAE, RMSE, bias or mean error, maximum error, error variance, and MAPE where zero values are not a concern. Add MASE when a comparison with a naïve benchmark is useful.
What success looks like: The report distinguishes average error, large misses, systematic direction, and instability.
Common mistake to avoid: Do not treat R² or a single average metric as a complete description of forecast quality.
Step 4: Segment errors by regime and time
What to do: Compare calm periods, shocks, recoveries, and structural breaks. Review residual distributions, rolling metrics, feature-space shifts, and the largest outliers.
What success looks like: You can identify when the model works, when it deteriorates, and which conditions produce the largest misses.
Common mistake to avoid: Do not allow strong performance in one stable period to conceal failure during a different operating regime.
Step 5: Test for lookahead bias and unstable relationships
What to do: Compare levels with differenced data, compare naive lookahead information with realistically lagged data, and review rolling correlations and coefficient stability.
What success looks like: Reported performance reflects the information and relationships that would genuinely have been available at prediction time.
Common mistake to avoid: Do not interpret trending levels or contemporaneous variables as evidence of reliable predictive power.
Step 6: Run baseline and stress scenarios
What to do: Change the assumptions most likely to affect the result, such as interest rates, operating performance, occupancy, inflation, or policy conditions. Record the effect on cash flow, coverage, break-even points, and cumulative outcomes.
What success looks like: Decision-makers can see the downside, the baseline, the recovery path, and the assumptions responsible for each variance.
Common mistake to avoid: Do not present a baseline forecast without quantifying how it changes under plausible adverse conditions.
Step 7: Document corrections and issue a verdict
What to do: Record validated errors, corrections, unresolved exceptions, evidence, and the final pass/fail result. Preserve the audit trail so the review can be repeated.
What success looks like: The final deliverable is complete, cited, reproducible, and understandable to a non-specialist reviewer.
Common mistake to avoid: Do not correct a number without recording why it changed and which source supports the correction.