Step 1: Define the audit scope
What to do: Identify the output being audited, the source documents it used, the important numbers and assertions it contains, and the consequences of an error.
What success looks like: Another reviewer can see exactly what is included, excluded, and considered material.
Common mistake to avoid: Treating every sentence as equally important instead of prioritizing decisions, calculations, and claims that must be defensible.
Step 2: Establish an independent reviewer
What to do: Use a second agent or separate review process that did not produce the original answer, so it can challenge the work without inheriting its assumptions.
What success looks like: The audit has a fresh basis for checking the deliverable rather than simply restating the first answer.
Common mistake to avoid: Asking the same agent to confirm its own work without introducing independent checks.
Step 3: Recompute the numbers
What to do: Recalculate totals, rates, comparisons, transformations, and other material figures from the available source data. For financial or operational work, compare the recomputed result with the number in the final deliverable.
What success looks like: Important numbers agree, or every difference is explicitly explained and classified.
Common mistake to avoid: Accepting a plausible-looking total without checking the underlying rows, filters, units, or formulas.
Step 4: Trace claims to source evidence
What to do: Connect each material claim to the exact source file, row, field, page, or reference that supports it. Record the chain so the result can be reviewed later.
What success looks like: A reviewer can move from a reported number back to its source without relying on memory or an opaque model explanation.
Common mistake to avoid: Citing a whole document when the specific field or row supporting the claim has not been identified.
Step 5: Test model and timing risk
What to do: Check whether the model behaves differently across periods, whether lookahead information has entered the analysis, and whether strong relationships are merely shared trends. The supplied macro diagnostics show why this matters: an in-sample levels R² of 98.1% fell to 19.0% after differencing, while realistic lagged-data R² was 18.6% compared with 80.0% for a naive lookahead specification.
What success looks like: The report distinguishes genuine predictive evidence from regime-specific fit, timing leakage, or spurious regression.
Common mistake to avoid: Treating one impressive aggregate metric as proof that the model is reliable in every period.
Step 6: Resolve exceptions and issue a verdict
What to do: Review failed checks, fix errors where possible, attach the evidence, and return a clear pass or fail result. If the same correction will recur, turn it into a reusable audit rule.
What success looks like: Stakeholders know what passed, what failed, what changed, and which evidence supports the decision.
Common mistake to avoid: Quietly editing the output without retaining the original finding and reason for the correction.