VALIDATION HARNESS

KOKON · audit · 5 synthetic companies × ground-truth manifests · engine self-audits on every load · Phase 524.23
DISCLOSURE · Self-validated against internally frozen ground truth. Status: regression mechanism, not external replication. The 5 fixtures + 44 manifest rows were authored by the engine team and live in internal/handlers/testdata/validation-fixtures/. Independent third-party replication is tracked separately and is NOT what this page proves. What this page DOES prove: every commit must keep the engine reconciled with the rules the team itself wrote down.
cumulative in-scope PASS

What this page does

Embedded in the binary are 5 synthetic companies (Northwind, Aurora, Apollo, Riverview, Pristine) with their seed-pinned engine payloads and ground-truth manifests. Every load runs all 5 through the live engine via the same RealDispatcher the diagnostic endpoint uses, then scores each row against the manifest. If a Phase 524 closure regresses, this page goes red.

Severity tolerance: exact (case-insensitive). Dollar tolerance: ±20%. Process-mining detectors silently no-fire when the payload supplies no process.* inputs — that is the correct restraint outcome.

Running engine on 5 fixtures…

FP-HELL · FORENSIC ADVERSARIAL BENCHMARK
SOURCE · Frozen offline sweep, not a live re-run. Live re-validation requires porting score_fphell.py ground-truth regeneration into the binary (tracked separately). This panel ships the evidence the engine produced on the recorded commit.
Fires the embedded seed-13 fixture (~6 MB, ~30 s) then opens /ui/forensic-case with the cytoscape graph rendered.

Loading FP-HELL scorecard…

FP-HELL · PER-AXIS DEFENSE BREAKDOWN
For each of the 8 adversarial axes: what poison was planted, what KOKON did, total violations across the 14-seed sweep. Companion to reports/fphell-axis-analysis-all-seeds.md + reports/fphell-axis-analysis-seed13.md in the repo.

Loading per-axis breakdown…

PAF · HEALTH (kill-rate, CI gates)
Pre-Adversarialized Findings (Phase 590.1-590.7) layer health. Plan §9 global invariant: red-team kill-rate must be > 0 on a non-trivial batch. Self-test corpus = 5 in-package planted fixtures (TBML / Sanctions / Structuring strong-true + planted-FP). If gate kill-rate fires FAIL, the v1 attack catalog or §A7 ruleset has regressed and PAF releases must be suspended.

Loading PAF health…

FORENSIC CALIBRATION · PER-FINDING-TYPE FP-RATE
Per-finding-type empirical FP-rate, Brier score, log-loss, and Wilson 95% CI computed from the recorded forensic outcome ledger (Phase 580). Zero-outcome freeze: per-type stats surface only at N_min = 20 matched observations (confirmed_scheme + innocent_collision). Between 20 and N_target = 250 the stats are tagged preliminary. Same release-management discipline as Phase 512 agent calibration.

Loading forensic calibration…