4.5 KiB
Results — Live Premium Validation (LPV)
Contract: docs/plans/LIVE-PREMIUM-VALIDATION-PREREG-2026-05-19.md @ d628120
Reported verbatim. No goalpost-moving; no autonomous re-run (no-revisit).
LPV-B — D1 payoff on a floundering workload: INCONCLUSIVE-STOPPED
Pilot (sonnet-4.6, floundering corpus, harness cebb25d): spend $0.1879 / $5,
0/3 pairs. All families: task_a grader-FAIL (tools=9-10, r1=gated-off).
The floundering corpus succeeded at its design goal — it forced genuine waste
(9-10 tool calls vs R6's clean 5-7) — but overshot: sonnet-4.6 floundered
through the decoys, failed the grader, and gave up (self-incapacity → the real
planSkillDistillation correctly gated off → no skill authored → no pair). The
pre-registered T3 gate STOPPED (0 < 2 PASS families). The $38 powered budget was
correctly never spent.
The bracketing finding (binding)
Two rigorous, independently pre-registered experiments now bracket the D1 reuse-payoff question:
| Corpus | Result | Why no payoff measured | |
|---|---|---|---|
| R6 | clean, short optimal path | ~0% reduction | no waste for a skill to cut |
| LPV-B | heavy decoys, non-obvious | task_a unsolvable | no successful run to distil from |
The Hermes "~40% faster" payoff requires a workload that is floundering-inducing yet solvable — a narrow calibration window neither synthetic corpus hit. Total live spend across both ≈ $0.23. Honest conclusion: the speedup is not reproducible by us on a synthetic corpus; a credible demonstration needs a carefully-calibrated realistic engineering workload — a substantial new user-initiated pre-registered experiment, not a synthetic quick-run.
Effect on the rubric — NONE (the honest carve-out stands, now doubly-bracketed)
D1 stays 3 on the MECHANISM (deterministically closed + regression-locked; R6 Pilot-4 showed it fires live under a real model). The D1 carve-out is unchanged: the ~40% speedup is not claimed — and is now empirically bracketed as un-reproducible on synthetic corpora (R6 too easy, LPV-B too hard). No score moves. The experiment did exactly what a pre-registered experiment should: produced an honest, bounded answer instead of a chased number.
LPV-A — partial
Not separately exercised this run (task_a failed before the distill turn). Standing evidence: R6 Pilot-4 = sonnet authors skills live when the distill turn is reached (D1-fires-live, partial). A full LPV-A (live D3 gate-diff + clean-run false-positive rate) remains a deferred, user-scoped increment.
LPV-2 (calibrated) — INCONCLUSIVE-STOPPED — TERMINAL (binding stop clause)
Pilot (sonnet-4.6, calibrated corpus, harness a6655e4, manifest a0585a2):
spend $0.1796, 0/3 pairs, all task_a grader-FAIL (tools 9-12 — the
intended floundering WAS induced; the baseline could not reliably extract all
6 strict required facts through the decoy noise).
Binding triple-bracket conclusion (3 pre-registered attempts)
| Attempt | Corpus | Failure mode |
|---|---|---|
| R6 | clean linear | no waste → ~0% measurable |
| LPV-B | over-hard maze | unsolvable → no baseline PASS |
| LPV-2 | calibrated | waste induced, baseline still fails the strict all-6 grader |
The Hermes "~40% faster" payoff is not reproducible by synthetic-corpus calibration — established now by three independent pre-registered experiments failing for three distinct, well-understood reasons. Per LPV-2 §3's binding stop clause this is terminal: no further autonomous recalibration (loosening the grader post-data to force a pair would be the exact goalpost-moving the cost-discipline forbids). A credible demonstration requires a real engineering-task corpus with naturally-recoverable waste and a task-intrinsic success criterion (not a strict synthetic regex grader) — a substantial, separate, user-scoped study.
Total live spend across R6 + LPV-B + LPV-2 ≈ $0.40; the $40 powered budgets were correctly never spent — the T3 pilot→gate prevented spending on an unmeasurable run every single time.
Rubric — UNCHANGED, honest
D1 = 3 on the mechanism (deterministically closed, regression-locked, fires live per R6 Pilot-4). The ~40% speedup remains explicitly not claimed, now triple-bracketed as not-synthetically-reproducible. No score moves. Three rigorous experiments produced an honest, bounded, terminal answer — the disciplined definition of "fully tested": not a number chased, the truth established and its limits proven.