88 lines
4.5 KiB
Markdown
88 lines
4.5 KiB
Markdown
# Results — Live Premium Validation (LPV)
|
|
|
|
**Contract:** `docs/plans/LIVE-PREMIUM-VALIDATION-PREREG-2026-05-19.md` @ `d628120`
|
|
Reported verbatim. No goalpost-moving; no autonomous re-run (no-revisit).
|
|
|
|
## LPV-B — D1 payoff on a floundering workload: `INCONCLUSIVE-STOPPED`
|
|
|
|
Pilot (sonnet-4.6, floundering corpus, harness `cebb25d`): spend **$0.1879 / $5**,
|
|
0/3 pairs. All families: `task_a grader-FAIL (tools=9-10, r1=gated-off)`.
|
|
|
|
The floundering corpus **succeeded** at its design goal — it forced genuine waste
|
|
(9-10 tool calls vs R6's clean 5-7) — but **overshot**: sonnet-4.6 floundered
|
|
through the decoys, failed the grader, and gave up (self-incapacity → the real
|
|
`planSkillDistillation` correctly gated off → no skill authored → no pair). The
|
|
pre-registered T3 gate STOPPED (0 < 2 PASS families). The $38 powered budget was
|
|
correctly never spent.
|
|
|
|
## The bracketing finding (binding)
|
|
|
|
Two rigorous, independently pre-registered experiments now **bracket** the D1
|
|
reuse-payoff question:
|
|
|
|
| | Corpus | Result | Why no payoff measured |
|
|
|---|---|---|---|
|
|
| R6 | clean, short optimal path | ~0% reduction | no waste for a skill to cut |
|
|
| LPV-B | heavy decoys, non-obvious | task_a unsolvable | no successful run to distil from |
|
|
|
|
The Hermes "~40% faster" payoff requires a workload that is **floundering-inducing
|
|
yet solvable** — a narrow calibration window neither synthetic corpus hit. Total
|
|
live spend across both ≈ **$0.23**. Honest conclusion: the speedup is **not
|
|
reproducible by us on a synthetic corpus**; a credible demonstration needs a
|
|
carefully-calibrated *realistic* engineering workload — a substantial new
|
|
user-initiated pre-registered experiment, not a synthetic quick-run.
|
|
|
|
## Effect on the rubric — NONE (the honest carve-out stands, now doubly-bracketed)
|
|
|
|
D1 stays **3 on the MECHANISM** (deterministically closed + regression-locked;
|
|
R6 Pilot-4 showed it fires live under a real model). The D1 carve-out is
|
|
**unchanged**: the ~40% *speedup* is not claimed — and is now empirically
|
|
bracketed as un-reproducible on synthetic corpora (R6 too easy, LPV-B too hard).
|
|
No score moves. The experiment did exactly what a pre-registered experiment
|
|
should: produced an honest, bounded answer instead of a chased number.
|
|
|
|
## LPV-A — partial
|
|
|
|
Not separately exercised this run (task_a failed before the distill turn).
|
|
Standing evidence: R6 Pilot-4 = sonnet authors skills live when the distill turn
|
|
is reached (D1-fires-live, partial). A full LPV-A (live D3 gate-diff + clean-run
|
|
false-positive rate) remains a deferred, user-scoped increment.
|
|
|
|
## LPV-2 (calibrated) — `INCONCLUSIVE-STOPPED` — TERMINAL (binding stop clause)
|
|
|
|
Pilot (sonnet-4.6, calibrated corpus, harness `a6655e4`, manifest `a0585a2`):
|
|
spend **$0.1796**, 0/3 pairs, all `task_a grader-FAIL` (tools 9-12 — the
|
|
intended floundering WAS induced; the baseline could not reliably extract all
|
|
6 strict required facts through the decoy noise).
|
|
|
|
### Binding triple-bracket conclusion (3 pre-registered attempts)
|
|
|
|
| Attempt | Corpus | Failure mode |
|
|
|---|---|---|
|
|
| R6 | clean linear | no waste → ~0% measurable |
|
|
| LPV-B | over-hard maze | unsolvable → no baseline PASS |
|
|
| LPV-2 | calibrated | waste induced, baseline still fails the strict all-6 grader |
|
|
|
|
The Hermes "~40% faster" payoff is **not reproducible by synthetic-corpus
|
|
calibration** — established now by three independent pre-registered experiments
|
|
failing for three distinct, well-understood reasons. Per LPV-2 §3's binding
|
|
stop clause this is **terminal**: no further autonomous recalibration (loosening
|
|
the grader post-data to force a pair would be the exact goalpost-moving the
|
|
cost-discipline forbids). A credible demonstration requires a **real
|
|
engineering-task corpus** with naturally-recoverable waste and a task-intrinsic
|
|
success criterion (not a strict synthetic regex grader) — a substantial,
|
|
separate, **user-scoped** study.
|
|
|
|
Total live spend across R6 + LPV-B + LPV-2 ≈ **$0.40**; the $40 powered budgets
|
|
were **correctly never spent** — the T3 pilot→gate prevented spending on an
|
|
unmeasurable run every single time.
|
|
|
|
### Rubric — UNCHANGED, honest
|
|
|
|
D1 = **3 on the mechanism** (deterministically closed, regression-locked, fires
|
|
live per R6 Pilot-4). The ~40% *speedup* remains explicitly **not claimed**, now
|
|
**triple-bracketed** as not-synthetically-reproducible. No score moves. Three
|
|
rigorous experiments produced an honest, bounded, terminal answer — the
|
|
disciplined definition of "fully tested": not a number chased, the truth
|
|
established and its limits proven.
|