This commit is contained in:
34
docs/plans/LPV2-PREREG-2026-05-19.md
Normal file
34
docs/plans/LPV2-PREREG-2026-05-19.md
Normal file
@@ -0,0 +1,34 @@
|
||||
# Pre-Registration — LPV-2 (calibrated solvable-yet-wasteful corpus)
|
||||
|
||||
**Date:** 2026-05-19 PM · **Status:** LOCKED before spend · **Repo @** `cebb25d`
|
||||
**User-initiated** ("experiment"). Builds on `LIVE-PREMIUM-VALIDATION-PREREG` @ `d628120`. Same discipline: strict pre-reg, no revisit, verbatim verdict.
|
||||
|
||||
## 0. Why — the bracketing told us exactly what to fix
|
||||
|
||||
| Experiment | Corpus | Failure mode |
|
||||
|---|---|---|
|
||||
| R6 | clean linear | optimal path short → **no waste** → ~0% measurable |
|
||||
| LPV-B | heavy decoys, circular | baseline **could not solve** task_a → no pair |
|
||||
|
||||
Both missed the Goldilocks: a workload the baseline **solves (PASS)** but only after **recoverable wasted exploration** a distilled skill can front-load. LPV-2 changes **only the corpus calibration** toward that window.
|
||||
|
||||
## 1. The single calibrated change (vs LPV-B `cebb25d`)
|
||||
|
||||
- Decoys per pipeline **4 → 2** (waste exists, but not a maze).
|
||||
- **No circular/dead-end decoy chains** (LPV-B's `see also → also deprecated → dead end` trap is what made it unsolvable). Decoys are single-hop and obviously inert once read.
|
||||
- The `loader.ts` indirection **stays** (this is the intended floundering: a fresh agent must discover loader.ts is the source of truth and that `see also:` is noise — exactly what a distilled skill front-loads).
|
||||
- The ACTIVE chain is **clean once on it** (LPV-B kept that; retained).
|
||||
|
||||
Net intended profile: baseline ≈ solvable in ~9-12 tool calls (PASS) with ~3-6 of those wasted on discovery/decoys; a skill encoding "loader.ts ACTIVE-only; ignore see-also; follow next:" → treatment skips the discovery → measurable reduction if the Hermes effect is real.
|
||||
|
||||
## 2. UNCHANGED (no goalpost move)
|
||||
|
||||
Metric (tool-calls to grader-correct), success bar (median reduction ≥0.40 AND one-sided sign test p<0.05), model (`anthropic/claude-sonnet-4.6`), arms (R6 two-phase distill), grader (required-facts), N (3 pilot → T3 gate → 20), caps ($5 pilot / $38 combined hard), outcomes, anti-p-hacking — **all identical to R6/LPV §§2-3,5-8**. Only §1's corpus calibration differs. Harness = same proven vehicle, env `LPV2=1`.
|
||||
|
||||
## 3. Pre-registered outcomes + the binding stop clause
|
||||
|
||||
- **PROVEN** (escalated, median ≥0.40 & p<0.05) — the Hermes ~40% reproduced on a solvable-yet-wasteful workload; D1's carve-out converts to a demonstrated payoff.
|
||||
- **NOT-PROVEN** (escalated, bar not met) — payoff real-but-below-40% or absent; reported honestly with effect size.
|
||||
- **INCONCLUSIVE-STOPPED** (pilot gate fails) — and **this is the third pre-registered synthetic attempt**. Per no-revisit, an INCONCLUSIVE here is **binding**: synthetic-corpus calibration is empirically not the path; a credible ~40% demonstration requires a *real engineering-task* corpus, which is a separate user-scoped study. **No further autonomous corpus recalibration** — that would be the goalpost-moving the discipline forbids.
|
||||
|
||||
This commit is the contract hash.
|
||||
Reference in New Issue
Block a user