Files
waggle-os/docs/plans/LPV2-PREREG-2026-05-19.md
Oleg Maslov 0c3e2ead3b
Some checks failed
Installer Smoke / installer-smoke (push) Has been cancelled
moving
2026-09-02 10:10:29 +02:00

2.9 KiB

Pre-Registration — LPV-2 (calibrated solvable-yet-wasteful corpus)

Date: 2026-05-19 PM · Status: LOCKED before spend · Repo @ cebb25d User-initiated ("experiment"). Builds on LIVE-PREMIUM-VALIDATION-PREREG @ d628120. Same discipline: strict pre-reg, no revisit, verbatim verdict.

0. Why — the bracketing told us exactly what to fix

Experiment Corpus Failure mode
R6 clean linear optimal path short → no waste → ~0% measurable
LPV-B heavy decoys, circular baseline could not solve task_a → no pair

Both missed the Goldilocks: a workload the baseline solves (PASS) but only after recoverable wasted exploration a distilled skill can front-load. LPV-2 changes only the corpus calibration toward that window.

1. The single calibrated change (vs LPV-B cebb25d)

  • Decoys per pipeline 4 → 2 (waste exists, but not a maze).
  • No circular/dead-end decoy chains (LPV-B's see also → also deprecated → dead end trap is what made it unsolvable). Decoys are single-hop and obviously inert once read.
  • The loader.ts indirection stays (this is the intended floundering: a fresh agent must discover loader.ts is the source of truth and that see also: is noise — exactly what a distilled skill front-loads).
  • The ACTIVE chain is clean once on it (LPV-B kept that; retained).

Net intended profile: baseline ≈ solvable in ~9-12 tool calls (PASS) with ~3-6 of those wasted on discovery/decoys; a skill encoding "loader.ts ACTIVE-only; ignore see-also; follow next:" → treatment skips the discovery → measurable reduction if the Hermes effect is real.

2. UNCHANGED (no goalpost move)

Metric (tool-calls to grader-correct), success bar (median reduction ≥0.40 AND one-sided sign test p<0.05), model (anthropic/claude-sonnet-4.6), arms (R6 two-phase distill), grader (required-facts), N (3 pilot → T3 gate → 20), caps ($5 pilot / $38 combined hard), outcomes, anti-p-hacking — all identical to R6/LPV §§2-3,5-8. Only §1's corpus calibration differs. Harness = same proven vehicle, env LPV2=1.

3. Pre-registered outcomes + the binding stop clause

  • PROVEN (escalated, median ≥0.40 & p<0.05) — the Hermes ~40% reproduced on a solvable-yet-wasteful workload; D1's carve-out converts to a demonstrated payoff.
  • NOT-PROVEN (escalated, bar not met) — payoff real-but-below-40% or absent; reported honestly with effect size.
  • INCONCLUSIVE-STOPPED (pilot gate fails) — and this is the third pre-registered synthetic attempt. Per no-revisit, an INCONCLUSIVE here is binding: synthetic-corpus calibration is empirically not the path; a credible ~40% demonstration requires a real engineering-task corpus, which is a separate user-scoped study. No further autonomous corpus recalibration — that would be the goalpost-moving the discipline forbids.

This commit is the contract hash.