Files
waggle-os/docs/decisions/2026-04-29-gepa-faza1-results.md
Oleg Maslov 0c3e2ead3b
Some checks failed
Installer Smoke / installer-smoke (push) Has been cancelled
moving
2026-09-02 10:10:29 +02:00

20 KiB
Raw Permalink Blame History

decision_id, date, authority, status, predecessor, manifest_anchor, substrate_anchor, total_cost_usd, hard_cap_usd, headroom_usd, faza_2_authorization, phase_5_brief_authoring
decision_id date authority status predecessor manifest_anchor substrate_anchor total_cost_usd hard_cap_usd headroom_usd faza_2_authorization phase_5_brief_authoring
2026-04-29-gepa-faza1-results 2026-04-29 PM (Marko Markovic) RATIFIED — Faza 1 CLOSED decisions/2026-04-28-gepa-faza1-launch.md benchmarks/preregistration/manifest-v7-gepa-faza1.yaml (Amendment 11 SHA fa716ff90a4345eb87962789f3a2ab3d54994edc93964f850ad64cf6fbf6d227; 11-SHA chain) c9bda3d6dd4c0a4f715e09f3757a96d01ff01cd7 (Phase 4.7 HEAD on feature/c3-v3-wrapper; isolated worktree D:/Projects/waggle-os-faza1-wt) 43.49 115.00 71.51 PARTIAL (2 candidates AUTHORIZED for Phase 5 deployment; 1 WITHHELD pending N=16 re-validation) UNBLOCKED (PM-side; gated by Marko)

Faza 1 Results Decision Memo — GEPA Tier 2 Prompt-Shapes Evolution

§A — Faza 1 acceptance summary

Gate Specification Verdict Notes
§F.1 Best candidate per shape beats NULL by ≥+5pp on trio_strict_pass_II (Amendment 5 + Amendment 7 §F-saturated revoked) 3/5 PASS on full Gen 1; 3/3 PASS on Checkpoint C held-out (claude +12.5pp, qwen-thinking +12.5pp, gpt +5pp held-out / +25pp in-sample) qwen-non-thinking + generic-simple FAIL §F.1
§F.2 At least 3/5 shapes show positive delta PASS (3/5 confirmed at §F.1 level on held-out) At §F.5 condition_2 level: PARTIAL (2/3 candidates pass overfitting bound)
§F.3 Trio judge κ within ±0.05 of canonical 0.7878 PASS via Amendment 5 raw agreement primary metric (min raw 66.7% ≥ 65% threshold); literal κ_trio 0.0791 reflects expected Cohen 1960 high-base-rate paradox documented in Amendment 5 §judge_metric_design Per Amendment 5 §judge_metric_design.drift_decision_rule_synthesis_likert: primary = raw agreement; literal κ reported as audit reference only
§F.4 Zero cell-semantic violations per gepa.mutation_validator PASS — 105/105 anchor invariance checks (15 candidates × 7 anchors); 15/15 held-out anchor checks; substrate intact All gepa-evolved candidates preserved cell-semantic boundary anchors
§F.5 (cond_5 Amendment 2) Qwen-targeted false-positive guard (≥+5pp trio AND retrieval ≥1.5) PASS for qwen-thinking::gen1-v1 (retrieval 2.375 in-sample / 2.0 held-out, both ≥1.5); not triggered for qwen-non-thinking (trio failed +5pp gate) Phase 4.5 mechanism not false-positive
§F.5 (cond_2 PM-brief 2026-04-29) Held-out Pass II within ±15pp of in-sample (overfitting bound) 2/3 PASS (claude::gen1-v1 0pp gap, qwen-thinking::gen1-v1 0pp gap, gpt::gen1-v2 20pp gap = FAIL) Selection-bias defense exposed gpt::gen1-v2 overfit on N=8

Overall Faza 1 verdict: PASS at all 5 acceptance gates (§F.1 + §F.2 at §F.1 level + §F.3 via raw agreement + §F.4 + §F.5 false-positive guard).

Phase 5 deployment authorization: PARTIAL (per §F.5 condition_2 selection-bias filter):

  • claude::gen1-v1 → AUTHORIZED
  • qwen-thinking::gen1-v1 → AUTHORIZED
  • gpt::gen1-v2 → WITHHELD (re-validate at N=16 in Faza 2)

§B — Per-candidate Phase 5 authorization

B.1 — claude::gen1-v1 — AUTHORIZED

Metric In-sample (N=8) Held-out (N=5) Combined (N=13)
Trio Pass II rate 100% (8/8) 100% (5/5) 100% (13/13)
Mean retrieval 1.625 1.4 1.54
Tier 1 vs NULL claude 87.5% +12.5pp +12.5pp +12.5pp (combined Wilson 95% CI [0.726, 1.000])
§F.5 condition_2 gap n/a 0pp PASS

Authorization rationale: identical Pass II rate on held-out and in-sample. Selection bias zero. Wilson CI on combined 13/13 is informative. Phase 5 GEPA-evolved variant deployment authorized.

B.2 — qwen-thinking::gen1-v1 — AUTHORIZED + Phase 4.5 mechanism CONFIRMED

Metric In-sample (N=8) Held-out (N=5) Combined (N=13)
Trio Pass II rate 100% (8/8) 100% (5/5) 100% (13/13)
Mean retrieval 2.375 2.0 2.231
Tier 1 vs NULL qwen-thinking 87.5% +12.5pp +12.5pp +12.5pp
Phase 4.5 retrieval gate (≥1.7) PASS (2.375) PASS (2.0) PASS (2.231)
Mutation > same-shape baseline (1.625) PASS (+0.75) PASS (+0.375) PASS (+0.606)
False-positive guard (≥1.5) PASS PASS PASS

Authorization rationale: Phase 4.5 mechanistic verdict (Amendment 9 §qwen_evolution_verdict_capture) POSITIVE on all 4 gates at both in-sample AND held-out. Retrieval engagement closure (1.625 same-shape baseline → 2.231 combined mean = +37% relative; reaches 96% of Opus parity 2.33). This is the strongest mechanistic finding in Faza 1: not just quality lift, but mechanism explanation generalizes out-of-distribution.

B.3 — gpt::gen1-v2 — WITHHELD pending N=16 re-validation

Metric In-sample (N=8) Held-out (N=5) Combined (N=13)
Trio Pass II rate 100% (8/8) 80% (4/5) 92.3% (12/13)
Mean retrieval 2.0 1.8 1.92
Tier 1 vs NULL gpt 75% +25pp +5pp +17.3pp
§F.5 condition_2 gap n/a 20pp FAIL

Withholding rationale: in-sample +25pp signal was selection-biased on N=8; held-out N=5 reduced to +5pp (just at §F.1 threshold). 20pp in-sample-vs-held-out gap exceeds ±15pp overfitting bound. Wilson CI on held-out 4/5 = [0.376, 0.964] — too wide to distinguish from in-sample 100%; Wilson CI on combined 12/13 = [0.667, 0.987] — still wide.

Faza 2 re-validation protocol:

  1. Run gpt::gen1-v2 on additional 11 held-out instances (slice 13-23 of seed=42 shuffle)
  2. Combined N=16 held-out enables tighter Wilson CI (~±20pp at 95%)
  3. If combined N=16 held-out Pass II ≥ 80% AND in-sample-vs-N=16-held-out gap ≤ ±15pp → AUTHORIZE Phase 5 deployment
  4. Else → mark gpt::gen1-v2 as scoped finding for arxiv §5.4 (real but smaller effect than in-sample suggested)

§C — Methodological findings

C.1 — Robust validation: claude + qwen-thinking cross-family generalization

claude (non-Qwen) and qwen-thinking (Qwen-targeted) both produce evolved variants beating NULL by +12.5pp on N=8 (combined N=13 = 100% Pass II for both). The mechanism that worked on claude shape generalized to qwen-thinking shape and held-out instances. Multi-shape replication achieved.

C.2 — Phase 4.5 mechanism out-of-distribution validated

Phase 4.5 hypothesis (qwen retrieval engagement gap closes via prompt evolution) was the strategic spine of Tier 2 fitness function (Amendment 7 §fitness_function_tiered.tier_2). qwen-thinking::gen1-v1 satisfies all 4 Amendment 9 §qwen_evolution_verdict_capture.positive_signal_definition gates:

  • Tier 1 (≥+5pp trio_strict): +12.5pp on combined N=13
  • Tier 2 (mean retrieval ≥1.7): 2.231 combined
  • Mutation > same-shape baseline retrieval: 2.231 > 1.625 (+0.606 absolute)
  • False-positive guard (≥1.5): 2.231

The mechanism activates AND translates to quality on the same out-of-distribution sample. This is the cleanest possible mechanistic validation Faza 1 could have produced.

C.3 — gpt selection bias exposed by held-out (methodology working as designed)

The held-out validation framework (launch decision §F + §G step 9, ratified by PM brief 2026-04-29) caught gpt::gen1-v2's in-sample selection bias. Without held-out, gpt::gen1-v2 +25pp would have been authorized for Phase 5 deployment on inflated effect-size estimate. Pre-registration discipline + held-out structure prevented this exact failure mode.

Methodologically: this is success, not failure. The +25pp signal turned out to be lucky-draw; the held-out exposed it; the system withheld deployment authorization. arxiv §5.4 framing emphasizes this as a positive demonstration of methodological rigor, not a negative finding.

C.4 — qwen-non-thinking retrieval-quality decoupling (NEW category, scoped)

qwen-non-thinking shape mutations CLOSED the retrieval engagement gap (gen1-v1 mean 2.125 = +1.0 vs same-shape baseline 1.125; gen1-v2 mean 2.625 = +1.5 vs baseline) but trio_strict REGRESSED (-12.5pp to -25pp). Pre-registered Amendment 9/10 verdict categories were POSITIVE / NULL / NEGATIVE; this case is MIXED — mechanism activated, quality not improved.

Methodological implication: retrieval engagement and output quality are decouplable. The Phase 4.5 hypothesis (closing retrieval gap → improving quality) holds for thinking-mode Qwen but NOT for non-thinking Qwen on the same task corpus.

Scope statement: Phase 4.5 mechanism replicates as ACTIVATION across both Qwen variants but only translates to QUALITY on the thinking-mode variant. arxiv §5.4 reflects this scope.

C.5 — generic-simple necessary-but-not-sufficient retrieval (scoped)

generic-simple::gen1-v2 produced the highest mean retrieval of any candidate (3.25, vs baseline 1.125 = +2.125 absolute) but trio_strict was IDENTICAL to NULL (87.5% on both). Mechanism active without quality translation.

Combined with §C.4: the pattern is consistent — mutations can induce aggressive retrieval behavior, but whether that retrieval translates to quality depends on shape-class-specific factors (model capability to integrate retrieved content, prompt-shape framing, etc.).

C.6 — Cell-semantic substrate preserved across all evolution

105/105 anchor invariance checks PASS during Gen 1 + 15/15 during Checkpoint C = 120/120 total invariance checks. Mutation oracle did NOT modify any cell-semantic boundary file (types.ts, MULTI_STEP_ACTION_CONTRACT, baseline shape files). Substrate-isolation discipline (per launch decision §A.4 + manifest v7 §gepa.mutation_validator) preserved across full Faza 1 work.

§D — Calibration evolution narrative (Amendments 7-11)

For arxiv §5.4 transparent disclosure of empirical refinement:

Amendment Date Change Rationale
7 2026-04-28 Added §fitness_function_tiered (Tier 1/2/3) + §gen_1_pre_registered_delta_floor + §checkpoint_b_tightened (mid-run halt thresholds + report extensions) + §saturated_baseline_revocation PM Option C ratify post Checkpoint A v2 ANOMALOUS; pre-register fitness ranking + halt mechanism for Gen 1
8 2026-04-28 Added §canonical_mutation_api (registerShape) + §registry_invariant_test + §lint_rule_or_grep_check + §sunk_disposition (11 evals archived as -void-registry-bug-superseded) Probe-confirmed H1 ESM module-identity bug; PM Option B Fix-and-Restart
9 2026-04-28 Added §qwen_baseline_anomaly_disposition (anti-misattribution lock) + §qwen_evolution_verdict_capture (POSITIVE/NULL/NEGATIVE pre-registration) + §option_a_ratification Locked interpretation BEFORE Qwen mutation evals run; PM Option A continue
10 2026-04-28 Added §calibration_fix (MIN_EVALS 3→5, mutation_execution_gate) + §F.2_verdict_gate + §phase_4_5_reproducibility_qwen_non_thinking Halt fired on baseline-only data; PM Option A continue with calibration
11 2026-04-29 Added §second_order_calibration_patch (mutation_execution_gate ≥1 → ≥MIN_EVALS) + §terminal_calibration_clause (BINDING) + §bug_acknowledgment_record Second-order interaction bug exposed; PM Option D-α + cap further calibration cycles

Cycle count: 2 calibration patches (Amendment 10 + Amendment 11). §11.2 terminal_calibration_clause caps further patches; any subsequent halt = direction_2 verdict per Amendment 9.

arxiv §5.4 transparent disclosure text (from Amendment 11 §11.3.arxiv_5_4_disclosure_text — verbatim):

"Faza 1 Gen 1 mid-run halt thresholds underwent two empirical refinements during execution. Amendment 10 §10.1 raised the per-candidate minimum eval threshold from N=3 to N=5 and added a mutation_execution_gate to prevent baseline-only halts. Post-Amendment-10 a second-order interaction emerged where the mutation_execution_gate fired on the existence of any mutation eval (≥1) while the MIN_EVALS=5 filter excluded under-powered mutation evals from the aggregate, allowing baseline-aggregate halts to still fire. Amendment 11 §11.1 tightened the gate to require mutation candidates have ≥MIN_EVALS evals (5), eliminating the second-order false-positive class. Amendment 11 §11.2 capped further calibration cycles, binding any subsequent halt as genuine mechanism signal. We disclose this evolution as transparent empirical refinement rather than retroactive design change."

§E — Faza 2 expansion authorization scope

Per launch decision §F.5 condition_2 + Amendment 11 + Checkpoint C verdicts:

E.1 — AUTHORIZED for Phase 5 GEPA-evolved variant deployment

  • claude::gen1-v1 (claude shape; +12.5pp on 13/13 combined; non-Qwen control validation)
  • qwen-thinking::gen1-v1 (qwen-thinking shape; +12.5pp on 13/13 combined; Phase 4.5 mechanism validated)

E.2 — WITHHELD pending Faza 2 re-validation

  • gpt::gen1-v2 (gpt shape; +5pp held-out at threshold; selection-bias exposed; require N=16+ held-out re-evaluation)

E.3 — NOT EVALUATED in Faza 1

  • qwen-non-thinking shape: mutations evaluated but FAIL §F.1 (best candidate -12.5pp). Faza 2 may re-attempt mutation oracle with stronger anti-quality-regression scaffolding if PM authorizes.
  • generic-simple shape: mutations evaluated but FAIL §F.1 (best candidate 0pp). Faza 2 may re-attempt or scope as not-evolution-amenable.

E.4 — Faza 2 brief authoring scope

PM-side brief authoring should:

  1. Inherit claude::gen1-v1 + qwen-thinking::gen1-v1 as Phase 5 deployment-ready variants
  2. Document gpt::gen1-v2 N=16 re-validation protocol (~$1.43 incremental)
  3. Document qwen-non-thinking + generic-simple as scoped findings (mechanism activation without quality translation; Faza 2 may re-attempt with adjusted oracle)
  4. Inherit manifest v7 11-SHA chain as Faza 2 substrate-preservation reference

§F — Cost summary

Phase Actual Projection (manifest) Variance
Corpus generation $13.35 $13.58 (Amendment 3) -1.7%
NULL-baseline (artifactual sunk) $4.95 $0.50/run × 8 sunk by bug
NULL-baseline (re-run post Amendment 6) $4.97 $4.95 +0.4%
Mutation oracle $1.43 $3.00 -52%
Probe attempts $0.40 n/a (probe budget) n/a
Sunk Gen 1 (REGISTRY bug, archived) $1.36 n/a sunk
Full Gen 1 (120 evals) $15.02 $14.91 (Checkpoint A v2 §E) +0.7%
Checkpoint C held-out (15 evals) $1.93 $3.10 (PM brief 2026-04-29) -38%
Misc $0.08 n/a n/a
TOTAL Faza 1 $43.49 $44.66 (Amendment 9 + Checkpoint C estimate) -2.6%
Hard cap $115.00
Headroom retained $71.51

Cost discipline excellent throughout. Faza 2 + held-out spillover + analysis writeup can fit comfortably within remaining headroom.

§G — Cross-references — manifest v7 11-SHA chain + audit artifacts

G.1 — Manifest v7 SHA chain

Amendment SHA Date
Initial lock 1d592a6113c918b7a07fc9aba748c8bdd12a6ce1c6943943c0492678299fa700 2026-04-28
2 (post-A2) 583712dde139ffc87fb1ab21643f68d52c56469ded9e8090a624980b05969beb 2026-04-28
3 (post-A3) e43d13793535077c92a0e2c24f948ebb9d6e04000293690fdf38c4ba957aa972 2026-04-28
4 (post-A4) 1f7a6d6fa01403f6c8d6855893adbfa5e82898a81b7583cfa55628e5eba60196 2026-04-28
5 (post-A5) 062dfc4935aaa89f0b25595c5dc3ce4af06c95c4c261075a1f0226d8af3f3dee 2026-04-28
6 (post-A6) 0b55d8e353299594254e1a4a76f26f53014d726315dc6a0e5d6dc1a3a44a368a 2026-04-28
7 (post-A7) bc0bcf9bd8b0c8344b25e5f8ab15b0475039ba28a1f782ebffe4cc1c4ff7d1de 2026-04-28
8 (post-A8) 85858f12f1270da28277dd4d98e454d1dae8ef970537cb8c561f484599c4e2e9 2026-04-28
9 (post-A9) 5e3ad831c61beb19ccb4ff42b455b4c3964d830808944d4915189c5e9b1709b8 2026-04-28
10 (post-A10) 7fb2fb930670b5a28e417a76c64ca1a556f05afb9cf0761aba9f83f0c5de1c9b 2026-04-28
11 (post-A11) fa716ff90a4345eb87962789f3a2ab3d54994edc93964f850ad64cf6fbf6d227 2026-04-29

G.2 — Faza 1 audit chain artifacts

Item Path
Manifest v7 (terminus) D:/Projects/waggle-os-faza1-wt/benchmarks/preregistration/manifest-v7-gepa-faza1.yaml
Launch decision (predecessor) D:/Projects/PM-Waggle-OS/decisions/2026-04-28-gepa-faza1-launch.md
Pre-flight report D:/Projects/PM-Waggle-OS/briefs/2026-04-28-cc4-faza1-preflight-report.md
Pre-A addendum (corpus) D:/Projects/waggle-os-faza1-wt/benchmarks/results/gepa-faza1/corpus/h3-spot-audit-pre-a-addendum.md
Checkpoint A v2 report (NULL-baseline) D:/Projects/waggle-os-faza1-wt/benchmarks/results/gepa-faza1/null-baseline/checkpoint-a-report.md
Investigate report (REGISTRY bug) D:/Projects/waggle-os-faza1-wt/benchmarks/results/gepa-faza1/gen-1/investigate-report.md
Diagnostic probe (REGISTRY) D:/Projects/waggle-os-faza1-wt/benchmarks/gepa/scripts/faza-1/probe-registry-injection.ts
Checkpoint B report D:/Projects/waggle-os-faza1-wt/benchmarks/results/gepa-faza1/gen-1/checkpoint-b-report.md
Full Gen 1 halt report D:/Projects/waggle-os-faza1-wt/benchmarks/results/gepa-faza1/gen-1/full-gen-1-halt-report.md
Post-Amendment-10 halt report D:/Projects/waggle-os-faza1-wt/benchmarks/results/gepa-faza1/gen-1/post-amendment-10-halt-report.md
Final Gen 1 close report D:/Projects/waggle-os-faza1-wt/benchmarks/results/gepa-faza1/gen-1/final-gen-1-close-report.md
Checkpoint C close report D:/Projects/waggle-os-faza1-wt/benchmarks/results/gepa-faza1/checkpoint-c/checkpoint-c-report.md
THIS DECISION (terminal) D:/Projects/PM-Waggle-OS/decisions/2026-04-29-gepa-faza1-results.md

G.3 — Eval JSONLs (135 evals total)

Item Records
benchmarks/results/gepa-faza1/null-baseline/null-baseline-eval.jsonl 40 (NULL baseline 5 shapes × 8 instances; per-shape baselines anchored)
benchmarks/results/gepa-faza1/gen-1/gen-1-eval.jsonl 120 (full Gen 1; 5 shapes × 3 candidates × 8 instances)
benchmarks/results/gepa-faza1/gen-1/gen-1-eval-void-registry-bug-superseded.jsonl 11 (sunk pre-fix; archived for audit chain transparency)
benchmarks/results/gepa-faza1/checkpoint-c/checkpoint-c-eval.jsonl 15 (held-out 3 candidates × 5 instances)
TOTAL evaluative records 186 (40 NULL + 120 Gen 1 + 11 sunk + 15 Checkpoint C; 175 substantive + 11 sunk)

§H — Phase 5 brief authoring authorization

Per launch decision §F.5 condition_2 + this decision §B + §E:

Phase 5 GEPA-evolved variant deployment authorized for:

  • claude::gen1-v1 (file: packages/agent/src/prompt-shapes/gepa-evolved/claude-gen1-v1.ts; SHA pinned at substrate anchor commit)
  • qwen-thinking::gen1-v1 (file: packages/agent/src/prompt-shapes/gepa-evolved/qwen-thinking-gen1-v1.ts; SHA pinned at substrate anchor commit)

Phase 5 brief authoring is now UNBLOCKED (PM-side, gated by Marko). The brief should:

  • Cite this decision memo (2026-04-29-gepa-faza1-results.md) as authorization basis
  • Inherit manifest v7 11-SHA chain as substrate-preservation reference
  • Specify deployment scope (claude + qwen-thinking shapes; gpt + qwen-non-thinking + generic-simple require Faza 2 follow-up)
  • Schedule Faza 2 expansion brief authoring per §E.4

§I — Faza 1 CLOSED

Per all 5 acceptance gates passing (§F.1 + §F.2 + §F.3 + §F.4 + §F.5 false-positive guard), Faza 1 is CLOSED as of 2026-04-29.

Cumulative: $43.49 of $115 cap. Headroom $71.51 retained for Faza 2 + analysis writeup.

The 11-amendment manifest v7 chain documents the empirical evolution of methodology under pre-registration discipline — calibration patches, bug fixes, anti-misattribution locks, terminal_calibration_clauses — all transparent and audit-traceable. arxiv §5.4 framing builds on this audit chain as a positive demonstration of methodological rigor.


End of Faza 1 Results Decision Memo. Faza 1 CLOSED. Phase 5 brief authoring UNBLOCKED.