16 KiB
CC-2 Faza 1 — Amendment 1 (PM ratification of pre-flight asks)
Date: 2026-04-28
Author: PM
Status: RATIFIED, binding upon paste-into-CC-2
Predecessor brief: briefs/2026-04-28-cc4-gepa-tier2-evolution-faza1-brief.md (266 lines)
Predecessor pre-flight: briefs/2026-04-28-cc4-faza1-preflight-report.md (239 lines)
Sesija executor: CC-2 (filename retains cc4 historical naming; CC-2 is operational executor)
Amendment scope: resolve 5 ratification asks (A-E) + 2 discovery items (max_tokens reconciliation, substrate HEAD pin)
§1 — Acknowledgment
Pre-flight rightly halt-ed at §6.3 ambiguity. Brief inherited Phase 4.3 hypothesis labels without verifying source corpus shape — this is exactly the failure class that feedback rule 6.1 (config inheritance audit) was created to surface. The catch saved $20 NULL-baseline burn on the wrong corpus + propagated downstream cost in Gen 1 + held-out validation that would have referenced inadequate-N source data.
Pre-flight report finding §3.4 (mutation validator anchored on MULTI_STEP_ACTION_CONTRACT in types.ts as the cell-semantic boundary linchpin) is also a strong design choice that goes beyond brief §3.4 specification. Ratify and adopt as standard.
§2 — Ratifications (asks A-E)
Ask A — H3 cell semantics disambiguation
RATIFIED: Option C (generate ≥40 net-new synthesis instances mirroring pilot NorthLane/CFO task family).
Rationale: only Option C preserves Phase 4.3 verdict anchor (which motivates entire GEPA work) and satisfies brief §6.3 ≥40 instances scope verification. Option A (3 instances) FAIL on statistical viability. Option B (LoCoMo agentic 400) disconnects from Phase 4.3 verdict — would force brief addendum reframing GEPA rationale away from agentic synthesis failure mode that empirically motivated the work. Option C is the only path that maintains methodological integrity.
Sub-asks (per pre-flight report §6 footer):
-
Target N for new H3 corpus = 50 — RATIFIED. Comfortable margin over §6.3 ≥40 threshold (8 NULL + 24 Gen 1 + 5 held-out + 13 buffer).
-
Subject model for instance generation = Opus 4.7 — RATIFIED. Consistent with mutation oracle (§3.3 of original brief), single-model anchor for entire pre-work + GEPA pipeline minimizes confounders.
-
Stratification axes — DEFER to CC-2 design judgment. Pre-flight report recommended "5 task families × 10 instances each, mirroring pilot's task-1/task-2/task-3 structure". Pilot had 3 task families; expansion to 5 requires authoring 2 net-new task families. PM does not specify which 5 axes — CC-2 designs stratification (with warm context on pilot artifact structure) and reports stratification design as part of manifest v7 LOCK §corpus_design block. Constraints:
- 5 task families minimum, all in NorthLane CFO synthesis domain (preserves anchor)
- Each task family yields 10 instances via persona/scenario/document-set variation
- Total stratification = 50 instances, deterministic-stratified via seed=42 for sampling
- Each instance must have ≥6 source documents (matching pilot ~5300-token CFO memo complexity)
- Each instance must have 6-dim Likert rubric (completeness, accuracy, synthesis, judgment, actionability, structure) per pilot pattern
-
Instance generation methodology — Opus 4.7 generates task scaffold (persona + scenario + 6-7 source document specs); PM does NOT review each instance pre-NULL-baseline (would balloon wall-clock). Instead: CC-2 spot-audits 5 random instances pre NULL-baseline kick (3% sample), reports any quality drift to PM in Checkpoint A halt. Trust mutation-oracle-as-task-generator pattern but with explicit spot-audit gate.
Ask B — Likert trio_strict_pass operationalization
RATIFIED: Operationalization (ii) — trio_mean ≥ T=4.0.
Rationale: pilot 2026-04-26 artifact already contains trio_strict_pass field with sample value true for trio_mean=4.583 and judge_minimax=failed. This empirically confirms (ii) operationalization with T probably = 4.0 as already-deployed pattern. Faza 1 reuses existing methodology rather than introducing new metric definition — preserves audit trail with pilot work and eventually with paper §5.4 conditional findings framing.
T=4.0 is also methodologically defensible: 4.0 on 1-5 Likert = "strong" rather than "passing", which is the threshold needed for fitness signal differentiation between GEPA candidates. T=3.5 would be too permissive (most candidates pass, low signal), T=4.5 too strict (most fail, low signal).
Manifest v7 must explicitly declare:
metric_operationalization:
trio_strict_pass:
method: aggregate_trio_mean_threshold
threshold: 4.0
citation: pilot_2026_04_26_artifact_pattern
Ask C — canonical κ=0.7878 source citation
RATIFIED: cite benchmarks/calibration/2026-04-24-trio-strict-recal.json, ratified by Stage 3 v6 Phase 1 trio judge ensemble pass commits 60d061e → 38a830e → 01f7ead (2026-04-24).
Manifest v6 §5.4 specifies the policy floor (κ ≥ 0.70 pass / [0.60, 0.70] borderline / <0.60 fail). The specific value 0.7878 is the empirically measured Phase 1 result.
Sub-ask CC-2 must verify:
- Confirm
benchmarks/calibration/2026-04-24-trio-strict-recal.jsonexists in waggle-os repo at HEAD (per memory entryproject_task25_stage3_v6_phase1_pass.md) - Compute SHA256 of file, pin in manifest v7 §canonical_kappa_anchor block
- If file is absent at HEAD: halt-and-PM (signal of repo state divergence — escalation)
Manifest v7 entry format:
canonical_kappa_anchor:
value: 0.7878
source_file: benchmarks/calibration/2026-04-24-trio-strict-recal.json
source_sha256: <CC-2 computes>
ratified_commits:
- 60d061e
- 38a830e
- 01f7ead
ratified_date: 2026-04-24
drift_threshold: 0.05 # per brief §4 condition 3
Ask D — path correction packages/core/ → packages/agent/
RATIFIED. Typo correction, no scope change. All future references in Faza 1 manifests, decision memos, GEPA outputs, and tests use packages/agent/src/prompt-shapes/ per pre-flight report §2.1 verified inventory.
Ask E — feedback_config_inheritance_audit.md reconstruction
NOT RATIFIED as proposed. Alternative path:
Original file lives at /sessions/inspiring-festive-lamport/mnt/.auto-memory/feedback_config_inheritance_audit.md (PM session memory, persists cross-sessions). It is NOT a waggle-os repo artifact and CC-2 cannot reach the path. Reconstructing in waggle-os would create a duplicate-but-stale copy that may drift from PM-side authoritative version.
Instead: embed 8 sub-rules inline in Faza 1 launch decision memo decisions/2026-04-28-gepa-faza1-launch.md under section §A — Inherited Pre-flight Rules. Source: brief §6.1-§6.8 verbatim. All Faza 1 audit references that would cite "feedback_config_inheritance_audit.md" instead cite "Faza 1 launch decision §A inherited pre-flight rules from PM brief §6".
This is cleaner: launch decision becomes self-contained binding contract for entire Faza 1 work, no external dependencies, reproducibility-ready for paper submission.
§3 — Discovery resolutions
Discovery 3.1 — judge max_tokens reconciliation
Brief §3.1 mandated max_tokens=3000 per judge "per Stage 3 v6 fix". This is a partial mis-citation in the brief. The Stage 3 v6 fix raised max_tokens specifically for Likert synthesis judging (judges need room to articulate per-dimension rationale across 6 dimensions), not for binary LoCoMo factoid judging (which does fine with 1024).
Manifest v6 §5.2 + §5.4 values (1024 / 1024 / 4096 for Opus / GPT / MiniMax) reflect LoCoMo factoid baseline, NOT synthesis Likert. Faza 1 is synthesis Likert (per Ask A Option C corpus type), so LoCoMo values are wrong inheritance.
RESOLUTION: CC-2 reads pilot 2026-04-26 judge config artifact (likely in benchmarks/results/pilot-2026-04-26/ or judge config YAML), extracts the actually-deployed max_tokens per judge for synthesis Likert. Pin those values in manifest v7 §judges block with explicit inherited_from: pilot_2026_04_26 cite.
If pilot artifact is missing or ambiguous on judge max_tokens, CC-2 halts pre manifest v7 LOCK and reports config archeology findings — PM ratifies values explicitly.
PM rec: expect values in 3000-4096 range for synthesis Likert (per intuitive scale of 6-dim rationale generation).
Discovery 4.5 — substrate HEAD pin (race condition guard)
Brief §3.5: "GEPA radi nad post-Phase 4.6 HEAD." Phase 4.7 commit c9bda3d is the actual post-Phase-4.6 anchor (commit be8f702 is Phase 4.6).
RESOLUTION: CC-2 pins manifest v7 substrate anchor on specific commit SHA c9bda3d (Phase 4.7 HEAD on feature/c3-v3-wrapper), NOT live HEAD. Race condition guard: CC-1 may commit Phase 4.4/4.5 work to feature/c3-v3-wrapper in parallel; CC-2 must operate on frozen Phase 4.7 anchor for entire Faza 1 to preserve reproducibility + apples-to-apples vs Phase 4.3 verdict.
CC-2 workflow:
git fetch origin feature/c3-v3-wrapper- Verify
c9bda3dis ancestor of branch HEAD (otherwise repo state divergence) git worktree add /tmp/faza1-worktree c9bda3d(isolated worktree on Phase 4.7 anchor)- All Faza 1 reads + GEPA evaluations use this worktree
- Final Faza 1 commits land back on feature/c3-v3-wrapper at HEAD via merge or cherry-pick (CC-2 designs final integration sequence and reports in Checkpoint C)
If c9bda3d is not ancestor (CC-1 force-pushed or branch rebased): halt-and-PM, escalation.
Manifest v7 entry:
substrate_anchor:
branch: feature/c3-v3-wrapper
commit_sha: c9bda3d
phase_label: Phase 4.7 (compression-engaged assertion test post-fold-in)
pin_method: git_worktree_isolated
rationale: race_condition_guard_vs_CC1_Phase_4_4_4_5_parallel_work
§4 — Updated cost projection (post Option C ratification)
Faza 1 LOCKED scope ($100 hard cap, $80 internal halt):
| Phase | Cost calc | Subtotal |
|---|---|---|
| Corpus generation (50 instances × Opus 4.7 generation oracle, ~$0.10/instance worst case) | 50 × $0.10 | $5.00 |
| NULL-baseline (5 shapes × 8 instances × $0.50 subject + judge cost per run) | 5 × 8 × $0.50 | $20.00 |
| GEPA Gen 1 (5 shapes × 3 candidates × 8 instances × $0.50) | 5 × 3 × 8 × $0.50 | $60.00 |
| Held-out validation (top-1 per shape × 5 instances × $0.50) | 5 × 1 × 5 × $0.50 | $12.50 |
| Mutation oracle (5 shapes × 2 mutations × 2 generations × $0.15 per mutation gen) | 5 × 2 × 2 × $0.15 | $3.00 |
| Total expected | ~$100.50 |
Tight against $100 cap. $80 internal halt remains. If actual mid-run cost exceeds projection by >30% (per brief §6.7 cost super-linear sub-rule), halt.
Cost discipline:
- Corpus generation completes BEFORE NULL-baseline kick (sequential, allows mid-checkpoint review)
- $5 corpus generation included in Checkpoint A scope (PM ratifies post corpus generation, pre NULL-baseline kick)
- Mutation oracle calls are cheaper than full evaluation calls (mutations don't run subject + judges, just produce candidate prompt)
If post-corpus-generation projection exceeds $100 cap: CC-2 halts, reports actual token costs, PM rerats to either reduce scope (e.g. drop generic-simple shape from Faza 1) or raise cap.
§5 — Updated halt-and-PM checkpoints (3 mandatory + 1 new pre-NULL)
| Checkpoint | Cumulative | Trigger | PM action |
|---|---|---|---|
| Pre-A (NEW) | ~$5 | Post corpus generation (50 instances + spot-audit 5 random) | Ratify corpus quality + NULL-baseline kick authorization |
| Checkpoint A | ~$25 | Post NULL-baseline 5 shapes × 8 instances | Ratify NULL trio-strict in 18-24% range + κ stability + Gen 1 kick |
| Checkpoint B | ~$50-65 | Mid-Gen 1 (after 30 evaluations) | Ratify intermediate κ + cell semantic violations review + complete Gen 1 |
| Checkpoint C | ~$100 | Post held-out validation | Acceptance verdict per brief §4 + Faza 2 expansion or PHF fallback |
Pre-A checkpoint added because corpus generation is non-trivial new step that didn't exist in original brief. Spot-audit 5 random instances at Pre-A is binding — PM must see sample quality before authorizing $95 downstream LLM run on the corpus.
§6 — Acceptance criteria update (post ratifications)
Brief §4 conditions remain binding except update §4 condition 1 prose:
§4 condition 1 (UPDATED): "Best GEPA candidate per shape beats NULL-baseline by ≥+5pp on trio_strict_pass rate (per H3 corpus, where trio_strict_pass = trio_mean ≥ 4.0 per Ask B ratification)"
Other conditions unchanged:
- §4.2: ≥3/5 shapes show positive delta
- §4.3: trio judge κ within ±0.05 of canonical 0.7878 (cite per Ask C)
- §4.4: zero cell semantic violations
§7 — Path forward (sequencing)
CC-2 next moves upon paste of Amendment 1 ratifications into session:
- Reconstruct sub-rule audit: ensure 8 sub-rules from brief §6 are accurately preserved in launch decision §A (per Ask E ratification)
- Read pilot judge config artifact: resolve Discovery 3.1 max_tokens
- Verify substrate anchor:
git fetch+ verifyc9bda3dancestry (per Discovery 4.5) - Verify κ anchor file: read
benchmarks/calibration/2026-04-24-trio-strict-recal.json, compute SHA256 (per Ask C) - Author manifest v7: with all explicit declarations (no inheritance gaps), pin all 4 anchors (corpus, κ, substrate, max_tokens)
- Author launch decision LOCK:
decisions/2026-04-28-gepa-faza1-launch.mdwith §A inherited rules + manifest v7 SHA + cost projection - Build GEPA harness scaffold + tests (≥80% coverage)
- Generate 50-instance H3 corpus + spot-audit 5 random
- Pre-A halt-and-PM: corpus quality review + NULL-baseline kick auth
- NULL-baseline run → Checkpoint A halt
No code authoring or LLM API calls outside this sequence.
§8 — Out-of-scope clarifications (post Amendment 1)
Still NOT in Faza 1 scope:
- H2 + H4 cells (Faza 2 expansion)
- More than 2 GEPA generations
- Population > 3 candidates per shape
- N > 8 per evaluation in Gen 1
- System prompt / cell semantics evolution (locked by §2 brief scope, enforced by §3.4 mutation validator)
- mind/ substrate modifications (locked by Discovery 4.5 substrate anchor)
- Apples-to-apples re-eval against pilot 2026-04-26 with original 12 instances (separate Korak 12 work; Faza 1 corpus is net-new per Ask A)
- Paper §5.4 framing update (post Phase 5 GEPA-evolved variant complete)
Newly in scope (Amendment 1):
- 50-instance H3 corpus generation (Ask A Option C)
- Pre-A halt-and-PM checkpoint (corpus quality gate)
- Pilot judge config archeology (Discovery 3.1)
- Substrate anchor pin via git worktree (Discovery 4.5)
- κ anchor SHA256 verification (Ask C)
- Inline §A inherited rules in launch decision (Ask E alternative)
§9 — Cross-references
- Predecessor brief:
briefs/2026-04-28-cc4-gepa-tier2-evolution-faza1-brief.md - Pre-flight report:
briefs/2026-04-28-cc4-faza1-preflight-report.md - Phase 4.3 verdict:
decisions/2026-04-28-phase-4-3-rescore-delta-report.md - Pilot artifact:
benchmarks/results/pilot-2026-04-26/pilot-task-{1,2,3}-C.jsonl - Manifest v6 anchor:
benchmarks/preregistration/manifest-v6-preregistration.yaml - κ anchor file:
benchmarks/calibration/2026-04-24-trio-strict-recal.json - κ ratification commits: 60d061e → 38a830e → 01f7ead
- Substrate anchor commit:
c9bda3d(Phase 4.7 HEAD on feature/c3-v3-wrapper) - Memory entry on Phase 1 κ ratification:
.auto-memory/project_task25_stage3_v6_phase1_pass.md(PM-side) - Feedback rules: brief §6 (canonical for Faza 1 work)
End of Amendment 1. Binding upon paste-into-CC-2. Proceed to manifest v7 + launch decision LOCK + corpus generation.