161 lines
8.5 KiB
Markdown
161 lines
8.5 KiB
Markdown
# CC Sprint 10 — Day 3 Brief
|
||
|
||
**Sprint:** 10 (Waggle-OS benchmarking, vector structure V1+V2)
|
||
**Day:** 3 (2026-04-22)
|
||
**Author:** PM (Claude Opus 4.7, Cowork mode)
|
||
**Status:** PM-ratified, ready for CC execution
|
||
**Predecessor:** `sessions/2026-04-22-sprint-10-day-2-status.md`
|
||
**Parent brief:** `briefs/2026-04-21-cc-sprint-10-tasks.md`
|
||
|
||
---
|
||
|
||
## 1. Day-3 ratification summary
|
||
|
||
Day-2 close is clean. 3/7 tasks CLOSED (Task 1.2, Task 1.3, Task 2.1) plus Task 1.1 scaffold. Headline: **tri-vendor κ = 0.7458 (substantial)**, Sonnet calibration 8/10 (borderline → triggered multi-vendor path, which is exactly what Day-2 Step 3 already executed). Zero anti-pattern #4 violations. $0.129 spent against $15 Sprint 10 ceiling (0.9%). 77/77 tests pass, tsc clean.
|
||
|
||
Per Day-2 §7 sequencing, Day-3 executes **one GO path** (Task 1.1 live run) and holds the remaining four tasks on external-input gates.
|
||
|
||
---
|
||
|
||
## 2. Day-3 execution plan
|
||
|
||
### 2.1 GO — Task 1.1 Qwen3.6 thinking-mode stability matrix (live run)
|
||
|
||
**Status transition:** scaffold CLOSED (Day-2) → live run (Day-3)
|
||
|
||
**Budget cap:** $1.50 (40 cells × Qwen rate via `qwen3.6-35b-a3b-via-openrouter` bridge route; Day-2 dry-run cost $0, live run estimate is ≤$1 but ceiling is $1.50 to absorb a retry).
|
||
|
||
**Route:** `qwen3.6-35b-a3b-via-openrouter` (OpenRouter bridge, DashScope direct route still pending Task 1.4 provisioning — live run uses the working bridge, not the canonical slug; this is per `project_target_model_qwen_35b.md` LOCKED 2026-04-21 policy).
|
||
|
||
**Scope:** Execute the 40-cell matrix verified in Day-2 dry-run. The scaffold rotates through all 4 outcome categories (`converged`, `divergent`, `timeout`, `parse-error`). Live run writes results to `preflight-results/qwen-stability-matrix-2026-04-22T<Z>.json`.
|
||
|
||
**Acceptance:**
|
||
- Matrix executes end-to-end without scaffold regression.
|
||
- At least one **safe config** emerges (≥70% converged across its row) — this is the Stage 2 kickoff gate condition per Sprint 10 brief §6.
|
||
- If **zero safe configs** surface, HARD STOP, PM review; do not auto-broaden matrix, do not re-tune thresholds post-hoc (anti-pattern #4).
|
||
- Report written to `docs/reports/qwen-stability-matrix-2026-04-22.md` with per-row converged/divergent/timeout/parse-error breakdown and recommended Stage 2 config.
|
||
- Commit + push.
|
||
|
||
**Timing:** ~30-60 minutes wall-clock per Day-2 §7 estimate.
|
||
|
||
**Post-close action:** CC writes brief update note to `sessions/2026-04-22-sprint-10-day-3-status.md` and returns to HOLD state awaiting Marko inputs for Tasks 2.2 / 1.5 / 1.4.
|
||
|
||
---
|
||
|
||
### 2.2 HOLD — Task 2.2 Fleiss' κ full 15-triple baseline
|
||
|
||
**Blocker:** Marko's 5 new PM-authored ground-truth triples (categories: temporal-scope, null-result, chain-of-anchor — or PM-selected equivalents).
|
||
|
||
**Ancillary blocker:** Instance #9 (`locomo_conv-50_q037`, Frank Ocean case) PM re-review. Per Day-2 §4.4, 4 of 5 non-Opus judgments flag F4 fabrication; only Opus 4.7 agrees with PM `correct/null`. This affects ground-truth stability for the 15-triple run. PM ratifies one of three options (see §5 of this brief and the re-review pack delivered separately).
|
||
|
||
**Execution trigger:** both inputs land from Marko → CC merges the 5 new triples with the 10-triple calibration set, re-runs ensemble on full 15, writes Fleiss' κ report.
|
||
|
||
**Pre-registered bands** (unchanged from parent brief §2.2):
|
||
- κ ≥ 0.80 → strong (ensemble-primary)
|
||
- 0.60 ≤ κ < 0.80 → substantial (tie-breaker policy required)
|
||
- 0.40 ≤ κ < 0.60 → moderate (PM review gate)
|
||
- κ < 0.40 → fair or worse (scope pivot)
|
||
|
||
**Budget:** $0.30 estimated per Day-2 §7.
|
||
|
||
**Non-action:** Do not run Task 2.2 on 10 triples as a placeholder. Day-2 κ=0.7458 is already indicative; running early on the smaller sample wastes budget and generates noise. Wait for full 15.
|
||
|
||
---
|
||
|
||
### 2.3 STANDBY — Task 1.4 DashScope dual-route
|
||
|
||
**Blocker:** Marko's classic DashScope API key.
|
||
|
||
**CC-side readiness:** Day-2 scaffold per brief §1.4 — LiteLLM config must contain both routes (`qwen3.6-35b-a3b` canonical DashScope direct + `qwen3.6-35b-a3b-via-openrouter` bridge) with failover policy documented (DashScope primary, OpenRouter retry-on-rate-limit fallback). Regression test must pass on both with byte-equivalent inference output on identical probe prompt.
|
||
|
||
**Execution trigger:** DashScope key lands → CC adds canonical route, writes regression, commits. ~30 min effort.
|
||
|
||
**Non-blocker reminder:** per parent brief §1.4, Task 1.4 does NOT block Stage 2 kickoff. OpenRouter bridge is sufficient for all Sprint 10 and Stage 2 budget projections. DashScope is on-prem parity hedge, not critical path.
|
||
|
||
---
|
||
|
||
### 2.4 STANDBY — Task 1.5 Harvest Claude artifacts adapter
|
||
|
||
**Blocker:** Marko's fresh Claude.ai export bundle.
|
||
|
||
**CC-side pre-work authorized:** CC can read `hive-mind/BACKLOG.md` commit `b3348fb` and refresh its understanding of the three source-path options (Option 1 current export bundles artifacts dir; Option 2 Claude.ai API listing; Option 3 Computer Use scraping). CC can also skeleton-start `hive-mind/packages/cli/src/commands/harvest-claude-artifacts.ts` with the UniversalImportItem type signature and the test file structure, **but must not commit before export inspection verifies which option is the correct primary.** Verification-first policy per parent brief §1.5.
|
||
|
||
**Execution trigger:** fresh export lands → CC inspects structure, confirms whether Option 1 (artifacts directory in export bundle) holds, implements adapter accordingly.
|
||
|
||
**Acceptance gate (from parent brief §1.5):**
|
||
- frame 421 (January 2026) + its artifacts (MASTER_PLAN_REVIZIJE.md and others) all accessible in chat-text substrate post-re-harvest
|
||
- 2 regression scenarios (artifact with valid parent, artifact without parent fallback)
|
||
- zero test regressions, tsc clean
|
||
|
||
---
|
||
|
||
## 3. Sequencing and fallback logic
|
||
|
||
```
|
||
Day-3 morning:
|
||
Task 1.1 live run (CC autonomous, ~30-60 min)
|
||
│
|
||
├─ safe config found → CLOSE, report, push, update Day-3 status
|
||
└─ zero safe configs → HARD STOP, PM review
|
||
|
||
Day-3 afternoon (parallel as Marko inputs land):
|
||
Triples arrive → Task 2.2 execute (~2h)
|
||
Export arrives → Task 1.5 begin verification (~1h inspect, then implement)
|
||
DashScope key arrives → Task 1.4 scaffold → regression → CLOSE (~30 min)
|
||
Instance #9 ratified → feeds Task 2.2 ground-truth set
|
||
```
|
||
|
||
**Day-3 floor outcome:** Task 1.1 CLOSED. That brings Sprint 10 to 4/7 CLOSED.
|
||
|
||
**Day-3 ceiling outcome (if all Marko inputs land by EOD):** 7/7 CLOSED, Sprint 10 moves to close-out briefing and Stage 2 kickoff memo.
|
||
|
||
---
|
||
|
||
## 4. Cost ceiling reminder
|
||
|
||
Sprint 10 budget: $15 hard stop.
|
||
Spent through Day-2: $0.129.
|
||
Day-3 projected: $1.50 (Task 1.1) + $0.30 (Task 2.2 if triggered) = $1.80 max.
|
||
Running total at Day-3 close: $1.93 (12.9% of ceiling).
|
||
|
||
Ample headroom for Stage 2 kickoff preparation in Sprint 11.
|
||
|
||
---
|
||
|
||
## 5. PM-ratified decisions bundled with this brief
|
||
|
||
The following decisions are delivered alongside this brief and carry PM authority for Day-3 execution:
|
||
|
||
**5a. Task 1.1 live run GO** — executes on `qwen3.6-35b-a3b-via-openrouter` bridge route without waiting for Task 1.4 DashScope provisioning. Rationale: OpenRouter bridge has stable inference path since 2026-04-21; Task 1.1 acceptance is model-behavior-independent of routing layer.
|
||
|
||
**5b. Instance #9 re-review policy** — decision ratified in separate response pack (Task 1.3/2.1 Instance #9 re-review, delivered as `decisions/2026-04-22-instance-9-reconciliation.md` once Marko picks option A/B/C).
|
||
|
||
**5c. Task 2.2 trigger condition** — Task 2.2 does not auto-start when 5 triples land alone. It requires BOTH (5 triples) AND (Instance #9 ratification) before kickoff. This protects ground-truth stability.
|
||
|
||
---
|
||
|
||
## 6. Out-of-scope for Day 3
|
||
|
||
- No landing copy work.
|
||
- No brand narrative work.
|
||
- No Stage 2 full-run execution (Stage 2 kickoff memo is Sprint 10 close deliverable, not Day-3).
|
||
- No scope expansion of Sprint 10 task list (7 tasks locked; no insertions without PM ratification via a new brief).
|
||
|
||
---
|
||
|
||
## 7. Reporting at Day-3 close
|
||
|
||
CC writes `sessions/2026-04-22-sprint-10-day-3-status.md` covering:
|
||
- Task 1.1 live run outcome + report link
|
||
- Which Marko inputs landed and which tasks that triggered
|
||
- Updated Sprint 10 close-criteria scorecard (per parent brief §10)
|
||
- Any anti-pattern flags surfaced during execution
|
||
- Delta against Day-2 κ, calibration, or acceptance bands (if Task 2.2 ran)
|
||
- Projected timeline for Sprint 10 close (Day-4 vs Day-5)
|
||
|
||
Same cadence and structure as Day-2 status doc.
|
||
|
||
---
|
||
|
||
**End of Day-3 brief. Awaiting CC execution on Task 1.1 live run.**
|