Files
Oleg Maslov 0c3e2ead3b
Some checks failed
Installer Smoke / installer-smoke (push) Has been cancelled
moving
2026-09-02 10:10:29 +02:00

65 lines
3.5 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Agentic Knowledge Work Pilot — N=3 Direction Validator
**Date:** 2026-04-26
**Type:** Pilot test (pre-full multiplier benchmark gate)
**Owner:** PM authoring brief, CC-1 executing
**Scope:** 3 tasks × 4 cells = 12 candidate runs + 36 judge calls (trio ensemble)
**Cost ceiling:** $5 hard cap, $4 halt
**Time budget:** 4-6 hours wall-clock
## Why this pilot exists
Stage 3 v6 N=400 LoCoMo proved the **memory substrate** thesis (oracle 74% > Mem0 66.9%). That's paper claim #1 — architecture beats peer-reviewed baseline on memory recall.
This pilot is paper claim #2**agentic knowledge work multiplier**. Question: does adding hive-mind memory + GEPA agent harness lift candidate model performance on real CEO/consultant work, not just memory recall?
PA V5 (April 2026) gave H1 PASS Opus 4.6 +5.2pp on knowledge work but on small N. This pilot replicates direction signal on N=3 across 4 cells. If pilot passes (H2/H3/H4 directional signs hold), full N=400 multiplier benchmark is authorized for paper claim #2 evidence.
If pilot fails (any of H2/H3/H4 reverses sign), we don't waste $150 on full benchmark — we go back to retrieval V2 work first.
## Files in this folder
| File | Purpose | Audience |
|---|---|---|
| `README.md` | This index — overview + sequencing | Marko, PM, CC-1 |
| `cc1-brief.md` | Technical execution brief | CC-1 primary |
| `task-1-strategic-synthesis.md` | Multi-document synthesis test materials | CC-1, judges |
| `task-2-cross-thread-coordination.md` | Cross-thread project coordination test | CC-1, judges |
| `task-3-decision-support.md` | Decision support under conflict test | CC-1, judges |
| `judge-rubric.md` | Likert 1-5 × 6 dimensions trio rubrika | Judge ensemble |
## Hypotheses pilot validates
- **H2:** Opus 4.7 + memory + harness > Opus 4.7 solo (multiplier on frontier model)
- **H3:** Qwen 3.6 35B-A3B + memory + harness > Qwen solo (multiplier on sovereign model)
- **H4:** Qwen + memory + harness ≥ Opus solo (SOTA-on-local proof, sovereignty bridge)
PASS criteria (binary):
- All 3 hypotheses show correct directional sign across ≥ 2 of 3 tasks (6/9 cells minimum)
- No catastrophic failure (any cell scoring < 2.0/5 overall on majority of judges)
If PASS → green-light full N=400 multiplier benchmark (Opus + Qwen + GPT-5.4 × 4 cells × N=400)
If FAIL → halt expansion, prioritize retrieval V2 work, schedule pilot retry post-V2
## Cost & time envelope
- Candidate model spend: ~$1.50 (12 runs, Opus dominates cost)
- Judge ensemble spend: ~$2.50 (36 calls × ~$0.07/call across Opus + GPT + MiniMax)
- Buffer: ~$1.00
- Total ceiling: $5.00, halt at $4.00
- Wall-clock target: 4-6 hours (parallel cell execution where possible)
## Sequencing
1. **PM** (you, now): generates pilot package — this folder
2. **Marko**: ratifies brief (1 review pass, optional adjustments)
3. **CC-1**: executes pilot — kicks runner, monitors halt rules, produces JSONL + summary
4. **PM**: adjudicates direction signal post-results, drafts go/no-go for full benchmark
5. **Marko**: ratifies go/no-go decision
## Notes on synthetic materials
All test materials in tasks 1-3 are **synthetic but realistic**, designed to mirror Marko's ICP work (CEO of mid-stage SaaS company, boutique consulting Partner, executive decision-maker). Documents are detailed enough to require genuine synthesis, not surface-level pattern matching.
Synthetic ≠ proxy. Each task has a clear "right answer shape" the judge rubric calibrates against — not a single correct answer, but a quality bar a real CEO/Partner would recognize as professional output.