This commit is contained in:
Oleg Maslov
2026-09-02 10:14:22 +02:00
parent 0c3e2ead3b
commit b20b138fe4
771 changed files with 161561 additions and 9027 deletions

View File

@@ -0,0 +1,149 @@
# Audit Response Plan: Clever Memory Loses
**Date:** 2026-07-10
**Input:** publication-strength audit (2026-07-10), verdict no-go arXiv/press, conditional-go corrected blog.
**Decision:** ACCEPT the audit's core finding and PIVOT the thesis. Defend only 3 sub-points (below). The audit refutes the "raw-only, zero structure" claim using our own `benchmarks/results/longmemeval/RESULTS.md` — that is not survivable in review, and internal records confirm it (LongMemEval final = observation extraction + category routing + voting + KG-ledger guard; LoCoMo final = seven-lane layered system).
---
## 1. Triage of the 10 blockers
| # | Blocker | Ruling | Action |
|---|---|---|---|
| 1 | Thesis contradicted by own evidence | **ACCEPT — fatal as written** | Pivot thesis (§2) |
| 2 | LongMemEval routes on `question_type` metadata | **ACCEPT** | No-oracle rerun = new headline (E1) |
| 3 | LongMemEval tie presented as win | **ACCEPT with partial defense** | Reword: "ties 468/500 micro; +0.14 on incumbent's own macro aggregation." Macro is Mastra's own published metric, so reporting it is fair — claiming SOTA on it is not. Post-QA draft already concedes noise; title/abstract/dossier do not. Fix all. |
| 4 | Adaptive test-set tuning invalidates confirmatory p-values | **ACCEPT** | Relabel campaign exploratory; frozen confirmatory reruns (E3) |
| 5 | "33 made it worse" overclaims | **ACCEPT** | Rename Engineering Intervention Log; two tables (BEAM / LME), per-row N, model, baseline, metric, execution status, CI, adoption rule. Correct slogan: "33 not adopted; 4 adopted." |
| 6 | Conflict effect not causally isolated (store vs prompt) | **ACCEPT experiment, CONTEST breadth claim both ways** | Run 2×2 ablation (E2). Also narrow our own incumbent claim: Zep/Graphiti is bitemporal and retains history — say "systems that reconcile at write time on the answer path," not "every incumbent." |
| 7 | "Reproduced each incumbent" false for 2 of 3 | **ACCEPT** | Use audit's replacement wording verbatim |
| 8 | "No memory content leaves the machine" misleading | **ACCEPT** | Use audit's replacement wording verbatim; fix press kit too |
| 9 | SOTA landscape stale (mem0 92.5 LoCoMo, Hindsight 73.9 BEAM-1M) | **ACCEPT with defense** | New numbers are protocol-incomparable (different models/configs) — handle with landscape table + protocol-compatibility column, not silent retitle. But unqualified "state of the art" in title is dead regardless. |
| 10 | Reproducibility uneven (LME pipeline not public) | **ACCEPT** | Publish scripts 3455 + per-question artifacts, or narrow Appendix B claim |
**Partial defenses to keep (write into rebuttal/limitations, do not overplay):**
- D1: `question_type` is benchmark-provided input, not a gold answer — but Mastra doesn't use it, so head-to-head is still unclean. Disclose + rerun; keep routed number as a labeled secondary result.
- D2: Macro aggregation is the incumbent's own leaderboard metric; we report both and lead with micro.
- D3: BEAM result IS the clean simple-substrate result — one benchmark where the raw-only story is fully true. The pivot thesis keeps it as the flagship.
---
## 2. Thesis pivot
**Old (dead):** one dumb raw-turn substrate, no distillation/graph/routing, wins all three.
**New:** *preserve dated raw evidence as the canonical store; make every derived view reversible; defer conflict resolution to read time.* Raw turns dominate detail- and contradiction-sensitive tasks (BEAM, all-raw win); derived observations and read-time aggregation help breadth/counting (LongMemEval, LoCoMo); nothing on the answer path irreversibly deletes evidence.
**Title candidates** (pick after E1/E2 results):
1. "Clever Memory Loses When It Deletes the Evidence" (keeps the sticky brand, now true)
2. "Preserve First, Transform Later: A Lossless Memory Substrate Across LoCoMo, LongMemEval, and BEAM"
**Contributions restated:** (1) lossless canonical-store architecture; (2) three protocol-matched studies, presented separately; (3) causal conflict-preservation ablation on BEAM; (4) protocol-fidelity audit (unchanged — strongest surviving section); (5) engineering intervention log, honestly labeled.
Bitter Lesson angle survives as: "do not irreversibly discard evidence," not "never build structure."
---
## 3. Phases
### Phase 0 — Verify audit citations (0.5 day, agents, no writes)
Audit is specific and matches memory, but confirm before rewriting on top of it:
- [ ] `RESULTS.md:43-49, 27-40, 64-76` say what audit says
- [ ] `hive-mind/benchmarks/longmemeval/42-compose-final.mjs:15,36-40` routes on `it.question_type`
- [ ] Ledger entries claimed flat/positive (+0.023 temporal-commit, count-hint +3, etc.) — recheck signs in source tables
- [ ] Mastra category counts sum to 468/500 (mastra.ai/research/observational-memory)
- [ ] mem0 memory-benchmarks repo current numbers; Hindsight BEAM-1M 73.9 blog post; LIGHT/Honcho results
- [ ] BEAM repo license split (CC BY-SA 4.0 data / MIT code)
- [ ] Supermemory 85.4 vs 85.9 inconsistency; MemR3 duplicate reference
- **Gate:** any audit claim that fails verification gets struck from the plan; rest proceeds.
### Phase 1 — P0 rewrite (12 days, no new compute)
1. Rewrite title/abstract/intro/conclusion around pivot thesis.
2. Kill four false slogans everywhere (draft, blog, deck, posts, press kit): "zero per-benchmark tuning," "swap and change nothing," "33 made it worse," "every transform loses."
3. LongMemEval: micro tie first, macro second, metadata-routing disclosed in results section, not a footnote.
4. Privacy wording per audit (§8). Reproduction wording per audit (§7).
5. Section 4 rewritten: one canonical store, three benchmark-specific read paths, presented as three configurations of one preservation principle.
6. Ledger → Engineering Intervention Log (two tables + qualitative synthesis).
7. Figures: fig1 redrawn as one store / three read paths (or labeled BEAM-only interim); fig2 split into two panels with CIs or adoption matrix.
8. Related work: Memori arXiv:2603.19935 proper cite; Zep bitemporal correction; landscape table with protocol-compatibility column incl. current mem0/Hindsight/LIGHT; fix Supermemory number; dedupe MemR3; real bibliography.
9. Strip "Phase B-1 draft" status line; fix page-count/table-count in arxiv-metadata; fix title duplication p.1; fix orphaned Table 6 / blank half-pages.
10. `DATA_LICENSES.md` (LoCoMo, LongMemEval CC?, BEAM CC BY-SA 4.0 data) + attribution in result JSONL release.
11. Archive `draft.v1.md`, both DOCX, old LaTeX to `docs/paper/archive/` with README note (they present the layered thesis — audit is right that leaving them loose invites "your own files disagree").
### Phase 2 — P1 experiments (compute; sequence by information value)
- **E1 — LongMemEval no-oracle rerun** (highest value, cheapest): frozen config, ONE uniform read policy across all 500 Q (arm A); optional arm B = NL-only question classifier, report its confusion matrix. New headline number = arm A. Routed 95.01 becomes labeled secondary. Risk handled in §4.
- **E2 — BEAM 2×2 store×prompt ablation**: {raw-versioned, reconciled-current-only} × {incumbent prompt, conflict-aware prompt}. Stage 1: contradiction-ability subset (~100 Q × 4 cells) — isolates the +23pp mechanism cheaply. Stage 2 (if stage 1 clean): full 700 × 4. Reconciled store = simulate write-time reconciliation over same turns (mem0-style ADD/UPDATE/DELETE pass).
- **E3 — Confirmatory frozen reruns**: BEAM + LongMemEval final configs, 3 independent answer/judge passes each, report run distributions. Fixes the "judge-noise SE" mislabel with actual re-judging variance.
- **E4 — LoCoMo paired test**: McNemar vs reproduced Memori per-question outcomes + paired CI on accuracy difference; demote one-sample z-test.
- **E5 — Stats hygiene**: paired bootstrap CIs everywhere; exact tests; label exploratory vs confirmatory endpoints; BEAM avg-score delta reported as tie (CI 0.019..+0.034).
### Phase 3 — Release ops (after 1+2)
1. Publish full LME pipeline (scripts 3455) + per-question artifacts to public repos; verify public-tree parity with Appendix B claims.
2. Regenerate ALL launch assets from ONE claim matrix (single source of truth: claim → evidence file → status). PDF, arxiv-metadata, blog, posts, press kit, deck.
3. Proper bibliography (BibTeX), consider LaTeX/Typst build instead of Chrome print.
4. Re-run internal QA gates (adversarial review, anti-paper) against the NEW draft.
### P2 (only if targeting main conference — defer)
Weaker answerer family + alternative judge; held-out confirmatory slice; costquality Pareto; real-world contradiction eval; one-command pinned repro env.
---
## 4. Risk register
| Risk | Handling |
|---|---|
| E1 no-oracle drops below 94.87 | Paper survives — pivot thesis does not require winning LME. Report honestly: "matches/near leader; routed variant reaches X with disclosed metadata routing." Tie-with-simpler-read-path is still a result. |
| E2 shows prompt (not store) carries the +23pp | Also survivable — thesis becomes "read-time conflict policy over preserved evidence"; store retention is the necessary precondition (prompt can't surface deleted history). Interaction cell measures exactly this. |
| Confirmatory reruns regress BEAM pass-rate significance | Report distribution; drop p-value claims to descriptive. BEAM avg was already a tie. |
| mem0/Hindsight newer numbers steal headline | Landscape table with protocol column; claims scoped "under incumbent's published protocol as of [date]." |
| Rewrite drifts back to hype | Claim matrix is the gate: no sentence in any launch asset without a matrix row. |
---
## 5. Go/no-go (mirrors audit gates)
| Target | Gate |
|---|---|
| Corrected blog | Phase 1 items 14 + slogan kill |
| Social/HN launch | Blog gate + landscape table |
| arXiv preprint | Phase 1 complete + E1 + E2-stage-1 |
| Workshop paper | + E3, E4 |
| Main conference | + P2 item(s) |
## 6. Suggested execution order
1. Phase 0 verification (today, parallel agents).
2. Decision checkpoint: confirm pivot + title with user.
3. E1 + E2-stage-1 launch (compute runs overnight) in parallel with Phase 1 rewrite.
4. Assemble claim matrix → regenerate assets → QA gates → arXiv.
Rough new-compute cost: E1 ~5001000 answer+judge calls (gpt-5-mini/gpt-4o); E2 stage 1 ~800 gpt-5 calls; E3 ~3×(700+500) both roles. Order of magnitude comparable to one prior full-700 run — low hundreds of dollars, not thousands.
---
## 7. Phase 0 RESULTS (2026-07-10, three independent verifiers)
**Verdict: audit confirmed on all internal citations and all statistics; 3 external claims softened in our favor.**
Internal (verify-internal): claims 17 ALL CONFIRMED with file:line quotes. Ledger decomposition of "33 lost": 21 strictly negative / 6 flat / 3 positive-unadopted / 3 analytical-only — all six audit-named entries executed A/Bs, flat-or-positive as audit said. No NL classifier anywhere in LME scripts 3455; routing purely on dataset `question_type`. Only audit slip: "6 tables" metadata was accurate (page count 14→16 still our error).
Stats (verify-stats): every number MATCH — BEAM ours 0.6482016 / 518/700; mem0 0.6408656 / 491/700 (gpt-5 answerer+judge, top_200 — like-for-like judge symmetry confirmed); McNemar 425/93/66/116, z=2.141, p=0.0323 (asymptotic), 0.0392 (continuity), 0.0389 (exact); paired delta +0.007336, CI [0.0193, +0.0340], bootstrap agrees; LoCoMo 1332/1540 recount exact; contradiction ours 0.5875 vs mem0 0.3571 (+0.2304). Audit's judge-noise-SE point confirmed: draft's "0.014 judge-noise band" is across-question sampling SE (0.01348/0.01359), not judge noise. Method note: ours↔mem0 pairing must join on question TEXT (id ordering differs; id-join collapses to 70 rows). **E4 unblocked: Memori per-question reproduction exists at `D:\Projects\memori-repo\benchmarks\results_gemma\eval_20260609T035542Z.json` (1540 entries, join-able).**
External (verify-external): Mastra 94.87 CONFIRMED = macro, gpt-5-mini answerer + gpt-4o judge (protocol-matched to ours; micro tie 468/500 exact). BEAM data license CC BY-SA 4.0 CONFIRMED → DATA_LICENSES.md required. Zep bitemporal CONFIRMED (edge invalidation, history preserved) → our "every incumbent deletes" claim dead. Memori cite arXiv:2603.19935 CONFIRMED. **Softened:** (a) Hindsight 73.9 BEAM-1M uses Llama-4-Maverick judge, unstated metric, headline actually 64.1%@10M — NOT comparable; (b) mem0 92.5 LoCoMo = Top-200 + GPT-5 judge — protocol-incomparable to our Memori-protocol 86.49; (c) no public evidence incumbents do/don't route on `question_type` — reporting per-category ≠ routing. Landscape table with protocol-compatibility column is the right instrument (audit agreed). LongMemEval `_abs` abstention marking (30 Q, id suffix not question_type) must be handled identically in E1.
**Decisions locked:** pivot thesis per §2; all Phase 1 items proceed; E1 arm A = last pre-routing ladder rung config, uniform for all 500 Q; final step after rewrite + E1/E2 = independent re-audit ("re-judge") of the new package.
---
## 8. RIVAL PROTOCOL FORENSICS (2026-07-10) — the "we lost SOTA" numbers dissected
**Mastra 94.87 (LME):** macro artifact. Micro = 468/500 = 93.60, EXACT TIE with our routed run. Same answerer (gpt-5-mini), same official gpt-4o judge. Our macro 95.01 > their 94.87 — but ours oracle-routed, theirs not. True deficit: production-legal only (~92.4 classifier-routed vs their 93.60). Their system: gemini-2.5-flash ingestion-time observation compression, one static ~30k-tok context, single pass, open source. Beat = close ~610 questions in temporal-reasoning + knowledge-update without labels. Multi-session already tied (116/133 both).
**mem0 92.5 (LoCoMo):** protocol-inflation stack, NOT a substrate win. Their own paper (arXiv:2504.19413) scored J≈67% on the same 1540. The 92.5 = gpt-5 answerer + gpt-5 judge with maximally lenient prompt (1-of-N list items = CORRECT; ±14-day dates; ±50% durations; same-valence emotions; abstention-banned CoT answerer; cat-3 gold truncated at semicolon; adversarial cat-5 excluded — same 1540 scope as ours) + Platform-v3 closed retriever, top-200 (top-k lever only +0.7pp vs top-50). Judge is directly reusable standalone: `benchmarks/locomo/prompts.get_judge_prompt` + `common/llm_client.LLMClient` — ~20-line script over our (category, question, gold, prediction) triples. Full comparable rerun config documented in forensic report.
**Eywa 81.45 (BEAM):** not a comparable number. Sonnet 4.6 as BOTH answerer AND judge, custom rubric harness (paper misleadingly says BEAM "introduced here"; official-nuggets-or-reauthored unverifiable — artifacts URL 403s, no code, single-author vendor self-report), ZERO in-harness baselines, undisclosed context budget. Answerer edge small (+1.4pp Sonnet-vs-gpt-4o by their own LoCoMo anchor); judge is the story. Triangulated: plain raw-turn substrate under their harness ≈ 0.700.74 avg → Eywa's true like-for-like edge ≈ 510 pts (concentrated in abstention 92.9, temporal 90.0; their weakness = summarization 64.1, same as ours; contradiction we already own via dated turns). **CRITICAL protocol note: Hindsight's 73.9 is the AMB harness (Gemini answerer + Gemini judge, vendor-run) — ALSO not comparable to our gpt-5/gpt-5.** Honest position: our 0.648/74.0% vs mem0 0.641/70.1% is the ONLY clean like-for-like BEAM-1M comparison in existence; no one has published a comparable number above ours. Landscape table needs judge column: Eywa (Sonnet self-judge) / Hindsight+Honcho (Gemini/Gemini AMB) / mem0+ours (gpt-5 or gpt-4o official-style) — three islands, not one leaderboard.
### Campaign menu (SOTA recovery)
- **C1 LoCoMo unqualified SOTA (cheapest, highest probability):** rerun our substrate under mem0's exact protocol (gpt-5 answerer, their judge prompt verbatim, top-200, cat 14). Expected 9295 given we score 86.49 under a FAR stricter judge. Stage 1 (cheap, no re-answering) = 3-pass judge decomposition over our EXISTING 1540 answers: (a) Memori judge baseline 86.49; (b) mem0 `_JUDGE_TEMPLATE` + gpt-4.1-mini → isolates prompt leniency; (c) mem0 `_JUDGE_TEMPLATE` + gpt-5 → isolates judge model. Residual to 92.5 after (c) = answerer + retrieval + their 7-step CoT answer prompt (abstention banned — third confound lever; for stage 2 rerun, decide ours-vs-theirs answer prompt explicitly). Adapter = ~30 lines importing `benchmarks/locomo/prompts.get_judge_prompt` + `common/llm_client.LLMClient` from D:\Projects\mem0-memory-benchmarks (cat-3 golds get `preprocess_answer` semicolon truncation). Total est. <$50, half a day.
- **C2 LME production-legal lead:** finish E1b exact number (was mid-run, ~$3), then target temporal-reasoning (84.2→) + knowledge-update + multi-session with label-free levers (uniform voting; observation-layer improvements à la Mastra). Need ≥469/500 micro no-oracle. Moderate difficulty.
- **C3 BEAM vs Hindsight 73.9:** abstention gate (biggest structural gap), knowledge-update latest-fact selection, summarization lane. Research campaign, days + iterative pilots. Judge caveat: Hindsight's judge identity (Llama-4-Maverick per their comparison page) still muddies exact comparability — verify before claiming.

View File

@@ -0,0 +1,34 @@
IMPORTANT: Do NOT read or execute any files under ~/.claude/, ~/.agents/, .claude/skills/, or agents/. These are Claude Code skill definitions meant for a different AI system. Do NOT modify agents/openai.yaml. Stay focused on this consultation only.
You are a brutally honest technical reviewer (think: hostile NeurIPS reviewer + benchmark methodologist). Another AI (Claude/Fable) orchestrating a memory-systems research program wants your independent verdict on its path to claiming SOTA on three conversational-memory benchmarks. Be direct, terse, no compliments. Challenge assumptions.
=== CONTEXT ===
System under test: "conflict-aware raw-turn memory" substrate (per-conversation stores, verbatim dated raw turns, hybrid dense+BM25 retrieval, conflict-preserving answer policy). Paper draft was audited; original "SOTA on all 3" claim collapsed. Verified current positions:
LOCOMO (1540 Q, cats 1-4, adversarial excluded):
- Ours: 86.49% under Memori protocol (gpt-4.1-mini answerer+judge, strict judge). Beats Memori 81.95 same protocol (paired artifacts exist).
- mem0 claims 92.5 (README) but their released per-question artifact reproduces 91.56 (1410/1540). Their protocol: gpt-5 answerer + gpt-5 judge with codified-lenient prompt (1-of-N list items = CORRECT, plus-minus 14-day dates, 50% duration tolerance, abstention-banned 7-step CoT answer prompt, cat-3 gold truncated at first semicolon) + top-200 memories. mem0's own 2025 paper scored ~67 on same set with stricter judge.
- Judge decomposition DONE on our existing answers: our answers + mem0 judge prompt on gpt-4.1-mini = 93.90; + mem0 judge prompt on gpt-5 = 90.13. So we sit 1.4pp under their artifact with a far weaker answerer (gpt-4.1-mini vs gpt-5) and top-70 retrieval.
- IN FLIGHT (C1 stage 2): rerun our substrate with gpt-5 answerer + top-200 parity + mem0 gpt-5 judge, McNemar vs their 1410/1540.
LONGMEMEVAL-S (500 Q, official gpt-4o judge, gpt-5-mini answerer both us and incumbent):
- Mastra (incumbent): 94.87 macro / 93.60 micro (468/500). No metadata used.
- Ours published: 95.01 macro / 93.60 micro (468/500) — micro EXACT TIE — but our routing read the benchmark-provided question_type annotation (oracle label). Disclosed now.
- Production-legal (text-only classifier routing, exact, just computed): 92.20 micro / 92.87 macro (461/500). Uniform no-routing: 88.20 micro.
- Miss analysis: 39 misses; classifier misroutes cost exactly 4; ~20 are hard substrate floor (multi-session counting/aggregation); union-of-all-methods ceiling 480/500. Levers: better classifier +2-4 (cap), abstention calibration +2-4, multi-method ensemble +3-6 at 2-5x cost. All exploratory; LongMemEval has NO held-out split; everything tunes on the same 500.
BEAM-1M (700 Q, 35 convs, nugget 0/0.5/1 judge):
- Ours: 0.6482 avg / 74.0% pass@>=0.5, gpt-5 answerer + gpt-5 judge (like-for-like with mem0's released run: 0.6409 / 70.14%). Paired McNemar on pass: p=0.0389. BUT: win is contingent on healing 56 empty-generation rows (pre-heal 0.6050/69.3 = loss; heal legitimate: literal empty completions, regenerated once blind, 24/56 still failed; mem0 file has 0 empties). AND the pass win is carried by the 0.5 partial-credit bucket (strict >0.5: we lose 427 vs 429). Contradiction ability: +23pp (0.588 vs 0.357, p<1e-4), heal-independent — the one bulletproof result. Token budget: our contexts ~21K tok/Q vs their ~5-8K (3-4x) — disclosed.
- Rival numbers all protocol-incomparable: Eywa 81.45 = Sonnet 4.6 SELF-judging with self-authored rubrics, zero baselines, artifacts 403; Hindsight 73.9 avg = vendor AMB harness, Gemini answerer+judge; Honcho 63.1 same. So ours-vs-mem0 is the only clean like-for-like pair in existence.
=== THE PLANNED PATH ===
C1 (in flight): LoCoMo under mem0's exact protocol (gpt-5 answerer, their lenient judge, top-200) -> expect 92-95 -> claim "SOTA under both strict (Memori) and lenient (mem0) protocols".
C2: LME label-free push: classifier upgrade + abstention calibration, frozen config, ONE final run, no iterating (test-set discipline) -> hope 465-467; ensemble variant could reach 468-472 at 2-5x cost. Claim: "ties/near-ties incumbent production-legally; exceeds under disclosed oracle routing".
C3: BEAM substrate work: abstention gate + knowledge-update latest-fact + summarization lane -> target clean avg-score lead over mem0 (currently tie) -> claim "leads the only like-for-like BEAM comparison".
Headline strategy: "SOTA under every protocol where a like-for-like comparison exists" + protocol-fidelity audit as a first-class contribution + falsification/intervention log.
=== QUESTIONS ===
(a) AGREE/DISAGREE per track (C1/C2/C3): is this the right way to defensible SOTA claims on all three?
(b) Where does a hostile reviewer still attack, in priority order?
(c) What would you do differently or additionally? Especially: test-set-tuning discipline when no held-out split exists (LME); making BEAM claims comparable across three protocol islands; whether C1's "win under a lenient judge" claim boomerangs (does beating mem0 under mem0's own inflated protocol undermine our protocol-fidelity story?).
(d) Is "SOTA under every protocol where a like-for-like comparison exists" a defensible headline? If not, what is the strongest defensible headline?
Return structured verdict: AGREE/DISAGREE per track with reasons + ranked additional proposals.

View File

@@ -0,0 +1,74 @@
# Fable × Codex consensus on the SOTA path (2026-07-11)
## 0. Second opinion: gpt-5.6-sol (high reasoning) — CONCURS, stricter on 3 points
Independent verdict before seeing prior reviewer, then adjudicated: agrees on all central conclusions (C1 → cross-judge protocol-sensitivity study; C2 SOTA chase dead — "frozen one final run is ceremony, not test discipline"; C3 contradiction result is the contribution; headline = protocol shopping, "an existentially quantified loophole"). Stricter additions:
1. **C1 2×2 insufficient** unless answerer, retrieval units, token budget, preprocessing, retry behavior are also controlled or explicitly factored. Report judge-flip transition counts + human-audit a stratified disagreement sample.
2. **BEAM McNemar likely INVALID**: 700 questions nested in 35 conversations — question-level McNemar assumes independent pairs, anti-conservative. Recompute with conversation-cluster bootstrap / cluster randomization; p=0.0389 may not survive. More fundamental than the partial-credit issue.
3. **LME oracle score gets no "exceeds incumbent" language anywhere** outside an ablation; disclosure does not repair task leakage.
Approved headline (both models): contradiction-handling gain + protocol-sensitivity audit. Performance-forward variant: "competitive aggregate performance and a robust contradiction advantage in controlled evaluations."
Codex (gpt-5-era, medium reasoning) consulted via /codex with full plan context (§7 Phase 0, §8 forensics, live C1/C2 results). Full verbatim output preserved below (§4). Prompt archived: `2026-07-11-codex-consult-prompt.txt`.
## 0.5 SMOKE TIER (agreed 2026-07-11, gpt-5.6-sol converged; <$15, 12 days, run BEFORE funding the package)
| # | Smoke | Kill criterion | Load-bearing? |
|---|---|---|---|
| S0 | Protocol/claim freeze doc + judge-calibration panel (30 frozen answers, official vs proposed judge, need ≥90% agreement, κ≥0.80, bias ≤2/30) | Any comparison inexpressible as one frozen protocol; calibration failure forbids cross-judge comparisons | YES |
| S1 | mem0-474 forensics ($0): their repo says GPT-5 answerer+judge for LME → 474 likely off-protocol; verify artifact-level | If 474 comparable → LME bar rises to 474 | bar-setting |
| S2 | Cluster bootstrap on existing BEAM data ($0, local, 10k paired conv bootstraps) | BEAM aggregate win dies unless cluster-aware 95% CI lower bound > 0 | YES |
| S7 | C1 stage 2 (sunk, in flight) | Green ≥+3pp vs best comparable lenient result; red if tie/lose/<1pp | YES |
| S4 | mem0-OSS under strict LoCoMo, 1 conv ~150 Q (~$3), sequential 2nd conv if ±5pp | Red if mem0 leads ≥5pp | YES |
| S6 | Paired LME 50 (stratified, frozen): Mastra exact config + ours, same 50 (~$2) | Mastra repro <41/50 kills comparability; ours trailing ≥3/50 = red | YES (paired part) |
| S5 | AMB harness spike, ours + BM25/oracle control, 40 Q (~$5, needs Gemini key) | Kill AMB plank unless ≥38/40 complete, no manual repair, ≤$0.15/Q | YES |
Settled en route: LongMemEval-V2 EXISTS (451 Q, agent-trajectory task, own harness) — excluded from scope. GO = all load-bearing green + S0 clean. FALLBACK to contradiction+audit paper if any load-bearing red, two grays, or judge calibration fails. No averaging failures across benchmarks.
## 1. Verdicts and agreement
| Track | Codex | Fable | Consensus |
|---|---|---|---|
| C1 LoCoMo parity rerun | AGREE with restrictions | agree | PROCEED, reframed as protocol-sensitivity experiment: report a 2×2 matrix (both systems × strict Memori judge and mem0 judge). The claim is ordering stability across protocols, not "we won under the lenient judge." We hold mem0's per-question answers → their cells cost only judge calls. |
| C2 LME lever push to beat Mastra | DISAGREE (adaptive tuning; "one final run" doesn't restore independence) | agree — flagged same risk before consult | PARK the lever push. Ship 92.20 production-legal + oracle-inflation quantification (93.60→92.20→88.20 ladder) as a methodology contribution. No LME SOTA claim without a genuinely unseen eval. |
| C3 BEAM substrate work for avg-score lead | DISAGREE as framed (retry asymmetry, threshold sensitivity, would be tuned on judged 700) | agree | REFRAME: contradiction result (+23pp, heal-independent, p<1e-4) is the headline contribution. Aggregate reported as full ordinal distribution + first-attempt primary / retry-normalized secondary under a pre-declared policy. Substrate improvements (abstention gate etc.) only with dev/eval separation. |
| Headline "SOTA under every protocol where like-for-like exists" | NOT defensible (movable denominator = protocol shopping) | accept | New headline: **conflict-preserving raw-turn memory delivers a large, reproducible contradiction-handling gain under controlled comparison, and a protocol-fidelity audit shows conversational-memory rankings are underidentified (judges, routing metadata, retries, budgets).** Conditional upgrade if C1 matrix shows stable ordering: "leads matched LoCoMo and BEAM comparisons." |
Fable's two disagreements with Codex (minor):
1. Codex: "competitor had no equivalent regeneration opportunity" on BEAM. mem0's released file has 0 empties — their managed pipeline either never failed or retried internally and invisibly. True symmetric rerun of their platform is impossible; the feasible fix is the pre-declared policy + first-attempt-primary reporting Codex also proposes.
2. Codex treats C1 as unfinished evidence; it is in flight, and stage-1 judge decomposition already delivered half the protocol-sensitivity matrix.
## 2. New actions proposed (need user approval)
| # | Action | Cost | Yield |
|---|---|---|---|
| A | Cross-judge matrix on frozen answers: ours + mem0's LoCoMo answers under BOTH judges (completes C1 2×2); ours + mem0's BEAM answers under a second judge (gpt-4o) + judge-agreement stats (kappa) | ~$2040 | Turns "protocol islands" into a first-class result; Codex proposals #3, #10 |
| B | BEAM reporting overhaul: pre-declared retry policy, first-attempt primary + healed secondary, full ordinal distribution (0/0.5/1 rates), cost/token disclosure | ~$0 (reporting) | Kills attacks #5, #6, #7 |
| C | Resume E2 causal ablation (store×prompt 2×2 on contradiction subset; was mid-flight: retrieval+reconciled store built, cells A/B partial) | ~$10 remaining | Codex proposal #4 — isolates conflict-preservation causally; upgrades the headline mechanism claim |
| D | Pre-registration doc: freeze metrics, comparisons, retry/exclusion rules, stat tests, stopping rules before any further runs | ~$0 (writing) | Codex proposal #5; discipline for everything after |
| E | Pareto costquality frontiers from existing artifacts | ~$05 | Codex proposal #6; converts token-budget liability into a result |
| F | External/unseen LME eval set (new generated conversations, independent adjudication) | days + $$ | Only path to any future LME SOTA claim; DEFER unless user wants it |
C1 stage 2 continues unchanged (already covers the ours×mem0-judge cells; A adds the mem0-answers×strict-judge cells).
## 3. Codex's ranked attack list (top 5 retained for the re-audit gate)
1. Test-set adaptation presented as evaluation. 2. SOTA defined after seeing favorable comparisons. 3. Attribution unsupported (answerer/judge/budget/substrate all vary). 4. Oracle-label contamination. 5. BEAM retry asymmetry.
## 4. Codex verbatim output
(preserved in full for the record)
You do not have a defensible "SOTA on all three benchmarks" path. You have: one likely protocol-specific LoCoMo win; one LongMemEval result below the incumbent under production-legal conditions; one narrow BEAM win on a permissive threshold, plus a much stronger contradiction-specific result. That can make a credible paper. Calling it universal SOTA would weaken it.
[C1 AGREE with restrictions — necessary; supports "best result in a controlled comparison under the released mem0 protocol"; does not establish dual-protocol SOTA; report paired outcomes + McNemar CI, strict judging of both systems' parity answers, factorial table (answerer/judge/top-k), category-level results, dedup sensitivity, cost parity; don't celebrate a lenient-judge win — call it "released mem0 evaluation protocol" and audit leniencies separately.]
[C2 DISAGREE — the process is not test-set discipline; classifier, abstention, method selection, miss analysis all informed by the same 500; "ONE final run" doesn't make it independent. 461/500 production-legal is below Mastra 468/500. Oracle result invalid for production-comparable claims; 0.14 macro is noise; "exceeds under oracle routing" is spin. Reframe as robustness/methodology.]
[C3 DISAGREE as framed — aggregate lead fragile: appears only after regenerating 56 failures, competitor had no equivalent opportunity, reverses under strict scoring, 34× context budget, significance tied to partial-credit threshold. Contradiction result is the real contribution: large, significant, mechanistically aligned, heal-independent. Center C3 on it. Substrate changes risk benchmark overfitting unless developed outside the 700 judged questions.]
[Attacks ranked: 1 test-set adaptation; 2 movable SOTA denominator / protocol shopping; 3 attribution unsupported; 4 oracle contamination; 5 retry asymmetry; 6 metric selection after outcome inspection; 7 cost-normalized inferiority; 8 permissive-judge optimization; 9 no uncertainty on tiny deltas; 10 benchmark saturation; 11 contradiction subset possibly post-hoc; 12 artifact asymmetry.]
[Do differently: stop targeting SOTA-on-all-three; three falsifiable claims instead. Genuinely untouched eval split for LME (external hold-out / pre-registration / nested CV by conversation; do NOT fabricate pseudo-held-out splits now). Separate comparison from causal attribution via factorial ablation — key question: does conflict preservation cause the contradiction gain. Normalize BEAM inference policy (pre-declare attempts/failure definitions/budgets), report first-attempt + retry-normalized + full ordinal distribution + cost. Bridge BEAM protocol islands with a cross-protocol judge matrix over the same frozen answer artifacts + judge-agreement stats (kappa, favoritism analysis, blinded human sample). Treat C1 as protocol-sensitivity: the result you want is ordering stability across protocols. Pre-register hypotheses and correction rules. Add cost-constrained Pareto frontiers.]
[Headline: "SOTA under every protocol where a like-for-like comparison exists" NOT defensible — LME production-legal is below incumbent, BEAM depends on retry+threshold, LoCoMo unfinished, denominator selective. Strongest defensible now: "Conflict-preserving raw-turn memory substantially improves contradiction handling, while a protocol-fidelity audit shows that benchmark rankings are highly sensitive to judges, routing metadata, retries, and inference budgets." If C1 succeeds under both judges matched: "Conflict-preserving raw-turn memory leads matched LoCoMo and BEAM comparisons, with a 23-point contradiction gain, while exposing substantial protocol-induced inflation in conversational-memory benchmarks."]
[Ranked proposals: 1 unseen external eval set; 2 BEAM symmetric retry rerun; 3 cross-protocol judge matrix on frozen artifacts; 4 causal ablations isolating conflict preservation; 5 pre-registration; 6 Pareto frontiers; 7 correct oracle-routed LME prominently; 8 contradiction/update as primary mechanistic result; 9 all BEAM ordinal outcomes; 10 protocol instability as central evaluation contribution. "The credible paper is not 'we won three leaderboards.' It is 'leaderboard claims in conversational memory are underidentified, and conflict preservation produces one large, reproducible capability gain under controlled comparison.'"]

View File

@@ -0,0 +1,99 @@
# Protocol & Claim Freeze (2026-07-11)
Pre-registration for the three planned races. Everything below is fixed **before**
any race result exists. Companion: `2026-07-11-fable-codex-consensus.md` (§0.5 smoke
tier, go/no-go). Calibration evidence: `KorroResearch/benchmarks/locomo-mem0-parity-2026-07/calib/`
(`panel_input.jsonl`, `panel_verdicts.jsonl`, `analyze_panel.py` — this smoke, Part 2).
**Global rules (all three races).**
- **One evaluation run per frozen config.** No reruns-until-win. If a run is voided
it is voided for a *disclosed operational reason* (crash, auth failure), not because
of its score, and the void is logged.
- **Retry policy.** Max **1** regeneration, triggered only on a *literally empty*
answer string (finish_reason truncation or empty content). Applied symmetrically
wherever we control the pipeline. **First-attempt result is the reported primary;**
retry-normalized is a disclosed secondary. Competitor artifacts we do not control
(mem0 released answers) get first-attempt-primary treatment with the asymmetry
disclosed — a true symmetric rerun of a managed platform is infeasible.
- **Statistics.** Paired tests only, **cluster-aware at the conversation level**
(questions are nested in conversations; question-level independence is false).
Report a 95% CI whose lower bound must clear the go threshold. Conversation-cluster
bootstrap (10k resamples) is the primary interval; McNemar is reported but its naive
p-value is treated as anti-conservative and never the sole basis for a claim.
- **Multiplicity.** Three benchmarks × one primary metric each = 3 primary tests.
HolmBonferroni across the 3 primaries; per-benchmark secondaries are descriptive,
not claim-bearing.
- **Systems enter only with a frozen, hashed answer artifact.** A system with no
reproducible per-question answer file does not enter the matrix.
---
## Race 1 — LoCoMo (4 systems × 2 judge protocols)
| Item | Freeze |
|---|---|
| Dataset | snap-research `locomo10.json` (Maharana et al., ACL 2024). Canonical build `locomo-1540` |
| Dataset hash | build `39e415e2f3a0fa1bd3cb1804a58d0b440b50d3070b2100698437e4ec402a5b24`; canonical instance_count 1531 |
| Evaluated set | **1540 rows** (mem0's own denominator: cat1=282, cat2=321, cat3=96, cat4=841). The 1531→1540 gap = 9 duplicate-question rows mem0 scores separately; we match their denominator for parity. Cat 5 adversarial (446) **excluded** |
| Answer artifacts | Ours = `locomo-7lane-w4-judgments-N1540.jsonl` (`answer_content`), hash `1252fbde90613ebb5622f6724def7e60a433f4d8bf754829c0a0592972f079fa`. mem0 = released per-question verdicts `mem0_perq_verdicts.json` (their answers held, joined on `conv_idx` + normalized question, gold tiebreak on 11 dup groups). Systems 3 & 4 enter only if a frozen answer file exists; else the matrix is 2×2 and reported as such |
| Answerer | Per-system, frozen in each answer artifact. **Not** re-answered for this race — judge-only |
| Protocol axis (the 2) | **P1 mem0-lenient:** `mem0-memory-benchmarks/benchmarks/locomo/prompts.py` `JUDGE_PROMPT` (partial-credit: ≥1 gold item ⇒ CORRECT), gold via `preprocess_answer` (cat-3 semicolon split). **P2 Memori-strict:** `memori-repo/benchmarks/02_run_benchmark.ipynb` cell 2 `ACCURACY_PROMPT` (verbatim), raw gold |
| Judge model | P1 = gpt-5 (arm c reference) and gpt-4.1-mini (production arm b), both frozen; P2 = gpt-4.1-mini. Judge model held constant within a protocol column |
| Primary metric | Judge accuracy (J-score), cats 1-4, on the 1540 |
| Statistic | Paired McNemar per system-pair **plus** conversation-cluster bootstrap 95% CI (10 conversations) |
| Claim shape | Ordering **stability across the two protocols**, not "we win under the lenient judge." A protocol-specific win is reported as protocol-specific |
The result this race is allowed to support: *ranking is (or is not) stable when the
judge protocol is swapped.* Cross-protocol number-vs-number comparison is licensed only
because S0 calibration cleared the substitution gate (below).
## Race 2 — LongMemEval (one harness: ours + Mastra + mem0)
| Item | Freeze |
|---|---|
| Dataset | `longmemeval_s_cleaned` (HF `xiaowu0162/longmemeval-cleaned`; Wu et al. 2024, arXiv:2410.10813) |
| Dataset hash | build `a8a99545d77a236e3c7aa1f5d0ccfd94d4bcc5c2d5adbd19f938aba844586c56`; jsonl `f21f62027a10e7e08fecdb3386c5ec83f409a5e92056a48be68bbb5473e7f262` |
| Evaluated set | **500 instances, including the 30 abstention (`_abs`) questions.** No-session instances excluded in canonical build. All three systems run on the identical 500 through **our** harness |
| Answerer | Frozen per system; production-legal config (no oracle routing). Oracle numbers, if shown, are ablation-only and never carry "exceeds incumbent" language |
| Judge | LongMemEval reference judge, single model+prompt frozen (source path recorded in run manifest), applied identically to all three systems |
| Retrieval depth/budget | Held identical across the three systems (same top-k, same token budget); disclosed |
| Primary metric | **Micro accuracy** over 500 |
| Statistic | Paired, cluster-aware (question-session) bootstrap 95% CI; Holm-corrected across the 3-benchmark family |
| Claim shape | Competitive aggregate + per-type breakdown. **No LME SOTA claim** — the 500 informed method development; one final run does not restore test-set independence |
## Race 3 — BEAM (AMB harness race)
| Item | Freeze |
|---|---|
| Dataset | `mohammadtavakoli78/BEAM` (Tavakoli et al. 2024, arXiv:2510.27246, ICLR 2026) |
| Dataset hash | `beam-1M`: `cc81ed8f7624261a2fa43a335eb46159b7665074b82ddfb4e2c0783fbf2caa46`, 700 questions across **35 conversations**, 1M chat size |
| Evaluated set | **700 questions, dedup-first.** Full ordinal outcome retained per question |
| Answerer | Ours + control lanes (BM25 / oracle) through the AMB harness; frozen config |
| Judge | Nugget-graded judge; rubric nuggets carried per question; judge model+prompt frozen in run manifest. Judge ported to grade 0 / 0.5 / 1 per nugget |
| Retry policy | Max-1-on-empty, first-attempt primary (global rule). BEAM's earlier 56-failure regeneration is **not** repeated |
| Primary metric | **Mean nugget score** (0/0.5/1) over 700 |
| Secondary | Pass rate; **full ordinal distribution** (rate of 0, 0.5, 1); cost/token per question |
| Statistic | **Conversation-cluster bootstrap** 95% CI over 35 clusters (question-level McNemar is invalid here — 700 nested in 35). Aggregate win requires cluster-aware CI lower bound > 0. The contradiction-subset result (+23pp) is the headline and is reported with its own cluster CI |
| Claim shape | Contradiction-handling gain is primary; aggregate is descriptive with full distribution + cost. No permissive-threshold aggregate claim |
---
## Go / No-Go (from consensus §0.5)
- **GO** = all load-bearing smokes green (S0, S2, S5, S6 paired, S7) **and** S0 calibration clean.
- **S7:** green ≥ +3pp vs best comparable lenient result; red if tie/lose/<1pp.
- **S6:** Mastra repro < 41/50 kills comparability; ours trailing ≥ 3/50 = red.
- **S4:** red if mem0-OSS leads ≥ 5pp under strict LoCoMo.
- **S2:** BEAM aggregate dies unless cluster-aware CI lower bound > 0.
- **FALLBACK** to the contradiction+audit paper if any load-bearing red, two grays, or
judge calibration fails. **No averaging failures across benchmarks.**
## Claim freeze
Headline is **not** "SOTA on three leaderboards." It is: *conflict-preserving raw-turn
memory delivers a large, reproducible contradiction-handling gain under controlled
comparison, while a protocol-fidelity audit shows conversational-memory rankings are
underidentified (judges, routing metadata, retries, budgets).* Conditional upgrade only
if the LoCoMo matrix shows stable ordering under both protocols: "leads matched LoCoMo
and BEAM comparisons." Any comparison that cannot be expressed as a single frozen
protocol above is out of scope for this package.