Files
waggle-os/docs/paper/2026-07-11-protocol-freeze.md
Oleg Maslov b20b138fe4 moving
2026-09-02 10:14:22 +02:00

7.8 KiB
Raw Permalink Blame History

Protocol & Claim Freeze (2026-07-11)

Pre-registration for the three planned races. Everything below is fixed before any race result exists. Companion: 2026-07-11-fable-codex-consensus.md (§0.5 smoke tier, go/no-go). Calibration evidence: KorroResearch/benchmarks/locomo-mem0-parity-2026-07/calib/ (panel_input.jsonl, panel_verdicts.jsonl, analyze_panel.py — this smoke, Part 2).

Global rules (all three races).

  • One evaluation run per frozen config. No reruns-until-win. If a run is voided it is voided for a disclosed operational reason (crash, auth failure), not because of its score, and the void is logged.
  • Retry policy. Max 1 regeneration, triggered only on a literally empty answer string (finish_reason truncation or empty content). Applied symmetrically wherever we control the pipeline. First-attempt result is the reported primary; retry-normalized is a disclosed secondary. Competitor artifacts we do not control (mem0 released answers) get first-attempt-primary treatment with the asymmetry disclosed — a true symmetric rerun of a managed platform is infeasible.
  • Statistics. Paired tests only, cluster-aware at the conversation level (questions are nested in conversations; question-level independence is false). Report a 95% CI whose lower bound must clear the go threshold. Conversation-cluster bootstrap (10k resamples) is the primary interval; McNemar is reported but its naive p-value is treated as anti-conservative and never the sole basis for a claim.
  • Multiplicity. Three benchmarks × one primary metric each = 3 primary tests. HolmBonferroni across the 3 primaries; per-benchmark secondaries are descriptive, not claim-bearing.
  • Systems enter only with a frozen, hashed answer artifact. A system with no reproducible per-question answer file does not enter the matrix.

Race 1 — LoCoMo (4 systems × 2 judge protocols)

Item Freeze
Dataset snap-research locomo10.json (Maharana et al., ACL 2024). Canonical build locomo-1540
Dataset hash build 39e415e2f3a0fa1bd3cb1804a58d0b440b50d3070b2100698437e4ec402a5b24; canonical instance_count 1531
Evaluated set 1540 rows (mem0's own denominator: cat1=282, cat2=321, cat3=96, cat4=841). The 1531→1540 gap = 9 duplicate-question rows mem0 scores separately; we match their denominator for parity. Cat 5 adversarial (446) excluded
Answer artifacts Ours = locomo-7lane-w4-judgments-N1540.jsonl (answer_content), hash 1252fbde90613ebb5622f6724def7e60a433f4d8bf754829c0a0592972f079fa. mem0 = released per-question verdicts mem0_perq_verdicts.json (their answers held, joined on conv_idx + normalized question, gold tiebreak on 11 dup groups). Systems 3 & 4 enter only if a frozen answer file exists; else the matrix is 2×2 and reported as such
Answerer Per-system, frozen in each answer artifact. Not re-answered for this race — judge-only
Protocol axis (the 2) P1 mem0-lenient: mem0-memory-benchmarks/benchmarks/locomo/prompts.py JUDGE_PROMPT (partial-credit: ≥1 gold item ⇒ CORRECT), gold via preprocess_answer (cat-3 semicolon split). P2 Memori-strict: memori-repo/benchmarks/02_run_benchmark.ipynb cell 2 ACCURACY_PROMPT (verbatim), raw gold
Judge model P1 = gpt-5 (arm c reference) and gpt-4.1-mini (production arm b), both frozen; P2 = gpt-4.1-mini. Judge model held constant within a protocol column
Primary metric Judge accuracy (J-score), cats 1-4, on the 1540
Statistic Paired McNemar per system-pair plus conversation-cluster bootstrap 95% CI (10 conversations)
Claim shape Ordering stability across the two protocols, not "we win under the lenient judge." A protocol-specific win is reported as protocol-specific

The result this race is allowed to support: ranking is (or is not) stable when the judge protocol is swapped. Cross-protocol number-vs-number comparison is licensed only because S0 calibration cleared the substitution gate (below).

Race 2 — LongMemEval (one harness: ours + Mastra + mem0)

Item Freeze
Dataset longmemeval_s_cleaned (HF xiaowu0162/longmemeval-cleaned; Wu et al. 2024, arXiv:2410.10813)
Dataset hash build a8a99545d77a236e3c7aa1f5d0ccfd94d4bcc5c2d5adbd19f938aba844586c56; jsonl f21f62027a10e7e08fecdb3386c5ec83f409a5e92056a48be68bbb5473e7f262
Evaluated set 500 instances, including the 30 abstention (_abs) questions. No-session instances excluded in canonical build. All three systems run on the identical 500 through our harness
Answerer Frozen per system; production-legal config (no oracle routing). Oracle numbers, if shown, are ablation-only and never carry "exceeds incumbent" language
Judge LongMemEval reference judge, single model+prompt frozen (source path recorded in run manifest), applied identically to all three systems
Retrieval depth/budget Held identical across the three systems (same top-k, same token budget); disclosed
Primary metric Micro accuracy over 500
Statistic Paired, cluster-aware (question-session) bootstrap 95% CI; Holm-corrected across the 3-benchmark family
Claim shape Competitive aggregate + per-type breakdown. No LME SOTA claim — the 500 informed method development; one final run does not restore test-set independence

Race 3 — BEAM (AMB harness race)

Item Freeze
Dataset mohammadtavakoli78/BEAM (Tavakoli et al. 2024, arXiv:2510.27246, ICLR 2026)
Dataset hash beam-1M: cc81ed8f7624261a2fa43a335eb46159b7665074b82ddfb4e2c0783fbf2caa46, 700 questions across 35 conversations, 1M chat size
Evaluated set 700 questions, dedup-first. Full ordinal outcome retained per question
Answerer Ours + control lanes (BM25 / oracle) through the AMB harness; frozen config
Judge Nugget-graded judge; rubric nuggets carried per question; judge model+prompt frozen in run manifest. Judge ported to grade 0 / 0.5 / 1 per nugget
Retry policy Max-1-on-empty, first-attempt primary (global rule). BEAM's earlier 56-failure regeneration is not repeated
Primary metric Mean nugget score (0/0.5/1) over 700
Secondary Pass rate; full ordinal distribution (rate of 0, 0.5, 1); cost/token per question
Statistic Conversation-cluster bootstrap 95% CI over 35 clusters (question-level McNemar is invalid here — 700 nested in 35). Aggregate win requires cluster-aware CI lower bound > 0. The contradiction-subset result (+23pp) is the headline and is reported with its own cluster CI
Claim shape Contradiction-handling gain is primary; aggregate is descriptive with full distribution + cost. No permissive-threshold aggregate claim

Go / No-Go (from consensus §0.5)

  • GO = all load-bearing smokes green (S0, S2, S5, S6 paired, S7) and S0 calibration clean.
  • S7: green ≥ +3pp vs best comparable lenient result; red if tie/lose/<1pp.
  • S6: Mastra repro < 41/50 kills comparability; ours trailing ≥ 3/50 = red.
  • S4: red if mem0-OSS leads ≥ 5pp under strict LoCoMo.
  • S2: BEAM aggregate dies unless cluster-aware CI lower bound > 0.
  • FALLBACK to the contradiction+audit paper if any load-bearing red, two grays, or judge calibration fails. No averaging failures across benchmarks.

Claim freeze

Headline is not "SOTA on three leaderboards." It is: conflict-preserving raw-turn memory delivers a large, reproducible contradiction-handling gain under controlled comparison, while a protocol-fidelity audit shows conversational-memory rankings are underidentified (judges, routing metadata, retries, budgets). Conditional upgrade only if the LoCoMo matrix shows stable ordering under both protocols: "leads matched LoCoMo and BEAM comparisons." Any comparison that cannot be expressed as a single frozen protocol above is out of scope for this package.