# @waggle/benchmarks-harness **Four-cell ablation harness** for Waggle memory + prompt-evolution benchmarks. Canonical model: **Qwen/Qwen3.6-35B-A3B** (see `docs/plans/BACKLOG-MASTER-2026-04-18.md` §H12). ## Four cells — the ablation grid Same dataset, same seed, same model across all four. Difference at report time between a baseline cell and a treatment cell isolates the causal contribution of the ablated component. | Cell | Memory | Prompt evolution | Purpose | |------|--------|------------------|---------| | `raw` | no | no | Baseline — LLM only, stateless per turn. | | `filtered` | yes | no | Isolates the memory-layer contribution. (Sprint 12 rename: was `memory-only`.) | | `compressed` | no | yes | Isolates the GEPA prompt-evolution contribution. (Sprint 12 rename: was `evolve-only`.) | | `full-context` | yes | yes | Joint contribution (memory × evolution). (Sprint 12 rename: was `full-stack`.) | ## Controls (not cells — diagnostic) | Control | Purpose | |---------|---------| | `verbose-fixed` | Day-1 sanity check. Tells the model to answer verbosely. On a short-factoid benchmark, this should **underperform** `raw`. If it doesn't, audit the harness. | Out-of-scope for Week 1 (deferred to Week 2 / Week 3 per brief): - `naive-rag` control - `oracle-memory` ceiling - Llama 3.1 8B + Opus 4.6 model integrations - Gemma 2 9B probe - Full τ-bench + LongMemEval loaders (synthetic fallback works today) ## CLI Run from repo root via the `bench` script, or invoke `tsx` directly: ```bash # Day 1 sanity check — one cell, one instance, dry-run (no LLM required) npm run bench -- --cell raw --dataset locomo --limit 1 --model qwen3.6-35b-a3b # Day 2 pre-flight smoke — all 4 cells, 50 instances each npm run bench -- --all-cells --dataset locomo --limit 50 --model qwen3.6-35b-a3b --budget 115 # Full run npm run bench -- --all-cells --dataset locomo --full --model qwen3.6-35b-a3b # Verbose-fixed control (Day 1 sanity) npm run bench -- --control verbose-fixed --dataset locomo --limit 50 --model qwen3.6-35b-a3b ``` ### Flags | Flag | Default | Notes | |------|---------|-------| | `--cell ` | — | One of `raw \| filtered \| compressed \| full-context`. | | `--all-cells` | — | Run all four sequentially with the same dataset + seed. | | `--control ` | — | Currently only `verbose-fixed`. | | `--dataset ` | `synthetic` | `synthetic \| locomo \| longmemeval`. External datasets throw `DatasetMissingError` if the canonical archive is absent; set `BENCH_SYNTHETIC_DATASET=1` to re-enable the dev-only synthetic fallback. | | `--limit N` | `10` | Cap instances. `--full` = no cap. | | `--model ` | `qwen3.6-35b-a3b` | Id from `config/models.json`. | | `--seed N` | `42` | Reproducibility — same seed → same instance order + dry-run output. | | `--budget USD` | `Infinity` | Hard USD cap. Run stops when cumulative cost exceeds. | | `--output ` | auto | JSONL output path. Default: `../results/--.jsonl`. | | `--dry-run` | auto | Stub LLM. Default-on when `LITELLM_URL` env is unset. | | `--live` | — | Force real LLM calls even if `LITELLM_URL` unset. | | `--help`, `-h` | — | Usage summary. | ### Environment | Var | Default | |-----|---------| | `LITELLM_URL` | `http://localhost:4000` | | `LITELLM_API_KEY` | `sk-waggle-dev` | ## Output — per-instance JSONL One line per instance, flat shape (friendly to `jq`, DuckDB, pandas): ```jsonl {"turnId":"9fcd1d25-979f-4094-9633-f9bc30471f08","cell":"raw","instance_id":"synthetic_001","model":"qwen3.6-35b-a3b","seed":42,"accuracy":1,"p50_latency_ms":1,"p95_latency_ms":1,"usd_per_query":0.00012,"failure_mode":null} ``` **turnId correlation** — the `turnId` field is a UUID v4 matching the per-turn trace ID generated by the production agent orchestrator (`packages/agent/src/turn-context.ts`, H-AUDIT-1). For `raw` and control cells, where the agent loop isn't exercised, the harness generates the turnId itself so every row has a correlation key. ## Aggregate summary Written alongside the JSONL as `.summary.json`: ```json { "run": { "kind": "cell", "name": "raw", "dataset": "synthetic", "model": "qwen3.6-35b-a3b", "seed": 42, ... }, "counts": { "total": 50, "completed": 50, "failed": 0, "budgetStoppedAt": null }, "metrics": { "meanAccuracy": 0.82, "p50LatencyMs": 230, "p95LatencyMs": 450, "totalUsd": 0.011, "meanUsdPerQuery": 0.00022 }, "failureModes": {} } ``` ## Datasets - `synthetic` — **built-in**, 60 instances, no external download required. Used by smoke tests and as a fallback when external data is missing. - `locomo` — expects `benchmarks/data/locomo/locomo.jsonl`. Gitignored. Downloaded separately (Week 1 work). - `longmemeval` — expects `benchmarks/data/longmemeval/longmemeval.jsonl`. Same pattern. ## Reproducibility `--seed N` fully determines: 1. The order instances are sampled from the dataset 2. The dry-run LLM stub's output (the stub is deterministic from the user prompt alone) For live LLM runs, `--seed` is still emitted into every JSONL record so downstream analysis can pin every row to a seed value; the model's own sampling is controlled via `temperature=0.0` in `src/llm.ts`. ## Smoke test ```bash npx vitest run benchmarks/harness/tests/smoke.test.ts ``` Covers every acceptance criterion from the Bucket 1 Task 7 brief.