5.2 KiB
@waggle/benchmarks-harness
Four-cell ablation harness for Waggle memory + prompt-evolution benchmarks.
Canonical model: Qwen/Qwen3.6-35B-A3B (see docs/plans/BACKLOG-MASTER-2026-04-18.md §H12).
Four cells — the ablation grid
Same dataset, same seed, same model across all four. Difference at report time between a baseline cell and a treatment cell isolates the causal contribution of the ablated component.
| Cell | Memory | Prompt evolution | Purpose |
|---|---|---|---|
raw |
no | no | Baseline — LLM only, stateless per turn. |
filtered |
yes | no | Isolates the memory-layer contribution. (Sprint 12 rename: was memory-only.) |
compressed |
no | yes | Isolates the GEPA prompt-evolution contribution. (Sprint 12 rename: was evolve-only.) |
full-context |
yes | yes | Joint contribution (memory × evolution). (Sprint 12 rename: was full-stack.) |
Controls (not cells — diagnostic)
| Control | Purpose |
|---|---|
verbose-fixed |
Day-1 sanity check. Tells the model to answer verbosely. On a short-factoid benchmark, this should underperform raw. If it doesn't, audit the harness. |
Out-of-scope for Week 1 (deferred to Week 2 / Week 3 per brief):
naive-ragcontroloracle-memoryceiling- Llama 3.1 8B + Opus 4.6 model integrations
- Gemma 2 9B probe
- Full τ-bench + LongMemEval loaders (synthetic fallback works today)
CLI
Run from repo root via the bench script, or invoke tsx directly:
# Day 1 sanity check — one cell, one instance, dry-run (no LLM required)
npm run bench -- --cell raw --dataset locomo --limit 1 --model qwen3.6-35b-a3b
# Day 2 pre-flight smoke — all 4 cells, 50 instances each
npm run bench -- --all-cells --dataset locomo --limit 50 --model qwen3.6-35b-a3b --budget 115
# Full run
npm run bench -- --all-cells --dataset locomo --full --model qwen3.6-35b-a3b
# Verbose-fixed control (Day 1 sanity)
npm run bench -- --control verbose-fixed --dataset locomo --limit 50 --model qwen3.6-35b-a3b
Flags
| Flag | Default | Notes |
|---|---|---|
--cell <name> |
— | One of raw | filtered | compressed | full-context. |
--all-cells |
— | Run all four sequentially with the same dataset + seed. |
--control <name> |
— | Currently only verbose-fixed. |
--dataset <id> |
synthetic |
synthetic | locomo | longmemeval. External datasets throw DatasetMissingError if the canonical archive is absent; set BENCH_SYNTHETIC_DATASET=1 to re-enable the dev-only synthetic fallback. |
--limit N |
10 |
Cap instances. --full = no cap. |
--model <id> |
qwen3.6-35b-a3b |
Id from config/models.json. |
--seed N |
42 |
Reproducibility — same seed → same instance order + dry-run output. |
--budget USD |
Infinity |
Hard USD cap. Run stops when cumulative cost exceeds. |
--output <path> |
auto | JSONL output path. Default: ../results/<cell>-<dataset>-<ts>.jsonl. |
--dry-run |
auto | Stub LLM. Default-on when LITELLM_URL env is unset. |
--live |
— | Force real LLM calls even if LITELLM_URL unset. |
--help, -h |
— | Usage summary. |
Environment
| Var | Default |
|---|---|
LITELLM_URL |
http://localhost:4000 |
LITELLM_API_KEY |
sk-waggle-dev |
Output — per-instance JSONL
One line per instance, flat shape (friendly to jq, DuckDB, pandas):
{"turnId":"9fcd1d25-979f-4094-9633-f9bc30471f08","cell":"raw","instance_id":"synthetic_001","model":"qwen3.6-35b-a3b","seed":42,"accuracy":1,"p50_latency_ms":1,"p95_latency_ms":1,"usd_per_query":0.00012,"failure_mode":null}
turnId correlation — the turnId field is a UUID v4 matching the
per-turn trace ID generated by the production agent orchestrator
(packages/agent/src/turn-context.ts, H-AUDIT-1). For raw and control
cells, where the agent loop isn't exercised, the harness generates the
turnId itself so every row has a correlation key.
Aggregate summary
Written alongside the JSONL as <name>.summary.json:
{
"run": { "kind": "cell", "name": "raw", "dataset": "synthetic", "model": "qwen3.6-35b-a3b", "seed": 42, ... },
"counts": { "total": 50, "completed": 50, "failed": 0, "budgetStoppedAt": null },
"metrics": { "meanAccuracy": 0.82, "p50LatencyMs": 230, "p95LatencyMs": 450, "totalUsd": 0.011, "meanUsdPerQuery": 0.00022 },
"failureModes": {}
}
Datasets
synthetic— built-in, 60 instances, no external download required. Used by smoke tests and as a fallback when external data is missing.locomo— expectsbenchmarks/data/locomo/locomo.jsonl. Gitignored. Downloaded separately (Week 1 work).longmemeval— expectsbenchmarks/data/longmemeval/longmemeval.jsonl. Same pattern.
Reproducibility
--seed N fully determines:
- The order instances are sampled from the dataset
- The dry-run LLM stub's output (the stub is deterministic from the user prompt alone)
For live LLM runs, --seed is still emitted into every JSONL record so
downstream analysis can pin every row to a seed value; the model's own
sampling is controlled via temperature=0.0 in src/llm.ts.
Smoke test
npx vitest run benchmarks/harness/tests/smoke.test.ts
Covers every acceptance criterion from the Bucket 1 Task 7 brief.