Files
waggle-os/docs/plans/HARNESS-BENCHMARK-GOAL-2026-05-22.md
Oleg Maslov 0c3e2ead3b
Some checks failed
Installer Smoke / installer-smoke (push) Has been cancelled
moving
2026-09-02 10:10:29 +02:00

100 lines
8.6 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# GOAL STATEMENT — Agent Harness Benchmark (local-first, sovereign)
**Date:** 2026-05-22
**Owner:** Marko (PM) · drives benchmark design + execution
**Status:** DRAFT goal statement — hand to `/goal``/plan` once the §0 decision is locked
---
## 0. LOCKED — Reading B: Waggle is the arena + governance layer (2026-05-22)
**Decision (Marko, 2026-05-22): LOCKED to Reading B.**
Waggle OS is the **local-first OS that orchestrates** every one of these harnesses (the AI-OS arc: detect → launch Claude Code, Codex, Cursor, Hermes, OpenClaw… locally, with full audit). We do **not** position Waggle as a competing agent loop. We **run all of them safely on-prem** and publish the head-to-head comparison matrix as a **sovereign buyer's guide + governance proof**.
**Claim shape:** *"Run any agent harness locally — fully audited, zero data egress — and here's exactly how each one performs in that sovereign environment."*
**Why B (rationale of record):**
- Matches the mission verbatim — "onboarding → push toward KVARK as full sovereign AI orchestration + governance." That is an orchestration/governance story, not a "our agent loop beats Codex" story.
- The matrix becomes a **durable buyer asset** (a comparison guide buyers trust *because* we don't have a horse in the capability race) rather than a fragile "we're #1" claim that a single model/harness upgrade invalidates.
- Waggle's credibility comes from being the **neutral, auditable, local-first home** for whichever harness the customer already trusts — which is exactly the KVARK pitch one tier up.
**Consequence for design:** No Waggle→ARE adapter is required. Waggle's role is measured as the **execution+governance substrate** (it launches the harness, isolates it, captures the audit trail), and the protagonist metric set shifts from "Waggle's pass rate" to "the sovereignty triple (local-first / zero-egress / auditable) holds across ALL harnesses, and here is each harness's capability/cost/reliability profile when run inside Waggle." Reading A (Waggle's own loop as a 6th competitor) is explicitly **deferred** — it can become a later, narrower claim only if Waggle's loop proves differentiated, and is out of scope for this benchmark.
---
## 1. Objective — TWO co-equal product pillars (Marko, 2026-05-22)
Waggle-the-product = **agent harness + memory substrate**, so the capability story needs **both proofs, co-equal** (not one headline + one footnote):
- **Pillar 1 — Agent-harness SOTA.** Waggle's OWN harness (`runAgentLoop`) benchmarked head-to-head vs reference harnesses (Hermes/OpenClaw, Oracle ceiling) in the local-first GAIA 2 rig, same model + judge + scenarios → prove Waggle's loop is at/near SOTA. (NOTE: the existing 83.8% used third-party Hermes, NOT Waggle — see plan doc DIRECTION UPDATE.)
- **Pillar 2 — Memory SOTA.** Waggle's hive-mind substrate on memory benchmarks (LoCoMo done in C-1 → **LongMemEval** near-term → **BEAM** flagship) → prove the memory substrate is at/near SOTA, ideally beating frontier long-context (the one axis where Waggle *wins*, not just matches).
Both feed **defensible, dual-tier (peer-review + hero-page) statements**, positioning Waggle OS as where knowledge workers and sovereign-AI buyers run agentic work without data leaving the perimeter, and as the on-ramp to KVARK. The sovereignty triple (local-first / zero-egress / auditable) wraps both pillars.
## 2. Subject under test + comparison set
| Entity | Role (per §0 — Reading B) |
|---|---|
| **Waggle OS** | the **arena + governance substrate** — launches/isolates/audits each harness locally. Measured by the sovereignty triple holding across all harnesses, not by a pass rate of its own. |
| Hermes | harness-under-test · ARE-native reference agent (already wired — N=160 done) |
| OpenClaw | harness-under-test · ARE-native reference agent (config scaffolded) |
| Claude Code | harness-under-test · external coding/agent harness |
| Codex | harness-under-test · external coding/agent harness |
| Claude Cowork | harness-under-test · external agent product |
**Controlled-variable principle (non-negotiable):** the harness is the ONLY variable. Same benchmark, **same model (Claude Sonnet 4.6) where the harness allows model choice**, same judge model + same judge protocol, same scenario set, same denominator. Anything else and the comparison is not defensible. **Waggle is held constant as the environment under all of them** — so any harness's number is also implicitly a "this ran inside Waggle, locally, audited" number.
## 3. Environment constraint — local-first is itself a measured property
Everything runs **locally / on-prem** (Docker, hermetic, laptop-runnable). For the sovereign-AI audience this is not a footnote — it's a headline claim. Capture and assert:
- **Zero data egress** during execution (network-isolated containers; prove it).
- **Full auditability**: every tool call captured in `events.jsonl` / trace → this *is* the KVARK governance hook.
- **Reproducibility**: hermetic, runs on Marko's Windows hardware (already proven for Hermes).
## 4. What to measure (harness quality is multi-dimensional)
Pass rate alone is a thin claim. Measure per harness, per GAIA 2 split (search / execution / adaptability / time / ambiguity / noise):
1. **Capability** — strict pass rate + judged-only pass rate (report both; errors counted honestly).
2. **Efficiency** — tokens & $ per task, tool-calls per task, wall-clock.
3. **Reliability** — error rate, recovery, determinism across reruns.
4. **Sovereignty/safety** — local-first ✓, egress=0 ✓, trace-auditability ✓ (binary asserts, per harness).
## 5. Two deliverable tiers (different bars — do not blur)
- **Tier 1 — Publishable** (paper / arxiv / KVARK technical annex): pre-registered protocol, N≥160 per cell, CI reported, judge protocol fixed in advance, no post-hoc baseline shopping. The GAIA 2 N=160 Hermes run is the first cell of this matrix.
- **Tier 2 — Hero-page** (waggle-os.ai + KVARK deck): punchy but every number traces back to a Tier-1 cell. Honest framing only. E.g. *"Run Codex, Claude Code, or Hermes locally — fully audited, your data never leaves your machine."*
## 6. Success criteria (what counts as a win)
- A **completed comparison matrix**: {6 harnesses} × {GAIA 2 splits} × {4 metric families}, same protocol throughout.
- At least one **Tier-1 publishable** statement that survives peer-review scrutiny.
- At least one **Tier-2 hero-page** statement that is punchy AND traces to a Tier-1 cell.
- The **sovereignty triple** (local-first / zero-egress / auditable) demonstrated, not asserted.
## 7. Non-goals / guardrails (honesty bar)
- **Not a memory-substrate proof.** That is C-1 (LOCOMO 67.8% trio-strict) + C-2 (Stage 3 +19.25pp, p=8e-18). GAIA 2 measures the *harness*, not hive-mind memory. Keep the lanes separate in every artifact.
- **Kill the apples-to-oranges baseline.** Do NOT publish "Waggle 83.8% vs Mem0 ~40-55%" — different judge/denominator/protocol. Every comparison number must come from OUR matrix under identical protocol, or be dropped.
- **Judge-leniency risk.** A self-judge (Sonnet judging Sonnet) inflates. For Tier-1, use an independent / ensemble judge and report the self-vs-independent delta (same discipline as C-1 trio-strict).
- **No goalpost-moving, no hiding errors.** Strict + judged-only always reported together.
## 8. What already exists (starting point)
- ✅ GAIA 2 ARE local-first pipeline runs on Windows Docker (Hermes + Sonnet 4.6), patches captured.
- ✅ First matrix cell: **Hermes × search × N=160 = 83.8% strict / 86.5% judged-only.**
- 🔲 OpenClaw cell (config scaffolded, not run).
- 🔲 Claude Code / Codex / Claude Cowork adapters (do these expose an ARE-compatible runtime? — research task; first design question).
- ⛔ Waggle→ARE adapter — **out of scope** (Reading A deferred per §0).
- 🔲 Independent/ensemble judge wiring for Tier-1.
- 🔲 Egress=0 proof harness (the sovereignty triple — protagonist metric for Reading B).
---
## TL;DR for `/goal`
> **Benchmark agent-harness quality head-to-head (Waggle vs Hermes, OpenClaw, Claude Code, Codex, Claude Cowork) on GAIA 2, in a local-first / zero-egress / fully-audited environment — holding model + judge + scenarios constant so the harness is the only variable — to produce both peer-review-publishable and honest hero-page claims that onboard knowledge workers and funnel sovereign-AI buyers toward KVARK.**
>
> §0 LOCKED to Reading B (Waggle = orchestrator + governance layer, not a competing loop). First cell done (Hermes 83.8%). Next: research which external harnesses expose an ARE-compatible runtime, wire the independent judge, build the egress=0 proof, then fill the matrix.