Files
waggle-os/docs/briefs/2026-04-29-benchmark-portfolio-refresh-2026-venues.md
Oleg Maslov 0c3e2ead3b
Some checks failed
Installer Smoke / installer-smoke (push) Has been cancelled
moving
2026-09-02 10:10:29 +02:00

234 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
brief_id: 2026-04-29-benchmark-portfolio-refresh-2026-venues
date: 2026-04-29
session: PM coordination (Cowork)
mission: Refresh hive-mind / Waggle benchmark portfolio against Q1-Q2 2026 venue landscape; preserve PHF launch posture; sequence post-launch tracks.
predecessor_decisions:
- decisions/2026-04-26-decision-matrix-self-judge-reframe.md # PHF binding
- decisions/2026-04-27-phase-2-acceptance-gate-PASS.md
- decisions/2026-04-29-gepa-faza1-results.md
predecessor_strategy: strategy/BENCHMARK-STRATEGY.txt # 2026-04-18; primary plan, NOT superseded
predecessor_evidence:
- benchmarks/results/v6-self-judge-rebench/apples-to-apples-memo.md
- benchmarks/results/v6-self-judge-rebench/self-judge-vs-trio-comparison.md
- benchmarks/results/stage3-n400-v6-final-memo.md
- gepa-phase-5/manifest.yaml # claude::gen1-v1 + qwen-thinking::gen1-v1
status: AMENDMENT-PROPOSAL (does NOT supersede 04-18 BENCHMARK-STRATEGY; supplements with 2026 venue refresh + competitive intel update)
authority_required: PM (Marko Marković) ratification on §7 ratification asks
horizon: 12 weeks (pre-launch finalization → 6 weeks post-launch sequencing)
---
# PM Brief — Benchmark Portfolio Refresh: 2026 Venue Landscape
## TL;DR
The 04-18 BENCHMARK-STRATEGY.txt remains the binding primary plan. PHF launch posture (substrate ceiling 74.0 % vs Mem0 peer-reviewed 66.9 %, methodology contribution +27.35 pp) holds and ships Day 0 unchanged.
This brief proposes **three additive amendments** anchored on Q1 2026 benchmark venue developments that postdate the 04-18 strategy:
1. **Add Gaia2 (Meta SuperIntelligence Labs, arxiv 2602.11964, 12 Feb 2026)** as Phase 3 primary agent-harness venue — replaces SWE-ContextBench as headline target.
2. **Add τ³-bench banking_knowledge (Sierra, 18 Mar 2026)** as Phase 4 KVARK-track venue — replaces proprietary BPMN-workflow benchmark.
3. **Inherit ERL methodology (ICLR 2026 MemAgents Workshop)** as the publication framing for Waggle self-evolution claim — eliminates need to invent new "self-improvement convergence" metric.
Plus one update to competitive intelligence: **Hermes Agent (Nous Research, 25 Feb 2026)** is now an architectural-philosophy competitor to Waggle, not in 03-March intel doc.
Five ratification asks in §7. No code or run actions before PM response.
---
## §1 — What does NOT change
The following are LOCKED and this brief does not propose modifications:
- **PHF claim and Day 0 narrative** — substrate ceiling 74.0 % vs Mem0 peer-reviewed 66.9 % / 68.4 %, +27.35 pp methodology bias quantification, V1 retrieval honest 48.25 %. Source: `decisions/2026-04-26-decision-matrix-self-judge-reframe.md`.
- **Coupled launch sequencing** — arxiv preprint + hive-mind public + Waggle landing + Stripe in a single Day 0 window.
- **Pricing** — Solo Free / Pro $19 / Teams $49 (LOCKED 04-18).
- **GEPA Phase 5 canary deployment** — claude::gen1-v1 + qwen-thinking::gen1-v1 in flight, scope LOCKED. Cost amendment ratified 04-30 ("stavi visi slobodno").
- **Stage 3 v6 N=400 LoCoMo** — closed PASS-WITH-HONEST-FRAMING. No re-run proposed.
- **arxiv paper structure** — `research/2026-04-26-arxiv-paper/` outline and skeleton remain primary author surface.
The portfolio refresh is **post-launch sequencing**, not pre-launch revision.
---
## §2 — Landscape changes since 04-18 (binding new evidence)
### 2.1 Gaia2 (Meta SuperIntelligence Labs)
**Anchor:** arxiv 2602.11964 (Froger et al., 12 Feb 2026); ARE platform repo `facebookresearch/meta-agents-research-environments`.
**What it measures:** asynchronous agent capability in a simulated mobile environment with 12 applications and 101 tools. Agents must operate under temporal constraints, adapt to noisy/dynamic events, resolve ambiguity, and collaborate. Pass@1 with write-action verifier per scenario.
**Current SOTA (Feb-Apr 2026):**
- GPT-5 (high): 42 % pass@1 (best overall; fails on time-sensitive tasks)
- Claude-4 Sonnet: trades accuracy/speed/cost
- Kimi-K2: **21 % pass@1 — open-source SOTA**
- No Anthropic dominance; no saturation.
**Why it replaces SWE-ContextBench in our portfolio:**
- Fresh venue (post-04-18); SWE-ContextBench is now Q4 2025 vintage and OpenClaw / Hermes have not engaged it either, so first-mover narrative is weaker.
- Agent capability domain matches Waggle product surface (general agentic tool use with persistent memory) better than SWE-ContextBench (code-context retrieval narrow scope).
- Universes architecture (isolated data partitions exposing identical tools but disjoint task content) provides clean substrate for self-evolution measurement — see §2.3.
- Open-source SOTA threshold of 21 % is realistic to beat with Qwen 3.6 35B + GEPA-evolved `qwen-thinking::gen1-v1` (+12.5 pp uplift validated in-sample n=8 + held-out n=5). Target band: 30-35 % pass@1, which enters Claude/GPT-5 reference zone.
**Cost estimate:** N=200 agent-task instances × ~$0.05/instance subject + judge harness = ~$25-40 per full run. Compute envelope manageable within existing GEPA Phase 5 cost amendment.
### 2.2 τ³-bench banking_knowledge (Sierra)
**Anchor:** Sierra Research blog 18 Mar 2026; `sierra-research/tau2-bench` repo; τ-Knowledge paper (Shi et al., arxiv 2603.04370).
**What it measures:** RAG-augmented customer service in banking domain. Configurable retrieval pipelines (keyword search, embedding-based, long-context, agentic shell-based). Task success measured by correctness of backend database state changes (dispute opened, card frozen, credit issued), not conversation polish. Pass^k metric for reliability.
**Current SOTA:**
- GPT-5.2 with high reasoning: ~25 % task success.
- Even with exact required documents provided: ~40 %. Bottleneck is reasoning/execution, not retrieval.
**Why it replaces proprietary BPMN-workflow benchmark:**
- Sierra is a credentialed third-party venue; community-driven leaderboard at taubench.com with verified submissions via S3 bucket trajectories.
- Banking domain is direct match for KVARK enterprise sales (regulated industry, RAG over policy documents, audit trail of agent actions).
- Bottleneck is exactly where hive-mind should add value (frame importance weighting, bitemporal validity, I/P/B distinction for hypothesis-vs-fact reasoning).
- Headroom is large (~25 % SOTA → ceiling ~40 %); a measurable lift here is the easiest-to-defend KVARK pitch artifact for regulated buyers.
- Proprietary BPMN-workflow benchmark in 04-18 strategy has zero adoption, zero comparison anchor, zero credibility — even if we publish it, no one cites it.
**Cost estimate:** N=200 instances × ~$0.10/instance (longer dialogues with retrieval round-trips) = ~$30-50. Same envelope class as Gaia2.
### 2.3 ERL methodology (ICLR 2026 MemAgents Workshop)
**Anchor:** "Experiential Reflective Learning for Self-Improving LLM Agents" (arxiv 2603.24639, March 2026). Published as conference paper at the ICLR 2026 MemAgents Workshop.
**What it does:** retrieval of heuristics from accumulated experience, injected into agent's system prompt before execution. No modification to core ReAct loop. Evaluated on Gaia2 Search + Execution splits and τ²-bench (all three customer service domains).
**Reported result:** +7.8 % success rate uplift over ReAct baseline on Gaia2; large gains in task completion reliability; outperforms prior experiential learning methods (ExpeL, AutoGuide, Reflexion).
**Why this matters for our portfolio:**
- The 04-18 strategy implies Waggle self-evolution claim needs a custom evaluation methodology. ERL provides the methodology already, with a published baseline (+7.8 %) to beat.
- Our hive-mind frame architecture (I/P/B, importance weighting) maps cleanly onto ERL's "selective retrieval of transferable heuristics" framing — this is publishable as an ERL extension, not as a separate framework.
- MemAgents Workshop venue exists and accepts work; we have a valid conference submission target instead of inventing a venue.
- Avoids the "we invented a metric to measure ourselves" credibility problem flagged in earlier prep work.
**Implication for Waggle launch comms:** the self-evolution claim moves from "trust us, internal benchmark shows X" to "validated against published ERL baseline on Gaia2". Order-of-magnitude credibility upgrade.
---
## §3 — Competitive intelligence update
The current `Waggle_Competitive_Intelligence_Full_Landscape_March_2026.docx` is dated. One material gap requires update before Day 0 comms freeze.
### 3.1 Hermes Agent (Nous Research)
**Launch date:** 25 February 2026. **Star count:** 110 K within 10 weeks of launch. **License:** open source.
**Architecture:** closed learning loop, prompt memory (MEMORY.md, USER.md), episodic archive (SQLite FTS5), procedural skills (auto-generated markdown). Internal benchmarks claim 40 % speedup on repeat tasks.
**Why this is material:** Hermes Agent occupies the same architectural-philosophy space as Waggle. The Hermes pitch is "agent that gets better over time at your specific workflows through closed learning loop". This is functionally identical to our self-evolution narrative.
**Defensible Waggle differentiators against Hermes (must appear in Day 0 comms):**
1. Bitemporal knowledge graph (Hermes uses flat SQLite FTS5).
2. I/P/B frame model with importance weighting and superseding-via-correction (Hermes does not distinguish hypothesis from fact).
3. MPEG-4 frame architecture and wiki compiler (Hermes has neither).
4. Apache 2.0 hive-mind as standalone npm package (`@hive-mind/core` etc.) — Hermes is monolithic.
5. EU AI Act audit triggers built-in (Hermes does not address).
6. **Published peer-reviewed-style benchmark results (apples-to-apples Mem0 + ERL methodology + Gaia2 + τ³)** — Hermes publishes only internal benchmarks.
Differentiator #6 is the moat. Hermes Agent has not engaged any standardized public benchmark venue. If we ship arxiv + Gaia2 + τ³ within Q2, the gap is unbridgeable for them in 2026.
### 3.2 OpenClaw security posture (no new evidence required)
OpenClaw March 2026 CVE cluster (9 CVEs in 4 days, including CVSS 9.9; Snyk flagged 1,467 malicious skills on ClawHub) is already in the existing intelligence doc per CC-1 audit. Confirming it remains in Day 0 narrative for regulated-industry pitches as "incumbent insecurity" framing.
---
## §4 — Recommended portfolio amendment
Replace BENCHMARK-STRATEGY.txt §3.4 (Phase 3) and §3.5 (Phase 4) primary venues. All other sections remain intact.
| Phase | 04-18 strategy | Proposed amendment | Rationale |
|---|---|---|---|
| Phase 0 (now → launch) | Stripe priority; nothing else | **Unchanged.** | PHF posture stable. |
| Phase 1 (post-Stripe, hive-mind alpha) | LoCoMo + bootstrap | **Unchanged.** Stage 3 v6 already complete. | Status quo. |
| Phase 2 (hive-mind launch) | LongMemEval + blog + GitHub public | **Unchanged.** | Coupled launch as PHF binds. |
| Phase 3 (Waggle benchmark integration) | SWE-bench sequential + SWE-ContextBench | **REPLACE with Gaia2 Search + Execution splits**, GEPA-evolved variants, ERL methodology framing. Target: 30-35 % pass@1 (open-source SOTA = 21 %). | Fresher venue, better domain fit, ERL publication target. |
| Phase 4 (KVARK milestones) | Proprietary BPMN-workflow + scale benchmarks | **REPLACE BPMN with τ³-bench banking_knowledge**, retain scale + multi-tenant + compliance latency benchmarks. Target: top-3 open-source on banking_knowledge. | Real venue, real comparison, regulated-industry sales artifact. |
Stretch targets in 04-18 strategy (BEAM 1M-token, SWE-ContextBench Memory track) deferred to Q3 2026 review.
---
## §5 — Sequencing (12-week horizon)
**Weeks 0-2 (now → hive-mind alpha):** PHF locked artifacts ship — arxiv preprint, hive-mind public, Waggle landing, Stripe. **No new benchmark work in this window.**
**Weeks 2-4 (post-launch consolidation):** Update `Waggle_Competitive_Intelligence_Full_Landscape_*` with Hermes Agent entry. Re-run any pitch deck slides that reference the outdated competitive landscape.
**Weeks 4-8 (Phase 3 Gaia2 sprint):**
- Week 4: Set up ARE platform locally; verify GEPA-evolved `qwen-thinking::gen1-v1` runs against Gaia2 Search split with no harness modification.
- Week 5: ERL-style heuristic retrieval wiring from hive-mind into agent system prompt (existing `retrieval-agent-loop.ts` is the integration point — 38.3 KB file already does adjacent work).
- Week 6: N=200 dry run on Search split; cost validation under $50.
- Week 7: Full Search + Execution split run, both ReAct baseline and ERL-augmented; trio-strict + self-judge dual reporting per PHF methodology lesson.
- Week 8: Results memo + arxiv submission to MemAgents Workshop or follow-on venue.
**Weeks 8-12 (Phase 4 τ³ sprint, KVARK track):**
- Week 8: Set up tau2-bench locally with banking_knowledge extras (`uv sync --extra knowledge`).
- Week 9: hive-mind retrieval pipeline integration as RAG provider; verify it satisfies tau2-bench `RetrievalProvider` interface.
- Week 10: N=100 dry run; calibrate per-task cost and latency.
- Week 11: Full N=200 run, frontier subject (Opus 4.7 + GPT-5.4) + Qwen subject for sovereignty story.
- Week 12: Submit results to taubench.com community leaderboard; produce KVARK enterprise sales one-pager with verified third-party broj.
---
## §6 — Risks and out-of-scope items
### 6.1 Risk: GEPA +12.5 pp uplift on N=13 may not generalize to Gaia2 task distribution
The held-out validation is statistically thin (N=5 held-out + N=8 in-sample). Gaia2 task distribution differs materially from in-sample evolution corpus (mobile environment, 12 apps, 101 tools vs analytical scenarios). Mitigation: Phase 3 sprint Week 6 dry run is the explicit checkpoint; if uplift collapses, halt and PM-escalate before Week 7 full run. Cost exposure if abort: ~$15.
### 6.2 Risk: ERL methodology reference point may shift before Workshop submission
The +7.8 % uplift is from the ERL paper as published. Other ERL extensions may publish between now and our submission window. Mitigation: framing should be "we extend ERL with bitemporal-KG-conditioned retrieval", not "we beat ERL by X". Defensible regardless of intermediate competitor work.
### 6.3 Risk: Hermes Agent or OpenClaw publish on Gaia2 / τ³ before us
Probability: low for OpenClaw (CVE remediation is consuming community bandwidth); medium for Hermes (Nous Research has paper-publishing track record). Mitigation: weeks 4-8 timeline above is aggressive; if Hermes publishes first, framing pivots to "Waggle vs Hermes head-to-head on Gaia2" rather than first-mover. Either way the published broj is the enterprise sales artifact.
### 6.4 Out of scope (explicitly)
- Any change to PHF claim, Day 0 narrative, coupling decision, or pricing.
- Frontier subject re-run of Stage 3 LoCoMo. (Earlier consideration deprecated by 04-25 self-judge re-eval evidence.)
- SWE-bench sequential learning curve experiment (Phase 3 in 04-18 strategy). Deferred to Q3 review pending Phase 3 Gaia2 results.
- New benchmark venue invention (StuLife, J-TTL, FieldWorkArena). Stick to community-recognized venues.
---
## §7 — Ratification asks
PM ratification required on the following five items before any Phase 3 or Phase 4 sprint kickoff. None blocks Day 0 launch.
1. **Ratify Gaia2 as Phase 3 primary agent-harness benchmark venue**, replacing SWE-ContextBench. (Y/N)
2. **Ratify τ³-bench banking_knowledge as Phase 4 KVARK-track primary venue**, replacing proprietary BPMN-workflow benchmark. (Y/N)
3. **Ratify ERL methodology inheritance** as the framing for Waggle self-evolution claim, with publication target = ICLR 2026 MemAgents Workshop or comparable venue. (Y/N)
4. **Ratify Hermes Agent competitive intelligence amendment** (§3.1) as binding update to `Waggle_Competitive_Intelligence_Full_Landscape_*` document. PM authorizes Marketing-side rewrite or assigns to CC. (Y/N + assignee)
5. **Ratify 12-week sequencing** in §5, with Weeks 4-12 Phase 3 + Phase 4 sprints contingent on successful Day 0 launch and post-launch consolidation Weeks 2-4. (Y/N)
After ratification, this brief becomes binding addendum to BENCHMARK-STRATEGY.txt; phase tables in §3.4 and §3.5 of that document are superseded by §4 of this brief. All other sections of 04-18 strategy remain primary.
---
## §8 — Cross-references
- 04-18 primary strategy: `strategy/BENCHMARK-STRATEGY.txt` (NOT superseded; supplemented).
- PHF binding decision: `decisions/2026-04-26-decision-matrix-self-judge-reframe.md`.
- Apples-to-apples Mem0 evidence: `benchmarks/results/v6-self-judge-rebench/apples-to-apples-memo.md` (in waggle-os repo).
- GEPA Phase 5 substrate: `gepa-phase-5/manifest.yaml` + `gepa-phase-5/preflight-evidence.md` (in waggle-os repo).
- arxiv paper anchor: `research/2026-04-26-arxiv-paper/00-paper-outline.md`.
- Existing competitive intel: `Waggle_Competitive_Intelligence_Full_Landscape_March_2026.docx` (in waggle-os repo root).
- Gaia2 paper: arxiv 2602.11964.
- τ³-bench / τ-Knowledge paper: arxiv 2603.04370.
- ERL paper: arxiv 2603.24639.
- Sierra leaderboard: taubench.com.
- ARE platform: github.com/facebookresearch/meta-agents-research-environments.
---
(2,847 words)