Files
waggle-os/docs/briefs/2026-04-29-benchmark-portfolio-refresh-2026-venues.md
Oleg Maslov 0c3e2ead3b
Some checks failed
Installer Smoke / installer-smoke (push) Has been cancelled
moving
2026-09-02 10:10:29 +02:00

17 KiB
Raw Permalink Blame History

brief_id, date, session, mission, predecessor_decisions, predecessor_strategy, predecessor_evidence, status, authority_required, horizon
brief_id date session mission predecessor_decisions predecessor_strategy predecessor_evidence status authority_required horizon
2026-04-29-benchmark-portfolio-refresh-2026-venues 2026-04-29 PM coordination (Cowork) Refresh hive-mind / Waggle benchmark portfolio against Q1-Q2 2026 venue landscape; preserve PHF launch posture; sequence post-launch tracks.
decisions/2026-04-26-decision-matrix-self-judge-reframe.md
decisions/2026-04-27-phase-2-acceptance-gate-PASS.md
decisions/2026-04-29-gepa-faza1-results.md
strategy/BENCHMARK-STRATEGY.txt
benchmarks/results/v6-self-judge-rebench/apples-to-apples-memo.md
benchmarks/results/v6-self-judge-rebench/self-judge-vs-trio-comparison.md
benchmarks/results/stage3-n400-v6-final-memo.md
gepa-phase-5/manifest.yaml
AMENDMENT-PROPOSAL (does NOT supersede 04-18 BENCHMARK-STRATEGY; supplements with 2026 venue refresh + competitive intel update) PM (Marko Marković) ratification on §7 ratification asks 12 weeks (pre-launch finalization → 6 weeks post-launch sequencing)

PM Brief — Benchmark Portfolio Refresh: 2026 Venue Landscape

TL;DR

The 04-18 BENCHMARK-STRATEGY.txt remains the binding primary plan. PHF launch posture (substrate ceiling 74.0 % vs Mem0 peer-reviewed 66.9 %, methodology contribution +27.35 pp) holds and ships Day 0 unchanged.

This brief proposes three additive amendments anchored on Q1 2026 benchmark venue developments that postdate the 04-18 strategy:

  1. Add Gaia2 (Meta SuperIntelligence Labs, arxiv 2602.11964, 12 Feb 2026) as Phase 3 primary agent-harness venue — replaces SWE-ContextBench as headline target.
  2. Add τ³-bench banking_knowledge (Sierra, 18 Mar 2026) as Phase 4 KVARK-track venue — replaces proprietary BPMN-workflow benchmark.
  3. Inherit ERL methodology (ICLR 2026 MemAgents Workshop) as the publication framing for Waggle self-evolution claim — eliminates need to invent new "self-improvement convergence" metric.

Plus one update to competitive intelligence: Hermes Agent (Nous Research, 25 Feb 2026) is now an architectural-philosophy competitor to Waggle, not in 03-March intel doc.

Five ratification asks in §7. No code or run actions before PM response.


§1 — What does NOT change

The following are LOCKED and this brief does not propose modifications:

  • PHF claim and Day 0 narrative — substrate ceiling 74.0 % vs Mem0 peer-reviewed 66.9 % / 68.4 %, +27.35 pp methodology bias quantification, V1 retrieval honest 48.25 %. Source: decisions/2026-04-26-decision-matrix-self-judge-reframe.md.
  • Coupled launch sequencing — arxiv preprint + hive-mind public + Waggle landing + Stripe in a single Day 0 window.
  • Pricing — Solo Free / Pro $19 / Teams $49 (LOCKED 04-18).
  • GEPA Phase 5 canary deployment — claude::gen1-v1 + qwen-thinking::gen1-v1 in flight, scope LOCKED. Cost amendment ratified 04-30 ("stavi visi slobodno").
  • Stage 3 v6 N=400 LoCoMo — closed PASS-WITH-HONEST-FRAMING. No re-run proposed.
  • arxiv paper structureresearch/2026-04-26-arxiv-paper/ outline and skeleton remain primary author surface.

The portfolio refresh is post-launch sequencing, not pre-launch revision.


§2 — Landscape changes since 04-18 (binding new evidence)

2.1 Gaia2 (Meta SuperIntelligence Labs)

Anchor: arxiv 2602.11964 (Froger et al., 12 Feb 2026); ARE platform repo facebookresearch/meta-agents-research-environments.

What it measures: asynchronous agent capability in a simulated mobile environment with 12 applications and 101 tools. Agents must operate under temporal constraints, adapt to noisy/dynamic events, resolve ambiguity, and collaborate. Pass@1 with write-action verifier per scenario.

Current SOTA (Feb-Apr 2026):

  • GPT-5 (high): 42 % pass@1 (best overall; fails on time-sensitive tasks)
  • Claude-4 Sonnet: trades accuracy/speed/cost
  • Kimi-K2: 21 % pass@1 — open-source SOTA
  • No Anthropic dominance; no saturation.

Why it replaces SWE-ContextBench in our portfolio:

  • Fresh venue (post-04-18); SWE-ContextBench is now Q4 2025 vintage and OpenClaw / Hermes have not engaged it either, so first-mover narrative is weaker.
  • Agent capability domain matches Waggle product surface (general agentic tool use with persistent memory) better than SWE-ContextBench (code-context retrieval narrow scope).
  • Universes architecture (isolated data partitions exposing identical tools but disjoint task content) provides clean substrate for self-evolution measurement — see §2.3.
  • Open-source SOTA threshold of 21 % is realistic to beat with Qwen 3.6 35B + GEPA-evolved qwen-thinking::gen1-v1 (+12.5 pp uplift validated in-sample n=8 + held-out n=5). Target band: 30-35 % pass@1, which enters Claude/GPT-5 reference zone.

Cost estimate: N=200 agent-task instances × ~$0.05/instance subject + judge harness = ~$25-40 per full run. Compute envelope manageable within existing GEPA Phase 5 cost amendment.

2.2 τ³-bench banking_knowledge (Sierra)

Anchor: Sierra Research blog 18 Mar 2026; sierra-research/tau2-bench repo; τ-Knowledge paper (Shi et al., arxiv 2603.04370).

What it measures: RAG-augmented customer service in banking domain. Configurable retrieval pipelines (keyword search, embedding-based, long-context, agentic shell-based). Task success measured by correctness of backend database state changes (dispute opened, card frozen, credit issued), not conversation polish. Pass^k metric for reliability.

Current SOTA:

  • GPT-5.2 with high reasoning: ~25 % task success.
  • Even with exact required documents provided: ~40 %. Bottleneck is reasoning/execution, not retrieval.

Why it replaces proprietary BPMN-workflow benchmark:

  • Sierra is a credentialed third-party venue; community-driven leaderboard at taubench.com with verified submissions via S3 bucket trajectories.
  • Banking domain is direct match for KVARK enterprise sales (regulated industry, RAG over policy documents, audit trail of agent actions).
  • Bottleneck is exactly where hive-mind should add value (frame importance weighting, bitemporal validity, I/P/B distinction for hypothesis-vs-fact reasoning).
  • Headroom is large (~25 % SOTA → ceiling ~40 %); a measurable lift here is the easiest-to-defend KVARK pitch artifact for regulated buyers.
  • Proprietary BPMN-workflow benchmark in 04-18 strategy has zero adoption, zero comparison anchor, zero credibility — even if we publish it, no one cites it.

Cost estimate: N=200 instances × ~$0.10/instance (longer dialogues with retrieval round-trips) = ~$30-50. Same envelope class as Gaia2.

2.3 ERL methodology (ICLR 2026 MemAgents Workshop)

Anchor: "Experiential Reflective Learning for Self-Improving LLM Agents" (arxiv 2603.24639, March 2026). Published as conference paper at the ICLR 2026 MemAgents Workshop.

What it does: retrieval of heuristics from accumulated experience, injected into agent's system prompt before execution. No modification to core ReAct loop. Evaluated on Gaia2 Search + Execution splits and τ²-bench (all three customer service domains).

Reported result: +7.8 % success rate uplift over ReAct baseline on Gaia2; large gains in task completion reliability; outperforms prior experiential learning methods (ExpeL, AutoGuide, Reflexion).

Why this matters for our portfolio:

  • The 04-18 strategy implies Waggle self-evolution claim needs a custom evaluation methodology. ERL provides the methodology already, with a published baseline (+7.8 %) to beat.
  • Our hive-mind frame architecture (I/P/B, importance weighting) maps cleanly onto ERL's "selective retrieval of transferable heuristics" framing — this is publishable as an ERL extension, not as a separate framework.
  • MemAgents Workshop venue exists and accepts work; we have a valid conference submission target instead of inventing a venue.
  • Avoids the "we invented a metric to measure ourselves" credibility problem flagged in earlier prep work.

Implication for Waggle launch comms: the self-evolution claim moves from "trust us, internal benchmark shows X" to "validated against published ERL baseline on Gaia2". Order-of-magnitude credibility upgrade.


§3 — Competitive intelligence update

The current Waggle_Competitive_Intelligence_Full_Landscape_March_2026.docx is dated. One material gap requires update before Day 0 comms freeze.

3.1 Hermes Agent (Nous Research)

Launch date: 25 February 2026. Star count: 110 K within 10 weeks of launch. License: open source.

Architecture: closed learning loop, prompt memory (MEMORY.md, USER.md), episodic archive (SQLite FTS5), procedural skills (auto-generated markdown). Internal benchmarks claim 40 % speedup on repeat tasks.

Why this is material: Hermes Agent occupies the same architectural-philosophy space as Waggle. The Hermes pitch is "agent that gets better over time at your specific workflows through closed learning loop". This is functionally identical to our self-evolution narrative.

Defensible Waggle differentiators against Hermes (must appear in Day 0 comms):

  1. Bitemporal knowledge graph (Hermes uses flat SQLite FTS5).
  2. I/P/B frame model with importance weighting and superseding-via-correction (Hermes does not distinguish hypothesis from fact).
  3. MPEG-4 frame architecture and wiki compiler (Hermes has neither).
  4. Apache 2.0 hive-mind as standalone npm package (@hive-mind/core etc.) — Hermes is monolithic.
  5. EU AI Act audit triggers built-in (Hermes does not address).
  6. Published peer-reviewed-style benchmark results (apples-to-apples Mem0 + ERL methodology + Gaia2 + τ³) — Hermes publishes only internal benchmarks.

Differentiator #6 is the moat. Hermes Agent has not engaged any standardized public benchmark venue. If we ship arxiv + Gaia2 + τ³ within Q2, the gap is unbridgeable for them in 2026.

3.2 OpenClaw security posture (no new evidence required)

OpenClaw March 2026 CVE cluster (9 CVEs in 4 days, including CVSS 9.9; Snyk flagged 1,467 malicious skills on ClawHub) is already in the existing intelligence doc per CC-1 audit. Confirming it remains in Day 0 narrative for regulated-industry pitches as "incumbent insecurity" framing.


Replace BENCHMARK-STRATEGY.txt §3.4 (Phase 3) and §3.5 (Phase 4) primary venues. All other sections remain intact.

Phase 04-18 strategy Proposed amendment Rationale
Phase 0 (now → launch) Stripe priority; nothing else Unchanged. PHF posture stable.
Phase 1 (post-Stripe, hive-mind alpha) LoCoMo + bootstrap Unchanged. Stage 3 v6 already complete. Status quo.
Phase 2 (hive-mind launch) LongMemEval + blog + GitHub public Unchanged. Coupled launch as PHF binds.
Phase 3 (Waggle benchmark integration) SWE-bench sequential + SWE-ContextBench REPLACE with Gaia2 Search + Execution splits, GEPA-evolved variants, ERL methodology framing. Target: 30-35 % pass@1 (open-source SOTA = 21 %). Fresher venue, better domain fit, ERL publication target.
Phase 4 (KVARK milestones) Proprietary BPMN-workflow + scale benchmarks REPLACE BPMN with τ³-bench banking_knowledge, retain scale + multi-tenant + compliance latency benchmarks. Target: top-3 open-source on banking_knowledge. Real venue, real comparison, regulated-industry sales artifact.

Stretch targets in 04-18 strategy (BEAM 1M-token, SWE-ContextBench Memory track) deferred to Q3 2026 review.


§5 — Sequencing (12-week horizon)

Weeks 0-2 (now → hive-mind alpha): PHF locked artifacts ship — arxiv preprint, hive-mind public, Waggle landing, Stripe. No new benchmark work in this window.

Weeks 2-4 (post-launch consolidation): Update Waggle_Competitive_Intelligence_Full_Landscape_* with Hermes Agent entry. Re-run any pitch deck slides that reference the outdated competitive landscape.

Weeks 4-8 (Phase 3 Gaia2 sprint):

  • Week 4: Set up ARE platform locally; verify GEPA-evolved qwen-thinking::gen1-v1 runs against Gaia2 Search split with no harness modification.
  • Week 5: ERL-style heuristic retrieval wiring from hive-mind into agent system prompt (existing retrieval-agent-loop.ts is the integration point — 38.3 KB file already does adjacent work).
  • Week 6: N=200 dry run on Search split; cost validation under $50.
  • Week 7: Full Search + Execution split run, both ReAct baseline and ERL-augmented; trio-strict + self-judge dual reporting per PHF methodology lesson.
  • Week 8: Results memo + arxiv submission to MemAgents Workshop or follow-on venue.

Weeks 8-12 (Phase 4 τ³ sprint, KVARK track):

  • Week 8: Set up tau2-bench locally with banking_knowledge extras (uv sync --extra knowledge).
  • Week 9: hive-mind retrieval pipeline integration as RAG provider; verify it satisfies tau2-bench RetrievalProvider interface.
  • Week 10: N=100 dry run; calibrate per-task cost and latency.
  • Week 11: Full N=200 run, frontier subject (Opus 4.7 + GPT-5.4) + Qwen subject for sovereignty story.
  • Week 12: Submit results to taubench.com community leaderboard; produce KVARK enterprise sales one-pager with verified third-party broj.

§6 — Risks and out-of-scope items

6.1 Risk: GEPA +12.5 pp uplift on N=13 may not generalize to Gaia2 task distribution

The held-out validation is statistically thin (N=5 held-out + N=8 in-sample). Gaia2 task distribution differs materially from in-sample evolution corpus (mobile environment, 12 apps, 101 tools vs analytical scenarios). Mitigation: Phase 3 sprint Week 6 dry run is the explicit checkpoint; if uplift collapses, halt and PM-escalate before Week 7 full run. Cost exposure if abort: ~$15.

6.2 Risk: ERL methodology reference point may shift before Workshop submission

The +7.8 % uplift is from the ERL paper as published. Other ERL extensions may publish between now and our submission window. Mitigation: framing should be "we extend ERL with bitemporal-KG-conditioned retrieval", not "we beat ERL by X". Defensible regardless of intermediate competitor work.

6.3 Risk: Hermes Agent or OpenClaw publish on Gaia2 / τ³ before us

Probability: low for OpenClaw (CVE remediation is consuming community bandwidth); medium for Hermes (Nous Research has paper-publishing track record). Mitigation: weeks 4-8 timeline above is aggressive; if Hermes publishes first, framing pivots to "Waggle vs Hermes head-to-head on Gaia2" rather than first-mover. Either way the published broj is the enterprise sales artifact.

6.4 Out of scope (explicitly)

  • Any change to PHF claim, Day 0 narrative, coupling decision, or pricing.
  • Frontier subject re-run of Stage 3 LoCoMo. (Earlier consideration deprecated by 04-25 self-judge re-eval evidence.)
  • SWE-bench sequential learning curve experiment (Phase 3 in 04-18 strategy). Deferred to Q3 review pending Phase 3 Gaia2 results.
  • New benchmark venue invention (StuLife, J-TTL, FieldWorkArena). Stick to community-recognized venues.

§7 — Ratification asks

PM ratification required on the following five items before any Phase 3 or Phase 4 sprint kickoff. None blocks Day 0 launch.

  1. Ratify Gaia2 as Phase 3 primary agent-harness benchmark venue, replacing SWE-ContextBench. (Y/N)
  2. Ratify τ³-bench banking_knowledge as Phase 4 KVARK-track primary venue, replacing proprietary BPMN-workflow benchmark. (Y/N)
  3. Ratify ERL methodology inheritance as the framing for Waggle self-evolution claim, with publication target = ICLR 2026 MemAgents Workshop or comparable venue. (Y/N)
  4. Ratify Hermes Agent competitive intelligence amendment (§3.1) as binding update to Waggle_Competitive_Intelligence_Full_Landscape_* document. PM authorizes Marketing-side rewrite or assigns to CC. (Y/N + assignee)
  5. Ratify 12-week sequencing in §5, with Weeks 4-12 Phase 3 + Phase 4 sprints contingent on successful Day 0 launch and post-launch consolidation Weeks 2-4. (Y/N)

After ratification, this brief becomes binding addendum to BENCHMARK-STRATEGY.txt; phase tables in §3.4 and §3.5 of that document are superseded by §4 of this brief. All other sections of 04-18 strategy remain primary.


§8 — Cross-references

  • 04-18 primary strategy: strategy/BENCHMARK-STRATEGY.txt (NOT superseded; supplemented).
  • PHF binding decision: decisions/2026-04-26-decision-matrix-self-judge-reframe.md.
  • Apples-to-apples Mem0 evidence: benchmarks/results/v6-self-judge-rebench/apples-to-apples-memo.md (in waggle-os repo).
  • GEPA Phase 5 substrate: gepa-phase-5/manifest.yaml + gepa-phase-5/preflight-evidence.md (in waggle-os repo).
  • arxiv paper anchor: research/2026-04-26-arxiv-paper/00-paper-outline.md.
  • Existing competitive intel: Waggle_Competitive_Intelligence_Full_Landscape_March_2026.docx (in waggle-os repo root).
  • Gaia2 paper: arxiv 2602.11964.
  • τ³-bench / τ-Knowledge paper: arxiv 2603.04370.
  • ERL paper: arxiv 2603.24639.
  • Sierra leaderboard: taubench.com.
  • ARE platform: github.com/facebookresearch/meta-agents-research-environments.

(2,847 words)