Benchmarks · view
Capabilities — where Waggle sits vs the field
How we're different

They automate tasks. Waggle remembers you.

The strong agents today are terminal coding tools for engineers. Waggle plays a different game: a persistent, local-first memory layer for knowledge workers that even runs those agents inside it. Here's the honest lay of the land.

The field

Task agents

Powerful, mostly terminal-based, model-locked, and built for developers. Each session starts fresh; the intelligence lives in the model, not in a memory of you.

Claude Code · Codex · Claude Cowork · Hermes · Odysseus
Our category

A memory layer + workspace

Knows you and your work across every session, runs on any model (even local), is built for non-technical experts — and can launch the task agents into its shared memory.

Waggle — complements them, doesn't compete head-on
Capability Waggle Claude CodeCLI CodexCLI CoworkAnthropic Hermesagent Odysseusagent
Yes Partial / limited No / not the focus

This is a positioning view, not a lab benchmark — capabilities as understood June 2026, focused on the dimensions Waggle is built around. Coding agents are genuinely excellent at deep code work (the last row), which is exactly why Waggle launches them rather than replacing them. The one head-to-head lab result is the memory benchmark → see "Memory SOTA".

State of the art · LoCoMo · June 2026

The best long-term memory on record.

On LoCoMo — the standard test for long-term conversational memory — Waggle's open-source substrate scores 87.66%, a new state of the art, measured under the prior leader's own protocol and judge.

Waggle · Hive Mindours · local
87.66
Memoriprev. SOTA
81.95
LangMemcorrected
78.05
Mem0baseline
62.47
+5.71
points over the prior best · z = 4.42, p < 10⁻⁵
92.75%
single-hop recall — ~1pt off the full-context ceiling
100%
local — warm recalls in 58–83 ms, on-device

LoCoMo, N = 1,540, GPT-4.1-mini as answerer & judge — the prior SOTA's exact published protocol, reproduced in-harness to 0.03 points before comparison. The intelligence lives in the memory layer, not the model, so it travels onto a small local model too. Honest caveat: higher token use per question than the leanest systems; efficiency work underway. Reproducible offline: github.com/marolinik/hive-mind.