35 KiB
Waggle / hive-mind — The Memory System, Explained
Audience: A developer taking over Waggle OS who needs a working mental model of how the agent remembers things — fast enough to be productive, deep enough to not break invariants.
Status: Written 2026-06-26 against packages/hive-mind-core/src/{mind,harvest} (the
substrate), packages/agent/src (cognify/recall wiring), packages/server/src/local/routes
(REST), and packages/memory-mcp/src (MCP surface).
⚠️ Path warning up front. The older
docs/memory-architecture.md(April 2026) and one of the source-reader passes describe the substrate as living underpackages/core/src/mind/. That path no longer exists. The memory substrate was moved topackages/hive-mind-core/src/{mind,harvest}/in the 2026-04-30 monorepo migration. Every file reference in this document points at the real, current location. See §8 for the full list of stale claims in the old doc.
1. Mental model
Waggle's memory is one SQLite database per "mind" (better-sqlite3 + the sqlite-vec
extension, WAL mode, foreign keys on). A mind is either the user's personal mind
(always loaded) or a workspace mind (lazy-loaded when a workspace is active). Every memory
layer — identity, awareness, the frame event-log, the knowledge graph, embeddings, full-text
index, audit logs — lives as tables in that same file.
The unit of memory is a frame: an append-only, immutable record of something the agent learned or did. Frames are deduplicated by content hash, embedded into a 1024-dim vector index, mirrored into an FTS5 full-text index, and mined for typed knowledge-graph entities. At recall time, a 10-lane retrieval pipeline pulls the most relevant frames (vector + keyword + importance + date-window + raw-turn lanes), fuses and re-ranks them, and appends the result to the system prompt before the LLM ever sees the user's message. The whole thing runs twice per turn: read before (recall), write after (cognify).
WRITE PATH
┌──────────────┐ ┌──────────────┐ ┌────────────────────────────────────┐
│ live turn │ │ pattern │ │ cognify(): │
│ user+asst │──▶│ write-back │──▶│ • createIFrame/PFrame (dedup'd) │
│ exchange │ │ (30+ regex) │ │ • extractEntities → KG upsert │
└──────────────┘ └──────────────┘ │ • co-occurrence + semantic rels │
│ • search.indexFrame() → vectors │
┌──────────────┐ ┌──────────────┐ └──────────────────┬─────────────────┘
│ bulk import │ │ 4-pass │ │
│ (ChatGPT, │──▶│ harvest │──────────────────────┤
│ Claude,PDF…) │ │ pipeline │ writes I/P/B frames │
└──────────────┘ └──────────────┘ ▼
┌────────────────────────────────────┐
│ ONE SQLite DB (per mind) │
│ memory_frames ── content_hash idx │
│ memory_frames_fts (FTS5) │
│ memory_frames_vec (sqlite-vec) │
│ memory_frame_chunks(_vec) │
│ knowledge_entities / _relations │
│ identity / awareness / sessions │
└────────────────────────────────────┘
▲
READ PATH │
┌──────────────┐ ┌────────────────────────────────────┴───────────────┐
│ user query │ │ recallMemory(): 10-lane pipeline │
│ (turn start) │──▶│ importance · semantic(BM25+vec via RRF) · date- │
└──────────────┘ │ window · profiles · facts · events · raw-detail · │
│ catch-up · workspace/personal split │
└──────────────┬─────────────────────────────────────┘
▼
cross-encoder rerank → scoring (recency/importance/popularity/context)
▼
dedup by frame-id → injection scan (all-or-nothing) → formatted text
▼
APPENDED to system prompt: [recalled memories] + [system prompt] + [user query]
▼
LLM
2. The core data model
What a "frame" is
A MemoryFrame is one immutable row in memory_frames. It is the atom of memory. Three kinds:
| Frame type | Meaning | base_frame_id |
|---|---|---|
| I (Index) | Full snapshot / initial fact. One per GOP to start. | null |
| P (Patch) | Incremental append referencing an I-frame; applied in t order. String concatenation, not a real diff — P-frames can only add, never retract. |
the I-frame's id |
| B (Branch) | Alternative to a frame without replacing it; allows divergent histories. Content serialized as JSON {description, references[]}. |
the branched frame's id |
A GOP ("Group of Pictures", borrowed from video codecs) is a session/conversation; gop_id
is a FK to sessions.gop_id, and t is a monotonic logical timestamp within that GOP.
Key columns: id, frame_type, gop_id, t, base_frame_id, content, importance, source, access_count, created_at, last_accessed, content_hash, metadata.
- importance ∈
critical | important | normal | temporary | deprecated - source ∈
user_stated | tool_verified | agent_inferred | import | system— these are the only 5 the SQLCHECKpermits (schema.ts:57). ⚠️ The TypeScriptFrameSourcetype (frames.ts:25) over-declares 3 more (personal | workspace | team_sync) that the DB would reject at insert;server/src/local/index.ts:1333actually passes'team_sync'— a latent constraint-violation bug. Fix by widening the CHECK or correcting the call site. - content_hash =
SHA256(stripHmPrefix(content).trim())— the provenance-insensitive dedup key - metadata = JSON
TEXT(NOT NULL DEFAULT '{}'); never full-text indexed; carries the Phase-2B "Memory Center" contract (kind/status/scope/tags/confidence/trace_id/…)
How the layers relate
| Layer | Table(s) | Shape | Role |
|---|---|---|---|
| Frames | memory_frames |
append-only I/P/B log | The substrate of truth — everything learned |
| Vectors | memory_frames_vec, memory_frame_chunks_vec |
vec0 virtual tables, 1024-dim |
Semantic search; rowid = frame id (or chunk id) |
| Full-text | memory_frames_fts |
FTS5 virtual table | Keyword (BM25) search; must stay in sync with frames |
| Chunks | memory_frame_chunks (+ _vec) |
per-frame paragraph chunks, ON DELETE CASCADE |
Chunk-level retrieval for long frames |
| Knowledge graph | knowledge_entities, knowledge_relations |
typed nodes + directional typed edges, bitemporal | Structured distillation of frames (person/project/file/…) |
| Identity | identity |
single row, CHECK(id=1) |
Who the agent is — injected into prompt, never a frame |
| Awareness | awareness |
≤10 rolling items, expiry-aware | What it's working on right now — injected, never a frame |
| Sessions | sessions |
GOP → project container | Groups frames; memory_frames.gop_id → sessions.gop_id |
| Audit (append-only) | install_audit, ai_interactions |
DDL-triggered immutable | EU AI Act compliance; DB refuses DELETE/UPDATE |
| Concept mastery | concept_mastery |
spaced-repetition rows | Orthogonal learning tracker — not linked to frames or KG |
Identity and awareness are out-of-band layers: they format themselves into markdown via
toContext() for prompt inclusion, but are never serialized as MemoryFrame records. Only
memory content (facts, events, profiles, decisions) becomes frames.
3. The WRITE path
There are two entry points that produce frames: live cognify (during a turn) and bulk
harvest (importing external corpora). Both converge on FrameStore.createIFrame/createPFrame
and the same content-hash dedup.
(a) Live cognify — during an agent turn
Fires after the agent loop succeeds (never on failure, so error traces don't become "ground truth"). Entry point in the chat route:
packages/server/src/local/routes/chat.ts:1456
const saved = await sessionOrch.autoSaveFromExchange(message, result.content, { traceId });
- Pattern write-back —
autoSaveFromExchange()(orchestrator.ts:839) delegates torunPatternWriteBack()(pattern-write-back.ts), which runs 30+ calibrated regexes over the (user, assistant) exchange:- Preferences (
"I prefer","call me","from now on"), corrections ("actually no","that's wrong"), decisions ("let's go with","decided to"), findings ("discovered that","turns out"), style signals ("keep it brief","bullets"). - Casual chatter (
"lunch","weather","thanks") is intentionally skipped. - Routing: preferences/corrections/style → personal mind; decisions/findings/ work-output → workspace mind (or personal if no workspace active).
- Preferences (
- Cognify pipeline — each surviving snippet runs
cognify()(cognify.ts:52):- Frame creation: I-frame if none exists for the GOP, else P-frame.
extractEntities(content)→ persons/orgs/products/concepts.upsertEntities()→ reuse existing entity id on case-insensitive type+name match, else create.- Co-occurrence relations: all-pairs
co_occurs_with(strength 0.8) among entities in the same text. - Semantic relations:
extractRelations()findsled_by/reports_to/depends_on/… and upserts edges. search.indexFrame(frameId, content)→ embeds + writes the vector row.- Optional
MemoryLinker.findRelated()→ related-frame links.
- Signal commit —
commitSurfacedSignals()(chat.ts:1447) marks M8 awareness signals "surfaced" so they don't re-show next turn.
(b) Bulk harvest — importing external corpora
HarvestPipeline.run() (harvest/pipeline.ts:100) converts ChatGPT/Claude/Gemini exports,
Markdown, PDFs, and URLs into frames + KG entities via a 4-pass distillation:
External source ─▶ Adapter ─▶ UniversalImportItem ─▶ Pass0..4 ─▶ DistilledKnowledge ─▶ Frames+KG
- Stage 1 — Adapter parsing. Each source has an adapter (
claude-adapter.ts,markdown-adapter.ts,pdf-adapter.ts, …) normalizing wildly different formats into aUniversalImportItem(id, source, type, title, content, messages[]). Adapters useraw-types.tshelpers (asRecord,getString) to safely narrow untrusted JSON. - Stage 2 — Injection scan (Pass 0). Every item is scanned (
pipeline.ts:111) withscanForInjection(probe, 'tool_output')over title + first 4KB before any LLM touches it. Hostile items are dropped entirely. - Stage 3 — 4-pass distillation:
- Pass 1 Classify (Haiku, cheap, batches of 20): assigns
value(skip/low/medium/high)- domain.
value='skip'items (greetings, loops) never reach Pass 2.
- domain.
- Pass 2 Extract (Sonnet): decisions/preferences/facts/knowledge/entities/relations.
- Pass 3 Synthesize (Sonnet): maps to
DistilledKnowledgewithtargetLayer(identity/frame/kg_entity/kg_relation) + importance + confidence + provenance. - Pass 4 Dedup (local, no LLM —
dedup.ts:70): see below.
- Pass 1 Classify (Haiku, cheap, batches of 20): assigns
- Stage 4 — Frame writing:
writeRawTurnFrames()(raw-turns.ts:112): each user/assistant message stored as[mind-rawturn conv:KEY turn:N speaker:S]frame for adjacency lookups (system messages skipped).writeMemoryLaneFrames()(extract-memory-lanes.ts:288): three parallel lanes —[mind-fact](importance normal),[mind-event] [YYYY-MM-DD](date resolved from a relative cue against the session date),[mind-profile NAME](replace-on-update — old profile for that speaker deleted first, unlike facts which accumulate).writeKgEntities()(extract-kg-entities.ts:230): exact-name dedup, bumpseen_counton hit.- Every extraction output is injection-scanned again before write (verbatim dialogue is the most injection-prone surface).
Deduplication (shared by both paths)
- Frame-level (
FrameStore.findDuplicate,frames.ts:264): O(1) indexed lookup —SELECT … WHERE content_hash = SHA256(stripHmPrefix(content).trim()). If found,touch()the existing frame (bump access count) and skip insertion.stripHmPrefixremoves the[hm session:… src:…]metadata prefix so the same turn captured from two different sources collapses into one frame. This hash is shared by insert, dedup lookup, update, compaction, and backfill so the semantics never drift. - Harvest-level (
dedup.ts:70): two-tier — (1) exact SHA256 on normalized content, then (2) fuzzy trigram similarity (threshold 0.75) against existing content. Contradictions (similarity 0.4–0.75 + importanceimportant) are logged but still written — the caller decides policy.
4. The READ path
Orchestrator.recallMemory(query, limit, opts) (orchestrator.ts:453–831) runs at the start
of every turn (chat.ts:767) and returns { text, count, recalled, recalledFrames }. It is
not simple keyword search — it's a 10-lane pipeline:
- Catch-up detection. Regex on the query (
"catch me up","where were we","what did we decide","brief me", …) → switches to an importance-based fetch instead of semantic search (catch-up should return important things, not semantically-close small talk). This lane skips the reranker. - Importance lane (K=5). Critical/Important frames surface on every query.
- Semantic search.
HybridSearch.search()runs FTS5 (BM25, stop-words stripped, OR-MATCH) andvec0vector distance in parallel, fused via Reciprocal Rank Fusion (RRF, K=60). - Date-window lane. If the query names a period (
"in May 2026"),parseDateWindow(query)(parse-date-window.ts) producessince/untilSQL filters. - Profile lane —
[mind-profile]frames. - Facts lane —
[mind-fact]frames, capped 60, oldest-first. - Events lane —
[mind-event]frames, chronological, capped 40; a dedicated "Events during X" section for windowed queries. - Raw-detail lane — verbatim
[mind-rawturn]excerpts, reranked via the cross-encoder (raw-detail-lane.ts,inprocess-reranker.ts). - Workspace / personal split — active mind renders first (visual precedence;
orchestrator.ts:692workspace,:700personal). - Injection scan (all-or-nothing) —
scanForInjection(joinedLines, 'tool_output')(orchestrator.ts:777). On a poisoned hit the entire recall returns empty. No partial recall.
Fusion → rerank → scoring. After RRF, results are re-ranked. The optional cross-encoder
reranker (inprocess-reranker.ts) soft-fails back to RRF ordering if unavailable. Final
ranking applies computeRelevance() (scoring.ts), four signals weighted by a profile
(balanced | recent | important | connected):
| Signal | How it scores |
|---|---|
| temporal | 7-day full strength, ~30-day half-life decay |
| popularity | log10 of access_count (1000× difference ≈ 0.3 score delta) |
| contextual | KG BFS distance 0/1/2/3 → 1.0/0.7/0.4/0.2 — currently 0 in production (see gotchas) |
| importance | critical 2.0 / important 1.5 / normal 1.0 / temporary 0.7 / deprecated 0.3 |
finalScore = rrfScore × relevanceScore.
Recall-context assembly & prompt injection. The lanes are merged, deduped by frame id,
the TEMPORAL_GUIDANCE constant is bundled in (orchestrator.ts:808), and the block is returned
as text. The chat route then does:
chat.ts:776 recalledContext = '\n\n' + recall.text;
chat.ts:905 systemPrompt = … orch.buildSystemPrompt() … (assembled ? '' : recalledContext);
producing the sandwich [recalled memories] + [system prompt] + [user query]. Crucially,
recall is NOT part of buildSystemPrompt() — it's appended separately. The exception:
if PromptAssembler is enabled, the assembler embeds recall internally (chat.ts:887) and the
route skips the manual append.
buildSystemPrompt() itself (orchestrator.ts:301) assembles three other sections: identity
(cached by content hash), self-awareness (improvement signals, uncached), and preloaded recent
context (uncached).
5. Layers reference
5.1 Storage & schema
- Purpose: the SQLite substrate — frames, embeddings, FTS, KG, audit, dedup, migrations.
- Files:
mind/db.ts,mind/schema.ts,mind/frames.ts,mind/content-hash.ts,mind/chunker.ts - Key functions:
MindDBctor (loadssqlite-vec, WAL, FK,initSchema),MindDB.initSchema(idempotent: SCHEMA_SQL + VEC_TABLE_SQL on fresh DB, elserunMigrations),MindDB.runMigrations(crash-recovery + guardedADD COLUMN+ content-hash backfill + append-only triggers),FrameStore.createIFrame/createPFrame/createBFrame,FrameStore.findDuplicate,FrameStore.update/delete/touch,hashFrameContent/stripHmPrefix,chunkText. - Chunking:
chunkText()(chunker.ts:58) splits on blank lines, greedily aggregates paragraphs to ≤2000 chars, sub-splits oversized paragraphs on sentence boundaries, applies 200-char overlap, and preserves absolutecharStart/charEndrelative to the parent frame.
5.2 Retrieval (hybrid search + scoring)
- Purpose: fuse keyword + vector, re-rank by profile.
- Files:
mind/search.ts,mind/scoring.ts,mind/inprocess-reranker.ts,mind/parse-date-window.ts,mind/raw-detail-lane.ts,mind/recall-context.ts,mind/resolve-relative-date.ts - Key functions:
HybridSearch.search(FTS5 + vec0 → RRF K=60 →computeRelevance),HybridSearch.indexFrame(must be called on every frame insert),computeRelevance,parseDateWindow.
5.3 Knowledge graph & semantic layers
- Purpose: typed entities + temporal relations distilled from frames; canonicalization.
- Files:
mind/knowledge.ts,mind/ontology.ts,mind/entity-normalizer.ts,mind/concept-tracker.ts,harvest/extract-kg-entities.ts - Key functions:
KnowledgeGraph.findEntityByName(exact, case-sensitive — the dedup primitive),KnowledgeGraph.searchEntities(fuzzy LIKE — not for dedup),dedupeByName(merge bynormalizeEntityName(name)::type, survivor = most-relations / lowest-id, repoint edges, retire dups, sumseen_count),traverse(BFS one edge type),bfsDistances,retireEntity(soft-delete viavalid_to=now()),normalizeEntityName(js→javascript,postgres→postgresql,k8s→kubernetes),isNoiseName,ConceptTracker.recordAnswer/getDueForReview. - Bitemporal: both entities and relations carry
valid_from/valid_to;valid_to IS NULL= active. Nothing is hard-deleted; rows are retired for audit history.
5.4 Identity / Awareness / Sessions
- Purpose: persistent agent identity, volatile working memory, and conversation containers — all feeding the prompt without being frames.
- Files:
mind/identity.ts,mind/awareness.ts,mind/sessions.ts,agent/src/orchestrator.ts,agent/src/context-loader.ts - Key functions:
IdentityLayer.update(column allowlist for SQLi defense) /toContext,AwarenessLayer.add/getAll(≤10, priority DESC, expiry via SQL WHERE) /updateMetadata(shallow JSON merge) /toContext,SessionStore.ensureActive(transaction +id DESCtiebreaker for same-second races) /ensure(idempotent for harvest),loadRecentContext. - Cross-workspace rule: identity is always from the personal mind; awareness is merged (personal first); frames/KG switch to the workspace mind when one is active.
5.5 Embeddings
- Purpose: multi-tier provider fallback with tier-gating, quota, and an embedder-lock. See §6.
- Files:
mind/embedding-provider.ts,mind/embeddings.ts,mind/inprocess-embedder.ts,mind/ollama-embedder.ts,mind/api-embedder.ts,mind/litellm-embedder.ts - Key functions:
createEmbeddingProvider,probeProvider,createInProcessEmbedder,ensureEmbeddingFingerprint(the embedder-lock),recreateVecTables(destructive),normalizeDimensions,maxEmbedCharsForModel,reembedPerText.
5.6 Harvest / ingestion
- Purpose: turn external conversations/docs into frames + KG via 4-pass distillation. See §3(b).
- Files:
harvest/pipeline.ts,harvest/types.ts,harvest/dedup.ts,harvest/raw-turns.ts,harvest/extract-memory-lanes.ts,harvest/extract-kg-entities.ts, plus per-source adapters (claude-adapter.ts,chatgpt-adapter.ts,gemini-adapter.ts,markdown-adapter.ts,pdf-adapter.ts,url-adapter.ts,plaintext-adapter.ts,perplexity-adapter.ts,claude-code-adapter.ts,universal-adapter.ts). - Key functions:
HarvestPipeline.run,dedup,writeRawTurnFrames,extractMemoryLanes,writeMemoryLaneFrames,extractKgEntities,writeKgEntities.
5.7 Cognify (live write wiring)
- Purpose: passively learn from each turn; the live read+write loop inside the agent.
- Files:
agent/src/cognify.ts,agent/src/orchestrator.ts,agent/src/pattern-write-back.ts,agent/src/entity-extractor.ts,server/src/local/routes/chat.ts - Key functions:
recallMemory,buildSystemPrompt,cognify,autoSaveFromExchange,commitSurfacedSignals,upsertEntities,createCoOccurrenceRelations,createSemanticRelations.
5.8 API / MCP surface
- Purpose: dual external surface — REST for the UI, MCP tools for Claude agents/clients.
- REST files:
server/src/local/routes/memory.ts(/api/memory/search,/api/memory/frames,/api/memory/stats),memory-center.ts(Phase-2B/api/memoryCRUD + merge + archive + trace),identity.ts(/api/identity),knowledge.ts(/api/memory/graph). - MCP files:
memory-mcp/src/tools/{memory,identity,awareness,knowledge,wiki,harvest,workspace, cleanup,ingest}.ts,memory-mcp/src/resources/memory.ts,memory-mcp/src/core/setup.ts. - Key functions:
normalizeToMemory(projects a frame + its JSON metadata blob into the sharedMemoryentity),save_memory/recall_memory(MCP),getWorkspaceMind(lazy per-workspaceMindDB),POST /api/memory/:id/merge(C11: concatenate + archive originals, never hard-delete).
6. Embedding provider chain & the embedder-lock rule
Why it matters operationally: if the embedding model changes, the vector index silently becomes meaningless or the DB refuses to boot. The system guards this aggressively.
Resolution chain (embedding-provider.ts, auto-probe in strict order, halts at first working):
inprocess ─▶ ollama ─▶ voyage ─▶ openai ─▶ litellm ─▶ mock
| Provider | What it is | Notes |
|---|---|---|
| inprocess | Xenova/all-MiniLM-L6-v2 via @huggingface/transformers |
384 native dims → normalized to 1024; ~23 MB, cached in ~/.waggle/models/, offline |
| ollama | POST localhost:11434/api/embed, nomic-embed-text |
30 s timeout |
| voyage | POST api.voyageai.com, voyage-3-lite |
needs API key; 15 s timeout |
| openai | POST api.openai.com, text-embedding-3-small |
needs API key; 15 s timeout |
| litellm | proxy /v1/embeddings |
optional Bearer; no explicit timeout |
| mock | deterministic byte-hash → Float32Array | zero semantic value, last resort |
- Tier enforcement:
TIER_CAPABILITIES[tier].embeddingProvidersgates which providers a user may reach; non-allowed providers are skipped during auto-probe. Only active whenuserTieris passed;WAGGLE_EVAL_MODE=1disables tier gates entirely (eval-harness measurement validity). - Quota: monthly count in
embedding_usage(user_id,year_month,count);checkQuota()throwsEmbeddingQuotaExceededError; resets at UTC month boundary. - Char capping:
maxEmbedCharsForModel()caps inputs (6K default, 8K for*-8kmodels) by model-NAME heuristic — not measured context. (D1 probe finding:nomic-berthas a hard 2048-token limit despite its name.) - Batch resilience: on a batch embed failure,
reembedPerText()retries one input at a time and degrades only the failing inputs to mock — the rest keep real vectors.
The embedder-lock (fingerprint guard)
ensureEmbeddingFingerprint() (db.ts) runs on the first vector op and records
{provider, model, dim} in the meta table. On subsequent ops:
| Condition | Result |
|---|---|
| same dim, same model | OK (match) |
| same dim, different model | update meta + warn (model-changed). Vectors stay numerically valid; cross-model semantic similarity is degraded, not undefined. |
| different dim | throws EmbeddingDimMismatchError — vec0 virtual tables cannot be ALTERed. |
Remediation for a dimension change is destructive: MindDB.recreateVecTables(newDim) DROPs
both vec tables and reinitializes, after which every frame must be re-embedded. Treat
1024 as effectively load-bearing.
7. Where it lives & how to call it
File map
packages/hive-mind-core/src/
├── mind/ ← the substrate (per-mind SQLite + layers)
│ ├── db.ts schema.ts MindDB, migrations, append-only triggers
│ ├── frames.ts FrameStore (I/P/B CRUD, dedup, compaction)
│ ├── content-hash.ts hashFrameContent / stripHmPrefix
│ ├── chunker.ts semantic paragraph chunking
│ ├── search.ts scoring.ts HybridSearch (RRF) + computeRelevance
│ ├── inprocess-reranker.ts cross-encoder rerank (soft-fail)
│ ├── parse-date-window.ts resolve-relative-date.ts temporal lanes
│ ├── raw-detail-lane.ts recall-context.ts recall assembly
│ ├── knowledge.ts ontology.ts entity-normalizer.ts KG + canonicalization
│ ├── concept-tracker.ts spaced-repetition (orthogonal)
│ ├── identity.ts awareness.ts sessions.ts out-of-band layers
│ ├── embedding-provider.ts embeddings.ts *-embedder.ts provider chain
│ ├── reconcile.ts multi-source reconciliation
│ └── evolution-runs.ts execution-traces.ts improvement-signals.ts (subsystems)
└── harvest/ ← ingestion (adapters + 4-pass pipeline)
├── pipeline.ts dedup.ts types.ts raw-types.ts
├── raw-turns.ts extract-memory-lanes.ts extract-kg-entities.ts
└── *-adapter.ts claude / chatgpt / gemini / markdown / pdf / url / …
packages/agent/src/ ← live wiring into the turn
orchestrator.ts (recallMemory, buildSystemPrompt, autoSaveFromExchange)
cognify.ts pattern-write-back.ts entity-extractor.ts context-loader.ts
packages/server/src/local/routes/ ← REST surface for the UI
memory.ts memory-center.ts identity.ts knowledge.ts
packages/memory-mcp/src/ ← MCP surface for Claude agents
tools/{memory,identity,awareness,knowledge,wiki,harvest,workspace,cleanup,ingest}.ts
resources/memory.ts core/setup.ts
How to call it
- From an agent (MCP):
save_memory(content, importance?, source?, workspace?)→ I-frame + index.recall_memory(query, limit?, workspace?, scope?, profile?)→ hybrid search (scope∈ personal/current/all). Alsoget_identity/set_identity,get_awareness/set_awareness,search_entities/save_entity,harvest_import,compile_wiki. - From the UI (REST):
GET /api/memory?status=active,GET /api/memory/search?q=…,POST /api/memory/frames,GET /api/memory/graph?scope=all|personal|workspace,GET/POST /api/identity. Workspace selectable via?workspace=id; mutations require an explicit?mind=personal|workspace(typos fail loudly with 400 — never silently widened). - Programmatically: open a
MindDB(dbPath), then constructFrameStore,HybridSearch,KnowledgeGraphagainstdb.raw. Lazy workspace minds come fromgetWorkspaceMind(id)(singleton per workspace viaMultiMindCache).
8. Gotchas & footguns
Cross-cutting invariants (break these and memory silently rots)
- Every
memory_framesINSERT must callHybridSearch.indexFrame. Writing the row directly bypasses FTS + vec; the frame becomes invisible to search and there is no cheap rebuild script. - Vectors are locked at 1024-dim.
vec0tables can't beALTERed. A dimension change means DROP + recreate + re-embed everything (recreateVecTables). The guard throws early at search time, but the fix is destructive. - Mock embedder = zero semantics. If the provider chain falls through to
mock, search still "works" but ranks randomly. Detect viagetStatus().activeProvider === 'mock'. - Model change at same dim is allowed but degrades cross-model similarity (a warning, not an error). Don't swap embedding models casually on a populated DB.
- Injection scan is all-or-nothing on recall. A single poisoned hit zeroes the entire recall block. Harvest drops hostile items entirely — no soft sanitization anywhere.
install_auditandai_interactionsare physically append-only via DDLBEFORE DELETE/UPDATEtriggers thatRAISE(ABORT)— the DB refuses mutation (EU AI Act Art. 12). Substrate changes confined toinstall_audithave nowhere to land on the OSS mirror — don't treat them as a pending port (see CLAUDE.md §7.5).
Frame/dedup subtleties
stripHmPrefixis load-bearing for dedup (OQ-6). Same-body captures from different sources must collapse regardless of the[hm …]prefix; using only.trim()regresses this. The hash is shared across insert/lookup/update/compaction/backfill so semantics never drift.- P-frames append, never retract. Patch application is string concatenation, not a diff. To
correct/retract,
update()the I-frame or mark itdeprecated. - content_hash INDEX is created in
runMigrations()(after a guardedADD COLUMN), not in SCHEMA_SQL. Creating it in SCHEMA_SQL crashes boot on pre-D3 DBs (2026-06-12 regression). setMetadata()does not invalidate FTS or vec — metadata is never indexed.createIFrameaccepts an optional ISO-8601createdAt(T separator + timezone). Invalid values fall back todatetime('now'); the harvest path validates + logs so exported timestamps preserve original ordering.
Knowledge-graph subtleties
- Use
findEntityByName()(exact) for dedup, neversearchEntities()(fuzzy LIKE). Fuzzy search drops the exact match from top-K once similar names accumulate — this caused 3506 duplicate "Phase" rows in the OSS repo before the fix. dedupeByNamekeys on name and type. "Marko" (person) and "Marko" (project) never merge.- Bitemporal soft-delete needs periodic cleanup.
valid_to IS NULL= active; nothing is hard- deleted, sodedupeByName/ retiring must run on-demand or rows accumulate. - The "connected"/contextual score is 0 in production (KNOWN GAP).
bfsDistances()computes entity→entity distances but there is no entity-id → frame-id bridge, so the contextual dimension contributes nothing today (scoring.ts §12). Akg_entity_framescross-ref table is hinted at inframes.delete()'s try-catch but is not in the schema. - Entity upsert is case-insensitive but type-specific.
John/johnmerge;john(PERSON) andjohn(PRODUCT) stay separate. isNoiseName()runs twice (parser + write seam) — don't assume one pass is enough.
Recall / multi-mind subtleties
- Recall is appended outside
buildSystemPrompt()(unless PromptAssembler is on). If you're hunting "why isn't memory in the prompt", look atchat.ts:767/776, not the orchestrator's prompt builder. - Catch-up mode skips the reranker and uses importance-fetch instead of semantic search.
- Workspace recall renders before personal (visual precedence).
- Frame IDs collide across personal + workspace minds (separate autoincrements). Multi-mind
reads tag results with
_mind/_workspace_name; mutations require an explicitmindparameter. - Identity cache key must hash the full JSON, not
updated_at— SQLite datetime has 1-second precision, so rapid edits within a second collide (orchestrator.ts:303). /api/memory/graph?scope=alloffsets IDs by +100k per workspace to avoid collisions in the merged viz; recomputed every request (expensive for many workspaces).- Awareness is hard-capped at 10 (
LIMIT 10in every read). Add an 11th and the oldest silently drops on next read; expiry is via SQL WHERE, not a background job.
Stale claims in the existing docs/memory-architecture.md (April 2026)
That doc predates the 2026-04-30 migration and should be read with these corrections:
- Wrong package path. It says the substrate lives in
packages/core/src/mind/. It does not — it lives inpackages/hive-mind-core/src/mind/. (packages/core/src/mind/does not exist on disk as of 2026-06-26.) Its own "Path discrepancy fixed" note is itself now stale. - "Schema version 1" / "five layers" undercount. The current schema has 10+ table groups (frames, FTS, vec, chunks+chunk-vec, KG, identity, awareness, sessions, improvement signals, install audit, procedures, AI interactions, execution traces, evolution runs, harvest sources, concept mastery). The "five memory layers" framing is a useful teaching simplification, not the physical schema.
- Reranker omitted. The doc describes RRF + scoring but predates the cross-encoder reranker
(
inprocess-reranker.ts) that the raw-detail and semantic lanes now use (soft-fails to RRF). - Embedding chain understated. It lists
inprocess → ollama → voyage → openai → mock; the current chain also includes a litellm stage before mock, plus tier-gating and quota.
Otherwise the old doc's descriptions of RRF fusion, scoring signals, dedup, and the dual-mind model remain conceptually accurate — only the paths and the layer/provider counts have drifted.