Files
waggle-os/docs/product-analysis/architecture-analysis.md
Oleg Maslov 0c3e2ead3b
Some checks failed
Installer Smoke / installer-smoke (push) Has been cancelled
moving
2026-09-02 10:10:29 +02:00

588 lines
30 KiB
Markdown

# Waggle OS -- Technical Architecture Analysis
**Date:** 2026-04-08
**Scope:** Full codebase analysis of packages/core, packages/agent, packages/waggle-dance, packages/shared
**Purpose:** Identify crown jewels, technical moat, differentiation, risks
---
## Architecture Diagram
```
+----------------------------------------------------------------------+
| DESKTOP SHELL |
| Tauri 2.0 (Rust) React 18 + Vite + Tailwind + shadcn/ui |
| IPC allowlist apps/web/ (desktop OS + dock UI) |
+-------------------------------+--------------------------------------+
|
Tauri IPC
|
+-------------------------------v--------------------------------------+
| FASTIFY SIDECAR (Node.js) |
| packages/server -- REST API -- Clerk JWT auth |
+---+------------+------------+------------------+---------------------+
| | | |
v v v v
+--------+ +----------+ +-----------+ +-------------+
| Agent | | Core | | Shared | | Waggle |
| Engine | | Mind | | Types + | | Dance |
| | | | | Tiers | | Protocol |
+---+----+ +----+-----+ +-----------+ +------+------+
| | |
| v |
| +------+-------+ |
| | SQLite + FTS5 | |
| | + sqlite-vec | |
| | (.mind files) | |
| +--------------+ |
| |
v v
+---+--------------------------------------------+---+
| LiteLLM Proxy (model-agnostic) |
| Anthropic | OpenAI | Ollama | LiteLLM gateway |
+------------------------------------------------+---+
|
+-----------v-----------+
| KVARK (Enterprise)|
| Sovereign AI platform|
| SharePoint/Jira/Slack|
+-----------------------+
```
### Data Flow
```
User Message
|
v
buildSystemPrompt() --> Identity + Self-Awareness + Preloaded Context
|
+-- recallMemory(query) --> HybridSearch (FTS5 keyword + sqlite-vec vector)
| |
| +--> RRF fusion + scoring (temporal/importance/graph)
|
v
runAgentLoop()
|-- LiteLLM /chat/completions (streaming SSE)
|-- Tool execution loop (max 10 turns)
| |-- scanForInjection() on every tool output
| |-- HookRegistry pre:/post: events
| |-- LoopGuard (duplicate call detection)
| |-- Governance policy enforcement
|
v
autoSaveFromExchange()
|-- Pattern matching: preferences, decisions, corrections, research
|-- CognifyPipeline: frame save + entity extraction + KG enrichment + vector index
|
v
Response to user (with recalled[] for UI badges)
```
---
## 1. Memory Architecture -- THE Crown Jewel
### 1.1 The MindDB
Each workspace and user gets a separate `.mind` SQLite database. The schema has 7 layers:
| Layer | Table(s) | Purpose |
|-------|----------|---------|
| 0 | `identity` | Single-row user profile (name, role, personality, system_prompt) |
| 1 | `awareness` | Active tasks, pending items, context flags (max 10, with expiry) |
| 2 | `memory_frames` + `memory_frames_fts` + `memory_frames_vec` | I/P/B frame memory with FTS5 keyword search and 1024-dim vector search |
| 3 | `knowledge_entities` + `knowledge_relations` | Entity-relation knowledge graph with temporal validity |
| 4 | `procedures` | GEPA-optimized prompt templates with success rate tracking |
| 5 | `improvement_signals` | Recurring behavioral patterns (capability gaps, corrections, workflow patterns) |
| 6 | `install_audit` | Capability install trust trail (proposed/approved/installed/rejected) |
### 1.2 Frame Architecture (I/P/B Model)
Memory uses a **video compression-inspired** frame model:
- **I-Frame (Intra):** Complete snapshot -- the baseline state of a memory topic. Self-contained.
- **P-Frame (Predictive):** Delta from an I-frame -- captures changes, updates, corrections. References a `base_frame_id`.
- **B-Frame (Bidirectional):** Cross-reference frame linking multiple other frames together. Stores `references[]` as JSON.
Frames are organized by **GOP (Group of Pictures)** mapped through sessions. Each session has a `gop_id`, and frames within it have a monotonically increasing `t` value. State reconstruction: take the latest I-frame for a GOP and apply all P-frames since.
**Importance levels** with multipliers: critical (2.0x), important (1.5x), normal (1.0x), temporary (0.7x), deprecated (0.3x).
**Source provenance:** Every frame tracks how it was created: `user_stated`, `tool_verified`, `agent_inferred`, `import`, `system`, `personal`, `workspace`.
**Deduplication:** SHA-256 content hashing on the last 500 frames prevents duplicate I-frames. Duplicates update access count instead.
### 1.3 Dual-Mind Architecture
The Orchestrator maintains **two simultaneous memory stores**:
- **Personal Mind:** User preferences, communication style, corrections. Persists across ALL workspaces.
- **Workspace Mind:** Project context, decisions, task progress, domain knowledge. Scoped to one workspace.
The `setWorkspaceMind()` method activates a workspace mind alongside the personal mind. Both are queried in parallel during `recallMemory()`. Personal preferences (detected by content prefix patterns like "User preference:", "Style note:", "Correction from user:") are always loaded regardless of active workspace.
**Memory routing rules in autoSaveFromExchange():**
- Preferences, corrections, style notes --> personal mind
- Decisions, research, work output --> workspace mind
### 1.4 HybridSearch -- Retrieval Engine
Search combines three signals using **Reciprocal Rank Fusion (RRF)**:
1. **FTS5 keyword search:** Porter stemming + unicode61 tokenizer, with OR-based matching for better recall. Stop word filtering. Falls back to LIKE on FTS5 parse errors.
2. **sqlite-vec vector search:** 1024-dimensional embeddings via pluggable EmbeddingProvider. Supports InProcess (Xenova/MiniLM), Ollama (nomic-embed-text), Voyage, OpenAI, LiteLLM, with deterministic mock fallback.
3. **Relevance scoring** with 4 configurable profiles:
- Temporal: exponential decay with 30-day half-life, 7-day recency boost
- Popularity: logarithmic access count scaling
- Contextual: knowledge graph BFS distance (0/1/2/3 hops = 1.0/0.7/0.4/0.2)
- Importance: multiplied by the frame's importance level
Four scoring profiles weight these differently: `balanced`, `recent`, `important`, `connected`.
### 1.5 CognifyPipeline -- Memory Extraction
The `cognify()` method is the write-path pipeline:
1. Ensure a session exists (create one if needed)
2. Save frame (I-frame if first in GOP, P-frame otherwise)
3. Extract entities from content via regex-based NER (persons, organizations, technologies, projects, concepts, tools)
4. Upsert entities into KnowledgeGraph
5. Create co-occurrence relations between entities in the same text
6. Extract semantic relations (led_by, reports_to, depends_on, maintained_by, affiliated_with, approved) via pattern matching
7. Index the frame for vector search
8. Optionally find related frames via MemoryLinker
### 1.6 autoSaveFromExchange -- Passive Memory Accumulation
After every user/assistant exchange, the Orchestrator scans for save-worthy signals:
- **Preferences:** 10 regex patterns ("I prefer...", "call me...", "from now on...", etc.)
- **Implicit style detection:** 6 behavioral patterns (bullet preference, concise preference, code-first, etc.)
- **Decisions:** 7 patterns ("let's go with...", "decided to...", "the plan is...", etc.)
- **Corrections:** User disagreements saved as important frames
- **Research findings:** Structured output with URLs saved with source attribution
- **Structured extraction (F29):** Inline decisions, user questions, key bullet points, work output summaries
This is the mechanism by which Waggle "learns" without explicit save commands.
### 1.7 What Makes This Different from ChatGPT Memory / Claude Projects
| Capability | Waggle OS | ChatGPT Memory | Claude Projects |
|------------|-----------|----------------|-----------------|
| Storage | Local SQLite per workspace | Cloud, opaque | Cloud, project-scoped files |
| Persistence | Permanent until deprecated | Session-scoped + background consolidation | Project file lifetime |
| Structure | I/P/B frames + knowledge graph | Flat facts | Flat documents |
| Search | Hybrid (keyword + vector + graph) | Unknown internal | Document-level retrieval |
| Dual scope | Personal + workspace minds | Single global | Per-project only |
| Entity extraction | Automatic with relations | No graph | No graph |
| Provenance | 7 source types tracked | No provenance | File-level only |
| Importance levels | 5 levels with scoring weights | Binary (remembered/not) | No importance |
| Conflict detection | CRITICAL protocol with user confirmation | Silent overwrite | No conflict handling |
| Data locality | User's machine, never leaves | OpenAI servers | Anthropic servers |
| Temporal scoring | Decay + recency boost | Unknown | No temporal weighting |
**The fundamental difference:** Waggle treats memory as a structured, queryable knowledge base with graph relations and temporal scoring. Competitors treat it as a flat fact store or document repository. The I/P/B frame model enables state reconstruction (like git), not just retrieval.
---
## 2. Agent Orchestration
### 2.1 The Agent Loop
`runAgentLoop()` is a clean, well-structured ReAct loop:
- LiteLLM-proxied chat completions (model-agnostic via OpenAI-compatible API)
- SSE streaming with tool call accumulation
- Rate limit handling: exponential backoff with retry cap (3 retries for 429, 502, 503, 504)
- Token budget enforcement (graceful termination when exceeded)
- AbortSignal support for client disconnection
- Injection scanning on every tool output via `scanForInjection()`
- Loop guard preventing identical tool calls from cycling
- Plugin tool merging at runtime
- Team governance policy enforcement (blocked tools list)
- Pre/post hook system (10 event types) for extensibility
### 2.2 System Prompt Construction
`buildSystemPrompt()` assembles 3 sections with section caching:
1. **Identity** (cached -- only recomputes when identity changes): User profile from IdentityLayer
2. **Self-Awareness** (uncached -- changes every call): Tool inventory, memory stats, improvement signals, skills list
3. **Preloaded Context** (uncached): Recent memories (importance-sorted), active awareness items, top knowledge entities, personal preferences
### 2.3 Behavioral Specification v3.0
Split into 5 named sections totaling approximately 290 lines of rules:
- **coreLoop:** 5-step thinking process (RECALL --> ASSESS --> ACT --> LEARN --> RESPOND) with CRITICAL memory conflict protocol
- **qualityRules:** Anti-hallucination discipline, structured output, context grounding, professional disclaimers
- **behavioralRules:** Memory-first, tool intelligence, narration heuristics, error recovery, planning
- **workPatterns:** Drafting from context, decision compression, research in context
- **intelligenceDefaults:** Tool catalog, capability acquisition, sub-agent delegation, workflow composition
### 2.4 Sub-Agent Orchestrator
`SubagentOrchestrator` implements a supervisor/worker pattern:
- **Dependency-ordered execution** with topological sorting
- **Context injection** between steps (step B can access step A's results)
- **Result aggregation** with 3 modes: concatenate, last, synthesize (the synthesize mode spawns a synthesizer sub-agent)
- **EventEmitter-based** status tracking for UI updates
- **7 role presets** with predefined tool sets (researcher, writer, coder, analyst, reviewer, planner, synthesizer)
- **Circular dependency detection** (breaks loops with error state)
### 2.5 Workflow Composer
Implements **lightest sufficient execution mode** selection:
1. `direct` -- Agent handles directly
2. `structured_single_agent` -- Agent follows a plan, no sub-agents
3. `skill_guided` -- Agent uses a loaded skill's workflow
4. `subagent_workflow` -- Full multi-agent orchestration
The composer analyzes task shape (type, phases) and picks the lightest mode that works. This prevents unnecessary sub-agent spawning for simple tasks.
### 2.6 Waggle Dance Protocol
Inter-agent communication protocol with typed messages:
- **Request types:** knowledge_check, task_delegation, skill_request, model_recommendation
- **Response types:** knowledge_match, task_claim
- **Broadcast types:** discovery, routed_share, skill_share, model_recipe
The `WaggleDanceDispatcher` routes messages to real handlers (memory search, worker spawning, skill installation, capability resolution).
---
## 3. Tool Intelligence
### 3.1 Dynamic Tool Filtering
Three filtering mechanisms:
- **Context-based:** Code tools vs research tools vs general (predefined sets)
- **Availability-based:** Runtime `checkAvailability()` on each tool
- **Offline-capable:** Tools tagged with `offlineCapable` for disconnected operation
- **Config-based:** Explicit `enabled_tools` / `disabled_tools` lists
### 3.2 Capability Router
When a tool is not found, `CapabilityRouter` resolves alternatives by searching across 6 sources:
1. Native tools (exact and partial match)
2. Installed skills (keyword matching in content)
3. Plugins (manifest matching)
4. MCP servers (name matching, health-aware)
5. Sub-agent roles (keyword mapping)
6. Connectors (service/action matching)
Falls back to a "missing" route with an install suggestion.
### 3.3 Context Compression
5-step pipeline for long conversations:
1. **Detect:** Estimate tokens (4 chars/token heuristic), check against threshold (default: 50% of 128K)
2. **Prune:** Replace old tool results with "[Cleared]" placeholders (free, no LLM)
3. **Protect:** Split into head (system + first N messages), tail (recent messages), middle (compressible)
4. **Summarize:** Call budget model on the middle using COMPACTION_PROMPT
5. **Inject:** Replace middle with summary message
Iterative: previous summaries are fed back for cumulative compression.
### 3.4 Credential Pool
Round-robin API key rotation with policy-based cooldowns:
- 429 (rate limit) --> 1 hour cooldown, auto-recovers
- 402 (payment required) --> 24 hour cooldown, auto-recovers
- 401 (unauthorized) --> permanently disabled
- Other errors --> 5 minute cooldown
Vault convention: `provider`, `provider-2`, `provider-3`, etc. Supports multiple keys per LLM provider for throughput maximization.
### 3.5 Injection Scanner
Three pattern categories with weighted scoring:
- **Role override patterns** (0.5 weight): "ignore previous instructions", "you are now", multi-language variants, memory wipe attempts
- **Prompt extraction patterns** (0.4 weight): "show your system prompt", "reveal your instructions"
- **Instruction injection patterns** (0.3/0.6 weight): "IMPORTANT: ignore", "[INST]", "<<SYS>>", fake authority claims
Tool outputs are scored higher (0.6) for instruction injection because they are more dangerous attack vectors. Threshold at 0.3 -- anything above is flagged.
### 3.6 Cost Tracker
Per-model pricing table with usage accumulation. Tracks input/output tokens per call with workspace-level cost attribution. Supports real-time daily totals and per-model breakdowns.
### 3.7 Improvement Detector
Three signal categories tracked in `improvement_signals` table:
- **Capability gaps:** When tools are missing, the gap is recorded. After repeated occurrences, it surfaces as an actionable suggestion.
- **Corrections:** User corrections are detected and tracked to prevent repeated mistakes.
- **Workflow patterns:** Recurring multi-step patterns that could benefit from templates.
Signals are surfaced once via the self-awareness system prompt and then marked as surfaced to avoid repetition.
---
## 4. Skill and Plugin System
### 4.1 Skills
Markdown files with optional YAML frontmatter:
```yaml
---
name: Deploy Helper
description: Helps deploy applications
permissions:
codeExecution: true
network: true
---
```
Skills are parsed by `parseSkillFrontmatter()` and loaded into the agent's context. They can be:
- Built-in (shipped with Waggle)
- User-created via the `create_skill` tool
- Shared between agents via Waggle Dance `skill_share` messages
- Discovered and installed via `acquire_capability` / `install_capability`
### 4.2 Hooks
10 lifecycle events with registry pattern:
- `pre:tool` / `post:tool` -- Before/after any tool execution
- `session:start` / `session:end` -- Session lifecycle
- `pre:response` / `post:response` -- Response generation
- `pre:memory-write` / `post:memory-write` -- Memory mutations (can cancel writes)
- `workflow:start` / `workflow:end` -- Workflow lifecycle
Hooks can be scoped to specific workspaces. Pre-hooks can cancel execution. Activity log maintained (last 50 events).
### 4.3 Install Audit Trail
Every capability installation is recorded:
- Timestamp, capability name, type (native/skill/plugin/mcp), source
- Risk level (low/medium/high)
- Trust source, approval class (standard/elevated/critical)
- Action (proposed/approved/installed/rejected/failed)
- Initiator (agent/user/system)
---
## 5. KVARK Integration
4 enterprise tools gated to Business/Enterprise tiers:
- `kvark_search` -- Full-text search across enterprise sources (SharePoint, Jira, Slack)
- `kvark_ask_document` -- Focused Q&A on a specific enterprise document
- `kvark_feedback` -- Relevance feedback loop for retrieval quality improvement
- `kvark_action` -- Governed enterprise actions (create Jira ticket, post Slack message) with audit trail
**Combined Retrieval** merges workspace memory, personal memory, and KVARK results:
- KVARK is only queried when local results are insufficient (< 3 results with score >= 0.7)
- Conflict detection between workspace memory and KVARK results using polarity analysis (positive vs negative status keywords)
- Every result carries explicit source attribution
---
## 6. Technology Decisions
### Why Tauri (not Electron)
- **Binary size:** Tauri 2.0 binaries are 5-15 MB vs Electron's 150+ MB (uses system WebView)
- **Memory footprint:** Significantly lower -- critical for a desktop AI app that already needs memory for embeddings and SQLite
- **Security:** Explicit IPC allowlist in `tauri.conf.json` (no "allow all" wildcard). Rust shell provides memory safety.
- **Cross-platform:** Windows + macOS from single codebase with Rust's cross-compilation
### Why SQLite + sqlite-vec (not Postgres + pgvector)
- **Desktop-first:** No database server to install. The `.mind` file IS the database. Zero config.
- **Portability:** Copy the file, move it between machines. Backup is a file copy.
- **Performance:** WAL mode for concurrent reads, FTS5 is compiled into SQLite. sqlite-vec provides HNSW-like approximate nearest neighbor search.
- **Privacy:** Data never leaves the user's machine. No connection string, no cloud database.
- **Cost:** Zero infrastructure cost. Perfect for a free-tier product.
### Why Fastify (not Express)
- **Performance:** Fastify is 2-5x faster than Express for JSON serialization, which matters for the streaming agent loop.
- **Schema validation:** Built-in JSON schema validation on routes.
- **Plugin system:** Clean plugin architecture for modular route registration.
### LiteLLM -- Model Agnostic Design
- **Single proxy endpoint:** Agent loop talks to one URL regardless of model provider.
- **Key rotation:** Combined with CredentialPool for multi-key management.
- **Model switching:** Users can change models without code changes. Personas declare `modelPreference` but users override.
- **Offline capability:** When LiteLLM is unavailable, offline-capable tools still work.
---
## 7. Technical Moat Assessment
### Strong Moats (Hard to Replicate)
| Component | Moat Strength | Why |
|-----------|--------------|-----|
| I/P/B Frame Model | **High** | Novel application of video compression concepts to memory. The frame-based state reconstruction with GOPs is architecturally unique. No competitor does this. |
| Dual-Mind Architecture | **High** | Personal + workspace memory with automatic routing is a system design insight. Requires deep thinking about scoping that simple RAG does not address. |
| autoSaveFromExchange | **Medium-High** | 30+ regex patterns for passive memory accumulation. The pattern library represents significant behavioral tuning that requires real user testing to calibrate. |
| HybridSearch with RRF | **Medium** | RRF fusion of keyword + vector + graph signals is well-known in IR research but uncommon in desktop AI. The 4 scoring profiles are a usability advantage. |
| CognifyPipeline | **Medium** | End-to-end write path from text to frames + entities + relations + vectors. Straightforward but well-integrated. |
| Behavioral Spec v3.0 | **Medium** | 290 lines of carefully tuned agent rules. The memory conflict protocol is a genuine innovation -- no other agent platform prevents memory drift this way. |
| Context Compression | **Medium** | 5-step pipeline with iterative summaries. The head/middle/tail splitting with budget model summarization is clever. |
| Improvement Signals | **Medium** | Self-correcting agent behavior via recurring pattern detection. Novel concept for consumer AI. |
### Weak Moats (Easily Replicated)
| Component | Moat Strength | Why |
|-----------|--------------|-----|
| Injection Scanner | **Low** | 20 regex patterns. Any team can build equivalent in a day. |
| Cost Tracker | **Low** | Simple pricing table + usage accumulation. |
| Credential Pool | **Low** | Standard round-robin with cooldowns. |
| Tool Filtering | **Low** | Predefined tool sets by context. |
| Entity Extractor | **Low** | Regex-based NER without ML. Accuracy is limited compared to spaCy or LLM-based extraction. |
### Compound Moat
The real moat is not any single component but **the integration of all of them into a coherent memory-first agent platform**. The combination of I/P/B frames + dual-mind + hybrid search + auto-save + behavioral spec + improvement signals creates a system where the agent genuinely gets better over time in a way that is structurally different from competitors.
---
## 8. Technical Debt and Risks
### High Priority
1. **Entity extraction is regex-only.** The `extractEntities()` function uses pattern matching with a hardcoded list of tech terms and proper noun heuristics. This will miss domain-specific entities and produce false positives. An LLM-based or spaCy-based extraction step would dramatically improve KnowledgeGraph quality.
2. **Knowledge graph queries are O(n) scans.** `getEntitiesByType('')` fetches ALL entities, then filters in JavaScript. For large knowledge bases, this will degrade. The graph needs indexed queries and possibly a proper graph traversal engine.
3. **No vector index maintenance.** sqlite-vec does not have automatic index rebuilding. As frames are deleted or updated, orphan vectors accumulate. No vacuum or reindexing mechanism exists.
4. **Token estimation is 4-chars-per-token heuristic.** The context compressor uses this approximation. For non-English text or code-heavy conversations, this can be off by 30-50%, causing premature or late compression.
5. **No memory compaction/consolidation.** Frames accumulate indefinitely. There is no mechanism to merge old P-frames into new I-frames, or to prune deprecated frames. Over months of use, the `.mind` file will grow unboundedly.
### Medium Priority
6. **Scoring profiles are static.** The 4 scoring profiles have hardcoded weights. There is no adaptive scoring that learns which profile works best for a given user or workspace.
7. **Conflict detection is keyword-based.** The `detectConflict()` function uses simple polarity word lists. It will miss semantic conflicts and produce false positives on keyword collisions.
8. **No embedding dimension migration.** If the embedding model changes (different dimension count), existing vectors in `memory_frames_vec` become incompatible. No migration path exists.
9. **Stripe integration is incomplete.** The tier system is defined but `stripePriceId` values come from environment variables. No billing webhook handling visible in the codebase.
10. **Waggle Dance protocol is partially implemented.** The dispatcher handles 4 of 8 message subtypes. `model_recommendation`, `knowledge_match`, `task_claim`, and `discovery` are defined but not dispatched.
### Low Priority
11. **No rate limiting on sidecar API routes.** The Fastify server exposes endpoints without throttling.
12. **Improvement signals are never pruned.** The `improvement_signals` table grows indefinitely with no archival.
---
## 9. Comparison to Competitors
### vs. ChatGPT (OpenAI)
| Dimension | Waggle OS | ChatGPT |
|-----------|-----------|---------|
| Memory model | Structured frames with I/P/B + knowledge graph | Flat fact store, opaque consolidation |
| Data location | Local (user's machine) | OpenAI cloud |
| Search | Hybrid (keyword + vector + graph) | Unknown internal |
| Multi-workspace | Dual-mind (personal + workspace) | Single global memory |
| Tool extensibility | Skills + plugins + MCP + connectors | GPT Actions (HTTP endpoints) |
| Enterprise bridge | KVARK integration with governed actions | No self-hosted option |
| Cost visibility | Per-model tracking with daily totals | Hidden in subscription |
| Offline | Partial (offline-capable tools) | None |
**Waggle advantage:** Memory depth, data sovereignty, enterprise bridge
**ChatGPT advantage:** Scale, model quality (GPT-4 family), ecosystem (millions of GPTs)
### vs. Claude Projects (Anthropic)
| Dimension | Waggle OS | Claude Projects |
|-----------|-----------|-----------------|
| Memory model | I/P/B frames + auto-save from conversations | Static files uploaded to project |
| Persistence | Permanent, cross-session, auto-enriched | File lifetime only |
| Context | Automatic memory recall per query | Full project files in context window |
| Desktop | Native Tauri app | Web-only |
| Agent tools | 30+ tools with plugin system | Limited tool use |
| Multi-agent | SubagentOrchestrator with dependency DAG | No multi-agent |
| Enterprise | KVARK with governed actions | No enterprise bridge |
**Waggle advantage:** Automatic memory, desktop, multi-agent, enterprise
**Claude advantage:** Model quality (Claude 4), massive context window (1M tokens), simpler UX
### vs. Cursor / Windsurf / Cline (AI Code Editors)
| Dimension | Waggle OS | AI Code Editors |
|-----------|-----------|-----------------|
| Scope | General-purpose workspace AI | Code-focused |
| Memory | Persistent knowledge graph | Code index only |
| Personas | 22 domain-specific roles | Single coding persona |
| Document output | DOCX generation, reports, briefs | Code output only |
| Enterprise | KVARK + governed actions | GitHub/GitLab integration |
**Waggle advantage:** Breadth (not just code), persistent memory, enterprise
**Code editor advantage:** Deeper code understanding, LSP integration, inline editing
### vs. Notion AI / Mem.ai
| Dimension | Waggle OS | Notion AI / Mem.ai |
|-----------|-----------|---------------------|
| Agent capability | Full ReAct loop with tools | Q&A over documents |
| Memory | Auto-extracted structured frames | Document-level |
| Privacy | Local-only SQLite | Cloud |
| Multi-agent | Yes | No |
| Extensibility | Skills + plugins + MCP | Limited |
**Waggle advantage:** True agent with tools, local data, extensibility
**Notion/Mem advantage:** Better collaborative editing, richer document UX
---
## 10. Summary of Crown Jewels
### Tier 1 -- Genuinely Innovative
1. **I/P/B Frame Model with GOP Sessions** -- Novel application of video compression to AI memory. Enables state reconstruction, importance-weighted retrieval, and provenance tracking in a way no competitor does.
2. **Dual-Mind Architecture** -- Separating personal identity/preferences from workspace knowledge, with automatic routing, is a system design insight that solves a real problem (cross-project preference continuity).
3. **Memory Conflict Protocol** -- The CRITICAL block in the behavioral spec that prevents memory drift through contradiction detection and user confirmation is a safety innovation absent from all competitors.
4. **autoSaveFromExchange** -- Passive memory accumulation from every conversation turn, with 30+ calibrated patterns for preferences, decisions, corrections, and research findings. This is what makes the memory system feel "alive."
### Tier 2 -- Well-Engineered Differentiators
5. **HybridSearch with Multi-Signal Scoring** -- RRF fusion of keyword + vector + knowledge graph with 4 configurable profiles and temporal decay. Solid IR engineering.
6. **Context Compression Pipeline** -- 5-step iterative compression that preserves critical information while managing context window limits. The head/middle/tail split with budget model summarization is well-designed.
7. **KVARK Combined Retrieval with Conflict Detection** -- Merging local memory with enterprise knowledge, only querying KVARK when local results are insufficient, with polarity-based conflict detection.
8. **Improvement Signal System** -- Self-correcting agent behavior through recurring pattern detection. The agent learns from its own failures.
### Tier 3 -- Solid Infrastructure
9. **Tier-Gated Capabilities** -- Clean tier architecture (SOLO/BASIC/TEAMS/ENTERPRISE) with per-capability enforcement including embedding quotas.
10. **CredentialPool with Policy-Based Cooldowns** -- Production-grade key rotation for multi-provider LLM access.
11. **Hook System** -- 10-event lifecycle with workspace scoping and cancellation support. Enables governance and extensibility.
12. **Waggle Dance Protocol** -- Inter-agent communication with typed messages for team collaboration.
---
*End of analysis.*