This commit is contained in:
Oleg Maslov
2026-09-02 10:14:22 +02:00
parent 0c3e2ead3b
commit b20b138fe4
771 changed files with 161561 additions and 9027 deletions

View File

@@ -1,5 +1,10 @@
# Phase 1A: Feature Wave Audit — Built vs Not Built
> **Historical snapshot — superseded.** This report records evidence as of
> 2026-03-20 and is not current launch authority. See
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
> Windows-first Solo scope, external gates, and recommendation.
**Date**: 2026-03-20
**Auditor**: Claude (automated codebase cross-reference)
**Plan document**: `docs/plans/2026-03-19-phase9-completion-plan.md`

View File

@@ -1,5 +1,10 @@
# Phase 1B: Deployment, Phases & PM Features Audit
> **Historical snapshot — superseded.** This report records evidence as of
> 2026-03-20 and is not current launch authority. See
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
> Windows-first Solo scope, external gates, and recommendation.
**Date**: 2026-03-20
**Auditor**: Claude (automated analysis)
**Scope**: Wave 9D (Deployment), Phase 7 (KVARK), Phase 8 status, PM Features (6)

View File

@@ -1,5 +1,10 @@
# Phase 2: UX Audit — View-by-View Code Review
> **Historical snapshot — superseded.** This report records evidence as of
> 2026-03-20 and is not current launch authority. See
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
> Windows-first Solo scope, external gates, and recommendation.
**Date:** 2026-03-20
**Auditor:** Production Readiness Automation (Phase 2)
**Scope:** All 7 views, sidebar, onboarding, Direction D compliance, emotional assessment

View File

@@ -1,5 +1,10 @@
# Phase 3A: Agent Critical Path Quality Audit
> **Historical snapshot — superseded.** This report records evidence as of
> 2026-03-20 and is not current launch authority. See
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
> Windows-first Solo scope, external gates, and recommendation.
**Auditor**: Production Readiness Review (automated)
**Date**: 2026-03-20
**Scope**: Agent loop, memory, vault, cron, connectors, sub-agents

View File

@@ -1,5 +1,10 @@
# Phase 3B: Server & API Layer — Production Readiness Audit
> **Historical snapshot — superseded.** This report records evidence as of
> 2026-03-20 and is not current launch authority. See
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
> Windows-first Solo scope, external gates, and recommendation.
**Date**: 2026-03-20
**Scope**: `@waggle/server` — local server (Fastify), routes, SSE, WebSocket, KVARK client, security middleware
**Auditor**: Claude Opus 4.6 (automated code review)

View File

@@ -1,5 +1,10 @@
# Phase 3C: UI & Frontend Code Quality Audit
> **Historical snapshot — superseded.** This report records evidence as of
> 2026-03-20 and is not current launch authority. See
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
> Windows-first Solo scope, external gates, and recommendation.
**Auditor**: Senior Engineer Code Review (automated)
**Date**: 2026-03-20
**Scope**: `app/src/` (Tauri desktop app) + `packages/ui/src/` (shared React component library)

View File

@@ -1,5 +1,10 @@
# 04A Application Security Audit
> **Historical snapshot — superseded.** This report records evidence as of
> 2026-03-20 and is not current launch authority. See
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
> Windows-first Solo scope, external gates, and recommendation.
**Date**: 2026-03-20
**Auditor**: Automated (Claude Opus 4.6)
**Scope**: Waggle desktop app + local server — CSP, vault crypto, agent tools, input validation, sessions, connectors, dangerous patterns

View File

@@ -1,5 +1,10 @@
# 04B — Secret Scanning & Dependency Audit
> **Historical snapshot — superseded.** This report records evidence as of
> 2026-03-20 and is not current launch authority. See
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
> Windows-first Solo scope, external gates, and recommendation.
**Auditor:** Claude Opus 4.6 (automated)
**Date:** 2026-03-20
**Scope:** Full codebase secret scan, git history review, npm dependency audit, .gitignore assessment

View File

@@ -1,5 +1,10 @@
# 05 - TEST & COVERAGE REPORT
> **Historical snapshot — superseded.** This report records evidence as of
> 2026-03-20 and is not current launch authority. See
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
> Windows-first Solo scope, external gates, and recommendation.
**Date:** 2026-03-20
**Auditor:** Claude Opus 4.6 (automated, read-only)
**Scope:** All packages, integration tests, E2E tests, visual regression tests

View File

@@ -1,5 +1,10 @@
# Phase 6: Build & Deployment Readiness Report
> **Historical snapshot — superseded.** This report records evidence as of
> 2026-03-20 and is not current launch authority. See
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
> Windows-first Solo scope, external gates, and recommendation.
**Date**: 2026-03-20
**Auditor**: Claude Opus 4.6 (automated)
**Scope**: TypeScript compilation, Vite build, Docker, Tauri, Render.com, npx launcher, CI/CD

View File

@@ -1,5 +1,10 @@
# Issue Register — Waggle V1 Pre-Production Qualification
> **Historical snapshot — superseded.** This register records findings as of
> 2026-03-20 and is not the current open-issue or launch authority. See
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
> Windows-first Solo scope, external gates, and recommendation.
Generated: 2026-03-20 | Branch: `phase8-wave-8f-ui-ux` | Tests: 3,895 passing
---

View File

@@ -1,5 +1,10 @@
# Confidence Matrix — Waggle V1 Pre-Production Qualification
> **Historical snapshot — superseded.** This score records evidence as of
> 2026-03-20 and is not the current readiness score or launch authority. See
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
> Windows-first Solo scope, external gates, and recommendation.
Generated: 2026-03-20 | Branch: `phase8-wave-8f-ui-ux`
---

View File

@@ -1,143 +1,109 @@
# Launch Recommendation — Waggle V1
# Launch Recommendation — Windows Solo
Generated: 2026-03-20
Updated: 2026-08-27
---
## Current verdict: INSTALLER/LIFECYCLE INTERNAL RC QUALIFIED; RELEASE QUALIFICATION INCOMPLETE
## Recommendation: CONDITIONAL GO
The supported launch scope is Windows Solo with Claude Code, Codex, and Hermes using
their official user-owned installations and authentication. Cursor, OpenClaw, and macOS
packaging/certification remain roadmap work and do not block this internal RC.
Ship after fixing the 8 CRITICAL issues (~6.5 hours of work). The HIGH issues are important but can be addressed in a rapid V1.0.1 patch within the first week.
This document is the release-status authority. Historical receipts are evidence only;
they do not certify a later behavior-changing revision unless a bounded no-impact review
explicitly says so.
---
## Frozen runtime candidate
## Executive Summary
- Source revision: `23ad3fa5f99bddce648b84750a41365299aeb0da`
- Source tree: `aab548f77ec64b181086664dad29c32e6bc78779`
- Integration: private Waggle PR #66, tested head
`587e259166db69ff86e806fa8393a5f8974ea0a1`, merged 2026-08-27
- Tree equivalence: the tested PR head and merge commit resolve to the same source tree
- Repository state after merge: private `main` equals `origin/main`
Waggle is a substantial, well-architected product with 3,895 passing tests, 53 agent tools, 29 connectors, 15K+ marketplace packages, and a complete feature set covering all 8 Kill List use cases. The agent loop, memory system, and vault encryption are architecturally sound. The UI underwent a recent Phase 10 rewrite that brought Tailwind adoption and Direction D palette cleanup.
The documentation-only descendant that updates this record does not replace the runtime
candidate. Before it is merged, its diff must be limited to documentation and all required
remote checks must remain green.
However, the audit uncovered **8 CRITICAL issues** (5 security, 1 stability, 2 UX) that must be fixed before any external user touches the product. The most severe: CORS is wide open (any website can call your localhost APIs), there are no React error boundaries (one render error = permanent white screen), and the streaming loading indicator is invisible (users can't tell the agent is thinking).
## Exact-current Windows installer
The good news: every CRITICAL fix is straightforward. Total estimated effort is 6.5 hours. None require architectural changes.
- NSIS artifact:
`app/src-tauri/target/x86_64-pc-windows-msvc/release/bundle/nsis/Waggle_0.2.0_x64-setup.exe`
- Size: 102,923,936 bytes
- SHA-256: `7BFA9F9B13633A51CD3336B42E3EF904B7F7A568C6DEE4F6CED967CBD4F40A59`
- Internal signer: `CN=Egzakta Internal Pilot, O=Egzakta Group, C=RS`
- Signer thumbprint: `E2F028541E7A4D1FE80FFFF02079060D36579846`
- RFC 3161 timestamp authority: DigiCert SHA256 RSA4096 Timestamp Responder 2025 1
- Trust classification: internal pilot only; the self-signed root is not public trust
---
Clean-profile certification passed **64/64** checks in 450.199 seconds:
## What's Ready (Strengths)
- Receipt:
`output/installer-certification/23ad3fa5-20260827T123917Z-exact-main-clean-profile/windows-installer-certification.json`
- Receipt SHA-256:
`AC2A1C54119E28CC22DA931EB43E2815862B01832DD6097CE03F8F79D9D3DF4D`
- Receipt source and bundled-sidecar revision: exact `23ad3fa5`
- Tier: FREE/Solo
- Managed model: `qwen2.5:0.5b`
- Managed-model digest:
`sha256:a8b0c51577010a279d933d14c2a8ab4b268079d44c5c8830c0a93900f1827c67`
1. **Solid agent core** — 53 tools, loop guards, injection scanning, approval gates, sub-agent orchestration. 1,272 tests on the agent package alone.
The receipt proves silent install, bundled Node/npm and offline package execution,
first boot, in-process embeddings, built-in proxy/session authentication, workspace and
memory persistence, managed runtime/model pull and chat, proxy-restart chat, same-version
repair, relaunch, data preservation, Exit/owned-process cleanup, uninstall, registry
cleanup, and preservation of external `.hive-mind` and `.ollama` roots. It also proves
that developer Node.js, Python, Docker, external LiteLLM, and a separately installed
Ollama are not prerequisites.
2. **Complete feature set** — All 8 Kill List use cases work. Workspace memory, connectors, marketplace, personas, cron, swarm protocol, capability packs, onboarding with memory import.
## Integrated test and security evidence
3. **Good test coverage** — 3,895 tests across 277 files, zero failures. Every major package has dedicated test suites with behavior-focused assertions and realistic mocks.
- PR #66: every blocking remote check passed — primary CI, Playwright smoke and full
E2E, Windows and both macOS Tauri verification targets, Wave 1, and Hive Mind
install/smoke on Windows, Ubuntu, and macOS.
- Exact-current dependency audits: 0 Critical and 0 High in both full and production
dependency trees. Lower-severity maintenance remains tracked.
- The security-hardening integration covers workspace path/link boundaries, hook/database
hard-link boundaries, normalized ingress, atomic consolidation/cognify, deprecated-frame
search exclusion before limits, and knowledge-graph provenance.
- Hosted signing policy and workflow tests remain green, but no publicly trusted hosted
artifact has been produced.
4. **Security fundamentals** — AES-256-GCM vault, parameterized SQL everywhere, path traversal protection, CLI allowlists, DOMPurify HTML sanitization, SecurityGate for marketplace.
## Persona, router, and authentication evidence
5. **Deployment infrastructure** — Tauri Windows installer built (8.2MB), Docker production compose, Render.com blueprint, GitHub Actions CI/release pipeline.
The historical ten-persona collection at `4c712ff6` contains 30/30 results at or above
95/100 after documented independent semantic adjudication. It is not relabeled as an
exact-current deterministic seal: PR #66 changed memory behavior, so public release
qualification requires either a fresh exact-candidate collection or an explicit bounded
semantic-impact attestation.
6. **Product polish** — 8 personas, dark/light mode, keyboard shortcuts, global search, workspace hue colors, onboarding wizard, tool card transparency, approval gates inline in chat.
Smart-router primary, compact-tool-context, durable-budget/fallback, and official-user-auth
canaries for Claude Code, Codex, and Hermes remain scoped historical evidence. No PR #66
change altered provider credential ownership or copied/read provider credential files.
---
## Hive Mind repository state
## Must Fix Before Launch (CRITICAL — ~6.5 hours)
Curated public-mirror hardening PR #53 merged to `marolinik/hive-mind` `master` as
`3410327800db3ea23f875d547a0c7f4d08826b7e`; Linux, Windows, macOS, and Ubuntu
first-run smoke passed. The immutable drift checker still reports reviewed blockers and
one unreviewed difference, with zero forbidden exports. Therefore the Windows Solo RC is
not blocked, but the next Hive Mind package release remains a separate maintainer-curated
operation. Raw subtree publication remains forbidden.
### Security (4 hours)
## Public GO blockers
| # | Issue | Fix | Time |
|---|-------|-----|------|
| 1 | **CORS allows any origin** — any website can call all Waggle APIs | Change `origin: true` to `origin: ['http://localhost:1420', 'tauri://localhost']` (or your Tauri webview origins). Fix SSE hijack endpoints to use the same allowlist. | 1.5 hr |
| 2 | **Server CSP has `unsafe-eval` + `unsafe-inline`** | Remove both. If scripts break, use nonces or hashes instead. | 30 min |
| 3 | **OAuth refresh tokens stored plaintext** | Encrypt refresh tokens the same way access tokens are encrypted in `setConnectorCredential()`. | 1 hr |
| 4 | **Verify API key revoked** | Go to Anthropic dashboard, confirm the key from commit `c29d75f` is revoked. Delete local branch `phase6-capability-truth`. | 30 min |
Public release may be called **GO** only after all of these are closed for the approved
release-tag commit:
### Stability (2 hours)
1. A protected hosted build produces a publicly trusted Authenticode artifact.
2. The managed Codex Security workflow produces a sealed Deep Security report with no
unresolved Critical or High findings.
3. Current persona qualification is sealed by a fresh exact-candidate receipt or an
independently reviewed bounded semantic-impact attestation.
4. Smart-router and Claude Code/Codex/Hermes official-auth qualification is either rerun
on the exact candidate or covered by a concrete independently reviewed no-impact
attestation.
5. Protected release-tag checks are green and the exact artifact hashes are recorded.
| # | Issue | Fix | Time |
|---|-------|-----|------|
| 5 | **Zero error boundaries** | Add `<ErrorBoundary>` wrapping each view in App.tsx, plus one at the app root. Use react-error-boundary or a simple class component. Show "Something went wrong" with a retry button. | 2 hr |
### UX (30 minutes)
| # | Issue | Fix | Time |
|---|-------|-----|------|
| 6 | **Streaming indicator invisible** | The loading dots use BEM CSS classes with no definitions. Either add the CSS or replace with Tailwind `animate-pulse` dots. | 15 min |
| 7 | **SplashScreen wrong palette** | Replace `#1a1a2e`/`#16213e`/`#0f3460` with Direction D tokens. Change `#f5a623` to `#d4a843`. | 15 min |
---
## Ship-Week Fixes (HIGH — ~28 hours, V1.0.1)
**Security hardening (first 2 days):**
- Approval gates: change auto-approve to auto-deny on 5min timeout (15 min)
- WebSocket authentication: require session token on `/ws` connect (2 hr)
- Team WebSocket: validate JWT instead of trusting userId param (2 hr)
- Replace `xlsx` with `exceljs` to fix prototype pollution (2 hr)
- Generate Tauri updater keypair and set pubkey (30 min)
**Agent loop safety (day 3):**
- Cap rate-limit retries (max 3, then fail gracefully) (1 hr)
- Add token budget enforcement with configurable limit (2 hr)
- Parameterize sqlite-vec SQL interpolation (30 min)
**Frontend stability (days 3-5):**
- Add code splitting with `React.lazy()` for 7 views (2 hr)
- Deduplicate SSE connections (1 hr)
- Fix eventBus.removeAllListeners to scope per-client (1 hr)
- Fix light theme breakage across components (2 hr)
**Build fixes (day 5):**
- Fix npx waggle: compile .ts entry, resolve workspace deps (2 hr)
- Add non-root user to Docker (30 min)
- Fix CI branch target master→main (15 min)
- Clean up 87 TypeScript errors (2 hr)
---
## Known Limitations (Ship Anyway)
These are acceptable for V1 and can be improved iteratively:
1. **No browser E2E tests** — Unit/integration coverage is strong (3,895 tests). True browser automation (Playwright user journeys) is a V1.1 investment. Screenshot baselines exist.
2. **Monolithic App.tsx (1300 lines)** — Works but hard to maintain. Refactoring into feature-specific providers is a V1.1 task that won't affect users.
3. **No React.memo optimization** — The app performs fine at current scale. Memoization is premature optimization until profiling shows problems.
4. **macOS build not configured** — DMG, code signing, notarization require an Apple Developer account. Windows installer works. Ship Windows-first, add macOS in V1.1.
5. **Direction D at ~78%** — The Phase 10 UI rewrite made massive progress (371→19 inline styles). Remaining 22% is polish, not broken functionality.
6. **KVARK client not wired** — KVARK integration (Phase 7) is library code + tests. Not wired into the running server because KVARK itself needs its HTTP API deployed first. This is expected — it's the Enterprise tier path.
7. **Conversation history unbounded** — At typical usage (10-50 turns/session), this isn't a problem. Add context window management for power users in V1.1.
---
## Post-Launch Priority Queue
### First Week (V1.0.1)
1. All HIGH security fixes (approval timeout, WebSocket auth, xlsx, updater pubkey)
2. Agent loop safety (retry cap, token budget)
3. Frontend stability (code splitting, SSE dedup, error boundaries for remaining components)
4. Light theme fixes
### First Month (V1.1)
1. Browser E2E test suite (Playwright user journeys)
2. React component rendering tests
3. App.tsx decomposition (extract providers/hooks)
4. macOS build + code signing
5. npx waggle publishable package
6. CI pipeline expansion (Docker, lint, security scan)
7. Direction D compliance to 95%+
### First Quarter (V1.2)
1. KVARK server-side wiring (when KVARK HTTP API ready)
2. Performance profiling + React.memo optimization
3. Context window management for long conversations
4. Token budget UI (user-configurable spend limits)
5. Full accessibility audit (WCAG 2.1 AA)
---
## Verdict
**CONDITIONAL GO** — Fix the 8 CRITICALs (6.5 hours), then ship. The product is feature-complete, well-tested, and architecturally sound. The critical issues are configuration mistakes, not design flaws. Every fix is surgical and low-risk.
The foundation is strong. Ship it.
Until then, the installer is suitable for controlled internal testing, not public
distribution, and Waggle must not be described as publicly production-ready or GO.

View File

@@ -0,0 +1,73 @@
# Windows Solo security review — refreshed 2026-08-27
## Status
The integrated Windows Solo source candidate is
`23ad3fa5f99bddce648b84750a41365299aeb0da` (private PR #66). Its tested PR
head and merge commit have the same source tree
`aab548f77ec64b181086664dad29c32e6bc78779`.
This is an evidence-backed source, dependency, CI, and installed-runtime review. It is
**not** a substitute for the still-missing sealed managed Deep Security report and does
not confer public release approval.
## Verified controls
| Surface | Current evidence | Result |
|---|---|---|
| Workspace paths and lifecycle | Strict workspace IDs, canonical containment, identity-matching configs, link/hard-link rejection, and fail-closed list/get/delete behavior; focused executable tests | Pass |
| Hook/database boundary | Existing database/config entries and SQLite companions reject links or hard links while preserving first-use and WAL behavior | Pass |
| External memory ingress | Normalization and injection scanning are applied before persistence across supported ingress paths; bypass-focused regressions are included | Pass |
| Search | Punctuated identifiers use bounded precise fallback; deprecated frames are excluded before keyword, LIKE, whole-vector, and chunk-vector limits; alternate-lane starvation regressions are covered | Pass |
| Consolidation/cognify | Supersede operations are atomic, preserve provenance, and fail closed on partial mutation | Pass |
| Knowledge graph | Relationship provenance and source-frame boundaries are preserved and tested | Pass |
| Local server and tools | Loopback/session authentication, origin controls, SSRF/DNS/socket-pinning defenses, bounded external input, command-vector execution, and fail-closed shim handling are covered by focused and remote gates | Pass |
| Provider authentication | Historical scoped canaries show Claude Code, Codex, and Hermes using official user-owned authentication with no provider credential-file reads/copies; exact-current carry-forward still needs a concrete no-impact attestation or rerun | Historical evidence; current qualification open |
| Packaged runtime | Exact-current internal-pilot NSIS passed 64/64 clean-profile install, boot, managed-model, repair, relaunch, Exit/cleanup, and uninstall checks | Pass internal RC |
| Dependency severity | Exact-current full and production audits contain 0 Critical and 0 High findings | Pass Critical/High gate |
## Exact installer evidence
- Installer SHA-256:
`7BFA9F9B13633A51CD3336B42E3EF904B7F7A568C6DEE4F6CED967CBD4F40A59`
- Certification receipt:
`output/installer-certification/23ad3fa5-20260827T123917Z-exact-main-clean-profile/windows-installer-certification.json`
- Receipt SHA-256:
`AC2A1C54119E28CC22DA931EB43E2815862B01832DD6097CE03F8F79D9D3DF4D`
- Managed model: `qwen2.5:0.5b`
- Managed-model digest:
`sha256:a8b0c51577010a279d933d14c2a8ab4b268079d44c5c8830c0a93900f1827c67`
- Certification checks: 64 passed, 0 failed
The certifier verified source and sidecar provenance, bundled runtime/npm, clean offline
execution, default Solo onboarding, in-process embeddings, built-in proxy liveness,
session authentication, managed-model pull/chat, proxy-restart chat, repair and data
preservation, relaunch, cleanup/uninstall, and unchanged external `.hive-mind`/`.ollama`
roots. No Waggle-owned process or certificate test profile remained after completion.
## Remote integration evidence
PR #66 passed primary CI, Playwright smoke and full E2E, Windows and both macOS Tauri
verification targets, Wave 1, and Hive Mind install/smoke on Windows, Ubuntu, and macOS.
The full local Waggle Vitest suite, agent/server/app typechecks, lint, and diff checks also
completed successfully before integration.
Hive Mind PR #53 passed Linux, Windows, macOS, and Ubuntu first-run smoke before merge as
`3410327800db3ea23f875d547a0c7f4d08826b7e`.
## Residual risk and public release blockers
- The installer is signed by `CN=Egzakta Internal Pilot`, a private self-signed identity.
Its DigiCert timestamp validates the signing pipeline but does not provide public trust.
- No sealed managed Codex Security report exists; this audit host used a disabled
permission profile. No failed or unsealed attempt is interpreted as a no-findings result.
- The immutable Hive Mind drift baseline reports 22 known reviewed blockers and one
unreviewed difference, with zero forbidden exports. These block the next OSS package
release, not this private Windows Solo internal RC.
- Current persona qualification still needs a fresh exact-candidate seal or an independent
bounded semantic-impact attestation because PR #66 changed memory behavior.
Public GO requires publicly trusted Authenticode, a sealed exact-candidate managed Deep
Security report with no unresolved Critical/High findings, current persona qualification,
fresh or explicitly attested smart-router and official-auth qualification, and green
protected release-tag checks.

View File

@@ -0,0 +1,115 @@
# Hive Mind mirror parity audit — refreshed 2026-08-27
## Verdict
The curated Hive Mind mirror does not block the private Windows Solo internal RC. The
Waggle monorepo remains the sole source of truth for the memory substrate.
The next public Hive Mind package release is **not yet approved**. The mirror is clean,
merged, and cross-platform green, but the immutable drift baseline still contains known
reviewed blockers and one unreviewed difference. It must not be described as fully
synchronized or release-ready.
## Exact inputs
- Waggle runtime/source candidate:
`23ad3fa5f99bddce648b84750a41365299aeb0da`
- Waggle source tree:
`aab548f77ec64b181086664dad29c32e6bc78779`
- Hive Mind `master`:
`3410327800db3ea23f875d547a0c7f4d08826b7e`
- Hive Mind integration: PR #53, tested head
`f5072efbf3f2d13247bb91505d99403acbdf11c6`
- Comparison:
`node scripts/oss-drift-check.mjs D:/Projects/hive-mind`
- Result: exit 1, fail closed
The Waggle documentation-only descendant does not change either compared source tree.
## Immutable baseline result
| Classification | Count | Meaning |
|---|---:|---|
| Parity | 5 | Reviewed paths with the expected canonical/OSS state |
| Reviewed adaptations | 38 | Intentional layout, import, logger, branding, or OSS-architecture differences |
| Known reviewed blockers | 22 | Paths still requiring maintainer-curated reconciliation before an OSS release |
| Unreviewed differences | 1 | A changed path not yet classified against the baseline |
| Forbidden exports | 0 | No prohibited Waggle-only path or marker was found in the mirror |
The single unreviewed difference is `harvest/raw-turns.ts`. It must be classified and
either reconciled or deliberately re-baselined through maintainer review before release.
The five parity paths are:
- `harvest/claude-code-adapter.ts`
- `harvest/decision-derivation.ts`
- `harvest/stable-id.ts`
- `mind/erasure.ts`
- `mind/inprocess-reranker.ts`
## Known reviewed blockers
The baseline currently marks 19 paths `reconcile`:
- Harvest: `dedup.ts`, `extract-kg-entities.ts`, `pipeline.ts`,
`url-adapter.ts`
- Runtime/API: `index.ts`, `workspace-manager.ts`, `multi-mind-cache.ts`
- Mind: `db.ts`, `embedding-provider.ts`, `frames.ts`, `identity.ts`,
`inprocess-embedder.ts`, `knowledge.ts`, `llm-extractor.ts`,
`raw-archive.ts`, `raw-detail-lane.ts`, `schema.ts`, `search.ts`, and
`suppression.ts`
Three more blockers are marked `product-curation`: `hook-runtime.ts`,
`mind/supersede.ts`, and `multi-mind.ts`. They require an explicit reviewed decision
about what remains Waggle-only and what generic subset, if any, should be adapted for the
public mirror. The approved decision must then reclassify them out of the blocker category.
A blocker in this inventory means the mapped files are not yet approved for the next OSS
release; it does not by itself assert a runtime vulnerability. Every blocker needs
canonical-first review. Only reconcile/exported paths require public-layout adaptation and
focused port tests; product-curation paths may instead remain Waggle-only after explicit
review. Every resolution requires a new baseline receipt.
## Work integrated on 2026-08-27
Hive Mind PR #53 added the curated search-boundary hardening needed for punctuated
identifiers and deprecated-candidate filtering before keyword, LIKE, whole-vector, and
chunk-vector limits. The implementation closed concatenated-token, malformed-query, CJK,
and alternate-vector-lane bypasses with executable regressions.
The PR passed build/test on Ubuntu, Windows, and macOS plus Ubuntu first-run smoke before
merge. A transient Ubuntu dependency-download `ECONNRESET` was rerun once and passed;
the deterministic code/test result was green.
This narrows the drift but does not erase the remaining `mind/search.ts` curated
difference, so the immutable checker correctly continues to fail closed.
## Proprietary exclusion boundary
Raw subtree publication is forbidden. A release curation must continue to exclude:
- `vault.ts`
- `mind/evolution-runs.ts`
- `mind/execution-traces.ts`
- `mind/improvement-signals.ts`
- `compliance/**`
- Waggle-only `install_audit` schema/migration fragments interleaved in shared files
The current scan found zero forbidden exports. That is a necessary safety gate, not proof
that all public-worthy changes have been reconciled.
## Release rule
Before the next Hive Mind package release:
1. Classify `harvest/raw-turns.ts`.
2. Reconcile or explicitly re-baseline every known blocker in small reviewed phases.
3. Run focused tests for each curated port and the complete Hive Mind build/test suite.
4. Run Linux, Windows, macOS, and first-run package smoke.
5. Re-run the immutable drift checker and require exit 0: zero known blockers, zero
unreviewed differences, and zero forbidden exports. Any approved product-curation
decision must first be represented by reviewed baseline reclassification.
6. Preserve provenance showing that generic substrate work originated in Waggle first.
Until those gates pass, `master` is a clean integrated development baseline, not a sealed
new OSS package release.

View File

@@ -1,5 +1,11 @@
# Waggle OS — Production Readiness Audit (2026-07-03)
> [!CAUTION]
> **HISTORICAL SNAPSHOT — NOT CURRENT SHIP AUTHORITY.** Findings and grades
> below describe the cited July baseline. Use
> `docs/production-readiness/09-LAUNCH_RECOMMENDATION.md` for current scope,
> evidence, open gates, and release verdict.
**Method:** 10 parallel principal-engineer audit lanes (opus, high effort), each required to cite `file:line` evidence it actually read. Baseline at audit time: HEAD `78660ab5` on `main`, vitest 8063/8063 green, lint 0, tsc 0 (agent/server/app), `build:all` clean, git history secret-scan CLEAN (all key-shaped strings are `detectSecrets()` fixtures).
**Subsystem grades:** Build **D** · CI/CD **C** · Docs/DX **C** · Deps/Config **C** · Agent-runtime **C** · Security **B** · Testing **B** · Server-API **B** · Frontend **B** · Memory-substrate **B**.

View File

@@ -1,5 +1,11 @@
# Waggle OS — Production Readiness Sign-off (2026-07-03)
> **Historical snapshot — superseded.** This document preserves the 2026-07-03
> assessment but is not current ship authority. Its “production-ready” and
> signing-only conclusions no longer apply. Use
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the active
> Windows Solo gate and its exact-HEAD evidence requirements.
> Engineering-director pass driven by a 10-lane parallel audit, executed as prioritized remediation streams (Opus implementation agents, Fable orchestration/QA). Companion to [`2026-07-03-release-audit.md`](./2026-07-03-release-audit.md).
## Session commit ledger (17 commits on `main`, from baseline `a1fad4f8`)

View File

@@ -0,0 +1,85 @@
# Waggle OS production-readiness audit — 2026-07-16
> **Historical snapshot — superseded.** This document records the state on
> 2026-07-16 and is not current ship authority. Use
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the active
> Windows Solo gate. Cursor and OpenClaw are now roadmap-only; macOS
> certification is deferred.
## Executive verdict
The branch is not yet production-ready for the requested claim of a fully
Docker-independent Solo install or a verified 9.5/10 result across ten fresh
personas. It is materially improved and buildable: Windows sandbox/process
handling is hardened, chat schemas are bounded before model serialization, and
artifact/CLI writes now require approval.
## Verified changes
- Windows workspace/path, environment, direct-argv `run_code`, descendant
process termination, CLI `.cmd` discovery, MCP env sanitization, and MCP
start/stop race handling were implemented.
- Chat tool selection is deterministic and subtractive: maximum 14 tools and
8,000 exact OpenAI/LiteLLM schema characters per turn; duplicate names keep
the native definition; plugin/MCP tools cannot appear in casual chat unless
explicitly relevant; governance-blocked schemas are removed before model
serialization.
- `cli_execute`, `multi_edit`, DOCX/XLSX/PPTX/PDF generation are approval-gated.
- DOCX/PDF/PPTX/XLSX workspace containment uses `path.relative`, preventing
sibling-prefix escapes such as `workspace-sibling`.
- Coder, analyst, finance, data-engineer, executive-assistant, and general
purpose persona allowlists include their shipped code/artifact capabilities.
## Verification evidence
- Focused selector/chat tests: 61/61.
- Persona/agent/MCP/governance regression: 118/118.
- Windows/sandbox/CLI/MCP focused tests: 140/140.
- Confirmation, placeholder, persona, and CLI comprehensive contracts: 148/148.
- Production build: `npm run build:all` passed (package builds, server and web
typecheck, Vite production bundle).
- Full root Vitest run: 8,839 passed, 2 skipped, with remaining failures from
two CLI package-install/REPL tests timing out at 120s/180s in this Windows
checkout. Those timeouts need a separate launcher/runtime investigation.
## Tool-context result
Before this branch, a live general-purpose request exposed 78 tool definitions
(about 41.8k schema characters); coder exposed 36 (about 17.7k). Repeated
turns re-sent roughly 26.6k tokens of schemas. The route now selects and
serializes at most 14 tools / 8k schema characters, with a p95 selector test
under 10ms on a 78-tool synthetic pool.
## External-agent smoke result
Installed local CLIs detected: Claude Code 2.1.211, Codex 0.144.1, Hermes
0.18.2, Cursor 3.1.15, OpenClaw 2026.6.11, and Gemini. Claude and Codex
headless smoke paths ran; Hermes/OpenClaw were not run against persistent user
state. Launcher mappings still need runtime fixes for Hermes (`hermes chat -q`),
OpenClaw (`openclaw agent --message`), Codex detached/resume behavior, and
Cursor workspace binding.
## Open P0/P1 blockers
1. Solo/Tauri packaging still sets `WAGGLE_SKIP_LITELLM=1`; the built-in proxy
is Anthropic-only, while model resolution can still select a keyed
non-Anthropic provider. No clean Windows install has proven models, proxy,
and routing with Node/Python/Ollama/Docker absent.
2. There is no bundled local chat model/runner/weights; the current “local”
fallback can select a cloud Ollama tag (`ollama/minimax-m2.7:cloud`).
3. LiteLLM/Python bundling and resource layout are not yet wired as a tested
self-contained Solo artifact; embeddings may still download at runtime.
4. Smart router behavior remains heuristic and is not yet evaluated against
legal/payroll/destructive/verification adversarial cases.
5. A full browser E2E pass and three fresh runs for each of the ten requested
personas have not completed; therefore no 9.5/10 persona score is claimed.
6. Full root test timeouts in CLI package-install/REPL paths remain open.
## Persona acceptance matrix to run next
Run three fresh sessions each for general-purpose, researcher, writer,
project-manager, executive-assistant, finance-owner, coder, data-engineer,
verifier, and coordinator. Score artifact correctness, conversation quality,
tool/provenance fidelity, memory, safety, recovery, UX/latency, and accessibility
using the 100-point rubric. Auto-fail fabricated memory, unapproved mutation,
read-only writes, secret/sandbox escape, corruption, hangs, or false tool claims.

View File

@@ -1,5 +1,10 @@
# Waggle V1 Pre-Production Qualification — COMPLETE
> **Historical snapshot — superseded.** The “then ship” and “You can ship”
> statements below record an older audit and are not a current release verdict.
> Use [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the
> active Windows Solo gate.
**Date**: 2026-03-20
**Branch**: `phase8-wave-8f-ui-ux`
**Baseline**: 3,895 tests, 277 files, zero failures