moving
This commit is contained in:
@@ -1,5 +1,10 @@
|
||||
# Phase 1A: Feature Wave Audit — Built vs Not Built
|
||||
|
||||
> **Historical snapshot — superseded.** This report records evidence as of
|
||||
> 2026-03-20 and is not current launch authority. See
|
||||
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
|
||||
> Windows-first Solo scope, external gates, and recommendation.
|
||||
|
||||
**Date**: 2026-03-20
|
||||
**Auditor**: Claude (automated codebase cross-reference)
|
||||
**Plan document**: `docs/plans/2026-03-19-phase9-completion-plan.md`
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
# Phase 1B: Deployment, Phases & PM Features Audit
|
||||
|
||||
> **Historical snapshot — superseded.** This report records evidence as of
|
||||
> 2026-03-20 and is not current launch authority. See
|
||||
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
|
||||
> Windows-first Solo scope, external gates, and recommendation.
|
||||
|
||||
**Date**: 2026-03-20
|
||||
**Auditor**: Claude (automated analysis)
|
||||
**Scope**: Wave 9D (Deployment), Phase 7 (KVARK), Phase 8 status, PM Features (6)
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
# Phase 2: UX Audit — View-by-View Code Review
|
||||
|
||||
> **Historical snapshot — superseded.** This report records evidence as of
|
||||
> 2026-03-20 and is not current launch authority. See
|
||||
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
|
||||
> Windows-first Solo scope, external gates, and recommendation.
|
||||
|
||||
**Date:** 2026-03-20
|
||||
**Auditor:** Production Readiness Automation (Phase 2)
|
||||
**Scope:** All 7 views, sidebar, onboarding, Direction D compliance, emotional assessment
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
# Phase 3A: Agent Critical Path Quality Audit
|
||||
|
||||
> **Historical snapshot — superseded.** This report records evidence as of
|
||||
> 2026-03-20 and is not current launch authority. See
|
||||
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
|
||||
> Windows-first Solo scope, external gates, and recommendation.
|
||||
|
||||
**Auditor**: Production Readiness Review (automated)
|
||||
**Date**: 2026-03-20
|
||||
**Scope**: Agent loop, memory, vault, cron, connectors, sub-agents
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
# Phase 3B: Server & API Layer — Production Readiness Audit
|
||||
|
||||
> **Historical snapshot — superseded.** This report records evidence as of
|
||||
> 2026-03-20 and is not current launch authority. See
|
||||
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
|
||||
> Windows-first Solo scope, external gates, and recommendation.
|
||||
|
||||
**Date**: 2026-03-20
|
||||
**Scope**: `@waggle/server` — local server (Fastify), routes, SSE, WebSocket, KVARK client, security middleware
|
||||
**Auditor**: Claude Opus 4.6 (automated code review)
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
# Phase 3C: UI & Frontend Code Quality Audit
|
||||
|
||||
> **Historical snapshot — superseded.** This report records evidence as of
|
||||
> 2026-03-20 and is not current launch authority. See
|
||||
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
|
||||
> Windows-first Solo scope, external gates, and recommendation.
|
||||
|
||||
**Auditor**: Senior Engineer Code Review (automated)
|
||||
**Date**: 2026-03-20
|
||||
**Scope**: `app/src/` (Tauri desktop app) + `packages/ui/src/` (shared React component library)
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
# 04A Application Security Audit
|
||||
|
||||
> **Historical snapshot — superseded.** This report records evidence as of
|
||||
> 2026-03-20 and is not current launch authority. See
|
||||
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
|
||||
> Windows-first Solo scope, external gates, and recommendation.
|
||||
|
||||
**Date**: 2026-03-20
|
||||
**Auditor**: Automated (Claude Opus 4.6)
|
||||
**Scope**: Waggle desktop app + local server — CSP, vault crypto, agent tools, input validation, sessions, connectors, dangerous patterns
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
# 04B — Secret Scanning & Dependency Audit
|
||||
|
||||
> **Historical snapshot — superseded.** This report records evidence as of
|
||||
> 2026-03-20 and is not current launch authority. See
|
||||
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
|
||||
> Windows-first Solo scope, external gates, and recommendation.
|
||||
|
||||
**Auditor:** Claude Opus 4.6 (automated)
|
||||
**Date:** 2026-03-20
|
||||
**Scope:** Full codebase secret scan, git history review, npm dependency audit, .gitignore assessment
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
# 05 - TEST & COVERAGE REPORT
|
||||
|
||||
> **Historical snapshot — superseded.** This report records evidence as of
|
||||
> 2026-03-20 and is not current launch authority. See
|
||||
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
|
||||
> Windows-first Solo scope, external gates, and recommendation.
|
||||
|
||||
**Date:** 2026-03-20
|
||||
**Auditor:** Claude Opus 4.6 (automated, read-only)
|
||||
**Scope:** All packages, integration tests, E2E tests, visual regression tests
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
# Phase 6: Build & Deployment Readiness Report
|
||||
|
||||
> **Historical snapshot — superseded.** This report records evidence as of
|
||||
> 2026-03-20 and is not current launch authority. See
|
||||
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
|
||||
> Windows-first Solo scope, external gates, and recommendation.
|
||||
|
||||
**Date**: 2026-03-20
|
||||
**Auditor**: Claude Opus 4.6 (automated)
|
||||
**Scope**: TypeScript compilation, Vite build, Docker, Tauri, Render.com, npx launcher, CI/CD
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
# Issue Register — Waggle V1 Pre-Production Qualification
|
||||
|
||||
> **Historical snapshot — superseded.** This register records findings as of
|
||||
> 2026-03-20 and is not the current open-issue or launch authority. See
|
||||
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
|
||||
> Windows-first Solo scope, external gates, and recommendation.
|
||||
|
||||
Generated: 2026-03-20 | Branch: `phase8-wave-8f-ui-ux` | Tests: 3,895 passing
|
||||
|
||||
---
|
||||
|
||||
@@ -1,5 +1,10 @@
|
||||
# Confidence Matrix — Waggle V1 Pre-Production Qualification
|
||||
|
||||
> **Historical snapshot — superseded.** This score records evidence as of
|
||||
> 2026-03-20 and is not the current readiness score or launch authority. See
|
||||
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the current
|
||||
> Windows-first Solo scope, external gates, and recommendation.
|
||||
|
||||
Generated: 2026-03-20 | Branch: `phase8-wave-8f-ui-ux`
|
||||
|
||||
---
|
||||
|
||||
@@ -1,143 +1,109 @@
|
||||
# Launch Recommendation — Waggle V1
|
||||
# Launch Recommendation — Windows Solo
|
||||
|
||||
Generated: 2026-03-20
|
||||
Updated: 2026-08-27
|
||||
|
||||
---
|
||||
## Current verdict: INSTALLER/LIFECYCLE INTERNAL RC QUALIFIED; RELEASE QUALIFICATION INCOMPLETE
|
||||
|
||||
## Recommendation: CONDITIONAL GO
|
||||
The supported launch scope is Windows Solo with Claude Code, Codex, and Hermes using
|
||||
their official user-owned installations and authentication. Cursor, OpenClaw, and macOS
|
||||
packaging/certification remain roadmap work and do not block this internal RC.
|
||||
|
||||
Ship after fixing the 8 CRITICAL issues (~6.5 hours of work). The HIGH issues are important but can be addressed in a rapid V1.0.1 patch within the first week.
|
||||
This document is the release-status authority. Historical receipts are evidence only;
|
||||
they do not certify a later behavior-changing revision unless a bounded no-impact review
|
||||
explicitly says so.
|
||||
|
||||
---
|
||||
## Frozen runtime candidate
|
||||
|
||||
## Executive Summary
|
||||
- Source revision: `23ad3fa5f99bddce648b84750a41365299aeb0da`
|
||||
- Source tree: `aab548f77ec64b181086664dad29c32e6bc78779`
|
||||
- Integration: private Waggle PR #66, tested head
|
||||
`587e259166db69ff86e806fa8393a5f8974ea0a1`, merged 2026-08-27
|
||||
- Tree equivalence: the tested PR head and merge commit resolve to the same source tree
|
||||
- Repository state after merge: private `main` equals `origin/main`
|
||||
|
||||
Waggle is a substantial, well-architected product with 3,895 passing tests, 53 agent tools, 29 connectors, 15K+ marketplace packages, and a complete feature set covering all 8 Kill List use cases. The agent loop, memory system, and vault encryption are architecturally sound. The UI underwent a recent Phase 10 rewrite that brought Tailwind adoption and Direction D palette cleanup.
|
||||
The documentation-only descendant that updates this record does not replace the runtime
|
||||
candidate. Before it is merged, its diff must be limited to documentation and all required
|
||||
remote checks must remain green.
|
||||
|
||||
However, the audit uncovered **8 CRITICAL issues** (5 security, 1 stability, 2 UX) that must be fixed before any external user touches the product. The most severe: CORS is wide open (any website can call your localhost APIs), there are no React error boundaries (one render error = permanent white screen), and the streaming loading indicator is invisible (users can't tell the agent is thinking).
|
||||
## Exact-current Windows installer
|
||||
|
||||
The good news: every CRITICAL fix is straightforward. Total estimated effort is 6.5 hours. None require architectural changes.
|
||||
- NSIS artifact:
|
||||
`app/src-tauri/target/x86_64-pc-windows-msvc/release/bundle/nsis/Waggle_0.2.0_x64-setup.exe`
|
||||
- Size: 102,923,936 bytes
|
||||
- SHA-256: `7BFA9F9B13633A51CD3336B42E3EF904B7F7A568C6DEE4F6CED967CBD4F40A59`
|
||||
- Internal signer: `CN=Egzakta Internal Pilot, O=Egzakta Group, C=RS`
|
||||
- Signer thumbprint: `E2F028541E7A4D1FE80FFFF02079060D36579846`
|
||||
- RFC 3161 timestamp authority: DigiCert SHA256 RSA4096 Timestamp Responder 2025 1
|
||||
- Trust classification: internal pilot only; the self-signed root is not public trust
|
||||
|
||||
---
|
||||
Clean-profile certification passed **64/64** checks in 450.199 seconds:
|
||||
|
||||
## What's Ready (Strengths)
|
||||
- Receipt:
|
||||
`output/installer-certification/23ad3fa5-20260827T123917Z-exact-main-clean-profile/windows-installer-certification.json`
|
||||
- Receipt SHA-256:
|
||||
`AC2A1C54119E28CC22DA931EB43E2815862B01832DD6097CE03F8F79D9D3DF4D`
|
||||
- Receipt source and bundled-sidecar revision: exact `23ad3fa5`
|
||||
- Tier: FREE/Solo
|
||||
- Managed model: `qwen2.5:0.5b`
|
||||
- Managed-model digest:
|
||||
`sha256:a8b0c51577010a279d933d14c2a8ab4b268079d44c5c8830c0a93900f1827c67`
|
||||
|
||||
1. **Solid agent core** — 53 tools, loop guards, injection scanning, approval gates, sub-agent orchestration. 1,272 tests on the agent package alone.
|
||||
The receipt proves silent install, bundled Node/npm and offline package execution,
|
||||
first boot, in-process embeddings, built-in proxy/session authentication, workspace and
|
||||
memory persistence, managed runtime/model pull and chat, proxy-restart chat, same-version
|
||||
repair, relaunch, data preservation, Exit/owned-process cleanup, uninstall, registry
|
||||
cleanup, and preservation of external `.hive-mind` and `.ollama` roots. It also proves
|
||||
that developer Node.js, Python, Docker, external LiteLLM, and a separately installed
|
||||
Ollama are not prerequisites.
|
||||
|
||||
2. **Complete feature set** — All 8 Kill List use cases work. Workspace memory, connectors, marketplace, personas, cron, swarm protocol, capability packs, onboarding with memory import.
|
||||
## Integrated test and security evidence
|
||||
|
||||
3. **Good test coverage** — 3,895 tests across 277 files, zero failures. Every major package has dedicated test suites with behavior-focused assertions and realistic mocks.
|
||||
- PR #66: every blocking remote check passed — primary CI, Playwright smoke and full
|
||||
E2E, Windows and both macOS Tauri verification targets, Wave 1, and Hive Mind
|
||||
install/smoke on Windows, Ubuntu, and macOS.
|
||||
- Exact-current dependency audits: 0 Critical and 0 High in both full and production
|
||||
dependency trees. Lower-severity maintenance remains tracked.
|
||||
- The security-hardening integration covers workspace path/link boundaries, hook/database
|
||||
hard-link boundaries, normalized ingress, atomic consolidation/cognify, deprecated-frame
|
||||
search exclusion before limits, and knowledge-graph provenance.
|
||||
- Hosted signing policy and workflow tests remain green, but no publicly trusted hosted
|
||||
artifact has been produced.
|
||||
|
||||
4. **Security fundamentals** — AES-256-GCM vault, parameterized SQL everywhere, path traversal protection, CLI allowlists, DOMPurify HTML sanitization, SecurityGate for marketplace.
|
||||
## Persona, router, and authentication evidence
|
||||
|
||||
5. **Deployment infrastructure** — Tauri Windows installer built (8.2MB), Docker production compose, Render.com blueprint, GitHub Actions CI/release pipeline.
|
||||
The historical ten-persona collection at `4c712ff6` contains 30/30 results at or above
|
||||
95/100 after documented independent semantic adjudication. It is not relabeled as an
|
||||
exact-current deterministic seal: PR #66 changed memory behavior, so public release
|
||||
qualification requires either a fresh exact-candidate collection or an explicit bounded
|
||||
semantic-impact attestation.
|
||||
|
||||
6. **Product polish** — 8 personas, dark/light mode, keyboard shortcuts, global search, workspace hue colors, onboarding wizard, tool card transparency, approval gates inline in chat.
|
||||
Smart-router primary, compact-tool-context, durable-budget/fallback, and official-user-auth
|
||||
canaries for Claude Code, Codex, and Hermes remain scoped historical evidence. No PR #66
|
||||
change altered provider credential ownership or copied/read provider credential files.
|
||||
|
||||
---
|
||||
## Hive Mind repository state
|
||||
|
||||
## Must Fix Before Launch (CRITICAL — ~6.5 hours)
|
||||
Curated public-mirror hardening PR #53 merged to `marolinik/hive-mind` `master` as
|
||||
`3410327800db3ea23f875d547a0c7f4d08826b7e`; Linux, Windows, macOS, and Ubuntu
|
||||
first-run smoke passed. The immutable drift checker still reports reviewed blockers and
|
||||
one unreviewed difference, with zero forbidden exports. Therefore the Windows Solo RC is
|
||||
not blocked, but the next Hive Mind package release remains a separate maintainer-curated
|
||||
operation. Raw subtree publication remains forbidden.
|
||||
|
||||
### Security (4 hours)
|
||||
## Public GO blockers
|
||||
|
||||
| # | Issue | Fix | Time |
|
||||
|---|-------|-----|------|
|
||||
| 1 | **CORS allows any origin** — any website can call all Waggle APIs | Change `origin: true` to `origin: ['http://localhost:1420', 'tauri://localhost']` (or your Tauri webview origins). Fix SSE hijack endpoints to use the same allowlist. | 1.5 hr |
|
||||
| 2 | **Server CSP has `unsafe-eval` + `unsafe-inline`** | Remove both. If scripts break, use nonces or hashes instead. | 30 min |
|
||||
| 3 | **OAuth refresh tokens stored plaintext** | Encrypt refresh tokens the same way access tokens are encrypted in `setConnectorCredential()`. | 1 hr |
|
||||
| 4 | **Verify API key revoked** | Go to Anthropic dashboard, confirm the key from commit `c29d75f` is revoked. Delete local branch `phase6-capability-truth`. | 30 min |
|
||||
Public release may be called **GO** only after all of these are closed for the approved
|
||||
release-tag commit:
|
||||
|
||||
### Stability (2 hours)
|
||||
1. A protected hosted build produces a publicly trusted Authenticode artifact.
|
||||
2. The managed Codex Security workflow produces a sealed Deep Security report with no
|
||||
unresolved Critical or High findings.
|
||||
3. Current persona qualification is sealed by a fresh exact-candidate receipt or an
|
||||
independently reviewed bounded semantic-impact attestation.
|
||||
4. Smart-router and Claude Code/Codex/Hermes official-auth qualification is either rerun
|
||||
on the exact candidate or covered by a concrete independently reviewed no-impact
|
||||
attestation.
|
||||
5. Protected release-tag checks are green and the exact artifact hashes are recorded.
|
||||
|
||||
| # | Issue | Fix | Time |
|
||||
|---|-------|-----|------|
|
||||
| 5 | **Zero error boundaries** | Add `<ErrorBoundary>` wrapping each view in App.tsx, plus one at the app root. Use react-error-boundary or a simple class component. Show "Something went wrong" with a retry button. | 2 hr |
|
||||
|
||||
### UX (30 minutes)
|
||||
|
||||
| # | Issue | Fix | Time |
|
||||
|---|-------|-----|------|
|
||||
| 6 | **Streaming indicator invisible** | The loading dots use BEM CSS classes with no definitions. Either add the CSS or replace with Tailwind `animate-pulse` dots. | 15 min |
|
||||
| 7 | **SplashScreen wrong palette** | Replace `#1a1a2e`/`#16213e`/`#0f3460` with Direction D tokens. Change `#f5a623` to `#d4a843`. | 15 min |
|
||||
|
||||
---
|
||||
|
||||
## Ship-Week Fixes (HIGH — ~28 hours, V1.0.1)
|
||||
|
||||
**Security hardening (first 2 days):**
|
||||
- Approval gates: change auto-approve to auto-deny on 5min timeout (15 min)
|
||||
- WebSocket authentication: require session token on `/ws` connect (2 hr)
|
||||
- Team WebSocket: validate JWT instead of trusting userId param (2 hr)
|
||||
- Replace `xlsx` with `exceljs` to fix prototype pollution (2 hr)
|
||||
- Generate Tauri updater keypair and set pubkey (30 min)
|
||||
|
||||
**Agent loop safety (day 3):**
|
||||
- Cap rate-limit retries (max 3, then fail gracefully) (1 hr)
|
||||
- Add token budget enforcement with configurable limit (2 hr)
|
||||
- Parameterize sqlite-vec SQL interpolation (30 min)
|
||||
|
||||
**Frontend stability (days 3-5):**
|
||||
- Add code splitting with `React.lazy()` for 7 views (2 hr)
|
||||
- Deduplicate SSE connections (1 hr)
|
||||
- Fix eventBus.removeAllListeners to scope per-client (1 hr)
|
||||
- Fix light theme breakage across components (2 hr)
|
||||
|
||||
**Build fixes (day 5):**
|
||||
- Fix npx waggle: compile .ts entry, resolve workspace deps (2 hr)
|
||||
- Add non-root user to Docker (30 min)
|
||||
- Fix CI branch target master→main (15 min)
|
||||
- Clean up 87 TypeScript errors (2 hr)
|
||||
|
||||
---
|
||||
|
||||
## Known Limitations (Ship Anyway)
|
||||
|
||||
These are acceptable for V1 and can be improved iteratively:
|
||||
|
||||
1. **No browser E2E tests** — Unit/integration coverage is strong (3,895 tests). True browser automation (Playwright user journeys) is a V1.1 investment. Screenshot baselines exist.
|
||||
|
||||
2. **Monolithic App.tsx (1300 lines)** — Works but hard to maintain. Refactoring into feature-specific providers is a V1.1 task that won't affect users.
|
||||
|
||||
3. **No React.memo optimization** — The app performs fine at current scale. Memoization is premature optimization until profiling shows problems.
|
||||
|
||||
4. **macOS build not configured** — DMG, code signing, notarization require an Apple Developer account. Windows installer works. Ship Windows-first, add macOS in V1.1.
|
||||
|
||||
5. **Direction D at ~78%** — The Phase 10 UI rewrite made massive progress (371→19 inline styles). Remaining 22% is polish, not broken functionality.
|
||||
|
||||
6. **KVARK client not wired** — KVARK integration (Phase 7) is library code + tests. Not wired into the running server because KVARK itself needs its HTTP API deployed first. This is expected — it's the Enterprise tier path.
|
||||
|
||||
7. **Conversation history unbounded** — At typical usage (10-50 turns/session), this isn't a problem. Add context window management for power users in V1.1.
|
||||
|
||||
---
|
||||
|
||||
## Post-Launch Priority Queue
|
||||
|
||||
### First Week (V1.0.1)
|
||||
1. All HIGH security fixes (approval timeout, WebSocket auth, xlsx, updater pubkey)
|
||||
2. Agent loop safety (retry cap, token budget)
|
||||
3. Frontend stability (code splitting, SSE dedup, error boundaries for remaining components)
|
||||
4. Light theme fixes
|
||||
|
||||
### First Month (V1.1)
|
||||
1. Browser E2E test suite (Playwright user journeys)
|
||||
2. React component rendering tests
|
||||
3. App.tsx decomposition (extract providers/hooks)
|
||||
4. macOS build + code signing
|
||||
5. npx waggle publishable package
|
||||
6. CI pipeline expansion (Docker, lint, security scan)
|
||||
7. Direction D compliance to 95%+
|
||||
|
||||
### First Quarter (V1.2)
|
||||
1. KVARK server-side wiring (when KVARK HTTP API ready)
|
||||
2. Performance profiling + React.memo optimization
|
||||
3. Context window management for long conversations
|
||||
4. Token budget UI (user-configurable spend limits)
|
||||
5. Full accessibility audit (WCAG 2.1 AA)
|
||||
|
||||
---
|
||||
|
||||
## Verdict
|
||||
|
||||
**CONDITIONAL GO** — Fix the 8 CRITICALs (6.5 hours), then ship. The product is feature-complete, well-tested, and architecturally sound. The critical issues are configuration mistakes, not design flaws. Every fix is surgical and low-risk.
|
||||
|
||||
The foundation is strong. Ship it.
|
||||
Until then, the installer is suitable for controlled internal testing, not public
|
||||
distribution, and Waggle must not be described as publicly production-ready or GO.
|
||||
|
||||
73
docs/production-readiness/10-SECURITY_REVIEW_2026-08-11.md
Normal file
73
docs/production-readiness/10-SECURITY_REVIEW_2026-08-11.md
Normal file
@@ -0,0 +1,73 @@
|
||||
# Windows Solo security review — refreshed 2026-08-27
|
||||
|
||||
## Status
|
||||
|
||||
The integrated Windows Solo source candidate is
|
||||
`23ad3fa5f99bddce648b84750a41365299aeb0da` (private PR #66). Its tested PR
|
||||
head and merge commit have the same source tree
|
||||
`aab548f77ec64b181086664dad29c32e6bc78779`.
|
||||
|
||||
This is an evidence-backed source, dependency, CI, and installed-runtime review. It is
|
||||
**not** a substitute for the still-missing sealed managed Deep Security report and does
|
||||
not confer public release approval.
|
||||
|
||||
## Verified controls
|
||||
|
||||
| Surface | Current evidence | Result |
|
||||
|---|---|---|
|
||||
| Workspace paths and lifecycle | Strict workspace IDs, canonical containment, identity-matching configs, link/hard-link rejection, and fail-closed list/get/delete behavior; focused executable tests | Pass |
|
||||
| Hook/database boundary | Existing database/config entries and SQLite companions reject links or hard links while preserving first-use and WAL behavior | Pass |
|
||||
| External memory ingress | Normalization and injection scanning are applied before persistence across supported ingress paths; bypass-focused regressions are included | Pass |
|
||||
| Search | Punctuated identifiers use bounded precise fallback; deprecated frames are excluded before keyword, LIKE, whole-vector, and chunk-vector limits; alternate-lane starvation regressions are covered | Pass |
|
||||
| Consolidation/cognify | Supersede operations are atomic, preserve provenance, and fail closed on partial mutation | Pass |
|
||||
| Knowledge graph | Relationship provenance and source-frame boundaries are preserved and tested | Pass |
|
||||
| Local server and tools | Loopback/session authentication, origin controls, SSRF/DNS/socket-pinning defenses, bounded external input, command-vector execution, and fail-closed shim handling are covered by focused and remote gates | Pass |
|
||||
| Provider authentication | Historical scoped canaries show Claude Code, Codex, and Hermes using official user-owned authentication with no provider credential-file reads/copies; exact-current carry-forward still needs a concrete no-impact attestation or rerun | Historical evidence; current qualification open |
|
||||
| Packaged runtime | Exact-current internal-pilot NSIS passed 64/64 clean-profile install, boot, managed-model, repair, relaunch, Exit/cleanup, and uninstall checks | Pass internal RC |
|
||||
| Dependency severity | Exact-current full and production audits contain 0 Critical and 0 High findings | Pass Critical/High gate |
|
||||
|
||||
## Exact installer evidence
|
||||
|
||||
- Installer SHA-256:
|
||||
`7BFA9F9B13633A51CD3336B42E3EF904B7F7A568C6DEE4F6CED967CBD4F40A59`
|
||||
- Certification receipt:
|
||||
`output/installer-certification/23ad3fa5-20260827T123917Z-exact-main-clean-profile/windows-installer-certification.json`
|
||||
- Receipt SHA-256:
|
||||
`AC2A1C54119E28CC22DA931EB43E2815862B01832DD6097CE03F8F79D9D3DF4D`
|
||||
- Managed model: `qwen2.5:0.5b`
|
||||
- Managed-model digest:
|
||||
`sha256:a8b0c51577010a279d933d14c2a8ab4b268079d44c5c8830c0a93900f1827c67`
|
||||
- Certification checks: 64 passed, 0 failed
|
||||
|
||||
The certifier verified source and sidecar provenance, bundled runtime/npm, clean offline
|
||||
execution, default Solo onboarding, in-process embeddings, built-in proxy liveness,
|
||||
session authentication, managed-model pull/chat, proxy-restart chat, repair and data
|
||||
preservation, relaunch, cleanup/uninstall, and unchanged external `.hive-mind`/`.ollama`
|
||||
roots. No Waggle-owned process or certificate test profile remained after completion.
|
||||
|
||||
## Remote integration evidence
|
||||
|
||||
PR #66 passed primary CI, Playwright smoke and full E2E, Windows and both macOS Tauri
|
||||
verification targets, Wave 1, and Hive Mind install/smoke on Windows, Ubuntu, and macOS.
|
||||
The full local Waggle Vitest suite, agent/server/app typechecks, lint, and diff checks also
|
||||
completed successfully before integration.
|
||||
|
||||
Hive Mind PR #53 passed Linux, Windows, macOS, and Ubuntu first-run smoke before merge as
|
||||
`3410327800db3ea23f875d547a0c7f4d08826b7e`.
|
||||
|
||||
## Residual risk and public release blockers
|
||||
|
||||
- The installer is signed by `CN=Egzakta Internal Pilot`, a private self-signed identity.
|
||||
Its DigiCert timestamp validates the signing pipeline but does not provide public trust.
|
||||
- No sealed managed Codex Security report exists; this audit host used a disabled
|
||||
permission profile. No failed or unsealed attempt is interpreted as a no-findings result.
|
||||
- The immutable Hive Mind drift baseline reports 22 known reviewed blockers and one
|
||||
unreviewed difference, with zero forbidden exports. These block the next OSS package
|
||||
release, not this private Windows Solo internal RC.
|
||||
- Current persona qualification still needs a fresh exact-candidate seal or an independent
|
||||
bounded semantic-impact attestation because PR #66 changed memory behavior.
|
||||
|
||||
Public GO requires publicly trusted Authenticode, a sealed exact-candidate managed Deep
|
||||
Security report with no unresolved Critical/High findings, current persona qualification,
|
||||
fresh or explicitly attested smart-router and official-auth qualification, and green
|
||||
protected release-tag checks.
|
||||
@@ -0,0 +1,115 @@
|
||||
# Hive Mind mirror parity audit — refreshed 2026-08-27
|
||||
|
||||
## Verdict
|
||||
|
||||
The curated Hive Mind mirror does not block the private Windows Solo internal RC. The
|
||||
Waggle monorepo remains the sole source of truth for the memory substrate.
|
||||
|
||||
The next public Hive Mind package release is **not yet approved**. The mirror is clean,
|
||||
merged, and cross-platform green, but the immutable drift baseline still contains known
|
||||
reviewed blockers and one unreviewed difference. It must not be described as fully
|
||||
synchronized or release-ready.
|
||||
|
||||
## Exact inputs
|
||||
|
||||
- Waggle runtime/source candidate:
|
||||
`23ad3fa5f99bddce648b84750a41365299aeb0da`
|
||||
- Waggle source tree:
|
||||
`aab548f77ec64b181086664dad29c32e6bc78779`
|
||||
- Hive Mind `master`:
|
||||
`3410327800db3ea23f875d547a0c7f4d08826b7e`
|
||||
- Hive Mind integration: PR #53, tested head
|
||||
`f5072efbf3f2d13247bb91505d99403acbdf11c6`
|
||||
- Comparison:
|
||||
`node scripts/oss-drift-check.mjs D:/Projects/hive-mind`
|
||||
- Result: exit 1, fail closed
|
||||
|
||||
The Waggle documentation-only descendant does not change either compared source tree.
|
||||
|
||||
## Immutable baseline result
|
||||
|
||||
| Classification | Count | Meaning |
|
||||
|---|---:|---|
|
||||
| Parity | 5 | Reviewed paths with the expected canonical/OSS state |
|
||||
| Reviewed adaptations | 38 | Intentional layout, import, logger, branding, or OSS-architecture differences |
|
||||
| Known reviewed blockers | 22 | Paths still requiring maintainer-curated reconciliation before an OSS release |
|
||||
| Unreviewed differences | 1 | A changed path not yet classified against the baseline |
|
||||
| Forbidden exports | 0 | No prohibited Waggle-only path or marker was found in the mirror |
|
||||
|
||||
The single unreviewed difference is `harvest/raw-turns.ts`. It must be classified and
|
||||
either reconciled or deliberately re-baselined through maintainer review before release.
|
||||
|
||||
The five parity paths are:
|
||||
|
||||
- `harvest/claude-code-adapter.ts`
|
||||
- `harvest/decision-derivation.ts`
|
||||
- `harvest/stable-id.ts`
|
||||
- `mind/erasure.ts`
|
||||
- `mind/inprocess-reranker.ts`
|
||||
|
||||
## Known reviewed blockers
|
||||
|
||||
The baseline currently marks 19 paths `reconcile`:
|
||||
|
||||
- Harvest: `dedup.ts`, `extract-kg-entities.ts`, `pipeline.ts`,
|
||||
`url-adapter.ts`
|
||||
- Runtime/API: `index.ts`, `workspace-manager.ts`, `multi-mind-cache.ts`
|
||||
- Mind: `db.ts`, `embedding-provider.ts`, `frames.ts`, `identity.ts`,
|
||||
`inprocess-embedder.ts`, `knowledge.ts`, `llm-extractor.ts`,
|
||||
`raw-archive.ts`, `raw-detail-lane.ts`, `schema.ts`, `search.ts`, and
|
||||
`suppression.ts`
|
||||
|
||||
Three more blockers are marked `product-curation`: `hook-runtime.ts`,
|
||||
`mind/supersede.ts`, and `multi-mind.ts`. They require an explicit reviewed decision
|
||||
about what remains Waggle-only and what generic subset, if any, should be adapted for the
|
||||
public mirror. The approved decision must then reclassify them out of the blocker category.
|
||||
|
||||
A blocker in this inventory means the mapped files are not yet approved for the next OSS
|
||||
release; it does not by itself assert a runtime vulnerability. Every blocker needs
|
||||
canonical-first review. Only reconcile/exported paths require public-layout adaptation and
|
||||
focused port tests; product-curation paths may instead remain Waggle-only after explicit
|
||||
review. Every resolution requires a new baseline receipt.
|
||||
|
||||
## Work integrated on 2026-08-27
|
||||
|
||||
Hive Mind PR #53 added the curated search-boundary hardening needed for punctuated
|
||||
identifiers and deprecated-candidate filtering before keyword, LIKE, whole-vector, and
|
||||
chunk-vector limits. The implementation closed concatenated-token, malformed-query, CJK,
|
||||
and alternate-vector-lane bypasses with executable regressions.
|
||||
|
||||
The PR passed build/test on Ubuntu, Windows, and macOS plus Ubuntu first-run smoke before
|
||||
merge. A transient Ubuntu dependency-download `ECONNRESET` was rerun once and passed;
|
||||
the deterministic code/test result was green.
|
||||
|
||||
This narrows the drift but does not erase the remaining `mind/search.ts` curated
|
||||
difference, so the immutable checker correctly continues to fail closed.
|
||||
|
||||
## Proprietary exclusion boundary
|
||||
|
||||
Raw subtree publication is forbidden. A release curation must continue to exclude:
|
||||
|
||||
- `vault.ts`
|
||||
- `mind/evolution-runs.ts`
|
||||
- `mind/execution-traces.ts`
|
||||
- `mind/improvement-signals.ts`
|
||||
- `compliance/**`
|
||||
- Waggle-only `install_audit` schema/migration fragments interleaved in shared files
|
||||
|
||||
The current scan found zero forbidden exports. That is a necessary safety gate, not proof
|
||||
that all public-worthy changes have been reconciled.
|
||||
|
||||
## Release rule
|
||||
|
||||
Before the next Hive Mind package release:
|
||||
|
||||
1. Classify `harvest/raw-turns.ts`.
|
||||
2. Reconcile or explicitly re-baseline every known blocker in small reviewed phases.
|
||||
3. Run focused tests for each curated port and the complete Hive Mind build/test suite.
|
||||
4. Run Linux, Windows, macOS, and first-run package smoke.
|
||||
5. Re-run the immutable drift checker and require exit 0: zero known blockers, zero
|
||||
unreviewed differences, and zero forbidden exports. Any approved product-curation
|
||||
decision must first be represented by reviewed baseline reclassification.
|
||||
6. Preserve provenance showing that generic substrate work originated in Waggle first.
|
||||
|
||||
Until those gates pass, `master` is a clean integrated development baseline, not a sealed
|
||||
new OSS package release.
|
||||
@@ -1,5 +1,11 @@
|
||||
# Waggle OS — Production Readiness Audit (2026-07-03)
|
||||
|
||||
> [!CAUTION]
|
||||
> **HISTORICAL SNAPSHOT — NOT CURRENT SHIP AUTHORITY.** Findings and grades
|
||||
> below describe the cited July baseline. Use
|
||||
> `docs/production-readiness/09-LAUNCH_RECOMMENDATION.md` for current scope,
|
||||
> evidence, open gates, and release verdict.
|
||||
|
||||
**Method:** 10 parallel principal-engineer audit lanes (opus, high effort), each required to cite `file:line` evidence it actually read. Baseline at audit time: HEAD `78660ab5` on `main`, vitest 8063/8063 green, lint 0, tsc 0 (agent/server/app), `build:all` clean, git history secret-scan CLEAN (all key-shaped strings are `detectSecrets()` fixtures).
|
||||
|
||||
**Subsystem grades:** Build **D** · CI/CD **C** · Docs/DX **C** · Deps/Config **C** · Agent-runtime **C** · Security **B** · Testing **B** · Server-API **B** · Frontend **B** · Memory-substrate **B**.
|
||||
|
||||
@@ -1,5 +1,11 @@
|
||||
# Waggle OS — Production Readiness Sign-off (2026-07-03)
|
||||
|
||||
> **Historical snapshot — superseded.** This document preserves the 2026-07-03
|
||||
> assessment but is not current ship authority. Its “production-ready” and
|
||||
> signing-only conclusions no longer apply. Use
|
||||
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the active
|
||||
> Windows Solo gate and its exact-HEAD evidence requirements.
|
||||
|
||||
> Engineering-director pass driven by a 10-lane parallel audit, executed as prioritized remediation streams (Opus implementation agents, Fable orchestration/QA). Companion to [`2026-07-03-release-audit.md`](./2026-07-03-release-audit.md).
|
||||
|
||||
## Session commit ledger (17 commits on `main`, from baseline `a1fad4f8`)
|
||||
|
||||
85
docs/production-readiness/2026-07-16-audit.md
Normal file
85
docs/production-readiness/2026-07-16-audit.md
Normal file
@@ -0,0 +1,85 @@
|
||||
# Waggle OS production-readiness audit — 2026-07-16
|
||||
|
||||
> **Historical snapshot — superseded.** This document records the state on
|
||||
> 2026-07-16 and is not current ship authority. Use
|
||||
> [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the active
|
||||
> Windows Solo gate. Cursor and OpenClaw are now roadmap-only; macOS
|
||||
> certification is deferred.
|
||||
|
||||
## Executive verdict
|
||||
|
||||
The branch is not yet production-ready for the requested claim of a fully
|
||||
Docker-independent Solo install or a verified 9.5/10 result across ten fresh
|
||||
personas. It is materially improved and buildable: Windows sandbox/process
|
||||
handling is hardened, chat schemas are bounded before model serialization, and
|
||||
artifact/CLI writes now require approval.
|
||||
|
||||
## Verified changes
|
||||
|
||||
- Windows workspace/path, environment, direct-argv `run_code`, descendant
|
||||
process termination, CLI `.cmd` discovery, MCP env sanitization, and MCP
|
||||
start/stop race handling were implemented.
|
||||
- Chat tool selection is deterministic and subtractive: maximum 14 tools and
|
||||
8,000 exact OpenAI/LiteLLM schema characters per turn; duplicate names keep
|
||||
the native definition; plugin/MCP tools cannot appear in casual chat unless
|
||||
explicitly relevant; governance-blocked schemas are removed before model
|
||||
serialization.
|
||||
- `cli_execute`, `multi_edit`, DOCX/XLSX/PPTX/PDF generation are approval-gated.
|
||||
- DOCX/PDF/PPTX/XLSX workspace containment uses `path.relative`, preventing
|
||||
sibling-prefix escapes such as `workspace-sibling`.
|
||||
- Coder, analyst, finance, data-engineer, executive-assistant, and general
|
||||
purpose persona allowlists include their shipped code/artifact capabilities.
|
||||
|
||||
## Verification evidence
|
||||
|
||||
- Focused selector/chat tests: 61/61.
|
||||
- Persona/agent/MCP/governance regression: 118/118.
|
||||
- Windows/sandbox/CLI/MCP focused tests: 140/140.
|
||||
- Confirmation, placeholder, persona, and CLI comprehensive contracts: 148/148.
|
||||
- Production build: `npm run build:all` passed (package builds, server and web
|
||||
typecheck, Vite production bundle).
|
||||
- Full root Vitest run: 8,839 passed, 2 skipped, with remaining failures from
|
||||
two CLI package-install/REPL tests timing out at 120s/180s in this Windows
|
||||
checkout. Those timeouts need a separate launcher/runtime investigation.
|
||||
|
||||
## Tool-context result
|
||||
|
||||
Before this branch, a live general-purpose request exposed 78 tool definitions
|
||||
(about 41.8k schema characters); coder exposed 36 (about 17.7k). Repeated
|
||||
turns re-sent roughly 26.6k tokens of schemas. The route now selects and
|
||||
serializes at most 14 tools / 8k schema characters, with a p95 selector test
|
||||
under 10ms on a 78-tool synthetic pool.
|
||||
|
||||
## External-agent smoke result
|
||||
|
||||
Installed local CLIs detected: Claude Code 2.1.211, Codex 0.144.1, Hermes
|
||||
0.18.2, Cursor 3.1.15, OpenClaw 2026.6.11, and Gemini. Claude and Codex
|
||||
headless smoke paths ran; Hermes/OpenClaw were not run against persistent user
|
||||
state. Launcher mappings still need runtime fixes for Hermes (`hermes chat -q`),
|
||||
OpenClaw (`openclaw agent --message`), Codex detached/resume behavior, and
|
||||
Cursor workspace binding.
|
||||
|
||||
## Open P0/P1 blockers
|
||||
|
||||
1. Solo/Tauri packaging still sets `WAGGLE_SKIP_LITELLM=1`; the built-in proxy
|
||||
is Anthropic-only, while model resolution can still select a keyed
|
||||
non-Anthropic provider. No clean Windows install has proven models, proxy,
|
||||
and routing with Node/Python/Ollama/Docker absent.
|
||||
2. There is no bundled local chat model/runner/weights; the current “local”
|
||||
fallback can select a cloud Ollama tag (`ollama/minimax-m2.7:cloud`).
|
||||
3. LiteLLM/Python bundling and resource layout are not yet wired as a tested
|
||||
self-contained Solo artifact; embeddings may still download at runtime.
|
||||
4. Smart router behavior remains heuristic and is not yet evaluated against
|
||||
legal/payroll/destructive/verification adversarial cases.
|
||||
5. A full browser E2E pass and three fresh runs for each of the ten requested
|
||||
personas have not completed; therefore no 9.5/10 persona score is claimed.
|
||||
6. Full root test timeouts in CLI package-install/REPL paths remain open.
|
||||
|
||||
## Persona acceptance matrix to run next
|
||||
|
||||
Run three fresh sessions each for general-purpose, researcher, writer,
|
||||
project-manager, executive-assistant, finance-owner, coder, data-engineer,
|
||||
verifier, and coordinator. Score artifact correctness, conversation quality,
|
||||
tool/provenance fidelity, memory, safety, recovery, UX/latency, and accessibility
|
||||
using the 100-point rubric. Auto-fail fabricated memory, unapproved mutation,
|
||||
read-only writes, secret/sandbox escape, corruption, hangs, or false tool claims.
|
||||
@@ -1,5 +1,10 @@
|
||||
# Waggle V1 Pre-Production Qualification — COMPLETE
|
||||
|
||||
> **Historical snapshot — superseded.** The “then ship” and “You can ship”
|
||||
> statements below record an older audit and are not a current release verdict.
|
||||
> Use [`09-LAUNCH_RECOMMENDATION.md`](./09-LAUNCH_RECOMMENDATION.md) for the
|
||||
> active Windows Solo gate.
|
||||
|
||||
**Date**: 2026-03-20
|
||||
**Branch**: `phase8-wave-8f-ui-ux`
|
||||
**Baseline**: 3,895 tests, 277 files, zero failures
|
||||
|
||||
Reference in New Issue
Block a user