19 KiB
FULL BACKLOG โ 2026-04-18
Purpose: Single surface of every open item across the polish sprint, consolidated backlog, PDF triage deferred items, Marko-side non-coding work, strategic decisions, and the newly identified GEPA wiring gaps. Merges POLISH-SPRINT-2026-04-18.md, BACKLOG-CONSOLIDATED-2026-04-17.md, and PDF-E2E-ISSUES-2026-04-17.md.
State at write-time: main @ 1c304cd, tree clean, 200 commits ahead of origin. Phase A of the polish sprint is 5/6 done; QW-3 remains.
Legend:
- โ DONE
- ๐ข PENDING (doable now)
- ๐ DEFERRED (needs design / bigger chunk)
- ๐ด BLOCKED (external โ Stripe / cert / Marko)
- โณ MARKO ACTION (non-engineering)
1. Polish Sprint 2026-04-18 โ phased plan (this week)
Phase A โ Quick Wins
| # | Item | Status | Commit |
|---|---|---|---|
| QW-1 | Prefill chat after onboarding | โ | 9d1c858 |
| QW-2 | Memory tab labels | โ | bb6ab50 |
| QW-3 | Skip boot on return visits (verify BOOT_KEY in Index.tsx:16) |
๐ข | โ |
| QW-4 | Back button onboarding 2-6 | โ | 47539ac |
| QW-5 | Dock tier rename + billing clarity | โ | 70c8d84 |
| CR-7 | CLAUDE.md ยง10 refresh | โ | 1c304cd |
Phase B โ Core bugs + light mode finish (~1 day)
| # | Item | Status |
|---|---|---|
| P35 | Spawn-agent "no models available" โ wire SpawnAgentPanel to live provider list (13 green) |
๐ข |
| P36 | Dock spawn-agent icon click โ verify, wire TaskCreate | ๐ข |
| P40 | BootScreen logo/animation renders in light mode | ๐ข |
| P41 | "Waggle AI" header text restyled for light theme | ๐ข |
| CR-2 | Remaining hive-950 โ semantic token sweep |
๐ข |
Phase C โ OW-6 PersonaSwitcher two-tier (0.5 day)
| # | Item | Status |
|---|---|---|
| OW-6 | UNIVERSAL MODES (8) + WORKSPACE SPECIALISTS split; hover tooltip with tagline / bestFor / wontDo. File: apps/web/src/components/os/overlays/PersonaSwitcher.tsx. Requires AgentPersona interface extensions per CLAUDE.md ยง5 (already shipped in personas.ts). |
๐ข |
Phase D โ Feature polish (~10 days)
Compliance UX (3.5d) โ Block 3b
| # | Task |
|---|---|
| 3b.1 | POST /api/compliance/export-pdf โ pdfmake buffer download |
| 3b.2 | Template system (sections, logo, branding, footer as JSON) |
| 3b.3 | Full-page ComplianceReport viewer + date picker + PDF button |
| 3b.4 | Custom branding (logo upload, org name, risk class override) |
| 3b.5 | KVARK template (IAM audit, data residency, department breakdown) |
Harvest UX (5d) โ Block 4
| # | Task | Status |
|---|---|---|
| 3.1 | Privacy headline | โ |
| 3.2 | Dedup summary | โ |
| 3.3 | SSE live progress streaming | ๐ข |
| 3.4 | Resumable harvests (checkpoint every 100 frames) | ๐ข |
| 3.5 | Identity auto-populate screen | ๐ข |
| 3.6 | Harvest-first onboarding tile | ๐ข |
Wiki v2 (5d) โ Block 3
| # | Task | Status |
|---|---|---|
| 2.1 | Markdown export | โ |
| 2.2 | Incremental recompile after harvest | ๐ข |
| 2.3 | Obsidian vault adapter | ๐ข |
| 2.4 | Notion structured export adapter | ๐ข |
| 2.5 | Wiki health report dashboard UI | ๐ข |
Medium UX fixes (1-4h each)
| # | Fix | Status |
|---|---|---|
| UX-1 | Reduce onboarding decisions (default Blank + General Purpose โ Ready) | ๐ข |
| UX-3 | Memory tab bar labels | โ (QW-2) |
| UX-4 | Dock text labels first 7d / 20 sessions | ๐ข |
| UX-5 | Hide token/cost behind dev mode | ๐ข |
| UX-6 | Chat header overflow menu | ๐ข |
| UX-7 | Tier-step copy clarify dock tier โ billing | โ (QW-5) |
Engagement features (half-day each)
| # | Feature |
|---|---|
| ENG-1 | "I just remembered" toast after 5th message |
| ENG-2 | WorkspaceBriefing collapsible sidebar |
| ENG-3 | Progressive dock unlock nudge at 10/50 sessions |
| ENG-4 | LoginBriefing on every launch (per-session + don't-show-again) |
| ENG-5 | Harvest-first onboarding โ move import pitch to step 2 |
| ENG-6 | Memory Score / Brain Health metric |
| ENG-7 | Suggested next actions after assistant response |
Responsive gaps
| # | Component | Issue |
|---|---|---|
| R-1 | Dock | Power tier (14 items) overflows < 768px |
| R-2 | StatusBar | 10+ items โ hide non-essential < 900px |
| R-3 | ChatApp | Session sidebar 192px โ collapse narrow |
| R-4 | OnboardingWizard | Template grid responsive columns |
| R-5 | AppWindow | Default sizes exceed mobile viewport |
Phase E โ Infra polish (~6 days)
| # | Item |
|---|---|
| CR-8 | Tauri binary verification on clean Windows VM |
| INST-1 | Ollama bundled installer (Install Ollama + pull Gemma 4) |
| INST-2 | Hardware scan (RAM/GPU โ model fit recommendation) |
| INST-3 | Ollama daemon auto-start (Windows service / macOS launchd) |
| CR-6 | hive-mind actual source extraction (scaffold exists, copy TODO) |
| CR-1 | MS Graph OAuth connector โ email / calendar / files harvest |
Phase F โ Content polish (~1 day)
| # | Item |
|---|---|
| CR-4 | Demo video script (90s + 5min) |
| CR-5 | LinkedIn launch posts (3-post sequence) |
| โ | Peer-reviewer outreach email (agent drafts, Marko sends) |
2. Marko โ non-coding items
| # | Action | Status | Blocks |
|---|---|---|---|
| M1 | ChatGPT export (OpenAI email) | โณ chase | Phase 1 harvest |
| M2 | Claude / Anthropic export | โ | โ |
| M3 | Google / Gemini export | โ | โ |
| M4 | Perplexity threads โ manual-only, skipped | โ | โ |
| M5 | API credit top-ups | โ | โ |
| M6 | Judge-model list revision (after w4/w25 proofs) | โณ later | Phase 5 |
| M7 | Stripe products (Pro $19, Teams $49/seat) | โณ today | Phase 7 |
| M8 | Windows EV code signing cert ($300-500/yr) | โณ Monday | Phase 7 |
| M9 | Apple Dev + Mac notarization | โณ Monday | Phase 7 |
| M10 | Greenlight launch date | โณ after proofs | Launch |
Strategic decisions pending
| # | Decision | Unlocks |
|---|---|---|
| C1 | hive-mind OSS timing โ ship-with or ship-before Waggle? | Launch sequence |
| C5 | Harvest-first onboarding โ replace step 2 vs parallel opt-in? | UX-1 / ENG-5 |
| C8 | Warm list 5-10 names to pre-email T-72h | Launch credibility |
| C9 | Papers โ single-author or dual-author? | Paper attribution |
| C11 | Marketplace model โ free / freemium / enterprise? | Skills monetization |
| ES | EvolveSchema attribution โ Mikhail vs ACE (Zhang et al.) | Paper 2 framing |
3. P0 Launch Blockers (beyond polish)
Block 2 โ Phase 1 Harvest Marko's real data (๐ด blocked on M1)
| # | Task |
|---|---|
| 1.1 | Import ChatGPT conversations โ harvest |
| 1.2 | Import Claude conversations โ harvest |
| 1.3 | Re-harvest Claude Code (fresh, all sessions) |
| 1.4 | Import Gemini conversations โ harvest |
| 1.5 | Import Perplexity threads โ harvest |
| 1.6 | Build Cursor adapter (0.5-1 day) |
| 1.7 | Post-harvest cognify on imported frames |
| 1.8 | Identity auto-populate from harvest |
| 1.9 | Wiki compile from real data |
| GATE | 10K-50K frames, dedup verified, KG populated |
Budget ~$50.
Block 5-8 โ proofs + papers
| Block | Name | Time | Budget |
|---|---|---|---|
| 5 | Phase 4 Memory Proof (MEMORY-HARVEST-TEST-PLAN.docx) |
10d | $300-500 |
| 6 | Phase 5 GEPA Full-System Proof (GEPA-EVOLUTION-TEST-PLAN.docx) |
18d | $1.5-2.5k |
| 7 | Phase 5b Combined Effect (COMBINED-EFFECT-TEST-PLAN.docx) |
6d | ~$500 |
| 8 | Phase 6 Write papers (2 arXiv) + Marko peer review | 5d | โ |
Block 9 โ Launch Prep
| # | Task | Status |
|---|---|---|
| 9.1 | Stripe dashboard + smoke test | ๐ด M7 |
| 9.2 | Code signing cert + updater keypair | ๐ด M8 |
| 9.3 | hive-mind source extraction (Apache 2.0) | ๐ข (scaffold done) |
| 9.4 | Binary build + clean Windows VM smoke | ๐ข |
| 9.5 | Clerk auth integration | ๐ด after 9.1 |
| 9.6 | Onboarding finalized (harvest-first) | ๐ข needs Block 4 |
| 9.7 | Mac notarization | โณ M9 |
| 9.8 | Landing page final polish | ๐ข |
Block 10 โ Launch Day ๐ด (gated)
Simultaneous: Waggle binary ยท hive-mind OSS ยท 2 arXiv papers ยท LinkedIn sequence ยท Pro/Teams live.
4. PDF E2E โ deferred 21 items (๐ )
Source: docs/plans/PDF-E2E-ISSUES-2026-04-17.md.
| # | Item | Effort |
|---|---|---|
| P4 | Permissions โ Mutation Gates merge with 3-level tool approval | ๐ big UX |
| P6 | Room feature โ verify 2 parallel agents visualization | ๐ |
| P8 | Agents vs Personas unify naming | ๐ก partial |
| P10 | Bee-style per-agent icons (dark + light) | ๐ design-heavy |
| P14 | Local browser only drive D โ multi-drive (C: required) | ๐ |
| P15 | Create Template modal overlaps Dashboard โ can't drag | ๐ |
| P16 | Files app local-folder create + explorer-style browse | ๐ big |
| P17 | App-wide tooltips on badges/options | ๐ broad |
| P18 | Waggle Dance โ display real discovery/handoff signals | ๐ |
| P21 | Timeline always empty โ wire to event stream | ๐ |
| P25 | Scheduled Jobs toggle stays off after trigger | ๐ |
| P26 | New scheduled-job creation unclear | ๐ |
| P28 | Marketplace empty | โ (fixed, 10 E2E green, 148 pkgs) |
| P29 | Skills & Apps cards not clickable โ no detail card | ๐ |
| P30 | MCP install CLI flow unclear | ๐ |
| P34 | Approvals app โ move to Ops or delete | ๐ |
| P35 | Spawn Agent "no models available" | ๐ โ Phase B above |
| P36 | Dock spawn-agent icon wiring | ๐ โ Phase B above |
| P39 | Status bar left shows static โ should be dynamic | ๐ก |
| P40 | Light-mode boot screen | ๐ โ Phase B above |
| P41 | Light-mode "Waggle AI" header text | ๐ โ Phase B above |
5. GEPA Wiring Closure โ NEW (4 items)
Context: Self-evolution library code is 100% present (357 evolution tests, full orchestrator, deploy callbacks, gates, compose, trace store, eval-dataset builder, makeRunningJudge, etc.). But four wiring gaps explain why a published evolution run hasn't produced a real agent-behavior improvement to date. Each is small but load-bearing.
G1 โ No autonomous evolution service / scheduler ๐ข
Claim: The server has optimizer-service.ts (the one-shot @waggle/optimizer wrapper) but no evolution-service.ts. There is no route that instantiates EvolutionOrchestrator on a schedule, no cron job that calls runOnce(), no daemon that mines traces into eval datasets. The full closed loop exists as library code that nothing automatically calls.
Evidence:
packages/server/src/local/services/โ containsoptimizer-service.ts, noevolution-service.ts.packages/server/src/local/routes/evolution.tsโ has/api/evolution/run(manual POST) that instantiates the orchestrator with a base judge + running judge, but it is only triggered by HTTP. The only cron reference is a comment on line 439:"backwards compat for tests + cron"โ no code.cron-service.tsexists in services but has no evolution-run registration.
Fix:
- Create
packages/server/src/local/services/evolution-service.tsthat owns a daemon loop and an auto-trigger policy. - Register an evolution cron in
cron-service.ts(configurable cadence, default daily at low-traffic hour) that callsrunOnce()with baseline auto-detection from the trace store. - Add a minimum-dataset gate so the scheduler skips runs when the trace table has fewer than N eligible examples (avoids burning API spend on no-op runs).
- Wire an on-demand trigger in the UI (Evolution tab โ Run now button already exists from Phase 8.5 โ ensure it reuses the same service).
Effort: 0.5-1 day.
G2 โ loadSystemPrompt ignores overrides ๐ข
Claim: prompt-loader.ts is a static file reader that reads {waggleDir}/system-prompt.md only. It does not integrate with loadBehavioralSpecOverrides or loadCustomPersonas. A deployed evolution writes overrides correctly via evolution-deploy.ts, but any consumer that reads the disk system prompt directly would not see those overrides.
Evidence:
packages/agent/src/prompt-loader.tsโ 26 lines total, onlyreadFileSyncofsystem-prompt.md. No override imports.- Override loaders live in
behavioral-spec.ts+custom-personas.ts, called byserver.activeBehavioralSpecdecorator (Phase 7.5) โ that chat path does work. - Gap: any other consumer (CLI, tests, future runtime integrations) that reads via
loadSystemPromptreceives the raw file without overrides.
Fix:
- Add
loadSystemPromptWithOverrides(waggleDir)that composes: base spec โ behavioral-spec overrides (viabuildActiveBehavioralSpec) โ persona system prompt (viagetPersona+ custom persona overrides) โ disk system-prompt.md append. - Migrate any remaining callers of
loadSystemPromptto the override-aware loader. - Deprecate the bare
loadSystemPrompt(keep export for test isolation only). - Add an assertion in agent-loop startup that logs a warning if overrides exist on disk but the active spec doesn't include them (catches wiring regressions).
Effort: 2-4 hours.
G3 โ Running judge not end-to-end on all eval paths ๐ข
Claim: Without makeRunningJudge, the judge compares prompt TEXT to expected output, turning GEPA into a prompt-text-similarity optimizer โ a meaningless gradient. The wrapper exists in evolution-llm-wiring.ts but not every runtime path assembles it.
Evidence:
/api/evolution/run(evolution.ts:377) โ CORRECTLY wrapsbaseJudgewithmakeRunningJudgefor GEPA instruction stage. This path is fine.iterative-optimizer.tsโ no matches formakeRunningJudgeorrunningJudgeinside the file. If anything uses this optimizer directly (not via the /run endpoint), it scores text similarity.scripts/evolution-hypothesis.mjsโ referenced in the grep as another consumer; needs audit.
Fix:
- Audit every consumer of
IterativeGEPA.run()(grepIterativeGEPA, inspect each caller). - For any caller that passes a bare judge for instruction evolution, wrap with
makeRunningJudge(base, llm). - Add a type guard / runtime check in
IterativeGEPA.run()that rejects judges which haven't been marked as running-capable (add a brand/phantom property tomakeRunningJudge's return soIterativeGEPAcan assert it). - Update
evolution-hypothesis.mjs+ any other standalone harnesses to use the running judge.
Effort: 2-4 hours including audit.
G4 โ Traces rarely finalized with success/verified/corrected ๐ข
Claim: EvalDatasetBuilder mines examples from execution_traces. If outcomes aren't consistently set to a terminal success value, buildExamplesFromTraces returns zero examples and the orchestrator skips with "no eligible traces". GEPA doesn't fail โ it just never runs.
Evidence:
packages/server/src/local/routes/chat.ts:1146โtraceRecorder.start()is called correctly at the start of each chat turn.packages/server/src/local/routes/chat.ts:1231โtraceRecorder.finalize(traceHandle, {...})is called. Need to verify the outcome argument always resolves to'success'/'verified'/'corrected'for turns that should be eligible, and audit what happens on tool-error / abort paths.harness-trace-bridge.ts:123โ only explicit'verified' | 'abandoned'literal found in agent src. Outcome coverage is thin in production code paths.
Fix:
- Audit
chat.tsfinalize paths โ what outcome do we emit on (a) successful final assistant message, (b) tool error mid-turn, (c) user abort / SSE disconnect, (d) rate-limit failure, (e) inner monologue / empty text? Document the matrix. - Ensure
'success'is emitted for turns that produced a valid final assistant message without fatal errors. - Backfill outcome on traces that have a valid final message but no explicit outcome (one-time migration script).
- Add a health metric in the Evolution dashboard: "Eligible traces available for next run: N" โ so the user sees the dataset pool size before kicking off a run.
- When
EvalDatasetBuilderreturns fewer thanminExamples, emit a structured error to the /run response body explaining WHY (current wording "no eligible traces to form dataset" is opaque to end users).
Effort: 0.5 day.
GEPA closure totals
4 items, ~2 engineering days, all unblocked. Ship order: G4 (makes runs possible) โ G2 (makes deploys consumable) โ G3 (audits correctness) โ G1 (autonomy).
After closure, the claim "Waggle self-evolves its agent behavior in production" becomes defensible โ today it is defensible only for library-level tests.
6. hive-mind OSS Integration ๐ข (7 days)
From docs/HIVE-MIND-INTEGRATION-DESIGN.md. 8 items across MCP resources, CLI, hooks, installer. Blocks C1 decision.
7. Accessibility (๐ข 1 day, post-launch OK)
| # | Fix | WCAG |
|---|---|---|
| A11Y-1 | Boot screen: screen-reader skip announce | 2.1.1 |
| A11Y-2 | Dock: 44x44px touch targets | 2.5.8 |
| A11Y-3 | Window title bar: icons on min/max buttons | 1.4.1 |
| A11Y-4 | PersonaSwitcher: aria-disabled on locked cards | 4.1.2 |
| A11Y-5 | Settings: role="switch" + aria-checked on toggles | 4.1.2 |
| A11Y-6 | Dashboard: health-dot shape differentiation | 1.4.1 |
| A11Y-7 | Chat feedback dropdown: focus trap + arrow keys | 2.1.1 |
| A11Y-8 | Global Search: role="dialog" | 1.3.1 |
| A11Y-9 | Memory: aria-label on importance slider | 1.3.1 |
8. Totals
| Category | Items | Eng days | Budget |
|---|---|---|---|
| Polish sprint Phases A-F | ~35 | ~20 | โ |
| P0 launch blockers (tests, papers, launch prep) | 50 | 50 | $2.65-4k |
| P1 ship quality (QW, OW-6, CR-*, 3b, INST) | ~25 | 6.5 | โ |
| P2 polish (Medium UX, ENG, wiki v2, harvest UX, responsive) | ~30 | 22 | โ |
| P3 future (A11Y, notarization, LinkedIn) | ~15 | 6.5 | โ |
| PDF deferred | 21 | ~7 | โ |
| GEPA wiring closure (NEW) | 4 | ~2 | โ |
| Total everything | ~150 | ~94 | ~$3-4k |
Calendar with parallelism: ~7-8 weeks to launch.
9. Critical path
Marko exports (M1) โโโบ Phase 1 Harvest (3d) โโโบ Phase 4 Memory Proof (10d) โโโบ Paper 1
โโ parallel โโบ Phase 2 Wiki v2 (7d) โ
โโ parallel โโบ Phase 3 Harvest UX (7d) Phase 5b Combined (7d) โโโบ Paper 2
โ
API credits (M5) โโโบ Phase 5 GEPA Proof (21d) โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
GEPA wiring closure (G1-G4, 2d) โ must ship before Phase 5
Stripe (M7) + Signing (M8) โโโบ Phase 7 Launch Prep โโโบ LAUNCH DAY
hive-mind extraction (CR-6) โโโโโโโโโโโโโโโโโโโโโโโโโโโบ LAUNCH DAY
GEPA closure (G1-G4) is now on the critical path for the Phase 5 GEPA proof โ without it, the proof would measure text similarity instead of real agent behavior.
10. Related docs
docs/plans/POLISH-SPRINT-2026-04-18.mdโ phased polish plandocs/plans/BACKLOG-CONSOLIDATED-2026-04-17.mdโ prior consolidated backlog (pre-GEPA-wiring audit)docs/plans/PDF-E2E-ISSUES-2026-04-17.mdโ PDF triagedocs/HIVE-MIND-INTEGRATION-DESIGN.mdโ OSS package designdocs/UX-ASSESSMENT-2026-04-16.mdโ UX findings sourcedocs/test-plans/*.docxโ Phase 4/5/7 protocolsdocs/REMAINING-BACKLOG-2026-04-16.mdโ 2026-04-16 master snapshotdocs/TOTAL-WORK-ESTIMATE.mdโ effort breakdown