moving
Some checks failed
Installer Smoke / installer-smoke (push) Has been cancelled

This commit is contained in:
Oleg Maslov
2026-09-02 10:10:29 +02:00
commit 0c3e2ead3b
3841 changed files with 970576 additions and 0 deletions

View File

@@ -0,0 +1,140 @@
# Judge 3 — Non-Developer Power User (Round 3)
## Persona
Ops lead. Lives in Notion, Airtable, Zapier. Builds automations for a living without writing code.
Learns every shortcut, opens every menu, and judges a product by one question: **is depth rewarded?**
I evaluated the full screenshot set (15 PNGs, cropped/zoomed where needed) and verified claims live
against the running app (`/api/home/briefing`, `/api/agents`, `/api/skills`, `/api/automations`,
`/api/memory/stats` with a Bearer session token).
## Scores
| # | Criterion | Score (15) |
|---|-----------|-------------|
| 1 | First-session clarity | **4** |
| 2 | "It knows me" | **4** |
| 3 | Visible agent growth | **3** |
| 4 | Desire to return | **4** |
| 5 | Absence of friction | **3** |
| | **Total** | **18 / 25** |
## Per-criterion reasoning
### 1. First-session clarity — 4
This is one of the better onboardings I've judged. Three steps, every one skippable, plain language
throughout ("Each workspace is its own brain"), and two genuinely smart moments:
- **Live greeting preview** on the "Tell us who you are" step ("Good evening, Marko — your work will
be remembered here") — the form pays off *while you fill it in*.
- **"Claude Code detected — Found 425 items"** with a one-click "Import my history" button. Auto-detecting
my existing AI usage and quantifying it is exactly the kind of respect-for-my-time a power user notices.
The novice dock (verified: 6 items — Home, Chat, Files, Vault, Settings, Spawn Agent) keeps day one calm,
and the plain-language hover descriptions in `dock-tiers.ts` are real ("Everything Waggle remembers about
you and your work").
Why not 5: the simple dock **omits Memory entirely**. Onboarding just sold me the memory superpower and
ingested 425 items — and then the novice sidebar has no Memory entry to go see what it learned. The #1
differentiator shouldn't be gated behind the power tier (complaint 4).
### 2. "It knows me" — 4
The welcome-back moment is genuinely felt. "Good evening, Marko — 23 memories · 271 people, projects &
things it knows · 2 awaiting your OK", followed by an "I REMEMBER" list quoting my actual working style
("I always work with a draft → critique → rewrite loop. The critique pass is the most important…"). The
cockpit headline "You've been away 10 days, Marko. Here's what happened:" with per-workspace context
summaries ("Q3 editorial direction: lean into skepticism, less hype") is the strongest continuity
experience in this product category. I verified the substance: 271 entities is the live API number,
the briefing API returns exactly the greeting, workspaces, pending counts, and scheduled commitments
shown on screen, and the workspace right-rail ("11 memories · 2 sessions", last activity `memory_write
2h ago`) is real data.
Why not 5: **the numbers don't reconcile, and the deep dive undersells the claim.** The modal says
"23 memories / 2 awaiting your OK"; `/api/memory/stats` reports `total.frameCount: 14` and the briefing
reports `needsReviewCount: 0`. Worse, clicking into Memory Center shows **one** memory card above a huge
empty honeycomb. The headline says it knows everything; the filing cabinet looks nearly empty. Trust in
a memory product is arithmetic — two surfaces calling different numbers "memories" is a credibility leak
(complaints 2, 6).
### 3. Visible agent growth — 3
The scaffolding is excellent; the evidence is thin. What works:
- The `presentation-design` skill carries an **"agent · review" provenance badge** (verified: its API
`source` is `chat-session` while all 19 other skills are unsourced) — a real artifact of the agent
authoring a skill from my conversations.
- The cockpit commits to future growth in first person: "I'll suggest a new skill for you at Jun 17,
10:00 AM" (verified in the briefing's `upNext`).
- The Memory → Evolution tab has the best plain-language framing of self-improvement I've seen:
"Your agent improves itself here… You review each proposal and accept or reject it — nothing changes
without you." Filters for proposed/accepted/deployed/rejected show someone designed for a real lifecycle.
Why 3 and not 4+: **the Evolution tab is an empty state.** Zero runs, "Select a run to review", and an
unexplained "New Run" button (run *what*? — a non-dev has no idea). The Agent Center shows one agent and
"avg success —" placeholder dashes. The single concrete growth artifact (the authored skill) is a tiny
chip on row 13 of a list — nothing on Home or in Evolution narrates "I created presentation-design from
your chat last week; here's what it does." The mission says growth must be *viscerally obvious*; today
it's a badge you must hunt for plus promissory copy (complaints 3, 6).
### 4. Desire to return — 4
For this persona, the return loop is the product's strongest muscle. The cockpit is a real "while you
were away" report with receipts: overnight digest (10 memories consolidated, 5 automations completed),
resume cards with pending counts, suggested next action, dated future commitments, and quick capture.
The Automation Center is honest Zapier-grade ops: 13 automations (verified live: 12 active + 1 paused,
real cron schedules, real last-run timestamps), next-3 / recent-3 panels, 100% recent success. The app
demonstrably works while I'm gone and shows me the ledger — that's the habit hook, and it's earned.
Why not 5: the trust required to hand an assistant more of my work was dented the moment I opened the
flagship workspace and saw a raw 404 error blob as the agent's reply (see criterion 5). Also a labeling
nit: the "Overnight" panel actually summarizes a 10-day absence (complaint 7).
### 5. Absence of friction — 3
The killer: in `07-workspace-resume.png`, the Writer workspace chat — the demo centerpiece — shows the
agent replying with a **raw JSON error dump**:
> `[spawn failed] LLM error (404): {"error":{"message":"Anthropic API error: {\"type\":\"not_found_error\",\"message\":\"model: auto\"}…request_id…}}`
I verified the cause live: the Editorial Critic agent's model is `"auto"` (`/api/agents`), which the LLM
router 404s on. That means a *default configuration* produces an unhandled developer-grade error rendered
straight into a non-technical user's conversation, with no recovery path. For the exact audience this
product targets, that is a session-ender (complaint 1).
Beyond that, friction is genuinely low — consistent left nav, every screen URL-routed, Ctrl+K palette
with suggested slash-commands and keyboard hints, skippable onboarding. But the polish cracks show in
exactly the places a power user looks: the palette footer says **⌘K** on a Windows build while the header
correctly says Ctrl+K (complaint 5), and three core screens (Memory, Agent Center, Evolution) are mostly
empty honeycomb wallpaper below a single row of content, which reads as "broken or abandoned" rather than
"new" (complaint 6).
## Numbered complaints (concrete, actionable)
1. **Raw API error JSON rendered in chat** (`07-workspace-resume.png`). Default agent model `"auto"`
404s through the router and the user sees `[spawn failed] LLM error (404): {…not_found_error…}`.
Fix: catch spawn/LLM failures and render a human message ("I couldn't start — my model isn't set up.
Open Settings → Models to fix it") with a retry button; never print raw payloads into conversation;
and make `"auto"` resolve to a working default so fresh installs can't hit this.
2. **Memory counts disagree across surfaces.** Welcome modal: "23 memories · 2 awaiting your OK".
Live APIs: `total.frameCount: 14`, `needsReviewCount: 0`. Pick one counting basis, one label, and
reuse the same endpoint for both surfaces.
3. **Evolution tab — the self-evolving showcase — is an unexplained empty state.** No runs, and "New Run"
means nothing to a non-developer. Seed it with the real story it already has: "Your agent authored
*presentation-design* from your chat session — review it here." Growth should be narrated, not implied.
4. **Novice dock omits Memory.** Onboarding imports 425 items into memory, then the 6-item simple sidebar
gives novices no way to visit it. Add Memory to the novice tier; it's the moat feature.
5. **⌘K glyph in the command palette footer on Windows** (header pill correctly shows `Ctrl K`).
Platform-detect the modifier hint.
6. **Dead honeycomb expanses on Memory Center, Agent Center, and Evolution** — a single content row above
~600px of empty wallpaper, plus an "avg success —" placeholder. Empty states should recruit
("Spawn a second agent from a template", "Import more history — 14+ sources") instead of showing blank hexes.
7. **"Overnight" digest header is wrong after long absences** — it summarized a 10-day gap. Make it
"While you were away (10 days)" when `lastActive` exceeds ~24h.
## Verdict in one line
The memory cockpit and the return loop are best-in-class for non-technical operators, but the raw 404 in
the flagship chat, the empty Evolution showcase, and the memory-count drift keep both superpowers from
landing without caveats: **18/25**.