BIN
judging/round3/crops/01-modal-top.png
Normal file
|
After Width: | Height: | Size: 85 KiB |
BIN
judging/round3/crops/02-band0.png
Normal file
|
After Width: | Height: | Size: 922 KiB |
BIN
judging/round3/crops/02-band1.png
Normal file
|
After Width: | Height: | Size: 1023 KiB |
BIN
judging/round3/crops/02-band2.png
Normal file
|
After Width: | Height: | Size: 481 KiB |
BIN
judging/round3/crops/02-main-bottom.png
Normal file
|
After Width: | Height: | Size: 473 KiB |
BIN
judging/round3/crops/02-main-top.png
Normal file
|
After Width: | Height: | Size: 550 KiB |
BIN
judging/round3/crops/02-sidebar.png
Normal file
|
After Width: | Height: | Size: 93 KiB |
BIN
judging/round3/crops/03-memory-center-c0.png
Normal file
|
After Width: | Height: | Size: 589 KiB |
BIN
judging/round3/crops/03-memory-center-c1.png
Normal file
|
After Width: | Height: | Size: 885 KiB |
BIN
judging/round3/crops/04-skills-hub-c0.png
Normal file
|
After Width: | Height: | Size: 817 KiB |
BIN
judging/round3/crops/04-skills-hub-c1.png
Normal file
|
After Width: | Height: | Size: 1.0 MiB |
BIN
judging/round3/crops/04-skills-hub-c2.png
Normal file
|
After Width: | Height: | Size: 528 KiB |
BIN
judging/round3/crops/04b-skills-agent-badge-c0.png
Normal file
|
After Width: | Height: | Size: 1.3 MiB |
BIN
judging/round3/crops/04b-skills-agent-badge-c1.png
Normal file
|
After Width: | Height: | Size: 1.2 MiB |
BIN
judging/round3/crops/05-agent-center-c0.png
Normal file
|
After Width: | Height: | Size: 772 KiB |
BIN
judging/round3/crops/05-agent-center-c1.png
Normal file
|
After Width: | Height: | Size: 910 KiB |
BIN
judging/round3/crops/05-agent-center-c2.png
Normal file
|
After Width: | Height: | Size: 345 KiB |
BIN
judging/round3/crops/05-agent-row.png
Normal file
|
After Width: | Height: | Size: 224 KiB |
BIN
judging/round3/crops/05-agent-row2.png
Normal file
|
After Width: | Height: | Size: 231 KiB |
BIN
judging/round3/crops/05-agent-row3.png
Normal file
|
After Width: | Height: | Size: 522 KiB |
BIN
judging/round3/crops/06-automation-center-c0.png
Normal file
|
After Width: | Height: | Size: 826 KiB |
BIN
judging/round3/crops/06-automation-center-c1.png
Normal file
|
After Width: | Height: | Size: 910 KiB |
BIN
judging/round3/crops/06-automation-center-c2.png
Normal file
|
After Width: | Height: | Size: 345 KiB |
BIN
judging/round3/crops/06-overview-zoom.png
Normal file
|
After Width: | Height: | Size: 364 KiB |
BIN
judging/round3/crops/12-command-center-c0.png
Normal file
|
After Width: | Height: | Size: 313 KiB |
BIN
judging/round3/crops/12-command-center-c1.png
Normal file
|
After Width: | Height: | Size: 245 KiB |
BIN
judging/round3/crops/12-footer.png
Normal file
|
After Width: | Height: | Size: 37 KiB |
BIN
judging/round3/crops/12-footer2.png
Normal file
|
After Width: | Height: | Size: 35 KiB |
BIN
judging/round3/crops/12-footer3.png
Normal file
|
After Width: | Height: | Size: 25 KiB |
BIN
judging/round3/crops/13-empty.png
Normal file
|
After Width: | Height: | Size: 461 KiB |
BIN
judging/round3/crops/13-left.png
Normal file
|
After Width: | Height: | Size: 427 KiB |
BIN
judging/round3/crops/13-memory-evolution-c0.png
Normal file
|
After Width: | Height: | Size: 767 KiB |
BIN
judging/round3/crops/13-memory-evolution-c1.png
Normal file
|
After Width: | Height: | Size: 935 KiB |
BIN
judging/round3/crops/13-memory-evolution-c2.png
Normal file
|
After Width: | Height: | Size: 346 KiB |
BIN
judging/round3/crops/14-sidebar.png
Normal file
|
After Width: | Height: | Size: 131 KiB |
145
judging/round3/judge-1-novice.md
Normal file
@@ -0,0 +1,145 @@
|
||||
# Round 3 — Judge 1: The Complete Novice
|
||||
|
||||
## Persona
|
||||
|
||||
I have never used anything like this. My phone is for messages, my laptop is for email and the
|
||||
web. I do not know what an "agent," an "MCP," a "workspace," or an "API key" is. I judged only
|
||||
what I could see and feel, and I judged my first-session path against the novice surface
|
||||
(the six-item sidebar in `14-novice-simple-dock.png`), treating the ~20-item sidebar screens as
|
||||
the place I might grow into. I verified the claimed plain-language hover descriptions exist in
|
||||
`apps/web/src/lib/dock-tiers.ts` (they do, and they are genuinely plain).
|
||||
|
||||
## Scores
|
||||
|
||||
| # | Criterion | Score (1–5) |
|
||||
|---|---|---|
|
||||
| 1 | First-session clarity | **4** |
|
||||
| 2 | "It knows me" feeling | **4** |
|
||||
| 3 | Visible agent growth | **2** |
|
||||
| 4 | Desire to return | **4** |
|
||||
| 5 | Absence of friction | **2** |
|
||||
|
||||
## Per-criterion reasoning
|
||||
|
||||
### 1. First-session clarity — 4
|
||||
|
||||
The onboarding is honestly good for someone like me. Three steps, a progress bar, "Skip setup"
|
||||
always visible, and the privacy line ("Your memory and data stay on your device") in words I
|
||||
understand. Step 1 asks things I can answer (name, what kind of work, "What do you want Waggle
|
||||
to help with?") and the live preview — "Good evening, Marko — your work will be remembered
|
||||
here" — instantly shows me what my answers buy me. Step 2's per-provider how-tos ("In ChatGPT:
|
||||
Settings → Data controls → Export data") are exactly the hand-holding I need, and the
|
||||
auto-detected "Claude Code detected — Found 425 items / Import my history" button is one-click
|
||||
magic for people who have it. The six-item sidebar is calm and mostly self-explanatory.
|
||||
|
||||
Why not 5: vocabulary leaks through at the worst moments. "Welcome to the Hive" and "YOUR AI
|
||||
OPERATING SYSTEM" tell me nothing (the small subtitle does all the work). Step 3 explains
|
||||
workspaces as "its own brain — memory, files, and agents stay isolated" — that's the first time
|
||||
the word "agents" appears, undefined, plus "isolated," a word I'd never use. And two of my six
|
||||
sidebar items are technical: "Vault" and "Spawn Agent" (complaints #2, #3 below).
|
||||
|
||||
### 2. "It knows me" feeling — 4
|
||||
|
||||
This is the app's strongest muscle. The returning-user panel (`01-home-welcome-back.png`)
|
||||
greets me by name, tells me "23 memories · 271 people, projects & things it knows · 2 awaiting
|
||||
your OK," and then — the genuinely visceral part — quotes my own working style back at me under
|
||||
"I REMEMBER": "I always work with a draft → critique → rewrite loop. The critique pass is the
|
||||
most important…" That is *me*, in my words. The Home header "You've been away 10 days, Marko.
|
||||
Here's what happened" plus "YOU WERE WORKING ON" cards with Continue buttons make the
|
||||
left-off-here promise concrete, and "Up next: I'll check in on quiet projects at Jun 15" makes
|
||||
it feel like it will keep knowing me.
|
||||
|
||||
Why not 5: the showcase list undercuts itself. The third "I REMEMBER" item is "Session
|
||||
(2026-04-30): What is sovereign AI — 4 messages" — a machine log label with a raw date sitting
|
||||
beside two beautifully human memories (complaint #7). And the numbers don't reconcile if I get
|
||||
curious: 23 memories in the header, "11 memories across 16 sessions" on the workspace card, and
|
||||
a Memory screen showing a single memory card (complaint #6). It feels slightly inflated.
|
||||
|
||||
### 3. Visible agent growth — 2
|
||||
|
||||
The promise is "it learns my workflows and upgrades its own skills." What I can actually *see*
|
||||
is almost nothing — and on my novice sidebar, literally nothing. The Skills Hub, Agent Center,
|
||||
and Evolution screens are not in my six-item tier, so my only growth signals are a future-tense
|
||||
promise ("I'll suggest a new skill for you at Jun 17") and "5 automations completed" overnight —
|
||||
which, on inspection, are mostly internal plumbing like "Index reconciliation" (complaint #10).
|
||||
On the power surfaces the story isn't better: the Evolution screen — the showcase for
|
||||
"improves itself" — is empty ("Select a run to review") and explains itself with "baseline vs
|
||||
winner, which gates fired" (complaint #4). The one skill the assistant actually built itself is
|
||||
marked only by a tiny "agent · review" badge that I would never decode as "I built this for you"
|
||||
(complaint #5). The left-panel copy "Your agent improves itself here… nothing changes without
|
||||
you" is lovely and reassuring — but it's a caption on an empty room. I *read about* growth;
|
||||
I never *saw* it.
|
||||
|
||||
### 4. Desire to return — 4
|
||||
|
||||
The return loop is well designed. "You've been away 10 days… here's what happened" reframes my
|
||||
absence as accumulated value. The Overnight panel (10 memories consolidated, 5 automations
|
||||
completed) says work happened while I slept. "Up next" gives me two concrete dated reasons to
|
||||
come back (Jun 15 check-in, Jun 17 new-skill suggestion). Quick capture ("Jot a note to
|
||||
remember…") invites me to deposit things, which is how habits form. I genuinely wanted to click
|
||||
"Continue" on the Writer demo card.
|
||||
|
||||
Why not 5: the impressive-stats panel shows "0 Artifacts created" — a deflating zero in
|
||||
jargon I don't know ("artifacts"? "consolidated"?) right where the app is trying to brag
|
||||
(complaint #11). And one of my two "you were working on" cards is "Default Workspace — 57d ago,"
|
||||
which feels stale rather than alive.
|
||||
|
||||
### 5. Absence of friction — 2
|
||||
|
||||
Onboarding and Home are smooth — skippable steps, sensible defaults, one-click import. But the
|
||||
single worst moment in the entire evidence set sits exactly where the magic is promised: the
|
||||
"resuming work" screen (`07-workspace-resume.png`). My message asks the assistant to run the
|
||||
critique pass, and its reply is, verbatim, a raw error dump in the chat bubble:
|
||||
|
||||
> `[spawn failed] LLM error (404): {"error":{"message":"Anthropic API error: {"type":"error","error": {"type":"not_found_error","message":"model: auto"},"request_id":"req_011CbyyFYCakcsK8rNTTEok1"}}`
|
||||
|
||||
"Spawn failed." A wall of braces. No "try again" button, no plain-language explanation, no
|
||||
recovery path I can see. As a novice this reads as "the app is broken and it's probably my
|
||||
fault," at the precise moment I was promised "pick up where you left off." Chat is core to my
|
||||
novice tier, so this is my path, not the power user's (complaint #1). Add the developer
|
||||
concepts pushed into my six-item world ("Vault — your API keys and secrets," "Spawn Agent") and
|
||||
friction earns a 2 despite the otherwise polished flow.
|
||||
|
||||
## Numbered complaints (concrete, actionable)
|
||||
|
||||
1. **Raw JSON error as the assistant's chat reply** (`07-workspace-resume.png`, Chat pane). The
|
||||
agent answers with `[spawn failed] LLM error (404): {"error":…"not_found_error"…"model: auto"…}`.
|
||||
Replace with a human card ("I couldn't connect just now — tap to retry"), put the raw error
|
||||
behind a "Show details" disclosure, and never show `request_id` JSON in a chat bubble.
|
||||
2. **"Spawn Agent" in the novice sidebar** (`14-novice-simple-dock.png`, bottom of sidebar).
|
||||
Both words are jargon to a first-timer. Rename for the simple tier (e.g., "New assistant").
|
||||
3. **"Vault" occupies one of six novice slots** (`14-novice-simple-dock.png`; tooltip per
|
||||
`dock-tiers.ts`: "Your API keys and secrets, stored locally"). A novice has no API keys.
|
||||
Worse: the simple dock has **no Memory entry** (confirmed in `dock-tiers.ts` lines 122–127),
|
||||
so the tier built for novices hides the app's #1 superpower. Swap Vault for Memory.
|
||||
4. **Evolution screen is an empty room with engineer copy** (`13-memory-evolution.png`).
|
||||
"Select a run to review… baseline vs winner, which gates fired." Seed one example run or hide
|
||||
the tab until a proposal exists; rewrite "gates fired" in plain words.
|
||||
5. **Self-built skill is invisible as an achievement** (`04b-skills-agent-badge.png`,
|
||||
presentation-design row). The only marker is a small "agent · review" badge. Add an explicit
|
||||
callout: "Waggle built this skill from your workflow — review and approve it."
|
||||
6. **Memory counts don't reconcile** (`01-home-welcome-back.png` "23 memories · 271 people…"
|
||||
vs workspace card "11 memories across 16 sessions" vs one visible card in
|
||||
`03-memory-center.png`). Make the headline number match what clicking through reveals.
|
||||
7. **Machine log entry in the "I REMEMBER" showcase** (`01-home-welcome-back.png`, third item:
|
||||
"Session (2026-04-30): What is sovereign AI — 4 messages"). Filter session-log labels out of
|
||||
the human-memory list, or rephrase them ("We talked about sovereign AI in April").
|
||||
8. **"Welcome to the Hive" / "YOUR AI OPERATING SYSTEM"** (`08-onboarding-welcome.png`) — the
|
||||
headline says nothing to a novice; the subtitle ("Remembers everything. Improves itself.")
|
||||
carries all the meaning. Lead with the plain-language promise.
|
||||
9. **First mention of "agents" is undefined** (`11-onboarding-workspace.png`: "memory, files,
|
||||
and agents stay isolated… Your agent learns each workspace's patterns"). One sentence earlier
|
||||
in the flow should introduce what an agent is ("your AI helper").
|
||||
10. **Overnight/automation brag is mostly plumbing** (`06-automation-center.png` /
|
||||
Home Overnight panel): "Memory consolidation," "Marketplace sync," "Index reconciliation"
|
||||
counted in "5 automations completed." Count only user-meaningful automations on Home, or
|
||||
label the rest "housekeeping."
|
||||
11. **Stat panel jargon + deflating zero** (`02-home-cockpit.png` / `14-novice-simple-dock.png`
|
||||
Overnight panel): "Memories consolidated," "0 Artifacts created." Use plain words
|
||||
("things it learned," "documents made for you") and hide zero-count stats.
|
||||
|
||||
## Verdict in one line
|
||||
|
||||
The memory promise lands — I felt greeted, remembered, and given reasons to come back — but the
|
||||
self-evolving promise is a caption on an empty room, and one raw JSON error at the
|
||||
resume-your-work moment would send a real novice straight back to their email tab.
|
||||
67
judging/round3/judge-2-casual-professional.md
Normal file
@@ -0,0 +1,67 @@
|
||||
# Judge 2 — Casual Non-Technical Professional (Round 3)
|
||||
|
||||
## Persona summary
|
||||
|
||||
I'm a marketing manager. I live in Outlook, Slack, and PowerPoint. I use ChatGPT a few times a week when I remember to. I will give a new app exactly one session before I decide whether it's worth my time. I do not know what an API, an MCP, or a "sub-agent" is, and I don't want to learn. I want the app to feel like a sharp assistant who remembers me, not like a developer console.
|
||||
|
||||
Evidence reviewed: all 15 screenshots under `judging/screenshots/` (read at full resolution, key regions cropped/zoomed), live app confirmed responding at http://localhost:8080 (200), and the sidebar hover descriptions verified in `apps/web/src/lib/dock-tiers.ts` (they are genuinely plain-language: "Everything Waggle remembers about you and your work", "Teach your agents new abilities").
|
||||
|
||||
## Scores
|
||||
|
||||
| # | Criterion | Score (1–5) |
|
||||
|---|---|---|
|
||||
| 1 | First-session clarity | 4 |
|
||||
| 2 | "It knows me" feeling | 4 |
|
||||
| 3 | Visible agent growth | 3 |
|
||||
| 4 | Desire to return | 3 |
|
||||
| 5 | Absence of friction | 2 |
|
||||
|
||||
## Per-criterion reasoning
|
||||
|
||||
### 1. First-session clarity — 4
|
||||
The onboarding is honestly better than most consumer apps I use. Three steps, plainly labeled ("Step 1 of 3"), with a "Skip setup" escape hatch always visible. "Tell us who you are" (09) speaks my language — role chips like Marketing, "What do you want Waggle to help with?" with options like "Draft documents & content", and a live preview of the greeting I'll get ("Good evening, Marko — your work will be remembered here") that makes the payoff tangible before I've invested anything. The import step (10) gives me per-tool instructions in plain steps ("In ChatGPT: Settings → Data controls → Export data. You'll get an email with the file") — that's exactly the hand-holding I need. The workspace step (11) explains itself in one sentence I understand: "Each workspace is its own brain." The novice six-item sidebar (14) — Home, Chat, Files, Vault, Settings, Spawn Agent — is not intimidating.
|
||||
|
||||
Why not 5: even on my simplified sidebar, two of the six items are jargon. "Spawn Agent" — spawn? That's a word from video games or programming, not from my world ("New assistant" or "New agent" would do). "Vault" is guessable but cold. And the very first screen (08) calls itself "YOUR AI OPERATING SYSTEM" / "Welcome to the Hive" — two metaphors stacked before I know what the app does. The subtitle ("Remembers everything. Improves itself.") saves it.
|
||||
|
||||
### 2. "It knows me" feeling — 4
|
||||
This is the app's best moment. The welcome-back modal (01) — "Good evening, Marko" / "23 memories · 271 people, projects & things it knows" / "2 awaiting your OK" — followed by "I REMEMBER" cards that quote my actual working style back to me ("I always work with a draft → critique → rewrite loop — the critique pass is the most important...") gave me a genuine small jolt. The Home screen (02/14) doubles down: "You've been away 10 days, Marko. Here's what happened" with per-workspace summaries and a Continue button. The restored workspace (07) shows my pinned editorial instruction still sitting at the top of the chat. This is the promise, delivered.
|
||||
|
||||
Why not 5 — two cracks in the spell:
|
||||
- The greeting modal (01) lists my workspace as "Writer demo — **Anua**" while the Home cockpit card (02/14) and the workspace header (07) call it "Writer demo — **Anya**". I zoomed in; it's unambiguous. An app whose whole pitch is "I remember everything about you" cannot misspell the name of my workspace in the very greeting that's supposed to prove it remembers. Trust in memory is binary for me.
|
||||
- The third "I REMEMBER" card reads "Session (2026-04-30): What is sovereign AI — 4 messages". That's a database row, not a memory. A person who remembered me would say "We talked about sovereign AI back in April."
|
||||
|
||||
### 3. Visible agent growth — 3
|
||||
The pieces exist, but as a casual user I would not actually *see* growth in this session:
|
||||
- The Memory → Evolution tab (13) — the screen literally dedicated to "self-evolving" — is an empty state: "Your agent improves itself here. When Waggle finds a better way... it proposes an upgrade." That's a promise, not evidence. Nothing proposed, nothing accepted, nothing deployed. The whole superpower is a placeholder card.
|
||||
- The agent-authored skill (04b, "presentation-design") is marked with a tiny amber "agent · review" pill on the far right of a dense ~20-row list. I had to zoom into the screenshot to find it. I would never notice it in real use, and if I did, "agent · review" means nothing to me. "Created by your agent — needs your OK" would.
|
||||
- What *does* work: Agent Center (05) shows a real agent, "Editorial Critic", with my real instruction as its description and "ran 14m ago" — concrete. And the Home "Up next" line "I'll suggest a new skill for you at Jun 17, 10:00 AM" at least dates the promise.
|
||||
Net: growth is asserted in three places and demonstrated in roughly half of one.
|
||||
|
||||
### 4. Desire to return — 3
|
||||
The return loop is genuinely well-designed on paper: the away-digest ("You've been away 10 days... Here's what happened"), overnight stats (30 memories consolidated, 5 automations completed), Automation Center (06) showing 12 active schedules at 100% success including a "Morning briefing", and Quick capture ("Jot a note to remember...") right on Home. That's a real reason to open it tomorrow morning — I want that briefing.
|
||||
|
||||
But the single workspace I'd actually return TO (07) greets me with a failed conversation (see criterion 5). The desire to return is built by the Home screen and destroyed by the work screen. Also, "2 awaiting your OK" / "1 pending" appears in three places, and nothing on the Home surface tells me what those are or where to click to deal with them — unresolved nags age badly.
|
||||
|
||||
### 5. Absence of friction — 2
|
||||
One screenshot decides this score. In 07-workspace-resume — the flagship "pick up where you left off" moment — the agent's most recent reply in my chat is, verbatim:
|
||||
|
||||
> `[spawn failed] LLM error (404): {"error":{"message":"Anthropic API error: {\"type\":\"error\",\"error\": {\"type\":\"not_found_error\",\"message\":\"model: auto\"},\"request_id\":\"req_011CbyyFYCacsK8rNTTEok1\"}"}}`
|
||||
|
||||
A raw JSON error blob, rendered as a normal chat bubble, with thumbs-up/thumbs-down buttons under it as if I might want to rate it. As a non-technical person, this reads as "the app is broken" — full stop. There is no plain-language explanation, no Retry button, no "we'll fix this" — nothing. This is precisely the moment a one-session-patience user closes the app and doesn't come back. Everything else is comparatively smooth (onboarding is friction-free, simple dock is calm, hover descriptions are good), but the core loop — talk to your agent — fails in the ugliest possible way in the evidence.
|
||||
|
||||
Secondary friction: Memory Center (03) shows a single memory card floating in a vast empty honeycomb, right after the greeting told me it holds "23 memories · 271 people, projects & things". Where are they? Filter defaults are hiding them or the count is inflated — either way it feels empty and contradicts the headline number. And the command palette (12) suggests "/spawn — Spawn a specialist sub-agent" to me, which is developer-speak.
|
||||
|
||||
## Numbered complaints (concrete, actionable)
|
||||
|
||||
1. **Raw API error shown as a chat reply** — 07-workspace-resume.png, main chat: `[spawn failed] LLM error (404): {"error":{"message":"Anthropic API error: ... model: auto ..."}}` rendered as a normal agent bubble with feedback buttons. Replace with a human message ("I couldn't reach my AI model just now — tap to retry") + Retry action; never show raw JSON to end users.
|
||||
2. **Workspace name inconsistency: "Anya" vs "Anua"** — Greeting modal (01-home-welcome-back.png) lists "Writer demo — Anua"; Home cockpit card (02/14) and workspace header (07) say "Writer demo — Anya". A memory product misspelling my workspace name in its "I remember you" greeting breaks the spell. Dedupe/fix the source data and render from one canonical name.
|
||||
3. **Memory Center contradicts the memory count** — 03-memory-center.png shows exactly one memory card ("Session (2026-04-30)...") on an empty honeycomb while the greeting claims "23 memories · 271 people, projects & things it knows". Default the view to show everything (or show "22 more in other workspaces/filters" affordance) so the flagship Memory screen never looks empty.
|
||||
4. **Evolution tab is an empty promise** — 13-memory-evolution.png: "Your agent improves itself here" with zero proposals across all filter chips (proposed/accepted/deployed/rejected/failed). The self-evolving superpower has no visible evidence. Seed it from real activity (e.g., surface the pending "presentation-design" agent skill here as a proposal card) so the first visit shows at least one concrete "I found a better way" item.
|
||||
5. **Agent-authored skill badge is invisible and cryptic** — 04b-skills-agent-badge.png: "agent · review" is a small amber pill at the right edge of a dense list row. Rename to plain language ("Created by your agent — needs your OK"), and surface it on Home ("Your agent built a new skill while you were away — review it"), since this is the proudest moment the product has.
|
||||
6. **Dev jargon on the novice surface** — "Spawn Agent" is one of only six items in my simplified sidebar (14-novice-simple-dock.png), and the command palette (12-command-center.png) suggests "/spawn — Spawn a specialist sub-agent". For this persona, say "New agent" / "Get help from a specialist". Also reconsider "Vault" → "Passwords & keys".
|
||||
7. **"Awaiting your OK" badges with no path to act** — "2 awaiting your OK" (01) and "1 pending" (02/14 workspace cards) appear repeatedly, but no visible element on Home explains what is pending or links to resolve it. Make the badge itself a button that opens the approval queue.
|
||||
8. **"Session (2026-04-30): What is sovereign AI — 4 messages" presented as a memory** — 01-home-welcome-back.png "I REMEMBER" card 3 reads like a database row. Rewrite session-derived memories in natural language ("In late April we explored what sovereign AI means").
|
||||
|
||||
## Verdict in one sentence
|
||||
|
||||
The memory greeting and the away-digest genuinely delivered the "it knows me" jolt — and then the one workspace I resumed answered me with a raw JSON error, which for someone like me is the difference between "magical assistant" and "broken software."
|
||||
140
judging/round3/judge-3-power-user.md
Normal file
@@ -0,0 +1,140 @@
|
||||
# Judge 3 — Non-Developer Power User (Round 3)
|
||||
|
||||
## Persona
|
||||
|
||||
Ops lead. Lives in Notion, Airtable, Zapier. Builds automations for a living without writing code.
|
||||
Learns every shortcut, opens every menu, and judges a product by one question: **is depth rewarded?**
|
||||
I evaluated the full screenshot set (15 PNGs, cropped/zoomed where needed) and verified claims live
|
||||
against the running app (`/api/home/briefing`, `/api/agents`, `/api/skills`, `/api/automations`,
|
||||
`/api/memory/stats` with a Bearer session token).
|
||||
|
||||
## Scores
|
||||
|
||||
| # | Criterion | Score (1–5) |
|
||||
|---|-----------|-------------|
|
||||
| 1 | First-session clarity | **4** |
|
||||
| 2 | "It knows me" | **4** |
|
||||
| 3 | Visible agent growth | **3** |
|
||||
| 4 | Desire to return | **4** |
|
||||
| 5 | Absence of friction | **3** |
|
||||
| | **Total** | **18 / 25** |
|
||||
|
||||
## Per-criterion reasoning
|
||||
|
||||
### 1. First-session clarity — 4
|
||||
|
||||
This is one of the better onboardings I've judged. Three steps, every one skippable, plain language
|
||||
throughout ("Each workspace is its own brain"), and two genuinely smart moments:
|
||||
|
||||
- **Live greeting preview** on the "Tell us who you are" step ("Good evening, Marko — your work will
|
||||
be remembered here") — the form pays off *while you fill it in*.
|
||||
- **"Claude Code detected — Found 425 items"** with a one-click "Import my history" button. Auto-detecting
|
||||
my existing AI usage and quantifying it is exactly the kind of respect-for-my-time a power user notices.
|
||||
|
||||
The novice dock (verified: 6 items — Home, Chat, Files, Vault, Settings, Spawn Agent) keeps day one calm,
|
||||
and the plain-language hover descriptions in `dock-tiers.ts` are real ("Everything Waggle remembers about
|
||||
you and your work").
|
||||
|
||||
Why not 5: the simple dock **omits Memory entirely**. Onboarding just sold me the memory superpower and
|
||||
ingested 425 items — and then the novice sidebar has no Memory entry to go see what it learned. The #1
|
||||
differentiator shouldn't be gated behind the power tier (complaint 4).
|
||||
|
||||
### 2. "It knows me" — 4
|
||||
|
||||
The welcome-back moment is genuinely felt. "Good evening, Marko — 23 memories · 271 people, projects &
|
||||
things it knows · 2 awaiting your OK", followed by an "I REMEMBER" list quoting my actual working style
|
||||
("I always work with a draft → critique → rewrite loop. The critique pass is the most important…"). The
|
||||
cockpit headline "You've been away 10 days, Marko. Here's what happened:" with per-workspace context
|
||||
summaries ("Q3 editorial direction: lean into skepticism, less hype") is the strongest continuity
|
||||
experience in this product category. I verified the substance: 271 entities is the live API number,
|
||||
the briefing API returns exactly the greeting, workspaces, pending counts, and scheduled commitments
|
||||
shown on screen, and the workspace right-rail ("11 memories · 2 sessions", last activity `memory_write
|
||||
2h ago`) is real data.
|
||||
|
||||
Why not 5: **the numbers don't reconcile, and the deep dive undersells the claim.** The modal says
|
||||
"23 memories / 2 awaiting your OK"; `/api/memory/stats` reports `total.frameCount: 14` and the briefing
|
||||
reports `needsReviewCount: 0`. Worse, clicking into Memory Center shows **one** memory card above a huge
|
||||
empty honeycomb. The headline says it knows everything; the filing cabinet looks nearly empty. Trust in
|
||||
a memory product is arithmetic — two surfaces calling different numbers "memories" is a credibility leak
|
||||
(complaints 2, 6).
|
||||
|
||||
### 3. Visible agent growth — 3
|
||||
|
||||
The scaffolding is excellent; the evidence is thin. What works:
|
||||
|
||||
- The `presentation-design` skill carries an **"agent · review" provenance badge** (verified: its API
|
||||
`source` is `chat-session` while all 19 other skills are unsourced) — a real artifact of the agent
|
||||
authoring a skill from my conversations.
|
||||
- The cockpit commits to future growth in first person: "I'll suggest a new skill for you at Jun 17,
|
||||
10:00 AM" (verified in the briefing's `upNext`).
|
||||
- The Memory → Evolution tab has the best plain-language framing of self-improvement I've seen:
|
||||
"Your agent improves itself here… You review each proposal and accept or reject it — nothing changes
|
||||
without you." Filters for proposed/accepted/deployed/rejected show someone designed for a real lifecycle.
|
||||
|
||||
Why 3 and not 4+: **the Evolution tab is an empty state.** Zero runs, "Select a run to review", and an
|
||||
unexplained "New Run" button (run *what*? — a non-dev has no idea). The Agent Center shows one agent and
|
||||
"avg success —" placeholder dashes. The single concrete growth artifact (the authored skill) is a tiny
|
||||
chip on row 13 of a list — nothing on Home or in Evolution narrates "I created presentation-design from
|
||||
your chat last week; here's what it does." The mission says growth must be *viscerally obvious*; today
|
||||
it's a badge you must hunt for plus promissory copy (complaints 3, 6).
|
||||
|
||||
### 4. Desire to return — 4
|
||||
|
||||
For this persona, the return loop is the product's strongest muscle. The cockpit is a real "while you
|
||||
were away" report with receipts: overnight digest (10 memories consolidated, 5 automations completed),
|
||||
resume cards with pending counts, suggested next action, dated future commitments, and quick capture.
|
||||
The Automation Center is honest Zapier-grade ops: 13 automations (verified live: 12 active + 1 paused,
|
||||
real cron schedules, real last-run timestamps), next-3 / recent-3 panels, 100% recent success. The app
|
||||
demonstrably works while I'm gone and shows me the ledger — that's the habit hook, and it's earned.
|
||||
|
||||
Why not 5: the trust required to hand an assistant more of my work was dented the moment I opened the
|
||||
flagship workspace and saw a raw 404 error blob as the agent's reply (see criterion 5). Also a labeling
|
||||
nit: the "Overnight" panel actually summarizes a 10-day absence (complaint 7).
|
||||
|
||||
### 5. Absence of friction — 3
|
||||
|
||||
The killer: in `07-workspace-resume.png`, the Writer workspace chat — the demo centerpiece — shows the
|
||||
agent replying with a **raw JSON error dump**:
|
||||
|
||||
> `[spawn failed] LLM error (404): {"error":{"message":"Anthropic API error: {\"type\":\"not_found_error\",\"message\":\"model: auto\"}…request_id…}}`
|
||||
|
||||
I verified the cause live: the Editorial Critic agent's model is `"auto"` (`/api/agents`), which the LLM
|
||||
router 404s on. That means a *default configuration* produces an unhandled developer-grade error rendered
|
||||
straight into a non-technical user's conversation, with no recovery path. For the exact audience this
|
||||
product targets, that is a session-ender (complaint 1).
|
||||
|
||||
Beyond that, friction is genuinely low — consistent left nav, every screen URL-routed, Ctrl+K palette
|
||||
with suggested slash-commands and keyboard hints, skippable onboarding. But the polish cracks show in
|
||||
exactly the places a power user looks: the palette footer says **⌘K** on a Windows build while the header
|
||||
correctly says Ctrl+K (complaint 5), and three core screens (Memory, Agent Center, Evolution) are mostly
|
||||
empty honeycomb wallpaper below a single row of content, which reads as "broken or abandoned" rather than
|
||||
"new" (complaint 6).
|
||||
|
||||
## Numbered complaints (concrete, actionable)
|
||||
|
||||
1. **Raw API error JSON rendered in chat** (`07-workspace-resume.png`). Default agent model `"auto"`
|
||||
404s through the router and the user sees `[spawn failed] LLM error (404): {…not_found_error…}`.
|
||||
Fix: catch spawn/LLM failures and render a human message ("I couldn't start — my model isn't set up.
|
||||
Open Settings → Models to fix it") with a retry button; never print raw payloads into conversation;
|
||||
and make `"auto"` resolve to a working default so fresh installs can't hit this.
|
||||
2. **Memory counts disagree across surfaces.** Welcome modal: "23 memories · 2 awaiting your OK".
|
||||
Live APIs: `total.frameCount: 14`, `needsReviewCount: 0`. Pick one counting basis, one label, and
|
||||
reuse the same endpoint for both surfaces.
|
||||
3. **Evolution tab — the self-evolving showcase — is an unexplained empty state.** No runs, and "New Run"
|
||||
means nothing to a non-developer. Seed it with the real story it already has: "Your agent authored
|
||||
*presentation-design* from your chat session — review it here." Growth should be narrated, not implied.
|
||||
4. **Novice dock omits Memory.** Onboarding imports 425 items into memory, then the 6-item simple sidebar
|
||||
gives novices no way to visit it. Add Memory to the novice tier; it's the moat feature.
|
||||
5. **⌘K glyph in the command palette footer on Windows** (header pill correctly shows `Ctrl K`).
|
||||
Platform-detect the modifier hint.
|
||||
6. **Dead honeycomb expanses on Memory Center, Agent Center, and Evolution** — a single content row above
|
||||
~600px of empty wallpaper, plus an "avg success —" placeholder. Empty states should recruit
|
||||
("Spawn a second agent from a template", "Import more history — 14+ sources") instead of showing blank hexes.
|
||||
7. **"Overnight" digest header is wrong after long absences** — it summarized a 10-day gap. Make it
|
||||
"While you were away (10 days)" when `lastActive` exceeds ~24h.
|
||||
|
||||
## Verdict in one line
|
||||
|
||||
The memory cockpit and the return loop are best-in-class for non-technical operators, but the raw 404 in
|
||||
the flagship chat, the empty Evolution showcase, and the memory-count drift keep both superpowers from
|
||||
landing without caveats: **18/25**.
|
||||
60
judging/round3/judge-4-junior-developer.md
Normal file
@@ -0,0 +1,60 @@
|
||||
# Judge 4 — Junior Developer Verdict (Round 3)
|
||||
|
||||
## Persona
|
||||
|
||||
Two years into the job. I live in VS Code, lean on Copilot all day, and keep a ChatGPT tab pinned. I'm the person who hears "AI agent OS" and immediately installs it, clicks every tab, opens DevTools when something looks off, and compares the polish to Raycast, Linear, and Cursor. I'm forgiving of rough edges in my own tooling but I judge a consumer-facing "habit-forming, delightful" promise by consumer standards.
|
||||
|
||||
What I reviewed: all 15 screenshots (full-res, cropped where needed), the dock hover-description source (`apps/web/src/lib/dock-tiers.ts`), and the live app (`GET /api/auth/session-token` → `GET /api/home/briefing` returned the real personalized briefing matching the screenshots — the memory claims aren't mocked pixels).
|
||||
|
||||
## Scores
|
||||
|
||||
| # | Criterion | Score (1–5) |
|
||||
|---|-----------|-------------|
|
||||
| 1 | First-session clarity | **4** |
|
||||
| 2 | "It knows me" | **4** |
|
||||
| 3 | Visible agent growth | **3** |
|
||||
| 4 | Desire to return | **4** |
|
||||
| 5 | Absence of friction | **3** |
|
||||
| | **Total** | **18 / 25** |
|
||||
|
||||
## Per-criterion reasoning
|
||||
|
||||
### 1. First-session clarity — 4
|
||||
|
||||
The onboarding is genuinely tight: 3 steps, progress dots, and the mission is stated up front — "Remembers everything. Improves itself. Built for knowledge work." nails both superpowers in eight words. The live greeting preview on the identity step ("Good evening, Marko — your work will be remembered here.") is a Raycast-grade touch. The "Claude Code detected — Found 425 items" auto-detect banner is the single most impressive moment in the whole flow; that's the kind of ambient smarts that makes a dev grin. The workspace step explicitly plants superpower #2 ("Your agent learns each workspace's patterns and can propose new skills — you approve every change."). The novice dock (14) at 6 items is exactly right, and I verified every nav entry carries a plain-language hover description in `dock-tiers.ts` ("Teach your agents new abilities", "Watch your agents work together live") — these are written for humans, not engineers.
|
||||
|
||||
Why not 5: the bee jargon stacks up fast (Hive, Waggle Dance, hexagon wallpaper everywhere) and "Spawn Agent" sits in the *novice* dock — "spawn" is process-table vocabulary, not novice vocabulary. And after creating the workspace, nothing in the captured flow hands me a first thing to *do*; the clarity ends at the threshold.
|
||||
|
||||
### 2. "It knows me" — 4
|
||||
|
||||
This is the strongest pillar. The welcome-back modal (01) is viscerally personal: "Good evening, Marko · 23 memories · 271 people, projects & things it knows · 2 awaiting your OK", then an "I REMEMBER" section quoting an actual learned workflow preference — "I always work with a draft → critique → rewrite loop. The critique pass is the most important…". That's not a counter, that's *my process reflected back at me*. The home cockpit (02) opens with "You've been away 10 days, Marko. Here's what happened:" and per-workspace context ("Q3 editorial direction: lean into skepticism, less hype") — and I confirmed via the live `/api/home/briefing` that this is real data, not a staged screenshot. The empty workspace's "Nothing here yet — start a chat and I'll remember it." is great voice.
|
||||
|
||||
Why not 5: I clicked through to Memory Center (03) expecting to see those 23 memories and found exactly **one** card in a sea of hexagons (workspace-scoped filtering, presumably — but nothing tells me that, so the headline claim doesn't visibly cash out where memories actually live). And the modal says "Writer demo — **Anua**" while every other surface says "**Anya**" — a memory product that misspells the name it remembers is a uniquely self-defeating typo.
|
||||
|
||||
### 3. Visible agent growth — 3
|
||||
|
||||
There is exactly one piece of real, present-tense evidence: the `agent · review` badge on the `presentation-design` skill (04b) — an agent-authored skill awaiting my approval. That's the right mechanic, made visible, and I like it. Around it, everything is future tense or empty: Home promises "I'll suggest a new skill for you at Jun 17, 10:00 AM"; the command palette footer echoes the same scheduled suggestion; and the dedicated **Evolution tab (13) — the flagship surface for "improves itself" — is an empty state**: a nice explainer card ("Your agent improves itself here… nothing changes without you") next to "Select a run to review" with zero runs and a "+ New Run" button. Agent Center (05) shows one agent (Editorial Critic, real lastRunAt, good instruction text) but "avg success —" as a dead em-dash. The Automation Center (06) is the healthiest growth-adjacent surface (12 active schedules, 100% recent success, real timestamps), but those are schedules I'd expect from any cron UI, not self-evolution. Concept: excellent. Proof on screen: one badge.
|
||||
|
||||
### 4. Desire to return — 4
|
||||
|
||||
The return loop is well-designed and I felt it. The away-briefing ("You've been away 10 days… here's what happened") plus overnight counters (10 memories consolidated, 5 automations completed) is the strategy-game daily-login pattern done tastefully. "Up next: I'll check in on quiet projects at Jun 15, 9:00 AM" creates a literal appointment with the app. Quick capture (Note/Task/Link/File) at the bottom of Home gives me a reason to open it even when I don't need an agent. The command palette (12) with `/catchup` pre-highlighted ("get up to speed instantly") is exactly the re-entry affordance a returning user wants.
|
||||
|
||||
Why not 5: the one chat transcript in evidence (07) ends with the agent failing (see complaint 1) — if my last memory of working here is an error dump, the briefing alone won't pull me back. And "0 artifacts created" overnight quietly undercuts the "things happened while you were away" story.
|
||||
|
||||
### 5. Absence of friction — 3
|
||||
|
||||
Navigation, onboarding, palette, and the cockpit are smooth and consistent. But the flagship resume screenshot (07) — the one demonstrating markdown-rendered chat — shows the agent's actual reply as a raw JSON blob: `[spawn failed] LLM error (404): {"error":{"message":"Anthropic API error: {\"type\":\"not_found_error\",\"message\":\"model: auto\"}…request_id…}}`. My pinned instruction renders in a beautiful orange callout, and the agent answers with a stack trace. I shrug at 404s for a living; the non-technical user this product targets closes the app. Add the Memory Center's 23-vs-1 mismatch, screens that are 70–80% hexagon wallpaper at low data volume, and the Anua/Anya inconsistency, and the polish floor is visibly below the (high) polish ceiling.
|
||||
|
||||
## Numbered complaints
|
||||
|
||||
1. **Raw API error rendered as agent chat output (07-workspace-resume.png).** The agent's reply is a verbatim JSON dump (`[spawn failed] LLM error (404) … "model: auto" … request_id`). Replace with a friendly error card ("I couldn't reach my model — check Settings → Models") plus a Retry button; never surface request IDs/JSON in the chat lane. Also fix the root cause: `model: auto` is 404ing against the Anthropic API in the demo path.
|
||||
2. **Evolution tab is an empty showcase (13-memory-evolution.png).** The dedicated "Your agent improves itself here" surface has zero runs — superpower #2's flagship screen proves nothing. Seed a first real evolution run from existing usage, or deep-link the `agent · review` skill proposal into this tab so there's always at least one reviewable item; hide the tab until then.
|
||||
3. **Memory claim doesn't cash out on click-through (01 vs 03).** Welcome modal says "23 memories · 271 people, projects & things it knows"; Memory Center displays one card with no explanation. Show the global count in the header and a one-click "Show all 23 across workspaces" when the active scope filters the list to near-empty.
|
||||
4. **"Spawn Agent" in the novice dock (14-novice-simple-dock.png).** Process-management jargon in the tier explicitly designed for novices. Rename to "New Agent" (or match the friendly register of the hover descriptions, which are otherwise excellent).
|
||||
5. **"Writer demo — Anua" vs "Writer demo — Anya" (01 modal vs 02/05/07).** A name inconsistency inside the "I remember you" modal reads as the app misremembering — the single worst polish bug a memory product can have. Audit demo/seed data for the typo.
|
||||
6. **Hexagon wallpaper dominates sparse screens (03, 05, 13).** At low data volume, 70–80% of Memory Center / Agent Center / Evolution is decoration. Replace dead space with functional empty-state content: suggested searches, sample memories, "create your first…" actions.
|
||||
7. **Dead-stat placeholder next to live data (05-agent-center.png).** "avg success —" sits beside an agent with a real lastRunAt. Compute the stat from the runs that exist, or hide it until N ≥ 1; an em-dash in a stats row reads as broken telemetry.
|
||||
|
||||
## Bottom line
|
||||
|
||||
The memory pillar is real and verifiable — the briefing API serves the same personalized context the screenshots show, and the "I REMEMBER" modal is the best moment in the product. The self-evolution pillar is currently one badge and two calendar promises wrapped in an empty flagship tab. And the single chat transcript offered as evidence contains an unhandled error. Fix complaints 1–3 and this is a 21–22/25 product; today it's a very promising 18.
|
||||
104
judging/round3/judge-5-senior-skeptic.md
Normal file
@@ -0,0 +1,104 @@
|
||||
# Judge 5 — Senior Engineer / Professional Skeptic (Round 3)
|
||||
|
||||
**Persona:** 15 years shipping software. Default position: "AI that learns" is marketing until I see it in the UI or the API. I read every screenshot at full resolution, pulled a session token, and curled every endpoint I doubted. Credit below is given only where the pixels and the JSON agree.
|
||||
|
||||
**Method:** 15 screenshots (`judging/screenshots/`, cropped via PIL for dense screens) cross-checked against live API at `http://localhost:8080/api` (`/auth/session-token` → Bearer): `/home/briefing`, `/home/overnight`, `/memory`, `/memory/stats` (global + per-workspace), `/identity`, `/skills`, `/agents`, `/automations`, `/evolution/runs`, `/audit/installs`. Also read `apps/web/src/lib/dock-tiers.ts` to verify the progressive-nav claim at source.
|
||||
|
||||
---
|
||||
|
||||
## Scores
|
||||
|
||||
| # | Criterion | Score (1–5) |
|
||||
|---|---|---|
|
||||
| 1 | First-session clarity | **4** |
|
||||
| 2 | "It knows me" (persisted, visible, correctable) | **4** |
|
||||
| 3 | Visible agent growth (substantiated or vapor) | **2** |
|
||||
| 4 | Desire to return (real value or dark patterns) | **3** |
|
||||
| 5 | Absence of friction | **2** |
|
||||
| | **Total** | **15 / 25** |
|
||||
|
||||
---
|
||||
|
||||
## Per-criterion reasoning
|
||||
|
||||
### 1. First-session clarity — 4
|
||||
|
||||
What earned credit (verified):
|
||||
|
||||
- Onboarding is a genuinely tight 3-step flow (08–11): plain-language identity step with a **live greeting preview** ("Good evening, Marko — your work will be remembered here"), a memory-import step, a workspace step. "Skip setup" is always visible. No tech jargon in step 1 ("What kind of work do you do?", chips, not schemas).
|
||||
- The import step (10) makes the memory pitch concrete instead of abstract: **"Claude Code detected — Found 425 items from Claude Code"** with a one-click import. That's the right way to sell memory — with the user's own data.
|
||||
- Progressive nav is real, not a claim: `dock-tiers.ts` `TIER_DOCK_CONFIG.simple` is exactly 6 entries (Home, Chat, Files, separator, Vault, Settings), matching `14-novice-simple-dock.png`, and **every entry carries a plain-language `description`** hover ("Everything Waggle remembers about you and your work", "Teach your agents new abilities"). Verified in source, lines 66–127.
|
||||
- The privacy line on the welcome screen ("Your memory and data stay on your device") is the correct first message for this product.
|
||||
|
||||
Why not 5:
|
||||
|
||||
- The first screen promises "Remembers everything. **Improves itself.**" and step 3 says "your agent … can propose new skills — you approve every change." Per criterion 3 below, the improves-itself half is essentially unbacked in this install (0 evolution runs). Front-loading an unproven superpower in the first 30 seconds is a clarity debt the user collects later.
|
||||
- The novice dock keeps a **"Spawn Agent"** CTA (jargon) while omitting Memory entirely (see Complaint 5) — the simplest tier hides the flagship and shows the most intimidating verb.
|
||||
|
||||
### 2. "It knows me" — 4
|
||||
|
||||
What earned credit (verified):
|
||||
|
||||
- **Persisted:** `/api/identity` → `{name: "Marko", role: "Founder", department: "Egzakta Group", created_at: 2026-04-16, updated_at: 2026-06-12}`. Two months of persistence, updated today. Personal mind frames (`/api/memory?limit=50`) hold real durable identity facts (id 1, created 2026-04-16: name, age, employer, favorite color, football club).
|
||||
- **Visible:** The home briefing greeting "You've been away **10 days**, Marko" is arithmetically honest — last real session activity is 2026-06-02 (frame timestamps), today is 2026-06-12. The welcome modal's "**271 people, projects & things it knows**" matches `/api/memory/stats` `total.entityCount: 271` exactly. Workspace resume cards carry real stored context ("Q3 editorial direction: lean into skepticism, less hype" is verbatim workspace memory, not template text). "2 awaiting your OK" = sum of the two workspaces' `pendingCount` (1+1). Numbers in the UI trace to the API.
|
||||
- **Scoped:** Per-workspace memory isolation is real: writer-demo-anya mind = 11 frames, default-workspace = 1, new-hive = 0 (per-workspace `/api/memory/stats`), and the Memory Center's "About you / About this work" scope chips plus the writer-demo right rail ("11 memories · 2 sessions") match exactly.
|
||||
- **Correctable:** The earlier duplicate-write pollution was cleaned via the product's own APIs and the trail is honest — writer-demo now shows 8 active + 2 `deprecated` + 1 `archived`, and the Memory Center exposes Active/Archived/Deprecated/Needs-review filters rather than silently disappearing data.
|
||||
|
||||
Why not 5:
|
||||
|
||||
- Strip away the staged workspace demo and the personal mind contains **2 frames** — both typed in by the user himself. The "it knows me" depth on display is mostly self-declared identity plus 271 graph entities of unexplained provenance (214 of them in a personal mind holding 2 frames). The 425-item Claude Code import that would prove harvest-scale knowledge was never run in this install.
|
||||
- No visible per-memory **edit** affordance in the Memory Center card (checkbox only) — correction at the UI level is archive/delete, not amend.
|
||||
|
||||
### 3. Visible agent growth — 2
|
||||
|
||||
This is the mission's second superpower and it is **the weakest evidenced claim in the product**.
|
||||
|
||||
- `/api/evolution/runs` → `{"runs": [], "count": 0}`. The Memory → Evolution tab (13) is an explainer empty state ("Your agent improves itself here … nothing changes without you") with a "+ New Run" button. Honest, well-written — and empty. Zero proposals, zero accepted, zero deployed, in an install whose identity dates to April.
|
||||
- The one genuine artifact: the **`presentation-design` skill carries `initiator: "agent", source: "chat-session"`** in `/api/skills`, surfaced in the UI as the amber **"agent · review"** badge (04/04b). UI and API agree; provenance and review-gating are real. This is the only substantiated self-evolution evidence in the entire install, and I credit it.
|
||||
- The Agent Center (05) holds exactly **one** agent, created **today** by the **user** (`createdBy: "user"`, `createdAt: 2026-06-12T18:59`), status Idle, **avg success "—"**. And the captured evidence of its only run (07) is a **failure**: a raw 404 dumped into chat because the literal string `"auto"` was sent to the Anthropic API as a model name. So the live demonstration of agent capability in this evidence set is an agent that cannot spawn.
|
||||
- "I'll suggest a new skill for you at Jun 17, 10:00 AM" (home cockpit) traces to automation id 7 "Capability suggestion", cron `0 10 * * 3` — a **scheduled prompt**, not learned behavior. Waggle Dance ("see what your agents learn from each other") was not demonstrable from the provided evidence.
|
||||
|
||||
Verdict: one real provenance badge does not substantiate "learns workflows, upgrades its own skills, shares knowledge across workspaces." As shipped here, the second superpower is ~90% vapor. 2/5 — the 2 is earned by the agent-authored skill with review gating.
|
||||
|
||||
### 4. Desire to return — 3
|
||||
|
||||
Real pull, honestly delivered:
|
||||
|
||||
- The away-gap briefing ("You've been away 10 days… Here's what happened") with resume cards, pending counts, and one-click Continue is the right return hook, and its numbers are API-backed (`/api/home/briefing` matches the screen field-for-field).
|
||||
- "Up next" commitments are real scheduler entries: "quiet projects at Jun 15, 9:00 AM" = Stale workspace check `nextRun 2026-06-15T07:00Z`; "new skill Jun 17, 10:00 AM" = Capability suggestion `nextRun 2026-06-17T08:00Z`. The app makes promises it has machinery to keep.
|
||||
- No dark patterns found: trial state is a quiet "Trial: 14d left" badge, the welcome modal has "Don't show again", quick capture is one field. Good.
|
||||
|
||||
Why only 3:
|
||||
|
||||
- The overnight story is weaker than it looks. "5 automations completed" is technically true, but `/api/automations` shows four of them sharing `lastRun ≈ 2026-06-12T01:05:57Z` — a **boot catch-up burst at 3:05 AM local**, not the advertised schedule. The Automation Center's own History column displays "Morning briefing — OK · 6/12/2026, 3:05:58 AM" against an 8:00 AM schedule. A user who checks whether the app kept its promise sees it kept it at the wrong time.
|
||||
- The day's one agent interaction ended in a raw error dump (Complaint 1). Value-on-return is promised well and delivered unevenly.
|
||||
|
||||
### 5. Absence of friction — 2
|
||||
|
||||
The consistency work that landed (audit taxonomy, provenance badge, count plumbing) shows — most UI numbers trace cleanly to APIs, which is rare. But the evidence set contains one disqualifying-grade defect and a cluster of trust-eroding inconsistencies; see complaints 1, 2, 4, 5, 6, 7.
|
||||
|
||||
---
|
||||
|
||||
## Numbered complaints (concrete, actionable)
|
||||
|
||||
1. **Raw API error dumped into chat** — `07-workspace-resume.png`, assistant message bubble, Writer demo — Anya chat: `[spawn failed] LLM error (404): {"error":{"message":"Anthropic API error: {\"type\":\"not_found_error\",\"message\":\"model: auto\"} ...}` . Two defects in one: (a) the model preference `"auto"` (visible as `model: "auto"` in `/api/agents`) is passed **literally** to the Anthropic API instead of being resolved to a concrete model; (b) the failure is rendered as nested raw JSON in a product aimed at non-technical users. Fix: resolve `auto` at dispatch in the spawn path; map provider errors to a human message ("I couldn't start the critic agent — model setup issue. Retry / Fix in Settings").
|
||||
|
||||
2. **Failures don't propagate to any health surface.** The Editorial Critic's only run (lastRunAt 2026-06-12T20:08Z) is the 404 failure above, yet `/api/home/overnight` reports `"failures": []`, Agent Center shows status "Idle" with **avg success "—"**, and the Automation Center claims "100% Success rate (recent runs)". The one thing that broke today is invisible everywhere except inside the chat transcript. Fix: spawn failures must write to the run/failure ledger that feeds `/api/home/overnight` and Agent Center stats.
|
||||
|
||||
3. **"Self-evolving" is front-loaded but unsubstantiated.** Onboarding (08: "Improves itself."; 11: "can propose new skills") and the Evolution tab promise self-improvement, but `/api/evolution/runs` returns `count: 0` and the Evolution UI (13) is an empty state with a manual "+ New Run" button. The only evidence is one agent-authored skill badge. Either seed a real first evolution run during onboarding/trial, or soften the first-screen copy until the loop demonstrably fires.
|
||||
|
||||
4. **Scheduler displays contradict their own schedules.** Automation Center (06) History: "Morning briefing — OK · 6/12/2026, 3:05:58 AM" and "Task reminder — OK · 3:05:58 AM" against 8:00/8:30 AM schedules (boot catch-up burst; four automations share `lastRun ≈ 01:05:57Z`). Additionally, 5 of 12 active automations had `nextRun` in the **past** at probe time (e.g., Memory compaction `next: 2026-06-12T01:30Z` observed at 20:30Z) and the "NEXT UP" panel silently omits them, so e.g. Memory compaction shows no next run anywhere. Fix: recompute `nextRun` after catch-up runs; label catch-up executions as such in History.
|
||||
|
||||
5. **The novice tier hides the flagship.** `dock-tiers.ts` `TIER_DOCK_CONFIG.simple` (the `DEFAULT_TIER`) contains no Memory entry — Home/Chat/Files/Vault/Settings only — while the dock still shows a "Spawn Agent" rocket (14-novice-simple-dock.png). The #1 superpower ("knows who you are") has no nav presence for exactly the non-technical audience the mission targets, but agent-spawning jargon does. Fix: swap Memory in (it has the best plain-language description in the file) and gate Spawn Agent to professional+.
|
||||
|
||||
6. **Unlabeled header count fluctuates across the session.** Topbar badge next to the model name reads 14 (21:30, 12/13-*.png) → 23 (22:16, 04-skills-hub.png) → 13 (22:22, 03/05-*.png) → 14 (22:25, 14-*.png). It tracks the global frame count through the duplicate-write/cleanup churn — i.e., it's truthful — but it has no label, sits inside a workspace-scoped breadcrumb ("Default Workspace · Memory · claude-sonnet-4-6 · ⓘ13") while counting **global** frames (default-workspace's own total is 3), and a 23→13 drop with no explanation reads as data loss to a user. Fix: tooltip + scope it to the breadcrumb's workspace, or move it to the Memory entry.
|
||||
|
||||
7. **Memory Center depth for the long-lived workspace is one stub.** `03-memory-center.png`: Default Workspace (57 days old, "1 pending") shows exactly one Active memory — "Session (2026-04-30): What is sovereign AI — 4 messages" — confirmed by API (1 active workspace frame). A returning user opening the moat feature in their default workspace sees a single 6-week-old session stub on an empty honeycomb. The harvest path that would fill this (the detected 425 Claude Code items) was never run. Fix: when a mind is near-empty, the empty space should carry the import CTA from onboarding step 2, not background art.
|
||||
|
||||
---
|
||||
|
||||
## Bottom line
|
||||
|
||||
The memory half of the pitch is more real than I expected to find: numbers on screen trace to APIs, scoping is genuine, the cleanup left an honest audit trail, and the briefing's "away 10 days" is true arithmetic. The self-evolution half is a well-designed empty room with one authentic exhibit (the `agent · review` skill badge). And the single captured attempt to actually use the agent ends in a raw 404 that no health surface admits happened. Ship-blocking for the "delightful for non-technical users" claim: complaints 1 and 2.
|
||||
|
||||
**Scores: 4 / 4 / 2 / 3 / 2 — total 15/25. 7 complaints.**
|
||||
125
judging/round3/verifier-report.md
Normal file
@@ -0,0 +1,125 @@
|
||||
# Round-3 Verifier Report — commit 8996f7e
|
||||
|
||||
**Verifier:** fresh-context verifier session, 2026-06-12
|
||||
**Commit under review:** `8996f7e238a8db040369188f9b5fa029394a56b9` — "feat(ux): judge round-2 fixes — store-level dedup roots, clean previews, honest totals, friendly agendas" (= current HEAD on main)
|
||||
|
||||
## VERDICT: PASS
|
||||
|
||||
All gates green (FE 944/944, server-local 920/920, weaver 31/31, tsc 0+0), every change traces to a
|
||||
named judge complaint or the habit-forming mission, no new dependencies, no flags or compat shims,
|
||||
and all four special-attention items check out. Three minor (non-blocking) findings below.
|
||||
|
||||
---
|
||||
|
||||
## 1. Diff audit — traceability, deps, flags
|
||||
|
||||
`git show --stat 8996f7e`: 79 files, +787/−55. Code surface: 10 FE files in `apps/web/src`,
|
||||
5 server files (`monthly-assessment.ts`, `routes/{home,memory,skills,workspace-context}.ts`),
|
||||
1 weaver file (`consolidation.ts`), 3 test files. The remainder is `judging/round2/*` evidence
|
||||
(judge verdicts, prior verifier report, screenshots/crops) — process artifacts of this mission.
|
||||
|
||||
- **Traceability:** every code change maps to a complaint named in the commit message
|
||||
(dedup root causes, honest totals, clean previews, jargon sweep, plural fix, future-only next-up,
|
||||
recency sorting, fresh-install card preservation). Nothing speculative; no abstractions added
|
||||
beyond a 7-line local `cleanPreview` helper and a 4-entry display-name map.
|
||||
- **No new dependencies:** no `package.json` touched anywhere in the commit. `renderChatMarkdown`,
|
||||
`parseSkillFrontmatter`, `deleteByContentPrefix` are all pre-existing in-repo utilities.
|
||||
- **No flags/shims:** no new `process.env` reads, no compat layers. `FRIENDLY_JOB_NAMES`
|
||||
(workspace-context.ts:126-132) is a presentation mapping, not a shim — verified its 4 keys
|
||||
exactly match the 4 seeded cron names in `packages/server/src/local/setup-crons.ts:23-26`,
|
||||
and unknown (user-created) names pass through unchanged.
|
||||
- **Boundary validation only:** no new internal validation layers introduced.
|
||||
|
||||
### Special-attention items
|
||||
|
||||
**(a) weaver `deleteByContentPrefix` — cross-session safety: OK.**
|
||||
`packages/weaver/src/consolidation.ts:218` deletes by prefix
|
||||
`` `Session (${sessionDate}): ${summary}` `` — date AND summary are both in the prefix, so a
|
||||
different session on the same day is untouched. The implementation
|
||||
(`packages/hive-mind-core/src/mind/frames.ts:343-354`, pre-existing W4.3 utility) escapes LIKE
|
||||
metacharacters (`\ % _`) and routes through `delete(id)` so FTS/vec/KG indexes are cleaned.
|
||||
Scope is the weaver's own per-mind `FrameStore` — cross-workspace deletion is structurally
|
||||
impossible. The new regression test explicitly distills a second, different-summary session on
|
||||
the same date and asserts it survives (consolidation.test.ts:212-226). ✔
|
||||
|
||||
**(b) memory.ts stats all-minds aggregation — error tolerance: OK.**
|
||||
`routes/memory.ts:423-434`: when no workspace is given, it loops `server.workspaceManager.list()`
|
||||
through `countMind(getWorkspaceMindDb(ws.id))`. `getWorkspaceMindDb` → `mindCache.getOrOpen()`
|
||||
(`multi-mind-cache.ts:40-82`) **returns `null` on any open failure** (caught + logged internally,
|
||||
incl. the path-traversal guard), and `countMind` early-returns on `!wsDb` — so an unavailable
|
||||
mind is skipped cleanly, per-workspace. The outer `try/catch` additionally covers enumeration
|
||||
failure, degrading to the personal-only total. The route cannot 500 from a bad workspace mind. ✔
|
||||
(Minor: a mind that opens but throws mid-query would abort counting of the *remaining*
|
||||
workspaces — partial total, still no crash. See finding M2.)
|
||||
|
||||
**(c) monthly-assessment zero-data skip — real months not skipped: OK.**
|
||||
`monthly-assessment.ts:298-304`: skip fires only when `totalInteractions === 0 &&
|
||||
skillsInstalled === 0`. Any month with real interactions (or with skills installed even at zero
|
||||
interactions) still writes its frame. The new test asserts a zero-data month writes nothing
|
||||
(`FrameStore.getRecent(10)` length 0), and the adjacent pre-existing tests (which write months
|
||||
with `totalInteractions: 142` etc.) all still pass in the 920/920 run. ✔
|
||||
|
||||
**(d) skills.ts cleanPreview — review-#3 frontmatter-leak protection: NO REGRESSION.**
|
||||
The structure is unchanged from the review-#3 fix: the on-disk read + `parseSkillFrontmatter`
|
||||
still overrides the raw-content default, so for every skill `loadSkills` returns (it only
|
||||
returns files actually present in `skillsDir`), the preview is derived from
|
||||
`frontmatter.description` (authored display text) or `parsed.body` — never the stamped
|
||||
frontmatter. The raw-content fallback fires only on a read race (file deleted between the two
|
||||
reads), identical to the pre-commit code. Decisively: the review-#3 regression test
|
||||
(`p5-skill-governance.test.ts:71-79` — preview must not contain `initiator:` or `---`, must
|
||||
contain the body heading text) is untouched by this commit and **passed** in the server-local
|
||||
run. `cleanPreview`'s `#`-stripping keeps `toContain('Real Heading')` true. ✔
|
||||
|
||||
## 2. Test runs (executed by this verifier, HEAD = 8996f7e)
|
||||
|
||||
| Suite | Command | Result |
|
||||
|---|---|---|
|
||||
| Frontend | `npx vitest run --root apps/web` | **944 passed (944)**, 91 files, 0 failed |
|
||||
| Server local | `npx vitest run packages/server/tests/local --root .` | **920 passed (920)**, 76 files, 0 failed |
|
||||
| Weaver | `npx vitest run packages/weaver/tests --root .` | **31 passed (31)**, 3 files, 0 failed |
|
||||
|
||||
Zero failures — nothing to adjudicate. (Counts match the commit message's claimed gates;
|
||||
"weaver 18/18" in the message referred to consolidation.test.ts alone — full weaver dir is 31.)
|
||||
|
||||
## 3. Typechecks
|
||||
|
||||
- `npx tsc --noEmit --project packages/server/tsconfig.json` → **0 errors** (exit 0)
|
||||
- `npx tsc -p apps/web/tsconfig.app.json --noEmit` → **0 errors** (exit 0)
|
||||
|
||||
## 4. New tests read — do they assert the new behavior?
|
||||
|
||||
1. **Weaver re-distill** (`packages/weaver/tests/consolidation.test.ts:212-226`): distills the
|
||||
same date+summary twice (key points evolving), plus a *different* summary same date; asserts
|
||||
exactly 2 distilled frames for the date, exactly 1 for the re-distilled summary, and that the
|
||||
survivor contains the updated `point B`. Asserts both replace-on-update AND no cross-session
|
||||
deletion. ✔
|
||||
2. **Assessment zero-data** (`packages/server/tests/local/monthly-assessment.test.ts:146-155`):
|
||||
saves an assessment with `totalInteractions: 0, skillsInstalled: 0` and asserts the
|
||||
FrameStore stays empty — exactly the skip behavior. ✔
|
||||
3. **Briefing-highlights status filter** (`apps/web/src/lib/briefing-highlights.test.ts:27-37`):
|
||||
feeds deprecated, archived, `User asked:`-echo, and one `active` frame; asserts only the
|
||||
living frame survives. The `make()` helper spreads overrides so `status` flows into
|
||||
`BriefingFrameLike` (which gained the `status` field in this commit). ✔
|
||||
|
||||
Brag-line tests (`login-briefing-brag.test.ts`) were also updated and assert the dropped
|
||||
"across N workspaces" clause everywhere except the zero-state.
|
||||
|
||||
## 5. Findings (all minor, non-blocking)
|
||||
|
||||
- **M1 — theoretical prefix-collision in weaver dedup:** if two distinct sessions on the *same
|
||||
date* have summaries where one is a strict string-prefix of the other ("Discussed Q2" vs
|
||||
"Discussed Q2 marketing strategy"), re-distilling the shorter one would delete the longer
|
||||
one's frame. LLM-generated summaries make exact prefix collisions unlikely; the cheap
|
||||
hardening would be including the `. `/end separator in the delete prefix. Not a spec
|
||||
violation — noted for awareness.
|
||||
- **M2 — partial-total on mid-loop throw in stats aggregation:** the all-minds `try/catch`
|
||||
wraps the whole loop, so one corrupt-but-openable mind aborts counting the remaining
|
||||
workspaces (silently smaller total). Unavailable (unopenable) minds are handled per-workspace
|
||||
via `getOrOpen → null`. Acceptable tolerance; per-workspace try would be stricter.
|
||||
- **M3 — untested new server paths:** the all-minds stats aggregation, the home.ts
|
||||
`hasContent` card filter, and `FRIENDLY_JOB_NAMES` have no direct tests (the commit's +5
|
||||
regression tests cover the four most behavior-critical changes; skills preview is covered
|
||||
indirectly by the pre-existing review-#3 test). Within the spirit of "add tests for new
|
||||
interactive logic" but not exhaustive.
|
||||
|
||||
None of these alter the verdict.
|
||||