moving
Some checks failed
Installer Smoke / installer-smoke (push) Has been cancelled

This commit is contained in:
Oleg Maslov
2026-09-02 10:10:29 +02:00
commit 0c3e2ead3b
3841 changed files with 970576 additions and 0 deletions

Binary file not shown.

After

Width:  |  Height:  |  Size: 107 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 321 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.1 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 958 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.1 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 951 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 262 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 937 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 808 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.2 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 994 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 693 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 922 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 432 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 329 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 933 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1004 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 412 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 313 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 85 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 134 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 430 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 742 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.4 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 593 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 429 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.4 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.5 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.5 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.8 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.8 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.3 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.1 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.1 MiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 485 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 422 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 534 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.1 MiB

View File

@@ -0,0 +1,80 @@
# Judge 1 — The Complete Novice (Round 2)
## Persona summary
I am not a computer person. I message on my phone, and on my laptop I do email and the web. I have never used an "AI desktop app". I do not know what an agent, a workspace, an MCP, or an API is. I judged only what I could see and feel in the screenshots, the way I would on my own at the kitchen table.
## Scores
| # | Criterion | Score (15) |
|---|---|---|
| 1 | First-session clarity | **4** |
| 2 | "It knows me" feeling | **4** |
| 3 | Visible agent growth | **3** |
| 4 | Desire to return | **4** |
| 5 | Absence of friction | **2** |
**Total: 17 / 25**
## Per-criterion reasoning
### 1. First-session clarity — 4
The setup is genuinely easy: three short steps, one obvious yellow button each time, and I can skip anything. Step 1 ("Tell us who you are", 09) asks things I can actually answer — my name, what kind of work I do, what I want help with — and the live preview line "Good evening, Marko — your work will be remembered here" instantly shows me what I'm getting. The privacy line on the first screen ("Your memory and data stay on your device. Nothing leaves without your say-so", 08) made me feel safe rather than spied on. Step 3's "Each workspace is its own brain" metaphor (11) actually helped me understand a word I didn't know.
What stops a 5: Step 2 ("Where do you use AI today?", 10) is the one step that assumes I'm already an AI person. Its main button says **"Harvest"** — I don't know what harvesting my computer means, and the line under it shows a raw folder path (`C:\Users\MarkoMarkovic\.claude`). If "Harvest" said "Import" I'd click it without fear. And the moment setup ends, I land facing a left sidebar of ~20 items I mostly can't read (see complaint 1), which dents my "I know what to do next" confidence.
### 2. "It knows me" feeling — 4
This is the app's best trick and it's delivered on screen, not just promised. The welcome panel (01) says "Good evening, Marko", counts "17 memories · 214 people, projects & things it knows across 3 workspaces", and then literally shows an "I REMEMBER" list with my own habits ("I always work with a draft → critique → rewrite loop…"). The Home screen (02) opens with "You've been away 10 days, Marko. Here's what happened:" and shows the exact things I was working on with a Continue button on each. When I resume the Writer workspace (07), the assistant picks up mid-project with "Decision Review & Next Steps — Anya's Content Strategy" without me re-explaining anything. The Memory screen (03) even knows my age, my team, my favorite color, and my football club. I genuinely felt remembered.
What stops a 5: on the very panel that delivers the magic, the first remembered item reads **"User asked: Review recent decisions and next steps"** (01). "User"? That's me — why is it talking about me in the third person like a machine log? And right under my name sits a yellow warning triangle saying **"⚠ 2 pending"** with no noun — pending *what*? A warning sign with no explanation is the one cold, slightly worrying note on an otherwise warm screen. (Small extra wobble: the header says 17 memories total across 3 workspaces, but two of the workspaces individually claim 11 memories each — the numbers don't obviously add up.)
### 3. Visible agent growth — 3
The *story* of growth is told beautifully in one place: the Evolution panel (13) says "Your agent improves itself here… You review each proposal and accept or reject it — nothing changes without you." That sentence is perfect — plain, reassuring, exciting. The Home screen's "Overnight: 9 memories consolidated, 5 automations completed" (02) made me feel it worked while I slept, and the Automation Center's "100% success rate" (06) looks healthy.
But the *evidence* of growth is thin or hidden everywhere I looked:
- The proudest possible moment — a skill the assistant built by itself — is marked with a tiny cryptic badge reading **"agent · review"** (04b, presentation-design row). I would scroll right past it. Nothing says "Your AI built this for you — take a look."
- The Evolution panel itself is **empty**: "Select a run to review… + New Run", and its explainer uses words like "baseline vs winner, which gates fired" (13). I would never press "New Run" because I don't know what a run is.
- The only agent I have, "Editorial Critic" (05), shows **"run never"**, "Idle", and "avg success —". My one helper looks like it has never done anything.
- The "Monthly Agent Assessment" memory cards (03) literally print "**Interactions**: 0" inside broken-looking text.
So I'm *told* it gets better, and I half-believe it, but I can't *see* it getting better.
### 4. Desire to return — 4
Real pulls exist: the away-greeting plus "Continue →" buttons on my unfinished work (02) make tomorrow's first click obvious. "Up next: … Jun 15, 9:00 AM / … Jun 17, 10:00 AM" gives me actual appointments, and the Overnight digest teaches me that things happen while I'm gone — the strongest reason to open it again ("what did it do last night?"). Quick capture ("Jot a note to remember…") invites a tiny daily habit.
What stops a 5: the appointments are written for engineers, not me — **"Stale workspace check"** sounds like something went moldy, and **"Capability suggestion"** is abstract. If those said "I'll tidy up your project Saturday morning" and "I'll suggest a new trick on Tuesday," I'd be excited instead of puzzled. The "2 pending" items also never tell me what reward awaits if I deal with them.
### 5. Absence of friction — 2
The core path (onboarding → home → chat) is mostly plain-spoken, and nothing felt scary — the privacy and "nothing changes without you" lines are calming. But jargon is everywhere I rest my eyes:
- The always-visible sidebar (02) is half alien: **MCP Hub, Connector Hub, Artifacts, Waggle Dance, Mission Control, Vault, Spawn Agent, Usage & Cost**. I recognize Home, Chat, Files, Settings — the rest is a foreign language I must stare at all day.
- Raw formatting symbols leak into the interface: skill descriptions show literal **"## What to do"** and **"**Check session history**"** (04, 04b), and memory cards show **"**Interactions**: 0 … ## Strengths"** (03). To me that looks broken.
- The chat header (07) shows a machine code, **"claude-sonnet-4-6"**, next to an unexplained **"Autopilot"** toggle.
- The very first screen's subtitle says **"Workspace-native"** (08) — insider speak before I've even clicked once.
- Smaller stumbles: "Harvest" (10), "agent · review" (04b), "⚠ 2 pending" (01), "Stale workspace check" (02), "baseline vs winner, which gates fired" (13), "Deprecated" / "Any confidence" filters (03), and "1 agents" (05 — grammar).
None of it is frightening, but a novice meets an unknown word on nearly every screen, so this can't score above 2.
## Numbered concrete complaints
1. **Sidebar jargon overload** (02, all screens, left nav): "MCP Hub", "Connector Hub", "Waggle Dance", "Mission Control", "Vault", "Spawn Agent", "Artifacts" — rename in plain words or hide behind an "Advanced" group; a novice can only parse Home/Chat/Files/Settings.
2. **"Harvest" button + raw file path** (10, onboarding step 2): the detected-history banner's action says "Harvest" and shows `C:\Users\MarkoMarkovic\.claude`. Say "Import my history" and demote the path to a tooltip.
3. **Unexplained "⚠ 2 pending"** (01, welcome panel header; also workspace cards): a warning triangle with no noun and no link. Say what is pending ("2 things need your OK") and make it clickable.
4. **Third-person "User asked:" in "I REMEMBER"** (01): the app calls me "User" in the very list meant to prove it knows me. Should read "You asked me to review recent decisions…".
5. **Raw markdown rendered as text** (04 and 04b skill rows; 03 "Monthly Agent Assessment" cards): literal `##` and `**` symbols visible in descriptions. Render the formatting or strip it.
6. **The self-built skill is uncelebrated** (04b, presentation-design row): the badge "agent · review" is the entire announcement that my assistant taught itself a new skill. Replace with explicit copy like "Built by your AI — review & approve" and consider surfacing it on Home.
7. **Agent Center reads as dead, not learning** (05): "1 agents" (grammar), "avg success —", and Editorial Critic showing "run never / Idle". The only agent looks like it has never worked; seed a first run or hide empty stats.
8. **Evolution panel is empty and jargon-gated** (13): great headline copy, but the action is "+ New Run" and the explainer says "baseline vs winner, which gates fired". A novice will never click. Offer "See how I'd improve myself" and translate gates/baseline into plain words.
9. **Machine model ID in chat header** (07): "claude-sonnet-4-6" dropdown and unexplained "Autopilot" pill. Hide the model string behind a friendly label ("Smart mode").
10. **System-speak schedule items** (02, "Up next"): "Stale workspace check" and "Capability suggestion" — reword as human promises ("I'll tidy up quiet projects", "I'll suggest a new skill").
11. **"Workspace-native" on the first screen** (08 subtitle): insider phrase at the single most novice-facing moment; say "Everything organized by project" or drop it.
12. **Memory filter jargon** (03): chips like "Deprecated" and a "Any confidence" dropdown, plus tabs "Weaver"/"Harvest" — meaningless to a novice; plain-word alternatives needed.
## Verdict in one line
The memory magic is real and visible — I felt greeted, remembered, and pulled back — but the "it keeps getting better" half of the promise is asserted in copy while the screens show empty runs, a never-run agent, and a cryptic badge, all wrapped in more engineer-speak than a novice can comfortably ignore.

View File

@@ -0,0 +1,163 @@
# Judge 2 — Casual Non-Technical Professional (Round 2)
## Persona
Marketing manager. I use ChatGPT a few times a week when I remember to. My real life is
Outlook, Slack, and PowerPoint. I did not read any documentation. I gave this app one
evening to prove it's worth a second evening. Evidence reviewed: 13 full-resolution
screenshots (onboarding flow, returning-user home, workspace resume, Memory Center,
Skills Hub, Agent Center, Automation Center, command palette, Evolution tab), plus a
liveness check against http://localhost:8080.
## Scores
| # | Criterion | Score (15) |
|---|-----------|-------------|
| 1 | First-session clarity | **4** |
| 2 | "It knows me" feeling | **4** |
| 3 | Visible agent growth | **2** |
| 4 | Desire to return | **4** |
| 5 | Absence of friction | **3** |
| | **Total** | **17 / 25** |
## Per-criterion reasoning
### 1. First-session clarity — 4
The good news first: this onboarding is the best part of the product, and it's written in
my language. Step 1 ("Tell us who you are", `09-onboarding-who-are-you.png`) asks things
I can actually answer — chips like "Marketing", "Draft documents & content", "Remember
everything I work on" — and the live preview line ("Good evening, Marko — your work will
be remembered here") told me what the product *is* before I ever saw the product. Three
steps, a visible "Skip setup" escape hatch, and a privacy promise in plain words ("Your
memory and data stay on your device"). After setup, the Home screen
(`02-home-cockpit.png`) tells me literally what to do: three "Continue" buttons, one
"Suggested next action", a quick-capture box. Five minutes in, I knew the pitch: it
remembers my work and picks up where I left off.
Why not 5: the very first words I read are "YOUR AI OPERATING SYSTEM / Welcome to the
Hive / Persistent memory. Workspace-native." (`08-onboarding-welcome.png`). "Operating
system," "Hive," and "workspace-native" are insider words — I briefly wondered if this
replaces something on my computer. And the moment I land in the app, the left sidebar
presents ~20 destinations (Agent Center, Skills Hub, Automation Center, Room, Connector
Hub, MCP Hub, Marketplace, Vault, Mission Control, Events & Logs…). I will never click
"MCP Hub." I don't know what an MCP is and I'm not going to find out.
### 2. "It knows me" feeling — 4
This is the product's strongest muscle and it flexes it everywhere. "Good evening, Marko
— 17 memories · 214 people, projects & things it knows across 3 workspaces"
(`01-home-welcome-back.png`) with an "I REMEMBER" list that includes an actual *working
preference* ("I always work with a draft → critique → rewrite loop. The critique pass is
the most important"). The home screen says "You've been away 10 days, Marko. Here's
what happened" and each workspace card carries a one-line memory of what I was doing
("Q3 editorial direction; lean into skepticism, less hype"). The resume screen
(`07-workspace-resume.png`) reconstructs a decision log with dates, rationale, and
stakeholders without me asking. That's the promise, delivered visibly. I felt it.
Why not 5 — a concrete one: I opened Memory Center (`03-memory-center.png`) to see "what
it remembers about me," and the first two cards are **"Monthly Agent Assessment —
2026-05 / 2026-04"** full of raw, unrendered markdown: `## *Interactions*: 0
**Correction Rate**: 0.0% *Improvement Trend*: 0%`. Robot diary entries — about the
agent, not about me, showing zeros, with literal `##` and `*` characters on screen — rank
*above* the genuinely charming "User's name is Marko Markovic, age 51… favorite color is
blue, supports Crvena Zvezda" card. The first shelf of my "memories" is machine
self-bookkeeping that looks broken. That one screen took the magic down a notch.
### 3. Visible agent growth — 2
The copy promises it; the screens don't show it. The Evolution tab
(`13-memory-evolution.png`) has lovely plain-language framing ("Your agent improves
itself here… You review each proposal and accept or reject it — nothing changes without
you") — and then it's an **empty state**. No runs, nothing proposed, nothing accepted,
and the call to action is a "+ New Run" button, which sounds like *I* am supposed to
operate the self-improvement machinery. The one real artifact is the "agent · review"
badge on the `presentation-design` skill (`04b-skills-agent-badge.png`) — which, if I
squint, means "the AI wrote itself a PowerPoint skill" (genuinely exciting for me!). But
the badge says only "agent · review" with no story, no "Waggle built this for you from
your deck work — take a look." I'd scroll past it. My one agent, Editorial Critic
(`05-agent-center.png`), shows "run never" and "avg success —". The overnight stats ("9
memories consolidated, 5 automations completed") read as system maintenance, not as "it
got better at MY job." Verdict: growth is asserted, not demonstrated. I could not tell
it's learning *for me specifically*.
### 4. Desire to return — 4
Honestly? Yes, I'd open it tomorrow — to see if the morning briefing trick works twice.
The return loop is well designed: it works overnight (Automation Center shows a 3:05 AM
morning briefing run that succeeded, `06-automation-center.png`), it greets me with what
changed, and "Continue" means I never pay the restart tax that makes me abandon ChatGPT
threads. The Editorial Critic concept — an agent that critiques every draft against my
agreed editorial direction — is exactly the kind of thing my job needs. "Up next:
Capability suggestion at Jun 17" even teases a reason to come back on a specific day.
Why not 5: my actual work lives in email, Slack, and PowerPoint, and nothing on the Home
screen connects to any of them. The habit only forms if I move my work *into* Waggle,
and after one session I haven't been given a reason to do that migration. The pull is
real but it's pulling against gravity.
### 5. Absence of friction — 3
Nothing made me want to slam the laptop shut, but several things made me sigh:
- The memory-import step (`10-onboarding-memory-import.png`) auto-detected **Claude Code**
(425 items) — a developer tool I've never opened. For *my* AI history (ChatGPT), the
instruction is: "Settings → Data controls → Export data. You'll get an email with the
file." Leave the app, do an export, wait for an email, download a JSON, come back,
upload. That is homework, on step 2 of 3, during the first run. It's skippable
(good), but the headline feature of onboarding only auto-works for developers.
- Raw markdown leaks everywhere a description appears: every row in Skills Hub
(`04-skills-hub.png`) shows fragments like "## What to do 1. *Identify the decision*
Clarif…". It reads as unfinished software.
- The Win+K palette (`12-command-center.png`) is slash-commands — `/catchup`, `/spawn`,
`/decide` — every row labeled "Command". The placeholder asks "What do you want to
do?" in my language and then answers exclusively in developer.
- Jargon tax: Hive, Harvest, Weaver, Frames, Vault, MCP Hub, "memory compaction",
"Memory lane extraction at 4:00 AM". I understand none of these and the UI doesn't
explain them.
None of this is fatal — the core paths (onboard, resume, chat) are smooth — hence a 3,
not lower.
## Numbered complaints (concrete & actionable)
1. **Onboarding step 2, ChatGPT card** (`10-onboarding-memory-import.png`): the only
path for a ChatGPT user is a manual export-and-wait-for-email errand outside the app.
Auto-detection worked only for Claude Code. Either make ChatGPT import painless or
move this ask to after first value, not step 2 of 3.
2. **Memory Center "About you" ordering + rendering** (`03-memory-center.png`): two
"Monthly Agent Assessment" cards with raw `## *Interactions*: 0 … *Correction Rate*:
0.0%` markdown rank above the actual about-me card. Render markdown and demote agent
self-assessments out of the default human-facing view.
3. **Skills Hub descriptions are raw skill-body fragments** (`04-skills-hub.png`): every
row shows "## What to do 1. **" with literal markdown symbols. Each skill needs a
one-line human description.
4. **"agent · review" badge is unexplained** (`04b-skills-agent-badge.png`): the single
on-screen proof of self-evolution has no plain-language story or tooltip. Say "Waggle
created this skill for you — review it" or the moment is lost on a novice.
5. **Evolution tab is an empty state with a dev-flavored CTA** (`13-memory-evolution.png`):
"Your agent improves itself here" followed by no runs and a "+ New Run" button puts
the burden of self-improvement on me. Seed it with a first proposal or hide it until
one exists.
6. **Editorial Critic has never run** (`05-agent-center.png`): "run never", "avg success
—", "0 running". My one agent is inert on the screen meant to showcase agents.
7. **Slash-command-only palette** (`12-command-center.png`): `/catchup`, `/spawn`,
`/skills`, all labeled "Command" — a developer idiom presented to someone who asked
"What do you want to do?" Plain-verb entries ("Catch me up", "Start a draft") should
lead.
8. **Sidebar overload + jargon naming** (`02-home-cockpit.png` left rail): ~20
destinations including MCP Hub, Vault, Weaver, Mission Control on first arrival.
A casual professional needs 5; tuck the rest behind "More" or a pro mode.
## Bottom line
The memory half of the promise is real and I felt it — the greeting, the
"away 10 days" recap, and the restored decision log are the best "it remembers me"
experience I've seen in an AI tool, and the onboarding that sets it up is genuinely
novice-friendly. The self-evolving half is currently a narrated promise: empty Evolution
screen, an unexplained badge, an agent that has never run. And the finish (raw markdown
in user-facing text, robot bookkeeping atop my memories, slash-command palette) keeps
whispering "built by developers, for developers" at exactly the moments the product is
trying to convince me otherwise.
**Scores: 4 / 4 / 2 / 4 / 3 — total 17/25. 8 complaints.**

View File

@@ -0,0 +1,75 @@
# Judge 3 — Non-Developer Power User (Ops Lead)
**Persona:** Operations lead who lives in Notion, Airtable, and Zapier. I don't code. I build automations, learn every keyboard shortcut, open every menu and tab on day one, and I judge a tool on whether going deep is rewarded — or whether the second layer is hollow. I notice when numbers don't reconcile across screens, because in my world a dashboard that contradicts itself is a dashboard I stop trusting.
**Evidence reviewed:** All 14 screenshots (read at full resolution via crops), plus live API verification against the running sidecar (`/api/home/briefing`, `/api/home/overnight`, `/api/skills`, `/api/agents`, `/api/automations`, `/api/memory/stats`).
---
## Scores
| # | Criterion | Score |
|---|-----------|:-----:|
| 1 | First-session clarity | **4** |
| 2 | "It knows me" feeling | **4** |
| 3 | Visible agent growth | **3** |
| 4 | Desire to return | **4** |
| 5 | Absence of friction | **2** |
**Total: 17/25 · 10 concrete complaints**
---
## Per-Criterion Reasoning
### 1. First-session clarity — 4/5
The onboarding is the best three-step flow I've seen in this category. Plain-language questions ("What do you want Waggle to help with?" with chips like *Automate repetitive work* — that's me), a live greeting preview that updates as I type my name ("Good evening, Marko — your work will be remembered here."), and the killer moment: **Step 2 auto-detected Claude Code with "Found 425 items at C:\Users\MarkoMarkovic\.claude" and a one-click Harvest button.** Per-source export how-tos for ChatGPT/Claude/Gemini/Perplexity (exact menu paths: "Settings → Data controls → Export data") are written for someone exactly like me. "Each workspace is its own brain" is the right one-line mental model. The privacy line on screen one ("Your memory and data stay on your device") earns trust immediately.
Why not 5: the sidebar I land in afterward has ~20 items, and a chunk of them are jargon a non-developer cannot parse from the label alone — **"MCP Hub", "Waggle Dance", "Weaver", "Room"** mean nothing on first read. And the Memory Center splits into **seven tabs** (Memories / Timeline / Graph / Harvest / Weaver / Wiki / Evolution) with no hint about which one I should care about first. The first session is clear; the first *deep dive* requires guessing.
### 2. "It knows me" feeling — 4/5
This is where the product is closest to its promise. "You've been away 10 days, Marko. Here's what happened:" with the actual date is exactly the greeting the mission describes. The welcome panel's **"I REMEMBER" section quoting a learned working preference back to me — "I always work with a draft → critique → rewrite loop. The critique pass is the most important" — is the single most convincing moment in the app.** The workspace resume (07) is genuinely excellent for an ops brain: a structured decision log ("DECISION 1: Q3 Editorial Pivot — 'Skepticism Over Hype', Approved by Marko (founder) on 2026-04-12", stakeholders with their authorities listed), plus a one-line workspace summary on the Home card ("Q3 editorial direction: lean into skepticism, less hype"). The memory card knowing my age, employer, and that I support Crvena Zvezda is the party trick that sells the demo.
Why not 5 — the numbers betray the magic. The modal header says **"17 memories … across 3 workspaces"**, but the Default Workspace card directly beneath it says **"11 memories"** and the Writer demo side panel (07) also says **"11 memories"** — 11+11+0 ≠ 17, and `/api/memory/stats` says 14 personal frames. A tool that claims to remember everything must not contradict itself about how much it remembers. Also, the welcome modal puts **Default Workspace (stale, last touched ~2 months ago) at the top** while Writer demo — the workspace with a pending item and a real summary — is collapsed at the bottom; the Home grid behind it sorts by recency. Two greeting surfaces, seconds apart, disagree about what I should care about.
### 3. Visible agent growth — 3/5
There IS real, verifiable growth evidence — I checked. The `presentation-design` skill carries an **"agent · review" provenance badge**, and the API confirms it (`"initiator": "agent", "source": "chat-session"`): the system wrote itself a skill and is honestly flagging it for my review. Workspace activity says "Created a skill · 10d ago". The Automation Center is the strongest power surface: 13 automations, 12 active schedules, **100% success rate with timestamped recent results** (Morning briefing OK 3:05:58 AM) and named next runs. Overnight: "9 memories consolidated".
But the marquee surface is hollow. **The Evolution tab — literally titled "Your agent improves itself here" — is completely empty: zero proposals across all six filter states (all/proposed/accepted/deployed/rejected/failed)**, even though "Prompt optimization" and "Monthly assessment" automations supposedly run. The Agent Center has exactly **one agent that has never run** ("Run never", "avg success —"). And the growth evidence that *does* surface in Memory is two "Monthly Agent Assessment" cards reporting **"Interactions: 0, Correction Rate: 0.0%"** — the system showing me a report card full of zeroes. I can see the *machinery* of self-evolution everywhere; I can only see one actual instance of it (the skill badge). Telling ≠ showing.
### 4. Desire to return — 4/5
The loop is real and it's built the way a Zapier user wants it: overnight digest with numbers (9 consolidated / 5 automations completed), an **"Up next" section with concrete dated items** ("Stale workspace check at Jun 15, 9:00 AM"), a suggested next action that deep-links into the right workspace ("Resume: Review recent decisions and next steps" — verified in the API as a per-workspace `next-action`), pending-count badges, and Quick capture (Note/Task/Link/File) so the cost of dumping a thought is near zero. `/catchup` being the top "Suggested for you" item in the Ctrl+K palette — with palette commands described in plain outcomes, not dev-speak — is a genuine fast path; the palette is discoverable via a visible "Ctrl+K" chip in the header. Depth is starting to be rewarded.
Why not 5: the overnight story is currently **housekeeping, not work product** — "0 Artifacts created", and the digest items (Memory compaction, Harvest sync, Index reconciliation) are the system doing chores on itself. "Up next" is likewise system maintenance ("Capability suggestion") framed as my agenda. I come back to Notion because something *for me* changed overnight; here, the agent mostly tidied its own room.
### 5. Absence of friction — 2/5
Nothing crashed, no dead-end navigation, every screen rendered, every API answered — but I open every menu, and nearly every menu had a rough edge. Four flagship surfaces are substantially empty at 1440×900, raw markdown leaks into user-facing text, and the numbers don't reconcile. Itemized below.
---
## Numbered Complaints (all concrete)
1. **Home cockpit has a dead column.** On 02-home-cockpit.png there is a vertical divider at ~x=1170 with a completely empty rail (~270px, ~19% of the viewport) to its right — nothing renders in it at all. That's prime real estate on the single most important screen, blank.
2. **The Evolution tab is an empty promise.** 13-memory-evolution.png: "Your agent improves itself here" + six filter chips + zero items in any state. The flagship "self-evolving" surface, on an account where evolution-adjacent automations (Prompt optimization, Monthly assessment) demonstrably run, shows nothing to review, accept, or reject.
3. **Machine exhaust pollutes Memory, ranked above the good stuff.** 03-memory-center.png: the top two memory cards are "Monthly Agent Assessment — 2026-05 / 2026-04" containing unrendered template markdown with all-zero stats ("# Monthly Agent Assessment … \*Interactions\*: 0 \*Correction Rate\*: 0.0% … ## Weaknesses ## Capability Gaps - None detecte…"), tagged FACT — while the genuinely personal "User's name is Marko Markovic, age 51…" card sits below them.
4. **Skills Hub descriptions leak raw markdown with mid-word truncation.** 04-skills-hub.png: list rows read "## What to do 1 \*\*"Identify the decision"\*\* — Clarif", "— Pull fro", "## When to use — Usa". Twenty skills and nearly every description line is a broken markdown fragment instead of a sentence.
5. **Memory counts don't reconcile across three surfaces.** Welcome modal header: "17 memories … across 3 workspaces". Default Workspace card in the same modal: "11 memories across 16 sessions". Writer demo side panel (07): "11 memories". `/api/memory/stats`: 14 personal frames. Four numbers, no arithmetic that connects them.
6. **The two greeting surfaces disagree on priority.** The welcome-back modal (01) lists Default Workspace (last active April, ~2 months stale) first and Writer demo (16d, 1 pending, has a summary) last; the Home grid behind it (02) sorts by recency. Same moment, contradictory ordering.
7. **Agent Center is a near-empty shell with a grammar bug.** 05-agent-center.png: header reads "1 agents · 0 running · avg success —"; the sole agent shows "Run never" and status Idle. Five tabs (All/Personal/Workspace/Team/Autonomous/Archive) over one never-executed item — depth is not yet rewarded here.
8. **Overnight digest reports chores, not output.** 02-home-cockpit.png: "0 Artifacts created"; Next up = Memory compaction, Memory lane extraction, Harvest sync; Up next = "Stale workspace check", "Capability suggestion". The system's self-maintenance is presented as my morning briefing.
9. **Unexplained jargon in primary navigation.** Sidebar labels "MCP Hub", "Waggle Dance", "Weaver", "Room" carry no tooltip-visible plain-language meaning for a non-developer, and Memory Center's 7 tabs (Memories/Timeline/Graph/Harvest/Weaver/Wiki/Evolution) overlap conceptually with no guidance on which to use when.
10. **Center screens don't use the canvas.** Memory Center (03), Agent Center (05), and Automation Center (06) all leave the bottom ~5060% of a 1440×900 window as bare honeycomb wallpaper; Memory Center shows only ~4 cards above the fold of an empty sea.
## What earned the points (for balance)
- Onboarding auto-detect ("Claude Code detected — Found 425 items") with one-click Harvest: best-in-class first-run moment.
- "I REMEMBER" quoting my learned draft→critique→rewrite preference back to me.
- Decision history with approver, date, rationale, and stakeholders in workspace resume — ops-grade.
- "agent · review" provenance badge, verified real in the API (`initiator: "agent"`).
- Automation Center's 100% success rate with timestamped runs and named next runs.
- Ctrl+K palette with contextual "Suggested for you" `/catchup` and plain-language command descriptions.

View File

@@ -0,0 +1,66 @@
# Judge 4 — Junior Developer Verdict (Round 2)
## Persona summary
Two years into the job. I live in VS Code, lean on Copilot all day, keep a ChatGPT tab pinned, and I will absolutely click every button in your app within ten minutes. I have opinions about Raycast's command palette and Linear's empty states, and I judge new tools against that bar. AI agents are the thing I'm most curious about right now — I want to see one actually do something, not read a card telling me it will.
Evaluated from 14 screenshots (1440x900, cropped/zoomed with PIL where needed) plus live verification against the running app at `localhost:8080` (session-token auth, then `/api/workspaces`, `/api/home/briefing`, `/api/memory/stats`, `/api/agents`, `/api/skills`, `/api/automations`, `/api/harvest/sources`).
## Scores
| # | Criterion | Score (1-5) |
|---|-----------|-------------|
| 1 | First-session clarity | **4** |
| 2 | "It knows me" feeling | **4** |
| 3 | Visible agent growth | **3** |
| 4 | Desire to return | **4** |
| 5 | Absence of friction | **3** |
**Total: 18/25 — Complaints: 10**
## Per-criterion reasoning
### 1. First-session clarity — 4
The onboarding is the best three steps in the product. "Each workspace is its own brain — memory, files, and agents stay isolated" (11-onboarding-workspace.png) teaches the core mental model in one sentence — better than most docs pages I've read. Step 1's live greeting preview ("Good evening, Marko — your work will be remembered here") updates as you type your name, which makes the memory promise concrete before you've even entered the app. Step 2 auto-detecting Claude Code ("Found 425 items at C:\Users\MarkoMarkovic\.claude") with a one-click Harvest button is the single biggest "whoa" moment for a developer — it found my actual workflow without me telling it anything. The privacy line on the welcome screen ("Your memory and data stay on your device") answers my first question unprompted.
But two of the three pillars get taught and one doesn't: workspaces and memory are explained; the self-evolving agent is never introduced. Nothing in steps 1-3 prepares you for Evolution, Agent Center, or the agent-review concept. And the moment onboarding ends, you land in a sidebar with roughly 20 items including unexplained jargon — "Room", "Waggle Dance", "Mission Control" (02-home-cockpit.png). I can navigate that because I navigate IDEs all day; the "especially non-technical users" in the mission statement cannot.
### 2. "It knows me" feeling — 4
This is the product's strongest muscle and most of it is real, not staged. "You've been away 10 days, Marko. Here's what happened:" (02) is exactly what I want from a tool I left running. The welcome modal (01) stacks specifics: "17 memories · 214 people, projects & things it knows across 3 workspaces" (the 214 matches `/api/memory/stats` entityCount exactly — I checked), an "I REMEMBER" section with an actual learned working preference ("I always work with a draft → critique → rewrite loop. The critique pass is the most important"), and per-workspace resume lines naming the last session topic. The workspace resume (07) is the payoff: the agent compiles "Decision Review & Next Steps — Anya's Content Strategy" from workspace memory with rendered markdown, dates, and stakeholders ("Marko — Approved the direction; has final say on brand positioning"). The memory card that knows I'm 51, work at Egzakta Group, and support Crvena Zvezda is the kind of detail that makes the greeting feel earned.
Two things stop the 5. First, the welcome modal misspells the workspace name — "Writer demo — Anua" (01) — while the cockpit behind it (02), the workspace panel (07), and the live API all say "Anya". A memory product that misremembers a name in its flagship "I remember you" surface undermines the exact feeling it's selling. Second, when you click through to Memory Center to see this famous memory, you get three cards, two of which are robot-generated assessment reports — the memory feels deep in the greeting and thin at the source.
### 3. Visible agent growth — 3
There is one genuinely excellent, verified artifact here: the `presentation-design` skill in Skills Hub carries an "agent · review" badge (04b), and the API confirms it is not paint — the skill record has `initiator: "agent"`, `source: "chat-session"`. An agent authored a skill, the system tracked provenance, and the UI gates it behind review. That is the self-evolution loop, real, end to end. The workspace panel's "LAST ACTIVITY: Created a skill — 10d ago" (07) reinforces it. The Editorial Critic agent (05) is also real (verified via `/api/agents`) and its goal is wired to remembered context — "apply the Q3 skepticism-over-hype editorial direction before anything ships" is literally Decision 1 from the workspace's memory. That memory-to-agent-config loop is visible and credible.
But everything else is scaffolding. The Evolution tab (13) — the marquee "agent improves itself" surface — is an empty state: zero runs, "Select a run to review", a "New Run" button, and explainer copy. Agent Center shows one agent, "0 running", "avg success —", never executed; the API shows it was created the same day as this evaluation and has no run history. All 13 automations (verified) are system-shipped maintenance jobs — memory compaction, harvest sync, marketplace sync — none learned from my workflows. The growth story today is one authored skill and a stack of promises. That's a real seed, not a visible garden.
### 4. Desire to return — 4
Honestly? Yes, I'd keep it running for a while, and that surprised me. The away-briefing loop (leave → come back → "here's what happened" → one-click Continue into the exact workspace with pending items flagged) is something neither Copilot nor ChatGPT does, and it's implemented, not mocked — `/api/home/briefing` returns the greeting, suggested next action, and per-workspace pending counts I saw on screen. The Claude Code harvest means it accumulates value from work I'm already doing. The Win+K palette (12) with `/catchup`, `/research`, `/spawn`, `/skills` is Raycast-literate and made me feel at home immediately; "/catchup — Workspace restart summary — get up to speed instantly" as the top suggestion is exactly the right default.
What stops the 5: the overnight report says "9 Memories consolidated, 0 Artifacts created, 5 Automations completed" (02). Nine consolidated memories is housekeeping; zero artifacts means the agent layer produced nothing for me while I was gone, and the one agent that could have (Editorial Critic) has never run. The return habit this app wants to build is "come back to finished work" — right now it's "come back to a well-organized summary of nothing having been done." I'd return daily for two weeks on the briefing alone; whether week three survives depends on that artifacts number going above zero.
### 5. Absence of friction — 3
No crashes, no broken layouts, navigation is coherent, and the visual identity (honey-on-dark hex theme) is consistent and genuinely attractive. But the rough edges are pervasive once you leave the happy path, and several are on flagship screens — see the numbered list. The two worst: raw markdown leaking as literal `#`/`##`/`**` text across both Memory Center cards and every Skills Hub row (this is table-stakes rendering, and the chat view proves the app can render markdown beautifully), and the Anua/Anya name inconsistency on the welcome modal. Add zero-data auto-generated memories polluting "About you" and three major screens that are 70-90% empty hexagon wallpaper, and the polish gap against the Linear/Raycast tier I compare everything to is clearly visible.
## Numbered complaints
1. **Workspace name misspelled in welcome modal** — 01-home-welcome-back.png shows "Writer demo — Anua"; the Home cockpit (02), the workspace side panel (07), and `GET /api/workspaces` all say "Writer demo — Anya". The memory product misremembers a name on its "I remember you" surface.
2. **Raw markdown rendered literally in Memory Center cards** — 03-memory-center.png: "# Monthly Agent Assessment — 2026-05", "**Interactions**: 0", "## Strengths - Low correction rate" displayed with literal hashes and asterisks instead of formatted text.
3. **Raw markdown in every Skills Hub row** — 04/04b: skill descriptions render as "# Brainstorm — ... ## What to do Run three" and "## Steps 1. Analyze audience needs and pr..." — every list row leaks frontmatter-style source instead of a clean one-line summary.
4. **Auto-generated noise crowds the "About you" memory** — 03: two of the three visible memories are "Monthly Agent Assessment" reports with Interactions: 0, Correction Rate: 0.0%, Improvement Trend: 0% — zero-data system output presented as things it "knows about me".
5. **Evolution tab has zero evidence of evolution** — 13-memory-evolution.png: the headline self-improvement surface is an empty state ("Select a run to review", no runs, "New Run" button). The superpower is an explainer card.
6. **Agent Center is one row and a void** — 05-agent-center.png: 1 agent, "0 running", "avg success —", never executed (API confirms no run history); ~85% of the screen is decorative hex background.
7. **Overnight digest reports "0 Artifacts created"** — 02-home-cockpit.png: the come-back-to-finished-work promise returns memory housekeeping (9 consolidated) and nothing produced.
8. **Onboarding never teaches the second superpower** — 08-11: identity, memory import, and workspace creation are covered; self-evolving agents, skill authorship, and the review gate are never introduced before the user encounters "agent · review" badges and the Evolution tab.
9. **Sidebar overload with unexplained jargon** — 02: ~20 nav items across 4 sections including "Room", "Waggle Dance", and "Mission Control" with no visible explanation — fine for me, hostile to the non-technical users in the mission.
10. **Welcome screen caption wraps awkwardly** — 08-onboarding-welcome.png: "Nothing leaves without your say-so." breaks as "say-" / "so." across two lines, a sloppy first impression on an otherwise immaculate first screen.
## Bottom line
The memory pillar is real and verified — greeting, briefing, resume, and provenance all check out against the live API. The evolution pillar has exactly one true artifact (the agent-authored skill with its review badge) surrounded by empty stages waiting for a performance. I'd run it next to VS Code this month. Whether it stays depends on the agents earning their tab.

View File

@@ -0,0 +1,119 @@
# Judge 5 — Senior Engineer / Professional Skeptic
**Persona:** 15 years shipping products. I assume "AI that learns" is inflated until the UI or API proves it. I read every screenshot at full resolution, then pulled a Bearer token and audited the live sidecar APIs (`/api/home/briefing`, `/api/home/overnight`, `/api/memory`, `/api/memory/stats`, `/api/identity`, `/api/skills`, `/api/skills/presentation-design`, `/api/agents`, `/api/automations`, `/api/evolution/runs`, `/api/audit/installs`) and cross-checked against `apps/web/src/lib/briefing-highlights.ts`, `login-briefing-brag.ts`, `LoginBriefing.tsx`, and `packages/server/src/local/routes/home.ts`. Credit is given below where the evidence is real. It often is. That makes the staged parts stand out more, not less.
## Scores
| # | Criterion | Score (15) |
|---|-----------|-------------|
| 1 | First-session clarity | **4** |
| 2 | "It knows me" feeling | **3** |
| 3 | Visible agent growth | **2** |
| 4 | Desire to return | **3** |
| 5 | Absence of friction | **2** |
**Total: 14/25 · 12 numbered complaints**
---
## Per-criterion reasoning
### 1. First-session clarity — 4/5
The onboarding (0811) is the strongest surface in the product, and the mental model it sells is honest:
- 3 steps, progress dots, `Skip setup` and `Skip this step` on every screen. No hostage-taking.
- Step 1 shows a live preview ("Good evening, Marko — your work will be remembered here") that the product actually delivers later — verified, the returning-user greeting matches.
- Step 2 ("Where do you use AI today?") performed a **real detection**: "Claude Code detected — Found 425 items at C:\Users\MarkoMarkovic\.claude". That's not a mock; that path exists on this machine. The 5 import cards (ChatGPT/Claude/Gemini/Perplexity/Other) include the actual export instructions per vendor.
- Step 3's framing — "Each workspace is its own brain — memory, files, and agents stay isolated" — is corroborated by the API: `/api/memory/stats?workspaceId=…` returns genuinely separate per-workspace minds (default-workspace: 11 frames/54 entities/312 relations; new-hive: 0/0/0).
- The privacy claim ("Your memory and data stay on your device") is at least architecturally consistent with a localhost sidecar.
- Win+K palette (12) with `/catchup`, `/decide`, `/review` etc. is discoverable and plainly described.
Why not 5: the first thing a returning user sees (the welcome panel) contains numbers that don't reconcile with each other or the API (complaint 3), and the cockpit invites you to "Continue" a workspace that the same panel says is empty (complaint 8). The model is graspable; the first screen's arithmetic isn't.
### 2. "It knows me" feeling — 3/5
The machinery is real. The lived evidence is half genuine, half staged, and the hero presentation shoots itself in the foot.
**Real (verified):**
- `/api/identity`: configured 2026-04-16, name/role/department persisted, updated today. The "214 people, projects & things it knows" headline equals the personal `entityCount` **exactly** (214) — that's real harvested knowledge-graph data, not a vanity number.
- Memory Center (03) statuses are not decoration: the store contains frames in `active`, `archived`, and `deprecated` states, and the transitions actually happened (frame 39, a junk "User preference" misclassification, was deprecated on 06-11 — the correction loop works).
- Workspace memories for writer-demo-anya are substantive: brand-voice rules, an editorial decision with approver and date, newsletter metrics with `tool_verified` source. The chat resume (07) renders a decision record consistent with those frames.
**Not earned:**
- Two of the three flagship "I REMEMBER" items on the welcome panel (01) are frames the system itself has **deprecated** (writer-demo frames 9 and 10, status=deprecated via API), and one of them is junk ("User asked: Review recent decisions and next steps" — a logged query, not a memory). The highlight ranker (`briefing-highlights.ts`) has no status field at all. The product leads with memories it has disowned.
- "You've been away 10 days, Marko" is contradicted by its own store: personal frames written 2026-06-11 13:47 and 14:17, and an agent created 2026-06-12T18:59 — 42 minutes before the briefing timestamp (19:41Z).
- The richest memories live in a workspace literally named "Writer demo — Anya", and all 11 of its seed frames were created in the **same second** (2026-05-27 23:49:49). Staged.
- The organically-grown personal mind is 14 frames, of which 10 are duplicate zero-data self-assessments (see criterion 3).
A 3: persistent, user-visible, correctable — proven. "It knows me" as a lived feeling — propped up by seeded demo data and undermined by the deprecated-highlights bug.
### 3. Visible agent growth — 2/5
The mission promises a self-evolving agent. Here is the full inventory of growth evidence on this install:
**Real (credit where due):**
- The `presentation-design` skill is genuinely agent-authored: frontmatter reads `initiator: agent`, `source: chat-session` (verified via `/api/skills/presentation-design`), and the Skills Hub renders an "agent · review" provenance badge (04b). One real, traceable, agent-created artifact. This is the single best piece of evidence in the product.
- `/api/audit/installs` is a real governance trail with risk/trust/approval taxonomy, and 4 of its 8 entries are **agent-initiated** capability proposals (`initiator: "agent"`, `action: "proposed"` — filesystem ×3, github connector). The agent demonstrably asks for capabilities.
**Vapor:**
- The Evolution screen (13) — the flagship self-evolution surface, copy: "Your agent improves itself here… nothing changes without you" — is an empty state. `/api/evolution/runs``{"runs":[],"count":0}`. Zero runs, ever, in a store whose data goes back to April.
- The Agent Center's only agent ("Editorial Critic") was created by the **user** (`createdBy: "user"`) at 2026-06-12T18:59 — minutes before judging — and has never run: status idle, runs never, avg success "—".
- The "Monthly Agent Assessment" memories are self-evaluation theater: every copy reads `Interactions: 0, Correction Rate: 0.0%`, and concludes "Strengths: Low correction rate". An agent grading itself A+ on a test it never sat. There are **ten duplicate copies** of this in a 14-frame personal mind.
- None of the agent's 4 capability proposals were ever approved or installed; the loop has never closed.
- "Shares knowledge across workspaces": no evidence found on any screen or endpoint.
Mechanism exists; growth has not happened. One real artifact keeps this off the floor: 2.
### 4. Desire to return — 3/5
The retention loop is engineered on the right axis — value, not dark patterns — but the value delivered is thin and partly self-referential.
**Real (verified):**
- "Overnight: 9 memories consolidated / 0 artifacts created / 5 automations completed" is computed from a real audit-event store (`home.ts` `readAuditCounts`, `memory_write` + file-write `tool_call` events over a 24h window). I reconciled the "5 automations": exactly 5 schedules have `lastRun` inside the window (Harvest sync, Memory compaction, Memory consolidation, Morning briefing, Task reminder). The honest "0 artifacts created" — displaying a zero rather than hiding it — is to this product's credit.
- The suggested action ("Resume: Review recent decisions and next steps") deep-links to a real pending task in writer-demo-anya (`pendingCount: 1` in the briefing API), and quick-capture is one keystroke away.
- No dark patterns anywhere: "Don't show again" on the welcome panel, skips throughout onboarding, autonomy is opt-in ("guided").
**Thin:**
- All 5 "overnight" completions fired in a single burst at 01:05:5758Z (3:05 AM local, same second) — a catch-up burst, not a humming overnight workforce. And what did the night shift produce? Memory compaction and consolidation whose visible output is… another duplicate zero-data Monthly Assessment frame (id 42, written 06-12 18:36). "9 memories consolidated" is a raw count of `memory_write` events, several of which were the agent re-writing its own junk.
- The thing that would actually pull a user back — an artifact, a drafted newsletter, a completed task — is exactly the number the panel honestly reports: 0.
Honest loop, weak payload: 3.
### 5. Absence of friction — 2/5
For roughly one hour of adversarial inspection across 9 screens and 11 endpoints, I logged 12 concrete defects (below), including same-screen numeric contradictions, a false hero greeting, past-due "next up" schedules under a 100% success banner, and one flaky API response. Each one is small; together they are exactly the credibility tax a memory product cannot afford. 2.
---
## Numbered complaints (all concrete, all actionable)
1. **Welcome panel showcases deprecated memories.** 01-home-welcome-back.png: 2 of 3 "I REMEMBER" highlights are writer-demo-anya frames 9 & 10, both `status: "deprecated"` (verified via `GET /api/memory?workspace=writer-demo-anya`); one is junk ("User asked: Review recent decisions and next steps"). Root cause: `apps/web/src/lib/briefing-highlights.ts``BriefingFrameLike` has no `status` field and `selectBriefingHighlights()` never filters; `LoginBriefing.tsx:128` feeds it a canned `searchMemory('important decision project plan', 'global')`. Filter `status === 'active'` before ranking.
2. **"You've been away 10 days, Marko" is false.** The store shows personal frames written 2026-06-11 13:47:56 and 14:17:58 (ids 39, 40), and the Editorial Critic agent created 2026-06-12T18:59:26 — 42 minutes before the briefing timestamp (2026-06-12T19:41:46Z). The greeting derives only from workspace chat `lastActive`. Either compute away-time from max(any activity) or say "last chat 10 days ago".
3. **Headline memory count doesn't reconcile with anything.** Welcome panel says "17 memories … across 3 workspaces" while its own cards show 11 (Default) + 0 (New Hive) (+11 writer-demo). `/api/memory/stats` gives personal=14, +default=25, all minds=36 — no combination yields 17. One screen, three mutually inconsistent numbers.
4. **Personal memory is 71% duplicate junk.** 10 of 14 personal frames are copies of "Monthly Agent Assessment" (5× 2026-04, 5× 2026-05; created 05-02 through 06-12), each reading `Interactions: 0 / Correction Rate: 0.0% / Strengths: Low correction rate`. The assessment automation re-writes duplicates on every run and grades itself on zero data. Two of these render as the top cards in Memory Center (03-memory-center.png).
5. **Provenance mislabeled.** Every automation-generated assessment frame carries `source: "user_stated"`. The user never stated them. This corrupts the exact trust signal the Memory Center's filter UI sells (the seeded demo data, ironically, gets it right with `tool_verified` on metrics frames).
6. **The self-evolution surface has never run.** 13-memory-evolution.png is an empty state under the copy "Your agent improves itself here"; `GET /api/evolution/runs``{"runs":[],"count":0}` on an install with two months of history. The superpower is a promise, not a record.
7. **Automation Center shows stale/past schedules under a "100%" banner.** 06-automation-center.png "NEXT UP" lists 6/12 3:30 AM / 4:00 AM / 5:00 AM — ~16 hours in the past at capture time. `GET /api/automations`: "Prompt optimization" `nextRun: 2026-04-17` (two months stale, `lastRun: null`); "Memory lane extraction" overdue with `lastRun: null`. Never-ran and overdue jobs are invisible to the "Success rate (recent runs): 100%" headline. Also UI says 12 active + 1 paused; the API returns 13 with no enabled/paused field exposed.
8. **Resume card to an empty workspace.** 02-home-cockpit.png: "YOU WERE WORKING ON — New Hive, 10d ago, Continue" for the same workspace the welcome panel calls "Nothing here yet — start a chat and I'll remember it" (0 memories, 0 sessions; `stats?workspaceId=new-hive` → 0 frames). "Working on" should require content.
9. **The only agent is judging-day staging.** Agent Center's "Editorial Critic": `createdBy: "user"`, `createdAt: 2026-06-12T18:59:26Z`, never executed (idle, runs never, avg success "—"). As evidence for "real agents," this is a prop placed on the set an hour before the audience arrived.
10. **Agent-authored skill bypasses the install audit.** `presentation-design` (initiator: agent — the product's best artifact) has no entry in `/api/audit/installs` (8 entries; only `smoke-test-skill`'s creation is audited). The governance trail advertised by the Audit tab doesn't cover the one capability the agent actually authored.
11. **Flaky memory listing.** My first `GET /api/memory?limit=50` returned `{"results":[],"count":0}`; the identical call minutes later returned all 14 frames. Observed once, not reproduced — but if the Memory Center hits this race, the user sees "no memories" in a memory product.
12. **Dedup misses live duplicates.** writer-demo frames 8 and 11 ("I always work with a draft → critique → rewrite loop…") are both `active` (created 05-27 and 06-02). The briefing code works around this with a first-line-hash dedup whose own comment admits "consolidation re-writes the same fact as a fresh frame" — the workaround is in the view layer instead of fixing the store.
---
## Bottom line
This is not vaporware — the substrate (per-workspace SQLite minds, a 214-entity knowledge graph, correctable memory statuses that have actually been exercised, real tool-detection at onboarding, a genuinely agent-authored skill with end-to-end provenance, an audit trail with agent-initiated proposals) is real and verifiable, which is more than most "AI that learns" products survive. But the two superpowers are unevenly proven: **memory** is real machinery presenting staged and self-polluted evidence through a hero panel that showcases its own deprecated frames; **self-evolution** is one real artifact standing in front of an evolution log with zero entries, an agent that has never run, and a self-assessment loop that praises itself on zero data. Ship the substrate's honesty all the way up to the welcome screen and criterion 2 and 3 become 5s. Today, the skeptic's verdict: the receipts exist in the database; the storefront oversells them.

View File

@@ -0,0 +1,91 @@
# Round-2 Verifier Report — commits 72fedf7 + 0ffd938
**Verifier:** fresh-context, 2026-06-12. HEAD at verification time = `0ffd938` (working tree matches the commits under review; only untracked `judging/crops/`, `judging/round2/`).
## VERDICT: PASS
All changes trace to judge complaints or the mission. No new dependencies, no flags/shims, no unrelated refactoring. All four gates green on a fresh run, including the marketplace-sync suite the commit message flagged as flaky. Security posture of the new chat-markdown renderer is sound and test-locked. Three minor, non-blocking notes below.
---
## 1. Scope tracing (every hunk → complaint or mission)
### 72fedf7 (42 files, +1251/202)
| Change | Traces to |
|---|---|
| `LoginBriefing.tsx``lastActive` from workspace store (`ws.lastActive`) not `ctx.lastActive`; honest empty-workspace nudge | Recency contradiction ("active yesterday" vs "away 10 days") — machine cron writes no longer count as user activity |
| `briefing-highlights.ts` — dedup by normalized first line, keep higher-importance then earliest timestamp | Same memory shown twice with two different ages |
| `AppShell.tsx` — LoginBriefing gated to `/home`; `OnboardingTooltips suppressed={ov.showGlobalSearch}` | Modal overlaying Memory/Skills; Ctrl+K tip painting over the open palette |
| `home.ts``SUGGESTION_MAX_IDLE_DAYS = 30` filter on suggested actions | Stale test prompts recommended as today's actions |
| `workspace-context.ts``SYSTEM_JOB_TYPES` filter on `buildUpcomingSchedules` | "Up next" showing the janitor's calendar (verified: "Marketplace sync" and "Index reconciliation" both run as `job_type='memory_consolidation'` per `setup-crons.ts:22,29`, so the filter catches every item the judges named) |
| `render-markdown.ts` + `TextBlock.tsx``renderChatMarkdown` | Literal `## DECISION 1` / `**bold**` noise in chat |
| `login-briefing-brag.ts` — "people, projects & things it knows" | entities/relations jargon |
| `ChatApp.tsx` / `AgentDetail.tsx` — Ask first / Trusted / Autopilot **display labels only**; internal `'normal'|'trusted'|'yolo'` values untouched | YOLO jargon — explicitly *not* a compat shim |
| `dock-tiers.ts` `description` + `AppShell` HintTooltip | Opaque nav labels ("Waggle Dance", "MCP Hub") |
| `activity-labels.ts` (new, 37 lines + test) | `tool_result: create_skill` machine vocabulary in the activity rail |
| `WorkspaceDesktopApp.tsx` — humanize summaries, "1 memory"/"1 session" plurals | Jargon + grammar complaints |
| `AgentsApp.tsx` — empty state lists built-in workspace assistants | "No agents yet" while an agent demonstrably worked (contradiction) |
| `EvolutionTab.tsx` — default filter `'all'` + plain-language primer | Agent-evolution invisibility |
| `AutomationCenterApp.tsx` — Next-up / Recent-results overview panels | Three stat tiles over a void; no answer to "what runs next / how did it go" |
| `skills.ts` + `types.ts` — absent provenance ⇒ `'built-in'` | Unfalsifiable provenance badge (stock skills attributed to the user) |
| `identity.ts` — merge-on-update | Partial identity write wiping stored fields |
| `profile.ts``deleteByContentPrefix('User identity: ')` before re-create | Profile-frame duplication root cause |
| `monthly-assessment.ts` — frames stamped source `'system'` | Provenance lie (agent report stamped user_stated) |
| `ImportStep.tsx` hints, `onboarding-profile.ts` greeting preview | Onboarding export how-tos + "real greeting preview" complaints |
| `HomeCockpit.tsx` `upNext ?? []` | Boundary hardening for an absent field (system-boundary validation — in-spec) |
| `judging/*.md` (5 judge verdicts, round1-fixes, verifier report) | Mission evidence artifacts, not code |
| Test updates (`phase3b/3c` MemoryRouter wrap, copy assertions) | Direct consequence of `useNavigate` in the new AgentsApp empty state |
### 0ffd938 (14 files)
- `truncateHighlight` strips `**`/`#`/`` ` `` tokens — highlights render as text nodes, so raw markdown showed literally (judge complaint).
- `notes/judge-round1-patterns.md` + 13 recaptured screenshots — evidence, in-mission.
### Negative checks
- **No new dependencies:** `git diff 72fedf7~1 0ffd938 --stat -- '**/package.json' package.json package-lock.json bun.lock` → empty. `renderChatMarkdown` is hand-rolled (~40 lines) instead of pulling a markdown lib — consistent with the no-new-deps constraint.
- **No flags/shims:** the only new props/params (`suppressed`, injectable `now` for test determinism) are direct fix mechanics. Autonomy rename is display-only.
- **No unrelated refactoring:** the `escapeHtml`/`applyInline` extraction in render-markdown.ts is the minimal factoring required for `renderChatMarkdown` to reuse the escape-first pipeline; `buildProfilePreview` rewrite *is* the greeting-preview complaint.
## 2. Special-attention items
### renderChatMarkdown security posture — SOUND
- **Escape-first confirmed:** both renderers run `escapeHtml()` (escapes `&`, `<`, `>`, `"`) over the *entire input* before any tag is emitted; block parsing in `renderChatMarkdown` operates on already-escaped lines and routes all inline content through `applyInline(escaped)`.
- **Href allowlist confirmed:** only `^https?:\/\//i` becomes `<a>`; anything else (javascript:, data:, vbscript:) renders as inert `label (url)` text.
- **Tests lock the defenses** (`apps/web/src/lib/render-markdown.test.ts`, read in full): raw `<script>` escaped before tag emission; `javascript:` link refused (asserts no `<a ` emitted); quote-escape blocks attribute breakout (`onmouseover="` absent); raw HTML escaped in every chat line shape (heading and bullet). These are real assertions against the real module, not snapshots.
- **Adversarial probe (no failure found):** a backtick code-span inside a link URL (`` [x](https://a`payload`) ``) yields a malformed `<a>` whose junk attributes come from the *fixed* code-span class string; the attacker payload lands in text position with `<` and `"` pre-escaped — no executable vector. Single quotes are not escaped, but every emitted attribute is double-quoted, so no breakout. Cosmetic quirk only.
- Minor: `data:`/`vbscript:` have no dedicated test case (the allowlist makes them inert by construction; `javascript:` is the representative lock).
### Identity merge-on-update — explicit empty string still clears: CONFIRMED
`body.role ?? existing?.role ?? ''``''` is non-nullish, so an explicit empty string passes through and clears; only *omitted* (undefined) fields fall back to stored values. Test-locked in `packages/server/tests/local/identity.test.ts` ("an explicit empty string still clears a field": writes `{name, role}`, then `{role: ''}`, asserts `role === ''` and `name === 'Marko'`) via real Fastify inject + in-memory MindDB — exercises the route *and* IdentityLayer.
### profile.ts deleteByContentPrefix — does not delete non-identity frames: CONFIRMED
`deleteByContentPrefix` is **pre-existing** (W4.3, `packages/hive-mind-core/src/mind/frames.ts:343`), not added by these commits. It escapes LIKE metacharacters (`\ % _`) and matches the exact literal prefix `'User identity: '` — only frames in the identity-card namespace match. Same established pattern as `monthly-assessment.ts:307`. Routes through `delete(id)` so FTS/vec/KG indexes stay consistent. Residual theoretical risk (a harvested frame whose content *literally begins* with `User identity: ` would be swept) is inherent to the pre-existing primitive, namespaced, and consistent with prior usage — not a regression introduced here.
## 3. Gates (fresh run by this verifier, 2026-06-12 21:3721:40)
| Gate | Result |
|---|---|
| `npx vitest run --root apps/web` | **943 passed (943)**, 91 files, 24.4s |
| `npx vitest run packages/server/tests/local --root .` | **919 passed (919)**, 76 files, 105.9s |
| `npx tsc --noEmit --project packages/server/tsconfig.json` | clean (exit 0) |
| `npx tsc -p apps/web/tsconfig.app.json --noEmit` | clean (exit 0) |
**Marketplace-sync adjudication:** `marketplace-sync.test.ts` **passed 12/12 in my run** (slow — 104s, network-dependent: "graceful errors" / multi-source aggregation cases each take 2036s). Grep of both full diffs for `marketplace`: matches are only (a) a tooltip `description` string on the dock's Marketplace nav entry (UI-only, no runtime logic), (b) commit-message and judging-report prose. **No marketplace server code, routes, sync logic, or test files are touched by either commit — a timeout in that file cannot be caused by these changes.** The commit message's 917/919 claim is consistent with a transient network flake.
## 4. Test spot-checks (read in full, assert the new behavior)
1. **`render-markdown.test.ts`** — 12 tests; XSS locks detailed above plus block rendering (headings→block strongs, bullets, numbered lists, hr, inline-inside-heading, blank-line spacing). Asserts on real renderer output strings.
2. **`identity.test.ts`** — 2 new merge-on-update tests against a real Fastify instance + `MindDB(':memory:')`: partial write preserves `role`/`department` while updating `name`; explicit `''` clears. Exactly the regression the fix targets.
3. **`briefing-highlights.test.ts`** — dedup test feeds two identical-content frames (timestamps 2026-05-01 / 2026-06-11) + one distinct; asserts exactly one survivor carrying the **earliest** (learned) timestamp. Also updated the limit test to use distinct contents so it still measures the limit, not the dedup — correct test hygiene.
Also verified: `p2-home-desktop.test.tsx` (Up next omitted for `[]` *and* `undefined`, rendered with ≥1 item), `phase3b-agent-center.test.tsx` (empty state must show "Already working for you" + workspace names), `p5-skill-governance.test.ts` (no-frontmatter skill ⇒ `'built-in'`).
## 5. Findings (non-blocking)
1. **Test-coverage gaps on new server logic (minor spec deviation):** the 30-day suggested-actions window (`home.ts`), the `SYSTEM_JOB_TYPES` up-next filter (`workspace-context.ts``buildUpcomingSchedules` has pre-existing unit tests that were *not* extended for the new filter), and the `profile.ts` replace-on-update call have **no new tests**. The spec's "add tests for new interactive logic" was honored for FE logic and the identity route, but these three server behaviors ship test-uncovered. Each is a small pure filter / one-line integration over a tested primitive, so risk is low — but the job-type filter in particular is behavioral and cheap to lock.
2. **Stale comment/type nits:** `apps/web/src/lib/types.ts:529` doc comment still says "Absent/legacy ⇒ 'user'" while the union and server now say `'built-in'`; `CapabilitiesApp.tsx:165` inline cast still narrows `initiator` to `'agent' | 'user'`. Runtime is correct ('built-in' flows through; `SkillRow` badges only `'agent'`), tsc is clean — documentation drift only.
3. **Cosmetic renderer quirk:** backtick-in-link-URL produces a malformed (but safe) anchor — see §2. Not exploitable; fix only if it ever surfaces visually.
## Bottom line
Both commits are tightly scoped to the judge complaints and the mission, the three special-attention risk areas (XSS posture, empty-string clear, prefix-scoped delete) all hold under direct inspection and test reads, and every gate passes fresh. The marketplace-sync flake is conclusively unrelated. PASS.