Eight independent review panels — Architecture, Product Strategy, Positioning, Strategy & Systems, Discovery & UX, Engineering Operations, Organizational Design, and a first-time visitor — were each asked to read the live site and return a structured critique. The panels ran in parallel without seeing each other's work. This is the synthesis.
"Implementation collapsed; clarity is now the constraint." Six of eight panels called this out as the strongest moment on the site — a genuine Rumelt-style diagnosis naming an observable external shock rather than manufactured urgency.
The weighted trust equation (clarity·0.30 + 1/blast_radius·0.20 + reversibility·0.20 + testability·0.20 + precedent·0.10) is falsifiable and testable. Rare in "AI governance" content. "Math, not politics" — but see the critique.
OTel-native from day one. Intent → trace, Spec → span, Contract → span. Backfill rules on cluster→intent promotion. Real cycle-time histograms. Not bolted on. This is the architectural depth that carries the rest of the site.
walkthrough.html traces SIG-010 through real file paths, real IDs, real timestamps. The one place the methodology feels concrete. Multiple panels pointed here as "the page that convinced me this was real."
Three commands, sixty seconds, real CLI, no API key for the first five steps. The cold visitor said it explicitly: "that's the moment I stopped suspecting this was vaporware." Most load-bearing page on the site.
43 signals, 19 specs, 19 decisions — with git-tracked receipts. Most pitch sites show a product tour; this one says "read our actual signal log." Senge-style learning loop visible. (But see the discovery critique — it's all internal.)
The concept-brief oscillates between "personal operating system," "team operating system," and "business operating system" — in the same paragraph. The pitch addresses senior PMs implicitly but never names them. Positioning found six different category framings across six pages. The cold visitor took 10 minutes to construct a mental model the site should have provided in the hero.
"A real PM leadership tool would know its buyer inside 30 seconds. This site doesn't." — Product Strategy panel (Wille voice)
The site claims continuous discovery via 43 signals, 19 specs, 19 decisions. Every signal is internal — Intent building Intent. SIG-010 ("Ari") is the single external practitioner voice, cited once, promoted to INT-003 as if one conversation were a pattern. This is the opposite of the Mom Test — confirmation bias with a dashboard.
"SIG-010 is a data point, not a pattern. Where are the 15–20 interviews with senior PMs, eng managers, and staff engineers describing the pain in their own words? Where's the opportunity solution tree?" — Discovery panel (Torres voice)
"The operating model for what comes next" (pitch) → "Operating System for AI-Augmented Teams" (concept-brief) → "A personal operating system" (concept-brief, same page) → "Transformation OS" (methodology) → "Four-step loop" (methodology) → "Product portfolio governed by Intent" (products). Dunford's category clarity test: 3/10.
"If the prospect has to do the work, you've already lost. The concept is strong; the category declaration is absent." — Positioning panel (Dunford voice)
Every sentence makes Intent the protagonist: "Intent brings it back," "Intent prescribes three layers," "Intent replaces the ceremony stack." StoryBrand violation. The reader is a spectator watching Intent perform rather than a hero being guided through a journey. No "you" anywhere above the fold.
"Current: 'Agile was right. Your toolchain isn't.' Replace with: 'Your AI writes code in an afternoon. Your team still plans it in a sprint. Intent closes that gap.' Place the reader, not the product, as the hero." — Positioning panel (Miller voice)
D17 cites Argyris and declares "Observe updates Layer 1 (domain understanding), not just Layer 3." But the only mechanism shown is spec-delta detection and signal suggestion. That's single-loop: observed vs. expected. True Argyris double-loop would question whether personas, journeys, or domain models themselves are wrong — and nothing in the loop challenges the premise. The 178-persona voice library is the latent mechanism, but it's framed as execution helper rather than premise-challenger.
"Label one persona pass 'Challenge the Intent' and you have real double-loop learning. Right now it's single-loop dressed in double-loop language." — Org Design panel (Argyris voice)
The 30/20/20/20/10 weighting of the trust formula is itself a political choice dressed as math. Who set those weights? Against what evidence? No calibration study, no sensitivity analysis, no mention of how the weights were validated. Argyris spots this instantly — the governing variable (how to weight trust) is hidden from examination.
"Politics in product decisions exists because stakeholders disagree on values, not because they can't compute weighted averages. A weighting is itself a political choice dressed as math." — Product Strategy panel (Gilad voice)
Four process boundaries for a single-practitioner tool is the opposite of D6's own "defer infrastructure until it's a measurable blocker" posture. This smells like accidental microservices-by-phase, not principled decomposition. Fowler refactor: start as one process with four modules, split only when deployment independence is proven necessary.
"Collapse the 4-server topology to one process with four modules until deployment independence is a proven need. Keep the MCP tool surface identical. Ship the monolith first." — ARB panel (Fowler/Hohpe voice)
No runbooks, no SLOs, no on-call story. events.jsonl is a single point of corruption with no health check, no backpressure, no DLQ. The deployment page literally says "Phase 4: persistence" — meaning the current production path has no durability guarantees. Known ID collision bug (SIG-022) sits unpatched in the signal stream. In-memory state in FastMCP Cloud containers.
"The observability page ships dashboards but not runbooks, SLOs, or error budgets. There isn't a single alert rule on the entire site. If you can't specify SLOs, you can't claim observability." — Engineering panel (Majors voice)
Notice→Spec→Execute→Observe is a rename of OODA (Boyd, 1976), PDCA (Deming), Build-Measure-Learn (Ries), and discovery/delivery (Torres/Cagan). The site nowhere acknowledges prior art — doubly strange given the 178-persona library already contains Torres, Cagan, Deming, and Ries. Intent as *synthesis* is a stronger story than Intent as *invention*.
"Add a Lineage section: OODA, PDCA, Lean Startup, Shape Up, Continuous Discovery, Torres's OST, Gilad's GIST. Credit the ancestors — your own persona library already contains every master this framework descends from." — Product Strategy panel (Cagan voice)
The site never addresses psychological safety once. Trust scoring is literally a scoring system applied to humans' clarity. If my specs keep scoring L1, am I the bottleneck? Signals capture friction with attribution — in a low-trust org this is a dossier mechanism. Who sees my signals? Can my manager read what I noticed? This is a performance-management shadow system waiting to happen.
"Intent is built by engineers for engineers and treats safety as emergent. It isn't. Write a Psychological Safety Contract: who sees signals, what trust scores are NOT used for, how disagreeing with a spec is protected." — Org Design panel (Edmondson voice)
Not a contradiction — they're grading different layers. ARB evaluated the schema; Engineering evaluated the operational surface. Both are right. The gap between them is the work to be done.
The reframe works as a hook but falls apart as a methodology claim. It's a great sentence and a weak taxonomy.
Change work is where methodologies die. The site never names who loses power — which guarantees resistance you never see coming. The Organizational Design panel mapped it:
"I get the thesis: AI made implementation cheap, so the bottleneck moved upstream to clarity. I can repeat it. I'm confused about what 'Intent' actually IS though — is it a framework? A CLI? A philosophy? A SaaS? Getting Started page saved the visit. Three commands, sixty seconds, real CLI, no API key. That's the moment I stopped suspecting this was vaporware. Then I hit Work System and the seven-level ontology knocked me sideways."
Best moment: Getting Started (the conversion point). Worst moment: Work System (the vocabulary cliff). One-sentence summary: "A git-backed CLI and work ontology that replaces tickets and sprints with a continuous signal→spec→agent-execute→observe loop, built for teams where AI writes the code and the real bottleneck is specification clarity."
"For staff engineers and principal PMs on 5–15 person teams using Claude Code daily, Intent cuts spec-to-working-code rework by X%." Everything else is a distraction until that sentence is defensible. Delete "personal OS," "business OS," "transformation OS" — all of them.
Not signals from dogfooding — external practitioners describing pain in their own words, with verbatim quotes. Replace the internal signal stream on dogfood.html with a second stream of external signals. Until this exists, Intent is a personal operating system marketed as a business one.
Current: "Agile was right. Your toolchain isn't." Better: "Your AI writes code in an afternoon. Your team still plans it in a sprint. Here's how to close that gap." Place the reader in the first sentence of every page.
The cold visitor explicitly said Getting Started is where they believed Intent was real. Three commands. Sixty seconds. Real CLI. Put that on the pitch page, not one click deep. Replace "See it working right now →" with "Capture your first signal in 10 minutes →".
Intent as synthesis is a stronger story than Intent as invention. The 178-persona library already contains every ancestor this framework descends from. Showing the lineage transforms Intent from "another methodology" into "the consolidation everyone's been waiting for."
Sequential SIG-NNN IDs will collide in multi-agent federation. This is a latent P0 data-integrity issue sitting in the signal stream. ULIDs or scoped prefixes. Fix this before a second engagement goes live.
Accidental microservices-by-phase. Keep the MCP tool surface identical. Ship the monolith first; extract intent-knowledge only when a team requests it standalone. Matches D6's own "defer infrastructure" posture and halves operational surface.
One page per server: SLI definitions, error budget policy, top 5 failure modes, diagnostic commands, escalation path. Add explicit SLOs for ingest lag, spec-delta detection, and contract verification freshness. Until this exists, "OTel-native" is a schema claim, not a capability.
State: who sees signals, what trust scores are NOT used for (performance reviews), how disagreeing with a spec is protected, how agent-output accountability is assigned. Without this page, Intent is unsafe to adopt in any organization with power dynamics — which is all of them.
Label one persona pass "Challenge the Intent" — Rumelt, Christensen, or Mintzberg asking "should we even want this outcome?" This is how the 178-persona library becomes the double-loop mechanism Argyris demands. Right now it's a marketing claim the tool surface cannot deliver.
| Finding | ARB | Product | Positioning | Strategy | Discovery | Engineering | Org Design | Cold Visitor | Severity |
|---|---|---|---|---|---|---|---|---|---|
| F1 · No target user | CRITICAL | ||||||||
| F2 · Discovery theater (N=1 external) | CRITICAL | ||||||||
| F3 · Category confusion (6 framings) | HIGH | ||||||||
| F4 · Reader is not the hero | HIGH | ||||||||
| F5 · Double-loop is asserted, not built | STRUCTURAL | ||||||||
| F6 · "Math replaces politics" is a tell | GAP | ||||||||
| F7 · 4-server MCP unjustified | GAP | ||||||||
| F8 · No runbooks/SLOs/on-call | CRITICAL | ||||||||
| F9 · Lineage unacknowledged (OODA/PDCA) | GAP | ||||||||
| F10 · Psychological safety not addressed | CRITICAL |
All eight panels agree on this much: the core reframe ("implementation collapsed; clarity is now the constraint") is a genuine insight, the architectural depth (trace model, contract semantics, OTel schema) is real, and the dogfooding discipline is unusual. That's the signal.
The noise: six of eight panels independently flagged that the site has no target user, no category declaration, no external discovery evidence, and makes the product (not the reader) the hero. Two panels flagged the architecture is under-justified and operationally immature. One panel (Org Design) flagged that the methodology's biggest latent failure mode — psychological safety — is never addressed.
The fix is not more content. It is subtraction and sharpening. Pick one user. Delete five framings. Put Getting Started's three commands above the fold. Run ten discovery interviews. Credit the ancestors. Fix the known ID collision. Write the runbooks. Ship the psychological safety contract. That is roughly four weeks of work — and it turns "interesting methodology essay" into "shippable platform with a real buyer."
The raw material is all here. The site is one ruthless edit away from working.