This zone is the honest ledger of Intent's evidence base — what we've actually built, what we've actually tested, what the panel review found, where discovery still has to happen, and where the evidence is thin. If the previous three zones describe what Intent is, this zone describes whether we've earned the right to say any of it yet.
Look at those four numbers carefully. 194 internal signals vs. 4 external signals. That asymmetry is still the single biggest gap in Intent's evidence base. The panels flagged it explicitly on 2026-04-09, when the ratio was 43 to 1; it has grown more lopsided since, not less. Three of the four external signals are 2026 convergences from published work (Cagan's context-engineering stack, Block's BuilderBot fleet, the MobAI team-coherence framing): they corroborate the problem shape, but none of them test Intent itself. The methodology has been dogfooded extensively by its author, which proves it runs, but that is not the same as proving it generalizes. Closing that gap requires structured discovery with the target user, and as of this writing it is still the open gate, not yet run. The gate is held deliberately at this boundary and only here: a public generalization claim is irreversible once made, and the operating commitment is to gate the irreversible and select the reversible; reversible internal builds are not gated. When it runs, the result is published here whether the hypothesis survives contact with real practitioners or not.
Goal: 10 structured discovery interviews with staff+ engineers and senior PMs on teams of 2–7 using Claude Code daily. Mom Test protocol, 45-minute format, explicit disconfirmation tracking.
Falsification criterion: Fewer than 7 of 10 participants describe the spec-clarity bottleneck in their own words, unprompted. If we fail that bar, the hypothesis is wrong and we publish that.
Current status: Protocol written. Participant seed list committed (4 named, 6 TBD). Outreach not yet started — discovery is the next gate before any generalization claim (or the worldview doorway's wording) goes live: an irreversible-boundary gate on outward claims, not a pre-build gate on the work itself. Synthesis publishes here when the wave completes.
Every signal captured while building Intent itself. Covers infrastructure gaps, methodology tensions, decisions made and revised, observations about how the loop runs on itself. This is what dogfooding a methodology actually looks like when the dogfood is visible.
The honest caveat: every one of these is internal. They prove Intent runs. They don't prove Intent generalizes. That's what the external discovery wave is for.
The interview wave: 10 structured discovery interviews with practitioners who don't know Intent and have no social reason to confirm its hypothesis. Mom Test protocol. Each interview produces a signal file with verbatim quotes, quality scoring, and an explicit disconfirmation section. That wave has not started.
The directory is no longer empty, though: three convergence proof-points from published work landed in mid-2026 in .intent/signals/external/ (Cagan's context-engineering coaching stack, Block's BuilderBot fleet profile, the MobAI team-coherence thesis). They are labeled as content analysis, not interviews, and they do not count toward the wave. The interview synthesis publishes here when the wave completes, whether it confirms or disconfirms.
Eight independent panels (48 persona voices) reviewed the previous version of this site in parallel. They flagged 10 cross-cutting findings including no target user (6/8 panels), category confusion (5/8), discovery theater (4/8), reader-not-as-hero (4/8), psychological safety never addressed (1/8 but critical). That review is the reason this entire v2-draft exists.
The review document itself was the most valuable deliverable of that session. You're reading the response to it right now.
Traces a single signal (SIG-010 from Ari, the original external voice) through the full loop: signal → cluster → intent → spec → execute → observe → loop closure. Real file paths, real IDs, real timestamps. The most concrete "here's what a real cycle looks like" artifact on the site.
Also a cautionary tale: this is ONE external voice treated as a pattern for a year. The panel review flagged this as the core discovery gap. We're not hiding it.
Running counts of signals, specs, decisions, MCP servers, and knowledge artifacts. Updated automatically from the git-tracked substrate. If a number on this page changes, you can trace it to the commit that changed it.
Currently: 194 signals captured, 45 spec documents, 14 ratified decisions, 4 MCP servers, 6 knowledge artifact types (census 2026-07-19). All internal to Intent building Intent.
Honest admission: we have not yet found a practitioner, team, or use case where the Intent hypothesis was clearly wrong. That is not a strength — it is a sign we haven't looked hard enough for disconfirmation. Confirmation is cheap; disconfirmation is the real signal.
The discovery interview protocol explicitly weights disconfirmation scoring highest. If we run 10 interviews and find zero disconfirmations, that itself is a disconfirmation — it means we did discovery wrong and produced an echo chamber.
Nearest candidate so far: Block's BuilderBot collapses capture surface and write surface into a single Slack thread, at fleet scale. That is a live counter-example to Intent's draft-only-chat doctrine, logged as an external signal in 2026-06. It challenges one doctrine, not yet the core hypothesis, and we are keeping it visible rather than explaining it away.
Hypothesis-level falsification: Fewer than 7 of 10 external discovery interviews describe the spec-clarity bottleneck in their own words, unprompted.
Architecture-level falsification: The P0 hardening backlog (SIG-022, events.jsonl persistence, cross-engagement leak test) cannot be completed in S1 without major redesign — meaning the architecture has deeper problems than the panel review found.
Safety-level falsification: The Psychological Safety Contract v1 cannot be signed by a team adopting Intent because their culture cannot honor the promises — AND this turns out to be the majority case in discovery interviews. That would mean Intent requires a cultural precondition most teams don't have.
Adoption-level falsification: When the rebuilt site re-runs the 8-panel review in S1, cross-cutting findings F1 / F3 / F4 do not drop substantially. That would mean we misunderstood what the panels were flagging and the rebuild didn't address the real problems.
The biggest insight from the 2026-04-09 session wasn't a finding — it was the realization that the panel-review mechanism itself is the most valuable primitive hiding inside Intent. Eight parallel persona panels produced deeper, more structured critique than any single-voice review in ~2 minutes of wall clock time. That's the async feedback loop Intent was supposed to have. We are shipping it as a first-class skill in S1.