Adapted from StartupAI source material dated March 1, 2026. This note explains the product judgment, not internal implementation details.
Source material: ADR-017
Opening thesis
Most validation tools are strong at the two ends — they will make you a plan and they will score your results. The hard part is the middle: actually getting real evidence out of the world. We decided not to hand you a plan and a grade and leave you alone for the part that matters most. Here is why.
The hollow middle
When the middle is vague, the final output is suspect. A tool can claim a recommendation is evidence-based, but if you cannot see how the evidence was supposed to be gathered, whether it was complete, or whether anyone noticed the quality problems before scoring — you are trusting a verdict with no visible chain behind it.
And it is a bad deal for you specifically. The system holds all the context — your assumptions, your target, your evidence gaps — and then asks you to go do the hardest, most method-dependent part on your own, only to reappear and judge the result. You came here because you were not sure how to validate. Being sent off to validate alone is the one thing that should not happen.
We give you the instruments, not just the assignment
So we treat collection as a real part of the product. The system turns your approved plan into the actual tools you need — interview scripts, survey drafts, research briefs, observation guides — each one tied to a specific assumption it is meant to test, so you always know why you are collecting it. You review that kit before you go out; it is your context, not ours, that has to be right.
When evidence comes back, intake is guided to catch the usual problems — missing sources, material dated before the plan, notes that do not connect to any hypothesis. And before anything gets scored, a quality pass flags the weak spots: thin samples, missing coverage, a pile of opinions with no observed behavior underneath.
Straight talk on where this is today: the first version leans on you running the conversations with the instruments we prepare, with more hands-off and expert-run help arriving in stages. The diagnostics flag problems; they do not silently fix them for you. We would rather show you a documented weakness than hide it inside a confident score.
Check the evidence route
Before you trust any discovery output, ask whether the evidence route was visible. Did the tool show what needed collecting? Give you the instruments? Track what came back? Flag the gaps before recommending anything?
Treat collection quality as part of the result, not a footnote. Five rushed chats, a stale spreadsheet, or a survey aimed at the wrong people should not carry the same weight as well-scoped evidence tied to a real assumption. The honest version of no is not “the system says no” — it is “here is what would make yes credible.” That is still useful, even when the answer is not ready.