Built around The Mom Test's argument that most customer interviews are worthless for two detectable reasons: leading questions, and answers that predict future behaviour instead of reporting past behaviour. The model proposes each follow-up and the rule gets a veto — because a model asked to run a discovery interview drifts toward exactly the politeness bias the method exists to defeat. Every answer is graded evidence, mixed or hypothetical by pattern, not by asking a model to mark its own homework, and a stage counts as covered only when something in it was a real memory.
The generation and the judgement are separated on purpose. A model asked to run a discovery interview drifts toward leading questions, because leading questions feel helpful and produce pleasant answers — the exact bias the method exists to defeat.