Buteo
Methodology

How Buteo validates a startup idea

Buteo, the AI idea validator, turns a startup idea into testable claims, researches each one on the open web, checks every finding against the page it cites, scores the evidence with a fixed formula, and recommends Continue, Pivot or Kill. A human makes the decision. This page is the whole method, including where it is weak.

By Buteo · Updated

What does Buteo do with an idea?

Buteo, the AI idea validator, treats a startup idea as a set of claims that can each be true or false, and then looks for published evidence on both sides of every one. A validation runs seven steps in order:

  1. Model the idea. Write one or more testable hypotheses for each of 6 lenses.
  2. Research in parallel. Up to 5 research passes search the open web at once, each with its own brief.
  3. Check every source. A finding is compared with the text of the page it cites.
  4. Score evidence quality. A fixed formula, not a model's opinion.
  5. Attack the strongest evidence. A final pass searches for reasons the best-supported claims are wrong.
  6. Write the verdict. Continue, Pivot or Kill, where every confirmed point cites a finding you can open.
  7. Hand the decision to you. The recommendation is Buteo's. The call is yours.

One measured first pass ran 55 searches, kept 52 findings, mapped 8 competitors and reached a first verdict in 13 minutes. That is one real run of one idea, so yours will differ.

How does an idea become testable claims?

Every idea makes a claim on each of 6 lenses, whether or not its author wrote it down. Buteo writes those claims out as hypotheses, and marks the ones the idea dies without as critical.

LensWhat a complete idea has to claim
problemthe pain exists, how often it bites, and why today's workaround is not good enough
marketenough buyers exist to matter, and there is a channel that actually reaches them
competitionwhat the named alternatives leave unsolved, and why a buyer would switch from what they use now
feasibilityit can be built and run by this team, and nothing it depends on (a platform, an API, a data source) can be withdrawn
financialthe buyer will pay this price, and winning and serving them costs less than they pay
riskno law, regulation, platform policy or structural force blocks it

Coverage is checked by counting, not by asking a model whether it did a good job: if any lens has no hypothesis after the first attempt, a second step writes one for exactly the lenses that are missing. The lens left without a claim is the one nothing would have tested.

Who does the research?

5 research passes run side by side. A planner picks which are relevant to the idea, with one restriction that matters: 2 of them are not the planner's to switch off, because a router that has just read a promising idea is the last thing that should decide whether to look for competitors and failure modes.

PassWhat it looks forRuns
Market & web researchMarket size, growth, trends and how buyers in the space are reached.When relevant
Competitor researchWho already does this, what they charge, and what their users complain about.Always
Customer researchWhether real people describe the problem in their own words, and how they cope today.When relevant
Business model researchWhat buyers pay for comparable things, and what winning and serving them costs.When relevant
Risks & contradictionsReasons this kind of idea fails: regulation, platform dependence, and failed predecessors.Always

Each pass writes up to 4 search queries, reads the 20 most relevant sources it gets back, and extracts findings from them. Community sites such as Reddit, Hacker News, G2, Capterra and Trustpilot are searched alongside the open web, because that is where people complain in their own words. No single page may contribute more than 3 findings to a run, so one long report cannot become the whole case.

Three narrower passes follow the first five:

  • Competitor profiles. Up to 4 named competitors are researched individually. A company the model merely remembers, and no searched page describes, is never listed.
  • The gap chase. Up to 3 hypotheses that nothing was found on are searched for again, critical ones first.
  • The claim audit. Described below.

How is a source checked?

A cited link is not proof that the page says what the claim says. Buteo checks each finding against the text of the page it cites, with two tests and no AI model involved:

  1. Every number in the claim must appear in the source. A tolerance of 5% lets “roughly 40%” stand against a page that says 39.6%, and “$1.2M” matches “1,200,000”. A percentage is never satisfied by a count.
  2. At least 50% of the claim's content words must appear in the source.

A finding can fail in two different ways, and they are not the same:

  • The address was never in the search results. The link is removed. There is nothing there to check.
  • The page is real but does not say this. The link is kept and the finding is labelled as unsupported by its source, so you can open the page and see for yourself.

Either way the finding's confidence is capped at 0.3, it scores as having no verified source, and the verdict is not allowed to cite it. It stays visible in the evidence ledger, counted and labelled, rather than being hidden.

How is evidence quality scored?

Every finding gets a quality score between 0 and 1 from a fixed weighted formula. It is deliberately not a model's judgement: quality is the number the verdict leans on, so it has to come out the same on every run and be impossible to argue into a better answer.

TermWeightHow it is measured
Source credibility40%The class of site the finding cites, from the table below.
Stated confidence25%How sure the extracting model said it was, capped at the source's credibility. A claim is not more certain than the page it came from is credible.
Recency20%Full marks inside 2 years, falling to zero at 6. A page with no recoverable date gets the neutral middle.
Corroboration15%Other research passes reaching substantially the same claim from a different page. Two passes quoting one page count as one source.

What each class of source is worth

Credibility is assigned by the class of site, matched on the registrable domain name so that a look-alike address cannot borrow a better site's standing. The tiers are coarse on purpose: they rank kinds of source, and nothing finer.

Class of sourceCredibility
Government0.95
Your experiment0.90
University0.90
Academic or scientific0.88
Research firm or statistics office0.80
Major news or business data0.75
Trade press0.65
Review or reference site0.60
Unrated site0.50
Community or social0.45
Company blog or stats roundup0.40
Self-published0.35
No verified source0.20

A site with an interest in the answer is rated below an unknown one: a company describing itself, a company blog, or a publisher of syndicated market reports. “Your experiment” is a result you report from a real-world test; it is rated alongside the most credible published sources, because it is direct observation of the real customer.

What tries to prove the idea wrong?

Two things, at two different moments. During research, the risk pass looks for reasons this kind of idea fails, and it always runs. After the evidence is scored, a claim audit ranks the supporting findings by quality and searches specifically to break the strongest 5 of them.

The second step exists because the first one cannot see what the other passes found: they all run at the same time. Only a pass that runs afterwards can attack a particular finding rather than the idea in general.

Finding nothing is a real outcome, and the audit is told so. A critic instructed to criticise with nothing to criticise will invent something. On many runs the audit adds no findings at all, and that is the price of the claims having been tested.

How is the verdict written?

The report sorts what was learned into three lists: confirmed, uncertain, and red flags. Each line names the findings it rests on, and one rule is enforced by the code rather than requested of the model: a “confirmed” line with no surviving citation is moved to “uncertain”. A citation survives only if the finding behind it passed source verification.

Confidence is reported per lens and means evidence strength: how well the sources found support the claims on that lens. It is not a rating of the idea. A strong idea with little published evidence scores low, and that is the honest reading.

Whether each hypothesis ends up supported, contradicted, mixed or untested is derived from the findings filed under it, not asserted by the model that wrote the report.

Where does a human decide?

A run stops and waits for a person at four points:

  1. Clarify. A few questions about the idea, before any research is paid for.
  2. Approve the tests. Buteo proposes real-world validation experiments for the questions research could not answer. You choose which to run, or none.
  3. Report what happened. The run pauses, for days if need be, until you bring back what an experiment found. That result enters the ledger as first-party evidence.
  4. Make the call. Continue, Pivot or Kill.

An idea is never killed automatically. Continuing sends the idea back through research with what is still open; the loop has a hard cap on passes and a ceiling on spend, and a run that reaches either goes straight to its report, marked as budget-limited. For how to make the decision itself, see Continue, pivot or kill.

What are the limits of this method?

A method that claims to check sources should say where it is weak. These are the limits we know about.

  • It reads what has been published. If nobody has written about whether your customer will pay, the hypothesis comes back untested. That is an honest answer, and it is what the validation experiments are for.
  • Search returns passages, not whole pages. A true claim drawn from a part of the page that was not retrieved can look unsupported. The check is calibrated to stay silent when it has too little text to judge fairly (under 200 characters), so it misses some unsupported claims rather than accusing supported ones.
  • The source tiers are coarse. A company's own pricing page is scored as a company describing itself, even though it is the best possible source for its own price.
  • Most web pages carry no date. When no date can be recovered, the recency term sits at its neutral middle and tells you nothing.
  • A research pass can fail. A language model sometimes returns something unusable. The run reports the lost pass instead of hiding it, but its findings are missing from that run.
  • The figures on this page come from one measured run. They describe the method's shape, not an average.
  • Verdicts have not been benchmarked. We have not yet scored Buteo's recommendations against a set of ideas whose outcomes are known. Treat a verdict as a well-sourced argument, not as a prediction.

More on what an AI can and cannot settle about an idea: Can AI validate a startup idea?

FAQ

Questions people ask about this

Does Buteo use an AI model to decide how trustworthy a source is?

No. Source credibility and evidence quality are a fixed formula over the class of site, the publication date, the stated confidence and cross-pass corroboration. A model is never asked how much to trust a page, because the answer would not be reproducible between runs and could be talked into a better one.

Can Buteo cite a source that does not say what it claims?

A language model can attach a claim to a real page that does not support it, so every finding is checked against the text of the page it cites: each number in the claim must appear there, and at least 50% of its content words. A finding that fails is labelled as unsupported by its source and cannot be cited in the verdict.

Why does Buteo not give one overall viability score?

Because nothing in the analysis computes one. Buteo scores how strong the evidence is on each of 6 lenses, which is a statement about the sources, not about the idea. A single number would read as a prediction of success, and no step in the method produces a prediction.

Does a Kill recommendation end my project?

No. Continue, Pivot or Kill is a recommendation, and a human makes the decision at a dedicated step. An idea is never killed automatically, and you can choose to continue against the recommendation.

Bring the idea you are least sure about.

Describe it in a paragraph. Your first validation is free, with no card, and every finding links to the page it came from.

Validate an idea