Evaluate the Result, Not the Architecture Diagram
Vendor architecture slides are unfalsifiable, and treating them as evidence is the most common mistake in AI procurement. What counts as evidence, and six questions for a vendor call.
Sit through enough AI vendor demos and you notice they share a slide. It has boxes and arrows. It explains an architecture — a pipeline, a graph, a layered something — and it is offered as the reason the product works.
That slide is unfalsifiable, and treating it as evidence is the most common mistake in AI procurement.
Why architecture explanations are not evidence
An architecture diagram tells you what a vendor believes about their own system. It does not tell you whether the system works, and there is no experiment you can run on a diagram.
Two vendors can show near-identical diagrams and ship products of wildly different quality, because everything that determines quality — the judgement in the implementation, what happens at the edges, whether it degrades gracefully on a codebase that does not look like the demo — is invisible at that altitude. Meanwhile a vendor with an unimpressive diagram and a strong result is the better purchase every time.
There is a second problem. A vendor who explains their mechanism in detail has told you what to build, which means either the mechanism is not the hard part or they have not thought about their own position. Neither inference is flattering, and neither has anything to do with whether the product will work for you.
The useful stance: the mechanism is the vendor's business. The result is yours.
What counts as evidence
Four things, in descending order of what they cost a vendor to fake.
A result you can reproduce. A benchmark with a published configuration — repository, commit, model and version, settings, the full question set, the scoring criteria and the evaluation procedure. If you can run it yourself and get their number, the number is real. If any element is missing, the result is a claim.
A live artefact you can interrogate. Something running, on a system you did not choose, that you can put your own questions to. Screenshots and recorded demos are directed by the vendor. A live thing answers what you ask it, including the questions they would rather you did not.
A negative result they published. This one is disproportionately informative. A company that publishes an experiment that did not work is a company reporting outcomes rather than curating them, and it means the positive results were probably not selected either. Almost nobody does this, which is exactly why it discriminates.
Named limits. "Here is where this does not help" is worth more than a page of capabilities. Vendors who cannot name a limit have either not tested at the edges or are not telling you.
Six questions for a vendor call
Skip the demo for ten minutes and ask these instead.
"What is your published benchmark and can I reproduce it?" Watch what happens when you ask for the commit hash and the exact model version. Those two details are trivial for a real result and awkward for a curated one.
"What was it measured against, and why that comparison?" Beware "the leading solution" and "the industry standard". The defensible phrasing is the highest-scoring alternative in this run, which is a statement about an experiment rather than a verdict on the market.
"Which of your numbers are measured and which are modelled?" Every vendor has both. The good answer distinguishes them without prompting. If a page presents an illustrative projection in the same visual language as a measurement, that is a deliberate choice and it tells you how the rest of the material was assembled.
"Show me a question it gets wrong." Every system has a failure mode. A vendor who can demonstrate theirs has tested at the edges. A vendor who cannot has only run the happy path — or is unwilling to, which is its own answer.
"What happens on a codebase that doesn't look like your demo?" Old, large, polyglot, undocumented, half-migrated. That is what your estate looks like. Their demo repository is not.
"Can I see an answer's evidence?" Every claim should trace back to the source it came from. An answer you cannot check is a second opinion from a source with no accountability.
Run your own test
Vendor benchmarks are run on repositories vendors chose. Yours is the only one whose result matters.
The protocol is not complicated: write fifty questions the way work actually arrives in your team, before you look at any product. Write reference answers from your own code. Hold the model, temperature, output limit, repository and commit constant across every system. Give each tool the context it normally supplies rather than equalising it — deciding what context to supply is the entire product. Score blind, against your reference answers, with the weights written down.
It costs about a day. The decision it settles usually costs six figures over a couple of years, and no vendor page can settle it for you.
The inversion worth internalising
Buyers have been trained to ask how does it work, because for most software the architecture predicted the outcome. For AI systems it frequently does not — the same broad approach produces very different results depending on execution, and the diagram cannot show you execution.
So invert it. Ask what it achieved, under what conditions, measured how, against what, and whether you can check it. A vendor who answers that crisply and declines to explain their mechanism is being commercially sensible and is giving you more to work with than one who reverses it.
Explanations feel like evidence. They are not. The result is the evidence, and it should survive your own test.
Put a number on your own workflow
Bring us the recurring workflow where AI still needs expensive people to supervise, review and correct. We baseline what it costs and put the target in writing before we build. Or connect a repository and see it on your own code first — 1,000 credits free, no card.
Build My Business Case → Start free →Frequently asked questions
How should I evaluate an AI development vendor?
On reproducible results rather than architecture explanations. Ask for a benchmark with the full configuration — repository, commit, model version, settings, question set and scoring — so you can run it yourself. Then run your own test on your own repository, which is the only result that describes your situation.
Why is a vendor's architecture diagram not enough?
Because it is unfalsifiable. Two vendors can show near-identical diagrams and ship products of very different quality, since everything that determines quality — implementation judgement, edge-case behaviour, how it degrades on a messy codebase — is invisible at that altitude. There is no experiment you can run on a diagram.
What are the strongest signals a vendor's numbers are honest?
A published configuration you can reproduce; a live artefact you can put your own questions to; a negative result they published anyway; and named limits. The negative result is the most informative, because a company that reports what did not work is probably not curating what did.
What should I ask when a vendor cites a comparison?
What exactly it was measured against and why. "The leading solution" is a market verdict nobody has earned. "The highest-scoring alternative in this run" describes an experiment, is much harder to dispute, and tells you the vendor knows the difference.



