Before you trust a system with numbers you don’t know, it has to reproduce numbers you already do. Here is the test in full — steal it, and run it on every AI vendor who calls you, including us.
By Greg Bush · Updated July 22, 2026 · 6 min read
The Mirror Test is a single rule for buying anything that produces numbers: before you trust a system with numbers you don’t know, it has to reproduce numbers you already do. You pick a baseline you’d stake money on — a bid you priced and won, a job you costed, last month’s orders. The system produces its own version blind, from the same source material. Then your expert compares it line by line. Passing means it gave you back your own truth. Anything else is a demo.
Every AI demo you will see this year looks perfect. That is what a demo is for. So how does an owner actually decide what to trust with their bids, their margins, their prices?
Never confuse the machine being impressive with the machine being right.
The demo shows you impressive. Only the mirror shows you right. This page is the test we run on ourselves, written out in full so you can steal it and run it on any vendor who calls you — including us.
A demo is a performance on material the vendor chose. It answers “can this software produce an impressive-looking output?” — a question whose answer has been yes for every product in the category since about 2023. It does not answer the only question that matters to someone who has to sign the bid: is the number correct?
The failure mode is specific and expensive. Modern systems are fluent, so a wrong quantity arrives looking exactly like a right one — same confidence, same formatting, same speed. There is no tell. You find out three weeks into the job, in the field, when the material runs short. Fluency is not accuracy, and no amount of watching a demo will separate them for you.
The only thing that separates them is ground truth you already own.
Three steps. One: pick a baseline you’d stake money on — something where you know the right answer because you lived it. Two: have the system produce its own version blind, from the same source material, never seeing your answers. Three: compare line by line with the person who knows the work — not “close enough overall,” but every line, with the deltas on the table.
A bid you priced and won. A job you completed and costed. Last month’s food cost. A document set you know cold. The requirement is not that the baseline is big — it is that you know the right answer, because you produced it. Your own past work is the only benchmark that cannot be gamed by a vendor, because the vendor has never seen it.
Same source material — the drawings, the invoices, the order history — but it never sees your answers. This is the step vendors quietly skip. If the system has been shown the target, or if the material is the vendor’s own sample set rather than your documents, you have learned nothing. Your documents, its numbers, no peeking.
Not a headline percentage. Line by line, with deltas visible, and the person who knows the work judging every gap. Some gaps will turn out to be the system’s error. Some will turn out to be yours — that happens more often than people expect, and it is one of the more valuable outcomes of the exercise. Either way, an expert decides. Never the software.
A Mirror Test is not a pilot, a trial, or a paid engagement. It is a comparison against work you have already done, which is why it can happen in a single conversation and why it costs you nothing but the time to dig out the file.
Here are our own results, as published. Each one is a system reproducing a number the client already trusted, checked by the expert who produced the original. Clients are anonymized; the figures are the ones our engagement records support.
| Engagement | The baseline they already trusted | What the system reproduced, blind |
|---|---|---|
| High-end residential painting contractor | An interior bid the estimator had priced himself, and his own manual takeoff for the largest scope item. | The priced bid to 97.5%, with the largest scope item within about 1% of his manual takeoff — produced from a 47-page architectural set for a ~37,000 sq ft residence in under 15 minutes, work that had been a 5–7 day outsourced wait. |
| Mid-Atlantic commercial general contractor | Six sealed bids the GC had already received and evaluated. | A conceptual budget across 21 CSI divisions in under two minutes, landing −0.2% against the sealed-bid aggregate. |
| Family-owned bakery-café (CoversIQ) | Their own register history — 43,501 real orders, at their prices. | A ranked weekly fix list with dollars attached: about $1,900 a month of specific, checkable opportunity, every line traceable back to their own tickets. |
| Scientific research organization | A literature base the team knew intimately — 16,500 papers. | Synthesis across the full corpus with zero uncited claims — every assertion traceable to a source document the researchers could open and check. |
Every engagement above: the client’s expert approves every line, and no staff were displaced. The machine does the reading and the measuring; the person does the deciding.
Notice what these have in common. In every case the impressive part — reading a 47-page drawing set, pricing 21 divisions, processing 43,501 orders, synthesizing 16,500 papers — is the cheap part. The expensive part, and the only part worth paying for, is the check that followed.
The Mirror Test is the first substantive conversation we have with any prospect, and if we can’t reproduce a number you already trust, we don’t propose. No pilot to buy time, no “it will improve with more data.” You shouldn’t buy — from us or from anyone — until somebody passes.
We publish that policy because it costs us something. A vendor who sells on the demo has every reason to avoid a blind comparison against your own records; a vendor who volunteers for one is telling you where their confidence comes from. That asymmetry is the most useful thing on this page, and it works just as well pointed at our competitors as it does pointed at us.
For the reading, often yes — general-purpose models are genuinely good at extracting and summarizing. The gap is verification and structure. A general assistant will give you a confident quantity with no audit trail, no line-by-line correspondence to your own estimate, and no way to tell a right answer from a fluent wrong one. What we build is the decomposition and the checking around the model — the inspection points where a person can catch an error before it reaches a bid. That is the part that takes engineering, and it is the part that makes the output usable.
One thing you already trust. For costing: a bid you priced and won, plus the drawing set it came from. For hospitality: an export of last month’s orders from your POS. For competitive intelligence: just your business name and the competitor that worries you. If you are not sure what qualifies, the rule is simple — bring the file you would use to settle an argument.
The call runs 30 to 45 minutes. Most of that is the comparison itself, because reading the source material is the fast part. You leave with a one-page document showing your number against ours, line by line, with the deltas — whether or not we passed.
We tell you, and we don’t propose. That is the entire point of running the test before anyone signs anything. A failed Mirror Test on a 45-minute call is a good outcome; the same discovery inside a live bid is an expensive one.
No — the method depends on them. The expert is the one who judges every line, and their judgment is what the system is measured against. Across every engagement we have published, no staff were displaced; the time recovered gets redeployed onto work that actually needs a person.
A demo proves a system can look right. Only a blind comparison against your own records proves it is right.
Three steps: a baseline you’d stake money on · the system’s version, produced blind · a line-by-line comparison judged by your expert.
Run it on every vendor who calls you, including us. The ones who decline have told you something.
Our policy: if we can’t reproduce a number you already trust, we don’t propose.
A priced bid, a completed job, last month’s orders — whatever you’d use to settle an argument. We’ll run it blind and show you our number against yours, line by line. If we can’t reproduce it, we’ll say so and we won’t propose.
Practical AI for small and mid-sized business. Live products, verified numbers, and your experts in charge of every decision.