Proof, not promises.
Client engagements and platform builds from the Blackfrog bench. We mark what’s measured versus what’s identified, we don’t post until the work is real, and we don’t name names unless the client opts in.
A five-to-seven-day outsourced takeoff → under fifteen minutes in-house, without changing pricing, scope, or who approves the work.
own manual takeoff
reproduced on a real project
headcount, more leverage
From a stack of drawings to a defensible line-item budget in under two minutes — calibrated against the estimator’s own sealed bids.
all 21 CSI divisions
bids on the calibrated sector
confirms every field
Products we’ve built and run ourselves — same discipline, different rooms. Every one keeps the expert in charge.
CoversIQ
A weekly profitability analyst for restaurants: 43,501 real orders analyzed in a live pilot, ~$1,900/mo of pricing opportunity identified — every number traceable to the query behind it.
Footprint
A local business’s whole digital footprint — and every competitor’s — audited and scored automatically every week, with a human reviewing every finding.
Literature synthesis
~16,500 scientific abstracts turned into fully cited answers in about thirty seconds — zero uncited claims allowed, the researcher stays the validator.
Client studies are anonymized at each firm’s request. Platform metrics are from production systems we operate.
How to read these.
Two kinds of evidence sit on this page and they carry different weight. The client engagements are paid work with an outside party who could contradict us: an estimator’s own priced bid reproduced to 97.5%, a conceptual budget landing within −0.2% of six sealed bids the contractor had already received. Those are the strongest claims we have, and both firms are anonymized at their own request.
The platform builds are products we run ourselves. The numbers there are from live production systems, which makes them real but not independent — we control both the system and the measurement. We label them as bench work rather than client results for exactly that reason.
Across both, the rule is the same one we apply before anything ships: the system has to reproduce a number you already know is true before it is trusted with one you don’t. Where a figure is opportunity identified rather than money banked, the case study says so. Where an aggregate could hide per-line error, it says that too. No engagement on this page displaced a single member of staff.
The next case study could be your numbers.
Bring something you already know the answer to — a priced bid, a completed job, last month’s numbers. The first conversation is the Mirror Test: the platform has to hand you back your own truth before you trust it with anything new. Measured, not projected.
Are these results real, or typical?
They’re real, from specific engagements, and reported as figures rather than promises — client studies anonymized at each firm’s request, platform metrics drawn from systems we operate in production. Every result carries how we know it’s right (reproduction against numbers the client already trusted) and the same guardrail: no staff displaced. Your results depend on your data; that’s exactly what the first Mirror Test conversation measures.
Practical AI for small and mid-sized business. Live products, verified numbers, and your experts in charge of every decision.