Every AI vendor in preconstruction leads with an accuracy percentage. We do too. Scope Agent runs at 97% match accuracy on scope item extraction against our internal validation dataset.
That number matters. It also isn't enough on its own. A percentage tells you how often a tool is right across a whole document set. It doesn't tell you which lines are the wrong ones. If you can't find them fast, you re-read everything, and the time you thought you bought is gone.
What turns accuracy into hours back is traceability: whether every line points to the drawing or spec section it came from. Ask about that second thing, not just the first.
Sit in on enough evaluation calls and the same sentence keeps coming back. A GC evaluating a hospital project told us the tool looked useful, "assuming it's accurate," and that this was the biggest part of the decision. The instinct is right. The way our industry answers it isn't.
Say a tool reviews a scope package and gets 90% of the lines right. That sounds like it removes 90% of the work. It doesn't. You don't know which 10% are wrong, so you check all of it. Your review time barely moves, and now you're second-guessing a machine on top of doing the job.
An accuracy percentage only converts into time saved if you can find the errors faster than you could have done the work yourself.
Two reasons.
The errors are expensive and they surface late. In North America, errors and omissions in contract documents has repeatedly ranked as the leading cause of construction disputes across the years Arcadis has tracked them, and the average North American dispute value reached $60.1 million in 2024. A missed line doesn't announce itself on bid day. It shows up as a change order.
Fabrications look like real findings. When a model can't find something, it will often produce a confident, correctly formatted line anyway. An invented scope item doesn't arrive flagged as suspect. It looks like every other line on the sheet.
The best measurement of this comes from another profession. Stanford researchers tested the leading AI legal research tools from LexisNexis and Thomson Reuters. Those products work the way ours does, grounding answers in a retrieved document set, and several had been marketed as eliminating hallucinations. Across 202 expert-scored queries they hallucinated 17% to 33% of the time. The errors weren't only invented citations. Real sources got mischaracterized, and inapplicable authority got cited as though it applied.
Grounding a tool in your project set does cut fabrication. It doesn't cut it to zero. Any vendor telling you the problem is solved is overselling, and that includes us. What we can do is show you where every line came from.
Ask for both numbers.
How often is it right? Get the figure, get the dataset it was measured on, and find out whether anyone measured a human baseline on the same exercise. On our internal validation set, Scope Agent hits 97% on scope item extraction and 96% subcontractor classification in the top two. A human estimator on the same exercise came in at 91.3%, and took about four days.
How fast can you find the lines that are wrong? The answer should be a mechanism. In Provision, every scope item carries a trace to the exact page and section that produced it. You click the line and land on the drawing detail or spec paragraph. Verification takes seconds.
If a vendor can't answer the second question, the first one doesn't mean much. Again, you cannot measure accuracy in isolation. You need to understand what recourse the user has when things are not right, just as much as it's important to understand how right the tool is.
A Director of Preconstruction at a mid-market GC told us he'd never want a PM to ask an AI system for a scope of work, print it, and issue it. He's right.
A senior estimating exec at a multifamily GC put the risk more bluntly. Some of what AI does now is borderline scary, because people get dependent on generated output without knowing what they're looking at.
The correct shape of the tool is workflow augmentation. It's a reviewable draft with citations, produced fast enough that a qualified estimator can audit it in a fraction of the time building it from scratch would take. Judgment stays with the person who owns the number. What changes is how much of their day goes to reading instead of deciding.
Volume without prioritization is its own failure. A GC estimator reviewing one of our early conflict outputs got 738 flagged items on a single project, and said the honest thing: discarding the irrelevant ones is easy, but there were a lot of them. Surfacing everything isn't the same as being useful. Ranking by cost impact is the work.
Get this wrong and the number is real. The Scope Gap Playbook, built from 200+ GC interviews, tracks what single misses cost. One example: $300K of lead-lined glass left out of a hospital imaging suite, absorbed by the GC under readily-inferable language. Nobody caught it because nobody could check it quickly.
That last one is the cheapest real test you have, and almost nobody does it. If you want to try it on a live set, book a walkthrough and bring your own documents. Risk Review applies the same citation requirement to contracts and specs, and the Chat Agent answers estimator questions with the source attached.
Run Scope Agent on your own documents and verify findings in seconds, not hours.
See Scope AgentMore Articles