Guide · Evaluation

How accurate is AI plan review?

An accuracy percentage is only useful when you know what was tested, which issues counted, and how the results were verified. Evaluate a review tool on a defined set of drawings and a record your reviewers can inspect.

Measure useful findings, missed issues, evidence quality, and verification effort separately. A long list of findings can contain duplicates or false alarms; a short list can omit important problems. Neither the list length nor the confidence of the writing measures accuracy.

Define the test before running the review

Choose the project types and document quality your team actually handles. Record the issue date, disciplines, page count, references, and review instructions. Keep the same input package when comparing repeated runs or different review methods.

Write down what counts as an in-scope issue. For example, a pilot may cover inconsistent equipment requirements across plans and specifications while excluding calculation checks. Without that boundary, a missed issue and an out-of-scope issue are easy to confuse.

Include more than a clean, digitally exported set if your normal work includes scans, addenda, or partial consultant packages. Record results by document condition so a strong result on one package is not presented as universal performance.

Create an independent reference list

Have qualified reviewers record known issues before they read the AI output. For each issue, include the location, source pair, reason it matters, and the context that establishes it as a real problem. This becomes the reference list for the pilot.

A human review is not automatically exhaustive. If the AI finds a new valid issue, add it to the adjudicated record and explain how it was verified. Keep the initial list and the final record distinct so the evaluation remains understandable.

Match issues by the underlying problem, not by identical wording. Three findings describing the same conflicting equipment duty should not count as three independent successes.

Separate precision from recall

Precision asks how many reported findings are valid. Recall asks how many known in-scope issues the review found. These answer different questions, as explained in Google’s guide to classification metrics.

Illustrative pilot only—not Groundbook performance data
MeasureExampleWhat it tells you
Precision12 valid findings out of 20 distinct reported findings = 60%How much of the output was useful after verification.
Recall against the reference list12 known issues found out of 15 in-scope reference issues = 80%How much of the known problem set was detected.
Missed reference issues3 of the 15 known issues were not foundWhich failure patterns need further review.

Define duplicate handling and uncertain findings before calculating these numbers. Do not describe recall against a limited reference list as proof that the tool found every actual error in the project.

Check the evidence, not just the conclusion

For each finding, inspect whether the cited sheet and location are correct, the cited text or table value was read correctly, and the linked sources describe the same condition. Then assess whether the reasoning follows from those sources.

  • Correct source, wrong interpretation: the equipment value is real, but it describes a different operating mode.
  • Wrong source relationship: the finding compares two similarly named rooms in different areas.
  • Missing context: an exception, detail, or current revision resolves the apparent conflict.
  • Unsupported conclusion: the issue sounds plausible but the cited documents do not establish it.

Code findings also need an applicability check: edition, amendments, occupancy, construction type, scope, and exceptions. A real code citation can still support the wrong conclusion for a particular design.

Measure reviewer effort and important misses

Track the time from opening a finding to reaching a review decision. Separate time spent navigating evidence from time spent on the engineering or design question. Count duplicate findings and questions that cannot be resolved from the supplied package.

Review high-consequence misses separately. An average across many minor findings can hide a serious coordination gap. Describe the actual consequence and affected scope instead of relying only on the software’s severity label.

A practical evaluation record includes: finding ID, source location, issue type, validity decision, duplicate group, severity after review, verification time, and reviewer notes. Keep unresolved items visible rather than forcing them into “correct” or “incorrect.”

Report a result your team can trust

Publish the scope and limits beside any metric: the number and types of projects, document conditions, issue categories, evaluation date, and reviewer method. State whether the reference issues were known in advance and how newly discovered issues were handled.

Groundbook’s published case studies report potential findings from particular reviews. Their totals are not a validated precision or recall score. Use the selected sample findings to understand the output format, then evaluate the software on your own representative package.

Repeat the pilot when a meaningful input or product change could affect the result. Keep the same recorded method so your team can tell whether the change improved useful detection or simply produced more output.

Questions, answered

FAQs

Does a higher finding count mean a more accurate review?

No. The count can include duplicates, false alarms, and minor observations. Check validity, coverage of known issues, evidence quality, and reviewer effort.

Can we measure recall without knowing every error in a project?

You can measure recall against a defined, independently reviewed reference list. State that limit clearly; the reference list may not include every actual issue.

Bring your next set into focus

Set up a review or talk through your project with our team.