The better AI gets, the more your expertise matters
AI tools are only as good as the structure they draw from and the people who check their work.

Whether or not your firm has an AI policy, AI is probably already part of your work: the meeting notes summarized this morning, the proposal drafted last week, or the contract a project manager asked a tool to scan for risk. The question is no longer whether AI arrives. It is how we realize its benefits while keeping documents trustworthy, because a construction project still has to be described precisely enough that a dozen parties who never meet can build the same thing.
A word about who I mean by specifier. Not only the dedicated specification professional, but also every architect, engineer, product representative, and owner’s representative who writes, edits, or relies on specifications.
I am not a specifier, but I often hear one concern from members, partners, and the organizations building these tools: that control of authoritative information may drift away from the professionals accountable for it. My argument runs the other way. The better these tools become, the more professional expertise matters.
What’s changed
First, the models guess less. An independent 6,000-question benchmark was launched in November 2025, rewarding accuracy and penalizing bad guesses. Then, the best model scored 4.8 out of 100. Today, the leading models score in the mid-40s, and the rate at which the top model gives a wrong answer rather than admitting uncertainty has fallen from 92 to 51 percent.
None of that is measured on construction documents, and the models still guess about half the time when they do not know.
Second, the tools that perform well do not answer from memory. They retrieve the relevant passage from the project’s own records (drawings, specifications, owner standards, prior decisions) and answer based on it, citing the source. A 2025 peer-reviewed study of this approach in construction management reported answer accuracy near 90 percent when the system finds the source before it writes.[1]
What AI looks like in practice
Adoption is broad and shallow. Deltek reports day-to-day generative AI use at 78 percent of architecture and engineering firms, but near 11 percent for scheduling and estimating. RIBA finds that 73 percent of UK users report better productivity, while only 17 percent say their designs are better for it.[2] The people closest to the documents are the most cautious, and that is not resistance. That is professional judgment doing its job.
Where AI helps most is in finding information: searching for answers that already exist in the project record, comparing drawing revisions and flagging changes for a person to review, testing design options while changes are still low-cost, and delivering cited answers to the field by text message.[3] One general contractor’s rule was three words: trust but verify.[4] Most of these figures are vendor-reported, and the evidence for lower project cost is still developing. What is solid is that experienced people spend less time locating information and more time evaluating it.
Where AI goes wrong
A generative model predicts the most likely next words based on patterns in its training data. It is not checking a fact; it is estimating one. RIBA suggests these models are “incentivized to promote plausibility over accuracy.”[5]
Practitioners and attorneys who review AI-assisted documents report the same short list: plausible-sounding reference standards that do not exist, generic sections that miss project-specific conditions, and no record of who reviewed the output or against what.[6]
A specification also coordinates design intent, owner requirements, codes, adjacent systems, procurement, and risk, so a recommendation can be plausible and still wrong for the project. Review is much more than proofreading.
GIRI estimates that avoidable error adds 10 to 25 percent to project cost, and Arcadis puts the average North American dispute at 56 million dollars and a year to resolve, with errors and omissions in contract documents again a leading cause.[7]
There is also a newer problem beneath the surface. AI tools can break the chain of custody for standards content, the documented trail showing where information came from and who handled it. A tool can retrieve a passage, blend it with an older edition or a forum post, and return a fluent paragraph that reads as authoritative. The user gets the paragraph, but not the trail. Information degrades every time it crosses an organizational boundary, and the reasoning behind a decision is usually the first thing lost.[8]
An AI summary can multiply that loss unless the underlying structure keeps the sources attached.
Why a shared language matters more now
When I say standards, I mean organizing standards: MasterFormat®, UniFormat®, and OmniClass®. They do not say how strong the concrete must be. They say where the concrete lives and what to call it, so the estimator, the submittal reviewer, the facility manager, and now the retrieval tool all look in the same place. Organizing standards structure the information; the project team supplies the data. This is different from the technical reference standards a specification cites, such as an ASTM test method. Shared structure is what makes an AI-generated section reviewable, and success depends on it.
A system that retrieves before it answers can only retrieve what was filed where it belongs and is named as everyone else calls it. Owners are asking for the same thing from the other end: “decision-grade” information that is current, connected, and traceable.[9] A 2026 Suffolk and MIT white paper calls clean, structured project data a precondition for AI’s gains, one that “the industry has not yet met.”[10] That quality is set during design and construction, one classified section at a time. My reading is that AI raises the value of a common vocabulary rather than lowering it.
Why expertise matters
Every technological shift in construction documentation has raised the same question: what should the tool do, and what must remain the responsibility of trained professionals? McKinsey suggests domain expertise is likely to matter more, not less, because the value lies in knowing when to override an AI recommendation.[11]
That is the case for education and certification. CSI’s courses and chapter programs, along with credentials such as the CDT, CCS, CCCA, and CCPR, build the judgment to translate owner objectives into requirements, coordinate drawings and specifications, and decide whether an AI recommendation fits the actual project. The person who signs is still accountable. Training is how that signature keeps meaning something.
Where this leaves us
Use the tools, and start where the evidence is: existing records, package review, and submittal matrices. Treat the output like a draft from a capable intern: useful, fast, and unsigned until a professional has checked the work.
CSI’s job is the same as it has been since 1948: keep the industry’s shared language clear and current, and keep the people who use it well-trained. AI makes that job even more important.
Notes
1Declan Jackson, William Keating, George Cameron, and Micah Hill-Smith, “AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models,” arXiv:2511.13029 (submitted November 17, 2025), abstract and 1, arxiv.org/abs/2511.13029; current scores at artificialanalysis.ai/evaluations/omniscience, accessed September 19, 2026 (GPT-6 Astra 44 at high effort, Claude Fable 5.1 43 at max effort); Artificial Analysis, “Benchmarking GPT-6 Astra,” September 9, 2026, AA-Omniscience section, artificialanalysis.ai/articles/benchmarking-gpt-6-astra (hallucination rate at max effort fell from 92 percent for GPT-5.6 Sol to 51 percent, while accuracy rose four points). For the earlier baseline, see Matthew Dahl, Varun Magesh, Mirac Suzgun, and Daniel E. Ho, “Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models,” Journal of Legal Analysis 16, no. 1 (2024): 64–93, doi.org/10.1093/jla/laae003, which found leading models hallucinated on 58 to 88 percent of verifiable questions about federal court cases.
2Chengke Wu, Wenjun Ding, Qisen Jin, Junjie Jiang, Rui Jiang, Qinge Xiao, Longhui Liao, and Xiao Li, “Retrieval Augmented Generation-Driven Information Retrieval and Question Answering in Construction Management,” Advanced Engineering Informatics 65 (May 2025): 103158, abstract, doi.org/10.1016/j.aei.2025.103158. The RAG4CM framework reported 0.898 answer accuracy and 0.924 top-3 retrieval accuracy.
3Deltek, 47th Annual Deltek Clarity Architecture & Engineering Industry Study (2026), 14–15, info.deltek.com/Clarity-AE; Royal Institute of British Architects, RIBA AI Report 2026 (London: RIBA, 2026), 16, 24, 25, riba.org/work/insights-and-resources/ai-report/. See also American Institute of Architects, “New Research Explores Perceptions and Opportunities of Artificial Intelligence in Architecture,” press release, March 11, 2025, aia.org/about-aia/press/new-research-explores-perceptions-and-opportunities-artificial-intelligence (regular use among architects is near 6 percent; roughly 90 percent are concerned about inaccurate output).
4Trunk Tools, “How Torcon Project Teams Save Time and Reduce Risk,” customer case study, March 12, 2026 (about 1,100 questions, an estimated 453 hours saved, and about 43 minutes saved per submittal review), and “How AMLI Residential Builds Smarter with Trunk Tools,” August 13, 2025, both published at trunk.tools; Autodesk, “Stantec | Autodesk Forma Site Design,” customer story, autodesk.com/customer-stories/stantec-forma-site-design-story, and Autodesk Forma blog, October 22, 2025, on the Hamdan Bin Rashid Cancer Hospital, Dubai; Get It Right Initiative, Artificial Intelligence and Error Reduction: The Opportunities and Challenges (July 2025), 14 and 18, getitright.uk.com/live/files/reports/14-giri-ai-and-error-reduction-report-15-07-25-334.pdf. The Trunk Tools and Autodesk figures are vendor-published customer estimates, not controlled studies.
5Matthew Thibault, “‘Trust but Verify’: How Novo Construction Compares Drawing Packages with AI,” Construction Dive, August 5, 2026, constructiondive.com/news/novo-construction-buildcheck-drawing-diffs-ai/826758/. Quotation from Colin Stoner, chief information officer, Novo Construction.
6Royal Institute of British Architects, RIBA AI Report 2026 (see note 3), 34.
7See, for example, Gotlaw STL, “When ‘Smart’ Tools Make Dumb Mistakes: AI’s Hidden Risks in Construction Documents,” gotlawstl.com/when-smart-tools-make-dumb-mistakes-ais-hidden-risks-in-construction-documents/, and Lucent Doors, “The Limits of AI in Architectural Door Specification,” lucentdoors.com/post/limits-of-ai-in-architectural-door-specification.
8Get It Right Initiative, Artificial Intelligence and Error Reduction (see note 4), 8 (root causes of error) and 32 (10 to 25 percent of project cost); Arcadis, Disruption to Innovation: Disputes Amid Technical & Economic Shifts, 16th Annual Construction Disputes Report (2026), United States findings: average dispute value $56.0 million and average duration 12.2 months in 2025; errors and/or omissions in contract documents and failure to understand or comply with contractual obligations ranked as the two leading causes, as in 2024.
9Keith Robinson, “The Owner’s Premium: Why Construction Contingency Reached 20%,” The Construction Standard, September 14, 2026, theconstructionstandard.com/blog/the-owners-premium; and TCS Team, “Measuring the Cost of Disconnected Information in Construction,” The Construction Standard, September 9, 2026, theconstructionstandard.com/blog/cost-of-disconnected-information.
10Tango Analytics, Portfolio Under Pressure: The 2026 Corporate Real Estate Decision Readiness Index (2026), 3, 10–11, and 18, tangoanalytics.com/landing/portfolio-under-pressure/; survey of 91 senior leaders at enterprises with revenue above $500 million.
11Suffolk and MIT Center for Real Estate and MIT Media Lab City Science Group, Construction in the Age of AI: An Industry White Paper and Research Roadmap (September 16, 2026), 22 and 27, suffolk.com/news/construction-in-the-age-of-ai/. The paper also asks manufacturers for machine-readable product data with classifications “using consistent industry definitions.”
12Daniel Ahmoye, Erik Sjödin, and Jose Luis Blanco, “How AI Is Reshaping the Future of the AEC Industry,” McKinsey & Company, July 15, 2026, mckinsey.com/industries/engineering-construction-and-building-materials/our-insights/how-ai-is-reshaping-the-future-of-the-aec-industry.
Author
Mark Dorsey, FASAE, CAE, has served as CEO of the Construction Specifications Institute (CSI) since 2015.

