Provo, Utah — September 16, 2026
On Constitution Day, observed on September 17, the day in 1787 that delegates to the Constitutional Convention signed the document in Philadelphia, BYU Law is giving researchers a significantly larger body of historical evidence to examine one of the most difficult questions in constitutional interpretation: What did the words of the Constitution mean to the people who wrote and ratified it?
Today the law school announced the release of a new searchable corpus containing nearly 7 million words from more than 14,000 primary-source texts drawn from The Documentary History of the Ratification of the Constitution (DHRC) — the sprawling collection maintained by the Wisconsin Historical Society and the University of Wisconsin-Madison's Center for the Study of the American Constitution.
The new material expands BYU Law's Corpus of Founding Era American English (COFEA) by roughly 20%. It's available now through the school's Law and Corpus Linguistics (LCL) platform at lawcorpus.byu.edu.
“These additions help researchers better understand the language and arguments that shaped the Constitution's adoption,” said David Armond, Assistant Dean for IT at BYU Law. “By making these sources searchable and accessible, we're providing scholars, attorneys and judges with a richer picture of how constitutional language was used and understood during one of the most important periods in American history.”

The project also illustrates a more contemporary question facing law schools and the legal profession: How should artificial intelligence be used when the person relying on it still needs enough expertise to know whether it is right?
That question arose directly during the creation of the new corpus.
A larger record of ratification
When BYU researchers began developing COFEA, one of their challenges was finding enough digitized founding-era material with reliable text and OCR. For the ratification period, one of the principal sources available to researchers was Elliot's Debates, a 19th-century compilation of debates from the state conventions that considered adoption of the Constitution.
It was useful, but it did not capture the full range of public debate surrounding ratification.
The Documentary History of the Ratification of the Constitution, a 41-volume project produced by the Wisconsin Historical Society and the Center for the Study of the Constitution at the University of Wisconsin-Madison, offered a much more comprehensive source.
Armond sat down with TechBuzz recently describing how about 30 BYU Law students worked through the documentary history, identifying actual primary-source documents and separating them from the substantial editorial material included in the volumes.
The resulting collection contains more than 14,000 texts and nearly 7 million words, expanding COFEA by roughly 20 percent. According to BYU Law, the DHRC material contains more than 17 times the amount of primary-source text found in Elliot's Debates.
The additional material also introduces voices that were less visible in the earlier corpus, including newspaper editors, pamphleteers and political advocates.
That diversity matters because corpus linguistics is designed to study language as it was actually used rather than relying solely on dictionaries, isolated quotations or modern interpretations.
A corpus functions much like a database of naturally occurring language. Researchers can search for a word or phrase, examine how frequently it occurred, study the words that surrounded it and compare usage across different collections.
For constitutional research, that provides an empirical way to investigate how particular words were being used during the period in which the Constitution was written and ratified. Researchers can then compare the ratification-era material with other founding-era sources to see where patterns converge or diverge.
BYU Law launched its corpus linguistics platform in 2016. Courts have subsequently cited COFEA in constitutional and statutory interpretation cases, including United States v. Escobar-Temal, in which Sixth Circuit Judge Amul Thapar relied on COFEA to examine 18th-century language usage.
The new corpus gives researchers another, much larger collection with which to conduct that kind of analysis.
AI helps find what humans might miss
Creating the corpus also presented a problem that would be familiar to anyone who has worked with a large historical archive: human beings make mistakes.
The students were going through thousands of pages, often distinguishing primary documents from editorial introductions and other material that was not supposed to be included in the corpus.

BYU Law students developed local AI models to help identify material whose language patterns suggested it might be modern editorial content rather than an authentic 18th-century source.
The important distinction is that the AI did not make the final decision. Rather, it was used to flag potential problems. Students then reviewed those flags against the original material. Armond said the researchers could be confident in the instances the system identified and humans verified, but they could not know how many errors the AI failed to detect.
That makes the project an unusual example of AI-assisted research: The technology was deliberately used in the negative. Rather than asking AI to determine what was correct, the researchers asked it to help identify what might be wrong. The underlying documents remained the authority.
BYU Law also linked corpus entries directly to the University of Wisconsin's digital library, allowing researchers to move from a search result back to the complete original source.
That emphasis on traceability and verification is central to the methodology. It also connects directly to Armond's broader concerns about AI in the legal profession.
The expertise problem
Armond describes the issue partly in terms of agency: When a professional delegates a task to AI, how does that person know the work was done correctly?
He does not see the answer as avoiding AI. Instead, he believes professionals need to understand which tasks AI can perform reliably, which require review and how to recognize when something is missing.
Armond has a framework for explaining the distinction, borrowed from a conversation with Utah artist Brian Kershisznik. Kershisznik once told him that becoming a good painter is roughly a twenty-step process. AI tools are very good at the first eleven of those steps, but not the last nine, the ones that require judgment, experience, and the kind of feel a machine can imitate but never quite originate.
"Humans have judgment and experience and emotions that the AI doesn't have," Armond said. "That's why most people can still detect AI-generated art versus" an original.
Law has a similar dynamic, Armond argues. AI can plow through the tedious, high-volume reading that law students have always had to slog through, but if students never build the underlying expertise themselves, "their agency will be given to the tool entirely."
If law students allow AI to perform too much of the underlying work, they may become proficient at receiving answers without developing the expertise needed to recognize when those answers are wrong or incomplete.
“The problem is that you need expertise at the end to evaluate whether or not AI actually did it correctly.”
That concern has implications beyond the classroom.
A Utah company doing it the hard way
Armond pointed to Salt Lake City-based Filevine, a legal tech company extensively covered by TechBuzz, as a company he thinks is applying that same philosophy well in a commercial legal-research product. In a recent demo, he said, Filevine's system was built to let AI accelerate the tedious parts of legal work while keeping an attorney visibly in control, showing exactly what the AI reviewed, why, and giving the attorney the ability to redirect it.
"Their system actually is uniquely suited to employ AI in a really powerful way," Armond said, calling it "an outstanding way for you to maintain control, maintain professional identity, and maintain skill."
That description lines up with what Filevine has already shipped. The company's Depo CoPilot tool, launched in 2025 as part of Depositions by Filevine, acts as an AI "second chair" during depositions. It tracks an attorney's stated goals in real time, flagging inconsistencies in testimony, and suggesting follow-up questions, without ever taking the questioning out of the attorney's hands. It's the same pattern Armond described in his own project: AI does the flagging, a human decides what to do with the flag.
More recently, Filevine's CEO Ryan Anderson, a BYU Law alumus, has talked has talked publicly about leading a broader overhaul of the company's AI architecture, after noticing competitors pulling ahead and growing uneasy with his own engineers' explanations for the gap — "when I hear gobbledygook explanations, even from a technician, I get very nervous," he said at a recent Reference Group event.
Armond also noted a pattern he's seen across the major legal research platforms broadly — Westlaw, Lexis, Bloomberg — nearly all of which have converged on citation verification as their primary guardrail against hallucination. Even so, he's found real limits: AI tools are excellent at surfacing a state's current, heavily used statutes, but weaker at catching older, still-binding law that regulators never got around to consolidating. "AI finds the most current stuff and the big stuff," he said. "It struggles to find that other stuff... it doesn't understand how humans work. We're illogical."
For Armond, that kind of human oversight is particularly important in legal education.
The question isn't whether a law student can get an answer from AI. It is whether the student knows enough law to recognize what the answer may have missed.
Armond described seeing AI research tools identify an important state statute while missing related provisions elsewhere in the code. An experienced lawyer might recognize that something is missing and know where to look. A less experienced researcher may not even realize that the answer is incomplete.
That is one reason, he said, law schools cannot simply teach students how to prompt AI systems. Students still need to develop the underlying expertise that allows them to evaluate the result.
Evidence versus answers
That distinction is particularly relevant to corpus linguistics.
A generative AI system can produce an answer to a question about what a constitutional term meant. But the researcher may not be able to reproduce exactly how the system arrived at that answer or determine which underlying sources drove it.
Corpus linguistics takes a different approach.
The researcher can document the search, examine the underlying language, see the frequency and context in which a term was used, and follow the result back to the original historical document.
That does not make corpus linguistics infallible. Armond said both AI and corpus linguistics can be used poorly. But the methodology provides a level of transparency that allows researchers to examine and challenge the evidence.
For constitutional interpretation, that distinction can be significant.
Judges generally do not introduce new evidence into a case; the parties do. But once corpus evidence is before a court, a judge can use it alongside other sources to investigate ordinary meaning or historical original public meaning. Armond said most judicial applications he has seen involve questions of definition, sometimes examining the words that modify or limit a particular term.
The new DHRC corpus gives lawyers and researchers a much larger pool of historical language from which to conduct those inquiries.
A broader research question
The Constitution Day release is therefore both an expansion of an existing research resource and part of a larger conversation at BYU Law about how technology should be used to understand legal meaning.
Armond would like to see BYU Law eventually host a conference examining AI and corpus linguistics as different approaches to determining meaning, potentially alongside other social-science methods such as surveys.
The goal would not necessarily be to determine which method is superior, but to understand where each is reliable, where each has limitations and what standards lawyers and judges should use when evaluating their results.
The law school is also looking ahead to other historical corpora. One potential project would focus on the Reconstruction era and the language surrounding the 13th, 14th and 15th Amendments. A preliminary version was previously made public but was withdrawn after researchers determined the source material was not sufficiently balanced.
For now, the Constitution Day release gives researchers a significantly larger window into the language surrounding the Constitution's ratification.
And the way BYU Law built it may offer a lesson about AI that extends well beyond historical research.
The researchers used AI to make a difficult, repetitive task more efficient. But they did not delegate the final judgment to the machine. They retained the human expertise necessary to examine the evidence, verify the result and understand what the technology might have missed.
For Armond, that may be the more important lesson as AI becomes increasingly capable of doing professional work. The value of expertise may increasingly lie not in doing every task yourself, but in knowing when the machine has done the task correctly.
Learn more at lawcorpus.byu.edu.