Transactional lawyers are trained to distrust unsupported assertions. When a counterparty proposes language and says "this is standard," the response is usually "standard where?" The same scrutiny should apply to AI-generated review comments. When the system flags a clause, the question "based on what?" has a real answer, and that answer matters.
In August, every flag and every drafted clause carries a citation. Not a vague reference to "firm precedent" or "comparable practice," but a specific reference to the deal from which the comparison derives: the matter name, the transaction type, the date it closed, and the specific clause provision that is being held up for comparison. The attorney can open that source and read it. This is not a minor product feature. It reflects a view about what AI output in legal work should be accountable for.
What Uncited Review Comments Actually Tell You
Most AI-assisted review tools produce comments that describe what they think is wrong without explaining how they reached that conclusion. "This indemnification clause is broader than standard." "Consider revising this limitation of liability." "This notice provision may not be enforceable as written."
These comments have a form that resembles legal analysis. They identify a clause, characterize a potential issue, and suggest a response. What they do not provide is the reasoning that underlies the characterization. Is the indemnification clause broader than what aggregate public data suggests is typical? Broader than what this firm has accepted before? Broader than what is appropriate given the risk profile of this particular client and transaction? The comment does not say.
An attorney receiving such a comment has three options: accept it on the basis of the system's apparent authority, reject it based on their own judgment, or verify it independently. Independent verification means pulling comparable past deals, which takes time and may not be feasible in the window available for review. In practice, time pressure often means the first two options dominate, and neither is satisfying.
Citation Changes the Verification Task
When August flags an indemnification clause, it says something like: "This indemnification scope extends to consequential damages. In 6 of 8 comparable deals in your bank, your firm used a mutual exclusion of consequential damages. See Meridian Capital APA 2024, Northgate Services MSA 2023, and four similar matters."
The attorney reviewing that comment now has a specific task. They can open the Meridian Capital APA and the Northgate Services MSA. They can read the actual indemnification provisions in those deals. They can assess whether those deals are in fact comparable to the current transaction in the ways that matter for this clause. They can decide whether the pattern holds in this context or whether this deal's specific risk allocation calls for a different approach.
This is different from evaluating an unsupported flag. The attorney is not asked to accept a characterization from an opaque system. They are asked to evaluate a specific comparison between the current draft and identified prior deals, using their own legal judgment. The AI's role is to surface the comparison, not to make the legal determination.
Citation as a Check on the System Itself
There is another function citation serves that is less obvious but equally important. When the system is wrong, citation makes the error visible and correctable.
AI systems make mistakes. They misclassify transaction types. They pull the wrong comparable deals for a given context. They find patterns in the precedent bank that are artifacts of how past deals were filed rather than genuine precedent for the current matter. Any system sophisticated enough to be useful is sophisticated enough to be wrong in interesting ways.
If August flags a clause and cites three past deals as the basis for the flag, and an attorney looks at those three deals and determines that two of them are not actually comparable to the current transaction for specific, identifiable reasons, that is useful information. The attorney can disregard the flag, and that feedback can improve how the system weights comparisons in subsequent matters. The error is surfaced rather than silently accepted.
By contrast, when an uncited system produces a flag that is wrong for subtle reasons, the error is much harder to detect and correct. The attorney either accepts a bad flag or overrides it on their own judgment, and in neither case is there a mechanism for the error to be identified and addressed. The system's mistakes are invisible because its reasoning is invisible.
What This Requires of the Precedent Bank
A citation is only as useful as the underlying source. For August's citations to be meaningful, the precedent bank has to contain deals that are actually comparable to the matters being reviewed, indexed accurately enough for the system to identify relevant ones, and accessible in a form the attorney can read and evaluate.
This has implications for how firms structure their precedent banks. A bank that contains only executed final agreements, without the negotiation history that led to them, provides a partial picture. A bank where deal metadata is inconsistent, where similar transactions are categorized differently by different practice groups or by different attorneys within the same group, produces noisier comparisons. A bank where old deals are mixed indiscriminately with recent ones, without date filtering, may produce citations to past practice that the firm has since moved away from.
These are solvable problems, and we work with firms on how to structure their ingest process to address them. But it is worth being honest that the quality of the citations depends on the quality of what the bank contains. A citation to a poorly categorized or outdated precedent is less useful than one to a well-characterized recent comparable.
The Broader Question of AI Accountability in Legal Work
The question of citation in AI-assisted legal work connects to a broader question about accountability. When AI output influences a legal document, and something in that document later causes a problem, the question of why the output said what it said has real consequences. A system that produces unsourced recommendations has no accountability trail. A system that cites its sources at least makes the reasoning available for examination.
We are not suggesting that AI-generated review comments with citations are immune to error or that citation alone is sufficient accountability. The cited precedent may be wrong for the situation. The comparison may be superficially similar but substantively different in a way the system did not detect. Citation is a necessary condition for accountability in AI-assisted legal review, not a sufficient one. What it does is ensure that when something goes wrong, there is a basis for understanding how and why, and that basis was visible at the time the decision was made.
Attorneys who have seen how hallucination in generic AI tools plays out in practice tend to find this point compelling. A confident recommendation about legal language, unsupported by any traceable source, is not professional-grade output in a field where "based on what?" is a foundational question. August's answer to that question is a specific citation the attorney can examine and evaluate. That is the baseline we think AI-assisted legal review should meet.