Verifying AI Output Before It Reaches a Client or a Court
Risk Management

Verifying AI Output Before It Reaches a Client or a Court

A model can produce a citation that looks perfect and points at nothing. Here is the verification workflow that catches it, who signs off, and what the matter file should record before anything leaves the firm.

SGSagnik G.

Most firms I talk to have already crossed the line on this without ever making a decision about it. Nobody sat in a partner meeting and voted to start drafting with a language model. What happened instead is that an associate used one to summarise a deposition transcript at eleven at night, it worked, and six weeks later three people are quietly using it for first drafts of motions, client update letters, and discovery responses. There is no policy, no verification step, and no record anywhere in the matter file that any of it happened. The work product looks completely normal from the outside, which is exactly the problem.

The risk here is not that AI writes badly. The risk is that it writes confidently. A model will produce a case citation with a plausible reporter volume, a plausible page number, a plausible year, and a holding stated in the same tone as every real citation in the same paragraph. It will summarise a contract clause that does not exist in the version you uploaded. It will state a statutory deadline that is correct in one state and wrong in yours. None of these failures announce themselves. They read as fluent, professional, and finished, which means the only thing standing between them and a filed document is whether somebody checked.

So this post is about the checking. Not the policy language, not the ethics opinions, not whether your bar has issued guidance yet. The concrete workflow: how citations get verified, what grounding a draft in a source of truth actually means in practice, who signs off, and what the matter file should record so that if anyone asks in two years what your firm did, the answer is not a memory. I have watched firms build this well and I have watched firms discover they needed it the hard way, and the difference is almost never intelligence. It is whether the verification step had a place to live.

Why unverified AI output is a competence problem before it is anything else

Competence duties across common law jurisdictions generally require a lawyer to bring adequate knowledge, skill, thoroughness and preparation to a matter, and in most places that duty has been read to include understanding the technology you use in the practice. The specific wording differs across the US, England and Wales, the Canadian provinces and the Australian states, so confirm what your own regulator says. But the underlying idea is consistent enough to build a workflow around. If you sign a document, you are representing that you exercised professional judgement over its contents. Delegating the drafting to a tool does not move that responsibility anywhere.

What makes this different from delegating to a junior is the failure mode. When a first year associate is out of their depth, the draft usually shows it. The reasoning wanders, the structure is off, the citations are thin, and a partner reading it can feel the weakness before they can articulate it. A model produces the opposite signal. The prose is clean, the structure is conventional, the transitions are smooth, and the fabricated authority sits in the middle of it wearing the same clothes as the real authority. Your instinct for spotting weak work has been trained on a correlation between confidence and correctness that no longer holds. That is the part firms underestimate, and it is why the verification has to be procedural rather than intuitive.

!
The signal you rely on is gone Weak human work usually reads as weak. Weak AI output reads as polished. Any verification process that depends on someone noticing that a passage "feels off" is already broken, because fluency is exactly what the tool is best at producing.

Candour to the tribunal is a separate duty, and it fails separately

Competence covers whether the work was done properly. Candour covers what you tell the court about it, and the two fail at different moments. A fabricated citation in a brief is a competence failure at the drafting stage. It becomes a candour problem the moment someone at the firm learns the citation is bad and the filing stays on the record without correction. Those are different obligations with different consequences, and firms tend to collapse them into one thing when they think about the risk.

The practical implication is that your verification workflow needs a second half. Most firms design a pre-filing check and stop there, which handles competence. What they do not design is the correction path: what happens when opposing counsel emails on a Thursday to say they cannot find one of your authorities, who at the firm gets told, how fast a corrected filing goes out, and what gets recorded about the sequence. Candour duties generally require reasonably prompt correction once you know, and the definition of prompt shrinks considerably once a court is involved. Decide in advance who makes that call, because deciding it in the moment is how a competence slip becomes something worse.

Citation checking is a mechanical task, so run it mechanically

Here is the rule I would put in writing if I were building this at a firm today. Every authority in an AI-assisted document gets independently retrieved from a primary source before the document leaves the building. Not confirmed as plausible. Not recognised by the reviewing attorney. Retrieved. Someone opens the reporter, the official court site, the legislation database, or the paid research platform, finds the actual text, and confirms three things: that the case or provision exists, that the citation details are correct, and that the proposition it is being cited for is one the source genuinely supports.

That third check is the one people skip, and it fails more often than outright fabrication does. A model will frequently cite a real case with correct details for a holding that case does not contain, or for a holding it contained before it was distinguished, narrowed, or superseded. Verification that stops at "the case exists" catches the loud failure and misses the quiet one. So the check has to include reading enough of the source to confirm the proposition, and it has to include the standard currency check your jurisdiction uses, whether that is a citator service, a noting-up tool, or a subsequent history search. The whole thing is tedious and unglamorous, which is why it needs to be a required step rather than a discretionary one.

  • Did someone retrieve every authority from a primary source rather than confirming it looked right
  • Does each source support the specific proposition it is cited for, not just the general topic
  • Has each authority been checked for subsequent history, appeal, or repeal
  • Do quoted passages match the source word for word, including ellipses and brackets
  • Are pinpoint citations pointing at the page or paragraph that contains the quoted material
  • Does every factual assertion trace back to a document in the matter file rather than to the model

Source of truth grounding is the step that prevents most of the damage

Citation checking catches invented law. Grounding catches invented facts, and in transactional and advisory work that is the bigger exposure. Grounding means the model is working from documents you supplied and verified rather than from its own recollection of how contracts of this type usually read. If you ask a tool to summarise an indemnity clause without giving it the executed agreement, you will get a summary of a typical indemnity clause, delivered with the same confidence as a summary of yours. The difference between those two outputs is invisible in the text and enormous in consequence.

In practice this means the working copy of every source document has to be the one document nobody disputes, which is a document management problem before it is an AI problem. If three versions of the same agreement are circulating in email and nobody is certain which one was executed, no verification workflow will save you, because the person checking the summary will check it against the wrong file. This is where storing everything against the matter earns its keep. In Casely every document sits in the matter file with AES-256 encryption under a per-firm key, and every document carries a comment field recording what changed and why, so the reviewer verifying an AI-generated summary can see which version they are looking at and what happened to it before they read a single line of the draft.

FeatureUngrounded draftGrounded and verified draft
What the model worked fromIts general sense of how documents like this readThe specific executed file stored against the matter
How a wrong fact surfacesWhen the client says that is not what we agreedDuring review, against the source document
What the reviewer checks againstTheir own memory of the dealThe version in the matter file with its change history
Time to reconstruct laterHours of email archaeologyOne matter file with the document and its comment trail
Failure modeConfident, fluent, and wrongCaught before it leaves the firm

Who signs off, and why it cannot be the person who generated it

Every AI-assisted document needs a named verifier, and that verifier should not be the person who wrote the prompt. This is not about seniority or trust. It is about the fact that the person who generated the output has already read it several times, has already accepted its framing, and is reviewing it with the specific blindness that comes from having watched it take shape. The same reason you do not proofread your own brief on the morning it is due applies here with more force, because the errors are better camouflaged.

For most small firms the practical answer is that the supervising attorney on the matter is the verifier, and that the verification is a distinct act rather than a general sense that the draft is fine. The verifier confirms the citations were independently retrieved, confirms the factual assertions trace to documents in the file, and confirms the legal reasoning is one they are prepared to defend as their own. If the firm has a designated technology partner or a practice group lead who owns the AI policy, they own the standard and the training, not the per-document check. The per-document check belongs to whoever is signing, because that is whose name goes on it.

  1. 01Draft generated with the source documents supplied from the matter file
  2. 02Every authority independently retrieved and checked for currency and proposition
  3. 03Every factual assertion traced back to a document stored on the matter
  4. 04Named verifier who did not generate the draft reviews and signs off
  5. 05Verification recorded against the matter with date, name and scope
  6. 06Correction path triggered immediately if a defect surfaces after filing

What the matter file should record about the verification

If you take one operational thing from this post, take this. The verification has to leave a trace, because a verification nobody can evidence is worth roughly what an unverified draft is worth once someone starts asking questions. What you want in the file is short and specific: that a tool was used in preparing this document, which parts it touched, who verified it, when, and what they checked. Four or five lines. It is not a compliance essay and it should not take anyone more than a minute.

The reason this matters is the same reason contemporaneous notes matter everywhere else in practice. Two years later, when a client questions an advice letter or a court raises a question about a filing, the difference between a firm that can point to a dated record of who verified what and a firm relying on an attorney's recollection is substantial. This is where a document comment field stops being a nice feature and starts being the record. In Casely every document carries that comment field recording what changed and why, so the verification note lives on the document itself rather than in an email thread that gets archived, and the matter stage tracker can carry an explicit verification step so a document cannot quietly move to the filing stage without one.

AES-256
encryption on every document, per-firm key
3K+
attorneys running their firm on Casely
15M+
billable hours tracked

Where the verification step belongs in the week

The most common way this fails is not refusal, it is timing. The firm agrees verification matters, everyone means it, and then the check gets scheduled for the last hour before a filing deadline, which is the one hour in the week when nobody has the patience to open a reporter and read three pages. Verification done under deadline pressure degrades into skimming, and skimming is exactly the mode in which fluent wrong text passes. So the step needs to sit earlier in the sequence, with the draft due internally well before the external deadline.

That is a calendaring decision more than a policy decision. If your deadline diary only holds the court date, the verification never gets its own slot and it will keep colliding with the filing. Attaching deadlines to the matter with next-date tracking, the way the deadline diary in Casely does, lets you put an internal verification date on the matter as its own dated obligation rather than as an intention. The firms that get this right treat the verification deadline as the real deadline and the court date as the consequence of missing it, which sounds dramatic until you have watched a document go out because there was no time left to check it.

The jurisdictional layer, which changes faster than anything else here

Disclosure obligations around AI-assisted work are genuinely unsettled and vary by jurisdiction, by court, and sometimes by individual judge. Some courts have issued standing orders requiring a certification about the use of generative tools in filings. Some bar regulators have published guidance on competence and confidentiality obligations. Others have said nothing yet. There is no universal rule you can adopt and there is no safe assumption that what applies in one court applies in the next one down the corridor, so check the standing orders for the specific court and the current guidance from your own regulator before you rely on any general statement, including this one.

What you can do is build the process so that any disclosure requirement is easy to satisfy rather than a scramble. If your file already records which documents were AI-assisted, which parts, and who verified them, then a certification requirement is a form-filling exercise. If your file records nothing, the same requirement becomes a firm-wide interrogation about what happened on a matter six months ago. Firms working across multiple jurisdictions feel this most sharply, because the standard has to be set at the strictest requirement they touch. Connected matters that link related files with the reason stated help here, because a disclosure question that lands on one matter usually needs answering across the group.

Client-facing output fails differently from court-facing output

Everything above focuses on filings, but the larger volume of AI-assisted text at most firms goes to clients rather than courts. Status updates, advice summaries, explanations of a settlement offer, answers to a question that came through the portal at nine at night. These carry no citation risk at all in many cases, which is exactly why they get less scrutiny, and the failure mode is subtler. A model summarising a matter will smooth over uncertainty, will state a likely outcome as an expected one, and will produce a tone of reassurance that nobody intended to convey. Clients read that as advice, because it arrived from their lawyer.

So client-facing output needs its own verification standard, focused on hedging rather than on sources. The verifier is checking that every prediction is framed as a prediction, that no timeline is stated more firmly than the facts support, and that nothing in the message creates an expectation the firm has not agreed to meet. This is also where privilege filtering matters, because a summary drafted from the full matter file may pull in material the client should never see. A client portal that filters by privilege automatically per document, the way Casely does, limits the blast radius of a mistake here, but it does not replace the read-through. Somebody still has to confirm that the message says what the firm meant to say.

Confidentiality is part of verification, not a separate conversation

There is a question that belongs at the front of this workflow rather than the back, which is what you put into the tool in the first place. Confidentiality duties apply to client information regardless of the tool, and the analysis depends heavily on the specific service, its data handling terms, whether inputs are used for training, and where the data is processed. Consumer chat products and enterprise deployments differ substantially on all four points, and the answer changes as vendors change their terms, so this is something to check periodically rather than once.

The internal control that matters here is access. If a matter is walled, the wall has to survive the AI workflow, which means a screened attorney cannot pull documents from that matter to feed a tool any more than they can read them directly. This is precisely why ethical walls in Casely are enforced at the server and data-access layer rather than hidden in the interface, so a walled user cannot reach a restricted matter through search, through the calendar, or through a forwarded link. A wall that only hides buttons is a wall that a copy-paste into a chat window walks straight through, and that is a confidentiality failure that started as a permissions failure.

!
Check what your tool does with the input Data handling terms differ enormously between consumer and enterprise offerings, and they change. Confirm whether inputs are retained, whether they are used for training, and where they are processed before any client material goes near a tool, and re-check when terms are updated.

Building the habit before you need it

The firms that handle this well are not the ones with the longest AI policy document. They are the ones where the verification step has a physical place in the workflow: a stage on the matter, a named person, a dated internal deadline, and a line in the file recording what was checked. Everything else is commentary. A policy that lives in a shared drive changes nobody's Thursday afternoon. A required stage that a document cannot move past without a verification note changes every Thursday afternoon, permanently.

Start narrow if you need to. Pick the two document types where the exposure is highest, usually anything going to a court and anything that reads as advice to a client, and require the full check on those before you worry about internal memos. Give the verification a real slot in the calendar rather than the last hour before the deadline. Make the record short enough that people will actually write it. And make sure the source documents the drafting works from are the ones sitting in the matter file rather than the ones circulating in email, because grounding fails long before verification does. Storing everything against the matter with a change history on each document is the foundation the rest of this rests on, and it is the reason legal document management software and a properly configured matter management software setup do more for AI risk than any policy memo will.

None of this is about being cautious with technology for its own sake. The tools are useful, the leverage is real, and firms that refuse to touch them are making a different mistake. But the value only holds if the output is verified, and verification only holds if it is a step in the system rather than a good intention. Put the check where the work happens, name the person who signs, and write the four lines that prove it was done. That is the whole discipline, and it costs far less than finding out you skipped it.

SG

WRITTEN BY

Sagnik G.

Writes on trust accounting, matter management, and the reporting side of a modern legal practice.

More about the team