How AI Is Actually Changing Legal Document Review
Legal Tech

How AI Is Actually Changing Legal Document Review

AI is good at ranking, clustering and flagging candidates in a document set. It is not good at deciding what privilege means in your matter. Here is where the line sits and why a human sign-off step stays mandatory.

SMSaumyajit M.Founder, Casely

Document review is the part of litigation practice most likely to be described in a vendor demo as "solved." It is not solved. What has genuinely changed over the last few years is the shape of the first two days of a review, and that change is real enough to matter to the economics of a small firm. What has not changed is who is accountable for the output. If a privileged email leaves your firm because a model scored it 0.31 and nobody looked, the model does not sit in front of the judge or the disciplinary panel. You do.

The unhelpful version of this conversation splits into two camps. One says AI will do review and attorneys will supervise a dashboard. The other says none of it is trustworthy and firms should keep doing linear review by eye. Both are wrong in the same way: they treat review as one task instead of a sequence of very different tasks, some of which are pattern-matching problems that machines are genuinely better at than tired humans at 11pm, and some of which are judgment problems that a model cannot even represent, let alone answer.

What follows is the operator's version. Where AI actually pulls weight in a review, what it produces that looks like an answer but is only a candidate, the specific ways it fails that demos never show you, and why the human sign-off step is not a compliance formality you can automate away later once you trust the tool more.

3K+
attorneys running their firm on Casely
15M+
billable hours tracked
AES-256
encryption on every document, per-firm key

What review actually looks like before any model touches it

Strip the software out and a review is a sorting problem wrapped in a judgment problem. Somebody collects a set of documents from custodians, mailboxes, shared drives, phone exports and whatever else is in scope. Somebody else has to decide, document by document, whether each one is responsive to the request, whether it is privileged, whether it is confidential in a way that requires a designation, and whether it contains anything that changes the theory of the case. Those are four separate determinations with four different standards, and firms routinely collapse them into one pass because they do not have the hours to do four.

That collapse is where most review errors are born, and it happens with or without AI. A junior does one read of a document and simultaneously decides relevance, privilege and confidentiality while trying to keep three coding standards in their head. Consistency degrades over the day, degrades further across a team, and degrades most of all between the first thousand documents and the last thousand when everyone has internalised shortcuts nobody wrote down. Any honest assessment of what AI adds has to be measured against that baseline, not against a theoretical perfect reviewer who never gets tired and never drifts.

Where AI genuinely earns its place: first-pass triage

The single most defensible use of a model in review is ranking. Given a seed set of documents an attorney has coded, a system can order the remaining population by likely relevance so the review team reads the dense material first and the noise last. This is not a new idea and it is not a language model innovation. Technology assisted review has been doing supervised ranking for well over a decade, and the newer generation of models mostly improves the recall you get from a smaller seed set and makes the interface less punishing. The value is straightforward: reviewers spend their first hours on documents that actually matter, and the case theory forms earlier because the good material surfaces on day one instead of day nine.

The economics of that shift are what make it worth a small firm's attention. A ranked queue means the partner who needs to make a settlement call has a usable picture of the exposure before the review is finished, not after. It also means the low-value tail of the set, the newsletters, the calendar invites, the automated system alerts, can be handled with a lighter touch and a sampling protocol rather than a full read. That is a real reduction in hours, and it shows up in your realisation numbers on the matter rather than in a marketing slide.

Clustering and near-duplicate detection, the quiet workhorse

The least glamorous AI capability in review is also the one that saves the most time, and it barely gets discussed because it does not sound impressive. Grouping near-identical documents means an email that appears in eleven custodian mailboxes with slightly different quoting gets reviewed once and the coding propagates to its siblings. Threading collapses a forty message chain into the inclusive branches so nobody reads the same paragraph nine times. Clustering by concept puts all the documents about one contract negotiation next to each other so a reviewer builds context instead of context switching every document.

This matters more than ranking for firms doing mid-sized reviews, because most document populations are enormously redundant and the redundancy is invisible when you review chronologically. A reviewer working a threaded, clustered, deduplicated set is doing genuinely different work from one working a raw chronological export, even with identical software otherwise. The failure mode here is mild and easy to catch: over-aggressive deduplication can collapse two documents that differ in a way that matters, such as a forwarded attachment that was edited between sends. That is why the near-duplicate threshold should be a decision an attorney makes on the matter, not a default a vendor set for everyone.

FeatureWhat the model does wellWhat the attorney still has to do
First passRanks a large set by likely relevance from a seed setDefines what relevant means for this matter and these requests
PrivilegeSurfaces candidates by participant, pattern and languageMakes the determination and owns it on the log
ConsistencyApplies one label the same way across a million documentsDecides whether the label was ever the right one

Privilege flagging is candidate generation, not a determination

Every serious review platform will now flag likely privileged documents, and the flags are useful. A model can spot that a document involves an attorney participant, that it discusses legal advice in recognisable language, that it sits in a thread with other flagged documents, or that it carries a work product marker. Used properly, that turns a needle-in-haystack privilege pass into a prioritised queue where a qualified attorney reviews the high-probability set first and samples the rest under a defined protocol.

What the model cannot do is make the privilege call, and this is not a matter of the model getting better. Privilege is a legal conclusion that depends on facts outside the document: whether the sender was acting in a legal capacity or a business capacity in that specific exchange, whether a third party's presence broke confidentiality, whether the underlying communication was for the purpose of obtaining legal advice or merely copied counsel for cover, and whether the privilege has been waived by prior conduct in this matter. A model reading the four corners of an email has no access to most of that. It is generating candidates. Treating a candidate list as a determination is how firms produce privileged material and then argue about clawback provisions afterwards.

!
A flag is not a finding Any privilege call that leaves your firm on a log should have a named attorney attached to it. A model score is an input to that decision, never a substitute for it.

What the model is actually doing, and why that matters

A classifier trained on your seed set is learning the statistical shape of the documents you coded, not the legal standard you had in mind when you coded them. If your seed set over-represents one custodian, the model learns that custodian's writing style as a relevance signal. If your early coding was inconsistent because two reviewers disagreed about a category and nobody resolved it, the model learns the inconsistency and applies it at scale with perfect confidence. The system is not reasoning about your discovery requests. It is finding correlations in the material you showed it and extending them.

Generative models add a second failure surface on top of that. When a large language model summarises a document or explains why it flagged something, that explanation is generated text, not an audit trail of the actual decision. It can be fluent, plausible and wrong, and it can cite a passage that does not exist in the document. This is the specific reason a summary produced by a model should never be the thing an attorney relies on when the underlying document is available. Read the document. Use the summary to decide reading order.

The error modes nobody puts in the demo

The failure that costs firms the most is quiet and systematic rather than loud and obvious. A model that has learned a subtly wrong notion of relevance does not produce a scattering of random mistakes you would notice. It produces a consistent blind spot, and the documents in that blind spot are the ones ranked lowest, which means they are the ones nobody reads. You do not discover the gap during review. You discover it when opposing counsel produces from their own collection a document that was sitting in your set the whole time, scored 0.08, never opened by a human.

The second error mode is drift across the life of a matter. Review scope changes when a new request lands, a custodian gets added, or the court narrows an issue. A model trained on the earlier scope carries the earlier assumptions forward silently unless somebody retrains and revalidates it. The third is more mundane and more common: reviewers stop reading carefully once they trust the ranking, because the top of the queue is reliably good and attention naturally follows reward. Automation bias in review is well documented in other high-volume decision domains, and there is no reason to think attorneys are immune to it.

  • Do you know what seed set your relevance model was trained on, and who coded it?
  • Has anyone sampled the low-ranked tail of this population and read it by hand?
  • If the scope changed mid-matter, was the model revalidated after the change?
  • Can you name the attorney who signed off on every privilege call on the log?

Why the human sign-off step is non-negotiable

The professional conduct rules across common law jurisdictions do not contain an exception for delegation to software. Your duty of competence covers the technology you choose to use, your duty of supervision covers the work product regardless of what produced it, and your duty of confidentiality follows the documents wherever you send them for processing. None of that changes because a vendor's accuracy claim is high. A tool cannot hold a duty, so the duty stays with the attorney by default, which means the sign-off step is not an optional layer of caution. It is where the responsibility actually lives.

There is also a practical argument that lands harder than the ethical one for most managing partners. A review you cannot explain is a review you cannot defend. If you are challenged on completeness, on privilege, or on the adequacy of your process, the answer "the software ranked it that way" is not an answer. The answer that works is a described protocol, a named attorney at each decision point, a documented sampling and validation approach, and a record of what changed and why when the process was adjusted mid-matter. That record does not build itself, and firms that treat AI as a way to skip building it are trading a known cost for an unbounded one.

Jurisdiction changes the answer more than the technology does

Courts in several common law jurisdictions have accepted technology assisted review in specific cases, most visibly in the United States following Da Silva Moore in 2012 and in England and Wales following Pyrrho Investments in 2016. That acceptance is real but it is narrower than vendors imply. Approval in a decided case is not a general licence, and the standards applied to disclosure differ substantially between jurisdictions. The English disclosure regime, the American federal and state discovery rules, the Canadian provincial rules, and the Australian federal court practice notes each set out different expectations for proportionality, for cooperation with the other side, and for what you must disclose about your own methodology.

The practical consequence is that the same technical process can be perfectly defensible in one forum and a problem in another, mostly because of what you are expected to tell opposing counsel and the court in advance. Some regimes lean heavily on discussing your search and review methodology with the other side before you run it. Others do not require that at all. Before you rely on any automated review process in a live matter, confirm the position in your own jurisdiction and, where relevant, in the specific court, because this is an area where the rules genuinely diverge and where local practice notes move faster than general commentary about them.

Building the review record you would actually want to defend

The record that protects a firm is not a log of model scores. It is a description of the protocol you followed, written before or during the review rather than reconstructed afterwards, including who defined relevance, what the seed set was, how validation samples were drawn, what the acceptance criteria were, and who reviewed the exceptions. Firms that build this as they go find it takes very little additional time. Firms that try to assemble it three months later, under pressure, discover that half the decisions were made verbally and nobody remembers the reasoning.

This is a matter management problem more than a review platform problem, which is why it tends to fall between the cracks. In Casely the review protocol, the coding standards memo and the validation samples live attached to the matter itself, and every document carries a comment field recording what changed and why, so the reasoning travels with the file rather than living in someone's memory. Connected matters link a review to related matters with the reason stated, which matters when the same custodian set has come up before and an earlier privilege position needs to stay consistent. The deadline diary attaches production dates to the matter with next-date auto-tracking so the validation pass has a hard date rather than a vague intention.

Confidentiality, ethical walls, and where the documents physically sit

Sending a client's document population to a third party model provider is a confidentiality decision, not an IT decision, and it deserves the same scrutiny as any other disclosure. The questions that matter are whether the provider trains on your data, how long they retain it, which jurisdiction the processing happens in, who at the provider can access it, and what happens to the copies when the matter ends. Some clients, particularly institutional and regulated ones, have outside counsel guidelines that answer these questions for you, and those guidelines are increasingly explicit about AI processing. Read them before the collection, not after.

Access control inside your own firm is the half of this that firms forget. A review often involves contract attorneys, a temporary team, or a paralegal who should see one matter and nothing else, and a document set is exactly the kind of concentrated confidential material that makes a screening failure serious. Casely enforces ethical walls at the server and data-access layer rather than hiding restricted matters in the interface, so a walled user cannot reach a restricted matter through search, the calendar, or a forwarded link. Documents carry AES-256 encryption with a per-firm key, and conflict checking searches the full contact and matter history including every role a party played and every closed matter, which is how you catch the reviewer who worked the other side of a related case two years ago.

98%
customer satisfaction
$0
to start, on the Free plan
0
extra logins needed for e-signatures

How to pilot this without betting a live matter on it

The way to find out whether an AI-assisted review works for your firm is to run it in parallel against a matter you have already completed. Take a closed review where you know the answers, feed it through the process, and compare what the ranking surfaced against what your team actually found. This costs a fixed amount of time and produces something no vendor demo can: a calibrated sense of where the tool helps you and where it quietly misses, on your documents, in your practice area, with your definition of relevance.

When you move to a live matter, keep the human pass on the full privileged-candidate set and keep a genuine random sample of the low-ranked tail, and treat both as fixed costs of the process rather than things to trim once you feel confident. Confidence is not evidence. The sample is the only thing standing between you and a systematic blind spot that has no other symptom until it becomes a problem in front of a judge.

  1. 01Run the process against a closed matter where you already know the answer
  2. 02Compare its ranking against what your team actually found and where it missed
  3. 03Write the review protocol and coding standards before the live matter starts
  4. 04Keep a full attorney pass on every privilege candidate
  5. 05Sample the low-ranked tail by hand and record the result on the matter

The honest summary

AI has changed legal document review in a narrow and genuinely useful way. It sorts, clusters, deduplicates, threads and ranks, and it does all of that faster and more consistently than a human team can. That is worth having, and for a small firm facing a review that would otherwise be economically impossible, it can be the difference between taking the matter and passing on it. Anyone telling you it does more than that is either selling something or has never had to defend a production.

What has not changed is the part that was always the actual work. Deciding what relevance means for this matter, making privilege determinations that depend on facts the document does not contain, catching the thing that is important for a reason nobody anticipated, and signing your name to the result. Those stay with an attorney, and the firms that get value out of these tools are the ones that put the sign-off step in the workflow deliberately and never let it erode. The ones that get burned are the ones that treated it as a temporary safeguard to remove once the accuracy numbers looked good enough.

If you want the surrounding infrastructure to hold that up, the review protocol, the version history, the per-document comment field, the access restrictions and the production deadlines all need to live on the matter rather than scattered across a review platform, a shared drive and an inbox. That is what legal document management software is for, and it is the part of this that has nothing to do with AI and everything to do with whether you can explain your process a year from now. Casely is cloud-native with no local install and a free plan at $0 to start, so you can put the record-keeping in place before you decide anything about the models.

SM

WRITTEN BY

Saumyajit M.Founder, Casely

Founder of Casely. Builds the practice management software the firm runs on, and writes about the operational side of running a legal practice.

More about the team