Performance Reviews for Legal Staff That Are Worth the Hour
Firm Management

Performance Reviews for Legal Staff That Are Worth the Hour

Most legal performance reviews run on impressions from the last three weeks. Here is how to measure attorneys, paralegals and admin staff on different things entirely, using data your system already holds, and why the raise conversation belongs on a separate day.

SMSaumyajit M.Founder, Casely

Sit in on enough law firm performance reviews and a pattern shows up fast. A partner blocks an hour, opens a document written the night before, and talks for forty minutes about impressions formed mostly in the previous three weeks. The associate hears that they are doing well, that they should try to be more proactive, and that the firm hopes to look at compensation at the end of the year. Everybody leaves the room slightly relieved it is over. Nothing measurable changes in the following twelve months, and next year the same hour gets spent the same way.

The problem is not that partners are lazy about this. The problem is that a review built on memory is structurally incapable of being useful. Memory over-weights the most recent matter, the loudest complaint, and the person whose desk is closest to yours. It cannot tell you that a paralegal quietly absorbed a forty percent increase in discovery volume in the second quarter, or that an associate's write-downs are concentrated in one practice area and one referring partner. Those facts sit in the system. Nobody pulls them, so the review runs on vibes and the firm makes staffing and pay decisions on vibes too.

There is a second failure layered on top of the first, and it is the more damaging one. Most firms run the development conversation and the compensation conversation inside the same hour. The moment a person suspects that what they say next affects their pay, they stop being honest about what they are struggling with. You have just destroyed the only part of the review that could have made them better. This piece is about fixing both problems: measuring the right things for three genuinely different kinds of role, and pulling the money out of the room.

3K+
attorneys running their firm on Casely
15M+
billable hours tracked
1-click
converts unbilled time into an invoice

Why impression-based reviews reward the wrong people

An impression-based review reliably rewards visibility rather than contribution. The associate who copies a partner on every email and stops by to narrate their week reads as high performing. The one who quietly closes matters early, catches conflicts before they mature, and never generates a crisis reads as unremarkable, because absence of drama does not register as achievement. Over a few review cycles this is not a small distortion. It shapes who gets the good work, who gets promoted, and who eventually leaves, and the person who leaves is often the one who was actually holding the operation together.

Recency bias compounds it. A review written in November is really a review of September and October. If a paralegal had a rough patch in March because they were covering two attorneys during a leave, that context has evaporated by the time anyone writes anything down. The fix is not a better memory or a longer form. It is a running evidence file assembled from the data the firm already generates every day, so the review document is a summary of a year rather than a snapshot of a month. Every time entry, matter stage change, deadline completion and invoice adjustment is a data point that was recorded contemporaneously, which is exactly what a fair review needs.

Pull the data before you write a single sentence

The sequence matters. Write the review after you look at the numbers, never before, because a review drafted from impressions and then decorated with a few supporting figures is still an impression-based review with better clothes. Start by exporting the year: time recorded and billed by person and by matter, matters opened and closed, average days in each matter stage, deadline completion against the diary, write-downs and write-offs by matter, and invoice turnaround from work performed to bill issued. In Casely, the matter stage tracker and the deadline diary carry most of this because both are attached to the matter rather than living in someone's calendar, so the history is queryable rather than anecdotal.

Then read the data before you interpret it. A number on its own is a prompt for a question, not a verdict. Long average days in the discovery stage might mean an associate is slow, or it might mean they inherited the two most document-heavy matters in the firm. A high write-down rate might mean padded time, or it might mean a partner is quoting fees that the work cannot be done inside. You will not know which until you ask, and asking is a far better use of the hour than reciting. The data's job is to make sure you ask about the right things instead of the memorable ones.

!
Never let billable hours be the only measure A person measured solely on hours will produce hours. You will get inflated narratives, work that could have been delegated staying with the expensive person, and quiet resistance to any process that makes the work faster.

What to actually measure for attorneys

For attorneys, the honest measures cluster around realization, cycle time and risk. Realization tells you what share of recorded time survives to a paid invoice, and read at the individual level over a full year it is one of the few numbers that reflects both efficiency and judgment. Cycle time by matter stage tells you whether matters are moving or parking. Risk shows up in the deadline record, in conflict checks run late or not at all, and in matters that sat in one stage for months without a client communication. Add matter origination and client development if the firm's economics depend on it, but keep it separate from delivery performance rather than blending both into a single score that explains nothing.

Then measure the things that are qualitative but still evidenced. Read a sample of their actual time entry narratives, because narratives are where you see whether someone is describing work in terms a client will accept or in terms that invite a write-down. Read a sample of their client communications through the portal. Look at whether their matters carry the connected matter links and contact labels that let the next person pick the file up cold. None of that is a metric, but all of it is evidence rather than impression, and an attorney can argue with it productively in a way they cannot argue with "I feel like your work has been a bit uneven."

What to actually measure for paralegals

The single most common review mistake in a small firm is grading a paralegal on a lightly modified version of the attorney scorecard. Billable hours are a poor primary measure for a paralegal even where their time is billed, because a large share of the most valuable work they do is structurally unbillable: chasing a signature, keeping the diary clean, rebuilding a document index, catching that a filing deadline moved. Measure a paralegal on throughput and reliability instead. How many matters did they carry, at what document volume, and what was the turnaround from request to draft ready for review?

Reliability is where the real signal sits. Look at deadline diary hygiene, meaning whether next dates were entered and updated rather than left to rot after the first hearing. Look at document management discipline, including whether the comment field on each document actually records what changed and why, since that field is the difference between a file another person can pick up and a file that has to be reconstructed. Look at portal responsiveness, meaning how long client documents sat unposted. A paralegal who scores well on these is worth more to the firm than one with higher recorded hours and a diary nobody trusts, and the review should say so in those terms.

FeatureAttorney measuresParalegal measures
Primary numberRealization on recorded timeMatters carried and turnaround time
Risk signalDeadlines and conflict checks run lateDiary entries missing a next date
Quality evidenceTime entry narratives and client commsDocument comments and file handover readiness
Growth measureOrigination and matter complexityPractice areas covered and autonomy level

What to actually measure for admin and intake staff

Admin and intake staff are the group most often reviewed on personality, which is close to useless and legally risky besides. Their work is highly measurable if you look at the right clocks. For intake, measure response time from first contact to a real human conversation, the completeness of the information captured, and the conversion rate from enquiry to opened matter, with the caveat that conversion depends heavily on the quality of the leads arriving. For billing admin, measure the lag from month end to invoices issued, the number of invoices that had to be reissued after a correction, and the age of outstanding receivables they are responsible for chasing.

The higher-stakes admin measures are compliance-shaped. Was the conflict check run against the full contact and matter history including closed matters, or against a quick search of open files only? Were client funds recorded to the correct matter ledger the day they arrived? In a system where trust accounting blocks any disbursement exceeding a matter's actual trust balance at the database transaction level, a person cannot create an overdraft, but they can still misallocate a receipt or delay a deposit, and both of those are worth reviewing. What you are really assessing is whether the operational spine of the firm is being maintained by someone who understands why it matters.

Separate the development conversation from the compensation one

Run them as two meetings on two different days, and say plainly at the start of the first one that money is not on the agenda. This is not a scheduling nicety. A development conversation only works if the person is willing to name what they find hard, where they are out of their depth, and what they want to learn. Nobody says any of that to a person who is about to decide their raise. Combine the two and you get a candidate performing an interview, which is why so many reviews consist of an employee agreeing enthusiastically with mild criticism and then changing nothing.

The practical sequence that works: development conversation first, ideally a few weeks ahead, focused entirely on evidence, growth and the next twelve months. Compensation conversation second, short, and decided beforehand on criteria the person already knows. The compensation meeting is where you explain the decision and the reasoning, not where you negotiate based on how the earlier conversation felt. Firms resist this because it looks like double the meetings. It is roughly the same total time, split so that each half can actually do its job, and the development half becomes something people prepare for rather than survive.

  1. 01Pull twelve months of data by person and role
  2. 02Write the evidence summary before forming an opinion
  3. 03Send the summary ahead so the person can prepare
  4. 04Hold the development conversation with money off the table
  5. 05Hold the compensation conversation separately with criteria stated

Build an evidence file, not a memory file

The reason most firms cannot do any of this is that the year's information was never captured anywhere retrievable. Fix that with a running file per person, updated quarterly, that takes about fifteen minutes each time. Log the matters they carried, the notable wins, the specific incidents with dates, and the data snapshot for the quarter. When review season arrives, the document is already ninety percent written and it covers the whole year rather than the last six weeks.

This is also where a practice management system earns its place. Contemporaneous records are more defensible than reconstructed ones, and every stage change, deadline completion, document comment and invoice adjustment is timestamped at the moment it happened. When you tell an associate that three matters sat in the same stage for over ninety days without client contact, you are describing recorded history, not an accusation. That changes the conversation from a dispute about perception to a discussion about what happened and what to do differently, which is the only kind of review conversation that ever produces a change in behaviour.

Watch for metrics that quietly punish good behaviour

Every measure you publish becomes a target, and some targets create damage. Measure attorneys purely on recorded hours and you discourage delegation, because handing work to a cheaper person costs the attorney their number. Measure paralegals on matters closed and you encourage them to push work back to attorneys. Measure intake purely on conversion and you get pressure to open matters the firm should have declined, which is the most expensive mistake on this list. The defence is to use a small basket of measures rather than one, and to look at any sharp movement in a single number as a question rather than a result.

Balance the efficiency measures with a quality measure that pulls the other way. Pair matters closed with write-down rate. Pair intake conversion with the proportion of those matters still active at ninety days. Pair speed of document turnaround with how often work came back for revision. A person who scores well on both sides of a pair is genuinely performing. A person who scores brilliantly on one side and badly on the other is optimising for the measure, and you would never have seen it looking at a single number.

  • Can you produce twelve months of performance data for each person without reconstructing it from memory?
  • Does each role have its own measures rather than a modified attorney scorecard?
  • Are development and compensation conversations held on separate days?
  • Does every measure you use have a counterweight that catches gaming?

Structure the hour so it is not a monologue

An hour spent well has a shape. Fifteen minutes for the person to walk through their own assessment first, before you say anything, because what someone raises unprompted tells you a great deal about their self-awareness and about problems you did not know existed. Twenty minutes for your evidence-based summary, working through the data and the specific incidents. Fifteen minutes on the next twelve months, agreeing two or three concrete things rather than a list of ten aspirations. Ten minutes for their questions about the firm, which is the part most partners skip and the part that most often surfaces something worth knowing.

Send the evidence summary at least two days ahead. People process criticism badly in real time and well after a night's sleep, and an ambush produces defensiveness rather than reflection. Finish by writing down what was agreed, in specific terms with dates attached, and send it. A commitment to "take on more complex matters" is not a commitment. A commitment to run the next two contested matters as first chair with a named supervising partner, reviewed in March, is something both of you can check against reality next time.

Keep the process legally defensible where you operate

Employment law around performance management varies sharply between jurisdictions and this is one area where a generic process can create real exposure. In the United Kingdom, dismissal for performance generally requires a documented process with warnings and an opportunity to improve, and unfair dismissal protections attach to most employees after a qualifying period. In much of the United States, at-will employment changes the picture considerably, but anti-discrimination law still makes inconsistent documentation dangerous. Australia's unfair dismissal regime and the Small Business Fair Dismissal Code impose their own requirements, and Canadian provinces differ from one another. Confirm the position for your own jurisdiction with local employment counsel before you build the process, not after a dispute starts.

!
Inconsistent documentation is the real exposure Two people in the same role reviewed on different criteria creates a comparison that is hard to defend later. Confirm your local employment law position with counsel before you finalise the process.

The practical rule that holds up almost everywhere is consistency. Use the same framework, the same measures and the same documentation standard for everyone in a given role. The moment two people in equivalent roles are reviewed on different criteria, you have created a comparison that is difficult to defend if it is ever examined. Keep the records, keep them factual, and keep them contemporaneous. A file of dated, specific, evidence-based reviews is both a better management tool and a far better position to be in than a folder of warm generalities followed by a sudden termination.

Where to start if your last review cycle was mostly guesswork

Do not attempt to build the perfect scorecard before the next cycle. Start by pulling twelve months of data for every person in the firm and simply reading it, without writing any reviews at all. You will find things you did not expect within an hour, and those surprises tell you which measures actually matter in your firm rather than which ones a template suggested. Then pick three or four measures per role, tell everyone what they are well before the cycle, and run one honest round. A slightly rough process that everyone understands beats an elegant one nobody knew about.

The infrastructure question underneath all of this is whether the data is retrievable at all. If time entries live in one place, deadlines in a shared calendar, documents on a drive and invoices in an accounting package, assembling a year of evidence per person is a project rather than an export, which is precisely why most firms skip it. When matters, time, deadlines, documents and billing sit in one system, the evidence file is a query. That is the case for keeping legal time tracking attached to the matter rather than in a separate tool, and for having reporting and analytics that can slice by person and by practice area without a manual rebuild. Casely is cloud-native with a free plan at $0 to start, so testing whether your firm's data is actually reviewable does not require a purchase decision first.

The hour is worth spending. It is only worth spending on evidence, on measures that fit the role in front of you, and on a conversation the person can be honest in. Strip out the money, bring the data, and the same sixty minutes stops being an annual ritual everyone tolerates and starts being the thing that keeps good people from quietly deciding to leave.

SM

WRITTEN BY

Saumyajit M.Founder, Casely

Founder of Casely. Builds the practice management software the firm runs on, and writes about the operational side of running a legal practice.

More about the team