By The LexText Team
Updated:

Time saved is easy to measure. The harder question is whether a legal AI tool improves the work that follows.
Key takeaways
Measure time to usable work, not time to first output.
Evaluate legal AI across four dimensions: speed, quality, verification, and how the saved time is used.
Test tools on real litigation workflows using consistent inputs and a written scorecard.
Every vendor can show you a fast draft. Few can show you what happens to it next.
A draft produced in three minutes may look impressive. But if an associate must reconstruct the analysis, replace weak authority, or spend hours verifying every proposition, the tool has not eliminated work. It has moved the work downstream.
That distinction matters. Thomson Reuters reports that legal professionals expect AI to free nearly 240 hours per year. The value of those hours, however, depends on what the technology produces and what lawyers do with the time it returns.
The better question is:
Does this tool help lawyers produce stronger, more reliable work and spend more of their time exercising legal judgment?
The following framework gives litigation teams a practical way to answer that question.
How quickly does the tool produce usable work?
“Time to first draft” is an incomplete metric.
Some tools produce text quickly but leave substantial reconstruction, source checking, and formatting to the lawyer. A more useful measure is time to usable work: the point at which the output is ready for substantive review rather than wholesale rebuilding.
For each workflow, record:
Total lawyer time
Number of substantial revisions
Time spent verifying the output
Senior-lawyer review and rework
Whether the output fits the team’s normal drafting process
A fast first answer is helpful. A fast answer that survives review is valuable.
Does the first pass improve the quality of the work?
Legal work cannot be evaluated by length or polish alone. A fluent draft can still miss a claim, misunderstand the procedural posture, rely on weak authority, or overlook the strongest counterargument.
For litigation work, assess:
Issue coverage: Did it identify the material claims, defenses, elements, and procedural issues?
Factual fit: Did it connect the law to the relevant allegations or evidence?
Arguments: Did it surface useful arguments and counterarguments?
Authority: Were the sources relevant and supportive of the propositions made?
Organization: Did it provide a structure a litigator could actually use?
Strategic value: Did it reveal a vulnerability or opportunity worth exploring?
Where practical, have an experienced lawyer compare AI-assisted and conventional work without knowing which is which. This produces better evidence than simply asking users whether they liked the tool.
How much verification work does it create?
Verification is part of the cost of producing legal work.
In a 2024 preregistered Stanford study, two leading AI legal research products hallucinated in more than 17% of the researchers’ test responses. The study evaluated specific products and tasks at a particular point in time, so the result should not be applied mechanically to every legal AI system. But it illustrates why polished output cannot be treated as verified output.
A pilot should test:
Whether cited authorities exist
Whether quotations are accurate
Whether each authority supports the stated proposition
Whether users can reach the underlying source quickly
How the tool handles incomplete information or a false premise
How much lawyer time is required to complete verification
The relevant question is not whether a system is described as “grounded.” It is how efficiently a lawyer can trace, test, and confirm the work.
Lawyers’ professional responsibilities remain unchanged. ABA Formal Opinion 512 emphasizes competence, confidentiality, supervision, candor, communication, and reasonable fees when generative AI is used.
What happens to the time that is saved?
A pilot often ends once a team calculates that a task became faster. That misses an important part of the analysis.
Ask what lawyers did with the returned time.
Did they:
Investigate another legal theory?
Develop a stronger factual narrative?
Test the argument against likely responses?
Give an associate more substantive feedback?
Advise the client earlier?
Prepare more thoroughly for a deposition or hearing?
Handle additional work without reducing quality?
Time saved is only potential value. What the firm gains depends on what fills the hours.
A practical litigation AI scorecard
Use the same scorecard across tools and workflows.
Dimension | What to measure | Warning sign |
Speed | Time to usable work and revision rounds | Fast output followed by substantial reconstruction |
Quality | Issue coverage, factual fit, arguments, authority, and organization | Polished writing that misses material issues |
Verification | Time required to confirm sources and factual support | Sources are difficult to trace or frequently need correction |
Consistency | Results across similar assignments and different users | Quality depends heavily on prompt-writing ability |
Workflow fit | Compatibility with templates, formats, and review processes | Output requires manual transfer or reformatting |
Better hours | How the saved time is reinvested | Savings are recorded, but no higher-value work follows |
How to run a useful pilot
Six questions to ask every legal AI vendor
How can lawyers trace generated statements and citations to their sources?
How does the system respond when information is incomplete or the user’s premise is wrong?
Is customer data retained or used to train models?
What security controls and independent audits are in place?
Can the system work with the firm’s templates and preferred work product?
How should the firm measure quality, accuracy, and time savings during a pilot?
A vendor’s willingness to discuss limitations is itself useful information.
The real measure of litigation AI
Litigation AI should not be purchased because it produces the longest brief or the fastest answer. It should be evaluated on whether it helps lawyers reach stronger, more defensible work with less avoidable effort.
That requires measuring the entire process: analysis, drafting, verification, revision, and the use of the time returned.
LexText was built around litigation-specific workflows. Teams can use it to analyze complaints element by element, develop arguments and counterarguments, draft pleadings, motions, and discovery, conduct legal research, and search case documents. The platform includes Hallucination Guard, zero data training, and SOC 2 Type II controls. Explore the LexText platform.
The most meaningful evaluation, however, is what the technology does with work your team already understands.
Bring a recent motion, complaint, or discovery request. In a private 30-minute session, a litigator on our team will run LexText on your work so you can evaluate its analysis, drafting, sourcing, and workflow fit for yourself.