I recently presented at a NAAG eDiscovery webinar on moving AI from pilot to practice inside state Attorney General offices. The session covered the operational groundwork: data and governance foundations, policy development, defensibility, and chain-of-custody requirements.
Nearly every follow-up question I've received since has come from attorneys and litigation support staff who have a use case, have run a pilot, and cannot get past the ROI question.
I spent years inside a state AG office before joining Relativity, first as a senior legal analyst and then as the e-discovery program manager at the California DOJ. I've been on both sides of the budget request for AI initiatives: building the case, and later, weighing whether to approve one.
The reason I’ve seen most AG-office AI initiatives stall isn't a weak business case. It's that we're asking our people the wrong question.
The Gap isn't Interest
At an earlier NAAG eDiscovery webinar this summer titled “Ethical and Practical Strategies for AI in eDiscovery,” NAAG polled its community of AGs coming from across the country looking to find out where AG offices stand on implementing AI in e-discovery. Forty-four percent of AG offices say they're watching AI closely but aren't using it yet. Thirty-one percent are experimenting or piloting. Seventeen percent aren't using it and don't know where to start. Eight percent are actively using AI in e-discovery matters.

Stephanie Caldwell, senior counsel for NAAG, who also presented on the webinar and manages NAAG's eDiscovery educational trainings, said: “Every office I talk to is somewhere on that curve, and almost none of them are stuck because they don't see the potential. They're stuck because no one has given them a way to make the case in terms their leadership weighs decisions in.”
Interest in AI is everywhere. What's missing is the path from a promising pilot to a workflow that is repeatable, scalable, and defensible. And the hurdle in that path is leadership approval and the ROI question attached to it.
The Three Currencies an AG Office Spends
Ask for return on investment and you've handed your team a problem they can't solve. We don't have a revenue line; our success metrics were written into statute, not a P&L. A staffer who tries to answer in corporate financial terms will often either fail or be tempted to stretch the numbers.
Mission impact is measurable, but not measured solely in dollars. Ask for it in the three currencies your office spends every day, and ask for the specific unit:
- Efficiency. Can a six-attorney enforcement bureau take on a matter that used to require sixteen? Measure it in attorney-hours per matter and matters carried per bureau attorney.
- Speed. Can we meet a statutory records deadline, respond to a Civil Investigative Demand (CID), or move a multistate production without adding headcount? Measure it in days to CID response and days to statutory deadline, tracked against the clock the statute sets.
- Risk. In a coalition where twenty or more offices review the same corporate production, consistency is itself a form of defensibility. Five human reviewers will code the ten-thousandth document differently than the first, even with a clear coding guide. Well-prompted AI applies the same logic to both. Measure it in coding agreement rate across reviewers, privilege error rate, and the age of the oldest request in the queue.
Every AI use case in an AG office likely maps to at least one of these. Require those numbers first, and dollars only where the dollar figure is honest.
What Proof of Mission Impact Looks Like
There's one artifact that demonstrates all three currencies at once: run the tool against a matter your office has already reviewed manually and put both results side by side.
At the California DOJ, we did exactly that. We received a production of 5 million records – the kind where standard search and review would have taken a projected 4-5 months across 12 staff. A subset had already been manually reviewed through a vendor. We ran that same subset through Relativity aiR as a pilot. It returned comparable results in ~8,500 records, required one person to operate, and produced citations supporting each document identification, which made verifying and correcting the output straightforward.
That comparison did something no vendor demo can. Attorneys saw accuracy. Budget saw hours. Executives saw a workflow that would survive a challenge. One artifact, three audiences, zero promises.
The same pattern stands up at scale, too. JND, a legal services firm supporting federal government clients, used AI-assisted review to narrow 1.3 million documents to 650,000, identify 66,000 tied to nine distinct case issues, and surface 122 documents of critical strategic importance – in one week, with three attorneys providing subject matter expertise, and in 20 percent of the time traditional methods would have required.
That's mission impact you can hold in your hand – a receipt that only costs you one already-reviewed matter to obtain.
The Defensibility Question has an Answer
For most of us, a gating concern has been what happens when opposing counsel challenges the methodology.
In July, a federal magistrate judge in the Northern District of California took it up. In Schulte v. LinkedIn Corp., No. 22-cv-00237-HSG (LB) (N.D. Cal. July 1, 2026), the court reviewed a workflow that used search strings to narrow the population, Relativity aiR to make responsiveness determinations, and human sampling for quality control. The court described aiR as a form of technology-assisted review (TAR) and evaluated it under the same reasonableness and proportionality standards courts have applied to TAR since Da Silva Moore in 2012. Plaintiffs' motion to compel aiR performance metrics was denied as discovery on discovery.
It's a magistrate's discovery order, not binding precedent, and your team should say so when they brief it to you. But the direction is unmistakable: generative AI review is being treated as the next step in an accepted line, not a novel category requiring new doctrine.
Three Things to Require Before You Approve AI
- The counterfactual. Not the cost of the tool; the cost of the status quo, covering metrics like staff hours on the target task, error and rework rates, and queue depth translated into statutory exposure. A 90-day public records backlog isn't an operational annoyance. Under Florida's Sunshine Law, Texas's TPIA with its mandatory AG opinion process, California's CPRA, or New York's FOIL, it's a compliance posture with a clock on it. When an office builds this baseline out, the case usually makes itself.
- The conservative scenario, not the ceiling. Require three models: 50 percent of vendor claims, 70–80 percent, and full potential. Make your team show all three. And make them say out loud that year one is net negative. Implementation, integration, and change management don't disappear because you signed off.
- The procurement clock. This is where AG-office AI momentum too often dies, and it's almost never rooted in the business case. Before you approve a timeline, require answers on StateRAMP authorization or a documented FedRAMP equivalency; in-state or CONUS data residency if your state mandates it; the purchasing vehicle, whether that's NASPO ValuePoint, a state cooperative agreement, or an existing IDIQ; and how long your state technology review board actually takes. Loop your state CIO/CISO in before vendor selection, not after.
Two Realities Worth Naming
Managing an AI implementation and setting realistic expectations is no small undertaking. Consider both of these when you embark on this journey.
- Your workforce. Attorneys and litigation support staff may hear "AI" and think job loss. In the JND matter, three attorneys did the substantive judgment work while AI handled the sorting. At California DOJ, our attorneys weren't replaced by AI – they were leveraged by it. The work wasn't feasible at pace without it. Say that early and specifically, or the quiet resistance will cost you more than the software.
- Your calendar. Most of us serve elected AGs. A program whose only payoff arrives in year three is politically exposed, and a transition can end it. Require milestones at six and twelve months that stand on their own. And agree in writing, before the pilot starts, on what "good enough to continue" and "not working" look like. A pilot that fails gracefully with documented learnings preserves far more institutional trust than one that was oversold and succeeded.
The First Thing to Change
The offices moving fastest on AI right now don't have the biggest budgets or the best technology. They have leadership that treated the justification process as seriously as the implementation.
If a promising initiative is sitting on your desk waiting for approval, the answer is rarely a better demo. So don't ask your people to prove a return. Ask them to prove the mission moved, and to show you the matter where it did.
Graphics for this article were created by Sarah Vachlon.
