Data Quality is the First Layer of Law Firm AI Governance
At a Glance
- Who this is for: CTOs and risk partners who've written AI controls and haven't yet checked the records those controls sit on.
- The problem: AI governance checks what a tool was asked and who reviewed the answer. It assumes the client and matter data underneath is right. In most firms some of it is wrong, and the firm hasn't measured how much.
- What changes with AI: A person who hits a bad record usually notices. A tool reading thousands of records passes the error on in a confident, well-formatted sentence.
- The takeaway: Audit the data your first AI use case will read, and give each part of it an owner, before you pick the tool.
What is data quality in a law firm? Data quality is how far the firm's records can be trusted for the job they're used for. In practice that means records that are complete, match between systems, are current, and can be traced back to where they came from. For AI, the job is being read at speed by software that won't stop and ask when something looks odd.
Where Governance Programmes Start
Our piece on AI governance for law firms set out 4 controls: grounded retrieval, a prompt and response log, a named reviewer, and a written exclusion list. We stand by all 4.
Every one of them sits downstream of the data.
Grounded retrieval ties an answer to your own documents. If the documents are wrong, you get a wrong answer with a citation attached, which is harder to spot.
The reviewer checks the output against the source. When the source is the problem, the check passes.
So I'd put data quality underneath all 4 controls. Call it layer 0.
What Bad Data Looks like Inside a Firm
These are the problems we run into most often when we open up a firm's systems:
- The same client held under 3 slightly different names across practice management, billing and the document management system, so a search finds 1 of them
- Matters marked open in one system and closed in another
- Fee earner, rate and matter type fields left blank, or filled with whatever default let the form save
- Documents migrated from an old system with their metadata stripped, so every file shows the migration date
- Parties and key dates that exist only inside the body of an engagement letter
Day to day, the firm copes with all of these. People know which client record is the real one. They know to ring accounts before trusting a matter status.
That knowledge lives in people's heads. An AI tool can't read it.
Why AI Makes the Problem Bigger
Ask an AI tool to summarise a client's live matters and it'll summarise the matters the records say are live. It won't know that 2 of them closed last year in a system it can't see. The answer will look finished and read well.
The SRA noted in its Risk Outlook report on AI in the legal market that "training and using AI effectively requires careful and sensitive use of data from both public and private data sources." I'd add one word to that: checked data.
There's a legal floor here too. Much of a law firm's data is personal data, and Article 5(1)(d) of the UK GDPR requires personal data to be "accurate and, where necessary, kept up to date." The ICO's guidance on the accuracy principle says organisations must "ensure that the source and status of personal data is clear."
An AI tool that copies an inaccurate record into 40 summaries has taken one accuracy problem and made 40. The ICO's separate guidance on AI and data protection covers accuracy in AI systems directly, and your risk team should read it before any rollout.
A Data Quality Audit Before You Deploy
Start with the fields your first AI use case will actually read. Leave the rest of the estate for later. Trying to audit everything at once is a reliable way to finish nothing.
| Check | What you're looking for | How to test it |
|---|---|---|
| Completeness | Required fields that are empty or holding a default value | Count blanks and default values on each field the AI will use |
| Consistency | The same client or matter recorded differently between systems | Match client and matter records across practice management, billing and the DMS, and count the mismatches |
| Currency | Stale statuses and departed fee earners still assigned to matters | Compare each matter's status with the date of its last time entry |
| Lineage | Records you can't trace back to a source | Pick 50 records at random and trace each one to its source system and date |
| Access | The AI can see things the person using it can't | Run the tool from a junior account against a restricted matter |
The access row catches firms out. Permission checks are usually tested for people. An AI tool connected with a service account can read far more than the lawyer asking it questions, and the answer can quietly include what that lawyer was never meant to see.
Give the Data an Owner
An audit is a snapshot. A dataset that's clean in March drifts by September unless keeping it clean is someone's actual job.
ISO 8000-150:2022 is the international standard for exactly this problem: the roles and responsibilities an organisation needs for data quality management. The idea works without going for certification.
Split the data into domains a person can own: clients, matters, time, documents. Give each owner the numbers from the audit. Review them monthly, and treat a rising mismatch count the same way you'd treat a rising lock-up figure.
In a lot of firms, these records belong to everyone. In practice that means IT inherits them on the day something goes wrong.
Use Your Next Migration
A firm moving practice management or case management systems already has to map, cleanse and reconcile every record it moves. That's most of a data quality audit, and the project is already paying for it.
In our case management consolidation for a top 20 global law firm, case records and documents moved across 2 jurisdictions in a month. Reconciling the records took most of the effort, and the verification steps were agreed before the migration ran.
If an AI rollout is on the plan within a year of a migration, build the data quality checks into the migration. Doing the work twice costs more than doing it once.
Common Pitfalls
- Choosing the AI tool first and finding the data problems halfway through the pilot.
- Auditing the whole estate at once and never getting to a result.
- Treating a migration as a copy job, so every old error arrives in the new system intact.
- Cleaning the data once with no named owner, so it drifts back within months.
- Testing what users can access and forgetting to test what the AI's service account can access.
How 3Rive Approaches This
We run the data quality audit against the first AI use case a firm has chosen, so the numbers are about something real. That work sits in our data service line.
Once the records are in shape, our AI service line builds the tool on top, with the governance controls designed in. If the firm is still deciding what the first use case should be, we start with Tech Advisory.