AI Governance for Law Firms: What a CTO Actually Has to Build
At a Glance
- Who this is for: Heads of IT, innovation leads and CTOs at law firms that already have AI tools in production and no clear answer when a partner asks who checked the output.
- The problem: AI policies get written as documents. The failures reaching the English courts are failures of process, where nobody checked a citation against a primary source and nobody was supervising the person who did not check.
- What is changing: The Law Society's April 2026 report on agentic AI finds solicitors remain responsible for outputs they cannot fully audit, and the judgments through 2026 keep making the same point about verification.
- The takeaway: Put the verification step inside the workflow rather than inside the policy.

AI governance in a law firm is an engineering problem before it is a policy problem. The controls that survive contact with a court are the ones built into the tools people actually use: a forced source check, a logged prompt, a named supervisor, and a written list of work that AI is not allowed to touch on its own.
Why This Matters Now
Two things happened in the last eighteen months that changed what a law firm's technology function is accountable for.
The first was judicial. In June 2025 a Divisional Court, presided over by Dame Victoria Sharp, heard two referred cases together under the Hamid jurisdiction. In one, five cited authorities did not exist. In the other, of 45 authorities relied on, 18 did not exist. The court sent its judgment to the Bar Council, the Law Society and the Council of the Inns of Court, and asked them to consider urgently what further steps they should take.
The second was a widening gap between guidance and build decisions. The Courts and Tribunals Judiciary updated its AI guidance for judicial office holders on 31 October 2025, expanding the sections on hallucinations and on confidentiality, and telling judges not to put private information into public AI tools. That guidance covers the bench, its clerks and its support staff. It does not tell a firm how to build anything, so the build decisions land on the technology function.
If you run IT or innovation at a firm, the practical consequence is that you now own a control that used to be somebody else's professional obligation.
Where the Process Broke
The most useful case for a CTO is not the famous one. It is Cork and another v Smith [2026] EWHC 1199 (Ch), decided in the Chancery Division on 22 May 2026.
A junior associate at a large UK firm asked an AI tool what a particular insolvency rule said. The tool produced text purporting to be Insolvency Rule 12.37(5), giving the court an express power it does not have. That text went into a letter to the court.
Read the judge's criticism closely, because it is a process criticism. The associate asked the AI what the provision said instead of reading the legislation. The tool itself flagged that its answer might be inaccurate. That warning was not acted on. When the court raised it, a second letter explained the wording as a "summary conclusion" rather than accepting that unverified AI output had gone out under the firm's name.
The judge described a concern about "a cavalier attitude" to accuracy. The firm was criticised in a published judgment, referred to the SRA, and agreed to meet the additional costs.
Now map that to your architecture. Each of those failures has a system-level counterpart. Asking the tool instead of the source is a retrieval design choice. Ignoring the tool's own uncertainty flag is an interface choice about whether a warning can be dismissed silently. The absence of supervision is a workflow choice about who sees a draft before it leaves the building. A written policy sits above all three and touches none of them.
Four Controls Worth Building Before You Buy Anything Else
These are ordered by how much risk they remove per week of engineering effort, which is not the order most firms build them in.
| # | Control | What it means in practice | What it prevents |
|---|---|---|---|
| 1 | Grounded retrieval, not recall | Research tools answer from a document set you control, and every assertion carries a link back to the source paragraph | A tool inventing a rule that has no source to link to |
| 2 | A logged prompt and response record | Prompts, responses, model version and user are written to a retention-managed store, readable by risk and by a supervisor | Having no answer when a court asks what the tool was asked |
| 3 | A named human in the workflow | The system will not release a document to a client or a court without an identified reviewer, and the reviewer sees the sources, not only the text | An unsupervised junior being the last person to look at it |
| 4 | A written exclusion list | Named work types where AI drafting is not permitted at all, agreed with the risk partner and enforced in the tool rather than in training | Quiet scope creep into the work with the worst downside |
Control 2 is the one firms skip, and it is the one that decides how a bad day goes. Without a log you cannot tell apart a person who checked and got unlucky from a person who never checked. The court in Cork had to reason from letters because that record did not exist.
Control 4 is the cheapest, and it is the one partners actually engage with. Ask a risk partner which five kinds of work they would not want an AI first draft on, write those down, and enforce them at the tool level. That conversation takes an hour and removes more exposure than a quarter of model tuning.
What Changes When the AI Acts on Its Own
The Law Society's April 2026 research on agentic AI in legal practice is worth reading in full, because it names a gap most firm architectures have not closed.
Its finding is that solicitors remain responsible for the outputs of agentic systems while not being able to fully audit them, and it calls that a widening regulatory and liability gap. It also notes that adoption is being driven by API integration moving faster than professional governance can keep up, which lets agentic behaviour appear without anyone deciding it should.
The report describes governance practice already shifting in response. Teams reviewing large document sets are moving to statistical sampling of AI performance rather than manual review of every output, so the question becomes whether the system behaves reliably inside defined bounds. That is a different discipline from checking work, and it needs someone who can design a sample and read the result.
There is a market effect too. The Law Society found that many solicitors confuse agentic AI with AI agents and with generative AI, which lets suppliers shape what firms expect to buy. Several in-house solicitors described being asked to deliver agentic solutions on technology that lacked the capability.
For a CTO the defensible position is narrow and fairly boring. Write down which decisions must stay with a person. Instrument the rest so you can prove how it behaved. Then buy against that list rather than against a demo.
Common Pitfalls
- Treating the AI policy as the control, when the policy has no way to stop the step it forbids.
- Letting the tool's own uncertainty warning be dismissed with no record that it was dismissed.
- Building verification as a training message rather than as a step the workflow will not skip.
- Buying an agentic product without asking what its audit output looks like when a regulator reads it.
- Logging prompts with no retention policy, which turns a control into a disclosure problem.
- Governing the enterprise tools while leaving public consumer AI unmanaged, so the risk simply moves off the estate.
How 3Rive Approaches This
We start AI work at the workflow rather than at the model, which is why our AI service line covers matter automation, document intelligence and AI workflows as delivered systems with the verification step designed in. Where a firm has not yet decided what it should own versus buy, that sits with Tech Advisory first. It is the same discipline behind our case management consolidation work for a global law firm, where the verification and reconciliation steps were agreed before the migration ran, not after it.
Key Takeaways
- English courts have criticised firms over unverified AI output in published judgments in both June 2025 and May 2026, and in each the missing step was a check against a primary source.
- Grounded retrieval, a prompt and response log, a named reviewer in the workflow and a written exclusion list remove more risk than any policy document.
- The Law Society's April 2026 agentic AI report identifies a liability gap: solicitors stay responsible for outputs they cannot fully audit.
- Governing agentic systems means validating that a system behaves inside defined bounds, which is a different skill from reviewing documents.
- A firm that cannot show what a tool was asked and who reviewed the answer has no defence available beyond one individual's word.