ArchitectureLegal Services
AI search over a firm's precedents and know-how: an architecture that lawyers can trust
AI knowledge search for a law firm is retrieval-augmented generation over the firm's own documents: it finds the right precedent clause, prior advice or practice note and answers with pinpoint citations to it. The hard parts are not the model. They are deciding which documents count as authoritative, carrying matter-level permissions and ethical walls into every query, keeping superseded templates out of answers and testing for invented authority before lawyers rely on it.
On this page
- What lawyers actually ask a knowledge search tool
- Five layers of a permission-aware precedent search system
- Choosing the corpus: curated know-how, all matter documents or both
- Ingesting from the document management system without losing context
- One precedent query, from question to cited answer
- Keeping answers current and testing for invented authority
- A hypothetical search for negotiated liability caps in one sector
- Design risks specific to legal knowledge search
- Questions and answers
- Sources
What lawyers actually ask a knowledge search tool
The questions are narrower than a general chat assistant suggests. Lawyers want the firm's standard clause for a particular point and the fallback positions it has agreed before; the last piece of advice the firm gave on a regulatory question; the practice note explaining how a procedure works; or the closing checklist used on a similar deal. They want the document itself, not a fluent paraphrase of it.
That shapes the architecture. Answers must point to a specific document and passage the lawyer can open. The tool must know which documents are the firm's approved positions and which are one-off negotiated outcomes. And it must never show a lawyer a document they could not open in the document management system, because an ethical wall breached through a search box is still a breach. Playbook-based review of an incoming contract is a different job, covered by the AI contract review use case.
Five layers of a permission-aware precedent search system
- Lawyer interface
Search and drafting panel inside the tools lawyers already use, showing answers with linked citations.
- Grounded answering
A language model drafts only from retrieved passages, cites each one and declines when support is weak.
- Permission-filtered search
Hybrid keyword and semantic search, filtered by the user's DMS access and ethical walls before ranking.
- Curated index
Chunks with metadata: practice area, document type, status, approval date, matter and access list.
- DMS and KM sources
The document management system, the know-how library and the matter system that hold the originals.
Choosing the corpus: curated know-how, all matter documents or both
Corpus scope decides answer quality and risk more than model choice does.
| Consideration | Curated KM corpus only | All matter documents | Tiered: curated first, matter documents labelled |
|---|---|---|---|
| Reliability of answers | High: approved precedents and notes | Mixed: drafts, superseded and negotiated versions | High for curated hits; matter hits marked as examples |
| Coverage | Limited to what KM has written up | Very broad, including rare points | Broad, with the trust level shown |
| Permission complexity | Low: mostly firm-wide access | High: matter access lists and ethical walls throughout | High for the matter tier |
| Effort to keep current | Owned by KM lawyers | Hard: nobody curates matter files | KM owns the top tier; status metadata drives the rest |
| Best fit | First release, or firms with strong KM libraries | Rarely the right starting point | Second phase, once permissions are proven |
Most firms will start with the curated corpus and add a labelled matter tier later.
Ingesting from the document management system without losing context
A firm's document management system, for example iManage or NetDocuments, holds versions, document types, authors, matter numbers and access controls. Ingestion should carry all of that into the index, not just the text. Version history tells the system which draft was final; document type separates an executed agreement from a markup; the matter number links each passage back to its access list.
Chunk documents along their legal structure, by clause, schedule or heading, rather than by fixed length, so a citation lands on a whole clause. Keep defined terms and cross-references with the chunk where possible, because a limitation clause without its definition of losses can mislead.
Permissions deserve one firm rule: the index stores each document's access list and the retrieval layer filters by the requesting user's identity and ethical-wall membership before anything is ranked or shown. The SRA's guidance treats systems that prevent access to confidential information, including separate access controls, as part of effective measures for information barriers1. The implementation pattern for filtering by identity at query time is covered in detail on our permission-aware retrieval page.
One precedent query, from question to cited answer
- Lawyer
Asks for the firm's position on a clause.
- Search service
Resolves the user's identity and wall memberships.
- Index
Returns only passages the user may see.
- Language model
Drafts an answer from the passages supplied.
- Audit log
Records the query, sources shown and user.
A hypothetical search for negotiated liability caps in one sector
Design risks specific to legal knowledge search
A negotiated concession presented as the firm's standard
Early signalLawyers quote one-off deal terms as house positions.
MitigationLabel trust tiers in every answer and rank curated precedents above matter documents.
Permission drift after a wall changes
Early signalA user removed from a matter still sees its documents in results.
MitigationSync access lists from the DMS on change events and test revocation before each release.
Fluent answers citing law that no source contains
Early signalAn answer names a case or statute that is not in any cited passage.
MitigationBlock display of uncited authorities and keep trap questions in the regression set.
Questions and answers
Should a firm build knowledge search on its document management system or on a separate index?
Usually a separate search index fed from the document management system. The DMS remains the system of record for documents and permissions, while the index adds chunking, embeddings and metadata such as document status. The key requirement is that the index keeps each document's access list synchronised with the DMS, so permission changes take effect in search quickly.
How is AI knowledge search different from AI legal research tools?
Research tools search published law: legislation, case law and commentary from legal publishers. Knowledge search covers the firm's own work product: precedents, prior advice, practice notes and executed documents. The two complement each other, but they carry different risks. Knowledge search must enforce client confidentiality and ethical walls; research output must be verified against authoritative sources before it is relied on.
Can the system draft from precedents as well as find them?
Yes, if drafting stays grounded. The safer pattern is to retrieve the firm's approved clause and adapt it to the facts the lawyer supplies, showing the source clause alongside the draft so the lawyer can see every change. Drafts produced without a retrieved source should be clearly marked as such or blocked for practice areas where the firm requires precedent-based drafting.
Sources
- Confidentiality of client information (guidance) — Solicitors Regulation Authority · checked 10 October 2026