Foundation models are already answering the public’s legal questions. This is the work of getting them to do that well across the tasks people actually need, and the shared standards, evidence, and tools that let any state build on it instead of starting alone.
These reports are working drafts shared for review. Enter the access code you received to browse the full catalog.
Incorrect code. Contact legaldesignlab@law.stanford.edu for access.
Explore these guides to learn about AI strategy, development, and evaluation for legal help. Start with the first section or jump to the topic you need.
What justice professionals are saying about AI, the real options for building or buying, and what makes the case for shared investment.
A survey of 74 justice professionals on how they approach AI: what scares them about vendors, what stops them from building, and why three-quarters are open to a shared commons.
What you'll take awayThere are only four real ways to change how these models behave, and owning your own model is the weakest of them.
What you'll take awayThe most durable parts of this work do not belong to any single state. This makes the case for national leadership on the shared layer: what it would own, why it must stay neutral, and why funding it once beats fifty states paying for it badly.
What you'll take awayHow to measure quality, what the measurements found, and how to cover fifty tasks.
Set up promptfoo and a reusable protocol like the field has done, measure any tool against the tasks you care about, and gather your own evidence before you deploy.
What you'll take awayAcross five rounds of testing, a verified list of your own resources lifted every model, while loading full articles helped only when they fit.
What you'll take awayA 60-case eval across Claude, Gemini, and GPT-4o on tagging the legal issue, the audience, the jurisdiction, and the urgency from a person's own words. The lesson: put your own category lists in the prompt.
What you'll take awayThe field has measured one task of fifty. This sets out the choices, the partners, and the work to build a real read on the rest, and how the issue-area agendas and the task taxonomy fit together.
What you'll take awayStart with the inventory, then how to get your own content AI-ready, then the specific tools to use or rebuild.
You can pick up the shared vocabularies, the datasets to build and test with, and the working tools, each with an honest note on how ready it is.
What you'll take awayA practical playbook for turning your website, guides, and trainings into a knowledge base an AI can use safely: how to gather and tier content, chunk it, build in safety, and grow from search to a knowledge graph.
What you'll take awayThe fuller, step-by-step draft companion to Guide 17. Every stage of the journey is laid out in detail, with the goal, the steps to take, the tools to use, a picture of what the output looks like, and the signs you have done it well.
What you'll take awayMany teams want to build one. This is a short note on what most decides whether a Q&A bot is safe and useful, with pointers to the guides that go deep.
What you'll take awayOut of an expert survey and a labeled legal document set: what counts as PII in legal records, how to judge whether a masking tool actually works, and how to test one without touching a real client's file.
What you'll take awayThe cheapest piece of safety infrastructure in the toolkit: a short deterministic check that catches the hallucinated links and misattributed numbers an LLM judge waves through. How to run it, and how to build your own.
What you'll take awayA FETCH-style ensemble classifier for the LIST taxonomy, built on Quinten Steenhuis's method. Whether you can use it, how to call it, how to build your own, and how its gaps grow LIST.
What you'll take awayA layer around the content that decides, for each kind of legal situation, whether to answer, what to ask first, and when to route to a human. How it works, how experts built the starter, and what is still unproven.
What you'll take awayThe standards and the cohort model the tools rest on, and where to start.
Shared field names and vocabularies let one team's knowledge base and directory be trusted and reused by another, and adopting them is a mapping exercise rather than a rebuild.
What you'll take awayA cohort either generates strategic clarity or carries working tools between organizations, and one that tries to do both at once tends to do neither well.
What you'll take awayYou locate your team by aim, content stage, and size, and the next move becomes clear, whether you are one person or five.
What you'll take awayThe projects, standards, and tools the Commons has published, gathered here.
The field's projects, datasets, benchmarks, and the legal help task taxonomy.
The fifty tasks across the justice journey that frame what AI should help with, and the rubric and evaluation work behind them.
The shared taxonomy of legal issues that ties content together across states.
The standards a knowledge base follows so its content can be trusted and reused.
The open agenda of tools the field is committing to build once and share.
The end-to-end help people need, built and shared by vertical cohorts.
The cohorts doing the building, and how to join one.