Four Ways to Make a Foundation Model Behave
The big models will attempt legal help work across the whole task taxonomy. We do not have a reliable, shared measure of how well they do it. Owning our own model is the weakest of the few real ways to shape how they behave. Here is the option space, and where the field should place its bets.
Common Legal Help AI. Working draft for comment. Margaret Hagan, Stanford Legal Design Lab.
Ask Claude, Gemini, or ChatGPT a legal question and it will answer. Upload a court summons and it will tell you what it thinks the document is and what to do next. These models are free, widely available, and willing to attempt almost any task in the legal help taxonomy you put to them.
What the field does not have is a reliable, shared measure of how well they do that work, or solid data on how often people rely on them for it. We can demonstrate that the models will attempt the tasks. We cannot yet say, with evidence, how often they get the jurisdiction right, how often they invent a citation, or how the quality holds up across the fifty different tasks people actually need. That absence of measurement is the starting point for everything that follows, and closing it is one of the four strategies below.
The work is also not one thing. The Legal Help Task Taxonomy on JusticeBench lists fifty tasks across seven stages of what people and providers do: reading a document and explaining it, spotting the issue inside a messy story, calculating a deadline, selecting and filling the right form, drafting a motion, researching the law, screening eligibility, routing a person to the organization that can help, keeping content current and the directory accurate. Legal question-answering is one tile. By the project counts on JusticeBench today, it has eighteen projects built on it. The document explainer has two. The deadline calculator has none. The field's building effort has concentrated on the chat box, and most of the taxonomy is barely touched.
So the question for the legal help field, and for the companies building these models, is not whether legal aid should train its own model. It is this: these foundation models are already capable of attempting this work, so how does the field get them good at the parts that matter, across the whole taxonomy, and how does it know when they are?
A note on where this comes from. A seven-state cohort came together to figure out how to harness generative AI to get more help to people in crisis, especially brief legal help: question and answer, and the other short interactions that meet someone at the start of a legal problem. The aim was to close some part of the justice gap. The original hypothesis was specific: take one of the foundation models, fine-tune it on legal help resources, and give the field a trained model of its own.
We are not going to fine-tune that model. Owning a model is the costliest of the four options below and the one that holds up worst over time, and the field's capacity is too scarce to spend on it. We also found that the work is wider than the brief-help question we started on, which is why this guide is about the whole taxonomy and not the chat box alone. The decision, and the widening, both generalize well past one grant.
There are only four places you can intervene on a foundation model's behavior. Here they are, ordered from the most control and the most cost to the most reach and the least.
1. Change the model itself
Train or fine-tune a legal help model of our own. This gives the most control over how the model behaves. It also costs the most, it goes stale as soon as a frontier lab ships its next version, and it requires permanent retraining and engineering capacity that no legal aid organization or court has. The frontier moves faster than any nonprofit's retraining cycle.
Picture the maintenance reality. A cohort fine-tunes a model on 2025 legal help resources and ships it. Six months later a new frontier model lands that beats the fine-tune on most tasks with no special training at all. Now the choice is to keep serving something already behind, or start the retraining over, with labeled data and engineering time that no one budgeted for. The people who would have to run that treadmill are the same content editors and managing attorneys already stretched thin across a statewide website.
There is a narrow exception, worth naming so this is not a strawman. A small, bounded, trained component can still earn its keep: a classifier that tags content with the right legal issue, for instance, where the task is specific and the model can be small and cheap to run. That is a different thing from training and maintaining a general legal help model. The cohort built exactly that kind of classifier and kept it narrow on purpose.
This is the option to set down, on purpose, so the field can spend its capacity on the three that hold up better.
2. Change what goes into the model when it works
Leave the model alone. Change what you put in front of it. Ground it in real local content and wrap it in real checks.
This looks different for each task, and that is the point. For Q&A it means the Justice Knowledge Base and retrieval. For the document explainer it means the model sees the rules of the right jurisdiction before it reads the summons. For referral routing it means a structured service directory with eligibility and geography underneath. For the standard document filler it means the actual form library and the rules that govern the form. The model stays general. The grounding is local. The grounding is what is meant to make the output right for this person, in this county, on this task.
Here is what this could look like in practice, described as an illustration rather than any one team's actual system. Picture a statewide legal help website team that runs a public legal information portal. Their guides and forms are already tagged by issue and jurisdiction and kept reasonably current. They add a brief-help assistant. When a user asks about an eviction notice, the system pulls that team's own guides for eviction in that county, hands them to a foundation model, and the model answers from that content with links back to the source. Before the answer reaches the user, a deterministic check confirms that every cited URL and phone number is real and serves that state. The model supplies the language. The team's content supplies the truth. The check supplies the floor.
There are two ways to deliver this, and they map to two kinds of teams. A team with in-house technical capacity can build and run its own tools on top of Claude, Gemini, or GPT. A team of one or two people cannot, and should not try. For them, the field exposes the shared pieces, the classifier, the content lookup, the safety checks, as MCP servers that any tool can call, so the small team wires up a hosted service instead of building a pipeline. MCP is the doorway, and a doorway is worth having. It is also worth saying plainly: a clean doorway in front of weak content still serves weak answers. MCP solves access. It does not solve whether the answer is correct. The teams doing this work compare notes through the KB Power Group, so no one is solving it alone.
This is the one place of the four where the cohort has run real evaluations, on synthetic queries under controlled conditions rather than deployed tools with live users. The grounding guide reports what they showed, and how far to trust them.
3. Change what the model makers put inside the model
This is the place that could reach people who never reach us. A legal aid website only helps the people who find it, and the justice gap research is consistent that most people with civil legal problems never get to legal aid at all. If the model itself carries the right content, accurate help can travel to people through tools they already have, without their ever landing on a legal aid site. That is the opportunity. It depends entirely on the model having good content to draw on.
There are two ways to pursue it. The first is to publish the resource and keep control of the terms. The field announces that the Justice Knowledge Base exists, alongside the task taxonomy and the knowledge standards, and publishes it as a structured, licensed feed that companies and developers can build on. The license is where the control lives. The core legal guidance and citations stay available so accurate help can travel into the tools people already use, and the field sets the terms for the rest, holds the standard and the source, and can build paid or premium tiers around the value-added layers for the institutions and companies that want them. As an analogy: transit agencies publish their schedules in a shared format and mapping apps consume them, and the agency still owns and updates the source and the standard underneath. The legal help version is a feed of current, jurisdiction-tagged content that a model or a developer pulls from on the field's terms.
The second is direct work with the companies: the early conversations with contacts at DeepMind in Google, and with the economic mobility team and Beneficial Deployments at Anthropic. The shape of that work is a memo and a standard, not a handshake. Here is structured legal help content across the task taxonomy. Here is what good behavior would look like on each task. Here is the evaluation that measures it. Here is the safety scaffolding for the high-risk tasks, the ones where an error is hard to undo: plea consequences for noncitizens, custody that crosses state lines, decisions that affect someone's housing or immigration status.
This is not a job for any single state, and that is the important part. No one legal aid organization can negotiate with Google, and none should try. The work belongs to a convener acting for the whole field. That role is only beginning to take shape, with Legal Help Commons newly launched to take it on, and over time it should grow into a national body with a mandate to hold these relationships. Picture how the labor splits: each state contributes its content to the shared feed and keeps stewardship of it, the convener publishes the feed and maintains the standard and carries the memo to the companies, and funders underwrite the stewardship because no state has a line item for it. The state does the content. The field does the negotiation.
These are conversations, not commitments. The companies have their own roadmaps and their own incentives. The feed it publishes, and the terms around it, are the part the field controls. The partnership is the part it asks for.
4. Change the incentives around the models
The fourth place to intervene is measurement. Publish a public, legal-help-specific benchmark across the whole taxonomy, not just Q&A: jurisdiction precision, citation safety, document understanding, triage accuracy, eligibility correctness, form selection. When the benchmark is public, the companies have a reason to score well on it, a procurement officer gets a number to ask a vendor for, and a two-person legal aid build and a frontier lab get measured the same way. The measurement guide is how that measurement works.
Here is what that could look like in use. Picture a public scoreboard on JusticeBench, with scores broken out by task. A legal aid director weighing three chatbot vendors asks each one for its score on the brief-help and referral suites, and treats a refusal to be tested as its own kind of answer. A funder writes "must pass the safety suite" into a request for proposals. A model company that sees its jurisdiction-precision score is low has a concrete and public reason to raise it.
For any of that to mean something, a neutral party has to run the benchmark. A vendor cannot grade the test it is sitting for. The credible owner is a convener working with academic and research partners, the kind of arrangement the cohort began with its evaluation partner, with subject-matter experts contributing the scenarios and the gold answers and funders covering the upkeep. The companies are the audience. The field holds the scorecard.
And the edge that is arriving fast: the models are being built into agents that act, not just chat boxes that answer. The question coming at the field is whether those agents can call legal-help MCP servers, and whether the field can verify what they call and whitelist what is safe. This has the least evidence of the four. It is also the clear direction of travel. Someone in the field needs to be watching the agent platforms now, deciding what to expose and what to whitelist, before the agents arrive and decide for us.
Where the bets go
These are not alternatives. The strategy is a portfolio across the four, weighted by what holds up and by what the field can hold onto.
The runtime content layer, place two, is where the cohort has data, and where teams keep their hands on their own tools. Publishing the feed plus a public benchmark, places three and four, is the broadest-reach and most durable combination, because it does not depend on any single company continuing to care. Direct partnerships, also place three, are worth having where the relationships already exist. They should not be the foundation. A partnership is revocable, and it seats the field inside someone else's product roadmap.
The state teams own the runtime content layer, supported by the KB Power Group and by shared MCP servers for the teams too small to build their own. The published feed, the partnerships, and the public benchmark are field-level work, held by a convener on behalf of everyone, because no single state can publish a national feed, negotiate with a model company, or keep a benchmark credible on its own. Funders underwrite that field-level layer. The model companies are the counterpart for places three and four, not the owner of any of it.
The line under all of it: the field keeps its footing when it controls the content and its terms and the benchmark stays public and neutral. The moment the strategy depends on one company's goodwill, a public good has been traded for a favor.
And the scope condition, worth repeating because the field keeps returning to the chat box: every one of these four has to work across the whole taxonomy. A model that answers an eviction question well but misreads the summons upload, routes the person to the wrong county, or fills the form with the wrong field has not solved the problem. The win is the model doing the work well across the taxonomy: the reading, the sorting, the drafting, the filing, the routing, and the answering.
What comes next
The pieces that follow are the evidence and the method behind these bets. Whether grounding actually makes the models better, and on which tasks. How the field can measure quality across fifty tasks instead of arguing about it. And, further out, who publishes the open feed and keeps the benchmark honest, because that is work no single state can carry alone.
These models will keep answering legal questions whether or not the field engages. The open question is whether they do it with the field's content and on the field's terms, or without them.