Does Grounding the Model in Our Content Actually Make It Better?
The strategy guide argued that the strongest place to shape a foundation model's behavior is the content the field puts in front of it. This guide is the evidence for that claim. Five rounds of testing in May 2026 showed that giving a general model a verified list of our own resources made it score higher on every model we tried, and that loading the full article text helped when the article fit the question but could hurt the strongest models when it did not. These were early results on synthetic questions rather than deployed tools, and that distinction matters.
Common Legal Help AI. Working draft for comment. Margaret Hagan, Stanford Legal Design Lab.
The strategy guide described four places where the field can change how a foundation model behaves, and it argued that the second place, which is changing what the field puts in front of the model at the moment it answers, is the one place where the cohort actually has evidence. This guide reports that evidence. The whole cohort bet is that grounding a general model in real state legal help content beats the model's general knowledge, and the Justice Knowledge Base exists because of that bet. In mid-May 2026 we tested the bet across five rounds, three models, and synthetic questions built from real legal help queries. The purpose of this guide is to report what we found and to be just as clear about what we did not test.
The questions were synthetic, drawn from a representative sample of legal issues using the LSC Justice Gap report and the HiiL and IAALS legal needs studies, and they were phrased from a database of real LiveChat queries that cohort partners shared and we rewrote into synthetic versions. The three models were Claude Sonnet 4.5, Gemini 2.5 Pro, and GPT-4o-mini, chosen to cover the range of models that teams actually deploy. Each round cost under twenty-five dollars and ran in under ninety minutes. Every answer was scored on a composite scale from zero to one, so a lift of a tenth of a point is a real move on that scale.
Three ways to ground an answer
When a person asks an AI legal help tool a question, there are three things the tool can do with the Justice Knowledge Base, and they line up with the choice the strategy guide described. The tool can ignore the knowledge base entirely, so that the model answers from its training alone with nothing in front of it. We called this condition C0, the bare baseline, and it is close to what the public gets today when it asks a general chatbot a legal question. The tool can instead hand the model a list of the relevant articles, giving it the titles, the links, and short descriptions, and then let the model answer from that list. We called this condition C1, directory grounding. Finally, the tool can hand the model the articles themselves, loading the full text of two or three relevant articles into its context so the model answers from the actual content. We called this condition C1.5, full-content grounding. The last two conditions are two versions of the same approach the strategy guide named, which is changing what goes into the model at the moment it works. The five rounds tested these three conditions at rising levels of difficulty, beginning with routine Illinois eviction questions and ending with the highest-risk edge cases the field worries about most.
The first finding: a verified list of your own resources helps
We started small. Round 1a was twenty-two Illinois eviction questions, and it pitted the bare model against directory grounding. Directory grounding won on both models we ran, lifting Claude by 0.073 and Gemini by 0.132. The bet paid off in the simplest case.
The obvious worry was that Illinois eviction is unusually well covered, and that the result would not travel. So Round 1b scaled the test to forty-eight questions across six states, which were Illinois, Oregon, West Virginia, Michigan, Ohio, and Texas, and across seven issue areas, on all three models. The lift held everywhere. Claude gained 0.132, Gemini gained 0.153, and GPT-4o-mini gained 0.180. Every state, every issue area, and every model moved in the same direction.
Two results from that round are worth pulling out. The first result is about cost. GPT-4o-mini with directory grounding scored 0.924, while Claude with no grounding scored 0.762. A small, inexpensive model with the right content in front of it outscored a premium frontier model working from memory. This is the clearest evidence behind the strategy guide's argument for setting the owned-model option aside, because a team does not need to train and maintain its own model to get good answers when a cheap general model with the right content already beats a premium model without it. For a legal aid organization choosing what to deploy on a real budget, that is the finding with the most immediate consequence.
The second result is about what kind of content matters. One state's content held 205 records in the JKB Content Index, while another state's held 2,414. The smaller, well-structured set produced a 0.239 lift, and the larger set produced a 0.034 lift. More records did not mean more help. The lift came from content that was well structured and relevant to the questions, not from raw volume. Quality and fit beat count.
| State | JKB index size | Average lift |
|---|---|---|
| Ohio | 205 | +0.239 |
| Illinois | 924 | +0.235 |
| Michigan | 363 | +0.215 |
| Oregon | 213 | +0.170 |
| West Virginia | 301 | +0.037 |
| Texas | 2,414 | +0.034 |
It is worth being honest about where the early lift came from. Most of it came from the deterministic verifier catching concrete errors in the bare model, such as links to pages that did not exist, an out-of-state aid organization cited in an answer for a state it does not serve, and a real phone number attached to the wrong organization. Directory grounding fixed those errors because it gave the model a verified, in-state list to anchor to. The measurement guide is about that scoring machinery and why it matters so much, so I will leave the full explanation there. For now the practical point stands, which is that the list helped.
Here is what this looks like in practice, and I am describing a scenario rather than a measured deployment. Picture a statewide legal help website team whose guides and forms are already tagged by issue and jurisdiction. Directory grounding is the cheapest and most reliable win available to that team, and it is the one a small team can actually reach. The team gives the model a current, verified list of its own resources for the issue and the county, and it checks the links and phone numbers before the answer goes out. The team that can do this is any team with a tagged content index and a service directory, which describes most of the cohort. The pieces that make it work, which are the content index and the list of known-real URLs, are reusable, so the next states do not have to build them from scratch.
The second finding: the full article helps when it fits, and can hurt when it does not
The harder question was whether loading the actual article text, rather than just the list, would help more. This is where the story stops being simple.
Round 2 put the full article text in front of the model on ten high-stakes Illinois housing questions. The first scoring showed a tie, because full-content grounding scored about the same as the directory list, around 0.95, on all three models. More content did not look like better answers. Then we read the answers side by side, by hand, and we saw that the full-content answers were better in ways the scoring had missed. They quoted the article instead of paraphrasing it, they named specific form numbers and deadline counts, and they admitted when the source did not cover something. The rubric we were using could not see any of that. So we sharpened the rubric, adding criteria for specific actionability, source fidelity, honest uncertainty, and whether quoted text actually appeared in the source. Rescored with the sharper rubric, full-content grounding separated from the list cleanly, lifting Claude by 0.151, Gemini by 0.092, and GPT-4o-mini by 0.125. The tie had been a measurement problem rather than a fact about the architecture. That lesson belongs to the measurement guide, but it changes how to read this one, because the first answer was wrong only because we were not measuring well enough to see the difference.
Round 3 took the sharpened test to the hard cases. It used forty high-risk questions across six states and sixteen scenarios, drawn from a 192-question dataset that subject-matter experts had flagged. The clean win came apart. Full-content grounding still helped GPT-4o-mini, by 0.132, but for Claude it was a tie at 0.002 below the list, and for Gemini it slightly hurt at 0.050 below the list. At the highest-risk tier, which was five questions covering criminal plea consequences for noncitizens, custody across state lines, and irreversible plea decisions, full-content grounding actively hurt the strong models. Claude dropped 0.228 below where it scored with the list alone, Gemini dropped 0.114, and only GPT-4o-mini gained.
The strategy guide warned that better access to content does not guarantee a correct answer, and this is the sharper version of that warning. Grounding the model in content that did not fit the question did not merely fail to help on the hardest questions, because it pulled the strongest models away from the better answers they would have given on their own. One example came from a custody article that covered ordinary modifications but said nothing about exclusive continuing jurisdiction under the UCCJEA, and the strong model, handed that article, anchored on it and missed the issue that actually decided the question.
The slice that resolves the picture is the one where experts confirmed that the article genuinely fit the scenario. There, every model gained from the full article, with Claude up 0.135, Gemini up 0.143, and GPT-4o-mini up 0.060. Where experts flagged the content as uncertain or gap-flagged, the full article hurt, dropping Claude by 0.07 and Gemini by 0.22. So the finding is not that full-content grounding is good or that it is bad. The finding is that full-content grounding helps when the content fits the question and can hurt when it does not, and that expert review of whether the content fits is the lever that decides which outcome you get.
I want to be honest about the size of that confirming slice, because it is only four questions. The pattern it shows is clear, but four questions is a small basis, and a larger expert-validated set is needed before anyone leans hard on it.
Here is what this means in practice, and again I am describing a scenario rather than something we observed. A vendor that says it grounds every answer in your content is not, by that fact alone, giving better answers, because on the questions that matter most, grounding in content that does not quite fit can make a strong model worse. Picture a legal aid director running her own highest-risk scenarios through a vendor's tool and checking whether the grounded answers actually beat the ungrounded ones. If they do not beat the ungrounded ones, the architecture is not ready, whatever the demo showed. The people who can act on this are procurement officers, funders writing requirements into a request for proposals, and any team choosing an architecture. For the content teams, the questions where the experts found gaps are not a failure but a roadmap, because they are a precise list of where state content needs work for the high-risk edge cases, and every hour of expert review buys quality that compounds across every tool that later uses the content.
What we did not test
The results above are descriptive means on modest samples, and they are not significance-tested effects. The scope was deliberately narrow, and it is worth stating the limits plainly.
We tested synthetic questions, not real users. We did not test retrieval, because in every round the articles were curated or matched outside the model, which means this work measures whether good content helps and not whether a system can find the right content on its own. We did not test whether answers stay stable over time. We did not test any model beyond the three we named. We did not test a deployed, end-to-end RAG system. We covered six states and a focused set of high-risk questions. The strongest signal, which is the confirmed-fit slice in Round 3, rests on four questions.
There is one more limit that ties directly back to the strategy guide. Every one of these tests was on the legal question-answering task. The strategy guide called that task one tile in a fifty-task taxonomy, and it warned against treating question-answering as the whole of the work. The grounding principle should extend to other tasks that depend on local content, such as explaining a document or routing a person to a service, but we have not tested that, and this evidence speaks to question-answering alone.
One operational note is worth flagging for anyone weighing Gemini for a deployed tool. Sixteen Gemini cells in Round 3 came back as Google API errors rather than answers, which is a four percent error rate, and we excluded those from the means. That error rate is its own kind of finding.
These results are enough to guide architecture and procurement decisions now. They are not enough to publish as settled fact, and we will say so plainly wherever the data is thin.
What comes next
The clear part holds up. Grounding a general model in real state legal help content made it measurably better on these synthetic questions, on every model we tested, and a cheap grounded model beat a premium ungrounded one. For directory grounding, the case is strong.
The careful part is the more useful finding. More content is not automatically better, and on the highest-risk questions, loading content that does not fit can pull a strong model off the right answer. That nuance was invisible until we measured well enough to see it, which is the subject of the measurement guide. The cases where full-content grounding backfired, which were the high-risk edge cases with incomplete content, are exactly the ones the Content Safety Wrapper is meant to catch, so that is where the architecture work goes next, and now there is evidence for building it rather than a hunch.
Read against the strategy guide, this is the evidence for its second strategy, and it carries a warning that runs through the rest of these guides. The field cannot know whether any of this is working, on question-answering or on the other forty-nine tasks, without measuring it carefully. That is where the measurement guide goes.
See the methods and materials: prompts, rubrics, scoring, datasets, and reproduction steps →