What We Need National Leadership On
The most durable parts of this work do not belong to any single state. The shared standard, the common tools, the public benchmark, the open feed to the model companies, and the evaluation and safety functions that get dropped when a team runs lean all point to a tier above the individual state. This guide makes the case for national leadership on that shared layer: what it would own, why it has to stay neutral, and why funding it once is cheaper than letting fifty states each pay for it badly.
Common Legal Help AI. Working draft for comment. Margaret Hagan, Stanford Legal Design Lab.
Across these guides, the same conclusion keeps surfacing. The strongest place to shape how the models behave is the content the field puts in front of them. The field can ground those models in real state content and measure whether the result is any good. The assets to do that work exist, and a team of any size can start. And the most valuable parts of all of it do not sit inside any one state.
The open feed that reaches the public through the model companies needs a publisher. The public benchmark that holds vendors accountable needs a neutral keeper, because a vendor cannot grade its own test. The shared standard that lets Illinois and Texas content work together needs an owner who keeps it current. The evaluation and safety functions that fall through the cracks at a lean team need somewhere to be carried. Every one of these points up, to a tier above the individual state, and that tier does not yet exist. This is the work that needs national leadership.
The floor that exists today
Through Legal Help Commons, the field has begun running common-tool working groups, on knowledge bases, on voice intake, and on OCR and data extraction, and it is learning from them. Each one forms around a shared problem, runs a focused workstream, and produces a shared artifact, on a monthly cadence with low overhead. The issue-area verticals, the groups organized around housing, debt, reentry, and public benefits, have just launched and are recruiting their first members. This working-group model is a floor rather than a building. It runs on volunteer time and convening goodwill, which means it can produce a standard, a classifier, and a set of rubrics, and it cannot guarantee that those things are maintained for the next decade, certified for the field to rely on, or funded so that the work does not depend on whoever has spare hours that month. The field needs a tier with the resources to do what a volunteer effort cannot.
What national leadership would own
A resourced national body would carry five functions, and each one is the long-term home for something these guides have already shown the field needs.
Standard-setting. It would own the Legal Help Knowledge Standard and its evolution, the LIST taxonomy, a shared way to record eligibility, which is still to be built, the rubric library, and the failure-mode registry. It would publish change notes when something updates and run a deprecation cycle when a field moves. This is the long-term home for the vocabularies the toolkit guide catalogs, and it matters because a taxonomy that no one maintains drifts out of date as the law changes, and a standard with no owner quietly forks into incompatible local versions.
Shared tooling. It would maintain the LIST classifier, the content safety wrapper schema and its starter scenarios, the synthetic document dataset, the federation gateway, and the evaluation harness templates. This is the home for the tools the toolkit guide catalogs, and it matters because a classifier needs retraining as the models underneath it change, and a safety wrapper needs new scenarios as new high-risk situations surface, and neither happens on its own.
Quality and certification. It would run the benchmark evaluations on the tools the field uses, publish the results, and certify tools and organizations against the standards. This is the home for the public benchmark the strategy and measurement guides describe, and it is the answer to the neutral-keeper problem, because the benchmark only means something if the body that runs it has no tool of its own in the rankings. It also gives a procurement officer a concrete number to require.
Coordination. It would convene the cohorts, run cross-cohort gatherings, keep the working groups connected, and host the publication of the artifacts. This is the connective tissue that keeps the field learning together instead of in parallel silos.
Funding flow. It would pool funding so that individual states are not each fundraising for the same shared infrastructure, and so that a single investment can be made once, at scale, with confidence that the work is being stewarded. This is the function that makes the other four sustainable, and it is the one the field most conspicuously lacks today.
None of this exists yet. It is a proposal for the field and its funders to weigh, not a plan anyone has adopted, and it does not assume any one organization as the home.
Why this should be funded once
The case is concrete, and it comes down to arithmetic. Right now, when a state wants its content tagged with LIST codes, or wants a safety wrapper for its chatbot, or wants to know whether a vendor's tool is any good, it either does without or it raises money to solve the problem for itself. Multiply that across the states, and the field is paying many times over for the same classifier, the same wrapper, and the same benchmark, and paying for them badly, because no single state has the scale to do any of them well. The shared layer is no one's job, so it is chronically underbuilt.
A national tier changes the arithmetic. A single investment, at scale, keeps the standard current, the tools maintained, the benchmark credible, and every state out of the business of reinventing the same infrastructure. The model companies are a counterpart in this, because the open feed and the public benchmark are what the field offers them, as the strategy guide described, and someone has to keep those alive for the offer to mean anything. The money that sustains the shared tier is the money that keeps the field at the table with the companies on its own terms, rather than dependent on whatever those companies decide to build without it.
What to do
For a state team reading this, the first moves are small and they are available this week. Locate yourself honestly on the journey from the getting-started guide, knowing that most teams are between Stage two and Stage three. Pick one function to invest in next quarter, and if you do not already have knowledge stewardship covered, start there. Pick one task to evaluate, with the document explainer as the easiest first AI task and brief advice as the highest-impact and highest-risk one. Adopt the shared standard on everything you publish from now on, even at the minimum of jurisdiction, LIST code, last-updated date, license, and provenance, and backfill the rest at your own pace. Join a working group, and tell the field what you find, because the practice of sharing what works and what does not, in a form another team can use, travels further than any single tool.
That invitation extends to your own projects, and not only to what you find using the shared assets. Several teams are building notable tools of their own, and those are theirs to describe, because the team that built the work is the one that should tell its story. If your team wants its project written up so the field can see it and learn from it, you are invited to write it up in your own words, and it can be featured so others can find it.
These guides opened on a plain fact, which is that the foundation models are answering the public's legal questions today, whether the field participates or not. Everything since has been about how the field gets those models to do that work well, across the whole range of tasks people need, and how it knows when they do. The field can build the content, the standard, the evidence, and the tools, and the work has started. What no one can do, state by state and grant by grant, is steward all of it for the long run. That is the work that needs national leadership and the resources to sustain it. The companies will keep doing this work either way. The only open question is whether the field meets them with a maintained standard, a credible benchmark, and an open feed it controls, or with a set of good prototypes and no one to keep them alive.