How Do We Find Out How Good AI Is, Across Fifty Tasks?
The field has a taxonomy of fifty legal help tasks and a real measurement of one of them. This guide sets out the honest gap, the choices for closing it, the partners the work needs, and the question of how the issue-area agendas and the task taxonomy fit together.
Common Legal Help AI. Working draft for comment. Margaret Hagan, Stanford Legal Design Lab.
The Legal Help Task Taxonomy lays out fifty tasks. The cohort measured one of them in full, which was legal question-answering, and even that evidence is early and rests on synthetic questions and five rounds of testing in May. The classification task is partly measured, against a 197-case ground-truth set. Four more tasks the cohort can speak to from building the tools, deploying some of them, and watching where they work, which is real knowledge and is not a scored evaluation. The other forty-four tasks have not been measured by anyone, as far as we know. That is the true state of the field's read on AI performance, and this guide is about how the field changes it.
What knowing one task actually takes
Before choosing how to measure fifty tasks, it helps to be precise about what measuring one took. For legal question-answering the cohort needed four things. It needed a clear task definition, which the taxonomy supplies. It needed a dataset of test cases, built synthetic so it could be shared. It needed a rubric for what a good answer looks like, written so a person and a machine can both apply it. And it needed a protocol for running models against the dataset and scoring them, which was promptfoo plus a deterministic verifier. Put those four together and you have an evidence pack for a task. JusticeBench is built to hold exactly this, because each task page can carry the task's definition, its datasets, its rubric, and its results. The inventory this guide is about is one of these packs per task, kept current.
The number is never just a number
A single score for a task hides the thing the grounding guide found, which is that AI performance depends on the content the model is grounded in and the jurisdiction it is answering for. A model may explain a standard federal form well and a county's local cover sheet badly. It may answer a housing question well in a state with a strong tagged knowledge base and badly in a state without one. So the real unit is a task, in an issue area, in a jurisdiction, against a particular model. Scoring the task on its own averages across all of it. That is a large matrix, and no one is going to fill all of it.
This is where the field's two ways of slicing the work meet. The Legal Help Commons runs issue-area agendas: housing, debt, reentry, public benefits. Those are vertical, and they ask what it takes to deliver good help in one domain from start to finish. The task taxonomy slices the other way. It names the task divorced from the issue area, so that document explanation or deadline calculation can be looked at across every domain at once. The two are complements, and the cleanest design connects them on purpose. The issue-area cohorts are where the cells of the matrix get filled, because a housing cohort building and testing its tools is already measuring the housing instances of classification, routing, and document explanation. The task taxonomy on JusticeBench is where those results get gathered by task, so the field can see how document explanation does across housing, debt, and family side by side. The vertical work produces the evidence. The horizontal taxonomy organizes it.
Who does the measuring
There are three ways to get the work done, and they are not exclusive.
The first is a central benchmark team. A national group runs evaluations across tasks and publishes them. This gives consistency and comparability, because the same hands run the same protocol every time. It also creates a bottleneck, because one team cannot keep up with fifty tasks across issue areas and jurisdictions while the model landscape changes every month.
The second is distributed contribution. Individual teams run their own evaluations on the tasks they care about, using the shared protocol and rubrics, and send their results back to a common registry. This is the model the measurement guide is built for, and it is the only one that scales, because the work spreads across everyone who has a stake in the answer. Its cost is comparability, because results are only poolable if teams use the same rubrics and report the same way, which means the standards have to come first.
The third is task stewardship. Each task, or a cluster of related tasks, gets an owner: an organization or a working group that builds the dataset, maintains the rubric, runs the benchmark, and keeps it current. This is how open standards and open-source projects survive. It gives depth and accountability for each task. Its cost is recruitment and sustainability, because a steward role needs a clear reason to exist and someone to fund it.
The honest recommendation is a hybrid. Hold the shared protocols and rubrics centrally, so results stay comparable. Assign task stewards for the priority tasks, so someone is accountable for depth. Use distributed contribution for everything else, so the work scales past what any one team could do alone. JusticeBench is the registry that ties the three together.
How often
A measurement is a snapshot of a model that will be replaced. An evaluation run in May can be stale by autumn. So the field has to choose a cadence, and there are two honest options that pair well.
The first is a yearly read. Once a year the field publishes a state of AI on legal help tasks: where each task sits, what improved, what slipped, what got measured for the first time. This is feasible, it is legible to leaders and funders, and it gives the field a rhythm to work to. The second is event-triggered re-runs on the priority tasks. When a major model is released, the stewards of the highest-stakes tasks re-run their benchmarks, because those are the tasks where a regression reaches a person with a real problem. The two together give a steady annual picture plus fast checks where the stakes are highest, without pretending the field can keep a live leaderboard running for all fifty tasks.
The partners the work needs
The measurement does not happen without people who are not only engineers. Each task needs subject-matter experts to write and validate its rubric and to do blind review, because a rubric for whether eviction advice is sound is a lawyer's judgment before it is a metric. The outcomes question, which is whether a better answer actually produces a better result for the person, needs researchers with causal methods, and there are credible homes for that in the Rhode Center, RAND's civil justice group, and the American Bar Foundation. Running evaluations across models needs access to those models, which is a place for provider relationships held so that no single provider sets the terms. And the whole inventory needs a neutral registry that belongs to the field rather than to a vendor, which is what JusticeBench is for. None of these is a solo job, and that is the case the national-leadership guide makes for a tier above any single state.
Where to start
The field cannot open all fifty tasks at once, so the choice is which few to take to full depth first. The sensible answer is to start where the issue-area agendas are already working, so the measurement rides on building that is already funded and underway. The housing, debt, reentry, and public benefits cohorts each touch a handful of the high-frequency tasks, and those instances are the first cells to fill. Take those to full depth, publish the protocol and the rubrics so the next team can extend them, and let teams add the tasks they care about through their own promptfoo runs. The forty-four untouched tasks are the invitation, and they are the reason to set up shared measurement now rather than after another year of tools shipping unmeasured.
What this guide is
This is a set of choices, not a plan the field has agreed to. Whether to fund task stewards, whether to commit to a yearly read, and how tightly to wire the issue-area cohorts to the task inventory are decisions for the field and its funders to make together. What the cohort can say with confidence is narrower and still useful. The method exists. The registry exists. One task has been measured end to end, which is the proof that the rest can be. The work now is deciding who measures what, how often, and with whom.