Working draft. Please do not circulate beyond this group yet.

Where Does Your Team Go From Here?

This is the guide a team reads to decide what to actually do next. Where a team goes depends on three things: what it is aiming for, how organized its content already is, and how many people it has. The honest starting point is that most teams are still getting their content in order, and AI built on top of disorganized content is just a faster way to be wrong. The useful news is that the shared assets from the toolkit guide let even a one-person team make real progress, and this guide is how to find the first move.

Common Legal Help AI. Working draft for comment. Margaret Hagan, Stanford Legal Design Lab.


The AI layer only works on top of content that has been kept current and a directory that has been structured. The toolkit guide laid out the shared assets that make that foundation cheaper to build. This guide is the practical one. It is about where a particular team should start, because a one-person team with a stale content library and a five-person team with a structured directory are not in the same place and should not do the same things first.

A team can locate itself on three coordinates. The first is the aim, which is how far the team wants to go. The second is the stage, which is how organized its content already is. The third is the size, which is how many people it has to do the work. Find yourself on all three, and the next move becomes clear.

Pick your aim first

A team should choose its aim before it chooses its tools, because the aim decides which work comes first. There are three tiers, and they stack.

Tier one is a reliable, well-organized foundation. The content is organized, current, comparable to peer states, and labeled well enough that other systems can use it, and the service directory is accurate. The team can point to its inventory and know what is in it. This is foundation work, it does not require any AI on top of it, and many teams are still partly here.

Tier two is AI-powered workflows. The team is building and evaluating AI features on top of the housekept content, such as a document explainer, a triage and referral router, a brief-advice chatbot with a safety wrapper, and a classifier that fills in the intake fields. The team is running evaluations and iterating, and the AI is improving the workflows for users and for staff.

Tier three is being the authoritative regional steward. The team becomes the supplier of authoritative, well-maintained, safety-wrapped content and service directory data for its region, so that other tools call it and other organizations rely on it. This is the tier that connects to the open feed the strategy guide described, because a Tier three team is exactly the kind of publisher that feed needs. It only makes sense once tier one is solid and tier two is working.

The tiers stack, and a team cannot skip the first one, and it should not jump to the third without the second. A team focused entirely on hotline or in-person service might have a different shape, and for any team that runs a public legal help website, these three are the aims that come up.

Find your stage

Inside each tier, a team moves through a four-stage journey, and this is the part the cohort watched play out across all seven states. It also matches what Heidi Behnke's RAILS work describes and what the KB Power Group has been mapping.

Stage one is wild content. The team has a lot of content, some of it good and some of it dated and some of it duplicated across pages, and it does not really know its own inventory. Here is the concrete test: a Stage one team cannot reliably pull everything it has on eviction in Cook County without someone searching by hand.

Stage two is good housekeeping. The team knows what it has, the content is organized and curated, the categories cover most of it, and the team can pull lists by topic and has a sense of its gaps. This is where many statewide legal help websites already are.

Stage three is a systematic pipeline. Stewardship is a real function with a cadence. Every piece of content carries its metadata, including jurisdiction, LIST code, language, last-updated date, license, provenance, authority level, and risk tier. Freshness review runs on a schedule, gaps are tracked as a backlog, new content gets its metadata at creation rather than in a later tagging sprint, and the data follows the Legal Help Knowledge Standard so it lines up with what peer states publish.

Stage four is AI-integrated. The AI tools run on top of the well-housekept content. Brief advice cites the right authority, triage routes to the right guide and the right service, the document explainer handles uploads, and the team iterates on the AI layer by adjusting prompts, adding safety wrapper rules, refining rubrics, and fixing failure modes, using evaluation results as the feedback signal.

Most cohort states are between Stage two and Stage three today, and a few are running Stage four experiments. The honest line bears repeating here, because it is the whole reason stage matters: AI without Stages two and three is just faster failure.

The six functions someone has to cover

Moving through the journey takes six functions, and every team already has an org chart, so the work is to locate these six inside it and decide how much time each one gets. The number of functions is fixed at six, and the number of people is not.

Knowledge stewardship owns the content metadata, the freshness review cadence, and the structural quality of how content gets chunked and tagged. In most teams this falls to the content editor or the managing attorney who owns the website.

Service directory coordination owns the structured listings, including eligibility, geography, capacity, hours, languages, and intake channel, and it verifies them on a cadence and records eligibility in structured fields. In most teams this is the intake supervisor or the referrals coordinator, often without dedicated time.

Brief advice interaction design designs the user-facing layer, including the chatbot prompts, the triage flow, the handoff logic, the escalation triggers, and the safety wrapper rules. In most teams this is the website manager working with an IT contractor, or it does not exist yet because brief-advice AI is not in production.

AI implementation and iteration actually builds, ships, and tunes the features, configuring the retrieval pipeline, writing and iterating the system prompts, wiring up the classifiers, connecting to the knowledge base, deploying, running experiments, and debugging when the model gives a wrong answer. In most teams this function is unstaffed, and teams rely on an outside developer or a vendor, and it is the function most often missing when AI adoption stalls.

DIY and guide resources produces and maintains the plain-language content, and in an AI era the writers work with structural awareness, so that action lists travel with their warnings, disclosures attach to the materials they qualify, and jurisdiction scope is built into the chunk. This is the traditional content writer role, evolved.

Quality, safety, and evaluation runs the rubrics, maintains the failure-mode registry, coordinates expert review, and decides what is pilot-ready, what is production-ready, and what needs to come down. In most teams this function is unfunded and falls through the cracks, and that gap is itself one of the cohort's findings.

The functions map onto the journey in a clean way. A team moving from Stage two to Stage three mostly needs knowledge stewardship and service directory coordination to get real time. A team moving from Stage three to Stage four needs to add brief advice interaction design, AI implementation, and quality and evaluation. The AI implementation function is where most teams hit a wall, because without someone who can build and ship, the design work and the evaluation work have nowhere to land.

How this works at one, three, or five people

Most teams will not staff up, so the path has to work at small sizes, and it does.

At one person, the work is possible because the shared assets from the toolkit guide substitute for capacity that the team does not have. The one-person team leans on the Legal Help Knowledge Standard, the LIST classifier, the safety wrapper starter, the evaluation rubrics, and the synthetic datasets, and it keeps a practitioner expert network it can pull on by issue area. AI implementation sits outside the team, with a contractor or a vendor. Here is what the rhythm looks like in practice, offered as an illustration. The one person cannot run a daily standup with themselves, so instead they run a monthly issue-area review: they pick one topic, such as eviction, do the full loop on it of content and directory and interaction and evaluation, ship it, and move to the next topic the following month. The cadence is slower, and the work still happens.

At three people, the team usually splits into a content lead covering knowledge stewardship and plain-language writing, an intake or directory lead covering the service directory and some quality work, and a technical or platform lead covering interaction design, AI implementation, and the federation work. The technical lead is the bottleneck role, and when that person is overloaded the AI work stalls, and shared infrastructure carries the rest.

At five people, the team adds a dedicated interaction designer and a dedicated evaluator, the technical lead can focus on implementation, and this is roughly where most cohort states are or want to be.

It is worth being honest about what gets dropped when resources are thin. The first thing dropped is usually evaluation, then sometimes safety review, then sometimes freshness review on the lower-volume topics, and AI implementation is more often outsourced than dropped, which brings its own tradeoffs around speed, cost, and institutional knowledge. The reason the shared infrastructure matters is that it makes these dropped and outsourced functions less catastrophic, because some of them can be carried at the national tier instead of falling on a team that cannot cover them. That national tier is the subject of the national-leadership guide.

What to do this quarter

The starting move depends on the stage, and it is a single concrete thing rather than a transformation.

A team at Stage one or two should put its real time into knowledge stewardship and service directory coordination, because that is the foundation everything else needs, and the Legal Help Knowledge Standard and the LIST classifier from the toolkit guide are the tools that make the tagging tractable. A team at Stage three that is ready to begin Stage four should run the evaluation methodology from the measurement guide on one real workflow, compare the answers with the knowledge base against the answers without it, publish the result, and join the working group forming for that issue area. The document explainer is a good first AI task because its failure modes are clear and its rubric is straightforward. Brief advice is the highest-impact task and also the highest-risk, so it is the one to wrap most carefully.

Here is what a quarter could look like for a three-person team sitting between Stage two and Stage three, offered as an illustration rather than a prescription. In the first month the content lead runs the LIST classifier over the untagged backlog and reviews the low-confidence cases, while the directory lead records eligibility for the top ten programs by composing the audience categories, the jurisdiction list, and a few plain fields such as income relative to the federal poverty level and case type. In the second month the technical lead stands up a document explainer on the now-tagged content and runs it against the synthetic document dataset. In the third month the team scores the explainer with the rubric library, brings in two practitioner experts to check the hardest cases, and decides whether it is ready for a limited pilot. At the end of the quarter the team has moved real content into Stage three and has one evaluated AI workflow, which is a genuine quarter of progress for three people who did not stop doing their day jobs.

What comes next

This guide was about what a single team does inside its own walls. The national-leadership guide is about the work no single team can do alone. The dropped functions, the evaluation and the safety review that fall through the cracks at low resource, the open feed that needs a publisher, and the benchmark that needs a neutral keeper, all point to a tier above the individual state. The question of who funds and stewards that shared layer, and who keeps it honest and current, is the one these guides have been building toward.

Where this lives in the Commons

Once you know your stage, the Commons' working groups are where the next step happens.