Working draft. Please do not circulate beyond this group yet.

Why Work From Shared Knowledge Standards

This is a guide to the groundwork that makes everything else interoperable: organizing your knowledge, your service directory, and your content from the same standards as your peers. It explains what the shared standards are, why working from the same ones pays off, and how to adopt them without rebuilding what you already have.

Common Legal Help AI. Working draft for comment. Margaret Hagan, Stanford Legal Design Lab.


A state can organize its content and its service directory however it likes, and most do, each in its own spreadsheet with its own column names and its own categories. That works inside one organization and breaks the moment anyone else tries to use the data, whether that is a peer state, an AI tool, or a partner platform. Shared knowledge standards are the fix. They are the agreement on what to name things and which values to use, so that one organization's knowledge base and directory can be trusted and reused by another. This guide covers what the cohort settled on and why it is worth adopting.

Read the full memo, Building a Federated Justice Knowledge Base: the value, the cost of not coordinating, the federation architecture, the content types, and the data contract →

What working from the same standard actually means

The Legal Help Knowledge Standard is the umbrella. It is at version 0.1, it is published under a CC BY 4.0 license, and it is designed to work in Airtable, Google Sheets, a database, or whatever a team already runs. It covers three connected databases: Organizations, which is who provides legal help; Services, which is what each organization offers and where matching happens; and the Content Index, which is the guides, forms, and tools available to the public. The full field lists and examples live at legalhelpcommons.org/standards. The point of it is that a team keeps its own system and still publishes data that lines up with what its peers publish.

The shared vocabularies underneath it

The standard rests on a few controlled vocabularies, and they are the part that makes data from two different states comparable. The LIST taxonomy is the shared vocabulary for what a legal problem is, so that eviction carries the same code everywhere. The seventeen standard audience categories describe who a service or a piece of content is for, so that referrals reach the right population. A team tags its directory and its content with these vocabularies, and its data becomes legible to anyone else using them.

Why working from the same standard pays off

The payoff shows up in four places. Referrals route correctly, because audience and geography use shared vocabularies a tool can read, and eligibility is recorded in structured fields rather than buried in free text. AI tools can actually use your content, because a model can pull the right passage for the right county when the content carries its jurisdiction and issue tags. Your work travels, because a peer state can trust and reuse your directory without re-cleaning it first. And you stop reinventing, because the field maintains the vocabularies and you adopt them instead of building your own. The cohort's stance is to build local and prepare to share: each team keeps control of its own data and publishes a small shared contract so that others can use it safely.

The service directory is where this matters most

The service directory carries the highest stakes, because a wrong referral sends a person to an organization that cannot help them. The standard service metadata covers the fields every service should carry, from legal issues and jurisdiction to languages, delivery mode, hours, capacity, and cost. Eligibility is the hardest of these and the least settled. HSDS and the Open Referral standard have openly acknowledged that eligibility is underspecified, and the cohort has not solved that: there is no finished eligibility standard. What a team can do today is compose the vocabularies that already exist, the audience categories for who a service is for and the jurisdiction list for where it reaches, and add a few plain fields it fills in directly, such as income relative to the federal poverty level, case type, and immigration status. Recording who qualifies that way, in structured fields rather than free text, is what turns a directory into something a routing tool can act on, and building a shared eligibility schema on top of it is open work for the field.

Federation, not one central database

None of this requires a national database that owns everyone's content. The federated knowledge base memo makes the case for federation over centralization, under one stance: build local, prepare to share. Each state keeps stewardship over its own data, the governance, editing, and quality assurance, and exposes a publisher connector that emits its content in a normalized format to a shared registry. Other tools, whether a state website, a cross-state agent, or a vetted partner, query a vendor-neutral retrieval API, with snapshots and update webhooks, rather than scraping or re-hosting anyone's content. The only thing standardized is the minimal data contract: jurisdiction, issue, audience, language, last-updated date, license, provenance, and citations. That aligns incentives. States stay the editors of record and set their own quality bar, and everyone gets one way to query everything with provenance carried through to each answer. The standard is what makes this possible, because a federation only works if every state's feed speaks the same language.

What it costs not to coordinate

The case for the standard is clearest in what happens without it. Every state and vendor builds its own silo, and the failures are predictable. Forms and deadlines drift until no one can say which page is current. Chatbots mix counties because no canonical jurisdiction tag stops them. Content goes stale with no owner and no way to see it. Answers carry no citation, date, or license, so errors cannot be audited or fixed. Teams re-prepare the same content for every new vendor, and knowledge gets locked into proprietary indexes. Cross-state agents and research cannot compose knowledge across incompatible schemas, and partners cannot consume content reliably without an API. The deepest risk is ceding the field's knowledge to big tech systems whose quality and business interests are unknown. A small shared contract prevents all of this: updates propagate, jurisdictions stay clean, and every answer shows its source and date. The memo lays out the full set of operational, ecosystem, and equity risks.

How to adopt it without rebuilding

Adoption is incremental. At a minimum, put five fields on every piece of content you publish from now on: jurisdiction, LIST code, last-updated date, license, and provenance. The Legal Help Knowledge Standard tells you exactly what to name the columns and which dropdown values to use, so this is a mapping exercise rather than a rebuild, and you backfill the rest at your own pace. Those five are the floor. The full per-record contract in the memo adds a stable id, a content type, the issue and jurisdiction codes, language, version, and a public-or-internal visibility flag, and it carries one rule that matters more than the rest: no procedural or deadline claim publishes without a supporting authority citation. The memo also sets out the steward roles that keep a knowledge base fresh and the initial cohort that agrees the contract, which is the governance all of this depends on. Record eligibility on your service directory next, composing the audience categories, the jurisdiction list, and a few plain fields, because that is where the routing payoff is largest. A team that does only this much is already interoperable with its peers, and the toolkit guide lists the vocabularies and the classifier that make the tagging tractable.

Where this lives in the Commons

These are the Commons' Knowledge Standards.