Working draft. Please do not circulate beyond this group yet.

Building a Federated Justice Knowledge Base

The full memo behind the standards guide. It lays out the value of a federated Justice Knowledge Base, the cost of not coordinating, the content that belongs inside, the technology to build and maintain it, the data contract each record must meet, and how a handful of states can start.

Common Legal Help AI. Working draft for comment. Margaret Hagan, Stanford Legal Design Lab.


Multiple teams are already building state-level knowledge bases, with different scopes (guides, rules, forms, services), formats, and tech stacks. Some need a strong knowledge base to power a RAG bot that answers visitors' questions. Some run a statewide legal help website with a lot of content and want it easier to manage. Some have realized that the new era of AI innovation depends on strong data management. For any of these reasons, local knowledge base development is happening everywhere, and that organic growth is healthy. It shows local leadership and a commitment to the infrastructure that powers access to justice technology. The worry is what comes next: local knowledge bases built to their own standards and setups, hard to maintain, and impossible to use outside the organization that built them.

There is an opening right now, before local work hardens in any one direction. If the public-interest legal community, especially the teams working on legal help websites, public education, forms, triage and referrals, and chat bots, aligns on a small set of interoperability rules, the field can build a far stronger and more sustainable body of legal knowledge. Local governance, workflows, content, and licensing all stay local. The one thing that travels is the output: each local knowledge base should be callable, combinable, and trustworthy across states, for search, for agents, and for external partners.

We are proposing that set of interoperability rules as the Justice Knowledge Base standard (JKB). This is not an attempt to take local content and control it nationally. The JKB standard is open, free guidance for any group organizing and structuring its access to justice content. A local group decides how to staff and manage its own knowledge base, and whether to attach licenses, costs, or other policies to it. This memo lays out an initial proposal: what value a local and federated knowledge base offers, what content belongs inside, what technology and standards house it, and how a team follows the standard.

The value of a federated knowledge base

A well-run local knowledge base is the fastest way for a state to turn scattered rules, forms, services, and how-to guidance into reliable public help. A federated knowledge base is how that same help travels safely across borders when it should. Locally, a knowledge base gives agencies, courts, and legal aid partners one place to keep the record straight: what the law says, what the current forms are, where to file and pay, who is eligible, and what the realistic next steps look like for a resident. Federated, it lets neighboring states, national tools, and external platforms speak to that local truth without copying or distorting it. Build once, use everywhere, always under the state's guardrails.

The guiding stance is build local, prepare to share. Each state keeps stewardship over its own knowledge base: governance, editing, quality assurance, and accountability. The only thing standardized is a minimal, shared data contract: jurisdiction tags, issue codes, audience, language, last-updated date, license, provenance, and citations to the underlying authorities. That small contract makes local outputs interoperable by default. A state structures its content however it likes internally and still publishes a predictable, machine-readable feed that other tools can consume safely.

Technically, the federation is local-first and interoperable by design. The pattern is simple. Local knowledge bases own their content and their processes. Each state exposes a publisher connector that emits normalized JSON to a shared registry or gateway. Consumers, whether state websites, cross-state agents, research tools, or vetted partners, query through a vendor-neutral retrieval API rather than scraping or re-hosting, with snapshots and update webhooks for bulk access and change alerts. The state's editorial flow stays intact, and others get a clean, contract-based way to call its content.

The model works because it aligns incentives. States keep control: they remain the editors of record and set the quality bar for their own residents. Everyone benefits from one way to query everything, with jurisdiction and issue guardrails enforced by the gateway, and provenance carried through to every answer. There is no need to replicate content into each model or vendor. Models and apps call the same retrieval service. That lowers operational risk, since there are fewer stale copies, reduces staff rework, since an update is seen everywhere, and improves public trust, since answers consistently cite state-owned sources. For each state, the payoff is concrete: a local knowledge base immediately upgrades your own website search, guides, and intake routing, and federating it turns that investment into a force multiplier across travelers, multi-jurisdiction issues, national research, and responsible partnerships, without ceding control.

If we don't coordinate

Without a federated standard and a small shared contract, every state and vendor builds its own silo. That means duplicate effort, inconsistent answers, and no safe way to combine or compare information across borders. The public pays in bad guidance and dead ends. Teams pay in rework, vendor lock-in, and stalled partnerships.

The operational and quality risks are concrete. Version drift and contradictions: forms, deadlines, and rules diverge across sites, and "which page is current" becomes unanswerable. Jurisdiction leakage: chatbots and search tools silently mix states or counties because there is no canonical filter to stop them. Staleness you cannot see: with no owners and no shared freshness view, outdated content persists unnoticed. No provenance trail: answers lack citations, dates, or licenses, so you cannot audit or fix upstream errors, and local tools get more answers wrong to the public.

The technical and ecosystem risks compound them. Reindex everywhere, forever: each model and vendor demands its own upload and format, so teams re-prepare the same content repeatedly. Vendor lock-in: knowledge gets trapped in proprietary indexes, switching costs rise, and quality assurance moves outside public institutions. Blocked federation: cross-state agents and research cannot safely compose knowledge across incompatible schemas. Broken partnerships: courts, platforms, and funders cannot consume content reliably without an API, snapshots, or webhooks. And ceding knowledge to others: consumers and legal teams become dependent on big tech companies' knowledge bases and models, whose investment in quality, and whose competing business interests, are unknown.

The equity and accountability risks land on the people least able to absorb them. Coverage goes uneven, as well-resourced states improve and others lag. Errors stay invisible without conformance tests like jurisdiction-precision and authority-backed-answer measurement. And missing per-record licenses chill reuse or prompt unauthorized copying. No federation means more work and less trust. A small shared contract and a common gateway give states control while making safe reuse possible, so updates propagate, jurisdictions stay clean, and every answer shows its source and date.

What belongs inside the knowledge base

The full list is below, and beginning with even some of it is meaningful.

  • Legal authorities: statutes and codes, court and local rules, standing and emergency orders, and key agency regulations, with pinpoint citations and effective dates. Ideally local laws too, including municipal ordinances, registries, inspections, and local aid programs.
  • Procedures and how-to guides: plain-language descriptions of reusable steps for common tasks, with prerequisites, exceptions, and detours such as mediation or administrative exhaustion.
  • Deadlines and counting rules: court versus calendar days, service-by-mail and tolling rules, triggers, and timers.
  • Forms, blank templates, and document assembly: official and local forms with IDs, revisions, acceptance rules, and languages, plus packets, field-level instructions, example filled forms, and the guided interviews and logic trees built to help people complete and file documents.
  • Example documents and filled-in templates: sample pleadings, letters, stipulations and settlements, and worksheets.
  • Filing and payment: where and how to file (e-file, mail, in person), fees and waivers, payment portals, notarization and certified copies, and e-filing and payment integrations.
  • Services and referral directory data: the legal aid organizations, courts, law libraries, community groups, and tools that can help, with geofences, eligibility, intake modes and hours, languages, and capacity signals, plus hotlines, self-help centers, emergency pathways, and specialized experts.
  • Case status, file access, and calendars: case-lookup connections, notification signups, remote-hearing rules, exhibit procedures, and calendar feeds.
  • Language and accessibility assets: translation memories, glossaries, approved phrases, plain-language variants by reading level, and accessible media scripts with alt text.

A knowledge base starts with the authorities that anchor everything else. Each authority item should carry a pinpoint citation, effective and expiration dates, a link to the source, and a short note on what it changes in practice. For example: "CA CCP section 1167, tenant must file an Answer within 10 court days," or "Cook County Rule X, mandatory mediation before filing." Anchoring every statement in primary sources prevents version drift and enables citation-backed answers, strict jurisdiction filters, and rapid updates when the law changes.

From there the knowledge base gathers guides, manuals, trainings, and other documentation, structured as atomic reusable steps: one chunk per how-to step, FAQ pair, rule subsection, or form-instruction block. Each step is tied to the authority that supports it, the deadline rule that governs its timing, the detours that can apply, and the prerequisites to meet. Normalizing deadlines and counting rules removes ambiguity about timing, so applications can calculate due dates and warn about them automatically.

Forms need a complete picture: statewide and local forms with IDs, revision dates, acceptance rules, languages, and example filled-in versions where shareable, with links to guided interviews and their underlying logic. Vetted examples show people what a good pleading, letter, or worksheet looks like. The service directory needs real-world constraints: coverage geofences, eligibility by income and case type and immigration status, intake modes and hours, languages, and capacity signals, paired with routing logic. Filing and payment content removes guesswork at the courthouse door. Case-status and calendar content supports the workflows that keep people informed and on time. And language and accessibility assets keep communication consistent, multilingual, and inclusive across every tool that draws on the knowledge base.

Where to draw content from

Where a team finds this content depends on its jurisdiction's setup. Good starting points: court and judiciary sites, clerk and division pages, and e-filing portals; legislature and agency portals and municipal code publishers; statewide legal help websites, law libraries, and bar referral sites; legal aid directories, community organizations, and 211; internal document management systems and document-assembly repositories and case-management exports; internal training decks and manuals; translation memories and glossaries and style guides; interviews; and selective sharing among trusted colleagues.

A checklist for state teams

  • Coverage: each content type has a plan to locate, collect, structure, revise, and refresh it.
  • Tagging: every item carries jurisdiction, issue (LIST), audience, language, last-updated, license, and provenance at a minimum.
  • Structure: each item is split and attached with citations as needed, at the level of chunking the content type requires.
  • Policy: an owner is assigned, a refresh cadence is recorded, and the item passes schema validation.

The technology for building and maintaining a JKB

Gathering the data

A durable federated knowledge base starts with disciplined intake. Content flows in three reliable ways: publisher connectors from each state's content or document management system, emitting normalized JSON; respectful crawling of public sites using sitemap-led extraction with HTML and PDF normalization; and managed uploads for internal training docs, templates, and translation assets. Each path feeds a bronze, silver, gold pipeline. Bronze stores raw content exactly as received. Silver normalizes and chunks it into legal units and adds citations. Gold has passed quality assurance and is safe to serve to apps and partners. This separation lets teams onboard any current system while converging on one clean knowledge base. As content is gathered, it is mapped to the core taxonomies the national group maintains, with a shared issue taxonomy such as LIST codes, crosswalks from local tags, and controlled vocabularies for jurisdictions, audiences, content types, and form IDs.

Storage and access

Storage works best on two shelves. The official shelf is a state-controlled database where the approved, current versions live, with dates, sources, and editor of record. The display shelf holds fast, searchable copies so websites and tools can find the right paragraph quickly. The original files, the PDFs, forms, and videos, stay in ordinary cloud folders and link back to the official records. Fix content once on the official shelf and the change appears everywhere people view it. For access, avoid vendor lock-in. Rather than uploading the state's content into many AI products, the knowledge base offers one doorway that any website, chatbot, or partner can use to ask for exactly what it needs: the steps for this issue in this county in Spanish, or the chunks that mention this form with their citations. It can also provide nightly downloads by state or topic and send alerts when something important changes.

Format

The format should be simple and consistent. Each item, whether a statute, a how-to step, a deadline rule, or a form, carries the same small set of fields: where it applies, what issue it covers, who it is for, what language it is in, when it was last updated, what license it carries, and where it came from. If an item tells someone to do something by a certain time, it must include the exact rule or statute that supports the instruction. A plain, web-standard format lets each state manage content however it likes internally while still speaking the same language when content is shared or combined.

Maintenance and refresh

Maintenance is a matter of ownership and cadence. Each high-traffic topic gets a named owner. The team sets reasonable refresh targets, for example forms updated within a week of a new revision, service directory entries checked monthly, and high-volume topics like eviction reviewed twice a year, and the system flags items that miss their target. Automatic checks help: link checkers, form-version monitors, and validators that stop an item from going live when a citation or jurisdiction tag is missing. With these habits, lawyers and website managers can focus on substance.

Security, access, and licensing

The knowledge base operates under least-privilege, data-minimizing principles. Public guidance is separated from sensitive operational data, and internal content like capacity signals is scoped to authenticated roles. Controls include role-based access, per-record visibility, audit logging, encryption in transit and at rest, and standardized data-processing agreements with contributors. Downstream tools receive only the fields allowed for their audience and purpose. For licensing, adopt clear, permissive licenses for public content, such as CC BY 4.0, to maximize reuse while preserving attribution. Internal or premium datasets may carry restricted licenses or memoranda of understanding. For sustainability, states may pursue service tiers, cost-sharing consortia, or foundation-backed stewardship, but licensing must never restrict the public's access to core legal guidance and citations.

How local teams get started

An initial cohort to define the data contract

Adopting the standard works like complying with a style manual: clear, prescriptive requirements for each record, with latitude for local practice. The first step is for a handful of interested state teams to agree on the basics: what content to include, which fields to attach and how, and how to collect and structure the knowledge in practice. The teams agree on how to segment long materials into reusable units so they perform well on the desired tasks. For any instruction or deadline, the supporting authority is attached at the pinpoint level, a steward is identified, and a review date is set. A cross-state working group maintains the shared contract, issues release notes, and runs compatibility checks before any change is adopted, with backward-compatible defaults and a trial period.

Building local steward roles

Sustained governance keeps the standard alive without slowing daily work. Each state designates a steward or steward team for local implementation and quality control. Human oversight approves citations, resolves conflicts, and reviews high-risk outputs like deadlines, service, and fee waivers. As scope grows, the steward gains authority to set freshness targets, adjudicate conflicts, and publish official corrections, with a documented resolution workflow that prioritizes controlling authority and public safety. Editorial checklists and subject-matter review are part of publication, and frontline feedback continuously improves quality, especially when expert judgment is captured as structured updates. Adoption also takes culture, training, and integration: clear roles, lightweight workflows, and tools that meet people where they work. Funders can anchor the work by supporting stewardship, shared infrastructure, and recurring training.

Tools to support the stewards

Light tooling helps. CMS extensions or structured forms enforce the required fields, building on existing Drupal or WordPress systems. Validators block publication when citations or jurisdiction tags are missing. PII maskers keep internal templates from exposing private data. Link-health monitors catch broken sources. Training emphasizes practical skills: turning narrative guidance into atomic chunks, recording pinpoint citations, and setting refresh cadences to match content risk. Expansion proceeds incrementally, starting with high-traffic scenarios like eviction answers or fee-waiver filings and publishing exemplary records others can adapt. Over time this builds a consistent, auditable foundation, with local stewardship preserved and interoperability assured. The knowledge base then becomes a living system: de-identified feedback, like "this did not apply here" or "the court rejected this form," becomes a structured quality signal that prompts editorial updates, within privacy-preserving and governance-approved learning loops.

A federated knowledge base must gather everything a person or provider needs to understand, decide, act, and finish a legal task, across law, procedure, forms, services, filing and payment, human help, follow-through, and negotiation, stitched together with citations, jurisdiction, and dates. If each state compiles locally to this contract, the field can pipe content together across states, power consistent agents, and support partners, without losing control, freshness, or provenance.

The data contract

Each record validates against this spec before it publishes to the shared gateway.

Required fields: a stable globally unique id; a content_type (authority, howto_step, deadline_rule, form, example_doc, filing_payment, service, case_status_calendar, or language_asset); a human-readable title; the canonical text or body; jurisdiction (a FIPS or court-district code, for example US-CA-Los_Angeles or US-TX-Travis-JP3); issue_type (a LIST code); user_type; language (BCP-47, for example en or es); last_updated (ISO date); provenance (source organization, source URL, and a hash); license (a URL or SPDX identifier, for example CC-BY-4.0); version (incrementing, with history kept); and visibility (public or internal).

Other fields where useful: risk_flags such as deadline, service_of_process, emergency, or immigration_consequence; owner_org or owner_person, the accountable steward; and review_due, the next review date computed from the cadence.

Relationship fields where applicable: authority_ids, the pinpoint citations that support claims; form_ids, the referenced forms with revision; service_ids, the matching referrals; and parent_id, the guide, FAQ, or packet an item belongs to.

Validation rules: no procedural or deadline claim publishes without at least one supporting authority_id; jurisdiction and issue_type must be present and valid or the record does not publish; and the license must be explicit for reuse and partner exports. A content spec and metadata contract like this lets every state build locally while publishing outputs that are interoperable by default, with jurisdiction discipline, citations, freshness, and provenance built in.

  • The knowledge data types brainstormed for the JKB: Airtable
  • Where the seven states' key data types live today: Airtable