Home · Writing · Multi-Agent Systems

When Multi-Agent Systems Become a Jury for Government Text

Who speaks for groups that were never written into the category list? MARS-Gov explores that question through open-set standpoint screening.

Every public-facing government document—a benefits policy, an administrative procedure, or a notice—quietly shapes who is included and who is overlooked. Bureaucratic bias is often structural rather than malicious: it can hide in default assumptions and seemingly neutral language.

MARS-Gov asks whether a group of LLM-based agents can examine government text as a careful jury would: by surfacing potential bias, grounding claims in evidence, and recognising affected groups that fall outside predefined categories.

Why one model is not enough

Giving a single model a document and asking “Is this biased?” leaves three gaps:

  • A single viewpoint. One model can systematically overlook people outside the perspective encoded in its training data.
  • Weak grounding. Without legal and policy sources, a judgment can become unsupported intuition.
  • Poor auditability. A bare conclusion cannot be challenged or reproduced.

MARS-Gov turns bias detection from a one-shot classification task into a multi-agent governance workflow.

A closed loop from evidence to verdict

  1. Normative evidence contextualisation assembles the legal and policy basis for evaluation.
  2. Legal and policy retrieval finds sources relevant to the text under review.
  3. Open-set screening identifies potentially affected groups, including groups beyond nine predefined categories.
  4. Dual-jury deliberation combines fixed jurors with a dynamic tenth juror.
  5. Risk routing decides whether further review is needed.
  6. Verdict rewriting and verification checks the conclusion against the evidence trail.
⚖️

Core idea: move from a single classification decision to an evidence-backed, deliberative, and reviewable governance process.

The “tenth juror”

Conventional bias-detection systems rely on a closed list of categories, such as gender, ethnicity, and age. Yet some of the most important harms affect groups that are not on that list: people with rare diseases, temporary workers, or people excluded by a particular digital-service design.

When MARS-Gov detects a possible impact on such a group, it creates a dynamic tenth juror to represent that standpoint in the deliberation. Fixed jurors offer stability; the dynamic juror keeps the system open to unfamiliar cases.

A disagreeing role in a multi-agent system can be more valuable than ten additional prompts. Diversity is not noise; it is information.

Making conclusions accountable

The final verdict is rewritten and checked against the retrieved evidence, producing a trace from legal basis to affected group, juror opinions, and conclusion. That trace is what makes a governance-oriented system explainable and auditable.

📄

This article accompanies the co-authored paper The “10th Juror”: Open-Set Standpoint Screening for Bureaucratic Bias Detection, accepted to EMNLP 2026. Get in touch to discuss multi-agent governance or AI fairness.

← Back to writing