Every public-facing government document—a benefits policy, an administrative procedure, or a notice—quietly shapes who is included and who is overlooked. Bureaucratic bias is often structural rather than malicious: it can hide in default assumptions and seemingly neutral language.
MARS-Gov asks whether a group of LLM-based agents can examine government text as a careful jury would: by surfacing potential bias, grounding claims in evidence, and recognising affected groups that fall outside predefined categories.
Why one model is not enough
Giving a single model a document and asking “Is this biased?” leaves three gaps:
- A single viewpoint. One model can systematically overlook people outside the perspective encoded in its training data.
- Weak grounding. Without legal and policy sources, a judgment can become unsupported intuition.
- Poor auditability. A bare conclusion cannot be challenged or reproduced.
MARS-Gov turns bias detection from a one-shot classification task into a multi-agent governance workflow.
A closed loop from evidence to verdict
- Normative evidence contextualisation assembles the legal and policy basis for evaluation.
- Legal and policy retrieval finds sources relevant to the text under review.
- Open-set screening identifies potentially affected groups, including groups beyond nine predefined categories.
- Dual-jury deliberation combines fixed jurors with a dynamic tenth juror.
- Risk routing decides whether further review is needed.
- Verdict rewriting and verification checks the conclusion against the evidence trail.
Core idea: move from a single classification decision to an evidence-backed, deliberative, and reviewable governance process.
The “tenth juror”
Conventional bias-detection systems rely on a closed list of categories, such as gender, ethnicity, and age. Yet some of the most important harms affect groups that are not on that list: people with rare diseases, temporary workers, or people excluded by a particular digital-service design.
When MARS-Gov detects a possible impact on such a group, it creates a dynamic tenth juror to represent that standpoint in the deliberation. Fixed jurors offer stability; the dynamic juror keeps the system open to unfamiliar cases.
A disagreeing role in a multi-agent system can be more valuable than ten additional prompts. Diversity is not noise; it is information.
Making conclusions accountable
The final verdict is rewritten and checked against the retrieved evidence, producing a trace from legal basis to affected group, juror opinions, and conclusion. That trace is what makes a governance-oriented system explainable and auditable.
This article accompanies the co-authored paper The “10th Juror”: Open-Set Standpoint Screening for Bureaucratic Bias Detection, accepted to EMNLP 2026. Get in touch to discuss multi-agent governance or AI fairness.
