How to Audit an Algorithm Used by a Government (October 2026)

You audit a government algorithm by mapping the decision it influences, establishing the legal authority behind it, obtaining the records that describe how it was built, testing its outputs across affected groups, checking whether human review and appeals are real, and publishing what you found with deadlines attached. Most people who want to run one are researchers, journalists, civil-society staff or municipal employees. The hard part is rarely the statistics; it is getting evidence out of an agency that bought the system from a vendor under a contract with a confidentiality clause.

An algorithmic audit is an independent, evidence-based examination of how an automated decision-making system actually behaves — its documentation, its inputs, its outputs and its real-world effects — carried out to determine whether it is lawful, fair, safe and accountable. It is not a review, an impact assessment or a penetration test. Those are separate exercises with separate outputs.

Last reviewed: October 2026. Regulatory references below reflect the position as of that date, and several are still moving.

Table of Contents

What You Need

An audit fails in the first month without six things in place. Work through them before you send a single records request.

A system, or at least a suspicion of one. The hardest practical problem is that the system often does not exist on paper. Internal audit teams describe the same pattern repeatedly: a tool gets deployed, no owner is assigned, no risk acceptance is recorded, and nobody can later say who decided what. Start from the procurement register, the capital projects log, any published AI inventory, and the department that handles the queues people complain about. Many agencies do not know what they own.

A named decision. “The benefits system” is too vague to audit. “The decision to flag a benefits claim for investigation before a caseworker sees it” is auditable. The narrower your framing, the more testable the output.

A records route. Know which freedom of information law or state public records act applies to the agency, what the response clock is, and whether your jurisdiction exempts deliberative material or trade secrets. See the checking authority step below.

Ground truth. To test whether a score is fair you need to know what actually happened to the people it scored: who was investigated, who was cleared, who won an appeal, who paid back a debt that was not owed. Without outcome labels you can still test process, documentation and procedure. Say so in the report rather than implying you measured accuracy.

Some technical capacity. A city with one data analyst can run a desk audit and an output test on a sample of a few hundred records. That is enough to establish whether an audit is warranted. It is not enough for a code inspection, and pretending otherwise produces a report nobody trusts.

An independence story. Who funds you, who reviewed the draft, and what you agreed not to publish. Practitioners treat the arrangement used by the Pymetrics audit with Northeastern University — grant agreed before results were known, published in full afterwards — as the credibility benchmark.

Step-by-Step

A government algorithm audit runs in seven stages: define the system, check its authority and transparency, audit the data and feature design, test performance and bias, evaluate human review and contestability, validate findings with affected people, then publish findings with owners and deadlines. Each stage produces a document you can hand to someone else.

  1. Define the system, its purpose and its scope
  2. Check legal authority, transparency and accountability
  3. Audit the data and the feature design
  4. Test performance, bias and real-world effects
  5. Evaluate automation, human review and appeals
  6. Validate findings with affected stakeholders
  7. Write findings, risk ratings and remediation steps
Step-by-Step

How to audit an algorithm used by a government: define the system

Start by writing down what the system does in one paragraph a caseworker could read. Inputs, outputs, who acts on the output, where the output appears in the record, and which vendor built it. Then separate the algorithm from the policy around it, because most confusion in public debate comes from treating a rule and a model as one thing.

You can tell them apart with two questions. Does the outcome change if you feed different data to it? If yes, there is a model. Does the outcome change if you feed identical data twice? If no, it is a rule, a workflow, or a human judgement written down. Plenty of “algorithmic” systems under public criticism are the second thing wearing the first thing’s marketing.

Fix your scope in writing before you gather evidence: the questions you will answer, the evidence standard you will hold, the limits of what you can reach, and what would count as a result. This protects you when the agency releases less than you asked for.

Four kinds of audit cover almost everything, and they prove different things:

Audit typeWhat it examinesEvidence it needsWhat it can prove
Documentation reviewPurpose, governance, impact assessments, contract termsPolicies, AIA, procurement recordsWhat the agency claimed it was doing
Output testingScores and decisions across groups and over timeDecision records with ground truthDisparate impact and error patterns
Code and dependency inspectionFeatures, training data, pretrained components, driftSource code, model files, data samplesWhy the output looks the way it does
Impact assessmentEffects on real people, workflows and rightsInterviews, case files, appeal outcomesWhether harms materialised in practice

Check authority, transparency, and accountability

Every automated decision a government makes needs a legal basis. Find the statute, ordinance, delegation or policy that authorises it, and write it down. A system with no traceable authority is a finding in itself, and it is one of the easiest findings to prove.

Alongside that, pull the records that should exist: the algorithmic impact assessment, the data protection impact assessment, the procurement file, the contract with any audit-rights clause, the vendor’s own fairness documentation, and the internal approval that named an accountable owner. Absence is informative, but record it carefully — the agency may hold a document you have not asked for by name.

Note what transparency the agency already publishes. Canada’s Directive on Automated Decision-Making requires an impact assessment tier for every automated system and publishes the list. The EU AI Act obliges providers and deployers differently depending on risk tier. US federal agencies inventory their AI use cases under OMB guidance. Where nothing is published, you are auditing in the dark, and the report should say so in plain words.

Audit the data and feature design

Ask where each field came from and when it was collected. Historical administrative data carries history: enforcement records reflect who was policed where, arrest data reflects who was arrested where, and both encode past differences in police presence. A model trained on that is learning the pattern, whether or not the protected attribute appears in the column list.

That is the proxy variable problem. Postcode, surname, school, prior address and name frequency all carry protected characteristics without naming them. Ask directly which features the system uses as proxies and whether anyone tested for that.

Check missing values. A blank field that gets encoded as a zero is a policy decision, and it usually penalises whichever group is recorded less completely. Ask how many records have gaps in each field, and whether any imputation rule differs by postcode, surname or caseworker. Check label quality too: if “fraudulent” was assigned historically by the same rules you are auditing, the outcomes cannot validate the model.

Finally, ask what the model version is, when it last retrained, and how the agency knows performance has not drifted since launch. A system scoring well in a vendor’s demo and poorly on live data is a common and entirely ordinary outcome.

Test performance, bias, and real-world effects

Start with selection rates. For each affected group, what share received the score, the referral, the flag or the denial? Compare the highest rate against the lowest: if the ratio falls below 0.8, that is the four-fifths rule used in US employment discrimination doctrine, and it is a screening trigger rather than proof.

Then check error rates, because a system can hit an acceptable selection ratio and still be badly wrong. Measure false positives and false negatives separately per group, and look at the metric that matches the harm. A false negative in fraud detection costs a caseworker an hour. A false positive in a benefits suspension costs someone rent money.

Present more than one fairness metric. They conflict by construction, and a report that quotes only the flattering one is the thing you are criticising other auditors for. Say which metric you chose and why it fits the decision’s consequences.

Look at time. Effects that are invisible in a single year can be stark across five. When testing without source code, generate counterfactual or synthetic records that vary one attribute at a time and compare scores, and treat the result as a probe rather than proof.

Evaluate automation, human review, and contestability

Find out what the system actually decides versus what it recommends. Ask a caseworker or officer to walk you through the last ten decisions in their queue, and watch where the score sits in the sequence. If overriding it is possible but never happens, or requires a written justification nobody writes, human review is decorative.

Read the appeal records. Count outcomes: how many appeals succeeded, how long each took from decision, and whether the person appealing ever learned that software was involved. A system that decides in seconds produces an appeal process measured in months. That gap is a finding.

Test whether the reasoning is understandable. Ask the agency to explain a specific score to you in language it could use with the person affected. If the honest answer is that the score cannot be explained in any form, that answer is publishable.

Validate findings with affected stakeholders

Interview the people the system scores. Residents who had a claim flagged, frontline workers who override or rubber-stamp the tool, civil-society caseworkers, and people who won an appeal after a long wait.

Run the sessions knowing the limits. Volunteers skew toward people with time and confidence, which skews against exactly the groups most exposed to the harm. Say in the report who you reached and who you did not. Where you are documenting individual experiences, get consent, strip identifiers, and describe patterns rather than cases where consent is thin.

Write findings, risk ratings, and remediation steps

A report nobody can act on is a press release. Structure yours as: executive summary, system description, methodology, limitations, findings with severity ratings, required corrections, monitoring measures, named owners, deadlines.

Severity ratings force a decision. Something like low, moderate, high and critical, with the reasoning shown, makes it obvious which findings stop the system and which get a fix inside the next release cycle.

Name a person or office for each corrective action and a date. Findings with no owner are absorbed silently, which is precisely what happened across years of automated fraud scoring until the public reporting arrived.

Before publishing, check what you hold. Quote the minimal amount of internal text that supports a finding, describe confidential or personal material in summary form rather than reproducing it, and separate what the evidence shows from what you infer. Legal exposure is real, and it is a reason to tighten claims, not a reason to sit on a finding that affects people’s benefits.

Common Mistakes

Treating fairness as one number. There is no universal fairness metric, and the ones that conflict are supposed to. Fix: report at least two, state which one matches the harm, and explain the trade-off instead of burying it.

Reviewing accuracy only. A model can be accurate and still impose a burden that falls entirely on one group, because the historical data already reflects it. Fix: test selection rates and error rates separately, per group.

Ignoring administrative context. Scores are read inside a queue with staffing limits, caseload targets and local politics. A low-impact score with heavy human review is not the same intervention as a high-impact score in an understaffed office. Fix: interview the front line and trace real cases, not just outputs.

Taking the vendor’s assurance at face value. The agency depends on the vendor for logs, model documentation and data pipelines, so the vendor controls what evidence exists. Fix: request the vendor’s own fairness and validation reports through the contract, and record the refusal if the contract has no audit-rights clause — that gap is a procurement finding.

Publishing a finding with no owner or date. Nothing gets fixed. Fix: every recommendation gets an office and a date, and the follow-up audit date goes in the report too.

Accepting an audit that only restates the documentation. Audit-washing is easy to recognise once you look for the tells: no data was tested, no subgroup was examined, the auditor was paid by the vendor, no limitations are stated, the document has no method section, and it says the system is fair without ever reporting a rate. Fix: score the report against the five questions above — what was tested, with what data, by whom, funded how, and what remains unknown. A published audit that answers none of them is a press release.

Exposing personal data in the evidence. Reproducing case files to prove harm helps nobody and ends the audit. Fix: publish the pattern and the count, keep the individuals protected, and store the underlying records under a confidentiality agreement.

Starting the technical work before naming the owner. If no office is accountable, no finding will ever be actioned, however well evidenced. Fix: make “who owns this system” step one, and report the absence of an owner as a finding in its own right.

Frequently Asked Questions

Can the public audit an algorithm used by a government?

Yes, though rarely from a position of strength. Members of the public, journalists, researchers and civil-society organisations audit government algorithms mainly through freedom of information and state public records requests, which force disclosure of impact assessments, procurement files and contracts. Access to decision-level data is harder and often refused or redacted. Most successful public audits are partial by design: they test one decision, one group, one period, and say clearly what they could not reach.

How can auditors assess an algorithm without source code?

Test the outputs instead. Request decision-level records with protected characteristics attached, then compare selection rates and error rates across groups. Pair that with counterfactual and synthetic records that vary one attribute at a time to see how much a single change moves a score. Interview the people who use the tool and read the case files. You can establish pattern and scale of effect without code; you cannot establish mechanism, and the report should say so.

Which fairness tests should a government algorithm audit use?

At minimum: selection rates per group and the ratio between the highest and lowest, error rates split into false positives and false negatives per group, and a time series showing whether outcomes diverge after launch. Add a counterfactual probe and a drift check where the same system keeps running. Report two or more fairness measures rather than one, because the standard measures conflict by design and the right choice depends on which error hurts more in that specific decision.

What should an audit do when the algorithm comes from a vendor?

Test the contract before anything else. Look for audit-rights clauses, documentation obligations, data-access terms, and whether the agency can disclose your findings without vendor approval. Request the vendor’s own validation and fairness reports through the agency. Document who holds the logs and who controls the data pipeline, because that determines what evidence exists at all. If the contract lets the vendor veto publication, that is a procurement finding worth reporting on its own.

When should a government pause an algorithm during an audit?

Pause when an audit finds a high-severity harm that is still happening, when a system is making decisions with no identified legal authority, when people cannot learn it was involved and cannot appeal in time, or when the agency cannot say who owns it. A pending court case or an unresolved discrimination complaint also justifies slowing use. A pause should come with a written condition for restart: the specific finding fixed, re-tested, and signed off by a named owner.

Conclusion

Start by naming one decision and one system, then find out who owns it and what document authorises it. Everything after that is evidence gathering: records, decision data, interviews, tests. Publish what you could not test as prominently as what you found, attach an owner and a deadline to every recommendation, and set a date to re-audit.

Leave a Comment