SkycrumbsSkycrumbs
AI Policy

How Governments Are Using AI — and Where It Goes Wrong

September 19, 2026·6 min read
How Governments Are Using AI — and Where It Goes Wrong

How Governments Are Using AI — and Where It Goes Wrong

Government agencies at every level are deploying AI — in hiring processes, welfare systems, law enforcement, tax administration, and public health. The scale and speed of this adoption often isn't visible to the public until something goes wrong. And things do go wrong, in ways that cause real harm to real people.

That doesn't mean AI has no place in government. There are genuine efficiency gains and service improvements that AI enables. But the public sector context introduces failure modes and accountability problems that don't appear in consumer applications. Understanding both sides is essential for anyone trying to make sense of where this is heading.

What Governments Are Actually Using AI For

The applications span a wide range:

Benefits administration. Automated systems screen applications for food assistance, unemployment insurance, disability benefits, and housing programs. Some systems flag applications for human review; others make initial eligibility determinations that are difficult for applicants to challenge.

Tax fraud detection. Tax authorities use pattern recognition to identify returns that warrant audit. This is one of the more uncontroversial applications — the core task is well-suited to pattern matching, the cost of false negatives is real, and false positives trigger human review rather than automatic penalty.

Public safety and law enforcement. Predictive policing tools estimate where crimes are likely to occur. Facial recognition is deployed at borders and in some urban surveillance systems. Risk assessment tools inform bail and sentencing decisions in some jurisdictions.

Healthcare and public health. Disease surveillance, hospital resource allocation, patient risk stratification, and emergency response planning all have active AI applications.

Infrastructure management. Traffic optimization, utility grid management, and public transit scheduling have been partly automated in many cities.

Permitting and licensing. Some jurisdictions have automated the routing and initial review of permit applications, reducing processing time and, in theory, inconsistency.

Where It Works Reasonably Well

Applications that share a few characteristics tend to perform better:

  • The task is well-defined with a clear success criterion
  • Errors are caught by a human review step before they cause harm
  • The system augments human judgment rather than replacing it
  • Performance can be monitored against ground truth
  • Affected people have accessible recourse when something goes wrong

Tax fraud detection fits this model reasonably well. So does optimizing traffic signal timing, where performance is measurable and errors have limited downside. Infrastructure monitoring with human oversight has delivered real operational improvements in multiple cities.

Where It Goes Wrong

High-stakes automated decisions without adequate oversight. Several countries and U.S. states have deployed welfare benefit systems that automatically denied or reduced benefits based on algorithmic determinations — sometimes due to data errors that applicants had no meaningful way to identify or challenge. Cases in the Netherlands and Australia involved large-scale wrongful debt collection pursued aggressively before courts intervened. Similar patterns have appeared in child protective services screening and disability assessment systems.

Predictive systems trained on historically biased data. When AI systems predict future behavior using data from past enforcement or decisions, they can encode historical disparities rather than reduce them. A predictive policing tool trained on past arrest data in neighborhoods where enforcement was already disproportionate will direct more resources to those neighborhoods, producing more arrests, which validates the initial prediction. The feedback loop is hard to break.

Facial recognition accuracy gaps. Error rates for facial recognition systems vary substantially by demographic group. In most evaluated systems, error rates are higher for women and for individuals with darker skin tones. When these systems are used in high-stakes contexts — border control, suspect identification — the error distribution has real consequences for the people misidentified.

Opacity that prevents accountability. Public sector AI systems are sometimes subject to procurement processes that result in black-box tools with confidential methodologies. People affected by the decisions can't understand why they were decided the way they were. Advocates can't assess whether the system is working fairly. Oversight bodies can't evaluate what they can't inspect.

Scope creep. Systems deployed for one purpose sometimes get extended to others without fresh evaluation. A document processing tool becomes a decision-making tool. A fraud detection system starts informing immigration enforcement. Each extension changes the risk profile in ways the original deployment analysis didn't consider.

What Accountability Mechanisms Exist

In the EU, the AI Act establishes requirements for high-risk AI systems — including many public sector applications — covering conformity assessments, transparency obligations, and human oversight requirements. Enforcement is still being developed, but the legal framework is more substantial than exists in most other jurisdictions.

In the United States, there's no comprehensive federal framework for government AI use. The executive branch has issued guidance through the Office of Management and Budget, but the requirements are limited and enforcement is weak. Individual agencies have adopted varying internal policies. Some states have passed narrower legislation addressing specific applications like facial recognition or automated decision-making in benefits.

Several cities have established algorithmic accountability offices or advisory boards, though their capacity varies widely. Some jurisdictions require algorithmic impact assessments before deploying high-stakes automated decision systems. These vary enormously in scope and rigor.

What Good Government AI Looks Like

The cases where AI deployments in government have worked well tend to share characteristics:

  • Human-in-the-loop by default for high-stakes decisions. The AI provides recommendations; humans with accountability make the decisions.
  • Clear recourse mechanisms. Affected people can find out how a decision was made and have a genuine pathway to contest it.
  • Ongoing performance auditing. Someone is responsible for monitoring accuracy, disparate impact, and error rates on a continuing basis — not just at deployment.
  • Procurement that favors transparency. Contracts that require vendors to allow independent audits, disclose training data characteristics, and produce explainable outputs.
  • Public disclosure. Registries of AI systems in use, with basic information about their purpose, scope, and performance, enable civil society oversight that would otherwise be impossible.

The Accountability Gap

The deepest problem with government AI isn't any particular technical failure — it's an accountability structure that doesn't match the risk. When a private company's algorithm performs poorly, competitive pressure and liability create incentives to fix it. When a government agency's algorithm performs poorly, the affected populations — often people with limited political and legal resources — face systems that are difficult to challenge and in which errors often don't generate immediate institutional consequences.

Fixing this requires legal frameworks that create real obligations, procurement practices that favor transparency, and oversight institutions with actual capacity to evaluate and enforce. Those are political and institutional questions as much as technical ones.


AI in government is neither a disaster to be stopped nor a straightforward improvement to be celebrated. It's a set of tools being applied in a context where the stakes are high, accountability is often weak, and the work of getting it right requires both technical and institutional investment.

Comments

Loading comments...

Leave a comment