1 engagement slot currently open
← The Agent Ledger
government benefits eligibility and case processing2026-07-31

780,000 Cases, 42,500 Hours: What the DWP's AI Eligibility Pilot Actually Measured

780,000 cases processed, 42,500 operational hours saved
Will Drewes
Will Drewes
Founder, Fern Strategy · 3 min read

One agency published the numbers. Most haven't.

The UK Department for Work and Pensions ran an AI-assisted eligibility workflow on Employment Support Allowance claims and disclosed the results publicly. That alone makes it worth examining. Most agencies piloting similar tools have not published anything this specific.

The headline figures: over 780,000 cases processed, with 42,500 operational hours saved. The system cross-references applicant health conditions against claim criteria and surfaces a recommendation. A human caseworker makes the final decision.

That human-in-the-loop structure matters, and the accuracy history explains why it's still there.

The accuracy gap is the real story

The DWP disclosed that the current system runs at 87% prediction accuracy. That sounds reasonable until you see what came before it. An earlier version, used from 2020 to 2024, hit only 35% accuracy. Caseworkers had to manually intervene on the remaining 65% of cases.

Think about what that earlier version actually cost in practice. If you're deploying a tool that's wrong nearly two-thirds of the time, you haven't reduced caseworker burden. You've added a review step on top of the original workload. The 87% figure represents a real operational improvement, but it also tells you the first four years were spent learning what didn't work.

The DWP's current design acknowledges this honestly: the system assists, it doesn't decide. That's not a hedge. Given the track record, it's the correct architecture.

What the research doesn't show yet

Beyond the DWP data, the public record on AI in government benefits is thin on hard numbers. Code for America has piloted tools to help workers answer policy questions and review eligibility documents for SNAP, but no outcome metrics have been published yet. Separate research testing LLMs against state-level SNAP and Medicaid rules found accuracy rates of 0% on state-specific numerical thresholds in some cases. Those tools were not deployed, but the finding is a useful reminder: a model that handles federal policy text can still fail completely on a state's specific dollar limits or household size tables.

The gap between "works in a demo" and "works on your actual ruleset" is where most of these projects stall.

What this means for benefits administrators and their oversight teams

The DWP case gives you a concrete benchmark and a realistic timeline. Four years to get from 35% to 87% accuracy is not a failure story, but it is an honest one. Any agency evaluating AI for eligibility processing should ask three questions before committing budget:

  1. What is the baseline accuracy on your specific rules, not a generic benchmark? State-level variation in benefits policy is significant. A system trained on federal guidelines may not handle your edge cases.
  2. Where does the human decision boundary sit, and is it auditable? The DWP's design keeps the final call with a caseworker. The audit trail for why a recommendation was made needs to be readable by a reviewer, not just a data scientist.
  3. What does the error mode look like? At 65% error rate, the DWP's earlier system created more work than it saved. Knowing how a system fails before you deploy it at scale is the difference between a pilot and a problem.

The 42,500 hours saved is a real number. So is the four-year accuracy climb that preceded it. Both belong in your business case.

Sources
  1. http://digitalgovernmenthub.org/publications/ai-powered-rules-as-code-experiments-with-public-benefits-policy-summary/
  2. https://www.ijsart.com/public/storage/paper/pdf_2/IJSARTV12I2104565.pdf
  3. https://www.yahoo.com/news/dwp-reveals-using-ai-benefit-113008027.html
  4. https://www.oecd.org/content/dam/oecd/en/publications/reports/2024/06/using-ai-to-manage-minimum-income-benefits-and-unemployment-assistance_b11ccde4/718c93a1-en.pdf
  5. https://www.oecd.org/en/publications/using-ai-to-manage-minimum-income-benefits-and-unemployment-assistance_718c93a1-en.html
  6. https://www.hhs.gov/sites/default/files/public-benefits-and-ai.pdf
  7. https://www.microsoft.com/en-us/microsoft-cloud/blog/public-health-social-services/2026/03/04/right-benefit-right-person-right-time-how-ai-is-reshaping-administration-of-benefits-programs-worldwide/
  8. https://www.microsoft.com/en-us/ai/use-case/improve-benefits-eligibility-payments-and-fraud-with-ai
  9. https://www.navapbc.com/case-studies/ai-tools-public-benefits
  10. https://servos.io/blog/aieligibility
  11. https://www.axios.com/2026/06/21/ai-snap-medicaid-unemployment-benefits
  12. https://www.govexec.com/technology/2026/05/anthropic-code-america-pilot-ai-tools-snap/413464/
  13. https://www.theguardian.com/society/2024/dec/06/revealed-bias-found-in-ai-system-used-to-detect-uk-benefits
  14. https://www.ncsl.org/technology-and-communication/artificial-intelligence-in-government-the-federal-and-state-landscape
  15. https://www.youtube.com/watch?v=WM0nsXoyaJI
  16. https://inthesetimes.com/article/ai-artificial-intelligence-food-stamps-snap-public-benefits
  17. https://datagrid.com/blog/ai-agents-public-benefits-application-review
  18. https://pmc.ncbi.nlm.nih.gov/articles/PMC12663265/
  19. https://bigbrotherwatch.org.uk/press-coverage/big-issue-unreliable-ai-usage-by-dwp-risks-vulnerable-people-being-treated-as-guinea-pigs/
  20. https://www.online.uc.edu/blog/artificial-intelligence-ai-benefits.html
Want an outcome like this in your workflow?

Fern's two-week Audit maps where a governed AI agent would pay off in your operation — and what has to be true to build it.

Book a scoping call →
Or just get new entries as they land. Real numbers only.