Validation Tax: What Agent ROI Actually Measures

7 min read
  • AI Agents
  • Workforce
  • Cost & Economics
  • Deep dive
Contents7 sections

The Six-Minute Draft That Took Two Hours

An agent drafts a determination letter in six minutes. A caseworker then spends most of an afternoon finding the governing policy version, checking three cited figures against the record system, and rewriting a paragraph that misread an exception. The slide presented to leadership records the six minutes.

That afternoon shows up in survey data. In the 2026 Hallucination Tax Report, commissioned by Collibra, a vendor of data and AI governance software, and fielded by The Harris Poll, "just over half of respondents (51%) report spending significant staff hours manually reviewing and correcting autonomous AI agent outputs before they go live". Collibra sells tooling for this problem, so we treat the figure as a vendor assertion until independent labor data supports it, which it does below.

EY, a professional services firm selling governance and risk advisory, reported in the same window that "Roughly half (49%) of respondents whose organization uses agentic AI say their organization's existing governance framework has not yet been updated to specifically include agentic AI requirements and risks," with 26 percent unable to detect unauthorized agents running internally. Two separately commissioned samples describe the same displacement: effort moving from producing output to confirming it.

Agents remove genuine work from genuine workflows, and the gain is real. Our August 26 piece, Agent Pilots Fail Because of Delivery Discipline, Not Technology, stopped at the pilot boundary. This one picks up after the pilot succeeds.

Where the Saved Hours Reappear

The measurement mistake we see most often in agent business cases is wall clock to completion. A four-hour task finishes in six minutes and the reduction gets logged at 97 percent. Output is finished when a person with authority is willing to release it. Measure time on task across the full workflow, and hours removed from producing the output turn up again as hours spent confirming it is correct.

Robert Half, a staffing firm with a commercial stake in hiring demand, has reported figures consistent with that second block, including that a large share of employers required more human oversight and quality control than expected when deploying AI, and that reviewing and refining AI-generated deliverables now consumes a meaningful share of affected employees' time. Academic work on LLM-assisted knowledge work names persistent verification burden as a recurring gap, noting that fabricated output "requires rigorous post-hoc human validation for scientific accuracy". An offset is not a cancellation. Much of the gain survives; the rest is spent where the pilot never instrumented, by people the business case never counted.

Design the Tool to Make Checking Cheap

A computer screen displaying a document with highlighted text passages and reference links to source material.

Validation cost is not fixed by the task. It is largely set by how the tool is built, and engineering effort spent there pays back faster than anything else once a pilot graduates.

The highest-return feature we build into agent workflows is inline provenance: every material claim carries a link to the record, policy section, or page it came from, so the reviewer confirms by clicking rather than searching. Next is confidence-tiered routing, where the agent scores its own certainty and the pipeline auto-releases routine output, queues borderline items for a quick look, and escalates high-consequence items with the conflict already surfaced. Third is structured output with machine-checkable fields, so arithmetic, date logic, eligibility thresholds, and formatting are validated by rules before a human sees anything. A diff view against the governing source shows where the agent's language departs from the authority it cites, and feedback capture at the point of correction records what was wrong and why, building the evaluation set that shows whether the agent is improving. Without that loop, an organization reviews the same class of error indefinitely.

An agent producing output no one can cheaply check has not automated the work, it has relocated it. Treat verification affordability as a design requirement alongside accuracy and latency, and the validation term shrinks instead of compounding with volume.

When the Math Still Works

Net efficiency equals elimination minus validation load, and a defensible business case shows both terms. In some workloads the second stays small on its own: high volume, low consequence per unit, output checkable against a clear source of truth. Classify these documents, extract these fields, draft these routine acknowledgments. Production cost falls toward zero while validation stays at the level of a glance.

Validation dominates where verification demands judgment rather than a lookup. A benefits eligibility determination, a procurement responsiveness finding, a clinical summary: checking any of these requires knowing which regulation, policy version, or record governs, and that knowledge lives in people. Gartner, an analyst firm with governance advisory offerings, has stated in a press release that it expects a substantial share of enterprises to demote or decommission autonomous agents by 2027 after governance gaps surface in production incidents, and argues that uniform governance regardless of autonomy level fails in both directions, burying low-risk agents in review they do not need while letting high-risk agents skip review they do.

Checking Is Its Own Skill

Two colleagues at a desk reviewing a printed report, one pointing at a paragraph on the page.

Reviewing agent output well draws on a different competency than producing the first draft. Validation asks for subject matter judgment, a working model of where this system tends to fail, and knowledge of which source governs a contested claim. Someone who used to do the work can usually learn to check it, though the competency should never be presumed present simply because the task is familiar. A reviewer who cannot catch the errors the agent is likely to make supplies a signature rather than a control. Peer-reviewed research on human-in-the-loop systems covers reviewer cognitive burden, time pressure on verification, and skill degradation from over-reliance on automated output. BCG, a consulting firm that sells transformation work, models that 50 to 55 percent of US jobs will be reshaped by AI over the next two to three years, drawing on labor data from Revelio Labs, with redesigned roles demanding greater expertise and accountability, and calls for "a scaled, strategic approach to upskilling and reskilling and the restructuring of career ladders".

For high-risk systems in the EU, competent oversight is statutory. Article 14 of the AI Act requires that high-risk systems "shall be designed and developed in such a way, including with appropriate human-machine interface tools, that they can be effectively overseen by natural persons during the period in which they are in use," and remote biometric identification decisions must be verified by at least two competent individuals. The regulation asks something of the builder as well as the employer: the interface has to make oversight possible. As we wrote in The Companies Selling AI Automation Just Hired 6,000 Human Engineers, four major AI providers committed more than $9 billion to people-heavy embedded engineering organizations.

Planning the Supervisor Role Before Deployment

Validation staffing and training belong in the business case as a priced line item, settled before deployment. In our AI readiness work, three questions predict whether an agent survives production. Who holds authority to declare which source governs a disputed answer. How reviewers are trained on that judgment before rollout instead of after an incident. How large the reviewer bench needs to be given expected volume and consequence, sized against the work rather than whatever headcount survived the elimination.

Pair those answers with a build that makes checking fast, and one reviewer covers several times the volume at the same standard of care. The organizations getting real returns design the supervision and the software together, rather than shipping the agent and discovering the supervision. An efficiency number showing only the elimination term is inflated, and the cost it conceals stays concealed until something ships that should not have. The AI supervisor is a role with standing, staffed and credentialed deliberately, and supported by tooling built to make its judgment possible to exercise.

Sources

  1. Collibra and The Harris Poll. 2026 Hallucination Tax Report
  2. EY. AI Risk and Governance Survey, September 15, 2026
  3. Spruce. Agent Pilots Fail Because of Delivery Discipline, Not Technology, August 26, 2026
  4. Robert Half. May 2026 Labor Market Update: For Employers and Job Seekers
  5. From Verification Burden to Trusted Collaboration: Design Goals for LLM-Assisted Literature Reviews. arXiv, December 12, 2025
  6. Gartner. Applying Uniform Governance Across AI Agents Will Lead to Enterprise AI Agent Failure, May 26, 2026
  7. Human-in-the-Loop Artificial Intelligence: A Systematic Review of Concepts, Methods, and Applications. MDPI Entropy, 2026
  8. Boston Consulting Group. AI Will Reshape More Jobs Than It Replaces, 2026
  9. EU AI Act Service Desk. Article 14: Human Oversight
  10. Spruce. The Companies Selling AI Automation Just Hired 6,000 Human Engineers, July 17, 2026

Want our take on your AI roadmap?

We help leaders turn strategy into production AI systems. Let's talk about what you're building.