From the Local Government edition of September 10, 2026
Procurement checklist research identifies expertise and review-coverage gaps
Tom Zick, Mason Kortz, David Eaves and Finale Doshi-Velez; Harvard University and University College London · Government AI procurement research · Canada, Brazil and Singapore discussions; qualified transfer to U.S. municipalities
- Publisher
- arXiv
- Original publication
- April 23, 2024, arXiv version 1
- Source retrieved
- 2026-09-11
What happened
The paper argues that checklists need expert interpretation and can miss low-cost, embedded and in-house AI.
Why it matters
New-to-archive historical scrutiny helps test whether municipal review processes reach actual systems. International examples are not U.S. legal requirements.
Evidence and measured results
The authors inspect two checklist frameworks, discuss practice with government officials, and describe a Harvard lab red-teaming exercise. Examples expose narrow supplier answers, purchasing thresholds and hidden components. The paper provides no representative interview sample size, causal effectiveness estimate or measured reduction in harm.
Limitations and uncertainty
Historical qualitative analysis; illustrative cases cannot estimate prevalence. Policy references describe 2024 circumstances and are not asserted as current law. Version 2 was inaccessible; version 1 was fully inspected.
Put this evidence to work
Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-11; this does not change the original publication date. Labels below come from the analysis itself.
Sales
Role takeaway
Ask a purchasing manager, CIO, risk lead and relevant service manager to walk through a recent approval. Can they identify who challenged supplier evidence and what happened when expertise was unavailable? A bounded engagement can examine a few completed assessments and map missing evidence to specific decisions. The value hypothesis is more defensible purchasing and clearer specialist demand, not certification or eliminated bias. The paper's international observations do not establish the customer's failure rate. Include low-cost tools and internal builds in discovery. For small governments, consider a shared expert-review arrangement only after confirming funding, access and practical turnaround requirements with the customer.
Pre-sales engineering
Role takeaway
Turn broad checklist questions into falsifiable tests for a selected service. Establish a system boundary, representative data, error categories and an explicit decision threshold before reviewing vendor submissions. Use adversarial examples and subgroup checks suited to the service; aggregate accuracy alone may conceal important mistakes. Validate the reviewer interface as well as model behavior, including the ability to reject an incorrect recommendation. Keep sensitive test data controlled and log evidence provenance. Because the paper provides no measured assurance effect, the proof of value should test whether reviewers detect planted evidence gaps and request appropriate follow-up, rather than assume that more checklist fields improve safety.
Delivery
Role takeaway
Appoint a service owner and a qualified review lead, then budget time for technical, domain and human-factors expertise. Build a repeatable evidence pack, escalation route and schedule for checking material changes. Dependencies include test access, supplier cooperation and staff who can act on findings. Train generalist buyers to identify uncertainty and seek help without implying they must become model auditors. Proposed acceptance criteria are detection of predefined evidence gaps, justified review decisions for all sampled acquisition routes, and a completed change-response exercise. These are proposed criteria. Monitor review backlog and unsupported approvals; an unfunded specialist requirement can otherwise become a procedural formality.
Implementation considerations
Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.
Architecture and integration
What must connect, and where does the AI sit in the workflow?
Include internally developed agents and newly embedded models in the same system map as purchased applications. No deployment topology was benchmarked.
Governance
Who approves, reviews and stays accountable for outcomes?
Test reviewer competence and decision authority, not just whether a form exists.
Security and privacy
What data, permissions and controls need testing?
Use qualified specialists to assess whether evidence about data handling and errors is sufficient for the local use case.
Accessibility and workforce
Who is affected, and what skills or accommodations follow?
Budget specialist support and resident input; generic staff training is not equivalent to technical assurance.
Procurement
What should contracts, pricing and exit terms secure?
Sample small purchases, feature upgrades and internal builds when testing review coverage.
Operating model
Which teams own the service once it runs?
Specify monitoring triggers, responsible reviewers and escalation when uncertainty remains.
What changed
New to the full 182-resource archive checked across offsets 0 and 100. No substantive source update or post-last-run event is claimed.
Publication history
- 2026-09-10Local Government · Issue 053 resources
Stable resource ID: ai-procurement-checklist-expertise-loopholes-2024