Lighthouse AdvisorySLED AI Adoption Intelligence
← Back to results

From the Local Government edition of September 10, 2026

Academic researchCautionaryNewly relevant · Apr 2024

Procurement checklist research identifies expertise and review-coverage gaps

Tom Zick, Mason Kortz, David Eaves and Finale Doshi-Velez; Harvard University and University College London · Government AI procurement research · Canada, Brazil and Singapore discussions; qualified transfer to U.S. municipalities

Publisher
arXiv
Original publication
April 23, 2024, arXiv version 1
Source retrieved
2026-09-11
Read original source

What happened

The paper argues that checklists need expert interpretation and can miss low-cost, embedded and in-house AI.

Why it matters

New-to-archive historical scrutiny helps test whether municipal review processes reach actual systems. International examples are not U.S. legal requirements.

Evidence and measured results

The authors inspect two checklist frameworks, discuss practice with government officials, and describe a Harvard lab red-teaming exercise. Examples expose narrow supplier answers, purchasing thresholds and hidden components. The paper provides no representative interview sample size, causal effectiveness estimate or measured reduction in harm.

Limitations and uncertainty

Historical qualitative analysis; illustrative cases cannot estimate prevalence. Policy references describe 2024 circumstances and are not asserted as current law. Version 2 was inaccessible; version 1 was fully inspected.

Put this evidence to work

Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-11; this does not change the original publication date. Labels below come from the analysis itself.

Sales

Role takeaway

Ask a purchasing manager, CIO, risk lead and relevant service manager to walk through a recent approval. Can they identify who challenged supplier evidence and what happened when expertise was unavailable? A bounded engagement can examine a few completed assessments and map missing evidence to specific decisions. The value hypothesis is more defensible purchasing and clearer specialist demand, not certification or eliminated bias. The paper's international observations do not establish the customer's failure rate. Include low-cost tools and internal builds in discovery. For small governments, consider a shared expert-review arrangement only after confirming funding, access and practical turnaround requirements with the customer.

Pre-sales engineering

Role takeaway

Turn broad checklist questions into falsifiable tests for a selected service. Establish a system boundary, representative data, error categories and an explicit decision threshold before reviewing vendor submissions. Use adversarial examples and subgroup checks suited to the service; aggregate accuracy alone may conceal important mistakes. Validate the reviewer interface as well as model behavior, including the ability to reject an incorrect recommendation. Keep sensitive test data controlled and log evidence provenance. Because the paper provides no measured assurance effect, the proof of value should test whether reviewers detect planted evidence gaps and request appropriate follow-up, rather than assume that more checklist fields improve safety.

Delivery

Role takeaway

Appoint a service owner and a qualified review lead, then budget time for technical, domain and human-factors expertise. Build a repeatable evidence pack, escalation route and schedule for checking material changes. Dependencies include test access, supplier cooperation and staff who can act on findings. Train generalist buyers to identify uncertainty and seek help without implying they must become model auditors. Proposed acceptance criteria are detection of predefined evidence gaps, justified review decisions for all sampled acquisition routes, and a completed change-response exercise. These are proposed criteria. Monitor review backlog and unsupported approvals; an unfunded specialist requirement can otherwise become a procedural formality.

Implementation considerations

Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.

Architecture and integration

What must connect, and where does the AI sit in the workflow?

Include internally developed agents and newly embedded models in the same system map as purchased applications. No deployment topology was benchmarked.

Governance

Who approves, reviews and stays accountable for outcomes?

Test reviewer competence and decision authority, not just whether a form exists.

Security and privacy

What data, permissions and controls need testing?

Use qualified specialists to assess whether evidence about data handling and errors is sufficient for the local use case.

Accessibility and workforce

Who is affected, and what skills or accommodations follow?

Budget specialist support and resident input; generic staff training is not equivalent to technical assurance.

Procurement

What should contracts, pricing and exit terms secure?

Sample small purchases, feature upgrades and internal builds when testing review coverage.

Operating model

Which teams own the service once it runs?

Specify monitoring triggers, responsible reviewers and escalation when uncertainty remains.

What changed

New to the full 182-resource archive checked across offsets 0 and 100. No substantive source update or post-last-run event is claimed.

Publication history

  1. 2026-09-10Local Government · Issue 053 resources
Read preserved resource versions (JSON)

Stable resource ID: ai-procurement-checklist-expertise-loopholes-2024