{"resourceId":"ai-procurement-checklist-expertise-loopholes-2024","versions":[{"version":"external-01c3792934278c4d7fa79e0b012a71b1fbd88c3bb1e951c134bb252e242241d2","resource":{"id":"ai-procurement-checklist-expertise-loopholes-2024","title":"Procurement checklist research identifies expertise and review-coverage gaps","organization":"Tom Zick, Mason Kortz, David Eaves and Finale Doshi-Velez; Harvard University and University College London","sector":"Government AI procurement research","geography":"Canada, Brazil and Singapore discussions; qualified transfer to U.S. municipalities","publishedAt":"April 23, 2024, arXiv version 1","publicationDate":"2024-04-23","eventDate":null,"sourceName":"arXiv","sourceLabel":"Academic qualitative analysis and red-teaming exercise; historical preprint","sourceUrl":"https://arxiv.org/html/2404.14660v1","evidenceClass":"academic-research","outcomeClass":"cautionary","topics":["developers-agents","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"The paper argues that checklists need expert interpretation and can miss low-cost, embedded and in-house AI.","sledRelevance":"New-to-archive historical scrutiny helps test whether municipal review processes reach actual systems. International examples are not U.S. legal requirements.","evidence":"The authors inspect two checklist frameworks, discuss practice with government officials, and describe a Harvard lab red-teaming exercise. Examples expose narrow supplier answers, purchasing thresholds and hidden components. The paper provides no representative interview sample size, causal effectiveness estimate or measured reduction in harm.","architectureImplications":"Interpretation: Include internally developed agents and newly embedded models in the same system map as purchased applications. No deployment topology was benchmarked.","governanceImplications":"Interpretation: Test reviewer competence and decision authority, not just whether a form exists.","securityPrivacyImplications":"Interpretation: Use qualified specialists to assess whether evidence about data handling and errors is sufficient for the local use case.","caveats":"Historical qualitative analysis; illustrative cases cannot estimate prevalence. Policy references describe 2024 circumstances and are not asserted as current law. Version 2 was inaccessible; version 1 was fully inspected.","streamIds":["local-government"],"roles":{"sales":"Interpretation: Ask a purchasing manager, CIO, risk lead and relevant service manager to walk through a recent approval. Can they identify who challenged supplier evidence and what happened when expertise was unavailable? A bounded engagement can examine a few completed assessments and map missing evidence to specific decisions. The value hypothesis is more defensible purchasing and clearer specialist demand, not certification or eliminated bias. The paper's international observations do not establish the customer's failure rate. Include low-cost tools and internal builds in discovery. For small governments, consider a shared expert-review arrangement only after confirming funding, access and practical turnaround requirements with the customer.","engineering":"Interpretation: Turn broad checklist questions into falsifiable tests for a selected service. Establish a system boundary, representative data, error categories and an explicit decision threshold before reviewing vendor submissions. Use adversarial examples and subgroup checks suited to the service; aggregate accuracy alone may conceal important mistakes. Validate the reviewer interface as well as model behavior, including the ability to reject an incorrect recommendation. Keep sensitive test data controlled and log evidence provenance. Because the paper provides no measured assurance effect, the proof of value should test whether reviewers detect planted evidence gaps and request appropriate follow-up, rather than assume that more checklist fields improve safety.","delivery":"Interpretation: Appoint a service owner and a qualified review lead, then budget time for technical, domain and human-factors expertise. Build a repeatable evidence pack, escalation route and schedule for checking material changes. Dependencies include test access, supplier cooperation and staff who can act on findings. Train generalist buyers to identify uncertainty and seek help without implying they must become model auditors. Proposed acceptance criteria are detection of predefined evidence gaps, justified review decisions for all sampled acquisition routes, and a completed change-response exercise. These are proposed criteria. Monitor review backlog and unsupported approvals; an unfunded specialist requirement can otherwise become a procedural formality."},"retrievedAt":"2026-09-11T03:02:04Z","enrichedAt":"2026-09-11T03:04:59Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Budget specialist support and resident input; generic staff training is not equivalent to technical assurance.","procurementImplications":"Interpretation: Sample small purchases, feature upgrades and internal builds when testing review coverage.","operatingModelImplications":"Interpretation: Specify monitoring triggers, responsible reviewers and escalation when uncertainty remains.","updateExplanation":"New to the full 182-resource archive checked across offsets 0 and 100. No substantive source update or post-last-run event is claimed.","sourceVerification":{"openedUrl":"https://arxiv.org/html/2404.14660v1","referenceExcerpt":"However, neither prescribes specific processes, metrics, or schedules for such monitoring.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}