{"resourceId":"modot-ai-ml-pilot-operational-fit-2024","versions":[{"version":"external-1086b215431c392a76d78a6f1ae1a1b7a44bdc5ce1ae4c340053a1008499fca9","resource":{"id":"modot-ai-ml-pilot-operational-fit-2024","title":"Missouri pilots show why technical accuracy can miss the operational target","organization":"Missouri Department of Transportation; High Street Consulting Group","sector":"State transportation","geography":"Missouri, United States","publishedAt":"July 2024; repository indexes July 1","publicationDate":"2024-07-01","eventDate":null,"sourceName":"MoDOT research report cmr 24-009","sourceLabel":"Government-sponsored contractor evaluation","sourceUrl":"https://rosap.ntl.bts.gov/view/dot/76741/dot_76741_DS1.pdf","evidenceClass":"government-evaluation","outcomeClass":"mixed","topics":["infrastructure","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"The median-inventory pilot missed operational precision and scope needs despite strong image metrics. Traffic factor grouping offered a more promising, partly prospective economic case.","sledRelevance":"Historical state-agency evidence newly inspected to qualify current AI scaling decisions; findings concern conventional ML, not generative copilots or autonomous agents.","evidence":"SparseInst used 1,100 annotated images and processed 10,000 tiles; reported mean pixel accuracy was 99.4% and correct-area prediction 93%. The $80,000 median pilot compared with an estimated $61,110 manual year, plus estimated six-to-nine-month correction work. AADT clustering/classification used 170 continuous stations and 19,828 short-term sites. Its $45,000 toolbox still needed estimated $10,000–$30,000 integration; projected savings compared with a richer performance-based manual process, not the minimal existing approach.","architectureImplications":"Interpretation: set spatial tolerances and downstream system contracts before choosing models; evaluate batch processing and GIS integration locally. Cloud, on-premises and hybrid cost comparisons remain unestablished.","governanceImplications":"Interpretation: require service-owner approval when scope narrows; retain the original business question in the acceptance test.","securityPrivacyImplications":"Interpretation: authorize imagery and warehouse access, scan dependencies, control exports and document retention. No security certification follows from this evaluation.","caveats":"June 2022–June 2024 project. Contractor findings, not a standard. Imagery covered southern Missouri; engineering precision was inadequate. Costs and counterfactual labor are estimates, not causal savings. Held-out performance details remain unclear.","streamIds":["state-government"],"roles":{"sales":"Interpretation: For transport operations, GIS, finance and IT leaders, qualify whether the desired result is approximate planning information or a maintained engineering asset. Ask how often the decision repeats, what omissions cost, and who funds correction work. A bounded engagement could compare one inventory segment with a manual reference and a commercial service option. The value hypothesis is lower total cost per accepted asset, including labeling, review and integration. The mixed pilot supports careful qualification; it does not establish universal ML economics or current pricing. Keep recurrent traffic analysis separate from a one-time inventory when discussing potential value.","engineering":"Interpretation: Fit depends on the downstream tolerance, not an aggregate image score. Build a geographically separated test set, include narrow medians and complex interchanges, and assess missing features alongside false detections. Prerequisites are licensed imagery, trusted labels, versioned GIS schemas and a named subject expert. Select hosting only after profiling transfer, compute and storage needs. Validate dependency security and service-account scopes before warehouse integration. A proposed proof of value should measure accepted geometry, correction effort and complete task cost against manual work. Reproduce the factor-grouping baseline separately; traffic-planning judgments need their own validation and cannot inherit computer-vision results.","delivery":"Interpretation: The asset owner should define acceptance and approve scope changes; GIS staff own labels and quality checks, while IT owns deployment and support. Start with a small mapped area and preserve manual fallback. Dependencies include imagery availability, annotation skills and capacity to investigate omissions. Train reviewers on difficult examples and record correction effort before expansion. Proposed acceptance criteria are an agreed spatial tolerance, a documented completeness threshold and a lower full cost per accepted asset than the local baseline. Review at pilot exit and after source-data changes. Risks include accepting attractive aggregate scores while losing the operational question or underfunding maintenance."},"retrievedAt":"2026-09-12T03:01:12Z","enrichedAt":"2026-09-12T03:02:59Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: test review tools with keyboard and assistive technology; budget specialist annotation and inspection time. No accessibility outcome was measured.","procurementImplications":"Interpretation: contract for accepted operational outputs, correction duties, source/code portability and support costs; avoid pricing solely per inference.","operatingModelImplications":"Interpretation: give the business owner authority to stop a technically successful pilot that fails the agreed use.","updateExplanation":"No matching source or Missouri finding in 217 archive records retrieved at offsets 0, 100 and 200, or targeted Missouri search. Historical counterevidence extends the archive beyond recent generative-AI rollout claims; no new September result asserted.","sourceVerification":{"openedUrl":"https://rosap.ntl.bts.gov/view/dot/76741/dot_76741_DS1.pdf","referenceExcerpt":"The outputs from this computer vision process both exceeded expectations and failed to meet expectations.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}