Public Sector & Government · Issue 06 ·
State Government
Two newly archived historical sources examine Missouri transport ML's operational and economic limits and Australian AI disclosure quality. Technical scores and published assurances do not establish service fitness. Role guidance and one cross-source interpretation propose task-specific acceptance evidence. No new September 11 measured service improvement was established; autonomous-agent, benefits-access and hosting-comparison evidence remain limited.
- Evidence records
- 2
- Cross-source patterns
- 1
- Evidence classes
- 1 government evaluation1 academic research
- Outcomes
- 1 mixed1 cautionary
- Source freshness
- 2 older, newly relevant
- Research completed
- 2026-09-12
Choose a role to see its takeaway beside every record in the ledger.
Synthesis · Lighthouse Advisory interpretation
Patterns across the evidence
Connect assurance evidence to the operational decision
Missouri's evaluation distinguishes technical performance from useful outputs; the Australian study distinguishes organizational disclosure from system assurance. For a state program, specify which decision each metric or statement can support and what operational evidence remains necessary. Neither source establishes failure in another state's deployment.
Operating questionWhat evidence shows that this particular workflow meets its service requirement, beyond model metrics or general assurance statements?
Supporting evidenceMissouri Department of Transportation; High Street Consulting GroupShidong Pan and coauthors
Full record · every source keeps its link and limitations
Evidence ledger
Missouri pilots show why technical accuracy can miss the operational target
The median-inventory pilot missed operational precision and scope needs despite strong image metrics. Traffic factor grouping offered a more promising, partly prospective economic case.
Why it matters, evidence and limitations
- Why it matters
- Historical state-agency evidence newly inspected to qualify current AI scaling decisions; findings concern conventional ML, not generative copilots or autonomous agents.
- Evidence and measured results
- SparseInst used 1,100 annotated images and processed 10,000 tiles; reported mean pixel accuracy was 99.4% and correct-area prediction 93%. The $80,000 median pilot compared with an estimated $61,110 manual year, plus estimated six-to-nine-month correction work. AADT clustering/classification used 170 continuous stations and 19,828 short-term sites. Its $45,000 toolbox still needed estimated $10,000–$30,000 integration; projected savings compared with a richer performance-based manual process, not the minimal existing approach.
- Limitations and uncertainty
- June 2022–June 2024 project. Contractor findings, not a standard. Imagery covered southern Missouri; engineering precision was inadequate. Costs and counterfactual labor are estimates, not causal savings. Held-out performance details remain unclear.
Australian disclosure study separates AI transparency from operational assurance
Public disclosures emphasize organizational assurance but often leave operational review mechanisms and shared-service responsibilities difficult to inspect.
Why it matters, evidence and limitations
- Why it matters
- Newly inspected historical research complements state governance coverage. Australian disclosure rules and institutional boundaries do not establish U.S. state requirements or failure rates.
- Evidence and measured results
- The November 2025 snapshot found 101 statements and 72 entities without one after a restructuring exclusion. Methods combine document coding, readability analysis and qualitative interpretation. Median Flesch–Kincaid grade level was 14.16; the separate GPT-5 lexical analysis used ten human spot checks with kappa 0.70. Disclosure-category presence is not a measure of risk mitigation.
- Limitations and uncertainty
- Preprint and historical document snapshot, not an audit of running systems or a test of reader comprehension. Lexical model errors, binary scoring and subjective annotation limit conclusions; missing public detail does not prove missing internal controls.
How to read this edition
Source findings, measured results and limitations come from the cited publications. Patterns, operating questions, role takeaways and implementation considerations are Lighthouse Advisory interpretation, stated as questions to validate locally rather than guaranteed outcomes. Vendor and operator claims are labeled as claims. Full research method.
- Government evaluation
- A public body’s measured evaluation or documented pilot.
- Academic research
- Research produced through an academic institution or peer-reviewed venue.