Public Sector & Government · Issue 02 ·
State Government
Three unarchived sources distinguish workflow evidence, staff perceptions and agent containment. Colorado's historical benefits experiment found no significant average assistance gain; NCDOT's pilot survey documents expectation changes without measured productivity outcomes. A September 7 UK statement adds current control-response lessons with jurisdictional and testing-context limits. Two cross-source interpretations support bounded validation. Current audit access and live-service outcome evidence remain gaps.
- Evidence records
- 3
- Cross-source patterns
- 2
- Evidence classes
- 2 academic research1 standards or public-body guidance
- Outcomes
- 2 mixed1 cautionary
- Source freshness
- 1 undated1 recent1 new this fortnight
- Research completed
- 2026-09-08
Choose a role to see its takeaway beside every record in the ledger.
Synthesis · Lighthouse Advisory interpretation
Patterns across the evidence
User sentiment needs a separate workflow outcome test
Colorado's concurrent comparison and NCDOT's changing perception scores show why approval should distinguish user attitudes from observed service performance. Neither source establishes statewide financial returns.
Operating questionWhich quality, effort and service measures must improve alongside user feedback before expanding the workflow?
Supporting evidenceMagesh, Martin, Surani, Perez, Rodolfa and Ho; collaboration with CDLE and U.S. DOLUniversity of North Carolina at Charlotte and North Carolina Department of Transportation
A sandbox must support both containment and effectiveness testing
Colorado illustrates the limits of sandbox effectiveness evidence; the UK statement illustrates the need to verify containment itself. A state evaluation plan should answer both questions before any live expansion, without treating one test as proof of the other.
Operating questionWhat evidence demonstrates both an enforced action boundary and incremental value in the intended human workflow?
Supporting evidenceMagesh, Martin, Surani, Perez, Rodolfa and Ho; collaboration with CDLE and U.S. DOLUK Cabinet Office; statement by Kanishka Narayan
Full record · every source keeps its link and limitations
Evidence ledger
Colorado benefits trial finds favorable user feedback without an average causal gain
AI-assisted fact-finding did not significantly improve average drafting time or quality against concurrent unaided adjudicators.
Why it matters, evidence and limitations
- Why it matters
- Direct state benefits evidence; historical sandbox tasks do not establish live eligibility or payment outcomes.
- Evidence and measured results
- Randomized crossover study: 8 non-randomly recruited adjudicators, 200 sampled historical claims, 788 drafts, and 6 internal QA reviewers. Average drafting-time reduction was 4 seconds (p=0.8); quality comparison p=0.37. Positive user feedback and favorable historical comparisons did not establish incremental workflow benefit.
- Limitations and uncertainty
- Small selected workforce; historical quit cases; possible sandbox behavior effects. No demonstrated reduction in claimant waiting time. Manuscript publication day and trial dates remain unverified. PDF figure screenshot failed; Figure 4 caption and results text were inspected.
NCDOT study finds staff expectations change after hands-on Copilot use
Perceived usefulness fell after the pilot; other aggregate acceptance constructs did not change significantly.
Why it matters, evidence and limitations
- Why it matters
- Direct state workforce evidence, relevant to shared productivity-suite rollout rather than traffic-system performance.
- Evidence and measured results
- Eight-week 2025 NCDOT pilot; 175 baseline respondents, 133 matched responses and 124 after screening. Table 3 reports usefulness 3.85 before versus 3.62 after, adjusted p<0.001. Paired survey analysis and exploratory clustering measure perceptions, not causal productivity effects.
- Limitations and uncertainty
- Single agency, short horizon, self-reports and no untreated comparison. Small persona transitions and keyword-based qualitative coding limit inference. Attrition checks cover measured baseline attitudes only. Preprint status; no verified net savings or accessibility effect.
UK September 7 statement reports tighter controls after agent-testing incidents
The minister reports that AISI is strengthening internet restrictions, monitoring and sandboxing after its own testing incident.
Why it matters, evidence and limitations
- Why it matters
- Relevant to state developer-agent and evaluation environments; UK institutional arrangements and policy do not govern U.S. states.
- Evidence and measured results
- Official policy response, not an independent incident reconstruction. The statement places the incidents in frontier-model testing or development, sometimes with deliberately reduced safeguards. It offers no measured prevention rate or state deployment sample.
- Limitations and uncertainty
- Ministerial attribution; underlying investigations were not independently re-inspected here. The claim that best-practice controls would almost certainly have prevented incidents is the minister's judgment, not a validated counterfactual. Statement date is not incident date.
How to read this edition
Source findings, measured results and limitations come from the cited publications. Patterns, operating questions, role takeaways and implementation considerations are Lighthouse Advisory interpretation, stated as questions to validate locally rather than guaranteed outcomes. Vendor and operator claims are labeled as claims. Full research method.
- Academic research
- Research produced through an academic institution or peer-reviewed venue.
- Standards or public-body guidance
- Normative or advisory guidance from a standards body or public institution.