Lighthouse Advisory · Research librarySLED AI Adoption Intelligence
Daily · 10:00 PM Central

Issue 01 · Evidence briefing

SLED AI Adoption Intelligence

Evidence, outcomes, and operating lessons for public-sector AI.

A decision-oriented read of what public institutions tried, what the evidence supports, and what leaders should design for next. Vendor claims are treated as claims, not outcomes.

10evidence records
5cross-source patterns
7topic lenses

Edition intelligence

Search the record over time

Enter one or more keywords to search the evidence record.
Edition archive7 published
  1. Issue 07September 3, 20266 records
  2. Issue 06September 2, 20265 records
  3. Issue 05September 1, 20266 records
  4. Issue 04August 31, 20264 records
  5. Issue 03August 30, 20265 records
  6. Issue 02August 29, 202614 records
  7. Issue 01August 27, 202610 records

Synthesis

Patterns across the evidence

01

Value is clearest in bounded, routine knowledge work

Measured and reported gains cluster around search, summarization, email, drafting, and coding support—not autonomous replacement of public judgment.

Which role-level baselines will prove the gain?

02

Public-facing assistants need bounded knowledge and continuous evaluation

GOV.UK’s source-bounded, measured rollout contrasts with MyCity’s audit dispute over accuracy, testing, and inconsistent answers.

Can the accuracy claim be reproduced independently?

03

Federated delivery works best with central assurance

Agency teams need room to develop use cases, while common intake, risk tiering, architecture review, training, and production gates remain centralized.

Who can approve, operate, pause, and retire the system?

04

Inventory and procurement learning remain operational weak points

Growing portfolios are difficult to govern when use cases arrive late in inventories and acquisition lessons are not captured for reuse.

Does every use case have an owner, risk tier, and exit path?

05

Education adoption is outpacing policy and learning evidence

Student and teacher use is rising faster than consistent rules, training, and evidence that distinguishes learning support from cognitive substitution.

Are rules consistent enough to support learning and equity?

Full record

Evidence ledger

Showing 10 of 10 records · All

Updated August 27, 2026

UK Department for Work and PensionsUnited Kingdom

Large workforce trial finds measurable—but smaller—time savings

A 3,549-person Microsoft 365 Copilot trial paired staff feedback with econometric analysis, separating perceived gains from measured time savings.

Government evaluationEffectiveGovernment operations
Read full analysis

What happened

A 3,549-person Microsoft 365 Copilot trial paired staff feedback with econometric analysis, separating perceived gains from measured time savings.

Evidence read

Ninety percent of participants reported saving time, while the econometric estimate was 19 minutes per person per day. Seventy-three percent perceived better work quality.

Why it matters for SLED

The gap between reported and measured gains gives state and local agencies a realistic basis for benefits cases and broad productivity-tool rollouts.

Architecture implications

Instrument adoption and time saved at the workflow level; connect approved copilots to governed document stores and preserve usage telemetry for benefits analysis.

Governance implications

Pair licensing with role-based training, baseline measures, acceptable-use rules, and a named benefits owner who can review continuation or expansion.

Security and privacy implications

Validate tenant boundaries, data-loss controls, retention, and access permissions before assistants can retrieve or summarize sensitive case material.

Limits of the evidence

The evaluation was not a randomized controlled trial, and several outcomes rely on participant self-reporting.

Department for Work and Pensions evaluation (opens in a new tab)
UK Government Digital ServiceUnited Kingdom

Public-sector developers report strong utility from coding assistants

A cross-government GitHub Copilot trial examined adoption, code acceptance, developer sentiment, and reported time savings across a large license cohort.

Government evaluationEffectiveDigital service delivery
Read full analysis

What happened

A cross-government GitHub Copilot trial examined adoption, code acceptance, developer sentiment, and reported time savings across a large license cohort.

Evidence read

Among 1,100 activated licenses, an average 418 users were active daily and participants averaged 2,298 chats. The code-line acceptance rate was 15.8%, and 58% said they would not return to working without an assistant.

Why it matters for SLED

Government technology teams can test assistants against delivery bottlenecks while monitoring actual usage, code acceptance, security, and developer experience.

Architecture implications

Treat approved coding assistants as part of the engineering platform, with repository boundaries, supported IDEs, code review, dependency scanning, and delivery telemetry.

Governance implications

Define permitted repositories and languages, require human review, and compare delivery and quality baselines before expanding access.

Security and privacy implications

Prevent sensitive source or secrets from entering unapproved models; validate data retention, prompt handling, and vendor training-use terms.

Limits of the evidence

Time savings were survey-reported, and code acceptance is not a direct measure of code quality or service outcomes.

Government Digital Service findings report (opens in a new tab)
GOV.UKUnited Kingdom

A public assistant improved accuracy through repeated live testing

Two GOV.UK Chat pilots used real questions, benchmark testing, and adversarial probes to improve a retrieval-based public information assistant.

Government evaluationEffectiveCitizen information services
Read full analysis

What happened

Two GOV.UK Chat pilots used real questions, benchmark testing, and adversarial probes to improve a retrieval-based public information assistant.

Evidence read

More than 10,000 users asked 26,000 questions; 73% rated the service useful and 64% were satisfied. Benchmark accuracy rose from 76% to 90%, 508 jailbreak attempts were blocked, the in-scope answer rate was 88%, and average response time was 10.7-second latency.

Why it matters for SLED

Public-facing assistants need a service-standard mindset: authoritative content retrieval, scoped questions, accessibility, safety testing, and ongoing quality measurement.

Architecture implications

Use source-bounded retrieval, an evaluation harness, latency monitoring, failure logging, and a service path that can degrade safely to authoritative content or human help.

Governance implications

Publish service measures, maintain escalation and content ownership, and require benchmark and red-team evidence before each material release.

Security and privacy implications

Log and analyze abuse without retaining unnecessary personal data; isolate retrieval sources and test jailbreak, prompt-injection, and disclosure risks.

Limits of the evidence

The results reflect a bounded pilot and benchmark; they do not establish accuracy for every topic or user circumstance.

Inside GOV.UK pilot lessons (opens in a new tab)
State of OhioOhio, United States

Federated governance moves approved state use cases into production

Ohio combined a central AI Council with agency participation, mandatory training, risk review, procurement checklists, and data classification for a growing statewide portfolio.

Public-sector association guidanceEmergingState government
Read full analysis

What happened

Ohio combined a central AI Council with agency participation, mandatory training, risk review, procurement checklists, and data classification for a growing statewide portfolio.

Evidence read

The association case study reports more than 100 approved use cases, 33 in use, and seven accessibility use cases in production.

Why it matters for SLED

A federated model can create a repeatable path from agency idea to approved deployment without forcing all domain decisions into one central team.

Architecture implications

Create common intake, reference architectures, data classification, reusable platform services, and higher-risk architecture review across agencies.

Governance implications

Define decision rights across the AI Council, agency owners, security, legal, accessibility, data governance, and procurement, with one risk-tiered production gate.

Security and privacy implications

Use data classification and security review to determine approved hosting, model access, logging, and human oversight for each risk tier.

Limits of the evidence

Deployment counts are association-supplied and do not independently establish benefit, equity, reliability, or cost-effectiveness.

NASCIO state innovation case study (opens in a new tab)
U.S. Internal Revenue ServiceUnited States

A large AI portfolio exposes inventory and workforce control gaps

GAO reviewed how the IRS manages a fast-growing AI portfolio spanning operational tools, voice bots, chatbots, and development-stage use cases.

Government auditMixedTax administration
Read full analysis

What happened

GAO reviewed how the IRS manages a fast-growing AI portfolio spanning operational tools, voice bots, chatbots, and development-stage use cases.

Evidence read

The IRS reported 126 use cases by June 2025, with about one-third operational. Eleven voice bots and two chatbots handled nearly 25 million sessions during the 2025 filing season; GAO also found inventory lag, mislabeling, and workforce gaps.

Why it matters for SLED

As portfolios grow, SLED leaders need an inventory connecting each use case to an owner, lifecycle status, data, risk tier, controls, and performance evidence.

Architecture implications

Make the use-case inventory a living system of record tied to architecture review, dependencies, production status, monitoring, and retirement.

Governance implications

Assign portfolio ownership, inventory-quality controls, workforce planning, outcome reporting, and deadlines for correcting lifecycle records.

Security and privacy implications

Link each use case to its data classification, privacy impact, authorization boundary, model monitoring, and incident owner.

Limits of the evidence

The federal tax environment differs from SLED operations, and session volume does not demonstrate answer quality or citizen benefit.

GAO-26-107522 (opens in a new tab)
U.S. Government Accountability OfficeUnited States

AI acquisition reviews rarely capture reusable lessons

GAO examined 13 AI acquisitions at four federal agencies and found that procurement processes did not systematically preserve lessons for future buyers.

Government auditCautionaryPublic procurement
Read full analysis

What happened

GAO examined 13 AI acquisitions at four federal agencies and found that procurement processes did not systematically preserve lessons for future buyers.

Evidence read

Four agencies lacked systematic lessons-learned requirements, missing reusable learning on data rights, testing, regional model accuracy, and discontinued solutions. All four concurred with GAO’s recommendations.

Why it matters for SLED

SLED buyers face similar risks when contracts omit evaluation data, performance thresholds, portability, audit access, or an exit path.

Architecture implications

Require pre-award test plans, integration boundaries, portability, performance acceptance criteria, observability, and a documented exit architecture.

Governance implications

Use AI-specific solicitation clauses and a shared lessons repository; assign procurement, legal, data, and technical owners to acceptance and renewal decisions.

Security and privacy implications

Contract for audit access, incident duties, data use and deletion, model-change notice, subcontractor controls, and security testing evidence.

Limits of the evidence

The sample is federal and limited to 13 acquisitions; local procurement statutes and market conditions vary.

GAO-26-107859 (opens in a new tab)
New York City ComptrollerNew York City, United States

City chatbot audit finds inconsistent answers and weak test evidence

An independent city audit identified inaccurate or inconsistent chatbot responses, performance delays, and insufficient detail behind reported accuracy claims.

Government auditCautionaryMunicipal citizen services
Read full analysis

What happened

An independent city audit identified inaccurate or inconsistent chatbot responses, performance delays, and insufficient detail behind reported accuracy claims.

Evidence read

Auditors tested system behavior and requested structured red-teaming. The Office of Technology and Innovation reported 95–99% accuracy but supplied insufficient test detail and disagreed with recommendations.

Why it matters for SLED

Public assistants create direct service and trust risk when answers are authoritative in tone but weakly tested or inconsistently documented.

Architecture implications

Build reproducible evaluation sets, performance monitoring, authoritative-source retrieval, failure capture, and a fallback path into the service architecture.

Governance implications

Require public error reporting, red-team protocols, content ownership, release gates, and independent assurance for high-impact performance claims.

Security and privacy implications

Test prompt injection, data disclosure, abusive use, and log handling as part of an auditable preproduction and recurring assurance program.

Limits of the evidence

The agency contested parts of the audit; readers should review both the findings and the response in the source.

New York City Comptroller audit (opens in a new tab)
RAND CorporationUnited States

School AI use grows faster than policy and professional learning

Nationally representative panels found rapidly growing AI use among students and teachers while training, school policy, and shared expectations remained uneven.

Independent researchMixedK–12 education
Read full analysis

What happened

Nationally representative panels found rapidly growing AI use among students and teachers while training, school policy, and shared expectations remained uneven.

Evidence read

Fifty-four percent of students and 53% of core-subject teachers reported using AI in 2025, each increasing by more than 15 percentage points. Training and policy lagged while stakeholder risk perceptions diverged.

Why it matters for SLED

Districts need operational guidance distinguishing instructional uses, student support, assessment integrity, accessibility, privacy, and staff responsibilities.

Architecture implications

Use an approved-tool catalog with identity, age-appropriate access, accessibility, integration, logging, and data minimization requirements.

Governance implications

Align acceptable-use rules, professional learning, assessment guidance, procurement, family communication, and outcome review at the district level.

Security and privacy implications

Apply student privacy law, parental and age safeguards, vendor data-use limits, retention controls, and non-AI access pathways.

Limits of the evidence

Survey responses describe reported behavior and perceptions, not causal effects on learning or teacher productivity.

RAND RRA4180-1 (opens in a new tab)
RAND CorporationUnited States

Students see convenience and critical-thinking risk at the same time

A survey of 1,214 youth tracked growing homework use and students’ own concerns about how generative AI may affect critical thinking.

Independent researchCautionarySecondary education
Read full analysis

What happened

A survey of 1,214 youth tracked growing homework use and students’ own concerns about how generative AI may affect critical thinking.

Evidence read

Reported homework use rose from 48% to 62% during 2025, while 67% believed greater AI use would harm critical-thinking skills. School rules often varied by teacher.

Why it matters for SLED

District policy must be legible at classroom level; inconsistent rules can undermine trust, equitable access, and meaningful assessment.

Architecture implications

Provide approved learning tools that support assignment-level disclosure, citation, accessibility, and non-AI completion paths without silent model changes.

Governance implications

Set consistent district expectations while giving educators task-level guidance; monitor learning, assessment validity, access, and student experience.

Security and privacy implications

Minimize student data, prohibit unapproved accounts, document vendor use of prompts and outputs, and preserve age-appropriate protections.

Limits of the evidence

Perceptions of harm do not prove a measured decline in critical thinking, and self-reported use may be imprecise.

RAND RRA4742-1 (opens in a new tab)
National Association of State Chief Information OfficersUnited States

Agentic systems shift state risk from answers to actions

NASCIO frames agentic AI as a move from drafting toward systems that take limited action across approvals, anomaly detection, and citizen-service workflows.

Public-sector association guidanceEmergingState government technology
Read full analysis

What happened

NASCIO frames agentic AI as a move from drafting toward systems that take limited action across approvals, anomaly detection, and citizen-service workflows.

Evidence read

The report identifies emerging opportunities and governance questions for multi-step automation; it does not present measured production outcomes.

Why it matters for SLED

The guidance is directly aimed at state technology leaders deciding where greater autonomy is useful and where it creates unacceptable operational risk.

Architecture implications

Use least-privilege tools, bounded action spaces, approval gates, tamper-resistant logs, transaction limits, rollback, and continuous evaluation.

Governance implications

Define who can authorize agent actions, approve higher-risk steps, pause execution, investigate incidents, and retire a workflow.

Security and privacy implications

Threat-model prompt injection, tool abuse, credential scope, cross-system data movement, unsafe chaining, and audit-log integrity before autonomous execution.

Limits of the evidence

This is emerging association guidance, not measured outcome evidence; treat the recommendations as an operating hypothesis to test.

NASCIO agentic AI report (opens in a new tab)

How to read this briefing

Methodology and definitions

Selection and freshness

This edition prioritizes primary government material, public audits, independent research, and relevant public-sector association guidance available for theAugust 27, 2026 run. Every surfaced item remains in the All view and keeps its original source.

Evidence classes

Government evaluation
A public body’s measured evaluation or documented pilot.
Government audit
An oversight review of performance, controls, or operations.
Academic research
Research produced through an academic institution or peer-reviewed venue.
Independent research
Research conducted outside the implementing organization.
Public-sector association guidance
Practitioner guidance or an association-supplied case; not independent outcome evidence.
Independent reporting
Independent reporting with attributable sources but without a formal evaluation design.
Standards or public-body guidance
Normative or advisory guidance from a standards body or public institution.
Vendor claim
A supplier-provided assertion that has not been upgraded to independent evidence.

Outcome labels

Effective
Evidence supports a useful result within the tested scope.
Mixed
Benefits and material limitations appear together.
Cautionary
The record surfaces failure, risk, or a control gap.
Emerging
A developing practice or claim without measured outcomes.

Claims discipline

Vendor, operator, and association claims are attributed and are not upgraded to independent evidence. Caveats identify self-reporting, bounded pilots, contested findings, and missing outcome measures.