Value is clearest in bounded, routine knowledge work
Measured and reported gains cluster around search, summarization, email, drafting, and coding support—not autonomous replacement of public judgment.
Issue 01 · Evidence briefing
Evidence, outcomes, and operating lessons for public-sector AI.
A decision-oriented read of what public institutions tried, what the evidence supports, and what leaders should design for next. Vendor claims are treated as claims, not outcomes.
Edition intelligence
Synthesis
Measured and reported gains cluster around search, summarization, email, drafting, and coding support—not autonomous replacement of public judgment.
GOV.UK’s source-bounded, measured rollout contrasts with MyCity’s audit dispute over accuracy, testing, and inconsistent answers.
Agency teams need room to develop use cases, while common intake, risk tiering, architecture review, training, and production gates remain centralized.
Growing portfolios are difficult to govern when use cases arrive late in inventories and acquisition lessons are not captured for reuse.
Student and teacher use is rising faster than consistent rules, training, and evidence that distinguishes learning support from cognitive substitution.
Full record
Showing 10 of 10 records · All
Updated August 27, 2026
A 3,549-person Microsoft 365 Copilot trial paired staff feedback with econometric analysis, separating perceived gains from measured time savings.
A 3,549-person Microsoft 365 Copilot trial paired staff feedback with econometric analysis, separating perceived gains from measured time savings.
Ninety percent of participants reported saving time, while the econometric estimate was 19 minutes per person per day. Seventy-three percent perceived better work quality.
The gap between reported and measured gains gives state and local agencies a realistic basis for benefits cases and broad productivity-tool rollouts.
Instrument adoption and time saved at the workflow level; connect approved copilots to governed document stores and preserve usage telemetry for benefits analysis.
Pair licensing with role-based training, baseline measures, acceptable-use rules, and a named benefits owner who can review continuation or expansion.
Validate tenant boundaries, data-loss controls, retention, and access permissions before assistants can retrieve or summarize sensitive case material.
The evaluation was not a randomized controlled trial, and several outcomes rely on participant self-reporting.
A cross-government GitHub Copilot trial examined adoption, code acceptance, developer sentiment, and reported time savings across a large license cohort.
A cross-government GitHub Copilot trial examined adoption, code acceptance, developer sentiment, and reported time savings across a large license cohort.
Among 1,100 activated licenses, an average 418 users were active daily and participants averaged 2,298 chats. The code-line acceptance rate was 15.8%, and 58% said they would not return to working without an assistant.
Government technology teams can test assistants against delivery bottlenecks while monitoring actual usage, code acceptance, security, and developer experience.
Treat approved coding assistants as part of the engineering platform, with repository boundaries, supported IDEs, code review, dependency scanning, and delivery telemetry.
Define permitted repositories and languages, require human review, and compare delivery and quality baselines before expanding access.
Prevent sensitive source or secrets from entering unapproved models; validate data retention, prompt handling, and vendor training-use terms.
Time savings were survey-reported, and code acceptance is not a direct measure of code quality or service outcomes.
Two GOV.UK Chat pilots used real questions, benchmark testing, and adversarial probes to improve a retrieval-based public information assistant.
Two GOV.UK Chat pilots used real questions, benchmark testing, and adversarial probes to improve a retrieval-based public information assistant.
More than 10,000 users asked 26,000 questions; 73% rated the service useful and 64% were satisfied. Benchmark accuracy rose from 76% to 90%, 508 jailbreak attempts were blocked, the in-scope answer rate was 88%, and average response time was 10.7-second latency.
Public-facing assistants need a service-standard mindset: authoritative content retrieval, scoped questions, accessibility, safety testing, and ongoing quality measurement.
Use source-bounded retrieval, an evaluation harness, latency monitoring, failure logging, and a service path that can degrade safely to authoritative content or human help.
Publish service measures, maintain escalation and content ownership, and require benchmark and red-team evidence before each material release.
Log and analyze abuse without retaining unnecessary personal data; isolate retrieval sources and test jailbreak, prompt-injection, and disclosure risks.
The results reflect a bounded pilot and benchmark; they do not establish accuracy for every topic or user circumstance.
Ohio combined a central AI Council with agency participation, mandatory training, risk review, procurement checklists, and data classification for a growing statewide portfolio.
Ohio combined a central AI Council with agency participation, mandatory training, risk review, procurement checklists, and data classification for a growing statewide portfolio.
The association case study reports more than 100 approved use cases, 33 in use, and seven accessibility use cases in production.
A federated model can create a repeatable path from agency idea to approved deployment without forcing all domain decisions into one central team.
Create common intake, reference architectures, data classification, reusable platform services, and higher-risk architecture review across agencies.
Define decision rights across the AI Council, agency owners, security, legal, accessibility, data governance, and procurement, with one risk-tiered production gate.
Use data classification and security review to determine approved hosting, model access, logging, and human oversight for each risk tier.
Deployment counts are association-supplied and do not independently establish benefit, equity, reliability, or cost-effectiveness.
GAO reviewed how the IRS manages a fast-growing AI portfolio spanning operational tools, voice bots, chatbots, and development-stage use cases.
GAO reviewed how the IRS manages a fast-growing AI portfolio spanning operational tools, voice bots, chatbots, and development-stage use cases.
The IRS reported 126 use cases by June 2025, with about one-third operational. Eleven voice bots and two chatbots handled nearly 25 million sessions during the 2025 filing season; GAO also found inventory lag, mislabeling, and workforce gaps.
As portfolios grow, SLED leaders need an inventory connecting each use case to an owner, lifecycle status, data, risk tier, controls, and performance evidence.
Make the use-case inventory a living system of record tied to architecture review, dependencies, production status, monitoring, and retirement.
Assign portfolio ownership, inventory-quality controls, workforce planning, outcome reporting, and deadlines for correcting lifecycle records.
Link each use case to its data classification, privacy impact, authorization boundary, model monitoring, and incident owner.
The federal tax environment differs from SLED operations, and session volume does not demonstrate answer quality or citizen benefit.
GAO examined 13 AI acquisitions at four federal agencies and found that procurement processes did not systematically preserve lessons for future buyers.
GAO examined 13 AI acquisitions at four federal agencies and found that procurement processes did not systematically preserve lessons for future buyers.
Four agencies lacked systematic lessons-learned requirements, missing reusable learning on data rights, testing, regional model accuracy, and discontinued solutions. All four concurred with GAO’s recommendations.
SLED buyers face similar risks when contracts omit evaluation data, performance thresholds, portability, audit access, or an exit path.
Require pre-award test plans, integration boundaries, portability, performance acceptance criteria, observability, and a documented exit architecture.
Use AI-specific solicitation clauses and a shared lessons repository; assign procurement, legal, data, and technical owners to acceptance and renewal decisions.
Contract for audit access, incident duties, data use and deletion, model-change notice, subcontractor controls, and security testing evidence.
The sample is federal and limited to 13 acquisitions; local procurement statutes and market conditions vary.
An independent city audit identified inaccurate or inconsistent chatbot responses, performance delays, and insufficient detail behind reported accuracy claims.
An independent city audit identified inaccurate or inconsistent chatbot responses, performance delays, and insufficient detail behind reported accuracy claims.
Auditors tested system behavior and requested structured red-teaming. The Office of Technology and Innovation reported 95–99% accuracy but supplied insufficient test detail and disagreed with recommendations.
Public assistants create direct service and trust risk when answers are authoritative in tone but weakly tested or inconsistently documented.
Build reproducible evaluation sets, performance monitoring, authoritative-source retrieval, failure capture, and a fallback path into the service architecture.
Require public error reporting, red-team protocols, content ownership, release gates, and independent assurance for high-impact performance claims.
Test prompt injection, data disclosure, abusive use, and log handling as part of an auditable preproduction and recurring assurance program.
The agency contested parts of the audit; readers should review both the findings and the response in the source.
Nationally representative panels found rapidly growing AI use among students and teachers while training, school policy, and shared expectations remained uneven.
Nationally representative panels found rapidly growing AI use among students and teachers while training, school policy, and shared expectations remained uneven.
Fifty-four percent of students and 53% of core-subject teachers reported using AI in 2025, each increasing by more than 15 percentage points. Training and policy lagged while stakeholder risk perceptions diverged.
Districts need operational guidance distinguishing instructional uses, student support, assessment integrity, accessibility, privacy, and staff responsibilities.
Use an approved-tool catalog with identity, age-appropriate access, accessibility, integration, logging, and data minimization requirements.
Align acceptable-use rules, professional learning, assessment guidance, procurement, family communication, and outcome review at the district level.
Apply student privacy law, parental and age safeguards, vendor data-use limits, retention controls, and non-AI access pathways.
Survey responses describe reported behavior and perceptions, not causal effects on learning or teacher productivity.
A survey of 1,214 youth tracked growing homework use and students’ own concerns about how generative AI may affect critical thinking.
A survey of 1,214 youth tracked growing homework use and students’ own concerns about how generative AI may affect critical thinking.
Reported homework use rose from 48% to 62% during 2025, while 67% believed greater AI use would harm critical-thinking skills. School rules often varied by teacher.
District policy must be legible at classroom level; inconsistent rules can undermine trust, equitable access, and meaningful assessment.
Provide approved learning tools that support assignment-level disclosure, citation, accessibility, and non-AI completion paths without silent model changes.
Set consistent district expectations while giving educators task-level guidance; monitor learning, assessment validity, access, and student experience.
Minimize student data, prohibit unapproved accounts, document vendor use of prompts and outputs, and preserve age-appropriate protections.
Perceptions of harm do not prove a measured decline in critical thinking, and self-reported use may be imprecise.
NASCIO frames agentic AI as a move from drafting toward systems that take limited action across approvals, anomaly detection, and citizen-service workflows.
NASCIO frames agentic AI as a move from drafting toward systems that take limited action across approvals, anomaly detection, and citizen-service workflows.
The report identifies emerging opportunities and governance questions for multi-step automation; it does not present measured production outcomes.
The guidance is directly aimed at state technology leaders deciding where greater autonomy is useful and where it creates unacceptable operational risk.
Use least-privilege tools, bounded action spaces, approval gates, tamper-resistant logs, transaction limits, rollback, and continuous evaluation.
Define who can authorize agent actions, approve higher-risk steps, pause execution, investigate incidents, and retire a workflow.
Threat-model prompt injection, tool abuse, credential scope, cross-system data movement, unsafe chaining, and audit-log integrity before autonomous execution.
This is emerging association guidance, not measured outcome evidence; treat the recommendations as an operating hypothesis to test.
How to read this briefing
This edition prioritizes primary government material, public audits, independent research, and relevant public-sector association guidance available for theAugust 27, 2026 run. Every surfaced item remains in the All view and keeps its original source.
Vendor, operator, and association claims are attributed and are not upgraded to independent evidence. Caveats identify self-reporting, bounded pilots, contested findings, and missing outcome measures.