Measured digital-service performance, student-use evidence, procurement memory, and security controls for public-sector AI.
A decision-oriented read of what public institutions tried, what the evidence supports, and what leaders should design for next. Vendor claims are treated as claims, not outcomes.
5evidence records
4cross-source patterns
7topic lenses
Edition intelligence
Search the record over time
Enter one or more keywords to search the evidence record.
Authoritative grounding works best with explicit service boundaries
GOV.UK Chat improved measured answer accuracy while remaining restricted to published government guidance and declining personal-case work. That combination—authoritative content, an in-scope answer threshold, source links, and human handoff—is a stronger public-service pattern than an unconstrained general assistant.
Which questions can the assistant answer from authoritative public content, and where must it refuse, clarify, or transfer to an authenticated human service?
AI governance needs reusable evidence, not repeated discovery
GAO found that four mature federal buyers did not systematically collect acquisition lessons, while OECD found weak and fragmented monitoring across government experimentation. SLED consortia and shared-service organizations can reduce repeated procurement mistakes by treating evaluation results, contract clauses, and operating lessons as managed institutional assets.
Where are contract terms, failure modes, evaluation results, model changes, and exit lessons stored so the next agency or institution can reuse them?
Student AI policy must govern the cognitive work, not only the tool
RAND's nationally representative youth survey found rising homework use alongside growing concern that AI harms critical thinking. The useful policy distinction is therefore cognitive augmentation versus cognitive offloading, supported by consistent course and school rules rather than a blanket approved-tool label.
For this assignment, which thinking must the student demonstrate independently, and which AI assistance deepens rather than replaces that thinking?
RAND's vulnerability analysis places the greatest structural exposure in training data and user-facing inference components such as context windows and retrieval pipelines. Those risks are often probabilistic and only partly patchable, so procurement and operations need compensating controls, provenance, monitoring, and repeated adversarial evaluation across model updates.
Which data, model, retrieval, context, interface, and tool components create risk, and which independent control contains each one when a conventional patch cannot?
Large public pilots improve accuracy while preserving a hard boundary around personal casework
Two GOV.UK Chat pilots involved more than 10,000 users and roughly 26,000 questions about taxes, benefits, visas, and other public services. The system was constrained to published GOV.UK guidance and did not access personal records or perform case-specific transactions.
Government evaluationMixedDigital government and citizen services
Read full analysis
What happened
Two GOV.UK Chat pilots involved more than 10,000 users and roughly 26,000 questions about taxes, benefits, visas, and other public services. The system was constrained to published GOV.UK guidance and did not access personal records or perform case-specific transactions.
Evidence read
Government evaluators report accuracy increasing from an early 76% benchmark to 90% across topics, an 88% answer rate for in-scope questions, and 73% usefulness and 64% satisfaction in an app follow-up survey. All 508 recorded jailbreak attempts were blocked. The evidence combines expert and automated scoring, usability tests, surveys, conversation reviews, and journey analysis, but the government evaluated its own service.
Why it matters for SLED
This is one of the larger measured public tests of a citizen-facing government assistant and provides practical evidence for state portals, municipal service navigation, and education-administration help desks.
Architecture implications
Use authoritative retrieval, source links, scope detection, clarification for ambiguous questions, and a clean handoff to authenticated human support. The implementation used Amazon Bedrock and Anthropic models behind a model-changeable service layer, illustrating how a public interface can remain stable while models are re-evaluated and replaced.
Governance implications
Publish the assistant's purpose and limits, define a strict answer-quality rubric, track accuracy by topic, and treat latency, refusals, escalation, and user trust as production service measures—not secondary UX concerns.
Security and privacy implications
Keep personal case data outside the public assistant until identity, authorization, data minimization, logging, and human escalation are designed. Re-test jailbreak and content safeguards whenever models, retrieval, or streaming behavior changes.
Limits of the evidence
The findings are operator-reported, not independently audited; accuracy scoring reflects GOV.UK's rubric and sampled conversations; satisfaction survey samples were much smaller than total participation; and a 90% score still leaves material room for incomplete or incorrect answers in consequential domains.
What happened
Two GOV.UK Chat pilots involved more than 10,000 users and roughly 26,000 questions about taxes, benefits, visas, and other public services. The system was constrained to published GOV.UK guidance and did not access personal records or perform case-specific transactions.
Evidence read
Government evaluators report accuracy increasing from an early 76% benchmark to 90% across topics, an 88% answer rate for in-scope questions, and 73% usefulness and 64% satisfaction in an app follow-up survey. All 508 recorded jailbreak attempts were blocked. The evidence combines expert and automated scoring, usability tests, surveys, conversation reviews, and journey analysis, but the government evaluated its own service.
Why it matters for SLED
This is one of the larger measured public tests of a citizen-facing government assistant and provides practical evidence for state portals, municipal service navigation, and education-administration help desks.
Architecture implications
Use authoritative retrieval, source links, scope detection, clarification for ambiguous questions, and a clean handoff to authenticated human support. The implementation used Amazon Bedrock and Anthropic models behind a model-changeable service layer, illustrating how a public interface can remain stable while models are re-evaluated and replaced.
Governance implications
Publish the assistant's purpose and limits, define a strict answer-quality rubric, track accuracy by topic, and treat latency, refusals, escalation, and user trust as production service measures—not secondary UX concerns.
Security and privacy implications
Keep personal case data outside the public assistant until identity, authorization, data minimization, logging, and human escalation are designed. Re-test jailbreak and content safeguards whenever models, retrieval, or streaming behavior changes.
Limits of the evidence
The findings are operator-reported, not independently audited; accuracy scoring reflects GOV.UK's rubric and sampled conversations; satisfaction survey samples were much smaller than total participation; and a 90% score still leaves material room for incomplete or incorrect answers in consequential domains.
U.S. Government Accountability OfficeUnited States
Federal audit finds AI buyers are not systematically capturing procurement lessons
GAO reviewed 13 AI acquisitions and 44 contracts or agreements across Defense, Homeland Security, GSA, and Veterans Affairs. None of the four agencies had policies requiring systematic collection of lessons learned for government-wide reuse.
Government auditCautionaryGovernment procurement and acquisition
Read full analysis
What happened
GAO reviewed 13 AI acquisitions and 44 contracts or agreements across Defense, Homeland Security, GSA, and Veterans Affairs. None of the four agencies had policies requiring systematic collection of lessons learned for government-wide reuse.
Evidence read
The performance audit found agencies were missing opportunities to capture practices such as data-rights terms and testing requirements or to avoid recurring mistakes. GAO made four policy recommendations, one to each agency, and all concurred. The sample was intentionally varied but nongeneralizable.
Why it matters for SLED
States, localities, districts, and public universities face the same fast-changing market with less contracting capacity. A shared acquisition memory can prevent every entity from rediscovering data-rights, testing, competition, monitoring, and exit problems independently.
Architecture implications
Procurement artifacts should identify model, infrastructure, data, integration, testing, monitoring, version-change, and portability responsibilities across the full AI stack rather than buying 'AI' as an undifferentiated capability.
Governance implications
Require an acquisition closeout and renewal record covering outcome evidence, failures, contract clauses, vendor performance, data rights, testing, cost behavior, model changes, and exit experience; publish reusable lessons through a statewide, systemwide, or consortium repository.
Security and privacy implications
Contracts should preserve audit and testing rights, define handling of government and constituent data, require change notification and ongoing performance monitoring, and establish deletion, portability, incident response, and termination obligations.
Limits of the evidence
The audit covers four federal agencies and a nongeneralizable sample selected partly for maturity and impact; it evaluates acquisition practice rather than the performance of the acquired AI systems and does not establish which contract approach yields the best return.
What happened
GAO reviewed 13 AI acquisitions and 44 contracts or agreements across Defense, Homeland Security, GSA, and Veterans Affairs. None of the four agencies had policies requiring systematic collection of lessons learned for government-wide reuse.
Evidence read
The performance audit found agencies were missing opportunities to capture practices such as data-rights terms and testing requirements or to avoid recurring mistakes. GAO made four policy recommendations, one to each agency, and all concurred. The sample was intentionally varied but nongeneralizable.
Why it matters for SLED
States, localities, districts, and public universities face the same fast-changing market with less contracting capacity. A shared acquisition memory can prevent every entity from rediscovering data-rights, testing, competition, monitoring, and exit problems independently.
Architecture implications
Procurement artifacts should identify model, infrastructure, data, integration, testing, monitoring, version-change, and portability responsibilities across the full AI stack rather than buying 'AI' as an undifferentiated capability.
Governance implications
Require an acquisition closeout and renewal record covering outcome evidence, failures, contract clauses, vendor performance, data rights, testing, cost behavior, model changes, and exit experience; publish reusable lessons through a statewide, systemwide, or consortium repository.
Security and privacy implications
Contracts should preserve audit and testing rights, define handling of government and constituent data, require change notification and ongoing performance monitoring, and establish deletion, portability, incident response, and termination obligations.
Limits of the evidence
The audit covers four federal agencies and a nongeneralizable sample selected partly for maturity and impact; it evaluates acquisition practice rather than the performance of the acquired AI systems and does not establish which contract approach yields the best return.
Student homework use rises as concern about critical-thinking harm also grows
A nationally representative RAND survey of 1,214 U.S. youth ages 12–29 found that reported homework use of AI increased substantially during 2025 while students remained uncertain about school rules and the effect on their own learning.
Independent researchCautionaryK–12 and higher education
Read full analysis
What happened
A nationally representative RAND survey of 1,214 U.S. youth ages 12–29 found that reported homework use of AI increased substantially during 2025 while students remained uncertain about school rules and the effect on their own learning.
Evidence read
Reported AI homework use rose from 48% in May 2025 to 62% in December 2025. Sixty-seven percent agreed that greater AI use for schoolwork would harm critical-thinking skills, more than 10 points higher than ten months earlier. Older students more often reported teacher-dependent rules and concern about being accused of cheating.
Why it matters for SLED
District and institution policy must respond to ordinary student behavior, not a hypothetical future. The evidence also supports designing assessment around cognitive work rather than relying primarily on detection or inconsistent teacher-by-teacher rules.
Architecture implications
Learning platforms and managed AI workspaces should make allowed modes visible at the assignment level, preserve process evidence where appropriate, and support AI-free as well as AI-augmented work rather than assuming one configuration fits every learning objective.
Governance implications
Create consistent schoolwide or institution-wide categories for cognitive augmentation, permitted assistance, disclosure, and independent work; involve students in policy design; and align assignments and appeals with those categories.
Security and privacy implications
Avoid turning concern about misuse into pervasive surveillance or automated accusations. Minimize collection of student prompts and drafts, restrict access, set retention limits, and keep detector scores out of consequential decisions without corroborating evidence and due process.
Limits of the evidence
The results are self-reported perceptions and behavior, not direct measures of learning loss or causal effects. The 12–29 age range spans substantially different educational settings, and concern about critical thinking does not demonstrate that harm occurred.
What happened
A nationally representative RAND survey of 1,214 U.S. youth ages 12–29 found that reported homework use of AI increased substantially during 2025 while students remained uncertain about school rules and the effect on their own learning.
Evidence read
Reported AI homework use rose from 48% in May 2025 to 62% in December 2025. Sixty-seven percent agreed that greater AI use for schoolwork would harm critical-thinking skills, more than 10 points higher than ten months earlier. Older students more often reported teacher-dependent rules and concern about being accused of cheating.
Why it matters for SLED
District and institution policy must respond to ordinary student behavior, not a hypothetical future. The evidence also supports designing assessment around cognitive work rather than relying primarily on detection or inconsistent teacher-by-teacher rules.
Architecture implications
Learning platforms and managed AI workspaces should make allowed modes visible at the assignment level, preserve process evidence where appropriate, and support AI-free as well as AI-augmented work rather than assuming one configuration fits every learning objective.
Governance implications
Create consistent schoolwide or institution-wide categories for cognitive augmentation, permitted assistance, disclosure, and independent work; involve students in policy design; and align assignments and appeals with those categories.
Security and privacy implications
Avoid turning concern about misuse into pervasive surveillance or automated accusations. Minimize collection of student prompts and drafts, restrict access, set retention limits, and keep detector scores out of consequential decisions without corroborating evidence and due process.
Limits of the evidence
The results are self-reported perceptions and behavior, not direct measures of learning loss or causal effects. The 12–29 age range spans substantially different educational settings, and concern about critical thinking does not demonstrate that harm occurred.
Organisation for Economic Co-operation and DevelopmentInternational; official guidance from 14 countries
Cross-country review finds experimentation widespread but monitoring and evaluation weak
OECD reviewed official experimentation guidance across 14 countries, academic literature, and case studies. It found rapid decentralized uptake, fragmented guidance, and few governments systematically measuring performance, impact, or compliance.
Standards or public-body guidanceEmergingPublic administration
Read full analysis
What happened
OECD reviewed official experimentation guidance across 14 countries, academic literature, and case studies. It found rapid decentralized uptake, fragmented guidance, and few governments systematically measuring performance, impact, or compliance.
Evidence read
The working paper synthesizes documented practices rather than testing one intervention. It proposes evaluation across five dimensions—performance, public value, feasibility, usability, and risk management—and identifies structured experimentation as a bridge between principles and scaled operation.
Why it matters for SLED
The gap mirrors SLED conditions: employees experiment before governance is mature, smaller organizations face inconsistent rules, and pilots are often counted without demonstrating service, learning, workforce, or compliance outcomes.
Architecture implications
Provide segregated sandboxes, approved data paths, reusable evaluation harnesses, model and prompt logging, and a governed route from experiment to production. Instrument quality, cost, latency, accessibility, and risk from the beginning rather than after a pilot is declared successful.
Governance implications
Use a common experiment charter with an accountable owner, hypothesis, baseline, success and stop criteria, affected-user review, and evidence package. Allow local experimentation within shared guardrails while central teams provide templates, expertise, procurement, and assurance services.
Security and privacy implications
Match experiment environments to data sensitivity; prohibit uncontrolled sensitive-data use; document model and vendor handling; and perform privacy, security, bias, and misuse testing before expanding access or authority.
Limits of the evidence
This is comparative guidance and synthesis, not causal outcome evidence. Official guidelines may differ from actual agency practice, and the review's international scope means legal and administrative assumptions do not transfer uniformly to U.S. SLED organizations.
What happened
OECD reviewed official experimentation guidance across 14 countries, academic literature, and case studies. It found rapid decentralized uptake, fragmented guidance, and few governments systematically measuring performance, impact, or compliance.
Evidence read
The working paper synthesizes documented practices rather than testing one intervention. It proposes evaluation across five dimensions—performance, public value, feasibility, usability, and risk management—and identifies structured experimentation as a bridge between principles and scaled operation.
Why it matters for SLED
The gap mirrors SLED conditions: employees experiment before governance is mature, smaller organizations face inconsistent rules, and pilots are often counted without demonstrating service, learning, workforce, or compliance outcomes.
Architecture implications
Provide segregated sandboxes, approved data paths, reusable evaluation harnesses, model and prompt logging, and a governed route from experiment to production. Instrument quality, cost, latency, accessibility, and risk from the beginning rather than after a pilot is declared successful.
Governance implications
Use a common experiment charter with an accountable owner, hypothesis, baseline, success and stop criteria, affected-user review, and evidence package. Allow local experimentation within shared guardrails while central teams provide templates, expertise, procurement, and assurance services.
Security and privacy implications
Match experiment environments to data sensitivity; prohibit uncontrolled sensitive-data use; document model and vendor handling; and perform privacy, security, bias, and misuse testing before expanding access or authority.
Limits of the evidence
This is comparative guidance and synthesis, not causal outcome evidence. Official guidelines may differ from actual agency practice, and the review's international scope means legal and administrative assumptions do not transfer uniformly to U.S. SLED organizations.
New vulnerability framework treats many AI weaknesses as structural rather than patchable
RAND decomposed generative AI architectures from training data through deployment interfaces and identified 31 vulnerability classes. Its highest aggregate risks clustered around training data and user-facing inference boundaries, including context windows and retrieval-augmented generation pipelines.
Independent researchCautionaryAI security and risk management
Read full analysis
What happened
RAND decomposed generative AI architectures from training data through deployment interfaces and identified 31 vulnerability classes. Its highest aggregate risks clustered around training data and user-facing inference boundaries, including context windows and retrieval-augmented generation pipelines.
Evidence read
The researchers combined literature review, public incident and attack sources monitored from August 2025 through March 2026, architectural decomposition, and structured threat and impact metrics. They conclude that some weaknesses persist across model versions and can be reduced but not eliminated through conventional patching.
Why it matters for SLED
SLED security teams need to integrate AI into vulnerability management without pretending probabilistic model behavior maps neatly to conventional CVEs or patch cycles. The framework gives architects and buyers a component-level way to assign controls and residual risk.
Architecture implications
Threat-model training and fine-tuning data, provenance, embeddings, context, retrieval, prompts, output interfaces, and any connected tools separately. Add input validation, retrieval isolation, provenance checks, least privilege, output controls, anomaly detection, and stochastic adversarial testing as compensating controls.
Governance implications
Require component-level risk assessments and residual-risk acceptance; connect AI findings to existing vulnerability, change, incident, and supplier-management processes; and re-evaluate after model, data, retrieval, or tool changes.
Security and privacy implications
Prioritize dataset provenance and controls at context and retrieval boundaries, where poisoned or injected content can affect confidentiality and integrity. Logging and monitoring must detect probabilistic exploitation and resource-exhaustion patterns, not only deterministic signatures.
Limits of the evidence
The taxonomy combines real-world and theoretical attack evidence and scores vulnerability classes rather than product-specific defects. It excludes bias harms, attacks that merely use AI, and external infrastructure or supply-chain vulnerabilities, and should be treated as an expandable baseline rather than a complete standard.
What happened
RAND decomposed generative AI architectures from training data through deployment interfaces and identified 31 vulnerability classes. Its highest aggregate risks clustered around training data and user-facing inference boundaries, including context windows and retrieval-augmented generation pipelines.
Evidence read
The researchers combined literature review, public incident and attack sources monitored from August 2025 through March 2026, architectural decomposition, and structured threat and impact metrics. They conclude that some weaknesses persist across model versions and can be reduced but not eliminated through conventional patching.
Why it matters for SLED
SLED security teams need to integrate AI into vulnerability management without pretending probabilistic model behavior maps neatly to conventional CVEs or patch cycles. The framework gives architects and buyers a component-level way to assign controls and residual risk.
Architecture implications
Threat-model training and fine-tuning data, provenance, embeddings, context, retrieval, prompts, output interfaces, and any connected tools separately. Add input validation, retrieval isolation, provenance checks, least privilege, output controls, anomaly detection, and stochastic adversarial testing as compensating controls.
Governance implications
Require component-level risk assessments and residual-risk acceptance; connect AI findings to existing vulnerability, change, incident, and supplier-management processes; and re-evaluate after model, data, retrieval, or tool changes.
Security and privacy implications
Prioritize dataset provenance and controls at context and retrieval boundaries, where poisoned or injected content can affect confidentiality and integrity. Logging and monitoring must detect probabilistic exploitation and resource-exhaustion patterns, not only deterministic signatures.
Limits of the evidence
The taxonomy combines real-world and theoretical attack evidence and scores vulnerability classes rather than product-specific defects. It excludes bias harms, attacks that merely use AI, and external infrastructure or supply-chain vulnerabilities, and should be treated as an expandable baseline rather than a complete standard.
This edition prioritizes primary government material, public audits, independent research, and relevant public-sector association guidance available for theAugust 30, 2026 run. Every surfaced item remains in the All view and keeps its original source.
Evidence classes
Government evaluation
A public body’s measured evaluation or documented pilot.
Government audit
An oversight review of performance, controls, or operations.
Academic research
Research produced through an academic institution or peer-reviewed venue.
Independent research
Research conducted outside the implementing organization.
Public-sector association guidance
Practitioner guidance or an association-supplied case; not independent outcome evidence.
Independent reporting
Independent reporting with attributable sources but without a formal evaluation design.
Standards or public-body guidance
Normative or advisory guidance from a standards body or public institution.
Vendor claim
A supplier-provided assertion that has not been upgraded to independent evidence.
Outcome labels
Effective
Evidence supports a useful result within the tested scope.
Mixed
Benefits and material limitations appear together.
Cautionary
The record surfaces failure, risk, or a control gap.
Emerging
A developing practice or claim without measured outcomes.
Claims discipline
Vendor, operator, and association claims are attributed and are not upgraded to independent evidence. Caveats identify self-reporting, bounded pilots, contested findings, and missing outcome measures.