Lighthouse Advisory · Research librarySLED AI Adoption Intelligence
Daily · 10:00 PM Central

Issue 03 · Evidence briefing

SLED AI Adoption Intelligence

Measured digital-service performance, student-use evidence, procurement memory, and security controls for public-sector AI.

A decision-oriented read of what public institutions tried, what the evidence supports, and what leaders should design for next. Vendor claims are treated as claims, not outcomes.

5evidence records
4cross-source patterns
7topic lenses

Synthesis

Patterns across the evidence

01

Authoritative grounding works best with explicit service boundaries

GOV.UK Chat improved measured answer accuracy while remaining restricted to published government guidance and declining personal-case work. That combination—authoritative content, an in-scope answer threshold, source links, and human handoff—is a stronger public-service pattern than an unconstrained general assistant.

Which questions can the assistant answer from authoritative public content, and where must it refuse, clarify, or transfer to an authenticated human service?

02

AI governance needs reusable evidence, not repeated discovery

GAO found that four mature federal buyers did not systematically collect acquisition lessons, while OECD found weak and fragmented monitoring across government experimentation. SLED consortia and shared-service organizations can reduce repeated procurement mistakes by treating evaluation results, contract clauses, and operating lessons as managed institutional assets.

Where are contract terms, failure modes, evaluation results, model changes, and exit lessons stored so the next agency or institution can reuse them?

03

Student AI policy must govern the cognitive work, not only the tool

RAND's nationally representative youth survey found rising homework use alongside growing concern that AI harms critical thinking. The useful policy distinction is therefore cognitive augmentation versus cognitive offloading, supported by consistent course and school rules rather than a blanket approved-tool label.

For this assignment, which thinking must the student demonstrate independently, and which AI assistance deepens rather than replaces that thinking?

Supporting evidenceRAND Corporation
04

AI security is component- and lifecycle-specific

RAND's vulnerability analysis places the greatest structural exposure in training data and user-facing inference components such as context windows and retrieval pipelines. Those risks are often probabilistic and only partly patchable, so procurement and operations need compensating controls, provenance, monitoring, and repeated adversarial evaluation across model updates.

Which data, model, retrieval, context, interface, and tool components create risk, and which independent control contains each one when a conventional patch cannot?

Full record

Evidence ledger

Showing 5 of 5 records · All

Updated August 30, 2026

UK Government Digital ServiceUnited Kingdom

Large public pilots improve accuracy while preserving a hard boundary around personal casework

Two GOV.UK Chat pilots involved more than 10,000 users and roughly 26,000 questions about taxes, benefits, visas, and other public services. The system was constrained to published GOV.UK guidance and did not access personal records or perform case-specific transactions.

Government evaluationMixedDigital government and citizen services
Read full analysis

What happened

Two GOV.UK Chat pilots involved more than 10,000 users and roughly 26,000 questions about taxes, benefits, visas, and other public services. The system was constrained to published GOV.UK guidance and did not access personal records or perform case-specific transactions.

Evidence read

Government evaluators report accuracy increasing from an early 76% benchmark to 90% across topics, an 88% answer rate for in-scope questions, and 73% usefulness and 64% satisfaction in an app follow-up survey. All 508 recorded jailbreak attempts were blocked. The evidence combines expert and automated scoring, usability tests, surveys, conversation reviews, and journey analysis, but the government evaluated its own service.

Why it matters for SLED

This is one of the larger measured public tests of a citizen-facing government assistant and provides practical evidence for state portals, municipal service navigation, and education-administration help desks.

Architecture implications

Use authoritative retrieval, source links, scope detection, clarification for ambiguous questions, and a clean handoff to authenticated human support. The implementation used Amazon Bedrock and Anthropic models behind a model-changeable service layer, illustrating how a public interface can remain stable while models are re-evaluated and replaced.

Governance implications

Publish the assistant's purpose and limits, define a strict answer-quality rubric, track accuracy by topic, and treat latency, refusals, escalation, and user trust as production service measures—not secondary UX concerns.

Security and privacy implications

Keep personal case data outside the public assistant until identity, authorization, data minimization, logging, and human escalation are designed. Re-test jailbreak and content safeguards whenever models, retrieval, or streaming behavior changes.

Limits of the evidence

The findings are operator-reported, not independently audited; accuracy scoring reflects GOV.UK's rubric and sampled conversations; satisfaction survey samples were much smaller than total participation; and a 90% score still leaves material room for incomplete or incorrect answers in consequential domains.

Government Digital Service pilot findings (opens in a new tab)
U.S. Government Accountability OfficeUnited States

Federal audit finds AI buyers are not systematically capturing procurement lessons

GAO reviewed 13 AI acquisitions and 44 contracts or agreements across Defense, Homeland Security, GSA, and Veterans Affairs. None of the four agencies had policies requiring systematic collection of lessons learned for government-wide reuse.

Government auditCautionaryGovernment procurement and acquisition
Read full analysis

What happened

GAO reviewed 13 AI acquisitions and 44 contracts or agreements across Defense, Homeland Security, GSA, and Veterans Affairs. None of the four agencies had policies requiring systematic collection of lessons learned for government-wide reuse.

Evidence read

The performance audit found agencies were missing opportunities to capture practices such as data-rights terms and testing requirements or to avoid recurring mistakes. GAO made four policy recommendations, one to each agency, and all concurred. The sample was intentionally varied but nongeneralizable.

Why it matters for SLED

States, localities, districts, and public universities face the same fast-changing market with less contracting capacity. A shared acquisition memory can prevent every entity from rediscovering data-rights, testing, competition, monitoring, and exit problems independently.

Architecture implications

Procurement artifacts should identify model, infrastructure, data, integration, testing, monitoring, version-change, and portability responsibilities across the full AI stack rather than buying 'AI' as an undifferentiated capability.

Governance implications

Require an acquisition closeout and renewal record covering outcome evidence, failures, contract clauses, vendor performance, data rights, testing, cost behavior, model changes, and exit experience; publish reusable lessons through a statewide, systemwide, or consortium repository.

Security and privacy implications

Contracts should preserve audit and testing rights, define handling of government and constituent data, require change notification and ongoing performance monitoring, and establish deletion, portability, incident response, and termination obligations.

Limits of the evidence

The audit covers four federal agencies and a nongeneralizable sample selected partly for maturity and impact; it evaluates acquisition practice rather than the performance of the acquired AI systems and does not establish which contract approach yields the best return.

GAO-26-107859 (opens in a new tab)
RAND CorporationUnited States

Student homework use rises as concern about critical-thinking harm also grows

A nationally representative RAND survey of 1,214 U.S. youth ages 12–29 found that reported homework use of AI increased substantially during 2025 while students remained uncertain about school rules and the effect on their own learning.

Independent researchCautionaryK–12 and higher education
Read full analysis

What happened

A nationally representative RAND survey of 1,214 U.S. youth ages 12–29 found that reported homework use of AI increased substantially during 2025 while students remained uncertain about school rules and the effect on their own learning.

Evidence read

Reported AI homework use rose from 48% in May 2025 to 62% in December 2025. Sixty-seven percent agreed that greater AI use for schoolwork would harm critical-thinking skills, more than 10 points higher than ten months earlier. Older students more often reported teacher-dependent rules and concern about being accused of cheating.

Why it matters for SLED

District and institution policy must respond to ordinary student behavior, not a hypothetical future. The evidence also supports designing assessment around cognitive work rather than relying primarily on detection or inconsistent teacher-by-teacher rules.

Architecture implications

Learning platforms and managed AI workspaces should make allowed modes visible at the assignment level, preserve process evidence where appropriate, and support AI-free as well as AI-augmented work rather than assuming one configuration fits every learning objective.

Governance implications

Create consistent schoolwide or institution-wide categories for cognitive augmentation, permitted assistance, disclosure, and independent work; involve students in policy design; and align assignments and appeals with those categories.

Security and privacy implications

Avoid turning concern about misuse into pervasive surveillance or automated accusations. Minimize collection of student prompts and drafts, restrict access, set retention limits, and keep detector scores out of consequential decisions without corroborating evidence and due process.

Limits of the evidence

The results are self-reported perceptions and behavior, not direct measures of learning loss or causal effects. The 12–29 age range spans substantially different educational settings, and concern about critical thinking does not demonstrate that harm occurred.

RAND American Youth Panel report (opens in a new tab)
Organisation for Economic Co-operation and DevelopmentInternational; official guidance from 14 countries

Cross-country review finds experimentation widespread but monitoring and evaluation weak

OECD reviewed official experimentation guidance across 14 countries, academic literature, and case studies. It found rapid decentralized uptake, fragmented guidance, and few governments systematically measuring performance, impact, or compliance.

Standards or public-body guidanceEmergingPublic administration
Read full analysis

What happened

OECD reviewed official experimentation guidance across 14 countries, academic literature, and case studies. It found rapid decentralized uptake, fragmented guidance, and few governments systematically measuring performance, impact, or compliance.

Evidence read

The working paper synthesizes documented practices rather than testing one intervention. It proposes evaluation across five dimensions—performance, public value, feasibility, usability, and risk management—and identifies structured experimentation as a bridge between principles and scaled operation.

Why it matters for SLED

The gap mirrors SLED conditions: employees experiment before governance is mature, smaller organizations face inconsistent rules, and pilots are often counted without demonstrating service, learning, workforce, or compliance outcomes.

Architecture implications

Provide segregated sandboxes, approved data paths, reusable evaluation harnesses, model and prompt logging, and a governed route from experiment to production. Instrument quality, cost, latency, accessibility, and risk from the beginning rather than after a pilot is declared successful.

Governance implications

Use a common experiment charter with an accountable owner, hypothesis, baseline, success and stop criteria, affected-user review, and evidence package. Allow local experimentation within shared guardrails while central teams provide templates, expertise, procurement, and assurance services.

Security and privacy implications

Match experiment environments to data sensitivity; prohibit uncontrolled sensitive-data use; document model and vendor handling; and perform privacy, security, bias, and misuse testing before expanding access or authority.

Limits of the evidence

This is comparative guidance and synthesis, not causal outcome evidence. Official guidelines may differ from actual agency practice, and the review's international scope means legal and administrative assumptions do not transfer uniformly to U.S. SLED organizations.

OECD Working Papers on Public Governance No. 93 (opens in a new tab)
RAND CorporationInternational relevance

New vulnerability framework treats many AI weaknesses as structural rather than patchable

RAND decomposed generative AI architectures from training data through deployment interfaces and identified 31 vulnerability classes. Its highest aggregate risks clustered around training data and user-facing inference boundaries, including context windows and retrieval-augmented generation pipelines.

Independent researchCautionaryAI security and risk management
Read full analysis

What happened

RAND decomposed generative AI architectures from training data through deployment interfaces and identified 31 vulnerability classes. Its highest aggregate risks clustered around training data and user-facing inference boundaries, including context windows and retrieval-augmented generation pipelines.

Evidence read

The researchers combined literature review, public incident and attack sources monitored from August 2025 through March 2026, architectural decomposition, and structured threat and impact metrics. They conclude that some weaknesses persist across model versions and can be reduced but not eliminated through conventional patching.

Why it matters for SLED

SLED security teams need to integrate AI into vulnerability management without pretending probabilistic model behavior maps neatly to conventional CVEs or patch cycles. The framework gives architects and buyers a component-level way to assign controls and residual risk.

Architecture implications

Threat-model training and fine-tuning data, provenance, embeddings, context, retrieval, prompts, output interfaces, and any connected tools separately. Add input validation, retrieval isolation, provenance checks, least privilege, output controls, anomaly detection, and stochastic adversarial testing as compensating controls.

Governance implications

Require component-level risk assessments and residual-risk acceptance; connect AI findings to existing vulnerability, change, incident, and supplier-management processes; and re-evaluate after model, data, retrieval, or tool changes.

Security and privacy implications

Prioritize dataset provenance and controls at context and retrieval boundaries, where poisoned or injected content can affect confidentiality and integrity. Logging and monitoring must detect probabilistic exploitation and resource-exhaustion patterns, not only deterministic signatures.

Limits of the evidence

The taxonomy combines real-world and theoretical attack evidence and scores vulnerability classes rather than product-specific defects. It excludes bias harms, attacks that merely use AI, and external infrastructure or supply-chain vulnerabilities, and should be treated as an expandable baseline rather than a complete standard.

RAND research report RR-A4983-1 (opens in a new tab)

How to read this briefing

Methodology and definitions

Selection and freshness

This edition prioritizes primary government material, public audits, independent research, and relevant public-sector association guidance available for theAugust 30, 2026 run. Every surfaced item remains in the All view and keeps its original source.

Evidence classes

Government evaluation
A public body’s measured evaluation or documented pilot.
Government audit
An oversight review of performance, controls, or operations.
Academic research
Research produced through an academic institution or peer-reviewed venue.
Independent research
Research conducted outside the implementing organization.
Public-sector association guidance
Practitioner guidance or an association-supplied case; not independent outcome evidence.
Independent reporting
Independent reporting with attributable sources but without a formal evaluation design.
Standards or public-body guidance
Normative or advisory guidance from a standards body or public institution.
Vendor claim
A supplier-provided assertion that has not been upgraded to independent evidence.

Outcome labels

Effective
Evidence supports a useful result within the tested scope.
Mixed
Benefits and material limitations appear together.
Cautionary
The record surfaces failure, risk, or a control gap.
Emerging
A developing practice or claim without measured outcomes.

Claims discipline

Vendor, operator, and association claims are attributed and are not upgraded to independent evidence. Caveats identify self-reporting, bounded pilots, contested findings, and missing outcome measures.