Lighthouse AdvisorySLED AI Adoption Intelligence
← Back to results

From the SLED-wide archive edition of August 27, 2026 and 1 later edition

Government evaluationMixedPublished · Mar 2026

Large public pilots improve accuracy while preserving a hard boundary around personal casework

UK Government Digital Service · Digital government and citizen services · United Kingdom

Publisher
5 things we learned testing GOV.UK Chat
Original publication
March 16, 2026
Source retrieved
Not recorded in the historical archive
Read original source

What happened

Two GOV.UK Chat pilots involved more than 10,000 users and roughly 26,000 questions about taxes, benefits, visas, and other public services. The system was constrained to published GOV.UK guidance and did not access personal records or perform case-specific transactions.

Why it matters

This is one of the larger measured public tests of a citizen-facing government assistant and provides practical evidence for state portals, municipal service navigation, and education-administration help desks.

Evidence and measured results

Government evaluators report accuracy increasing from an early 76% benchmark to 90% across topics, an 88% answer rate for in-scope questions, and 73% usefulness and 64% satisfaction in an app follow-up survey. All 508 recorded jailbreak attempts were blocked. The evidence combines expert and automated scoring, usability tests, surveys, conversation reviews, and journey analysis, but the government evaluated its own service.

Limitations and uncertainty

The findings are operator-reported, not independently audited; accuracy scoring reflects GOV.UK's rubric and sampled conversations; satisfaction survey samples were much smaller than total participation; and a 90% score still leaves material room for incomplete or incorrect answers in consequential domains.

Put this evidence to work

Lighthouse Advisory interpretation, grounded in this source as summarized in the preserved archive. Enriched 2026-09-05; this does not change the original publication date. Labels below come from the analysis itself.

Sales

Role takeaway

Problem and stakeholders: Portal and contact-center owners may need to help residents navigate published guidance without drawing a public assistant into personal casework.

Discovery
Which questions have authoritative answers, who maintains them, and where can users reach authenticated support?
Value hypothesis
A bounded information service could improve navigation and reduce avoidable handoffs, subject to local measurement.
Potential engagement
A content-readiness review and limited citizen-service pilot involving digital services, communications, accessibility, security, and frontline staff.
Evidence boundary
The archived UK operator evaluation supplies a rubric and scale example. Its reported 90% accuracy and blocked jailbreak attempts do not establish local accuracy, cost savings, resistance to new attacks, or suitability for benefits decisions.

Pre-sales engineering

Role takeaway
Fit
Use this pattern for public information grounded in approved content.
Architecture
Separate retrieval, scope detection, model access, citations, and authenticated casework so model changes do not redefine service authority.
Prerequisites
Maintained sources, representative local questions, an answer-quality rubric, and functioning human support.
Constraints
Topic coverage, ambiguity, latency, and accessibility require local assessment; the archived Bedrock/Anthropic implementation is an example rather than a requirement.
Security
Keep personal records outside public chat and test injected content, jailbreaks, and accidental sensitive input.
Proposed validation
Independently score answers by topic, inspect citation support and refusals, and compare navigation and handoff performance with the existing service before expansion.

Delivery

Role takeaway

Work and dependencies: Establish curated content, a question benchmark, escalation procedures, and regression tests before a resident pilot.

Ownership
The service owner accepts outcomes; content stewards maintain sources, security reviews boundaries, and contact-center staff own escalations.
Skills and adoption
Train editors and support agents to distinguish unanswered, ambiguous, and out-of-scope requests; involve residents using assistive technology.
Governance checkpoints
Approve scope and data handling before launch and re-evaluate model or retrieval changes.
Proposed acceptance
Agreed topic-level quality, supported source links, accessible task completion, and measured escalation response must be demonstrated against local baselines.
Risks
Stale guidance, smaller satisfaction samples, and remaining answer errors make headline accuracy unsuitable as an automatic release threshold.

Implementation considerations

Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.

Architecture and integration

What must connect, and where does the AI sit in the workflow?

Use authoritative retrieval, source links, scope detection, clarification for ambiguous questions, and a clean handoff to authenticated human support. The implementation used Amazon Bedrock and Anthropic models behind a model-changeable service layer, illustrating how a public interface can remain stable while models are re-evaluated and replaced.

Governance

Who approves, reviews and stays accountable for outcomes?

Publish the assistant's purpose and limits, define a strict answer-quality rubric, track accuracy by topic, and treat latency, refusals, escalation, and user trust as production service measures—not secondary UX concerns.

Security and privacy

What data, permissions and controls need testing?

Keep personal case data outside the public assistant until identity, authorization, data minimization, logging, and human escalation are designed. Re-test jailbreak and content safeguards whenever models, retrieval, or streaming behavior changes.

The preserved archive analysis covered architecture, governance and security. Not assessed for this record: accessibility and workforce, procurement, operating model.

Publication history

  1. 2026-08-30SLED-wide archive · Issue 035 resources
  2. 2026-08-27SLED-wide archive · Issue 0110 resources
Read preserved resource versions (JSON)

Stable resource ID: govuk-chat-pilots