From the SLED-wide archive edition of August 27, 2026 and 1 later edition
Large public pilots improve accuracy while preserving a hard boundary around personal casework
UK Government Digital Service · Digital government and citizen services · United Kingdom
- Publisher
- 5 things we learned testing GOV.UK Chat
- Original publication
- March 16, 2026
- Source retrieved
- Not recorded in the historical archive
What happened
Two GOV.UK Chat pilots involved more than 10,000 users and roughly 26,000 questions about taxes, benefits, visas, and other public services. The system was constrained to published GOV.UK guidance and did not access personal records or perform case-specific transactions.
Why it matters
This is one of the larger measured public tests of a citizen-facing government assistant and provides practical evidence for state portals, municipal service navigation, and education-administration help desks.
Evidence and measured results
Government evaluators report accuracy increasing from an early 76% benchmark to 90% across topics, an 88% answer rate for in-scope questions, and 73% usefulness and 64% satisfaction in an app follow-up survey. All 508 recorded jailbreak attempts were blocked. The evidence combines expert and automated scoring, usability tests, surveys, conversation reviews, and journey analysis, but the government evaluated its own service.
Limitations and uncertainty
The findings are operator-reported, not independently audited; accuracy scoring reflects GOV.UK's rubric and sampled conversations; satisfaction survey samples were much smaller than total participation; and a 90% score still leaves material room for incomplete or incorrect answers in consequential domains.
Put this evidence to work
Lighthouse Advisory interpretation, grounded in this source as summarized in the preserved archive. Enriched 2026-09-05; this does not change the original publication date. Labels below come from the analysis itself.
Sales
Role takeaway
Problem and stakeholders: Portal and contact-center owners may need to help residents navigate published guidance without drawing a public assistant into personal casework.
- Discovery
- Which questions have authoritative answers, who maintains them, and where can users reach authenticated support?
- Value hypothesis
- A bounded information service could improve navigation and reduce avoidable handoffs, subject to local measurement.
- Potential engagement
- A content-readiness review and limited citizen-service pilot involving digital services, communications, accessibility, security, and frontline staff.
- Evidence boundary
- The archived UK operator evaluation supplies a rubric and scale example. Its reported 90% accuracy and blocked jailbreak attempts do not establish local accuracy, cost savings, resistance to new attacks, or suitability for benefits decisions.
Pre-sales engineering
Role takeaway
- Fit
- Use this pattern for public information grounded in approved content.
- Architecture
- Separate retrieval, scope detection, model access, citations, and authenticated casework so model changes do not redefine service authority.
- Prerequisites
- Maintained sources, representative local questions, an answer-quality rubric, and functioning human support.
- Constraints
- Topic coverage, ambiguity, latency, and accessibility require local assessment; the archived Bedrock/Anthropic implementation is an example rather than a requirement.
- Security
- Keep personal records outside public chat and test injected content, jailbreaks, and accidental sensitive input.
- Proposed validation
- Independently score answers by topic, inspect citation support and refusals, and compare navigation and handoff performance with the existing service before expansion.
Delivery
Role takeaway
Work and dependencies: Establish curated content, a question benchmark, escalation procedures, and regression tests before a resident pilot.
- Ownership
- The service owner accepts outcomes; content stewards maintain sources, security reviews boundaries, and contact-center staff own escalations.
- Skills and adoption
- Train editors and support agents to distinguish unanswered, ambiguous, and out-of-scope requests; involve residents using assistive technology.
- Governance checkpoints
- Approve scope and data handling before launch and re-evaluate model or retrieval changes.
- Proposed acceptance
- Agreed topic-level quality, supported source links, accessible task completion, and measured escalation response must be demonstrated against local baselines.
- Risks
- Stale guidance, smaller satisfaction samples, and remaining answer errors make headline accuracy unsuitable as an automatic release threshold.
Implementation considerations
Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.
Architecture and integration
What must connect, and where does the AI sit in the workflow?
Use authoritative retrieval, source links, scope detection, clarification for ambiguous questions, and a clean handoff to authenticated human support. The implementation used Amazon Bedrock and Anthropic models behind a model-changeable service layer, illustrating how a public interface can remain stable while models are re-evaluated and replaced.
Governance
Who approves, reviews and stays accountable for outcomes?
Publish the assistant's purpose and limits, define a strict answer-quality rubric, track accuracy by topic, and treat latency, refusals, escalation, and user trust as production service measures—not secondary UX concerns.
Security and privacy
What data, permissions and controls need testing?
Keep personal case data outside the public assistant until identity, authorization, data minimization, logging, and human escalation are designed. Re-test jailbreak and content safeguards whenever models, retrieval, or streaming behavior changes.
The preserved archive analysis covered architecture, governance and security. Not assessed for this record: accessibility and workforce, procurement, operating model.
Publication history
- 2026-08-30SLED-wide archive · Issue 035 resources
- 2026-08-27SLED-wide archive · Issue 0110 resources
Stable resource ID: govuk-chat-pilots