{"resourceId":"govuk-chat-pilots","versions":[{"version":"legacy/2026-08-30/govuk-chat-public-pilots","resource":{"id":"govuk-chat-public-pilots","title":"Large public pilots improve accuracy while preserving a hard boundary around personal casework","organization":"UK Government Digital Service","sector":"Digital government and citizen services","geography":"United Kingdom","publishedAt":"March 16, 2026","sourceName":"5 things we learned testing GOV.UK Chat","sourceLabel":"Government Digital Service pilot findings","sourceUrl":"https://insidegovuk.blog.gov.uk/2026/03/16/5-things-we-learned-testing-gov-uk-chat-an-ai-assistant-for-government/","evidenceClass":"government-evaluation","outcomeClass":"mixed","topics":["knowledge-work","developers-agents","infrastructure","data-security","accessibility-workforce","operating-model"],"finding":"Two GOV.UK Chat pilots involved more than 10,000 users and roughly 26,000 questions about taxes, benefits, visas, and other public services. The system was constrained to published GOV.UK guidance and did not access personal records or perform case-specific transactions.","sledRelevance":"This is one of the larger measured public tests of a citizen-facing government assistant and provides practical evidence for state portals, municipal service navigation, and education-administration help desks.","evidence":"Government evaluators report accuracy increasing from an early 76% benchmark to 90% across topics, an 88% answer rate for in-scope questions, and 73% usefulness and 64% satisfaction in an app follow-up survey. All 508 recorded jailbreak attempts were blocked. The evidence combines expert and automated scoring, usability tests, surveys, conversation reviews, and journey analysis, but the government evaluated its own service.","architectureImplications":"Use authoritative retrieval, source links, scope detection, clarification for ambiguous questions, and a clean handoff to authenticated human support. The implementation used Amazon Bedrock and Anthropic models behind a model-changeable service layer, illustrating how a public interface can remain stable while models are re-evaluated and replaced.","governanceImplications":"Publish the assistant's purpose and limits, define a strict answer-quality rubric, track accuracy by topic, and treat latency, refusals, escalation, and user trust as production service measures—not secondary UX concerns.","securityPrivacyImplications":"Keep personal case data outside the public assistant until identity, authorization, data minimization, logging, and human escalation are designed. Re-test jailbreak and content safeguards whenever models, retrieval, or streaming behavior changes.","caveats":"The findings are operator-reported, not independently audited; accuracy scoring reflects GOV.UK's rubric and sampled conversations; satisfaction survey samples were much smaller than total participation; and a 90% score still leaves material room for incomplete or incorrect answers in consequential domains."}},{"version":"legacy/2026-08-27/govuk-chat-pilots","resource":{"id":"govuk-chat-pilots","title":"A public assistant improved accuracy through repeated live testing","organization":"GOV.UK","sector":"Citizen information services","geography":"United Kingdom","publishedAt":"March 16, 2026","sourceName":"5 things we learned testing GOV.UK Chat","sourceLabel":"Inside GOV.UK pilot lessons","sourceUrl":"https://insidegovuk.blog.gov.uk/2026/03/16/5-things-we-learned-testing-gov-uk-chat-an-ai-assistant-for-government/","evidenceClass":"government-evaluation","outcomeClass":"effective","topics":["knowledge-work","infrastructure","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"Two GOV.UK Chat pilots used real questions, benchmark testing, and adversarial probes to improve a retrieval-based public information assistant.","sledRelevance":"Public-facing assistants need a service-standard mindset: authoritative content retrieval, scoped questions, accessibility, safety testing, and ongoing quality measurement.","evidence":"More than 10,000 users asked 26,000 questions; 73% rated the service useful and 64% were satisfied. Benchmark accuracy rose from 76% to 90%, 508 jailbreak attempts were blocked, the in-scope answer rate was 88%, and average response time was 10.7-second latency.","architectureImplications":"Use source-bounded retrieval, an evaluation harness, latency monitoring, failure logging, and a service path that can degrade safely to authoritative content or human help.","governanceImplications":"Publish service measures, maintain escalation and content ownership, and require benchmark and red-team evidence before each material release.","securityPrivacyImplications":"Log and analyze abuse without retaining unnecessary personal data; isolate retrieval sources and test jailbreak, prompt-injection, and disclosure risks.","caveats":"The results reflect a bounded pilot and benchmark; they do not establish accuracy for every topic or user circumstance."}},{"version":"enrichment/2026-09-05T02:42:45.193Z/govuk-chat-public-pilots","resource":{"id":"govuk-chat-pilots","title":"Large public pilots improve accuracy while preserving a hard boundary around personal casework","organization":"UK Government Digital Service","sector":"Digital government and citizen services","geography":"United Kingdom","publishedAt":"March 16, 2026","publicationDate":"2026-03-16","eventDate":null,"sourceName":"5 things we learned testing GOV.UK Chat","sourceLabel":"Government Digital Service pilot findings","sourceUrl":"https://insidegovuk.blog.gov.uk/2026/03/16/5-things-we-learned-testing-gov-uk-chat-an-ai-assistant-for-government/","evidenceClass":"government-evaluation","outcomeClass":"mixed","topics":["knowledge-work","developers-agents","infrastructure","data-security","accessibility-workforce","operating-model"],"finding":"Two GOV.UK Chat pilots involved more than 10,000 users and roughly 26,000 questions about taxes, benefits, visas, and other public services. The system was constrained to published GOV.UK guidance and did not access personal records or perform case-specific transactions.","sledRelevance":"This is one of the larger measured public tests of a citizen-facing government assistant and provides practical evidence for state portals, municipal service navigation, and education-administration help desks.","evidence":"Government evaluators report accuracy increasing from an early 76% benchmark to 90% across topics, an 88% answer rate for in-scope questions, and 73% usefulness and 64% satisfaction in an app follow-up survey. All 508 recorded jailbreak attempts were blocked. The evidence combines expert and automated scoring, usability tests, surveys, conversation reviews, and journey analysis, but the government evaluated its own service.","architectureImplications":"Use authoritative retrieval, source links, scope detection, clarification for ambiguous questions, and a clean handoff to authenticated human support. The implementation used Amazon Bedrock and Anthropic models behind a model-changeable service layer, illustrating how a public interface can remain stable while models are re-evaluated and replaced.","governanceImplications":"Publish the assistant's purpose and limits, define a strict answer-quality rubric, track accuracy by topic, and treat latency, refusals, escalation, and user trust as production service measures—not secondary UX concerns.","securityPrivacyImplications":"Keep personal case data outside the public assistant until identity, authorization, data minimization, logging, and human escalation are designed. Re-test jailbreak and content safeguards whenever models, retrieval, or streaming behavior changes.","caveats":"The findings are operator-reported, not independently audited; accuracy scoring reflects GOV.UK's rubric and sampled conversations; satisfaction survey samples were much smaller than total participation; and a 90% score still leaves material room for incomplete or incorrect answers in consequential domains.","streamIds":["state-government","local-government"],"roles":{"sales":"Interpretation — Problem and stakeholders: Portal and contact-center owners may need to help residents navigate published guidance without drawing a public assistant into personal casework. Discovery: Which questions have authoritative answers, who maintains them, and where can users reach authenticated support? Value hypothesis: A bounded information service could improve navigation and reduce avoidable handoffs, subject to local measurement. Potential engagement: A content-readiness review and limited citizen-service pilot involving digital services, communications, accessibility, security, and frontline staff. Evidence boundary: The archived UK operator evaluation supplies a rubric and scale example. Its reported 90% accuracy and blocked jailbreak attempts do not establish local accuracy, cost savings, resistance to new attacks, or suitability for benefits decisions.","engineering":"Interpretation — Fit: Use this pattern for public information grounded in approved content. Architecture: Separate retrieval, scope detection, model access, citations, and authenticated casework so model changes do not redefine service authority. Prerequisites: Maintained sources, representative local questions, an answer-quality rubric, and functioning human support. Constraints: Topic coverage, ambiguity, latency, and accessibility require local assessment; the archived Bedrock/Anthropic implementation is an example rather than a requirement. Security: Keep personal records outside public chat and test injected content, jailbreaks, and accidental sensitive input. Proposed validation: Independently score answers by topic, inspect citation support and refusals, and compare navigation and handoff performance with the existing service before expansion.","delivery":"Interpretation — Work and dependencies: Establish curated content, a question benchmark, escalation procedures, and regression tests before a resident pilot. Ownership: The service owner accepts outcomes; content stewards maintain sources, security reviews boundaries, and contact-center staff own escalations. Skills and adoption: Train editors and support agents to distinguish unanswered, ambiguous, and out-of-scope requests; involve residents using assistive technology. Governance checkpoints: Approve scope and data handling before launch and re-evaluate model or retrieval changes. Proposed acceptance: Agreed topic-level quality, supported source links, accessible task completion, and measured escalation response must be demonstrated against local baselines. Risks: Stale guidance, smaller satisfaction samples, and remaining answer errors make headline accuracy unsuitable as an automatic release threshold."},"retrievedAt":null,"enrichedAt":"2026-09-05T02:42:45.193Z","enrichmentBasis":"archived evidence"}},{"version":"enrichment/2026-09-05T02:33:27.019Z/govuk-chat-pilots","resource":{"id":"govuk-chat-pilots","title":"A public assistant improved accuracy through repeated live testing","organization":"GOV.UK","sector":"Citizen information services","geography":"United Kingdom","publishedAt":"March 16, 2026","publicationDate":"2026-03-16","eventDate":null,"sourceName":"5 things we learned testing GOV.UK Chat","sourceLabel":"Inside GOV.UK pilot lessons","sourceUrl":"https://insidegovuk.blog.gov.uk/2026/03/16/5-things-we-learned-testing-gov-uk-chat-an-ai-assistant-for-government/","evidenceClass":"government-evaluation","outcomeClass":"effective","topics":["knowledge-work","infrastructure","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"Two GOV.UK Chat pilots used real questions, benchmark testing, and adversarial probes to improve a retrieval-based public information assistant.","sledRelevance":"Public-facing assistants need a service-standard mindset: authoritative content retrieval, scoped questions, accessibility, safety testing, and ongoing quality measurement.","evidence":"More than 10,000 users asked 26,000 questions; 73% rated the service useful and 64% were satisfied. Benchmark accuracy rose from 76% to 90%, 508 jailbreak attempts were blocked, the in-scope answer rate was 88%, and average response time was 10.7-second latency.","architectureImplications":"Use source-bounded retrieval, an evaluation harness, latency monitoring, failure logging, and a service path that can degrade safely to authoritative content or human help.","governanceImplications":"Publish service measures, maintain escalation and content ownership, and require benchmark and red-team evidence before each material release.","securityPrivacyImplications":"Log and analyze abuse without retaining unnecessary personal data; isolate retrieval sources and test jailbreak, prompt-injection, and disclosure risks.","caveats":"The results reflect a bounded pilot and benchmark; they do not establish accuracy for every topic or user circumstance.","streamIds":["state-government","local-government"],"roles":{"sales":"Interpretation — Customer problem: residents need usable answers from authoritative government content without an unreliable new service channel. Stakeholders: digital-service and contact-center owners, content publishers, accessibility, privacy, and security teams. Discovery: which questions have stable official answers; who owns updates; what falls outside scope; and how do residents reach human help? Value hypothesis: a bounded retrieval assistant may improve information access if quality and fallback are demonstrated locally. Potential engagement: content readiness, question-set design, and a limited public-service pilot. The GOV.UK testing trajectory provides a useful operating example. Unsupported claims: its 90% benchmark accuracy, satisfaction results, and blocked jailbreak count do not promise universal accuracy, equivalent local performance, or comprehensive attack resistance.","engineering":"Interpretation — Fit: use retrieval over a controlled official-content collection for scoped information questions. Architecture and integration: retain source links, versioned benchmarks, failure logging, latency monitoring, and a fallback to official pages or the service desk. Prerequisites: accountable content owners, representative questions, expected-answer scoring, accessibility requirements, and an escalation channel. Constraints: benchmark coverage and changing content limit the meaning of aggregate accuracy; the reported latency also warrants local usability testing. Security: isolate approved retrieval sources, minimize personal data in logs, and test prompt injection, jailbreaks, disclosure, and abusive use. Proposed proof: reproduce answer-quality, in-scope handling, latency, and fallback tests on the agency's own questions, including accessible interaction and out-of-scope cases.","delivery":"Interpretation — Work: prepare the content collection, establish a benchmark, integrate escalation, launch a bounded pilot, and review failures and user feedback. Dependencies: current authoritative content, service-desk capacity, and an accessible interface. Ownership: content publishers maintain answers; the service owner accepts performance; engineering operates monitoring and security handles abuse. Skills and adoption: train support staff to explain scope and redirect residents appropriately. Governance checkpoints: initial safety/accessibility review and repeat benchmark/adversarial tests before material releases. Proposed acceptance: meet locally agreed quality and latency thresholds, demonstrate working fallback for all designated unsupported scenarios, and resolve identified accessibility blockers. Risks include stale source content, unrepresentative questions, and presenting pilot accuracy as a guarantee for every resident circumstance."},"retrievedAt":null,"enrichedAt":"2026-09-05T02:33:27.019Z","enrichmentBasis":"archived evidence"}}]}