{"resourceId":"oecd-government-experimentation","versions":[{"version":"legacy/2026-08-30/oecd-government-experimentation","resource":{"id":"oecd-government-experimentation","title":"Cross-country review finds experimentation widespread but monitoring and evaluation weak","organization":"Organisation for Economic Co-operation and Development","sector":"Public administration","geography":"International; official guidance from 14 countries","publishedAt":"July 20, 2026","sourceName":"Generative AI experimentation in government: Learning from emerging guidelines","sourceLabel":"OECD Working Papers on Public Governance No. 93","sourceUrl":"https://www.oecd.org/en/publications/generative-ai-experimentation-in-government_42815683-en.html","evidenceClass":"standards-guidance","outcomeClass":"emerging","topics":["knowledge-work","developers-agents","infrastructure","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"OECD reviewed official experimentation guidance across 14 countries, academic literature, and case studies. It found rapid decentralized uptake, fragmented guidance, and few governments systematically measuring performance, impact, or compliance.","sledRelevance":"The gap mirrors SLED conditions: employees experiment before governance is mature, smaller organizations face inconsistent rules, and pilots are often counted without demonstrating service, learning, workforce, or compliance outcomes.","evidence":"The working paper synthesizes documented practices rather than testing one intervention. It proposes evaluation across five dimensions—performance, public value, feasibility, usability, and risk management—and identifies structured experimentation as a bridge between principles and scaled operation.","architectureImplications":"Provide segregated sandboxes, approved data paths, reusable evaluation harnesses, model and prompt logging, and a governed route from experiment to production. Instrument quality, cost, latency, accessibility, and risk from the beginning rather than after a pilot is declared successful.","governanceImplications":"Use a common experiment charter with an accountable owner, hypothesis, baseline, success and stop criteria, affected-user review, and evidence package. Allow local experimentation within shared guardrails while central teams provide templates, expertise, procurement, and assurance services.","securityPrivacyImplications":"Match experiment environments to data sensitivity; prohibit uncontrolled sensitive-data use; document model and vendor handling; and perform privacy, security, bias, and misuse testing before expanding access or authority.","caveats":"This is comparative guidance and synthesis, not causal outcome evidence. Official guidelines may differ from actual agency practice, and the review's international scope means legal and administrative assumptions do not transfer uniformly to U.S. SLED organizations."}},{"version":"enrichment/2026-09-05T02:42:45.193Z/oecd-government-experimentation","resource":{"id":"oecd-government-experimentation","title":"Cross-country review finds experimentation widespread but monitoring and evaluation weak","organization":"Organisation for Economic Co-operation and Development","sector":"Public administration","geography":"International; official guidance from 14 countries","publishedAt":"July 20, 2026","publicationDate":"2026-07-20","eventDate":null,"sourceName":"Generative AI experimentation in government: Learning from emerging guidelines","sourceLabel":"OECD Working Papers on Public Governance No. 93","sourceUrl":"https://www.oecd.org/en/publications/generative-ai-experimentation-in-government_42815683-en.html","evidenceClass":"standards-guidance","outcomeClass":"emerging","topics":["knowledge-work","developers-agents","infrastructure","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"OECD reviewed official experimentation guidance across 14 countries, academic literature, and case studies. It found rapid decentralized uptake, fragmented guidance, and few governments systematically measuring performance, impact, or compliance.","sledRelevance":"The gap mirrors SLED conditions: employees experiment before governance is mature, smaller organizations face inconsistent rules, and pilots are often counted without demonstrating service, learning, workforce, or compliance outcomes.","evidence":"The working paper synthesizes documented practices rather than testing one intervention. It proposes evaluation across five dimensions—performance, public value, feasibility, usability, and risk management—and identifies structured experimentation as a bridge between principles and scaled operation.","architectureImplications":"Provide segregated sandboxes, approved data paths, reusable evaluation harnesses, model and prompt logging, and a governed route from experiment to production. Instrument quality, cost, latency, accessibility, and risk from the beginning rather than after a pilot is declared successful.","governanceImplications":"Use a common experiment charter with an accountable owner, hypothesis, baseline, success and stop criteria, affected-user review, and evidence package. Allow local experimentation within shared guardrails while central teams provide templates, expertise, procurement, and assurance services.","securityPrivacyImplications":"Match experiment environments to data sensitivity; prohibit uncontrolled sensitive-data use; document model and vendor handling; and perform privacy, security, bias, and misuse testing before expanding access or authority.","caveats":"This is comparative guidance and synthesis, not causal outcome evidence. Official guidelines may differ from actual agency practice, and the review's international scope means legal and administrative assumptions do not transfer uniformly to U.S. SLED organizations.","streamIds":["state-government","local-government"],"roles":{"sales":"Interpretation — Problem and stakeholders: Innovation teams, CIOs, procurement, risk leaders, and program managers may count pilots without knowing whether they improve public service. Discovery: Does each experiment have an owner, baseline, public-value hypothesis, stop rule, and route to operational funding? Value hypothesis: Shared experiment discipline could make results comparable and prevent unsupported scaling. Potential engagement: Review the pilot portfolio and apply an evidence charter to a small set of varied workflows. Evidence boundary: OECD synthesizes guidance and cases across 14 countries rather than evaluating one controlled intervention. Its five dimensions provide useful questions; they do not prove a sandbox, central team, or governance model will improve outcomes under U.S. SLED conditions.","engineering":"Interpretation — Fit: Use governed experimentation for bounded knowledge-work, copilot, or agent trials before granting production authority. Architecture: Reuse segregated sandboxes, approved data paths, versioned prompts and models, evaluation harnesses, and monitoring. Prerequisites: Defined use cases, data classification, baseline tasks, and agreed performance, public-value, feasibility, usability, and risk measures. Constraints: International administrative assumptions require local translation, and sandbox results may not transfer to production integrations. Security: Limit sensitive data, tool permissions, and network access to approved scope. Proposed validation: Run comparable tasks with and without assistance, record quality, latency, cost, accessibility, and failures, then demonstrate a controlled production handoff or an evidence-based decision to stop.","delivery":"Interpretation — Work and dependencies: Inventory experiments and introduce a charter, evaluation record, and graduation decision through existing portfolio processes. Ownership: Programs own hypotheses and public-service outcomes; central technology, procurement, privacy, and security teams supply reusable controls and assurance. Skills and adoption: Train pilot leads to measure baselines and involve affected users in usability and accessibility review. Governance checkpoints: Approve data and scope before testing, inspect evidence before integration, and confirm funding and support before production. Proposed acceptance: Reviewed pilots have results across the five OECD dimensions, limitations, and reasoned continue, change, or stop decisions. Risks: Fragmented guidance, weak measurement, and unfunded operations can survive a technically successful sandbox; international guidance does not resolve local staffing or legal constraints."},"retrievedAt":null,"enrichedAt":"2026-09-05T02:42:45.193Z","enrichmentBasis":"archived evidence"}}]}