{"resourceId":"code-for-america-ai-landscape-2026-impact","versions":[{"version":"external-47d3e3a511883e3a96d296a170171ead006b174a424d8402d58bc51724d65b95","resource":{"id":"code-for-america-ai-landscape-2026-impact","title":"State AI assessment finds impact reporting trails experimentation","organization":"Code for America","sector":"State AI adoption and benefits administration","geography":"United States","publishedAt":"May 2026; exact day not established on inspected assessment","publicationDate":null,"eventDate":null,"sourceName":"Code for America","sourceLabel":"Independent nonprofit desk-research assessment","sourceUrl":"https://codeforamerica.org/explore/government-ai-landscape-assessment/","evidenceClass":"independent-research","outcomeClass":"mixed","topics":["knowledge-work","developers-agents","infrastructure","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"The assessment finds widespread experimentation but limited impact reporting and continuous learning. It distinguishes readiness, piloting, implementation and impact.","sledRelevance":"Direct state-government portfolio evidence with a benefits-access lens. Use it to frame evaluation questions, not to infer any state's present performance.","evidence":"Research ended in March 2026. Methods combine public-document desk research, state feedback opportunities and advisory rubric review. This is a maturity assessment, not a controlled test of AI effectiveness.","architectureImplications":"Interpretation: require reusable evaluation telemetry and inventory identifiers across platforms; do not equate a shared data platform with safe integration.","governanceImplications":"Interpretation: link each use case to outcome measures, review dates and decisions to continue, change or retire.","securityPrivacyImplications":"Interpretation: collect minimal evaluation data and restrict access to benefits records and sensitive employee prompts.","caveats":"Public reporting can miss internal work. Rendered state totals were unavailable; no counts or rankings are asserted. Full PDF was email-gated and not accessed. Findings are not a September status census.","streamIds":["state-government"],"roles":{"sales":"Interpretation: Ask agency executives, benefits directors and budget owners what counts as public value and who can supply the baseline. A bounded engagement could inventory selected pilots and design a measurable continuation decision. The value hypothesis is better allocation of implementation effort, not guaranteed savings from adding AI. Ask whether success means faster processing, fewer corrections, more completed applications or improved access. Keep those outcomes separate from license uptake. Do not use maturity language as a competitive ranking or claim a customer has weak performance merely because public reporting is sparse.","engineering":"Interpretation: Establish a common evaluation record containing workflow version, data boundary, human involvement, task outcome and costs. Map connectors and legacy interfaces before selecting a deployment model; this assessment supplies no local capacity specification. Prerequisites include baseline access, representative cases and consent or other approved handling for evaluation data. Test a limited workflow across ordinary and difficult cases, recording both system output and human corrections. Include accessibility failures and unauthorized disclosure attempts. Require reproducible comparisons across model changes. The proposed design addresses evidence gaps without treating a desk-research maturity rubric as technical assurance.","delivery":"Interpretation: Put an agency service owner in charge of recurring evaluation, supported by analysts, frontline staff, accessibility specialists and security. First document the current process, then establish a test cohort and reporting cadence. Dependencies include reliable operational data and time for human review. Train staff to log corrections and route incidents without discouraging honest reporting. Proposed acceptance criteria: each pilot has a baseline, named owner, cost report and explicit continuation threshold, with a completed review before expansion. Risks include measuring only usage, missing excluded residents, inconsistent metrics and losing institutional knowledge when pilot staff leave."},"retrievedAt":"2026-09-07T03:00:54Z","enrichedAt":"2026-09-07T03:03:44Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: measure outcomes for users with disabilities and language needs; include reviewer learning and correction workload.","procurementImplications":"Interpretation: contracts should preserve measurement access, portability and ability to stop an ineffective pilot.","operatingModelImplications":"Interpretation: give service owners recurring responsibility for performance review rather than making evaluation a one-time launch task.","updateExplanation":"New canonical source URL in searched archive. Adds a methodological counterweight to unarchived deployment announcements; older research is explicitly dated and not presented as a fresh event.","sourceVerification":{"openedUrl":"https://codeforamerica.org/explore/government-ai-landscape-assessment/","referenceExcerpt":"This analysis was compiled primarily through public data.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}