{"resourceId":"colorado-ui-fact-finding-randomized-sandbox","versions":[{"version":"external-aab1a3767aa79d68a882b3725be106eb7e199fea1952417b47db77ca4fad9dcb","resource":{"id":"colorado-ui-fact-finding-randomized-sandbox","title":"Colorado benefits trial finds favorable user feedback without an average causal gain","organization":"Magesh, Martin, Surani, Perez, Rodolfa and Ho; collaboration with CDLE and U.S. DOL","sector":"State unemployment insurance administration","geography":"Colorado, United States","publishedAt":"Undated manuscript in Yale's January 2026 repository path; exact publication date unverified","publicationDate":null,"eventDate":null,"sourceName":"Evaluating Generative AI in Benefits Administration: A Demonstration Project","sourceLabel":"Academic manuscript hosted by Yale Tobin Center","sourceUrl":"https://tobin.yale.edu/sites/default/files/2026-01/CDLE_AI_Jan13_2026.pdf","evidenceClass":"academic-research","outcomeClass":"mixed","topics":["knowledge-work","infrastructure","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"AI-assisted fact-finding did not significantly improve average drafting time or quality against concurrent unaided adjudicators.","sledRelevance":"Direct state benefits evidence; historical sandbox tasks do not establish live eligibility or payment outcomes.","evidence":"Randomized crossover study: 8 non-randomly recruited adjudicators, 200 sampled historical claims, 788 drafts, and 6 internal QA reviewers. Average drafting-time reduction was 4 seconds (p=0.8); quality comparison p=0.37. Positive user feedback and favorable historical comparisons did not establish incremental workflow benefit.","architectureImplications":"The prototype used open models inside CDLE's secure cloud sandbox. Interpretation: keep case records authoritative and make generated questions editable before release.","governanceImplications":"Interpretation: require a concurrent workflow comparison and a separate decision on permission to enter live service.","securityPrivacyImplications":"Interpretation: approve training-data reuse, isolate evaluation identities, and minimize sensitive prompt logs.","caveats":"Small selected workforce; historical quit cases; possible sandbox behavior effects. No demonstrated reduction in claimant waiting time. Manuscript publication day and trial dates remain unverified. PDF figure screenshot failed; Figure 4 caption and results text were inspected.","streamIds":["state-government"],"roles":{"sales":"Interpretation: Focus discovery on the agency's actual bottleneck: drafting, missing information, staff review or claimant response delays. Bring the benefits director, QA lead, frontline adjudicators, privacy staff and procurement together. Ask which baseline can be measured and which claimant outcomes matter. A bounded evaluation engagement could compare one assistance workflow with the current process and produce a continuation decision. The value hypothesis is better targeting of improvement effort, potentially including consistency, rather than promised staffing savings. This study's result limits any sales claim about automatic time savings. Do not imply that a successful question generator proves eligibility accuracy or faster benefit payment.","engineering":"Interpretation: Fit is strongest where a discrete drafting step can be instrumented without giving the model adjudication authority. Map the case-management interface, approved compute boundary, access identities and permitted training records before choosing a model. Prerequisites include representative historical cases, staff availability and a quality rubric agreed with program specialists. Build a proof of value that records full human editing time, not just generation latency, and compares blinded output quality with an unaided condition. Test irrelevant questions, sensitive-data leakage and recovery after model changes. The sandbox result warrants later monitored workflow validation; it supplies no production capacity estimate or general cloud-versus-on-premises cost advantage.","delivery":"Interpretation: Start with process mapping, data approval and a protected test environment. Benefits operations owns scope and outcomes; QA owns assessment; platform and security teams own access, deployment and incident handling. Dependencies include adjudicator release time, representative cases and approval of record retention. Train reviewers to reject or rewrite drafts and document why. Proposed acceptance criteria: every test case has timing and quality evidence, all critical disclosure failures are resolved, and a named owner approves the pilot review before live expansion. Risks include overestimating gains from historical baselines, inadequate time for correction and failing to measure additional work imposed on claimants."},"retrievedAt":"2026-09-08T03:02:23Z","enrichedAt":"2026-09-08T03:06:00Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: test question comprehension, language access and reviewer workload; a longer questionnaire can increase claimant burden.","procurementImplications":"Interpretation: condition expansion on locally observed quality and cost evidence, with evaluation access and an exit option.","operatingModelImplications":"Interpretation: assign benefits operations the outcome decision and QA the independent review workflow; model maintenance requires funded ownership.","updateExplanation":"No matching URL or CDLE record in the full 85-resource archive or targeted search. Historical evidence newly added for its concurrent-control counterpoint to prior benefits and procurement coverage; not described as a new September trial.","sourceVerification":{"openedUrl":"https://tobin.yale.edu/sites/default/files/2026-01/CDLE_AI_Jan13_2026.pdf","referenceExcerpt":"a statistically insignificant reduction of only 4 seconds (p=0.8).","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}