{"resourceId":"utah-tax-rag-comparative-pilot-2025","versions":[{"version":"external-7d164862411ee663d8ba1aeb48feec09c9e096dca0c4c1c174b588df05e25bc3","resource":{"id":"utah-tax-rag-comparative-pilot-2025","title":"Utah tax pilot improves benchmark scores, with generalization and production limits","organization":"Utah Division of Technology Services and Utah State Tax Commission","sector":"State tax administration and shared technology services","geography":"Utah, United States","publishedAt":"2025 award submission; exact publication day unverified","publicationDate":null,"eventDate":null,"sourceName":"Utah State Tax Commission AI Pilot","sourceLabel":"State-authored NASCIO award submission","sourceUrl":"https://www.nascio.org/wp-content/uploads/2025/09/UT_Artificial-Intelligence.pdf","evidenceClass":"government-evaluation","outcomeClass":"mixed","topics":["knowledge-work","infrastructure","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"The state reports improved RAG answer scores after platform tuning; this is an operator evaluation, not proof of live service improvement.","sledRelevance":"Direct state tax-assistance evidence, newly added as a historical procurement and validation case.","evidence":"July 2024–February 2025 project: four vendor platforms; 366 initial questions; expert rubric. Scores of 3 or 4 rose from 73% to 97%; top scores rose from 61% to 83%. Phase II reused lower-scoring questions, and one platform's testing stopped.","architectureImplications":"Reported progression from vendor environments to the DTS stack encountered provisioning delays. Interpretation: validate the intended hosting configuration before comparing options.","governanceImplications":"Interpretation: distinguish tuning success from acceptance on an untouched test set.","securityPrivacyImplications":"Interpretation: public grounding documents do not establish protection for sensitive production prompts; test identity, logging and leakage separately.","caveats":"No independent evaluation, held-out sample size or measured call-handling benefit established. Production was in progress when written; present status is unknown. Scores concern rubric categories, not universal accuracy. Chart text was inspected; screenshot yielded no inspectable image.","streamIds":["state-government"],"roles":{"sales":"Interpretation: Explore the tax service owner's actual bottleneck with call-center managers, content specialists, shared IT, finance and procurement. Ask whether staff lose time locating guidance, checking answers or correcting outdated material. A bounded comparative evaluation could establish which workflow deserves investment. The value hypothesis is more reliable access to approved guidance, subject to local testing. Offer a test plan and vendor comparison with a documented stop decision. Do not promise the published benchmark gain, lower staffing needs or a production return. Ask who can supply representative questions and release expert reviewers without impairing service.","engineering":"Interpretation: Fit is internal assistance over controlled tax content. Map the knowledge source, retrieval index, staff interface and approved hosting boundary before integration. Prerequisites include versioned authoritative documents, expert reviewers and a held-out question set. Keep tuning questions separate from acceptance cases, test missing or contradictory guidance, and record abstention and citation correctness. Compare full human verification time with the current search process. Validate permission boundaries and prompt retention independently of factual accuracy. Test deployment in the intended environment before commitment; no cloud-versus-on-premises performance advantage or autonomous-agent authority is established here.","delivery":"Interpretation: Assign tax operations the acceptance decision, content experts the knowledge refresh process and shared IT the runtime. First baseline the current search workflow, then configure retrieval, train reviewers and run a controlled pilot. Dependencies include specialist time, hosting readiness and support funding. Proposed acceptance criteria: every sampled answer has a review outcome, all critical misleading-answer failures are resolved, and quality plus verification time meet pre-agreed thresholds on untouched cases. Require documented approval before expansion and repeat tests after content or model changes. Risks include test-set overfitting, stale guidance and hidden support effort."},"retrievedAt":"2026-09-09T03:01:07Z","enrichedAt":"2026-09-09T03:01:41Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: measure staff verification effort and assistive-technology usability alongside answer quality.","procurementImplications":"Interpretation: require comparable scope, reproducible tests, data portability and full support costs before vendor selection.","operatingModelImplications":"Interpretation: fund tax-content ownership and routine regression testing, not only platform setup.","updateExplanation":"No matching URL or related tax-pilot record found across the 119-resource archive and targeted search. Historical source adds comparative procurement and test-set limitations absent from the latest edition; no new September event claimed.","sourceVerification":{"openedUrl":"https://www.nascio.org/wp-content/uploads/2025/09/UT_Artificial-Intelligence.pdf","referenceExcerpt":"ROI was not a primary element of the project scope.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}