From the State Government edition of September 8, 2026
Utah tax pilot improves benchmark scores, with generalization and production limits
Utah Division of Technology Services and Utah State Tax Commission · State tax administration and shared technology services · Utah, United States
- Publisher
- Utah State Tax Commission AI Pilot
- Original publication
- 2025 award submission; exact publication day unverified
- Source retrieved
- 2026-09-09
What happened
The state reports improved RAG answer scores after platform tuning; this is an operator evaluation, not proof of live service improvement.
Why it matters
Direct state tax-assistance evidence, newly added as a historical procurement and validation case.
Evidence and measured results
July 2024–February 2025 project: four vendor platforms; 366 initial questions; expert rubric. Scores of 3 or 4 rose from 73% to 97%; top scores rose from 61% to 83%. Phase II reused lower-scoring questions, and one platform's testing stopped.
Limitations and uncertainty
No independent evaluation, held-out sample size or measured call-handling benefit established. Production was in progress when written; present status is unknown. Scores concern rubric categories, not universal accuracy. Chart text was inspected; screenshot yielded no inspectable image.
Put this evidence to work
Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-09; this does not change the original publication date. Labels below come from the analysis itself.
Sales
Role takeaway
Explore the tax service owner's actual bottleneck with call-center managers, content specialists, shared IT, finance and procurement. Ask whether staff lose time locating guidance, checking answers or correcting outdated material. A bounded comparative evaluation could establish which workflow deserves investment. The value hypothesis is more reliable access to approved guidance, subject to local testing. Offer a test plan and vendor comparison with a documented stop decision. Do not promise the published benchmark gain, lower staffing needs or a production return. Ask who can supply representative questions and release expert reviewers without impairing service.
Pre-sales engineering
Role takeaway
Fit is internal assistance over controlled tax content. Map the knowledge source, retrieval index, staff interface and approved hosting boundary before integration. Prerequisites include versioned authoritative documents, expert reviewers and a held-out question set. Keep tuning questions separate from acceptance cases, test missing or contradictory guidance, and record abstention and citation correctness. Compare full human verification time with the current search process. Validate permission boundaries and prompt retention independently of factual accuracy. Test deployment in the intended environment before commitment; no cloud-versus-on-premises performance advantage or autonomous-agent authority is established here.
Delivery
Role takeaway
Assign tax operations the acceptance decision, content experts the knowledge refresh process and shared IT the runtime. First baseline the current search workflow, then configure retrieval, train reviewers and run a controlled pilot. Dependencies include specialist time, hosting readiness and support funding.
- Proposed acceptance criteria
- every sampled answer has a review outcome, all critical misleading-answer failures are resolved, and quality plus verification time meet pre-agreed thresholds on untouched cases. Require documented approval before expansion and repeat tests after content or model changes. Risks include test-set overfitting, stale guidance and hidden support effort.
Implementation considerations
Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.
Architecture and integration
What must connect, and where does the AI sit in the workflow?
Reported progression from vendor environments to the DTS stack encountered provisioning delays. Interpretation: validate the intended hosting configuration before comparing options.
Governance
Who approves, reviews and stays accountable for outcomes?
Distinguish tuning success from acceptance on an untouched test set.
Security and privacy
What data, permissions and controls need testing?
Public grounding documents do not establish protection for sensitive production prompts; test identity, logging and leakage separately.
Accessibility and workforce
Who is affected, and what skills or accommodations follow?
Measure staff verification effort and assistive-technology usability alongside answer quality.
Procurement
What should contracts, pricing and exit terms secure?
Require comparable scope, reproducible tests, data portability and full support costs before vendor selection.
Operating model
Which teams own the service once it runs?
Fund tax-content ownership and routine regression testing, not only platform setup.
What changed
No matching URL or related tax-pilot record found across the 119-resource archive and targeted search. Historical source adds comparative procurement and test-set limitations absent from the latest edition; no new September event claimed.
Publication history
- 2026-09-08State Government · Issue 034 resources
Stable resource ID: utah-tax-rag-comparative-pilot-2025