Lighthouse AdvisorySLED AI Adoption Intelligence
← Back to results

From the State Government edition of September 8, 2026

Government evaluationMixedUndated source

Utah tax pilot improves benchmark scores, with generalization and production limits

Utah Division of Technology Services and Utah State Tax Commission · State tax administration and shared technology services · Utah, United States

Publisher
Utah State Tax Commission AI Pilot
Original publication
2025 award submission; exact publication day unverified
Source retrieved
2026-09-09
Read original source

What happened

The state reports improved RAG answer scores after platform tuning; this is an operator evaluation, not proof of live service improvement.

Why it matters

Direct state tax-assistance evidence, newly added as a historical procurement and validation case.

Evidence and measured results

July 2024–February 2025 project: four vendor platforms; 366 initial questions; expert rubric. Scores of 3 or 4 rose from 73% to 97%; top scores rose from 61% to 83%. Phase II reused lower-scoring questions, and one platform's testing stopped.

Limitations and uncertainty

No independent evaluation, held-out sample size or measured call-handling benefit established. Production was in progress when written; present status is unknown. Scores concern rubric categories, not universal accuracy. Chart text was inspected; screenshot yielded no inspectable image.

Put this evidence to work

Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-09; this does not change the original publication date. Labels below come from the analysis itself.

Sales

Role takeaway

Explore the tax service owner's actual bottleneck with call-center managers, content specialists, shared IT, finance and procurement. Ask whether staff lose time locating guidance, checking answers or correcting outdated material. A bounded comparative evaluation could establish which workflow deserves investment. The value hypothesis is more reliable access to approved guidance, subject to local testing. Offer a test plan and vendor comparison with a documented stop decision. Do not promise the published benchmark gain, lower staffing needs or a production return. Ask who can supply representative questions and release expert reviewers without impairing service.

Pre-sales engineering

Role takeaway

Fit is internal assistance over controlled tax content. Map the knowledge source, retrieval index, staff interface and approved hosting boundary before integration. Prerequisites include versioned authoritative documents, expert reviewers and a held-out question set. Keep tuning questions separate from acceptance cases, test missing or contradictory guidance, and record abstention and citation correctness. Compare full human verification time with the current search process. Validate permission boundaries and prompt retention independently of factual accuracy. Test deployment in the intended environment before commitment; no cloud-versus-on-premises performance advantage or autonomous-agent authority is established here.

Delivery

Role takeaway

Assign tax operations the acceptance decision, content experts the knowledge refresh process and shared IT the runtime. First baseline the current search workflow, then configure retrieval, train reviewers and run a controlled pilot. Dependencies include specialist time, hosting readiness and support funding.

Proposed acceptance criteria
every sampled answer has a review outcome, all critical misleading-answer failures are resolved, and quality plus verification time meet pre-agreed thresholds on untouched cases. Require documented approval before expansion and repeat tests after content or model changes. Risks include test-set overfitting, stale guidance and hidden support effort.

Implementation considerations

Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.

Architecture and integration

What must connect, and where does the AI sit in the workflow?

Reported progression from vendor environments to the DTS stack encountered provisioning delays. Interpretation: validate the intended hosting configuration before comparing options.

Governance

Who approves, reviews and stays accountable for outcomes?

Distinguish tuning success from acceptance on an untouched test set.

Security and privacy

What data, permissions and controls need testing?

Public grounding documents do not establish protection for sensitive production prompts; test identity, logging and leakage separately.

Accessibility and workforce

Who is affected, and what skills or accommodations follow?

Measure staff verification effort and assistive-technology usability alongside answer quality.

Procurement

What should contracts, pricing and exit terms secure?

Require comparable scope, reproducible tests, data portability and full support costs before vendor selection.

Operating model

Which teams own the service once it runs?

Fund tax-content ownership and routine regression testing, not only platform setup.

What changed

No matching URL or related tax-pilot record found across the 119-resource archive and targeted search. Historical source adds comparative procurement and test-set limitations absent from the latest edition; no new September event claimed.

Publication history

  1. 2026-09-08State Government · Issue 034 resources
Read preserved resource versions (JSON)

Stable resource ID: utah-tax-rag-comparative-pilot-2025