Lighthouse AdvisorySLED AI Adoption Intelligence
← Back to results

From the Local Government edition of September 11, 2026

Academic researchEmergingNewly relevant · Mar 2025

Municipal budget chatbot prototype reports better answers, with incomplete validation

Jerry Xu, Justin Wang, Joley Leung and Jasmine Gu, Lexington High School · Municipal finance and resident information · Lexington, Massachusetts, United States

Publisher
arXiv
Original publication
March 30, 2025; arXiv version 1
Source retrieved
2026-09-12
Read original source

What happened

The authors report improved budget-answer accuracy from a domain-specific retrieval and agent workflow; resident benefit remains unmeasured.

Why it matters

New-to-archive historical evidence for municipal budget access. Author school affiliation does not make this a K12 use case.

Evidence and measured results

The abstract reports 78% accurate responses versus 60% for GPT-4o and 35% for Gemini. Queries were manually checked against official budget documents. The inspected text does not state the query count, independent grading or uncertainty estimates; the results figure could not be retrieved. No resident time-saving or comprehension trial is reported.

Limitations and uncertainty

Historical author evaluation, not a verified municipal production deployment. Figure access failed despite HTML and PDF-text inspection. Current model comparisons, generalizability, costs and public accessibility remain unestablished.

Put this evidence to work

Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-12; this does not change the original publication date. Labels below come from the analysis itself.

Sales

Role takeaway

Start with the finance director, communications lead, clerk and resident-service manager. Ask which recurring budget questions are difficult, who validates fiscal-year distinctions, and how residents currently obtain corrections. A bounded engagement could test a public-document assistant against the existing website and staff-assisted service. The value hypothesis is easier access to verifiable answers; this prototype does not support guaranteed savings, head-count reductions or increased public trust. Establish demand before proposing a purchase. For a smaller municipality, identify the staff time needed to maintain the corpus and answer escalations. Include residents who use assistive technology and those who prefer a non-chat channel in discovery.

Pre-sales engineering

Role takeaway

Build a read-only proof of value using approved budgets and a separately maintained answer set. Evaluate retrieval, calculations, fiscal-year selection, projected-versus-actual distinctions and multi-turn ambiguity independently. Show the exact supporting passage and provide a fallback when evidence is missing. The prerequisite is a versioned corpus with a finance owner and representative resident questions. Compare the assisted workflow with ordinary document search, using current models rather than assuming the historical ranking persists. Test prompt injection, log retention and tool access. Measure error severity, response latency and cost, and assess whether finance reviewers can detect incorrect answers. Neither code availability nor citation presence is sufficient acceptance evidence.

Delivery

Role takeaway

Assign finance as content owner and IT as service operator. Prepare documents, construct the evaluation set, train service staff and define a correction route before a limited pilot. Dependencies include document accessibility, fiscal-year metadata, reviewer availability and maintenance funding. Proposed acceptance criteria are correct year and value on every critical test, accurate citations on the agreed sample, a demonstrated escalation path and no material accessibility regression against the existing service. These are proposed criteria, not study results. Track net staff effort and resident task completion alongside adoption. Pause publication of unsupported answers, and rerun checks after budget revisions, prompt changes or model upgrades.

Implementation considerations

Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.

Architecture and integration

What must connect, and where does the AI sit in the workflow?

The described prototype combines LangChain, OpenAI embeddings, ChromaDB, ReAct tools and page references. Interpretation: Test fiscal-year metadata and arithmetic separately; choose hosting only after reviewing data flows and operating cost.

Governance

Who approves, reviews and stays accountable for outcomes?

Finance staff should approve the answer corpus and define when uncertainty requires a human response.

Security and privacy

What data, permissions and controls need testing?

Restrict the initial corpus to approved public documents, redact user logs and test malicious document instructions and excessive tool permissions.

Accessibility and workforce

Who is affected, and what skills or accommodations follow?

Test screen-reader navigation and plain-language comprehension; preserve phone and document access. No disability-outcome evidence was established.

Procurement

What should contracts, pricing and exit terms secure?

Require exportable documents, prompts and evaluations, transparent usage costs and a maintainable exit plan.

Operating model

Which teams own the service once it runs?

Budget ownership, corpus refresh and resident correction handling require funded staff responsibilities.

What changed

New to the full 217-resource archive checked at offsets 0, 100 and 200. Historical evidence newly added for budget-assistant validation; no source update or post-last-run event is claimed.

Publication history

  1. 2026-09-11Local Government · Issue 063 resources
Read preserved resource versions (JSON)

Stable resource ID: grasp-lexington-budget-chatbot-prototype-2025