{"resourceId":"grasp-lexington-budget-chatbot-prototype-2025","versions":[{"version":"external-8b5577b56104e82aec12694c1f2e431c0e395a67192ceeb15d99f685e766fac7","resource":{"id":"grasp-lexington-budget-chatbot-prototype-2025","title":"Municipal budget chatbot prototype reports better answers, with incomplete validation","organization":"Jerry Xu, Justin Wang, Joley Leung and Jasmine Gu, Lexington High School","sector":"Municipal finance and resident information","geography":"Lexington, Massachusetts, United States","publishedAt":"March 30, 2025; arXiv version 1","publicationDate":"2025-03-30","eventDate":null,"sourceName":"arXiv","sourceLabel":"Author-evaluated research prototype; historical preprint","sourceUrl":"https://arxiv.org/html/2503.23299v1","evidenceClass":"academic-research","outcomeClass":"emerging","topics":["knowledge-work","developers-agents","data-security","accessibility-workforce","operating-model"],"finding":"The authors report improved budget-answer accuracy from a domain-specific retrieval and agent workflow; resident benefit remains unmeasured.","sledRelevance":"New-to-archive historical evidence for municipal budget access. Author school affiliation does not make this a K12 use case.","evidence":"The abstract reports 78% accurate responses versus 60% for GPT-4o and 35% for Gemini. Queries were manually checked against official budget documents. The inspected text does not state the query count, independent grading or uncertainty estimates; the results figure could not be retrieved. No resident time-saving or comprehension trial is reported.","architectureImplications":"The described prototype combines LangChain, OpenAI embeddings, ChromaDB, ReAct tools and page references. Interpretation: Test fiscal-year metadata and arithmetic separately; choose hosting only after reviewing data flows and operating cost.","governanceImplications":"Interpretation: Finance staff should approve the answer corpus and define when uncertainty requires a human response.","securityPrivacyImplications":"Interpretation: Restrict the initial corpus to approved public documents, redact user logs and test malicious document instructions and excessive tool permissions.","caveats":"Historical author evaluation, not a verified municipal production deployment. Figure access failed despite HTML and PDF-text inspection. Current model comparisons, generalizability, costs and public accessibility remain unestablished.","streamIds":["local-government"],"roles":{"sales":"Interpretation: Start with the finance director, communications lead, clerk and resident-service manager. Ask which recurring budget questions are difficult, who validates fiscal-year distinctions, and how residents currently obtain corrections. A bounded engagement could test a public-document assistant against the existing website and staff-assisted service. The value hypothesis is easier access to verifiable answers; this prototype does not support guaranteed savings, head-count reductions or increased public trust. Establish demand before proposing a purchase. For a smaller municipality, identify the staff time needed to maintain the corpus and answer escalations. Include residents who use assistive technology and those who prefer a non-chat channel in discovery.","engineering":"Interpretation: Build a read-only proof of value using approved budgets and a separately maintained answer set. Evaluate retrieval, calculations, fiscal-year selection, projected-versus-actual distinctions and multi-turn ambiguity independently. Show the exact supporting passage and provide a fallback when evidence is missing. The prerequisite is a versioned corpus with a finance owner and representative resident questions. Compare the assisted workflow with ordinary document search, using current models rather than assuming the historical ranking persists. Test prompt injection, log retention and tool access. Measure error severity, response latency and cost, and assess whether finance reviewers can detect incorrect answers. Neither code availability nor citation presence is sufficient acceptance evidence.","delivery":"Interpretation: Assign finance as content owner and IT as service operator. Prepare documents, construct the evaluation set, train service staff and define a correction route before a limited pilot. Dependencies include document accessibility, fiscal-year metadata, reviewer availability and maintenance funding. Proposed acceptance criteria are correct year and value on every critical test, accurate citations on the agreed sample, a demonstrated escalation path and no material accessibility regression against the existing service. These are proposed criteria, not study results. Track net staff effort and resident task completion alongside adoption. Pause publication of unsupported answers, and rerun checks after budget revisions, prompt changes or model upgrades."},"retrievedAt":"2026-09-12T03:02:15Z","enrichedAt":"2026-09-12T03:05:08Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Test screen-reader navigation and plain-language comprehension; preserve phone and document access. No disability-outcome evidence was established.","procurementImplications":"Interpretation: Require exportable documents, prompts and evaluations, transparent usage costs and a maintainable exit plan.","operatingModelImplications":"Interpretation: Budget ownership, corpus refresh and resident correction handling require funded staff responsibilities.","updateExplanation":"New to the full 217-resource archive checked at offsets 0, 100 and 200. Historical evidence newly added for budget-assistant validation; no source update or post-last-run event is claimed.","sourceVerification":{"openedUrl":"https://arxiv.org/html/2503.23299v1","referenceExcerpt":"The questions, while diverse, may not encompass all potential real-world scenarios, limiting the ability to generalize the findings.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}