{"resourceId":"urban-minneapolis-zoning-retrieval-benchmark-2026","versions":[{"version":"external-6bd7edfd816e6ea2ce78297fb20acbf5780fcabb574aa5ca70f406c971b1014a","resource":{"id":"urban-minneapolis-zoning-retrieval-benchmark-2026","title":"Zoning benchmark exposes retrieval and useful-answer limitations","organization":"Urban Institute","sector":"Local government","geography":"Minneapolis, Minnesota, United States","publishedAt":"2026-03-19","publicationDate":"2026-03-19","eventDate":null,"sourceName":"Urban Institute","sourceLabel":"Zoning benchmark exposes retrieval and useful-answer limitations","sourceUrl":"https://www.urban.org/urban-wire/how-can-local-governments-use-ai-answer-community-members-questions-about-zoning-and","evidenceClass":"independent-research","outcomeClass":"cautionary","topics":["knowledge-work","developers-agents","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"A local-code exercise found poor retrieval and unhelpful answers despite customization.","sledRelevance":"Historical evidence newly added for municipal resident-information testing; not a September outcome.","evidence":"Researchers used developer and homeowner personas, manual expert review and five evaluation dimensions. Abstention reduced fabrication but did not ensure usefulness. No service-time baseline or causal deployment result is reported.","architectureImplications":"Interpretation: Test retrieval separately from answer generation; preserve ordinance versions and citations. Compare current approaches on local questions before choosing hosting or agents.","governanceImplications":"Interpretation: Require planner approval of the evaluation rubric and a maintained route for disputed answers.","securityPrivacyImplications":"Interpretation: Keep public-code testing separate from private application files; threat-test instructions embedded in retrieved documents.","caveats":"One-city exercise; article lacks aggregate scores and a full sample count. Linked benchmark access failed and workbook contents were not inspected. Findings are model- and configuration-specific.","streamIds":["local-government"],"roles":{"sales":"Interpretation: Ask the planning director, resident-service manager and CIO whether repeated zoning questions or misunderstood requirements consume material staff effort. What can residents resolve unaided today, and which questions require professional judgment? Offer a bounded discovery and benchmark engagement on one ordinance area. The value hypothesis is fewer avoidable handoffs with reliable answers, subject to local measurement. Include community representatives when defining useful responses. Do not sell a general model as a substitute for planning advice or promise fewer housing delays. Establish the baseline and escalation workload before estimating value; a technically correct refusal can still leave the resident's problem unresolved.","engineering":"Interpretation: Start with a read-only assistant over approved public ordinances. Map document extraction, section relationships, retrieval, generation and citation display; verify version synchronization with the authoritative source. Prerequisites include a planner-reviewed question set and expected supporting passages. Test multi-section questions, obsolete rules, unsupported assumptions and malicious document instructions. Compare retrieval recall, answer correctness and useful abstention independently. Current cloud, on-premises and hybrid options require local cost and privacy assessment; this exercise does not rank them. Keep agents from submitting applications during validation. Use blinded expert scoring and repeated runs before allowing resident access.","delivery":"Interpretation: Assign a planning-service owner and a technical indexing owner, then build the test corpus with frontline staff. Budget ordinance cleanup, accessible user testing and correction handling before rollout. Train staff to inspect citations and preserve conventional advice channels. Proposed acceptance criteria are improvement in resident task completion, no deterioration in expert-scored accuracy and successful refresh tests after a code change; set numerical thresholds before testing. These are future criteria, not observed results. Review misses and refusals weekly during the pilot. Stop expansion if maintenance or escalation demand exceeds funded capacity, and require renewed approval after retrieval or model changes."},"retrievedAt":"2026-09-10T03:03:45Z","enrichedAt":"2026-09-10T03:04:50Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Include nonexpert residents and assistive-technology users in proposed testing; measure successful task completion as well as correctness.","procurementImplications":"Interpretation: Ask bidders to demonstrate locally selected questions and export evidence; include code-update maintenance in pricing.","operatingModelImplications":"Interpretation: Planning owns authoritative answers; IT maintains indexing; contact staff handle unresolved questions.","updateExplanation":"New to the full archive; no source update claimed.","sourceVerification":{"openedUrl":"https://www.urban.org/urban-wire/how-can-local-governments-use-ai-answer-community-members-questions-about-zoning-and","referenceExcerpt":"AI tools often provided unhelpful answers.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}