{"resourceId":"nyc-mycity-audit","versions":[{"version":"legacy/2026-08-27/nyc-mycity-audit","resource":{"id":"nyc-mycity-audit","title":"City chatbot audit finds inconsistent answers and weak test evidence","organization":"New York City Comptroller","sector":"Municipal citizen services","geography":"New York City, United States","publishedAt":"December 30, 2025","sourceName":"Audit Report on the New York City Office of Technology and Innovation’s MyCity System","sourceLabel":"New York City Comptroller audit","sourceUrl":"https://comptroller.nyc.gov/reports/audit-report-on-the-new-york-city-office-of-technology-and-innovations-mycity-system/","evidenceClass":"government-audit","outcomeClass":"cautionary","topics":["infrastructure","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"An independent city audit identified inaccurate or inconsistent chatbot responses, performance delays, and insufficient detail behind reported accuracy claims.","sledRelevance":"Public assistants create direct service and trust risk when answers are authoritative in tone but weakly tested or inconsistently documented.","evidence":"Auditors tested system behavior and requested structured red-teaming. The Office of Technology and Innovation reported 95–99% accuracy but supplied insufficient test detail and disagreed with recommendations.","architectureImplications":"Build reproducible evaluation sets, performance monitoring, authoritative-source retrieval, failure capture, and a fallback path into the service architecture.","governanceImplications":"Require public error reporting, red-team protocols, content ownership, release gates, and independent assurance for high-impact performance claims.","securityPrivacyImplications":"Test prompt injection, data disclosure, abusive use, and log handling as part of an auditable preproduction and recurring assurance program.","caveats":"The agency contested parts of the audit; readers should review both the findings and the response in the source."}},{"version":"enrichment/2026-09-05T02:33:27.019Z/nyc-mycity-audit","resource":{"id":"nyc-mycity-audit","title":"City chatbot audit finds inconsistent answers and weak test evidence","organization":"New York City Comptroller","sector":"Municipal citizen services","geography":"New York City, United States","publishedAt":"December 30, 2025","publicationDate":"2025-12-30","eventDate":null,"sourceName":"Audit Report on the New York City Office of Technology and Innovation’s MyCity System","sourceLabel":"New York City Comptroller audit","sourceUrl":"https://comptroller.nyc.gov/reports/audit-report-on-the-new-york-city-office-of-technology-and-innovations-mycity-system/","evidenceClass":"government-audit","outcomeClass":"cautionary","topics":["infrastructure","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"An independent city audit identified inaccurate or inconsistent chatbot responses, performance delays, and insufficient detail behind reported accuracy claims.","sledRelevance":"Public assistants create direct service and trust risk when answers are authoritative in tone but weakly tested or inconsistently documented.","evidence":"Auditors tested system behavior and requested structured red-teaming. The Office of Technology and Innovation reported 95–99% accuracy but supplied insufficient test detail and disagreed with recommendations.","architectureImplications":"Build reproducible evaluation sets, performance monitoring, authoritative-source retrieval, failure capture, and a fallback path into the service architecture.","governanceImplications":"Require public error reporting, red-team protocols, content ownership, release gates, and independent assurance for high-impact performance claims.","securityPrivacyImplications":"Test prompt injection, data disclosure, abusive use, and log handling as part of an auditable preproduction and recurring assurance program.","caveats":"The agency contested parts of the audit; readers should review both the findings and the response in the source.","streamIds":["local-government"],"roles":{"sales":"Interpretation — Customer problem: a municipal assistant can create service and trust risk when accuracy claims cannot be reproduced. Stakeholders: digital-service leadership, content and contact-center owners, technology, accessibility, privacy, and independent oversight. Discovery: who scores accuracy; can another reviewer reproduce the result; which inconsistent answers matter most; and how are errors corrected? Value hypothesis: independent evaluation and clear fallback may produce a more defensible service-quality case. Potential engagement: review the current benchmark, reproduce representative questions, and close assurance gaps. Unsupported claims: the reported 95–99% accuracy is a disputed agency claim in this record, not an established performance result. The contested audit also does not justify assuming every public chatbot has the same defects.","engineering":"Interpretation — Fit: assess an existing or planned public information assistant with reproducible service tests. Architecture and integration: use authoritative-source retrieval, versioned questions and expected answers, performance monitoring, failure capture, and fallback into existing service channels. Prerequisites: benchmark ownership, scoring rules, content access, and records of model and configuration versions. Constraints: repeated prompts may yield inconsistent answers, and opaque scoring prevents meaningful comparison. Security: test prompt injection, disclosure, abusive inputs, and log retention as part of the evaluation. Proposed proof: rerun representative prompts across repeated trials, independently score correctness and consistency, measure delay, and verify fallback. Preserve both observed results and the operating team's response so disagreement is visible rather than averaged away.","delivery":"Interpretation — Work: establish the evaluation record, assign content owners, introduce public error reporting, and integrate recurring independent assurance into release decisions. Dependencies: access to prompts, scoring, versions, and service logs, plus owner cooperation when findings are disputed. Ownership: the service owner accepts residual risk; content teams correct answers; engineering monitors behavior; an independent reviewer assesses claims. Skills and adoption: train support staff in escalation and evaluators in reproducible scoring. Governance checkpoints: preproduction review, material release, and follow-up on unresolved findings. Proposed acceptance: independent reviewers can reproduce the reported customer benchmark, known severe errors are resolved or safely redirected, and performance/accessibility limits are documented. Risks include defensive reporting, selective test sets, and unresolved disagreements obscuring resident impact."},"retrievedAt":null,"enrichedAt":"2026-09-05T02:33:27.019Z","enrichmentBasis":"archived evidence"}}]}