From the SLED-wide archive edition of August 27, 2026
City chatbot audit finds inconsistent answers and weak test evidence
New York City Comptroller · Municipal citizen services · New York City, United States
- Publisher
- Audit Report on the New York City Office of Technology and Innovation’s MyCity System
- Original publication
- December 30, 2025
- Source retrieved
- Not recorded in the historical archive
What happened
An independent city audit identified inaccurate or inconsistent chatbot responses, performance delays, and insufficient detail behind reported accuracy claims.
Why it matters
Public assistants create direct service and trust risk when answers are authoritative in tone but weakly tested or inconsistently documented.
Evidence and measured results
Auditors tested system behavior and requested structured red-teaming. The Office of Technology and Innovation reported 95–99% accuracy but supplied insufficient test detail and disagreed with recommendations.
Limitations and uncertainty
The agency contested parts of the audit; readers should review both the findings and the response in the source.
Put this evidence to work
Lighthouse Advisory interpretation, grounded in this source as summarized in the preserved archive. Enriched 2026-09-05; this does not change the original publication date. Labels below come from the analysis itself.
Sales
Role takeaway
- Customer problem
- a municipal assistant can create service and trust risk when accuracy claims cannot be reproduced.
- Stakeholders
- digital-service leadership, content and contact-center owners, technology, accessibility, privacy, and independent oversight.
- Discovery
- who scores accuracy; can another reviewer reproduce the result; which inconsistent answers matter most; and how are errors corrected?
- Value hypothesis
- independent evaluation and clear fallback may produce a more defensible service-quality case.
- Potential engagement
- review the current benchmark, reproduce representative questions, and close assurance gaps.
- Unsupported claims
- the reported 95–99% accuracy is a disputed agency claim in this record, not an established performance result. The contested audit also does not justify assuming every public chatbot has the same defects.
Pre-sales engineering
Role takeaway
- Fit
- assess an existing or planned public information assistant with reproducible service tests.
- Architecture and integration
- use authoritative-source retrieval, versioned questions and expected answers, performance monitoring, failure capture, and fallback into existing service channels.
- Prerequisites
- benchmark ownership, scoring rules, content access, and records of model and configuration versions.
- Constraints
- repeated prompts may yield inconsistent answers, and opaque scoring prevents meaningful comparison.
- Security
- test prompt injection, disclosure, abusive inputs, and log retention as part of the evaluation. Proposed proof: rerun representative prompts across repeated trials, independently score correctness and consistency, measure delay, and verify fallback. Preserve both observed results and the operating team's response so disagreement is visible rather than averaged away.
Delivery
Role takeaway
- Work
- establish the evaluation record, assign content owners, introduce public error reporting, and integrate recurring independent assurance into release decisions.
- Dependencies
- access to prompts, scoring, versions, and service logs, plus owner cooperation when findings are disputed.
- Ownership
- the service owner accepts residual risk; content teams correct answers; engineering monitors behavior; an independent reviewer assesses claims.
- Skills and adoption
- train support staff in escalation and evaluators in reproducible scoring.
- Governance checkpoints
- preproduction review, material release, and follow-up on unresolved findings.
- Proposed acceptance
- independent reviewers can reproduce the reported customer benchmark, known severe errors are resolved or safely redirected, and performance/accessibility limits are documented. Risks include defensive reporting, selective test sets, and unresolved disagreements obscuring resident impact.
Implementation considerations
Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.
Architecture and integration
What must connect, and where does the AI sit in the workflow?
Build reproducible evaluation sets, performance monitoring, authoritative-source retrieval, failure capture, and a fallback path into the service architecture.
Governance
Who approves, reviews and stays accountable for outcomes?
Require public error reporting, red-team protocols, content ownership, release gates, and independent assurance for high-impact performance claims.
Security and privacy
What data, permissions and controls need testing?
Test prompt injection, data disclosure, abusive use, and log handling as part of an auditable preproduction and recurring assurance program.
The preserved archive analysis covered architecture, governance and security. Not assessed for this record: accessibility and workforce, procurement, operating model.
Publication history
- 2026-08-27SLED-wide archive · Issue 0110 resources
Stable resource ID: nyc-mycity-audit