Lighthouse AdvisorySLED AI Adoption Intelligence
← Back to results

From the SLED-wide archive edition of August 27, 2026

Government auditCautionaryPublished · Dec 2025

City chatbot audit finds inconsistent answers and weak test evidence

New York City Comptroller · Municipal citizen services · New York City, United States

Publisher
Audit Report on the New York City Office of Technology and Innovation’s MyCity System
Original publication
December 30, 2025
Source retrieved
Not recorded in the historical archive
Read original source

What happened

An independent city audit identified inaccurate or inconsistent chatbot responses, performance delays, and insufficient detail behind reported accuracy claims.

Why it matters

Public assistants create direct service and trust risk when answers are authoritative in tone but weakly tested or inconsistently documented.

Evidence and measured results

Auditors tested system behavior and requested structured red-teaming. The Office of Technology and Innovation reported 95–99% accuracy but supplied insufficient test detail and disagreed with recommendations.

Limitations and uncertainty

The agency contested parts of the audit; readers should review both the findings and the response in the source.

Put this evidence to work

Lighthouse Advisory interpretation, grounded in this source as summarized in the preserved archive. Enriched 2026-09-05; this does not change the original publication date. Labels below come from the analysis itself.

Sales

Role takeaway
Customer problem
a municipal assistant can create service and trust risk when accuracy claims cannot be reproduced.
Stakeholders
digital-service leadership, content and contact-center owners, technology, accessibility, privacy, and independent oversight.
Discovery
who scores accuracy; can another reviewer reproduce the result; which inconsistent answers matter most; and how are errors corrected?
Value hypothesis
independent evaluation and clear fallback may produce a more defensible service-quality case.
Potential engagement
review the current benchmark, reproduce representative questions, and close assurance gaps.
Unsupported claims
the reported 95–99% accuracy is a disputed agency claim in this record, not an established performance result. The contested audit also does not justify assuming every public chatbot has the same defects.

Pre-sales engineering

Role takeaway
Fit
assess an existing or planned public information assistant with reproducible service tests.
Architecture and integration
use authoritative-source retrieval, versioned questions and expected answers, performance monitoring, failure capture, and fallback into existing service channels.
Prerequisites
benchmark ownership, scoring rules, content access, and records of model and configuration versions.
Constraints
repeated prompts may yield inconsistent answers, and opaque scoring prevents meaningful comparison.
Security
test prompt injection, disclosure, abusive inputs, and log retention as part of the evaluation. Proposed proof: rerun representative prompts across repeated trials, independently score correctness and consistency, measure delay, and verify fallback. Preserve both observed results and the operating team's response so disagreement is visible rather than averaged away.

Delivery

Role takeaway
Work
establish the evaluation record, assign content owners, introduce public error reporting, and integrate recurring independent assurance into release decisions.
Dependencies
access to prompts, scoring, versions, and service logs, plus owner cooperation when findings are disputed.
Ownership
the service owner accepts residual risk; content teams correct answers; engineering monitors behavior; an independent reviewer assesses claims.
Skills and adoption
train support staff in escalation and evaluators in reproducible scoring.
Governance checkpoints
preproduction review, material release, and follow-up on unresolved findings.
Proposed acceptance
independent reviewers can reproduce the reported customer benchmark, known severe errors are resolved or safely redirected, and performance/accessibility limits are documented. Risks include defensive reporting, selective test sets, and unresolved disagreements obscuring resident impact.

Implementation considerations

Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.

Architecture and integration

What must connect, and where does the AI sit in the workflow?

Build reproducible evaluation sets, performance monitoring, authoritative-source retrieval, failure capture, and a fallback path into the service architecture.

Governance

Who approves, reviews and stays accountable for outcomes?

Require public error reporting, red-team protocols, content ownership, release gates, and independent assurance for high-impact performance claims.

Security and privacy

What data, permissions and controls need testing?

Test prompt injection, data disclosure, abusive use, and log handling as part of an auditable preproduction and recurring assurance program.

The preserved archive analysis covered architecture, governance and security. Not assessed for this record: accessibility and workforce, procurement, operating model.

Publication history

  1. 2026-08-27SLED-wide archive · Issue 0110 resources
Read preserved resource versions (JSON)

Stable resource ID: nyc-mycity-audit