Lighthouse AdvisorySLED AI Adoption Intelligence
← Back to results

From the Public Safety edition of September 6, 2026

Academic researchMixedNewly relevant · Dec 2025

AI feedback changes officer speech scores, with different results across two agencies

Ian T. Adams, Kyle McLean and Geoffrey P. Alpert · Law enforcement oversight · Aurora, Colorado, and Richland County, South Carolina, United States

Publisher
Criminology
Original publication
First published December 22, 2025; 2026 journal issue
Source retrieved
2026-09-07
Read original source

What happened

AI-generated feedback changed algorithm-defined speech scores, with benefits differing by agency and feedback route.

Why it matters

New-to-archive evidence for evaluating police coaching tools; not proof of reduced misconduct or improved public trust.

Evidence and measured results

Two six-month randomized trials enrolled 219 Aurora and 165 Richland officers, using no-feedback controls, self-assessment and supervisor-mediated arms. Analyses covered 124,443 and 65,172 videos. Aurora's two intervention arms reduced substandard scores; Richland's self-assessment arm increased high scores. Richland's supervisor-mediated result was not significant.

Limitations and uncertainty

The proprietary linguistic outcome lacked independent human validation. Scores are proxies, not established measures of procedural justice. Gaming and organizational context limit interpretation. Funding is attributed to the Laura and John Arnold Foundation.

Put this evidence to work

Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-07; this does not change the original publication date. Labels below come from the analysis itself.

Sales

Role takeaway

Oversight leaders and training commanders may need more systematic coaching, while officers and community representatives need assurance that scoring is meaningful. Ask which behavior is the target, who can contest a flag, and whether the department intends coaching or discipline. A bounded engagement could independently review sample interactions and compare feedback routes. The value hypothesis is more useful coaching after classification and workflow validation. Do not convert score improvements into claims about fewer rights violations, complaints or use-of-force incidents. This evidence provides a reason to evaluate a specific feedback mechanism, not a guarantee of organizational reform or a justification for indiscriminate surveillance.

Pre-sales engineering

Role takeaway

A candidate design links each flag to an authorized recording segment and records model version, feedback delivery and reviewer disposition. Prerequisites include representative recordings, independent annotators, permissions and agreement on what the outcome means. Test false flags, speaker attribution, accents, noise and missing recordings before any personnel integration. Keep analytical workspaces separate from authoritative case and employment records. Validate against human-reviewed behavior and independent outcomes, not just the same model's scores. Compare feedback delivery modes using a controlled pilot. The source does not specify a validated accuracy guarantee or establish suitability for an autonomous disciplinary agent.

Delivery

Role takeaway

Training leadership should run the pilot with privacy, labor, oversight and evaluation partners. Document permitted use, explain the process to officers, and provide a route to challenge incorrect classifications. Dependencies include reliable recording ingestion, supervisor capacity and independent assessment skills. Governance checkpoints should precede launch, changes to scoring and any extension into personnel decisions. Proposed acceptance criteria include tested access restrictions, an agreed error threshold on independently labeled samples, documented feedback delivery and no unexplained disparity by tested language group. Track complaints and staff experience separately from tool scores. Stop expansion if score optimization displaces substantive coaching or meaningful review.

Implementation considerations

Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.

Architecture and integration

What must connect, and where does the AI sit in the workflow?

Separate BWC evidence storage, scoring, feedback and personnel decisions. Preserve model version and segment references; size ingestion and retention against local video volumes rather than trial volumes.

Governance

Who approves, reviews and stays accountable for outcomes?

Validate the construct before using scores in discipline, promotion or public performance claims.

Security and privacy

What data, permissions and controls need testing?

Limit review access, protect bystander audio, audit supervisor queries and prevent secondary reuse without approval.

Accessibility and workforce

Who is affected, and what skills or accommodations follow?

Test accents and language differences; involve labor representatives in surveillance and appeal policies.

Procurement

What should contracts, pricing and exit terms secure?

Require independent validation access, change notice and auditable exports of scores and supporting segments.

Operating model

Which teams own the service once it runs?

Coaching ownership and feedback routes are implementation choices to test, not interchangeable features.

Publication history

  1. 2026-09-06Public Safety · Issue 014 resources
Read preserved resource versions (JSON)

Stable resource ID: aurora-richland-ai-feedback-randomized-trials