From the Public Safety edition of September 6, 2026
AI feedback changes officer speech scores, with different results across two agencies
Ian T. Adams, Kyle McLean and Geoffrey P. Alpert · Law enforcement oversight · Aurora, Colorado, and Richland County, South Carolina, United States
- Publisher
- Criminology
- Original publication
- First published December 22, 2025; 2026 journal issue
- Source retrieved
- 2026-09-07
What happened
AI-generated feedback changed algorithm-defined speech scores, with benefits differing by agency and feedback route.
Why it matters
New-to-archive evidence for evaluating police coaching tools; not proof of reduced misconduct or improved public trust.
Evidence and measured results
Two six-month randomized trials enrolled 219 Aurora and 165 Richland officers, using no-feedback controls, self-assessment and supervisor-mediated arms. Analyses covered 124,443 and 65,172 videos. Aurora's two intervention arms reduced substandard scores; Richland's self-assessment arm increased high scores. Richland's supervisor-mediated result was not significant.
Limitations and uncertainty
The proprietary linguistic outcome lacked independent human validation. Scores are proxies, not established measures of procedural justice. Gaming and organizational context limit interpretation. Funding is attributed to the Laura and John Arnold Foundation.
Put this evidence to work
Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-07; this does not change the original publication date. Labels below come from the analysis itself.
Sales
Role takeaway
Oversight leaders and training commanders may need more systematic coaching, while officers and community representatives need assurance that scoring is meaningful. Ask which behavior is the target, who can contest a flag, and whether the department intends coaching or discipline. A bounded engagement could independently review sample interactions and compare feedback routes. The value hypothesis is more useful coaching after classification and workflow validation. Do not convert score improvements into claims about fewer rights violations, complaints or use-of-force incidents. This evidence provides a reason to evaluate a specific feedback mechanism, not a guarantee of organizational reform or a justification for indiscriminate surveillance.
Pre-sales engineering
Role takeaway
A candidate design links each flag to an authorized recording segment and records model version, feedback delivery and reviewer disposition. Prerequisites include representative recordings, independent annotators, permissions and agreement on what the outcome means. Test false flags, speaker attribution, accents, noise and missing recordings before any personnel integration. Keep analytical workspaces separate from authoritative case and employment records. Validate against human-reviewed behavior and independent outcomes, not just the same model's scores. Compare feedback delivery modes using a controlled pilot. The source does not specify a validated accuracy guarantee or establish suitability for an autonomous disciplinary agent.
Delivery
Role takeaway
Training leadership should run the pilot with privacy, labor, oversight and evaluation partners. Document permitted use, explain the process to officers, and provide a route to challenge incorrect classifications. Dependencies include reliable recording ingestion, supervisor capacity and independent assessment skills. Governance checkpoints should precede launch, changes to scoring and any extension into personnel decisions. Proposed acceptance criteria include tested access restrictions, an agreed error threshold on independently labeled samples, documented feedback delivery and no unexplained disparity by tested language group. Track complaints and staff experience separately from tool scores. Stop expansion if score optimization displaces substantive coaching or meaningful review.
Implementation considerations
Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.
Architecture and integration
What must connect, and where does the AI sit in the workflow?
Separate BWC evidence storage, scoring, feedback and personnel decisions. Preserve model version and segment references; size ingestion and retention against local video volumes rather than trial volumes.
Governance
Who approves, reviews and stays accountable for outcomes?
Validate the construct before using scores in discipline, promotion or public performance claims.
Security and privacy
What data, permissions and controls need testing?
Limit review access, protect bystander audio, audit supervisor queries and prevent secondary reuse without approval.
Accessibility and workforce
Who is affected, and what skills or accommodations follow?
Test accents and language differences; involve labor representatives in surveillance and appeal policies.
Procurement
What should contracts, pricing and exit terms secure?
Require independent validation access, change notice and auditable exports of scores and supporting segments.
Operating model
Which teams own the service once it runs?
Coaching ownership and feedback routes are implementation choices to test, not interchangeable features.
Publication history
- 2026-09-06Public Safety · Issue 014 resources
Stable resource ID: aurora-richland-ai-feedback-randomized-trials