{"resourceId":"aurora-richland-ai-feedback-randomized-trials","versions":[{"version":"external-85f042b2f8a4861420234196a1b9b011a046285b9cfc8df113dd30172713d1e0","resource":{"id":"aurora-richland-ai-feedback-randomized-trials","title":"AI feedback changes officer speech scores, with different results across two agencies","organization":"Ian T. Adams, Kyle McLean and Geoffrey P. Alpert","sector":"Law enforcement oversight","geography":"Aurora, Colorado, and Richland County, South Carolina, United States","publishedAt":"First published December 22, 2025; 2026 journal issue","publicationDate":"2025-12-22","eventDate":null,"sourceName":"Criminology","sourceLabel":"Peer-reviewed randomized trials","sourceUrl":"https://onlinelibrary.wiley.com/doi/full/10.1111/1745-9125.70028","evidenceClass":"academic-research","outcomeClass":"mixed","topics":["data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"AI-generated feedback changed algorithm-defined speech scores, with benefits differing by agency and feedback route.","sledRelevance":"New-to-archive evidence for evaluating police coaching tools; not proof of reduced misconduct or improved public trust.","evidence":"Two six-month randomized trials enrolled 219 Aurora and 165 Richland officers, using no-feedback controls, self-assessment and supervisor-mediated arms. Analyses covered 124,443 and 65,172 videos. Aurora's two intervention arms reduced substandard scores; Richland's self-assessment arm increased high scores. Richland's supervisor-mediated result was not significant.","architectureImplications":"Interpretation: Separate BWC evidence storage, scoring, feedback and personnel decisions. Preserve model version and segment references; size ingestion and retention against local video volumes rather than trial volumes.","governanceImplications":"Interpretation: Validate the construct before using scores in discipline, promotion or public performance claims.","securityPrivacyImplications":"Interpretation: Limit review access, protect bystander audio, audit supervisor queries and prevent secondary reuse without approval.","caveats":"The proprietary linguistic outcome lacked independent human validation. Scores are proxies, not established measures of procedural justice. Gaming and organizational context limit interpretation. Funding is attributed to the Laura and John Arnold Foundation.","streamIds":["public-safety"],"roles":{"sales":"Interpretation: Oversight leaders and training commanders may need more systematic coaching, while officers and community representatives need assurance that scoring is meaningful. Ask which behavior is the target, who can contest a flag, and whether the department intends coaching or discipline. A bounded engagement could independently review sample interactions and compare feedback routes. The value hypothesis is more useful coaching after classification and workflow validation. Do not convert score improvements into claims about fewer rights violations, complaints or use-of-force incidents. This evidence provides a reason to evaluate a specific feedback mechanism, not a guarantee of organizational reform or a justification for indiscriminate surveillance.","engineering":"Interpretation: A candidate design links each flag to an authorized recording segment and records model version, feedback delivery and reviewer disposition. Prerequisites include representative recordings, independent annotators, permissions and agreement on what the outcome means. Test false flags, speaker attribution, accents, noise and missing recordings before any personnel integration. Keep analytical workspaces separate from authoritative case and employment records. Validate against human-reviewed behavior and independent outcomes, not just the same model's scores. Compare feedback delivery modes using a controlled pilot. The source does not specify a validated accuracy guarantee or establish suitability for an autonomous disciplinary agent.","delivery":"Interpretation: Training leadership should run the pilot with privacy, labor, oversight and evaluation partners. Document permitted use, explain the process to officers, and provide a route to challenge incorrect classifications. Dependencies include reliable recording ingestion, supervisor capacity and independent assessment skills. Governance checkpoints should precede launch, changes to scoring and any extension into personnel decisions. Proposed acceptance criteria include tested access restrictions, an agreed error threshold on independently labeled samples, documented feedback delivery and no unexplained disparity by tested language group. Track complaints and staff experience separately from tool scores. Stop expansion if score optimization displaces substantive coaching or meaningful review."},"retrievedAt":"2026-09-07T03:01:50Z","enrichedAt":"2026-09-07T03:01:50Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Test accents and language differences; involve labor representatives in surveillance and appeal policies.","procurementImplications":"Interpretation: Require independent validation access, change notice and auditable exports of scores and supporting segments.","operatingModelImplications":"Interpretation: Coaching ownership and feedback routes are implementation choices to test, not interchangeable features.","sourceVerification":{"openedUrl":"https://onlinelibrary.wiley.com/doi/full/10.1111/1745-9125.70028","referenceExcerpt":"The findings are mixed but positive","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}