{"resourceId":"virginia-beach-truststat-replication-2026","versions":[{"version":"external-bf25cbbb21744e6f8a5fe021eaaaefa25397ac809c220d49bbccfdf70b48a2a9","resource":{"id":"virginia-beach-truststat-replication-2026","title":"Body-camera AI reproduces a broad training conclusion with unresolved measurement limits","organization":"Kyle McLean, Jeffrey Rojek and Justin Nix","sector":"Law enforcement evaluation","geography":"Virginia Beach, Virginia, United States","publishedAt":"May 16, 2026","publicationDate":"2026-05-16","eventDate":null,"sourceName":"Journal of Experimental Criminology","sourceLabel":"Peer-reviewed original research; full article and PDF","sourceUrl":"https://link.springer.com/article/10.1007/s11292-026-09754-4","evidenceClass":"academic-research","outcomeClass":"mixed","topics":["knowledge-work","infrastructure","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"TrustStat scores supported a prior de-escalation training study's broad conclusion, but did not measure identical constructs to human observation.","sledRelevance":"Historical evidence newly archived for departments evaluating automated review of officer interactions.","evidence":"Virginia Beach footage and RMS data were analyzed with the vendor blinded to training assignment and human scores. AI regressions used 181 pre-test and 232 post-test observations. Calm and tone showed post-test differences; respect had a similar pre-test coefficient. Human-observation comparison tables used smaller samples.","architectureImplications":"Interpretation: Separate evidence ingestion, scoring and researcher analysis; preserve sample membership across versions.","governanceImplications":"Interpretation: Validate each intended use before scores influence personnel decisions.","securityPrivacyImplications":"Interpretation: Limit video access and secondary reuse, with auditable transfer and retention controls.","caveats":"One agency and training program; proprietary scoring details withheld. Vendor supplied AI scoring free; authors declared no financial relationship. Different constructs and sample sizes limit equivalence claims. No measured cost saving or causal benefit from deploying AI itself.","streamIds":["public-safety"],"roles":{"sales":"Interpretation: Training commanders, research partners, oversight officials and workforce representatives need timely evaluation of officer communication. Ask which behavior matters, whether a usable human-coded baseline exists, and what decisions a score would influence. A bounded engagement could assess one training cohort with independent reviewers and a documented comparison plan. The value hypothesis is faster access to trustworthy feedback if local validation supports it. Do not sell the study as proof of reduced misconduct, fewer complaints or a substitute for reviewers. Clarify whether the buyer seeks research assistance, coaching or discipline because those purposes require different evidence and authorization.","engineering":"Interpretation: Fit is a protected batch-analysis pilot. Establish stable links among footage, RMS metadata, coded observations and model outputs; restrict identities to authorized staff. Prerequisites include representative recordings, independent coding expertise and a clear definition of each target behavior. Evaluate missing files, subgroup performance, sample exclusions and score stability across versions. Compare the same encounters under each method before interpreting aggregate agreement. Obtain technical documentation sufficient to understand failure modes and test current software rather than assuming continuity with the study. Hosting requirements remain unestablished; compare approved cloud processing with local governance needs before choosing an architecture.","delivery":"Interpretation: A training-evaluation lead should coordinate analysts, privacy staff, supervisors and workforce representatives. Agree on permissible uses, select recordings, train coders, reconcile disagreements and review pilot findings before expansion. Dependencies include recording quality, approved access and sufficient evaluation capacity. Proposed acceptance: every analyzed encounter has traceable inputs and exclusion reasons, critical disagreement is reviewed, and predefined behavior-specific validity thresholds are met on held-out local material. These are proposals rather than demonstrated outcomes. Offer an accessible review and challenge route. Monitor adoption and review workload; pause if staff begin treating coaching scores as disciplinary findings without separate authorization."},"retrievedAt":"2026-09-13T03:01:48Z","enrichedAt":"2026-09-13T03:05:28Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Test communication variation and disability-related differences; provide staff a contextual review route.","procurementImplications":"Interpretation: Negotiate documentation, version notices, export and audit rights before accepting proprietary scores.","operatingModelImplications":"Interpretation: Keep evaluation ownership independent of vendor success reporting; distinguish coaching from discipline.","updateExplanation":"New-to-archive May 2026 study adds a human-observation replication test, distinct from previously archived speech-feedback trials.","sourceVerification":{"openedUrl":"https://link.springer.com/article/10.1007/s11292-026-09754-4","referenceExcerpt":"the measures being compared are not the same beyond the broad construct of improved communication.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}