From the Public Safety edition of September 7, 2026
Blinded police-report study distinguishes perceived quality from factual verification
Ian T. Adams and coauthors · Public safety · United States; one police agency
- Publisher
- CrimRxiv
- Original publication
- May 8, 2026
- Source retrieved
- 2026-09-08
What happened
AI-assisted reports were less readable, while the primary overall perceived-quality difference was not statistically significant.
Why it matters
New-to-archive quality evidence complements yesterday's Manchester timing trial using related underlying reports; it is not an independent deployment or a new September experiment.
Evidence and measured results
The sample contained 20 assisted and 60 conventional reports; 92 raters supplied 354 evaluations covering 79 reports. Flesch scores were 52.28 versus 57.92. Overall quality p=.094; the accuracy-rating subscale p=.038. These are perceptions, not verified factual-error rates.
Limitations and uncertainty
Preprint, single agency/tool, small assisted sample, subjective ratings and generic readability metrics. Multiple subscales warrant caution. Methods and discussion differ on when raters were primed about AI; omit strong detection claims.
Put this evidence to work
Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-08; this does not change the original publication date. Labels below come from the analysis itself.
Sales
Role takeaway
Records supervisors, patrol leadership and legal consumers need reports that communicate reliably. Ask whether current complaints concern omissions, readability or correction work, and whether anyone checks reports against original evidence. A bounded engagement could compare assisted and conventional reports for one incident category using blinded reviewers and source verification. The value hypothesis is a better-informed purchase decision, not guaranteed quality improvement. Explain that the observed accuracy result is a rating, not a counted fabrication rate. Because these reports relate to a previously studied agency, do not present the paper as independent multi-agency replication.
Pre-sales engineering
Role takeaway
Fit is a drafting copilot with a verifiable evidence path. Map transcription, generation, officer additions, review and RMS export; record model version and intermediate artifacts under approved access controls. Prerequisites include representative authorized recordings and independently checked reference facts. Test missing visual details, speaker attribution, unsupported additions and accessible presentation. Compare total correction effort and reader comprehension alongside factual completeness. Cloud, hybrid and local alternatives require separate data-handling and integration assessment; the preprint validates none of them. Any agent should prepare a draft while an accountable officer retains submission authority.
Delivery
Role takeaway
The records unit should own a limited pilot with training, IT and prosecutor/defender input. Establish baseline quality, recruit independent readers, train supervisors on distinct error types and document reasons for revisions. Dependencies include source access, annotation skills and protected reviewer time. Proposed acceptance requires no critical unsupported facts in the agreed test set, complete provenance, and a locally agreed comprehension and correction-effort target. These are proposed gates, not study results. Recheck after model changes and track non-adoption. Risks include rating fluent text generously, treating a nonsignificant result as equivalence and overlooking burdens shifted to downstream readers.
Implementation considerations
Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.
Architecture and integration
What must connect, and where does the AI sit in the workflow?
Pair drafts and final reports with authorized underlying audio, visual observations and explicit officer additions.
Governance
Who approves, reviews and stays accountable for outcomes?
Separate readability, perceived accuracy and ground-truth correctness in acceptance rubrics.
Security and privacy
What data, permissions and controls need testing?
Limit access to recordings and drafts, maintaining approved retention and discoverable provenance.
Accessibility and workforce
Who is affected, and what skills or accommodations follow?
Test comprehension with downstream readers and language-diverse reviewers; polished vocabulary is not an accessibility measure.
Procurement
What should contracts, pricing and exit terms secure?
Require a local quality evaluation and exportable draft history before expansion; this study is no guarantee about newer models.
Operating model
Which teams own the service once it runs?
Records supervisors and downstream legal readers should jointly own quality definitions.
Publication history
- 2026-09-07Public Safety · Issue 024 resources
Stable resource ID: police-report-quality-blinded-evaluation-2026