From the Public Safety edition of September 8, 2026
Police-report research warns that officer review can miss material errors
Federation of American Scientists; Jon Peha · Public safety · United States
- Publisher
- Federation of American Scientists
- Original publication
- June 9, 2026; describes research conducted in 2025
- Source retrieved
- 2026-09-09
What happened
Peha reports material inaccuracies in generated police reports and missed errors when experienced officers reviewed deliberately flawed reports.
Why it matters
Historical source newly added to the archive supplies a review-training lens for U.S. police drafting pilots; it is not a new September experiment.
Evidence and measured results
The author describes a 2025 CMU exercise using three kinds of generative AI and a separate officer-editing exercise involving hallucinations, omissions and event-order errors. Officers had not received specific AI-editing training. No participant count, numeric error rate or controlled training-effect estimate is supplied.
Limitations and uncertainty
This is an organizer's account in a policy memo, not a complete peer-reviewed methods report. A university exercise cannot establish field error rates or prove training fixes the problem. Proposed NIJ programs are recommendations, not verified current services.
Put this evidence to work
Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-09; this does not change the original publication date. Labels below come from the analysis itself.
Sales
Role takeaway
Police records leaders, training staff, prosecutors and procurement teams need to know whether assisted drafting leaves trustworthy reports. Ask how reviewers check original material, which incident types dominate work, and whether error correction is measured alongside time. A bounded engagement could design a review benchmark and training pilot. The value hypothesis is fewer undetected errors at an acceptable total workload. The qualitative exercise supports testing that hypothesis; it does not establish an error rate, cost saving or training benefit for a customer. Qualify source access and legal review capacity before proposing operational deployment.
Pre-sales engineering
Role takeaway
Fit is a controlled drafting and comparison workspace. Link each factual assertion to approved source material where feasible and retain authorized versions for audit. Prerequisites include representative incidents, adjudicated reference facts, permission boundaries and a records policy. Test omissions, wrong speakers, chronology changes and unsupported certainty, measuring errors remaining after review as well as total completion time. Compare trained and current-process reviewers using an agreed design. Prevent cross-case retrieval and unapproved provider reuse. Do not infer that a confidence score or second model establishes accuracy; the source lacks quantitative evidence that any review architecture works.
Delivery
Role takeaway
A records-quality owner should coordinate training, supervisors, legal reviewers and IT. Build a consented or properly de-identified test set, define an error rubric, establish baseline workload and rehearse correction of a circulated report. Prerequisites are original evidence and staff time for independent adjudication. Proposed acceptance is complete traceability for sampled pilot reports and meeting locally agreed material-error and workload thresholds before expansion. Train with subtle errors rather than obvious placeholders, and monitor adoption separately from quality. Risks include nominal approval, reviewer fatigue, inaccessible recordings and performance changes after product updates.
Implementation considerations
Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.
Architecture and integration
What must connect, and where does the AI sit in the workflow?
Keep source recordings, transcription, drafts and edits traceable in an approved records workflow. Select cloud, on-premises or hybrid processing only after testing data flows and retention; the memo validates no hosting design. Autonomous filing agents have limited fit.
Governance
Who approves, reviews and stays accountable for outcomes?
Require source-based review and a stop-use route; a signature alone is an inadequate proof-of-value metric.
Security and privacy
What data, permissions and controls need testing?
Test case separation, provider retention and restrictions on training with sensitive inputs.
Accessibility and workforce
Who is affected, and what skills or accommodations follow?
Train reviewers using realistic omissions and attribution errors, including diverse speech and accessibility needs.
Procurement
What should contracts, pricing and exit terms secure?
Require a representative evaluation set and version-specific quality evidence before licensing.
Operating model
Which teams own the service once it runs?
Assign police records leadership responsibility for ongoing quality, with prosecutors and defense-access procedures involved.
Publication history
- 2026-09-08Public Safety · Issue 033 resources
Stable resource ID: fas-police-report-review-research-2026