Lighthouse AdvisorySLED AI Adoption Intelligence
← Back to results

From the Public Safety edition of September 8, 2026

Independent researchCautionaryNewly relevant · Jun 2026

Police-report research warns that officer review can miss material errors

Federation of American Scientists; Jon Peha · Public safety · United States

Publisher
Federation of American Scientists
Original publication
June 9, 2026; describes research conducted in 2025
Source retrieved
2026-09-09
Read original source

What happened

Peha reports material inaccuracies in generated police reports and missed errors when experienced officers reviewed deliberately flawed reports.

Why it matters

Historical source newly added to the archive supplies a review-training lens for U.S. police drafting pilots; it is not a new September experiment.

Evidence and measured results

The author describes a 2025 CMU exercise using three kinds of generative AI and a separate officer-editing exercise involving hallucinations, omissions and event-order errors. Officers had not received specific AI-editing training. No participant count, numeric error rate or controlled training-effect estimate is supplied.

Limitations and uncertainty

This is an organizer's account in a policy memo, not a complete peer-reviewed methods report. A university exercise cannot establish field error rates or prove training fixes the problem. Proposed NIJ programs are recommendations, not verified current services.

Put this evidence to work

Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-09; this does not change the original publication date. Labels below come from the analysis itself.

Sales

Role takeaway

Police records leaders, training staff, prosecutors and procurement teams need to know whether assisted drafting leaves trustworthy reports. Ask how reviewers check original material, which incident types dominate work, and whether error correction is measured alongside time. A bounded engagement could design a review benchmark and training pilot. The value hypothesis is fewer undetected errors at an acceptable total workload. The qualitative exercise supports testing that hypothesis; it does not establish an error rate, cost saving or training benefit for a customer. Qualify source access and legal review capacity before proposing operational deployment.

Pre-sales engineering

Role takeaway

Fit is a controlled drafting and comparison workspace. Link each factual assertion to approved source material where feasible and retain authorized versions for audit. Prerequisites include representative incidents, adjudicated reference facts, permission boundaries and a records policy. Test omissions, wrong speakers, chronology changes and unsupported certainty, measuring errors remaining after review as well as total completion time. Compare trained and current-process reviewers using an agreed design. Prevent cross-case retrieval and unapproved provider reuse. Do not infer that a confidence score or second model establishes accuracy; the source lacks quantitative evidence that any review architecture works.

Delivery

Role takeaway

A records-quality owner should coordinate training, supervisors, legal reviewers and IT. Build a consented or properly de-identified test set, define an error rubric, establish baseline workload and rehearse correction of a circulated report. Prerequisites are original evidence and staff time for independent adjudication. Proposed acceptance is complete traceability for sampled pilot reports and meeting locally agreed material-error and workload thresholds before expansion. Train with subtle errors rather than obvious placeholders, and monitor adoption separately from quality. Risks include nominal approval, reviewer fatigue, inaccessible recordings and performance changes after product updates.

Implementation considerations

Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.

Architecture and integration

What must connect, and where does the AI sit in the workflow?

Keep source recordings, transcription, drafts and edits traceable in an approved records workflow. Select cloud, on-premises or hybrid processing only after testing data flows and retention; the memo validates no hosting design. Autonomous filing agents have limited fit.

Governance

Who approves, reviews and stays accountable for outcomes?

Require source-based review and a stop-use route; a signature alone is an inadequate proof-of-value metric.

Security and privacy

What data, permissions and controls need testing?

Test case separation, provider retention and restrictions on training with sensitive inputs.

Accessibility and workforce

Who is affected, and what skills or accommodations follow?

Train reviewers using realistic omissions and attribution errors, including diverse speech and accessibility needs.

Procurement

What should contracts, pricing and exit terms secure?

Require a representative evaluation set and version-specific quality evidence before licensing.

Operating model

Which teams own the service once it runs?

Assign police records leadership responsibility for ongoing quality, with prosecutors and defense-access procedures involved.

Publication history

  1. 2026-09-08Public Safety · Issue 033 resources
Read preserved resource versions (JSON)

Stable resource ID: fas-police-report-review-research-2026