{"resourceId":"fas-police-report-review-research-2026","versions":[{"version":"external-d5b0c9fbee05a818e777ccc4a55c15257ad6ca0fe1f11cb8dcfa18071f1caffc","resource":{"id":"fas-police-report-review-research-2026","title":"Police-report research warns that officer review can miss material errors","organization":"Federation of American Scientists; Jon Peha","sector":"Public safety","geography":"United States","publishedAt":"June 9, 2026; describes research conducted in 2025","publicationDate":"2026-06-09","eventDate":null,"sourceName":"Federation of American Scientists","sourceLabel":"Research organizer's qualitative account and policy proposal","sourceUrl":"https://fas.org/publication/safe-ai-police-reports/","evidenceClass":"independent-research","outcomeClass":"cautionary","topics":["knowledge-work","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"Peha reports material inaccuracies in generated police reports and missed errors when experienced officers reviewed deliberately flawed reports.","sledRelevance":"Historical source newly added to the archive supplies a review-training lens for U.S. police drafting pilots; it is not a new September experiment.","evidence":"The author describes a 2025 CMU exercise using three kinds of generative AI and a separate officer-editing exercise involving hallucinations, omissions and event-order errors. Officers had not received specific AI-editing training. No participant count, numeric error rate or controlled training-effect estimate is supplied.","architectureImplications":"Interpretation: Keep source recordings, transcription, drafts and edits traceable in an approved records workflow. Select cloud, on-premises or hybrid processing only after testing data flows and retention; the memo validates no hosting design. Autonomous filing agents have limited fit.","governanceImplications":"Interpretation: Require source-based review and a stop-use route; a signature alone is an inadequate proof-of-value metric.","securityPrivacyImplications":"Interpretation: Test case separation, provider retention and restrictions on training with sensitive inputs.","caveats":"This is an organizer's account in a policy memo, not a complete peer-reviewed methods report. A university exercise cannot establish field error rates or prove training fixes the problem. Proposed NIJ programs are recommendations, not verified current services.","streamIds":["public-safety"],"roles":{"sales":"Interpretation: Police records leaders, training staff, prosecutors and procurement teams need to know whether assisted drafting leaves trustworthy reports. Ask how reviewers check original material, which incident types dominate work, and whether error correction is measured alongside time. A bounded engagement could design a review benchmark and training pilot. The value hypothesis is fewer undetected errors at an acceptable total workload. The qualitative exercise supports testing that hypothesis; it does not establish an error rate, cost saving or training benefit for a customer. Qualify source access and legal review capacity before proposing operational deployment.","engineering":"Interpretation: Fit is a controlled drafting and comparison workspace. Link each factual assertion to approved source material where feasible and retain authorized versions for audit. Prerequisites include representative incidents, adjudicated reference facts, permission boundaries and a records policy. Test omissions, wrong speakers, chronology changes and unsupported certainty, measuring errors remaining after review as well as total completion time. Compare trained and current-process reviewers using an agreed design. Prevent cross-case retrieval and unapproved provider reuse. Do not infer that a confidence score or second model establishes accuracy; the source lacks quantitative evidence that any review architecture works.","delivery":"Interpretation: A records-quality owner should coordinate training, supervisors, legal reviewers and IT. Build a consented or properly de-identified test set, define an error rubric, establish baseline workload and rehearse correction of a circulated report. Prerequisites are original evidence and staff time for independent adjudication. Proposed acceptance is complete traceability for sampled pilot reports and meeting locally agreed material-error and workload thresholds before expansion. Train with subtle errors rather than obvious placeholders, and monitor adoption separately from quality. Risks include nominal approval, reviewer fatigue, inaccessible recordings and performance changes after product updates."},"retrievedAt":"2026-09-09T03:01:16Z","enrichedAt":"2026-09-09T03:03:37Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Train reviewers using realistic omissions and attribution errors, including diverse speech and accessibility needs.","procurementImplications":"Interpretation: Require a representative evaluation set and version-specific quality evidence before licensing.","operatingModelImplications":"Interpretation: Assign police records leadership responsibility for ongoing quality, with prosecutors and defense-access procedures involved.","sourceVerification":{"openedUrl":"https://fas.org/publication/safe-ai-police-reports/","referenceExcerpt":"We observed that officers missed many problems, including those that might matter in legal proceedings","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}