{"resourceId":"jil-prosecution-memo-bias-experiment-2025","versions":[{"version":"external-6883521f27faf9a39d5d68e393c81ef1cd7c3ea5062fa00014fdbca93471c213","resource":{"id":"jil-prosecution-memo-bias-experiment-2025","title":"Historical prosecution experiment finds adverse recommendations despite exculpatory facts","organization":"Justice Innovation Lab; Rory Pulvino, Dan Sutton and JJ Naddeo","sector":"Prosecution and criminal courts","geography":"United States; reports from Oklahoma City, Buffalo and Seattle","publishedAt":"July 29, 2025","publicationDate":"2025-07-29","eventDate":null,"sourceName":"Justice Innovation Lab Knowledge Hub","sourceLabel":"Original independent research report and methods page","sourceUrl":"https://knowledgehub.justiceinnovationlab.org/reports/ai-in-prosecution1","evidenceClass":"independent-research","outcomeClass":"cautionary","topics":["knowledge-work","developers-agents","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"A GPT-3.5-Turbo experiment found a tendency toward prosecution, including legally deficient scenarios; no racial disparity in recommendations was detected in this test.","sledRelevance":"New archive coverage of U.S. prosecutorial memo assistance, distinct from police report writing and judicial interviews. This is historical evidence, not a current-model evaluation.","evidence":"Researchers used 20 arrest reports and paired altered versions, varying role, prompt detail and racial identifiers. Repeated requests produced over 144,000 responses. More context changed recommendations and variability. Methods: https://knowledgehub.justiceinnovationlab.org/reports/ai-in-prosecution1/data.","architectureImplications":"Interpretation: Isolate drafting from filing, preserve input and prompt versions, and require source-linked treatment of contradictory evidence.","governanceImplications":"Interpretation: Legal supervisors should define which tasks may use assistance and review missed exculpatory facts separately from prose quality.","securityPrivacyImplications":"Interpretation: Use authorized, minimized test records; confirm retention, training use and access terms before uploading case material.","caveats":"Small underlying case sample despite many responses; no prosecutor comparison group or observed case outcomes. Older model and uneven flaw severity limit generalization. No claim that present systems share these rates or that racial fairness is established.","streamIds":["public-safety"],"roles":{"sales":"Interpretation: Prosecutors, public defenders, legal supervisors and justice IT leaders need reliable assistance under heavy caseloads. Ask whether staff use open-ended memo prompts, which decisions the drafts influence, and how contradictory facts are reviewed. A bounded engagement could inventory one drafting workflow and develop an expert-adjudicated test set before any integration. The value hypothesis is discovering unsuitable uses and measuring review burden, not proving that automation improves charging decisions. Explain that repeated responses are not thousands of independent criminal cases. Do not claim current products have the same behavior or promise fewer wrongful prosecutions from a training package. Applicability depends on the actual model, task and local legal context.","engineering":"Interpretation: Fit is an isolated evaluation of the proposed assistant, not automated charging. Preserve source documents, relevant approved legal references, model identifiers and prompt versions in an access-controlled environment. Prerequisites include attorney-created reference assessments and representative examples of missing elements and conflicting evidence. Repeat identical inputs to measure variability, then vary prompts and model versions deliberately. Compare assisted and unassisted reviewers because this study lacks a prosecutor baseline. Test whether a fluent draft draws attention away from exculpatory material. Cloud or local deployment does not itself solve this failure mode. Any agent integration needs separate permission boundaries around case updates, correspondence and filings.","delivery":"Interpretation: A designated legal practice lead should coordinate prosecutors, defense-informed reviewers, paralegals, IT and privacy counsel. Map where drafts influence action, agree on permissible use, establish an unassisted baseline and conduct a reversible shadow evaluation. Train users on omission detection and escalation, with accessible review instructions. Proposed acceptance: all test cases with decisive contradictory facts reach attorney review, every accepted recommendation records an independent rationale, and unresolved critical errors stop progression. These criteria are proposals rather than research results. Dependencies include protected review time and lawful access to case records. Risks include selective testing, automation bias, unstable outputs and treating better formatting as sound judgment."},"retrievedAt":"2026-09-12T03:02:21Z","enrichedAt":"2026-09-12T03:04:22Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Train attorneys and paralegals on variability and missing information; provide accessible review templates and a non-AI route.","procurementImplications":"Interpretation: Require model/version disclosure and permission to run adverse-case tests before purchasing a prosecution assistant.","operatingModelImplications":"Interpretation: Assign legal ownership of evaluation and correction; IT alone cannot judge charging-memo quality.","updateExplanation":"New-to-archive historical source fills a prosecution-specific evaluation gap. July 2025 findings and old-model limitations are preserved; no claim of a fresh event.","sourceVerification":{"openedUrl":"https://knowledgehub.justiceinnovationlab.org/reports/ai-in-prosecution1","referenceExcerpt":"this testing does not have a prosecutor comparison group to measure against.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}