Lighthouse AdvisorySLED AI Adoption Intelligence
← Back to results

From the Public Safety edition of September 11, 2026

Independent researchCautionaryNewly relevant · Jul 2025

Historical prosecution experiment finds adverse recommendations despite exculpatory facts

Justice Innovation Lab; Rory Pulvino, Dan Sutton and JJ Naddeo · Prosecution and criminal courts · United States; reports from Oklahoma City, Buffalo and Seattle

Publisher
Justice Innovation Lab Knowledge Hub
Original publication
July 29, 2025
Source retrieved
2026-09-12
Read original source

What happened

A GPT-3.5-Turbo experiment found a tendency toward prosecution, including legally deficient scenarios; no racial disparity in recommendations was detected in this test.

Why it matters

New archive coverage of U.S. prosecutorial memo assistance, distinct from police report writing and judicial interviews. This is historical evidence, not a current-model evaluation.

Evidence and measured results

Researchers used 20 arrest reports and paired altered versions, varying role, prompt detail and racial identifiers. Repeated requests produced over 144,000 responses. More context changed recommendations and variability. Methods: https://knowledgehub.justiceinnovationlab.org/reports/ai-in-prosecution1/data.

Limitations and uncertainty

Small underlying case sample despite many responses; no prosecutor comparison group or observed case outcomes. Older model and uneven flaw severity limit generalization. No claim that present systems share these rates or that racial fairness is established.

Put this evidence to work

Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-12; this does not change the original publication date. Labels below come from the analysis itself.

Sales

Role takeaway

Prosecutors, public defenders, legal supervisors and justice IT leaders need reliable assistance under heavy caseloads. Ask whether staff use open-ended memo prompts, which decisions the drafts influence, and how contradictory facts are reviewed. A bounded engagement could inventory one drafting workflow and develop an expert-adjudicated test set before any integration. The value hypothesis is discovering unsuitable uses and measuring review burden, not proving that automation improves charging decisions. Explain that repeated responses are not thousands of independent criminal cases. Do not claim current products have the same behavior or promise fewer wrongful prosecutions from a training package. Applicability depends on the actual model, task and local legal context.

Pre-sales engineering

Role takeaway

Fit is an isolated evaluation of the proposed assistant, not automated charging. Preserve source documents, relevant approved legal references, model identifiers and prompt versions in an access-controlled environment. Prerequisites include attorney-created reference assessments and representative examples of missing elements and conflicting evidence. Repeat identical inputs to measure variability, then vary prompts and model versions deliberately. Compare assisted and unassisted reviewers because this study lacks a prosecutor baseline. Test whether a fluent draft draws attention away from exculpatory material. Cloud or local deployment does not itself solve this failure mode. Any agent integration needs separate permission boundaries around case updates, correspondence and filings.

Delivery

Role takeaway

A designated legal practice lead should coordinate prosecutors, defense-informed reviewers, paralegals, IT and privacy counsel. Map where drafts influence action, agree on permissible use, establish an unassisted baseline and conduct a reversible shadow evaluation. Train users on omission detection and escalation, with accessible review instructions.

Proposed acceptance
all test cases with decisive contradictory facts reach attorney review, every accepted recommendation records an independent rationale, and unresolved critical errors stop progression. These criteria are proposals rather than research results. Dependencies include protected review time and lawful access to case records. Risks include selective testing, automation bias, unstable outputs and treating better formatting as sound judgment.

Implementation considerations

Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.

Architecture and integration

What must connect, and where does the AI sit in the workflow?

Isolate drafting from filing, preserve input and prompt versions, and require source-linked treatment of contradictory evidence.

Governance

Who approves, reviews and stays accountable for outcomes?

Legal supervisors should define which tasks may use assistance and review missed exculpatory facts separately from prose quality.

Security and privacy

What data, permissions and controls need testing?

Use authorized, minimized test records; confirm retention, training use and access terms before uploading case material.

Accessibility and workforce

Who is affected, and what skills or accommodations follow?

Train attorneys and paralegals on variability and missing information; provide accessible review templates and a non-AI route.

Procurement

What should contracts, pricing and exit terms secure?

Require model/version disclosure and permission to run adverse-case tests before purchasing a prosecution assistant.

Operating model

Which teams own the service once it runs?

Assign legal ownership of evaluation and correction; IT alone cannot judge charging-memo quality.

What changed

New-to-archive historical source fills a prosecution-specific evaluation gap. July 2025 findings and old-model limitations are preserved; no claim of a fresh event.

Publication history

  1. 2026-09-11Public Safety · Issue 063 resources
Read preserved resource versions (JSON)

Stable resource ID: jil-prosecution-memo-bias-experiment-2025