Public Sector & Government · Issue 06 ·
Public Safety
Three historical sources new to the archive cover the UK PoliceAI pilot programme, a U.S. prosecution-memo experiment, and a COMPAS dataset reanalysis. Operator projections remain distinct from measured results; old-model and single-county limitations are explicit. One cross-source interpretation addresses multiple dimensions of bias. No fresh institutional-prison outcomes or new-since-last-run operational results were verified.
- Evidence records
- 3
- Cross-source patterns
- 1
- Evidence classes
- 1 vendor claim1 independent research1 academic research
- Outcomes
- 1 emerging1 cautionary1 mixed
- Source freshness
- 3 older, newly relevant
- Research completed
- 2026-09-12
Choose a role to see its takeaway beside every record in the ledger.
Synthesis · Lighthouse Advisory interpretation
Patterns across the evidence
Test multiple dimensions of bias rather than relying on one reassuring result
JIL found adverse prosecutorial recommendations without detecting racial disparities in that experiment; the COMPAS reanalysis found improved aggregate performance alongside persistent group error differences. These distinct tests show why neither a single parity result nor an aggregate metric settles suitability for a justice task. They do not establish equivalent mechanisms or outcomes across tools.
Operating questionDoes evaluation separately test decision-specific errors, group disparities and output stability against a locally justified reference?
Supporting evidenceJustice Innovation Lab; Rory Pulvino, Dan Sutton and JJ NaddeoSarah Kienzle, Gissel Velarde and Brian Gannon; IU International University of Applied Sciences
Full record · every source keeps its link and limitations
Evidence ledger
PoliceAI launch sets national evidence-processing pilots, with benefits still to be validated
The Home Office announced a national AI centre and digital-evidence pilots, alongside planned independent testing and a public tool registry.
Why it matters, evidence and limitations
- Why it matters
- Historical implementation evidence newly added to the archive. U.S. agencies can examine shared evaluation capacity; British funding, legal powers and national coordination do not transfer directly.
- Evidence and measured results
- The June announcement allocates £75 million over three years to PoliceAI and describes pilots in up to 10 forces during 2026–27. It reports 800 hours of kidnapping-case footage reviewed in three hours, but supplies no comparator, error assessment or reproducible evaluation. Future savings and rollout statements are projections.
- Limitations and uncertainty
- The allowed vendor-claim category is used for operator claims, not to imply the Home Office is a vendor. Current rollout and registry completion are unverified. No demonstrated general savings or crime reduction.
Historical prosecution experiment finds adverse recommendations despite exculpatory facts
A GPT-3.5-Turbo experiment found a tendency toward prosecution, including legally deficient scenarios; no racial disparity in recommendations was detected in this test.
Why it matters, evidence and limitations
- Why it matters
- New archive coverage of U.S. prosecutorial memo assistance, distinct from police report writing and judicial interviews. This is historical evidence, not a current-model evaluation.
- Evidence and measured results
- Researchers used 20 arrest reports and paired altered versions, varying role, prompt detail and racial identifiers. Repeated requests produced over 144,000 responses. More context changed recommendations and variability. Methods: https://knowledgehub.justiceinnovationlab.org/reports/ai-in-prosecution1/data.
- Limitations and uncertainty
- Small underlying case sample despite many responses; no prosecutor comparison group or observed case outcomes. Older model and uneven flaw severity limit generalization. No claim that present systems share these rates or that racial fairness is established.
COMPAS dataset reanalysis finds technical gains do not remove unequal error patterns
The reanalysis reports classifier-specific tradeoffs: optimization improved some predictive results, while unequal racial error patterns persisted.
Why it matters, evidence and limitations
- Why it matters
- Historical research newly added for state/local court and corrections risk-tool scrutiny. It does not evaluate a current deployed COMPAS system.
- Evidence and measured results
- The study uses 7,214 historical records and stratified ten-fold validation, comparing logistic regression, SVM and XGBoost configurations. PDF Table 6 reports XGBoost F1 of 60.05% for vanilla and 64.95% with optimization plus correlation removal. Table 5 shows reduced logistic-regression F1 under that mitigation. These are retrospective model metrics, not justice outcomes.
- Limitations and uncertainty
- Old single-county data, uncertain source matching and different preprocessing from earlier work limit replication and transfer. Proprietary model internals were unavailable. Re-arrest is not identical to offending. No causal harm-reduction result.
How to read this edition
Source findings, measured results and limitations come from the cited publications. Patterns, operating questions, role takeaways and implementation considerations are Lighthouse Advisory interpretation, stated as questions to validate locally rather than guaranteed outcomes. Vendor and operator claims are labeled as claims. Full research method.
- Vendor claim
- A supplier-provided assertion that has not been upgraded to independent evidence.
- Independent research
- Research conducted outside the implementing organization.
- Academic research
- Research produced through an academic institution or peer-reviewed venue.