{"resourceId":"compas-reanalysis-mitigation-tradeoffs-2026","versions":[{"version":"external-6883521f27faf9a39d5d68e393c81ef1cd7c3ea5062fa00014fdbca93471c213","resource":{"id":"compas-reanalysis-mitigation-tradeoffs-2026","title":"COMPAS dataset reanalysis finds technical gains do not remove unequal error patterns","organization":"Sarah Kienzle, Gissel Velarde and Brian Gannon; IU International University of Applied Sciences","sector":"Criminal justice risk assessment; courts and corrections","geography":"Germany-led analysis of Broward County, Florida historical data","publishedAt":"April 23, 2026; experiments June–September 2024","publicationDate":"2026-04-23","eventDate":null,"sourceName":"AI and Ethics, Springer Nature","sourceLabel":"Peer-reviewed research; full article, methods and PDF tables inspected","sourceUrl":"https://link.springer.com/article/10.1007/s43681-026-01116-0","evidenceClass":"academic-research","outcomeClass":"mixed","topics":["data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"The reanalysis reports classifier-specific tradeoffs: optimization improved some predictive results, while unequal racial error patterns persisted.","sledRelevance":"Historical research newly added for state/local court and corrections risk-tool scrutiny. It does not evaluate a current deployed COMPAS system.","evidence":"The study uses 7,214 historical records and stratified ten-fold validation, comparing logistic regression, SVM and XGBoost configurations. PDF Table 6 reports XGBoost F1 of 60.05% for vanilla and 64.95% with optimization plus correlation removal. Table 5 shows reduced logistic-regression F1 under that mitigation. These are retrospective model metrics, not justice outcomes.","architectureImplications":"Interpretation: Keep versioned datasets, feature transformations and evaluation outputs reproducible; separate experimental scores from operational case records.","governanceImplications":"Interpretation: Define outcome labels, acceptable errors and decision authority before evaluating a model.","securityPrivacyImplications":"Interpretation: Restrict linked justice records and auditing attributes to approved evaluators; test access and retention.","caveats":"Old single-county data, uncertain source matching and different preprocessing from earlier work limit replication and transfer. Proprietary model internals were unavailable. Re-arrest is not identical to offending. No causal harm-reduction result.","streamIds":["public-safety"],"roles":{"sales":"Interpretation: Court administrators, probation leaders, statisticians, civil-rights reviewers and procurement staff need to know what an assessment actually predicts and whose errors it increases. Ask which decision the score changes, how its target is defined, and whether validation matches current local practice. Offer a bounded data-and-evaluation review before considering replacement software. The value hypothesis is clearer procurement and operating decisions, with explicit tradeoffs rather than a universal fairness claim. This reanalysis supplies a reason to examine error distribution and data provenance. It does not establish an opportunity to deploy its experimental classifier, guaranteed reductions in recidivism, or that every actuarial tool performs alike.","engineering":"Interpretation: Fit is a controlled retrospective validation pipeline around a precisely defined outcome. Prerequisites include lawful record access, reliable linkage, time-stamped features and a local reference cohort. Compare a transparent baseline with candidate models using held-out periods, subgroup errors and calibration; check preprocessing for leakage. Report missing labels and changes in supervision or policing that alter the target. Cloud, on-premises and hybrid hosting can all support this work if controls are enforceable; the paper does not choose among them. Generative copilots and autonomous agents have limited relevance to this tabular experiment. A useful proof of value tests reproducibility and local validity before any score reaches decision-makers.","delivery":"Interpretation: An accountable justice programme owner should partner with records specialists, a statistician, IT security and affected-community representatives. Document the existing decision process, reconcile data, agree on error definitions and run a shadow evaluation. Train staff to distinguish risk estimates from facts and to record overrides. Governance checkpoints should approve data use, evaluation design and any operational progression. Proposed acceptance: all reported metrics can be reproduced from versioned inputs, subgroup limitations are documented, and leadership explicitly resolves tradeoffs before use. Set numerical thresholds locally in advance. These are proposed criteria, not observed results. Key risks are stale labels, hidden linkage errors and converting a better aggregate score into an unjustified decision policy."},"retrievedAt":"2026-09-12T03:02:21Z","enrichedAt":"2026-09-12T03:04:22Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Provide plain-language explanations and a usable correction route for affected people; train staff to interpret uncertainty.","procurementImplications":"Interpretation: Require current local validation and auditable transformations; do not accept removal of sensitive features as proof of fairness.","operatingModelImplications":"Interpretation: Maintain separate responsibility for data quality, statistical evaluation and consequential decisions.","updateExplanation":"New-to-archive source adds a classifier-mitigation comparison to prior corrections guidance. Historical data and experiment dates are explicit; no post-last-run development is asserted.","sourceVerification":{"openedUrl":"https://link.springer.com/article/10.1007/s43681-026-01116-0","referenceExcerpt":"better predictive performance did not completely eliminate uneven error distribution across racial groups.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}