{"resourceId":"madukoma-athletics-anomaly-benchmark-2026","versions":[{"version":"external-cfa902fb7b96c82d391c0661907ddd340d848077712bf8f2e6b54286b990e17b","resource":{"id":"madukoma-athletics-anomaly-benchmark-2026","title":"Sprint-screening preprint shows low confirmed-sanction precision and contextual gaps","organization":"Carnegie Mellon University Africa; Blessed Madukoma and Prasenjit Mitra","sector":"Collegiate athletics","geography":"International athletics data; authors based in Rwanda","publishedAt":"April 23, 2026","publicationDate":"2026-04-23","eventDate":null,"sourceName":"arXiv:2604.21953v1","sourceLabel":"Original academic preprint; not a collegiate deployment evaluation","sourceUrl":"https://arxiv.org/html/2604.21953v1","evidenceClass":"academic-research","outcomeClass":"cautionary","topics":["knowledge-work","infrastructure","data-security","governance-procurement","operating-model"],"finding":"Retrospective screening results support caution about treating performance anomalies as evidence of wrongdoing.","sledRelevance":"Transferable to college track analytics oversight; the paper includes an NCAA race illustration but no NCAA-specific validation.","evidence":"Table III benchmarks eight methods on 31,604 100-m athletes, 25 with recorded sanctions. Excess Performance flags 226, including 2 sanctioned athletes: precision .009, recall .080, F1 .016. Five methods find no sanctioned athletes. The implemented baseline is static, despite broader trajectory language.","architectureImplications":"Interpretation: preserve source records, missingness and model versions in a reviewable analysis service; benchmark simple methods before complex ones.","governanceImplications":"Interpretation: confine any exploratory use to authorized expert review; never turn an anomaly into an automatic accusation or eligibility decision.","securityPrivacyImplications":"Interpretation: public results joined to sensitive inferences still require access limits and a correction process.","caveats":"Preprint; incomplete sanction labels, zero-imputed missing wind and unadjusted altitude. No prospective NCAA evaluation. The authors' claim that incomplete labels make recall a conservative lower bound is not established; no guilt inference is warranted.","streamIds":["college-athletics"],"roles":{"sales":"Interpretation: The relevant problem is whether an analytics service can support careful review without generating disproportionate investigative burden. Engage athletics research, athlete representatives, compliance leadership and institutional privacy staff before discussing a use case. Ask what decision is authorized, which independent evidence is available and whether a screening tool is needed at all. A bounded engagement could audit historical data quality and evaluation methodology without producing individual accusations. Its value hypothesis is better understanding of error and workload, not more sanctions. The retrospective benchmark does not establish savings, NCAA suitability or a market opportunity at any named institution.","engineering":"Interpretation: Start with a reproducible offline benchmark and a restricted review interface. Preserve athlete-identity provenance, event context, label dates and explicit missing-value indicators. Validate separately by event and data completeness, and compare simple baselines under the same protocol. Prevent future information from leaking into a purported prospective test. Measure review workload at a fixed alert budget, report uncertainty from sparse labels and seek independent expert adjudication. Access to public records does not justify unrestricted dissemination of inferences. Infrastructure placement remains an institutional choice; autonomous investigative agents and automatic sanctions have no supported role in this evidence.","delivery":"Interpretation: A designated research lead should own any feasibility study, with privacy and athlete-welfare review before data linkage. Establish a reproducible dataset snapshot, contextual-quality checks and a written restriction on downstream use. Train reviewers to distinguish anomaly, recorded sanction and unverified status. Proposed acceptance criteria include full provenance for reviewed records, documented missingness, reproducible aggregate metrics and no automated consequential action. Stop if reviewers cannot explain errors or if lawful authority and data rights remain unresolved. Evaluate adoption through reviewer understanding, not alert volume. Incomplete labels and reputational harm are central risks; retrospective success cannot authorize production use."},"retrievedAt":"2026-09-12T03:01:12Z","enrichedAt":"2026-09-12T03:01:56Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: make uncertainty understandable and fund qualified reviewers rather than merely exposing a score.","procurementImplications":"Interpretation: demand prospective validation and contextual-data coverage before considering operational acquisition.","operatingModelImplications":"Interpretation: assign accountable research oversight and maintain a strict separation between analysis and authorized enforcement.","updateExplanation":"Absent from all 217 canonical archive records scanned at offsets 0, 100 and 200, including URL and related-title checks. Newly archived historical evidence; no new-since-last-run publication or update to an existing resource is claimed.","sourceVerification":{"openedUrl":"https://arxiv.org/html/2604.21953v1","referenceExcerpt":"This per-athlete normalization accounts for inter-athlete variability but uses a static baseline, ignoring career progression.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}