{"resourceId":"cusp-scientific-forecasting-judge-limits-2026","versions":[{"version":"external-681497beb13c36967593dbf32170fc8ed6a802f78d7e0d74972d93c677bfe144","resource":{"id":"cusp-scientific-forecasting-judge-limits-2026","title":"CUSP exposes scientific forecasting limits and uncertainty in automated judging","organization":"University of Oxford, Stanford University, Allen Institute for AI, Sakana AI and collaborators","sector":"AI-assisted scientific research","geography":"United Kingdom, United States and Japan collaboration","publishedAt":"May 21, 2026, arXiv v1; foundational context","publicationDate":"2026-05-21","eventDate":null,"sourceName":"arXiv","sourceLabel":"Author preprint; methods, result tables and human-judge appendix inspected","sourceUrl":"https://arxiv.org/html/2605.22681v1","evidenceClass":"academic-research","outcomeClass":"cautionary","topics":["knowledge-work","developers-agents","governance-procurement","operating-model"],"finding":"Scientific approach recognition and forecasting reliability differ; automated grading also requires scrutiny.","sledRelevance":"Relevant to university research planning and evaluation; international benchmark findings do not establish local grant-selection or discovery outcomes.","evidence":"CUSP includes 4,760 milestones. Table 2 reports merged binary accuracy of 0.453–0.519 against a stated 0.50 chance baseline. Appendix E.2 reports Pearson r=0.34 between AI and human free-response scores on 60 examples reviewed by three evaluators.","architectureImplications":"Interpretation: separate literature retrieval, forecast generation and external evaluation.","governanceImplications":"Interpretation: require domain adjudication before using forecasts in research prioritization.","securityPrivacyImplications":"Interpretation: use authorized corpora and keep unpublished proposals out of unapproved retrieval or judging endpoints.","caveats":"Preprint and retrospective, selectively sourced benchmark with generated tasks. Per-task denominators vary; Appendix A.4 label-ratio wording is unclear. Human judge validation is small and covers two models. No prospective accuracy or institutional productivity result.","streamIds":["research"],"roles":{"sales":"Interpretation: Engage research development and principal investigators who use assistants to assess promising directions. Ask what decisions a forecast influences, which errors are costly and how claims are reviewed today. Offer a bounded evaluation of one planning workflow using public historical questions. The value hypothesis is better identification of unsupported confidence. Do not promise reliable discovery prediction or use this benchmark to justify automated funding decisions. Include domain-review cost and explain that a fluent rationale is not observed scientific progress.","engineering":"Interpretation: Build an evaluation harness with dated input snapshots, explicit retrieval boundaries and retained model versions. Keep proposed mechanisms separate from forecasts of realization. Use blinded domain review alongside deterministic scoring, and test unavailable-information cases. Proposed proof should compare the assisted workflow with an expert-led baseline on a preregistered local corpus, report task-specific denominators and measure calibration and reviewer disagreement. Audit the judging system independently; restricted hosting alone cannot resolve invalid questions or inappropriate acceptance criteria.","delivery":"Interpretation: A principal investigator should own scientific acceptance with research software support and independent domain reviewers. Agree the corpus, rubric, stopping rules and data permissions before adoption. Train staff to preserve uncertainty and log disagreements. Proposed acceptance criteria include a traceable source for every factual input, reproducible scoring, adjudication of every consequential disagreement and a documented human release decision. Track correction time alongside output quality. Risks include retrospective leakage, selection bias, overstated confidence and treating generated grades as authoritative."},"retrievedAt":"2026-09-14T03:03:16Z","enrichedAt":"2026-09-14T03:03:16Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: teach users to read uncertainty and preserve methodological review skills.","procurementImplications":"Interpretation: require replayable evaluation evidence rather than a single model ranking.","operatingModelImplications":"Interpretation: principal investigators own research decisions; software teams maintain evaluation traces.","updateExplanation":"Identifier and URL absent from the full archive. Older May evidence newly covered for research-prioritization and evaluator-calibration implications, not September news or an updated source.","sourceVerification":{"openedUrl":"https://arxiv.org/html/2605.22681v1","referenceExcerpt":"Because ground truth is not yet available, the Time Capsule is not used for accuracy evaluation","promptVersion":"sled-research-v3.2","model":null,"basis":"agent-reported inspection"}}}]}