{"resourceId":"reproducibility-cost-documentation-proxy-manuscript-2026","versions":[{"version":"external-7d9d13ad6288d7325517d08fbec43f073a0b1eaf0ccb861b1dd5b6ea27f17557","resource":{"id":"reproducibility-cost-documentation-proxy-manuscript-2026","title":"Reproducibility study separates documentation quality from actual reproduction cost","organization":"Anonymous authors in inspected review manuscript","sector":"AI research methods and reproducibility","geography":"International AI/ML publication sample; author affiliations withheld in manuscript","publishedAt":"Undated manuscript marked under review at ICLR 2026","publicationDate":null,"eventDate":null,"sourceName":"OpenReview","sourceLabel":"Public academic review manuscript; methods and tables inspected","sourceUrl":"https://openreview.net/pdf/88fdc3b21a41a3c4e7f068ccd07708fdc42dad7c.pdf","evidenceClass":"academic-research","outcomeClass":"cautionary","topics":["developers-agents","knowledge-work","infrastructure","data-security","governance-procurement","operating-model"],"finding":"Documentation scoring reveals barriers to reuse but cannot establish actual reproduction time or scientific validity.","sledRelevance":"Useful for university artifact-review services; the international publication sample is not a U.S. institutional deployment evaluation.","evidence":"The study analyzes 918 empirical papers from 1,061 sampled across seven venues in 2022–2024. A second review covers 46 papers. Expertise scoring was dropped for weak reliability; venue comparisons use documentation rubrics, not timed reproductions.","architectureImplications":"Interpretation: preserve execution environments and artifacts alongside documentation scores.","governanceImplications":"Interpretation: maintain separate acceptance decisions for readable documentation and independently reproduced conclusions.","securityPrivacyImplications":"Interpretation: reproducibility does not authorize unrestricted release of private datasets; document lawful access paths.","caveats":"Older anonymous manuscript, not the inaccessible journal version. One primary reviewer and limited second review constrain inference. No causal estimate of checklist effectiveness or measured labor savings. Screenshot failed; PDF text, tables and limitations were readable.","streamIds":["research"],"roles":{"sales":"Interpretation: Engage research integrity officers, library repository teams and research software engineers where published analyses are expensive to reconstruct. Ask what usually blocks a rerun, who pays for correction, and whether current checks assess documents or executed results. Offer an artifact-readiness assessment followed by a small independent reproduction pilot. The value hypothesis is identifying preventable handoff failures before publication. Do not turn the paper's rubric differences into percentage labor savings or rank local departments by venue. Establish the institution's own baseline for reviewer time, unavailable inputs and unresolved discrepancies before proposing expansion.","engineering":"Interpretation: Build a manifest linking data permissions, code, configuration, environment and expected outputs. Verify that links resolve and that the package can run in a clean approved environment. Treat generated documentation as an unverified draft, and capture actual commands and outputs outside the agent's editable workspace. Compare documentation scores with time and error outcomes on local packages to test whether the proxy is useful. Retain missing-input controls and check data export restrictions. The manuscript's limited rater coverage makes a locally calibrated rubric and independent adjudication more useful than copying its aggregate percentages.","delivery":"Interpretation: Assign a repository steward for long-term artifacts and a research integrity owner for the acceptance rubric. Dependencies include permitted materials, versioned environments and reviewers with relevant methods expertise. Train investigators to prepare a handoff that another team can follow, then schedule an independent replay before release. Proposed acceptance criteria include retrievable artifacts, an explained outcome for every attempted reproduction, and agreement on any scientific tolerance. Track correction effort separately from compute expense. Risks include brittle external links, inaccessible private data, inconsistent reviewer scoring and treating successful execution as confirmation of the underlying hypothesis."},"retrievedAt":"2026-09-11T03:01:22Z","enrichedAt":"2026-09-11T03:04:01Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: budget reviewer expertise and supply accessible artifact inventories; a score cannot substitute for disciplinary judgment.","procurementImplications":"Interpretation: require exportable environments and durable artifact storage; evaluate actual local recovery effort before estimating value.","operatingModelImplications":"Interpretation: research software staff maintain replay tools; independent domain reviewers accept scientific claims.","updateExplanation":"Manuscript URL and study title absent from full archive. Recent journal discovery motivated inspection of the older public manuscript; no claim of version equivalence or newly published manuscript.","sourceVerification":{"openedUrl":"https://openreview.net/pdf/88fdc3b21a41a3c4e7f068ccd07708fdc42dad7c.pdf","referenceExcerpt":"Thus, this is an indirect measure of cost","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}