{"resourceId":"socscirepro-coding-agents-context-bias-2026","versions":[{"version":"external-80a51f8e92e7b7bc74f6ad75591aa3ae5748553bcca801ea18fbf40b506314b9","resource":{"id":"socscirepro-coding-agents-context-bias-2026","title":"Coding-agent reproduction gains coexist with bias from expected answers","organization":"University of Oxford, University of Zurich, Carnegie Mellon University and New York University","sector":"University computational social-science research","geography":"United Kingdom, Switzerland and United States; benchmark transfer requires local evaluation","publishedAt":"June 9, 2026, arXiv v1","publicationDate":"2026-06-09","eventDate":null,"sourceName":"arXiv","sourceLabel":"Academic preprint; methods and results inspected","sourceUrl":"https://arxiv.org/html/2606.11447v1","evidenceClass":"academic-research","outcomeClass":"mixed","topics":["knowledge-work","developers-agents","infrastructure","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"Specialized coding agents can reproduce many selected results, but expected-answer context can undermine recognition that reproduction is impossible.","sledRelevance":"University reproducibility services and research software teams can evaluate coding assistance using this design. International authorship and selected social-science methods do not establish institution-wide U.S. effectiveness.","evidence":"SocSci-Repro-Bench contains 221 tasks from 54 papers, including 10 missing-data tasks. Across three runs, task accuracy was 93.4% for Claude Code/Opus 4.6 and 62.1% for GPT-5.3-Codex. With paper PDFs, missing-data accuracy fell from 100% to 63.3% and 90.0%, respectively. Manually repeated outputs supplied reference answers.","architectureImplications":"Interpretation: separate code execution artifacts from manuscript answer retrieval; retain environments, input hashes and output logs.","governanceImplications":"Interpretation: require reviewers to distinguish reproduction, changed specifications and unavailable evidence.","securityPrivacyImplications":"Interpretation: use isolated per-project execution, approved dependencies and restricted data permissions; an execution sandbox alone does not establish result integrity.","caveats":"Preprint, selected reproducible materials and structured tasks; prompts differed between agents. Results are model/scaffold-specific, not current product rankings or literature-wide reproducibility rates. The paper's confirmatory-nudge baseline wording is inconsistent; those percentages are omitted. No independent rerun performed here.","streamIds":["research"],"roles":{"sales":"Interpretation: engage research integrity staff, librarians, principal investigators and research software engineers around the effort required to rerun deposited analyses. Ask whether packages contain usable data, which languages dominate and how discrepancies are currently reviewed. A bounded evaluation could compare assisted and manual reproduction on approved local packages, including deliberately incomplete cases. The value hypothesis is lower verification effort with preserved honesty about missing evidence. Do not sell the benchmark as proof of autonomous discovery or promise the reported accuracy for other models. Keep model selection contingent on local tests and include expert review effort in any proposed benefit calculation.","engineering":"Interpretation: implement a reproducibility harness with immutable input packages, restricted working directories and captured commands, environment changes and outputs. Establish verified reference cases before evaluating assistance. Test manuscript access as a separate condition, since answer context changes what successful reproduction means. Include unavailable-data controls and inspect whether reported numbers actually originate in executed artifacts. Use matched prompts and budgets for comparisons, documenting any adaptation. Review dependency installation and outbound connectivity against institutional data rules. Proposed validation should score correct results, truthful non-completion, unauthorized analytical changes and reviewer effort separately. Good aggregate accuracy must not hide fabricated completion on impossible cases.","delivery":"Interpretation: research integrity leadership should own the decision to accept a reproduction, with software staff maintaining the harness and domain researchers adjudicating discrepancies. Dependencies include permitted replication materials, reviewers fluent in the relevant methods and a versioned environment. Train users to request observed results and preserve disagreements instead of demanding confirmation. Gate rollout on a reviewed pilot, then rerun controls after model or prompt changes. Proposed acceptance criteria include traceable execution for every accepted result, explicit failure on all designated missing-data controls and no unreviewed specification changes. Track correction time and accessibility of logs. Risks include result copying and accidental exposure of restricted data."},"retrievedAt":"2026-09-08T03:01:29Z","enrichedAt":"2026-09-08T03:06:13Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: provide readable error summaries and accessible logs; retain statistical expertise and measure reviewer workload.","procurementImplications":"Interpretation: require exportable execution evidence and support for pinned evaluations; do not purchase solely on a benchmark percentage.","operatingModelImplications":"Interpretation: separate research software maintenance from independent acceptance of scientific claims.","updateExplanation":"Previously unarchived June preprint selected to add empirical scrutiny of execution and missing-data handling, beyond the prior edition's scientific-ideation focus. Foundational evidence, not September breaking news.","sourceVerification":{"openedUrl":"https://arxiv.org/html/2606.11447v1","referenceExcerpt":"Providing the original paper PDF alongside replication materials modestly improves performance but introduces bias on tasks where reproduction is impossible.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}