From the Research edition of September 7, 2026
Coding-agent reproduction gains coexist with bias from expected answers
University of Oxford, University of Zurich, Carnegie Mellon University and New York University · University computational social-science research · United Kingdom, Switzerland and United States; benchmark transfer requires local evaluation
- Publisher
- arXiv
- Original publication
- June 9, 2026, arXiv v1
- Source retrieved
- 2026-09-08
What happened
Specialized coding agents can reproduce many selected results, but expected-answer context can undermine recognition that reproduction is impossible.
Why it matters
University reproducibility services and research software teams can evaluate coding assistance using this design. International authorship and selected social-science methods do not establish institution-wide U.S. effectiveness.
Evidence and measured results
SocSci-Repro-Bench contains 221 tasks from 54 papers, including 10 missing-data tasks. Across three runs, task accuracy was 93.4% for Claude Code/Opus 4.6 and 62.1% for GPT-5.3-Codex. With paper PDFs, missing-data accuracy fell from 100% to 63.3% and 90.0%, respectively. Manually repeated outputs supplied reference answers.
Limitations and uncertainty
Preprint, selected reproducible materials and structured tasks; prompts differed between agents. Results are model/scaffold-specific, not current product rankings or literature-wide reproducibility rates. The paper's confirmatory-nudge baseline wording is inconsistent; those percentages are omitted. No independent rerun performed here.
Put this evidence to work
Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-08; this does not change the original publication date. Labels below come from the analysis itself.
Sales
Role takeaway
Engage research integrity staff, librarians, principal investigators and research software engineers around the effort required to rerun deposited analyses. Ask whether packages contain usable data, which languages dominate and how discrepancies are currently reviewed. A bounded evaluation could compare assisted and manual reproduction on approved local packages, including deliberately incomplete cases. The value hypothesis is lower verification effort with preserved honesty about missing evidence. Do not sell the benchmark as proof of autonomous discovery or promise the reported accuracy for other models. Keep model selection contingent on local tests and include expert review effort in any proposed benefit calculation.
Pre-sales engineering
Role takeaway
Implement a reproducibility harness with immutable input packages, restricted working directories and captured commands, environment changes and outputs. Establish verified reference cases before evaluating assistance. Test manuscript access as a separate condition, since answer context changes what successful reproduction means. Include unavailable-data controls and inspect whether reported numbers actually originate in executed artifacts. Use matched prompts and budgets for comparisons, documenting any adaptation. Review dependency installation and outbound connectivity against institutional data rules. Proposed validation should score correct results, truthful non-completion, unauthorized analytical changes and reviewer effort separately. Good aggregate accuracy must not hide fabricated completion on impossible cases.
Delivery
Role takeaway
Research integrity leadership should own the decision to accept a reproduction, with software staff maintaining the harness and domain researchers adjudicating discrepancies. Dependencies include permitted replication materials, reviewers fluent in the relevant methods and a versioned environment. Train users to request observed results and preserve disagreements instead of demanding confirmation. Gate rollout on a reviewed pilot, then rerun controls after model or prompt changes. Proposed acceptance criteria include traceable execution for every accepted result, explicit failure on all designated missing-data controls and no unreviewed specification changes. Track correction time and accessibility of logs. Risks include result copying and accidental exposure of restricted data.
Implementation considerations
Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.
Architecture and integration
What must connect, and where does the AI sit in the workflow?
Separate code execution artifacts from manuscript answer retrieval; retain environments, input hashes and output logs.
Governance
Who approves, reviews and stays accountable for outcomes?
Require reviewers to distinguish reproduction, changed specifications and unavailable evidence.
Security and privacy
What data, permissions and controls need testing?
Use isolated per-project execution, approved dependencies and restricted data permissions; an execution sandbox alone does not establish result integrity.
Accessibility and workforce
Who is affected, and what skills or accommodations follow?
Provide readable error summaries and accessible logs; retain statistical expertise and measure reviewer workload.
Procurement
What should contracts, pricing and exit terms secure?
Require exportable execution evidence and support for pinned evaluations; do not purchase solely on a benchmark percentage.
Operating model
Which teams own the service once it runs?
Separate research software maintenance from independent acceptance of scientific claims.
What changed
Previously unarchived June preprint selected to add empirical scrutiny of execution and missing-data handling, beyond the prior edition's scientific-ideation focus. Foundational evidence, not September breaking news.
Publication history
- 2026-09-07Research · Issue 023 resources
Stable resource ID: socscirepro-coding-agents-context-bias-2026