Lighthouse AdvisorySLED AI Adoption Intelligence
← Back to results

From the Research edition of September 7, 2026

Academic researchMixedNewly relevant · Jun 2026

Coding-agent reproduction gains coexist with bias from expected answers

University of Oxford, University of Zurich, Carnegie Mellon University and New York University · University computational social-science research · United Kingdom, Switzerland and United States; benchmark transfer requires local evaluation

Publisher
arXiv
Original publication
June 9, 2026, arXiv v1
Source retrieved
2026-09-08
Read original source

What happened

Specialized coding agents can reproduce many selected results, but expected-answer context can undermine recognition that reproduction is impossible.

Why it matters

University reproducibility services and research software teams can evaluate coding assistance using this design. International authorship and selected social-science methods do not establish institution-wide U.S. effectiveness.

Evidence and measured results

SocSci-Repro-Bench contains 221 tasks from 54 papers, including 10 missing-data tasks. Across three runs, task accuracy was 93.4% for Claude Code/Opus 4.6 and 62.1% for GPT-5.3-Codex. With paper PDFs, missing-data accuracy fell from 100% to 63.3% and 90.0%, respectively. Manually repeated outputs supplied reference answers.

Limitations and uncertainty

Preprint, selected reproducible materials and structured tasks; prompts differed between agents. Results are model/scaffold-specific, not current product rankings or literature-wide reproducibility rates. The paper's confirmatory-nudge baseline wording is inconsistent; those percentages are omitted. No independent rerun performed here.

Put this evidence to work

Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-08; this does not change the original publication date. Labels below come from the analysis itself.

Sales

Role takeaway

Engage research integrity staff, librarians, principal investigators and research software engineers around the effort required to rerun deposited analyses. Ask whether packages contain usable data, which languages dominate and how discrepancies are currently reviewed. A bounded evaluation could compare assisted and manual reproduction on approved local packages, including deliberately incomplete cases. The value hypothesis is lower verification effort with preserved honesty about missing evidence. Do not sell the benchmark as proof of autonomous discovery or promise the reported accuracy for other models. Keep model selection contingent on local tests and include expert review effort in any proposed benefit calculation.

Pre-sales engineering

Role takeaway

Implement a reproducibility harness with immutable input packages, restricted working directories and captured commands, environment changes and outputs. Establish verified reference cases before evaluating assistance. Test manuscript access as a separate condition, since answer context changes what successful reproduction means. Include unavailable-data controls and inspect whether reported numbers actually originate in executed artifacts. Use matched prompts and budgets for comparisons, documenting any adaptation. Review dependency installation and outbound connectivity against institutional data rules. Proposed validation should score correct results, truthful non-completion, unauthorized analytical changes and reviewer effort separately. Good aggregate accuracy must not hide fabricated completion on impossible cases.

Delivery

Role takeaway

Research integrity leadership should own the decision to accept a reproduction, with software staff maintaining the harness and domain researchers adjudicating discrepancies. Dependencies include permitted replication materials, reviewers fluent in the relevant methods and a versioned environment. Train users to request observed results and preserve disagreements instead of demanding confirmation. Gate rollout on a reviewed pilot, then rerun controls after model or prompt changes. Proposed acceptance criteria include traceable execution for every accepted result, explicit failure on all designated missing-data controls and no unreviewed specification changes. Track correction time and accessibility of logs. Risks include result copying and accidental exposure of restricted data.

Implementation considerations

Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.

Architecture and integration

What must connect, and where does the AI sit in the workflow?

Separate code execution artifacts from manuscript answer retrieval; retain environments, input hashes and output logs.

Governance

Who approves, reviews and stays accountable for outcomes?

Require reviewers to distinguish reproduction, changed specifications and unavailable evidence.

Security and privacy

What data, permissions and controls need testing?

Use isolated per-project execution, approved dependencies and restricted data permissions; an execution sandbox alone does not establish result integrity.

Accessibility and workforce

Who is affected, and what skills or accommodations follow?

Provide readable error summaries and accessible logs; retain statistical expertise and measure reviewer workload.

Procurement

What should contracts, pricing and exit terms secure?

Require exportable execution evidence and support for pinned evaluations; do not purchase solely on a benchmark percentage.

Operating model

Which teams own the service once it runs?

Separate research software maintenance from independent acceptance of scientific claims.

What changed

Previously unarchived June preprint selected to add empirical scrutiny of execution and missing-data handling, beyond the prior edition's scientific-ideation focus. Foundational evidence, not September breaking news.

Publication history

  1. 2026-09-07Research · Issue 023 resources
Read preserved resource versions (JSON)

Stable resource ID: socscirepro-coding-agents-context-bias-2026