From the Research edition of September 12, 2026
AgentActionBench exposes the gap between generated code and executed research
Hanhua Hong and colleagues; University of Manchester and collaborating institutions · Academic research reproducibility · UK, China and Vietnam collaborations; international paper sample
- Publisher
- arXiv
- Original publication
- September 10, 2026
- Source retrieved
- 2026-09-13
What happened
Action-based evaluation reveals execution and result-verification weaknesses.
Why it matters
Transferable evaluation design for university research software teams; no direct U.S. campus deployment evidence.
Evidence and measured results
150 papers: 120 ML and 30 AI4Science; 15 form the human-annotated development subset. Three teams plus a Codex–GPT-5.4 baseline are evaluated. Table 3 reports weighted rubric scores of 49.64% for the leader and 24.19% for the baseline, not percentages of fully reproduced papers. The leaderboard excludes the development subset.
Limitations and uncertainty
Preprint; GPT-4o-mini judging can hallucinate. Narrow scientific coverage, generated rubrics and limited human validation constrain generalization. No institutional labor baseline or production outcome.
Put this evidence to work
Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-13; this does not change the original publication date. Labels below come from the analysis itself.
Sales
Role takeaway
Discuss unreliable paper-to-code handoffs with research computing leaders, principal investigators and research integrity staff. Ask how often generated implementations actually run, what counts as a faithful result and who resolves disputed evaluations. Offer a bounded reproduction assessment on a small, permitted local paper set. The value hypothesis is identifying failure stages before researchers build on unsupported outputs. Do not present rubric scores as discovery success rates or promise reduced staffing. Include reviewer time and dependency repair in the scope, and establish a manual baseline for total effort before estimating economic value.
Pre-sales engineering
Role takeaway
Build a sandbox with mediated file and command access, immutable external logging and versioned environments. Integrate campus identity and approved storage; select cloud, local or hybrid execution according to data permissions and workload requirements. Preserve exit status, outputs and failed attempts separately from the agent's narrative. A proof of value should compare trace-based judgments with independent human reruns, deliberately including plausible code that was never executed. Calibrate any automated judge locally because agreement between rubric sources cannot guarantee correct judgments. Restrict outbound access and confirm logs cannot disclose confidential data.
Delivery
Role takeaway
Assign a research software engineering lead and an independent disciplinary reviewer. Dependencies include lawful dataset access, runnable environments, an agreed rubric and protected review time. Onboard researchers through an accessible example showing the difference between writing code and validating its result. Proposed acceptance criteria are a complete action trace for every attempted task, documented treatment of each execution failure and an independent comparison against the expected scientific result. Track elapsed time, reviewer effort and unresolved discrepancies. These are proposed measures, not benchmark observations. Risks include judge errors, environment drift and overfitting to checklist wording.
Implementation considerations
Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.
Architecture and integration
What must connect, and where does the AI sit in the workflow?
Record actual execution outside agent-editable artifacts and retain dependency versions.
Governance
Who approves, reviews and stays accountable for outcomes?
Distinguish rubric compliance, successful reruns and independently accepted scientific conclusions.
Security and privacy
What data, permissions and controls need testing?
Isolate generated code; redact secrets and restricted research data from retained logs.
Accessibility and workforce
Who is affected, and what skills or accommodations follow?
Provide readable trace summaries and train domain reviewers to assess execution evidence.
Procurement
What should contracts, pricing and exit terms secure?
Require exportable action logs and price human adjudication alongside model and compute costs.
Operating model
Which teams own the service once it runs?
Research software engineers own replay infrastructure; domain scientists own result acceptance.
What changed
URL and AgentActionBench title absent from the 247-resource full archive, paginated at offsets 0, 100 and 200. Newly covered September preprint; not claimed to have appeared since the last completed run.
Publication history
- 2026-09-12Research · Issue 073 resources
Stable resource ID: agentactionbench-execution-evidence-202609