Lighthouse AdvisorySLED AI Adoption Intelligence
← Back to results

From the Research edition of September 12, 2026

Academic researchCautionaryNew this fortnight

AgentActionBench exposes the gap between generated code and executed research

Hanhua Hong and colleagues; University of Manchester and collaborating institutions · Academic research reproducibility · UK, China and Vietnam collaborations; international paper sample

Publisher
arXiv
Original publication
September 10, 2026
Source retrieved
2026-09-13
Read original source

What happened

Action-based evaluation reveals execution and result-verification weaknesses.

Why it matters

Transferable evaluation design for university research software teams; no direct U.S. campus deployment evidence.

Evidence and measured results

150 papers: 120 ML and 30 AI4Science; 15 form the human-annotated development subset. Three teams plus a Codex–GPT-5.4 baseline are evaluated. Table 3 reports weighted rubric scores of 49.64% for the leader and 24.19% for the baseline, not percentages of fully reproduced papers. The leaderboard excludes the development subset.

Limitations and uncertainty

Preprint; GPT-4o-mini judging can hallucinate. Narrow scientific coverage, generated rubrics and limited human validation constrain generalization. No institutional labor baseline or production outcome.

Put this evidence to work

Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-13; this does not change the original publication date. Labels below come from the analysis itself.

Sales

Role takeaway

Discuss unreliable paper-to-code handoffs with research computing leaders, principal investigators and research integrity staff. Ask how often generated implementations actually run, what counts as a faithful result and who resolves disputed evaluations. Offer a bounded reproduction assessment on a small, permitted local paper set. The value hypothesis is identifying failure stages before researchers build on unsupported outputs. Do not present rubric scores as discovery success rates or promise reduced staffing. Include reviewer time and dependency repair in the scope, and establish a manual baseline for total effort before estimating economic value.

Pre-sales engineering

Role takeaway

Build a sandbox with mediated file and command access, immutable external logging and versioned environments. Integrate campus identity and approved storage; select cloud, local or hybrid execution according to data permissions and workload requirements. Preserve exit status, outputs and failed attempts separately from the agent's narrative. A proof of value should compare trace-based judgments with independent human reruns, deliberately including plausible code that was never executed. Calibrate any automated judge locally because agreement between rubric sources cannot guarantee correct judgments. Restrict outbound access and confirm logs cannot disclose confidential data.

Delivery

Role takeaway

Assign a research software engineering lead and an independent disciplinary reviewer. Dependencies include lawful dataset access, runnable environments, an agreed rubric and protected review time. Onboard researchers through an accessible example showing the difference between writing code and validating its result. Proposed acceptance criteria are a complete action trace for every attempted task, documented treatment of each execution failure and an independent comparison against the expected scientific result. Track elapsed time, reviewer effort and unresolved discrepancies. These are proposed measures, not benchmark observations. Risks include judge errors, environment drift and overfitting to checklist wording.

Implementation considerations

Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.

Architecture and integration

What must connect, and where does the AI sit in the workflow?

Record actual execution outside agent-editable artifacts and retain dependency versions.

Governance

Who approves, reviews and stays accountable for outcomes?

Distinguish rubric compliance, successful reruns and independently accepted scientific conclusions.

Security and privacy

What data, permissions and controls need testing?

Isolate generated code; redact secrets and restricted research data from retained logs.

Accessibility and workforce

Who is affected, and what skills or accommodations follow?

Provide readable trace summaries and train domain reviewers to assess execution evidence.

Procurement

What should contracts, pricing and exit terms secure?

Require exportable action logs and price human adjudication alongside model and compute costs.

Operating model

Which teams own the service once it runs?

Research software engineers own replay infrastructure; domain scientists own result acceptance.

What changed

URL and AgentActionBench title absent from the 247-resource full archive, paginated at offsets 0, 100 and 200. Newly covered September preprint; not claimed to have appeared since the last completed run.

Publication history

  1. 2026-09-12Research · Issue 073 resources
Read preserved resource versions (JSON)

Stable resource ID: agentactionbench-execution-evidence-202609