{"resourceId":"agentactionbench-execution-evidence-202609","versions":[{"version":"external-c9aa39e1d34b8205ab2c7838900cfa56ae70bddc1e22076e88b47cbd35c22a77","resource":{"id":"agentactionbench-execution-evidence-202609","title":"AgentActionBench exposes the gap between generated code and executed research","organization":"Hanhua Hong and colleagues; University of Manchester and collaborating institutions","sector":"Academic research reproducibility","geography":"UK, China and Vietnam collaborations; international paper sample","publishedAt":"September 10, 2026","publicationDate":"2026-09-10","eventDate":null,"sourceName":"arXiv","sourceLabel":"Academic shared-task preprint; methods, tables and limitations inspected","sourceUrl":"https://arxiv.org/html/2609.11117v1","evidenceClass":"academic-research","outcomeClass":"cautionary","topics":["developers-agents","knowledge-work","infrastructure","operating-model"],"finding":"Action-based evaluation reveals execution and result-verification weaknesses.","sledRelevance":"Transferable evaluation design for university research software teams; no direct U.S. campus deployment evidence.","evidence":"150 papers: 120 ML and 30 AI4Science; 15 form the human-annotated development subset. Three teams plus a Codex–GPT-5.4 baseline are evaluated. Table 3 reports weighted rubric scores of 49.64% for the leader and 24.19% for the baseline, not percentages of fully reproduced papers. The leaderboard excludes the development subset.","architectureImplications":"Interpretation: record actual execution outside agent-editable artifacts and retain dependency versions.","governanceImplications":"Interpretation: distinguish rubric compliance, successful reruns and independently accepted scientific conclusions.","securityPrivacyImplications":"Interpretation: isolate generated code; redact secrets and restricted research data from retained logs.","caveats":"Preprint; GPT-4o-mini judging can hallucinate. Narrow scientific coverage, generated rubrics and limited human validation constrain generalization. No institutional labor baseline or production outcome.","streamIds":["research"],"roles":{"sales":"Interpretation: Discuss unreliable paper-to-code handoffs with research computing leaders, principal investigators and research integrity staff. Ask how often generated implementations actually run, what counts as a faithful result and who resolves disputed evaluations. Offer a bounded reproduction assessment on a small, permitted local paper set. The value hypothesis is identifying failure stages before researchers build on unsupported outputs. Do not present rubric scores as discovery success rates or promise reduced staffing. Include reviewer time and dependency repair in the scope, and establish a manual baseline for total effort before estimating economic value.","engineering":"Interpretation: Build a sandbox with mediated file and command access, immutable external logging and versioned environments. Integrate campus identity and approved storage; select cloud, local or hybrid execution according to data permissions and workload requirements. Preserve exit status, outputs and failed attempts separately from the agent's narrative. A proof of value should compare trace-based judgments with independent human reruns, deliberately including plausible code that was never executed. Calibrate any automated judge locally because agreement between rubric sources cannot guarantee correct judgments. Restrict outbound access and confirm logs cannot disclose confidential data.","delivery":"Interpretation: Assign a research software engineering lead and an independent disciplinary reviewer. Dependencies include lawful dataset access, runnable environments, an agreed rubric and protected review time. Onboard researchers through an accessible example showing the difference between writing code and validating its result. Proposed acceptance criteria are a complete action trace for every attempted task, documented treatment of each execution failure and an independent comparison against the expected scientific result. Track elapsed time, reviewer effort and unresolved discrepancies. These are proposed measures, not benchmark observations. Risks include judge errors, environment drift and overfitting to checklist wording."},"retrievedAt":"2026-09-13T03:01:05Z","enrichedAt":"2026-09-13T03:03:17Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: provide readable trace summaries and train domain reviewers to assess execution evidence.","procurementImplications":"Interpretation: require exportable action logs and price human adjudication alongside model and compute costs.","operatingModelImplications":"Interpretation: research software engineers own replay infrastructure; domain scientists own result acceptance.","updateExplanation":"URL and AgentActionBench title absent from the 247-resource full archive, paginated at offsets 0, 100 and 200. Newly covered September preprint; not claimed to have appeared since the last completed run.","sourceVerification":{"openedUrl":"https://arxiv.org/html/2609.11117v1","referenceExcerpt":"execution and result verification emerging as the primary bottlenecks","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}