{"resourceId":"ai-scientist-workshop-selection-limits-2026","versions":[{"version":"external-c9aa39e1d34b8205ab2c7838900cfa56ae70bddc1e22076e88b47cbd35c22a77","resource":{"id":"ai-scientist-workshop-selection-limits-2026","title":"AI Scientist workshop result shows bounded automation with human selection","organization":"Sakana AI, University of Oxford, University of British Columbia and collaborating researchers","sector":"AI-assisted computational science","geography":"Japan, United Kingdom and Canada; machine-learning workshop setting","publishedAt":"March 25, 2026","publicationDate":"2026-03-25","eventDate":null,"sourceName":"Nature","sourceLabel":"Peer-reviewed developer-authored study; commercial interests disclosed","sourceUrl":"https://www.nature.com/articles/s41586-026-10265-5","evidenceClass":"academic-research","outcomeClass":"mixed","topics":["knowledge-work","developers-agents","infrastructure","governance-procurement","operating-model"],"finding":"A selected generated manuscript cleared a workshop review threshold, with important limits on autonomy and generality.","sledRelevance":"Relevant to university computational research pilots; international ML results do not establish benefits in physical laboratories or research administration.","evidence":"Three manually selected manuscripts underwent blinded workshop review. One met the acceptance bar; all were withdrawn under the protocol. Authors judged none suitable for the main conference. Human filtering considered topic fit, implementation and formatting. Methods describe template-based and template-free pipelines.","architectureImplications":"Interpretation: keep ideation, code execution, manuscript generation and acceptance as separately inspectable stages.","governanceImplications":"Interpretation: disclose human selection and generated content; obtain applicable review authorization before involving external reviewers.","securityPrivacyImplications":"Interpretation: approve data and literature endpoints before providing unpublished material to external models.","caveats":"Small selected sample; developer affiliations and commercial interests. Computational experiments only. No matched institutional productivity baseline or proof of reliable autonomous science. Exact workshop event date is not established here.","streamIds":["research"],"roles":{"sales":"Interpretation: Engage principal investigators, research development leaders and computing services where experimental iteration consumes scarce time. Ask which computational steps are repeatable, how candidate ideas are selected and what evidence an investigator needs before trusting a manuscript. Offer a bounded pilot around a known public dataset and an established scientific question. The value hypothesis is faster exploration with accountable selection, to be tested against the existing workflow. Avoid equating workshop review with transformative discovery or promising autonomous publication. Budget discarded candidates, domain review and correction effort; the selected sample cannot establish campus-wide savings.","engineering":"Interpretation: Implement separate modules for idea proposals, authorized data access, code execution and evidence-linked write-up. Use resource caps, reproducible containers and external experiment records. Treat generated citations and figures as testable claims. For a proof of value, compare a scientist-led baseline with the assisted workflow using the same data, compute budget and acceptance rubric; retain all candidates to expose selection effects. Verify numerical outputs independently and evaluate total reviewer time. External model services require approved handling of unpublished data, while local execution still requires sandboxing and dependency controls.","delivery":"Interpretation: Name the principal investigator as scientific owner and research computing as execution support. Agree on the question, data permissions, baseline and stopping rules before adoption. Train researchers to document both human selection and automated steps. Proposed acceptance criteria include a replayable experiment, verified citations, disclosure of every excluded candidate and independent sign-off on each retained conclusion. Separate manuscript quality from scientific contribution in the final review. Record compute costs and correction time. Risks include selective reporting, confident unsupported claims and staff mistaking polished outputs for validated results; the observed workshop result does not remove these responsibilities."},"retrievedAt":"2026-09-13T03:01:05Z","enrichedAt":"2026-09-13T03:03:17Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: retain scientific mentoring and accessible output review; automation changes review tasks rather than proving staff replacement.","procurementImplications":"Interpretation: assess full candidate-generation and selection costs, not only the successful manuscript.","operatingModelImplications":"Interpretation: a named investigator remains accountable for methods, claims and release decisions.","updateExplanation":"Article URL and study absent from the full 247-resource archive. Explicitly older contextual evidence for interpreting the newly covered reproduction benchmark; no claim of a September publication or substantive source update.","sourceVerification":{"openedUrl":"https://www.nature.com/articles/s41586-026-10265-5","referenceExcerpt":"We manually filtered the most promising outputs at each stage","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}