Education · Issue 03 ·
Research
Three newly archived sources cover a September 8 public-university research announcement, research-administration failure reporting and independent scientific-agent integrity tests. Planned benefits, operator fixes and benchmark findings remain distinct. The edition adds operational controls and scrutiny of the evaluator itself, with no institution-wide productivity claim. U.S. university coverage is paired with international research; measured administrative benefits and independent deployment validation remain gaps.
- Evidence records
- 3
- Cross-source patterns
- 1
- Evidence classes
- 2 vendor claim1 academic research
- Outcomes
- 2 emerging1 cautionary
- Source freshness
- 2 new this fortnight1 older, newly relevant
- Research completed
- 2026-09-09
Choose a role to see its takeaway beside every record in the ledger.
Synthesis · Lighthouse Advisory interpretation
Patterns across the evidence
Preserve failure status through the evidence chain
Vandalizer's reported error-state changes and SciIntegrity-Bench's integrity tests address different places where unsupported completion can arise. A reviewable workflow must preserve input, processing and final-report status together. Neither source validates the other system or proves that visible warnings prevent misconduct.
Operating questionCan a reviewer trace every completed claim through successful input processing and observed execution, and distinguish failed processing from absent evidence?
Supporting evidenceUniversity of Idaho AI4RAZonglin Yang, Xingtong Liu and Xinyan Xu; Readraft Lab and university affiliations
Full record · every source keeps its link and limitations
Evidence ledger
UW Genesis projects outline research integration work, with outcomes still prospective
UW announces participation in four Genesis projects. The described scientific and infrastructure benefits remain goals.
Why it matters, evidence and limitations
- Why it matters
- Direct public-university research relevance; broader institutional transfer requires workload and collaboration assessment.
- Evidence and measured results
- Planned work spans sensing, protein design and astronomy. Astronomy infrastructure would support multiple data types across cloud and HPC; another project would connect AI-guided design with fabrication and measurement. No outcome sample, baseline or measured gain is reported.
- Limitations and uncertainty
- Announcement, not a completed deployment evaluation. Exact award dates are unspecified. The schema lacks an institutional-announcement class; vendor-claim denotes interested-party attribution, not that UW is a vendor.
Vandalizer release targets silent truncation and misleading extraction status
The operator reports changes that distinguish failed processing from absent evidence and incomplete reports from completed work.
Why it matters, evidence and limitations
- Why it matters
- Research administrators reviewing proposals need to recognize when document processing failed before relying on compliance or extraction results.
- Evidence and measured results
- Release notes describe larger-model routing for oversized documents, estimated citation-page labels, unreadable-file warnings and explicit extraction failures. No error-rate sample, controlled baseline, performance measurement or independent test is supplied.
- Limitations and uncertainty
- Operator claims were not tested in the application. Routing destinations, deployment configuration and cost effects are unspecified. Publication date is known; a separate release-event date is not established from the inspected notes.
Scientific-agent benchmark exposes integrity failures, with reporting inconsistencies limiting inference
Controlled integrity dilemmas reveal failures that ordinary task-completion scoring can miss; the paper itself requires cautious reading.
Why it matters, evidence and limitations
- Why it matters
- A test-design reference for university research software and integrity teams, not a measured U.S. deployment outcome.
- Evidence and measured results
- Table 2 reports 36 Fail labels in 231 evaluations: seven models, 33 synthetic scenarios, minimal ReAct scaffold and common prompt. Authors manually assessed reports and execution traces against predefined checklists. There is no human-workflow baseline.
- Limitations and uncertainty
- Preprint; three scenarios per category and no measured annotator agreement. Universal fabrication wording conflicts with Appendix H; its disclosure totals also have ambiguous row coding. Omit those claims and ablation effect sizes. Table 2 counts are author-reported, not independently rerun.
How to read this edition
Source findings, measured results and limitations come from the cited publications. Patterns, operating questions, role takeaways and implementation considerations are Lighthouse Advisory interpretation, stated as questions to validate locally rather than guaranteed outcomes. Vendor and operator claims are labeled as claims. Full research method.
- Vendor claim
- A supplier-provided assertion that has not been upgraded to independent evidence.
- Academic research
- Research produced through an academic institution or peer-reviewed venue.