Education · Issue 04 ·
Research
Three newly archived sources cover planned U.S. university research capacity, a peer-reviewed clinical-data agent evaluation and an exploratory scientific-discovery benchmark. Benefits, failure modes and evaluator limitations remain distinct. September 8 and September 4 studies and an undated system page are explicitly dated; no claim of same-day deployment or autonomous discovery. Role guidance emphasizes local validation and review capacity. Independent production measurements and research-administration outcomes remain gaps.
- Evidence records
- 3
- Cross-source patterns
- 0
- Evidence classes
- 2 academic research1 standards or public-body guidance
- Outcomes
- 1 mixed1 emerging1 cautionary
- Source freshness
- 2 new this fortnight1 undated
- Research completed
- 2026-09-10
Choose a role to see its takeaway beside every record in the ledger.
Synthesis · Lighthouse Advisory interpretation
Patterns across the evidence
The evidence in this edition did not support a cross-source pattern. Each record below stands on its own.
Full record · every source keeps its link and limitations
Evidence ledger
Clinical-data agent study separates plausible plans from correct execution
Plans and code quality diverged; successful execution did not establish analytical validity.
Why it matters, evidence and limitations
- Why it matters
- Relevant to university research analysis assistants, with specialty-specific validation required.
- Evidence and measured results
- Sonnet 4.6 was tested in three modes and three prompt conditions, repeated three times: 27 runs against a public ophthalmology dataset and reference R analysis. Eight of 17 narrative summaries were fully satisfactory; two had clinically meaningful errors. Main-text tables document code and reporting failures.
- Limitations and uncertainty
- Single agent and precleaned dataset; possible text-level contamination; reference analysis itself had diagnostic shortcomings. Main text and tables inspected; supplement not independently inspected and code not rerun. No measured net labor saving.
Horizon installation page describes planned capacity, not achieved research outcomes
TACC describes installation and anticipates Phase 1 production in Fall 2026; operational scientific benefits remain prospective.
Why it matters, evidence and limitations
- Why it matters
- Direct U.S. public-university research-computing capacity planning relevance.
- Evidence and measured results
- The specifications describe Grace Blackwell and Vera platforms, InfiniBand, solid-state storage and liquid cooling. Performance and efficiency figures are operator claims without a workload-level evaluation method, sample or complete baseline on this page.
- Limitations and uncertainty
- Undated first-party page, not independent evaluation. No production acceptance result or measured discovery improvement. Classified as standards-guidance for its system-documentation function; promotional specifications remain operator claims.
Discovery benchmark exposes scrutiny gaps while leaving its own validation incomplete
Reported execution strengths exceeded control and robustness performance under a bounded automated rubric.
Why it matters, evidence and limitations
- Why it matters
- Useful for designing university scientific-agent evaluations; not proof of real-world discovery readiness.
- Evidence and measured results
- Forty tasks across ten domains tested four scaffolds once each using the same DeepSeek-V4-Flash base model with thinking disabled. Mean scores were 58.4–60.3/100, without reliable pairwise separation. A fixed quantized GLM-5.1 judge scored artifacts; aggregation alone was deterministic.
- Limitations and uncertainty
- One model and run per task; uncertain contamination; human calibration deferred. Prompts, limits and evaluated run artifacts are not public. Single-phase data constrain generalization scoring. No independent reproduction or repository execution performed.
How to read this edition
Source findings, measured results and limitations come from the cited publications. Patterns, operating questions, role takeaways and implementation considerations are Lighthouse Advisory interpretation, stated as questions to validate locally rather than guaranteed outcomes. Vendor and operator claims are labeled as claims. Full research method.
- Academic research
- Research produced through an academic institution or peer-reviewed venue.
- Standards or public-body guidance
- Normative or advisory guidance from a standards body or public institution.