Lighthouse AdvisorySLED AI Adoption Intelligence
← Back to results

From the K–12 edition of September 6, 2026

Academic researchMixedNewly relevant · Dec 2025

Supervised LearnLM mathematics trial offers short-term promise with uncertain transfer advantage

LearnLM Team, Google and Eedi · K–12 primary and secondary education · United Kingdom; five secondary schools

Publisher
AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms
Original publication
December 29, 2025 (arXiv v1); trial May–June 2025
Source retrieved
2026-09-07
Read original source

What happened

A supplier-authored trial supports a bounded human-supervised tutoring workflow; it does not establish autonomous tutoring effectiveness or durable learning.

Why it matters

Older research newly relevant through August tutoring scrutiny; no archive match found. Transfer requires local curriculum, staffing and safeguarding validation.

Evidence and measured results

165 students aged 13–15, seven weeks: student assignment to hints or tutoring, then session-level assignment to human or supervised LearnLM. Baseline-adjusted Bayesian next-topic success was 66.2% versus 60.7%; the 5.5-percentage-point difference had a 95% credible interval of -1.4 to 12.4. Transfer was same-day.

Limitations and uncertainty

Preprint, provider involvement, one platform and subject; every draft was reviewed. Session crossover prevents estimating cumulative effects; throughput was not rigorously measured. Selected transfers exclude students not continuing that day. Long-term retention and US applicability remain unproven.

Put this evidence to work

Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-07; this does not change the original publication date. Labels below come from the analysis itself.

Sales

Role takeaway

Tutoring directors, mathematics leads and procurement staff may want to expand support while retaining instructional quality. Ask whether trained reviewers are available, what curriculum the platform covers, and how independent learning will be measured. Offer a small supervised mathematics pilot with a human-tutoring comparison. The value hypothesis is maintaining useful support under measured staffing constraints. Do not sell the point estimate as proven superiority, extend it to long-term retention, or promise labor savings. This supplier-authored exploratory evidence supports testing a workflow, not replacing tutors or assuming another model, subject or student population will yield the same result.

Pre-sales engineering

Role takeaway

Build a draft-review-release queue that fails closed when a reviewer is unavailable and falls back to human support. Prerequisites are validated curriculum content, reliable learner identity and supplier-approved evaluation access. Test malformed context, answer leakage, cross-student data exposure, accessible rendering and audit completeness with synthetic cases before use. For proof of value, compare same-day performance and a separately designed delayed assessment while recording edits and review time. Keep model and prompt versions fixed during comparison. Do not copy the paper's persona instructions into production; clearly communicating AI assistance should be part of the local design review.

Delivery

Role takeaway

The tutoring service owner should recruit qualified reviewers, train them in pacing and escalation, and coordinate with curriculum, safeguarding and evaluation leads. Dependencies include consent procedures, approved records handling and scheduled human coverage. Before launch, approve the evaluation design and stop criteria.

Proposed acceptance
every released message has a reviewer record, fallback works in an outage exercise, and both delayed learning and staff effort are reported against the chosen comparison. Include students needing accommodations in usability review. Risks include review becoming a rubber stamp, frustration that is invisible in correctness scores, and expanding before results or staffing needs are understood.

Implementation considerations

Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.

Architecture and integration

What must connect, and where does the AI sit in the workflow?

Observed: a custom API connected a pedagogically tuned Gemini 2.0 Flash derivative to Eedi, with a tutor approval gate. Interpretation: preserve that gate; do not assume current models reproduce the tested system.

Governance

Who approves, reviews and stays accountable for outcomes?

Define independent follow-up assessment and explicitly disclose AI assistance. The experimental interface did not distinguish session type.

Security and privacy

What data, permissions and controls need testing?

Restrict prompt context, protect student transcripts, require retention and secondary-use terms, and audit access. The paper's safety reporting is not a privacy certification.

Accessibility and workforce

Who is affected, and what skills or accommodations follow?

Test language, disability access and frustration escalation locally; reserve qualified tutor time instead of presuming staffing reductions.

Procurement

What should contracts, pricing and exit terms secure?

Buy a validated instructional workflow with human coverage, not an unsupported learning or throughput guarantee.

Operating model

Which teams own the service once it runs?

Human tutors remain accountable for messages and escalation; measure their review burden and fallback capacity.

What changed

Not an archive repeat. Older supplier-authored trial was retrieved to inspect the underlying evidence surfaced by the August 2026 Stanford synthesis; retains its December 2025 publication and May–June 2025 trial dates.

Publication history

  1. 2026-09-06K–12 · Issue 014 resources
Read preserved resource versions (JSON)

Stable resource ID: learnlm-eedi-supervised-rct-2025