{"resourceId":"learnlm-eedi-supervised-rct-2025","versions":[{"version":"external-6a699404547feef3c4d29c3c0b5064bf83a07b824e424933ab1c1b7e3bf9557d","resource":{"id":"learnlm-eedi-supervised-rct-2025","title":"Supervised LearnLM mathematics trial offers short-term promise with uncertain transfer advantage","organization":"LearnLM Team, Google and Eedi","sector":"K–12 primary and secondary education","geography":"United Kingdom; five secondary schools","publishedAt":"December 29, 2025 (arXiv v1); trial May–June 2025","publicationDate":"2025-12-29","eventDate":null,"sourceName":"AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms","sourceLabel":"Supplier-authored exploratory randomized trial preprint","sourceUrl":"https://arxiv.org/html/2512.23633v1","evidenceClass":"academic-research","outcomeClass":"mixed","topics":["knowledge-work","developers-agents","infrastructure","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"A supplier-authored trial supports a bounded human-supervised tutoring workflow; it does not establish autonomous tutoring effectiveness or durable learning.","sledRelevance":"Interpretation: older research newly relevant through August tutoring scrutiny; no archive match found. Transfer requires local curriculum, staffing and safeguarding validation.","evidence":"165 students aged 13–15, seven weeks: student assignment to hints or tutoring, then session-level assignment to human or supervised LearnLM. Baseline-adjusted Bayesian next-topic success was 66.2% versus 60.7%; the 5.5-percentage-point difference had a 95% credible interval of -1.4 to 12.4. Transfer was same-day.","architectureImplications":"Observed: a custom API connected a pedagogically tuned Gemini 2.0 Flash derivative to Eedi, with a tutor approval gate. Interpretation: preserve that gate; do not assume current models reproduce the tested system.","governanceImplications":"Interpretation: define independent follow-up assessment and explicitly disclose AI assistance. The experimental interface did not distinguish session type.","securityPrivacyImplications":"Interpretation: restrict prompt context, protect student transcripts, require retention and secondary-use terms, and audit access. The paper's safety reporting is not a privacy certification.","caveats":"Preprint, provider involvement, one platform and subject; every draft was reviewed. Session crossover prevents estimating cumulative effects; throughput was not rigorously measured. Selected transfers exclude students not continuing that day. Long-term retention and US applicability remain unproven.","streamIds":["k12"],"roles":{"sales":"Interpretation: Tutoring directors, mathematics leads and procurement staff may want to expand support while retaining instructional quality. Ask whether trained reviewers are available, what curriculum the platform covers, and how independent learning will be measured. Offer a small supervised mathematics pilot with a human-tutoring comparison. The value hypothesis is maintaining useful support under measured staffing constraints. Do not sell the point estimate as proven superiority, extend it to long-term retention, or promise labor savings. This supplier-authored exploratory evidence supports testing a workflow, not replacing tutors or assuming another model, subject or student population will yield the same result.","engineering":"Interpretation: Build a draft-review-release queue that fails closed when a reviewer is unavailable and falls back to human support. Prerequisites are validated curriculum content, reliable learner identity and supplier-approved evaluation access. Test malformed context, answer leakage, cross-student data exposure, accessible rendering and audit completeness with synthetic cases before use. For proof of value, compare same-day performance and a separately designed delayed assessment while recording edits and review time. Keep model and prompt versions fixed during comparison. Do not copy the paper's persona instructions into production; clearly communicating AI assistance should be part of the local design review.","delivery":"Interpretation: The tutoring service owner should recruit qualified reviewers, train them in pacing and escalation, and coordinate with curriculum, safeguarding and evaluation leads. Dependencies include consent procedures, approved records handling and scheduled human coverage. Before launch, approve the evaluation design and stop criteria. Proposed acceptance: every released message has a reviewer record, fallback works in an outage exercise, and both delayed learning and staff effort are reported against the chosen comparison. Include students needing accommodations in usability review. Risks include review becoming a rubber stamp, frustration that is invisible in correctness scores, and expanding before results or staffing needs are understood."},"retrievedAt":"2026-09-07T02:48:23Z","enrichedAt":"2026-09-07T02:52:08Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: test language, disability access and frustration escalation locally; reserve qualified tutor time instead of presuming staffing reductions.","procurementImplications":"Interpretation: buy a validated instructional workflow with human coverage, not an unsupported learning or throughput guarantee.","operatingModelImplications":"Interpretation: human tutors remain accountable for messages and escalation; measure their review burden and fallback capacity.","updateExplanation":"Not an archive repeat. Older supplier-authored trial was retrieved to inspect the underlying evidence surfaced by the August 2026 Stanford synthesis; retains its December 2025 publication and May–June 2025 trial dates.","sourceVerification":{"openedUrl":"https://arxiv.org/html/2512.23633v1","referenceExcerpt":"Measuring substantive, longer-term effects on learning will require a different approach.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}