Lighthouse AdvisorySLED AI Adoption Intelligence
← Back to results

From the K–12 edition of September 10, 2026

Vendor claimMixedRecent

Khanmigo experiments show why faster tutoring needs multidimensional evaluation

Khan Academy · K–12 education · Khan Academy platform; U.S.-based operator, experiment population not fully specified

Publisher
Methodologies for Improving the Quality of AI Tutoring in K-12 Education
Original publication
arXiv version submitted August 7, 2026; metadata reports conference version first online June 25, 2026
Source retrieved
2026-09-11
Read original source

What happened

Operator experiments reveal trade-offs among response speed, answer disclosure and engagement; these are not independent evidence of retained learning.

Why it matters

Newly added to this archive for fall 2026 district decisions. Supports technical scrutiny of tutoring components and supplier change management.

Evidence and measured results

Limiting Math Agent guidance reportedly reduced answer disclosure 85.5% ±5.56% while cognitive engagement fell 18.09% ±7.45%. Reported margins correspond to 95% intervals. Most experiments assign conversation threads; the comparator is the prior workflow. Per-experiment sample counts and absolute baselines are not provided in the inspected results.

Limitations and uncertainty

Same users can encounter multiple conditions. LLM judges are imperfect; offline tests are single-turn. Moderation and bias evaluation are outside scope. The paper does not establish delayed learning or district-wide savings.

Put this evidence to work

Lighthouse Advisory interpretation, grounded in this source. Enriched 2026-09-11; this does not change the original publication date. Labels below come from the analysis itself.

Sales

Role takeaway

District curriculum, IT and procurement teams need to understand whether a tutoring release improves the learning workflow they are buying. Ask what independent work students must complete, which metrics matter at renewal and who approves supplier experiments. Offer a bounded version-review and evaluation design engagement using the actual licensed feature. The value hypothesis is a more defensible deployment decision. Do not present operator percentages as guaranteed learning gains or staffing savings; results depend on the product, workflow and measurement method.

Pre-sales engineering

Role takeaway

Require an observable pipeline with explicit model, moderation, context and tool boundaries. Validate a proposed change against the incumbent using a frozen test set, then an authorized supervised pilot with independent student assessment. Test code-execution isolation, timeout handling, privacy controls and rollback. Track latency together with mathematical correctness, assistance level and independently completed work. A faster component is insufficient if downstream behavior worsens. Cloud, local and hybrid options require separate infrastructure and support estimates; the study is not a district deployment specification.

Delivery

Role takeaway

Implementation includes supplier change review, evaluation instrumentation, educator training and a student-support fallback. The curriculum lead owns instructional acceptance; IT operates the service; privacy and safeguarding staff approve collection and escalation. Proposed acceptance is a documented comparison against the prior workflow, with locally chosen learning and safety thresholds met and a rehearsed rollback. Review results after initial adoption and subsequent releases. Risks include unreliable automatic judges, interaction between experiments, hidden support labor and learners adapting their behavior after launch.

Implementation considerations

Lighthouse Advisory interpretation across the operating dimensions a public-sector buyer must settle before this evidence becomes a design. Each note answers the question under its heading for this specific source.

Architecture and integration

What must connect, and where does the AI sit in the workflow?

Test each agent step and downstream conversation as one service, with traceable configuration and rollback. Do not add agent calls where a deterministic rule meets the requirement.

Governance

Who approves, reviews and stays accountable for outcomes?

District approval should cover experimentation and version changes, including who can authorize exposure to changed behavior.

Security and privacy

What data, permissions and controls need testing?

Review tool sandbox isolation, student-context minimization and logging access; traceability must not become indefinite storage of pupil conversations.

Accessibility and workforce

Who is affected, and what skills or accommodations follow?

Check reading demands and response waits with learners who need accommodations; engagement counts do not establish accessible learning.

Procurement

What should contracts, pricing and exit terms secure?

Ask suppliers for denominators, absolute rates, judge calibration and change notices before accepting percentage improvements.

Operating model

Which teams own the service once it runs?

Assign a curriculum owner to weigh conflicting metrics and an IT owner to operate rollback and incident response.

Publication history

  1. 2026-09-10K–12 · Issue 054 resources
Read preserved resource versions (JSON)

Stable resource ID: khanmigo-quality-methods-2026