{"resourceId":"khanmigo-quality-methods-2026","versions":[{"version":"external-5720f70db75cc4bd8af9b029f1d7158cb5855b9a849a6fc3a81b31eb867430c4","resource":{"id":"khanmigo-quality-methods-2026","title":"Khanmigo experiments show why faster tutoring needs multidimensional evaluation","organization":"Khan Academy","sector":"K–12 education","geography":"Khan Academy platform; U.S.-based operator, experiment population not fully specified","publishedAt":"arXiv version submitted August 7, 2026; metadata reports conference version first online June 25, 2026","publicationDate":"2026-08-07","eventDate":null,"sourceName":"Methodologies for Improving the Quality of AI Tutoring in K-12 Education","sourceLabel":"Operator-authored AIED paper, arXiv full text","sourceUrl":"https://arxiv.org/pdf/2608.11259","evidenceClass":"vendor-claim","outcomeClass":"mixed","topics":["knowledge-work","developers-agents","infrastructure","data-security","operating-model"],"finding":"Operator experiments reveal trade-offs among response speed, answer disclosure and engagement; these are not independent evidence of retained learning.","sledRelevance":"Interpretation: Newly added to this archive for fall 2026 district decisions. Supports technical scrutiny of tutoring components and supplier change management.","evidence":"Limiting Math Agent guidance reportedly reduced answer disclosure 85.5% ±5.56% while cognitive engagement fell 18.09% ±7.45%. Reported margins correspond to 95% intervals. Most experiments assign conversation threads; the comparator is the prior workflow. Per-experiment sample counts and absolute baselines are not provided in the inspected results.","architectureImplications":"Interpretation: Test each agent step and downstream conversation as one service, with traceable configuration and rollback. Do not add agent calls where a deterministic rule meets the requirement.","governanceImplications":"Interpretation: District approval should cover experimentation and version changes, including who can authorize exposure to changed behavior.","securityPrivacyImplications":"Interpretation: Review tool sandbox isolation, student-context minimization and logging access; traceability must not become indefinite storage of pupil conversations.","caveats":"Same users can encounter multiple conditions. LLM judges are imperfect; offline tests are single-turn. Moderation and bias evaluation are outside scope. The paper does not establish delayed learning or district-wide savings.","streamIds":["k12"],"roles":{"sales":"Interpretation: District curriculum, IT and procurement teams need to understand whether a tutoring release improves the learning workflow they are buying. Ask what independent work students must complete, which metrics matter at renewal and who approves supplier experiments. Offer a bounded version-review and evaluation design engagement using the actual licensed feature. The value hypothesis is a more defensible deployment decision. Do not present operator percentages as guaranteed learning gains or staffing savings; results depend on the product, workflow and measurement method.","engineering":"Interpretation: Require an observable pipeline with explicit model, moderation, context and tool boundaries. Validate a proposed change against the incumbent using a frozen test set, then an authorized supervised pilot with independent student assessment. Test code-execution isolation, timeout handling, privacy controls and rollback. Track latency together with mathematical correctness, assistance level and independently completed work. A faster component is insufficient if downstream behavior worsens. Cloud, local and hybrid options require separate infrastructure and support estimates; the study is not a district deployment specification.","delivery":"Interpretation: Implementation includes supplier change review, evaluation instrumentation, educator training and a student-support fallback. The curriculum lead owns instructional acceptance; IT operates the service; privacy and safeguarding staff approve collection and escalation. Proposed acceptance is a documented comparison against the prior workflow, with locally chosen learning and safety thresholds met and a rehearsed rollback. Review results after initial adoption and subsequent releases. Risks include unreliable automatic judges, interaction between experiments, hidden support labor and learners adapting their behavior after launch."},"retrievedAt":"2026-09-11T03:01:41Z","enrichedAt":"2026-09-11T03:04:43Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Check reading demands and response waits with learners who need accommodations; engagement counts do not establish accessible learning.","procurementImplications":"Interpretation: Ask suppliers for denominators, absolute rates, judge calibration and change notices before accepting percentage improvements.","operatingModelImplications":"Interpretation: Assign a curriculum owner to weigh conflicting metrics and an IT owner to operate rollback and incident response.","sourceVerification":{"openedUrl":"https://arxiv.org/pdf/2608.11259","referenceExcerpt":"Currently our offline evaluations only evaluate a single turn of a model.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}