Education · Issue 05 ·
K–12
Four newly archived sources examine assessment reliability, tutoring workflow trade-offs, LAUSD student-access restrictions and school staff data handling. Original research distinguishes scoring accuracy from agreement; an operator paper reports mixed tutoring metrics, not retained learning. Reporting and UK guidance support operational questions without proving policy benefits. Two cross-source patterns connect measurement validity and separate user/data decisions. These are older sources newly relevant to fall decisions, not September 10 news. Long-term learning, subgroup/accessibility outcomes, independent costs and security assurance remain gaps.
- Evidence records
- 4
- Cross-source patterns
- 2
- Evidence classes
- 1 academic research1 vendor claim1 independent reporting1 standards or public-body guidance
- Outcomes
- 2 mixed2 emerging
- Source freshness
- 2 undated1 recent1 new this fortnight
- Research completed
- 2026-09-11
Choose a role to see its takeaway beside every record in the ledger.
Synthesis · Lighthouse Advisory interpretation
Patterns across the evidence
Validate the AI measure before using it to judge the intervention
The assessment study’s uneven agreement and Khan Academy’s reliance on automatic judges make measurement quality an implementation dependency. Include independent human checks and distinct learning outcomes before interpreting a dashboard improvement as educational progress.
Operating questionHow do we know the scoring or engagement measure is reliable for this task and learner group?
Supporting evidencePeczuh, Kumar and colleagues; university, ETS and AERDF collaboratorsKhan Academy
Define student access and staff data handling as separate operating decisions
LAUSD’s reported student restrictions and DfE’s staff-workflow guidance concern different users and data paths. A district needs explicit decisions for both; a student-access rule alone does not settle how employees handle pupil records.
Operating questionWhich user, feature and data path is authorized, and who owns the unresolved cases?
Supporting evidenceLos Angeles Unified School District; reported by K-12 DiveUK Department for Education
Full record · every source keeps its link and limitations
Evidence ledger
Essay-scoring study exposes large differences between critical-thinking subskills
Fine-tuned Llama performed best overall, but aggregate accuracy concealed weak agreement on some subskills.
Why it matters, evidence and limitations
- Why it matters
- Newly added to this archive for fall 2026 district decisions. Supports validation of AI-assisted assessment before consequential use.
- Evidence and measured results
- Exploratory comparison of prompting and fine-tuning against human labels on 500 U.S. grade 6–12 essays, averaged over five random essay splits. Table 6 reports evidence-strength accuracy .768 but Krippendorff’s alpha .234; source-synthesis accuracy .780 and alpha .740. These are scoring metrics, not learning gains.
- Limitations and uncertainty
- Prompts and source materials were unavailable; labels were imbalanced. The exploratory human reliability threshold was .6, with an exception for facts/opinions. No classroom intervention or durable-learning effect was tested.
Khanmigo experiments show why faster tutoring needs multidimensional evaluation
Operator experiments reveal trade-offs among response speed, answer disclosure and engagement; these are not independent evidence of retained learning.
Why it matters, evidence and limitations
- Why it matters
- Newly added to this archive for fall 2026 district decisions. Supports technical scrutiny of tutoring components and supplier change management.
- Evidence and measured results
- Limiting Math Agent guidance reportedly reduced answer disclosure 85.5% ±5.56% while cognitive engagement fell 18.09% ±7.45%. Reported margins correspond to 95% intervals. Most experiments assign conversation threads; the comparator is the prior workflow. Per-experiment sample counts and absolute baselines are not provided in the inspected results.
- Limitations and uncertainty
- Same users can encounter multiple conditions. LLM judges are imperfect; offline tests are single-turn. Moderation and bias evaluation are outside scope. The paper does not establish delayed learning or district-wide savings.
LAUSD access restrictions put interim controls ahead of longer-term policy decisions
K-12 Dive reports LAUSD confirmed restrictions on student generative-AI access for 2026–27 while reviewing instructional uses and safeguards.
Why it matters, evidence and limitations
- Why it matters
- Newly added to this archive for fall 2026 district decisions. Adds a distinct district implementation case to existing NYC coverage.
- Evidence and measured results
- The reporter attributes confirmation to a district spokesperson’s September 2 email. Previous access depended on age and required preparation. A committee is expected to bring recommendations to the board by school-year end.
- Limitations and uncertainty
- This is attributed reporting, not a directly inspected policy instrument or technical audit. The district meeting transcript could not be opened. Exceptions, enforcement completeness and educational effects remain unverified.
School data guidance extends review beyond text entered into AI tools
Guidance addresses approved tools, data-officer review, training-data use, age restrictions and transparency about metadata as well as prompts.
Why it matters, evidence and limitations
- Why it matters
- Newly added to this archive for fall 2026 district decisions. Adds staff-workflow and metadata handling detail to prior product-safety coverage.
- Evidence and measured results
- Its administrative example drafts a parent message without pupil identifiers, adding them afterward. It also identifies location, IP, system and browser information as potential collected data. This is guidance, not an evaluated workload intervention.
- Limitations and uncertainty
- The visible dates apply to the manual and do not establish when this section changed. UK obligations must not be restated as U.S. law. No effect size, baseline or evaluation sample is supplied.
How to read this edition
Source findings, measured results and limitations come from the cited publications. Patterns, operating questions, role takeaways and implementation considerations are Lighthouse Advisory interpretation, stated as questions to validate locally rather than guaranteed outcomes. Vendor and operator claims are labeled as claims. Full research method.
- Academic research
- Research produced through an academic institution or peer-reviewed venue.
- Vendor claim
- A supplier-provided assertion that has not been upgraded to independent evidence.
- Independent reporting
- Independent reporting with attributable sources but without a formal evaluation design.
- Standards or public-body guidance
- Normative or advisory guidance from a standards body or public institution.