Education · Issue 05 ·
Student Success
Two newly archived sources examine programming-tutor feedback and learner judgment. A Berkeley deployment comparison finds mixed correct application despite stronger feedback uptake; a recent critical review provides unvalidated diagnostic questions, not causal evidence of harm. Both are explicit evidence backfill, not September 10 events. No cross-source patterns are asserted. Access failures limited trial coverage; current advising outcomes, durable learning, disability-specific effects and total cost remain gaps.
- Evidence records
- 2
- Cross-source patterns
- 0
- Evidence classes
- 2 academic research
- Outcomes
- 1 mixed1 cautionary
- Source freshness
- 1 older, newly relevant1 new this fortnight
- Research completed
- 2026-09-11
Choose a role to see its takeaway beside every record in the ledger.
Synthesis · Lighthouse Advisory interpretation
Patterns across the evidence
The evidence in this edition did not support a cross-source pattern. Each record below stands on its own.
Full record · every source keeps its link and limitations
Evidence ledger
Programming tutor comparison finds stronger feedback uptake but mixed correct application
A misconception-focused tutor shows more feedback uptake, but correct application varies by assignment. This is not a learning-gain estimate.
Why it matters, evidence and limitations
- Why it matters
- Direct public-university evidence for evaluating programming feedback. Other disciplines need their own observable success criteria.
- Evidence and measured results
- 10,235 submissions: 681 students in Fall 2024 and 958 in Fall 2025. Both tutors use GPT-4 with different prompts. Table 3 favors the newer tutor for uptake throughout; correct application improves on assignments 1–2, reverses on 4, and has nonsignificant differences on 3 and 5.
- Limitations and uncertainty
- Nonrandom cross-semester comparison; cohort confounding, indirect attribution of edits and no delayed learning test. LLM judging has limited human validation. Optional helpfulness ratings cover about 38% of sampled submissions.
Critical review offers questions for preserving learner judgment, not a validated dependency scale
The review distinguishes useful assistance from delegation that displaces learner judgment; it does not establish that frequent AI use causes harm.
Why it matters, evidence and limitations
- Why it matters
- Useful for U.S. college assessment design discussions, with no local prevalence or causal effect to transfer.
- Evidence and measured results
- Purposive synthesis of 60 works, searches updated July 2026. No new dataset, experimental baseline or effect estimate. Proposed diagnostic criteria require validation.
- Limitations and uncertainty
- Nonexhaustive conceptual review; hypotheses from adjacent domains are not demonstrated educational effects. Detailed table endpoint failed, so no table-only claims are used; main-text methods and limitations were accessible.
How to read this edition
Source findings, measured results and limitations come from the cited publications. Patterns, operating questions, role takeaways and implementation considerations are Lighthouse Advisory interpretation, stated as questions to validate locally rather than guaranteed outcomes. Vendor and operator claims are labeled as claims. Full research method.
- Academic research
- Research produced through an academic institution or peer-reviewed venue.