{"resourceId":"berkeley-tutor-feedback-uptake-2026","versions":[{"version":"external-fdbf3eb2e7843c09d9919ef450db4d58a27be4310a16aedc20b79e69b011eda7","resource":{"id":"berkeley-tutor-feedback-uptake-2026","title":"Programming tutor comparison finds stronger feedback uptake but mixed correct application","organization":"University of California, Berkeley and research collaborators","sector":"Higher education teaching and learning","geography":"California, United States; one undergraduate programming course","publishedAt":"May 7, 2026","publicationDate":"2026-05-07","eventDate":null,"sourceName":"The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness","sourceLabel":"Original academic preprint; methods, tables and limitations inspected","sourceUrl":"https://arxiv.org/html/2605.05648v1","evidenceClass":"academic-research","outcomeClass":"mixed","topics":["knowledge-work","developers-agents","data-security","governance-procurement","accessibility-workforce","operating-model"],"finding":"A misconception-focused tutor shows more feedback uptake, but correct application varies by assignment. This is not a learning-gain estimate.","sledRelevance":"Interpretation: Direct public-university evidence for evaluating programming feedback. Other disciplines need their own observable success criteria.","evidence":"10,235 submissions: 681 students in Fall 2024 and 958 in Fall 2025. Both tutors use GPT-4 with different prompts. Table 3 favors the newer tutor for uptake throughout; correct application improves on assignments 1–2, reverses on 4, and has nonsignificant differences on 3 and 5.","architectureImplications":"Interpretation: Connect code versions, feedback and test results using pseudonymous identifiers. A tutor copilot should explain suggested changes without production access. Hosting alternatives are unevaluated; compare cloud egress, hybrid storage and local inference requirements before selecting infrastructure.","governanceImplications":"Interpretation: Register pedagogical quality, feedback uptake and independent transfer as distinct endpoints. Validate automated scoring with educators before using it for student decisions.","securityPrivacyImplications":"Interpretation: Remove credentials and identifying comments from code sent to models. Isolate code execution, restrict evaluator access and set deletion rules for revision histories.","caveats":"Nonrandom cross-semester comparison; cohort confounding, indirect attribution of edits and no delayed learning test. LLM judging has limited human validation. Optional helpfulness ratings cover about 38% of sampled submissions.","streamIds":["student-success"],"roles":{"sales":"Interpretation — Computing departments may know that students submit passing code without knowing whether automated feedback helped them reason. Include course faculty, learning technology staff, institutional research and students. Ask which concepts remain difficult after a correction, how much feedback review costs and what would justify changing the tutor. A bounded engagement could evaluate one assignment sequence using a current approved tool. The value hypothesis is better diagnosis of unusable or misleading feedback, subject to local testing. Do not convert submission volume into student reach, feedback uptake into mastery, or this comparison into promised retention gains. The source supports evaluation design, not a procurement winner or staffing reduction.","engineering":"Interpretation — Fit the evaluation to a sandboxed programming environment with stable assignment identifiers and versioned feedback. Prerequisites include an educator-authored rubric, lawful access to revision data and an independent test task. Keep execution separate from model inference and deny access to production secrets or student-record writes. Sample judge disagreements for human adjudication, including large rewrites and partially adopted advice. Proposed proof of value: compare tutor configurations within the same cohort, report correctness and answer leakage separately, and assess an unaided transfer task. Validate latency and recovery under course deadlines. The source's observational comparison cannot establish that changing prompts alone will improve the local course.","delivery":"Interpretation — The course lead should own instructional decisions, with an evaluator owning the analysis and IT owning service reliability. Prepare a data map, review code-handling permissions and train teaching assistants to identify harmful feedback. Dependencies include accessible coding tools, marking capacity and a manual support route. Proposed acceptance criteria: every reported endpoint has a reconciled denominator; sampled automated scores receive human review; critical feedback errors have an owner; and independent transfer results are reviewed before expansion. These are proposed gates. Track support effort and voluntary participation as well as usage. Risks include treating repeated submissions as independent learners, review fatigue and optimizing for passing tests instead of understanding."},"retrievedAt":"2026-09-11T03:01:11Z","enrichedAt":"2026-09-11T03:03:23Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Test screen-reader navigation through code diffs and feedback; budget teaching-assistant adjudication. No accessibility benefit is demonstrated.","procurementImplications":"Interpretation: Require exportable revision logs, version notices, test access and bounded inference charges; include educator review in cost estimates.","operatingModelImplications":"Interpretation: Faculty own feedback standards; learning technology owns integration; institutional research independently evaluates outcomes.","updateExplanation":"New in the full 182-record archive and identifier-specific search. Explicit May evidence backfill adds behavioral evaluation of a deployed public-university tutor, not a September event.","sourceVerification":{"openedUrl":"https://arxiv.org/html/2605.05648v1","referenceExcerpt":"Our engagement-based metrics measure immediate feedback uptake but do not assess longer-term learning outcomes","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}