{"resourceId":"akgun-toker-ai-three-week-retention-2025","versions":[{"version":"external-45807b2217d7a6000864e9fd06f2edd0633c36878617b3080a1f9c593c1d6a81","resource":{"id":"akgun-toker-ai-three-week-retention-2025","title":"Task-specific AI learning gains weaken at follow-up; reporting limits qualify the retention claim","organization":"Mahir Akgun and Sacip Toker; Penn State University and Atilim University","sector":"Higher education student learning and support","geography":"Single unnamed private university; study country not established in inspected text","publishedAt":"July 10, 2025","publicationDate":"2025-07-10","eventDate":null,"sourceName":"Short-Term Gains, Long-Term Gaps: The Impact of GenAI and Search Technologies on Retention","sourceLabel":"Original arXiv manuscript, listed as forthcoming AIED 2025 proceedings","sourceUrl":"https://arxiv.org/html/2507.07357v1","evidenceClass":"academic-research","outcomeClass":"mixed","topics":["knowledge-work","governance-procurement","accessibility-workforce","operating-model"],"finding":"ChatGPT improved immediate lower-order task assessment relative to control, but the study does not establish a general durable-learning advantage.","sledRelevance":"Interpretation: Useful backfill for college pilots needing delayed independent assessment. The institution is unnamed, so author affiliations cannot establish geographic transferability.","evidence":"Final n=123 from 152 volunteers; randomized four-tool comparison with unaided quizzes and follow-up three weeks after the final task. Tables 4–5 show Task 1 ChatGPT means 82.6 then 65.5, versus control 60.2 then 59.3; time-by-group p<.01. Higher-order immediate group differences were nonsignificant (p=.514).","architectureImplications":"Interpretation: Separate supported practice from restricted assessment in the learning platform. Version tool configurations and preserve assessable work. Infrastructure choice, model generation, agents and cloud versus on-premises performance are not evaluated.","governanceImplications":"Interpretation: Register endpoints, exclusions and analysis before a local trial. Do not turn nonsignificant differences into proof of equivalence or harm.","securityPrivacyImplications":"Interpretation: Keep student identifiers out of public prompts, minimize research logs and separate grading access from service administration. Privacy-themed tasks are not a product security assessment.","caveats":"Single site, post-assignment exclusions, fixed task order and restricted tools limit generalization. Cluster counts total 153 despite 152 volunteers; Task 2 prose conflicts with its table. Do not infer higher-order harm from nonsignificance. Model version and event dates are unspecified.","streamIds":["student-success"],"roles":{"sales":"Interpretation — Teaching leaders need to know whether improved performance persists beyond supported practice. Bring faculty, tutors, institutional research and student representatives into discovery. Ask what skill should endure, when it can be reassessed and how withdrawal will affect the estimate. Offer a bounded evaluation of an existing approved assistant with a delayed checkpoint. The value hypothesis is better evidence for a course decision, rather than a guaranteed positive result. Reporting inconsistencies make this a design prompt, not a sales benchmark. Do not promise lasting learning gains, assert that AI universally harms cognition or use the study to justify staffing reductions.","engineering":"Interpretation — Fit a local evaluation to a learning platform that can distinguish practice access from independent testing. Prerequisites include equivalent assessment forms, content approval, consent arrangements and a documented model configuration. Record availability and tool use without collecting unnecessary personal content. Test answer leakage, inaccessible questions and failure recovery before enrollment. Proposed proof of value: compare baseline-adjusted immediate and delayed independent performance, reporting missing data and uncertainty. Keep supported outputs separate from scores used to infer competence. A mixed-resource comparison may better match local study habits, but it would be a new design rather than a replication of this restricted-tool experiment.","delivery":"Interpretation — A course lead should own the educational decision and an evaluator should own the analysis plan. Schedule follow-up before students leave the course, train markers and document exclusions consistently. Dependencies include assessment capacity, accessible alternatives and privacy approval. Proposed acceptance criteria: all planned endpoints reported, denominators reconciled, independent scoring completed and delayed results reviewed before expansion. These are local gates, not observed achievements. Track attrition, support time and deviations alongside outcomes. The source's inconsistent counts illustrate why delivery teams should reconcile participant flow before communicating effect sizes. Retain human teaching support while evaluating whether the intervention changes learning."},"retrievedAt":"2026-09-10T03:01:53Z","enrichedAt":"2026-09-10T03:02:41Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Supply equivalent accessible assessment formats and human support; no disability-specific benefit or staffing saving is established.","procurementImplications":"Interpretation: Require evaluation access, configuration records and usage-cost reporting before renewing tutor licenses against learning claims.","operatingModelImplications":"Interpretation: Faculty own tasks and assessment, institutional research owns analysis, and learning technology staff own reliable access.","updateExplanation":"New URL in full archive and candidate search. Explicit 2025 evidence backfill addresses the archive's delayed-retention gap; no new September event or revised study version is claimed.","sourceVerification":{"openedUrl":"https://arxiv.org/html/2507.07357v1","referenceExcerpt":"Consequently, the final sample consisted of 123 students.","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}