{"resourceId":"generative-ai-can-harm-teaching","versions":[{"version":"legacy/2026-08-29/generative-ai-can-harm-teaching","resource":{"id":"generative-ai-can-harm-teaching","title":"Teacher-facing AI trial finds lower student motivation and uneven academic harm","organization":"University of Pennsylvania and partner researchers","sector":"Middle and high school education","geography":"Türkiye","publishedAt":"June 25, 2026","sourceName":"Generative AI Can Harm Teaching","sourceLabel":"SSRN randomized field experiment working paper","sourceUrl":"https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7007339","evidenceClass":"academic-research","outcomeClass":"cautionary","topics":["knowledge-work","governance-procurement","accessibility-workforce","operating-model"],"finding":"A semester-long randomized field experiment assigned 193 teachers across 14 middle and high schools to business as usual, a curriculum-grounded GPT-4o teaching assistant, or the assistant plus weekly reminders and usage feedback, covering 2,816 students and 14,198 student-course observations.","sledRelevance":"The study directly tests a prominent K–12 procurement proposition: that giving teachers a generative assistant for lesson planning, assessments, feedback, differentiation, and communications will benefit students as well as save staff time.","evidence":"Teacher AI access reduced student intrinsic motivation by 0.11 standard deviations and produced no statistically significant average academic gain. Students of lower-performing teachers scored 0.129 standard deviations worse; two-thirds of teacher conversations focused on producing materials, and the median interaction was only two prompts.","architectureImplications":"Design teacher tools to require curriculum context, iterative adaptation, reflection, and classroom feedback rather than one-click content generation; capture usage patterns that distinguish output substitution from instructional support.","governanceImplications":"Evaluate teacher workload and student outcomes separately, stratify results by teacher and student context, and require implementation supports that preserve professional voice and pedagogical judgment.","securityPrivacyImplications":"Keep student-identifiable information out of prompts unless the environment is specifically approved for it, and govern conversation logs used for coaching or evaluation as potentially sensitive personnel and education records.","caveats":"The paper is a working paper rather than a peer-reviewed journal article; it covers one private-school network, one semester, and one custom tool. Subgroup effects and proposed mechanisms need replication."}},{"version":"enrichment/2026-09-05T02:33:27.019Z/generative-ai-can-harm-teaching","resource":{"id":"generative-ai-can-harm-teaching","title":"Teacher-facing AI trial finds lower student motivation and uneven academic harm","organization":"University of Pennsylvania and partner researchers","sector":"Middle and high school education","geography":"Türkiye","publishedAt":"June 25, 2026","publicationDate":"2026-06-25","eventDate":null,"sourceName":"Generative AI Can Harm Teaching","sourceLabel":"SSRN randomized field experiment working paper","sourceUrl":"https://papers.ssrn.com/sol3/papers.cfm?abstract_id=7007339","evidenceClass":"academic-research","outcomeClass":"cautionary","topics":["knowledge-work","governance-procurement","accessibility-workforce","operating-model"],"finding":"A semester-long randomized field experiment assigned 193 teachers across 14 middle and high schools to business as usual, a curriculum-grounded GPT-4o teaching assistant, or the assistant plus weekly reminders and usage feedback, covering 2,816 students and 14,198 student-course observations.","sledRelevance":"The study directly tests a prominent K–12 procurement proposition: that giving teachers a generative assistant for lesson planning, assessments, feedback, differentiation, and communications will benefit students as well as save staff time.","evidence":"Teacher AI access reduced student intrinsic motivation by 0.11 standard deviations and produced no statistically significant average academic gain. Students of lower-performing teachers scored 0.129 standard deviations worse; two-thirds of teacher conversations focused on producing materials, and the median interaction was only two prompts.","architectureImplications":"Design teacher tools to require curriculum context, iterative adaptation, reflection, and classroom feedback rather than one-click content generation; capture usage patterns that distinguish output substitution from instructional support.","governanceImplications":"Evaluate teacher workload and student outcomes separately, stratify results by teacher and student context, and require implementation supports that preserve professional voice and pedagogical judgment.","securityPrivacyImplications":"Keep student-identifiable information out of prompts unless the environment is specifically approved for it, and govern conversation logs used for coaching or evaluation as potentially sensitive personnel and education records.","caveats":"The paper is a working paper rather than a peer-reviewed journal article; it covers one private-school network, one semester, and one custom tool. Subgroup effects and proposed mechanisms need replication.","streamIds":["k12"],"roles":{"sales":"Interpretation — Customer problem: a teacher-productivity purchase can overlook student motivation and learning even when materials are produced faster. Stakeholders: curriculum, assessment, teacher development, school leadership, student services, privacy, and teachers. Discovery: will the tool support reflection or substitute for instructional judgment; which student outcomes are tracked; and could effects differ by teacher context? Value hypothesis: a supported, carefully evaluated workflow may address workload without sacrificing educational goals, but this record argues for testing that hypothesis. Potential engagement: instructional-use review and a monitored, limited evaluation. Unsupported claims: the single-network working paper does not prove all teacher AI is harmful, while its lack of average academic gain cannot support a promise of improved attainment or universal time-to-learning conversion.","engineering":"Interpretation — Fit: assess teacher-facing assistance for curriculum-grounded planning, assessment, and feedback while preserving professional adaptation. Architecture and integration: require curriculum context, iterative review, and classroom feedback in the workflow; capture enough usage context to distinguish material substitution from instructional support. Prerequisites: instructional rubrics, approved data handling, and a separate student-outcome evaluation design. Constraints: a median two-prompt interaction and material-production focus in the study suggest usage quality needs examination; one custom tool and one semester limit transfer. Security: exclude student identifiers from unapproved prompts and govern coaching logs as potentially sensitive education/personnel information. Proposed proof: evaluate reviewed materials and patterns of adaptation, then assess motivation and academic outcomes separately from workload and tool usage.","delivery":"Interpretation — Work: develop teacher coaching, build curriculum and reflection checkpoints, and run an evaluation that separates workload from student motivation and attainment. Dependencies: assessment expertise, teacher participation, privacy-approved records, and sufficient time to observe classroom effects. Ownership: curriculum and school leaders own instructional quality; an evaluation lead defines measures; teachers retain professional judgment; privacy/HR owners govern sensitive coaching data. Skills and adoption: practice adapting outputs to students rather than accepting one-shot materials. Governance checkpoints: evaluation design, classroom readiness, interim harm review, and expansion decision. Proposed acceptance: report workload and student measures separately, examine relevant subgroup patterns cautiously, and apply agreed pause criteria if educational outcomes deteriorate. Risks include hidden subgroup harm, surveillance-like coaching, and overgeneralizing an unreplicated working paper."},"retrievedAt":null,"enrichedAt":"2026-09-05T02:33:27.019Z","enrichmentBasis":"archived evidence"}}]}