{"resourceId":"edm-critical-thinking-subskills-2026","versions":[{"version":"external-5720f70db75cc4bd8af9b029f1d7158cb5855b9a849a6fc3a81b31eb867430c4","resource":{"id":"edm-critical-thinking-subskills-2026","title":"Essay-scoring study exposes large differences between critical-thinking subskills","organization":"Peczuh, Kumar and colleagues; university, ETS and AERDF collaborators","sector":"K–12 education","geography":"United States student essays; international research collaboration","publishedAt":"June 2026 proceedings; exact publication day unknown","publicationDate":null,"eventDate":null,"sourceName":"Toward LLM-supported Automated Assessment of Critical Thinking Subskills","sourceLabel":"EDM 2026 original research paper","sourceUrl":"https://educationaldatamining.org/wp-content/uploads/2026/proceedings/2026.EDM.full-papers/2026.EDM.full-papers.174.pdf","evidenceClass":"academic-research","outcomeClass":"mixed","topics":["knowledge-work","developers-agents","governance-procurement","accessibility-workforce"],"finding":"Fine-tuned Llama performed best overall, but aggregate accuracy concealed weak agreement on some subskills.","sledRelevance":"Interpretation: Newly added to this archive for fall 2026 district decisions. Supports validation of AI-assisted assessment before consequential use.","evidence":"Exploratory comparison of prompting and fine-tuning against human labels on 500 U.S. grade 6–12 essays, averaged over five random essay splits. Table 6 reports evidence-strength accuracy .768 but Krippendorff’s alpha .234; source-synthesis accuracy .780 and alpha .740. These are scoring metrics, not learning gains.","architectureImplications":"Interpretation: Keep rubric, task materials, model version and scoring service separately versioned. Compare hosted inference with managed open-weight deployment on total support effort, not model size alone.","governanceImplications":"Interpretation: Assessment owners should approve each scored construct and an appeal route before scores influence pupils.","securityPrivacyImplications":"Interpretation: Redact essay identifiers, restrict reviewer access and prohibit unintended training reuse; local hosting also needs access and retention controls.","caveats":"Prompts and source materials were unavailable; labels were imbalanced. The exploratory human reliability threshold was .6, with an exception for facts/opinions. No classroom intervention or durable-learning effect was tested.","streamIds":["k12"],"roles":{"sales":"Interpretation: Assessment and curriculum leaders may need faster feedback without mischaracterizing student reasoning. Ask which subskills, grades and consequences are in scope, who handles contested scores, and whether teachers have usable task context. A bounded engagement could compare a proposed scoring workflow with independently adjudicated local work. The value hypothesis is better evidence for a purchasing decision, not guaranteed grading savings. Do not sell the reported accuracy as mastery detection, fairness assurance or improved learning. Applicability is limited to a deliberately defined assessment task.","engineering":"Interpretation: Build a shadow-scoring proof of value with no automated grade write-back. Require consented or otherwise authorized essays, a local rubric, independent labels and held-out tasks. Capture model and prompt versions and test hosted versus managed inference under district security controls. Measure agreement beyond accuracy and errors for rare proficiency levels; investigate disagreements with task materials visible to reviewers. Define latency and support-cost requirements locally. Do not extrapolate a model ranking to other grades or constructs without new validation.","delivery":"Interpretation: The assessment lead owns the rubric and adjudication; IT owns the scoring service and privacy staff approve data use. Train reviewers, create an appeal workflow and schedule a decision checkpoint before classroom exposure. Proposed acceptance requires completing a pre-agreed validation sample, reporting every subskill separately and meeting locally approved error thresholds, with no automatic consequential decisions. Track review effort and teacher adoption alongside scores. Main risks are unrepresentative labels, rubric drift and overreliance on an apparently precise output."},"retrievedAt":"2026-09-11T03:01:41Z","enrichedAt":"2026-09-11T03:04:43Z","enrichmentBasis":"retrieved source","accessibilityWorkforceImplications":"Interpretation: Include multilingual writing and accommodated tasks in local validation; budget specialist time to adjudicate disagreements.","procurementImplications":"Interpretation: Request per-subskill confusion matrices and external validation, not one overall accuracy claim.","operatingModelImplications":"Interpretation: Keep teacher judgment authoritative and review disagreement trends after every material model change.","sourceVerification":{"openedUrl":"https://educationaldatamining.org/wp-content/uploads/2026/proceedings/2026.EDM.full-papers/2026.EDM.full-papers.174.pdf","referenceExcerpt":"We did not have access to the essay prompt, writing task, or source materials","promptVersion":"sled-research-v3.1","model":null,"basis":"agent-reported inspection"}}}]}