GenAI-HITL Development and Validity Review of Reading-Writing Constructed-Response Tasks
Park Goun
Korea National University of Education, Ph.D. Student
Korean Language Education Research Vol. 60 No. 4 pp.129-174 (2025)
Abstract
Park Goun This study developed constructed-response tasks aligned with the 2022 revised Korean Language Arts (“Reading and Writing”) achievement standards using a Generative AI-Human-in-the-Loop (HITL) approach and reported preliminary validity evidence from expert-criterion ratings. A three-stage protocol integrating context engineering and Chain-of- Thought generated a linked package of texts, prompts, rubrics, and ex- planations. Eighteen in-service Korean language teachers from 13 regions rated the outputs on 12 items across three domains. The overall mean was high (M = 4.32), with the strongest ratings for standard alignment and structural coherence; internal consistency and inter-rater agreement were acceptable. Learner-level appropriateness was lower, indicating limits in capturing non-formal factors (developmental stage, classroom context) despite effective operationalisation of formalised curriculum elements. The findings suggest AI outputs can serve as teacher-adjustable drafts, and that human-AI collaboration may strengthen the balance between formal and substantive validity.
Keywords
Automatic item generationgenerative AIhuman–AI collaborationdescriptive assessmentreading and writing
