Shamya Karumbaiah
2026
LLM-Generated Stories for Students with Significant Cognitive Disabilities: Promise, Gaps, and Evaluation Framework
Pragati Maheshwary | Ananya Ganesh | Shamya Karumbaiah
Proceedings of the Joint Workshop on Readability and Text Simplification (READIxTSAR) @ LREC 2026
Pragati Maheshwary | Ananya Ganesh | Shamya Karumbaiah
Proceedings of the Joint Workshop on Readability and Text Simplification (READIxTSAR) @ LREC 2026
Students with significant cognitive disabilities (SCD) require specially designed accessible stories for reading comprehension assessments, yet creating such content is labor-intensive and difficult to scale. This preliminary study investigates whether large language models (LLMs) can generate short accessible stories for alternate assessment system. Using an 8-fold cross-validation design, we generated 120 stories with GPT-4o via one-shot prompting with human-written exemplars and evaluated them against a test set comprising 7 expert-human written stories as baselines across three dimensions: simplicity, fluency & coherence, and thematic adherence. Cross-validation results show that generated stories meet surface-level simplicity targets, with approximately two-thirds falling within the human baseline range for readability metrics. However, generated stories exhibited a systematic coherence gap where only 5% fell within the human range for adjacent sentence similarity, a pattern consistent across all folds. Thematic adherence was moderate, with adequate diversity across stories. These findings suggest LLMs can serve as a drafting tool within accessible content generation pipelines, but human expert review remains essential to ensure coherence, testability, and alignment with quality standards required for high-stakes alternate assessments.
2025
Identifying Biases in Large Language Model Assessment of Linguistically Diverse Texts
Lionel Hsien Meng | Shamya Karumbaiah | Vivek Saravanan | Daniel Bolt
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
Lionel Hsien Meng | Shamya Karumbaiah | Vivek Saravanan | Daniel Bolt
Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress
The development of Large Language Models (LLMs) to assess student text responses is rapidly progressing but evaluating whether LLMs equitably assess multilingual learner responses is an important precursor to adoption. Our study provides an example procedure for identifying and quantifying bias in LLM assessment of student essay responses.