Institution profile

Pearson School Research Network

Industry researchnorthamerica · us
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Calibrating Generative AI to Produce Realistic Essays for Data Augmentation

Feb 06, 2026

This study addresses the performance bottleneck in machine learning–based automated essay scoring systems caused by limited training data by proposing the use of large language models to generate synthetic student essays for data augmentation. The work presents the first empirical evaluation of three prompting strategies—“next-sentence prediction,” “sentence-level prompting,” and “25-shot exemplars”—systematically comparing their effectiveness in generating text that preserves original essay quality and exhibits human-like authenticity. Results indicate that the next-sentence prediction strategy achieves the highest scoring consistency and, alongside sentence-level prompting, best retains the quality of the source essays. Moreover, texts generated via next-sentence prediction and the 25-shot approach demonstrate the greatest authenticity. This research provides both effective strategies and empirical evidence supporting the use of synthetic data augmentation in automated essay scoring.

0 citationsRead paper

Automatic Detection of Inauthentic Templated Responses in English Language Assessments

Sep 10, 2025

In high-stakes English proficiency testing, low-proficiency test-takers frequently resort to memorized essay templates to circumvent automated scoring systems, thereby compromising scoring fairness and validity. This paper formally introduces the task of Automated Detection of Template-Driven Responses (AuDITR), aimed at identifying inauthentic, highly formulaic responses. Methodologically, we design a set of multidimensional textual features—including syntactic repetitiveness, semantic rigidity, and lexical distribution anomalies—and integrate them into a lightweight machine learning classifier optimized for efficiency and adaptability to emerging template strategies. Experimental evaluation on authentic high-stakes test data yields an F1-score of 0.86, substantially outperforming established baselines. Our work contributes both a novel, well-defined detection task and a deployable technical framework to enhance the adversarial robustness of automated scoring systems against template-based cheating.

0 citationsRead paper

Toward Subtrait-Level Model Explainability in Automated Writing Evaluation

Sep 10, 2025

This study addresses the lack of interpretability and transparency in subtrait scoring models within automated writing evaluation (AWE), where “black-box” decisions hinder educators’ and students’ understanding of fine-grained scoring rationales. To bridge this gap, we propose the first application of generative language models (GLMs) for interpretable subtrait-level modeling: GLMs generate natural-language explanations for subtrait scores and quantify alignment with human judgments. Correlation analyses demonstrate moderate agreement between GLM-predicted subtrait scores and human ratings (r ≈ 0.4–0.6), while all subtraits exhibit statistically significant correlations with holistic scores (p < 0.01). Our approach enhances AWE interpretability by grounding explanations in linguistically coherent, human-aligned reasoning, and provides actionable, fine-grained diagnostic feedback to support human-AI collaborative assessment.

0 citationsRead paper

Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi

Apr 28, 2025

The feasibility and performance limits of large language models (LLMs) in replacing human experts for qualitative coding—a core task in systematic reviews—remain empirically unestablished. Method: This study conducts the first empirical comparison of GPT-4 and Kimi on real-world systematic review data, evaluating inter-coder agreement (measured by Cohen’s κ) across varying data scales and question complexities, with human expert coding as the benchmark. Contribution/Results: Both LLMs achieve near-human consistency (κ ≥ 0.85) on small-scale, structured coding tasks but exhibit substantial degradation (κ ≤ 0.50) on open-ended, high-complexity problems. The findings delineate an evidence-based applicability threshold for LLM-assisted evidence synthesis and propose a task-characteristic–driven framework for model selection and human–AI collaboration. This work provides methodological guidance and empirical validation for the trustworthy integration of LLMs into systematic review workflows.

0 citationsRead paper
Recent publications

Latest Papers

Calibrating Generative AI to Produce Realistic Essays for Data Augmentation

Feb 06, 2026

This study addresses the performance bottleneck in machine learning–based automated essay scoring systems caused by limited training data by proposing the use of large language models to generate synthetic student essays for data augmentation. The work presents the first empirical evaluation of three prompting strategies—“next-sentence prediction,” “sentence-level prompting,” and “25-shot exemplars”—systematically comparing their effectiveness in generating text that preserves original essay quality and exhibits human-like authenticity. Results indicate that the next-sentence prediction strategy achieves the highest scoring consistency and, alongside sentence-level prompting, best retains the quality of the source essays. Moreover, texts generated via next-sentence prediction and the 25-shot approach demonstrate the greatest authenticity. This research provides both effective strategies and empirical evidence supporting the use of synthetic data augmentation in automated essay scoring.

0 citationsRead paper

Automatic Detection of Inauthentic Templated Responses in English Language Assessments

Sep 10, 2025

In high-stakes English proficiency testing, low-proficiency test-takers frequently resort to memorized essay templates to circumvent automated scoring systems, thereby compromising scoring fairness and validity. This paper formally introduces the task of Automated Detection of Template-Driven Responses (AuDITR), aimed at identifying inauthentic, highly formulaic responses. Methodologically, we design a set of multidimensional textual features—including syntactic repetitiveness, semantic rigidity, and lexical distribution anomalies—and integrate them into a lightweight machine learning classifier optimized for efficiency and adaptability to emerging template strategies. Experimental evaluation on authentic high-stakes test data yields an F1-score of 0.86, substantially outperforming established baselines. Our work contributes both a novel, well-defined detection task and a deployable technical framework to enhance the adversarial robustness of automated scoring systems against template-based cheating.

0 citationsRead paper

Toward Subtrait-Level Model Explainability in Automated Writing Evaluation

Sep 10, 2025

This study addresses the lack of interpretability and transparency in subtrait scoring models within automated writing evaluation (AWE), where “black-box” decisions hinder educators’ and students’ understanding of fine-grained scoring rationales. To bridge this gap, we propose the first application of generative language models (GLMs) for interpretable subtrait-level modeling: GLMs generate natural-language explanations for subtrait scores and quantify alignment with human judgments. Correlation analyses demonstrate moderate agreement between GLM-predicted subtrait scores and human ratings (r ≈ 0.4–0.6), while all subtraits exhibit statistically significant correlations with holistic scores (p < 0.01). Our approach enhances AWE interpretability by grounding explanations in linguistically coherent, human-aligned reasoning, and provides actionable, fine-grained diagnostic feedback to support human-AI collaborative assessment.

0 citationsRead paper

Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi

Apr 28, 2025

The feasibility and performance limits of large language models (LLMs) in replacing human experts for qualitative coding—a core task in systematic reviews—remain empirically unestablished. Method: This study conducts the first empirical comparison of GPT-4 and Kimi on real-world systematic review data, evaluating inter-coder agreement (measured by Cohen’s κ) across varying data scales and question complexities, with human expert coding as the benchmark. Contribution/Results: Both LLMs achieve near-human consistency (κ ≥ 0.85) on small-scale, structured coding tasks but exhibit substantial degradation (κ ≤ 0.50) on open-ended, high-complexity problems. The findings delineate an evidence-based applicability threshold for LLM-assisted evidence synthesis and propose a task-characteristic–driven framework for model selection and human–AI collaboration. This work provides methodological guidance and empirical validation for the trustworthy integration of LLMs into systematic review workflows.

0 citationsRead paper