Institution profile

Educational Testing Service

Academic institutionnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs

Mar 02, 2026

This study addresses the threat posed by high-quality AI-generated essays to the authenticity of writing assessment and the limited generalization capability of existing detectors in cross–large language model (LLM) scenarios. It presents the first systematic evaluation of mainstream AI text detectors on essays generated by multiple LLMs, constructing a multi-source dataset based on publicly available GRE prompts to conduct an empirical analysis of cross-model generalization. The findings reveal a significant performance drop in current detectors when applied across different LLMs. Building on these insights, the work proposes actionable retraining strategies and responsible deployment guidelines to enhance the robustness and practical utility of detection tools in real-world educational settings.

0 citationsRead paper

Can ChatGPT Code Communication Data Fairly?: Empirical Evidence from Multiple Collaborative Tasks

Oct 23, 2025

This study investigates whether large language models (LLMs), specifically ChatGPT, exhibit gender or racial bias in the automated coding of collaborative communication data. Method: Drawing on three canonical collaborative tasks—negotiation, problem solving, and decision making—we develop a human-annotated, rule-based multimodal coding framework and conduct the first systematic fairness evaluation of ChatGPT’s coding performance across multiple collaborative tasks. Contribution/Results: Experimental results show no statistically significant bias along gender or racial dimensions (p > 0.05); ChatGPT’s coding accuracy is not significantly different from human annotators (Δ < 2.1%, p = 0.12). This work fills a critical empirical gap in fairness research on LLMs for collaborative assessment and demonstrates ChatGPT’s reliability and scalability for high-throughput, large-scale measurement of collaborative competence.

0 citationsRead paper

A Survey of Idiom Datasets for Psycholinguistic and Computational Research

Aug 15, 2025

Idiomatic expressions’ non-compositionality poses persistent modeling and experimental challenges in both psycholinguistics and computational linguistics, exacerbated by a longstanding lack of coordinated data resources across disciplines. This paper conducts the first systematic meta-analysis of 53 cross-disciplinary idiom datasets, evaluating them along three dimensions: annotation schemas (e.g., familiarity, transparency, idiomaticity), task designs (e.g., comprehension, generation, acceptability judgment), and linguistic coverage. Results reveal a fundamental misalignment: psycholinguistic datasets prioritize controlled experimental variables and fine-grained behavioral annotations, whereas computational datasets emphasize downstream task compatibility and model-oriented evaluation—leading to significant gaps in annotation granularity, evaluation metrics, and multilingual support. The study identifies critical deficiencies in cross-disciplinary resource integration and proposes a unified framework for constructing idiomatic data that supports multi-task learning, multi-paradigm evaluation, and multilingual generalization. This work provides an empirical foundation and methodological roadmap for advancing interdisciplinary research on idioms.

0 citationsRead paper
Recent publications

Latest Papers

Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMs

Mar 02, 2026

This study addresses the threat posed by high-quality AI-generated essays to the authenticity of writing assessment and the limited generalization capability of existing detectors in cross–large language model (LLM) scenarios. It presents the first systematic evaluation of mainstream AI text detectors on essays generated by multiple LLMs, constructing a multi-source dataset based on publicly available GRE prompts to conduct an empirical analysis of cross-model generalization. The findings reveal a significant performance drop in current detectors when applied across different LLMs. Building on these insights, the work proposes actionable retraining strategies and responsible deployment guidelines to enhance the robustness and practical utility of detection tools in real-world educational settings.

0 citationsRead paper

Can ChatGPT Code Communication Data Fairly?: Empirical Evidence from Multiple Collaborative Tasks

Oct 23, 2025

This study investigates whether large language models (LLMs), specifically ChatGPT, exhibit gender or racial bias in the automated coding of collaborative communication data. Method: Drawing on three canonical collaborative tasks—negotiation, problem solving, and decision making—we develop a human-annotated, rule-based multimodal coding framework and conduct the first systematic fairness evaluation of ChatGPT’s coding performance across multiple collaborative tasks. Contribution/Results: Experimental results show no statistically significant bias along gender or racial dimensions (p > 0.05); ChatGPT’s coding accuracy is not significantly different from human annotators (Δ < 2.1%, p = 0.12). This work fills a critical empirical gap in fairness research on LLMs for collaborative assessment and demonstrates ChatGPT’s reliability and scalability for high-throughput, large-scale measurement of collaborative competence.

0 citationsRead paper

A Survey of Idiom Datasets for Psycholinguistic and Computational Research

Aug 15, 2025

Idiomatic expressions’ non-compositionality poses persistent modeling and experimental challenges in both psycholinguistics and computational linguistics, exacerbated by a longstanding lack of coordinated data resources across disciplines. This paper conducts the first systematic meta-analysis of 53 cross-disciplinary idiom datasets, evaluating them along three dimensions: annotation schemas (e.g., familiarity, transparency, idiomaticity), task designs (e.g., comprehension, generation, acceptability judgment), and linguistic coverage. Results reveal a fundamental misalignment: psycholinguistic datasets prioritize controlled experimental variables and fine-grained behavioral annotations, whereas computational datasets emphasize downstream task compatibility and model-oriented evaluation—leading to significant gaps in annotation granularity, evaluation metrics, and multilingual support. The study identifies critical deficiencies in cross-disciplinary resource integration and proposes a unified framework for constructing idiomatic data that supports multi-task learning, multi-paradigm evaluation, and multilingual generalization. This work provides an empirical foundation and methodological roadmap for advancing interdisciplinary research on idioms.

0 citationsRead paper