A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol

📅 2026-07-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study systematically evaluates the stability and applicability of a validated instructional feedback classification protocol across evolving text representation methods—ranging from sparse features and frozen Transformer embeddings to prompting large language models—and under English–Spanish cross-lingual transfer. Through stratified cross-validation and human annotation consistency analysis, the research finds that while state-of-the-art large models circa 2026 achieve the highest F1 scores on Spanish topic classification, they do not significantly outperform lightweight models in sentiment classification or English tasks. These results demonstrate the protocol’s robustness over time and across languages, suggesting that model selection should prioritize deployment constraints over methodological novelty. This work provides the first empirical evidence of the resilience of an instructional feedback classification framework amid technological advances and multilingual scenarios.
📝 Abstract
Institutions collect far more open-ended teaching-evaluation feedback than they read. A prior study introduced a validated protocol for classifying such comments by thematic category and sentiment, built from a documented annotation guide, an intra-annotator reliability measurement, stratified cross-validation, and a held-out evaluation on a Spanish institutional corpus with a frozen-encoder design. Two questions limit its reuse: whether a protocol fixed to 2019-era frozen embeddings stays competitive as representation methods advance, and whether it transfers to a second language. We re-run it on the original Spanish data across three representation generations, sparse lexical features, frozen transformer embeddings, and prompted large language models, and transfer its sentiment task to English with a balanced 45,000-comment corpus checked against an aspect-labeled education dataset. Treating paired comparisons as descriptive, we find the protocol durable: a 2026 frontier model posts the highest thematic F1 on the hardest Spanish task, yet shows no sentiment advantage over a cheap model and no descriptive separation from it on English, so model choice is a deployment decision, not a property of the method.
Problem

Research questions and friction points this paper is trying to address.

durability
cross-language transfer
teaching-feedback classification
representation methods
sentiment analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

durability
cross-language transfer
teaching-feedback classification
frozen-encoder
large language models
💼 Related Jobs
No related jobs found.
E
Esteban U. Vega Barajas
Universidad de Guadalajara