Institution profile

Eedi

Industry research
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs

Mar 03, 2026

This study addresses the challenge of efficiently and accurately predicting student response behavior in educational platforms by conducting the first systematic comparison between specialized Knowledge Tracing (KT) models and large language models (LLMs) on real-world student interaction data. Through quantitative evaluation of prediction accuracy, inference latency, and deployment cost, the research demonstrates that KT models significantly outperform LLMs in both accuracy and F1 score, achieve inference speeds several orders of magnitude faster, and incur substantially lower deployment costs. These findings reveal that, for educational prediction tasks, domain-specific models offer marked advantages over general-purpose large language models in both performance and economic efficiency, thereby providing empirical support and practical guidance for model selection in personalized learning interventions.

0 citationsRead paper

Misconception Diagnosis From Student-Tutor Dialogue: Generate, Retrieve, Rerank

Feb 02, 2026

This study addresses the challenge of timely and accurate identification of students’ misconceptions in educational dialogues by proposing a novel three-stage approach that integrates generation, retrieval, and re-ranking. The method first leverages a large language model to generate potential misconceptions, then retrieves a candidate set based on embedding similarity, and finally employs a fine-tuned small model to re-rank candidates for improved relevance. To the best of our knowledge, this work is the first to synergistically apply these techniques to misconception diagnosis. Evaluated on real-world instructional dialogue data, the approach significantly outperforms baseline models. Experimental results demonstrate that fine-tuning effectively enhances generation quality, ablation studies confirm the necessity of each component, and the fine-tuned small model surpasses larger closed-source counterparts, highlighting the method’s efficiency and practicality.

0 citationsRead paper

Uncertainty-Aware Knowledge Tracing Models

Sep 25, 2025

Knowledge tracing (KT) models often fail to detect students’ erroneous selections of distractors, leaving latent misconceptions undiagnosed. To address this, we introduce predictive uncertainty modeling into KT for the first time, leveraging probabilistic deep learning to quantify per-prediction confidence. Experiments demonstrate a statistically significant positive correlation between high uncertainty and model misclassification (p < 0.01), enabling effective identification of student cognitive biases. Our approach requires no additional annotations or pedagogical interventions, yielding interpretable instructional signals that support precise diagnosis and adaptive remediation—even under resource constraints. Key contributions are: (1) the first uncertainty-aware KT framework; (2) empirical validation that uncertainty serves as a reliable proxy for prediction errors; and (3) a novel, operationally viable analytical dimension for trustworthy educational AI—balancing reliability with practical deployability.

0 citationsRead paper

PIIvot: A Lightweight NLP Anonymization Framework for Question-Anchored Tutoring Dialogues

May 22, 2025

To address the recall-precision trade-off in PII anonymization for educational tutoring dialogues, this paper proposes a lightweight, context-aware anonymization framework leveraging question-answer structural priors. Methodologically, it introduces question-answer anchoring to simplify PII identification, integrating rule-guided named entity recognition (NER), context-sensitive pattern matching, and a lightweight sequence labeling model. Key contributions include: (1) releasing QATD-2k—the largest publicly available real-world educational dialogue dataset for PII anonymization research; (2) achieving a 12.6% F1-score improvement over prior methods in educational dialogue contexts, with inference speed of 320 tokens/second; and (3) enabling end-to-end, low-latency de-identification, already deployed in multiple educational AI data governance pipelines. The framework balances accuracy, efficiency, and practical deployability without compromising anonymization robustness.

0 citationsRead paper
Recent publications

Latest Papers

Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs

Mar 03, 2026

This study addresses the challenge of efficiently and accurately predicting student response behavior in educational platforms by conducting the first systematic comparison between specialized Knowledge Tracing (KT) models and large language models (LLMs) on real-world student interaction data. Through quantitative evaluation of prediction accuracy, inference latency, and deployment cost, the research demonstrates that KT models significantly outperform LLMs in both accuracy and F1 score, achieve inference speeds several orders of magnitude faster, and incur substantially lower deployment costs. These findings reveal that, for educational prediction tasks, domain-specific models offer marked advantages over general-purpose large language models in both performance and economic efficiency, thereby providing empirical support and practical guidance for model selection in personalized learning interventions.

0 citationsRead paper

Misconception Diagnosis From Student-Tutor Dialogue: Generate, Retrieve, Rerank

Feb 02, 2026

This study addresses the challenge of timely and accurate identification of students’ misconceptions in educational dialogues by proposing a novel three-stage approach that integrates generation, retrieval, and re-ranking. The method first leverages a large language model to generate potential misconceptions, then retrieves a candidate set based on embedding similarity, and finally employs a fine-tuned small model to re-rank candidates for improved relevance. To the best of our knowledge, this work is the first to synergistically apply these techniques to misconception diagnosis. Evaluated on real-world instructional dialogue data, the approach significantly outperforms baseline models. Experimental results demonstrate that fine-tuning effectively enhances generation quality, ablation studies confirm the necessity of each component, and the fine-tuned small model surpasses larger closed-source counterparts, highlighting the method’s efficiency and practicality.

0 citationsRead paper

Uncertainty-Aware Knowledge Tracing Models

Sep 25, 2025

Knowledge tracing (KT) models often fail to detect students’ erroneous selections of distractors, leaving latent misconceptions undiagnosed. To address this, we introduce predictive uncertainty modeling into KT for the first time, leveraging probabilistic deep learning to quantify per-prediction confidence. Experiments demonstrate a statistically significant positive correlation between high uncertainty and model misclassification (p < 0.01), enabling effective identification of student cognitive biases. Our approach requires no additional annotations or pedagogical interventions, yielding interpretable instructional signals that support precise diagnosis and adaptive remediation—even under resource constraints. Key contributions are: (1) the first uncertainty-aware KT framework; (2) empirical validation that uncertainty serves as a reliable proxy for prediction errors; and (3) a novel, operationally viable analytical dimension for trustworthy educational AI—balancing reliability with practical deployability.

0 citationsRead paper

PIIvot: A Lightweight NLP Anonymization Framework for Question-Anchored Tutoring Dialogues

May 22, 2025

To address the recall-precision trade-off in PII anonymization for educational tutoring dialogues, this paper proposes a lightweight, context-aware anonymization framework leveraging question-answer structural priors. Methodologically, it introduces question-answer anchoring to simplify PII identification, integrating rule-guided named entity recognition (NER), context-sensitive pattern matching, and a lightweight sequence labeling model. Key contributions include: (1) releasing QATD-2k—the largest publicly available real-world educational dialogue dataset for PII anonymization research; (2) achieving a 12.6% F1-score improvement over prior methods in educational dialogue contexts, with inference speed of 320 tokens/second; and (3) enabling end-to-end, low-latency de-identification, already deployed in multiple educational AI data governance pipelines. The framework balances accuracy, efficiency, and practical deployability without compromising anonymization robustness.

0 citationsRead paper