Institution profile

National Board of Medical Examiners

Academic institutionnorthamerica · us
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

FinRAG-12B: A Production-Validated Recipe for Grounded Question Answering in Banking

May 06, 2026

This study addresses the challenges of accuracy, regulatory compliance, and verifiability in deploying large language models (LLMs) within the banking sector by proposing a data-efficient, end-to-end framework. The framework integrates LLM-as-a-Judge filtering, citation annotation, curriculum learning, and a calibrated rejection mechanism, while supporting quantized deployment across the entire pipeline—from data construction to efficient inference. A domain-specific model trained on only 143 million tokens surpasses GPT-4.1 in citation accuracy and demonstrates substantially improved rejection behavior on unanswerable queries. Deployed across more than 40 financial institutions, the system increases query resolution rates by 7.1 percentage points (p<0.001), accelerates response times by 3–5×, and reduces inference costs by 20–50×.

0 citationsRead paper

Enhancing Security and Strengthening Defenses in Automated Short-Answer Grading Systems

Apr 30, 2025

This study exposes the severe vulnerability of Transformer-based automated short-answer scoring systems in medical education under adversarial gaming. We systematically identify three novel classes of short-answer gaming strategies targeting such systems. To enhance robustness, we propose a defense framework integrating adversarial training, ensemble voting, and ridge regression. Furthermore, we introduce a GPT-4–driven multi-prompt mechanism for detecting gaming behavior—a first in this domain. Experimental results demonstrate that our framework significantly reduces misclassification rates; the ensemble component improves defense success rate by over 40%. Leveraging diverse prompt engineering, GPT-4 achieves 89.2% accuracy in identifying gaming attempts. This work provides the first systematic security analysis and robustness enhancement solution tailored specifically to short-answer scoring in AI-powered educational assessment, advancing trustworthy AI for high-stakes medical education evaluation.

0 citationsRead paper
Recent publications

Latest Papers

FinRAG-12B: A Production-Validated Recipe for Grounded Question Answering in Banking

May 06, 2026

This study addresses the challenges of accuracy, regulatory compliance, and verifiability in deploying large language models (LLMs) within the banking sector by proposing a data-efficient, end-to-end framework. The framework integrates LLM-as-a-Judge filtering, citation annotation, curriculum learning, and a calibrated rejection mechanism, while supporting quantized deployment across the entire pipeline—from data construction to efficient inference. A domain-specific model trained on only 143 million tokens surpasses GPT-4.1 in citation accuracy and demonstrates substantially improved rejection behavior on unanswerable queries. Deployed across more than 40 financial institutions, the system increases query resolution rates by 7.1 percentage points (p<0.001), accelerates response times by 3–5×, and reduces inference costs by 20–50×.

0 citationsRead paper

Enhancing Security and Strengthening Defenses in Automated Short-Answer Grading Systems

Apr 30, 2025

This study exposes the severe vulnerability of Transformer-based automated short-answer scoring systems in medical education under adversarial gaming. We systematically identify three novel classes of short-answer gaming strategies targeting such systems. To enhance robustness, we propose a defense framework integrating adversarial training, ensemble voting, and ridge regression. Furthermore, we introduce a GPT-4–driven multi-prompt mechanism for detecting gaming behavior—a first in this domain. Experimental results demonstrate that our framework significantly reduces misclassification rates; the ensemble component improves defense success rate by over 40%. Leveraging diverse prompt engineering, GPT-4 achieves 89.2% accuracy in identifying gaming attempts. This work provides the first systematic security analysis and robustness enhancement solution tailored specifically to short-answer scoring in AI-powered educational assessment, advancing trustworthy AI for high-stakes medical education evaluation.

0 citationsRead paper