A Nationwide Japanese Medical Claims Foundation Model: Balancing Model Scaling and Task-Specific Computational Efficiency

📅 2026-04-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the relationship between model scale and downstream task performance in structured healthcare data, an aspect that remains poorly understood. Leveraging insurance claims data from 519 hospitals in Japan, the authors develop and evaluate five Encoder-only Transformer foundation models ranging from 2.2M to 101M parameters for predicting disease onset and medication use. The work reveals, for the first time, a task-dependent performance saturation effect: disease prediction benefits from larger models, whereas medication prediction achieves optimal performance at just 11M parameters—reducing pretraining time by 178 hours. Across all tasks, the best-performing models consistently outperform a LightGBM baseline as measured by PR-AUC, offering empirical guidance for the efficient deployment of foundation models in healthcare settings.

Technology Category

Application Category

📝 Abstract
Clinical risk prediction using longitudinal medical data supports individualized care. Self-supervised foundation models have emerged as a promising approach for leveraging large-scale unlabeled healthcare records. In natural language processing, scaling laws suggest that larger models achieve predictably lower pretraining losses, supporting the foundation model paradigm. However, for structured medical data, characterized by a limited vocabulary and sparse observations, whether increasing model size consistently improves downstream predictions is unclear, as most studies evaluate only a single model scale. In this study, we evaluated the relationship between model scale and downstream task performance for structured medical foundation models. Using a random sample (2.3 million patients, 32 hospitals) from a nationwide 519-hospital Japanese claims database, we pretrained encoder-only Transformers at five scales (2.2M-101M parameters) for disease incidence and medication prediction. Downstream performance saturated at task-dependent thresholds: disease prediction benefited from larger models (32M-101M), whereas medication prediction saturated at 11M, reducing pretraining time by 178 h. Across all tasks, the best-performing model consistently outperformed a Light Gradient Boosting Machine baseline in the area under the precision-recall curve. These findings indicate that, unlike the monotonically decreasing pretraining loss, the optimal model size varied depending on task characteristics. This task-dependent saturation provides practical guidance for balancing predictive performance and computational cost in structured medical foundation models.
Problem

Research questions and friction points this paper is trying to address.

foundation models
model scaling
structured medical data
clinical risk prediction
computational efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

foundation model
model scaling
structured medical data
task-dependent saturation
computational efficiency
🔎 Similar Papers
No similar papers found.
N
Nanae Aratake
Graduate School of Medicine, Kyoto University, Kyoto, Japan
T
Taisei Tosaki
Graduate School of Medicine, Kyoto University, Kyoto, Japan
Y
Yuji Okamoto
Graduate School of Medicine, Kyoto University, Kyoto, Japan
E
Eiichiro Uchino
Graduate School of Medicine, Kyoto University, Kyoto, Japan
M
Masaki Nakamura
Medical Data Vision Co., Ltd., Tokyo, Japan
N
Nobutomo Matsui
IQVIA Solutions Japan G.K., Tokyo, Japan
A
Akiko Hatakama
DeSC Healthcare, Inc., Tokyo, Japan
Y
Yasushi Okuno
Graduate School of Medicine, Kyoto University, Kyoto, Japan