PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us

📅 2026-09-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过建立PhenoBench基准,利用深度表型队列解决健康相关问题的多模态数据分析不直接可比性的问题,评估了多种模型在90个临床任务上的表现。
📝 Abstract
Deeply phenotyped cohorts combine clinical, imaging, molecular, and wearable observations across timescales from seconds to years. This breadth can reveal which measurements inform which health-related questions, but heterogeneous analyses are not directly comparable. We present PhenoBench, an executable benchmark built around the Human Phenotype Project, in which more than 13,000 participants have completed the initial visit. Each question fixes the target, eligible population, timing, and allowed information; its evaluation contract specifies the split, metric, baseline, and claim boundary. The benchmark defines 90 clinically grounded tasks across 15 domains and 26 input modalities. Measurements showed question- and representation-dependent predictive value, including positive, near-zero, and negative changes in held-out performance relative to matched baselines. We used PhenoBench to evaluate emerging tabular foundation models across 160 matched regression comparisons spanning 52 tasks. These models ranked above standard task-specific models in aggregate but, averaged across the three pretrained models within each cell, improved on ridge by a median of only 0.004 $R^2$ (95% CI, 0.002--0.006). We then used the same cohort data and evaluation contracts to evaluate 14 language models, collectively covering 40 tasks spanning phenotype recovery, classification, follow-up forecasting, and participant ordering. Without cohort-specific fitting, language models made informative predictions on some tasks, but showed task-specific capability gaps, shared failures of scale, and rarely surpassed models fitted on the same fields. PhenoBench turns a multimodal longitudinal cohort into a versioned, auditable evaluation system where new questions, measurements, and models can be added without redefining existing comparisons.
Problem

Research questions and friction points this paper is trying to address.

deeply phenotyped cohorts
multimodal data
health-related questions
benchmark system
model evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

PhenoBench
executable benchmark
deeply phenotyped cohort
multimodal longitudinal data
standardized evaluation
🔎 Similar Papers
No similar papers found.
G
Gal Sapir
Pheno.AI, Tel Aviv, Israel; Department of Computer Science and Applied Mathematics, Weizmann Institute of Science, Rehovot, Israel
Alon Diament
Alon Diament
Pheno.AI, Tel Aviv, Israel; Department of Computer Science and Applied Mathematics, Weizmann Institute of Science, Rehovot, Israel
A
Adva Wolf
Pheno.AI, Tel Aviv, Israel; Department of Computer Science and Applied Mathematics, Weizmann Institute of Science, Rehovot, Israel
D
Doron Yaya-Stupp
Department of Computer Science and Applied Mathematics, Weizmann Institute of Science, Rehovot, Israel
D
Dikla Gelbard Solodkin
Pheno.AI, Tel Aviv, Israel; Department of Computer Science and Applied Mathematics, Weizmann Institute of Science, Rehovot, Israel
D
Dana Azouri
Pheno.AI, Tel Aviv, Israel; Department of Computer Science and Applied Mathematics, Weizmann Institute of Science, Rehovot, Israel
A
Anat Etzion-Fuchs
Pheno.AI, Tel Aviv, Israel; Department of Computer Science and Applied Mathematics, Weizmann Institute of Science, Rehovot, Israel
G
Guy Lutsker
Department of Computer Science and Applied Mathematics, Weizmann Institute of Science, Rehovot, Israel; Department of Molecular Cell Biology, Weizmann Institute of Science, Rehovot, Israel
Eran Segal
Eran Segal
Professor of Computer Science, Weizmann Institute of Science
Computational biology
Hagai Rossman
Hagai Rossman
Mohamed bin Zayed University of Artificial Intelligence (MBZUAI)
Biomedical AIEpidemiologyMachine Learning for healthcare