Flower Hub: A Reproducible Benchmarking Platform for Federated Learning in Simulation and Deployment
本文针对联邦学习中难以复现、比较和扩展的问题,提出Flower Hub平台,通过将基准测试打包为可执行的应用程序,支持在模拟和实际部署中统一评估。
本文针对联邦学习中难以复现、比较和扩展的问题,提出Flower Hub平台,通过将基准测试打包为可执行的应用程序,支持在模拟和实际部署中统一评估。
This study investigates whether AI scientist agents can achieve effective in-context learning through real experimental feedback, with a focus on their sensitivity to such feedback in scientific experimental design. Leveraging 800 independent Cell Painting high-content screening experiments, we compare the performance of large language models (Claude Sonnet 4.5/4.6) with and without access to real feedback, employing randomized label controls to verify that observed improvements depend on the structure of the feedback. Results show that real feedback increases the average number of discoveries per feature by 53.4% (p = 0.003). Furthermore, model upgrades reduce gene hallucination rates to 3–9% and significantly enhance in-context learning efficacy (+11.0 hits, p = 0.003), highlighting the critical role of model capability thresholds in enabling effective learning from feedback.
Existing generative foundation models lack support for histopathology and struggle with tasks such as virtual staining. To address this gap, we propose CytoSyn—the first latent diffusion foundation model tailored for histopathology—trained via large-scale self-supervision on over 10,000 whole-slide images from The Cancer Genome Atlas (TCGA), augmented with sampling optimization and anti-hallucination slide-level regularization techniques. The refined CytoSyn-v2 further advances performance, achieving state-of-the-art results in both photorealism and diversity of generated images. Notably, even when trained exclusively on tumor data, it can generate high-quality hematoxylin-and-eosin (H&E)-stained images across disease domains, including inflammatory bowel disease. Both the model and dataset are publicly released to facilitate a broad range of computational pathology applications.
Automated and interpretable microscopic inflammation assessment in hematoxylin-eosin (H&E) whole-slide images (WSIs) remains a critical unmet need in inflammatory bowel disease (IBD) pathology diagnosis. Method: We propose an end-to-end interpretable prediction framework integrating multi-instance learning (MIL) with a dual-module explainability architecture: HistoPLUS performs cell-level immune/epithelial classification, EpiSeg executes epithelial tissue semantic segmentation, and tile-level feature attribution guides region localization. Contribution/Results: This is the first framework to achieve biologically consistent, cross-dataset robust, and pathologically interpretable prediction in IBD. On the primary cohort, it achieves an AUC of 0.83; external validation yields AUCs of 0.99 and 0.84 on two independent cohorts. Attribution analysis consistently identifies high-scoring regions as immune-cell–rich and low-scoring regions as predominantly normal epithelium—demonstrating strong biological plausibility and generalizability.
Current large language models (LLMs) exhibit limited performance on biomedical reasoning tasks—including target druggability assessment, therapeutic modality matching, and drug perturbation effect prediction—hindering translational medicine progress. To address this, we propose a verifiability-guided reinforcement learning framework that performs post-training on open-source small-scale LLMs using a self-constructed dataset of over 300,000 verifiable biomedical question-answer pairs, yielding the OwkinZero model. Our method significantly enhances cross-task generalization, achieving— for the first time—a small-model superiority over larger commercial LLMs on standardized biomedical reasoning benchmarks. A hybrid training variant further attains state-of-the-art performance across all evaluated tasks. This work establishes a new paradigm for AI-driven biological discovery: efficient, verifiable, and fully reproducible.
本文针对联邦学习中难以复现、比较和扩展的问题,提出Flower Hub平台,通过将基准测试打包为可执行的应用程序,支持在模拟和实际部署中统一评估。
This study investigates whether AI scientist agents can achieve effective in-context learning through real experimental feedback, with a focus on their sensitivity to such feedback in scientific experimental design. Leveraging 800 independent Cell Painting high-content screening experiments, we compare the performance of large language models (Claude Sonnet 4.5/4.6) with and without access to real feedback, employing randomized label controls to verify that observed improvements depend on the structure of the feedback. Results show that real feedback increases the average number of discoveries per feature by 53.4% (p = 0.003). Furthermore, model upgrades reduce gene hallucination rates to 3–9% and significantly enhance in-context learning efficacy (+11.0 hits, p = 0.003), highlighting the critical role of model capability thresholds in enabling effective learning from feedback.
Existing generative foundation models lack support for histopathology and struggle with tasks such as virtual staining. To address this gap, we propose CytoSyn—the first latent diffusion foundation model tailored for histopathology—trained via large-scale self-supervision on over 10,000 whole-slide images from The Cancer Genome Atlas (TCGA), augmented with sampling optimization and anti-hallucination slide-level regularization techniques. The refined CytoSyn-v2 further advances performance, achieving state-of-the-art results in both photorealism and diversity of generated images. Notably, even when trained exclusively on tumor data, it can generate high-quality hematoxylin-and-eosin (H&E)-stained images across disease domains, including inflammatory bowel disease. Both the model and dataset are publicly released to facilitate a broad range of computational pathology applications.
Automated and interpretable microscopic inflammation assessment in hematoxylin-eosin (H&E) whole-slide images (WSIs) remains a critical unmet need in inflammatory bowel disease (IBD) pathology diagnosis. Method: We propose an end-to-end interpretable prediction framework integrating multi-instance learning (MIL) with a dual-module explainability architecture: HistoPLUS performs cell-level immune/epithelial classification, EpiSeg executes epithelial tissue semantic segmentation, and tile-level feature attribution guides region localization. Contribution/Results: This is the first framework to achieve biologically consistent, cross-dataset robust, and pathologically interpretable prediction in IBD. On the primary cohort, it achieves an AUC of 0.83; external validation yields AUCs of 0.99 and 0.84 on two independent cohorts. Attribution analysis consistently identifies high-scoring regions as immune-cell–rich and low-scoring regions as predominantly normal epithelium—demonstrating strong biological plausibility and generalizability.
Current large language models (LLMs) exhibit limited performance on biomedical reasoning tasks—including target druggability assessment, therapeutic modality matching, and drug perturbation effect prediction—hindering translational medicine progress. To address this, we propose a verifiability-guided reinforcement learning framework that performs post-training on open-source small-scale LLMs using a self-constructed dataset of over 300,000 verifiable biomedical question-answer pairs, yielding the OwkinZero model. Our method significantly enhances cross-task generalization, achieving— for the first time—a small-model superiority over larger commercial LLMs on standardized biomedical reasoning benchmarks. A hybrid training variant further attains state-of-the-art performance across all evaluated tasks. This work establishes a new paradigm for AI-driven biological discovery: efficient, verifiable, and fully reproducible.