APICURON: a reactive infrastructure for credit attribution across distributed research data ecosystems
为解决生物数据管理中贡献未被充分认可的问题,APICURON平台通过实时记录和转换策展事件为可验证的工作单元,提供了一种信用归属机制。
为解决生物数据管理中贡献未被充分认可的问题,APICURON平台通过实时记录和转换策展事件为可验证的工作单元,提供了一种信用归属机制。
本文提出Traceable Trust框架,以评估和设计AI输出到生物科学研究行动的转化过程,确保其可信度。
This work addresses the issue of representation collapse in existing self-supervised learning methods, which rely heavily on within-batch interactions and are particularly vulnerable under small-batch training—especially when applied to high-dimensional scientific data constrained by memory limitations and class imbalance. To overcome this, the authors propose IConE, a novel framework that decouples collapse prevention from batch-level statistics by introducing globally learnable auxiliary instance embeddings and explicit diversity regularization. This enables stable training without dependence on batch statistics, supporting batch sizes as small as one. IConE consistently outperforms state-of-the-art contrastive and non-contrastive methods across diverse 2D and 3D biomedical datasets, maintaining high intrinsic dimensionality and robustness even under extremely small batch settings (B=1–64), thereby effectively mitigating representation collapse.
In zero-shot transfer for microscopy image segmentation, selecting appropriate pre-trained models is challenging due to the absence of target labels and source training data. Method: We propose the first fully unsupervised, source-data-free, and label-free transferability estimation method, grounded in generalization theory and the perturbation consistency assumption. Our framework quantifies output stability under input perturbations, integrating feature-space consistency and prediction confidence calibration. Contribution/Results: This is the first approach enabling truly zero-shot, source-independent transferability assessment for both semantic and instance segmentation. Evaluated on multimodal microscopy datasets, it achieves high correlation (Spearman ρ > 0.85) between estimated model rankings and actual target-domain performance—significantly outperforming existing unsupervised baselines.
为解决生物数据管理中贡献未被充分认可的问题,APICURON平台通过实时记录和转换策展事件为可验证的工作单元,提供了一种信用归属机制。
本文提出Traceable Trust框架,以评估和设计AI输出到生物科学研究行动的转化过程,确保其可信度。
This work addresses the issue of representation collapse in existing self-supervised learning methods, which rely heavily on within-batch interactions and are particularly vulnerable under small-batch training—especially when applied to high-dimensional scientific data constrained by memory limitations and class imbalance. To overcome this, the authors propose IConE, a novel framework that decouples collapse prevention from batch-level statistics by introducing globally learnable auxiliary instance embeddings and explicit diversity regularization. This enables stable training without dependence on batch statistics, supporting batch sizes as small as one. IConE consistently outperforms state-of-the-art contrastive and non-contrastive methods across diverse 2D and 3D biomedical datasets, maintaining high intrinsic dimensionality and robustness even under extremely small batch settings (B=1–64), thereby effectively mitigating representation collapse.
In zero-shot transfer for microscopy image segmentation, selecting appropriate pre-trained models is challenging due to the absence of target labels and source training data. Method: We propose the first fully unsupervised, source-data-free, and label-free transferability estimation method, grounded in generalization theory and the perturbation consistency assumption. Our framework quantifies output stability under input perturbations, integrating feature-space consistency and prediction confidence calibration. Contribution/Results: This is the first approach enabling truly zero-shot, source-independent transferability assessment for both semantic and instance segmentation. Evaluated on multimodal microscopy datasets, it achieves high correlation (Spearman ρ > 0.85) between estimated model rankings and actual target-domain performance—significantly outperforming existing unsupervised baselines.