Multiple Myeloma Lesion Segmentation on Whole-Body Diffusion-Weighted Imaging via Efficient Anatomical Anticipation and Multimodal Confirmation
本文针对多发性骨髓瘤病灶在全身扩散加权成像上自动分割的难题,提出了一种结合高效解剖预测和多模态确认的两阶段框架。
本文针对多发性骨髓瘤病灶在全身扩散加权成像上自动分割的难题,提出了一种结合高效解剖预测和多模态确认的两阶段框架。
研究解决了3D医学图像问答中的视觉冗余问题,通过SeVeR框架选择性地压缩和检索多模态信息,减少了不必要的视觉暴露,提高了性能。
为解决病理基础模型在全切片细胞表示能力上的评估问题,通过CellPath-Bench基准测试使用标准化多类线性探针评价不同模型的细胞类型解码能力和跨域泛化性能。
This work addresses the need for clinical intelligence to infer latent patient states from incomplete observations, rather than merely learning isolated mappings from images to answers. To this end, the authors propose the first CT-centric multimodal world model that constructs a unified implicit representation of patient state, jointly handling three tasks: readout, reconstruction, and simulation. Built upon a 3-billion-parameter shared Transformer architecture, the model incorporates zero-initialized CT adapters and Hounsfield window sampling to enable bidirectional generation between language and volumetric data. Its effectiveness is validated on HounsBench, a newly introduced benchmark. Experimental results demonstrate superior performance across all three tasks, significantly advancing CT-based clinical understanding.
This work addresses the memory capacity bottleneck in large language model (LLM) serving caused by the limited high-bandwidth memory (HBM), which struggles to accommodate growing model weights, KV caches, and multiple model variants. For the first time, it systematically explores high-bandwidth flash (HBF) as a capacity-extension layer for HBM, preserving high-bandwidth compute pathways while integrating HBF into the GPU memory hierarchy. The authors design model state residency policies tailored to the read-heavy, write-light access patterns of LLM inference, supported by system-level modeling. Experimental results demonstrate that this approach significantly increases the number of expert replicas in mixture-of-experts (MoE) architectures, reduces loading overhead in multi-model serving, and enables replication of popular models, thereby substantially improving overall serving efficiency.
本文针对多发性骨髓瘤病灶在全身扩散加权成像上自动分割的难题,提出了一种结合高效解剖预测和多模态确认的两阶段框架。
研究解决了3D医学图像问答中的视觉冗余问题,通过SeVeR框架选择性地压缩和检索多模态信息,减少了不必要的视觉暴露,提高了性能。
为解决病理基础模型在全切片细胞表示能力上的评估问题,通过CellPath-Bench基准测试使用标准化多类线性探针评价不同模型的细胞类型解码能力和跨域泛化性能。
This work addresses the need for clinical intelligence to infer latent patient states from incomplete observations, rather than merely learning isolated mappings from images to answers. To this end, the authors propose the first CT-centric multimodal world model that constructs a unified implicit representation of patient state, jointly handling three tasks: readout, reconstruction, and simulation. Built upon a 3-billion-parameter shared Transformer architecture, the model incorporates zero-initialized CT adapters and Hounsfield window sampling to enable bidirectional generation between language and volumetric data. Its effectiveness is validated on HounsBench, a newly introduced benchmark. Experimental results demonstrate superior performance across all three tasks, significantly advancing CT-based clinical understanding.
This work addresses the memory capacity bottleneck in large language model (LLM) serving caused by the limited high-bandwidth memory (HBM), which struggles to accommodate growing model weights, KV caches, and multiple model variants. For the first time, it systematically explores high-bandwidth flash (HBF) as a capacity-extension layer for HBM, preserving high-bandwidth compute pathways while integrating HBF into the GPU memory hierarchy. The authors design model state residency policies tailored to the read-heavy, write-light access patterns of LLM inference, supported by system-level modeling. Experimental results demonstrate that this approach significantly increases the number of expert replicas in mixture-of-experts (MoE) architectures, reduces loading overhead in multi-model serving, and enables replication of popular models, thereby substantially improving overall serving efficiency.