Institution profile

Hupan Lab

Research institutionasia · cn
Research library9linked papers
Opportunities0open roles
Selected work

Representative Papers

HounsWorld: A Multimodal World Model for Hidden Patient-State Readout, Reconstruction, and Simulation

Aug 13, 2026

This work addresses the need for clinical intelligence to infer latent patient states from incomplete observations, rather than merely learning isolated mappings from images to answers. To this end, the authors propose the first CT-centric multimodal world model that constructs a unified implicit representation of patient state, jointly handling three tasks: readout, reconstruction, and simulation. Built upon a 3-billion-parameter shared Transformer architecture, the model incorporates zero-initialized CT adapters and Hounsfield window sampling to enable bidirectional generation between language and volumetric data. Its effectiveness is validated on HounsBench, a newly introduced benchmark. Experimental results demonstrate superior performance across all three tasks, significantly advancing CT-based clinical understanding.

0 citationsRead paper

Potential Applications of HBF in LLM Serving Systems

Aug 13, 2026

This work addresses the memory capacity bottleneck in large language model (LLM) serving caused by the limited high-bandwidth memory (HBM), which struggles to accommodate growing model weights, KV caches, and multiple model variants. For the first time, it systematically explores high-bandwidth flash (HBF) as a capacity-extension layer for HBM, preserving high-bandwidth compute pathways while integrating HBF into the GPU memory hierarchy. The authors design model state residency policies tailored to the read-heavy, write-light access patterns of LLM inference, supported by system-level modeling. Experimental results demonstrate that this approach significantly increases the number of expert replicas in mixture-of-experts (MoE) architectures, reduces loading overhead in multi-model serving, and enables replication of popular models, thereby substantially improving overall serving efficiency.

0 citationsRead paper
Recent publications

Latest Papers

HounsWorld: A Multimodal World Model for Hidden Patient-State Readout, Reconstruction, and Simulation

Aug 13, 2026

This work addresses the need for clinical intelligence to infer latent patient states from incomplete observations, rather than merely learning isolated mappings from images to answers. To this end, the authors propose the first CT-centric multimodal world model that constructs a unified implicit representation of patient state, jointly handling three tasks: readout, reconstruction, and simulation. Built upon a 3-billion-parameter shared Transformer architecture, the model incorporates zero-initialized CT adapters and Hounsfield window sampling to enable bidirectional generation between language and volumetric data. Its effectiveness is validated on HounsBench, a newly introduced benchmark. Experimental results demonstrate superior performance across all three tasks, significantly advancing CT-based clinical understanding.

0 citationsRead paper

Potential Applications of HBF in LLM Serving Systems

Aug 13, 2026

This work addresses the memory capacity bottleneck in large language model (LLM) serving caused by the limited high-bandwidth memory (HBM), which struggles to accommodate growing model weights, KV caches, and multiple model variants. For the first time, it systematically explores high-bandwidth flash (HBF) as a capacity-extension layer for HBM, preserving high-bandwidth compute pathways while integrating HBF into the GPU memory hierarchy. The authors design model state residency policies tailored to the read-heavy, write-light access patterns of LLM inference, supported by system-level modeling. Experimental results demonstrate that this approach significantly increases the number of expert replicas in mixture-of-experts (MoE) architectures, reduces loading overhead in multi-model serving, and enables replication of popular models, thereby substantially improving overall serving efficiency.

0 citationsRead paper