Institution profile

Vipshop

Industry researchasia · cn
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

HierarchicalKV: A GPU Hash Table with Cache Semantics for Continuous Online Embedding Storage

Mar 17, 2026

This work addresses severe memory inefficiency in conventional GPU hash tables when embedding tables exceed the capacity of a single GPU’s high-bandwidth memory (HBM), as these structures retain all key-value pairs regardless of access patterns. To overcome this limitation, the authors propose HierarchicalKV—the first GPU hash table that treats caching semantics as a first-class operation. It replaces traditional dictionary semantics with a policy-driven eviction mechanism that either updates entries in place or rejects insertions, thereby avoiding costly rehashing and overflow failures. Key innovations include cache-line-aligned buckets, inline score-driven upserts, dynamic dual-bucket selection, three-level concurrency control, and a hierarchical key-value separation architecture. Evaluated on an NVIDIA H100 NVL, HierarchicalKV achieves up to 3.9 billion key-value operations per second, maintains load factors between 0.50 and 1.00 with less than 5% throughput variation, outperforms WarpCore by 1.4×, and surpasses indirect-addressing baselines by 2.6–9.4×, with integration already adopted in multiple open-source recommendation frameworks.

0 citationsRead paper

Intrinsic Concept Extraction Based on Compositional Interpretability

Mar 12, 2026

Existing unsupervised methods struggle to fully disentangle and reconstruct composable intrinsic concepts from a single image. To address this limitation, this work introduces the CI-ICE task—the first framework aimed at extracting hierarchical intrinsic concepts at both object and attribute levels. We propose HyperExpress, a novel approach that integrates diffusion generative models with hyperbolic space embeddings to achieve semantic structure-preserving and relation-aware disentangled learning in a concept-level embedding space. HyperExpress substantially outperforms current state-of-the-art methods, enabling highly interpretable and strongly compositional reconstructions from individual images.

0 citationsRead paper

Subject-Consistent and Pose-Diverse Text-to-Image Generation

Jul 11, 2025

Addressing the challenge of simultaneously preserving subject identity consistency and enabling pose/compositional diversity in text-to-image (T2I) generation, this paper introduces CoDi—a two-stage diffusion control framework. In the early denoising stage, a pose-aware optimal transport mechanism enables identity-preserving subject transfer; in the later stage, saliency-guided feature selection and explicit enhancement of critical identity features jointly disentangle and optimize consistency and diversity. CoDi requires no additional training and supports parameter-free inference-time optimization. Quantitative and qualitative evaluations demonstrate that CoDi outperforms state-of-the-art methods across key metrics—including subject identity consistency, pose diversity, and prompt fidelity—particularly in complex visual narrative tasks.

0 citationsRead paper
Recent publications

Latest Papers

HierarchicalKV: A GPU Hash Table with Cache Semantics for Continuous Online Embedding Storage

Mar 17, 2026

This work addresses severe memory inefficiency in conventional GPU hash tables when embedding tables exceed the capacity of a single GPU’s high-bandwidth memory (HBM), as these structures retain all key-value pairs regardless of access patterns. To overcome this limitation, the authors propose HierarchicalKV—the first GPU hash table that treats caching semantics as a first-class operation. It replaces traditional dictionary semantics with a policy-driven eviction mechanism that either updates entries in place or rejects insertions, thereby avoiding costly rehashing and overflow failures. Key innovations include cache-line-aligned buckets, inline score-driven upserts, dynamic dual-bucket selection, three-level concurrency control, and a hierarchical key-value separation architecture. Evaluated on an NVIDIA H100 NVL, HierarchicalKV achieves up to 3.9 billion key-value operations per second, maintains load factors between 0.50 and 1.00 with less than 5% throughput variation, outperforms WarpCore by 1.4×, and surpasses indirect-addressing baselines by 2.6–9.4×, with integration already adopted in multiple open-source recommendation frameworks.

0 citationsRead paper

Intrinsic Concept Extraction Based on Compositional Interpretability

Mar 12, 2026

Existing unsupervised methods struggle to fully disentangle and reconstruct composable intrinsic concepts from a single image. To address this limitation, this work introduces the CI-ICE task—the first framework aimed at extracting hierarchical intrinsic concepts at both object and attribute levels. We propose HyperExpress, a novel approach that integrates diffusion generative models with hyperbolic space embeddings to achieve semantic structure-preserving and relation-aware disentangled learning in a concept-level embedding space. HyperExpress substantially outperforms current state-of-the-art methods, enabling highly interpretable and strongly compositional reconstructions from individual images.

0 citationsRead paper

Subject-Consistent and Pose-Diverse Text-to-Image Generation

Jul 11, 2025

Addressing the challenge of simultaneously preserving subject identity consistency and enabling pose/compositional diversity in text-to-image (T2I) generation, this paper introduces CoDi—a two-stage diffusion control framework. In the early denoising stage, a pose-aware optimal transport mechanism enables identity-preserving subject transfer; in the later stage, saliency-guided feature selection and explicit enhancement of critical identity features jointly disentangle and optimize consistency and diversity. CoDi requires no additional training and supports parameter-free inference-time optimization. Quantitative and qualitative evaluations demonstrate that CoDi outperforms state-of-the-art methods across key metrics—including subject identity consistency, pose diversity, and prompt fidelity—particularly in complex visual narrative tasks.

0 citationsRead paper