Institution profile

National Institute of Healthcare Data Science

Academic institution
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement

Sep 01, 2025

Addressing the challenges of identity preservation, reliance on fine-tuning, and data scarcity in text-to-video generation, this paper proposes a training-free triple-enhancement framework. First, GPT-4o–driven face-aware prompt enhancement bridges the semantic gap between textual descriptions and visual content. Second, a prompt-aware reference image optimization mechanism improves input consistency. Third, a unified gradient-guided strategy jointly optimizes identity fidelity and spatiotemporal coherence during diffusion model sampling—enabling inference-time refinement without architectural modification. The method requires no model training or fine-tuning. Extensive evaluation on a thousand-video benchmark demonstrates significant improvements in character identity consistency and video quality, outperforming state-of-the-art approaches in both automated metrics and human assessment. It ranked first in the ACM Multimedia 2025 Challenge, validating its strong generalizability and practical applicability.

0 citationsRead paper

Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration

Aug 05, 2025

Open-vocabulary human-object interaction (HOI) detection requires generalization to unseen interaction categories, yet existing vision-language model (VLM)-based approaches suffer from a mismatch between image-level pretraining and fine-grained region-level interaction modeling, and textual descriptions often fail to capture discriminative visual appearance details. To address this, we propose an interaction-aware prompt generator that dynamically constructs scene-adaptive prompts to facilitate knowledge transfer among semantically similar interactions. We further introduce a language-model-guided concept calibration mechanism to enhance the discriminability of interaction representations. Our framework integrates region-level feature modeling, cross-modal similarity optimization, and hard negative sampling to improve fine-grained relational reasoning and zero-shot generalization. Extensive experiments demonstrate state-of-the-art performance on SWIG-HOI and HICO-DET, validating both effectiveness and strong generalization capability.

0 citationsRead paper
Recent publications

Latest Papers

Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement

Sep 01, 2025

Addressing the challenges of identity preservation, reliance on fine-tuning, and data scarcity in text-to-video generation, this paper proposes a training-free triple-enhancement framework. First, GPT-4o–driven face-aware prompt enhancement bridges the semantic gap between textual descriptions and visual content. Second, a prompt-aware reference image optimization mechanism improves input consistency. Third, a unified gradient-guided strategy jointly optimizes identity fidelity and spatiotemporal coherence during diffusion model sampling—enabling inference-time refinement without architectural modification. The method requires no model training or fine-tuning. Extensive evaluation on a thousand-video benchmark demonstrates significant improvements in character identity consistency and video quality, outperforming state-of-the-art approaches in both automated metrics and human assessment. It ranked first in the ACM Multimedia 2025 Challenge, validating its strong generalizability and practical applicability.

0 citationsRead paper

Open-Vocabulary HOI Detection with Interaction-aware Prompt and Concept Calibration

Aug 05, 2025

Open-vocabulary human-object interaction (HOI) detection requires generalization to unseen interaction categories, yet existing vision-language model (VLM)-based approaches suffer from a mismatch between image-level pretraining and fine-grained region-level interaction modeling, and textual descriptions often fail to capture discriminative visual appearance details. To address this, we propose an interaction-aware prompt generator that dynamically constructs scene-adaptive prompts to facilitate knowledge transfer among semantically similar interactions. We further introduce a language-model-guided concept calibration mechanism to enhance the discriminability of interaction representations. Our framework integrates region-level feature modeling, cross-modal similarity optimization, and hard negative sampling to improve fine-grained relational reasoning and zero-shot generalization. Extensive experiments demonstrate state-of-the-art performance on SWIG-HOI and HICO-DET, validating both effectiveness and strong generalization capability.

0 citationsRead paper