Institution profile

Yuanshi Technology

Industry researchasia · cn
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection

Sep 27, 2025

To address the dual challenges of poor generalization to unseen categories and low robustness under adverse conditions (e.g., low illumination, occlusion) in open-world object detection, this paper proposes a curriculum-based cross-modal contrastive learning framework—first integrating RGB-thermal (RGBT) multimodal perception with vision-language alignment. To mitigate catastrophic forgetting in two-stage training, exponential moving average (EMA) is adopted, providing theoretical guarantees for preserving prior knowledge. Jointly leveraging RGBT pretraining and cross-modal contrastive learning, our method simultaneously enhances category openness and environmental robustness. Extensive experiments on FLIR, OV-COCO, and OV-LVIS benchmarks yield 80.1 AP⁵⁰, 48.6 AP⁵⁰ₙₒᵥₑₗ, and 35.7 mAPᵣ, respectively—outperforming state-of-the-art methods by significant margins.

0 citationsRead paper

IGD: Instructional Graphic Design with Multimodal Layer Generation

Jul 14, 2025

Existing automated graphic design methods face two key bottlenecks: traditional two-stage pipelines lack intelligence and creativity, while diffusion-based approaches generate only non-editable pixel-level images with blurry text rendering and limited practicality. This paper proposes the first natural language–driven, editable multimodal layer generation framework, integrating multimodal large language models (MLLMs) and diffusion models in an end-to-end jointly trained architecture. We introduce a novel paradigm of parameterized rendering coupled with image asset co-generation: the MLLM parses user instructions to predict layer attributes and layout structure, while the diffusion model synthesizes high-fidelity visual content. Experiments across diverse design scenarios demonstrate significant improvements over state-of-the-art methods, enabling efficient generation of high-fidelity, semantically aligned, fully editable vector and layered design files—effectively bridging creative flexibility and engineering practicality.

0 citationsRead paper

Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation

Mar 25, 2025

To address the challenges of fine-grained text-video alignment and cross-segment visual inconsistency in multi-scene long-video generation, this paper proposes Dual-Mask Diffusion Transformer (Dual-Mask DiT). Methodologically, it introduces a novel synergistic mechanism of symmetric binary masking and segment-level conditional masking to achieve precise segment-wise text-to-video alignment within the DiT architecture; it further integrates text-visual cross-attention with an autoregressive scene expansion strategy to ensure temporal coherence. The core contribution is the first joint modeling of segment-level semantic alignment and long-term temporal consistency within a diffusion Transformer framework. Experiments on multi-scene video generation demonstrate that our method significantly improves both cross-segment visual consistency and semantic alignment accuracy. Both qualitative and quantitative evaluations outperform existing state-of-the-art approaches.

0 citationsRead paper
Recent publications

Latest Papers

C3-OWD: A Curriculum Cross-modal Contrastive Learning Framework for Open-World Detection

Sep 27, 2025

To address the dual challenges of poor generalization to unseen categories and low robustness under adverse conditions (e.g., low illumination, occlusion) in open-world object detection, this paper proposes a curriculum-based cross-modal contrastive learning framework—first integrating RGB-thermal (RGBT) multimodal perception with vision-language alignment. To mitigate catastrophic forgetting in two-stage training, exponential moving average (EMA) is adopted, providing theoretical guarantees for preserving prior knowledge. Jointly leveraging RGBT pretraining and cross-modal contrastive learning, our method simultaneously enhances category openness and environmental robustness. Extensive experiments on FLIR, OV-COCO, and OV-LVIS benchmarks yield 80.1 AP⁵⁰, 48.6 AP⁵⁰ₙₒᵥₑₗ, and 35.7 mAPᵣ, respectively—outperforming state-of-the-art methods by significant margins.

0 citationsRead paper

IGD: Instructional Graphic Design with Multimodal Layer Generation

Jul 14, 2025

Existing automated graphic design methods face two key bottlenecks: traditional two-stage pipelines lack intelligence and creativity, while diffusion-based approaches generate only non-editable pixel-level images with blurry text rendering and limited practicality. This paper proposes the first natural language–driven, editable multimodal layer generation framework, integrating multimodal large language models (MLLMs) and diffusion models in an end-to-end jointly trained architecture. We introduce a novel paradigm of parameterized rendering coupled with image asset co-generation: the MLLM parses user instructions to predict layer attributes and layout structure, while the diffusion model synthesizes high-fidelity visual content. Experiments across diverse design scenarios demonstrate significant improvements over state-of-the-art methods, enabling efficient generation of high-fidelity, semantically aligned, fully editable vector and layered design files—effectively bridging creative flexibility and engineering practicality.

0 citationsRead paper

Mask$^2$DiT: Dual Mask-based Diffusion Transformer for Multi-Scene Long Video Generation

Mar 25, 2025

To address the challenges of fine-grained text-video alignment and cross-segment visual inconsistency in multi-scene long-video generation, this paper proposes Dual-Mask Diffusion Transformer (Dual-Mask DiT). Methodologically, it introduces a novel synergistic mechanism of symmetric binary masking and segment-level conditional masking to achieve precise segment-wise text-to-video alignment within the DiT architecture; it further integrates text-visual cross-attention with an autoregressive scene expansion strategy to ensure temporal coherence. The core contribution is the first joint modeling of segment-level semantic alignment and long-term temporal consistency within a diffusion Transformer framework. Experiments on multi-scene video generation demonstrate that our method significantly improves both cross-segment visual consistency and semantic alignment accuracy. Both qualitative and quantitative evaluations outperform existing state-of-the-art approaches.

0 citationsRead paper