Institution profile

Lovart AI

Industry researchnorthamerica · us
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers

May 17, 2026

Existing video style transfer methods are often hindered by the scarcity of large-scale triplet data and effective modeling paradigms, leading to temporal inconsistency, fragile handling of occlusions, and flickering artifacts. To address these limitations, this work introduces VISTA-1000, the first large-scale synthetic dataset with aligned style, content, and motion, encompassing 1,000 distinct artistic styles. Building upon this dataset, we propose a context-aware transfer framework based on diffusion Transformers, augmented with a lightweight style adapter for robust style representation. By integrating joint modeling with a disentanglement strategy, our approach significantly outperforms existing methods in terms of style fidelity, temporal coherence, and content preservation, effectively suppressing flickering and drift artifacts.

0 citationsRead paper

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration

May 17, 2026

This work addresses the challenges of identity drift, background inconsistency, and semantic degradation in long-form stylized or actor-replaced cinematic content, which arise from frequent shot transitions and viewpoint changes. To tackle these issues, the authors propose a multi-agent collaborative framework that leverages a scene-level JSON script as a semantic backbone, integrating dynamic visual reference anchors with a grid-based batch keyframe generation mechanism. Building upon shared contextual modeling in latent space, the framework enables joint keyframe synthesis and incorporates closed-loop verification with selective regeneration for rigorous identity and alignment auditing. Its core innovation lies in the novel Dual-Bridge Consistency mechanism, which effectively enforces long-term language–vision coherence across hundreds of shots. Evaluated on the SoapBench benchmark, the method significantly outperforms existing commercial video generation APIs, demonstrating superior narrative fidelity and long-range consistency.

0 citationsRead paper

Unlocking the Latent Canvas: Eliciting and Benchmarking Symbolic Visual Expression in LLMs

Mar 15, 2026

This work addresses the underexplored potential of large language models (LLMs) in symbolic visual representation by proposing SVE-ASCII, a framework that systematically investigates and enhances LLMs’ intrinsic ability to generate and comprehend visual content purely within textual space. Departing from conventional approaches that rely on external rendering or code execution, SVE-ASCII leverages a “Seed-and-Evolve” data synthesis strategy, context-aware style editing, and unified instruction tuning to jointly optimize text-to-ASCII art generation and ASCII-to-text understanding. The study introduces ASCIIArt-7K, a high-quality dataset, and ASCIIArt-Bench, a comprehensive benchmark. Experimental results demonstrate that training on generation substantially improves visual understanding performance, revealing a bidirectional enhancement between perception and generation in symbolic visual processing. All code, data, and models are publicly released.

0 citationsRead paper

SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokens

Feb 07, 2026

This work addresses the limitation of existing unified diffusion models, which typically support only a single conditioning input and struggle to flexibly integrate heterogeneous visual conditions. To overcome this, we propose a post-training framework that introduces multi-attribute tokens—representing style, content, subject, and identity—into a diffusion Transformer, along with a novel selective interleaving mechanism. This enables compositional editing, selective attribute transfer, and fine-grained multimodal alignment. Built upon the Bagel unified backbone, our approach leverages 700k interleaved text–image sequences for post-training to construct efficient multi-attribute embedding and conditional fusion modules. Experiments demonstrate that our method significantly outperforms Bagel in compositional generation tasks, achieving notable improvements in controllability, cross-condition consistency, and visual quality.

0 citationsRead paper

OmniPSD: Layered PSD Generation with Diffusion Transformer

Dec 09, 2025

This work addresses the generation and decomposition of hierarchical PSD files with transparent alpha channels. Methodologically: (i) we design a spatial-attention-driven multi-layer compositing mechanism that jointly models semantic structure, spatial relationships, and transparency; (ii) we propose an iterative context-erasure decomposition strategy to enable editable layer parsing from a single input image; and (iii) we introduce an RGBA-VAE encoder to ensure lossless alpha-channel reconstruction. Evaluated on a newly constructed RGBA hierarchical dataset, our approach significantly outperforms existing image layering methods in generation fidelity, inter-layer structural consistency, and alpha-channel accuracy. To the best of our knowledge, this is the first end-to-end framework capable of generating fully editable, transparency-aware PSD files directly from text or image inputs—leveraging a diffusion Transformer architecture based on the Flux design.

0 citationsRead paper
Recent publications

Latest Papers

VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers

May 17, 2026

Existing video style transfer methods are often hindered by the scarcity of large-scale triplet data and effective modeling paradigms, leading to temporal inconsistency, fragile handling of occlusions, and flickering artifacts. To address these limitations, this work introduces VISTA-1000, the first large-scale synthetic dataset with aligned style, content, and motion, encompassing 1,000 distinct artistic styles. Building upon this dataset, we propose a context-aware transfer framework based on diffusion Transformers, augmented with a lightweight style adapter for robust style representation. By integrating joint modeling with a disentanglement strategy, our approach significantly outperforms existing methods in terms of style fidelity, temporal coherence, and content preservation, effectively suppressing flickering and drift artifacts.

0 citationsRead paper

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration

May 17, 2026

This work addresses the challenges of identity drift, background inconsistency, and semantic degradation in long-form stylized or actor-replaced cinematic content, which arise from frequent shot transitions and viewpoint changes. To tackle these issues, the authors propose a multi-agent collaborative framework that leverages a scene-level JSON script as a semantic backbone, integrating dynamic visual reference anchors with a grid-based batch keyframe generation mechanism. Building upon shared contextual modeling in latent space, the framework enables joint keyframe synthesis and incorporates closed-loop verification with selective regeneration for rigorous identity and alignment auditing. Its core innovation lies in the novel Dual-Bridge Consistency mechanism, which effectively enforces long-term language–vision coherence across hundreds of shots. Evaluated on the SoapBench benchmark, the method significantly outperforms existing commercial video generation APIs, demonstrating superior narrative fidelity and long-range consistency.

0 citationsRead paper

Unlocking the Latent Canvas: Eliciting and Benchmarking Symbolic Visual Expression in LLMs

Mar 15, 2026

This work addresses the underexplored potential of large language models (LLMs) in symbolic visual representation by proposing SVE-ASCII, a framework that systematically investigates and enhances LLMs’ intrinsic ability to generate and comprehend visual content purely within textual space. Departing from conventional approaches that rely on external rendering or code execution, SVE-ASCII leverages a “Seed-and-Evolve” data synthesis strategy, context-aware style editing, and unified instruction tuning to jointly optimize text-to-ASCII art generation and ASCII-to-text understanding. The study introduces ASCIIArt-7K, a high-quality dataset, and ASCIIArt-Bench, a comprehensive benchmark. Experimental results demonstrate that training on generation substantially improves visual understanding performance, revealing a bidirectional enhancement between perception and generation in symbolic visual processing. All code, data, and models are publicly released.

0 citationsRead paper

SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokens

Feb 07, 2026

This work addresses the limitation of existing unified diffusion models, which typically support only a single conditioning input and struggle to flexibly integrate heterogeneous visual conditions. To overcome this, we propose a post-training framework that introduces multi-attribute tokens—representing style, content, subject, and identity—into a diffusion Transformer, along with a novel selective interleaving mechanism. This enables compositional editing, selective attribute transfer, and fine-grained multimodal alignment. Built upon the Bagel unified backbone, our approach leverages 700k interleaved text–image sequences for post-training to construct efficient multi-attribute embedding and conditional fusion modules. Experiments demonstrate that our method significantly outperforms Bagel in compositional generation tasks, achieving notable improvements in controllability, cross-condition consistency, and visual quality.

0 citationsRead paper

OmniPSD: Layered PSD Generation with Diffusion Transformer

Dec 09, 2025

This work addresses the generation and decomposition of hierarchical PSD files with transparent alpha channels. Methodologically: (i) we design a spatial-attention-driven multi-layer compositing mechanism that jointly models semantic structure, spatial relationships, and transparency; (ii) we propose an iterative context-erasure decomposition strategy to enable editable layer parsing from a single input image; and (iii) we introduce an RGBA-VAE encoder to ensure lossless alpha-channel reconstruction. Evaluated on a newly constructed RGBA hierarchical dataset, our approach significantly outperforms existing image layering methods in generation fidelity, inter-layer structural consistency, and alpha-channel accuracy. To the best of our knowledge, this is the first end-to-end framework capable of generating fully editable, transparency-aware PSD files directly from text or image inputs—leveraging a diffusion Transformer architecture based on the Flux design.

0 citationsRead paper