Institution profile

Eyeline Studios

Industry researchnorthamerica · us
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

Safe-Control: A Safety Patch for Mitigating Unsafe Content in Text-to-Image Generation Models

Aug 28, 2025

Existing text-to-image (T2I) models are vulnerable to misuse for generating unsafe content, while mainstream safety mechanisms exhibit poor robustness under distribution shifts or adversarial attacks and typically require model fine-tuning. To address this, we propose a plug-and-play safety patching framework that operates without modifying the original model’s weights. Our method introduces learnable, data-driven safety-aware conditioning signals into intermediate layers of the frozen diffusion model during the denoising process. It supports multi-strategy fusion and cross-model transferability, significantly enhancing resilience against both distribution shifts and adversarial prompts. Evaluated on six state-of-the-art T2I models, our approach reduces unsafe image generation to 7%, outperforming seven existing SOTA safety methods, while preserving high-fidelity image quality and strong text–image alignment.

0 citationsRead paper

MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling

Aug 11, 2025

Existing long-sequence video generation methods suffer from weak auxiliary capabilities, low visual fidelity, and insufficient narrative expressiveness. This paper proposes a multi-agent collaborative end-to-end text-to-video narrative generation framework encompassing the full pipeline: scriptwriting, shot design, character modeling, keyframe generation, animation synthesis, and audio generation. We innovatively introduce the 3E principle—Exploration, Evaluation, and Enhancement—to ensure output completeness at each stage, and pioneer a multimodal co-generation mechanism enabling synchronized production of video, voiceover, and background music. The framework is highly extensible and compatible with mainstream generative models. Experiments demonstrate state-of-the-art performance across auxiliary capability, visual fidelity, and narrative expressiveness; it generates high-quality, long-duration, expressive narrative videos from only brief textual prompts.

0 citationsRead paper

FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution

Apr 09, 2025

Addressing the challenge of jointly achieving high accuracy, inter-frame consistency, and low latency in real-time depth estimation for high-resolution video, this paper proposes a lightweight temporal enhancement framework. Built upon a pre-trained single-image depth model, it incorporates explicit temporal consistency constraints and a resolution-adaptive decoding architecture, requiring only minimal video data for fine-tuning. To our knowledge, this is the first approach to enable real-time streaming inference at 2K resolution (2044×1148) and 24 FPS under lightweight fine-tuning, while maintaining both high depth accuracy and strong inter-frame consistency. Evaluated on multiple unseen datasets, the method improves boundary sharpness by 32%, achieves 2.1× faster inference speed than the state-of-the-art, and retains top-tier accuracy.

0 citationsRead paper

Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset

Mar 18, 2025

Video portrait relighting faces a fundamental trade-off between photorealism and temporal stability. Existing approaches rely heavily on high-quality paired multi-illumination video data, severely limiting their generalizability and practical applicability. To address this, we propose the first conditional video diffusion model specifically designed for portrait video relighting. Our method introduces a novel dynamic illumination embedding mechanism and adopts a hybrid training paradigm combining static One-Light-At-a-Time (OLAT) data with uncurated single-illumination in-the-wild videos—eliminating the need for paired multi-illumination sequences. Leveraging spatiotemporal consistency losses and lightweight conditional adaptation of a pre-trained diffusion backbone, our framework enables end-to-end relighting under arbitrary target lighting conditions. Extensive experiments demonstrate state-of-the-art performance in both photorealism and temporal coherence, significantly outperforming existing methods.

0 citationsRead paper

Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction

Feb 13, 2025

This paper addresses the challenge of jointly optimizing camera parameters, lens distortion, and 3D Gaussian representations in wide-field-of-view (FOV) image reconstruction. Methodologically: (1) it proposes a hybrid distortion modeling module integrating an invertible residual network with explicit grid-based fusion, enabling robust modeling for arbitrary ultra-wide-angle lenses; (2) it introduces a cube-map resampling strategy to preserve geometric fidelity and rendering quality under large FOV; and (3) it combines Gaussian Splatting rasterization acceleration with end-to-end optimization. Evaluated on both synthetic and real-world datasets, the method achieves state-of-the-art performance: it significantly reduces the number of required input images, outperforms conventional camera models in reconstruction accuracy, and produces high-resolution, distortion-free, dense scene representations without artifacts.

0 citationsRead paper
Recent publications

Latest Papers

Safe-Control: A Safety Patch for Mitigating Unsafe Content in Text-to-Image Generation Models

Aug 28, 2025

Existing text-to-image (T2I) models are vulnerable to misuse for generating unsafe content, while mainstream safety mechanisms exhibit poor robustness under distribution shifts or adversarial attacks and typically require model fine-tuning. To address this, we propose a plug-and-play safety patching framework that operates without modifying the original model’s weights. Our method introduces learnable, data-driven safety-aware conditioning signals into intermediate layers of the frozen diffusion model during the denoising process. It supports multi-strategy fusion and cross-model transferability, significantly enhancing resilience against both distribution shifts and adversarial prompts. Evaluated on six state-of-the-art T2I models, our approach reduces unsafe image generation to 7%, outperforming seven existing SOTA safety methods, while preserving high-fidelity image quality and strong text–image alignment.

0 citationsRead paper

MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling

Aug 11, 2025

Existing long-sequence video generation methods suffer from weak auxiliary capabilities, low visual fidelity, and insufficient narrative expressiveness. This paper proposes a multi-agent collaborative end-to-end text-to-video narrative generation framework encompassing the full pipeline: scriptwriting, shot design, character modeling, keyframe generation, animation synthesis, and audio generation. We innovatively introduce the 3E principle—Exploration, Evaluation, and Enhancement—to ensure output completeness at each stage, and pioneer a multimodal co-generation mechanism enabling synchronized production of video, voiceover, and background music. The framework is highly extensible and compatible with mainstream generative models. Experiments demonstrate state-of-the-art performance across auxiliary capability, visual fidelity, and narrative expressiveness; it generates high-quality, long-duration, expressive narrative videos from only brief textual prompts.

0 citationsRead paper

FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution

Apr 09, 2025

Addressing the challenge of jointly achieving high accuracy, inter-frame consistency, and low latency in real-time depth estimation for high-resolution video, this paper proposes a lightweight temporal enhancement framework. Built upon a pre-trained single-image depth model, it incorporates explicit temporal consistency constraints and a resolution-adaptive decoding architecture, requiring only minimal video data for fine-tuning. To our knowledge, this is the first approach to enable real-time streaming inference at 2K resolution (2044×1148) and 24 FPS under lightweight fine-tuning, while maintaining both high depth accuracy and strong inter-frame consistency. Evaluated on multiple unseen datasets, the method improves boundary sharpness by 32%, achieves 2.1× faster inference speed than the state-of-the-art, and retains top-tier accuracy.

0 citationsRead paper

Lux Post Facto: Learning Portrait Performance Relighting with Conditional Video Diffusion and a Hybrid Dataset

Mar 18, 2025

Video portrait relighting faces a fundamental trade-off between photorealism and temporal stability. Existing approaches rely heavily on high-quality paired multi-illumination video data, severely limiting their generalizability and practical applicability. To address this, we propose the first conditional video diffusion model specifically designed for portrait video relighting. Our method introduces a novel dynamic illumination embedding mechanism and adopts a hybrid training paradigm combining static One-Light-At-a-Time (OLAT) data with uncurated single-illumination in-the-wild videos—eliminating the need for paired multi-illumination sequences. Leveraging spatiotemporal consistency losses and lightweight conditional adaptation of a pre-trained diffusion backbone, our framework enables end-to-end relighting under arbitrary target lighting conditions. Extensive experiments demonstrate state-of-the-art performance in both photorealism and temporal coherence, significantly outperforming existing methods.

0 citationsRead paper

Self-Calibrating Gaussian Splatting for Large Field of View Reconstruction

Feb 13, 2025

This paper addresses the challenge of jointly optimizing camera parameters, lens distortion, and 3D Gaussian representations in wide-field-of-view (FOV) image reconstruction. Methodologically: (1) it proposes a hybrid distortion modeling module integrating an invertible residual network with explicit grid-based fusion, enabling robust modeling for arbitrary ultra-wide-angle lenses; (2) it introduces a cube-map resampling strategy to preserve geometric fidelity and rendering quality under large FOV; and (3) it combines Gaussian Splatting rasterization acceleration with end-to-end optimization. Evaluated on both synthetic and real-world datasets, the method achieves state-of-the-art performance: it significantly reduces the number of required input images, outperforms conventional camera models in reconstruction accuracy, and produces high-resolution, distortion-free, dense scene representations without artifacts.

0 citationsRead paper