Institution profile

Runway

Industry researchnorthamerica · us
Official website
Research library12linked papers
Opportunities13open roles
Selected work

Representative Papers

Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs Supplementary

Mar 11, 2026

Current vision-language benchmarks conflate object orientation with spatial concepts such as position and scene context, hindering accurate evaluation of multimodal large language models’ orientation reasoning capabilities. This work proposes DORI, a cognitively inspired hierarchical benchmark that, for the first time, decomposes orientation into four dimensions grounded in human cognitive development. By leveraging bounding-box isolation, standardized spatial reference frames, and structured multiple-choice questions, DORI constructs a large-scale evaluation set of 33,656 questions spanning both coarse-grained (categorical) and fine-grained (metric) levels. Experiments across 24 state-of-the-art models reveal that current systems perform near chance on object-centric orientation tasks (best accuracy: 54.2%/45.0%), with pronounced failures in compound rotations and reference frame transformations, exposing their fundamental reliance on category-based heuristics rather than genuine geometric reasoning.

0 citationsRead paper

VideoSketcher: Video Models Prior Enable Versatile Sequential Sketch Generation

Feb 17, 2026

This work addresses the limitation of existing generative models that treat sketches as static images, thereby ignoring their intrinsic sequential drawing process. The authors propose a two-stage fine-tuning approach that leverages a large language model for semantic-driven planning of drawing steps and employs a pre-trained text-to-video diffusion model to generate temporally coherent, high-fidelity sketch sequences. By decoupling stroke-order learning from appearance generation, the method achieves richly detailed, semantically ordered dynamic sketch synthesis using only seven hand-drawn examples. It further supports controllable brush-style rendering and autoregressive interactive drawing, significantly enhancing both generation quality and user controllability under extremely low-data conditions.

0 citationsRead paper

Inspiration Seeds: Learning Non-Literal Visual Combinations for Generative Exploration

Feb 09, 2026

This work addresses the limitations of current generative models, which rely heavily on precise textual prompts and thus struggle to support the open-ended and ambiguous visual exploration typical of early-stage creative ideation. The authors propose a prompt-free, feedforward generative framework that synthesizes semantically meaningful and visually coherent image combinations from only two input images. By eliminating dependence on language, the method constructs training data exclusively from visual triplets and leverages a CLIP-based sparse autoencoder to extract disentangled editing directions from the CLIP latent space, enabling non-literal recombination of visual concepts. This approach empowers designers to conduct intuitive and efficient visual exploration during the initial phases of the creative process, fostering inspiration without the constraints of explicit textual guidance.

0 citationsRead paper

ShapeUP: Scalable Image-Conditioned 3D Editing

Feb 05, 2026

Existing 3D editing methods struggle to simultaneously achieve visual controllability, geometric consistency, and scalability, often suffering from slow inference, visual drift, or reliance on fixed priors. This work proposes an image-conditioned 3D editing framework that formulates editing as a supervised latent-to-latent transformation within a native 3D representation, leveraging pretrained 3D foundation models for fine-grained control. Notably, it enables the first mask-free, implicitly localized editing approach, preserving structural consistency while overcoming the scalability limitations of prior training-free methods. By supervising a 3D diffusion Transformer (DiT) on triplets of source 3D shapes, edited 2D images, and target 3D shapes, the method outperforms both trainable and training-free baselines in identity preservation and editing fidelity, demonstrating efficient, robust, and scalable 3D content editing capabilities.

0 citationsRead paper
Recent publications

Latest Papers

Seeing Isn't Orienting: A Cognitively Grounded Benchmark Reveals Systematic Orientation Failures in MLLMs Supplementary

Mar 11, 2026

Current vision-language benchmarks conflate object orientation with spatial concepts such as position and scene context, hindering accurate evaluation of multimodal large language models’ orientation reasoning capabilities. This work proposes DORI, a cognitively inspired hierarchical benchmark that, for the first time, decomposes orientation into four dimensions grounded in human cognitive development. By leveraging bounding-box isolation, standardized spatial reference frames, and structured multiple-choice questions, DORI constructs a large-scale evaluation set of 33,656 questions spanning both coarse-grained (categorical) and fine-grained (metric) levels. Experiments across 24 state-of-the-art models reveal that current systems perform near chance on object-centric orientation tasks (best accuracy: 54.2%/45.0%), with pronounced failures in compound rotations and reference frame transformations, exposing their fundamental reliance on category-based heuristics rather than genuine geometric reasoning.

0 citationsRead paper

VideoSketcher: Video Models Prior Enable Versatile Sequential Sketch Generation

Feb 17, 2026

This work addresses the limitation of existing generative models that treat sketches as static images, thereby ignoring their intrinsic sequential drawing process. The authors propose a two-stage fine-tuning approach that leverages a large language model for semantic-driven planning of drawing steps and employs a pre-trained text-to-video diffusion model to generate temporally coherent, high-fidelity sketch sequences. By decoupling stroke-order learning from appearance generation, the method achieves richly detailed, semantically ordered dynamic sketch synthesis using only seven hand-drawn examples. It further supports controllable brush-style rendering and autoregressive interactive drawing, significantly enhancing both generation quality and user controllability under extremely low-data conditions.

0 citationsRead paper

Inspiration Seeds: Learning Non-Literal Visual Combinations for Generative Exploration

Feb 09, 2026

This work addresses the limitations of current generative models, which rely heavily on precise textual prompts and thus struggle to support the open-ended and ambiguous visual exploration typical of early-stage creative ideation. The authors propose a prompt-free, feedforward generative framework that synthesizes semantically meaningful and visually coherent image combinations from only two input images. By eliminating dependence on language, the method constructs training data exclusively from visual triplets and leverages a CLIP-based sparse autoencoder to extract disentangled editing directions from the CLIP latent space, enabling non-literal recombination of visual concepts. This approach empowers designers to conduct intuitive and efficient visual exploration during the initial phases of the creative process, fostering inspiration without the constraints of explicit textual guidance.

0 citationsRead paper

ShapeUP: Scalable Image-Conditioned 3D Editing

Feb 05, 2026

Existing 3D editing methods struggle to simultaneously achieve visual controllability, geometric consistency, and scalability, often suffering from slow inference, visual drift, or reliance on fixed priors. This work proposes an image-conditioned 3D editing framework that formulates editing as a supervised latent-to-latent transformation within a native 3D representation, leveraging pretrained 3D foundation models for fine-grained control. Notably, it enables the first mask-free, implicitly localized editing approach, preserving structural consistency while overcoming the scalability limitations of prior training-free methods. By supervising a 3D diffusion Transformer (DiT) on triplets of source 3D shapes, edited 2D images, and target 3D shapes, the method outperforms both trainable and training-free baselines in identity preservation and editing fidelity, demonstrating efficient, robust, and scalable 3D content editing capabilities.

0 citationsRead paper