Thinking with Cameras: Active Visual Reasoning via Dynamic Viewpoint Control for Surveillance Video Understanding
为解决监控视频理解中目标远、小、遮挡或移出视野的问题,本文通过动态视角控制和强化学习优化视角策略的方法,提高视觉证据获取能力。
为解决监控视频理解中目标远、小、遮挡或移出视野的问题,本文通过动态视角控制和强化学习优化视角策略的方法,提高视觉证据获取能力。
Existing CAD research typically addresses individual tasks in isolation, lacking a unified benchmark for multimodal multitask learning. This work introduces the first comprehensive multimodal benchmark encompassing point cloud reconstruction, text- and image-to-CAD generation, and CAD-based question answering. Furthermore, we propose UniCAD-MLLM, an end-to-end general-purpose multimodal large language model that, for the first time, integrates textual, visual, sketch, and point cloud inputs within a single unified framework to enable collaborative modeling and understanding across diverse tasks. Evaluated on both the newly introduced UniCAD benchmark and the established Fusion360 dataset, UniCAD-MLLM consistently outperforms existing specialized and multitask approaches, achieving state-of-the-art performance across all evaluated tasks.
Multimodal large language models (MLLMs) remain limited in fine-grained image understanding—particularly for deformable object keypoint localization. To address this, we propose the first general-purpose keypoint understanding framework, introducing a novel “identify-then-detect” paradigm and structured chain-of-thought reasoning to achieve unified, cross-scene and cross-category keypoint localization. Our method jointly leverages instruction-driven semantic parsing and pixel-level keypoint regression, trained on a large-scale, multi-category dataset comprising over 500K samples. Evaluated on multiple benchmarks, it achieves state-of-the-art performance, significantly improving localization accuracy and generalization under complex occlusions and diverse object appearances. Moreover, it enhances semantic controllability in human–AI collaborative interaction by enabling precise, instruction-guided keypoint interpretation.
To address the challenge of unifying multimodal understanding and generation within a single framework, this paper introduces OpenUni—the first fully open-source, minimalist unified architecture. Its core design modularly couples off-the-shelf multimodal large language models (MLLMs) and diffusion models via learnable queries and a lightweight Transformer connector, activating only 1.1B or 3.1B parameters. Crucially, OpenUni avoids end-to-end training, drastically reducing computational overhead while enabling strong cross-modal synergy. Experiments demonstrate state-of-the-art performance on instruction-aligned image generation tasks and top-tier results across multiple benchmarks—including GenEval, DPG-Bench, and WISE. To foster reproducibility and community advancement, the project releases all model weights, training code, and a high-quality dataset comprising 23 million image-text pairs.
Visual understanding and generation tasks suffer from misaligned representation granularities, hindering joint optimization within unified multimodal frameworks; existing approaches prioritize low-level visual features at the expense of semantic comprehension. To address this, we propose Harmon—a novel framework featuring a shared Masked Autoregressive (MAR) encoder, the first to simultaneously achieve strong semantic representation and high-fidelity generation capabilities. Harmon introduces a three-stage progressive co-training paradigm that intrinsically unifies understanding and generation. Evaluated on multiple benchmarks—including GenEval, MJHQ30K, and WISE—Harmon achieves state-of-the-art performance in image generation while matching the visual understanding accuracy of dedicated semantic encoders (e.g., Janus). Crucially, it attains optimal trade-offs between both tasks using a single, unified encoder.
为解决监控视频理解中目标远、小、遮挡或移出视野的问题,本文通过动态视角控制和强化学习优化视角策略的方法,提高视觉证据获取能力。
Existing CAD research typically addresses individual tasks in isolation, lacking a unified benchmark for multimodal multitask learning. This work introduces the first comprehensive multimodal benchmark encompassing point cloud reconstruction, text- and image-to-CAD generation, and CAD-based question answering. Furthermore, we propose UniCAD-MLLM, an end-to-end general-purpose multimodal large language model that, for the first time, integrates textual, visual, sketch, and point cloud inputs within a single unified framework to enable collaborative modeling and understanding across diverse tasks. Evaluated on both the newly introduced UniCAD benchmark and the established Fusion360 dataset, UniCAD-MLLM consistently outperforms existing specialized and multitask approaches, achieving state-of-the-art performance across all evaluated tasks.
Multimodal large language models (MLLMs) remain limited in fine-grained image understanding—particularly for deformable object keypoint localization. To address this, we propose the first general-purpose keypoint understanding framework, introducing a novel “identify-then-detect” paradigm and structured chain-of-thought reasoning to achieve unified, cross-scene and cross-category keypoint localization. Our method jointly leverages instruction-driven semantic parsing and pixel-level keypoint regression, trained on a large-scale, multi-category dataset comprising over 500K samples. Evaluated on multiple benchmarks, it achieves state-of-the-art performance, significantly improving localization accuracy and generalization under complex occlusions and diverse object appearances. Moreover, it enhances semantic controllability in human–AI collaborative interaction by enabling precise, instruction-guided keypoint interpretation.
To address the challenge of unifying multimodal understanding and generation within a single framework, this paper introduces OpenUni—the first fully open-source, minimalist unified architecture. Its core design modularly couples off-the-shelf multimodal large language models (MLLMs) and diffusion models via learnable queries and a lightweight Transformer connector, activating only 1.1B or 3.1B parameters. Crucially, OpenUni avoids end-to-end training, drastically reducing computational overhead while enabling strong cross-modal synergy. Experiments demonstrate state-of-the-art performance on instruction-aligned image generation tasks and top-tier results across multiple benchmarks—including GenEval, DPG-Bench, and WISE. To foster reproducibility and community advancement, the project releases all model weights, training code, and a high-quality dataset comprising 23 million image-text pairs.
Visual understanding and generation tasks suffer from misaligned representation granularities, hindering joint optimization within unified multimodal frameworks; existing approaches prioritize low-level visual features at the expense of semantic comprehension. To address this, we propose Harmon—a novel framework featuring a shared Masked Autoregressive (MAR) encoder, the first to simultaneously achieve strong semantic representation and high-fidelity generation capabilities. Harmon introduces a three-stage progressive co-training paradigm that intrinsically unifies understanding and generation. Evaluated on multiple benchmarks—including GenEval, MJHQ30K, and WISE—Harmon achieves state-of-the-art performance in image generation while matching the visual understanding accuracy of dedicated semantic encoders (e.g., Janus). Crucially, it attains optimal trade-offs between both tasks using a single, unified encoder.