Institution profile

Tetras AI

Industry researchnorthamerica · us
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD

Jun 03, 2026

Existing CAD research typically addresses individual tasks in isolation, lacking a unified benchmark for multimodal multitask learning. This work introduces the first comprehensive multimodal benchmark encompassing point cloud reconstruction, text- and image-to-CAD generation, and CAD-based question answering. Furthermore, we propose UniCAD-MLLM, an end-to-end general-purpose multimodal large language model that, for the first time, integrates textual, visual, sketch, and point cloud inputs within a single unified framework to enable collaborative modeling and understanding across diverse tasks. Evaluated on both the newly introduced UniCAD benchmark and the established Fusion360 dataset, UniCAD-MLLM consistently outperforms existing specialized and multitask approaches, achieving state-of-the-art performance across all evaluated tasks.

0 citationsRead paper

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model

Jul 15, 2025

Multimodal large language models (MLLMs) remain limited in fine-grained image understanding—particularly for deformable object keypoint localization. To address this, we propose the first general-purpose keypoint understanding framework, introducing a novel “identify-then-detect” paradigm and structured chain-of-thought reasoning to achieve unified, cross-scene and cross-category keypoint localization. Our method jointly leverages instruction-driven semantic parsing and pixel-level keypoint regression, trained on a large-scale, multi-category dataset comprising over 500K samples. Evaluated on multiple benchmarks, it achieves state-of-the-art performance, significantly improving localization accuracy and generalization under complex occlusions and diverse object appearances. Moreover, it enhances semantic controllability in human–AI collaborative interaction by enabling precise, instruction-guided keypoint interpretation.

0 citationsRead paper

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

May 29, 2025

To address the challenge of unifying multimodal understanding and generation within a single framework, this paper introduces OpenUni—the first fully open-source, minimalist unified architecture. Its core design modularly couples off-the-shelf multimodal large language models (MLLMs) and diffusion models via learnable queries and a lightweight Transformer connector, activating only 1.1B or 3.1B parameters. Crucially, OpenUni avoids end-to-end training, drastically reducing computational overhead while enabling strong cross-modal synergy. Experiments demonstrate state-of-the-art performance on instruction-aligned image generation tasks and top-tier results across multiple benchmarks—including GenEval, DPG-Bench, and WISE. To foster reproducibility and community advancement, the project releases all model weights, training code, and a high-quality dataset comprising 23 million image-text pairs.

0 citationsRead paper

Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Mar 27, 2025

Visual understanding and generation tasks suffer from misaligned representation granularities, hindering joint optimization within unified multimodal frameworks; existing approaches prioritize low-level visual features at the expense of semantic comprehension. To address this, we propose Harmon—a novel framework featuring a shared Masked Autoregressive (MAR) encoder, the first to simultaneously achieve strong semantic representation and high-fidelity generation capabilities. Harmon introduces a three-stage progressive co-training paradigm that intrinsically unifies understanding and generation. Evaluated on multiple benchmarks—including GenEval, MJHQ30K, and WISE—Harmon achieves state-of-the-art performance in image generation while matching the visual understanding accuracy of dedicated semantic encoders (e.g., Janus). Crucially, it attains optimal trade-offs between both tasks using a single, unified encoder.

0 citationsRead paper
Recent publications

Latest Papers

UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD

Jun 03, 2026

Existing CAD research typically addresses individual tasks in isolation, lacking a unified benchmark for multimodal multitask learning. This work introduces the first comprehensive multimodal benchmark encompassing point cloud reconstruction, text- and image-to-CAD generation, and CAD-based question answering. Furthermore, we propose UniCAD-MLLM, an end-to-end general-purpose multimodal large language model that, for the first time, integrates textual, visual, sketch, and point cloud inputs within a single unified framework to enable collaborative modeling and understanding across diverse tasks. Evaluated on both the newly introduced UniCAD benchmark and the established Fusion360 dataset, UniCAD-MLLM consistently outperforms existing specialized and multitask approaches, achieving state-of-the-art performance across all evaluated tasks.

0 citationsRead paper

KptLLM++: Towards Generic Keypoint Comprehension with Large Language Model

Jul 15, 2025

Multimodal large language models (MLLMs) remain limited in fine-grained image understanding—particularly for deformable object keypoint localization. To address this, we propose the first general-purpose keypoint understanding framework, introducing a novel “identify-then-detect” paradigm and structured chain-of-thought reasoning to achieve unified, cross-scene and cross-category keypoint localization. Our method jointly leverages instruction-driven semantic parsing and pixel-level keypoint regression, trained on a large-scale, multi-category dataset comprising over 500K samples. Evaluated on multiple benchmarks, it achieves state-of-the-art performance, significantly improving localization accuracy and generalization under complex occlusions and diverse object appearances. Moreover, it enhances semantic controllability in human–AI collaborative interaction by enabling precise, instruction-guided keypoint interpretation.

0 citationsRead paper

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

May 29, 2025

To address the challenge of unifying multimodal understanding and generation within a single framework, this paper introduces OpenUni—the first fully open-source, minimalist unified architecture. Its core design modularly couples off-the-shelf multimodal large language models (MLLMs) and diffusion models via learnable queries and a lightweight Transformer connector, activating only 1.1B or 3.1B parameters. Crucially, OpenUni avoids end-to-end training, drastically reducing computational overhead while enabling strong cross-modal synergy. Experiments demonstrate state-of-the-art performance on instruction-aligned image generation tasks and top-tier results across multiple benchmarks—including GenEval, DPG-Bench, and WISE. To foster reproducibility and community advancement, the project releases all model weights, training code, and a high-quality dataset comprising 23 million image-text pairs.

0 citationsRead paper

Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Mar 27, 2025

Visual understanding and generation tasks suffer from misaligned representation granularities, hindering joint optimization within unified multimodal frameworks; existing approaches prioritize low-level visual features at the expense of semantic comprehension. To address this, we propose Harmon—a novel framework featuring a shared Masked Autoregressive (MAR) encoder, the first to simultaneously achieve strong semantic representation and high-fidelity generation capabilities. Harmon introduces a three-stage progressive co-training paradigm that intrinsically unifies understanding and generation. Evaluated on multiple benchmarks—including GenEval, MJHQ30K, and WISE—Harmon achieves state-of-the-art performance in image generation while matching the visual understanding accuracy of dedicated semantic encoders (e.g., Janus). Crucially, it attains optimal trade-offs between both tasks using a single, unified encoder.

0 citationsRead paper