Institution profile

Theta Labs

Industry researchnorthamerica · us
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

Vector-Bench: Can Models Surgically Edit SVG Code?

Jul 21, 2026

This study addresses the challenge of precisely repairing SVG code under visual instruction guidance, requiring models to modify only specified regions while preserving all other protected content. To this end, the authors introduce a benchmark comprising 40 high-difficulty tasks and propose a novel dual-norm reward mechanism that integrates semantic invariance with attribute-aware tolerance. They also introduce new evaluation dimensions, including validity-gated repair progress and Unintended Change Rate (UCR). Through rigorous assessment involving SVG parsing-rendering validation, structural-semantic consistency checks, and deterministic norm scoring across 34 models, they find that even the strongest model achieves a full-norm success rate of merely 15.0%, with an average repair progress of 43.7%, revealing significant limitations in current approaches to faithful, constrained editing.

0 citationsRead paper

FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow

Mar 19, 2026

Existing indoor scene generation methods struggle to simultaneously achieve high photorealism, fine-grained object-level control, and global style consistency. To address this challenge, this work proposes a tri-branch collaborative generation model that uniquely integrates multimodal graph conditioning with rectified flow mechanisms. Specifically, a multimodal graph neural network models inter-object relationships, while tightly coupled rectified flows across layout, shape, and texture branches enable dynamic interaction of object information and style alignment during generation. The proposed approach significantly outperforms current language- or graph-conditioned baselines in terms of photorealism, style coherence, and human preference, achieving synergistic optimization between object-level precision and scene-level stylistic unity.

0 citationsRead paper

DARA: Few-shot Budget Allocation in Online Advertising via In-Context Decision Making with RL-Finetuned LLMs

Jan 21, 2026

This work addresses the challenge of optimizing advertisers’ cumulative value under strict budget constraints in low-data regimes within online advertising. To this end, the authors propose DARA, a two-stage framework: the first stage leverages the in-context learning capability of large language models (LLMs) to generate an initial campaign plan, while the second stage refines this plan through feedback-driven reasoning for precise numerical optimization. The approach innovatively combines the few-shot generalization strength of LLMs with reinforcement learning fine-tuning, introducing a GRPO-Adaptive policy that dynamically optimizes the reference strategy. By decoupling the decision process into distinct reasoning and optimization phases, DARA achieves superior performance over existing baselines on both real-world and synthetic datasets, consistently enhancing advertisers’ cumulative value under stringent budget limitations.

0 citationsRead paper

Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion

Nov 23, 2025

Existing 3D urban generation methods rely on monolithic diffusion models, limiting both personalization and scalable expansion. To address this, we propose a top-down hierarchical planning framework—“City–District–Grid”—that integrates large language model (LLM)-driven reasoning to enable user-guided, customizable design and continuous urban evolution. Our key innovation is a relation-guided interactive expansion mechanism, incorporating scene-graph-aware distance constraints and semantic layout optimization to ensure spatial coherence. We further introduce a multi-dimensional evaluation benchmark covering semantic fidelity, geometric accuracy, texture quality, and layout合理性, with six quantitative metrics. Leveraging a “generate–optimize–evaluate” image synthesis loop and image-to-3D reconstruction, our method jointly synthesizes hierarchical structure and local details. Experiments demonstrate state-of-the-art performance across generation quality, scalability, and user controllability.

0 citationsRead paper
Recent publications

Latest Papers

Vector-Bench: Can Models Surgically Edit SVG Code?

Jul 21, 2026

This study addresses the challenge of precisely repairing SVG code under visual instruction guidance, requiring models to modify only specified regions while preserving all other protected content. To this end, the authors introduce a benchmark comprising 40 high-difficulty tasks and propose a novel dual-norm reward mechanism that integrates semantic invariance with attribute-aware tolerance. They also introduce new evaluation dimensions, including validity-gated repair progress and Unintended Change Rate (UCR). Through rigorous assessment involving SVG parsing-rendering validation, structural-semantic consistency checks, and deterministic norm scoring across 34 models, they find that even the strongest model achieves a full-norm success rate of merely 15.0%, with an average repair progress of 43.7%, revealing significant limitations in current approaches to faithful, constrained editing.

0 citationsRead paper

FlowScene: Style-Consistent Indoor Scene Generation with Multimodal Graph Rectified Flow

Mar 19, 2026

Existing indoor scene generation methods struggle to simultaneously achieve high photorealism, fine-grained object-level control, and global style consistency. To address this challenge, this work proposes a tri-branch collaborative generation model that uniquely integrates multimodal graph conditioning with rectified flow mechanisms. Specifically, a multimodal graph neural network models inter-object relationships, while tightly coupled rectified flows across layout, shape, and texture branches enable dynamic interaction of object information and style alignment during generation. The proposed approach significantly outperforms current language- or graph-conditioned baselines in terms of photorealism, style coherence, and human preference, achieving synergistic optimization between object-level precision and scene-level stylistic unity.

0 citationsRead paper

DARA: Few-shot Budget Allocation in Online Advertising via In-Context Decision Making with RL-Finetuned LLMs

Jan 21, 2026

This work addresses the challenge of optimizing advertisers’ cumulative value under strict budget constraints in low-data regimes within online advertising. To this end, the authors propose DARA, a two-stage framework: the first stage leverages the in-context learning capability of large language models (LLMs) to generate an initial campaign plan, while the second stage refines this plan through feedback-driven reasoning for precise numerical optimization. The approach innovatively combines the few-shot generalization strength of LLMs with reinforcement learning fine-tuning, introducing a GRPO-Adaptive policy that dynamically optimizes the reference strategy. By decoupling the decision process into distinct reasoning and optimization phases, DARA achieves superior performance over existing baselines on both real-world and synthetic datasets, consistently enhancing advertisers’ cumulative value under stringent budget limitations.

0 citationsRead paper

Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion

Nov 23, 2025

Existing 3D urban generation methods rely on monolithic diffusion models, limiting both personalization and scalable expansion. To address this, we propose a top-down hierarchical planning framework—“City–District–Grid”—that integrates large language model (LLM)-driven reasoning to enable user-guided, customizable design and continuous urban evolution. Our key innovation is a relation-guided interactive expansion mechanism, incorporating scene-graph-aware distance constraints and semantic layout optimization to ensure spatial coherence. We further introduce a multi-dimensional evaluation benchmark covering semantic fidelity, geometric accuracy, texture quality, and layout合理性, with six quantitative metrics. Leveraging a “generate–optimize–evaluate” image synthesis loop and image-to-3D reconstruction, our method jointly synthesizes hierarchical structure and local details. Experiments demonstrate state-of-the-art performance across generation quality, scalability, and user controllability.

0 citationsRead paper