Institution profile

Beike Inc.

Industry researchasia · cn
Official website
Research library9linked papers
Opportunities0open roles
Selected work

Representative Papers

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training

Jun 29, 2026

This work addresses the challenge that large language models struggle to stably self-improve without external supervision, as existing methods often lack dynamic awareness of problem difficulty and consequently over-optimize on easy samples while neglecting hard or boundary cases. To overcome this, the paper proposes DRIFT, a novel framework that introduces, for the first time, a synergistic mechanism combining problem-level difficulty routing with token-level pacing gating. This mechanism dynamically allocates self-distillation and reinforcement learning signals and guides exploration at critical reasoning positions. DRIFT further incorporates a success experience buffer and a two-stage curriculum learning strategy to enable fine-grained control over the self-improvement process. Evaluated across five benchmarks and three model scales, DRIFT significantly outperforms GRPO and SDPO, achieving an average score of 79.5% and a ToolUse task accuracy of 79.2%, establishing new state-of-the-art results.

0 citationsRead paper

Do We Need Distinct Representations for Every Speech Token? Unveiling and Exploiting Redundancy in Large Speech Language Models

Apr 08, 2026

This work addresses the excessive computational overhead in large spoken language models caused by high sampling rates, which result in unnecessarily long input sequences far exceeding semantic requirements. Through inter-layer oracle intervention analysis, the study reveals that while shallow layers retain fine-grained acoustic details, deeper layers exhibit highly structured redundancy. Leveraging this insight, the authors propose Affinity Pooling—a training-free method that dynamically merges tokens based on similarity, both at the input and within deep layers. This approach preserves semantic accuracy while significantly improving efficiency: it reduces prefill FLOPs by 27.48%, decreases memory consumption by approximately 1.7×, and accelerates first-token latency by about 1.1×.

0 citationsRead paper

FloorplanVLM: A Vision-Language Model for Floorplan Vectorization

Feb 06, 2026

This work proposes a novel end-to-end method for converting raster floorplans into engineering-grade vector graphics that satisfy complex topological and geometric constraints. By reframing vectorization as an image-conditioned sequence modeling task, the approach directly outputs a JSON sequence encoding the global structural layout, enabling precise holistic representation of non-Manhattan elements such as slanted walls and arcs. Adopting a “pixel-to-sequence” paradigm, the method circumvents the fragmentation and heuristic fragility of conventional pipelines. Trained on two newly curated datasets—Floorplan-2M and Floorplan-HQ-300K—using a vision-language model architecture with supervised fine-tuning and Group Relative Policy Optimization (GRPO), the model achieves a 92.52% exterior wall IoU on the FPBench-2K benchmark, demonstrating substantial improvements in structural validity and generalization capability.

0 citationsRead paper

DreamHome-Pano: Design-Aware and Conflict-Free Panoramic Interior Generation

Feb 06, 2026

This work addresses the challenge in multi-condition indoor panoramic image generation where stylistic preferences often conflict with architectural constraints, leading to geometric distortions in spatial layouts. To resolve this, the authors propose a controllable generation framework that bridges layout and style conditions through semantic prompts and employs a Prompt-LLM module to achieve cross-modal alignment. By integrating structure-aware geometric priors with a multi-condition disentanglement mechanism, the framework establishes a conflict-free control architecture that effectively isolates stylistic influences from spatial layout during generation. A multi-stage training strategy combining supervised fine-tuning and reinforcement learning enables the model to simultaneously preserve high aesthetic quality and significantly enhance structural consistency, offering a reliable solution for professional-grade panoramic indoor visualization.

0 citationsRead paper
Recent publications

Latest Papers

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training

Jun 29, 2026

This work addresses the challenge that large language models struggle to stably self-improve without external supervision, as existing methods often lack dynamic awareness of problem difficulty and consequently over-optimize on easy samples while neglecting hard or boundary cases. To overcome this, the paper proposes DRIFT, a novel framework that introduces, for the first time, a synergistic mechanism combining problem-level difficulty routing with token-level pacing gating. This mechanism dynamically allocates self-distillation and reinforcement learning signals and guides exploration at critical reasoning positions. DRIFT further incorporates a success experience buffer and a two-stage curriculum learning strategy to enable fine-grained control over the self-improvement process. Evaluated across five benchmarks and three model scales, DRIFT significantly outperforms GRPO and SDPO, achieving an average score of 79.5% and a ToolUse task accuracy of 79.2%, establishing new state-of-the-art results.

0 citationsRead paper

Do We Need Distinct Representations for Every Speech Token? Unveiling and Exploiting Redundancy in Large Speech Language Models

Apr 08, 2026

This work addresses the excessive computational overhead in large spoken language models caused by high sampling rates, which result in unnecessarily long input sequences far exceeding semantic requirements. Through inter-layer oracle intervention analysis, the study reveals that while shallow layers retain fine-grained acoustic details, deeper layers exhibit highly structured redundancy. Leveraging this insight, the authors propose Affinity Pooling—a training-free method that dynamically merges tokens based on similarity, both at the input and within deep layers. This approach preserves semantic accuracy while significantly improving efficiency: it reduces prefill FLOPs by 27.48%, decreases memory consumption by approximately 1.7×, and accelerates first-token latency by about 1.1×.

0 citationsRead paper

FloorplanVLM: A Vision-Language Model for Floorplan Vectorization

Feb 06, 2026

This work proposes a novel end-to-end method for converting raster floorplans into engineering-grade vector graphics that satisfy complex topological and geometric constraints. By reframing vectorization as an image-conditioned sequence modeling task, the approach directly outputs a JSON sequence encoding the global structural layout, enabling precise holistic representation of non-Manhattan elements such as slanted walls and arcs. Adopting a “pixel-to-sequence” paradigm, the method circumvents the fragmentation and heuristic fragility of conventional pipelines. Trained on two newly curated datasets—Floorplan-2M and Floorplan-HQ-300K—using a vision-language model architecture with supervised fine-tuning and Group Relative Policy Optimization (GRPO), the model achieves a 92.52% exterior wall IoU on the FPBench-2K benchmark, demonstrating substantial improvements in structural validity and generalization capability.

0 citationsRead paper

DreamHome-Pano: Design-Aware and Conflict-Free Panoramic Interior Generation

Feb 06, 2026

This work addresses the challenge in multi-condition indoor panoramic image generation where stylistic preferences often conflict with architectural constraints, leading to geometric distortions in spatial layouts. To resolve this, the authors propose a controllable generation framework that bridges layout and style conditions through semantic prompts and employs a Prompt-LLM module to achieve cross-modal alignment. By integrating structure-aware geometric priors with a multi-condition disentanglement mechanism, the framework establishes a conflict-free control architecture that effectively isolates stylistic influences from spatial layout during generation. A multi-stage training strategy combining supervised fine-tuning and reinforcement learning enables the model to simultaneously preserve high aesthetic quality and significantly enhance structural consistency, offering a reliable solution for professional-grade panoramic indoor visualization.

0 citationsRead paper