Institution profile

Hello Group Inc.

Industry researchnorthamerica · us
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation

Aug 11, 2026

This work addresses the unreliability of teacher supervision signals in on-policy distillation by introducing, for the first time, a prompt-level teacher consistency reliability metric \( R \), and empirically validates its positive correlation with distillation performance. To efficiently estimate \( R \) without extensive teacher inference, the authors propose ROUGE-5 F1 as a proxy metric, enabling prompts to be ranked in descending order of reliability and integrated into a reliability-aware prompt scheduling mechanism. The approach combines independently sampled student trajectories with teacher trajectories filtered by a verifier, achieving consistent and significant improvements over existing baselines across mathematical and code generation tasks on Qwen3 and Gemma4 models, and demonstrating robust gains in all six configurations of FiRe-OPD and ExOPD.

0 citationsRead paper

Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems

Mar 30, 2026

This work addresses the degradation in generalization and catastrophic forgetting observed in current mainstream multimodal models under content moderation and adversarial scenarios, primarily due to insufficient fine-grained visual perception and weak modeling of long-tailed noise. To mitigate these limitations, the authors propose a data-training co-optimization paradigm that integrates a compact architecture—comprising InternViT-300M, an MLP head, and Qwen3-1.7B—with a three-stage progressive training pipeline (pre-training, mid-training, and post-training). This approach effectively balances general-purpose capability retention and domain-specific adaptability within a constrained parameter budget. The resulting model achieves an average score of 67.90 across seven multimodal benchmarks on OpenCompass, an average recall of 94.38% on seven content moderation tasks, and a weighted recall of 82.82% on adversarial OCR-based violation detection, outperforming Gemini-2.5-Pro.

0 citationsRead paper

Amber-Image: Efficient Compression of Large-Scale Diffusion Transformers

Feb 18, 2026

This work addresses the high computational cost and deployment challenges of large-scale diffusion Transformer models in text-to-image generation. The authors propose an efficient compression framework that requires no retraining from scratch, leveraging temporally aware depth pruning and a hybrid single-stream architecture. By integrating localized weight averaging, inter-layer knowledge distillation, and progressive fine-tuning, the method substantially compresses a 60-layer dual-stream MMDiT model. With less than 2,000 GPU hours of training cost, the approach achieves a 70% reduction in parameters while preserving high-fidelity image synthesis and strong text rendering capabilities, as demonstrated on both DPG-Bench and LongText-Bench benchmarks.

0 citationsRead paper
Recent publications

Latest Papers

ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation

Aug 11, 2026

This work addresses the unreliability of teacher supervision signals in on-policy distillation by introducing, for the first time, a prompt-level teacher consistency reliability metric \( R \), and empirically validates its positive correlation with distillation performance. To efficiently estimate \( R \) without extensive teacher inference, the authors propose ROUGE-5 F1 as a proxy metric, enabling prompts to be ranked in descending order of reliability and integrated into a reliability-aware prompt scheduling mechanism. The approach combines independently sampled student trajectories with teacher trajectories filtered by a verifier, achieving consistent and significant improvements over existing baselines across mathematical and code generation tasks on Qwen3 and Gemma4 models, and demonstrating robust gains in all six configurations of FiRe-OPD and ExOPD.

0 citationsRead paper

Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems

Mar 30, 2026

This work addresses the degradation in generalization and catastrophic forgetting observed in current mainstream multimodal models under content moderation and adversarial scenarios, primarily due to insufficient fine-grained visual perception and weak modeling of long-tailed noise. To mitigate these limitations, the authors propose a data-training co-optimization paradigm that integrates a compact architecture—comprising InternViT-300M, an MLP head, and Qwen3-1.7B—with a three-stage progressive training pipeline (pre-training, mid-training, and post-training). This approach effectively balances general-purpose capability retention and domain-specific adaptability within a constrained parameter budget. The resulting model achieves an average score of 67.90 across seven multimodal benchmarks on OpenCompass, an average recall of 94.38% on seven content moderation tasks, and a weighted recall of 82.82% on adversarial OCR-based violation detection, outperforming Gemini-2.5-Pro.

0 citationsRead paper

Amber-Image: Efficient Compression of Large-Scale Diffusion Transformers

Feb 18, 2026

This work addresses the high computational cost and deployment challenges of large-scale diffusion Transformer models in text-to-image generation. The authors propose an efficient compression framework that requires no retraining from scratch, leveraging temporally aware depth pruning and a hybrid single-stream architecture. By integrating localized weight averaging, inter-layer knowledge distillation, and progressive fine-tuning, the method substantially compresses a 60-layer dual-stream MMDiT model. With less than 2,000 GPU hours of training cost, the approach achieves a 70% reduction in parameters while preserving high-fidelity image synthesis and strong text rendering capabilities, as demonstrated on both DPG-Bench and LongText-Bench benchmarks.

0 citationsRead paper