Institution profile

Fabu Inc.

Industry researchasia · jp
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output Space

Sep 26, 2025

To address the high memory overhead of Mixture-of-Experts (MoE) models due to their large parameter count and the significant accuracy degradation caused by existing expert-level pruning methods, this paper proposes a fine-grained atomic expert pruning framework. Its core innovation lies in the first-ever mapping of Hessian information from the expert parameter space to the atomic expert output space—reducing computational complexity to *O(d²)*—and leveraging the Optimal Brain Surgeon principle to estimate atomic expert importance via forward/backward passes on a small calibration dataset. Evaluated on DeepSeek-MoE and Qwen-MoE, the method achieves 20–25% model compression with negligible accuracy loss while reducing inference FLOPs by approximately 20%, substantially improving MoE model deployment efficiency.

0 citationsRead paper

Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking

Aug 15, 2025

To address the insufficient robustness of 3D multi-object tracking (MOT) in autonomous driving under crowded scenes and imperfect detection outputs, this paper proposes DSC-Track: an end-to-end online tracking framework leveraging dynamic scene cue consistency. The core innovations include a spatiotemporal Point-Pair Feature (PPF) encoder and a cue-consistency Transformer module, which explicitly model inter-frame geometric relationships to suppress interference from irrelevant objects and enable stable feature matching and trajectory association. By integrating point-pair feature encoding, Transformer-based feature alignment, and dynamic feature updating, DSC-Track significantly improves association reliability under high object density and low-quality detections. Evaluated on nuScenes and Waymo Open Dataset, it achieves state-of-the-art performance: AMOTA scores of 73.2% (validation) and 70.3% (test) on nuScenes—substantially outperforming existing methods.

0 citationsRead paper

SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization

Apr 20, 2025

Current supervised fine-tuning (SFT) of text-to-image diffusion models optimizes only pixel-level MSE loss, failing to jointly ensure global perceptual quality and structural consistency. To address this, we propose Self-supervised Direct Preference Optimization (SUDO), a novel paradigm that (i) introduces the first annotation-free, self-supervised mechanism for generating preference image pairs—jointly modeling local details and global semantics—and (ii) seamlessly integrates Direct Preference Optimization (DPO) into the diffusion training pipeline, jointly optimizing DPO loss, pixel-wise MSE, and text-conditioned sampling. Evaluated on Stable Diffusion 1.5 and XL, SUDO achieves significantly lower FID scores, alongside consistent improvements in CLIP Score and human evaluation metrics. Our method simultaneously enhances structural coherence and fine-grained realism, without requiring human-annotated preferences.

0 citationsRead paper

SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement

Feb 10, 2025

To address the low quality of coarse-grain segmentation masks and high annotation costs, this paper proposes SAMRefiner++, a general-purpose and efficient mask refinement framework that is plug-and-play compatible with any pre-trained segmentation model (e.g., SAM) without requiring additional annotations. Methodologically: (i) it introduces a novel noise-robust prompting mechanism integrating distance-guided points, context-aware elastic bounding boxes, and Gaussian-style masks; (ii) it designs a “Split–Tackle–Merge” (STM) pipeline to handle multi-object segmentation challenges; and (iii) it incorporates an unsupervised IoU-adaptive optimization module combining iterative calibration with adaptive prompt engineering. Extensive experiments on multiple benchmarks demonstrate significant improvements over state-of-the-art methods in both accuracy and inference efficiency, achieving new SOTA performance. The source code is publicly available.

0 citationsRead paper
Recent publications

Latest Papers

HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output Space

Sep 26, 2025

To address the high memory overhead of Mixture-of-Experts (MoE) models due to their large parameter count and the significant accuracy degradation caused by existing expert-level pruning methods, this paper proposes a fine-grained atomic expert pruning framework. Its core innovation lies in the first-ever mapping of Hessian information from the expert parameter space to the atomic expert output space—reducing computational complexity to *O(d²)*—and leveraging the Optimal Brain Surgeon principle to estimate atomic expert importance via forward/backward passes on a small calibration dataset. Evaluated on DeepSeek-MoE and Qwen-MoE, the method achieves 20–25% model compression with negligible accuracy loss while reducing inference FLOPs by approximately 20%, substantially improving MoE model deployment efficiency.

0 citationsRead paper

Delving into Dynamic Scene Cue-Consistency for Robust 3D Multi-Object Tracking

Aug 15, 2025

To address the insufficient robustness of 3D multi-object tracking (MOT) in autonomous driving under crowded scenes and imperfect detection outputs, this paper proposes DSC-Track: an end-to-end online tracking framework leveraging dynamic scene cue consistency. The core innovations include a spatiotemporal Point-Pair Feature (PPF) encoder and a cue-consistency Transformer module, which explicitly model inter-frame geometric relationships to suppress interference from irrelevant objects and enable stable feature matching and trajectory association. By integrating point-pair feature encoding, Transformer-based feature alignment, and dynamic feature updating, DSC-Track significantly improves association reliability under high object density and low-quality detections. Evaluated on nuScenes and Waymo Open Dataset, it achieves state-of-the-art performance: AMOTA scores of 73.2% (validation) and 70.3% (test) on nuScenes—substantially outperforming existing methods.

0 citationsRead paper

SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization

Apr 20, 2025

Current supervised fine-tuning (SFT) of text-to-image diffusion models optimizes only pixel-level MSE loss, failing to jointly ensure global perceptual quality and structural consistency. To address this, we propose Self-supervised Direct Preference Optimization (SUDO), a novel paradigm that (i) introduces the first annotation-free, self-supervised mechanism for generating preference image pairs—jointly modeling local details and global semantics—and (ii) seamlessly integrates Direct Preference Optimization (DPO) into the diffusion training pipeline, jointly optimizing DPO loss, pixel-wise MSE, and text-conditioned sampling. Evaluated on Stable Diffusion 1.5 and XL, SUDO achieves significantly lower FID scores, alongside consistent improvements in CLIP Score and human evaluation metrics. Our method simultaneously enhances structural coherence and fine-grained realism, without requiring human-annotated preferences.

0 citationsRead paper

SAMRefiner: Taming Segment Anything Model for Universal Mask Refinement

Feb 10, 2025

To address the low quality of coarse-grain segmentation masks and high annotation costs, this paper proposes SAMRefiner++, a general-purpose and efficient mask refinement framework that is plug-and-play compatible with any pre-trained segmentation model (e.g., SAM) without requiring additional annotations. Methodologically: (i) it introduces a novel noise-robust prompting mechanism integrating distance-guided points, context-aware elastic bounding boxes, and Gaussian-style masks; (ii) it designs a “Split–Tackle–Merge” (STM) pipeline to handle multi-object segmentation challenges; and (iii) it incorporates an unsupervised IoU-adaptive optimization module combining iterative calibration with adaptive prompt engineering. Extensive experiments on multiple benchmarks demonstrate significant improvements over state-of-the-art methods in both accuracy and inference efficiency, achieving new SOTA performance. The source code is publicly available.

0 citationsRead paper