Institution profile

Shopee

Industry researchasia · sg
Official website
Research library44linked papers
Opportunities0open roles
Selected work

Representative Papers

Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation

Aug 03, 2026

This work addresses the challenge of transferring and continually updating pre-trained knowledge in recommender systems under behavioral distribution shifts. To this end, the authors propose a knowledge–geometry disentanglement framework that extracts transferable knowledge via Behavior Multi-Token Prediction (BMTP) and decouples knowledge encoding from task-specific geometric learning through read-only cross-attention, Anchored Calibration Residuals (ACR), and orthogonal embedding spaces. This design enables interference-free knowledge updates and efficient downstream adaptation. The method consistently outperforms strong baselines by 4–12% across eight public benchmarks and demonstrates sustained effectiveness on Shopee’s 90-day online streaming data, with A/B tests showing a 1.75% increase in per-user GMV and a 1.53% boost in ad revenue.

0 citationsRead paper

RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

Jul 30, 2026

This work addresses the limitations of existing representation engineering approaches, which rely on synthetic data and suffer from irreproducible evaluations and susceptibility to superficial patterns. The authors construct the first large-scale, multi-source aligned capability representation framework grounded in real-world benchmarks, curated from over 10,000 academic papers and hundreds of public datasets, spanning 94 distinct capabilities. This framework enables cross-benchmark aggregation of capability vectors and transferable evaluation, effectively mitigating bias from any single data source. Experiments across 12 large language models reveal that benchmark-pooled capability vectors exhibit stable clustering structures; differential mean achieves the best performance in 10 models, while logistic regression outperforms others across the greatest number of capability–model combinations, underscoring the critical influence of both evaluation dimensions and readout methodologies.

0 citationsRead paper

OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

Jul 21, 2026

This work addresses the challenges of long-form audiovisual reasoning, where critical evidence is sparse and distributed across modalities, and high-fidelity inputs incur prohibitive computational costs that existing full-modality large language models struggle to handle efficiently. The authors propose a tool-augmented post-training framework in which the model first constructs a low-cost global preview and then selectively invokes localized high-fidelity zoom-in tools for fine-grained analysis as needed. A novel TimeAnchor mechanism ensures temporal consistency across multiple granularities, while a temporally enhanced data engine automatically generates tool-use trajectories without manual annotation. Through joint optimization via supervised fine-tuning and reinforcement learning, the method significantly improves answer accuracy and temporal localization on multiple audiovisual benchmarks, while concentrating high-fidelity computation on information-dense regions.

0 citationsRead paper

MM-TRELLIS: Point-Cloud Guided Multi-Modal 3D Vehicle Generation in Autonomous Driving

Jun 23, 2026

Existing methods struggle to fuse multi-view images and LiDAR point clouds for generating geometrically accurate and high-fidelity 3D vehicle models in real-world driving scenarios. This work proposes MM-TRELLIS, the first approach to incorporate LiDAR point clouds as test-time guidance within a native 3D diffusion generative model. By conditioning on multi-view images and enforcing geometric alignment during the denoising process, MM-TRELLIS achieves highly consistent generation through multimodal fusion. Additionally, it introduces a voxel filtering strategy based on 3D Gaussian Splatting opacities to effectively suppress floating artifacts. Evaluated on the Waymo dataset, the method significantly outperforms existing approaches, achieving state-of-the-art performance in fidelity, geometric accuracy, and cross-view consistency of the generated vehicle models.

0 citationsRead paper

EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent

Jun 16, 2026

This work addresses the limitations of existing shopping agent benchmarks, which fail to capture the realistic distribution of users’ hidden intentions across queries, profiles, and interactive clarifications, nor assess agents’ ability to uncover needs in long-horizon tasks. The authors introduce a new benchmark comprising 662 shopping tasks grounded in real Amazon products and reviews, requiring agents to infer and satisfy hidden user needs through up to 100 tool calls and ultimately recommend a single item. They explicitly decouple user intent into three components: observable queries, tool-gated profiles, and scripted clarifications, and propose a fine-grained scoring scheme with type- and source-aware labels to enable failure attribution. The dataset is generated via an automated pipeline with answers fixed prior to textual generation to ensure reliability. Evaluations on seven state-of-the-art models reveal that even the strongest achieves only 57.1% overall accuracy and performs significantly worse on hidden than explicit intents, exposing a critical bottleneck in current systems.

0 citationsRead paper
Recent publications

Latest Papers

Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation

Aug 03, 2026

This work addresses the challenge of transferring and continually updating pre-trained knowledge in recommender systems under behavioral distribution shifts. To this end, the authors propose a knowledge–geometry disentanglement framework that extracts transferable knowledge via Behavior Multi-Token Prediction (BMTP) and decouples knowledge encoding from task-specific geometric learning through read-only cross-attention, Anchored Calibration Residuals (ACR), and orthogonal embedding spaces. This design enables interference-free knowledge updates and efficient downstream adaptation. The method consistently outperforms strong baselines by 4–12% across eight public benchmarks and demonstrates sustained effectiveness on Shopee’s 90-day online streaming data, with A/B tests showing a 1.75% increase in per-user GMV and a 1.53% boost in ad revenue.

0 citationsRead paper

RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

Jul 30, 2026

This work addresses the limitations of existing representation engineering approaches, which rely on synthetic data and suffer from irreproducible evaluations and susceptibility to superficial patterns. The authors construct the first large-scale, multi-source aligned capability representation framework grounded in real-world benchmarks, curated from over 10,000 academic papers and hundreds of public datasets, spanning 94 distinct capabilities. This framework enables cross-benchmark aggregation of capability vectors and transferable evaluation, effectively mitigating bias from any single data source. Experiments across 12 large language models reveal that benchmark-pooled capability vectors exhibit stable clustering structures; differential mean achieves the best performance in 10 models, while logistic regression outperforms others across the greatest number of capability–model combinations, underscoring the critical influence of both evaluation dimensions and readout methodologies.

0 citationsRead paper

OmniReasoner: Thinking with Long Audio-Video via Native Tool Use

Jul 21, 2026

This work addresses the challenges of long-form audiovisual reasoning, where critical evidence is sparse and distributed across modalities, and high-fidelity inputs incur prohibitive computational costs that existing full-modality large language models struggle to handle efficiently. The authors propose a tool-augmented post-training framework in which the model first constructs a low-cost global preview and then selectively invokes localized high-fidelity zoom-in tools for fine-grained analysis as needed. A novel TimeAnchor mechanism ensures temporal consistency across multiple granularities, while a temporally enhanced data engine automatically generates tool-use trajectories without manual annotation. Through joint optimization via supervised fine-tuning and reinforcement learning, the method significantly improves answer accuracy and temporal localization on multiple audiovisual benchmarks, while concentrating high-fidelity computation on information-dense regions.

0 citationsRead paper

MM-TRELLIS: Point-Cloud Guided Multi-Modal 3D Vehicle Generation in Autonomous Driving

Jun 23, 2026

Existing methods struggle to fuse multi-view images and LiDAR point clouds for generating geometrically accurate and high-fidelity 3D vehicle models in real-world driving scenarios. This work proposes MM-TRELLIS, the first approach to incorporate LiDAR point clouds as test-time guidance within a native 3D diffusion generative model. By conditioning on multi-view images and enforcing geometric alignment during the denoising process, MM-TRELLIS achieves highly consistent generation through multimodal fusion. Additionally, it introduces a voxel filtering strategy based on 3D Gaussian Splatting opacities to effectively suppress floating artifacts. Evaluated on the Waymo dataset, the method significantly outperforms existing approaches, achieving state-of-the-art performance in fidelity, geometric accuracy, and cross-view consistency of the generated vehicle models.

0 citationsRead paper

EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent

Jun 16, 2026

This work addresses the limitations of existing shopping agent benchmarks, which fail to capture the realistic distribution of users’ hidden intentions across queries, profiles, and interactive clarifications, nor assess agents’ ability to uncover needs in long-horizon tasks. The authors introduce a new benchmark comprising 662 shopping tasks grounded in real Amazon products and reviews, requiring agents to infer and satisfy hidden user needs through up to 100 tool calls and ultimately recommend a single item. They explicitly decouple user intent into three components: observable queries, tool-gated profiles, and scripted clarifications, and propose a fine-grained scoring scheme with type- and source-aware labels to enable failure attribution. The dataset is generated via an automated pipeline with answers fixed prior to textual generation to ensure reliability. Evaluations on seven state-of-the-art models reveal that even the strongest achieves only 57.1% overall accuracy and performs significantly worse on hidden than explicit intents, exposing a critical bottleneck in current systems.

0 citationsRead paper