Institution profile

Alipay

Industry researchasia · cn
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Beyond Retrieval: A Multitask Benchmark and Model for Code Search

May 06, 2026

Code search has usually been evaluated as first-stage retrieval, even though production systems rely on broader pipelines with reranking and developer-style queries. Existing benchmarks also suffer from data contamination, label noise, and degenerate binary relevance. In this paper, we introduce \textsc{CoREB}, a contamination-limited, multitask \underline{co}de \underline{r}etrieval and r\underline{e}ranking \underline{b}enchmark, together with a fine-tuned code reranker, that goes beyond retrieval to cover the full code search pipeline. \textsc{CoREB} is built from counterfactually rewritten LiveCodeBench problems in five programming languages and delivered as timed releases with graded relevance judgments. We benchmark eleven embedding models and five rerankers across three tasks: text-to-code, code-to-text, and code-to-code. Our experiments reveal that: \circone code-specialised embeddings dominate code-to-code retrieval (${\sim}2{\times}$ over general encoders), yet no single model wins all three tasks; \circtwo short keyword queries, the format closest to real developer search, collapse every model to near-zero nDCG@10; \circthree off-the-shelf rerankers are task-asymmetric, with a 12-point swing on code-to-code and no baseline net-positive across all tasks; \circfour our fine-tuned \textsc{CoREB-Reranker} is the first to achieve consistent gains across all three tasks. The data and model are released.

0 citationsRead paper

EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation

Jul 05, 2025

Human animation generation faces two major bottlenecks: slow inference and task fragmentation—existing video generation models incur high computational costs, while lip-syncing, audio-driven animation, and keyframe interpolation require separate specialized models. This paper proposes a unified multi-task human animation generation framework centered on spatiotemporal local reconstruction for holistic task modeling. We introduce a multimodal disentangled cross-attention module and a MAE-inspired input reconstruction mechanism to enable flexible conditional control from text, audio, or keyframes. Training employs an alternating SFT + reward-based reinforcement learning strategy. Our 1.3B-parameter model surpasses ten-times-larger baselines in facial and upper-body video generation, achieving superior fidelity, strong generalization across diverse inputs, and over 100× inference speedup—significantly enhancing practical deployability.

0 citationsRead paper
Recent publications

Latest Papers

Beyond Retrieval: A Multitask Benchmark and Model for Code Search

May 06, 2026

Code search has usually been evaluated as first-stage retrieval, even though production systems rely on broader pipelines with reranking and developer-style queries. Existing benchmarks also suffer from data contamination, label noise, and degenerate binary relevance. In this paper, we introduce \textsc{CoREB}, a contamination-limited, multitask \underline{co}de \underline{r}etrieval and r\underline{e}ranking \underline{b}enchmark, together with a fine-tuned code reranker, that goes beyond retrieval to cover the full code search pipeline. \textsc{CoREB} is built from counterfactually rewritten LiveCodeBench problems in five programming languages and delivered as timed releases with graded relevance judgments. We benchmark eleven embedding models and five rerankers across three tasks: text-to-code, code-to-text, and code-to-code. Our experiments reveal that: \circone code-specialised embeddings dominate code-to-code retrieval (${\sim}2{\times}$ over general encoders), yet no single model wins all three tasks; \circtwo short keyword queries, the format closest to real developer search, collapse every model to near-zero nDCG@10; \circthree off-the-shelf rerankers are task-asymmetric, with a 12-point swing on code-to-code and no baseline net-positive across all tasks; \circfour our fine-tuned \textsc{CoREB-Reranker} is the first to achieve consistent gains across all three tasks. The data and model are released.

0 citationsRead paper

EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation

Jul 05, 2025

Human animation generation faces two major bottlenecks: slow inference and task fragmentation—existing video generation models incur high computational costs, while lip-syncing, audio-driven animation, and keyframe interpolation require separate specialized models. This paper proposes a unified multi-task human animation generation framework centered on spatiotemporal local reconstruction for holistic task modeling. We introduce a multimodal disentangled cross-attention module and a MAE-inspired input reconstruction mechanism to enable flexible conditional control from text, audio, or keyframes. Training employs an alternating SFT + reward-based reinforcement learning strategy. Our 1.3B-parameter model surpasses ten-times-larger baselines in facial and upper-body video generation, achieving superior fidelity, strong generalization across diverse inputs, and over 100× inference speedup—significantly enhancing practical deployability.

0 citationsRead paper