Institution profile

Qifu Technology

Industry researchasia · cn
Official website
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

Hypothesis-Driven Skill Optimization for LLM Agents

Jun 21, 2026

This work addresses the vulnerability of external skill updating to sparse or noisy execution trajectories, which can lead to the entrenchment of ineffective or even detrimental policies. To mitigate this issue, the authors propose a training-free skill optimization framework that leverages a falsifiable hypothesis-driven mechanism. By integrating controlled experiments, behavioral discrepancy analysis, and progressive skill disclosure, the method enables auditable and noise-resilient skill curation and execution at frozen model inference endpoints. Evaluated on ALFWorld, the approach yields substantial performance gains: average success rates improve by 6.9 and 4.0 percentage points for Qwen3-8B and Qwen3.6-27B, respectively. Notably, it maintains a +7.1-point advantage even under 20% erroneous feedback and demonstrates strong cross-run and cross-model transferability.

0 citationsRead paper

Semantic Triplet Restoration: A Novel Protocol for Hierarchical Table Understanding in Large Language Models

May 29, 2026

This work addresses the challenge that large language models struggle to interpret the semantic relationships embedded in two-dimensional table layouts, merged cells, and hierarchical headers. To overcome this, the authors propose the Semantic Triplet Restoration (STR) protocol, which explicitly reconstructs tables into atomic factual triplets of the form ⟨item path, attribute path, value⟩. A lightweight, query-aware routing module, TripletQL, is introduced to dynamically select relevant triplets during inference. By abandoning layout-centric representations such as HTML or Markdown, STR substantially reduces input sequence length while achieving competitive or superior performance compared to existing HTML-based approaches across four English and Chinese table question answering benchmarks. The method demonstrates particularly pronounced efficiency gains for smaller models and long tables.

0 citationsRead paper

Out-of-Distribution Detection Based on Total Variation Estimation

Jan 22, 2026

This work addresses the challenge of detecting out-of-distribution (OOD) inputs in deployed machine learning models, which often suffer from performance degradation under distributional shift. The authors propose TV-OOD, a novel OOD detection method that, for the first time, leverages Total Variation (TV) as a core principle. By employing a Total Variation network estimator to quantify each input’s contribution to the overall total variation, TV-OOD constructs a detection criterion that requires neither additional training nor generative models. Extensive experiments across multiple image classification architectures and standard benchmark datasets demonstrate that TV-OOD achieves performance comparable to or better than current state-of-the-art methods under common OOD evaluation metrics, thereby confirming its effectiveness and broad applicability.

0 citationsRead paper

QvTAD: Differential Relative Attribute Learning for Voice Timbre Attribute Detection

Aug 21, 2025

Voice timbre attribute detection (vTAD) faces two key challenges: high subjectivity in attribute descriptions and severe label imbalance—hindering fine-grained modeling and cross-speaker generalization. To address these, we propose QvTAD: (1) a graph-structured data augmentation strategy leveraging directed acyclic graphs and disjoint-set union to automatically mine high-quality paired speech samples; (2) a relative timbre offset-aware differential attention module that explicitly models attribute-level contrastive relationships; and (3) integration of speaker embeddings from pretrained FACodec, enhanced by differential denoising and contrastive amplification mechanisms. Evaluated on the VCTK-RVA benchmark, QvTAD achieves significant improvements across multiple timbre descriptors over state-of-the-art methods, with particularly pronounced gains in cross-speaker generalization. Our framework establishes a new paradigm for vTAD—interpretable, robust, and scalable.

0 citationsRead paper
Recent publications

Latest Papers

Hypothesis-Driven Skill Optimization for LLM Agents

Jun 21, 2026

This work addresses the vulnerability of external skill updating to sparse or noisy execution trajectories, which can lead to the entrenchment of ineffective or even detrimental policies. To mitigate this issue, the authors propose a training-free skill optimization framework that leverages a falsifiable hypothesis-driven mechanism. By integrating controlled experiments, behavioral discrepancy analysis, and progressive skill disclosure, the method enables auditable and noise-resilient skill curation and execution at frozen model inference endpoints. Evaluated on ALFWorld, the approach yields substantial performance gains: average success rates improve by 6.9 and 4.0 percentage points for Qwen3-8B and Qwen3.6-27B, respectively. Notably, it maintains a +7.1-point advantage even under 20% erroneous feedback and demonstrates strong cross-run and cross-model transferability.

0 citationsRead paper

Semantic Triplet Restoration: A Novel Protocol for Hierarchical Table Understanding in Large Language Models

May 29, 2026

This work addresses the challenge that large language models struggle to interpret the semantic relationships embedded in two-dimensional table layouts, merged cells, and hierarchical headers. To overcome this, the authors propose the Semantic Triplet Restoration (STR) protocol, which explicitly reconstructs tables into atomic factual triplets of the form ⟨item path, attribute path, value⟩. A lightweight, query-aware routing module, TripletQL, is introduced to dynamically select relevant triplets during inference. By abandoning layout-centric representations such as HTML or Markdown, STR substantially reduces input sequence length while achieving competitive or superior performance compared to existing HTML-based approaches across four English and Chinese table question answering benchmarks. The method demonstrates particularly pronounced efficiency gains for smaller models and long tables.

0 citationsRead paper

Out-of-Distribution Detection Based on Total Variation Estimation

Jan 22, 2026

This work addresses the challenge of detecting out-of-distribution (OOD) inputs in deployed machine learning models, which often suffer from performance degradation under distributional shift. The authors propose TV-OOD, a novel OOD detection method that, for the first time, leverages Total Variation (TV) as a core principle. By employing a Total Variation network estimator to quantify each input’s contribution to the overall total variation, TV-OOD constructs a detection criterion that requires neither additional training nor generative models. Extensive experiments across multiple image classification architectures and standard benchmark datasets demonstrate that TV-OOD achieves performance comparable to or better than current state-of-the-art methods under common OOD evaluation metrics, thereby confirming its effectiveness and broad applicability.

0 citationsRead paper

QvTAD: Differential Relative Attribute Learning for Voice Timbre Attribute Detection

Aug 21, 2025

Voice timbre attribute detection (vTAD) faces two key challenges: high subjectivity in attribute descriptions and severe label imbalance—hindering fine-grained modeling and cross-speaker generalization. To address these, we propose QvTAD: (1) a graph-structured data augmentation strategy leveraging directed acyclic graphs and disjoint-set union to automatically mine high-quality paired speech samples; (2) a relative timbre offset-aware differential attention module that explicitly models attribute-level contrastive relationships; and (3) integration of speaker embeddings from pretrained FACodec, enhanced by differential denoising and contrastive amplification mechanisms. Evaluated on the VCTK-RVA benchmark, QvTAD achieves significant improvements across multiple timbre descriptors over state-of-the-art methods, with particularly pronounced gains in cross-speaker generalization. Our framework establishes a new paradigm for vTAD—interpretable, robust, and scalable.

0 citationsRead paper