Institution profile

UniDT Technology

Industry research
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Reconciling Contradictory Views on the Effectiveness of SFT in LLMs: An Interaction Perspective

May 18, 2026

This study investigates why supervised fine-tuning (SFT) is consistently effective for small models yet yields inconsistent or even detrimental results in large language models (LLMs). To address this, the work introduces a novel perspective by analyzing the evolution of inter-token interactions during SFT, leveraging interaction-based interpretability techniques to quantify and track dynamic changes in interaction strength. The findings reveal that in LLMs, SFT primarily acts as a denoising mechanism within an extremely short training window, after which it rapidly overfits, leading to performance degradation. This insight offers a new understanding of the boundaries of SFT effectiveness and empirically validates the necessity of early stopping across multiple LLMs and datasets, providing practical guidance for optimizing fine-tuning protocols.

0 citationsRead paper

Unilaw-R1: A Large Language Model for Legal Reasoning with Reinforcement Learning and Iterative Inference

Oct 11, 2025

To address three key challenges in complex legal reasoning—insufficient domain-specific knowledge, unreliable logical consistency, and poor generalization across legal tasks—this paper introduces UniLaw-R1, the first open-source 7B-parameter large language model explicitly optimized for legal reasoning. Methodologically, we propose a two-stage training paradigm: (i) supervised fine-tuning (SFT) on 17K high-quality legal chain-of-thought samples, followed by (ii) reinforcement learning (RL)-based joint optimization, augmented with an iterative reasoning mechanism. We further construct Unilaw-R1-Eval, a dedicated benchmark for rigorous evaluation. Experimental results demonstrate that UniLaw-R1 achieves an average 6.6% improvement over Qwen-2.5-7B-Instruct on LawBench and LexEval, matching the performance of the 32B-parameter DeepSeek-R1-Distill-Qwen-32B (54.9%). The model significantly enhances both accuracy and interpretability in legal reasoning tasks.

0 citationsRead paper

EmoHead: Emotional Talking Head via Manipulating Semantic Expression Parameters

Mar 25, 2025

This work addresses the challenge of emotion-controllable audio-driven talking-head video generation, where abstract and ambiguous emotional semantics hinder precise facial expression control. We propose a label-specifiable audio-expression mapping module, integrated with a pretrained hyperplane orthogonal probing mechanism, enabling—for the first time—fine-grained, disentangled modulation of neural radiance fields (NeRFs) via interpretable emotion parameters. Our method jointly models audio and facial expressions while explicitly disentangling semantic emotion parameters, facilitating emotion-aware head animation synthesis within the NeRF rendering framework. Extensive experiments demonstrate significant improvements in expression reconstruction fidelity and emotion consistency across multi-emotion control tasks. Both visual quality and emotion controllability achieve state-of-the-art performance.

0 citationsRead paper
Recent publications

Latest Papers

Reconciling Contradictory Views on the Effectiveness of SFT in LLMs: An Interaction Perspective

May 18, 2026

This study investigates why supervised fine-tuning (SFT) is consistently effective for small models yet yields inconsistent or even detrimental results in large language models (LLMs). To address this, the work introduces a novel perspective by analyzing the evolution of inter-token interactions during SFT, leveraging interaction-based interpretability techniques to quantify and track dynamic changes in interaction strength. The findings reveal that in LLMs, SFT primarily acts as a denoising mechanism within an extremely short training window, after which it rapidly overfits, leading to performance degradation. This insight offers a new understanding of the boundaries of SFT effectiveness and empirically validates the necessity of early stopping across multiple LLMs and datasets, providing practical guidance for optimizing fine-tuning protocols.

0 citationsRead paper

Unilaw-R1: A Large Language Model for Legal Reasoning with Reinforcement Learning and Iterative Inference

Oct 11, 2025

To address three key challenges in complex legal reasoning—insufficient domain-specific knowledge, unreliable logical consistency, and poor generalization across legal tasks—this paper introduces UniLaw-R1, the first open-source 7B-parameter large language model explicitly optimized for legal reasoning. Methodologically, we propose a two-stage training paradigm: (i) supervised fine-tuning (SFT) on 17K high-quality legal chain-of-thought samples, followed by (ii) reinforcement learning (RL)-based joint optimization, augmented with an iterative reasoning mechanism. We further construct Unilaw-R1-Eval, a dedicated benchmark for rigorous evaluation. Experimental results demonstrate that UniLaw-R1 achieves an average 6.6% improvement over Qwen-2.5-7B-Instruct on LawBench and LexEval, matching the performance of the 32B-parameter DeepSeek-R1-Distill-Qwen-32B (54.9%). The model significantly enhances both accuracy and interpretability in legal reasoning tasks.

0 citationsRead paper

EmoHead: Emotional Talking Head via Manipulating Semantic Expression Parameters

Mar 25, 2025

This work addresses the challenge of emotion-controllable audio-driven talking-head video generation, where abstract and ambiguous emotional semantics hinder precise facial expression control. We propose a label-specifiable audio-expression mapping module, integrated with a pretrained hyperplane orthogonal probing mechanism, enabling—for the first time—fine-grained, disentangled modulation of neural radiance fields (NeRFs) via interpretable emotion parameters. Our method jointly models audio and facial expressions while explicitly disentangling semantic emotion parameters, facilitating emotion-aware head animation synthesis within the NeRF rendering framework. Extensive experiments demonstrate significant improvements in expression reconstruction fidelity and emotion consistency across multi-emotion control tasks. Both visual quality and emotion controllability achieve state-of-the-art performance.

0 citationsRead paper