Institution profile

Soochow University

Academic institutionasia · cn
Official website
Research library406linked papers
Opportunities0open roles
Selected work

Representative Papers

The Box is in the Pen: Evaluating Commonsense Reasoning in Neural Machine Translation

Nov 01, 2020Findings

This work investigates neural machine translation (NMT) models’ capacity to resolve lexical and syntactic ambiguities via commonsense reasoning. To this end, we introduce CReM—the first commonsense reasoning evaluation benchmark tailored for NMT—comprising 1,200 triplets spanning seven categories of commonsense knowledge. We propose a dual-dimensional evaluation framework assessing both accuracy (where mainstream models achieve only 60.1%) and cross-context consistency (with merely 31% consistency), enabling the first systematic quantification of NMT’s commonsense reasoning capability. Extensive comparative experiments are conducted using models including BERT and GPT-2; statistical analysis and error attribution reveal that contextual modeling and effective commonsense integration remain critical bottlenecks. The CReM benchmark is publicly released to serve as a standardized evaluation tool for future research.

24 citations1 influentialRead paper

Eye-tracked Virtual Reality: A Comprehensive Survey on Methods and Privacy Challenges

May 23, 2023arXiv.org

This paper addresses privacy leakage risks arising from the correlation between eye-tracking data and visual stimuli in VR environments. We systematically survey full-stack VR eye-tracking technologies—from pupil detection and gaze estimation to cognitive modeling—published between 2012 and 2022, alongside their associated privacy threats. First, we establish the first cross-disciplinary survey framework bridging VR eye-tracking and privacy protection, identifying three privacy-centric research directions. Second, we propose a novel co-design paradigm integrating eye movement authentication with data anonymization, synergizing computer vision, human-computer interaction modeling, differential privacy, adversarial generation, and biometric encryption. Third, we clarify the technological evolution trajectory and privacy threat landscape, and introduce quantifiable evaluation metrics and an implementable defense roadmap. Our work provides both theoretical foundations and practical guidelines for developing secure and trustworthy VR systems. (149 words)

21 citationsRead paper

Bounded and Uniform Energy-based Out-of-distribution Detection for Graphs

Apr 18, 2025International Conference on Machine Learning

Graph Neural Networks (GNNs) suffer from distorted out-of-distribution (OOD) node scoring due to unbounded negative-energy scores and logit shift, impairing reliability in node-level OOD detection. Method: We propose a dual-optimization energy calibration framework that jointly enforces bounded negative-energy constraints and suppresses logit shift, integrating energy-based OOD discrimination, graph embedding enhancement, and adaptive logit calibration. Contribution/Results: The framework significantly improves OOD detection robustness and cross-graph consistency. On structural perturbation OOD benchmarks, it reduces False Positive Rate at 95% True Positive Rate (FPR95) by 28.4% (without OOD exposure) and 22.7% (with OOD exposure) over current state-of-the-art methods. This establishes a new paradigm for trustworthy, safety-critical GNN-based OOD detection.

3 citationsRead paper

Continuous-time q-Learning for Jump-Diffusion Models under Tsallis Entropy

Jul 04, 2024arXiv.org

This paper investigates continuous-time reinforcement learning under jump-diffusion dynamics. Methodologically, it introduces Tsallis entropy regularization into the Q-learning framework for the first time—overcoming the inherent limitations of Gibbs-type policies arising from Shannon entropy—and establishes a martingale characterization of the Q-function via martingale representation theory and stochastic optimal control, incorporating Lagrange multipliers to derive non-Gibbsian optimal policies with compact support. Two algorithmic variants—explicit and implicit Lagrange multiplier-based continuous-time Q-learning—are proposed, enabling Actor-Critic-style alternating updates. Analytical solutions are obtained for two financial problems: optimal portfolio liquidation and nonlinear quadratic control. Numerical experiments demonstrate high stability and rapid convergence. The core contribution is the first rigorous Tsallis entropy-driven continuous-time Q-learning theoretical framework, validated for its superior modeling capability and computational efficacy under jump-diffusion dynamics.

3 citationsRead paper

Optimal payoff under Bregman-Wasserstein divergence constraints

Nov 27, 2024

This paper studies the optimal payoff selection problem for expected utility maximizers under a Bregman–Wasserstein (BW) divergence constraint, designed to control deviation from a reference payoff while allowing asymmetric penalties for upside and downside deviations—better aligning with real-world investment objectives. Methodologically, it provides the first analytical solution to the optimal payoff structure under BW divergence constraints, employing a convex function φ to flexibly encode directional deviation preferences and thereby overcoming the symmetry limitation inherent in classical Wasserstein distance. By integrating convex analysis, optimal transport theory, and stochastic optimization, the authors formulate a utility maximization framework regularized by a Bregman penalty term. Theoretically, they derive a closed-form expression for the optimal payoff. Numerical experiments demonstrate that tuning φ enables precise calibration of risk attitudes and significantly improves alignment between payoff allocation and investor-specific goals.

2 citationsRead paper
Recent publications

Latest Papers

Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue

Aug 11, 2026

This work addresses the challenge that existing large language models struggle to consistently track users’ dynamically evolving emotions in multi-turn empathetic dialogues, and conventional reinforcement learning suffers from a mismatch between policy and experience due to fixed interaction distributions. To overcome these limitations, the authors propose a dual-loop self-evolution framework: an inner loop optimizes the empathetic policy using verifiable emotion-based rewards, while an outer loop dynamically reshapes the interaction experience distribution based on policy performance, enabling co-evolution of policy and data. The approach innovatively integrates boundary-capability-prioritized sampling, uncertainty-guided exploration, and uniform replay mechanisms to achieve sample-efficient training under limited rollout budgets. Evaluated on the SAGE benchmark, the method boosts Qwen3-8B’s performance from 53.87 to 79.24, substantially outperforming protocol-matched baselines by +7.23 points.

0 citationsRead paper

REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems

Aug 11, 2026

Current safety evaluations of large language model (LLM) agents predominantly rely on single-metric attack success rates, which inadequately capture the real-world risk of policy violations during environmental interaction. This work proposes an executable red-teaming framework that generates attacks grounded in explicit safety constraints, executes them within an isolated sandbox, and validates actual harm through service credentials and final-state changes. The study introduces a novel state-anchored diagnostic mechanism to uncover the agent’s “recognition–execution gap” and designs a training-free policy reminder that substantially reduces policy violations. Evaluated across 1,661 test cases involving six models and three agent frameworks, the macro-average attack success rate reaches 65.69%; notably, the policy reminder reduces confirmed violation rates by over 70 percentage points.

0 citationsRead paper

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

Aug 10, 2026

This work addresses the limited cost-effectiveness of existing routing methods that merely assign simple tasks to small models without enhancing their capabilities. To overcome this, the authors propose a multi-cycle adaptation mechanism operating at the granularity of single inference calls. The approach leverages a teacher model to generate verification demonstrations from the small model’s failures, integrating skill distillation and LoRA fine-tuning to continuously improve its competence. Joint optimization is performed over a dynamic skill library, task-specific adapters, and a cost-calibrated routing policy, complemented by a verifier-supported fallback mechanism. Experiments show that Qwen2.5-Coder-1.5B achieves a pass rate increase from 28.7% to 49.7% on HumanEval+MBPP; the deployment strategy attains 88.3% of peak performance at only 60.8% of the cost; and Qwen3.5-2B matches the performance of an unadapted 4B model on TAU-2.

0 citationsRead paper

RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation

Aug 09, 2026

This work addresses the challenge of KV cache memory constraints in long-context large language model inference, where existing methods struggle to allocate limited cache budgets effectively across layers. The authors propose a dynamic cache allocation strategy grounded in perturbation propagation effects: by injecting norm-adaptive perturbations into value caches and measuring their impact on the final output distribution via KL divergence, they quantify each layer’s sensitivity to compression. This yields a normalized sensitivity profile, which is then mapped exponentially to distribute the cache budget across layers. Notably, this approach is the first to assess inter-layer sensitivity based on predictive perturbations, overcoming limitations of methods relying on layer depth or static attention statistics. Evaluated on the LongBench benchmark, it significantly outperforms current KV compression techniques under identical cache budgets, achieving state-of-the-art average performance.

0 citationsRead paper

PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling

Aug 09, 2026

Existing tool-calling evaluation benchmarks suffer from insufficient coverage of long-tail scenarios, a lack of adversarial negative samples, and reliance on large language model (LLM) annotations未经 execution verification, hindering fine-grained error attribution. This work proposes a two-stage diagnostic benchmark construction framework: first decoupling candidate tool-call generation from deterministic validation based on actual API execution, and then generating multi-difficulty queries—including adversarial negatives targeting omission, hallucination, and parameter errors—through decomposition along plugin capability, intent, and boundary dimensions, iteratively refined until convergence. By integrating execution-verified labels and an LLM-as-judge mechanism, the approach overcomes the limitations of conventional benchmarks that rely solely on end-to-end accuracy. Evaluations across five major model families demonstrate that the proposed benchmark precisely characterizes error profiles, reveals performance disparities across difficulty levels and error types, and achieves high agreement between LLM judgments and human annotations.

0 citationsRead paper