Institution profile

Yangzhou University

Academic institutionasia · cn
Official website
Research library51linked papers
Opportunities0open roles
Selected work

Representative Papers

Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs

Jul 08, 2026

Large language models are vulnerable to adversarial prompts and jailbreaking attacks, yet the precise mechanisms by which their internal reasoning is perturbed remain unclear. This work proposes a causal diagnostic framework grounded in internal computation graphs, which constructs and compares attribution graphs under clean and adversarial prompts to enable structural alignment and path analysis. The approach identifies invariant, suppressed, and emergent computational motifs, revealing a strong association between vulnerability motifs and unsafe model behaviors. It further supports node-level causal interventions. Experiments across multiple open-source large language models and jailbreaking benchmarks demonstrate that targeted interventions significantly enhance model robustness, confirming that structural biases in the computation graph are a key factor enabling successful attacks.

0 citationsRead paper

Knowledge-Enhanced Agentic Vulnerability Repair

Jul 01, 2026

Existing automated program repair approaches struggle to precisely identify root causes and underutilize historical repair knowledge. This work proposes KeaRepair, a knowledge-enhanced, agent-driven repair framework that uniquely integrates multidimensional historical vulnerability knowledge with tool-augmented ReAct-style reasoning to establish a closed-loop, knowledge-driven process for both diagnosis and patch generation. By combining a multi-perspective knowledge base, retrieval-augmented generation, and a multi-layer validation mechanism leveraging compilation, proof-of-concept exploits, and test suites, KeaRepair achieves cross-language generalization. Evaluated on 55 real-world C/C++ vulnerabilities, it attains a repair rate of 83.64% (46/55), including six unique cases that baseline methods fail to resolve.

0 citationsRead paper
Recent publications

Latest Papers

Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs

Jul 08, 2026

Large language models are vulnerable to adversarial prompts and jailbreaking attacks, yet the precise mechanisms by which their internal reasoning is perturbed remain unclear. This work proposes a causal diagnostic framework grounded in internal computation graphs, which constructs and compares attribution graphs under clean and adversarial prompts to enable structural alignment and path analysis. The approach identifies invariant, suppressed, and emergent computational motifs, revealing a strong association between vulnerability motifs and unsafe model behaviors. It further supports node-level causal interventions. Experiments across multiple open-source large language models and jailbreaking benchmarks demonstrate that targeted interventions significantly enhance model robustness, confirming that structural biases in the computation graph are a key factor enabling successful attacks.

0 citationsRead paper

Knowledge-Enhanced Agentic Vulnerability Repair

Jul 01, 2026

Existing automated program repair approaches struggle to precisely identify root causes and underutilize historical repair knowledge. This work proposes KeaRepair, a knowledge-enhanced, agent-driven repair framework that uniquely integrates multidimensional historical vulnerability knowledge with tool-augmented ReAct-style reasoning to establish a closed-loop, knowledge-driven process for both diagnosis and patch generation. By combining a multi-perspective knowledge base, retrieval-augmented generation, and a multi-layer validation mechanism leveraging compilation, proof-of-concept exploits, and test suites, KeaRepair achieves cross-language generalization. Evaluated on 55 real-world C/C++ vulnerabilities, it attains a repair rate of 83.64% (46/55), including six unique cases that baseline methods fail to resolve.

0 citationsRead paper