Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks
针对复杂多步骤创意任务中自动评估的不稳定性及偏见问题,提出CreaEval方法,通过分析与评判分离的方式提高评估可靠性。
针对复杂多步骤创意任务中自动评估的不稳定性及偏见问题,提出CreaEval方法,通过分析与评判分离的方式提高评估可靠性。
本文提出REP-LIE方法,通过低秩矩阵梯度估计权重重要性并进行迭代剪枝,以减少Transformer模型在微调过程中的资源消耗。
本文提出了一种新的弹幕生成框架Genda及基于弹幕的多模态假新闻检测模型DM-FEND,解决了实时检测中的时间不一致性问题。
Large language models are vulnerable to adversarial prompts and jailbreaking attacks, yet the precise mechanisms by which their internal reasoning is perturbed remain unclear. This work proposes a causal diagnostic framework grounded in internal computation graphs, which constructs and compares attribution graphs under clean and adversarial prompts to enable structural alignment and path analysis. The approach identifies invariant, suppressed, and emergent computational motifs, revealing a strong association between vulnerability motifs and unsafe model behaviors. It further supports node-level causal interventions. Experiments across multiple open-source large language models and jailbreaking benchmarks demonstrate that targeted interventions significantly enhance model robustness, confirming that structural biases in the computation graph are a key factor enabling successful attacks.
Existing automated program repair approaches struggle to precisely identify root causes and underutilize historical repair knowledge. This work proposes KeaRepair, a knowledge-enhanced, agent-driven repair framework that uniquely integrates multidimensional historical vulnerability knowledge with tool-augmented ReAct-style reasoning to establish a closed-loop, knowledge-driven process for both diagnosis and patch generation. By combining a multi-perspective knowledge base, retrieval-augmented generation, and a multi-layer validation mechanism leveraging compilation, proof-of-concept exploits, and test suites, KeaRepair achieves cross-language generalization. Evaluated on 55 real-world C/C++ vulnerabilities, it attains a repair rate of 83.64% (46/55), including six unique cases that baseline methods fail to resolve.
针对复杂多步骤创意任务中自动评估的不稳定性及偏见问题,提出CreaEval方法,通过分析与评判分离的方式提高评估可靠性。
本文提出REP-LIE方法,通过低秩矩阵梯度估计权重重要性并进行迭代剪枝,以减少Transformer模型在微调过程中的资源消耗。
本文提出了一种新的弹幕生成框架Genda及基于弹幕的多模态假新闻检测模型DM-FEND,解决了实时检测中的时间不一致性问题。
Large language models are vulnerable to adversarial prompts and jailbreaking attacks, yet the precise mechanisms by which their internal reasoning is perturbed remain unclear. This work proposes a causal diagnostic framework grounded in internal computation graphs, which constructs and compares attribution graphs under clean and adversarial prompts to enable structural alignment and path analysis. The approach identifies invariant, suppressed, and emergent computational motifs, revealing a strong association between vulnerability motifs and unsafe model behaviors. It further supports node-level causal interventions. Experiments across multiple open-source large language models and jailbreaking benchmarks demonstrate that targeted interventions significantly enhance model robustness, confirming that structural biases in the computation graph are a key factor enabling successful attacks.
Existing automated program repair approaches struggle to precisely identify root causes and underutilize historical repair knowledge. This work proposes KeaRepair, a knowledge-enhanced, agent-driven repair framework that uniquely integrates multidimensional historical vulnerability knowledge with tool-augmented ReAct-style reasoning to establish a closed-loop, knowledge-driven process for both diagnosis and patch generation. By combining a multi-perspective knowledge base, retrieval-augmented generation, and a multi-layer validation mechanism leveraging compilation, proof-of-concept exploits, and test suites, KeaRepair achieves cross-language generalization. Evaluated on 55 real-world C/C++ vulnerabilities, it attains a repair rate of 83.64% (46/55), including six unique cases that baseline methods fail to resolve.