Institution profile

University of Akron

Academic institutionnorthamerica · us
Official website
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

Preventing Error Propagation in Multi-Agent AI through Runtime Monitoring

Jun 27, 2026

This work addresses the dual-edged nature of reasoning trace exchange in multi-agent AI systems, where sharing intermediate reasoning steps can both correct errors and inadvertently propagate inaccuracies, potentially steering correct answers toward incorrect ones. The authors propose a runtime monitoring framework in which multiple language models first answer multiple-choice questions independently, then exchange reasoning traces to revise their responses. This approach enables the first systematic quantification of the proportions of beneficial versus detrimental answer revisions and identifies key conditions under which error propagation occurs. Empirical evaluations across domains including cybersecurity, networking, and general knowledge demonstrate that multi-agent reasoning can enhance accuracy under specific conditions, while also delineating scenarios where it introduces significant risks.

0 citationsRead paper

Memory as an Attack Surface in LLM Agents: A Study on Multiple-Choice Question Answering

Jun 27, 2026

This work demonstrates that the external memory mechanisms of large language model (LLM) agents constitute a novel attack surface: even when provided with clean inputs, maliciously injected memories can induce erroneous outputs. The study introduces an LLM agent architecture with controllable external memory and systematically investigates, for the first time, how memory manipulation affects multiple-choice question-answering behavior. Through carefully designed memory injection and evaluation experiments, quantitative analysis reveals that only a small number of misleading memory entries significantly degrade answer accuracy and effectively steer the agent toward selecting predetermined incorrect options. These findings establish that the memory component poses a tangible security risk, highlighting the need for robust safeguards against adversarial memory tampering in LLM-based systems.

0 citationsRead paper

Classification of Single and Mixed Partial Discharges under Switching Voltage Using an AWA-CNN Framework

May 20, 2026

This study addresses the challenge of partial discharge (PD) identification under switching voltage excitation, where discharges concentrate around voltage transitions and are significantly harder to classify than under sinusoidal conditions. To tackle this, the authors propose an amplitude–width–area (AWA) visualization method that maps time-domain pulse features into a two-dimensional image: amplitude and area serve as spatial coordinates, while pulse width is encoded by color, effectively revealing distinct distribution patterns of different PD sources. Leveraging this representation, convolutional neural networks—including InceptionV3 and ResNet-18—are employed to achieve high-accuracy classification of six single- and mixed-source PD types. Experimental results demonstrate a classification accuracy exceeding 96%, substantially outperforming a Random Forest baseline (73.33%), thereby validating the effectiveness and superiority of the proposed approach for multi-class PD identification in complex switching voltage environments.

0 citationsRead paper

Adversarial Reframing: A Framework for Targeted Generation in Language Models

May 20, 2026

Large language models are vulnerable to prompt-based attacks (jailbreaking) that circumvent safety mechanisms and elicit harmful content. This work proposes the THREAT framework, which formalizes adversarial prompt generation as a non-convex optimization problem for the first time. By integrating multi-agent collaborative reasoning, iterative adversarial search, and language model redirection techniques, THREAT efficiently produces highly stealthy jailbreaking prompts. Experimental results demonstrate that the method significantly outperforms existing attack strategies across multiple models and datasets, achieving higher attack success rates with lower computational overhead. Notably, fewer than 1% of the generated prompts are flagged as harmful—despite an original refusal rate of approximately 50%—thereby exposing previously undetected security vulnerabilities in current alignment approaches.

0 citationsRead paper
Recent publications

Latest Papers

Preventing Error Propagation in Multi-Agent AI through Runtime Monitoring

Jun 27, 2026

This work addresses the dual-edged nature of reasoning trace exchange in multi-agent AI systems, where sharing intermediate reasoning steps can both correct errors and inadvertently propagate inaccuracies, potentially steering correct answers toward incorrect ones. The authors propose a runtime monitoring framework in which multiple language models first answer multiple-choice questions independently, then exchange reasoning traces to revise their responses. This approach enables the first systematic quantification of the proportions of beneficial versus detrimental answer revisions and identifies key conditions under which error propagation occurs. Empirical evaluations across domains including cybersecurity, networking, and general knowledge demonstrate that multi-agent reasoning can enhance accuracy under specific conditions, while also delineating scenarios where it introduces significant risks.

0 citationsRead paper

Memory as an Attack Surface in LLM Agents: A Study on Multiple-Choice Question Answering

Jun 27, 2026

This work demonstrates that the external memory mechanisms of large language model (LLM) agents constitute a novel attack surface: even when provided with clean inputs, maliciously injected memories can induce erroneous outputs. The study introduces an LLM agent architecture with controllable external memory and systematically investigates, for the first time, how memory manipulation affects multiple-choice question-answering behavior. Through carefully designed memory injection and evaluation experiments, quantitative analysis reveals that only a small number of misleading memory entries significantly degrade answer accuracy and effectively steer the agent toward selecting predetermined incorrect options. These findings establish that the memory component poses a tangible security risk, highlighting the need for robust safeguards against adversarial memory tampering in LLM-based systems.

0 citationsRead paper

Classification of Single and Mixed Partial Discharges under Switching Voltage Using an AWA-CNN Framework

May 20, 2026

This study addresses the challenge of partial discharge (PD) identification under switching voltage excitation, where discharges concentrate around voltage transitions and are significantly harder to classify than under sinusoidal conditions. To tackle this, the authors propose an amplitude–width–area (AWA) visualization method that maps time-domain pulse features into a two-dimensional image: amplitude and area serve as spatial coordinates, while pulse width is encoded by color, effectively revealing distinct distribution patterns of different PD sources. Leveraging this representation, convolutional neural networks—including InceptionV3 and ResNet-18—are employed to achieve high-accuracy classification of six single- and mixed-source PD types. Experimental results demonstrate a classification accuracy exceeding 96%, substantially outperforming a Random Forest baseline (73.33%), thereby validating the effectiveness and superiority of the proposed approach for multi-class PD identification in complex switching voltage environments.

0 citationsRead paper

Adversarial Reframing: A Framework for Targeted Generation in Language Models

May 20, 2026

Large language models are vulnerable to prompt-based attacks (jailbreaking) that circumvent safety mechanisms and elicit harmful content. This work proposes the THREAT framework, which formalizes adversarial prompt generation as a non-convex optimization problem for the first time. By integrating multi-agent collaborative reasoning, iterative adversarial search, and language model redirection techniques, THREAT efficiently produces highly stealthy jailbreaking prompts. Experimental results demonstrate that the method significantly outperforms existing attack strategies across multiple models and datasets, achieving higher attack success rates with lower computational overhead. Notably, fewer than 1% of the generated prompts are flagged as harmful—despite an original refusal rate of approximately 50%—thereby exposing previously undetected security vulnerabilities in current alignment approaches.

0 citationsRead paper