Institution profile

United States Military Academy

Academic institutionnorthamerica · us
Official website
Research library13linked papers
Opportunities0open roles
Selected work

Representative Papers

CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning

Sep 04, 2026

Aggregating heterogeneous vision-language models (VLMs) can improve multimodal reasoning, but neither an individual model's confidence nor that of the aggregated answer measures reliability at the system level. We present CUSP (Collective Uncertainty through Semantic Opinion Pooling), a training-free uncertainty quantification framework that maps multiple VLM responses to a shared semantic response space, pools them into a pooled semantic opinion, and reports two complementary system-level signals: collective uncertainty, the dispersion of the pooled opinion, and Jensen-Shannon divergence (JSD), the conflict among the model-level opinions. Within this pooled semantic opinion, the unnormalized collective entropy decomposes exactly into the mean of the models'individual semantic entropies and the JSD, separating total dispersion from model conflict. Requiring neither token logits nor calibration labels, CUSP applies to open-weight and commercial VLMs alike. In static multi-VLM ensembles, collective uncertainty is the strongest signal in the small-model regime (0.764 AUROC for prediction-error detection, 0.889 AUARC for abstention), outperforming uncertainty baselines majority voting and naive selection by 4.7 to 15.8 points and widening its margin as the ensemble grows; JSD is strongest in the evaluated commercial regime (0.819 AUROC, 0.910 AUARC) and ranks hard-answer model conflict with AUROC up to 0.982. The pooled prediction also improves accuracy over the average single model by 5.6 to 13.0 points. Over the full trajectory of a multi-step, multi-agent system, subagent collective uncertainty ranks system failures above chance (0.619 AUROC) and gives the best abstention ordering among the evaluated signals (0.699 AUARC).

0 citationsRead paper

A Vision for the Future of an AI-Integrated Research Ecosystem

Aug 05, 2026

This study addresses the challenges posed by generative AI to scholarly publishing—particularly ambiguous authorship, inadequate disclosure, and systemic pressures—by moving beyond conventional individual-level disclosure approaches. It proposes a trustworthy research infrastructure framework centered on provenance, calibration, and accountability. Through policy analysis, multi-stakeholder practice investigations, and dual-scenario foresight exercises projecting to 2036, the work systematically explores future configurations of paper functionality, peer review mechanisms, reviewer roles, and incentive structures. Notably, it shifts the AI governance paradigm from restrictive control toward infrastructural redesign, aiming to embed trustworthiness as the default state of academic communication. The study identifies three critical challenges and offers a forward-looking roadmap for evolving scientific communication paradigms in the age of AI.

0 citationsRead paper

Synthetic Network Packet Generation through Statistical Learning and Genetic Algorithms

Jun 18, 2026

This study addresses the limitations of existing IoT intrusion detection datasets, which commonly suffer from fixed attack categories and extreme class imbalance, as well as the inability of current generative models to guarantee the physical validity of synthesized packets. To overcome these challenges, the paper proposes two novel synthesis approaches that embed hard validity constraints directly into the generation process: a statistical learning method based on PCA and dual anomaly detection boundaries, and a genetic algorithm that formulates data generation as a multi-objective optimization problem. Innovatively integrating dual anomaly gating, feature-range clamping, and an independent validation mechanism, both methods significantly enhance the fidelity and validity of synthetic data. Evaluated on the ACI IoT 2023 dataset, they achieve average anomaly rates of 1.20% (at 1,091 pkt/s) and 0.62% (at 5.7 pkt/s), respectively, and successfully expand ARP spoofing samples to 1,000 instances—a 200-fold increase.

0 citationsRead paper

Categorical Robustness Assessment for Machine Learning based Network Intrusion Detection Systems

Jun 10, 2026

This study addresses the lack of systematic evaluation of model robustness in existing network intrusion detection systems under adversarial attacks. Within a unified framework and using the ACI-IoT-2023 dataset, the authors conduct a cross-architecture and cross-attack robustness comparison of 1D CNN, LSTM, and Random Forest models against FGSM and PGD attacks with perturbation budgets ε ∈ [0.01, 0.1]. The findings reveal that high baseline accuracy does not guarantee strong adversarial robustness: although Random Forest achieves 99.98% accuracy under benign conditions, its performance drops by 73 percentage points under minimal perturbations. In contrast, the 1D CNN demonstrates superior robustness, maintaining 95.5% accuracy at ε = 0.01 with gradual degradation. This work is the first to systematically uncover disparities in adversarial vulnerability among mainstream models in normalized feature spaces, offering critical insights for real-world deployment.

0 citationsRead paper
Recent publications

Latest Papers

CUSP: Decomposable Collective Uncertainty for Multi-Agent Multimodal Reasoning

Sep 04, 2026

Aggregating heterogeneous vision-language models (VLMs) can improve multimodal reasoning, but neither an individual model's confidence nor that of the aggregated answer measures reliability at the system level. We present CUSP (Collective Uncertainty through Semantic Opinion Pooling), a training-free uncertainty quantification framework that maps multiple VLM responses to a shared semantic response space, pools them into a pooled semantic opinion, and reports two complementary system-level signals: collective uncertainty, the dispersion of the pooled opinion, and Jensen-Shannon divergence (JSD), the conflict among the model-level opinions. Within this pooled semantic opinion, the unnormalized collective entropy decomposes exactly into the mean of the models'individual semantic entropies and the JSD, separating total dispersion from model conflict. Requiring neither token logits nor calibration labels, CUSP applies to open-weight and commercial VLMs alike. In static multi-VLM ensembles, collective uncertainty is the strongest signal in the small-model regime (0.764 AUROC for prediction-error detection, 0.889 AUARC for abstention), outperforming uncertainty baselines majority voting and naive selection by 4.7 to 15.8 points and widening its margin as the ensemble grows; JSD is strongest in the evaluated commercial regime (0.819 AUROC, 0.910 AUARC) and ranks hard-answer model conflict with AUROC up to 0.982. The pooled prediction also improves accuracy over the average single model by 5.6 to 13.0 points. Over the full trajectory of a multi-step, multi-agent system, subagent collective uncertainty ranks system failures above chance (0.619 AUROC) and gives the best abstention ordering among the evaluated signals (0.699 AUARC).

0 citationsRead paper

A Vision for the Future of an AI-Integrated Research Ecosystem

Aug 05, 2026

This study addresses the challenges posed by generative AI to scholarly publishing—particularly ambiguous authorship, inadequate disclosure, and systemic pressures—by moving beyond conventional individual-level disclosure approaches. It proposes a trustworthy research infrastructure framework centered on provenance, calibration, and accountability. Through policy analysis, multi-stakeholder practice investigations, and dual-scenario foresight exercises projecting to 2036, the work systematically explores future configurations of paper functionality, peer review mechanisms, reviewer roles, and incentive structures. Notably, it shifts the AI governance paradigm from restrictive control toward infrastructural redesign, aiming to embed trustworthiness as the default state of academic communication. The study identifies three critical challenges and offers a forward-looking roadmap for evolving scientific communication paradigms in the age of AI.

0 citationsRead paper

Synthetic Network Packet Generation through Statistical Learning and Genetic Algorithms

Jun 18, 2026

This study addresses the limitations of existing IoT intrusion detection datasets, which commonly suffer from fixed attack categories and extreme class imbalance, as well as the inability of current generative models to guarantee the physical validity of synthesized packets. To overcome these challenges, the paper proposes two novel synthesis approaches that embed hard validity constraints directly into the generation process: a statistical learning method based on PCA and dual anomaly detection boundaries, and a genetic algorithm that formulates data generation as a multi-objective optimization problem. Innovatively integrating dual anomaly gating, feature-range clamping, and an independent validation mechanism, both methods significantly enhance the fidelity and validity of synthetic data. Evaluated on the ACI IoT 2023 dataset, they achieve average anomaly rates of 1.20% (at 1,091 pkt/s) and 0.62% (at 5.7 pkt/s), respectively, and successfully expand ARP spoofing samples to 1,000 instances—a 200-fold increase.

0 citationsRead paper

Categorical Robustness Assessment for Machine Learning based Network Intrusion Detection Systems

Jun 10, 2026

This study addresses the lack of systematic evaluation of model robustness in existing network intrusion detection systems under adversarial attacks. Within a unified framework and using the ACI-IoT-2023 dataset, the authors conduct a cross-architecture and cross-attack robustness comparison of 1D CNN, LSTM, and Random Forest models against FGSM and PGD attacks with perturbation budgets ε ∈ [0.01, 0.1]. The findings reveal that high baseline accuracy does not guarantee strong adversarial robustness: although Random Forest achieves 99.98% accuracy under benign conditions, its performance drops by 73 percentage points under minimal perturbations. In contrast, the 1D CNN demonstrates superior robustness, maintaining 95.5% accuracy at ε = 0.01 with gradual degradation. This work is the first to systematically uncover disparities in adversarial vulnerability among mainstream models in normalized feature spaces, offering critical insights for real-world deployment.

0 citationsRead paper