Institution profile

DePaul University

Academic institutionnorthamerica · us
Official website
Research library23linked papers
Opportunities0open roles
Selected work

Representative Papers

PAS-QFL: Personalized Ansatz Selection for Quantum Federated Learning under Client Data Heterogeneity

Aug 14, 2026

This study addresses the performance instability and unfairness caused by data heterogeneity and class imbalance in Quantum Federated Learning (QFL) by proposing a personalized ansatz selection framework. Unlike traditional parameter-tuning approaches, this method decomposes quantum neural networks into globally shared and client-private ansätze. A stability criterion ensures reliable aggregation, while private decision heads are customized via macro-F1 optimization to adapt to local data distributions, achieving structured personalization by uploading only shared parameters. Experimental results demonstrate that this framework significantly outperforms fixed-ansatz baselines in macro-F1 across heterogeneous tasks. These findings validate the effectiveness of structural personalization in enhancing QFL performance, overcoming the limitations of conventional methods that rely solely on hyperparameter adjustment.

0 citationsRead paper

Algebraic Decomposition Theory for Transformer Length Generalization

Aug 13, 2026

This work addresses the lack of a precise characterization of Transformers’ length generalization capability on regular languages. By extending classical finite semigroup decomposition theory to the additive group of integers, the authors construct an algebraic framework based on the C-RASP formal system, thereby providing the first complete characterization of the class of regular languages amenable to length generalization by Transformers. This approach overcomes limitations of Krohn–Rhodes theory through the introduction of iterated cyclic decompositions and a polynomial-time decision algorithm. The resulting method not only efficiently determines whether any given regular language supports length generalization but also demonstrates significantly higher prediction accuracy than existing classification approaches in empirical evaluations.

0 citationsRead paper

Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI

Aug 06, 2026

This study addresses the systemic failures of automatic speech recognition (ASR) systems on low-resource, Indigenous, and non-standard language varieties, which reproduce colonial linguistic hierarchies. Integrating theories of linguistic capital, raciolinguistic ideologies, and decolonial computing, the work proposes a participatory framework that positions affected communities as co-designers, evaluators, and governance partners. It introduces an innovative “triple harm” taxonomy—comprising misrecognition, misalignment, and mistrust—and a seven-layer contextual model of linguistic diversity, embedding critical language policy perspectives directly into ASR design and evaluation. The resulting culturally competent ASR framework, accompanied by a minimal auditing protocol, significantly enhances the visibility, representation, and agency of marginalized language communities in speech technologies.

0 citationsRead paper

ML in a Box: Analyzing Containerization Practices in Open Source ML Projects

Jul 11, 2026

This study addresses the limited understanding of containerization practices in machine learning (ML) projects, particularly the lack of systematic investigation into how iterative ML workflows affect Docker image size, build performance, and caching behavior. Through a large-scale empirical analysis of Dockerfiles from 1,993 open-source ML projects—integrating static parsing, build log tracing, cache behavior monitoring, and semantic mining of code commits—the work reveals ML-specific container usage patterns and proposes seven ML-tailored Dockerfile refactoring strategies. The findings show that ML images average 10.27 GB in size and require 8.84 minutes to build; 44.4% of commits trigger rebuilds, with 96.4% caused by context changes and 71% exhibiting computational redundancy. The proposed methods substantially reduce image size and improve build efficiency.

0 citationsRead paper
Recent publications

Latest Papers

PAS-QFL: Personalized Ansatz Selection for Quantum Federated Learning under Client Data Heterogeneity

Aug 14, 2026

This study addresses the performance instability and unfairness caused by data heterogeneity and class imbalance in Quantum Federated Learning (QFL) by proposing a personalized ansatz selection framework. Unlike traditional parameter-tuning approaches, this method decomposes quantum neural networks into globally shared and client-private ansätze. A stability criterion ensures reliable aggregation, while private decision heads are customized via macro-F1 optimization to adapt to local data distributions, achieving structured personalization by uploading only shared parameters. Experimental results demonstrate that this framework significantly outperforms fixed-ansatz baselines in macro-F1 across heterogeneous tasks. These findings validate the effectiveness of structural personalization in enhancing QFL performance, overcoming the limitations of conventional methods that rely solely on hyperparameter adjustment.

0 citationsRead paper

Algebraic Decomposition Theory for Transformer Length Generalization

Aug 13, 2026

This work addresses the lack of a precise characterization of Transformers’ length generalization capability on regular languages. By extending classical finite semigroup decomposition theory to the additive group of integers, the authors construct an algebraic framework based on the C-RASP formal system, thereby providing the first complete characterization of the class of regular languages amenable to length generalization by Transformers. This approach overcomes limitations of Krohn–Rhodes theory through the introduction of iterated cyclic decompositions and a polynomial-time decision algorithm. The resulting method not only efficiently determines whether any given regular language supports length generalization but also demonstrates significantly higher prediction accuracy than existing classification approaches in empirical evaluations.

0 citationsRead paper

Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI

Aug 06, 2026

This study addresses the systemic failures of automatic speech recognition (ASR) systems on low-resource, Indigenous, and non-standard language varieties, which reproduce colonial linguistic hierarchies. Integrating theories of linguistic capital, raciolinguistic ideologies, and decolonial computing, the work proposes a participatory framework that positions affected communities as co-designers, evaluators, and governance partners. It introduces an innovative “triple harm” taxonomy—comprising misrecognition, misalignment, and mistrust—and a seven-layer contextual model of linguistic diversity, embedding critical language policy perspectives directly into ASR design and evaluation. The resulting culturally competent ASR framework, accompanied by a minimal auditing protocol, significantly enhances the visibility, representation, and agency of marginalized language communities in speech technologies.

0 citationsRead paper

ML in a Box: Analyzing Containerization Practices in Open Source ML Projects

Jul 11, 2026

This study addresses the limited understanding of containerization practices in machine learning (ML) projects, particularly the lack of systematic investigation into how iterative ML workflows affect Docker image size, build performance, and caching behavior. Through a large-scale empirical analysis of Dockerfiles from 1,993 open-source ML projects—integrating static parsing, build log tracing, cache behavior monitoring, and semantic mining of code commits—the work reveals ML-specific container usage patterns and proposes seven ML-tailored Dockerfile refactoring strategies. The findings show that ML images average 10.27 GB in size and require 8.84 minutes to build; 44.4% of commits trigger rebuilds, with 96.4% caused by context changes and 71% exhibiting computational redundancy. The proposed methods substantially reduce image size and improve build efficiency.

0 citationsRead paper