Institution profile

University of Alberta

Academic institutionnorthamerica · ca
Official website
Research library642linked papers
Opportunities0open roles
Selected work

Representative Papers

Clustering-based anomaly detection in multivariate time series data

Mar 01, 2021Applied Soft Computing

Addressing the challenge of unsupervised anomaly detection in multivariate time series—where both temporal dynamics and inter-variable dependencies must be jointly captured—this paper proposes a temporally aware unsupervised clustering framework. Methodologically, it integrates sliding-window segmentation, dynamic time warping (DTW)-driven time-series clustering, and autoencoder-based feature learning, jointly optimizing reconstruction error and cluster-based outlier scores to simultaneously discriminate anomalies in magnitude and shape. Its key innovation lies in embedding temporal structural priors directly into the clustering process, substantially enhancing modeling capacity for complex cross-variable dependencies and nonlinear temporal evolution. Extensive experiments on multiple benchmark datasets demonstrate that the method significantly outperforms state-of-the-art unsupervised and semi-supervised baselines, achieving superior detection accuracy, robustness to noise and distribution shifts, and inherent interpretability through interpretable cluster assignments and reconstruction residuals.

195 citations1 influentialRead paper

Multivariate time series anomaly detection: A framework of Hidden Markov Models

Nov 01, 2017Applied Soft Computing

This paper addresses the challenge of multivariate time-series anomaly detection by proposing a unified probabilistic framework based on Hidden Markov Models (HMMs). Unlike conventional univariate approaches, the framework explicitly models both the dynamics of state transitions and cross-dimensional dependencies among variables, jointly learning normal behavioral patterns via a probabilistic graphical structure. Parameters are efficiently estimated using the Expectation-Maximization (EM) algorithm, and anomalies are scored via a likelihood-ratio-based mechanism. Extensive experiments on benchmark multivariate time-series datasets—including SMD, MSL, and SMAP—demonstrate that the method achieves an average 12.6% improvement in F1-score over strong baselines such as Isolation Forest, LSTM-VAE, and DeepAR. The approach delivers high accuracy, robustness to noise and distribution shifts, and inherent interpretability through its probabilistic formulation and explicit state-transition modeling.

128 citations1 influentialRead paper

Application-Driven Innovation in Machine Learning

Mar 26, 2024International Conference on Machine Learning

Application-driven machine learning (ML) research has been systematically undervalued in academia, leading to a growing disconnect between algorithmic innovation and real-world needs; this marginalization is reinforced by structural biases in peer review, faculty hiring, and pedagogy. Method: This paper introduces, for the first time, a formally defined “application-driven ML research paradigm,” elucidating its complementary relationship with the dominant methodology-driven paradigm. Drawing on interdisciplinary frameworks from education theory, research governance, and ML practice—and substantiated by empirical case studies and institutional critique—it diagnoses three systemic barriers hindering such research. Contribution/Results: The core contribution is a set of actionable, process-level interventions to reform academic evaluation systems, grounded in both theoretical analysis and pragmatic implementation pathways. These proposals have already catalyzed curricular reforms in AI education and adjustments to national funding review criteria across multiple universities, fostering cross-domain collaboration and methodological feedback loops between application domains and core ML research.

22 citations2 influentialRead paper

Towards Understanding Retrieval Accuracy and Prompt Quality in RAG Systems

Nov 29, 2024arXiv.org

The impact of key design decisions—RAG activation, retrieval granularity, and knowledge integration strategy—on RAG system performance remains poorly understood. Method: We conduct systematic ablation studies across three code/qa benchmarks and two state-of-the-art LLMs, quantitatively evaluating how document type, recall rate, document selection strategy, and prompt engineering jointly affect answer correctness and confidence via multi-dimensional analysis, cross-model/dataset comparison, and joint prompt-retrieval analysis. Contribution/Results: We identify precise interaction patterns and operational boundaries among these factors and propose nine actionable, empirically grounded guidelines for diagnosing and optimizing RAG failures. Our findings significantly improve RAG system stability, debuggability, and reliability, offering rigorous empirical evidence and a principled methodology to support the engineering deployment of LLM-augmented systems.

7 citationsRead paper

An Integrated Fusion Framework for Ensemble Learning Leveraging Gradient-Boosting and Fuzzy Rule-Based Models

Nov 01, 2024IEEE Transactions on Artificial Intelligence

Fuzzy rule models offer strong interpretability but suffer from poor scalability and susceptibility to overfitting in complex tasks and large-scale data scenarios. To address these limitations, this paper proposes a novel ensemble framework integrating gradient boosting with fuzzy rule-based base learners. We introduce a dynamic control factor that adaptively adjusts the weights of fuzzy base models in each boosting iteration, simultaneously serving as a regularizer and performance optimizer. Additionally, we design a validation-set-driven, sample-level correction mechanism to enhance generalization and ensemble diversity. Experimental results demonstrate that our approach significantly mitigates overfitting, reduces rule complexity (e.g., fewer rules and shorter antecedents), and preserves high model interpretability and maintainability. The method thus provides a practical pathway for deploying interpretable AI in complex industrial applications.

6 citationsRead paper
Recent publications

Latest Papers