Institution profile

Polish Academy of Sciences

Academic institutioneurope · pl
Official website
Research library146linked papers
Opportunities0open roles
Selected work

Representative Papers

Multivariate time series anomaly detection: A framework of Hidden Markov Models

Nov 01, 2017Applied Soft Computing

This paper addresses the challenge of multivariate time-series anomaly detection by proposing a unified probabilistic framework based on Hidden Markov Models (HMMs). Unlike conventional univariate approaches, the framework explicitly models both the dynamics of state transitions and cross-dimensional dependencies among variables, jointly learning normal behavioral patterns via a probabilistic graphical structure. Parameters are efficiently estimated using the Expectation-Maximization (EM) algorithm, and anomalies are scored via a likelihood-ratio-based mechanism. Extensive experiments on benchmark multivariate time-series datasets—including SMD, MSL, and SMAP—demonstrate that the method achieves an average 12.6% improvement in F1-score over strong baselines such as Isolation Forest, LSTM-VAE, and DeepAR. The approach delivers high accuracy, robustness to noise and distribution shifts, and inherent interpretability through its probabilistic formulation and explicit state-transition modeling.

128 citations1 influentialRead paper

PLLuM: A Family of Polish Large Language Models

Nov 05, 2025

To address the scarcity and inadequate cultural adaptation of large language models (LLMs) for non-English languages, the PLLuM project introduces the first open-source, transparent, Polish-language LLM family. Methodologically: (1) it curates a high-quality, 100-billion-token Polish pretraining corpus and a dedicated instruction-following dataset; (2) it employs a Transformer-based architecture integrating pretraining, supervised fine-tuning, and preference alignment; and (3) it incorporates hybrid output correction and multi-layer safety filtering, grounded in a responsible AI governance framework. The primary contributions are: (i) the first publicly released series of open-weight PLLuM models; (ii) state-of-the-art performance on downstream tasks—including public administration—significantly surpassing existing baselines; and (iii) bridging the critical gap in Polish LLMs to advance a sovereign, trustworthy, and culturally grounded open AI ecosystem.

2 citationsRead paper

Quantifying patterns of punctuation in modern Chinese prose.

Feb 01, 2025Chaos

This study investigates the universality and structural complexity of punctuation distribution in modern Chinese prose. Using corpora from three contemporary Chinese novels, we apply statistical modeling, discrete Weibull distribution fitting, Zipf’s law validation, and multifractal analysis. We first demonstrate that inter-punctuation distances for non-terminal punctuation strictly follow a discrete Weibull distribution—supporting cross-linguistic universality of punctuation patterns. In contrast, terminal-punctuation intervals significantly deviate from this distribution, reflecting the high syntactic variability and structural complexity inherent in Chinese sentence architecture. Furthermore, Gao Xingjian’s *Soul Mountain* exhibits pronounced multifractal characteristics, corroborating the role of elevated narrative freedom in shaping hierarchical textual structure. Collectively, these findings provide quantifiable, punctuation-based empirical evidence for syntactic complexity in written Chinese, advancing formal linguistic analysis of discourse-level structure.

1 citationsRead paper

Detecting dependence structure: visualization and inference

Oct 08, 2024

This paper addresses the problem of interpretable detection of dependency structures among random variables. We propose a novel framework integrating a rank-transform-based estimator for the quantile dependence function with a local acceptance region. The method constructs robust quantile dependence measures via rank standardization and employs local hypothesis testing to enable visual diagnostic assessment of dependency patterns and rigorous independence testing under finite samples. Key contributions include: (1) the first nonparametric estimation and theoretical derivation of the quantile dependence function; (2) guaranteed validity of statistical tests at any sample size, balancing high global power with precise localization of heterogeneous dependencies; (3) superior empirical power across diverse alternative models and successful identification of heterogeneous non-independence in real-world data; and (4) a computationally efficient algorithm supporting intuitive, graphical diagnostic interpretation.

1 citationsRead paper
Recent publications

Latest Papers