Institution profile

Polish-Japanese Academy of Information Technology

Academic institutioneurope · pl
Official website
Research library15linked papers
Opportunities0open roles
Selected work

Representative Papers

Entropy of Ukrainian

Apr 30, 2026

This study addresses a gap in information-theoretic research on Ukrainian by applying Shannon’s 1951 human prediction experiment methodology to this underexplored language. Recruiting 184 volunteers via crowdsourcing, the authors conducted an online character-level prediction task to estimate an upper bound on the language’s entropy. The experiment yielded an entropy upper bound of approximately 1.201 bits per character. The work presents the first human-prediction-based estimate of information entropy for Ukrainian and provides a fully reproducible framework, including open-sourced experimental protocols, code, and detailed challenge analyses. Furthermore, the results enable direct comparison with large language model performance, establishing a replicable paradigm for information-theoretic investigations of low-resource languages.

0 citationsRead paper

Element-wise Modulation of Random Matrices for Efficient Neural Layers

Dec 15, 2025

Fully connected (FC) layers suffer from high memory and computational overhead due to parameter redundancy. Existing compression methods often sacrifice accuracy or introduce engineering complexity. This paper proposes the Parameterized Random Projection (PRP) layer: it employs a fixed random projection matrix for efficient feature mixing, augmented by lightweight, element-wise learnable scaling and bias modulation—thereby decoupling linear transformation from adaptive modeling. PRP is the first approach to unify fixed random projection with fine-grained, element-level modulation, reducing parameter complexity to linear while preserving strong generalization. Theoretical analysis, supported by low-rank approximation modeling, substantiates its efficacy. Experiments across multiple benchmarks demonstrate zero accuracy degradation, significant inference speedup, and substantial memory reduction—enabling deployment on resource-constrained edge devices.

0 citationsRead paper

Rough Sets for Explainability of Spectral Graph Clustering

Dec 13, 2025

Graph spectral clustering (GSC) applied to text suffers from poor interpretability due to semantic disconnection between the embedded spectral space and original semantics, interference from noisy documents, and algorithmic randomness. To address this, we propose the first unsupervised, content-aware interpretability enhancement framework for GSC grounded in rough set theory. Our method first obtains cluster structure via graph Laplacian eigendecomposition and k-means spectral clustering; it then constructs upper and lower approximations under term-frequency–guided semantic constraints to quantify cluster semantic stability and structural uncertainty. Evaluated on multiple text datasets, our approach improves explanation fidelity by 32% and user comprehension consistency by 27%, without compromising clustering quality. The core innovation lies in the deep integration of rough set boundary region analysis with spectral clustering—enabling semantically traceable and quantifiably stable interpretable clustering.

0 citationsRead paper

Copyright in AI Pre-Training Data Filtering: Regulatory Landscape and Mitigation Strategies

Nov 26, 2025

Widespread reliance on unlicensed web-scale data for AI pretraining poses escalating global copyright infringement risks, yet current regulatory approaches remain predominantly reactive, lacking proactive, pre-training compliance mechanisms. Method: We systematically analyze copyright governance frameworks across the EU, U.S., and major Asia-Pacific jurisdictions, identifying structural deficiencies in licensing acquisition, content filtering, and enforcement oversight. To address delayed risk detection and incomplete technical coverage during pretraining, we propose a “Proactive Multi-Layer Filtering Framework” integrating access control, perceptual hashing, ML-based classifiers, dynamic database matching, and transparency tools. Contribution/Results: The framework enables end-to-end identification, blocking, and verifiable mitigation of copyright risks in training data pipelines. It is the first to embed copyright compliance intrinsically into the AI training frontend—establishing a technically feasible, governance-aligned pathway that balances creator rights with sustainable AI development.

0 citationsRead paper

Global AI Governance Overview: Understanding Regulatory Requirements Across Global Jurisdictions

Nov 26, 2025

Current global AI training data governance—across the EU, U.S., and Asia-Pacific—relies predominantly on reactive enforcement, lacking proactive copyright filtering during pretraining, thereby undermining creator rights and threatening AI’s long-term sustainability. Addressing two core challenges—difficult license acquisition and unverifiable filter efficacy—the paper proposes a multi-tiered *pre-ingestion* filtering framework integrating access control, perceptual hashing, ML-based classifiers, and real-time cross-referencing against dynamic copyright databases to identify and block high-risk content prior to training. Unlike existing approaches relying solely on transparency tools or post-hoc detection, this framework shifts copyright protection to the earliest data intake stage, ensuring scalability and auditability. Empirical analysis demonstrates its capacity to systematically close regulatory gaps, offering a practical, globally applicable governance paradigm that balances AI innovation with creator rights protection.

0 citationsRead paper
Recent publications

Latest Papers

Entropy of Ukrainian

Apr 30, 2026

This study addresses a gap in information-theoretic research on Ukrainian by applying Shannon’s 1951 human prediction experiment methodology to this underexplored language. Recruiting 184 volunteers via crowdsourcing, the authors conducted an online character-level prediction task to estimate an upper bound on the language’s entropy. The experiment yielded an entropy upper bound of approximately 1.201 bits per character. The work presents the first human-prediction-based estimate of information entropy for Ukrainian and provides a fully reproducible framework, including open-sourced experimental protocols, code, and detailed challenge analyses. Furthermore, the results enable direct comparison with large language model performance, establishing a replicable paradigm for information-theoretic investigations of low-resource languages.

0 citationsRead paper

Element-wise Modulation of Random Matrices for Efficient Neural Layers

Dec 15, 2025

Fully connected (FC) layers suffer from high memory and computational overhead due to parameter redundancy. Existing compression methods often sacrifice accuracy or introduce engineering complexity. This paper proposes the Parameterized Random Projection (PRP) layer: it employs a fixed random projection matrix for efficient feature mixing, augmented by lightweight, element-wise learnable scaling and bias modulation—thereby decoupling linear transformation from adaptive modeling. PRP is the first approach to unify fixed random projection with fine-grained, element-level modulation, reducing parameter complexity to linear while preserving strong generalization. Theoretical analysis, supported by low-rank approximation modeling, substantiates its efficacy. Experiments across multiple benchmarks demonstrate zero accuracy degradation, significant inference speedup, and substantial memory reduction—enabling deployment on resource-constrained edge devices.

0 citationsRead paper

Rough Sets for Explainability of Spectral Graph Clustering

Dec 13, 2025

Graph spectral clustering (GSC) applied to text suffers from poor interpretability due to semantic disconnection between the embedded spectral space and original semantics, interference from noisy documents, and algorithmic randomness. To address this, we propose the first unsupervised, content-aware interpretability enhancement framework for GSC grounded in rough set theory. Our method first obtains cluster structure via graph Laplacian eigendecomposition and k-means spectral clustering; it then constructs upper and lower approximations under term-frequency–guided semantic constraints to quantify cluster semantic stability and structural uncertainty. Evaluated on multiple text datasets, our approach improves explanation fidelity by 32% and user comprehension consistency by 27%, without compromising clustering quality. The core innovation lies in the deep integration of rough set boundary region analysis with spectral clustering—enabling semantically traceable and quantifiably stable interpretable clustering.

0 citationsRead paper

Copyright in AI Pre-Training Data Filtering: Regulatory Landscape and Mitigation Strategies

Nov 26, 2025

Widespread reliance on unlicensed web-scale data for AI pretraining poses escalating global copyright infringement risks, yet current regulatory approaches remain predominantly reactive, lacking proactive, pre-training compliance mechanisms. Method: We systematically analyze copyright governance frameworks across the EU, U.S., and major Asia-Pacific jurisdictions, identifying structural deficiencies in licensing acquisition, content filtering, and enforcement oversight. To address delayed risk detection and incomplete technical coverage during pretraining, we propose a “Proactive Multi-Layer Filtering Framework” integrating access control, perceptual hashing, ML-based classifiers, dynamic database matching, and transparency tools. Contribution/Results: The framework enables end-to-end identification, blocking, and verifiable mitigation of copyright risks in training data pipelines. It is the first to embed copyright compliance intrinsically into the AI training frontend—establishing a technically feasible, governance-aligned pathway that balances creator rights with sustainable AI development.

0 citationsRead paper

Global AI Governance Overview: Understanding Regulatory Requirements Across Global Jurisdictions

Nov 26, 2025

Current global AI training data governance—across the EU, U.S., and Asia-Pacific—relies predominantly on reactive enforcement, lacking proactive copyright filtering during pretraining, thereby undermining creator rights and threatening AI’s long-term sustainability. Addressing two core challenges—difficult license acquisition and unverifiable filter efficacy—the paper proposes a multi-tiered *pre-ingestion* filtering framework integrating access control, perceptual hashing, ML-based classifiers, and real-time cross-referencing against dynamic copyright databases to identify and block high-risk content prior to training. Unlike existing approaches relying solely on transparency tools or post-hoc detection, this framework shifts copyright protection to the earliest data intake stage, ensuring scalability and auditability. Empirical analysis demonstrates its capacity to systematically close regulatory gaps, offering a practical, globally applicable governance paradigm that balances AI innovation with creator rights protection.

0 citationsRead paper