Institution profile

National Institute of Informatics

Academic institutionasia · jp
Official website
Research library472linked papers
Opportunities0open roles
Selected work

Representative Papers

Differentiable Rule Induction from Raw Sequence Inputs

Feb 14, 2026International Conference on Learning Representations

Existing differentiable inductive logic programming (ILP) approaches struggle to learn symbolic rules directly from raw continuous data—such as time-series or images—primarily due to the explicit label leakage problem: without supervision from feature-level labels, they cannot reliably map continuous inputs to symbolic variables. This work proposes an end-to-end neuro-symbolic framework that integrates self-supervised differentiable clustering with a novel differentiable ILP formulation, enabling direct learning of interpretable symbolic rules from raw data without requiring explicit labels. By circumventing the label leakage bottleneck for the first time, the method preserves rule interpretability while substantially improving generalization and applicability. Experiments on both temporal and visual tasks demonstrate its ability to discover accurate and intuitively meaningful symbolic rules.

2 citationsRead paper

UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation

Apr 29, 2025

This study addresses the lack of universality in large language model (LLM) detoxification by proposing the first unified detoxification framework that requires no model-specific hyperparameter tuning. Methodologically, it employs contrastive decoding to generate high-quality synthetic detoxified data, followed by instruction fine-tuning to achieve cross-architecture generalization across GPT-2, OPT, Falcon, and LLaMA-2. Key contributions include: (1) the first general-purpose data distillation paradigm tailored for detoxification; (2) the first demonstration of a single hyperparameter configuration effective across multiple generations and architectures; and (3) the discovery of an intrinsic connection between detoxification and political bias mitigation. Experiments show an average 38.7% reduction in toxicity across multiple benchmarks, with negligible impact on language modeling capability (perplexity increases only by 1.2%). Notably, detoxified data distilled from GPT-2 transfers effectively to larger models such as LLaMA-2.

2 citationsRead paper

Post-Quantum Cryptography-Based Bidirectional Authentication Key Exchange Protocol and Industry Applications: A Case Study of Instant Messaging

Apr 09, 2026

This work addresses the dual requirements of mutual authentication and key agreement in post-quantum secure environments, particularly for applications such as instant messaging. The authors propose a mutual authentication key exchange protocol based on ML-KEM, integrating post-quantum digital signatures (PQC-DSA) with a key encapsulation mechanism (KEM). To unify PQC public keys and enable efficient bidirectional authentication and key negotiation, they introduce three novel dual-use X.509 certificate types—composite, catalytic, and chameleon. Experimental evaluation demonstrates that the proposed scheme achieves practical performance while significantly reducing communication overhead. Furthermore, its post-quantum security and deployment feasibility are validated in real-world instant messaging scenarios.

1 citationsRead paper

Even GPT-5.2 Can't Count to Five: The Case for Zero-Error Horizons in Trustworthy LLMs

Jan 22, 2026

This work addresses the persistent challenge that large language models (LLMs) still commit errors even on fundamental tasks, lacking reliable zero-error performance guarantees and thereby limiting their deployment in safety-critical applications. To this end, the paper introduces the “Zero-Error Horizon” (ZEH)—a metric quantifying the maximum problem scope an LLM can handle without committing any errors. The authors develop a framework combining systematic zero-error testing, tree-structured accelerated search, and online Softmax optimization to substantially reduce the computational cost of ZEH evaluation. Experiments reveal that state-of-the-art models such as GPT-5.2 still fail on ostensibly simple tasks like short-string parity and bracket balancing. While ZEH correlates with overall accuracy, it provides finer-grained insight into model capabilities and achieves up to a tenfold improvement in evaluation efficiency.

1 citationsRead paper

The Third VoicePrivacy Challenge: Preserving Emotional Expressiveness and Linguistic Content in Voice Anonymization

Jan 17, 2026

This work addresses the challenge of voice anonymization by preserving linguistic content and emotional expression while concealing speaker identity. It introduces the first systematic evaluation framework that explicitly incorporates emotional fidelity as a core assessment dimension, establishing a multi-objective optimization paradigm that jointly optimizes privacy protection, semantic preservation, and emotional consistency. By integrating techniques such as speaker embedding perturbation, voice conversion, and generative modeling, and by introducing objective metrics based on adversarial attack models, the framework enables comprehensive evaluation of various baseline and submitted anonymization systems. Experimental results demonstrate that the proposed approach effectively balances privacy guarantees with speech utility, offering a new benchmark and guiding direction for future research in voice privacy.

1 citationsRead paper
Recent publications

Latest Papers