Institution profile

Waseda University

Academic institutionasia · jp
Official website
Research library381linked papers
Opportunities0open roles
Selected work

Representative Papers

BoostDream: Efficient Refining for High-Quality Text-to-3D Generation from Multi-View Diffusion

Jan 30, 2024International Joint Conference on Artificial Intelligence

To address the longstanding trade-off between quality and efficiency in text-to-3D generation, this paper proposes a plug-and-play, efficient 3D refinement framework that elevates coarse, feedforward-generated 3D assets to high-fidelity levels within seconds. Methodologically, we introduce the first 3D model distillation mechanism, design a multi-view-aware Score Distillation Sampling (SDS) loss, and incorporate joint guidance from normal maps and text prompts—thereby overcoming the “Janus dilemma” of SDS, where geometric accuracy and rendering speed are conventionally at odds. The framework supports diverse differentiable 3D representations—including NeRF and Gaussian Splatting—without requiring retraining. Extensive experiments demonstrate consistent superiority over state-of-the-art baselines across geometric completeness, texture realism, and inference speed, achieving synergistic improvements in both quality and efficiency.

9 citationsRead paper

Development and Validation of Engagement and Rapport Scales for Evaluating User Experience in Multimodal Dialogue Systems

May 20, 2025

Existing multimodal foreign language learning dialogue systems lack validated instruments for assessing user experience. Method: This study develops and validates a dual-dimensional scale measuring user engagement and rapport, integrating theories from educational psychology, social psychology, and second language acquisition to uniquely distinguish human tutor versus AI agent experiences. Rigorous psychometric evaluation—including Cronbach’s α analysis, confirmatory factor analysis (CFA), and human–AI comparative experiments—was conducted. Contribution/Results: The scale demonstrates excellent reliability (α > 0.90) and construct validity (CFI > 0.95). Empirical findings reveal systematic differences between human and AI interactions across subdimensions—including task focus, affective responsiveness, and trust formation—highlighting critical design implications. The validated instrument provides a reusable theoretical framework and measurement benchmark for iterative UX optimization and evaluation of multimodal educational dialogue systems.

3 citationsRead paper

OpenDecoder: Open Large Language Model Decoding to Incorporate Document Quality in RAG

Jan 13, 2026

This work proposes OpenDecoder, a novel framework that addresses the challenge of inconsistent retrieval quality in Retrieval-Augmented Generation (RAG) systems, which often undermines answer accuracy. OpenDecoder is the first approach to explicitly integrate multi-dimensional document quality signals—including relevance scores, ranking positions, and query performance prediction metrics—directly into the decoding process of large language models (LLMs), enabling quality-aware generation control. The framework is highly flexible, allowing seamless incorporation of arbitrary external quality indicators and compatibility with various LLM post-training objectives. Extensive experiments across five benchmark datasets demonstrate that OpenDecoder significantly outperforms existing baselines, substantially enhancing the robustness of RAG systems against noisy contexts and improving the reliability of generated responses.

2 citationsRead paper

Lessons Learned from the URGENT 2024 Speech Enhancement Challenge

Jun 02, 2025

This paper addresses long-overlooked bottlenecks in speech enhancement (SE): (1) bandwidth mismatch and implicit label noise in training corpora; (2) insufficient robustness under extreme conditions (e.g., speaker overlap, high noise/reverberation) and lack of quantifiable metrics for hard samples; and (3) poor correlation between single objective metrics and subjective perceptual quality. We propose a data quality diagnostic framework with bandwidth consistency verification, revealing—for the first time—systematic effective bandwidth deviations and >15% label noise across mainstream SE corpora. Furthermore, we introduce a difficulty-aware, multi-metric fusion evaluation framework that integrates objective measures with MOS-mapped weighted aggregation. Experiments demonstrate a 32% improvement in Pearson correlation (r) between automatic assessment and human judgments, significantly enhancing the reliability and interpretability of SE system development.

1 citations1 influentialRead paper

Variable Splitting Binary Tree Models Based on Bayesian Context Tree Models for Time Series Segmentation

Jan 22, 2026

This work proposes a variational segmented binary tree model based on Bayesian context trees to address the challenges of flexibly modeling changepoint locations and achieving compact tree representations in time series segmentation. The method employs recursive logistic regression to dynamically characterize interval partitions over the time domain, jointly inferring both segmentation positions and tree depth. By integrating local variational approximation with the Context Tree Weighting (CTW) algorithm, the approach enables efficient posterior inference. Experimental results demonstrate that the model effectively recovers segmental structures in synthetic data while significantly enhancing the compactness and generalization capability of the resulting tree representation.

1 citationsRead paper
Recent publications

Latest Papers