Institution profile

National Tsing Hua University

Academic institutionasia · tw
Official website
Research library193linked papers
Opportunities0open roles
Selected work

Representative Papers

Energy-Efficient Prediction in Textile Manufacturing: Enhancing Accuracy and Data Efficiency With Ensemble Deep Transfer Learning

Jan 19, 2026IEEE Access

This study addresses the challenge of high energy consumption in traditional textile manufacturing and the limited applicability of deep neural networks (DNNs) in production output prediction due to data scarcity caused by the high cost of sensor deployment. To overcome this, the authors propose an Ensemble Deep Transfer Learning (EDTL) framework that uniquely integrates ensemble learning with transfer learning. EDTL leverages models pretrained on data-rich production lines and incorporates a feature alignment layer to enhance cross-line generalization, enabling effective knowledge transfer to data-scarce lines. Evaluated on a real-world textile factory dataset, EDTL achieves a 5.66% improvement in prediction accuracy and a 3.96% gain in robustness compared to conventional DNNs when only 20%–40% of training data is available, significantly enhancing both data efficiency and model performance.

2 citationsRead paper

Incorporating Eye-Tracking Signals Into Multimodal Deep Visual Models For Predicting User Aesthetic Experience In Residential Interiors

Jan 23, 2026

Predicting users’ subjective aesthetic experiences of residential interior design is highly challenging due to its strong individual variability and reliance on complex visual perception. This work proposes a dual-branch CNN-LSTM framework that, for the first time, incorporates eye-tracking signals—including gaze fixations and pupillary responses—as privileged information, enabling end-to-end multimodal fusion with visual features extracted from interior design videos. Notably, the model maintains strong performance even when deployed using only visual inputs at inference time, achieving accuracies of 72.2% and 66.8% on objective dimensions (e.g., lighting) and subjective dimensions (e.g., perceived relaxation), respectively—significantly outperforming existing video-based baselines. Ablation studies further reveal that pupillary responses contribute most substantially to the evaluation of objective attributes.

1 citationsRead paper

uLayout: Unified Room Layout Estimation for Perspective and Panoramic Images

Mar 27, 2025

This work addresses the challenge of separately modeling perspective and panoramic images for indoor layout geometry estimation. To this end, we propose the first end-to-end unified framework. Our method projects both image types into a common equirectangular space and introduces a modulated shared CNN backbone, complemented by latitude-adaptive feature allocation and 1D convolutional domain conditioning—enabling field-of-view-invariant feature extraction and column-wise layout regression. Evaluated on real-world benchmarks including LSUN and Matterport3D, our approach achieves state-of-the-art performance in geometric layout estimation while significantly improving cross-modal generalization. The source code is publicly available.

1 citationsRead paper

TIPO: Text to Image with Text Presampling for Prompt Optimization

Nov 12, 2024arXiv.org

To address the manual dependency, high computational cost, and poor scalability of prompt engineering in text-to-image generation, this paper proposes a lightweight, distribution-aware prompt optimization framework. Methodologically, it abandons large language models and reinforcement learning, introducing the novel “prompt pre-sampling” paradigm: modeling the statistical distribution of training-set prompts and performing differentiable reparameterized sampling and guided optimization in the text embedding space. The approach incurs negligible overhead during inference while enabling end-to-end prompt enhancement under semantic fidelity constraints. Experiments demonstrate substantial improvements: +12.3% in aesthetic score, −31.7% in distortion rate, and enhanced alignment between generated images and target data distributions. The method achieves superior efficiency, scalability, and generalization across diverse prompts and models, without requiring architectural modifications or additional training data.

1 citationsRead paper
Recent publications

Latest Papers