Institution profile

University of Rochester

Academic institutionnorthamerica · us
Official website
Research library391linked papers
Opportunities0open roles
Selected work

Representative Papers

Enhancing Financial Report Question-Answering: A Retrieval-Augmented Generation System with Reranking Analysis

Feb 18, 2026

Financial analysts face significant challenges extracting information from lengthy 10-K reports, which often exceed 100 pages. This paper presents a Retrieval-Augmented Generation (RAG) system designed to answer questions about S&P 500 financial reports and evaluates the impact of neural reranking on system performance. Our pipeline employs hybrid search combining full-text and semantic retrieval, followed by an optional reranking stage using a cross-encoder model. We conduct systematic evaluation using the FinDER benchmark dataset, comprising 1,500 queries across five experimental groups. Results demonstrate that reranking significantly improves answer quality, achieving 49.0 percent correctness for scores of 8 or above compared to 33.5 percent without reranking, representing a 15.5 percentage point improvement. Additionally, the error rate for completely incorrect answers decreases from 35.3 percent to 22.5 percent. Our findings emphasize the critical role of reranking in financial RAG systems and demonstrate performance improvements over baseline methods through modern language models and refined retrieval strategies.

7 citationsRead paper

Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions

Sep 25, 2024arXiv.org

Existing emotional TTS systems are constrained by discrete emotion labels and sparse annotation, limiting their ability to capture the continuity and complexity of human affect. This paper proposes the first method to seamlessly integrate the psychological PAD (Pleasure-Arousal-Dominance) three-dimensional emotion model into a language-model-driven TTS framework—enabling unsupervised disentanglement and learning of continuous emotional styles directly from expressive speech, without requiring explicit emotion labels. Key innovations include: (1) a classification-based emotion dimension predictor trained on labeled speech data, and (2) an end-to-end LM-TTS architecture jointly modeling linguistic and psychometric representations. Experiments demonstrate that our approach significantly improves emotional naturalness and spectral coverage of synthesized speech under zero-shot emotion-label supervision. Both objective metrics (e.g., F0 variance, spectral contrast) and subjective MOS scores surpass those of state-of-the-art baselines.

4 citations1 influentialRead paper

Diffusion Index Forecast with Tensor Data

Nov 04, 2025

This paper addresses diffusion index forecasting with both tensor and non-tensor predictors, proposing a factor-augmented regression framework that preserves the intrinsic tensor structure. Methodologically: (1) it constructs a latent factor model via CP decomposition to explicitly capture the multilinear structure of high-dimensional tensor data; (2) it derives asymptotically efficient prediction intervals accounting for estimation uncertainty in latent factors; and (3) it designs a cross-sectionally robust thresholded covariance estimator, integrated with multi-source sparse penalized regression to tackle high-dimensional variable selection under small-sample settings. Theoretical analysis establishes consistency and robustness, corroborated by extensive simulations. In an empirical application to U.S. trade flow data, the method significantly outperforms conventional factor models and LASSO-based approaches, delivering superior forecasting accuracy and statistical reliability for structured heterogeneous data.

2 citationsRead paper

SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks

Feb 06, 2026

Existing multi-turn jailbreaking attacks suffer from high exploration complexity and intent drift, making it difficult to generate coherent and goal-consistent adversarial dialogues. This work proposes SEMA, a novel framework that achieves the first open-loop multi-turn jailbreak attack without requiring feedback from the victim model. SEMA employs self-generated multi-turn adversarial prompts for supervised fine-tuning prefilling and integrates a reinforcement learning reward mechanism based on intent alignment, compliance risk, and level of detail, unifying single-turn and multi-turn attack formulations. Evaluated on benchmarks such as AdvBench, SEMA attains an average ASR@1 of 80.1%, surpassing current single-turn and multi-turn baselines—as well as SFT and DPO variants—by over 33.9%. The approach significantly reduces exploration complexity while preserving harmful intent consistency, demonstrating strong generalization and reproducibility.

1 citations1 influentialRead paper

PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement

Dec 02, 2025

Latent diffusion models (LDMs) suffer from pixel-level inconsistencies—including color shifts, texture mismatches, and boundary seams—in local image editing due to aggressive latent-space compression. Existing approaches struggle to balance generality and fidelity. To address this, we propose a universal, pixel-level refinement framework comprising three key components: (1) a differentiable discriminative pixel-space modeling module that explicitly enforces fine-grained detail consistency; (2) a background-aware latent decoding mechanism coupled with adversarial artifact simulation during training to suppress structural artifacts; and (3) a plug-and-play direct pixel-space optimization module compatible with diverse implicit representations and editing tasks. Extensive experiments on image inpainting, object removal, and object insertion demonstrate substantial improvements in visual fidelity. Our method achieves state-of-the-art performance across multiple benchmarks while exhibiting strong generalization and practical applicability.

1 citationsRead paper
Recent publications

Latest Papers