Institution profile

University at Albany, SUNY

Academic institutionnorthamerica · us
Official website
Research library44linked papers
Opportunities0open roles
Selected work

Representative Papers

Supervised Mixed-Frequency Learning for Macro-Financial Forecasting When Factors are Weak

Aug 12, 2026

This study addresses the limitations of traditional MIDAS models, which rely on strong factor assumptions and underperform in macro-financial forecasting when factors are weak. To overcome this, the authors propose the SsPCA-MIDAS model, which integrates supervised scaled principal component analysis (SsPCA) into the mixed-data sampling framework. This approach achieves, for the first time under weak factor conditions, consistent estimation and asymptotic normality, thereby enabling valid statistical inference. Moreover, the model can be combined with machine learning techniques such as Boosting to enhance predictive accuracy. Empirical results demonstrate that SsPCA-MIDAS significantly outperforms existing methods in forecasting key U.S. macroeconomic and financial indicators—including GDP growth, inflation, unemployment, asset prices, and volatility—and successfully identifies economically meaningful predictors.

0 citationsRead paper

A Hybrid LSTM--Vision Transformer Architecture for Predicting HRRR Forecast Errors

Jun 17, 2026

This study addresses significant errors in high-resolution numerical weather prediction models—such as the HRRR—arising from complex boundary layer dynamics and convective processes, which existing methods struggle to characterize due to insufficient representation of vertical atmospheric structure. To overcome this limitation, the authors introduce a novel hybrid architecture that incorporates vertical atmospheric profiles (e.g., radiosonde soundings) into error prediction for the first time. The model combines an LSTM to capture temporal patterns in surface observations with a Vision Transformer that leverages attention mechanisms to represent vertical atmospheric evolution. Trained end-to-end, this approach enables hourly error prediction for precipitation, 10-meter wind speed, and 2-meter temperature. Compared to a pure LSTM baseline, the proposed method yields substantial improvements across all three variables, particularly during short lead times and periods of intense boundary layer activity, nearly doubling skill in precipitation error prediction and markedly enhancing the representation of convection-driven errors.

0 citationsRead paper

Robust LLM Watermarking with Minimal Semantic Distortion for IP Protection

May 21, 2026

Proprietary large language models are vulnerable to intellectual property theft, yet existing watermarking methods often suffer from semantic distortion, factual inaccuracies, and insufficient robustness, while lacking key-conditioned mechanisms that support multi-user and cross-provider scenarios. This work proposes SAFESEAL, a novel framework that integrates key-conditioned tournament sampling with context-aware synonym substitution to embed high-fidelity watermarks while preserving named entities. It further introduces a key-conditioned contrastive detector for owner-specific verification. SAFESEAL achieves, for the first time, a compelling balance of high detection accuracy (98.2%), minimal semantic distortion (BERTScore 0.983; entity similarity 0.963), and strong robustness. Human evaluations rank its output quality as best among competitors, with inference latency matching the fastest baseline. The authors also release the first public watermarking leaderboard and an interactive demo platform.

0 citationsRead paper

Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification

May 17, 2026

This study addresses the challenge of automatically identifying negative citations in legal case law, a task highly susceptible to high-risk misclassifications that conventional accuracy metrics fail to adequately capture. To bridge this gap, the authors construct the first fine-grained, multi-label expert-annotated dataset specifically tailored to this task, comprising 239 real-world legal citations. They further propose “average severity error” as a more legally meaningful evaluation metric aligned with practical judicial concerns. A systematic benchmarking of prominent large language models—including Gemini 2.5 Flash and GPT-5-mini—reveals that Gemini 2.5 Flash achieves 79.1% accuracy in coarse-grained classification, while GPT-5-mini performs best on fine-grained tasks with 67.7% accuracy. These findings highlight nuanced differences in models’ capabilities for complex legal reasoning and establish a reliable baseline for future research in this domain.

0 citationsRead paper
Recent publications

Latest Papers

Supervised Mixed-Frequency Learning for Macro-Financial Forecasting When Factors are Weak

Aug 12, 2026

This study addresses the limitations of traditional MIDAS models, which rely on strong factor assumptions and underperform in macro-financial forecasting when factors are weak. To overcome this, the authors propose the SsPCA-MIDAS model, which integrates supervised scaled principal component analysis (SsPCA) into the mixed-data sampling framework. This approach achieves, for the first time under weak factor conditions, consistent estimation and asymptotic normality, thereby enabling valid statistical inference. Moreover, the model can be combined with machine learning techniques such as Boosting to enhance predictive accuracy. Empirical results demonstrate that SsPCA-MIDAS significantly outperforms existing methods in forecasting key U.S. macroeconomic and financial indicators—including GDP growth, inflation, unemployment, asset prices, and volatility—and successfully identifies economically meaningful predictors.

0 citationsRead paper

A Hybrid LSTM--Vision Transformer Architecture for Predicting HRRR Forecast Errors

Jun 17, 2026

This study addresses significant errors in high-resolution numerical weather prediction models—such as the HRRR—arising from complex boundary layer dynamics and convective processes, which existing methods struggle to characterize due to insufficient representation of vertical atmospheric structure. To overcome this limitation, the authors introduce a novel hybrid architecture that incorporates vertical atmospheric profiles (e.g., radiosonde soundings) into error prediction for the first time. The model combines an LSTM to capture temporal patterns in surface observations with a Vision Transformer that leverages attention mechanisms to represent vertical atmospheric evolution. Trained end-to-end, this approach enables hourly error prediction for precipitation, 10-meter wind speed, and 2-meter temperature. Compared to a pure LSTM baseline, the proposed method yields substantial improvements across all three variables, particularly during short lead times and periods of intense boundary layer activity, nearly doubling skill in precipitation error prediction and markedly enhancing the representation of convection-driven errors.

0 citationsRead paper

Robust LLM Watermarking with Minimal Semantic Distortion for IP Protection

May 21, 2026

Proprietary large language models are vulnerable to intellectual property theft, yet existing watermarking methods often suffer from semantic distortion, factual inaccuracies, and insufficient robustness, while lacking key-conditioned mechanisms that support multi-user and cross-provider scenarios. This work proposes SAFESEAL, a novel framework that integrates key-conditioned tournament sampling with context-aware synonym substitution to embed high-fidelity watermarks while preserving named entities. It further introduces a key-conditioned contrastive detector for owner-specific verification. SAFESEAL achieves, for the first time, a compelling balance of high detection accuracy (98.2%), minimal semantic distortion (BERTScore 0.983; entity similarity 0.963), and strong robustness. Human evaluations rank its output quality as best among competitors, with inference latency matching the fastest baseline. The authors also release the first public watermarking leaderboard and an interactive demo platform.

0 citationsRead paper

Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classification

May 17, 2026

This study addresses the challenge of automatically identifying negative citations in legal case law, a task highly susceptible to high-risk misclassifications that conventional accuracy metrics fail to adequately capture. To bridge this gap, the authors construct the first fine-grained, multi-label expert-annotated dataset specifically tailored to this task, comprising 239 real-world legal citations. They further propose “average severity error” as a more legally meaningful evaluation metric aligned with practical judicial concerns. A systematic benchmarking of prominent large language models—including Gemini 2.5 Flash and GPT-5-mini—reveals that Gemini 2.5 Flash achieves 79.1% accuracy in coarse-grained classification, while GPT-5-mini performs best on fine-grained tasks with 67.7% accuracy. These findings highlight nuanced differences in models’ capabilities for complex legal reasoning and establish a reliable baseline for future research in this domain.

0 citationsRead paper