Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment

📅 2026-09-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过分析金融情绪工具与人类标签的一致性,发现两者在不同时间框架下的预测有效性差异,从而挑战了传统评估方法的有效性。
📝 Abstract
Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where both can be measured at once: a corpus of securities class actions (2002-2025) linking 70,500 X messages to abnormal stock returns, with a single-annotator human labelled gold sample. Running five instruments (VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator) through one identical pipeline, we find that the relationship between construct and predictive validity depends on the sampling convention and score representation. Under conventional method-specific sampling, human agreement aligns more closely with graded same-day associations than with one-day leads. On a fixed-n panel, however, agreement has similar graded rank correlations at both horizons, while the coarse ordering remains weak. Benchmark agreement therefore establishes semantic validity but does not by itself determine predictive rankings. In a conversation that is 17.6% spam, message volume predicts neither market damage nor settlement size.
Problem

Research questions and friction points this paper is trying to address.

Financial NLP
sentiment analysis
predictive validity
construct validity
securities class actions
Innovation

Methods, ideas, or system contributions that make the work stand out.

construct validity
predictive validity
financial NLP
sentiment analysis
sampling convention
🔎 Similar Papers
No similar papers found.