Empirical Evaluation of Open-Source Large Language Models for Retrieval-Augmented Generation in ESG Domain
本文评估了七种开源大型语言模型在ESG报告中的表现,使用了498份真实报告和100对合成QA进行测试,发现模型在检索方面表现良好,但在生成准确性上需要改进。
本文评估了七种开源大型语言模型在ESG报告中的表现,使用了498份真实报告和100对合成QA进行测试,发现模型在检索方面表现良好,但在生成准确性上需要改进。
本文构建了一个针对ESG领域的基准数据集,评估了14种嵌入模型在检索和检索增强生成任务中的表现,以解决ESG文本处理的有效性问题。
为解决次季节到季节尺度的海洋预测问题,提出Neptune模型,结合CNNs和SFNOs方法,以高时空分辨率模拟全球海洋状态。
This study addresses the challenges in time series anomaly detection posed by extreme class imbalance and scarce labeled data, which hinder supervised approaches and lead to high false positive rates in unsupervised methods. To overcome these limitations, the authors propose an unsupervised detection framework that integrates Haar discrete wavelet transform with a tailored t-test. By decomposing the signal across multiple scales and applying statistically grounded significance testing, the method effectively identifies anomalies without requiring labeled data. This work is the first to synergistically combine Haar wavelets with theoretically justified t-tests, substantially reducing false positives while enhancing detection accuracy. Extensive experiments on 343 real-world datasets demonstrate that the proposed approach outperforms current state-of-the-art unsupervised and self-supervised methods in both detection speed and accuracy.
Data assimilation provides a systematic framework for combining dynamical models with partial and noisy observations to infer the evolving state of a system. In this work, we undertake a comparative study of Data Assimilation with Transfer Operators (DATO) and Quantum Mechanical Data Assimilation (QMDA), focusing on their mathematical formulation, algorithmic structure, and empirical performance. Both methods are first cast within a common operator-theoretic framework, which makes it possible to compare, on a unified basis, their representations of uncertainty, forecast propagation, and assimilation updates. We then analyse their principal similarities and differences with respect to state-space structure, update mechanisms, structural preservation properties, and computational cost. To complement the theoretical analysis, we assess both approaches on benchmark dynamical systems across a range of observational settings, including noisy, sparse, and partially observed regimes. Our results show that, despite their shared operator-theoretic motivation, DATO and QMDA embody substantially different assimilation paradigms, leading to distinct advantages and limitations in terms of interpretability, robustness, and scalability. The present study helps delineate the regimes in which each framework is most effective and offers broader insight into the design of operator-based methodologies for data assimilation.
本文评估了七种开源大型语言模型在ESG报告中的表现,使用了498份真实报告和100对合成QA进行测试,发现模型在检索方面表现良好,但在生成准确性上需要改进。
本文构建了一个针对ESG领域的基准数据集,评估了14种嵌入模型在检索和检索增强生成任务中的表现,以解决ESG文本处理的有效性问题。
为解决次季节到季节尺度的海洋预测问题,提出Neptune模型,结合CNNs和SFNOs方法,以高时空分辨率模拟全球海洋状态。
This study addresses the challenges in time series anomaly detection posed by extreme class imbalance and scarce labeled data, which hinder supervised approaches and lead to high false positive rates in unsupervised methods. To overcome these limitations, the authors propose an unsupervised detection framework that integrates Haar discrete wavelet transform with a tailored t-test. By decomposing the signal across multiple scales and applying statistically grounded significance testing, the method effectively identifies anomalies without requiring labeled data. This work is the first to synergistically combine Haar wavelets with theoretically justified t-tests, substantially reducing false positives while enhancing detection accuracy. Extensive experiments on 343 real-world datasets demonstrate that the proposed approach outperforms current state-of-the-art unsupervised and self-supervised methods in both detection speed and accuracy.
Data assimilation provides a systematic framework for combining dynamical models with partial and noisy observations to infer the evolving state of a system. In this work, we undertake a comparative study of Data Assimilation with Transfer Operators (DATO) and Quantum Mechanical Data Assimilation (QMDA), focusing on their mathematical formulation, algorithmic structure, and empirical performance. Both methods are first cast within a common operator-theoretic framework, which makes it possible to compare, on a unified basis, their representations of uncertainty, forecast propagation, and assimilation updates. We then analyse their principal similarities and differences with respect to state-space structure, update mechanisms, structural preservation properties, and computational cost. To complement the theoretical analysis, we assess both approaches on benchmark dynamical systems across a range of observational settings, including noisy, sparse, and partially observed regimes. Our results show that, despite their shared operator-theoretic motivation, DATO and QMDA embody substantially different assimilation paradigms, leading to distinct advantages and limitations in terms of interpretability, robustness, and scalability. The present study helps delineate the regimes in which each framework is most effective and offers broader insight into the design of operator-based methodologies for data assimilation.