Benchmarking Embedding Models for ESG Data

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文构建了一个针对ESG领域的基准数据集,评估了14种嵌入模型在检索和检索增强生成任务中的表现,以解决ESG文本处理的有效性问题。
📝 Abstract
The use of Environmental, Social, and Governance (ESG) data is fundamental for modern corporate accountability, sustainability reporting, and financial decision-making. Embedding models have emerged as a powerful approach for transforming unstructured ESG text into numerical representations suitable for downstream natural language processing (NLP) tasks. However, their effectiveness in these ESG-specific tasks has not been systematically studied. In this paper, we construct a benchmark dataset specifically tailored to the ESG domain. We benchmark fourteen models, both open-source and closed-source embedding models, comparing their performance with respect to retrieval, and Retrieval-Augmented Generation (RAG). The results demonstrate performance variations across different models, with Qwen3-based models achieving the highest overall performance. This study provides practical insights into which models are better suited for ESG RAG tasks.
Problem

Research questions and friction points this paper is trying to address.

ESG data
embedding models
natural language processing
benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmarking
ESG Data
Embedding Models
Retrieval-Augmented Generation (RAG)
Qwen3
🔎 Similar Papers
No similar papers found.