Beyond Vector Search: Comparing Classical RAG with Hybrid GraphRAG for Climate Science Q\&A
论文提出一种结合向量搜索、GraphRAG、Leiden社区检测和交叉编码重排序的混合架构,以解决传统RAG系统在处理气候科学领域复杂文献时无法捕捉概念间层级关系的问题。
论文提出一种结合向量搜索、GraphRAG、Leiden社区检测和交叉编码重排序的混合架构,以解决传统RAG系统在处理气候科学领域复杂文献时无法捕捉概念间层级关系的问题。
To address low exploration efficiency and poor convergence in sparse-reward reinforcement learning, this paper proposes an intrinsic motivation mechanism that jointly leverages variational state novelty and prior knowledge from large language models (LLMs). Methodologically, it introduces LLM-encoded task-semantic priors—novelly integrated into intrinsic reward computation—for the first time, synergistically modeling them with VAE-based state-novelty rewards to achieve dynamic exploration-exploitation trade-offs; the overall framework is embedded within the A2C actor-critic architecture. Evaluated on the MiniGrid DoorKey task, our approach significantly improves sample efficiency and policy performance: while standard A2C fails entirely, our method achieves stable convergence. The core contribution lies in pioneering the injection of LLM-derived, structured world knowledge into intrinsic reward design—establishing a transferable, semantics-guided paradigm for sparse-reward RL.
Traditional data splitting methods ignore intrinsic instance quality, undermining model validation robustness. This paper introduces Item Response Theory (IRT)—a psychometric framework—into the machine learning validation phase for the first time. We model instance heterogeneity using IRT’s three parameters: difficulty, discrimination, and guessing. Based on this, we propose an IRT-guided data partitioning method that explicitly accounts for instance-level reliability. Key findings reveal that high-guessing instances significantly degrade model performance and identify interpretable subgroups affecting the bias–variance trade-off. Experiments across four tabular datasets demonstrate that our optimized splitting improves validation accuracy by over 20 percentage points (e.g., rising from <50% to >70% in certain cases), substantially enhancing assessment reliability. This work establishes a novel, data-quality-aware paradigm for model validation.
Traditional ML classifier evaluation overlooks data complexity and model robustness, leading to inflated performance estimates. To address this, we propose a novel, fairness-oriented evaluation paradigm that decouples *capability* from *robustness*. Our method integrates Item Response Theory (IRT) with the Glicko-2 dynamic rating system to construct a difficulty-aware instance-response model, and quantifies true classifier performance via classifier adversarial tournaments. Experiments on the OpenML-CC18 benchmark reveal that only 15% of datasets exhibit substantive difficulty; a 50%-reduced subset preserves full evaluation fidelity; and Random Forest achieves the highest capability score. This framework transcends static, population-averaged assessment by enabling interpretable, dynamically updated, dual-dimensional evaluation—simultaneously measuring intrinsic discriminative ability and resilience to adversarial instance selection—thereby advancing ML benchmarking toward greater ecological validity and diagnostic precision.
论文提出一种结合向量搜索、GraphRAG、Leiden社区检测和交叉编码重排序的混合架构,以解决传统RAG系统在处理气候科学领域复杂文献时无法捕捉概念间层级关系的问题。
To address low exploration efficiency and poor convergence in sparse-reward reinforcement learning, this paper proposes an intrinsic motivation mechanism that jointly leverages variational state novelty and prior knowledge from large language models (LLMs). Methodologically, it introduces LLM-encoded task-semantic priors—novelly integrated into intrinsic reward computation—for the first time, synergistically modeling them with VAE-based state-novelty rewards to achieve dynamic exploration-exploitation trade-offs; the overall framework is embedded within the A2C actor-critic architecture. Evaluated on the MiniGrid DoorKey task, our approach significantly improves sample efficiency and policy performance: while standard A2C fails entirely, our method achieves stable convergence. The core contribution lies in pioneering the injection of LLM-derived, structured world knowledge into intrinsic reward design—establishing a transferable, semantics-guided paradigm for sparse-reward RL.
Traditional data splitting methods ignore intrinsic instance quality, undermining model validation robustness. This paper introduces Item Response Theory (IRT)—a psychometric framework—into the machine learning validation phase for the first time. We model instance heterogeneity using IRT’s three parameters: difficulty, discrimination, and guessing. Based on this, we propose an IRT-guided data partitioning method that explicitly accounts for instance-level reliability. Key findings reveal that high-guessing instances significantly degrade model performance and identify interpretable subgroups affecting the bias–variance trade-off. Experiments across four tabular datasets demonstrate that our optimized splitting improves validation accuracy by over 20 percentage points (e.g., rising from <50% to >70% in certain cases), substantially enhancing assessment reliability. This work establishes a novel, data-quality-aware paradigm for model validation.
Traditional ML classifier evaluation overlooks data complexity and model robustness, leading to inflated performance estimates. To address this, we propose a novel, fairness-oriented evaluation paradigm that decouples *capability* from *robustness*. Our method integrates Item Response Theory (IRT) with the Glicko-2 dynamic rating system to construct a difficulty-aware instance-response model, and quantifies true classifier performance via classifier adversarial tournaments. Experiments on the OpenML-CC18 benchmark reveal that only 15% of datasets exhibit substantive difficulty; a 50%-reduced subset preserves full evaluation fidelity; and Random Forest achieves the highest capability score. This framework transcends static, population-averaged assessment by enabling interpretable, dynamically updated, dual-dimensional evaluation—simultaneously measuring intrinsic discriminative ability and resilience to adversarial instance selection—thereby advancing ML benchmarking toward greater ecological validity and diagnostic precision.