How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation
研究评估了九种自动化事实核查系统在不同数据集上的表现,发现系统性能受领域和指标影响大,证据检索是主要瓶颈。
研究评估了九种自动化事实核查系统在不同数据集上的表现,发现系统性能受领域和指标影响大,证据检索是主要瓶颈。
This work addresses the lack of standardized SPARQL query annotations in academic knowledge graphs by proposing a zero-shot natural language to SPARQL query generation method. Built upon the Qwen3-1.7B language model, the approach integrates symbolic prompting with an execution feedback mechanism and introduces Group-Relative Policy Optimization (GRPO) to this task for the first time. The model is trained via reinforcement learning that combines structural constraints with answer-level rewards. Experimental results demonstrate that GRPO significantly outperforms zero-shot baselines and exhibits strong template generalization capabilities. Further improvements in overall accuracy are achieved when GRPO is combined with DoRA fine-tuning at the same model scale. Ablation studies confirm that execution feedback is a critical factor driving performance gains.
This study addresses the challenge of inconsistent construct definitions in information systems research, which impedes cumulative knowledge development. To resolve this issue, the authors propose a novel approach that leverages task-adapted text embeddings and clustering to generate candidate construct groupings. They introduce the first custom loss function explicitly balancing semantic purity against model parsimony, enabling the construction of unified and interpretable structural equation models. The method further supports dynamic analysis of how construct groupings evolve under varying optimization objectives. Empirical validation on two information systems datasets demonstrates the approach’s effectiveness in achieving semantically coherent integration of constructs and their interrelationships.
This work proposes a novel architecture based on the Model Context Protocol (MCP) for complex question answering over multiple knowledge graphs, integrating SPARQL federated query capabilities into large language model agents for the first time. The approach employs a three-stage collaborative mechanism—endpoint discovery, schema exploration, and query generation—to enable automatic federated querying across heterogeneous knowledge graph sources. The study extends existing federated KGQA benchmarks and systematically evaluates the effectiveness of various agent designs, demonstrating that SPARQL-MCP achieves superior performance and exhibits strong practical potential in complex, multi-source question-answering scenarios.
Existing neural combinatorial optimization (NCO) self-improvement methods suffer from low sample efficiency—requiring extensive sampling to generate a single expert trajectory—and neglect multi-agent permutation symmetry (e.g., vehicle interchangeability in the Vehicle Routing Problem), hindering generalization and cooperative learning. This work proposes the first self-improvement framework for NCO operating directly in the joint action space. It introduces a set-prediction-based loss function that explicitly enforces agent-permutation invariance, and adopts a proxy-task assignment architecture to generate multi-agent actions in parallel within a single step. Evaluated on multiple standard combinatorial optimization benchmarks, our method achieves superior solution quality, reduces inference latency by 30–50%, improves training efficiency by 2.1×, and—critically—unifies high sample efficiency with strong cooperative capability for the first time.
研究评估了九种自动化事实核查系统在不同数据集上的表现,发现系统性能受领域和指标影响大,证据检索是主要瓶颈。
This work addresses the lack of standardized SPARQL query annotations in academic knowledge graphs by proposing a zero-shot natural language to SPARQL query generation method. Built upon the Qwen3-1.7B language model, the approach integrates symbolic prompting with an execution feedback mechanism and introduces Group-Relative Policy Optimization (GRPO) to this task for the first time. The model is trained via reinforcement learning that combines structural constraints with answer-level rewards. Experimental results demonstrate that GRPO significantly outperforms zero-shot baselines and exhibits strong template generalization capabilities. Further improvements in overall accuracy are achieved when GRPO is combined with DoRA fine-tuning at the same model scale. Ablation studies confirm that execution feedback is a critical factor driving performance gains.
This study addresses the challenge of inconsistent construct definitions in information systems research, which impedes cumulative knowledge development. To resolve this issue, the authors propose a novel approach that leverages task-adapted text embeddings and clustering to generate candidate construct groupings. They introduce the first custom loss function explicitly balancing semantic purity against model parsimony, enabling the construction of unified and interpretable structural equation models. The method further supports dynamic analysis of how construct groupings evolve under varying optimization objectives. Empirical validation on two information systems datasets demonstrates the approach’s effectiveness in achieving semantically coherent integration of constructs and their interrelationships.
This work proposes a novel architecture based on the Model Context Protocol (MCP) for complex question answering over multiple knowledge graphs, integrating SPARQL federated query capabilities into large language model agents for the first time. The approach employs a three-stage collaborative mechanism—endpoint discovery, schema exploration, and query generation—to enable automatic federated querying across heterogeneous knowledge graph sources. The study extends existing federated KGQA benchmarks and systematically evaluates the effectiveness of various agent designs, demonstrating that SPARQL-MCP achieves superior performance and exhibits strong practical potential in complex, multi-source question-answering scenarios.
Existing neural combinatorial optimization (NCO) self-improvement methods suffer from low sample efficiency—requiring extensive sampling to generate a single expert trajectory—and neglect multi-agent permutation symmetry (e.g., vehicle interchangeability in the Vehicle Routing Problem), hindering generalization and cooperative learning. This work proposes the first self-improvement framework for NCO operating directly in the joint action space. It introduces a set-prediction-based loss function that explicitly enforces agent-permutation invariance, and adopts a proxy-task assignment architecture to generate multi-agent actions in parallel within a single step. Evaluated on multiple standard combinatorial optimization benchmarks, our method achieves superior solution quality, reduces inference latency by 30–50%, improves training efficiency by 2.1×, and—critically—unifies high sample efficiency with strong cooperative capability for the first time.