Rethinking Sign Language Translation: The Impact of Signer Dependence on Model Evaluation
研究针对手语翻译模型过度依赖特定手语者的问题,采用签名者独立验证方法评估模型,揭示了现有评估方式可能高估模型能力。
研究针对手语翻译模型过度依赖特定手语者的问题,采用签名者独立验证方法评估模型,揭示了现有评估方式可能高估模型能力。
研究通过改进TabFMs与CoxPH和DeepHit的结合方法,解决了结构化数据中被删失的时间到事件预测问题,并评估了不同适应接口的有效性。
To address the proliferation of misinformation, existing large language model (LLM)-based fact-checking approaches suffer from high computational overhead, severe hallucination risks, and poor deployability. This paper proposes DeReC, a lightweight and efficient fact verification framework that pioneers the integration of general-purpose text embeddings with dense retrieval—replacing LLM-based generative reasoning—and introduces a dedicated classifier for end-to-end verification. By preserving semantic understanding while eliminating autoregressive generation, DeReC significantly reduces computational cost. Experiments show that DeReC achieves an F1 score of 65.58% on RAWFC, outperforming the state-of-the-art L-Defense (61.20%) and accelerating inference by 20× (95% runtime reduction); on LIAR-RAW, it achieves 12× speedup (92% reduction). This work is the first to empirically validate the superiority of non-generative dense retrieval for fact-checking, establishing a new paradigm for scalable, low-cost, and robust verification systems.
Current evaluations of multilingual large language models’ (LLMs’) cultural reasoning capabilities rely predominantly on answer accuracy, neglecting interpretability and cross-linguistic comparability. To address this, we propose CRaFT—the first explanation-based framework for cross-cultural reasoning assessment. CRaFT introduces a four-dimensional explanatory quality metric: cultural fluency, deviation, consistency, and linguistic adaptability. Leveraging the World Values Survey, we construct a culturally grounded, multilingual question–explanation dataset covering Arabic, Bengali, and Spanish (2,100+ instances). Empirical analysis reveals salient language-specific patterns: Arabic responses exhibit lower cultural fluency; Bengali reasoning achieves higher overall quality; GPT-4 demonstrates strong linguistic adaptability but weak consistency; conversely, FANAR shows high stability yet limited flexibility. CRaFT establishes a novel, interpretable, decomposable, and cross-linguistically comparable paradigm for evaluating culturally intelligent multilingual LLMs.
To address inaccurate and poorly interpretable job-title matching in resume recommendation systems—caused by low lexical overlap or semantic ambiguity—this paper proposes a hierarchical matching method integrating semantic modeling with domain knowledge. Methodologically: (1) it introduces a self-supervised hybrid architecture coupling fine-tuned SBERT with a graph neural network, explicitly injecting domain knowledge graphs into semantic matching; (2) it designs a hierarchical evaluation strategy that performs fine-grained analysis across semantic relevance intervals, revealing model-behavior discrepancies obscured by global metrics. Experiments show a 25% reduction in RMSE over strong baselines on the high-relevance subset, significantly improving both matching accuracy and decision interpretability. The core contribution is a knowledge-enhanced hierarchical semantic alignment framework that jointly improves matching robustness and explainability.
研究针对手语翻译模型过度依赖特定手语者的问题,采用签名者独立验证方法评估模型,揭示了现有评估方式可能高估模型能力。
研究通过改进TabFMs与CoxPH和DeepHit的结合方法,解决了结构化数据中被删失的时间到事件预测问题,并评估了不同适应接口的有效性。
To address the proliferation of misinformation, existing large language model (LLM)-based fact-checking approaches suffer from high computational overhead, severe hallucination risks, and poor deployability. This paper proposes DeReC, a lightweight and efficient fact verification framework that pioneers the integration of general-purpose text embeddings with dense retrieval—replacing LLM-based generative reasoning—and introduces a dedicated classifier for end-to-end verification. By preserving semantic understanding while eliminating autoregressive generation, DeReC significantly reduces computational cost. Experiments show that DeReC achieves an F1 score of 65.58% on RAWFC, outperforming the state-of-the-art L-Defense (61.20%) and accelerating inference by 20× (95% runtime reduction); on LIAR-RAW, it achieves 12× speedup (92% reduction). This work is the first to empirically validate the superiority of non-generative dense retrieval for fact-checking, establishing a new paradigm for scalable, low-cost, and robust verification systems.
Current evaluations of multilingual large language models’ (LLMs’) cultural reasoning capabilities rely predominantly on answer accuracy, neglecting interpretability and cross-linguistic comparability. To address this, we propose CRaFT—the first explanation-based framework for cross-cultural reasoning assessment. CRaFT introduces a four-dimensional explanatory quality metric: cultural fluency, deviation, consistency, and linguistic adaptability. Leveraging the World Values Survey, we construct a culturally grounded, multilingual question–explanation dataset covering Arabic, Bengali, and Spanish (2,100+ instances). Empirical analysis reveals salient language-specific patterns: Arabic responses exhibit lower cultural fluency; Bengali reasoning achieves higher overall quality; GPT-4 demonstrates strong linguistic adaptability but weak consistency; conversely, FANAR shows high stability yet limited flexibility. CRaFT establishes a novel, interpretable, decomposable, and cross-linguistically comparable paradigm for evaluating culturally intelligent multilingual LLMs.
To address inaccurate and poorly interpretable job-title matching in resume recommendation systems—caused by low lexical overlap or semantic ambiguity—this paper proposes a hierarchical matching method integrating semantic modeling with domain knowledge. Methodologically: (1) it introduces a self-supervised hybrid architecture coupling fine-tuned SBERT with a graph neural network, explicitly injecting domain knowledge graphs into semantic matching; (2) it designs a hierarchical evaluation strategy that performs fine-grained analysis across semantic relevance intervals, revealing model-behavior discrepancies obscured by global metrics. Experiments show a 25% reduction in RMSE over strong baselines on the high-relevance subset, significantly improving both matching accuracy and decision interpretability. The core contribution is a knowledge-enhanced hierarchical semantic alignment framework that jointly improves matching robustness and explainability.