Scaling Hindi Quantum Natural Language Processing through Automatic Pregroup Supertagging
本文通过将印地语pregroup标注视为分类任务,并评估多种方法,解决印地语量子自然语言处理中的手动类型分配问题,提高自动标注准确性。
本文通过将印地语pregroup标注视为分类任务,并评估多种方法,解决印地语量子自然语言处理中的手动类型分配问题,提高自动标注准确性。
This work addresses the challenge that the performance of quantum reservoir computing (QRC) is highly sensitive to architectural design, yet its vast hyperparameter space lacks efficient automated design methods. The authors formulate QRC architecture search as a constrained black-box optimization problem and propose a novel hybrid search framework that, for the first time, integrates a large language model (LLM) as a high-level controller within a reproducible search loop. This framework leverages memory mechanisms, mutation, crossover, deduplication, and exploration strategies to efficiently guide architecture generation without gradient information. Experimental results on NARMA10, Mackey-Glass prediction, and temporal parity tasks demonstrate that the method significantly outperforms random search within a budget of 25 evaluations, achieving a 23.6% relative error reduction on Mackey-Glass, thereby validating the potential of generative models in orchestrating quantum machine learning architectures.
Large language models (LLMs) often generate harmful content, posing serious risks to AI safety and public trust. Existing neuron-level intervention methods suffer from poor stability, strong context dependency, and unintended degradation of linguistic capabilities. To address these limitations, we propose EigenShift—a fine-tuning-free, low-overhead layer-wise intervention framework. EigenShift decouples the toxicity generation mechanism via inter-layer feature aggregation and output-layer feature decomposition. It innovatively separates toxicity detection and generation into distinct expert modules, enabling structured intervention through eigenvalue decomposition and generation-alignment analysis. Evaluated on the Jigsaw and ToxiCN benchmarks, EigenShift achieves significant and robust suppression of toxic outputs while preserving language modeling performance. The method demonstrates strong interpretability, cross-dataset generalizability, and deployment efficiency—offering a practical, principled solution for safe LLM inference.
Existing ethical and clinical decision-making benchmarks inadequately assess LLMs’ capacity to navigate intertwined ethical dilemmas—such as confidentiality, autonomy, and fairness—in mental health contexts. To address this gap, we introduce EthicsMH, the first fine-grained, mental health–specific ethical reasoning evaluation framework, accompanied by an open-source benchmark comprising 125 real-world ethically conflicting scenarios. Leveraging model-assisted generation augmented by multi-round expert validation, our methodology integrates moral psychology and clinical practice knowledge to design a structured annotation schema that supports multidimensional evaluation of AI decision justification, explanation quality, and alignment with professional standards. Key contributions include: (i) incorporation of multi-stakeholder perspectives; (ii) expert-aligned reasoning pathways; and (iii) standardized, clinically grounded decision options—collectively establishing a scalable, reproducible evaluation standard for responsible alignment of sensitive healthcare AI and fostering community-driven ethical assessment infrastructure.
Detecting black humor in internet memes is challenging due to its reliance on implicit, sensitive, and highly culture-dependent multimodal cues. To address this, we introduce the first large-scale Chinese meme dataset for black humor analysis (4,379 samples), supporting three tasks: black humor detection, target category identification, and intensity grading. Methodologically, we propose a Tri-stream Cross-Reasoning Network that jointly fuses OCR-extracted text, ViT-derived visual features, and structured reasoning sequences generated by a large vision-language model. We further innovate with a role-reversal self-cycling mechanism to better model cultural context and ironic logic. Experiments demonstrate significant improvements over strong baselines across all three tasks. Both the dataset and source code are publicly released to advance research in content safety and multimodal humor understanding.
本文通过将印地语pregroup标注视为分类任务,并评估多种方法,解决印地语量子自然语言处理中的手动类型分配问题,提高自动标注准确性。
This work addresses the challenge that the performance of quantum reservoir computing (QRC) is highly sensitive to architectural design, yet its vast hyperparameter space lacks efficient automated design methods. The authors formulate QRC architecture search as a constrained black-box optimization problem and propose a novel hybrid search framework that, for the first time, integrates a large language model (LLM) as a high-level controller within a reproducible search loop. This framework leverages memory mechanisms, mutation, crossover, deduplication, and exploration strategies to efficiently guide architecture generation without gradient information. Experimental results on NARMA10, Mackey-Glass prediction, and temporal parity tasks demonstrate that the method significantly outperforms random search within a budget of 25 evaluations, achieving a 23.6% relative error reduction on Mackey-Glass, thereby validating the potential of generative models in orchestrating quantum machine learning architectures.
Large language models (LLMs) often generate harmful content, posing serious risks to AI safety and public trust. Existing neuron-level intervention methods suffer from poor stability, strong context dependency, and unintended degradation of linguistic capabilities. To address these limitations, we propose EigenShift—a fine-tuning-free, low-overhead layer-wise intervention framework. EigenShift decouples the toxicity generation mechanism via inter-layer feature aggregation and output-layer feature decomposition. It innovatively separates toxicity detection and generation into distinct expert modules, enabling structured intervention through eigenvalue decomposition and generation-alignment analysis. Evaluated on the Jigsaw and ToxiCN benchmarks, EigenShift achieves significant and robust suppression of toxic outputs while preserving language modeling performance. The method demonstrates strong interpretability, cross-dataset generalizability, and deployment efficiency—offering a practical, principled solution for safe LLM inference.
Existing ethical and clinical decision-making benchmarks inadequately assess LLMs’ capacity to navigate intertwined ethical dilemmas—such as confidentiality, autonomy, and fairness—in mental health contexts. To address this gap, we introduce EthicsMH, the first fine-grained, mental health–specific ethical reasoning evaluation framework, accompanied by an open-source benchmark comprising 125 real-world ethically conflicting scenarios. Leveraging model-assisted generation augmented by multi-round expert validation, our methodology integrates moral psychology and clinical practice knowledge to design a structured annotation schema that supports multidimensional evaluation of AI decision justification, explanation quality, and alignment with professional standards. Key contributions include: (i) incorporation of multi-stakeholder perspectives; (ii) expert-aligned reasoning pathways; and (iii) standardized, clinically grounded decision options—collectively establishing a scalable, reproducible evaluation standard for responsible alignment of sensitive healthcare AI and fostering community-driven ethical assessment infrastructure.
Detecting black humor in internet memes is challenging due to its reliance on implicit, sensitive, and highly culture-dependent multimodal cues. To address this, we introduce the first large-scale Chinese meme dataset for black humor analysis (4,379 samples), supporting three tasks: black humor detection, target category identification, and intensity grading. Methodologically, we propose a Tri-stream Cross-Reasoning Network that jointly fuses OCR-extracted text, ViT-derived visual features, and structured reasoning sequences generated by a large vision-language model. We further innovate with a role-reversal self-cycling mechanism to better model cultural context and ironic logic. Experiments demonstrate significant improvements over strong baselines across all three tasks. Both the dataset and source code are publicly released to advance research in content safety and multimodal humor understanding.