PuzzleMate: Benchmarking MLLMs for Egocentric Puzzle Assistance
本文通过PuzzleMate框架研究了多模态大语言模型在拼图任务中的实时指导能力,揭示了现有模型如GPT-5.2和Gemini-2.5-Pro在精准空间推理和序列逻辑上的不足。
本文通过PuzzleMate框架研究了多模态大语言模型在拼图任务中的实时指导能力,揭示了现有模型如GPT-5.2和Gemini-2.5-Pro在精准空间推理和序列逻辑上的不足。
本文提出GenQAS,通过结合张量网络引导的强化学习和生成重放方法来解决量子架构搜索中的样本匮乏问题。
为解决工业云中多变高维工作负载预测问题,提出了一种结合量子力学和神经网络的QB-HNN模型,并通过量子黑洞双相优化算法进行训练优化。
This work addresses the gap between benchmark-driven embedding model selection and real-world deployment constraints by introducing the first framework to evaluate embedding models within a complete retrieval pipeline. It systematically compares the end-to-end performance of T3EM’s commercial API against leading open-source models across diverse tasks—including retrieval, classification, clustering, and semantic similarity—as covered by the MTEB benchmark, while jointly accounting for latency, cost, task type, and deployment conditions. The study develops a comprehensive, end-to-end model selection guide encompassing embedding generation, indexing, search, and chunking strategies, revealing significant performance discrepancies that emerge only in full-system contexts. These insights provide practitioners with actionable, empirically grounded criteria for embedding model adoption in real-world applications.
This work addresses the limited generalization of conventional quantum error correction decoders across diverse code families and noise environments. The authors propose the first universal meta-decoding framework, which jointly optimizes a classical Meta-MLP teacher model and a hardware-aware variational quantum circuit (VQC) through meta-learning. A confidence-gated mechanism is introduced to enable selective recovery, thereby avoiding blind replacement of decoding decisions. This approach achieves, for the first time, unified decoding across multiple stabilizer codes and noise types. Experimental results demonstrate that confidence gating significantly reduces logical error rates across five evaluation scenarios. Notably, on the most challenging Planar 5×5 code, the VQC-based decoder lowers the logical failure ratio from 25.91 to 1.11, substantially outperforming the non-gated baseline.
本文通过PuzzleMate框架研究了多模态大语言模型在拼图任务中的实时指导能力,揭示了现有模型如GPT-5.2和Gemini-2.5-Pro在精准空间推理和序列逻辑上的不足。
本文提出GenQAS,通过结合张量网络引导的强化学习和生成重放方法来解决量子架构搜索中的样本匮乏问题。
为解决工业云中多变高维工作负载预测问题,提出了一种结合量子力学和神经网络的QB-HNN模型,并通过量子黑洞双相优化算法进行训练优化。
This work addresses the gap between benchmark-driven embedding model selection and real-world deployment constraints by introducing the first framework to evaluate embedding models within a complete retrieval pipeline. It systematically compares the end-to-end performance of T3EM’s commercial API against leading open-source models across diverse tasks—including retrieval, classification, clustering, and semantic similarity—as covered by the MTEB benchmark, while jointly accounting for latency, cost, task type, and deployment conditions. The study develops a comprehensive, end-to-end model selection guide encompassing embedding generation, indexing, search, and chunking strategies, revealing significant performance discrepancies that emerge only in full-system contexts. These insights provide practitioners with actionable, empirically grounded criteria for embedding model adoption in real-world applications.
This work addresses the limited generalization of conventional quantum error correction decoders across diverse code families and noise environments. The authors propose the first universal meta-decoding framework, which jointly optimizes a classical Meta-MLP teacher model and a hardware-aware variational quantum circuit (VQC) through meta-learning. A confidence-gated mechanism is introduced to enable selective recovery, thereby avoiding blind replacement of decoding decisions. This approach achieves, for the first time, unified decoding across multiple stabilizer codes and noise types. Experimental results demonstrate that confidence gating significantly reduces logical error rates across five evaluation scenarios. Notably, on the most challenging Planar 5×5 code, the VQC-based decoder lowers the logical failure ratio from 25.91 to 1.11, substantially outperforming the non-gated baseline.