WeaveMark: Robust and Scalable Multi-bit LLM Watermarking via Coded Payload Spreading
为解决多比特LLM水印技术在提取精度、文本质量和载荷容量间的权衡问题,提出WeaveMark方案,通过编码载荷扩散等方法提高性能。
为解决多比特LLM水印技术在提取精度、文本质量和载荷容量间的权衡问题,提出WeaveMark方案,通过编码载荷扩散等方法提高性能。
本文提出逻辑神经置信传播(L-NBP)解码器,通过将解码目标从物理层转向逻辑层,并结合神经网络权重训练,实现线性复杂度下的高精度量子纠错。
为解决混合大语言模型在异构芯片系统上高效运行的问题,HYDRA框架通过探索芯片组合、放置、带宽分配及动态调度等方法优化硬件加速。
This work addresses the challenge of balancing computational efficiency, area overhead, and numerical accuracy in hardware implementations of activation functions. The authors propose a novel algorithm-hardware co-design framework that introduces, for the first time, a cross-activation mixed-order, non-uniform piecewise polynomial approximation strategy. This approach is jointly optimized with an RTL-level area cost model to achieve high accuracy and hardware reuse under a unified configuration. Experimental results demonstrate that the proposed design attains a mean squared error below 8.22×10⁻⁸ and incurs no more than a 1.02% Top-1 accuracy loss across over 700 neural network variants and three NLP models. Implemented in 22 nm technology, the hardware footprint is only 6,800 μm².
Current large language model (LLM)-generated clinical research manuscripts commonly suffer from fabricated citations, data drift, and omissions of reporting guidelines, yet existing tools lack effective validation mechanisms. This work proposes an integrated generation-and-verification architecture that decomposes the writing process into 43 skill modules—including 21 deterministic detectors—orchestrated by a unified coordinator. It introduces a “maximally deterministic” completeness gating mechanism to enforce structured, traceable audits and re-execution checks at each stage. Evaluated on the STARD, PRISMA, and STROBE benchmark datasets, the approach successfully identified all 27 injected defects with zero false positives, substantially outperforming general-purpose LLM-based review methods.
为解决多比特LLM水印技术在提取精度、文本质量和载荷容量间的权衡问题,提出WeaveMark方案,通过编码载荷扩散等方法提高性能。
本文提出逻辑神经置信传播(L-NBP)解码器,通过将解码目标从物理层转向逻辑层,并结合神经网络权重训练,实现线性复杂度下的高精度量子纠错。
为解决混合大语言模型在异构芯片系统上高效运行的问题,HYDRA框架通过探索芯片组合、放置、带宽分配及动态调度等方法优化硬件加速。
This work addresses the challenge of balancing computational efficiency, area overhead, and numerical accuracy in hardware implementations of activation functions. The authors propose a novel algorithm-hardware co-design framework that introduces, for the first time, a cross-activation mixed-order, non-uniform piecewise polynomial approximation strategy. This approach is jointly optimized with an RTL-level area cost model to achieve high accuracy and hardware reuse under a unified configuration. Experimental results demonstrate that the proposed design attains a mean squared error below 8.22×10⁻⁸ and incurs no more than a 1.02% Top-1 accuracy loss across over 700 neural network variants and three NLP models. Implemented in 22 nm technology, the hardware footprint is only 6,800 μm².
Current large language model (LLM)-generated clinical research manuscripts commonly suffer from fabricated citations, data drift, and omissions of reporting guidelines, yet existing tools lack effective validation mechanisms. This work proposes an integrated generation-and-verification architecture that decomposes the writing process into 43 skill modules—including 21 deterministic detectors—orchestrated by a unified coordinator. It introduces a “maximally deterministic” completeness gating mechanism to enforce structured, traceable audits and re-execution checks at each stage. Evaluated on the STARD, PRISMA, and STROBE benchmark datasets, the approach successfully identified all 27 injected defects with zero false positives, substantially outperforming general-purpose LLM-based review methods.