A Cone-Constrained Bilinear Decomposition for Total Scaled-Gradient Variation Models
为解决TSGV正则化器的计算难题,提出了一种双线性分解方法,并通过交替最小化法求解,保证了收敛性和边缘保持性能。
为解决TSGV正则化器的计算难题,提出了一种双线性分解方法,并通过交替最小化法求解,保证了收敛性和边缘保持性能。
本文探讨了簇图编辑距离的计算复杂性和度量几何,通过将q*表示为一个仿射函数来提供两种明确的l1模型,并提出了一种O(n log n)时间算法以找到成本低于2q*的对齐方式。
This study addresses the misalignment between existing visual token selection criteria and reconstruction quality under fixed bandwidth constraints. We propose Gated Counterfactual Rectification (GCR-C), a method that constructs candidate sets and performs full-budget counterfactual evaluations to dynamically replace baseline actions only when positive gains are confirmed. This approach effectively bridges the gap between selection strategies and final reconstruction outcomes. Experiments demonstrate that GCR-C significantly improves reconstruction quality at low-to-medium bitrates across diverse datasets and channel conditions without increasing actual bitrate consumption. Furthermore, the method exhibits robust generalization capabilities, establishing a novel paradigm for communication-aware reconstruction tasks.
This work addresses the challenge of balancing semantic fidelity and transmission efficiency in resource-constrained visual Internet-of-Things systems by proposing a semantic-aware generative image transmission framework. The approach integrates instance segmentation–driven semantic scoring with prediction entropy–guided recoverability assessment to intelligently sample discrete VQ tokens, and leverages MaskGIT for reconstructing missing content at the edge or cloud. Spatially dispersed scheduling via Halton sequences is introduced to enhance generation quality. At a bitrate of 0.074 bpp—only 44.6% of that required by DeepJSCC/WITT—the method achieves a PSNR of 29.9 dB, while downstream detection tasks demonstrate that its semantic masking strategy significantly outperforms random masking.
Structured output often degrades reasoning performance, yet the underlying cause remains unclear. This work disentangles the effects of output format from prompt-length confounds through carefully designed natural language controls and a four-level complexity framework, evaluated across multiple models (Sonnet, Haiku, GPT-4o-mini, Opus) and five benchmarks, including MATH-Hard and AIME. The study introduces a “capacity competition” mechanism, demonstrating that performance loss stems not from the structured format itself but from insufficient residual model capacity: high-capacity models handle JSON output without degradation, whereas capacity-constrained models suffer substantial drops (Haiku ↓36.2 pp, GPT-4o-mini ↓28.0 pp). To mitigate this, the authors propose a “reason-then-format” strategy, which recovers 80–87% of lost accuracy, effectively alleviating the issue.
为解决TSGV正则化器的计算难题,提出了一种双线性分解方法,并通过交替最小化法求解,保证了收敛性和边缘保持性能。
本文探讨了簇图编辑距离的计算复杂性和度量几何,通过将q*表示为一个仿射函数来提供两种明确的l1模型,并提出了一种O(n log n)时间算法以找到成本低于2q*的对齐方式。
This study addresses the misalignment between existing visual token selection criteria and reconstruction quality under fixed bandwidth constraints. We propose Gated Counterfactual Rectification (GCR-C), a method that constructs candidate sets and performs full-budget counterfactual evaluations to dynamically replace baseline actions only when positive gains are confirmed. This approach effectively bridges the gap between selection strategies and final reconstruction outcomes. Experiments demonstrate that GCR-C significantly improves reconstruction quality at low-to-medium bitrates across diverse datasets and channel conditions without increasing actual bitrate consumption. Furthermore, the method exhibits robust generalization capabilities, establishing a novel paradigm for communication-aware reconstruction tasks.
This work addresses the challenge of balancing semantic fidelity and transmission efficiency in resource-constrained visual Internet-of-Things systems by proposing a semantic-aware generative image transmission framework. The approach integrates instance segmentation–driven semantic scoring with prediction entropy–guided recoverability assessment to intelligently sample discrete VQ tokens, and leverages MaskGIT for reconstructing missing content at the edge or cloud. Spatially dispersed scheduling via Halton sequences is introduced to enhance generation quality. At a bitrate of 0.074 bpp—only 44.6% of that required by DeepJSCC/WITT—the method achieves a PSNR of 29.9 dB, while downstream detection tasks demonstrate that its semantic masking strategy significantly outperforms random masking.
Structured output often degrades reasoning performance, yet the underlying cause remains unclear. This work disentangles the effects of output format from prompt-length confounds through carefully designed natural language controls and a four-level complexity framework, evaluated across multiple models (Sonnet, Haiku, GPT-4o-mini, Opus) and five benchmarks, including MATH-Hard and AIME. The study introduces a “capacity competition” mechanism, demonstrating that performance loss stems not from the structured format itself but from insufficient residual model capacity: high-capacity models handle JSON output without degradation, whereas capacity-constrained models suffer substantial drops (Haiku ↓36.2 pp, GPT-4o-mini ↓28.0 pp). To mitigate this, the authors propose a “reason-then-format” strategy, which recovers 80–87% of lost accuracy, effectively alleviating the issue.