Contextrast++: Robust Multi-Scale Contextual Contrastive Learning for Semantic Segmentation
针对语义分割中上下文捕捉及长尾分布问题,提出Contextrast++方法,通过多尺度特征融合与边界感知负采样提升性能。
针对语义分割中上下文捕捉及长尾分布问题,提出Contextrast++方法,通过多尺度特征融合与边界感知负采样提升性能。
该论文介绍了一个名为OpenSCvx的开源Python框架,用于解决轨迹优化问题,通过提供符号建模接口自动生成并求解轨迹优化问题。
This work addresses the high inference latency of Vision Transformers in multi-view 3D object detection, which stems from dense processing of 2D image tokens and 3D queries, and the inability of existing sparse methods to jointly optimize both. The authors propose a correlation-aligned sparsification framework that, for the first time, enables joint sparsification of 2D tokens and 3D queries within a ViT architecture. Their approach employs a 2D–3D cross-correlation head to dynamically assess relevance and co-select critical tokens and queries, complemented by a feature caching and reactivation mechanism to reuse filtered features. Evaluated on nuScenes and a newly introduced nuScenes-Relevance benchmark, the method achieves up to 3× speedup with only marginal accuracy degradation, substantially improving inference efficiency and enabling scalable, real-time 3D detection.
Existing diffusion language models struggle to effectively improve generation quality at test time through increased inference computation, primarily because conventional best-of-K sampling is constrained by repeatedly drawing from the same distribution. This work proposes Hierarchical Scaling Search ($S^3$), which introduces verifier-guided, trajectory-level search into diffusion language models for the first time. During the denoising process, $S^3$ dynamically expands multiple candidate trajectories and employs a lightweight, reference-free verifier to evaluate and selectively resample high-potential paths while preserving diversity. Notably, this approach approximates a reward-weighted distribution without modifying the base model or decoding schedule. Experiments demonstrate that $S^3$ significantly enhances performance on mathematical reasoning benchmarks such as MATH-500 and GSM8K when applied to LLaDA-8B-Instruct, validating the efficacy of test-time scaling.
Traditional binary preference supervision struggles to capture fine-grained quality in reasoning processes, limiting the alignment efficacy of large language models on complex reasoning tasks. This work proposes CU-DPO, a novel framework that replaces binary labels with continuous utility scores to enable fine-grained alignment across diverse prompt-driven cognitive strategies. The approach employs a two-stage decoupled training procedure: first selecting strategies via a best-vs-all mechanism, then refining strategy execution through margin-stratified contrastive learning combined with entropy regularization. Theoretical analysis reveals a sample complexity improvement of Θ(K log K). Empirical results demonstrate that, across seven base models, strategy selection accuracy improves from 35–46% to 68–78%, with in-distribution mathematical reasoning performance gaining up to 6.6 points and strong generalization observed on out-of-distribution tasks.
针对语义分割中上下文捕捉及长尾分布问题,提出Contextrast++方法,通过多尺度特征融合与边界感知负采样提升性能。
该论文介绍了一个名为OpenSCvx的开源Python框架,用于解决轨迹优化问题,通过提供符号建模接口自动生成并求解轨迹优化问题。
This work addresses the high inference latency of Vision Transformers in multi-view 3D object detection, which stems from dense processing of 2D image tokens and 3D queries, and the inability of existing sparse methods to jointly optimize both. The authors propose a correlation-aligned sparsification framework that, for the first time, enables joint sparsification of 2D tokens and 3D queries within a ViT architecture. Their approach employs a 2D–3D cross-correlation head to dynamically assess relevance and co-select critical tokens and queries, complemented by a feature caching and reactivation mechanism to reuse filtered features. Evaluated on nuScenes and a newly introduced nuScenes-Relevance benchmark, the method achieves up to 3× speedup with only marginal accuracy degradation, substantially improving inference efficiency and enabling scalable, real-time 3D detection.
Existing diffusion language models struggle to effectively improve generation quality at test time through increased inference computation, primarily because conventional best-of-K sampling is constrained by repeatedly drawing from the same distribution. This work proposes Hierarchical Scaling Search ($S^3$), which introduces verifier-guided, trajectory-level search into diffusion language models for the first time. During the denoising process, $S^3$ dynamically expands multiple candidate trajectories and employs a lightweight, reference-free verifier to evaluate and selectively resample high-potential paths while preserving diversity. Notably, this approach approximates a reward-weighted distribution without modifying the base model or decoding schedule. Experiments demonstrate that $S^3$ significantly enhances performance on mathematical reasoning benchmarks such as MATH-500 and GSM8K when applied to LLaDA-8B-Instruct, validating the efficacy of test-time scaling.
Traditional binary preference supervision struggles to capture fine-grained quality in reasoning processes, limiting the alignment efficacy of large language models on complex reasoning tasks. This work proposes CU-DPO, a novel framework that replaces binary labels with continuous utility scores to enable fine-grained alignment across diverse prompt-driven cognitive strategies. The approach employs a two-stage decoupled training procedure: first selecting strategies via a best-vs-all mechanism, then refining strategy execution through margin-stratified contrastive learning combined with entropy regularization. Theoretical analysis reveals a sample complexity improvement of Θ(K log K). Empirical results demonstrate that, across seven base models, strategy selection accuracy improves from 35–46% to 68–78%, with in-distribution mathematical reasoning performance gaining up to 6.6 points and strong generalization observed on out-of-distribution tasks.