GeoRay: Gauge-Aware Feed-Forward Satellite 3D Reconstruction in the Geodetic Frame
本文提出GeoRay方法,解决卫星摄影测量中绝对地理框架下的密集表面高度重建问题,通过轻量级射线一致适配器和显式基准机制实现。
本文提出GeoRay方法,解决卫星摄影测量中绝对地理框架下的密集表面高度重建问题,通过轻量级射线一致适配器和显式基准机制实现。
Existing vision-language models exhibit limited performance in drone image object detection due to significant domain discrepancies, and conventional parameter-efficient fine-tuning (PEFT) approaches struggle to address the unique challenges of aerial imagery—namely, the bird’s-eye view, background-dominated scenes, and extremely small objects. To overcome these limitations, this work proposes DroneFINE, a novel domain-aware dual-module framework. It introduces a dynamic multi-path HyperAdapter for flexible feature adaptation and a text-guided SemanticGate to effectively suppress irrelevant background clutter. By transcending the static architectural constraints of traditional PEFT methods, DroneFINE achieves substantial performance gains on the VisDrone and UAVDT benchmarks, closely approaching the accuracy of full fine-tuning while requiring only a minimal number of trainable parameters.
Existing online data mixing methods are limited to a single optimization objective, making them ill-suited for the multidimensional and dynamic data composition requirements of large language model pretraining. This work formulates data scheduling as a reinforcement learning problem in a continuous control space and introduces the first multi-objective reward function that integrates data-driven, loss-driven, and model-driven perspectives. Leveraging the Soft Actor-Critic algorithm, the approach enables efficient exploration and policy optimization. Experimental results demonstrate that the proposed method achieves the validation perplexity of the best baseline on The Pile using only 56% of the training steps, while also delivering a 7.2% improvement on zero-shot MMLU performance and consistent gains across multiple benchmark evaluations.
Existing approaches struggle to effectively detect multimodal misinformation in real-world scenarios characterized by multilingual long-form text, multiple images, heterogeneous sources, and fine-grained textual-visual inconsistencies. To address this challenge, this work introduces ReMMDBench, the first real-world benchmark supporting multilingualism, multi-image inputs, fine-grained labels, and evidence provenance. Furthermore, it proposes an agent-based verification framework endowed with persistent memory, which decomposes claims into atomic assertions and caches structured evidence for reuse. By integrating large vision-language models with cross-modal fusion strategies, the framework achieves efficient and accurate misinformation detection. On a five-class veracity assessment task, the method attains 41.80% accuracy and 39.12% macro F1-score, reducing verification costs by up to 79.9% compared to baseline approaches.
This work addresses the challenge of predicting viewer sentiment in video advertisements, where fine-grained emotion-relevant behaviors and visual cues are difficult to capture from full-frame inputs. The authors propose an action-centric, structured evidence enhancement framework that extracts temporal subject-predicate-object triplets, crops visual patches of participating entities, and integrates visible text to construct explicit, spatially localizable multimodal reasoning cues. This approach uniquely combines action triplets with entity-specific visual crops to guide interpretable sentiment reasoning in multimodal large language models (Qwen2.5-VL/Qwen3-VL). Evaluated on the Pitts dataset, the method significantly outperforms baseline approaches, and transfer experiments on AdsQA and TVQA subsets demonstrate its strong generalization capability.
本文提出GeoRay方法,解决卫星摄影测量中绝对地理框架下的密集表面高度重建问题,通过轻量级射线一致适配器和显式基准机制实现。
Existing vision-language models exhibit limited performance in drone image object detection due to significant domain discrepancies, and conventional parameter-efficient fine-tuning (PEFT) approaches struggle to address the unique challenges of aerial imagery—namely, the bird’s-eye view, background-dominated scenes, and extremely small objects. To overcome these limitations, this work proposes DroneFINE, a novel domain-aware dual-module framework. It introduces a dynamic multi-path HyperAdapter for flexible feature adaptation and a text-guided SemanticGate to effectively suppress irrelevant background clutter. By transcending the static architectural constraints of traditional PEFT methods, DroneFINE achieves substantial performance gains on the VisDrone and UAVDT benchmarks, closely approaching the accuracy of full fine-tuning while requiring only a minimal number of trainable parameters.
Existing online data mixing methods are limited to a single optimization objective, making them ill-suited for the multidimensional and dynamic data composition requirements of large language model pretraining. This work formulates data scheduling as a reinforcement learning problem in a continuous control space and introduces the first multi-objective reward function that integrates data-driven, loss-driven, and model-driven perspectives. Leveraging the Soft Actor-Critic algorithm, the approach enables efficient exploration and policy optimization. Experimental results demonstrate that the proposed method achieves the validation perplexity of the best baseline on The Pile using only 56% of the training steps, while also delivering a 7.2% improvement on zero-shot MMLU performance and consistent gains across multiple benchmark evaluations.
Existing approaches struggle to effectively detect multimodal misinformation in real-world scenarios characterized by multilingual long-form text, multiple images, heterogeneous sources, and fine-grained textual-visual inconsistencies. To address this challenge, this work introduces ReMMDBench, the first real-world benchmark supporting multilingualism, multi-image inputs, fine-grained labels, and evidence provenance. Furthermore, it proposes an agent-based verification framework endowed with persistent memory, which decomposes claims into atomic assertions and caches structured evidence for reuse. By integrating large vision-language models with cross-modal fusion strategies, the framework achieves efficient and accurate misinformation detection. On a five-class veracity assessment task, the method attains 41.80% accuracy and 39.12% macro F1-score, reducing verification costs by up to 79.9% compared to baseline approaches.
This work addresses the challenge of predicting viewer sentiment in video advertisements, where fine-grained emotion-relevant behaviors and visual cues are difficult to capture from full-frame inputs. The authors propose an action-centric, structured evidence enhancement framework that extracts temporal subject-predicate-object triplets, crops visual patches of participating entities, and integrates visible text to construct explicit, spatially localizable multimodal reasoning cues. This approach uniquely combines action triplets with entity-specific visual crops to guide interpretable sentiment reasoning in multimodal large language models (Qwen2.5-VL/Qwen3-VL). Evaluated on the Pitts dataset, the method significantly outperforms baseline approaches, and transfer experiments on AdsQA and TVQA subsets demonstrate its strong generalization capability.