Institution profile

China Electronics Technology Group Corporation

Industry researchasia · cn
Official website
Research library38linked papers
Opportunities0open roles
Selected work

Representative Papers

DroneFINE: Domain-Aware Parameter-Efficient Fine-Tuning of Vision-Language Detectors for Drone Images

Jun 30, 2026

Existing vision-language models exhibit limited performance in drone image object detection due to significant domain discrepancies, and conventional parameter-efficient fine-tuning (PEFT) approaches struggle to address the unique challenges of aerial imagery—namely, the bird’s-eye view, background-dominated scenes, and extremely small objects. To overcome these limitations, this work proposes DroneFINE, a novel domain-aware dual-module framework. It introduces a dynamic multi-path HyperAdapter for flexible feature adaptation and a text-guided SemanticGate to effectively suppress irrelevant background clutter. By transcending the static architectural constraints of traditional PEFT methods, DroneFINE achieves substantial performance gains on the VisDrone and UAVDT benchmarks, closely approaching the accuracy of full fine-tuning while requiring only a minimal number of trainable parameters.

0 citationsRead paper

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning

Jun 23, 2026

Existing online data mixing methods are limited to a single optimization objective, making them ill-suited for the multidimensional and dynamic data composition requirements of large language model pretraining. This work formulates data scheduling as a reinforcement learning problem in a continuous control space and introduces the first multi-objective reward function that integrates data-driven, loss-driven, and model-driven perspectives. Leveraging the Soft Actor-Critic algorithm, the approach enables efficient exploration and policy optimization. Experimental results demonstrate that the proposed method achieves the validation perplexity of the best baseline on The Pile using only 56% of the training steps, while also delivering a 7.2% improvement on zero-shot MMLU performance and consistent gains across multiple benchmark evaluations.

0 citationsRead paper

ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection

Jun 22, 2026

Existing approaches struggle to effectively detect multimodal misinformation in real-world scenarios characterized by multilingual long-form text, multiple images, heterogeneous sources, and fine-grained textual-visual inconsistencies. To address this challenge, this work introduces ReMMDBench, the first real-world benchmark supporting multilingualism, multi-image inputs, fine-grained labels, and evidence provenance. Furthermore, it proposes an agent-based verification framework endowed with persistent memory, which decomposes claims into atomic assertions and caches structured evidence for reuse. By integrating large vision-language models with cross-modal fusion strategies, the framework achieves efficient and accurate misinformation detection. On a five-class veracity assessment task, the method attains 41.80% accuracy and 39.12% macro F1-score, reducing verification costs by up to 79.9% compared to baseline approaches.

0 citationsRead paper

GRACE: Boosting Video MLLMs with Grounded Action-Centric Evidence for Viewer Sentiment Prediction

Jun 15, 2026

This work addresses the challenge of predicting viewer sentiment in video advertisements, where fine-grained emotion-relevant behaviors and visual cues are difficult to capture from full-frame inputs. The authors propose an action-centric, structured evidence enhancement framework that extracts temporal subject-predicate-object triplets, crops visual patches of participating entities, and integrates visible text to construct explicit, spatially localizable multimodal reasoning cues. This approach uniquely combines action triplets with entity-specific visual crops to guide interpretable sentiment reasoning in multimodal large language models (Qwen2.5-VL/Qwen3-VL). Evaluated on the Pitts dataset, the method significantly outperforms baseline approaches, and transfer experiments on AdsQA and TVQA subsets demonstrate its strong generalization capability.

0 citationsRead paper
Recent publications

Latest Papers

DroneFINE: Domain-Aware Parameter-Efficient Fine-Tuning of Vision-Language Detectors for Drone Images

Jun 30, 2026

Existing vision-language models exhibit limited performance in drone image object detection due to significant domain discrepancies, and conventional parameter-efficient fine-tuning (PEFT) approaches struggle to address the unique challenges of aerial imagery—namely, the bird’s-eye view, background-dominated scenes, and extremely small objects. To overcome these limitations, this work proposes DroneFINE, a novel domain-aware dual-module framework. It introduces a dynamic multi-path HyperAdapter for flexible feature adaptation and a text-guided SemanticGate to effectively suppress irrelevant background clutter. By transcending the static architectural constraints of traditional PEFT methods, DroneFINE achieves substantial performance gains on the VisDrone and UAVDT benchmarks, closely approaching the accuracy of full fine-tuning while requiring only a minimal number of trainable parameters.

0 citationsRead paper

Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning

Jun 23, 2026

Existing online data mixing methods are limited to a single optimization objective, making them ill-suited for the multidimensional and dynamic data composition requirements of large language model pretraining. This work formulates data scheduling as a reinforcement learning problem in a continuous control space and introduces the first multi-objective reward function that integrates data-driven, loss-driven, and model-driven perspectives. Leveraging the Soft Actor-Critic algorithm, the approach enables efficient exploration and policy optimization. Experimental results demonstrate that the proposed method achieves the validation perplexity of the best baseline on The Pile using only 56% of the training steps, while also delivering a 7.2% improvement on zero-shot MMLU performance and consistent gains across multiple benchmark evaluations.

0 citationsRead paper

ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection

Jun 22, 2026

Existing approaches struggle to effectively detect multimodal misinformation in real-world scenarios characterized by multilingual long-form text, multiple images, heterogeneous sources, and fine-grained textual-visual inconsistencies. To address this challenge, this work introduces ReMMDBench, the first real-world benchmark supporting multilingualism, multi-image inputs, fine-grained labels, and evidence provenance. Furthermore, it proposes an agent-based verification framework endowed with persistent memory, which decomposes claims into atomic assertions and caches structured evidence for reuse. By integrating large vision-language models with cross-modal fusion strategies, the framework achieves efficient and accurate misinformation detection. On a five-class veracity assessment task, the method attains 41.80% accuracy and 39.12% macro F1-score, reducing verification costs by up to 79.9% compared to baseline approaches.

0 citationsRead paper

GRACE: Boosting Video MLLMs with Grounded Action-Centric Evidence for Viewer Sentiment Prediction

Jun 15, 2026

This work addresses the challenge of predicting viewer sentiment in video advertisements, where fine-grained emotion-relevant behaviors and visual cues are difficult to capture from full-frame inputs. The authors propose an action-centric, structured evidence enhancement framework that extracts temporal subject-predicate-object triplets, crops visual patches of participating entities, and integrates visible text to construct explicit, spatially localizable multimodal reasoning cues. This approach uniquely combines action triplets with entity-specific visual crops to guide interpretable sentiment reasoning in multimodal large language models (Qwen2.5-VL/Qwen3-VL). Evaluated on the Pitts dataset, the method significantly outperforms baseline approaches, and transfer experiments on AdsQA and TVQA subsets demonstrate its strong generalization capability.

0 citationsRead paper