Institution profile

CloudWalk Technology

Industry researchasia · cn
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls

Jul 13, 2026

This work investigates how coordination content—such as shared states or tool outputs—in multi-agent or memory-augmented large language models consumes tokens within a fixed context budget, thereby displacing task instructions and evidence and degrading performance. To quantify this “task budget displacement” effect, the authors introduce the Roundtable Context Window Test (RCWT), which systematically controls total token budget, content ordering, task type, and evaluation metrics to disentangle the impact of coordination content volume from competition for contextual resources. Through controlled prompting, context window scaling, ablation studies, and experiments across multiple models—including GPT-4.1-mini, Claude Haiku 4.5, and Gemini 2.5 Flash—they find that performance sharply declines when fewer than a few hundred tokens remain for task evidence. However, if the evidence is fully preserved, models can achieve perfect task completion even when coordination content occupies up to 95% of the context window.

0 citationsRead paper

How Far are Modern Trackers from UAV-Anti-UAV? A Million-Scale Benchmark and New Baseline

Dec 08, 2025

This work addresses a critical gap in anti-drone research: prior studies focus predominantly on static ground-based platforms and neglect dynamic adversarial tracking between mobile UAVs. We thus introduce UAV-Anti-UAV multimodal visual tracking—a novel task wherein a pursuing UAV must localize and continuously track a hostile target UAV in real-time video streams. The task is severely challenged by dual-motion-induced dynamic interference. To enable systematic study, we construct the first large-scale, fully annotated dataset comprising 1,810 video sequences, each accompanied by natural-language prompts and 15 fine-grained tracking attributes. We further propose MambaSTS, a baseline method integrating the Mamba state-space model with Transformer architecture to jointly model spatial, temporal, and semantic information over long sequences. Evaluation on our dataset reveals significant room for improvement, establishing a new benchmark and technical foundation for mobile-platform anti-drone tracking.

0 citationsRead paper

COST: Contrastive One-Stage Transformer for Vision-Language Small Object Tracking

Apr 02, 2025

To address redundancy in multi-stage fusion, cross-modal representation inconsistency, and the difficulty of tracking small targets (<32×32) due to weak visual appearance in vision-language (VL) tracking, this paper proposes COST: a Contrastive One-stage Transformer fusion framework. COST employs an end-to-end single-stage architecture integrated with contrastive alignment and mutual information maximization to achieve semantically consistent modeling between video and language descriptions. We introduce VL-SOT500—the first dedicated benchmark for small-object VL tracking—comprising VL-SOT230 and VL-SOT270 subsets. Extensive experiments demonstrate that linguistic cues significantly enhance weak visual feature representation. COST achieves state-of-the-art performance across five mainstream VL tracking benchmarks and VL-SOT500, with particularly notable improvements in small-target tracking accuracy.

0 citationsRead paper
Recent publications

Latest Papers

RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls

Jul 13, 2026

This work investigates how coordination content—such as shared states or tool outputs—in multi-agent or memory-augmented large language models consumes tokens within a fixed context budget, thereby displacing task instructions and evidence and degrading performance. To quantify this “task budget displacement” effect, the authors introduce the Roundtable Context Window Test (RCWT), which systematically controls total token budget, content ordering, task type, and evaluation metrics to disentangle the impact of coordination content volume from competition for contextual resources. Through controlled prompting, context window scaling, ablation studies, and experiments across multiple models—including GPT-4.1-mini, Claude Haiku 4.5, and Gemini 2.5 Flash—they find that performance sharply declines when fewer than a few hundred tokens remain for task evidence. However, if the evidence is fully preserved, models can achieve perfect task completion even when coordination content occupies up to 95% of the context window.

0 citationsRead paper

How Far are Modern Trackers from UAV-Anti-UAV? A Million-Scale Benchmark and New Baseline

Dec 08, 2025

This work addresses a critical gap in anti-drone research: prior studies focus predominantly on static ground-based platforms and neglect dynamic adversarial tracking between mobile UAVs. We thus introduce UAV-Anti-UAV multimodal visual tracking—a novel task wherein a pursuing UAV must localize and continuously track a hostile target UAV in real-time video streams. The task is severely challenged by dual-motion-induced dynamic interference. To enable systematic study, we construct the first large-scale, fully annotated dataset comprising 1,810 video sequences, each accompanied by natural-language prompts and 15 fine-grained tracking attributes. We further propose MambaSTS, a baseline method integrating the Mamba state-space model with Transformer architecture to jointly model spatial, temporal, and semantic information over long sequences. Evaluation on our dataset reveals significant room for improvement, establishing a new benchmark and technical foundation for mobile-platform anti-drone tracking.

0 citationsRead paper

COST: Contrastive One-Stage Transformer for Vision-Language Small Object Tracking

Apr 02, 2025

To address redundancy in multi-stage fusion, cross-modal representation inconsistency, and the difficulty of tracking small targets (<32×32) due to weak visual appearance in vision-language (VL) tracking, this paper proposes COST: a Contrastive One-stage Transformer fusion framework. COST employs an end-to-end single-stage architecture integrated with contrastive alignment and mutual information maximization to achieve semantically consistent modeling between video and language descriptions. We introduce VL-SOT500—the first dedicated benchmark for small-object VL tracking—comprising VL-SOT230 and VL-SOT270 subsets. Extensive experiments demonstrate that linguistic cues significantly enhance weak visual feature representation. COST achieves state-of-the-art performance across five mainstream VL tracking benchmarks and VL-SOT500, with particularly notable improvements in small-target tracking accuracy.

0 citationsRead paper