Institution profile

Wyze Labs

Industry researchnorthamerica · us
Official website
Research library8linked papers
Opportunities0open roles
Selected work

Representative Papers

AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition

Jul 28, 2026

This work addresses the challenges of on-device speech emotion recognition, where large self-supervised models incur prohibitive computational costs, and existing knowledge distillation approaches suffer from unreliable teacher predictions and neglect of inter-sample relational structures. To overcome these limitations, the authors propose an adaptive multi-teacher relational distillation framework. This framework incorporates a one-class SVM–based mechanism to assess teacher reliability and dynamically weight their predictions, while also introducing a relational distillation loss that preserves cross-sample semantic structure by aligning similarity matrices between teachers and students. Experiments on IEMOCAP and CREMA-D demonstrate that four lightweight student models consistently outperform single-teacher baselines, and ablation studies confirm the complementary benefits of the two core components.

0 citationsRead paper

Comment and Control: Hijacking Agentic Workflows via Context-Grounded Evolution

May 11, 2026

This work demonstrates that large language model (LLM) agents integrated into automation platforms are vulnerable to manipulation via malicious inputs—such as GitHub comments—leading to risks including credential leakage and arbitrary command execution. To systematically uncover these novel attack surfaces in agent workflows, the authors propose “Context-Anchored Evolution,” a method that combines static path feasibility analysis, dynamic prompt provenance tracing, and runtime capability assessment within a three-stage contextual framework, enhanced by evolutionary input generation to achieve targeted hijacking of LLM agents. The resulting JAW framework successfully exploits 4,714 GitHub workflows and 8 n8n templates, compromising 15 widely used GitHub Actions and 2 official n8n nodes. Vulnerabilities identified through this approach have been acknowledged and patched by vendors including GitHub, Google, and Anthropic.

0 citationsRead paper

IIR-VLM: In-Context Instance-level Recognition for Large Vision-Language Models

Jan 20, 2026

This work addresses the challenge that large vision-language models struggle to distinguish individual instances—such as specific people or objects—thereby limiting their applicability in personalized scenarios. To overcome this limitation, the authors propose an auxiliary visual encoding mechanism that leverages a pre-trained instance-level recognition expert model to provide specialized features to the large vision-language model. This enables the model to perform few-shot, context-aware learning and achieve fine-grained understanding of novel instances from a single example, without requiring extensive instance-specific data or additional training. The approach is the first to enable in-context one-shot instance-level recognition in large vision-language models and supports cross-category instance perception. Evaluated on both existing and newly curated multi-category benchmarks—including faces, persons, pets, and general objects—the method significantly outperforms current state-of-the-art approaches.

0 citationsRead paper
Recent publications

Latest Papers

AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition

Jul 28, 2026

This work addresses the challenges of on-device speech emotion recognition, where large self-supervised models incur prohibitive computational costs, and existing knowledge distillation approaches suffer from unreliable teacher predictions and neglect of inter-sample relational structures. To overcome these limitations, the authors propose an adaptive multi-teacher relational distillation framework. This framework incorporates a one-class SVM–based mechanism to assess teacher reliability and dynamically weight their predictions, while also introducing a relational distillation loss that preserves cross-sample semantic structure by aligning similarity matrices between teachers and students. Experiments on IEMOCAP and CREMA-D demonstrate that four lightweight student models consistently outperform single-teacher baselines, and ablation studies confirm the complementary benefits of the two core components.

0 citationsRead paper

Comment and Control: Hijacking Agentic Workflows via Context-Grounded Evolution

May 11, 2026

This work demonstrates that large language model (LLM) agents integrated into automation platforms are vulnerable to manipulation via malicious inputs—such as GitHub comments—leading to risks including credential leakage and arbitrary command execution. To systematically uncover these novel attack surfaces in agent workflows, the authors propose “Context-Anchored Evolution,” a method that combines static path feasibility analysis, dynamic prompt provenance tracing, and runtime capability assessment within a three-stage contextual framework, enhanced by evolutionary input generation to achieve targeted hijacking of LLM agents. The resulting JAW framework successfully exploits 4,714 GitHub workflows and 8 n8n templates, compromising 15 widely used GitHub Actions and 2 official n8n nodes. Vulnerabilities identified through this approach have been acknowledged and patched by vendors including GitHub, Google, and Anthropic.

0 citationsRead paper

IIR-VLM: In-Context Instance-level Recognition for Large Vision-Language Models

Jan 20, 2026

This work addresses the challenge that large vision-language models struggle to distinguish individual instances—such as specific people or objects—thereby limiting their applicability in personalized scenarios. To overcome this limitation, the authors propose an auxiliary visual encoding mechanism that leverages a pre-trained instance-level recognition expert model to provide specialized features to the large vision-language model. This enables the model to perform few-shot, context-aware learning and achieve fine-grained understanding of novel instances from a single example, without requiring extensive instance-specific data or additional training. The approach is the first to enable in-context one-shot instance-level recognition in large vision-language models and supports cross-category instance perception. Evaluated on both existing and newly curated multi-category benchmarks—including faces, persons, pets, and general objects—the method significantly outperforms current state-of-the-art approaches.

0 citationsRead paper