Institution profile

Guangdong Institute of Intelligence Science and Technology

Academic institutionasia · cn
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents

Jun 03, 2026

This work addresses the limitation of existing embodied language model agents, which respond only to literal queries and fail to capture users’ implicit intentions—such as inferring availability or emotional state from a question like “Where is Lin Wei?” To bridge this gap, the authors propose IntentFrame, a framework that introduces structured intent reasoning between scene perception and tool invocation to explicitly model latent user needs. IntentFrame incorporates a gap-calibration mechanism that dynamically allocates exploration budgets and selects appropriate tools. This approach enables the first targeted probing for implicit intents without relying on memorized answers. Evaluated on a benchmark of 100 queries across four scenarios, IntentFrame improves implicit intent coverage by 0.07 over ReAct (p < 10⁻⁶), reduces probing actions by 82%, and entirely eliminates privacy-violating tool calls.

0 citationsRead paper

TCM-DiffRAG: Personalized Syndrome Differentiation Reasoning Method for Traditional Chinese Medicine based on Knowledge Graph and Chain of Thought

Feb 26, 2026

This work addresses the limitations of conventional retrieval-augmented generation (RAG) approaches in handling the complex reasoning and individualized variability inherent in traditional Chinese medicine (TCM) syndrome differentiation and treatment. To bridge this gap, the authors propose an enhanced RAG framework that integrates a structured TCM knowledge graph with chain-of-thought (CoT) reasoning, achieving, for the first time, effective alignment between general TCM knowledge and personalized clinical inference. By synergistically combining knowledge graphs, CoT prompting, RAG, and large language models, the proposed method significantly outperforms native large language models, supervised fine-tuned models, and other RAG baselines across multiple TCM datasets. Notably, it substantially improves the performance of non-Chinese large language models on TCM-specific tasks, demonstrating its effectiveness in contextualizing domain-specific reasoning within a linguistically diverse setting.

0 citationsRead paper

PhysDrape: Learning Explicit Forces and Collision Constraints for Physically Realistic Garment Draping

Feb 08, 2026

This work addresses the limitations of existing deep learning approaches for garment draping simulation, which rely on soft constraints to handle collisions and often fail to simultaneously ensure geometric feasibility and physical plausibility, leading to mesh distortion or body penetration. To overcome these issues, we propose PhysDrape, a hybrid framework that integrates a physics-informed graph neural network with a differentiable explicit physical solver. Our method introduces, for the first time, a two-stage differentiable solving pipeline—combining force equilibrium optimization with hard projection constraints—and incorporates the Saint Venant–Kirchhoff material model to strictly enforce non-penetration and quasi-static equilibrium during end-to-end training. Experiments demonstrate that PhysDrape significantly reduces strain energy, nearly eliminates body penetration, and achieves superior physical fidelity and real-time robustness compared to current state-of-the-art methods.

0 citationsRead paper

Agri-R1: Empowering Generalizable Agricultural Reasoning in Vision-Language Models with Reinforcement Learning

Jan 08, 2026arXiv.org

This study addresses the limitations of existing vision-language models in agricultural disease diagnosis—namely, their reliance on strong annotations, poor interpretability, and weak generalization in open-ended scenarios. The authors propose a novel method that automatically generates reasoning data without manual labeling by integrating vision-language synthesis with large language model filtering, constructing a high-quality training set using only 19% of the original samples. They further introduce a new reward function combining domain-specific lexicons and fuzzy matching, enabling structured reasoning through Group Relative Policy Optimization (GRPO). Evaluated on CDDMBench, their 3B-parameter model substantially outperforms 7B–13B baselines, achieving a 23.2% gain in disease identification accuracy, a 33.3% improvement in agricultural question answering, and a 26.10-point increase in cross-domain generalization.

0 citationsRead paper
Recent publications

Latest Papers

AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents

Jun 03, 2026

This work addresses the limitation of existing embodied language model agents, which respond only to literal queries and fail to capture users’ implicit intentions—such as inferring availability or emotional state from a question like “Where is Lin Wei?” To bridge this gap, the authors propose IntentFrame, a framework that introduces structured intent reasoning between scene perception and tool invocation to explicitly model latent user needs. IntentFrame incorporates a gap-calibration mechanism that dynamically allocates exploration budgets and selects appropriate tools. This approach enables the first targeted probing for implicit intents without relying on memorized answers. Evaluated on a benchmark of 100 queries across four scenarios, IntentFrame improves implicit intent coverage by 0.07 over ReAct (p < 10⁻⁶), reduces probing actions by 82%, and entirely eliminates privacy-violating tool calls.

0 citationsRead paper

TCM-DiffRAG: Personalized Syndrome Differentiation Reasoning Method for Traditional Chinese Medicine based on Knowledge Graph and Chain of Thought

Feb 26, 2026

This work addresses the limitations of conventional retrieval-augmented generation (RAG) approaches in handling the complex reasoning and individualized variability inherent in traditional Chinese medicine (TCM) syndrome differentiation and treatment. To bridge this gap, the authors propose an enhanced RAG framework that integrates a structured TCM knowledge graph with chain-of-thought (CoT) reasoning, achieving, for the first time, effective alignment between general TCM knowledge and personalized clinical inference. By synergistically combining knowledge graphs, CoT prompting, RAG, and large language models, the proposed method significantly outperforms native large language models, supervised fine-tuned models, and other RAG baselines across multiple TCM datasets. Notably, it substantially improves the performance of non-Chinese large language models on TCM-specific tasks, demonstrating its effectiveness in contextualizing domain-specific reasoning within a linguistically diverse setting.

0 citationsRead paper

PhysDrape: Learning Explicit Forces and Collision Constraints for Physically Realistic Garment Draping

Feb 08, 2026

This work addresses the limitations of existing deep learning approaches for garment draping simulation, which rely on soft constraints to handle collisions and often fail to simultaneously ensure geometric feasibility and physical plausibility, leading to mesh distortion or body penetration. To overcome these issues, we propose PhysDrape, a hybrid framework that integrates a physics-informed graph neural network with a differentiable explicit physical solver. Our method introduces, for the first time, a two-stage differentiable solving pipeline—combining force equilibrium optimization with hard projection constraints—and incorporates the Saint Venant–Kirchhoff material model to strictly enforce non-penetration and quasi-static equilibrium during end-to-end training. Experiments demonstrate that PhysDrape significantly reduces strain energy, nearly eliminates body penetration, and achieves superior physical fidelity and real-time robustness compared to current state-of-the-art methods.

0 citationsRead paper

Agri-R1: Empowering Generalizable Agricultural Reasoning in Vision-Language Models with Reinforcement Learning

Jan 08, 2026arXiv.org

This study addresses the limitations of existing vision-language models in agricultural disease diagnosis—namely, their reliance on strong annotations, poor interpretability, and weak generalization in open-ended scenarios. The authors propose a novel method that automatically generates reasoning data without manual labeling by integrating vision-language synthesis with large language model filtering, constructing a high-quality training set using only 19% of the original samples. They further introduce a new reward function combining domain-specific lexicons and fuzzy matching, enabling structured reasoning through Group Relative Policy Optimization (GRPO). Evaluated on CDDMBench, their 3B-parameter model substantially outperforms 7B–13B baselines, achieving a 23.2% gain in disease identification accuracy, a 33.3% improvement in agricultural question answering, and a 26.10-point increase in cross-domain generalization.

0 citationsRead paper