Institution profile

Lionrock AI Lab

Industry researchnorthamerica · us
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

PVI: Plug-in Visual Injection for Vision-Language-Action Models

Mar 13, 2026

Existing vision-language-action models struggle with multi-stage manipulation tasks due to their reliance on semantically abstracted pretrained vision models, which often neglect geometric details and lack explicit temporal modeling. To address this, this work proposes a lightweight, encoder-agnostic, plug-and-play module that injects video-level visual representations—such as those from V-JEPA2 or DINOv2—into flow-matching action experts via a zero-initialized residual path. This enables single-stage fine-tuning without modifying the backbone architecture. The approach provides the first empirical validation that video-level features significantly outperform static image features in long-horizon manipulation tasks. Consistent performance gains are demonstrated on both simulated and real-world dual-arm cloth-folding benchmarks, with particularly pronounced improvements in multi-stage scenarios requiring state tracking.

0 citationsRead paper

Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent

Sep 22, 2025

Existing GUI agents face two critical challenges: unreliable reward signals and limited capacity for online trajectory generation—leading to unreliable reasoning and low data efficiency. To address these, we propose Orcust, a novel framework introducing two key mechanisms: (1) Principle-Constrained Reward Modeling (PCRM), which incorporates interpretable, human-aligned rules into the reinforcement learning reward function to ensure policy adherence to domain principles; and (2) Online Virtual Machine-Grounded Trajectory Construction (OVTC), which leverages a lightweight virtual machine to autonomously generate high-quality, structured interaction trajectories enabling fine-grained, stepwise training. Integrating chain-of-thought reasoning with LLM-driven rule feedback, Orcust achieves +22.2% and +23.9% improvements on ScreenSpot and ScreenSpot-Pro benchmarks, respectively. The framework significantly enhances reasoning reliability, task adaptability, and system scalability.

0 citationsRead paper

L0: Reinforcement Learning to Become General Agents

Jun 30, 2025

To address scalability and training efficiency bottlenecks in deploying large language models (LLMs) as autonomous agents for multi-turn, long-horizon tasks, this paper introduces L0—a fully end-to-end reinforcement learning framework. Methodologically, L0 features: (1) NB-Agent, an agent architecture adopting a “code-as-action” REPL execution paradigm; (2) Reinforcement Learning with Verifiable Rewards (RLVR), which directly elicits problem-solving capabilities from base models without supervised fine-tuning; and (3) a lightweight sandboxed concurrent agent pool enabling high-throughput, low-cost environment interaction. Evaluated on Qwen2.5-7B-Instruct, L0 achieves substantial improvements: SimpleQA accuracy rises from 30% to 80%, and HotpotQA from 22% to 41%. The framework is open-sourced, establishing a novel paradigm for scalable, efficient training of LLM-based autonomous agents.

0 citationsRead paper
Recent publications

Latest Papers

PVI: Plug-in Visual Injection for Vision-Language-Action Models

Mar 13, 2026

Existing vision-language-action models struggle with multi-stage manipulation tasks due to their reliance on semantically abstracted pretrained vision models, which often neglect geometric details and lack explicit temporal modeling. To address this, this work proposes a lightweight, encoder-agnostic, plug-and-play module that injects video-level visual representations—such as those from V-JEPA2 or DINOv2—into flow-matching action experts via a zero-initialized residual path. This enables single-stage fine-tuning without modifying the backbone architecture. The approach provides the first empirical validation that video-level features significantly outperform static image features in long-horizon manipulation tasks. Consistent performance gains are demonstrated on both simulated and real-world dual-arm cloth-folding benchmarks, with particularly pronounced improvements in multi-stage scenarios requiring state tracking.

0 citationsRead paper

Orcust: Stepwise-Feedback Reinforcement Learning for GUI Agent

Sep 22, 2025

Existing GUI agents face two critical challenges: unreliable reward signals and limited capacity for online trajectory generation—leading to unreliable reasoning and low data efficiency. To address these, we propose Orcust, a novel framework introducing two key mechanisms: (1) Principle-Constrained Reward Modeling (PCRM), which incorporates interpretable, human-aligned rules into the reinforcement learning reward function to ensure policy adherence to domain principles; and (2) Online Virtual Machine-Grounded Trajectory Construction (OVTC), which leverages a lightweight virtual machine to autonomously generate high-quality, structured interaction trajectories enabling fine-grained, stepwise training. Integrating chain-of-thought reasoning with LLM-driven rule feedback, Orcust achieves +22.2% and +23.9% improvements on ScreenSpot and ScreenSpot-Pro benchmarks, respectively. The framework significantly enhances reasoning reliability, task adaptability, and system scalability.

0 citationsRead paper

L0: Reinforcement Learning to Become General Agents

Jun 30, 2025

To address scalability and training efficiency bottlenecks in deploying large language models (LLMs) as autonomous agents for multi-turn, long-horizon tasks, this paper introduces L0—a fully end-to-end reinforcement learning framework. Methodologically, L0 features: (1) NB-Agent, an agent architecture adopting a “code-as-action” REPL execution paradigm; (2) Reinforcement Learning with Verifiable Rewards (RLVR), which directly elicits problem-solving capabilities from base models without supervised fine-tuning; and (3) a lightweight sandboxed concurrent agent pool enabling high-throughput, low-cost environment interaction. Evaluated on Qwen2.5-7B-Instruct, L0 achieves substantial improvements: SimpleQA accuracy rises from 30% to 80%, and HotpotQA from 22% to 41%. The framework is open-sourced, establishing a novel paradigm for scalable, efficient training of LLM-based autonomous agents.

0 citationsRead paper