M$^3$P-R1: Reinforcement Learning for Large Language Model Guided Multi-Modal Motion Planning via MIP Code Generation
研究提出M$^3$P-R1方法,通过强化学习调优大语言模型生成MIP代码,解决多模态运动规划中的连续动作与离散模式转换联合推理难题。
研究提出M$^3$P-R1方法,通过强化学习调优大语言模型生成MIP代码,解决多模态运动规划中的连续动作与离散模式转换联合推理难题。
论文提出GameXpert-Bench,通过三个阶段(游戏生成、修复和优化)评估大语言模型作为编码代理在游戏开发中的表现。
This work addresses the high memory overhead and neglect of local structure in gradient computation for implicit nonlinear solvers within differentiable simulation. The authors propose a solver-level differentiation method that constructs an adjoint algorithm symmetric to the forward solve by reverse-scanning a block-structured implicit solver, entirely avoiding the assembly of a global Jacobian matrix. For the first time, adjoint computation is aligned with the block structure of the forward solver, combining vertex-block descent with reverse-colored Gauss–Seidel sweeps to enable efficient backpropagation using only local 3×3 adjoint solves. This approach leverages operator-view approximations of the inverse and its transpose. On a single GPU, it achieves a 33× speedup and 71× reduction in memory compared to unrolled automatic differentiation, enabling, for the first time, differentiable elastic dynamics simulation of million-contact coupled soft bodies with up to 8 million vertices.
This work addresses the challenges of efficient retrieval and context compression faced by long-horizon large language model agents when processing extensive interaction histories, which often lead to high inference costs and performance bottlenecks. The authors propose MemoryCPT, the first end-to-end trainable memory framework that unifies offline memory construction with online query-conditioned context generation. By jointly optimizing Query-Agnostic Distillation (QAD) and Query-Aware Retrieval-summarization (QAR), MemoryCPT enables efficient memory management. The study introduces a novel Quality-per-Cost (QPC) metric and integrates reciprocal rank fusion (RRF), LoRA fine-tuning, and Group Relative Policy Optimization (GRPO). Evaluated on the LoCoMo and LongMemEval benchmarks, MemoryCPT significantly outperforms existing methods by improving response quality while controlling inference cost, with ablation studies confirming the effectiveness of each component.
This work addresses the limitations of existing 3D variational autoencoders (VAEs) in high-fidelity reconstruction, which suffer either from the high computational cost of voxel-based representations or from detail loss due to sparsity and global smoothing in set-based approaches. To overcome these challenges, the authors propose a hierarchical set-based VAE that achieves efficient, high-fidelity reconstruction through progressive densification of anchor-based VecSet latent variables and geometry-aware local decoding. Key innovations include hierarchical point shuffle upsampling for latent densification, an AVS-Conv local aggregation operator replacing global attention, and a multi-scale query decoding mechanism that fuses features across granularities. Experiments demonstrate that the proposed method significantly outperforms state-of-the-art approaches on Objaverse, ABO, and real-world datasets, achieving approximately 10× faster decoding than prior set-based methods and nearly 10× greater model compactness compared to voxel-based baselines.
研究提出M$^3$P-R1方法,通过强化学习调优大语言模型生成MIP代码,解决多模态运动规划中的连续动作与离散模式转换联合推理难题。
论文提出GameXpert-Bench,通过三个阶段(游戏生成、修复和优化)评估大语言模型作为编码代理在游戏开发中的表现。
This work addresses the high memory overhead and neglect of local structure in gradient computation for implicit nonlinear solvers within differentiable simulation. The authors propose a solver-level differentiation method that constructs an adjoint algorithm symmetric to the forward solve by reverse-scanning a block-structured implicit solver, entirely avoiding the assembly of a global Jacobian matrix. For the first time, adjoint computation is aligned with the block structure of the forward solver, combining vertex-block descent with reverse-colored Gauss–Seidel sweeps to enable efficient backpropagation using only local 3×3 adjoint solves. This approach leverages operator-view approximations of the inverse and its transpose. On a single GPU, it achieves a 33× speedup and 71× reduction in memory compared to unrolled automatic differentiation, enabling, for the first time, differentiable elastic dynamics simulation of million-contact coupled soft bodies with up to 8 million vertices.
This work addresses the challenges of efficient retrieval and context compression faced by long-horizon large language model agents when processing extensive interaction histories, which often lead to high inference costs and performance bottlenecks. The authors propose MemoryCPT, the first end-to-end trainable memory framework that unifies offline memory construction with online query-conditioned context generation. By jointly optimizing Query-Agnostic Distillation (QAD) and Query-Aware Retrieval-summarization (QAR), MemoryCPT enables efficient memory management. The study introduces a novel Quality-per-Cost (QPC) metric and integrates reciprocal rank fusion (RRF), LoRA fine-tuning, and Group Relative Policy Optimization (GRPO). Evaluated on the LoCoMo and LongMemEval benchmarks, MemoryCPT significantly outperforms existing methods by improving response quality while controlling inference cost, with ablation studies confirming the effectiveness of each component.
This work addresses the limitations of existing 3D variational autoencoders (VAEs) in high-fidelity reconstruction, which suffer either from the high computational cost of voxel-based representations or from detail loss due to sparsity and global smoothing in set-based approaches. To overcome these challenges, the authors propose a hierarchical set-based VAE that achieves efficient, high-fidelity reconstruction through progressive densification of anchor-based VecSet latent variables and geometry-aware local decoding. Key innovations include hierarchical point shuffle upsampling for latent densification, an AVS-Conv local aggregation operator replacing global attention, and a multi-scale query decoding mechanism that fuses features across granularities. Experiments demonstrate that the proposed method significantly outperforms state-of-the-art approaches on Objaverse, ABO, and real-world datasets, achieving approximately 10× faster decoding than prior set-based methods and nearly 10× greater model compactness compared to voxel-based baselines.