Institution profile

LightSpeed Studios

Industry researchasia · cn
Official website
Research library17linked papers
Opportunities0open roles
Selected work

Representative Papers

Differentiate the Solver, Not the Equation: Reverse-Sweep Adjoints for Block Implicit Simulation

Aug 09, 2026

This work addresses the high memory overhead and neglect of local structure in gradient computation for implicit nonlinear solvers within differentiable simulation. The authors propose a solver-level differentiation method that constructs an adjoint algorithm symmetric to the forward solve by reverse-scanning a block-structured implicit solver, entirely avoiding the assembly of a global Jacobian matrix. For the first time, adjoint computation is aligned with the block structure of the forward solver, combining vertex-block descent with reverse-colored Gauss–Seidel sweeps to enable efficient backpropagation using only local 3×3 adjoint solves. This approach leverages operator-view approximations of the inverse and its transpose. On a single GPU, it achieves a 33× speedup and 71× reduction in memory compared to unrolled automatic differentiation, enabling, for the first time, differentiable elastic dynamics simulation of million-contact coupled soft bodies with up to 8 million vertices.

0 citationsRead paper

MemoryCPT: An End-to-End Agent Memory Framework for Cost-Performance Trade-off

Aug 05, 2026

This work addresses the challenges of efficient retrieval and context compression faced by long-horizon large language model agents when processing extensive interaction histories, which often lead to high inference costs and performance bottlenecks. The authors propose MemoryCPT, the first end-to-end trainable memory framework that unifies offline memory construction with online query-conditioned context generation. By jointly optimizing Query-Agnostic Distillation (QAD) and Query-Aware Retrieval-summarization (QAR), MemoryCPT enables efficient memory management. The study introduces a novel Quality-per-Cost (QPC) metric and integrates reciprocal rank fusion (RRF), LoRA fine-tuning, and Group Relative Policy Optimization (GRPO). Evaluated on the LoCoMo and LongMemEval benchmarks, MemoryCPT significantly outperforms existing methods by improving response quality while controlling inference cost, with ablation studies confirming the effectiveness of each component.

0 citationsRead paper

MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction

Jul 27, 2026

This work addresses the limitations of existing 3D variational autoencoders (VAEs) in high-fidelity reconstruction, which suffer either from the high computational cost of voxel-based representations or from detail loss due to sparsity and global smoothing in set-based approaches. To overcome these challenges, the authors propose a hierarchical set-based VAE that achieves efficient, high-fidelity reconstruction through progressive densification of anchor-based VecSet latent variables and geometry-aware local decoding. Key innovations include hierarchical point shuffle upsampling for latent densification, an AVS-Conv local aggregation operator replacing global attention, and a multi-scale query decoding mechanism that fuses features across granularities. Experiments demonstrate that the proposed method significantly outperforms state-of-the-art approaches on Objaverse, ABO, and real-world datasets, achieving approximately 10× faster decoding than prior set-based methods and nearly 10× greater model compactness compared to voxel-based baselines.

0 citationsRead paper
Recent publications

Latest Papers

Differentiate the Solver, Not the Equation: Reverse-Sweep Adjoints for Block Implicit Simulation

Aug 09, 2026

This work addresses the high memory overhead and neglect of local structure in gradient computation for implicit nonlinear solvers within differentiable simulation. The authors propose a solver-level differentiation method that constructs an adjoint algorithm symmetric to the forward solve by reverse-scanning a block-structured implicit solver, entirely avoiding the assembly of a global Jacobian matrix. For the first time, adjoint computation is aligned with the block structure of the forward solver, combining vertex-block descent with reverse-colored Gauss–Seidel sweeps to enable efficient backpropagation using only local 3×3 adjoint solves. This approach leverages operator-view approximations of the inverse and its transpose. On a single GPU, it achieves a 33× speedup and 71× reduction in memory compared to unrolled automatic differentiation, enabling, for the first time, differentiable elastic dynamics simulation of million-contact coupled soft bodies with up to 8 million vertices.

0 citationsRead paper

MemoryCPT: An End-to-End Agent Memory Framework for Cost-Performance Trade-off

Aug 05, 2026

This work addresses the challenges of efficient retrieval and context compression faced by long-horizon large language model agents when processing extensive interaction histories, which often lead to high inference costs and performance bottlenecks. The authors propose MemoryCPT, the first end-to-end trainable memory framework that unifies offline memory construction with online query-conditioned context generation. By jointly optimizing Query-Agnostic Distillation (QAD) and Query-Aware Retrieval-summarization (QAR), MemoryCPT enables efficient memory management. The study introduces a novel Quality-per-Cost (QPC) metric and integrates reciprocal rank fusion (RRF), LoRA fine-tuning, and Group Relative Policy Optimization (GRPO). Evaluated on the LoCoMo and LongMemEval benchmarks, MemoryCPT significantly outperforms existing methods by improving response quality while controlling inference cost, with ablation studies confirming the effectiveness of each component.

0 citationsRead paper

MSVS-VAE: Multi-Scale Anchored VecSet for High-Fidelity 3D Reconstruction

Jul 27, 2026

This work addresses the limitations of existing 3D variational autoencoders (VAEs) in high-fidelity reconstruction, which suffer either from the high computational cost of voxel-based representations or from detail loss due to sparsity and global smoothing in set-based approaches. To overcome these challenges, the authors propose a hierarchical set-based VAE that achieves efficient, high-fidelity reconstruction through progressive densification of anchor-based VecSet latent variables and geometry-aware local decoding. Key innovations include hierarchical point shuffle upsampling for latent densification, an AVS-Conv local aggregation operator replacing global attention, and a multi-scale query decoding mechanism that fuses features across granularities. Experiments demonstrate that the proposed method significantly outperforms state-of-the-art approaches on Objaverse, ABO, and real-world datasets, achieving approximately 10× faster decoding than prior set-based methods and nearly 10× greater model compactness compared to voxel-based baselines.

0 citationsRead paper