Institution profile

Altos Labs

Industry researchnorthamerica · us
Official website
Research library5linked papers
Opportunities0open roles
Selected work

Representative Papers

VIScore: Diagnosing Planning-Relevant Quality in Latent World Models

Aug 11, 2026

Existing evaluation metrics struggle to effectively link latent space characteristics with planning success, particularly lacking diagnostic capability on out-of-distribution data. This work proposes VIScore, the first unified, quantifiable metric that jointly models the three stages of encoding, prediction, and search-based planning through three dimensions: Veracity, Influence, and Sobriety, comprehensively assessing a world model’s support for planning tasks. Experimental results demonstrate that VIScore achieves a Spearman correlation exceeding 0.75 across a cross-task success pool and exhibits significantly lower calibration error than constant fitting baselines. It attains state-of-the-art performance on both seen and unseen models and datasets, thereby transcending traditional evaluation paradigms that focus solely on latent representations.

0 citationsRead paper

VISReg: Variance-Invariance-Sketching Regularization for JEPA training

Jun 01, 2026

This work addresses the challenge in self-supervised learning of simultaneously modeling the full shape of embedding distributions and maintaining training stability while preventing representation collapse. The authors propose VISReg, which decouples distributional shape constraints from scale control for the first time: it employs the sliced Wasserstein distance instead of covariance regularization to capture the distribution shape, complemented by an independent variance term to regulate scale. By integrating the flexibility of VICReg with the distributional rigor of sketching, VISReg effectively mitigates collapse and provides stable gradients. Evaluated within the JEPA framework and pretrained on ImageNet-1K and ImageNet-22K, VISReg achieves state-of-the-art performance across out-of-distribution, low-quality, long-tailed, and low-rank data settings.

0 citationsRead paper

MedGym:A Unified Continuous-Time Benchmark for Dynamic Medical Treatment Reinforcement Learning

May 31, 2026

Existing reinforcement learning (RL) approaches struggle to address key challenges in clinical settings, including the continuous-time dynamics of patient physiology, irregular intervention intervals, and heterogeneous individual treatment responses. To bridge this gap, this work proposes MedGym, the first continuous-time unified benchmark for medical RL. MedGym leverages physics-informed neural networks to construct configurable simulation environments from real-world clinical data, enabling rigorous evaluation of both offline and online RL algorithms. The framework natively supports irregular observation and intervention schedules, facilitates personalized treatment recommendations, allows trajectory-level safety assessment, and enables systematic analysis of the performance discrepancy between offline training and online deployment. By incorporating richer temporal dynamics and clinical realism, MedGym provides a more faithful and informative evaluation platform for advancing safe and effective RL applications in healthcare.

0 citationsRead paper

Interaction-Limited Safe Continuous-Time RL for Dynamical Medical Treatment

May 31, 2026

This work addresses the challenge of dynamic medical treatment, which requires joint optimization of treatment intensity and interaction timing. Existing approaches often rely on fixed interaction intervals or enforce safety only at discrete time points, failing to account for continuous state evolution and intermediate risks. The authors formulate the problem as an options-based semi-Markov decision process with trajectory-level safety constraints, where each option comprises a continuous-time treatment policy and its duration. Key contributions include a safety tightening mechanism that provably ensures trajectory-wide safety with high probability by imposing appropriate constraints at interaction times, a finite-sample policy learning theory grounded in logged data, and a data-driven conservative surrogate method. Experiments demonstrate that the proposed adaptive interaction mechanism significantly outperforms fixed-interval strategies across multiple safety policies, enhancing both treatment safety and efficacy.

0 citationsRead paper

PRiMeFlow: Capturing Complex Expression Heterogeneity in Perturbation Response Modelling

Apr 15, 2026

Modeling cellular responses to perturbations is highly challenging due to the heterogeneity of single-cell gene expression and complex implicit dependencies among genes. To address this, this work introduces, for the first time, a flow matching framework into the field, proposing an end-to-end method that directly models the effects of genetic and small-molecule perturbations on cell states within the native gene expression space. The approach employs a U-Net architecture to parameterize the velocity field and achieves high-fidelity predictions by fitting single-cell expression distributions. Evaluated on the PerturBench benchmark, the model demonstrates superior performance and was awarded first place in the general track of the inaugural ARC Virtual Cell Challenge, confirming its effectiveness and state-of-the-art capability in capturing both expression heterogeneity and perturbation-induced changes.

0 citationsRead paper
Recent publications

Latest Papers

VIScore: Diagnosing Planning-Relevant Quality in Latent World Models

Aug 11, 2026

Existing evaluation metrics struggle to effectively link latent space characteristics with planning success, particularly lacking diagnostic capability on out-of-distribution data. This work proposes VIScore, the first unified, quantifiable metric that jointly models the three stages of encoding, prediction, and search-based planning through three dimensions: Veracity, Influence, and Sobriety, comprehensively assessing a world model’s support for planning tasks. Experimental results demonstrate that VIScore achieves a Spearman correlation exceeding 0.75 across a cross-task success pool and exhibits significantly lower calibration error than constant fitting baselines. It attains state-of-the-art performance on both seen and unseen models and datasets, thereby transcending traditional evaluation paradigms that focus solely on latent representations.

0 citationsRead paper

VISReg: Variance-Invariance-Sketching Regularization for JEPA training

Jun 01, 2026

This work addresses the challenge in self-supervised learning of simultaneously modeling the full shape of embedding distributions and maintaining training stability while preventing representation collapse. The authors propose VISReg, which decouples distributional shape constraints from scale control for the first time: it employs the sliced Wasserstein distance instead of covariance regularization to capture the distribution shape, complemented by an independent variance term to regulate scale. By integrating the flexibility of VICReg with the distributional rigor of sketching, VISReg effectively mitigates collapse and provides stable gradients. Evaluated within the JEPA framework and pretrained on ImageNet-1K and ImageNet-22K, VISReg achieves state-of-the-art performance across out-of-distribution, low-quality, long-tailed, and low-rank data settings.

0 citationsRead paper

MedGym:A Unified Continuous-Time Benchmark for Dynamic Medical Treatment Reinforcement Learning

May 31, 2026

Existing reinforcement learning (RL) approaches struggle to address key challenges in clinical settings, including the continuous-time dynamics of patient physiology, irregular intervention intervals, and heterogeneous individual treatment responses. To bridge this gap, this work proposes MedGym, the first continuous-time unified benchmark for medical RL. MedGym leverages physics-informed neural networks to construct configurable simulation environments from real-world clinical data, enabling rigorous evaluation of both offline and online RL algorithms. The framework natively supports irregular observation and intervention schedules, facilitates personalized treatment recommendations, allows trajectory-level safety assessment, and enables systematic analysis of the performance discrepancy between offline training and online deployment. By incorporating richer temporal dynamics and clinical realism, MedGym provides a more faithful and informative evaluation platform for advancing safe and effective RL applications in healthcare.

0 citationsRead paper

Interaction-Limited Safe Continuous-Time RL for Dynamical Medical Treatment

May 31, 2026

This work addresses the challenge of dynamic medical treatment, which requires joint optimization of treatment intensity and interaction timing. Existing approaches often rely on fixed interaction intervals or enforce safety only at discrete time points, failing to account for continuous state evolution and intermediate risks. The authors formulate the problem as an options-based semi-Markov decision process with trajectory-level safety constraints, where each option comprises a continuous-time treatment policy and its duration. Key contributions include a safety tightening mechanism that provably ensures trajectory-wide safety with high probability by imposing appropriate constraints at interaction times, a finite-sample policy learning theory grounded in logged data, and a data-driven conservative surrogate method. Experiments demonstrate that the proposed adaptive interaction mechanism significantly outperforms fixed-interval strategies across multiple safety policies, enhancing both treatment safety and efficacy.

0 citationsRead paper

PRiMeFlow: Capturing Complex Expression Heterogeneity in Perturbation Response Modelling

Apr 15, 2026

Modeling cellular responses to perturbations is highly challenging due to the heterogeneity of single-cell gene expression and complex implicit dependencies among genes. To address this, this work introduces, for the first time, a flow matching framework into the field, proposing an end-to-end method that directly models the effects of genetic and small-molecule perturbations on cell states within the native gene expression space. The approach employs a U-Net architecture to parameterize the velocity field and achieves high-fidelity predictions by fitting single-cell expression distributions. Evaluated on the PerturBench benchmark, the model demonstrates superior performance and was awarded first place in the general track of the inaugural ARC Virtual Cell Challenge, confirming its effectiveness and state-of-the-art capability in capturing both expression heterogeneity and perturbation-induced changes.

0 citationsRead paper