Institution profile

InstaDeep

Industry researcheurope · uk
Official website
Research library17linked papers
Opportunities0open roles
Selected work

Representative Papers

Overconfident Oracles: Limitations of In Silico Sequence Design Benchmarking

Feb 24, 2025

This paper identifies a critical unreliability in machine learning–based oracle benchmarks for biological sequence design: over 70% of methods exhibit rank reversal across different oracle models—due to architectural or training stochasticity—revealing a fundamental deficiency in out-of-distribution generalization. Method: The authors systematically characterize the detrimental impact of oracle inconsistency on benchmark validity and propose a hybrid evaluation framework integrating multiple biophysically grounded metrics—including stability, foldability, and solubility—while constraining the oracle’s scoring domain to enhance design robustness. Contribution/Results: Through large-scale reproduction of 12 state-of-the-art design methods, cross-oracle consistency analysis, and out-of-distribution generalization diagnostics, the framework significantly improves wet-lab validation rates. It establishes a new paradigm for constructing trustworthy, AI-driven benchmarks for protein and nucleic acid design—grounded in empirical feasibility and physical realism.

2 citations1 influentialRead paper

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation

May 13, 2026

This work addresses the limitations of existing contrastive reinforcement learning methods, which are largely confined to off-policy settings and continuous action spaces, and lack effective self-supervised mechanisms for on-policy frameworks, discrete actions, and multi-agent scenarios. The paper proposes Contrastive Proximal Policy Optimization (CPPO), the first approach to integrate contrastive learning directly with on-policy optimization. CPPO constructs Q-values and derives policy advantages through contrastive representations of state-action-goal triplets, eliminating the need for handcrafted rewards or experience replay. The method unifies support for both continuous and discrete action spaces as well as single- and multi-agent settings. Evaluated across 18 tasks, CPPO significantly outperforms existing contrastive RL methods in 14 and matches or exceeds the performance of reward-intensive PPO in 12.

0 citationsRead paper

Force-Aware Neural Tangent Kernels for Scalable and Robust Active Learning of MLIPs

May 13, 2026

This work addresses key challenges in active learning for machine-learned interatomic potentials (MLIPs)—namely, the large scale of candidate pools, the need to jointly leverage energy and force supervision signals, and poor robustness under distribution shift—by proposing a linearly scalable acquisition framework. The method extends the neural tangent kernel (NTK) to force-aware settings for the first time, introducing force-NTK and joint energy–force NTK as natural similarity measures for vector field prediction. By integrating block-wise feature-space posterior variance filtering with embeddings from a pretrained MLIP, it avoids explicit computation of large kernel matrices while enhancing both efficiency and robustness. Experiments demonstrate that the approach achieves the lowest energy and force errors on OC20, matches or exceeds state-of-the-art performance on T1x, PMechDB, and RGD benchmarks with greater efficiency, and significantly outperforms ensemble-based methods under distribution shift in the candidate pool.

0 citationsRead paper

Pretrained Model Representations as Acquisition Signals for Active Learning of MLIPs

May 05, 2026

This work addresses the challenge of inefficient training of machine-learned interatomic potentials (MLIPs) in reaction chemistry, where high-cost quantum chemical labels and scarcity of transition-state configurations severely limit data availability. The authors propose a novel active learning strategy that eschews auxiliary uncertainty modules, Bayesian training, or ensemble methods, instead leveraging the latent representations of a pretrained MACE-based MLIP to construct an acquisition function guided by the finite-width neural tangent kernel (NTK) and activation kernel. They demonstrate for the first time that the latent space of a pretrained MLIP inherently provides efficient and reliable acquisition signals, whose geometric structure aligns with model error while preserving chemical interpretability. On reaction chemistry benchmarks, the method reduces by 38% and 28%, respectively, the amount of data required to achieve baseline energy and force errors, substantially improving data efficiency and uncertainty estimation.

0 citationsRead paper
Recent publications

Latest Papers

Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation

May 13, 2026

This work addresses the limitations of existing contrastive reinforcement learning methods, which are largely confined to off-policy settings and continuous action spaces, and lack effective self-supervised mechanisms for on-policy frameworks, discrete actions, and multi-agent scenarios. The paper proposes Contrastive Proximal Policy Optimization (CPPO), the first approach to integrate contrastive learning directly with on-policy optimization. CPPO constructs Q-values and derives policy advantages through contrastive representations of state-action-goal triplets, eliminating the need for handcrafted rewards or experience replay. The method unifies support for both continuous and discrete action spaces as well as single- and multi-agent settings. Evaluated across 18 tasks, CPPO significantly outperforms existing contrastive RL methods in 14 and matches or exceeds the performance of reward-intensive PPO in 12.

0 citationsRead paper

Force-Aware Neural Tangent Kernels for Scalable and Robust Active Learning of MLIPs

May 13, 2026

This work addresses key challenges in active learning for machine-learned interatomic potentials (MLIPs)—namely, the large scale of candidate pools, the need to jointly leverage energy and force supervision signals, and poor robustness under distribution shift—by proposing a linearly scalable acquisition framework. The method extends the neural tangent kernel (NTK) to force-aware settings for the first time, introducing force-NTK and joint energy–force NTK as natural similarity measures for vector field prediction. By integrating block-wise feature-space posterior variance filtering with embeddings from a pretrained MLIP, it avoids explicit computation of large kernel matrices while enhancing both efficiency and robustness. Experiments demonstrate that the approach achieves the lowest energy and force errors on OC20, matches or exceeds state-of-the-art performance on T1x, PMechDB, and RGD benchmarks with greater efficiency, and significantly outperforms ensemble-based methods under distribution shift in the candidate pool.

0 citationsRead paper

Pretrained Model Representations as Acquisition Signals for Active Learning of MLIPs

May 05, 2026

This work addresses the challenge of inefficient training of machine-learned interatomic potentials (MLIPs) in reaction chemistry, where high-cost quantum chemical labels and scarcity of transition-state configurations severely limit data availability. The authors propose a novel active learning strategy that eschews auxiliary uncertainty modules, Bayesian training, or ensemble methods, instead leveraging the latent representations of a pretrained MACE-based MLIP to construct an acquisition function guided by the finite-width neural tangent kernel (NTK) and activation kernel. They demonstrate for the first time that the latent space of a pretrained MLIP inherently provides efficient and reliable acquisition signals, whose geometric structure aligns with model error while preserving chemical interpretability. On reaction chemistry benchmarks, the method reduces by 38% and 28%, respectively, the amount of data required to achieve baseline energy and force errors, substantially improving data efficiency and uncertainty estimation.

0 citationsRead paper

Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective

Apr 27, 2026

In weak-to-strong alignment, strong models often confidently err on samples lying in the blind spots of weak teachers, leading to alignment failure. This work addresses this issue by analyzing the problem through the lens of bias–variance–covariance decomposition, integrating mismatch theory with practical post-training pipelines. The authors derive a mismatch-based upper bound on the overall risk and introduce a “blind-spot deception” metric to characterize the mechanism of alignment breakdown. Empirical evaluation across SFT, RLHF, and RLAIF on PKU-SafeRLHF and HH-RLHF datasets reveals that the variance of the strong model is the strongest predictor of blind-spot deception across settings, with covariance providing supplementary signals. Furthermore, blind-spot assessment effectively distinguishes whether alignment failures stem from inherited weak supervision or from regions of high uncertainty in the weak model.

0 citationsRead paper