Institution profile

Cradle

Industry research
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

EvoFlows: Evolutionary Edit-Based Flow-Matching for Protein Engineering

Mar 12, 2026

Traditional protein engineering approaches struggle to perform controllable and non-trivial sequence edits on template proteins while preserving their native-like properties. This work proposes a variable-length sequence-to-sequence modeling framework based on edit flows, which uniquely integrates edit operations with flow matching to jointly predict both the location and type of mutations by learning evolutionary trajectories among related proteins. By combining evolutionary information with a controllable editing mechanism, the method diverges from conventional autoregressive or masked language models. Evaluated on UniRef and OAS datasets, it demonstrates sequence distribution modeling capabilities comparable to state-of-the-art masked language models while significantly improving the generation of natural yet structurally novel protein variants.

0 citationsRead paper

g-DPO: Scalable Preference Optimization for Protein Language Models

Oct 22, 2025

DPO faces scalability bottlenecks in aligning protein language models with experimental design objectives: the number of training preference pairs grows quadratically with sequence count, rendering training prohibitively expensive even on moderate-scale datasets. To address this, we propose Cluster-DPO—the first efficient preference optimization framework tailored to protein sequence space. It reduces redundancy by clustering sequence embeddings to prune non-informative preference pairs and introduces intra-cluster likelihood amortization to substantially lower computational overhead while preserving gradient consistency. Evaluated on three protein engineering tasks—enzyme activity, thermostability, and binding affinity—Cluster-DPO achieves in vitro and in vivo performance comparable to standard DPO, with 1.8–3.7× faster training. Crucially, speedup scales favorably with dataset size. Cluster-DPO thus establishes a scalable new paradigm for aligning large-scale protein language models with experimental objectives.

0 citationsRead paper
Recent publications

Latest Papers

EvoFlows: Evolutionary Edit-Based Flow-Matching for Protein Engineering

Mar 12, 2026

Traditional protein engineering approaches struggle to perform controllable and non-trivial sequence edits on template proteins while preserving their native-like properties. This work proposes a variable-length sequence-to-sequence modeling framework based on edit flows, which uniquely integrates edit operations with flow matching to jointly predict both the location and type of mutations by learning evolutionary trajectories among related proteins. By combining evolutionary information with a controllable editing mechanism, the method diverges from conventional autoregressive or masked language models. Evaluated on UniRef and OAS datasets, it demonstrates sequence distribution modeling capabilities comparable to state-of-the-art masked language models while significantly improving the generation of natural yet structurally novel protein variants.

0 citationsRead paper

g-DPO: Scalable Preference Optimization for Protein Language Models

Oct 22, 2025

DPO faces scalability bottlenecks in aligning protein language models with experimental design objectives: the number of training preference pairs grows quadratically with sequence count, rendering training prohibitively expensive even on moderate-scale datasets. To address this, we propose Cluster-DPO—the first efficient preference optimization framework tailored to protein sequence space. It reduces redundancy by clustering sequence embeddings to prune non-informative preference pairs and introduces intra-cluster likelihood amortization to substantially lower computational overhead while preserving gradient consistency. Evaluated on three protein engineering tasks—enzyme activity, thermostability, and binding affinity—Cluster-DPO achieves in vitro and in vivo performance comparable to standard DPO, with 1.8–3.7× faster training. Crucially, speedup scales favorably with dataset size. Cluster-DPO thus establishes a scalable new paradigm for aligning large-scale protein language models with experimental objectives.

0 citationsRead paper