Novel model for sampling the structural space of proteins - applied to nanobodies.
Protein structures are encoded into latent tokens using a discrete variational auto-encoder (dVAE), which captures residue-level local geometric features in a roto-translation invariant manner.
The model combines a dVAE encoder-decoder with a Structure Language Model (SLM) to model sequence-to-structure relationships.
One can sample alternate conformations by providing an amino acid sequence as input and use the SLM to generate latent tokens representing potential conformations. Decode these latent tokens back into 3D structures using the dVAE decoder to generate diverse structural ensembles.
Altogether, though they develop ESMdiff in the context of the paper, it is rather the structural language in the context of the broader framework of language models that forms the core of the message, rather than any single method.
The model was evaluated on tasks like generating equilibrium dynamics, fold-switching, and intrinsically disordered proteins, using metrics like JS-divergence, TM-score, and RMSD. It outperformed existing methods in speed and accuracy, generating structures 20-100× faster.
An evolution of the ESM model family that scales up in terms of parameters, data, and computational power compared to ESM2, which allows it to improve on sequence, structure, and function representations of proteins.
ESM3 is a multimodal, bidirectional transformer that models sequence, structure, and function using discrete token representations for each modality. It merges these representations into a single latent space and is trained with a masked language model objective, allowing it to generate and predict across different modalities.
ESM3's largest model has 98 billion parameters.
The model was trained with 1.07 × 10²⁴ floating point operations (FLOPs) over a dataset of 771 billion tokens from 2.78 billion proteins.
Structural tokens in ESM3 are encoded by a discrete autoencoder that compresses three-dimensional protein structures into a sequence of discrete tokens. This is done by encoding local atomic environments around each amino acid and representing them in a simplified form that captures geometric properties.
There are a total of 4096 structural tokens to be had.
The structural autoencoder tokenizes protein structures by encoding local neighborhoods around each amino acid into discrete tokens. It uses a geometric attention mechanism that operates in local reference frames, based on bond geometry. This mechanism encodes and reconstructs the atomic structure, supervised by a geometric loss that preserves distances and orientations of bonds and atoms.
ESM3 can be used to generate novel sequences/proteins. generates protein sequences and structures by prompting the model with sequence or structural tokens. It uses iterative sampling, starting from a masked context, where tokens are predicted and unmasked progressively until a full sequence or structure is generated. This allows the model to create novel proteins that respect the given prompts or constraints.
The model was verified experimentally by generating novel proteins, including a green fluorescent protein, that was synthesized and tested for fluorescence in laboratory conditions. The novel protein had a sequence identity of 58% to the nearest known fluorescent protein.
Study evaluates a number of generative models on datasets of antibodies with reported affinities.
The methods tested were: MEAN, dyMEAN, IgBLEND, Ablang, Ablang2, AntiBerty, ESM, Antifold, ESM-IF, AbX, Diffab + their own version of Diffab.
Datasets used were the Absci HER2 datasets (100s of binders) and a number of datasets with tens of binders each.
All models have some correlation with the affinity data, though weak.
Adding epitope information is not a game changer, showing that information that is mostly captured is fitness of antibody first and antigen second, if at all.
Employing structural information helps as compared to purely sequence approaches.
Authors describe how using a structure predictor one can re-design the binding site, to maintain binding.
They use a proprietary GaluxDesign method, the method achieves 1.4 Å Ca RMSD in predicting CDR-H3 loop structures, leveraging a unique scoring metric (G-pass rate) that assesses both confidence and structural consistency for antibody design.
The method outperforms AlphaFold 2.3, ABlooper, and ImmuneBuilder in predicting CDR-H3 loop structures, with significantly lower RMSD values (1.4 Å compared to 2.4-3.7 Å), particularly on a more challenging, time-separated dataset.
The binding propensity to HER2 was evaluated using a large mutant library and calculated via the G-pass rate, outperforming AlphaFold's PAE-based scoring. The model showed strong discrimination with an AUROC of 0.758, compared to 0.529 for AlphaFold. The novel loop is scored using their metric (G-pass rate) in complex with Her2.
Novel antibody sequences were designed by predicting six CDR loops in antibody-protein complexes, using GaluxDesign models. These designs were experimentally tested, achieving high success rates, including a 13.2% success rate for HER2 antibody designs using yeast display methods.
Authors demonstrate that using scores from DeepAb one can sort mutations in an antibody that improve affinity and a series of other properties.
The authors used the DeepAb structure prediction mode model to rank mutations based on their impact on structure prediction confidence, leading to the design of 200 novel anti-hen egg lysozyme (HEL) antibody variants.
Single-point mutations from a deep mutational scanning (DMS) dataset (Warszawski et al.) were combined into multi-mutation variants (up to 7 mutations), and these variants were selected based on DeepAb scores for experimental testing.
The designed variants were expressed and tested for thermostability, colloidal stability, and binding affinity to HEL.
Large percentage of the variants showed improved thermostability (91%) and affinity (94%), with 10% showing significant increases in binding affinity.
A subset of 27 high-performing variants was further tested for developability characteristics, including nonspecific binding, aggregation propensity, and self-association, ensuring their practical usability.
Novel generative model for antibody sequences that supports Vh/Vl pairing and generation of developable sequences.
Three models were created, IgGen (unpaired model), p-IgGen (unpaired fine-tuned on pairs) and developable p-IgGen (paired fine-tuned on developable sequences).
They used ca. 250m unpaired sequences and 1.8m paired sequences for training.
The model is based on GPT-2 but with rotary position embedding.
Developable sequences were defined as structural models of the 1.8m that had good TAP metrics (900,000 in total).
The model is much smaller than many of the models out there, (17m params), so it is more lightweight in training and application.
The model performs better on immunogenicity prediction than other models but worse on expression prediction.
Novel language model AntiBARTy with demonstration of how to use it to diffuse novel antibodies with favorable solubility properties.
The core model is a BART-based transformer, with 16m parameters.
It was firstly trained on all human heavy and light chains from OAS (254m heavies and 342m lights <- yes, more lights). This was followed by fine tuning on the higher quality paired data from OAS.
The diffusion model was based on U-net (CNN used for segmentation of medical images), totaling 3m parameters.
They define low and high solubility classes as predicted by protein-sol on paired OAS, with roughly 20k samples for each class.
Overall, one can sample from multivariate to get a vector in Antibarty latent space and use it to get an antibody sequence that is either high or low protein-sol predicted solubility.
Authors demonstrate that using inverse folding, one can affinity mature antibodies, confirmed experimentally.
Authors employ ESM-IF as the inverse folding algorithm.
They take two existing antibodies, bebletovimab and BD55-5840, both instrumental in COVID-19.
They introduce all possible single point mutations to the Vh and Vl regions (about 4300). They pick the best perplexity for experimental characterization.
The best perplexity ones have many framework mutations (bebletovimab 10/14 and BD55 5840 3/6). There was only one mutation to CDR-H3 in Bebletovimab.
Inverse folding mother achieves much better performance when antigen is used as well.
Authors introduce a large antibody-specific language model, Fabcon (2.4B) that demonstrably improves on predicting antibody specificity.
Model is based on the Falcon LLM and is trained no the CLM objective (predict next amino acid going N-to-C terminal)
The model was trained on paired (2.5m) and unpaired (821m) data from OAS.
The pre-trained model was tested on its ability to fine tune on binder prediction using three datasets anti-her2, anti-sars-cov2 and anti-il6.
When comparing against multiple other models on the binders prediction, the largest Fabcon model fares best, showing the benefit of overparametrization.
Since the model was trained on CLS objective, it can be used for sequence generation, producing sequences that are very human-like as compared to human PBMCs.
They employed ESM-IF as a base model for fine tuning.
They made one pass through ABodyBuilder2 models of paired OAS sequences (~150k sequences) and then ~2000 crystal structures.
They tested whether shotgun masking (random residues) is better than span-masking. Though shotgun performed better in general, span-masking is better in case the the entire CDRs are obscured (realistic case for design).
AntiFold improves upon author’s earlier ab-specific inverse folding method AbMPNN (fine-tuned ProteinMPNN), 43 % vs 60% sequence recovery on CDR-H3.
Authors took a handful of native structures, sampled sequences using different methods and modeled them using ABodyByuilder2 to see if the sampled sequences maintain the same fold. AntiFold achieves better (0.67) RMSD to the original backbone than AbMPNN (0.74) and ESM-IF (0.75).