Computational Antibody Papers

Filter by tags
protein design
Filter by published year
All
TitleKey points
  • 2025-09-30

    mBER: Controllable de novo antibody design with million-scale experimental screening

    • binding prediction
    • generative methods
    • protein design
    • experimental techniques
    • Novel de novo antibody design method with massive experimental testing.
    • The computational method involves integration, not retraining, of existing tools. It combines AlphaFold-Multimer, protein language models (ESM2/AbLang2), and NanoBodyBuilder2 with templating/sequence priors to design/filter antibody-format binders.
    • They perform massive testing. >1.1 million VHH binders designed across 436 targets (145 tested); ~330k experimentally screened.
    • Hit rates look low per binder (~0.5–1%) but that’s ~50× above random libraries, and still yields thousands of validated binders.
    • Target-level success is 45%, for how many targets we got binders; some epitopes reached 30–38% hit rates after filtering.
    • The big caveat is the specificity of epitopes- it really makes a difference, with some epitopes producing nought.
    • Introduces a novel diffusion-based inverse folding method (RL-DIF) that improves the foldable diversity of generated sequences—i.e., it can generate more diverse sequences that still fold into the desired structure.
    • The model uses categorical denoising diffusion for sequence generation, followed by reinforcement learning (DDPO) to improve structural consistency with the target fold.
    • During reinforcement learning, ESMFold is used to predict the 3D structure of generated sequences, which is then compared (via TM-score) to the structure predicted from the native sequence to ensure they fold similarly.
    • Compared to baselines like PiFold and ProteinMPNN, RL-DIF achieves similar sequence recovery and structural consistency but significantly better foldable diversity—a critical advantage in protein design.
    • Novel inverse folding algorithm based on a discrete diffusion framework.
    • Unlike earlier methods that focused on masked language modeling (MLM) (e.g., LM-Design) or autoregressive sequence generation (e.g., ProteinMPNN), this work introduces a discrete denoising diffusion model (MapDiff) to iteratively refine protein sequences toward the native sequence. The method incorporates an IPA-based refinement step that selectively re-predicts low-confidence residues.
    • Structural input is limited to the protein backbone only, represented as residue-level graphs. All-atom information is not used for either masked or unmasked residues.
    • On the CATH 4.2 full test set, their method achieves the best sequence recovery rate of 61.03%, outperforming baselines such as: ProteinMPNN: 48.63% PiFold: 51.40% LM-Design: 53.19% GRADE-IF: 52.63%
    • MapDiff also achieves the lowest perplexity (3.46) across models.
  • 2025-07-03

    Antibody Design using Chai-2

    • generative methods
    • binding prediction
    • protein design
    • Introduced a novel model, Chai-2, that shows over 100× improvement in de novo antibody design success rates compared to prior methods.
    • The model is prompted with the structure of the target, epitope residues, and desired antibody format (e.g., scFv or VHH).
    • Benchmarking was performed on 52 antigens that had no known antibodies in the PDB, ensuring evaluation on novel, unbiased targets.
    • Generated antibodies were structurally and sequentially dissimilar to any known antibodies, indicating that Chai-2 designs novel binders, not memorized ones.
    • For VHH (nanobody) formats, the model achieved an experimental hit rate of 20%, validated in a single experimental round.
    • Novel method to design antibodies based on boltz-1.
    • They added a sequence head to boltz-1 to perform simultaneous sequence/structure co-design.
    • They employed data from SAbDab to fine-tune boltz-1 on antibody-antigen complexes.
    • They compared to dyMEAN and DiffAB looking at amino acid recovery, RMSD and Rosetta InterfaceAnalyzer energy - their model does better on these computational benchmarks.
  • 2025-06-05

    Adapting ProteinMPNN for antibody design without retraining

    • protein design
    • generative methods
    • Novel method to bias ProteinMPNN for antibody design, without modifying model weights.
    • Logits from protein-general ProteinMPNN and antibody-specific AbLANG are added and softmaxed. Addition of AbLANG is supposed to push the model into the antibody-acceptable space.
    • On in-silico experiments ProteinMPNN+AbLang outperformed ProteinMPNN alone and rivalled antibody-specific AbMPNN.
    • Authors designed 96 variants of Trastuzumab CDR-H3 using ProteinMPNN, AbLang and ProteinMPNN+AbLang each. AbLANG and ProteinMPNN produced 1 and 3 successful variants respecitively (both out of 96) whereas their combination produced 36 successful variants.
    • None of the variants were better variants than WT Trastuzumab.
  • 2025-06-05

    AbBFN2: A flexible antibody foundation model based on Bayesian Flow Networks

    • developability
    • generative methods
    • protein design
    • Novel generative modeling framework (AbBFN2) using Bayesian Flow Networks (BFNs) for antibody sequence optimization.
    • Trains on sequences from Observed Antibody Space (OAS) combined with genetic and biophysical annotations, leveraging a denoising approach for both conditional and unconditional sequence generation. Targets include optimizing Therapeutic Antibody Profiler (TAP) annotations.
    • Computationally validated for germline assignment accuracy, species prediction (humanness), and TAP parameter optimization.
    • Combines multiple antibody design objectives into a unified, single-step optimization process, unlike existing software methods which are typically specialized for individual tasks.
    • Protein design method based on Boltz-1.
    • Boltz-1 is an open-source reproduction of AlphaFold3, which uses a diffusion module to co-fold molecular structures (proteins, ligands, etc.).
    • For design purposes, BoltzDesign1 sidesteps the full structure generation step and instead uses only the Pairformer (which outputs a distogram — a probabilistic representation of all pairwise residue distances). This allows broader exploration of sequence space, as it optimizes over the distribution of possible structures rather than committing to a single conformation.
    • Given a target (such as a small molecule or protein), they weakly initialize a binder sequence using random logits. This sequence is then iteratively refined by backpropagating loss through the Pairformer (and optionally through the Confidence module) to increase the predicted quality of the binder–target interaction.
    • A full 3D structure can be generated at the end using the Boltz-1 structure module, but this is not part of the optimization loop.
    • They benchmarked their method in silico on small molecule targets and a set of protein–protein interactions from the BindCraft benchmark, comparing performance to RfDiffusion All-Atom.
  • 2025-04-28

    BindCraft: one-shot design of functional protein binders

    • protein design
    • non-antibody stuff
    • BindCraft is an easy-to-use pipeline for computational protein binder design.
    • It employs AlphaFold2-Multimer to hallucinate binders via backpropagation.
    • Given a target structure and binder parameters (e.g., sequence length), the binder sequence is initialized with random logits and iteratively optimized via gradient descent through the AF2-Multimer network.
    • After binder hallucination, the sequence and surface residues are further optimized using MPNNsol, and AF2-Monomer is used to repredict and filter high-confidence designs.
    • Binder designs were validated experimentally through in vitro assays, X-ray crystallography, and cryo-EM.
    • Reported success rates ranged from 25% to 100%, with most binders in the nanomolar affinity range, a few in the micromolar range, and backbone RMSDs of ~1.7 Å to 3.1 Å between design models and solved structures.
  • 2025-04-28

    Atom level enzyme active site scaffolding using RFdiffusion2

    • protein design
    • non-antibody stuff
    • Improvement upon earlier RFDiffusion, enhancing stability and accuracy in designing enzyme active sites.
    • Catalytic sites can now be specified at the atomic level instead of the residue backbone level used previously. This eliminates the need to explicitly enumerate side-chain rotamers.
    • Training uses flow matching, a technique that simplifies and stabilizes the diffusion training process.
    • Benchmarked on a set of 41 diverse enzyme active sites; RFdiffusion2 succeeded in all 41 cases, significantly outperforming the earlier RFDiffusion, which succeeded in only 16.