Computational Antibody Papers

Filter by tags
All
Filter by published year
All
TitleKey points
    • Contrastive-learned sequence embeddings that predict antibody complementarity-determining region (CDR) structural similarity (RMSD) without explicit 3D structure generation, enabling ultra-fast structural retrieval across multi-million sequence databases.
    • AbSlang is trained on SAbDab structure pairs- it predicts CDR RMSD with mean absolute errors (MAE = 0.42--0.56A) comparable to state-of-the-art 3D structure predictors like AlphaFold3 and ABodyBuilder3, while outperforming sequence-identity baselines.
    • Using a twin neural network with 8 FlashAttention transformer encoder layers, AbSLang projects protein language model embeddings (IgT5 for paired VH/VL; ESM-C for single-domain VH/VHH) into 512-dimensional vectors trained against dynamic time warping (DTW) loop alignments.
    • Integrated with FAISS product quantization, AbSLang-search queries 2.5 million paired sequences in 7 seconds using an 879 MB index (compared to 18 minutes and 508 GB of storage for ABodyBuilder3 PDB files) and searches 10 million heavy chains in 1.22 seconds.
    • AbSLang-search matches explicit structure prediction workflows in identifying structurally similar antibodies (top-20 neighbor overlap of 13.60/20 vs. 14.40/20), enabling scalable mining of structurally convergent antibodies across distinct genetic lineages.
    • AbSlang was not benchmarked against the likes of Foldseek.
  • 2026-09-29

    Structure-based Antibody Renumbering

    • annotation/numbering
    • Numbering software for antibodies that is purely structural.
    • Structure-based Antibody Renumbering (SAbR) infers standardized residue indices strictly from backbone heavy-atom coordinates.
    • It matches the accuracy of sequence-based tools on standard antibodies while dramatically outperforming them on engineered antibodies with unusually long or unnatural CDR loops.
    • The algorithm generates per-residue embeddings using a fine-tuned ProteinMPNN encoder and aligns query backbones to position-averaged reference embeddings built from over 100,000 IMGT-numbered ImmuneBuilder models, followed by deterministic framework anchor refinement for CDR loops.
    • SAbR achieved a 97.8% framework accuracy on SAbDab test structures and >99% accuracy on de novo RFantibody designs, demonstrating robust handling of complex multidomain constructs (e.g., scFvs) and structural insertions that confound sequence alignment methods.
    • Set of open source models (DELPHI) for self-association and polyreactivity trained on a sizeable set of proprietary sequences.
    • DELPHI benchmarks 25 protein language model (PLM) and classifier combinations to enable sequence-based prediction of antibody developability liabilities with residue-level interpretability.
    • Evaluated using CDR H3-cluster-stratified cross-validation to prevent sequence-similarity leakage, DELPHI achieves high within-distribution performance for polyreactivity (mean AUC 0.959) and size-exclusion chromatography (SEC) monomer purity (mean AUC 0.933), while demonstrating strong zero-shot transfer to clinical cohorts and a 246,293-antibody public library (AUC up to 0.950).
    • Benchmarks show that language model choice is a stronger performance determinant than downstream classifier architecture with paired-chain PLMs (AbLang2, IgBert) leading cross-library generalization and automated learning curves demonstrate that predictive performance plateaus around 5,000 CDR H3-diverse training sequences.
    • Integrating SHAP, Integrated Gradients, and in silico CDR3 loop mutagenesis, DELPHI maps attributions to individual positions and uncovers a shared CDR H3 electrostatic risk signature (arginine/lysine enrichment, aspartate depletion, and net positive charge) that simultaneously drives polyreactivity and SEC non-monomeric failure
    • Five major pharmaceutical companies collaboratively fine-tuned the OpenFold3 model across ~20,000 private structures using federated learning, ensuring sensitive data and IP never left their secure environments.
    • The resulting model (AISB-1-Fed) significantly outperformed public baseline models, increasing high-quality interface predictions from 35.6% to 52.1% and correct ligand poses from 28.9% to 46.8%.
    • Public databases typically lack dense, drug-relevant medicinal-chemistry series. By pooling private data, the initiative effectively tripled the high-quality protein-ligand data available for training.
    • Although explicitly trained to improve protein-ligand co-folding (small molecule binding), the model also yielded a substantial, unexpected improvement in protein-protein interface (PPI) prediction quality.
    • Prediction of non-specific binding using ML from a diverse dataset.
    • Evaluated on a dataset of 1,563 VHH molecules measured via a high-throughput baculovirus particle (BVP) assay, the workflow uses a Leave-One-Group-Out cross-validation scheme to assess predictive performance in both high-similarity Lead Optimization (LO) and diverse Lead Isolation (LI) screening contexts.
    • A soft-voting ensemble classifier combining Logistic Regression, Random Forest, and LightGBM using AlphaFold2-derived structural descriptors achieved an AUROC of 0.73 in cross-validation and 0.79 on an independent test set, identifying surface-exposed hydrophobicity and positive charge patches in the complementarity-determining regions (CDRs) as primary drivers of polyreactivity.
    • Predictive performance sharply degrades for VHH sequences beyond a Levenshtein edit distance of 20 from the training dataset, establishing clear domain-of-applicability boundaries and demonstrating the necessity of sequence diversity for out-of-domain generalization.
    • Novel protein design algorithm - TorchCraft.
    • TorchCraft addresses antibody design through framework-conditioned VHH (nanobody) design, where CDR regions (CDR1, CDR2, and CDR3) are optimized while keeping a selected antibody framework sequence fixed.
    • To enforce antibody-specific sequence preferences, the framework incorporates a CDR-restricted IgLM language-model objective that uses the fixed framework and surrounding CDRs as infilling context to guide CDR optimization toward natural antibody distributions.
    • Experimental wet-lab campaigns validated raw TorchCraft VHH designs without post hoc sequence redesign against four targets (BHRF1, EFNA5, PDGFR beta, and S100A4), yielding measured apparent dissociation constants (K_D) ranging from 24.3 nM to 354 nM.
    • In computational benchmarks across 15 targets and five distinct nanobody frameworks (5JDS, 7EOW, 7XLO, 8COH, and 8Z8V), TorchCraft achieved high computational pass rates using a joint criterion evaluating AF3-derived interface confidence (TorchScore) and an ESM2-150M perplexity expression threshold
    • MD simulation study showing that paratopes rigidify upon maturation.
    • Through over 8.5 milliseconds of molecular dynamics simulations across seven antibody lineages, the authors demonstrate that affinity maturation selectively tunes paratope dynamics by rigidifying protein-contacting regions while enhancing flexibility at glycan-contacting interfaces.
    • Global antibody flexibility displays no uniform trend across lineages. Intermediate antibodies often exhibit non-monotonic dynamic changes, localized conformational entropy is specifically adapted depending on the target antigen interface.
    • Variable region dynamics remain largely consistent regardless of whether the constant region is present or whether light chain isotypes (Kappa vs. Lambda) are swapped, proving that computational costs for all-atom simulations can be cut by at least half by simulating variable regions alone without sacrificing accuracy.
    • Novel algorithm for sequence-structure co-design - SimpleDesign.
    • SimpleDesign introduces a single-stage framework for joint protein sequence and structure co-design that operates directly in data space, eliminating the need for complex structure tokenizers or multi-stage training.
    • It pairs discrete masked sequence recovery trained via cross-entropy with continuous C_alpha coordinate denoising trained via velocity-matching MSE, implemented using standard Transformer or Mixture-of-Transformer backbones.
    • On benchmarks spanning 100–500 amino acids, SimpleDesign matches or exceeds the sequence-structure consistency and structural diversity of complex tokenized protein language models such as ESM3 and DPLM2.
    • Compared to specialized geometric flow and diffusion models, it produces substantially higher-quality sequences with significantly lower ProGen2 perplexity and higher predicted foldability (pLDDT).
    • Review recounting insights from an EMBL-EBI workshop on how AI/ML are applied alongside in vitro assays and Quantitative Systems Pharmacology (QSP) to evaluate and mitigate preclinical immunogenicity risks for biotherapeutics.
    • Advanced in silico tools (such as NetMHCIIpan 4.3, Graph-pMHC, and HLAIIPred) accurately predict HLA Class II peptide presentation using mass-spectrometry immunopeptidomics data, though predicting T-cell receptor (TCR) binding for unseen epitopes and B-cell epitopes remains a bottleneck.
    • Pharmaceutical workflows integrate computational screening early in candidate selection to guide deimmunization, while QSP "middle-out" modeling combines computational predictions with empirical assay data to simulate clinical anti-drug antibody (ADA) formation and pharmacokinetic impacts.
    • The risk assessment framework is expanding beyond standard monoclonal antibodies to address unique immunological challenges in complex modalities, including AAV viral vectors, CAR-T cell therapies, and CRISPR-Cas9 gene editing components.
    • Future improvements in clinical ADA prediction rely on overcoming data fragmentation, assay non-standardization, and small dataset sizes, with collaborative initiatives like federated learning offering a path forward.
    • RFOptimization - a training-free framework that converts initial 3D biomolecular designs into high-confidence candidates by interleaving RoseTTAFold3 gradient-guided MCMC sequence search with Boltz and MPNN inverse-folding cycling.
    • The tool requires an initial 3D complex (PDB or mmCIF) as an input seed and yields a complete trajectory of optimized candidate sequences, predicted structures, and confidence metrics within 1–10 GPU minutes.
    • Benchmarked across general mini-protein binders, cyclic peptides, small-molecule biosensors, and catalytic enzymes, the pipeline uses customizable residue masks that allow users to target specific regions (such as interface residues or antibody CDR loops) while keeping functional motifs fixed.
    • Without wet-lab testing, performance was evaluated strictly in silico, achieving up to a 4.4x improvement in held-out AlphaFold3 refolding pass rates and outperforming existing baselines under a 3-model consensus filter (AF3, RF3, Boltz) at ~26 GPU-minutes per passing design.