Computational Antibody Papers

Filter by tags
All
Filter by published year
All
TitleKey points
    • Review of B and T cell epitope predictors
    • The review unifies AI-driven B- and T-cell epitope identification by establishing a novel 12-task taxonomy mapped across a shared five-stage predictive pipeline, systematically evaluating 155 studies (2015–2026) to address severe methodological fragmentation, database underutilization, and benchmark overfitting in computational immunoinformatics.
    • The proposed taxonomy organizes the field into linear and conformational B-cell prediction alongside ten sequential T-cell sub-tasks (T1–T10; spanning upstream HLA typing, antigen processing, MHC class I/II binding, immunogenicity, TCR specificity, and structural modeling to downstream neoantigen, tumor antigen, and vaccine design applications) while providing a multidimensional landscape analysis of 43 dedicated immunology databases and 144 benchmark datasets.
    • Methodological synthesis across 272 representation learning encodings and 148 classifiers indicates that deep learning and protein language models (PLMs) such as ESM and BERT have become the dominant paradigm in structural, MHC binding, and TCR specificity tasks, whereas classical machine learning remains heavily over-represented in translational applications like tumor T-cell antigen classification (T9) and multi-epitope vaccine construct design (T10).
    • The analysis exposes critical field-wide bottlenecks, including steep prediction accuracy collapse on unseen epitopes/TCRs, an immunogenicity (T5) accuracy ceiling that severely degrades neoantigen prioritization, an absence of uncertainty quantification, and poor practical deployment practices where only 12.5% of audited B/T-cell predictors offer dual code repository and operational web server availability.
    • Novel de novo antibody design framework (not new model), based on mimicking the binding partner of the target.
    • Authors introduce MIMOSA (MIMic-Oriented Structural Antibody generation), an inference-time, model-agnostic framework that guides diffusion-based de novo antibody design by steering CDR backbones toward known target-partner interaction fingerprints, achieving 20–26% experimental hit rates and single-digit nanomolar binders on targets where unconstrained generative models fail.
    • The framework applies a differentiable "mimic potential" during diffusion to dynamically steer CDR backbone coordinates, followed by optional sequence-fixing during inverse folding to preserve essential hotspot chemistry. It flexibly accommodates both contiguous single-loop motifs using a sliding-window search and spatially dispersed multi-CDR hotspots via Hungarian per-residue assignment without requiring model retraining.
    • Validated using RFAntibody and BoltzGen backbones across three targets (KEAP1, uPA, and IL-8), MIMOSA successfully generated nanobody binders reaching affinities down to 4.1 nM for IL-8 and 37 nM for KEAP1. The method focused search space without collapsing diversity, maintaining high CDR sequence variation (53–61% CDR differences between designs) across multiple framework scaffolds and loop lengths.
    • Experimental evaluation revealed that mimic fidelity complements standard folding confidence metrics (ipSAE), providing a critical selection signal on targets like IL-8 where folding confidence alone cannot differentiate binders from non-binders.
    • Novel epitope-specific antibody design method.
    • TiDE-Ab is a conditional SE(3) flow-matching framework for de novo epitope-specific antibody design that utilizes dynamic guidance scheduling to discover global binding poses from scratch while generating structurally valid, clash-free CDR conformations.
    • Operating on unpaired antigen and antibody scaffold structures without pre-aligned frames, the model introduces Time-Dependent Classifier-Free Guidance (TD-CFG), applying strong epitope conditioning early to establish global pose orientation, followed by cosine annealing and early truncation to permit undisturbed local CDR refinement.
    • Across 55 non-redundant benchmark target complexes, TiDE-Ab demonstrated state-of-the-art performance by achieving higher epitope recall than RFantibody (0.935 vs. 0.878) while eliminating over 95% of steric clashes (8.14 vs. 168.20 clashes per complex).
    • In therapeutic case studies on TGF-β and IL-17A, TiDE-Ab reliably generated candidates reproducing both isoform-selective and cross-reactive clinical binding profiles using realistic partial-epitope inputs, whereas baseline diffusion models consistently failed to yield viable structures.
    • Study of the bovine ultra long CDR-H3s.
    • Cattle produce unique immunoglobulins featuring ultralong CDR-H3 (uCDR-H3) loops that project a modular, disulfide-rich "knob" paratope from a conserved stalk, enabling access to recessed or cryptic epitopes and enabling reformatting into diverse therapeutic architectures.
    • Generated through restricted VDJ gene usage (IGHV1-7, IGHD8-2, and IGHJ2-4), cysteine-rich somatic hypermutation, and diverse intra-knob disulfide pairing topologies, uCDR-H3 regions achieve extensive structural diversity while relying on highly conserved V30 lambda light chains to provide framework stability and favorable biophysical traits.
    • Because antigen binding depends exclusively on the 4-6 kDa knob mini-domain, paratopes isolated via display technologies can be engineered into common light chain bispecific antibodies, loop-grafted Fc-knob variants, Knobbodies, framework-inserted Fabs/VHHs, artificial coiled-coil arrays, or autonomous solid-phase synthesized peptides.
    • While uCDR-H3 constructs exhibit favorable developability traits like high thermal stability and limited aggregation across broad pH ranges, translational challenges remain regarding potential non-human immunogenicity, sequence-dependent self-association tendencies, and structural instability when minimal knobs lack terminal spatial constraints.
    • Contrastive-learned sequence embeddings that predict antibody complementarity-determining region (CDR) structural similarity (RMSD) without explicit 3D structure generation, enabling ultra-fast structural retrieval across multi-million sequence databases.
    • AbSlang is trained on SAbDab structure pairs- it predicts CDR RMSD with mean absolute errors (MAE = 0.42--0.56A) comparable to state-of-the-art 3D structure predictors like AlphaFold3 and ABodyBuilder3, while outperforming sequence-identity baselines.
    • Using a twin neural network with 8 FlashAttention transformer encoder layers, AbSLang projects protein language model embeddings (IgT5 for paired VH/VL; ESM-C for single-domain VH/VHH) into 512-dimensional vectors trained against dynamic time warping (DTW) loop alignments.
    • Integrated with FAISS product quantization, AbSLang-search queries 2.5 million paired sequences in 7 seconds using an 879 MB index (compared to 18 minutes and 508 GB of storage for ABodyBuilder3 PDB files) and searches 10 million heavy chains in 1.22 seconds.
    • AbSLang-search matches explicit structure prediction workflows in identifying structurally similar antibodies (top-20 neighbor overlap of 13.60/20 vs. 14.40/20), enabling scalable mining of structurally convergent antibodies across distinct genetic lineages.
    • AbSlang was not benchmarked against the likes of Foldseek.
  • 2026-09-29

    Structure-based Antibody Renumbering

    • annotation/numbering
    • Numbering software for antibodies that is purely structural.
    • Structure-based Antibody Renumbering (SAbR) infers standardized residue indices strictly from backbone heavy-atom coordinates.
    • It matches the accuracy of sequence-based tools on standard antibodies while dramatically outperforming them on engineered antibodies with unusually long or unnatural CDR loops.
    • The algorithm generates per-residue embeddings using a fine-tuned ProteinMPNN encoder and aligns query backbones to position-averaged reference embeddings built from over 100,000 IMGT-numbered ImmuneBuilder models, followed by deterministic framework anchor refinement for CDR loops.
    • SAbR achieved a 97.8% framework accuracy on SAbDab test structures and >99% accuracy on de novo RFantibody designs, demonstrating robust handling of complex multidomain constructs (e.g., scFvs) and structural insertions that confound sequence alignment methods.
    • Set of open source models (DELPHI) for self-association and polyreactivity trained on a sizeable set of proprietary sequences.
    • DELPHI benchmarks 25 protein language model (PLM) and classifier combinations to enable sequence-based prediction of antibody developability liabilities with residue-level interpretability.
    • Evaluated using CDR H3-cluster-stratified cross-validation to prevent sequence-similarity leakage, DELPHI achieves high within-distribution performance for polyreactivity (mean AUC 0.959) and size-exclusion chromatography (SEC) monomer purity (mean AUC 0.933), while demonstrating strong zero-shot transfer to clinical cohorts and a 246,293-antibody public library (AUC up to 0.950).
    • Benchmarks show that language model choice is a stronger performance determinant than downstream classifier architecture with paired-chain PLMs (AbLang2, IgBert) leading cross-library generalization and automated learning curves demonstrate that predictive performance plateaus around 5,000 CDR H3-diverse training sequences.
    • Integrating SHAP, Integrated Gradients, and in silico CDR3 loop mutagenesis, DELPHI maps attributions to individual positions and uncovers a shared CDR H3 electrostatic risk signature (arginine/lysine enrichment, aspartate depletion, and net positive charge) that simultaneously drives polyreactivity and SEC non-monomeric failure
    • Five major pharmaceutical companies collaboratively fine-tuned the OpenFold3 model across ~20,000 private structures using federated learning, ensuring sensitive data and IP never left their secure environments.
    • The resulting model (AISB-1-Fed) significantly outperformed public baseline models, increasing high-quality interface predictions from 35.6% to 52.1% and correct ligand poses from 28.9% to 46.8%.
    • Public databases typically lack dense, drug-relevant medicinal-chemistry series. By pooling private data, the initiative effectively tripled the high-quality protein-ligand data available for training.
    • Although explicitly trained to improve protein-ligand co-folding (small molecule binding), the model also yielded a substantial, unexpected improvement in protein-protein interface (PPI) prediction quality.
    • Prediction of non-specific binding using ML from a diverse dataset.
    • Evaluated on a dataset of 1,563 VHH molecules measured via a high-throughput baculovirus particle (BVP) assay, the workflow uses a Leave-One-Group-Out cross-validation scheme to assess predictive performance in both high-similarity Lead Optimization (LO) and diverse Lead Isolation (LI) screening contexts.
    • A soft-voting ensemble classifier combining Logistic Regression, Random Forest, and LightGBM using AlphaFold2-derived structural descriptors achieved an AUROC of 0.73 in cross-validation and 0.79 on an independent test set, identifying surface-exposed hydrophobicity and positive charge patches in the complementarity-determining regions (CDRs) as primary drivers of polyreactivity.
    • Predictive performance sharply degrades for VHH sequences beyond a Levenshtein edit distance of 20 from the training dataset, establishing clear domain-of-applicability boundaries and demonstrating the necessity of sequence diversity for out-of-domain generalization.
    • Novel protein design algorithm - TorchCraft.
    • TorchCraft addresses antibody design through framework-conditioned VHH (nanobody) design, where CDR regions (CDR1, CDR2, and CDR3) are optimized while keeping a selected antibody framework sequence fixed.
    • To enforce antibody-specific sequence preferences, the framework incorporates a CDR-restricted IgLM language-model objective that uses the fixed framework and surrounding CDRs as infilling context to guide CDR optimization toward natural antibody distributions.
    • Experimental wet-lab campaigns validated raw TorchCraft VHH designs without post hoc sequence redesign against four targets (BHRF1, EFNA5, PDGFR beta, and S100A4), yielding measured apparent dissociation constants (K_D) ranging from 24.3 nM to 354 nM.
    • In computational benchmarks across 15 targets and five distinct nanobody frameworks (5JDS, 7EOW, 7XLO, 8COH, and 8Z8V), TorchCraft achieved high computational pass rates using a joint criterion evaluating AF3-derived interface confidence (TorchScore) and an ESM2-150M perplexity expression threshold