Computational Antibody Papers

Filter by tags
developability
Filter by published year
2025
TitleKey points
  • 2025-11-26

    Drug-like antibody design against challenging targets with atomic precision

    • protein design
    • generative methods
    • developability
    • Update on earlier Chai-2 results adding developability and structural validations.
    • Previously generated scFv hits were reformatted into full-length IgGs; ~93% retained binding.
    • Developability was assessed using NanoDSF (Tm), HIC-HPLC, BVP ELISA, and AC-SINS, with Jain-style green-flag thresholds.
    • Most reformatted IgGs passed ≥3 of 4 developability flags, indicating good biophysical properties without further optimization.
    • Newly designed antibodies were generated for the GPCR benchmarks, showing successful in silico design against challenging targets.
    • Introduced a novel pairing predictor for VhVl chains with a clever strategy to sample negative pairs.
    • Defines three negative sampling strategies:
    • Random pairing, where heavy and light chains are shuffled without constraints.
    • V-gene mismatching, where non-native pairs are generated by combining VH and VL sequences drawn from different V-gene families, but within biologically plausible V-gene segments. This captures realistic but unobserved combinations that could occur during recombination.
    • Full V(D)J mismatching, where heavy and light chains are paired using completely distinct germline origins across V, D, and J gene segments. This produces negative examples that are maximally diverse yet biologically meaningful, reflecting combinations never seen in natural repertoires.
    • Shows that the space of possible VH–VL germline combinations is far larger than what is observed in public datasets, revealing non-random biological constraints on pairing.
    • Demonstrates that models trained on V-gene and especially VDJ mismatched datasets achieve the highest and most generalizable performance, outperforming existing methods such as ImmunoMatch, p-IgGen, and Humatch — confirming that biologically grounded negative sampling is key to robust VH–VL pairing prediction.
    • Benchmarking of computational models for predicting antibody aggregation propensity (developability) using size-exclusion chromatography (SEC) readouts.
    • Developed an experimental dataset of ~1,200 IgG1 antibodies, measured for monomer percentage and ΔRT (difference in retention time) relative to a reference.
    • Evaluated four main prediction pipelines: Sequence + structure-based features (hand-crafted biophysical features from Schrödinger, using AlphaFold2 or ImmuneBuilder for structure). PLM (protein language model) pipeline (e.g., ESM2-8M, fine-tuned or LoRA-adapted). GNN (graph neural network) pipeline using residue graphs from predicted structures. PLM + GNN hybrid pipeline combining sequence embeddings with structural graphs.
    • Two structure prediction tools were benchmarked: AlphaFold2 (high accuracy, slow) and ImmuneBuilder (faster, antibody-optimized, slightly less accurate).
    • The sequence + structure feature model achieved the highest accuracy overall, but low sensitivity (missed many problematic antibodies).
    • The PLM-only pipeline performed nearly as well and offered a much faster, high-throughput solution, making it attractive for early screening.
    • The GNN and PLM + GNN approaches performed comparably, with GNN slightly better for ΔRT predictions but more variable.
    • Using ImmuneBuilder instead of AlphaFold2 reduced sensitivity slightly but greatly improved speed without major loss of accuracy.
    • So all pipelines performed similarly within a narrow performance range, but faster, less resource-intensive approaches (PLM and ImmuneBuilder-based pipelines) offer strong trade-offs for early-stage developability screening.
    • They introduce a template-free diffusion model for antibody humanization.
    • It receives CDR sequences, reconstructing the framework regions without needing humanized templates.
    • Benchmarked against Sapiens, Humatch, Llamanade, and AbNatiV across multiple datasets (e.g., HuAb348, Humab25, Nano300), showing improved humanness, germline identity, and binding retention.
    • Demonstrates preserved or enhanced binding and stability in vitro, though no direct ADA correlation analysis was performed.
    • Introduces TNP, a nanobody-specific developability profiler inspired by TAP.
    • Uses six metrics: total CDR length, CDR3 length, CDR3 compactness, and patch scores for hydrophobicity, positive charge, and negative charge.
    • Thresholds are calibrated to 36 clinical-stage nanobodies.
    • In vitro assays on 108 nanobodies (36 clinical-stage + 72 proprietary) show partial agreement with TNP flags, indicating complementary—but not perfectly correlated—assessments.
    • Combined in vitro/in silico method for optimization of binders.
    • Start from a wild-type scFv (heavy chain), build a random-mutant library, FACS-sort on multiple antigens, deep-sequence bins + input, and use per-sequence enrichment (bin/library) as the supervised target for (antibody, antigen) training pairs.
    • Train uncertainty-aware regressors (xGPR or ByteNet-SNGP) on those enrichment targets; run in-silico directed evolution (ISDE) from the WT, proposing single mutations and auto-rejecting moves with high predictive uncertainty while optimizing the worst-case score across antigens.
    • Binding is protected by the multi-antigen objective + uncertainty gating during ISDE; risky proposals are discarded before they enter the candidate set.
    • Filter candidates for humanness with SAM/AntPack and for solubility with CamSol v2.2 (framework is extensible to add other gates); final wet-lab set kept 29 designs after applying these filters and uncertainty checks.
    • Beyond large in-silico tests, yeast-display across 10 SARS-CoV-2 RBDs shows most designs outperform WT; a representative clone (Delta-63) improves KD on 8/10 variants and competes with ACE2.
    • Novel method to assess antibody immunogenicity.
    • Created two reference libraries: a positive set from human proteins and antibodies (OAS + proteome) and a negative set from murine antibody sequences (OAS).
    • Antibody sequences are fragmented into 8–12-mer peptides.
    • Peptide fragments are scored: +1.0 if matching the positive reference, −0.2 if matching the negative reference.
    • Validated on 217 therapeutic antibodies with known clinical ADA incidence, showing strong negative correlation between hit rate and ADA.
    • On 25 humanized antibody pairs, ImmunoSeq correctly predicted reduced immunogenicity after humanization, consistent with experimental results.
    • Mutational analysis of Trastuzumab framework (FW) regions to modulate antibody stability and function, moving beyond the traditional focus on CDRs.
    • Authors evaluated antibody-specific language models (AbLang2, AntiBERTy, etc.), a general protein language model (ESM-2), and a structure-based Rosetta approach. While the language models showed limited utility in suggesting beneficial FW mutations, Rosetta provided more reliable predictions based on structural stability.
    • Language model-derived mutation suggestions were generally less informative than Rosetta’s, which successfully identified stabilizing FW mutations not biased toward germline residues.
    • Authors experimentally characterized selected mutants in vitro, assessing thermostability, antigen (HER2) binding, and functional effects such as ADCC and tumor cell viability. Some mutations preserved function, while others decoupled binding from downstream activity.
    • Demonstration of generation of novel natural antibodies biased towards favorable biophysical scores.
    • Developed a masked discrete diffusion–based generative model by retraining an ESM-2 (8M) architecture on paired heavy- and light-chain sequences from the Observed Antibody Space (OAS), using an order-agnostic diffusion objective to capture natural repertoire features.
    • Built ridge-regression predictors on ESM-2 embeddings using experimental developability measurements for 246 clinical-stage antibodies, focusing on hydrophobicity (HIC RT) and self-association (AC-SINS pH 7.4), and achieved cross-validated Spearman’s ρ of 0.42 for HIC and 0.49 for AC-SINS.
    • Showed that unconditionally generated sequences maintain high naturalness, scoring with AbLang-2 and p-IgGen language models to yield log-likelihood distributions comparable to natural and clinical repertoires.
    • Applied Soft Value-based Decoding in Diffusion (SVDD) guidance during sampling to bias generation toward sequences with low predicted hydrophobicity and self-association, enriching the fraction of candidates in the high-developability quadrant.
  • 2025-06-05

    AbBFN2: A flexible antibody foundation model based on Bayesian Flow Networks

    • developability
    • generative methods
    • protein design
    • Novel generative modeling framework (AbBFN2) using Bayesian Flow Networks (BFNs) for antibody sequence optimization.
    • Trains on sequences from Observed Antibody Space (OAS) combined with genetic and biophysical annotations, leveraging a denoising approach for both conditional and unconditional sequence generation. Targets include optimizing Therapeutic Antibody Profiler (TAP) annotations.
    • Computationally validated for germline assignment accuracy, species prediction (humanness), and TAP parameter optimization.
    • Combines multiple antibody design objectives into a unified, single-step optimization process, unlike existing software methods which are typically specialized for individual tasks.