Computational Antibody Papers

Filter by tags
protein design
Filter by published year
All
TitleKey points
    • Novel protein design algorithm - TorchCraft.
    • TorchCraft addresses antibody design through framework-conditioned VHH (nanobody) design, where CDR regions (CDR1, CDR2, and CDR3) are optimized while keeping a selected antibody framework sequence fixed.
    • To enforce antibody-specific sequence preferences, the framework incorporates a CDR-restricted IgLM language-model objective that uses the fixed framework and surrounding CDRs as infilling context to guide CDR optimization toward natural antibody distributions.
    • Experimental wet-lab campaigns validated raw TorchCraft VHH designs without post hoc sequence redesign against four targets (BHRF1, EFNA5, PDGFR beta, and S100A4), yielding measured apparent dissociation constants (K_D) ranging from 24.3 nM to 354 nM.
    • In computational benchmarks across 15 targets and five distinct nanobody frameworks (5JDS, 7EOW, 7XLO, 8COH, and 8Z8V), TorchCraft achieved high computational pass rates using a joint criterion evaluating AF3-derived interface confidence (TorchScore) and an ESM2-150M perplexity expression threshold
    • Novel algorithm for sequence-structure co-design - SimpleDesign.
    • SimpleDesign introduces a single-stage framework for joint protein sequence and structure co-design that operates directly in data space, eliminating the need for complex structure tokenizers or multi-stage training.
    • It pairs discrete masked sequence recovery trained via cross-entropy with continuous C_alpha coordinate denoising trained via velocity-matching MSE, implemented using standard Transformer or Mixture-of-Transformer backbones.
    • On benchmarks spanning 100–500 amino acids, SimpleDesign matches or exceeds the sequence-structure consistency and structural diversity of complex tokenized protein language models such as ESM3 and DPLM2.
    • Compared to specialized geometric flow and diffusion models, it produces substantially higher-quality sequences with significantly lower ProGen2 perplexity and higher predicted foldability (pLDDT).
    • RFOptimization - a training-free framework that converts initial 3D biomolecular designs into high-confidence candidates by interleaving RoseTTAFold3 gradient-guided MCMC sequence search with Boltz and MPNN inverse-folding cycling.
    • The tool requires an initial 3D complex (PDB or mmCIF) as an input seed and yields a complete trajectory of optimized candidate sequences, predicted structures, and confidence metrics within 1–10 GPU minutes.
    • Benchmarked across general mini-protein binders, cyclic peptides, small-molecule biosensors, and catalytic enzymes, the pipeline uses customizable residue masks that allow users to target specific regions (such as interface residues or antibody CDR loops) while keeping functional motifs fixed.
    • Without wet-lab testing, performance was evaluated strictly in silico, achieving up to a 4.4x improvement in held-out AlphaFold3 refolding pass rates and outperforming existing baselines under a 3-model consensus filter (AF3, RF3, Boltz) at ~26 GPU-minutes per passing design.
    • Discovery and engineering of bispecific single-domain antibody (sdAb)-based IL-21 mimetics (surrogate agonists) that target the IL-21R and IL-2Rgamma receptor subunits to activate downstream STAT3 signaling and induce Granzyme B expression in immune cells.
    • ColabFold (AlphaFold2) was used to model complex structures of IL-21R and IL-2Rgamma bound to VHH paratopes, generating structural hypotheses on how distinct epitope bins and paratope orientations dictate productive receptor signaling geometry.
    • ProteinMPNN was applied to perform structure-based framework engineering by aligning VHH backbones to established VH dimer templates (PDB 7LU9 and 7L6M), yielding novel mutation sets (dsdAb(3)–(5)) designed to induce noncovalent VHH:VHH intramolecular dimerization.
    • Candidate designs were assembled onto antibody scaffolds and energy-minimized using Molecular Operating Environment (MOE), followed by binding interface evaluation and selection using PRODIGY.
    • The computationally engineered framework mutations enforced spatial rigidity and proximity between paratopes, successfully converting weak or inactive bispecific formats into highly potent cytokine mimetics without altering individual antigen-binding affinities.
    • End-to-end framework using generative deep learning to design de novo single-domain antibodies against the snake neurotoxin alpha-cobratoxin, leading to the experimental validation of low-nanomolar binders that achieved 100% in vivo survival in mice.
    • Conducted a head-to-head in silico comparison of three vhh-capable generative design tools (Germinal, RFantibody, and BoltzGen) across three structural scaffolds (9GCN, 7XL0, and 3EAK), fixing CDR loop lengths and targeting five specific epitope hotspot residues (D27, R33, K35, R36, and V37).
    • Evaluated complex predictions using AlphaFold3 (AF3) interface predicted TM-scores iptm combined with a target-aligned vhh structural self-consistency filter RMSD < 6Å), where Germinal generated a substantially higher fraction of passing candidates (46/300) than RFantibody (6/900) or BoltzGen (12/900).
    • Profiling against SAbDab, OAS, and INDI antibody databases demonstrated that Germinal generated more novel CDR3 sequences and broader loop conformation sampling, which significantly narrowed the compute-time required per successful candidate despite Germinal's higher raw GPU runtime per design.
    • Scaled the Germinal pipeline to ~8,000 trajectories using multi-stage filtering (including pDockQ2, PAE, spatial aggregation propensity, and AF3 re-prediction) to select candidates for experimental testing, finding retrospectively that AF3 ipSAE_min ranked true binders more effectively than standard ipTM
    • A restriction free reproduction of the antibody design workflow germinal
    • Replaces proprietary dependencies (PyRosetta, IgLM) with an open-source toolchain (OpenMM, AbLang1, sc-rs) and fixes multi-chain bugs, enabling unrestricted academic and commercial deployment.
    • Demonstrates that AbLang1-guided hallucination significantly increases initial cofolding pass rates (e.g., 33.7% vs. 18.6% for PD-L1) with equal or higher structural confidence, at the cost of a ~1.5x increase in per-trajectory compute time.
    • Uses hard-coded placeholder values for three energy metrics which degrades ensemble selection and disables the interface hydrogen-bond filter and lacks wet-lab experimental validation of the generated binders.
    • Novel de novo protein binder design pipeline from Boltz that pairs an optimized candidate generation engine with BoltzPPI, a novel, interaction-aware scoring model built to rank designs based on binding confidence rather than just geometric plausibility.
    • To train BoltzPPI, the authors created a set of positive labels (PDB and patent complexes) and negative labels (synthetic non-interacting protein pairs co-localized via Boltz-2). This input pool (token, pair, distance, and mask features) is processed through a Pairformer stack using specific optimization tricks (co-trained focal loss, multi-view representation dropping, and Gaussian noise injection) to predict a binary binding-confidence score.
    • Significantly improves experimental VHH design performance, boosting the confirmed-binder hit rate from 3.3% to 8.0% on novel targets and successfully discovering screening hits for 7 out of 10 benchmark targets from the Chai-2 dataset.
    • Yields highly manufacturable binders, with 58% of its confirmed candidates passing a comprehensive panel of strict biophysical filters (such as thermal stability, purity, and low aggregation Propensity), outperforming both BoltzGen (40%) and clinical-stage VHH controls (21%).
    • Generates structurally distinct binders whose CDR loops remain highly distinct from any known entries in the SAbDab database, with the entire pipeline made available to researchers via the Boltz API and Lab platforms.
    • Case study of JAM-2, a generative biomolecular design model that engineered drug-like VHH antibodies against five distinct peptide-MHC class I (pMHC-I) targets across two HLA alleles using only target amino acid sequences as input.
    • When reformatted into bispecific T-cell engagers (TCEs) without experimental optimization, the designs mediated potent, sub-nanomolar T-cell activation and successfully directed primary human T cells to kill target-presenting cells.
    • Binders achieved stringent selectivity, showing a minimum 216-fold preference for NY-ESO-1 over highly similar human self-peptides and safely discriminating mutant KRAS variants (G12V and G12C) from wild-type sequences. A cryo-EM structure verified this atomic accuracy with a whole-complex Ca RMSD of 0.93 Å.
    • Bypassing the need for downstream engineering, 78% of the designed antibodies passed five core industry-standard biophysical developability criteria, exhibiting strong expression titers, favorable monomericity, thermal stability, and low polyreactivity
  • 2026-05-29

    Language Modeling Materializes a World Model of Protein Biology

    • language models
    • protein design
    • structure prediction
    • New versions of ESM and ESMFold, ESMC (ESM Cambrian) and ESMFold2 to model protein sequence, structure, and function.
    • Like previous versions, ESMC is trained entirely on sequences using a masked language modeling (MLM) objective. However, it scales up to 2.8 billion metagenomic sequences, nearly a 100-fold increase over the 50 million sequences used for ESM2.
    • ESMFold2 achieves state-of-the-art atomic resolution and directly outperforms AlphaFold3 on complex antibody-antigen predictions (even when AlphaFold3 is given multiple sequence alignments and ESMFold2 operates from sequence alone).
    • To design therapeutic antibody fragments (scFvs), they input the target sequence and lock in a stable, known antibody framework template, leaving only the target-recognizing loops (CDRs) to be filled in.
    • AI-guided optimization: The system uses mathematical backpropagation through both ESMC and ESMFold2 to iteratively optimize the CDR sequences. It automatically mutates the loops to maximize structural interface confidence scores.
    • They synthesized and tested these computational designs in the wet lab, achieving high experimental hit rates and discovering entirely novel binders with therapeutically relevant nanomolar affinities.
    • Novel protein design model, revisiting the SE(3)architecture.
    • Genie 3 is an all-atom, SE(3)-equivariant structure diffusion model that treats proteins as branched polymers to capture sidechain details, utilizing a Latent Transformer with bidirectional layer updates and an Invariant Point Attention structural decoder.
    • The authors did not test the model on therapeutic formats such as antibodies or nanobodies; instead, they focused entirely on generating generic de novo protein binders, unconditional monomers, and functional motif scaffolds.
    • The method was computationally benchmarked using self-consistency pipelines (ProteinMPNN/ESMFold), MotifBench for functional sites, and a strict AF2M+ binder interface metric, alongside real-world experimental validation that yielded a 12.5% hit rate against the Nipah virus Glycoprotein G