Discovery and engineering of bispecific single-domain antibody (sdAb)-based IL-21 mimetics (surrogate agonists) that target the IL-21R and IL-2Rgamma receptor subunits to activate downstream STAT3 signaling and induce Granzyme B expression in immune cells.
ColabFold (AlphaFold2) was used to model complex structures of IL-21R and IL-2Rgamma bound to VHH paratopes, generating structural hypotheses on how distinct epitope bins and paratope orientations dictate productive receptor signaling geometry.
ProteinMPNN was applied to perform structure-based framework engineering by aligning VHH backbones to established VH dimer templates (PDB 7LU9 and 7L6M), yielding novel mutation sets (dsdAb(3)–(5)) designed to induce noncovalent VHH:VHH intramolecular dimerization.
Candidate designs were assembled onto antibody scaffolds and energy-minimized using Molecular Operating Environment (MOE), followed by binding interface evaluation and selection using PRODIGY.
The computationally engineered framework mutations enforced spatial rigidity and proximity between paratopes, successfully converting weak or inactive bispecific formats into highly potent cytokine mimetics without altering individual antigen-binding affinities.
End-to-end framework using generative deep learning to design de novo single-domain antibodies against the snake neurotoxin alpha-cobratoxin, leading to the experimental validation of low-nanomolar binders that achieved 100% in vivo survival in mice.
Conducted a head-to-head in silico comparison of three vhh-capable generative design tools (Germinal, RFantibody, and BoltzGen) across three structural scaffolds (9GCN, 7XL0, and 3EAK), fixing CDR loop lengths and targeting five specific epitope hotspot residues (D27, R33, K35, R36, and V37).
Evaluated complex predictions using AlphaFold3 (AF3) interface predicted TM-scores iptm combined with a target-aligned vhh structural self-consistency filter RMSD < 6Å), where Germinal generated a substantially higher fraction of passing candidates (46/300) than RFantibody (6/900) or BoltzGen (12/900).
Profiling against SAbDab, OAS, and INDI antibody databases demonstrated that Germinal generated more novel CDR3 sequences and broader loop conformation sampling, which significantly narrowed the compute-time required per successful candidate despite Germinal's higher raw GPU runtime per design.
Scaled the Germinal pipeline to ~8,000 trajectories using multi-stage filtering (including pDockQ2, PAE, spatial aggregation propensity, and AF3 re-prediction) to select candidates for experimental testing, finding retrospectively that AF3 ipSAE_min ranked true binders more effectively than standard ipTM
Novel dataset of 160 VHH-Fc profiled across 10 biophysical assays.
They demonstrated that tabular neural networks (TabICLv2, TabPFN v2.5) trained solely on 559 IgG heavy chains outperform intra-format VHH-Fc models in zero-shot predictions, showing that training data scale outweighs structural divergence between scaffolds.
Surface-driven properties transfer with high accuracy, heparin binding (rho=.82), hydrophobicity (HIC, rho=.63), and self-association (AC-SINS, rho=0.62), whereas thermostability (Tm2, rho=0.16) remains scaffold-dependent and requires direct measurement.
Augmenting models with simple surface property inputs (HIC and HAC) boosts prediction accuracy for complex liabilities like polyreactivity (PR-CHO, rho = 0.40 to 0.51).
Benchmarking the ability of co-folding models to distinguish nanobody binders and non-binders.
Evaluated four state-of-the-art structure prediction models (AlphaFold3, Boltz-2, Chai-1, and IntFold) across true binder ranking, out-of-distribution (OOD) sequence detection, and mutational sensitivity in nanobody–antigen complexes.
Discovered that no single confidence score excels across all tasks: while Boltz-2 achieved the highest median accuracy for true binder identification, local metrics like PLDDT (particularly in AlphaFold3) were far superior at catching OOD alanine-substituted sequences.
Contributed a original in vivo camelid immunization dataset targeting CD33, revealing that all evaluated models struggle to generalize when discriminating enriched binders from realistic immune repertoire background sequences.
Demonstrated that commonly used hard filtering thresholds (e.g., pAE < 10) can erroneously discard up to 75% of true binders, underscoring the need to combine complementary global interface and local CDR metrics rather than relying on single confidence scores.
Novel experimental and computational pipeline designed to characterize nanobody immune repertoires following immunization and phage display selection - NanoMAP.
It introduces a flexible clustering method that identifies clonal families by grouping sequences with similar V/J segments and CDR lengths, then applying a unique merging step that allows for minor CDR variations.
When benchmarked against MMseqs2 and Immcantation (SCOPer), NanoMAP scored higher on computational metrics (Silhouette, phenotypic quality, and stability) and showed better alignment with expert-curated "ground truth" labels.
AnewOmni, foundation model that unifies the design of small molecules, peptides, and antibodies into a single framework.
The team evaluated approximately 3,000 candidates for the "undruggable" KRAS G12D target by alternating between AnewOmni for CDR design and AlphaFold3 for structural validation.
Out of 7 synthesized nanobodies, the model achieved a 75% success rate (3 out of 4) when using a conservative structural consistency filter.
The most successful nanobody design demonstrated a high binding affinity with a Kd of 587 nM
The authors evaluated AlphaFold3, Boltz-2, and Chai-1 on their ability to distinguish cognate (correct) nanobody-antigen pairs from incorrect, non-binding pairings.
They used 106 experimental complexes and generated a combinatorial matrix of 11,132 shuffled non-cognate pairings to serve as ground-truth "incorrect" decoys.
Internal confidence scores (specifically ipTM) were very weakly predictive of true binding. In terms of Average Precision (PR-AUC), AF3 performed best, followed by Chai-1 and then Boltz-2.
Increased sampling improves structural geometry but does not help models "select" the correct binder. Most quality gains occur within 10–25 samples; deeper sampling primarily increases the number of plausible-looking false positives.
Authors propose a new method to train a nanobody structure predictor by using ‘blueprints’.
They developed a classifier (NbFrame) to identify whether the HCDR3 loop adopts a kinked (framework-contacting) or extended (solvent-exposed) conformation. This allows the model to use sequence-encoded priors and explicit constraints during the folding process.
The model itself is very lightweight and runs significantly faster than "heavy" models like AF or Boltz. NbForge achieves sub-second inference speeds, predicting structures in less than a second on both CPU and GPU. In comparison, models like AlphaFold3 or Boltz1 typically require tens of seconds to minutes per structure - MSA is a different story altogether.
The model matches the HCDR3 prediction quality of heavy models while being much more efficient. While AF3 and Boltz1 are more accurate at modeling the rigid framework, NbForge achieves parity in the hypervariable HCDR3 region—the part most critical for binding. Its speed and high recovery of disulphide bonds make it ideal for triaging millions of candidates in large-scale discovery campaigns.
Benchmarking of pretrained protein, antibody, and nanobody language model representations on a comprehensive suite of nanobody-specific tasks.
The authors introduce eight tasks spanning variable-region annotation, CDR infilling, antigen binding prediction, paratope prediction, affinity prediction, polyreactivity, thermostability, and nanobody type classification (e.g. VHH, VNAR, conventional antibody chains).
They evaluate generic protein LMs, antibody-specific LMs, and nanobody-specific LMs under a unified and standardized benchmark.
All backbone models are kept frozen, with task-specific lightweight heads trained on top to isolate representational quality.
No single model consistently outperforms others across all tasks, showing that nanobody-specific pretraining alone does not guarantee superior performance over antibody-specific or generic protein language models.