Analysis of developability data from 33 internal Biogen programs, covering 18,540 antibodies.
Focused on three dimensions: hydrophobicity (HIC), polyspecificity (PSR), and self-association (AC-SINS).
Labeled subsets included 4,594 (PSR), 1,792 (HIC), and 7,727 (AC-SINS) sequences.
Benchmarked three PLMs: ESM2 (general-purpose), plus IgBert and IgT5 (antibody-specific).
Domain-adaptive fine-tuning consistently boosted antibody-specific PLMs, but often degraded ESM2 performance.
Antibody-specific PLMs generally provided better embeddings for PSR and AC-SINS, while ESM2 remained highly competitive for HIC.
Perplexity was only weakly correlated in aggregate, but showed significant association with PSR/AC-SINS failure when controlled for a fixed light chain
Benchmarked 113 teams on predicting five key developability traits: hydrophobicity, thermostability, self-association, expression titer, and polyreactivity.
Models were trained on the GDPa1 dataset (246 antibodies) and blindly tested on GDPa3 (80 diverse antibodies from OAS).
While cross-validation (CV) results were promising, performance plummeted on the test set e.g., self-association dropped from a 0.653 CV Spearman's rho to 0.356.
Hydrophobicity was the most predictable (rho = 0.708), while expression titer was the most challenging (rho = 0.310).
Winning models varied by assay; for example, team AbDevelop won for self-interaction, while microcrisprtm2 led in thermostability.
Prompt-based, in-context prediction of antibody developability properties using large language models, rather than training separate predictors per property.
As a baseline, they evaluate TxGemma, a therapeutics-specific multimodal LLM that supports task switching via prompts and is fine-tuned using LoRA.
The study relies on a very large antibody dataset (~876k heavy chains) with in-silico–computed biophysical developability properties, combining sequence-based and structure-based predictors.
Models are trained and evaluated using prompts that include antibody sequences together with partially observed property/value pairs, asking the model to infer a missing property for a query sequence.
To prevent shortcut learning where the model ignores context and relies only on sequence, the authors introduce AB-context-aware training, which applies a random latent transformation jointly to context properties and targets during training, forcing explicit use of contextual information.
By simulating batch effects, they show that standard fine-tuned TxGemma degrades sharply as batch bias increases (from ~0.99 Spearman ρ with no bias to ~0.95 with moderate bias and ~0.58 with strong bias), whereas context-aware training remains robust even under strong batch effects.
Faster free to use version of NetMHCPIIpan for deimmunization.
The model uses a small neural network (MLP) trained on one-hot encoded 15-mer peptides. To improve accuracy, it identifies the strongest 9-residue ‘binding core’ within those peptides and aligns them before scoring.
The training data isn't experimental directly but distilled from NetMHCIIpan-4.3. The authors took 75,000 peptides, ran them through the original tool, and weighted the results by how common 97 different DRB1 alleles are in North America to create a single risk score.
By predicting the final risk score in one pass, rather than calculating 97 individual allele bindings, it runs 300,000x faster while keeping a 95% correlation with the original tool's results.
To prove it works on real drugs, they tested it against MAPPs data (physical peptide presentation) from vatreptacog alfa, a drug that failed clinical trials due to immune reactions. It successfully flagged the same high-risk mutations as the much slower original software.
Its main value is the speed and differentiability wrt NetMHCIIpan, so it can sit in generative AI pipelines. It allows designers to screen millions of protein variants for "self" vs "non-self" peptides in minutes rather than weeks.
FLAb2 substantially expands existing antibody benchmarks, introducing the largest public dataset to date with a strong focus on developability rather than binding alone.
A broad spectrum of models is evaluated, including generic protein language models, antibody-specific models, structure-aware predictors, and simple physics-based baselines such as charge and pI calculations.
Zero-shot predictions from pretrained protein models are generally weak and unreliable for antibody developability. Surprisingly, simple charge-based features often outperform large models for properties such as aggregation, polyreactivity, and pharmacokinetics.
Intrinsic properties (e.g. thermostability, expression) are substantially easier to predict than extrinsic or context-dependent properties such as polyreactivity, pharmacokinetics, or immunogenicity.
Few-shot learning improves performance, but even the best models typically achieve only moderate correlations (ρ ≈ 0.4–0.6) on statistically robust datasets, highlighting the difficulty of the task.
Incorporating structural information improves predictions, particularly in the zero-shot setting, and helps reduce biases present in sequence-only models.
Many pretrained models primarily capture evolutionary signal, effectively measuring distance from germline rather than true developability. Encouragingly, this germline bias largely disappears once models are fine-tuned in a few-shot setting.
Scaling model size alone provides limited benefit. Given sufficient training data, simple one-hot encodings paired with small neural networks can match or outperform billion-parameter protein language models, emphasizing that data quality and quantity matter more than model scale.
Introduced a novel pairing predictor for VhVl chains with a clever strategy to sample negative pairs.
Defines three negative sampling strategies:
Random pairing, where heavy and light chains are shuffled without constraints.
V-gene mismatching, where non-native pairs are generated by combining VH and VL sequences drawn from different V-gene families, but within biologically plausible V-gene segments. This captures realistic but unobserved combinations that could occur during recombination.
Full V(D)J mismatching, where heavy and light chains are paired using completely distinct germline origins across V, D, and J gene segments. This produces negative examples that are maximally diverse yet biologically meaningful, reflecting combinations never seen in natural repertoires.
Shows that the space of possible VH–VL germline combinations is far larger than what is observed in public datasets, revealing non-random biological constraints on pairing.
Demonstrates that models trained on V-gene and especially VDJ mismatched datasets achieve the highest and most generalizable performance, outperforming existing methods such as ImmunoMatch, p-IgGen, and Humatch — confirming that biologically grounded negative sampling is key to robust VH–VL pairing prediction.
Benchmarking of computational models for predicting antibody aggregation propensity (developability) using size-exclusion chromatography (SEC) readouts.
Developed an experimental dataset of ~1,200 IgG1 antibodies, measured for monomer percentage and ΔRT (difference in retention time) relative to a reference.
Evaluated four main prediction pipelines: Sequence + structure-based features (hand-crafted biophysical features from Schrödinger, using AlphaFold2 or ImmuneBuilder for structure). PLM (protein language model) pipeline (e.g., ESM2-8M, fine-tuned or LoRA-adapted). GNN (graph neural network) pipeline using residue graphs from predicted structures. PLM + GNN hybrid pipeline combining sequence embeddings with structural graphs.
Two structure prediction tools were benchmarked: AlphaFold2 (high accuracy, slow) and ImmuneBuilder (faster, antibody-optimized, slightly less accurate).
The sequence + structure feature model achieved the highest accuracy overall, but low sensitivity (missed many problematic antibodies).
The PLM-only pipeline performed nearly as well and offered a much faster, high-throughput solution, making it attractive for early screening.
The GNN and PLM + GNN approaches performed comparably, with GNN slightly better for ΔRT predictions but more variable.
Using ImmuneBuilder instead of AlphaFold2 reduced sensitivity slightly but greatly improved speed without major loss of accuracy.
So all pipelines performed similarly within a narrow performance range, but faster, less resource-intensive approaches (PLM and ImmuneBuilder-based pipelines) offer strong trade-offs for early-stage developability screening.
Introduces TNP, a nanobody-specific developability profiler inspired by TAP.
Uses six metrics: total CDR length, CDR3 length, CDR3 compactness, and patch scores for hydrophobicity, positive charge, and negative charge.
Thresholds are calibrated to 36 clinical-stage nanobodies.
In vitro assays on 108 nanobodies (36 clinical-stage + 72 proprietary) show partial agreement with TNP flags, indicating complementary—but not perfectly correlated—assessments.