Benchmarking of pretrained protein, antibody, and nanobody language model representations on a comprehensive suite of nanobody-specific tasks.
The authors introduce eight tasks spanning variable-region annotation, CDR infilling, antigen binding prediction, paratope prediction, affinity prediction, polyreactivity, thermostability, and nanobody type classification (e.g. VHH, VNAR, conventional antibody chains).
They evaluate generic protein LMs, antibody-specific LMs, and nanobody-specific LMs under a unified and standardized benchmark.
All backbone models are kept frozen, with task-specific lightweight heads trained on top to isolate representational quality.
No single model consistently outperforms others across all tasks, showing that nanobody-specific pretraining alone does not guarantee superior performance over antibody-specific or generic protein language models.