Benchmarking agents on making decisions on experimental workflows in biologics discovery.
TxBench-Antibody Discovery benchmarks AI agents on 100 decision-focused evaluations derived from public experimental data across 10 biologics competencies, ranging from target selection and binding kinetics to cellular pharmacology and candidate de-risking.
The benchmark evaluated 20 model-harness configurations, combining LLMs (e.g., Opus 5, Grok 4.6, Gemini 3.7 Flash, GPT-5.6) with execution scaffolds like Claude Code across 6,000 total runs using autonomous, deterministically graded workflows.
Frontier agents remain unreliable decision-makers, with the top configuration (Opus 5 with Claude Code) achieving only a 53.0% pass rate, while additional cost, token usage, or tool activity failed to consistently improve accuracy.
82.3% of agent failures stemmed from flawed scientific judgment and biological context misinterpretation rather than computational execution errors (17.7%), with model rankings fluctuating significantly depending on the specific biological discipline.
Discovery and engineering of bispecific single-domain antibody (sdAb)-based IL-21 mimetics (surrogate agonists) that target the IL-21R and IL-2Rgamma receptor subunits to activate downstream STAT3 signaling and induce Granzyme B expression in immune cells.
ColabFold (AlphaFold2) was used to model complex structures of IL-21R and IL-2Rgamma bound to VHH paratopes, generating structural hypotheses on how distinct epitope bins and paratope orientations dictate productive receptor signaling geometry.
ProteinMPNN was applied to perform structure-based framework engineering by aligning VHH backbones to established VH dimer templates (PDB 7LU9 and 7L6M), yielding novel mutation sets (dsdAb(3)–(5)) designed to induce noncovalent VHH:VHH intramolecular dimerization.
Candidate designs were assembled onto antibody scaffolds and energy-minimized using Molecular Operating Environment (MOE), followed by binding interface evaluation and selection using PRODIGY.
The computationally engineered framework mutations enforced spatial rigidity and proximity between paratopes, successfully converting weak or inactive bispecific formats into highly potent cytokine mimetics without altering individual antigen-binding affinities.
End-to-end framework using generative deep learning to design de novo single-domain antibodies against the snake neurotoxin alpha-cobratoxin, leading to the experimental validation of low-nanomolar binders that achieved 100% in vivo survival in mice.
Conducted a head-to-head in silico comparison of three vhh-capable generative design tools (Germinal, RFantibody, and BoltzGen) across three structural scaffolds (9GCN, 7XL0, and 3EAK), fixing CDR loop lengths and targeting five specific epitope hotspot residues (D27, R33, K35, R36, and V37).
Evaluated complex predictions using AlphaFold3 (AF3) interface predicted TM-scores iptm combined with a target-aligned vhh structural self-consistency filter RMSD < 6Å), where Germinal generated a substantially higher fraction of passing candidates (46/300) than RFantibody (6/900) or BoltzGen (12/900).
Profiling against SAbDab, OAS, and INDI antibody databases demonstrated that Germinal generated more novel CDR3 sequences and broader loop conformation sampling, which significantly narrowed the compute-time required per successful candidate despite Germinal's higher raw GPU runtime per design.
Scaled the Germinal pipeline to ~8,000 trajectories using multi-stage filtering (including pDockQ2, PAE, spatial aggregation propensity, and AF3 re-prediction) to select candidates for experimental testing, finding retrospectively that AF3 ipSAE_min ranked true binders more effectively than standard ipTM
Novel dataset of 160 VHH-Fc profiled across 10 biophysical assays.
They demonstrated that tabular neural networks (TabICLv2, TabPFN v2.5) trained solely on 559 IgG heavy chains outperform intra-format VHH-Fc models in zero-shot predictions, showing that training data scale outweighs structural divergence between scaffolds.
Surface-driven properties transfer with high accuracy, heparin binding (rho=.82), hydrophobicity (HIC, rho=.63), and self-association (AC-SINS, rho=0.62), whereas thermostability (Tm2, rho=0.16) remains scaffold-dependent and requires direct measurement.
Augmenting models with simple surface property inputs (HIC and HAC) boosts prediction accuracy for complex liabilities like polyreactivity (PR-CHO, rho = 0.40 to 0.51).
Evaluated ten co-folding protocols on a benchmark of 412 human monomeric antigen complexes
They demonstrate that recent architectures like Protenix v2 (52% top-1 success on post-cutoff Fv complexes) substantially outperform earlier methods like AlphaFold-Multimer (20%) and Protenix v0.5, which serves as the open-source AlphaFold3 reproduction baseline.
Modeling accuracy scaled inversely with CDR-H3 loop length. Short loops (<11 residues) were predicted accurately across all methods (0.4–0.7 Å Calpha RMSD), whereas long loops (>16 residues) remained challenging but showed distinct improvement in Protenix v2 (median 2.67 Å RMSD vs. 3.22–3.88 Å in older methods).
CDR-H3 accuracy was identified as the primary structural feature distinguishing successful complex predictions DockQ \ge 0.49 from failures, demonstrating a strong inverse correlation with overall DockQ scores (Spearman rho = -0.74) and high metric discrimination (AUROC = 0.91).
Prediction failures in unconstrained models were dominated by sampling limitations (failure to generate native-like poses) rather than ipTM ranking failures; providing idealized epitope constraints or increasing seed depth successfully rescued many of these failures by guiding the search space.
Training specificity models across several domains.
Models were trained strictly on public 1D sequence data (amino acids, DNA/RNA nucleotides, and chemical SMILES) across six biological domains, requiring no 3D structural information.
Base sequence encoders were kept frozen while small 5–7 million parameter projection heads were trained in under an hour per fold on a single GPU using thermodynamic contrastive learning.
Development was driven entirely via natural-language prompts by a domain expert with zero coding experience, with all numerical claims verified by an independent AI auditor.
Top-1 target retrieval accuracy reached up to 98.0% from pools of 512 candidates, successfully generalizing to unseen rare HLA alleles and non-canonical binding targets that rule-based tools miss.
Used as re-ranking filters alongside existing computational workflows, the models dramatically boosted precision, such as raising CRISPR off-target prediction precision from 33.2% to 94.0%.
Novel way (ProteinDPO) to apply pre-trained models to biophysical readouts.
ProteinDPO is the first framework to apply Direct Preference Optimization (introducing a novel scalar-weighted DPO objective) to align protein generative models with experimental biophysical data, without overfitting like traditional supervised fine-tuning.
Despite training exclusively on small monomeric stabilities, ProteinDPO generalizes zero-shot to accurately rank the thermal melting temperatures of multichain antibodies and score antibody-antigen binding affinities.
Applied to H5N1 influenza hemagglutinin, it generated stabilized variants yielding up to a 32C improvement in thermal stability while retaining strong nanomolar binding affinity to broadly neutralizing anti-HA antibodies.
A restriction free reproduction of the antibody design workflow germinal
Replaces proprietary dependencies (PyRosetta, IgLM) with an open-source toolchain (OpenMM, AbLang1, sc-rs) and fixes multi-chain bugs, enabling unrestricted academic and commercial deployment.
Demonstrates that AbLang1-guided hallucination significantly increases initial cofolding pass rates (e.g., 33.7% vs. 18.6% for PD-L1) with equal or higher structural confidence, at the cost of a ~1.5x increase in per-trajectory compute time.
Uses hard-coded placeholder values for three energy metrics which degrades ensemble selection and disables the interface hydrogen-bond filter and lacks wet-lab experimental validation of the generated binders.
Benchmarking of some structural prediction methods on antibody-related tasks.
Benchmarked five computational tools across two core tasks: AlphaFold3 (AF3), ImmuneBuilder (ABodyBuilder2/ABB2), and IgFold for antibody variable fragment (Fv) structure prediction, as well as AF3, dyMEAN, and GRAMM for antigen–antibody complex structure prediction and docking.
Tools were tested on 50 non-redundant humanized antibody–antigen Fab complexes from the Protein Data Bank (PDB), filtered for resolution < 3.0Å and released after December 19, 2023, to eliminate training data overlap.
All Fv predictors achieved high backbone accuracy (mean TM-score > 0.97), but AF3 demonstrated statistically significant advantages in global Fv geometry and hypervariable CDR-H3 loop modeling (median CDR-H3 RMSD of 0.86 Å, compared to 1.34 Å for ABB2 and 1.52 Å for IgFold). In complex prediction, AF3 substantially outperformed dyMEAN and GRAMM, generating reliable docking poses for 46% of complexes, whereas classical rigid-body docking (GRAMM) and epitope-guided modeling (dyMEAN) almost completely failed.
Paratope and epitope residue recovery is strictly dependent on initial docking accuracy. When docking succeeds, AF3 reliably identifies interface residues (F1 score ~ 0.85/0.87), salt bridges (84% recall), and non-bonded contacts (75% recall).
Novel method, DeepAAAssembly, to model antibody antigen complexes, beating AF3.
DeepAAAssembly introduces a hybrid framework that pairs deep learning inter-chain distance predictions with flexible Monte Carlo sampling, overcoming the lack of strong co-evolutionary signals in antibody-antigen interfaces.
The pipeline feeds Voronoi tessellation contact geometry, AntiBERTy embeddings, and a Triangular Awareness module into a multi-column CNN, creating continuous multi-peak distance energy landscapes to guide both global orientation sampling and local CDR loop refinement.
Evaluated on Docking Benchmark 5.5, it surpassed AF3 with a 12.9% higher average DockQ score on best-generated models (0.454 vs. 0.402) and boosted the overall modeling success rate (DockQ > 0.23) from 47.8% to 56.7%.
Demonstrates strong generalization on CASP15/16 blind targets by successfully correcting severe initial orientation errors, elevating previously failed structural predictions into acceptable, physically plausible models