Comprehensive repository supporting computational research in antibody engineering and therapeutic design, with a focus on antibody-antigen interactions
Introduction
ASD DB (Antigen Specific Antibody Database) is a comprehensive repository supporting computational research in antibody engineering and therapeutic design, with a focus on antibody-antigen interactions. This database addresses the challenge of dispersed binding data by aggregating and standardizing antibody-antigen interaction information from multiple sources.
The current release is version 2.0, which grows the database from 25 to 29 integrated public sources and stores every chain separately. Versions 1.0 and 1.1 remain available for anyone reproducing earlier work.
ASD DB is freely available for non-commercial research use under CC BY-NC 4.0; the licence text ships as LICENSE.txt next to the release folders. The database can be accessed through the Download section, where the files of each release can be downloaded from Google Drive, or via Google Colab, which provides a ready-to-use notebook with step-by-step guidance for loading and working with the database directly in Colab. The Colab notebook showcases the basics of downloading, loading, and filtering of the database, along with the necessary dependencies to make it happen. The overview of the dataset statistics, schema, and dataset composition is presented in the sections below.
Data sources
The primary dataset sources creating the database:
1. Genbank Database
2. SKEMPI 2.0
3. Peer-reviewed publications
4. Patent datasets
Five new sources absci, flab_hie2023, htm, scmc and seha join the 24 sources carried over from version 1.1. Selected for diversity: experimentally validated computational design and affinity maturation of known binders. Release figures are counted over the whole database; the statistics below exclude the buzz dataset. Full notes in the changelog.
The statistics describing the dataset exclude the buzz dataset, whose 524,346 samples of HER2 variants dominate every other source in size. Everything else, patents included, is counted. The figures below are for release 2.0, with 1.1 shown beneath each.
653,473
1.1: 675,413Total Antibody-Antigen
Records
343,848
1.1: 340,807Unique Antibodies
108,134
1.1: 104,587Complete Heavy/Light
Chain Pairs
9,985
1.1: 9,573Unique Antigens
Antigen Distribution: Primary focus on infectious diseases and cancer targets
Affinity Measurement Types: Includes quantitative metrics (Gibbs free energy changes, kinetic constants, IC₅₀) and qualitative binding assessments
Antibody Structure: The majority of entries include both heavy and light chains
The dataset contains 70,823 nanobodies among non-patent databases (110,853 with patents) and 132,399 scFv antibodies.
Nanobodies:
Alphaseq: 67,058 entries
Patents: 40,030 entries
Literature: 2,374 entries
Structures: 1,258 entries
AATP: 93 entries
OSH: 30 entries
RMNA: 10 entries
scFv:
Alphaseq: 131,645
SCMC: 754
The dataset was split into 3 confidence categories: medium, high and very_high, with very_high being the most dominant (52.0%) group:
The “very high” category means that both the sequences and methodology used for calculating the affinity were robust, an example of such a dataset is AbDesign.
The “high” confidence represents datasets which have to be manually curated or contain only the name and/or mutations of the target antigens/antibodies. The flab dataset is a member of this group, as it includes antigen names, not their respective sequences.
“Medium” category is a product of automated data discovery, which, although validated, can contain a degree of uncertainty, like the patents database; in version 2.0 htm is medium too.
ASD DB is freely available for non-commercial organisations for non-commercial research under CC BY-NC 4.0. Commercial inquiries are welcome via contact us.
Version 2.0 is the current release. Earlier versions stay available so published results can be reproduced against the exact files they used.
All releases and the licence sit side by side in the Google Drive folder. The notebooks download their release and walk through loading and filtering it; version 1.0 has no notebook of its own and is downloaded straight from its folder.
Statistics of each dataset can be seen below
Dataset
Is naturalantibody dataset
Records 2.0
Records 1.1
unique antigens
unique antibodies
affinity_type
confidence
absci
False
1,208
—
4
849
Kd (nM), bool
high
flab_hie2023
False
146
—
7
104
-log Kd (nM)
high
genbank
True
2,989
2,989
347
638
bool
high
htm
False
385
—
10
385
EC50 (nM), Kd (nM), ka (1/Ms)
medium
literature
True
8,226
5,636
1,788
6,919
bool
high
scmc
False
754
—
1
740
Kd (M)
high
seha
False
15
—
2
15
KD (M)
high
Confidence groups in version 2.0. Every row in both releases carries one; the three groups add up to the 1,177,819 rows of the release.
very_high
612,266 rows · 52.0%1.1: 612,266
high
374,684 rows · 31.8%1.1: 370,030
medium
190,869 rows · 16.2%1.1: 217,463
Data processing pipeline
1. Aggregation: Collection from distinct sources, resulting in 29 integrated datasets in version 2.0, up from 25 in version 1.1. buzz is one of the 29; it is held out of the statistics above, where its size would swamp every other source
2. Curation: Multi-stage pipeline combining automated extraction, normalization, and manual verification
3. Standardization: Common structure implemented across all studies
4. Validation: Automated feasibility checks and manual verification of critical datasets

Data structure
The dataset follows a standardized schema for all studies, containing Antibody-Antigen Pairs along with affinity measures.
From version 2.0 the sequences live in a chains struct rather than in flat columns: chains.antibody_chains holds the heavy and light chain under their own keys, chains.target_chains holds the antigen, and chains.antibody_chains_numbered holds the RIOT numbering of every antibody chain. The schema supports one entry per antigen chain and a heavy-only antibody; in the 2.0 files every row carries a single target chain under the key full, and nanobody rows carry a light entry that is empty or null. Versions 1.0 and 1.1 ship flat heavy_sequence, light_sequence and antigen_sequence columns instead of the chains maps; the 2.0 notebook shows how to derive them.
Fields in version 2.0
Name
Description
Example
dataset
Dataset abbreviation that allows for reading about its characteristics in this documentation
flab_koenig2017
affinity_type
Type of measurement performed to obtain the binding affinity. 16 distinct values in version 2.0, listed per dataset in the table above
-log Kd (M)
affinity
Value obtained using the method in affinity_type. Stored as a string; 8,714 rows hold the literal "nan"
9.348612972
processed_measurement
The measurement as a number, where affinity could be parsed
9.348612972
chains.antibody_chains
map<string,string> — keys "heavy" and "light"
{heavy: EVQLVESGGGLVQPGGSLRL…, light: DIQMTQSPSSLSASVGDRVT…}
chains.target_chains
map<string,string> — key "full" on every row today
{full: GQNHHEVVKFMDVYQRSYCH…}
chains.antibody_chains_numbered
map<string,map<string,string>> — full RIOT AIRR record per chain
{heavy: {locus: igh, v_call: IGHV3-23*04, cdr3_aa: ARFVFFLPYAMDY, …}, light: {locus: igk, v_call: IGKV1D-39*01, cdr3_aa: QQSYTTPPT, …}}
metadata.target_name
Name or label for the antigen target (e.g., "covid_gamma").
vascular endothelial growth factor (vegf)
metadata.target_pdb
Id of an antigen connecting it to a specific PDB database entry
2fjg
metadata.target_uniprot
Id of an antigen connecting it to a specific Uniprot database entry
—
metadata.source_url
Provenance: where the measurements were taken from
https://github.com/Graylab/FLAb
metadata.article_url
Provenance: the publication behind the measurement
https://www.pnas.org/doi/ 10.1073/pnas.1613231114
confidence
Level of confidence in binding of exact antibody sequence to an exact antigen sequence
high
nanobody
Boolean value specifying whether the record is a single-domain (VHH) antibody
false
scfv
Boolean value specifying whether records containing only heavy sequence contain a single-chain variable fragments, or antibodies
false
Format: Delta table. Version 2.0 — 20 parquet parts, 2.4 GB. Version 1.1 — 16 parts, 500 MB.
Dataset composition
The database includes multiple specialized datasets with diverse characteristics:
RMNA: Binary binding data for 10 antibodies against 2 antigen targets (RiVax and Apolipoprotein A1),
BUZZ: Affinity data for Trastuzumab mutations binding to HER2 (524,346 entries)
DLGO: Examines antibody mutations affecting binding against COVID-19 variants
aatp: KD affinity measurements for HER2 VHH (nanobody) mutations.
AAE: KD binding affinities for VH antibody variants binding to hen egg lysozyme
AntiBinder: the upstream project supplying the covid-19, hiv, met and biomap datasets below.
COVID-19: Neutralization data for SARS-CoV-2 with full VH/VL sequences (Cov-AbDab): 27,301 records, 32 antigens, 6,759 unique antibodies.
HIV: Heavy/light chain antibody sequences targeting HIV (LANL database): 48,008 records, 940 antigens, 192 unique antibodies.
BioMap: Binding ΔG values for antigen–antibody complexes across 8 species: 2,725 records, 594 antigens.
MET: Affinity data for 4,000 emibetuzumab variants targeting the MET receptor.
FLAB: Aggregated data from five publications:
FLAB_Hie2023 (146 records, 7 antigens; new in 2.0): Binding data for antibodies targeting viral glycoproteins including Ebolavirus glycoprotein, Influenza A hemagglutinin (HA), and its Group A subtype. flab_hie2022 (55 rows) was dropped, not renamed.
FLAB_Koenig2017 (4,275 datapoints): Extensive binding data for antibodies targeting growth factors such as Vascular Endothelial Growth Factor (VEGF).
FLAB_Rosace2023 (19 datapoints; 25 before version 2.0): Binding data focused on antibodies targeting Tumor Necrosis Factor-alpha (TNF-α) and SARS-CoV spike proteins.
FLAB_Shanehsazzadeh2023 (448 datapoints): Binding data for antibodies targeting the HER2 receptor, including multiple and zero mutation variants.
FLAB_Warszawski2019 (2,048 datapoints): Binding data associated with antibodies targeting hen egg lysozyme.
Patent databases: Paired antibody-antigen sequences extracted using NLP techniques: 190,484 records spanning 5,891 antigens.
Structure dataset: Binders extracted from structures in the PDB containing 3,969 antigen-antibody pairs (2,711 antibodies and 1,258 nanobodies).
ABDesign: Systematic point mutations at binding residues of CDR-H3: 672 records, 13 antigens.
Literature: Semi-manual dataset containing antibodies and antigens extracted from selected research articles: 8,226 records, 1,788 antigens, including 2,374 nanobodies.
AlphaSeq: Dataset involving mutations of antibodies to achieve better binding characteristics across 3 distinct antigen sequences: Human TIGIT, SARS-CoV2_RBD and Human PD-1 (the HER2 data sit in buzz).
OSH: Affinity binding data using Human NKp30 as a target.
ABBD: This dataset offers eight antibody-antigen cases enriched through heavy chain mutations (7 distinct antigen sequences).
Genbank: Derived from the previously published Genbank dataset: 2,989 records, 347 antigens.
absci: SPR data for de novo designed and lead-optimised antibodies: 1,208 records, 4 antigens (COL6A3, AZGP1, CHI3L2, IL36RA), binary call plus Kd (nM).
htm: High-throughput biophysical characterisation of an antibody library: 385 records, 10 targets (PD-L1, PD-L2, TIGIT, DCC, ROBO2, LOX1, IL23R-Fc, syncytin-2 and others), as Kd (nM), EC50 (nM) and association rate.
scmc: 754 SPR affinities in Kd (M) for single-chain variants binding HER2 (PDB 1N8Z); the only new source flagged as scFv.
seha: 15 KD (M) measurements against 2 antigens.
Dataset-Specific Notes
Due to the fact that all datasets in the ASD database represent unique research, their characteristics may differ. Below are important characteristics of some of the datasets:
abbd: The maximum -log Kd value is capped at 6.0, which generally indicates a non-binding interaction.
Patents: This database does not contain all interactions from the NaturalAntibody's patent database. Only complete and verified antibody-antigen pairs have been included.
Inclusion criteria
Transparency and completeness of data
Relevance to human health
Quantitative binding affinity measurements
Complete amino acid sequences for all biomolecules
Changelog
Version 2.0
The main goal of the release is better data and annotation quality: additional quality checks on top of version 1.1, all chains re-numbered with the current release of RIOT, and every measurement now carrying its source publication and data location.
Five new sources
Rows added per source. Selected for diversity: experimentally validated computational design and affinity maturation of known binders.
Changes in existing sources
The 26,979 patent records were removed on quality grounds.
Data model: chains are now normalised and kept separate
Antigens: the schema supports one entry per antigen chain in chains.target_chains; the 2.0 files carry a single chain per row, under the key full.
Nanobodies: the schema supports a heavy-only entry in chains.antibody_chains; in the 2.0 files nanobody rows carry a light entry that is empty or null.
Single-chain fragments: the two fused domains are annotated separately.
Numbering: every antibody chain carries its RIOT numbering in chains.antibody_chains_numbered, taken from the current RIOT release so records stay comparable across sources.
Provenance: every measurement carries its source publication and data location.
Version 1.1
Removed a duplicate copy of covid-19, read from two source files holding the same measurements. The copy kept names the SARS-CoV-2 variant each antigen belongs to. Affected 27,324 rows, leaving 27,301. The table now holds 1,199,759 rows, down from 1,227,083.
Resolved an issue with scFv constructs being numbered as one domain, so the heavy numbering actually described the light chain. Affected 48,683 rows in alphaseq.
Fixed an issue with the light chain being part of heavy_sequence. Affected 4 rows in structures-antibodies.
Addressed an issue with the antibody being part of antigen_sequence. Affected 446 rows in flab.
Fixed an issue with scFv and nanobody both set on the same row. Affected 31 rows in literature dataset.
Addressed a problem where a single light chain row was classified as a nanobody. Affected 1 row in literature dataset.
Corrected the antigen sequence in flab_koenig2017: the dataset now uses VEGF (8–109) instead of the previous 395-aa P15692 sequence. Affected 4,275 rows.
ASD DB is freely available for non-commercial organisations for non-commercial research under CC BY-NC 4.0. Commercial inquiries are welcome via contact us.
Version 2.0 is the current release. Earlier versions stay available so published results can be reproduced against the exact files they used.
All releases and the licence sit side by side in the Google Drive folder. The notebooks download their release and walk through loading and filtering it; version 1.0 has no notebook of its own and is downloaded straight from its folder.
Citing this work
The Antigen Specific Antibody Database accompanies the paper below. Please cite it when you use the data.
ASD: antigen-specific antibody database
Arkadiusz Czerwiński, Paweł Dudzic, Konrad Wójtowicz, Igor Jaszczyszyn, Weronika Bielska, Sonia Wrobel, Samuel Demharter, Roberto Spreafico, Victor Greiff, Konrad Krawczyk
mAbs 18(1):2623330, 2026. doi:10.1080/19420862.2026.2623330
@article{czerwinski2026asd,
title = {ASD: antigen-specific antibody database},
author = {Czerwiński, Arkadiusz and Dudzic, Paweł and Wójtowicz, Konrad and Jaszczyszyn, Igor and Bielska, Weronika and Wrobel, Sonia and Demharter, Samuel and Spreafico, Roberto and Greiff, Victor and Krawczyk, Konrad},
journal = {mAbs},
volume = {18},
number = {1},
pages = {2623330},
year = {2026},
doi = {10.1080/19420862.2026.2623330}
}