Build of 11 September 2026. Every figure on this page is transcribed from the build artifacts, not from notes.
Measurements come from ChEMBL together with additional curated data sources. The training endpoints are IC50, Ki, Kd, EC50 and Kb, so "more potent" spans binding affinity and functional potency together.
Activity is derived from the reported value and unit rather than read from a precomputed field, because in this data the precomputed column has been observed present only on erroneous rows. Duplicate readings collapse by median per receptor, structure and endpoint, never by the most potent value, because taking the best value reliably selects unit errors rather than the best measurement.
| Stage | Rows |
|---|---|
| Measurements read from all sources | 1,129,336 |
| After scope, endpoint, assay class and unit filters, then duplicate collapse | 353,120 |
| Fitting split | 333,727 |
| Held-out split | 19,393 |
This build applies no assay-quality threshold. An earlier build required an assay to carry at least ten compounds and a pActivity standard deviation of at least 0.5, which discarded about a third of the rows; removing that threshold is most of the growth in the table above. The split is temporal, on the publication year of the assay's source document: measurements published before 2023 fit the model and 2023 onward test it.
| Endpoint | Fitting | Held out |
|---|---|---|
| Ki | 152,648 | 7,650 |
| IC50 | 101,511 | 5,847 |
| EC50 | 72,220 | 5,432 |
| Kd | 4,425 | 282 |
| Kb | 2,923 | 182 |
Comparisons are matched by endpoint, so a functional reading is never ranked directly against a binding one.
This is the part of the method that most affects whether a result means anything. The bundle recognises 796 receptor names, but a name being recognised is not the same as the model having been fitted on it.
| Model | Receptors fitted | With a held-out score | Names recognised | Ligands in fitting |
|---|---|---|---|---|
| Potency | 251 | 182 | 796 | 185,102 |
| Selectivity | 255 | 170 | 796 | 61,742 |
The ranking pages group fitted receptors separately from merely recognised ones, and show each fitted receptor's own held-out accuracy beside its name. The full list, with per-receptor measurement counts and accuracies, is on the receptors page.
| Model | Receptors scored | Median | With 50+ comparisons | Median of those | Of those, below 0.60 |
|---|---|---|---|---|---|
| Potency | 182 | 0.638 | 137 | 0.619 | 57 |
| Selectivity | 170 | 0.688 | 75 | 0.675 | 21 |
A headline accuracy is a comparison-weighted average, so receptors carrying many held-out comparisons dominate it. Both headline figures reconcile exactly with the per-receptor table above.
A comparison is one row of three concatenated blocks. Nothing is scored on its own and subtracted.
| Block | Dimensions | Contents |
|---|---|---|
| Ligand | 1,038 | Morgan count fingerprint, radius 2, 1,024 bits, then 14 descriptors: molecular weight, heavy atom count, bond count, rotatable bonds, ring count, and per-element atom counts for C, N, O, S, F, Cl, Br, I and P |
| Sequence | 480 | ESM2 esm2_t12_35M_UR50D, mean pooled over residues with the start and end tokens excluded, full-length UniProt canonical sequence rather than a domain slice |
Potency rows are ligand, sequence, ligand, so 2×1,038 + 480 = 2,556 columns. Selectivity rows are sequence, ligand, sequence, so 2×480 + 1,038 = 1,998. The width is checked when a model loads and a mismatch raises rather than scoring silently.
Random forests, 300 trees, square-root feature sampling. Minimum leaf size 20 for potency and 8 for selectivity, reflecting the very different comparison counts.
Every comparison is entered twice, once as given and once with the two outer blocks exchanged and the label inverted, so no answer can depend on which side a ligand happened to be written on. At prediction time the same symmetry is applied as an average over both argument orders, which is why comparing A with B and comparing B with A return probabilities that sum to exactly 1.
| Model | Comparisons fitted | Rows after the swap | Held-out comparisons | Accuracy |
|---|---|---|---|---|
| Potency | 2,146,498 | 4,292,996 | 413,181 | 0.640 |
| Selectivity | 188,273 | 376,546 | 15,802 | 0.694 |