GPCR Foundation Model

Limitations

The things most likely to make a result misleading, stated plainly.

It only knows the receptors it was fitted on

Potency was fitted on 251 receptors and selectivity on 255. The bundle recognises 796 receptor names, so most names it accepts were never in training. Every accuracy figure published here covers comparisons between receptors in the fitted set. Nothing measures what happens outside it, and no claim is made.

The ranking pages separate the two groups and flag a run that used a pasted sequence or an unfitted receptor.

It is a comparator, not a potency predictor

It returns which of two things wins, never how tightly either binds. It has no opinion on whether either compound is active at all, so a confident ranking of two inactive compounds is still a confident ranking. Use it to order a set you already have reason to care about.

Accuracy is an average hiding a wide spread

Potency scores 0.640 overall. Across the 137 receptors carrying at least 50 held-out comparisons the median is 0.619 and 57 of them fall below 0.60. Selectivity scores 0.694 overall, and across its 75 receptors at that threshold the median is 0.675 with 21 below 0.60.

Receptors with only a handful of held-out comparisons can read 0.000 or 1.000 and neither figure means anything, which is why the counts are shown beside every accuracy on the receptors page.

Check the accuracy shown beside the receptor you picked before acting on its result.

Closely related receptors are the hard case

This is exactly where a selectivity answer would be most valuable, and it is where the model is weakest. Muscarinic receptors sit near chance, between 0.53 and 0.55, because M1 through M5 differ by little that a sequence embedding sees. Opioid receptors, which are easier to tell apart, reach 0.91.

Strength is a population statistic

The strength returned with a comparison is not the probability that the particular answer is right. It is a ranking signal whose reliability has been measured across a large held-out set. A comparison in the 0.70 to 0.80 band belongs to a population that is right 86% of the time for potency and 78% for selectivity. It does not mean that answer is 70% likely to be correct.

For potency the relationship also stops being monotonic above 0.80, falling from 0.887 to 0.869 on 1,537 comparisons, so no higher cutoff is offered. Selectivity stays monotonic all the way to its top band.

The held-out split is temporal, not compound-disjoint

Fitting uses older measurements and testing uses newer ones, which is the realistic setup: it asks whether the model helps on what comes next. A consequence is that a compound measured in both eras appears on both sides. That is deliberate, but it means the held-out figures are not a test on unseen chemistry, and they should not be read as one.

Functional and binding data are pooled

"More potent" spans IC50, Ki, Kd, EC50 and Kb together. Comparisons are matched by endpoint, so a functional reading is never ranked directly against a binding one, but the single headline accuracy averages over both kinds of question.

This build applies no assay-quality threshold

An earlier build required an assay to carry at least ten compounds and a pActivity spread of at least 0.5. This one admits every assay. That is a deliberate trade: it adds about a third more measurements and lifts both accuracies, but it also admits assays too small or too flat to order anything, and those contribute comparisons the model cannot do better than guess on.

These are the measured figures for the models retrained on the full measurement set on 11 September 2026, read from the build artifacts.