Blog James Okafor

Building a lean compound library for a target class with sparse data

Building a lean compound library for a target class with sparse data

Sparse data is a different problem, not an absent one

When a medicinal chemistry program targets a protein with limited published binding data, the instinct is sometimes to wait until more SAR is available before using computational scoring. That instinct is understandable but usually backward. The programs that benefit most from computational prioritisation are often those with the least data, because they have the fewest resources to absorb wasted synthesis runs.

Sparse data does not mean no data. It means the available signal is narrower and the transfer assumptions matter more. Working with that signal productively requires a different approach to library construction than you would use for a target with dense ChEMBL coverage.

Start with what the target class tells you

Even when direct activity data for your specific target is limited, the broader target class often has well-characterised pharmacophoric requirements. A poorly-characterised serine hydrolase may have only a handful of published inhibitors, but the serine hydrolase family as a whole has established binding geometry preferences: a nucleophilic reactive group or tight-binding warhead positioned to interact with the catalytic serine, a hydrophobic cap occupying the acyl chain binding tunnel, and a leaving group region tolerant of polar substitution.

A lean initial library built to systematically explore variation across these three regions, while staying within physicochemical property bounds appropriate for the assay format, gives a scoring model something useful to work with even before you have any measured activity from this program. The fragments in your library are not random: they are hypotheses about what the binding geometry permits, framed in terms the model can evaluate.

When we work with programs at this stage, the library construction question is usually: given the target class pharmacophore, what molecular diversity are we deliberately covering in this set, and what are we choosing to defer? A library of 120 compounds with deliberate variation across three pharmacophoric vectors is more useful than 300 compounds that cluster around a single starting hit.

Transfer learning from related scaffolds: where it helps and where it fails

Scoring models trained on related scaffolds can produce useful rank-ordering for a sparse-data target, but the degree of transfer depends on two factors: how similar the binding site geometry is across the related set, and how close the query compounds are (in descriptor space) to the training set.

For a target with close structural homologues, graph neural network-based scoring models can generalise binding affinity predictions reasonably well because the features they have learned, hydrogen bond donor acceptor geometry, ring system burial, hydrophobic contact area, are partially conserved. The confidence bands on these predictions will be wider than for a target with dense direct training data, and you should treat them accordingly: as rank-ordering signals within a series, not as absolute binding energy estimates.

The transfer fails when the binding site geometry differs substantially from the training set. An allosteric site on a target that has no close structural homologue in the training data will produce low-confidence, unreliable scores. Here the right response is not to discard the scoring entirely, but to weight synthesizability and ADMET predictions more heavily in the prioritisation, since those property predictions are more target-agnostic. A compound that is synthesizable, ADMET-clean, and occupies underexplored chemical space relative to your assay hits is a reasonable synthesis candidate even when the binding score is uncertain.

Designing for learning, not just for activity

In a sparse-data program, each synthesis run has a dual purpose: testing a compound hypothesis and generating SAR data that improves future decisions. This means library composition should explicitly consider structural diversity across scaffolds, not just potency optimisation within a single scaffold family.

A common mistake is to rapidly converge on the highest-scoring compounds from the first round and build a second-generation library as close analogues of those. If the first-round scoring was uncertain, convergence on a narrow series amplifies the uncertainty. A better approach: take your top 10 scoring compounds, group them by scaffold, and ensure your second-generation library includes at least two or three scaffold classes from that set rather than just the one with the highest individual score.

This is not saying you should synthesize less potent compounds for the sake of diversity. It is saying that when your scoring confidence is low, diversifying the synthesis queue hedges against the model having learned the wrong features from sparse data. The experimental results from a structurally diverse set give you much better signal for model calibration than a dozen close analogues of a single hit.

Practical construction: size and property bounds

For a sparse-data program with a small synthesis capacity, a lean initial library of 80 to 150 compounds is usually more defensible than 300. The value of a compound in this context is not just its potential binding activity but the SAR information it generates per gram of synthetic effort. Compounds that are easy to make, easy to assay, and structurally distinct from other library members carry higher information value than an additional close analogue of an existing hit.

Property filtering at the library construction stage is more aggressive in sparse-data programs than in mature SAR programs. Since you cannot rely on experimental ADMET data from previous rounds, computational ADMET flags carry more weight. Compounds with predicted CYP liability, poor aqueous solubility, or high P-glycoprotein efflux scores should be deprioritised at the library stage, not rescued by strong binding predictions. The binding score in a sparse-data setting carries too much uncertainty to justify synthesizing a compound with likely assay liabilities.

Tanimoto diversity filtering against your existing set is useful here. Before finalising library selections, check that no two compounds in the queue have a Tanimoto coefficient above 0.7 on ECFP4 fingerprints. If two compounds are that similar and the scoring model has low confidence on both, you are essentially running the same experiment twice.

Interpreting rank order from a sparse-data scoring run

When you receive a ranked shortlist from scoring a sparse-data library, the rank ordering is more informative than the absolute scores. A binding affinity prediction of 7.4 kcal/mol on an uncertain target should not be interpreted as a likely Ki of around 50 nM. It should be interpreted as: this compound's structural features are consistent with those of reasonably potent binders in the related training data, placing it above the median of this submission batch.

The compounds worth synthesizing first are those that appear in the top tier across multiple scoring axes. A compound that ranks in the top quartile for predicted binding, has a synthesizability score above 0.65, and carries no high-confidence ADMET flags is a reasonable first synthesis candidate even in a sparse-data setting, because multiple independent model signals are converging.

Conversely, a compound at rank 1 for predicted binding alone, with a synthesizability score below 0.4 and a medium hERG flag, should move down the queue. In a mature program with dense SAR, a compelling predicted binding score might justify working around the synthesis difficulty. In a sparse program where the binding prediction itself carries high uncertainty, that reasoning does not hold.

When to revisit the scoring model as data accumulates

As your experimental rounds proceed and you accumulate IC50 data for 20 to 30 compounds, you reach a point where you can start calibrating the scoring model against your own program data. This is the moment when the sparse-data constraint lifts, at least partially. Feeding your measured binding data back into the scoring run, as a constraint set for rescoring the next library batch, dramatically improves rank-ordering accuracy for your specific target.

At this stage, the library construction strategy shifts: you have enough measured SAR to support more targeted analogue synthesis within a validated scaffold series, while maintaining some exploration of the structural periphery. The lean 80-150 compound initial library was designed to get you to this point efficiently. What comes next is a different problem.