Blog Daniel Scott

How scoring reduces your synthesis queue without reducing your hit diversity

How scoring reduces your synthesis queue without reducing your hit diversity

The standard concern when someone suggests filtering a compound series by computational score before synthesis is this: you will discard the weird ones. The outliers. The compound that does not fit the model's learned patterns but turns out to be the most interesting structural lead in the whole set.

This concern is legitimate. It is also solvable, and the solution is not to skip filtering. It is to use the confidence band alongside the point estimate to decide which compounds to keep for their uncertainty value rather than their score value.

Why the concern is valid

Computational scoring models are trained on existing data. They learn which structural features correlate with binding, acceptable ADMET properties, and synthesizability based on compounds that have already been made and measured. A structurally novel compound, one that sits in a sparse region of chemical space relative to the training data, may score low not because it is a poor candidate but because the model is extrapolating from distant neighbors and the extrapolation is conservative.

If you filter your synthesis queue to keep only compounds above the 70th percentile composite score, you will disproportionately discard compounds in novel structural territory. The model's uncertainty about novel structures manifests as lower mean predictions, because the conservative prior in a well-calibrated model pulls predictions toward the mean when evidence is weak. You are not discarding bad compounds; you are discarding the compounds the model knows least about.

For programs that have already explored the well-precedented part of the scaffold's SAR, the novel outliers are often exactly where the interesting differentiation is. Filtering them out based on score alone defeats the purpose of the synthesis campaign.

How confidence bands solve the problem

The confidence band on each Alkira score reflects the model's epistemic uncertainty about that prediction. A narrow band means the model has seen many structurally similar compounds in training; the prediction is well-constrained. A wide band means the compound is structurally novel relative to training data; the actual property value could be substantially above or below the mean prediction.

This creates two distinct populations in a ranked compound list:

High score, narrow band: The model is confident this compound is a good candidate. Prioritize for synthesis. The score is a reliable signal.

Low score, wide band: The model is uncertain. The compound might be genuinely poor, or it might be a novel structure the model is underestimating. These compounds warrant a manual structural review before exclusion.

High score, wide band: The model predicts good properties but is uncertain. Could be a genuine high-quality candidate the model is seeing through the lens of structurally similar but not identical precedents. Worth including in the synthesis batch with awareness that the confidence is lower than the score implies.

Low score, narrow band: The model is confident this compound performs poorly. These are the safest cuts from the synthesis queue. If the model has high-confidence evidence across many similar compounds that this structural class underperforms, deprioritization is well-founded.

A practical workflow for retaining diversity

When we build synthesis queues using Alkira's output, we apply a two-pass filter rather than a single score cutoff.

Pass one: set a hard floor on composite score for the main synthesis batch. Compounds above the score floor and with confidence bands below a width threshold go into the primary queue. These are the well-characterized, highly ranked candidates.

Pass two: from the compounds that fell below the score floor, identify those with wide confidence bands. Sort this subset by structural diversity relative to the primary queue. Keep the top 20 to 30 percent of this set based on Tanimoto distance from primary queue compounds. These are the high-uncertainty structural outliers worth investigating.

The result is a synthesis batch that covers both the high-confidence best candidates and a structurally diverse set of uncertain outliers. Total queue size is typically 20 to 30 percent larger than it would be with a strict score cutoff, but the additional compounds are selected specifically because they represent structural territory the primary queue does not cover.

The queue reduction that actually matters

The 62% median queue reduction figure in Alkira's early data does not mean programs are synthesizing 62% fewer compounds. It means the compounds being deprioritized are the low-score, narrow-band compounds: those where the model is confidently predicting poor performance across multiple properties. These are not interesting outliers; they are well-characterized poor candidates.

The diversity concern applies to compounds with wide confidence bands, and those compounds are not being deprioritized at the same rate. A compound that scores 0.45 composite with a confidence band of +/- 0.18 looks different from a compound scoring 0.45 with a band of +/- 0.04. The first is genuinely uncertain; the second is reliably mid-range. The queue reduction comes disproportionately from the second type.

This is not a trivial distinction. It is the core of the answer to the diversity concern: filtering based on score alone discards interesting outliers, but filtering based on score plus confidence band retains them. The model's uncertainty is informative, and using it correctly changes the character of what gets deprioritized.

When the two-pass approach is not sufficient

There are program types where even the two-pass approach requires additional thought. Fragment merging campaigns, where the synthesis queue contains compounds designed to test specific fragment combination hypotheses, often include compounds that are intentionally novel and low-scoring. For these, the scoring output functions primarily as a property-risk flag, not a go/no-go signal. The synthesis decision is driven by the fragment-merging hypothesis; the score tells you what property liabilities you are taking on with that structural hypothesis.

Similarly, programs exploring mechanisms of action where the structural biology is poorly understood may have low confidence scores across the board simply because the target class is underrepresented in training data. In these cases, the scoring is most useful for ADMET and synthesizability assessment, where training data coverage is broader, and less useful for binding affinity prediction, where the structural biology gap limits reliability.

The tool is most powerful when used with an understanding of its training data coverage. Where that coverage is strong, let the score guide queue composition. Where coverage is sparse, treat the score as one input among several and weight the confidence band as the more informative signal.