Blog Daniel Scott

Binding affinity vs synthesizability: what to do when they point in opposite directions

Binding affinity vs synthesizability: what to do when they point in opposite directions

What the conflict is telling you

A compound with a predicted binding affinity of 8.9 kcal/mol and a synthesizability score of 0.31 is a specific kind of problem. The binding prediction is saying: this structure has features consistent with potent binding at your target. The synthesizability score is saying: the retrosynthetic graph for this molecule does not resolve into accessible building blocks through known reaction pathways with the route lengths and step counts that predict practical bench synthesis.

Both scores can be correct simultaneously. The conflict is not a model error. It is information about where the compound sits in the space of what is chemically desirable versus what is currently chemically practical. How you use that conflict productively determines whether the compound represents a research direction worth developing or a dead end to deprioritise.

What a synthesizability score below 0.4 actually means

Synthesizability scores are typically derived from retrosynthesis graph analysis, trained on large sets of compounds with known synthetic routes. A score of 0.31 does not mean the compound cannot be made. It means the retrosynthetic analysis, working backward from the target SMILES through known reaction transforms, did not identify routes meeting the practical criteria: reaction steps within a manageable range, intermediates accessible from commercial building blocks, no functional group collisions requiring extensive protecting group strategy.

Several structural patterns consistently produce low synthesizability scores. Dense stereochemical complexity, particularly compounds with three or more chiral centres where the relative configuration matters for biological activity, creates route length and protecting group overhead that pushes scores down. Macrocycles and bridged bicyclics often score poorly because known efficient ring-forming reactions are limited. Unusual heterocyclic motifs that are not well-represented in the training data for retrosynthetic transforms also produce uncertain low scores.

It is worth distinguishing between a score that is low because the route is genuinely difficult and a score that is low because the molecule contains a structural motif that is underrepresented in the training data. The model has a blind spot for chemistry it has seen rarely, not just chemistry that is objectively hard.

Using the conflict as a scaffold hypothesis rather than a dead end

A compound with exceptional predicted binding but low synthesizability is most useful as a structural hypothesis that defines a binding geometry worth targeting. The question to ask is not "how do we make this specific molecule" but "what is the core structural feature driving the binding prediction, and can that feature be embedded in a more synthetically tractable scaffold?"

In practice, this means running a set of bioisosteric or analogue queries around the high-binding, low-synth compound. Which features of the SMILES are most responsible for the binding score: the ring system geometry, the hydrogen bond acceptor vectors, the hydrophobic burial? If the binding prediction is driven primarily by a specific ring fusion geometry, are there related ring systems with similar geometry that score better on synthesizability?

Scaffold hopping in this direction, using the high-binding, low-synth compound as the binding geometry reference and exploring structurally adjacent scaffolds for synthesizability, is more productive than either discarding the compound or attempting the difficult synthesis. The difficult molecule is pointing you at a binding hypothesis. Your job is to find a way to test that hypothesis with a compound you can actually make in a reasonable number of steps.

When a hard synthesis is justified anyway

There are programs where committing bench time to a low synthesizability compound is genuinely defensible. If your program has exhausted all the high-binding, tractable analogues within the current series and the difficult compound represents the only predicted path to meaningful potency improvement, the synthesis may be worth attempting. If you have access to a CRO with specialist expertise in the relevant chemistry, the practical difficulty may be lower than the synthesizability score implies.

The question is resource proportionality. A synthesizability score of 0.31 on a target with limited other options is different from a synthesizability score of 0.31 when your queue contains fifteen compounds scoring above 0.65 and within 0.5 kcal/mol of the difficult compound's binding prediction. In the second scenario, the difficult compound is deprioritised by resource logic even if its absolute binding prediction is marginally better.

We are not saying low synthesizability scores are an automatic disqualifier. We are saying they carry weight in proportion to how constrained your synthesis capacity is and how many alternatives exist at comparable predicted binding.

The case for simultaneous synthesis queue ranking

One pattern we see in practice is that programs evaluate binding and synthesizability as sequential filters rather than simultaneous scoring axes. First they select for binding above a threshold, then they check whether the top binders are synthesizable. This sequential approach discards compounds that would rank well on a combined score.

Consider two compounds: one predicting at 8.7 kcal/mol with a synthesizability score of 0.72, and another at 9.1 kcal/mol with a synthesizability score of 0.31. The sequential filter gives the second compound to the queue because it passes the binding threshold first and only then runs into the synthesizability problem. A simultaneous multi-axis ranking would weight both scores from the start and likely place the first compound higher, since its combined profile is better for a program that needs to move through synthesis rounds efficiently.

The practical implication is to treat binding and synthesizability as parallel inputs to prioritisation, not as a binding-first gate followed by a synthesizability check.

Documenting the conflict for program memory

A compound that registers exceptional binding but fails the synthesis queue should be recorded as a program reference, not silently dropped. As the program progresses and structural understanding of the binding site deepens, a previously intractable synthesis may become approachable through a new route or building block. Or a new structural analogue may emerge in a library screen that captures the binding geometry more efficiently.

Keeping a log of high-binding, low-synth compounds as potential future starting points means your computational work from round one remains accessible for round five, when you might have better chemistry to test those hypotheses. The binding score identified something real. Preserving that signal, even when the compound does not make the immediate synthesis queue, is part of using computational scoring as a cumulative research tool rather than a one-shot filter.

The conflict between binding and synthesizability is not a failure of the scoring system. It is the system doing its job: surfacing the tension between what is chemically desirable and what is practically buildable, so you can make that tradeoff deliberately rather than by default.