Blog James Okafor

How to read synthesizability scores before you book bench time

How to read synthesizability scores before you book bench time

The synthesizability score in your ranked output is not a promise. It is a probability estimate based on retrosynthetic graph search, and understanding what that means changes how you use it at the bench.

What the model is actually computing

When Alkira computes a synthesizability score, the underlying process traces a retrosynthesis graph from your candidate structure back toward commercially available building blocks. At each bond disconnection, the model evaluates whether the implied transformation is precedented in its training set, whether the intermediate is chemically stable, and whether the forward reaction would proceed with acceptable selectivity under standard conditions.

The score aggregates those individual step-plausibilities into a single normalized value between 0 and 1. A score of 0.85 means the model found a retrosynthetic route with a sequence of individually plausible disconnections. It does not mean the synthesis is fast or simple. It means a route was found that draws on precedented chemistry at each step.

This distinction matters because "precedented" and "practical in your lab right now" are not the same thing.

Why 0.7 is a starting point, not a threshold

Many teams treat scores above 0.7 as a green light for synthesis. That is a reasonable first-pass heuristic, but it conflates two separate questions: whether a viable route exists and whether that route is feasible in your specific program context.

A score of 0.78 on a compound with four stereocenters tells you a retrosynthetic path was found. It does not tell you that the diastereoselective steps in that path are achievable without protecting group choreography that takes weeks to develop. The retrosynthesis model treats each stereocenter as a manageable complexity factor. It cannot weigh your specific reagent availability, your team's depth in that reaction class, or the time cost of characterizing each chiral intermediate by chiral HPLC.

The score is computed on structure. Practical synthesis time is computed on context.

Three patterns that produce optimistic scores

Chiral center density. Compounds with three or more stereocenters frequently score in the 0.65 to 0.75 range, which looks moderate to low. The score reflects that each center is individually achievable. It does not capture the combinatorial difficulty of setting multiple centers simultaneously with the correct relative configuration. A molecule scoring 0.71 with five chiral centers may demand more bench time than a molecule scoring 0.60 with two centers but a known, efficient synthesis route.

Orthogonal protection requirements. Reactive functional groups that require sequential protection and deprotection steps often score better than they synthesize. The retrosynthesis graph can disconnect the target molecule cleanly, but the protecting group strategy is an implicit cost the model does not represent as explicit synthetic steps. A target that requires carbamate, silyl ether, and benzyl protection in sequence to reach a late intermediate might score 0.72 while actually requiring eight steps in practice.

Constrained ring systems. Macrocycle-like geometries and strained bicyclic scaffolds are a third case. The model may find disconnections, but the ring-closure step that enforces conformational constraint typically has low effective molarity, making the cyclization sluggish and sensitive to dilution, temperature, and concentration. This does not show up in the disconnection score; it shows up in the actual yield at the key ring-forming step.

We are not saying scores become uninformative in these structural contexts. We are saying the number should start your analysis, not finish it.

Reading the confidence band alongside the score

Alkira reports a confidence band with each synthesizability score. A compound at 0.82 +/- 0.04 and a compound at 0.82 +/- 0.14 represent different epistemic situations, even with identical point estimates.

Narrow bands indicate the model has seen many structurally similar compounds in its training data. The estimate is well-constrained. Wide bands indicate structural novelty: the model is extrapolating from more distant structural neighbors, and the actual synthesizability could land meaningfully above or below the estimate.

When you see a high score with a wide band, that compound is worth a closer structural look. The model may be pattern-matching on a superficially similar but mechanistically different precedent. When you see a moderate score with a very narrow band, the model is telling you, with some confidence, that this class of structure consistently falls in that range.

A practical scenario: PARP inhibitor analogs

Consider a scenario familiar to programs running PARP inhibitor analogs. The parent phthalazinone scaffold scores 0.81. A set of 34 bioisosteric replacements at the phthalazinone nitrogen shows a spread from 0.53 to 0.89. The highest-scoring analogs introduce simple alkyl or aryl substitution at that nitrogen. The lowest-scoring analog replaces the nitrogen with a spirocyclic ring connection.

A prioritization based purely on synthesis score would direct effort toward the 0.89 compounds. But looking at the structures, those top-scoring compounds are simple modifications with routes nearly identical to the parent. They provide incremental SAR coverage. The 0.62 spirocyclic compound is harder to make but introduces genuine three-dimensional structural novelty that could differentiate binding selectivity.

The score is telling you about route access. It is not telling you which compound is worth the synthesis effort for your program's specific SAR objectives. Those are different questions.

Calibrating your read over time

After working with synthesizability scores over several compound series, most medicinal chemists develop a personal calibration: which score ranges match their own reaction-feasibility intuition, and where the model tends to be systematically optimistic or conservative for the structural classes they work with regularly.

That calibration is worth keeping. When the model says a compound is accessible and your reaction memory says otherwise, you have two options worth exploring. Either the model has found a route you had not considered, which is genuinely useful, or the model is over-generalizing from structurally adjacent but mechanistically distinct precedents. Both are informative for how you trust the next score in that structural class.

The score is fastest when it agrees with you. It is most valuable when it does not.