Blog Daniel Scott

The synthesis cost of picking the wrong lead compound first

The synthesis cost of picking the wrong lead compound first

The cost of a failed synthesis run does not usually appear in any budget line. It shows up in the calendar: a medicinal chemist spending 3 to 6 bench days on a compound that fails the first biochemical screen at 50 micromolar, then starting the cycle again with the next compound in the queue.

When we started building Alkira's scoring pipeline, one of the first things we did was map out what this failure pattern actually costs a program over a typical 12-month hit-to-lead campaign. The numbers were uncomfortable to look at.

What a failed synthesis run actually costs

The direct cost is labor and materials for a single synthesis attempt: reagents, solvent, analytical time for characterization, typically HPLC purification, and NMR confirmation of the target structure. For a moderately complex small molecule, this runs AUD 800 to 2,000 per compound depending on reagent complexity.

The indirect cost is what most programs undercount: the 3 to 6 bench days of a trained medicinal chemist. At fully loaded labor rates for a chemist at an early-stage biotech in Melbourne or Sydney, that translates to AUD 1,800 to 4,500 per compound attempt. Materials are a rounding error compared to time.

Now consider a program running a synthesis queue of 30 compounds per quarter, with 40% of them failing early screens at step one. That is 12 compounds per quarter generating no useful SAR data. Over a year: 48 synthesis attempts producing nothing actionable, at a combined direct plus labor cost of approximately AUD 125,000 to 310,000, with no data to show for it.

This estimate does not include the opportunity cost of the SAR coverage you did not build because the failed compounds occupied queue slots that better candidates could have filled.

Why the first compound matters disproportionately

In early drug discovery, the first synthesized lead compound anchors the SAR model. The team builds intuition around that structure: which modifications improved potency, which killed selectivity, which introduced tox flags. When the first compound turns out to be a poor lead, because it has a binding mode that does not generalize, or a reactive group that is easily missed in computational work, or a poor synthetic handle for the analogs the program needs, that anchoring effect works against you.

You do not just lose the time spent synthesizing the wrong first compound. You lose the weeks spent building SAR around the wrong scaffold, generating data that becomes less relevant once you realize the scaffold needs to change. The cost is multiplicative, not additive.

We saw this pattern clearly in a scenario we reconstructed during Alkira's validation work: a kinase program where the initial hit compound scored well on binding affinity prediction but carried a synthesizability score of 0.42 and a tox flag for CYP3A4 inhibition. The program synthesized it anyway, generated 15 analogs over three months, and then made a scaffold change. The three-axis ranking would have placed this compound at position 28 in the prioritized queue, below seven compounds with better balanced profiles across binding, tox, and synthesizability.

The hidden cost of queue management without scoring

Most early-stage programs manage their synthesis queue through a combination of docking scores, medicinal chemist intuition, and availability of starting materials. That is not a bad approach. Experienced medicinal chemists develop reliable intuition about which compounds are likely to be tractable. The problem is systematic blind spots.

Reactive functional groups that look innocent in 2D structure but behave as pan-assay interference compounds, PAINS, only become obvious after the biochemical data comes back. Compounds with high docking scores but poor aqueous solubility, which would fail any assay at relevant concentrations, are easily missed when the ranking decision relies on docking alone. Synthesizability edge cases, where a compound looks simple but requires a chiral resolution that takes two extra weeks to develop, are almost never caught before synthesis.

These are not individual mistakes. They are systematic gaps in any prioritization approach that evaluates each property independently rather than jointly.

What changes with three-axis pre-screening

Alkira's approach to this problem is to provide a joint rank on all three axes, with confidence bands, before any synthesis decision is made. The intent is not to replace the medicinal chemist's judgment but to give that judgment a better-structured input.

When a compound ranks in the top 10 on binding affinity but falls to position 23 after tox and synthesizability are included, the chemist now has a specific question to answer: is the binding advantage sufficient to justify the added synthesis complexity and the tox flag? That is a different decision from the original one, which was based on binding alone and did not surface the tradeoff explicitly.

The 62% median synthesis queue reduction we see across the teams using Alkira reflects this. Programs are not synthesizing fewer compounds in total; they are shifting the queue toward compounds that are more likely to be synthesizable and less likely to fail early screens. The synthesis success rate improves because the failure modes were visible before the queue decision, not after.

A note on what scoring does not fix

We should be direct about where the cost argument has limits. Computational pre-screening reduces synthesis queue failures. It does not eliminate them. A compound can clear all three scoring axes and still fail for reasons a scoring model cannot predict: unexpected polymorphism in the crystalline form, an impurity at the limit of detection that turns out to be a potent inhibitor, a solvent-dependent side reaction in the final deprotection step.

The scoring pipeline is a prioritization tool, not a prediction oracle. The goal is to shift the distribution of what gets synthesized toward better-balanced candidates, reducing the frequency of costly early failures. The irreducible uncertainty in synthesis and early screening does not disappear; it becomes a smaller fraction of the total program cost because the easily avoidable failures are no longer filling the queue.

The value is real but bounded. For teams where synthesis queue failures are systematic and recurring, the bound is meaningful. For programs with very short, tightly controlled compound series, the marginal impact is smaller. Knowing which category your program falls into is the first step in deciding how much weight to give the prioritization decision.