Blog Priya Mehta

When a tox flag should override a strong binding score

When a tox flag should override a strong binding score

The conflict is real on both sides

A compound predicted at 9.1 kcal/mol binding affinity is worth noticing. That score places it among the strongest candidates in most submission batches. But a concurrent hERG inhibition flag at high confidence is also real information. The temptation is to reason that a good binding score overrides the tox concern, or that the tox model has a high false positive rate and can be discounted. Both arguments are sometimes correct and sometimes wrong.

What we have found, working through these conflicts with medicinal chemistry teams, is that the resolution almost always depends on three factors: the severity and confidence of the tox flag, the reversibility of the tox liability, and the degree of chemical headroom available to address it. Getting clear on all three before deciding whether to synthesize is faster and cheaper than synthesizing and finding out.

Understanding what the hERG flag is actually saying

The hERG potassium channel is implicated in cardiac repolarisation. Inhibition of hERG can prolong the QT interval, and at sufficient exposure, this creates arrhythmia risk. As a result, hERG inhibition is one of the most consistently measured safety endpoints in early drug discovery, and training data for hERG inhibition models is unusually abundant relative to other tox endpoints.

This means hERG predictions are generally among the more reliable tox flags a computational scoring tool returns. The false positive rate is not negligible, but it is lower than for endpoints with sparse training data. When Alkira returns a high-confidence hERG flag, the prior probability that the compound does inhibit hERG is meaningfully elevated above baseline.

That said, the hERG flag is not a single piece of information. The structural correlates of hERG activity are reasonably well understood: basic nitrogen atoms at a specific distance from a hydrophobic aromatic group are a common pharmacophoric pattern. A flag driven by a clearly present basic amine and an aromatic ring separated by the right linker carries different weight than a flag on a compound with ambiguous structural features. Looking at what drove the flag, ideally at the feature attribution level if your scoring tool provides it, helps calibrate how much weight to give it.

When the tox flag should move the compound down the queue

There are cases where the tox flag clearly overrides the binding score, regardless of how strong the binding prediction is.

First, if the flag is for a genotoxic or mutagenic liability rather than hERG, the bar for proceeding is much higher. Reactive electrophiles, Michael acceptors, or alerts consistent with DNA intercalation are structural concerns that are expensive to engineer around and carry program-level risk if they surface in subsequent assays. A strong binding score on a compound with a predicted covalent-genotoxic alert is not a reason to prioritise it. It is a reason to deprioritise it and look for close structural analogues that retain the binding features without the reactive moiety.

Second, if multiple independent tox flags are flagged at high confidence simultaneously, the probability of at least one being a true positive rises. A compound with concurrent hERG, phospholipidosis, and reactive metabolite flags is telling you that several structural features are flagging across different models trained on different underlying data. The binding score is not a sufficient counter-argument.

Third, if the target indication involves dosing at high exposure levels over long durations, the risk from a given tox liability is higher than for a target requiring low doses or short-course administration. Even a moderate hERG prediction is worth more caution if the compound profile suggests it will achieve high systemic concentrations.

When the binding score can outweigh the flag

We are not saying a tox flag always ends a compound's candidacy. There are situations where proceeding to synthesis and experimental measurement is the right call, with the flag on record.

If the hERG flag is at medium rather than high confidence, and the compound has exceptional predicted binding with a clean profile on every other axis, it is reasonable to synthesize and measure the hERG patch-clamp result before making a final call. Many hERG flags at medium confidence do not translate to measurable channel inhibition in vitro. The computational model is giving you a prior; the experimental result is the update.

If the tox flag is for a liability that has an established medicinal chemistry fix, and the fix is compatible with retaining the binding features, the compound is better treated as a template for a second-generation design than as a dead end. A basic amine that is driving a hERG prediction can sometimes be replaced by a less basic bioisostere, or the linker distance modified to disrupt the pharmacophoric pattern. If that modification is straightforward, the compound with the flag is still valuable: it tells you the binding geometry is right, and the scaffold is worth developing.

The key question is whether the flag points to a structural feature that is intrinsic to the binding pharmacophore or incidental to it. If the hERG-active feature overlaps directly with the binding-active feature, engineering it out will likely compromise binding. If it is a substituent that could be modified without touching the binding contacts, there is a path forward.

The decision logic we apply at Alkira

In practice, when we see a high binding score alongside a tox flag, our ranked output places the compound in the shortlist with the tox flag displayed prominently rather than filtering it out entirely. The chemist making the synthesis decision sees both pieces of information: the binding rank and the flag severity.

The practical heuristic is: a compound with a high-confidence tox flag for a serious liability (genotoxic, highly confident hERG above a threshold, reactive metabolite at primary structural alert level) does not make the first-round synthesis queue. It is not discarded from the program, it is marked for structural modification before synthesis. A compound with a medium-confidence hERG flag and a clean rest-of-profile makes the queue with the flag documented, and the hERG result is measured in the first round of assaying.

The logic is asymmetric by design. False negatives (missing a real safety issue) are more costly to a program than false positives (delaying a good compound for a round of structural assessment). The scoring tool's job is to surface the conflict clearly. The medicinal chemist's judgment governs the final decision.

Using the conflict to generate structural hypotheses

A conflict between a strong binding score and a tox flag is not just a prioritisation problem. It is also a structural design question. The compound telling you that it binds well but triggers hERG is giving you a starting point for analogue design: what is the minimum modification to address the flag while preserving the binding?

This is where scoring is most useful as a design tool rather than a filter. If you can run a quick round of in silico analogue scoring, varying the substituents most likely responsible for the hERG pharmacophore, you can identify whether there are close analogues that retain the binding score while reducing the tox flag. Scoring those analogues before any synthesis happens is far cheaper than synthesizing the original compound, measuring the hERG liability, then designing the second-generation analogue from scratch.

The binding score and the tox flag together define the design space worth exploring. That is more useful than either piece of information alone.