Research note: predicting cancer cytotoxicity from multi-omics signals
The most valuable signal in early cancer drug design is not raw potency, it is selectivity. A note on why predicting cytotoxicity well means predicting biology, and how multi-omics nudges models toward mechanism instead of memorisation.
Killing a cancer cell in a dish is trivial. Bleach manages it. The hard and far more interesting problem is killing the diseased cell while sparing the healthy one, reliably, across the genetic chaos of real tumours. That is why we think the signal that matters most in early cancer drug design is selectivity, not potency, and why predicting cytotoxicity properly means predicting biology rather than a single tidy number.
Potency is a scalar, cytotoxicity is a context
A compound’s effect on a cell depends on the cell: its lineage, its mutations, its expression state, and which pathways it leans on to stay alive. Train a model to predict "is this compound cytotoxic?" from structure alone and it will happily learn shortcuts, memorising which chemotypes were toxic in the training lines and overfitting to lineage instead of mechanism. It looks brilliant on familiar cells and quietly falls apart on the unfamiliar ones you actually care about.
Why multi-omics changes the question
Transcriptomic and proteomic signals hand the model something closer to a mechanistic view: not just the molecule, but the machinery it acts on and the state of the cell it acts in. Proteomic abundance is an especially honest readout of the proteins a compound can actually engage, more so than expression alone. Fusing these modalities behaves like a regularizer, nudging the model toward mechanism and away from lineage memorisation, which is precisely what you need to generalise to cell lines it has never seen. This is the thesis behind Gnosis II (disease-cell selectivity) and Gnosis IV (downstream gene-expression response): model the response, not just the endpoint.
The prediction worth having is the selectivity window
Framed properly, the quantity you want is a difference: diseased-cell effect minus healthy-cell effect. Optimising that window changes what generative chemistry is even asked to do. It stops rewarding enthusiastic poisons and starts rewarding molecules with an actual therapeutic margin. It also reframes failure in a useful way: a molecule that is potent everywhere is not a lead, it is a liability, and the sooner a system can say so, the more synthesis and screening budget it quietly rescues.
Downstream response as an early-warning system
Gene-expression response is interesting for a second reason. It can catch on-target-but-wrong-outcome effects and hint at resistance before the bench does. Two compounds can hit the same target and produce completely different transcriptional programs, one that marches the cell toward death and one the cell shrugs off and routes around. A model that predicts the downstream program, not just the binding event, lets chemists ask a better question of every candidate: not "does it bind?" but "does the cell respond the way we wanted?"
The obligatory honesty section
These are early capabilities and we treat them as such. Predictions are for research use only. Every score should travel with a calibrated sense of when it can be trusted. And a benchmark only means something when it is reported on held-out, out-of-distribution data, not on the examples the model already had memorised. The purpose of an integrated system is not to replace the experiment. It is to make the next experiment a smarter bet.
Integrated evidence beats isolated predictions. The interesting frontier is not a more potent molecule, it is a more selective one, chosen because we modelled how the cell actually responds.
Vecentra Research
Vecentra outputs are intended for research use only and are not validated for clinical diagnosis or therapeutic decision-making.