6Target Discovery
Every drug starts with a bet on a molecule: if we push on this protein, the disease gets better. Target discovery is the science of choosing that molecule well. The hard part is not finding genes that look involved in a disease — modern genomics hands you thousands. The hard part is finding a gene that is causal (perturbing it actually moves the disease, not just a bystander that lights up alongside it) and druggable (a real molecule can bind it and change what it does). Get this choice right and everything downstream has a chance; get it wrong and no amount of clever chemistry saves the program. This chapter follows the spine of the whole part: the problem, the models that attack it, and what is still genuinely hard.
6.1The problem: which gene or protein to drug
Start with a disease and a wish list of genes that correlate with it. A genome-wide association study (GWAS, a scan that tests millions of common genetic variants for statistical association with a trait; see the genetics primer, Chapter 5) will hand you hundreds of loci for a common disease. Expression studies add thousands more genes that are turned up or down in patient tissue. This is the trap the field spent a decade learning to avoid: association is cheap and causation is rare. A gene can be differentially expressed because it drives the disease, or because the disease changed it, or because both share an upstream cause. Only the first kind is a target. Drug a bystander and your compound may work beautifully in the assay and do nothing in a patient.
So a usable target must clear two gates. The first is causality: intervening on the gene product must change the disease, ideally in the direction you can achieve with a drug. The second is druggability — the capacity of a protein to be modulated by a drug-like molecule with enough affinity and selectivity to matter. Classic druggable proteins have well-formed binding pockets: kinases, G-protein-coupled receptors, ion channels, nuclear receptors. Estimates put the "druggable genome" at roughly 3,000 to 4,500 of the ~20,000 human protein-coding genes, and fewer than 700 are actually hit by an approved drug. Transcription factors and scaffolding proteins are often called "undruggable," though that label is partly a statement about where industry has looked, not a law of chemistry — new modalities (PROTACs, molecular glues, antibodies, oligonucleotides) keep moving the line.
Intuition
A target is a lever, not a symptom. You are not looking for the gene most correlated with the disease; you are looking for the one that, when you pull it, the disease actually moves — and that you can build a lever for.
Collaborator
A wet-lab partner asks: "This gene is the top hit in our patient RNA-seq — isn't that our target?" Differential expression tells you a gene is involved, not that it is upstream. It could be responding to the disease rather than causing it. Before we commit, we want an orthogonal line of evidence that pushing the gene changes the phenotype — human genetics, or a direct perturbation in cells.
6.2Models that attempt it
No single model outputs "here is your target." What the field has built instead is a stack of methods that each supply one kind of evidence, plus an aggregator that reconciles them.
Genetics-informed prioritization is the anchor, because a genetic variant is nature's own perturbation experiment, assigned roughly at random at conception and fixed before the disease begins. If a variant that changes a gene also changes disease risk, the arrow points from gene to disease, not the other way — this is what makes genetics a causal signal where expression is only correlational. Turning a GWAS locus into a specific gene is itself a modeling problem, because the associated variant is usually not in the gene it acts on; the Open Targets Platform's locus-to-gene (L2G) model is a trained classifier that scores which nearby gene a locus most likely acts through, using distance, molecular-QTL colocalization, and functional genomics. The Platform then combines that gene assignment with rare-variant burden tests, animal models, and known drugs into a single target-disease evidence score (Buniello et al., 2025). Mendelian randomization (MR) sharpens the same logic into a quasi-experiment: it uses variants that mimic a drug's action (say, variants in PCSK9 that lower LDL cholesterol) as an instrument to estimate the causal effect of modulating that target, effectively a natural randomized trial run over a lifetime. MR is how genetics can forecast both efficacy and side effects before a molecule exists. How to read a specific variant's mechanism — coding vs regulatory, gain vs loss of function — is the subject of variant-to-mechanism (Chapter 13).
Network and omics embeddings attack a different gap: most disease genes are not yet in any GWAS, so you want to generalize from known biology. Represent genes and proteins as nodes in an interaction or pathway graph, learn embeddings, and score an unknown gene by its proximity to known disease genes ("guilt by association"). This surfaces plausible novel targets but inherits the correlation problem — proximity in a network is a hypothesis, not a cause.
Perturbation readouts supply the experimental arm. Perturb-seq couples a pooled CRISPR screen to single-cell RNA sequencing: knock out or knock down each of thousands of genes, then read the full transcriptome of each perturbed cell to see what that gene actually controls (Replogle et al., 2022). This is causal by construction — you did the intervention — and genome-scale versions now profile all expressed genes across millions of cells. It tells you a gene's downstream program in a dish, which is exactly the functional evidence a network embedding can only guess at. Turning those readouts into predictions for unseen perturbations is a modeling frontier (GEARS, scGPT-perturb) central to cell engineering (Chapter 11); the single-cell foundation models that back them are covered in the appendix.
Literature and LLM-assisted synthesis sits on top. The evidence for any target is scattered across decades of papers, and reading it is what a biologist spends weeks doing. Retrieval-augmented and agentic LLM systems now traverse biomedical knowledge graphs and literature to assemble and rank target hypotheses; OriGene, for instance, nominated and then experimentally validated previously underexplored targets in liver and colorectal cancer (Zhang et al., 2025). These systems are fast and tireless synthesizers, but they hallucinate confident mechanisms and inherit every bias in the literature, so their output is a lead to check, not a verdict.
The empirical payoff justifies the emphasis on genetics. Drug targets with human genetic support are roughly twice as likely to survive clinical development to approval — the finding that reoriented the industry (Nelson et al., 2015; King et al., 2019). The most careful recent estimate, with a decade more genetic data, puts the boost at about 2.6-fold and shows it grows with confidence in the causal gene, but does not depend on the variant's effect size (Minikel et al., 2024).
Collaborator
A skeptical statistician asks: "If genetic support doubles success, why isn't every approved drug genetically supported?" Because genetics is a strong filter, not a requirement. Many good targets have no common-variant signal (the biology is essential, or the variation is too rare to detect), and plenty of genetically supported targets still fail for reasons genetics can't see — toxicity, delivery, or a disease that has already progressed past the target. Doubling a base rate that starts around 10% still leaves most programs failing.
6.3What they do well and what is still hard
These methods are genuinely good at three things: converting messy association into a ranked shortlist, forcing a causal question ("does perturbing this move the phenotype?") to the front, and surfacing candidates a single expert would miss. The genetics-first strategy is the closest thing the field has to a validated prior on success.
What remains hard starts with the same word the chapter opened on. Causality is still only approximated. MR assumes its instrument affects the disease only through the target (no pleiotropy), which is often untestable; L2G scores are calibrated probabilities, not proofs; a Perturb-seq hit in an immortalized cell line may not fire in the tissue and disease context that matters. Each method reduces the correlation-causation gap without closing it.
Novelty is the deeper limit. Every method here leans on prior knowledge — GWAS needs common variants of detectable effect, network embeddings need a gene to sit near known biology, LLMs can only synthesize what has been written. So the systems are strongest exactly where we already suspected the answer and weakest for the truly new biology that would open an undrugged disease. The genetic-support advantage is itself evidence of this: it rewards targets that human variation happens to illuminate, which is not the same set as the targets that matter.
Finally, validation is expensive and slow. A nominated target is a hypothesis; confirming it means knockouts, animal models, and eventually a molecule and a trial, at a cost of years and millions. Because in-silico nomination is now cheap and validation is not, the bottleneck has moved: the constraint is no longer generating candidate targets but affording to test them, and choosing which few to test is where a wrong call is most costly. This dish-to-patient translation gap — a target that behaves in cells or mice but not in people — is the recurring theme of the "doing it for real" chapters (Part V) and the reason a doubled success rate still means most programs fail.
Common trap
Treating an aggregated target score as a decision. Open Targets, an MR p-value, and an LLM's summary are inputs to a judgment, not the judgment. Two targets with the same score can differ wildly in tractability, safety, and how confidently the causal gene is even identified — read the underlying evidence, not just the number.
References
- Buniello, A., Suveges, D., Cruz-Castillo, C., Bernal Llinares, M., Cornu, H., Lopez, I., Tsukanov, K., Roldan-Romero, J. M., et al. (2025). Open Targets Platform: facilitating therapeutic hypotheses building in drug discovery. Nucleic Acids Research.
- King, E. A., Davis, J. W., & Degner, J. F. (2019). Are drug targets with genetic support twice as likely to be approved? Revised estimates of the impact of genetic support for drug mechanisms on the probability of drug approval. PLOS Genetics.
- Minikel, E. V., Painter, J. L., Dong, C. C., & Nelson, M. R. (2024). Refining the impact of genetic evidence on clinical success. Nature.
- Nelson, M. R., Tipney, H., Painter, J. L., Shen, J., Nicoletti, P., Shen, Y., Floratos, A., Sham, P. C., et al. (2015). The support of human genetic evidence for approved drug indications. Nature Genetics.
- Replogle, J. M., Saunders, R. A., Pogson, A. N., Hussmann, J. A., Lenail, A., Guna, A., Mascibroda, L., Wagner, E. J., et al. (2022). Mapping information-rich genotype-phenotype landscapes with genome-scale Perturb-seq. Cell.
- Zhang, Z., Qiu, Z., Wu, Y., Li, S., Wang, D., Liu, Y., Zhou, Z., Hu, Y., et al. (2025). OriGene: A Self-Evolving Virtual Disease Biologist Automating Therapeutic Target Discovery. bioRxiv.
Check yourself
Questions a sharp collaborator might ask on this chapter. Pick an answer to see whether it holds up.
-
A gene is the single most strongly upregulated transcript in patient tissue versus healthy controls. Why is this, on its own, weak evidence that it is a good drug target?
The core problem is direction of causation, not measurement. A gene can be upregulated because it causes the disease, because the disease changed it, or because both share an upstream driver — only the first makes it a target. This is exactly why human genetics and direct perturbation, which carry causal information, are prized over expression correlation. -
Why is a genetic variant treated as a stronger causal signal for a target than a case-control expression difference?
The temporal ordering is what buys causality: because the variant is assigned quasi-randomly and precedes the disease, the disease cannot have caused the variant. Note the third option is actually false in a way that matters — the causal variant usually sits outside its target gene, which is why the locus-to-gene mapping problem exists at all. -
Human genetic support roughly doubles a target's probability of clinical success. What does this NOT imply?
Genetic support is a strong filter, not a requirement: the base success rate is low, so even a 2 to 2.6-fold boost leaves most programs failing, and many good targets have no detectable common-variant signal. Minikel and colleagues also showed the effect grows with confidence in the causal gene but is largely independent of the variant's effect size. -
What makes a Perturb-seq readout a fundamentally different kind of evidence from a network-embedding prediction?
Perturb-seq actually knocks the gene out and reads the downstream transcriptome, so the effect it reports is caused by your intervention. A network embedding infers 'guilt by association' from proximity to known disease genes, which is a testable hypothesis, not a demonstrated effect. Both still face the dish-to-patient gap: a causal effect in an immortalized cell line may not hold in the relevant tissue. -
Why are LLM-agent systems that synthesize the target literature best treated as producing leads rather than decisions?
These systems are fast, tireless synthesizers but inherit the literature's biases and can assert mechanisms that are not supported, so a nomination is a hypothesis to check. Systems like OriGene close part of this loop by pairing nomination with experimental validation, which is what turns a plausible lead into evidence. -
A collaborator argues that because in-silico target nomination is now cheap and fast, the main bottleneck in target discovery has been solved. What is the flaw?
Cheap nomination moves the bottleneck rather than removing it: because generating candidates is now abundant and experimental validation is costly, the binding constraint is choosing the few hypotheses worth the years and millions to confirm. A wrong pick is most expensive precisely at this validation gate.