9Protein and Binder Design
Prediction and design are mirror images. A structure predictor (Chapter 8) reads a sequence you hand it and tells you the shape it folds into. Design runs the arrow backwards: you specify what you want — a shape, a surface that grips a chosen target, an active site that performs a reaction — and the model invents a molecule that meets the spec. This is de novo design: building a protein or small molecule from scratch rather than editing one nature already made. The mental model for the whole field fits in one line. Generative models now propose novel molecules by the thousand, cheap silicon filters throw most of them away, and the survivors go to a bench that decides which handful actually work. The gap between "the computer likes it" and "it works in a tube" is the whole story of the chapter.
9.1The problem: designing new molecules
The forward problem has one right answer and you can check it: fold the sequence, compare to the crystal structure. The inverse problem has many right answers and no cheap way to check any of them. A thousand different sequences can fold to the same backbone; countless backbones could present the surface you want. You are searching an astronomically large space (twenty amino acids at every position, hundreds of positions) for the rare members that fold, stay soluble, and do a job. And "does it do the job" cannot be read off a file — it needs a physical measurement.
The field is organized by what you specify. Fix a backbone shape and you want a de novo protein that folds to it. Fix a target (a disease protein) and you want a binder, a small designed protein that clamps onto a chosen patch of its surface with high affinity. Fix a chemical reaction and you want an enzyme, the hardest ask, because you are designing a precise geometric arrangement of catalytic residues, not just a complementary surface. A parallel track designs small molecules to sit in a protein's pocket, the generative face of structure-based drug design, which gets its own chapter next (Chapter 10).
Intuition
Prediction asks "what does this sequence do?" Design asks "what sequence does this?" — same physics, but the second question has a haystack of answers and you can only test a few straws.
Collaborator
"Why not just mutate a natural protein that already sort of works?" That is directed evolution or engineering, and it is often the right call. De novo design earns its keep when nature has no good starting point — a binder to a target with no known antibody, an enzyme for a reaction biology never invented, a fold built to spec. You trade a warm start for an unconstrained blank page.
9.2Models that attempt it
The workhorse for structure is diffusion, the same denoising idea as in image generators applied to atomic coordinates. RFdiffusion starts from a cloud of random 3D points and iteratively denoises them into a plausible protein backbone, and you can condition it: hold a target fixed and it grows a binder against a chosen patch; scaffold a set of catalytic residues and it builds a protein around them (Watson et al., 2023). Chroma is a comparable programmable backbone generator that can be steered by symmetry, shape, and other constraints (Ingraham et al., 2023). These models give you a backbone — a shape — but no sequence.
Turning a shape into a sequence is inverse folding, and the standard tool is ProteinMPNN (Dauparas et al., 2022). Given fixed backbone coordinates, it predicts which amino acids fold to that backbone, decoding residues with a message-passing graph network over the atoms.
Common trap
ProteinMPNN is not a masked-language-model encoder like ESM (Chapter 8), and it does not read a sequence. It reads 3D coordinates and writes a sequence. ESM learns amino-acid statistics from millions of sequences; ProteinMPNN learns the structure-to-sequence map from the PDB. Different input, different job. Confusing the two is a common interview stumble.
For small molecules, the analogue is target-aware generation. TargetDiff diffuses atom coordinates and types inside a protein pocket, generating 3D molecules shaped to fit (Guan et al., 2023), while docking models like DiffDock frame pose-finding — where and how a molecule sits — as diffusion over the ligand's translations, rotations, and torsions (Corso et al., 2022). These share the design pipeline's spirit but inherit small-molecule headaches: synthesizability and the crudeness of scoring binding from structure alone, which the next chapter takes up in full (Chapter 10).
The pieces only become a method when you chain them into a loop.
The middle step, the self-consistency filter, is what makes the whole thing tractable. You designed a sequence for a backbone; now feed that sequence to an independent structure predictor (AlphaFold2 or ESMFold, Chapter 8) and ask whether it folds back to the shape you intended. If the predicted structure matches — low RMSD between design and refold, high confidence (pLDDT) — you trust it; if not, you discard it. Because you generate thousands and predict cheaply, you can throw away 99% and still have plenty to order. A newer twist skips the separate backbone step entirely: BindCraft runs backpropagation through AlphaFold2 itself, optimizing a binder sequence directly toward a high-confidence complex ("hallucination"), and reports experimental success rates of 10-100% depending on target (Pacesa et al., 2025). DeepMind's AlphaProteo is a closed-system generator reporting high hit rates and nanomolar affinities across several targets in a single screening round (Zambaldi et al., 2024).
9.3What they do well and what is still hard
Read the numbers with the pipeline in mind. When a paper says "90% of designs bind," it means 90% of the survivors after self-consistency and developability filtering, screened in an assay tuned to that method. The in-silico funnel is spectacular and the raw wet-lab hit rate is far lower — many designs never fold, never express, or bind too weakly to matter. Aggressive filtering is what closes the gap, which is exactly why the loop, not any single model, is the unit of progress.
The sharpest divide is binding versus function. Designing a surface that sticks to a target is now close to routine for many targets, because sticking is a matter of shape and chemical complementarity — the thing diffusion and self-consistency are built to get right. Designing genuine catalysis is far harder: an enzyme must hold several residues in near-perfect geometry, stabilize a fleeting transition state, and cycle substrate in and product out, all while the protein breathes. Progress is real and recent — designed serine hydrolases now show meaningful catalytic efficiency and crystal structures matching the design, using RFdiffusion plus ensemble scoring of active-site preorganization (Lauko et al., 2025), and RFdiffusion2 scaffolds active sites from bare functional-group geometry — but designed enzymes still land orders of magnitude below evolved ones in turnover, and each success tests only a handful of the many designs made.
Collaborator
"Your model loves this binder. Will it survive my lab?" Passing self-consistency is necessary, not sufficient. Developability is the rest: does it express (can cells actually manufacture it), stay soluble instead of aggregating, tolerate storage, avoid triggering an immune response? A predictor scores none of these directly. Treat a passing design as a hypothesis to be measured, and budget for the ones that fold on-screen but clump in the tube.
Note
The static-structure limitation from Chapter 8 bites twice as hard here. A predictor that returns one rigid snapshot cannot fully judge a design whose function is motion — an enzyme's catalytic cycle, a binder that must accommodate a flexible target. Self-consistency checks the fold, not the dynamics, which is one reason function lags behind shape.
The takeaway to carry forward: generative design has made proposing plausible novel molecules cheap and fast, and the binding problem is well on its way to solved for accessible targets. What remains hard is everything the wet lab measures and the file cannot — expression, developability, and above all genuine function. The bench, not the benchmark, still writes the verdict, which is the theme of the lab-in-the-loop workflow in Chapter 17.
References
- Corso, G., Stärk, H., Jing, B., Barzilay, R., & Jaakkola, T. (2022). DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking. ICLR 2023. arXiv:2210.01776.
- Dauparas, J., Anishchenko, I., Bennett, N., Bai, H., Ragotte, R. J., Milles, L. F., Wicky, B. I. M., Courbet, A., et al. (2022). Robust deep learning-based protein sequence design using ProteinMPNN. Science.
- Guan, J., Qian, W. W., Peng, X., Su, Y., Peng, J., & Ma, J. (2023). 3D Equivariant Diffusion for Target-Aware Molecule Generation and Affinity Prediction. ICLR. arXiv:2303.03543.
- Ingraham, J. B., Baranov, M., Costello, Z., Barber, K. W., Wang, W., Ismail, A., Grigoryan, G., et al. (2023). Illuminating protein space with a programmable generative model. Nature.
- Lauko, A., Pellock, S. J., Sumida, K. H., Anishchenko, I., Juergens, D., Ahern, W., Jeung, J., Shida, A. F., et al. (2025). Computational design of serine hydrolases. Science.
- Pacesa, M., Nickel, L., Schellhaas, C., Schmidt, J., Pyatova, E., Kissling, L., Barendse, P., Choudhury, J., et al. (2025). One-shot design of functional protein binders with BindCraft. Nature.
- Watson, J. L., Juergens, D., Bennett, N. R., Trippe, B. L., Yim, J., Eisenach, H. E., Ahern, W., Borst, A. J., et al. (2023). De novo design of protein structure and function with RFdiffusion. Nature.
- Zambaldi, V., La, D., Chu, A. E., Patani, H., Danson, A. E., Kwan, T. O. C., Frerix, T., Schneider, R. G., et al. (2024). De novo design of high-affinity protein binders with AlphaProteo. arXiv preprint. arXiv:2409.08022.
Check yourself
Questions a sharp collaborator might ask on this chapter. Pick an answer to see whether it holds up.
-
A colleague calls ProteinMPNN 'basically ESM for design.' What is the sharpest correction?
ProteinMPNN is an inverse-folding model: its input is backbone geometry and its output is a compatible sequence, learned from PDB structure-sequence pairs via message passing. ESM is a protein language model trained on sequences alone. Neither predicts structure from sequence in the AlphaFold sense, and ProteinMPNN uses no diffusion and far fewer training examples than ESM's sequence corpora. -
A design paper reports '90% of designs bind the target.' What does that number most accurately describe?
Headline success rates are measured late in the funnel, after aggressive in-silico filtering, and in an assay calibrated to that pipeline. The end-to-end yield from raw generation is far lower, because most designs never survive the self-consistency and developability filters. This is why the design loop, not any single model, is the real unit of progress. -
Why is designing an enzyme that catalyzes a reaction much harder than designing a binder to a target?
Binding is largely a matter of shape and chemical complementarity, exactly what diffusion plus self-consistency handle well. Catalysis demands sub-angstrom positioning of multiple catalytic residues, transition-state stabilization, and substrate turnover in a moving protein. Designed enzymes exist and improve, but still fall orders of magnitude below evolved turnover. Size is not the barrier, and self-consistency scoring applies to both. -
In the generate-filter-validate loop, what does the self-consistency filter actually check?
Self-consistency closes the loop between design and prediction: you designed a sequence for a backbone, so you refold that sequence with an independent predictor (AlphaFold, ESMFold) and keep it only if it returns to the intended shape. It scores foldability, not developability, synthesizability, or measured affinity, which is why passing designs are still hypotheses for the bench. -
A design passes self-consistency with excellent pLDDT but fails in your collaborator's lab. Which explanation is consistent with what the filter can and cannot see?
Self-consistency and pLDDT judge whether a sequence folds to the intended static shape; they say nothing about expression yield, aggregation, storage stability, or immunogenicity, and a single rigid snapshot cannot capture function that depends on motion. Developability is measured at the bench, not scored by the predictor, so a fold-perfect design can still clump in the tube.