Generative AI for Protein Design: How RFdiffusion, ProteinMPNN, and Diffusion Models Are Engineering New Proteins
Generative AI protein design now uses RFdiffusion, ProteinMPNN, and diffusion models to build therapeutics from scratch.
Generative AI protein design has moved past predicting how existing proteins fold and into building entirely new ones from a functional specification. Tools such as RFdiffusion and ProteinMPNN now generate protein backbones and sequences to order, producing candidate binders and enzymes that reach laboratory testing within weeks rather than the years that physics-based design once required.
Key takeaways
- Generative AI protein design creates new protein sequences and structures to a functional specification, rather than predicting the shape of an existing protein.
- RFdiffusion generates protein backbones by adapting a structure-prediction network into a diffusion model, and ProteinMPNN then assigns amino acid sequences to those backbones.
- Diffusion-based tools including Chroma extend generative design to full protein complexes, while newer sequence-based frameworks explore similar territory without requiring a solved structure first.
- Applications now include de novo binders, enzymes, and therapeutic candidates, with several designs validated by X-ray crystallography, cryo-electron microscopy, or direct biochemical assay.
- Validating a design requires both computational self-consistency checks and laboratory testing, since a plausible in silico structure does not guarantee that a protein folds or functions as intended.
How generative AI models design new proteins
Generative protein design borrows its core mechanism from image-generating diffusion models: a network learns to reverse a gradual noising process, starting from randomly arranged atoms and iteratively refining them into a realistic structure. Applied to proteins, this means beginning with noise and denoising it, step by step, into a backbone with plausible bond geometry and a foldable shape.
This approach runs in the opposite direction from structure prediction. A structure-prediction model like AlphaFold takes an existing amino acid sequence and predicts the three-dimensional shape it adopts. A generative design model instead starts from a target function, shape, or binding site and produces a new backbone and sequence that could satisfy it. The two approaches are complementary, and most design pipelines link them directly: a generative model proposes a backbone, a sequence design network assigns amino acids, and a structure predictor checks whether the resulting sequence is likely to fold as intended.
That checking step matters because a generated backbone is only a hypothesis. Researchers commonly predict the structure of a newly designed sequence with a separate tool and compare it back against the original design, a practice known as self-consistency, before committing any resources to laboratory synthesis.
RFdiffusion and structure-based de novo protein design
RFdiffusion, developed by the Institute for Protein Design at the University of Washington, fine-tunes the RoseTTAFold structure-prediction network on a protein structure denoising task, turning a model built to predict structure into one that generates it. The result is a diffusion model that can produce protein backbones unconditionally or under specific constraints, including target binding sites, symmetric assemblies, and enzyme active sites.
The tool's developers reported designing protein binders against five separate targets, including influenza hemagglutinin, and used cryo-electron microscopy to confirm that one such binder's actual structure closely matched the computational design model. In a separate demonstration scaffolding the p53 helix that binds the MDM2 protein, the strongest designed binder bound roughly three orders of magnitude more tightly than the natural peptide it was built to outcompete. This kind of target-conditioned generation, in which the model receives a binding partner and a set of interface hotspots rather than working unconditionally, has made de novo binder design one of RFdiffusion's most widely adopted applications.
RFdiffusion also scaffolds enzyme active sites, arranging a protein backbone around a fixed geometry of catalytic side chains. A follow-up model, RFdiffusion2, extended this active site scaffolding to work directly from atomic-level active site geometries rather than requiring the residues to be pre-indexed along the backbone, solving all 41 cases in a diverse benchmark compared with 16 for the original method, and producing working retroaldolase and hydrolase enzymes after screening fewer than 96 candidate sequences per reaction.
ProteinMPNN and sequence design for generated backbones
ProteinMPNN is a graph-based neural network that predicts the amino acid sequence most likely to fold into a given protein backbone, one residue at a time, conditioned on its structural neighbors. It solves the step that RFdiffusion leaves open: a generated backbone has a shape but no sequence, and ProteinMPNN is what fills that gap. Developed by the same Institute for Protein Design group behind RFdiffusion, it has become the standard sequence design step in most modern generative pipelines.
Its developers reported a sequence recovery rate of 52.4% for ProteinMPNN, well above the 32.9% achieved by the physics-based Rosetta method it was benchmarked against, and found that it could rescue several previously failed designs originally generated with Rosetta or AlphaFold. Because RFdiffusion generates structure without assigning sequence, most modern design pipelines pair the two tools directly: RFdiffusion proposes a backbone, and ProteinMPNN designs several candidate sequences for that backbone before a structure predictor filters the results.
ProteinMPNN's sequence recovery advantage extends across single-chain monomers, symmetric homo-oligomers, and hetero-oligomeric interfaces, and later extensions such as LigandMPNN adapt the same architecture to design sequences around bound small molecules, metals, or nucleic acids rather than protein backbones alone.
Chroma, EvoDiff, and other protein engineering AI models
Chroma and EvoDiff extend generative protein design beyond the pipeline formed by RFdiffusion and ProteinMPNN, applying diffusion in different ways to reach complexes and disordered sequences that the original tools were not built to handle. Chroma, developed by Generate Biomedicines, is a diffusion model that samples full protein complexes directly, using a graph neural network architecture built for sub-quadratic computational scaling rather than fine-tuning an existing structure predictor. Its developers reported experimentally testing 310 Chroma-designed proteins and confirmed two crystal structures with a backbone deviation of roughly 1.0 angstrom from the computational design, evidence that the sampled structures are physically realizable rather than merely plausible on paper.
A separate line of work applies diffusion directly to protein sequences rather than structures. EvoDiff, a framework from Microsoft Research, trains a discrete diffusion model on evolutionary-scale sequence data, aiming to generate viable proteins, including those with disordered regions that structure-based tools struggle to represent, without ever requiring a solved backbone as a starting point. This sequence-first approach remains an active research direction rather than an established production tool, and its developers have described further scaling and experimental validation as necessary next steps.
Commercial and academic groups including Cradle Bio, EvolutionaryScale, and Profluent are developing related generative and language-model-based approaches to protein engineering, reflecting how quickly this field has moved from a handful of academic tools toward a broader ecosystem. Sequence-based generative methods sit alongside protein language models, which extract evolutionary signal from existing sequences to predict function rather than generate new ones, underscoring how large language models in biology now span prediction, function annotation, and outright generation within the same broad toolkit.
Table 1: A qualitative comparison of generative protein design tools by core approach and primary output.
| Tool | Core approach | Primary output | Typical use case |
| RFdiffusion | Structure-prediction network fine-tuned as a diffusion model | Protein backbones | Binders, symmetric assemblies, enzyme scaffolds |
| ProteinMPNN | Graph neural network for inverse folding | Amino acid sequences for a given backbone | Sequence design following backbone generation |
| Chroma | Diffusion model with sub-quadratic scaling | Full protein complexes, structure and sequence | Large assemblies, programmable design constraints |
| EvoDiff | Discrete diffusion over evolutionary sequence data | Protein sequences | Sequence-first design, disordered regions |

Figure 1: A labeled diagram of the generative protein design pipeline, from backbone generation through sequence design to structural validation. Credit: AI-generated image created using Google Gemini (2026).
Binders, AI enzyme design, and validating generated proteins
De novo binder design is the application generating the most immediate interest, since a protein that binds a specific target with high affinity has direct value in diagnostics, research tools, and therapeutics. In one striking example, researchers used RFdiffusion to design proteins that neutralize three-finger toxins from elapid snake venom, and the resulting binders protected mice from a lethal neurotoxin challenge, pointing toward cheaper, more stable alternatives to antibody-based antivenom.
Enzyme design remains a harder problem, since catalysis depends on precisely positioning multiple functional groups around a reaction's transition state rather than simply matching a static binding surface. Even so, atom-level scaffolding methods have produced functional retroaldolases and metal-dependent hydrolases directly from computed reaction geometries. Some of these designs reach catalytic efficiencies that exceed those of previously designed enzymes for the same reaction, though generated enzymes still generally fall short of the catalytic efficiency of enzymes refined by natural evolution.
Regardless of application, no generated protein is considered validated until it clears both computational and experimental checks. A practical framework for evaluating a candidate design typically follows a consistent sequence:
- Generate a backbone conditioned on the desired function, shape, or binding target using a tool such as RFdiffusion or Chroma.
- Design one or more candidate sequences for that backbone with ProteinMPNN or a comparable inverse-folding method.
- Predict the structure of each candidate sequence independently and compare it against the original design model, discarding candidates that diverge substantially.
- Express and purify the strongest computational candidates and confirm folding and stability with an assay such as circular dichroism or size-exclusion chromatography.
- Test the specific function the protein was designed for, whether binding affinity, catalytic activity, or thermal stability, before considering the design successful.
Even designs that pass every computational filter can fail at the bench, since a structure predictor's confidence reflects geometric plausibility rather than a guarantee of correct folding, and this gap is precisely why every generative protein design campaign still ends in the laboratory rather than at a computer screen.
Generative AI protein design is becoming an engineering discipline
Generative AI protein design has shifted the central question in protein engineering from what a given protein does to what a protein could be built to do for a specific job. RFdiffusion and ProteinMPNN, used together, form the backbone-then-sequence pipeline that most current design campaigns rely on, while diffusion-based alternatives such as Chroma and emerging sequence-first frameworks extend the same underlying logic to complexes and disordered regions.
The pace of change in this field means that today's leading tools are unlikely to be the last word, but the underlying pattern (generate a candidate, filter it computationally, and confirm it experimentally) is likely to persist even as the specific models improve. That pattern sits within the wider shift already reshaping AI and data science tools across life science research, in which prediction, annotation, and now generation increasingly share the same computational foundations.
This content includes text that has been created with the assistance of generative AI and has undergone editorial review before publishing. Technology Networks' AI policy can be found here.