
Custom enzyme. From the Nobel Prize to atomic precision
Stockholm, December 10, 2024
David Baker receives the Nobel Prize in Chemistry. On stage at Stockholm's Konserthuset, in a white tie and black tailcoat, he talks about protein design. About algorithms that learn the language of evolution - the same language that nature has been writing with for four billion years. The Swedish royal family is listening. Cameras broadcast all over the world.
Baker has been working on this for thirty years. He came to biology from physics - an unusual path. As a postdoc at Berkeley in the 1990s, he programmed the first versions of Rosetta, which was intended to solve the problem of protein folding: how to predict three-dimensional structure from amino acid sequences.
The problem was then considered computationally unsolvable. Baker solved it.
Three months after the Nobel ceremony, his team publishes a paper that proves that the Nobel Prize was just a stopover. Woody Ahern, Jason Yim and Doug Tischer - three young scientists from the Institute for Protein Design in Seattle - look at the results of the latest experiment at 11:47 p.m. RFdiffusion2, a new deep learning model, has just designed an enzyme from scratch. Not by fitting into existing scaffolding. Not by imitating nature. The algorithm got the coordinates of atoms in space — and invented a protein to hold them.
Baker sits next to him. He looks at the screen. He knows that the circle he opened - and almost closed prematurely - in 2008 has now been completed.
The first enzyme. And the first failure
In 2008, Baker's team published wNaturea work that made history in structural biology: the first enzyme designed entirely by a computer. Kemp elimination reaction catalyst - not used by any known organism. Nature never evolved this protein. They were designed by a human.
There was only one problem. The enzyme worked - but terribly slowly. Catalytic value five orders of magnitude below natural counterparts. `kcat/Km` of 0.1–1 M⁻¹s⁻¹ - for comparison, the average enzyme in your stomach operates at a rate of ~10⁵ M⁻¹s⁻¹.
It's as if an architect designed a house that stands - but is impossible to live in. The foundations hold, the walls don't crack, but the doors don't open.
Baker spent the next decade trying to fix the house. Directed evolution in the laboratory—a method that mimics Darwinian selection but in a test tube—resulted in a 200-fold increase in activity. Still not enough. Each subsequent experiment: new mutations, new tests, new disappointments. Five years. Ten.
"I told my team to stop wasting time on this enzyme," Baker later admitted. He said the words out loud. At a group meeting, in front of everyone. He's run out of ideas.
Other researchers have had similar experiences. David Hilvert from ETH Zurich designed an alternative catalytic pathway and got... another slow enzyme. The field of enzyme design - still promising in 2015, funded by DARPA and the Gates Foundation - began to smell like failure.
The problem was fundamental, not accidental. Previous methods, such as RosettaMatch, searched millions of existing protein structures for a scaffold for the active site. They found something - and stuffed the enzyme in there. But the geometry was never perfect. There was always something sticking out. There was always one hydrogen bond that was 0.3 ångström too far away. In enzymatic catalysis, 0.3 Å is a gap.
Baker understood this better than anyone. And it hurt him.
An architect who got a new tool
The breakthrough came from an unexpected direction: from diffusion models, the same architecture that powers image generators. AlphaFold2 solved the structure prediction problem in 2020 — but for Baker, something else was more important. If a neural network can predict how a given chain of amino acids will fold, can it also work the other way around? Generate a new string for a given structure?
In 2022, Baker's lab showed that it did. RFdiffusion — named after its predecessor RoseTTAFold — acted like an architect who no longer had to choose from a catalog of ready-made homes. He takes a piece of paper and draws from scratch. The diffusion model, trained on hundreds of thousands of known structures, gradually denoised random atomic coordinates until a stable protein chain emerged.
Except he was still drawing at the skeleton level - the main chain. He couldn't place specific atoms in specific places.
Baker needed something more. An enzyme is not just a shape. It is the precise arrangement of several—exactly several—functional groups around a reaction transition state. Oxygen atom here. Amine group at 109 degrees. A metal ion at a distance of 2.1 Å from the substrate. If any of these elements shifts by a fraction of ångström, the enzyme stops working. Just like in 2008.
RFdiffusion2, described in January 2026 inNature Methods, removes this restriction.
Instead of designing a framework and then adding side chains to it - a standard procedure in all previous methods - the new model operates immediately at the atomic level. He was given a description of the active center as a set of XYZ coordinates. No hints as to which amino acids fit there. No list of allowed conformations.
And from this description it generates the entire protein.
Magnitude of difference: Previous methods were able to design a scaffold for 16 of the 41 active sites in the test set. RFdiffusion2 - for all 41. Ahern and Tischer repeated the test three times. 41/41 every time.
Baker reportedly smiled for the first time in weeks as he read the manuscript before submitting it to the editor.
Three enzymes. Less than a hundred attempts each
Baker's team went further: they designed enzymes for three completely different catalytic mechanisms. Ester hydrolase. Cysteine protease - The same family as papain, a digestive enzyme from papaya fruit. And an enzyme that uses a metal ion as a cofactor.
In each case, they tested fewer than 96 sequences before finding active variants. 96 is the number of wells on a standard enzyme assay plate. This isn't a coincidence - it's an intentional demonstration of scalability.
For comparison, classic enzyme designs from the RosettaMatch era required testing of thousands of candidates, most of which showed no activity. It's like an architect building ten houses, nine of which are falling down and the tenth standing - but the doors still won't open.
Ahern, the paper's first author, tested ester hydrolase, an enzyme that breaks down ester bonds. This is a key reaction for plastic recycling, because PET and polyurethane are polyesters. First 48 sequences: no activity. Another 24: trace. Last 24: `kcat/Km` at 10² M⁻¹s⁻¹.
Still three orders of magnitude away from natural enzymes. But Baker looks at these numbers differently than he did in 2008. Then he saw failure. Now he sees the starting point. And it knows exactly where it's going: directed evolution, optimization, scaling.
A market waiting for an architect
The global market for industrial enzymes – from laundry detergents to pharmaceutical catalysts – was worth $14.1 billion in 2024. It is expected to reach 20 billion by 2030. It is driven by three trends: green chemistry (enzymes replacing toxic solvents), pharmaceutical biotechnology and - increasingly - enzymatic recycling of plastic.
Except that all commercial enzymes come from nature. None were designed from scratch by a computer.
Novozymes, a global leader with 48 percent market share, employs teams that scour soil, compost and hot springs for organisms with enzymes of interest. They then optimize the proteins they find through directed evolution. It works. But it's like looking for a specific sentence in the library instead of writing it yourself.
RFdiffusion2 changes this logic. Instead of asking "what has nature already done", he asks "what do we need?" If you need an enzyme that breaks down PET at 70 degrees - you describe the geometry of the active site and you get a protein. It does not matter whether such an enzyme existed before.
Baker does not compete with Novozymes. He changes the rules of the game.
The competition is keeping up: DeepMind Spin-off Isomorphic Labs, Meta with its ESM model, ByteDance with Protenix all working on protein design. But Baker has the first mover advantage. Its Institute for Protein Design combines computational modeling with a wet lab that tests designs outright. The "design → synthesis → test → model correction" loop takes weeks, not years.
The architect has a studio in the same building as the construction site.
Poland: draftsmen without a construction site
Poland's problem with enzyme design is not a lack of scientists. There are them - scattered over several centers. The problem is that they are in separate buildings.
On the one hand, we have computational chemistry groups: at the Warsaw University of Technology, at the Jagiellonian University, at the Center for New Technologies of the University of Warsaw. They can model active sites with quantum precision - DFT, coupled clusters, QM/MM dynamics. They publish inJACSandAngewandte Chemie
On the other hand, wet biology laboratories: in Łódź (Molecure), in Gdańsk (Polpharma Biologics), in Wrocław (Captor Therapeutics, Ryvu). They can express, purify and test proteins on the milligram scale.
There is no pipeline between these two worlds. There is no platform that takes a design from a computational lab and throws it into a fermenter. One thing is missing: an integrator.
RFdiffusion2 — whose code is available under the CC-BY-NC-ND license — could be that glue. The model doesn't require a million-dollar GPU cluster. It runs on a single NVIDIA A100 card, which every Polish computing center has. Cyfronet AGH has five such cards as part of the PL-Grid infrastructure.
The FENG program - European Funds for a Modern Economy - reserves PLN 300 million for "biotechnology and bioinformatics" in the 2024-2027 perspective. The grant path is simple: a consortium of universities + biotechnology companies + access to Cyfronet's computing power. Purpose: not published inNature Methods, only a working enzyme for an industrial process. For example: esterase that breaks down polyurethane from furniture waste.
Time is short. Isomorphic Labs spends $200 million a year. Meta is hiring hundreds of engineers for the ESM project. If Poland does not build a bridge between the computational laboratory and the wet laboratory within 18 months, the gap will become impossible to bridge. Not because there is a lack of talent. Just because there's not enough time.
Baker proved that it is possible to make an enzyme from scratch. Now the question is who will be the first to make a product out of it. And will Poland be at the start at all?
Sources
Ahern W., Yim J., Tischer D. et al.,Atom-level enzyme active site scaffolding using RFdiffusion2, Nature Methods 23, 96–105 (2026), DOI: 10.1038/s41592-025-02975-x.
Röthlisberger D. et al.,Kemp elimination catalysts by computational enzyme design, Nature 453, 190–195 (2008), DOI: 10.1038/nature06879.
Watson J.L. et al.,De novo design of protein structure and function with RFdiffusion, Nature 620, 1089–1100 (2023), DOI: 10.1038/s41586-023-06415-8.
MarketsandMarkets,Industrial Enzymes Market - Global Forecast to 2030, 2024.
Comments· 0
No comments yet. Be the first.