Biotechnology · Artificial Intelligence · Deep Tech

42 times better in 24 hours. Anatomy of an enzyme design machine

Readiness level4 / 9Validated in the lab
AuthorTETRL09 editorial team
Published
Reading time10 min

Take a plastic plate with 96 holes. In each of them there is a transparent liquid. No cells. No bacteria. Only water, amino acids, DNA and transcription machinery isolated fromE. coli. Within 24 hours, this fluid produces an enzyme that has never existed in nature before. And it does it 42 times more efficiently than its wild ancestor.

This is not a scenario from the laboratory of the future. This is the platform that Michael Jewett's team from Northwestern University and Stanford described in January 2025 inNature Communications. Its anatomy consists of four layers. Each represents a separate breakthrough - but together they create a machine that cuts enzyme engineering from months to hours.

Fig. 1. Architecture of the DBTL platform: from DNA through cell-free expression to the ML model. Source: Landwehr et al., Nature Communications 16, 865 (2025).

Layer 1: DNA - a design that fits on a piece of paper

The first layer of the platform is the most abstract and least expensive. It starts with selecting the amino acids to be replaced in the enzyme.

The target was McbA, an amide synthetase from a marine bacteriumMarinactinospora thermotolerans, discovered in 2012 in a sediment sample from the South China Sea. This enzyme naturally forms amide bonds, a chemical motif present in about 25 percent of all drugs on the market, from the painkiller acetaminophen to the anticancer imatinib. It works like a molecular welder: it grabs a carboxylic acid and an amine and combines them at the cost of one ATP molecule. No toxic solvents, no high temperatures. Just water, buffer and adenosine triphosphate.

But wild McbA is a poor welder for special tasks. In a screen of 1,100 reactions with substrates ranging from simple aliphatic amines to complex heterocycles, most products appeared in trace amounts, detectable only by mass spectrometry. Only 11 of the 21 pharmaceutical molecules tested could be seen in the UV chromatogram. The best variants achieved a conversion of around 12 percent.

Jewett's team didn't try to guess which amino acids to replace. He used a combination of three computational tools: ROSETTA for structural modeling, EVmutation to analyze evolutionary trends from hundreds of related protein sequences, and PROSS to predict thermodynamic stability. Of the 500 amino acids that make up McbA, 64 positions were selected for mutagenesis. It was done in one day.

Each mutation—each of the 64—was designed as a single PCR primer. A few nanograms of DNA. Cost: several dozen cents per starter. At this point, the platform does not need any advanced hardware yet. All he needs is the McbA gene sequence and access to public evolutionary databases.

Layer 2: Expression - the soup that replaces the cell

The second layer is the most radical element of the entire architecture. And one that distinguishes it from the classical protein engineering practiced for the last 30 years.

The standard path goes like this: you introduce the mutated gene intoE. coliby transformation, you wait for colonies (overnight), multiply the culture (another night), induce expression (a few hours), lyse the cells, purify the protein on a chromatographic column. Then you test. If it doesn't work, you go back to point one. One iteration: 3-5 days.

Jewett's team missed living cells completely.

The system works like this: a linear DNA fragment with a mutation (PCR product) is placed in the test tube and assembled into a full gene using the Gibson Assembly method - without the participation of cells. Cell extract containing ribosomes, RNA polymerases, tRNA, amino acids and energy buffers is then added. It's like taking an entire protein factory out of a cell and putting it in a test tube. No cell wall. No bacterial metabolism competing for substrates. Just a biochemical reaction - transcription and translation occurring in the same vessel.

The whole process from PCR to active enzyme takes less than 24 hours. One quarter of the time of the classic track. And - most importantly - it does not require sterile breeding conditions. The test tube is a closed biochemical system, insensitive to contamination.

A number that should stop you for a moment: 1,217. That's how many McbA variants the team generated, each defined by a DNA sequence, each produced in the same cell-free "soup." None required a living organism. None required chromatographic purification - the enzymes were tested directly in the reaction mixture.

Layer 3: Test - 10,953 reactions and zero cells

The third layer is the moment of truth. The enzyme variants go into the reaction mixture with the acid and amine. Enzyme concentration: approximately 1 micromolar - so low that in a classical cell lysate system it would be indistinguishable from host proteins. Substrate concentration: 25 millimoles. After incubation - HPLC analysis.

Among the 1,100 reactions with wild-type McbA, a clear preference pattern was discovered. Aryl, benzoic and cinnamic acids - the enzyme took them efficiently. Aliphatic and fatty acids - virtually no activity. Primary aliphatic amines - very good. Aromatic amines - barely.

This step also revealed unexpected stereoselectivity. For sulpiride - an antipsychotic drug from the benzamide group - McbA strongly preferred the S enantiomer over the R. The difference was visible to the naked eye in the chromatogram: one peak was high, the other was practically absent. The enzyme selected the correct product geometry itself, without any additional chiral catalysts. This is the same selectivity that the pharmaceutical industry pays millions for, using expensive ligands and low-temperature syntheses.

Of the 1,217 variants, approximately 80 individual mutants - from the first round - were included in the training set. Each with an exact conversion percentage. Also those that made the activity worse. Also those that didn't change anything. Every failure is a data point. Each neutral mutation provides information about which positions in the protein are evolutionarily frozen and which are plastic.

This is a fundamental difference from classic guided evolution. There you select only the "winners" - the best variants from a given round. You throw away the rest. Here, each variant, including the one worse than the wild type, is a valuable point on the fitness map. Because the ML model learns not only from successes, but also from failures.

Layer 4: Model - 1970's regression that beats neural networks

The fourth layer is anticlimactic in the best sense of the word. No transformers. No GPUs. No billions of parameters. Ridge-only regression - a statistical technique introduced in 1970 by Hoerl and Kennard - with an additional feature: zero-shot prediction from the EVmutation evolutionary model.

Each enzyme variant was encoded as a vector: one-hot encoding of the amino acid position (20 bits per position - one for each of the 20 amino acids) plus a scalar with EVmutation - a number estimating the evolutionary compatibility of a given mutation with the McbA protein family.

The model was trained on 77 mutants. It didn't need the cloud - it all fit on the CPU of a standard laptop. The team tested different variants: regression alone, regression with features from ESM-1b (protein language model from Meta), regression with EVmutation. The best results - as measured by NDCG, a metric that assesses a model's ability to correctly rank variants from best to worst - were achieved by regression augmented by EVmutation. Simpler models generalized better.

Importantly: the model operated on a small training set on purpose. The team tested whether 77 mutants was the minimum or whether even less data would be enough - because this directly translates into the cost of the entire pipeline. It turned out that with 30-40 mutants, the quality of prediction decreases, but it still allows the identification of the best variants. At 77 - works reliably.

The results on moclobenide - an antidepressant from the MAO inhibitor group - were the most spectacular. The best predicted variant achieved a 96 percent conversion. For the wild type it was 12 percent. The catalytic efficiency - the ratio of the catalytic constant to the Michaelis constant, i.e. how many product molecules per enzyme molecule per second - increased from 18.2 to 764 M⁻¹s⁻¹. That's a 42-fold improvement.

Not every mutation was necessary. In the moclobemide variant, the first mutation did practically all the work. The second one (A323F) improved thermal stability by 5.81 degrees Celsius - which is important for industrial processes where the enzyme works for hours at elevated temperatures - but did not add catalytic activity.

For itopride, a prokinetic drug, the improvement was almost 30-fold. The median activity improvement for all nine pharmaceutical targets exceeded 16 percent, with each compound getting its own dedicated McbA variant.

Figure 2. Catalytic activity of McbA variants for 9 target pharmaceuticals. Source: Landwehr et al., Nature Communications 16, 865 (2025).

A race that is already underway

In January 2025, when Jewett released these results, the market was not waiting. Codexis — the California company that supplied Merck with transaminase for the synthesis of sitagliptin (Januvia) in 2010 — had revenues of more than $80 million a year from enzyme engineering alone. Novonesis, the industrial enzyme giant resulting from the merger of Novozymes and Chr. Hansen, served thousands of customers in 30 countries. Ginkgo Bioworks—a Boston startup that has raised more than $2 billion for automated biology engineering foundries—designed custom enzymes for pharmaceutical, cosmetics and food companies.

Caltech's Frances Arnold won the Nobel Prize in 2018 for directed evolution, a technique that defined enzyme engineering for 25 years. David Baker of the University of Washington - Nobel 2024 for protein designde novo. Jewett showed a third way: neither evolution by selection nor design from scratch. Just a platform that integrates data from a small number of experiments and uses simple statistics to predict what has not yet been measured.

The difference between Jewett's platform and Arnold's or Baker's approach is like testing every possible bridge by building it and seeing if it collapses, and using an engineering model that tells you in advance which designs have a chance. Guided evolution is the former. Projectsde novothe latter - but it requires computationally expensive physical simulations and works for relatively simple structures. Jewett's platform is a third way: statistics on a modest data set that finds the optimum without understanding the physics of the problem.

Fig. 3. Three paths of enzyme engineering - comparison: directed evolution vs. de novo design vs. ML-cell-free platform. Source: own study based on Landwehr et al. (2025).

Poland: know every layer, have no platform

Polish biotechnology has pieces of the puzzle. The Institute of Organic Chemistry of the Polish Academy of Sciences in Warsaw is developing enzyme catalysis. Center for New Technologies, University of Warsaw - computational protein modeling. Warsaw University of Technology - scaling bioprocesses. IBIB PAN - metabolic engineering. Wrocław University of Science and Technology - enzyme immobilization. Each of these places has competences corresponding to one layer of Jewett's architecture.

There is a lack of integration. No one has combined cell-free expression, high-throughput screening, and ML modeling into a single DBTL platform.

The Polish pharmaceutical industry knows this. Selvita - PLN 324 million in revenue, 900 employees, listed on the WSE - invests in structural biology. Molecure and Ryvu Therapeutics are developing new drugs. Polpharma produces hundreds of tons of API annually in Starogard Gdański. None of these companies have access to an integrated biocatalyst design platform on-site.

Funding paths exist. NCN MAESTRO: PLN 2-3 million for basic research on cell-free protein synthesis. PARP SMART Path: PLN 10-15 million for the implementation of the DBTL platform in a university-company consortium. FNP TEAM: another few million for the international team. The amounts are adequate.

The problem lies elsewhere. Jewett's team of six authors combines expertise in computational biology, biochemical engineering, synthetic biology and analytical chemistry. In Poland, these disciplines are located in separate institutes, in separate faculties, financed from separate grants. No one has submitted an application that would combine them in one laboratory.

And time flies. Codexis and Ginkgo Bioworks will not stop so that Poland can catch up. In three years, the Polish pharmaceutical industry will need biocatalysts for new syntheses - and will buy them from someone who is already building a four-layer platform.

Źródła

  1. Landwehr G.M., Bogart J.W., Magalhaes C., Hammarlund E.G., Karim A.S., Jewett M.C.,Accelerated enzyme engineering by machine-learning guided cell-free expression, Nature Communications, 16, 865 (2025).DOI: 10.1038/s41467-024-55399-0

Comments· 0

No comments yet. Be the first.

Add a comment