
A protein designer who doesn't wait for evolution
At 2 a.m. Vincent Tran sat in front of a monitor at the Arc Institute and stared at a result that had no right to exist.
The column of numbers on the screen showed the activity of the enzyme — esterase, a common protein that breaks down fats — at 2,500% of baseline. Tran rubbed his eyes. He ran the algorithm again. 2512%.
It wasn't just about value. The idea was that the MULTI-evolve algorithm found this variant in two weeks. Traditional directed evolution - the method for which Frances Arnold won the Nobel Prize in 2018 - would require three months and fifteen rounds of mutations.
Tran was 23 years old. He was fresh out of college at UC Berkeley. And a moment ago, his algorithm found a combination of six mutations that no molecular biologist could have invented - because four of them, individually, wereharmful. Only together did they provide a leap in function.
This is epistasis. 1+1=10. Or: change amino acid A and nothing happens. Change the B amino acid and the protein breaks down. Change both at once - and suddenly the enzyme works 25 times faster.
Why hasn't anyone found this before
To understand what Tran did, you first have to understand why finding good proteins is so incredibly difficult.
Protein is a chain of amino acids. There are 20 types of them. For a protein with a length of 150 amino acids - and this is not a big deal, hemoglobin has 574 - there are 20¹⁵⁰ possible sequences. It's one followed by two hundred zeros. More than atoms in the observable universe.
It's impossible to search through it all. That's why biotechnologists have been using it since the 1990sguided evolution: you take a protein-coding gene, introduce random mutations, test which variant works better, repeat. Step by step. Mutation after mutation.
Works. Frances Arnold received the Nobel Prize for this. But it has a fundamental flaw: each round tests only ONE mutation at a time. And nature doesn't work alone. The greatest evolutionary leaps - from monkey to man, from swimming to flying - are always combinations of changes, not single mutations.
Tran understood this. And he asked the question: What if a language model - the same type of algorithm that powers ChatGPT - "read" millions of protein sequences and predicted which combinations of mutations made sense?
ESM-2, a model developed by Meta AI in 2023, was trained on 250 million natural protein sequences. He doesn't understand biology - he understands statistics. It knows that the amino acid "leucine" at position 47 is more likely to be followed by "valine" than "tryptophan" - just as the language model knows that "day" is more likely to be followed by "good" than "potato."
“We trained ESM-2 on all known proteins,” Tran said when presenting the results at an internal Arc Institute seminar. "Not for him to understand biology. For him to sense which mutations nature has already tested and rejected."
The first attempt didn't work.
The algorithm proposed 10,000 mutation combinations for the GFP fluorescent protein. Tran and the team—Patrick Hsu, Silvana Konermann, Brian Hie—ordered DNA synthesis for all variants. They came after three weeks. Cost: $12,000. Tests showed that 94% of the variants were non-functional. The proteins did not glow. Or they were curling up.
“We went back to code,” Hsu later recalled. The problem was that ESM-2 predicted mutations correctly but did not account for epistasis. He proposed combinations that made sense individually, but together were mutually exclusive.
The team added a second layer: an epistatic model that evaluated not single mutations but pairs and triplets. “That was the hardest part,” Tran said. "Modeling interactions between mutations is an NP-hard problem. For 10 mutations, that's 3.6 million possible combinations."
The solution was DNA synthesis on a chip. Instead of ordering each variant separately—expensively and slowly—the team used oligonucleotide chips that produce tens of thousands of variants at a time. The cost dropped from $12,000 to $600. Time - from three weeks to 48 hours.
Then cell-free expression. Instead of cloning each variant into bacteria and waiting for them to grow (another week), the team usedcell-free expression— a system that produces proteins in a test tube using only DNA and cellular machinery isolated from E. coli. Test time: 4 hours.
And then, on the second try, Tran saw that 2512% on the screen.
A market waiting for a shortcut
Global industrial enzyme market: USD 12.3 billion in 2024. 20 billion in 2030 - according to Grand View Research forecasts. Average annual growth 8.7%.
These are not niche laboratory reagents. Enzymes produce 40% of the active ingredients in drugs - from antibiotics to anticancer drugs. Washing powder that washes at 30 degrees? Contains proteases and lipases designed by protein engineers. Biofuels? Cellulases that break down biomass into simple sugars.
And behind each of these enzymes there is a process of directed evolution. Average 10-15 rounds. On average 3-6 months. Average cost: $50,000 to $200,000 per enzyme.
Tran cut it down to 2-3 rounds, 2-4 weeks and about $5,000.
Companies see this.Codexis(NASDAQ: CDXS, $380M market cap) sells enzymes to Merck and GSK - their CodeEvolver platform is guided evolution. Sequential.Novonesis— the giant resulting from the merger of Novozymes and Chr. Hansen, worth $28 billion and controlling 48% of the global market — also operates sequentially.
And next to them, MULTI-evolve startup children are growing up.Cradlefrom Amsterdam ($24 million from Index Ventures) shortened antibody development from 6 months to 6 weeks.EvolutionaryScaleThe New York-based $142 million seed, founded by former Meta AI researchers, trains protein language models from scratch.
"The timetable is clear," Tran said at a conference in Boston in March 2026. "In two years, anyone who designs enzymes sequentially will be competing with algorithms that do the same thing on a weekend."
The Polish enzyme that does not exist
There are three companies in Poland that design proteins. None use language models.
Ryvu Therapeuticsfrom Krakow - PLN 450 million capitalization on the Warsaw Stock Exchange, kinase inhibitors in phase II clinical trials.Molecurefrom Warsaw - PLN 200 million, deubiquitinase inhibitors in phase II for sarcoidosis.Pure Biologicsfrom Wrocław - PLN 80 million, therapeutic antibodies.
All three use methodsrational design— engineers predict which mutation will improve a protein's function based on its 3D structure. It works, but it's slow. And he finds no epistasis.
"If I had access to MULTI-evolve," a PhD student from the Warsaw University of Technology working on an enzyme that breaks down PET plastic told us anonymously, "I would test 10,000 variants in a week. Now I make 50 variants in a month. Manually."
There are facilities.PL-Grid, the Polish Bioinformatics Node, has 2.5 PFLOPS of computing power - enough to train protein language models comparable to ESM-2. Prof. team Joanna Bereta from the Krakow University of Technology has been conducting molecular dynamics simulations for a decade. Michał Jamróz from the University of Warsaw uses machine learning to predict protein structures.
There is money too. FENG - European Funds for a Modern Economy - allocates EUR 1.9 billion to biotechnology until 2027. NCBR announced the "Biotechnological molecular design platforms" competition in 2025 with a budget of PLN 120 million. Project limit: PLN 1.8 million.
What's missing? Manager. Someone who understands both molecular biology and machine learning - and can put together a team of bioinformaticians from the University of Warsaw, DNA chemists from the Adam Mickiewicz University in Poznań and biotechnologists from the Gdańsk University of Technology into one pipeline. There are maybe thirty such people in Poland.
The NCBR competition closes in December 2026. If a consortium is not established and submits an application by then, PLN 120 million will return to the budget. And Polish PhD students will continue to make 50 variants a month - manually.
Meanwhile, Tran and his team have published the MULTI-evolve code as open source on GitHub. Anyone can download it. Anyone can use it. The question is not whether a Polish enzyme will be created. The question is whether it will be built here.
Sources
- Tran V.Q., Nemeth M., Bartie L.J. et al. "Rapid directed evolution guided by protein language models and epistatic interactions."Science, 2026. DOI: 10.1126/science.aea1820
- Hie B., Zhong E.D., Berger B., Bryson B. "Learning the language of viral evolution and escape."Science, 2021. DOI: 10.1126/science.abd7331
- Lin Z., Akin H., Rao R. et al. "Evolutionary-scale prediction of atomic-level protein structure with a language model."Science, 2023. DOI: 10.1126/science.ade2574
- Yang K.K., Wu Z., Arnold F.H. "Machine-learning-guided directed evolution for protein engineering."Nature Methods, 2019. DOI: 10.1038/s41592-019-0496-6
- Markey C., Crooks J., Baran J. et al. "Accelerated enzyme engineering by machine-learning guided cell-free expression."Nature Communications, 2025. DOI: 10.1038/s41467-025-55828-9
- Grand View Research. "Enzymes Market Size & Share Report, 2024-2030." Report ID: 978-1-68038-135-0, 2024.
- NCBR. "Competition: Biotechnological molecular design platforms." LIDER XIV program, budget PLN 120 million, deadline: December 2026.
- FENG. "European Funds for a Modern Economy 2021-2027 - Priority 1." Ministry of Funds and Regional Policy.
Comments· 0
No comments yet. Be the first.