MIT Language Model Boosts Protein Production in Yeast, Cutting Drug Development Costs

MIT engineers trained a language model on yeast codon usage to design genes that boost protein production, outperforming commercial tools for five of six proteins. The approach could lower drug development costs by improving yields of protein-based medicines.

Massachusetts Institute of Technology chemical engineers have developed a language model that learns a yeast's codon usage patterns to design genes that boost protein production, potentially lowering the cost of drug development. The model outperformed four commercial codon optimization tools for five of six tested proteins, including the monoclonal antibody trastuzumab. The findings were published this week in the Proceedings of the National Academy of Sciences.

The researchers focused on Komagataella phaffii, a yeast widely used for making recombinant proteins. They trained an encoder-decoder style large language model on amino acid sequences and matching DNA coding sequences from roughly 5,000 proteins naturally produced by the yeast, using a publicly available dataset from the National Center for Biotechnology Information. The model learns the syntax of how codons are used, accounting for neighboring codons and longer-range relationships across a gene.

After training, the model designed codon-optimized sequences for six proteins: human growth hormone (hGH), human growth colony-stimulating factor (hGCSF), a VHH nanobody called 3B2, an engineered variant of a SARS-CoV-2 receptor binding domain (RBD), human serum albumin (HSA), and the IgG1 monoclonal antibody trastuzumab. These designs were compared with sequences produced by four commercial tools: Azenta, IDT, GenScript, and Thermo Fisher. Each version was inserted into K. phaffii cells, and target protein production was measured.

Across the six proteins, the MIT model produced the best titer for five and ranked second for the sixth. Codon optimization improved production for some molecules more than others: hGH and hGCSF saw about a 25% improvement, while HSA showed about a threefold improvement compared to the native coding sequence. Using native sequences, HSA reached a titer of 45 mg/L, while bovine serum albumin (BSA) and mouse serum albumin (MSA) reached 60 mg/L and 100 mg/L, respectively. Codon optimization increased BSA and MSA titers by an additional 25%, to 75 mg/L and 135 mg/L.

The commercial tools varied in consistency. GenScript produced the best titer for trastuzumab but often landed between 80% and 100% of the top titer across the set. Thermo produced best titers for three of the six proteins but performed poorly on two others, including lower results for HSA and trastuzumab. IDT ranked lowest on the study's two performance metrics and did not produce the best titer for any tested protein.

Related Entities

Related Articles

References

  1. How their mother's death spurred two brothers to speed up the drug discovery process through AI · financialpost.com
  2. A cheaper, more sustainable way to manufacture breakthrough HIV drug Lenacapavir · phys.org
  3. AI breakthrough could dramatically lower the cost of drug development · thebrighterside.news