
- Sponsor
- Department of Civil and Environmental Engineering
- Originating Calendar
- CEE Seminars and Conferences
Artificial intelligence for enzymatic plastic biodegradation: Data extraction, function
prediction, and novel enzyme discovery
Advisor: Professor Na Wei
Zoom Meeting (ID: 832 9787 6779 Password: 756228)
Abstract
This dissertation develops a closed-loop pipeline for discovering plastic-degrading enzymes to
address the long-standing challenge of plastic accumulation in the environment by integrating
machine learning, artificial intelligence (ML/AI), and wet-lab experiments. The pipeline consists
of a data extraction framework (termed PEZy-Extract) to automatically extract enzymatic plastic
degradation data from scientific literature, a function prediction framework (termed PED) to
predict whether an input enzyme-plastic pair is degradable or non-degradable, an enzyme mining
framework (termed PEZy-Miner) to rank potentially degradable enzyme-plastic pairs from a
large candidate pool, and, ultimately, an experimental characterization framework to
comprehensively analyze the top-ranked enzymes for degrading plastics of interest.
The pipeline forms a closed Design-Build-Test-Learn (DBTL) loop, in which each framework
provides the foundation for downstream stages while incorporating feedback from upstream
framework. In the first framework, PEZy-Extract, AI agents were developed to create a data set
integrating information from a wide range of experimentally verified enzymes and various
plastic substrates. The compiled data set can be used to uncover domain knowledge as well as
train downstream ML models. In the second framework, PED, protein language models were
used to learn the abundant contextual information in enzyme sequences, and feature extraction
was performed at both the amino acid level and the global sequence level. Binary classification
was subsequently performed, achieving an overall accuracy of 90.2% and outperforming
sequence-based protein classification models reported in the existing literature. In the third
framework, PEZy-Miner, confidence and uncertainty estimation were developed and
comprehensively analyzed, resulting in a collection of top-ranked enzyme candidates
recommended as promising candidates. These candidates were further characterized through a
series of experiments, achieving a success rate of >60%, with several novel enzymes
outperforming benchmark enzymes. The validated enzymes were subsequently used to expand
the data set created in the first framework, thereby closing the DBTL loop.
Comprehensive studies were conducted to evaluate model performance and characterize enzyme
activities. Major tasks included developing ML/AI models or agents for each framework,
analyzing the accuracy of each ML/AI model, comparing the models with benchmark methods,
developing synthetic biology and biochemistry experiments, and incorporating bioinformatics
into the discovery pipeline.