PhD Final Defense – Renjing Jiang

Aug 10, 2026   10:00 am  
CEEB 3019
Sponsor
Department of Civil and Environmental Engineering

Artificial intelligence for enzymatic plastic biodegradation: Data extraction, function

prediction, and novel enzyme discovery

Advisor: Professor Na Wei

Zoom Meeting (ID: 832 9787 6779 Password: 756228)

Abstract

This dissertation develops a closed-loop pipeline for discovering plastic-degrading enzymes to

address the long-standing challenge of plastic accumulation in the environment by integrating

machine learning, artificial intelligence (ML/AI), and wet-lab experiments. The pipeline consists

of a data extraction framework (termed PEZy-Extract) to automatically extract enzymatic plastic

degradation data from scientific literature, a function prediction framework (termed PED) to

predict whether an input enzyme-plastic pair is degradable or non-degradable, an enzyme mining

framework (termed PEZy-Miner) to rank potentially degradable enzyme-plastic pairs from a

large candidate pool, and, ultimately, an experimental characterization framework to

comprehensively analyze the top-ranked enzymes for degrading plastics of interest.

The pipeline forms a closed Design-Build-Test-Learn (DBTL) loop, in which each framework

provides the foundation for downstream stages while incorporating feedback from upstream

framework. In the first framework, PEZy-Extract, AI agents were developed to create a data set

integrating information from a wide range of experimentally verified enzymes and various

plastic substrates. The compiled data set can be used to uncover domain knowledge as well as

train downstream ML models. In the second framework, PED, protein language models were

used to learn the abundant contextual information in enzyme sequences, and feature extraction

was performed at both the amino acid level and the global sequence level. Binary classification

was subsequently performed, achieving an overall accuracy of 90.2% and outperforming

sequence-based protein classification models reported in the existing literature. In the third

framework, PEZy-Miner, confidence and uncertainty estimation were developed and

comprehensively analyzed, resulting in a collection of top-ranked enzyme candidates

recommended as promising candidates. These candidates were further characterized through a

series of experiments, achieving a success rate of >60%, with several novel enzymes

outperforming benchmark enzymes. The validated enzymes were subsequently used to expand

the data set created in the first framework, thereby closing the DBTL loop.

Comprehensive studies were conducted to evaluate model performance and characterize enzyme

activities. Major tasks included developing ML/AI models or agents for each framework,

analyzing the accuracy of each ML/AI model, comparing the models with benchmark methods,

developing synthetic biology and biochemistry experiments, and incorporating bioinformatics

into the discovery pipeline.

link for robots only