Browser-based deep learning

O-GlcNAcPRED-DL

Predict human and mouse protein O-GlcNAcylation sites from FASTA sequences with an ensemble deep-learning model.

Protein sequence passing through a neural network to site scores for O-GlcNAcPRED-DL

About the model

Sequence-based candidate site prediction

O-GlcNAcylation functions in a protein- and site-specific manner. Despite substantial progress in experimentally mapping O-GlcNAc sites, prediction remains challenging. O-GlcNAcPRED-DL applies the published species-specific deep-learning ensembles to human or mouse protein sequences.

Compared with existing methods, O-GlcNAcPRED-DL showed improved sensitivity and accuracy in the published study. It can help expedite discovery of candidate O-GlcNAc sites, especially in human and mouse proteins, and support study of their functions in physiology and disease. Predictions should prioritize experimental work, not replace experimental evidence.

O-GlcNAcAtlas data are split for training and testing; Word2Vec, one-hot, BLOSUM62, and AAindex features feed CNN-BiLSTM models whose selected outputs vote in O-GlcNAcPRED-DL
Model-development workflow preserved from the publication concept: Atlas train/test splits, four feature representations, CNN–BiLSTM models, model selection, and ensemble voting. The published species-specific ensemble now runs in the browser.

Cite O-GlcNAcPRED-DL

Fengzhu Hu, Weiyu Li, Yaoxiang Li, Chunyan Hou, Junfeng Ma, and Cangzhi Jia. O-GlcNAcPRED-DL: prediction of protein O-GlcNAcylation sites based on an ensemble model of deep learning. Journal of Proteome Research. 2024;23(1):95–106. doi:10.1021/acs.jproteome.3c00458.