Protein feature engineering framework for AMPylation site prediction

Hardik Prabhu; Hrushikesh Bhosale; Aamod Sane; Renu Dhadwal; Vigneshwar Ramakrishnan; Jayaraman Valadi

doi:10.1038/s41598-024-58450-8

Protein feature engineering framework for AMPylation site prediction

Sci Rep. 2024 Apr 15;14(1):8695. doi: 10.1038/s41598-024-58450-8.

Authors

Hardik Prabhu^{1

2}, Hrushikesh Bhosale¹, Aamod Sane¹, Renu Dhadwal¹, Vigneshwar Ramakrishnan³, Jayaraman Valadi⁴

Affiliations

¹ Computing and Data Sciences, FLAME University, Pune, 412115, India.
² Robert Bosch Centre for Cyber Physical Systems, Indian Institute of Science, Bengaluru, 560012, India.
³ Bioinformatics Center, School of Chemical and Biotechnology, SASTRA Deemed to be University, Thanjavur, 613401, India.
⁴ Computing and Data Sciences, FLAME University, Pune, 412115, India. jayaraman.vk@flame.edu.in.

PMID: 38622194
DOI: 10.1038/s41598-024-58450-8

Abstract

AMPylation is a biologically significant yet understudied post-translational modification where an adenosine monophosphate (AMP) group is added to Tyrosine and Threonine residues primarily. While recent work has illuminated the prevalence and functional impacts of AMPylation, experimental identification of AMPylation sites remains challenging. Computational prediction techniques provide a faster alternative approach. The predictive performance of machine learning models is highly dependent on the features used to represent the raw amino acid sequences. In this work, we introduce a novel feature extraction pipeline to encode the key properties relevant to AMPylation site prediction. We utilize a recently published dataset of curated AMPylation sites to develop our feature generation framework. We demonstrate the utility of our extracted features by training various machine learning classifiers, on various numerical representations of the raw sequences extracted with the help of our framework. Tenfold cross-validation is used to evaluate the model's capability to distinguish between AMPylated and non-AMPylated sites. The top-performing set of features extracted achieved MCC score of 0.58, Accuracy of 0.8, AUC-ROC of 0.85 and F1 score of 0.73. Further, we elucidate the behaviour of the model on the set of features consisting of monogram and bigram counts for various representations using SHapley Additive exPlanations.

MeSH terms

Adenosine Monophosphate / metabolism
Amino Acid Sequence
Protein Processing, Post-Translational*
Threonine / metabolism
Tyrosine* / metabolism

Substances

Tyrosine
Adenosine Monophosphate
Threonine