本平台为互联网非涉密平台,严禁处理、传输国家秘密、工作秘密或敏感信息

基于深度学习探究胰蛋白酶催化蛋白质的特异性酶切预测
CSTR:
作者:
作者单位:

作者简介:

刘洋(1998-),男,硕士,研究方向:食品生物信息技术与多组学技术,E-mail:huhaoaaq@163.com 通讯作者:徐巨才(1991-),男,博士,讲师,研究方向:食品生物信息技术与智造技术,E-mail:xujucai.happy@163.com;共同通讯作者:刘磊(1982-),男,博士,研究员,研究方向:农产品加工与品质调控,E-mail:liulei309@tom.com

通讯作者:

中图分类号:

基金项目:

国家自然科学基金青年科学基金项目(32402276);广东省科技创新战略专项市县科技创新支撑(大专项+任务清单)项目-暨2023年江门市关键核心技术“揭榜挂帅”制项目(2023780200060009632);广东省基础与应用基础研究基金联合基金青年基金项目(2022A1515110711);广东省“百千万工程”项目(BQW2024001)


Prediction of Trypsin-catalyzed Protein Cleavage Specificity Based on Deep Learning
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    胰蛋白酶催化蛋白质的特异性酶切预测对于蛋白质的理论降解和结构分析具有重要意义。该研究采用卷积神经网络和长短期记忆网络构建了一种基于深度学习的胰蛋白酶催化蛋白质酶解预测模型,并探讨了不同超参数对模型性能的影响。结果表明:较低的学习率、较高的Batch size和较浅的卷积层数对模型训练的稳定性和收敛性有利,可保障较好的预测效果,模型的较优工作参数为:学习率0.001、Batch size 512、卷积层数1。在此条件下,模型在PXD010627数据集上的准确率为0.950、特异性为0.987、精确度为0.986、召回率为0.961、F1分数为0.973,显示出良好的预测能力和稳定性。进一步应用于不同物种的公开数据集,模型的准确率均保持在0.920以上,AUC值、精确率、召回率和F1分数均在0.900~0.989之间,说明模型在胰蛋白酶催化蛋白质特异性酶切预测体系中具备较强的泛化能力,可有效提升酶切位点预测的准确性和可靠性。本研究将有望为蛋白质组学鉴定及空间结构分析提供了一种新思路,推动蛋白质酶解研究与生物信息学的发展。

    Abstract:

    The prediction of trypsin-catalyzed specific proteolysis is vital for guiding the theoretical degradation and structural analysis of proteins. In this study, a deep learning-based model was constructed to predict trypsin-catalyzed protein cleavage using convolutional neural networks (CNN) and long short-term memory (LSTM) networks. The impact of various hyperparameters on the model's performance was also explored. The results showed that a lower learning rate, higher batch size, and fewer convolutional layers were beneficial for the model's training stability, ensuring superior predictive outcomes. The optimal parameters were identified as a learning rate of 0.001, batch size of 512, and one convolutional layer. In that case, the model achieved an accuracy of 0.950, specificity of 0.987, precision of 0.986, recall of 0.961, and an F1 score of 0.973 on dataset PXD010627, demonstrating excellent predictive capability and stability. Furthermore, when applied to publicly available datasets from different species, the model maintained an accuracy >0.920, with AUC values, precision, recall, and F1 scores ranging between 0.900 and 0.989. This indicates a strong generalization ability of the model in predicting specific cleavage sites of trypsin-catalyzed proteins which could considerably enhance the accuracy and reliability of cleavage site predictions. We hope this work could offer new ideas for protein identification and spatial structure analysis in proteomics, promoting advancements in proteolytic research and bioinformatics.

    参考文献
    相似文献
    引证文献
引用本文

刘洋,黄峻洪,岑幸仪,肖楚乔,李谆淳,陈选,李桂香,徐巨才,刘磊.基于深度学习探究胰蛋白酶催化蛋白质的特异性酶切预测[J].现代食品科技,2025,41(11):68-76.

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2024-08-04
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2025-12-05
  • 出版日期:
文章二维码