Chinese Journal of Tissue Engineering Research ›› 2026, Vol. 30 ›› Issue (36): 9589-9596.doi: 10.12307/2026.899

Previous Articles     Next Articles

Diabetes prediction and analysis of influencing factors integrating genetic information

Liu Yi1, Lu Jiarong2, Wu Jianyong1    

  1. 1Key Laboratory of Special Environment and Health Research in Xinjiang, School of Public Health, Xinjiang Medical University, Urumqi 830017, Xinjiang Uygur Autonomous Region, China; 2School of Data and Statistical Sciences, Xinjiang University of Finance and Economics, Urumqi 830012, Xinjiang Uygur Autonomous Region, China; 3College of Statistics and Data Science, Xinjiang University of Finance and Economics, Urumqi 830012, Xinjiang Uygur Autonomous Region, China
  • Received:2025-10-17 Revised:2026-03-05 Online:2026-12-28 Published:2026-05-25
  • Contact: Wu Jianyong, PhD, Doctoral supervisor, Associate researcher, Key Laboratory of Special Environment and Health Research in Xinjiang, School of Public Health, Xinjiang Medical University, Urumqi 830017, Xinjiang Uygur Autonomous Region, China
  • About author:Liu Yi, PhD candidate, Key Laboratory of Special Environment and Health Research in Xinjiang, School of Public Health, Xinjiang Medical University, Urumqi 830017, Xinjiang Uygur Autonomous Region, China
  • Supported by:
    The Project of the Key Laboratory of Special Environment and Health Research in Xinjiang, No. SKL-SEHR-2024-13 (to LY) 

Abstract: BACKGROUND: Early risk assessment and accurate diagnosis are crucial for the clinical prevention and management of diabetes mellitus. Genetic factors play a significant role in the pathogenesis of diabetes; however, most current studies insufficiently integrate genetic information into risk modeling.
OBJECTIVE: To construct a comprehensive feature dataset integrating genetic factors, anthropometric indicators, and insulin metabolism metrics, and to develop an interpretable diabetes mellitus prediction model for early risk assessment and prediction. 
METHODS: A publicly available diabetes prediction dataset from Kaggle was used, comprising 5 070 valid samples, including 1 936 diabetic and 3 134 non-diabetic individuals. The model was trained on the integrated feature dataset using the Histogram-based Gradient Boosting Decision Tree (HistGBDT) algorithm. Hyperparameters were optimized via grid search. Model performance was evaluated using accuracy, precision, recall, F1-score, and area under the receiver operating characteristic curve. The SHapley Additive exPlanations (SHAP) framework was applied to identify key features influencing diabetes risk and to enhance model interpretability. Ablation studies were conducted to validate the contribution of genetic factors.
RESULTS AND CONCLUSION: The proposed interpretable diabetes prediction model outperformed existing mainstream models across all evaluation metrics, achieving an accuracy of 98.03%, precision of 97.66%, recall of 97.16%, F1-score of 97.41%, and the area under the receiver operating characteristic curve of 97.86%, representing an improvement of 1%–4% in overall performance. Ablation experiments demonstrated that integrating genetic factors enables more comprehensive and effective capture of diabetes risk characteristics. SHAP analysis identified triceps skinfold thickness, insulin release test, body mass index, oral glucose tolerance test, and diastolic blood pressure as the top five influential factors. The interpretability analysis provides a theoretical foundation and technical support for early diabetes identification and personalized health management.


Key words: diabetes mellitus, prediction model, machine learning, genetic factors, interpretable model, risk assessment

CLC Number: