Convergent Feature Selection and Predictive Modeling Identify AIM2 as a Potential Molecular Diagnostic Marker in Acute Myeloid Leukemia

Authors

  • Shehzad Khalil Institute of Biotechnology and Genetic Engineering, The University of Agriculture, Peshawar, Khyber Pakhtunkhwa, Pakistan
  • Fayaz Ahmad Khan Department of Biotechnology, Abdul Wali Khan University, Mardan 23200, Khyber Pakhtunkhwa, Pakistan
  • Ahmed Mujtaba Veesar Institute of Biotechnology and Genetic Engineering, The University of Sindh, Jamshoro, Sindh, Pakistan
  • Zain Uddin Department of Biotechnology, Balochistan University of Information Technology, Engineering and Management Sciences, Pakistan
  • Halaar Hussain Institute of Biotechnology and Genetic Engineering, The University of Sindh, Jamshoro, Sindh, Pakistan
  • Abubakar Basiru Department of Medical Laboratory Science, Federal University of Health Sciences, Ila Orangun, Osun State, Nigeria
  • Abbas Ahmad Department of Biotechnology, Abdul Wali Khan University, Mardan 23200, Khyber Pakhtunkhwa, Pakistan

DOI:

https://doi.org/10.63056/ijair.2.3.2026.301

Keywords:

Acute myeloid leukemia, AIM2, Biomarker, Transcriptomics, Machine learning, Random Forest, SVM-RFE, Inflammasome, RNA-sequencing, Differential expression

Abstract

Background: Acute myeloid leukemia (AML) is an aggressive hematologic malignancy that occurs by clonal expansion of myeloid progenitor cells, is highly molecularly heterogeneous, and has a poor prognosis and low precision of diagnosis. Defining strong transcriptomic signatures that can differentiate AML from normal hematopoietic progenitors would be critical to increase the accuracy of diagnosis and the therapeutic stratification.

Methods: RNA-sequencing data were obtained from GEO (accession number GSE235686) for 15 AML patients and 2 CD34+ healthy hematopoietic stem/progenitor cells (HSPC) controls. The limma-voom framework was then used for differential expression analysis after data preprocessing and data quality filtering. A composite scoring system using multiple metrics was built to rank candidate biomarkers. The functional enrichment analysis was performed with clusterProfiler and ShinyGO v0.8. The Wilcoxon rank-sum test with AUC filtering, Random Forest feature importance, and Support Vector Machine with Recursive Feature Elimination (SVM-RFE) were used as three complementary machine learning algorithms to determine a final biomarker panel. Leave one out cross validation (LOOCV) was used to assess the classification performance on six machine learning models.

Results: There were 857 significantly differentially expressed genes (DEGs) identified (|log₂FC| > 1, FDR < 0.05); 631 of these (73.6%) were up-regulated in AML. Functional enrichment identified key pathways of immune response, cytokine signaling and DNA replication as dysregulated. A composite score was created that favoured 30 candidate biomarkers that were all up-regulated in patients with AML with complete consistency. A multi-algorithm feature selection led to a set of 13 final biomarkers (selected by two or more algorithms). The only gene that was selected by all three algorithms (AUC = 1.000) was Absent in Melanoma 2 (AIM2). Random Forest classification had a perfect LOOCV performance in terms of Accuracy (1.000), AUC (1.000), Sensitivity (1.000) and Specificity (1.000) compared with Logistic Regression (AUC = 0.733).

Conclusions: Incorporation of transcriptomic profiling and multi-algorithm machine learning is an effective strategy for discovering strong AML biomarkers. AIM2 is the most consistently reported candidate as it has been known to play a role in innate immune signaling through the inflammasome, and has not been found in normal CD34+ HSPCs. These results may be used to establish that AIM2 is a diagnostic and therapeutic target in AML which could be tested in larger cohorts in the future

Downloads

Published

2025-09-10