Synthetic Data Generation and AI Model Performance: The Moderating Role of Data Quality

Authors

  • Abdul Basit Department of Computer Science, Habib University, Karachi, Pakistan

DOI:

https://doi.org/10.63056/tljet.2.2.2026.273

Keywords:

synthetic data, artificial intelligence, machine learning, data quality, generative AI, synthetic-data generation, model performance, data augmentation, predictive performance, moderation analysis

Abstract

With the growing need for big and diverse datasets, the use of artificial data has become a viable alternative for artificial intelligence (AI) and machine learning applications. But a higher number of synthetic data does not necessarily lead to better model performance, as useful observations depend on the quality of the generated observations. This study assessed the effect of synthetic data generation on the performance of AI models and explored the moderation effect of data quality. The data sets were generated from a controlled experimental machine-learning design with varying amounts of synthetic observations (0%, 25%, 50%, 75%, and 100%). The model architecture, hyperparameters, model pre-processing and model testing were kept fixed for all experimental conditions and the machine-learning models were trained under each experimental condition. Synthetic-data quality was evaluated using statistical similarity, distributional similarity, completeness, consistency, diversity and the maintenance of feature-level relationships. The accuracy, precision, recall, F1-score and ROC-AUC were applied to assess the performance of the AI models on an independent real-world testing set. The results showed that the level of synthetic-data augmentation that results in moderate added data had the highest overall model performance, with the 50% synthetic-data condition showing the highest performance compared to the baseline real-data-only condition and the complete synthetic-data condition. Additionally, the moderation analysis revealed that the data quality had a significant effect on enhancing the relationship between the use of synthetic data and AI model performance. The use of lower-quality synthetic observations reduced the benefits of using synthetic observations more, while high-quality synthetic data provided useful information for training of the models. The study found that when it came to the use of synthetic data, it was important for them not only to be adequate in number, but to be high quality and representative of the observations generated. The results emphasized the need for a multi-dimensional synthetic-data evaluation prior to their use in developing and deploying AI models.

Downloads

Published

2026-05-26

How to Cite

Basit, A. (2026). Synthetic Data Generation and AI Model Performance: The Moderating Role of Data Quality. Turing Ledger Journal of Engineering & Technology, 2(2), 35–48. https://doi.org/10.63056/tljet.2.2.2026.273