Javascript must be enabled to continue!
Can Synthetic Data Replace Real Data? Evaluating the Efficacy, Limits, and Label Trustworthiness of Generative Data in Supervised Learning
View through CrossRef
The scarcity of high-quality, human-annotated training data has made synthetic data generation an increasingly important component of supervised machine learning pipelines, yet the conditions under which synthetic data can safely substitute for real data remain poorly understood, particularly in tabular healthcare domains where privacy constraints restrict data access. This study presents an empirical and theoretical investigation into the efficacy and limits of CTGAN-generated synthetic data using a real-world medical appointment dataset of 110,527 records with a 4:1 class imbalance. Applying the Train-Synthetic-Test-Real (TSTR) protocol across five supervised classifiers and six synthetic-to-real mixing ratios, we find that AUC-ROC degrades consistently as synthetic proportion increases, with all classifiers entering a critical degradation zone beyond 60% synthetic data. Tree-based architectures exhibit average AUC-ROC reductions of 8.72%, while parametric models show 10.55% reduction at 100% synthetic, both measured against real-data baselines. Beyond distributional mismatch, we identify and formally define synthetic label corruption-a distinct failure mode in which automatically assigned synthetic labels carry an independent error distribution not attributable to generative distributional drift-and demonstrate that 96.6% of these corrupted labels reflect a systematic minority-class overgeneration bias in CTGAN's conditional sampling mechanism. To address this, we propose the Synthetic Label Trustworthiness Estimation (SLTE) framework, a pre-training pipeline comprising Label Confidence Score computation via a real-data anchor model, topology-aware k-nearest neighbour label correction, and uncertainty-stratified sample routing. SLTE reduces tree-based AUC degradation from 8.72% to 3.84% and parametric degradation from 10.55% to 7.94%, recovering 66.6% of Gradient Boosting's lost discriminative performance at full synthetic substitution, offering a principled foundation for label-aware data augmentation in privacy-critical supervised learning environments.
Title: Can Synthetic Data Replace Real Data? Evaluating the Efficacy, Limits, and Label Trustworthiness of Generative Data in Supervised Learning
Description:
The scarcity of high-quality, human-annotated training data has made synthetic data generation an increasingly important component of supervised machine learning pipelines, yet the conditions under which synthetic data can safely substitute for real data remain poorly understood, particularly in tabular healthcare domains where privacy constraints restrict data access.
This study presents an empirical and theoretical investigation into the efficacy and limits of CTGAN-generated synthetic data using a real-world medical appointment dataset of 110,527 records with a 4:1 class imbalance.
Applying the Train-Synthetic-Test-Real (TSTR) protocol across five supervised classifiers and six synthetic-to-real mixing ratios, we find that AUC-ROC degrades consistently as synthetic proportion increases, with all classifiers entering a critical degradation zone beyond 60% synthetic data.
Tree-based architectures exhibit average AUC-ROC reductions of 8.
72%, while parametric models show 10.
55% reduction at 100% synthetic, both measured against real-data baselines.
Beyond distributional mismatch, we identify and formally define synthetic label corruption-a distinct failure mode in which automatically assigned synthetic labels carry an independent error distribution not attributable to generative distributional drift-and demonstrate that 96.
6% of these corrupted labels reflect a systematic minority-class overgeneration bias in CTGAN's conditional sampling mechanism.
To address this, we propose the Synthetic Label Trustworthiness Estimation (SLTE) framework, a pre-training pipeline comprising Label Confidence Score computation via a real-data anchor model, topology-aware k-nearest neighbour label correction, and uncertainty-stratified sample routing.
SLTE reduces tree-based AUC degradation from 8.
72% to 3.
84% and parametric degradation from 10.
55% to 7.
94%, recovering 66.
6% of Gradient Boosting's lost discriminative performance at full synthetic substitution, offering a principled foundation for label-aware data augmentation in privacy-critical supervised learning environments.
Related Results
Woningcorporaties en Vastgoedontwikkeling
Woningcorporaties en Vastgoedontwikkeling
This summary highlights the findings of the PhD-thesis ‘Woningcorporaties en Vastgoedontwikkeling: Fit for Use’ (‘Housing associations and Real Estate Development: Fit for Use?’). ...
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
CREATING LEARNING MEDIA IN TEACHING ENGLISH AT SMP MUHAMMADIYAH 2 PAGELARAN ACADEMIC YEAR 2020/2021
The pandemic Covid-19 currently demands teachers to be able to use technology in teaching and learning process. But in reality there are still many teachers who have not been able ...
Safety and Efficacy of Atezolizumab in Ovarian Cancer
Safety and Efficacy of Atezolizumab in Ovarian Cancer
Abstract
Introduction
Although the efficacy of PD-L1 blockade has been evaluated in analyses that combine pharmacologically distinct antibodies, the specific efficacy and safety of...
Off-label use in Palliative Care – more common than expected. A retrospective chart review
Off-label use in Palliative Care – more common than expected. A retrospective chart review
Abstract
Objective
Off-label drug use seems to be integral to palliative care pharmacotherapy. Balancing potential risks and be...
Breast Self-Examination System Using Multifaceted Trustworthiness: Observational Study (Preprint)
Breast Self-Examination System Using Multifaceted Trustworthiness: Observational Study (Preprint)
BACKGROUND
Breast cancer is the leading cause of mortality among women worldwide. However, female patients often feel reluctant and embarrassed about meetin...
Self-Supervised Transformer Networks: Unlocking New Possibilities for Label-Free Data
Self-Supervised Transformer Networks: Unlocking New Possibilities for Label-Free Data
In machine learning, self-supervised transformer networks have become a new way of doing things, especially when it comes to handling and understanding huge amounts of data that ha...
Hubungan Pengetahuan terkait Label Gizi dengan Kebiasaan Membaca Label Gizi pada Siswa SMA Al-Islam
Hubungan Pengetahuan terkait Label Gizi dengan Kebiasaan Membaca Label Gizi pada Siswa SMA Al-Islam
Latar Belakang: Masih sedikit konsumen yang dapat memahami dan menggunakan label gizi sesuai dengan fungsinya. Hal ini dikarenakan masih rendahnya kesadaran masyarakat terkait pent...
ANALISA KESESUAIAN STANDAR LABEL PANGAN PADA KEMASAN PRODUK BISKUIT LOKAL DAN IMPOR TEREGISTRASI DI BADAN PENGAWAS OBAT DAN MAKANAN
ANALISA KESESUAIAN STANDAR LABEL PANGAN PADA KEMASAN PRODUK BISKUIT LOKAL DAN IMPOR TEREGISTRASI DI BADAN PENGAWAS OBAT DAN MAKANAN
Abstrak
Hal yang mendasari penelitian ini karena banyak ditemukan label kemasan yang beredar tidak sesuai standar Peraturan Pemerintah No. 69 Tahun 1999. Berdasarkan Peraturan Pem...

