Abstract
This article surveys the practical challenges of using synthetic data generation for tabular microdata, the format most common in social, economic, and health statistics. It examines the persistent tension between fidelity and utility — keeping synthetic records statistically faithful and analytically useful — and the disclosure risks that remain even when no real record is copied. The review spans method families from statistical models to deep generative approaches and the difficulty of evaluating them fairly. Published in Applied Sciences, it offers practitioners a structured map of the pitfalls to anticipate before adopting synthetic data in production.
synthetic data tabular microdata data generation