Skip to content
SwissAnon
Back to research
2024 Peer-reviewed

Assessing the Potentials of LLMs and GANs as State-of-the-Art Tabular Synthetic Data Generation Methods

M. Miletic, Murat Sariyar — Privacy in Statistical Databases (Lecture Notes in Computer Science)

Abstract

This paper assesses two families of deep generative models — large language models (LLMs) and generative adversarial networks (GANs) — as methods for producing synthetic tabular data. It evaluates how well each reproduces the statistical structure of the original data (fidelity and utility) while considering the disclosure risk that synthetic records can still carry. By placing these newer approaches alongside established statistical synthesis methods, the study clarifies where deep learning genuinely helps and where it falls short for tabular microdata. Presented at Privacy in Statistical Databases, it adds evidence to the ongoing debate over deep generative models in official and health statistics.

synthetic data large language models GANs tabular data