Abstract
Synthetic data are often assumed to be safe because no record corresponds to a real person, but that assumption needs to be tested. This paper assesses the disclosure risk of fully synthetic population data, examining whether realistic synthetic records can nonetheless leak information about the real population they are modelled on. Using synthetic EU-SILC populations, it provides early evidence that carefully generated synthetic data can sharply reduce re-identification risk while remaining useful. Presented at Privacy in Statistical Databases, the study helped establish disclosure-risk evaluation as a necessary step for synthetic data, not an afterthought.
synthetic data synthetic populations disclosure risk EU-SILC