Pseudonymization and Anonymization of Tabular Data — Methods, Tools, and Applications
A two-day, hands-on deepening of the one-day course: legal context, disclosure-risk measurement, the full sdcMicro anonymization toolbox, utility assessment, and synthetic-data generation with synthpop — applied end-to-end on real tabular microdata. Recommended for teams who want to go a step further than the one-day overview.
Syllabus
Introduction, identity & regulation
Pseudonymization vs anonymization; the revised Swiss FADP (revDSG) and GDPR; what counts as identifying; data-sharing modes (PUF / SUF / secure use).
Overview of the SDC workflow
The end-to-end pipeline: classify variables, define the disclosure scenario, measure risk, apply targeted methods, re-measure, document and release.
Anonymization methods
Global recoding and generalization, local suppression (k-anonymity via sdcMicro kAnon), PRAM, microaggregation, noise addition, swapping and shuffling — when each applies.
Disclosure-risk measurement
Individual and global re-identification risk, frequency counts, l-diversity for sensitive attributes, SUDA — measuring risk per record to target treatment.
Data utility
Information-loss measures and the risk–utility trade-off; choosing parameters against a utility target instead of by guesswork.
Hands-on: anonymizing real microdata
End-to-end sdcMicro session on a real survey dataset (EU-SILC): build the SDC object, recode, suppress, perturb, and produce an SDC report.
Synthetic data
Generating synthetic tabular data with synthpop; evaluating fidelity and disclosure risk with riskutility; when synthesis beats classical anonymization.
From course to practice
Applying the workflow to participants' own data-sharing problems; documentation that stands up to a regulator.
Instructors Matthias Templ