Anonymization and Synthetic Data — Methods, Tools, and Applications
A one-day, hands-on workshop spanning the full arc of data anonymization and synthetic data — legal foundations, identity and risk, the core SDC techniques, hands-on sdcMicro, utility measurement, and modern synthetic-data generation.
Syllabus
Introduction & legal foundations
Why anonymization matters and the legal frame it operates in.
The concept of identity in anonymization
What identity and re-identification actually mean for microdata.
General methodological overview
The landscape of anonymization and synthetic-data methods.
Key anonymization techniques
Microaggregation, local recoding, local suppression, and classical synthetic data.
Re-identification risk
Attacker scenarios and risk metrics, worked through examples.
Utility metrics
General and use-case-specific utility — concepts and examples.
Hands-on with sdcMicro
Two practical sessions applying the methods in R with sdcMicro.
Synthetic data generation
Statistical vs. deep-learning and LLM approaches to synthetic data.
Instructors Murat Sariyar (Bern University of Applied Sciences) · Matthias Templ (FHNW)