Software
Open-source, peer-reviewed, auditable
Our open-source R packages for risk estimation, anonymization, and synthetic data — on CRAN and GitHub — with download statistics for the CRAN packages.
Packages
CRAN · since 2007
sdcMicro
R package for statistical disclosure control of microdata — risk estimation, anonymization methods, and utility measurement. Peer-reviewed in the Journal of Statistical Software.
CRAN · since 2014
simPop
Simulation of synthetic populations from survey and census data.
CRAN · since 2010
RecordLinkage
R package for probabilistic and deterministic record linkage — deduplication and entity resolution across data sources, with comparison, classification, and evaluation tools.
GitHub
robSynth
Robust synthetic microdata generation when the training data are contaminated — a drop-in alternative to synthpop using MM-estimators, weighted logistic regression, and tree-based robust methods so outliers and coding errors don't propagate into the synthetic output.
In development
synvey
Design-aware, robust synthetic data generation — the successor to robSynth. Replaces the conditional models in sequential synthesis with MM-estimators and Huber-weighted logistic regression, keeping the synthetic data close to the clean data-generating process even when the source is contaminated.
CRAN downloads
Daily downloads of sdcMicro, simPop, and RecordLinkage recorded at the Posit (RStudio) CRAN mirror via cranlogs — the only CRAN mirror that publishes download logs, so the true total across all mirrors is higher. Refreshed at each build.
Granularity
Range
Time frame Last 180 days · 15 Jan – 13 Jul
7 days1821 days
Total downloads 22,186
Daily average 123
Peak day 301 31 Mar 2026
vs previous 180d
-46%
By package · selected window
sdcMicro 6,664
simPop 2,564
RecordLinkage 12,958
Counts are from the Posit (RStudio) CRAN mirror via cranlogs — one of many CRAN mirrors, so the true total is higher. Refreshed at each site build.
Last refreshed: 15 Jul 2026, 05:45 UTC
Try the methods
Interactive, in-browser demos — anonymise a table, generate synthetic data, or walk the full pipeline. Everything runs in your browser; no data leaves the page.
Open the demos