Abstract
This work extends synthetic-population generation with machine learning, using gradient-boosted trees (XGBoost) to model the conditional distributions from which synthetic individuals and households are drawn. By capturing non-linear relationships and interactions that parametric models can miss, the approach produces synthetic data that stay close to the structure of the real population. A simulated-annealing calibration step then adjusts the generated population so that it matches known aggregate margins from official sources, keeping it consistent with published totals. Published in Algorithms, the method targets privacy-compliant population data for microsimulation and official-statistics applications.
synthetic data population data XGBoost calibration