Abstract
Deduplicating electronic patient records usually requires costly manual review of uncertain record pairs. This paper develops active-learning strategies built on classification trees that select the most informative pairs for human labelling, so that high linkage quality is reached with far less manual effort. By focusing expert attention where it most improves the classifier, the approach makes large-scale deduplication more practical. Published in the Journal of Biomedical Informatics, it contributes methods directly relevant to maintaining clean, de-duplicated medical databases.
record linkage deduplication active learning classification trees