Privacy preserving record linkage in the presence of missing values

Yuan Chi*, Jun Hong, Anna Jurek, Weiru Liu, Dermot O'Reilly

*Corresponding author for this work

Research output: Contribution to journalArticle (Academic Journal)peer-review

13 Citations (Scopus)
275 Downloads (Pure)

Abstract

The problem of record linkage is to identify records from two datasets, which refer to the same entities (e.g. patients). A particular issue of record linkage is the presence of missing values in records, which has not been fully addressed. Another issue is how privacy and confidentiality can be preserved in the process of record linkage. In this paper, we propose an approach for privacy preserving record linkage in the presence of missing values. For any missing value in a record, our approach imputes the similarity measure between the missing value and the value of the corresponding field in any of the possible matching records from another dataset. We use the k-NNs (k Nearest Neighbours in the same dataset) of the record with the missing value and their distances to the record for similarity imputation. For privacy preservation, our approach uses the Bloom filter protocol in the settings of both standard privacy preserving record linkage without missing values and privacy preserving record linkage with missing values. We have conducted an experimental evaluation using three pairs of synthetic datasets with different rates of missing values. Our experimental results show the effectiveness and efficiency of our proposed approach.

Original languageEnglish
Pages (from-to)199-210
Number of pages12
JournalInformation Systems
Volume71
Early online date5 Jul 2017
DOIs
Publication statusPublished - 1 Nov 2017

Research Groups and Themes

  • Jean Golding

Keywords

  • Record linkage
  • Probabilistic record linkage
  • Privacy preserving record linkage
  • Missing values
  • k-nearest neighbours
  • Data encryption

Fingerprint

Dive into the research topics of 'Privacy preserving record linkage in the presence of missing values'. Together they form a unique fingerprint.

Cite this