Loading

Reviewing Candidates

Pairs scoring at or above the review threshold are raised as candidates. Working that queue is what keeps the data clean.

Where to find it

Architect Panel → Data:

  • Match Policies — the policies producing candidates
  • Merges — the record of what was merged

Architect Panel → Automation:

  • Tasks — Duplicate Detection Sweep, daily

The sweep

The Duplicate Detection Sweep task runs daily and compares records under each enabled policy. It ships disabled — enable it under Automation → Tasks, or candidates are only ever produced when something else triggers a comparison.

Work it regularly

Duplicates get harder to resolve with age. Two records created yesterday differ by a typo; two records a year old have separate activity, separate correspondence and possibly separate transactions, and merging them means reconciling all of it.

A short weekly session on a small queue is far less work than an annual clean-up, and produces better decisions.

What to check on a pair

  1. Which fields matched, and what they scored. A high score from one strong field is different from a high score accumulated from several weak ones.
  2. The fields that did not match. Two records agreeing on name and address but with different dates of birth are probably two people.
  3. The activity on each. Records with genuinely separate histories deserve more scrutiny.
  4. When each was created, and by what route. Two records created minutes apart through a form are almost certainly one submission twice.

Family members are the classic false match

Same surname, same address, similar or shared contact details. Two people at one household are not a duplicate, and merging them is both a data error and a data protection one — you have combined two individuals' records.

Where a policy raises these repeatedly, it needs a distinguishing field weighted higher rather than a reviewer being more careful.

When you cannot tell

Leave it. An unresolved candidate is a small annoyance; an incorrect merge combines two people's data.

Where it matters, find out — contact them, or check another source. Where it does not, leaving it is a legitimate outcome.

Feed the queue back into the policy

Patterns in what you dismiss are telling you the policy is wrong. Repeatedly rejecting the same shape of pair means a weight or a threshold needs adjusting, not that reviewers should keep rejecting them.

Review the policy after the first few weeks of real use.

Decide who owns the queue

It needs to be somebody's job, with enough knowledge of the data to judge an ambiguous pair. An unowned queue simply grows, and a growing queue eventually gets cleared carelessly.

Worked example

A weekly review of about twenty candidates resolves fifteen as clear duplicates from a web form submitted twice. Three are family members at one address and are dismissed. Two are genuinely unclear and left. After a month, the recurring family-member pattern prompts adding date of birth to the policy with a high weight, and those stop appearing.

Recommendations

  • Enable the sweep and work the queue weekly.
  • Look at what did not match, not only what did.
  • Leave genuinely ambiguous pairs alone.
  • Adjust the policy when you dismiss the same pattern repeatedly.