Technical Notes / Editorial draft
When two names are not one person
A proposed approach to historical entity resolution that preserves competing identities and the evidence behind each match.
Prepared for editorial review · Proposed approach · No empirical results reported
Identity claims across records
A transcription identifies what a page appears to say. It does not establish that a name belongs to a person already found in another collection. Our design treats those as separate decisions: first a source-bound mention, then a proposed identity match.
Consider two fictional records naming Joannes Martin. Matching names are a reason to investigate. They are not enough to combine the records. Dates, places, recorded relatives and the purpose of each document need to be examined together.
Source mentions and normalisation
The proposed data model retains the original spelling, a normalized search form, the source reference and the location of the mention on the page. Normalization should improve discovery without overwriting the evidence.
Candidate matches should carry an explanation: which observations support the match, which contradict it and what remains unknown. A researcher should be able to reject a match without losing either record.
Review decisions and evaluation
A review decision should record the sources consulted, the claim assessed and the reason for acceptance or rejection. Proposed matches remain distinct from researcher-confirmed identities. Where evidence is missing or conflicting, the question remains unresolved. Later evidence may require an earlier decision to be revised.
A future evaluation should measure false merges separately from missed matches, using a documented review set and a stated scope. This note proposes a method; it reports no measured accuracy or production entity-resolution capability.