Three governments use the same software to decide who you are. I asked each of them how often it gets that wrong. None of them could tell me.
In June I sent the same question, in two languages, to three national institutions on three continents. The question was simple. When your software decides that two records belong to the same person, how often is it wrong?
The Ministry of Justice in London answered on 15 July. The Australian Bureau of Statistics answered on 8 September. Germany’s Federal Statistical Office answered on 26 August. All three confirmed they use a probabilistic record-linkage library called Splink. All three declined, in different ways, to give a number.
The differences between the three refusals are the story.
Splink is an open-source Python library built by a team inside the UK Ministry of Justice. It implements a method statisticians have used since 1969, the Fellegi-Sunter model. Given two records with a name, a date of birth, an address and perhaps a sex, it computes a weight for each field, adds them up, and produces a probability that the records describe the same human being. Above a threshold, the records are joined. Below it, they are not.
It is good software. It is fast, well documented, and free. That is why it has spread. The Australian paper that recommended adopting it describes a spine build that took 73 days under the old method and 1.5 days under Splink. Germany’s statistical office used it for the test run of its 2031 register-based census. In Britain it runs live inside the justice system.
None of that is the problem. The problem is what comes after the threshold.
Every probabilistic linkage produces two kinds of error. A false match joins two different people into one record. A missed match leaves one person split into two. The rates of both are the single most important fact about any linkage system, because they tell you how many people are being misidentified. The method is designed to make those rates estimable. Splink ships evaluation tools for exactly this purpose.
I asked each institution for the rates. Here is what came back.
The MoJ’s own transparency record says Splink runs a real-time Core Person Record linking individuals across courts, prisons and probation, including the DELIUS, NOMIS, Common Platform, LIBRA, FamilyMan and CaseMan case-management systems. A pilot shares Police National Computer numbers with police on the strength of those links. Of the three systems in this piece it is the only one where a wrong answer has a direct operational consequence for a named individual.
My request (reference 260616044) asked for five things: the governance documents, the model configuration, the decision threshold, the observed error rates, and the point at which a human reviews a match.
The Department’s response opens with the sentence “We can confirm that the MoJ holds the information that you have requested.”
It then declined to release the Data Protection Impact Assessment under section 31 of the Freedom of Information Act, the exemption for the prevention and detection of crime. For the model configuration, the threshold and the error performance, it cited section 21, information already reasonably accessible, and gave the same GitHub link for all three.
The repository contains model files. Model files describe what the model assumes about names, dates and addresses. They do not and cannot contain an observed false-match rate, which is a measurement taken after the model runs. Either the Department holds that measurement somewhere other than the repository, in which case section 21 was the wrong exemption, or it holds no such measurement, in which case the correct answer was “not held,” and the opening sentence of the letter is wrong. I have asked for an internal review on that point. A response is due in mid-September.
The letter did volunteer three things I had not been told before. There is no clerical review when records are first matched. A record enters a review queue only if it has already been linked and later falls below a “fracture threshold of 18.” And, in the Department’s words, “there are no complaint procedures specific to Splink itself, as Splink does not directly make decisions about individuals.”
So the sequence is: the software joins two records without a person looking; the join is used in court and probation workflows; it may be shared with police; and if it is wrong, there is no route of complaint about the software, because the software is not considered to have decided anything.
The ABS Person Linkage Spine joins three national datasets: the Medicare Consumer Directory, the DOMINO Centrelink welfare file, and personal income tax records from the Australian Taxation Office. The June 2025 build covers 39.87 million people and was the first to use Splink. It sits under the Person Level Integrated Data Asset, PLIDA, which is the foundation for a large share of Australian social research.
My request (reference FOI2025-26_87) asked for the same things. The ABS released twelve documents, six in part, four in full, and refused two.
On error rates it was direct. “The ABS advises that false positive and false negative rates are not measured for the Person Linkage Spine.”
The released documents make that sentence hard to read charitably.
Document 8 is the internal paper that persuaded the Bureau’s Methodology Design Committee to endorse Splink in December 2024. It lists, as a principal reason to move from deterministic to probabilistic linking, that “having a probabilistic model allows us to straightforwardly get estimates of precision (the proportion of all links that are true links).” It goes on: estimated match probabilities “are particularly useful in the absence of ground truth data (this is the case for ABS linkages) because they enable the calculation of several key metrics (such as precision and recall).” It warns that “model specification and potentially some form of calibration of the estimated match probabilities is essential to avoid biased estimates.”
Document 3, a June 2024 presentation to the Data Integration Program Board, includes a timeline item for June to December 2024: “Develop quality measures from Splink outputs to inform future linkage quality assurance.”
Document 10 is the October 2025 methodology and quality assessment for the June 2025 Spine. It reports linkage rates, missingness and duplicate rates. It does not report a false-positive or false-negative estimate anywhere in its 24 pages.
It does report the thresholds. Stage 1 links require a match probability above 0.99. Stage 2 links, roughly 500,000 across the three datasets, are accepted at a match weight of 1.5 or higher. Under Splink’s standard definition a match weight of 1.5 corresponds to a posterior probability of about 74 per cent. If the model is even approximately calibrated, which the Bureau’s own paper says is essential, then somewhere in the region of one in four Stage 2 links would be expected to be wrong, and nobody is checking. The document says the threshold “was validated by reviewing links formed at that level.” A review of that kind produces a count of true and false links. No such count was released.
The governance minutes add tone. At the June 2024 oversight forum, one member observed that 98 per cent of Splink’s links matched the old method and asked what was in the other 2 per cent, whether it was “heavily biased in a certain area,” and whether “a bias correction [is] required.” Another member said he did not think the Bureau needed permission to change method and suggested it “could publicise a change in method without raising public alarm.” On the question of public engagement, the recorded recommendation was “informing rather than consulting.”
The two withheld documents are described in the decision as containing “detailed information concerning internal statistical quality assurance and operational assessment processes.” They were refused under section 47G of the FOI Act, the exemption for the business affairs of a person or organisation. The organisation whose business affairs the Bureau is protecting is the Bureau.
The Statistisches Bundesamt ran a Methodentest Bevölkerung, a population methods test, under the Register Census Testing Act, as preparation for the 2031 census, which is intended to be built from administrative registers rather than a questionnaire. My inquiry was logged as a request under the Informationsfreiheitsgesetz (reference A31-210-26).
The office confirmed that Splink was used. On thresholds, it explained that two cut-offs were set per search space, that there were several search spaces per data source, and that it was “nicht sinnvoll” to name any of them. On the number of records that could not be linked, it said the figure “differed greatly by source” and did not provide it. On error rates it said this: false links “can only be determined with the help of a reference set that reflects the underlying truth. They therefore cannot readily be quantified.” Unlinked records, it added, had been quantified, but by source, and no figure was given.
Two things stand out. First, an unlinked record is not a missed match. An unlinked record is one with no partner; a missed match is a true pair the software failed to join. The answer substitutes a coverage statistic for an error statistic. Second, the office did have a reference set. It manually reviewed selected cases, which is what a labelled sample is, and the same methods test included a Wohnsitzanalyse, a residence survey of up to 100,000 people, which is about as close to ground truth as a statistical office ever gets. It had the material to compute a rate. It reported none.
I also asked whether the Federal Data Protection Commissioner had been involved in the impact assessment. The answer was that the assessment was carried out “according to the guidelines” of the Commissioner. That is not the same thing, and I have asked the Commissioner directly.
Evaluation reports, the office says, are being written now, with first publication expected in 2027. The refusals in the decision do not cite any of the statutory exemptions the IFG requires. I have lodged an objection.
Put the three responses in a row.
London says the number is derivable from the published model.
Canberra says the number is not measured.
Wiesbaden says the number cannot be measured.
The London position implies it is a matter of reading the files. The Canberra position, backed by the Bureau’s own endorsement paper, is that the tool produces the number and the Bureau simply has not produced it. The Wiesbaden position is that it is impossible without a reference set the office in fact possesses. The organisation that built the software published evaluation functions for precisely this measurement. The organisations that adopted the software adopted it partly because of that capability.
What none of them adopted is the discipline of using it. And the one that built it, the one whose links go to police in real time, is the one that withheld its impact assessment and told me the software makes no decisions.
It does not show that any specific individual has been harmed by a mislink in any of the three systems. I have no such case, and the documents contain none.
It does not show that the software is defective. The evidence points the other way: Splink is competent, the ABS found it more accurate than its predecessor on manual review, and the errors here are institutional, not technical.
It does not show the actual error rates. That is the point. Nobody has produced them, and for three institutions with a combined reach of well over a hundred million people, that is the finding.
An internal review is pending at the Ministry of Justice. An internal review has been lodged with the Australian Bureau of Statistics on the withheld quality-assurance documents, the missing methodology paper referenced in its own minutes, and the validation review its own report describes. An objection has been lodged with the Statistisches Bundesamt.
On 9 September 2026 the three responses and their attachments were sent to the following:
House of Commons committees: Justice; Home Affairs; Science, Innovation and Technology; Public Administration and Constitutional Affairs.
House of Lords committees: Justice and Home Affairs; Public Services; Constitution; Communications and Digital.
Joint committee: Human Rights.
Regulators and the Department: the Office for Statistics Regulation; the Ministry of Justice Data Protection Officer.
Civil society: Big Brother Watch; Open Rights Group.
Press: The Guardian; Computer Weekly; The Bureau of Investigative Journalism.
If any of the three institutions produces an error rate, I will publish it here.
The full responses from all three institutions are published below. This piece will be updated as the reviews conclude.
These are the responses themselves, as received. Nothing in the substance has been altered: the only change is that my own email address has been removed from the letterheads. The officials’ published contact details, which appear on the institutions’ own letterhead, are left as they were.
Read them against the piece above. The sentences quoted here are short, and the surrounding paragraphs are where an institution’s reasoning either holds up or does not.