Is this already public?
Reconciliation against the standard radiocarbon compilations
A compilation that merely restates public data is not worth much. So: are these determinations already in the aggregators everyone uses?
Checked against calpal_2020_08_20.tsv, nerd.csv, p3k14c_scrubbed_fuzzed.csv, radonb_radondaily.txt, xronos_data.csv — 217,251 distinct laboratory codes in total.
| Status | Determinations | Share |
|---|---|---|
| Already in the public compilations | 6 | 0% |
| Not found upstream | 1338 | 90% |
| No usable lab code — uncheckable either way | 147 | 10% |
The answer
Almost none of this register is in the public compilations.
That result was suspicious enough to be worth attacking before believing, since it flatters the project. Two checks:
The join demonstrably works. It finds matches when matches exist — three
Oxford determinations resolved correctly across databases that write the same
laboratory as OxA, OXA and oxa.
The laboratory numbers interleave. This is the strong evidence. Take
Mannheim: this register holds MAMS codes numbered 11649–55552; the
aggregators hold MAMS codes numbered 7217–111837. The ranges sit inside one
another. The same laboratory measured both sets, in the same period — and the
two sets share not one determination.
The same pattern holds for Belfast, Novosibirsk, Penn State and Moscow. These are not dates the aggregators have under a different spelling. They are dates the aggregators do not have.
Why the gap exists
The aggregators are built from English-language literature and from archaeogenetics papers. Independent survey of their Eurasian coverage found roughly 150–200 Russian and Central Asian steppe Bronze Age determinations across every one of them combined, with Kazakhstan close to absent — dozens of rows, not thousands. A large share trace to ancient-DNA publications rather than to site excavation reports.
The determinations here come from Russian- and Kazakh-language journals that the aggregators do not ingest.
What this does and does not license
It licenses saying the extraction contributes determinations, not only reservoir flags.
It does not license calling the number exact. Matching is by laboratory code alone, so a determination republished under a different code cannot be found by this method — treat “not upstream” as an upper bound on novelty. And 371 of the unscreenable determinations are absent upstream, which means the reservoir problem this project documents is, for the most part, not visible in the compilations at all.
A further tenth of the register carries no usable laboratory code, because the source published none. Those cannot be checked against anything, by anyone, ever — which is its own small finding about the record.