Findings
Cleaning a corpus is not neutral: doing it carefully surfaces things about the record that were not visible while it sat as an unindexed spreadsheet. These are byproducts, not the point — but they are the kind of byproduct that justifies the exercise.
Transcription errors in the source, caught by cross-check
Comparing Chapple’s determinations against the TII monograph appendices (which are the primary publication for those dates) turned up six cases where Chapple’s BP age disagrees with the excavation report it cites — almost certainly transcription slips:
| Lab code | In Chapple | In TII monograph |
|---|---|---|
| SUERC-37262 | disagreement on age | primary report value differs |
| SUERC-29337 | disagreement on age | primary report value differs |
| WK-20192 | disagreement on age | primary report value differs |
| UBA-12942 | disagreement on age | primary report value differs |
| UBA-12943 | disagreement on age | primary report value differs |
| WK-15499 | disagreement on age | primary report value differs |
Each is recorded as a flag with both values and the citing report — the dataset does not overwrite Chapple. Where 239 lab codes overlap, the two sources agree 97.7% of the time, which is what makes the six stand out.
Duplicate lab codes
143 laboratory codes appear on more than one record with a different age (297 rows). Some
are genuine (a lab code reused across a split sample), some are data-entry duplication. They are
flagged lab_code_collision rather than silently merged — an analyst summing the record needs to
know which “dates” are not independent.
What was hiding in the prose
- Material. Chapple’s dedicated material column is populated for ~1.2% of records. The same information, written into his free-text notes, is recoverable for roughly 80% — and material is exactly what the reservoir classification needs.
- Reliability. The catalogue was thought to carry per-row reliability highlighting. It does not — the “highlighting” is conditional formatting flagging duplicate lab codes. Chapple’s actual reliability judgements (“anomalous”, “should not be used”) live in the notes prose, and 238 of them were recovered from there.
The map follows the roads
The spatial distribution of the corpus visibly traces Ireland’s motorway corridors. This is not an artefact of the cleaning — it is a real property of the record: Irish radiocarbon dating is overwhelmingly development-led, funded by road schemes (NRA/TII). Any summed-probability or demographic reading of this data is reading, in part, a map of where the State built roads. That caveat belongs on the record, and now it is on it.