The Ethnos Corpus – documentation
This document outlines the ethnos database and its pipeline: what the Ethnos data comprises, its origins, collection methods, repair processes, scoring, and omissions. It sits beneath ethnos.app and api.ethnos.app – distinct from the site and API, which are covered elsewhere.
It addresses two audiences with contrasting gaps:
the technically fluent reader unfamiliar with the data landscape in anthropology and the social sciences, and
the anthropologist or researcher who knows the literature but not the machinery.
Each chapter begins with the concept in plain language and continues with the mechanism. No chapter depends on the previous one, yet they follow the data’s flow.
The short version
There is no field-wide bibliographic database for anthropology. Not an open one, not a paid one, not a partial one. What exists are general-purpose indexes built around the journal article, and library catalogues built around the book, each holding a different fragment of the field and each confidently wrong about the parts it does not cover.
So this corpus is assembled rather than downloaded: source by source, with an explicit rule for what each source is allowed to say, a repair layer for the contradictions that result, and a scoring model that says how much evidence exists that a given record belongs to this field. It currently holds 7,698,445 works and 7,786,681 publications, distilled from 22,678,823 fetched source records – 297 GB of cached responses – in a 48 GB database.
It is the work of one person, on their own hardware, at their own cost.
Chapters
| # | File | What it covers |
|---|---|---|
| 01 | The problem: there is no source | Why no field-wide source exists, what each available source can and cannot say, and the trust rule that follows from it |
| 02 | The shape of the data | Works vs. publications, the entity model, what is actually in the tables |
| 03 | Sources | Every source in use, what it is authoritative for, what it gets wrong, and one that was assessed and refused |
| 04 | Collection | Worklists, API cost as a design constraint, caching, the source archive |
| 05 | Cleaning and repair | The four recurring corruptions and the rule adopted against each |
| 06 | Relevance: scoring and filtering | Tiers, the score, the four constraints, the classes – and why nothing is deleted |
| 07 | What is missing | The measured gaps, and what each one is the boundary of |
| 08 | Files and availability | What an availability record is, what is not stored, and what is not verified |
| 09 | Operations | The daily run, its ordering constraints, reversibility, and scale |
Reading the figures
Every count in this folder was read directly from the live database on 2026-08-28 and describes its state on that date. The corpus is re-collected, re-cleaned and re-scored daily, so the numbers move. The claims should not.
A count from this corpus means "as recorded here, from these sources, as of this date". It does not mean "as published in the world". The absence of a work is not evidence that it does not exist; it is evidence that no reachable source described it in a way this pipeline could ingest.