The Ethnos Corpus – documentation

This document outlines the ethnos database and its pipeline: what the Ethnos data comprises, its origins, collection methods, repair processes, scoring, and omissions. It sits beneath ethnos.app and api.ethnos.app – distinct from the site and API, which are covered elsewhere.

It addresses two audiences with contrasting gaps:

the technically fluent reader unfamiliar with the data landscape in anthropology and the social sciences, and
the anthropologist or researcher who knows the literature but not the machinery.

Each chapter begins with the concept in plain language and continues with the mechanism. No chapter depends on the previous one, yet they follow the data’s flow.


The short version

There is no field-wide bibliographic database for anthropology. Not an open one, not a paid one, not a partial one. What exists are general-purpose indexes built around the journal article, and library catalogues built around the book, each holding a different fragment of the field and each confidently wrong about the parts it does not cover.

So this corpus is assembled rather than downloaded: source by source, with an explicit rule for what each source is allowed to say, a repair layer for the contradictions that result, and a scoring model that says how much evidence exists that a given record belongs to this field. It currently holds 7,698,445 works and 7,786,681 publications, distilled from 22,678,823 fetched source records – 297 GB of cached responses – in a 48 GB database.

It is the work of one person, on their own hardware, at their own cost.


Chapters

#FileWhat it covers
01The problem: there is no sourceWhy no field-wide source exists, what each available source can and cannot say, and the trust rule that follows from it
02The shape of the dataWorks vs. publications, the entity model, what is actually in the tables
03SourcesEvery source in use, what it is authoritative for, what it gets wrong, and one that was assessed and refused
04CollectionWorklists, API cost as a design constraint, caching, the source archive
05Cleaning and repairThe four recurring corruptions and the rule adopted against each
06Relevance: scoring and filteringTiers, the score, the four constraints, the classes – and why nothing is deleted
07What is missingThe measured gaps, and what each one is the boundary of
08Files and availabilityWhat an availability record is, what is not stored, and what is not verified
09OperationsThe daily run, its ordering constraints, reversibility, and scale

Reading the figures

Every count in this folder was read directly from the live database on 2026-08-28 and describes its state on that date. The corpus is re-collected, re-cleaned and re-scored daily, so the numbers move. The claims should not.

A count from this corpus means "as recorded here, from these sources, as of this date". It does not mean "as published in the world". The absence of a work is not evidence that it does not exist; it is evidence that no reachable source described it in a way this pipeline could ingest.