02 – The shape of the data

← The problem · Index · next → Sources


One distinction governs everything else in this database, and it is worth getting straight before anything else.

A work is not a publication

A work is the intellectual thing: an argument, a study, a book, written once.

A publication is a particular issued version of it: this printing, in this journal issue, with this DOI, at this page range, in this binding.

One work can have many publications. A book issued in cloth, paperback and PDF is one work and three publications. An article that later appears as a chapter in an edited volume is one work and two publications. A translated article registered separately in each language is one work and several publications.

In plain words. Think of a library catalogue card versus the copies on the shelf. The card is the work. Each edition, printing and format is a publication. This database keeps both, and keeps them linked, because most confusion in bibliographic data comes from systems that keep only one of the two.

This is not pedantry. It is what makes it possible to count a book's citations once instead of three times, to let a reader who found one edition see that the others exist, and to record that two records describing "the same thing" really are the same thing without destroying the differences between them.

The database is type-agnostic at the work level: a work has a title, a subtitle, an abstract and a language, and nothing else. Whether the thing is an article, a book or a chapter is a property of the publication, because it is the issued object that has a carrier.


The entity model

                        work_references
                     (97,343,498 citation links)
                              ↑   ↓
    ┌───────────┐        ┌─────────────┐        ┌────────────┐
    │  PERSONS  │───────▶│    WORKS    │◀───────│  SUBJECTS  │
    │ 4,902,477 │credited│  7,698,445  │ tagged │  217,828   │
    │           │   on   │             │  with  │            │
    └───────────┘        └──────┬──────┘        └────────────┘
     via authorships            │ 1 → n           via work_subjects
     (16,870,149, with role)    ▼                 (85,332,494 links)
    ┌───────────┐        ┌─────────────┐        ┌────────────┐
    │  VENUES   │◀───────│PUBLICATIONS │───────▶│   FILES    │
    │  339,662  │appears │  7,786,681  │ avail- │ 7,173,322  │
    │           │   in   │             │ able as│            │
    └───────────┘        └──────┬──────┘        └────────────┘
                                │ published / funded by
                                ▼
                       ┌──────────────────┐
                       │  ORGANIZATIONS   │
                       │    1,099,971     │
                       └──────────────────┘
TableRowsWhat it holds
works7,698,445Title, subtitle, abstract, language. No type.
publications7,786,681The issued version: carrier type, DOI, ISBN, dates, venue, pages, licence, source tag
persons4,902,477Contributors. 1,689,443 (34.5%) carry an ORCID
authorships16,870,149Who is credited on what, with a role – author, editor, translator, reviewer
venues339,662Journals, series, repositories – and container books (see below)
subjects217,828Topical vocabulary across six independent schemes
work_subjects85,332,494Which subject is attached to which work
work_references97,343,498Citation edges. 58,233,380 (59.8%) resolve to a work held here
organizations1,099,971855,892 institutes, 219,068 funders, 21,314 publishers, 3,697 universities
funding1,092,290Grant links between works and funder organizations
files7,173,322Availability records – see chapter 08

The database occupies 48.0 GB across 30 tables – 21.4 GB of data and 26.6 GB of indexes. That ratio is deliberate: nearly every access path in the pipeline and the API is an indexed lookup, because a full scan over a 97-million-row table is not a thing one person's hardware can afford to do casually.


The strange one: a book is its own venue

An article's venue is its journal. A chapter's venue is the book it sits in. So the database models a container book as a venue in its own right, typed SOURCE_BOOK.

This is how a chapter reaches its container's publisher, language, subject headings and identity without those being stamped onto the chapter itself. A chapter is never given the book's identifiers – that would make the chapter look like the book – so the shared venue is the only route it has to its container.

It is also why the venue table is dominated not by journals but by books:

Venue typeCountWhat it is
SOURCE_BOOK313,860A container book – the venue of its own chapters
JOURNAL24,262A serial
REPOSITORY653A preprint server or archive
BOOK_SERIES537A monograph series
CONFERENCE242Proceedings
OTHER108Identity unresolved

Container books are keyed in layers – by a work-level book identifier where one is known, by ISBN-13 otherwise, and by name only as a last resort – so that the different editions of one book consolidate onto one container instead of fragmenting. Getting that consolidation right is the single largest cleaning problem in the project, and it is documented in chapter 05.


What kind of things the corpus holds

CarrierPublicationsNote
Article7,119,575Journal articles – the best-described part of the corpus
Book394,030Monographs, edited volumes, reference works
Chapter241,116Contributions inside a container book
Thesis11,059Dissertations, where a registry carries them
Review4,584Book reviews – assigned by repair, never by a source (ch. 05)
Report2,767
Preprint1,547
Other12,003Datasets, conference papers, unclassifiable carriers

The carrier vocabulary deliberately collapses genre distinctions the sources make and the database does not need: a monograph is a book, a reference entry is a chapter, a posted preprint is a preprint. What matters downstream is the physical shape of the thing, because that is what decides how it is identified, where its metadata lives, and which container it belongs to.


Coverage in time and language

Publications by decade:

DecadePublicationsDecadePublications
1900s21,9601970s377,229
1910s29,5731980s522,079
1920s49,8091990s744,471
1930s83,7922000s1,140,202
1940s93,9522010s1,917,013
1950s154,7842020s2,373,317
1960s244,704

The oldest record is dated 1400. The steep recent rise is partly real growth in publishing and partly a coverage effect: registries describe recent decades far better than old ones. That is a finding about the sources, not about the discipline – and it is a reason to be careful with any time-series drawn from this corpus.

Language coverage reflects both the field and the sources' own bias:

English6,422,531Italian25,311
French449,909Turkish18,817
Portuguese289,838Polish10,370
Spanish260,781Catalan8,882
German183,550Dutch7,319

Every work carries a language code. Where no source stated one, it is inferred from the text itself – one of the few places the pipeline derives a fact rather than recording one.


next → 03 – Sources