About Weft

Weft is a library of interlinear editions. Each page sets a passage of a classical text in its original form and, under every word, two further rows: how the word most likely sounded when the text was first written, and a literal word-for-word English gloss. Under each line come one or more published translations. Notes attach to words and lines. The library runs from the Pyramid Texts of Unas (about 2350 BC) to the General Act of the Berlin Conference (1885), in twenty-four languages.

It is dedicated to Neith, goddess of the loom, and to the author's mother. The weft is the thread carried back and forth across the warp; the source text is the warp, and sound, gloss and sense are the weft.

Reading a page

Every page is one self-contained file. It opens from disk without a server and prints cleanly; the fonts load from Google Fonts when you are online, and some scripts need them to display.

Why the editions are built like software

A printed interlinear fixes one editor's judgments in one state. Weft keeps every layer of a page as a plain text file in a public repository, under version control. Four consequences follow.

  1. Every judgment is attributed. A correction is a small entry that names who made it, when, and why. The page shows where each value came from.
  2. Machine work and human work are kept apart. The pipeline generates a first layer (the text, lemma and grammar from a treebank where one exists, the sound from rules) into gen/. Human corrections go in a separate overlay, curated/, which wins where the two differ. The generated layer can be rebuilt at any time without losing a single human decision.
  3. Every change is reviewable and reversible. A correction arrives as a proposal, is discussed in the open, and can be undone. The history of each reading is kept.
  4. Rules improve the whole library at once. When a pronunciation rule is corrected, the works in that language are regenerated from it, and the change in each generated file can be inspected before it is accepted. A specialist's correction to a rule reaches every line it governs.

What the sound row claims

The first pronunciation scheme on every page is the reconstruction Weft judges best supported of how the text most likely sounded when it was first written: Dante in the Florentine of about 1307, Luther's theses in Latin as it was read in Saxony, the Rigveda with its pitch accent. Where the evidence for a text's own period is thin, the nearest well-studied stage stands in, and the page says so: Homer is read in the reconstructed classical Attic of the fifth century BC, about three centuries after the poems took shape. Later traditions (the Erasmian Greek of schools, the ecclesiastical Latin of the church, modern Icelandic) are offered second, and the reader can switch between them.

The evidence differs greatly by language. For classical Latin and Greek it is strong and well studied. Egyptian writing records no vowels at all, so any vocalized reading is a conjecture; Sumerian's sounds are known only indirectly, through Akkadian scribal usage, and its stress and vowel length are unknown. Pages whose first scheme rests on indirect evidence say so at the top and invite specialists to correct it. Languages and pronunciation sets out, language by language, what each scheme rests on and where it is weakest.

The respelling approximates the reconstruction for an English reader. It is not a claim that the original speakers sounded like English speakers.

Draft and reviewed

Much of the library is draft. Where no treebank exists (Eddic poetry, the Bayeux captions, Hammurabi), the lemma, grammar and gloss were annotated by hand against a named dictionary, and the notes were written for this edition. All of it is marked draft on the page. Only a human reviewer lifts that status, by checking the material and recording the review under their name.

Generated material is labelled by what produced it. Material produced by rule or by a treebank is not reviewed by default either; where a value is doubtful, the corrections overlay is where the better reading goes.

Sources and licences

How to take part

You do not need to use git. Each kind of contribution has a short form on GitHub that asks the right questions; a maintainer turns the answer into a correction with your name on it.

Each page links to the correction, pronunciation and recording forms from its foot and, where the pronunciation is reconstructed, from its top. Contributors who do use git can submit the correction directly; Contributing describes both routes.

Further reading