Languages and pronunciation

This page sets out, for each language in the library, what the sound row on the page claims, what the code says it rests on, and where it is weakest. It is written for specialists deciding whether to trust the pronunciation row for their language and how to correct it. Every statement here is taken from the module in pipeline/weft/, its data files, or the manifests and notes of the works that use it. Where the repository names no source, this page says so.

What the sound row is

How to propose a change

A pronunciation change is made only with a citation: a grammar, an article or a corpus.

  1. The Pronunciation issue form (.github/ISSUE_TEMPLATE/pronunciation.yml), linked from each page with the work and scheme filled in. It asks for the scope (one word, a rule, the stress or accent rule, or a new scheme alongside the existing ones), the words affected, what the sound should be (IPA if possible), the citation, and how you wish to be credited. No git is needed.
  2. A pull request against pipeline/weft/<module>.py and its data files under pipeline/weft/data/. For languages where the reading is entered per word (Egyptian, Akkadian, Sumerian, Vedic, Tamil, Persian, the runes, Japanese), the per-word form lives in the work's edition file, texts/<work>/edition.yaml. After a rule change, run weft acquire and weft draft for each work in that language, then weft check <work> and uv run --with pytest pytest -q tests before submitting. weft say <Greek words> prints Greek words in the restored and Erasmian schemes; other languages have no command-line phonemizer yet.

A different reconstruction can be added as a new scheme rather than replacing the current one; schemes sit side by side and the reader chooses.

Summary

"First" marks the page default. Confidence is the manifest's sound_confidence; "not set" means the page shows no invitation.

LanguageModuleWorksSchemesConfidence
Old EgyptianegyptianPyramid Texts of Unasold-egyptian (first), egyptologicallow
SumeriansumerianEnheduanna, Temple Hymn 5recon-2300 (first), classroomlow
AkkadianakkadianHammurabi, prologueob-1750 (first), classroomlow
Vedic SanskritsanskritRigveda 1.1vedic (first), modernmedium
Ancient GreekgreekIliad, Odyssey, Aristotlerestored (first), erasmiannot set
John, Beatitudes, 1 Corinthians 13koine (first), erasmiannot set
Epictetus, Marcus Aureliuskoine (first), restored, erasmiannot set
Biblical HebrewhebrewGenesistiberian (first), modern-israelinot set
Classical ChinesechineseSunzi, Daodejingold-chinese (first), tang, mandarinmedium
Li Baitang (first), mandarin, cantonesemedium
Sahidic CopticcopticGospel of Mark 1sahidic (first), bohairicmedium
LatinlatinOvid, Res Gestaeclassical (first), ecclesiasticalnot set
Bayeux Tapestryanglo-norman (first), classicalmedium
Magna Cartaanglo-latin (first), classicalnot set
Saer de Quincy, two chartersanglo-latin (first), classicalmedium
Picoitalian-humanist (first), classicalnot set
Inter caeteraitalian-humanist (first), classicalmedium
Erasmuslow-countries (first), classicalnot set
More, Utopiatudor-english (first), classicalnot set
Luther, Ninety-five Thesesgerman-humanist (first), ecclesiastical, classicalnot set
Descartes and Newtonas-first-read (first), classicalnot set
Old TamiltamilTirukkuralold-tamil (first), modernmedium
RunicrunicKylver, Gallehus, Rökas-carved (only)low
Old EnglisholdenglishBeowulfwest-saxon (only)not set
Old NorsenorseVöluspá, Hávamál, Þrymskviða, Grettis sagaold-norse (first), modern-icelandicnot set
Snorri, Gylfaginningold-norse (first), modern-icelandicmedium
Old East SlavicoldeastslavicPrimary Chronicle, 859-862orv1100 (first), rulow
PersianpersianRubaiyatearly (first), modernmedium
Middle MongolianmongolianSecret History of the Mongols, 1-10mm-1250 (only)low
JapanesejapaneseBashōedo-1686 (first), modernnot set
Franco-ItalianoldfrenchMarco Polo, Cipangufr1300 (first), it1300medium
ItalianitalianDante, Petrarch, Machiavelliflorentine (first), modernnot set
SpanishspanishColumbus, 1493c1492 (first), modernnot set
DutchdutchLinschotenh1596 (first), modernnot set
FrenchfrenchMontaignem1580 (first), modernnot set
Berlin Act, 1885fr1885 (first), modernnot set
GermangermanKant, Nietzschenorthern (first), modernnot set
Luther's Bibleecg1545 (first), modernnot set
SwahiliswahiliSteere, The Kites and the Crowsz1870 (first), modernmedium
Quenya, Sindarinelvishprivate works only (Tolkien, in copyright)tolkien (only)not set

The registry that maps a manifest's language code to a module is PHON in pipeline/weft/draft.py.

Old Egyptian

Schemes.

What it rests on. The consonant values follow "the outline of Loprieno (1995) and Allen (2013)" (module docstring; the notes cite Loprieno, Ancient Egyptian: A Linguistic Introduction, chapter 3, and Allen, The Ancient Egyptian Language: An Historical Study). The ejective reading of d and ḏ is attributed to Loprieno. Each of the fourteen vocalized forms names its evidence in the vocalization map of texts/pyramid-texts-unas/edition.yaml (for example Coptic ⲛⲟⲩⲧⲉ for nā́ṯir, Greek Ὄννος for the king's name), and the page shows it in the word's provenance. The module states that these are Weft's inferences by the sound correspondences in those grammars, not forms quoted from them.

Where it is weakest. The module calls this "the most reconstructed sound row in the library" and "conjectural word by word". Only fourteen forms are vocalized; every other word falls back to the Egyptological convention inside the first scheme, in lower case without stress, and is flagged per token. Each consonant value in the list above is marked as debated; the key notes that ꜣ may have weakened to a glottal stop or y by the Middle Kingdom and that some scholars take ꜥ as a d-like stop in the earliest texts. Several vocalization entries call individual vowels uncertain or a guess.

What a specialist could improve. Add or correct entries in the vocalization map and the n field of the tokens, with the evidence named; revise consonant values in egyptian.py; propose an alternative consonant system as a second reconstruction.

Sumerian

Schemes.

What it rests on. Jagersma, A Descriptive Grammar of Sumerian (2010), "in broad agreement with Edzard, Sumerian Grammar (2003)" (module docstring). The notes add Foxvog, Introduction to Sumerian Grammar, on pronunciation. Lemma and morphology come from the Oracc lemmatization (epsd2/literary); the cuneiform row from the Oracc Sign List value table.

Where it is weakest. The module describes the language as an isolate whose sound "is known only indirectly" and the row as "among the most reconstructed" in the library. Stress and vowel length are unknown and not shown. The phoneme conventionally written dr or ř is not applied to any word (manifest and module). One broken sign has no sound.

What a specialist could improve. The rules in sumerian.py (stop series, the treatment of written doubling, whether and where to apply ř); per-token readings in the optional n field of texts/enheduanna-temple-hymns/edition.yaml where the spelling misleads (the module's example is a-a for aya).

Akkadian

Schemes.

What it rests on. The sibilant view is attributed to "Streck and Kogan among others" (module); the notes cite M. P. Streck, "Sibilants in the Old Babylonian texts of Hammurapi and of the governors in Qaṭṭunān" (2006). Stress in both schemes follows the rule in Huehnergard, A Grammar of Akkadian (3rd edition, 2011). The sound is computed from a hand normalization (n) in the edition; the cuneiform row is generated from a modern sign reading (sg) through the Oracc Sign List (pipeline/weft/data/akk_osl_values.tsv, CC0).

Where it is weakest. The manifest: the values of š, s, z and the emphatics are debated, and stress is inferred by a modern rule. The key says the exact sound of the emphatics is not known. The transliteration is Harper's of 1904, with older conventions; the sound does not depend on it, since it reads n.

What a specialist could improve. The normalizations in texts/akkadian-hammurabi/edition.yaml (vowel length and contraction drive both sound and stress); the consonant values and the stress rule in akkadian.py; a second reconstruction that keeps š as sh, which the notes name as the competing view.

Vedic Sanskrit

Schemes.

What it rests on. The ancient phonetic treatises (prātiśākhyas) and Pāṇini's statement on short a (module). The notes cite W. S. Allen, Phonetics in Ancient India (1953), and Macdonell, A Vedic Grammar for Students (1916), appendix III; and Arnold, Vedic Metre (1905), for the metre. The accent is hand-entered in the IAST form n; decode_marks reads the Devanagari accent strokes back into raised syllables so the two can be checked against each other (tested in tests/test_golden.py).

Where it is weakest. The manifest: the treatises are several centuries later than the hymns. Some pādas count seven syllables as written where recitation restored a lost syllable; the metre row shows the written count. The treebank's words are unaccented, so the accent comes from the edition alone.

What a specialist could improve. Vowel and consonant values in sanskrit.py; the accented forms in texts/rigveda-1-1/edition.yaml; a rule for the restored syllables, which the metre row does not yet count.

Ancient Greek

Schemes.

Homer and Aristotle open in restored; the New Testament books, Epictetus and Marcus Aurelius open in koine. There is no separate Homeric or archaic scheme. The Iliad and Odyssey manifests relabel restored for their pages: "classical Attic of the 5th century BC with its pitch accent; for Homer, the nearest well-studied stage, not the poet's own (approximate)".

What it rests on. W. S. Allen, Vox Graeca, for restored (module); Randall Buth for koine (module label and the New Testament manifests); the Epictetus notes cite Horrocks, Greek: A History of the Language and its Speakers (2nd edition, 2010), chapter 5. The length of α, ι, υ, which the spelling does not show, comes from quantity tables: grc_quantities.yaml (shared), grc_quantities_iliad.yaml and grc_quantities_marcus.yaml, each entry from the LSJ headword and, for verse, the foot where the hexameter shows it.

The hexameter scanner. greek.scan_hexameter (version 0.1) scores every arrangement of five dactyls or spondees plus a final foot, with a cost for each rule it leans on (epic correption, muta cum liquida, metrical lengthening, an unlisted long α ι υ, synizesis, neglected digamma), and reports each costed choice as a failure class. test_hexameter_scanner_matches_hand_scansion in tests/test_golden.py requires its output to equal every hand scansion in the curated overlays: 62 lines (Odyssey 1.1-10 and Iliad 1.1-52). Not modelled: digamma beyond a short list of stems, lengthening before initial liquids except as a costed option, and synizesis beyond word-final -εω.

Where it is weakest.

What a specialist could improve. Entries in the quantity tables; the digamma stem list and the costs in the scanner; the Koine vowel and consonant tables in greek.py; a scheme for the language of the Homeric poems, which the repository does not yet have.

Biblical Hebrew

Schemes.

What it rests on. Geoffrey Khan, The Tiberian Pronunciation Tradition of Biblical Hebrew (2020) (module). The pointing gives the vowels, dagesh gives gemination and hard stops, and the cantillation accent gives the stressed syllable, so no quantity table is used. Shewa and begadkefat follow "the standard grammars" (module; no grammar is named).

Where it is weakest. The manifest: the vowels are the Masoretes' of about AD 900, so "as first written" reaches the pointing, not the period of composition. Vocal against silent shewa and qamets gadol against qatan are guessed where the pointing is ambiguous. The verse-final silluq is recovered from the sof pasuq.

What a specialist could improve. The shewa and qamets heuristics in hebrew.py, with a named grammar; an earlier reconstruction as a scheme placed before tiberian, if one can be stated per word.

Classical Chinese

Schemes.

What it rests on. Old Chinese readings are taken per character from Wiktionary's data modules, each pinned to a revision id, into pipeline/weft/data/lzh_oc_bs.yaml (CC BY-SA 4.0). A character Baxter-Sagart lacks falls back to Zhengzhang Shangfang (2003) and says so. Tang readings are Unihan kTang, which follows Hugh M. Stimson, T'ang Poetic Vocabulary (1976). A character with no reading of its own borrows one from a glyph variant (kZVariant, then kSemanticVariant) and says so.

Where it is weakest.

What a specialist could improve. Entries in lzh_oc_bs.yaml (a reading, or a better fallback for the four characters outside Baxter-Sagart); Tang readings for the missing characters, which need a source other than Unihan; Mandarin choices for polyphonic characters.

Latin

All Latin schemes keep the classical penultimate stress rule (except the French method, below), so every word needs its vowel length from a quantity table: pipeline/weft/data/lat_quantities.yaml, plus per-work tables for the Bayeux Tapestry, Inter caetera, the Ninety-five Theses, the Res Gestae and the Saer de Quincy charters. Entries follow Lewis and Short headword quantities plus inflectional endings, confirmed against the hexameter for Ovid. A word missing from the table is reported as quantity-unknown.

Schemes.

IdLabelUsed byWhat it represents
classicalClassical: restored pronunciation of Cicero's and Ovid's Romefirst for Ovid and the Res Gestae; third for Luther; second elsewhereafter W. S. Allen, Vox Latina: c and g hard, v as w, ae as ai, length audible
ecclesiasticalEcclesiastical: Italianate church Latinsecond for Ovid and the Res Gestae; second of three for Luthersoft c and g before front vowels, v as v, ae as e, length heard only in stress
anglo-normanAs first read: Latin in Normandy and Norman England around 1070 (approximate)Bayeux Tapestryc before front vowels ts, g and j dʒ, h silent, u as French u, s between vowels z
anglo-latinAnglo-Latin: as a clerk in England read Latin around 1215 (approximate)Magna Carta; the Saer de Quincy charterssoft c ts, soft g and j dʒ, h silent, v as v, s between vowels z
italian-humanistAs first read: Latin in northern Italy in the 1480s (approximate)Pico; Inter caeteraItalian vowels, soft c ch, gn ny, sc sh, ti ts, h silent, s between vowels z
low-countriesAs first read: Latin in the Low Countries around 1500 (approximate)ErasmusDutch vowel values long in open syllables, u as Dutch uu, g a fricative, ch kh, ti ts
tudor-englishAs first read: Latin in England around 1516 (approximate)MoreEnglish long values in stressed open syllables at an earlier stage of the Great Vowel Shift, ti as si
german-humanistAs first read: Latin in Saxony around 1517 (approximate)Luther, Ninety-five ThesesGerman lengthening rule, c before front vowels ts, g hard, qu kv, v as f, final devoicing
as-first-readAs first read: Newton in the English manner, Descartes in the French (approximate)Descartes and Newtonper section dialect: english (1680s English method) or french (1640s French method, final stress, nasal vowels)

Manifests relabel some of these for their page (Erasmus, Pico, Inter caetera, More, Luther, the Res Gestae).

What it rests on.

Where it is weakest.

What a specialist could improve. Quantity-table entries, especially hidden quantities and medieval words keyed by their classical stems; the rule functions in latin.py (_humanist, _german, _norman, _national) for a named period grammar; an eu diphthong; a separate southern Italian or Roman reading for the papal chancery.

Sahidic Coptic

Schemes.

What it rests on. Peust, Egyptian Phonology (1999), for the letter values; Layton for the reading of a doubled vowel; the letter table published by the Coptic Orthodox Diocese of the Southern United States for bohairic (module). The accent of Greek loanwords and the list of unstressed particles are in pipeline/weft/data/cop_stress.yaml. Lemma and morphology come from UD_Coptic-Scriptorium (CC BY 4.0), which annotates a different digital text of Mark; spelling differences are normalized before the two are compared, and what remains is reported as edition-differs-from-treebank (manifest). Glosses are hand-written against the treebank's lemmas, with Crum's Coptic Dictionary (1939) and Lambdin's Introduction to Sahidic Coptic (1983) as references (manifest).

Where it is weakest. The edition prints no supralinear strokes and no punctuation, so the syllabic vowel is supplied by rule rather than read from the manuscript (manifest). That Greek loanwords keep the Greek accent is an assumption (manifest). The value of ⲩ standing alone in Greek words may already have shifted (key). Each word is phonemized alone.

What a specialist could improve. The stress table and the particle list in cop_stress.yaml; the syllabic-consonant rule and the Greek-letter values in coptic.py; the normalization list used to compare the edition with the treebank's text.

Old Tamil

Schemes.

No syllable is capitalized: the module states that Tamil stress does not distinguish meaning.

What it rests on. "The description of sounds in the Tolkāppiyam and the usual reconstructions of Old Tamil" (module). No modern reconstruction is named. The sound is computed from a hand ISO 15919 transliteration, checked against the script by tamil_to_iso.

Where it is weakest. The date of the text is uncertain (4th to 6th century AD, with earlier estimates), so the reading is "an approximation of an approximate period" (manifest). The values of c, ṟ, ṟṟ, ḻ and the āytam are the least certain. Sandhi across words is written but not modelled.

What a specialist could improve. Consonant allophony and the āytam in tamil.py, with a named reconstruction; the transliterations in texts/tamil-tirukkural/edition.yaml.

Runic inscriptions (Proto-Norse and Old East Norse)

Scheme. as-carved, "As carved: Proto-Norse around 400, Old East Norse around 800". One scheme; each token's dialect (pn or oen) selects the rules. The Proto-Norse z is a buzzing sound; Old Norse ʀ is "between z and r"; v and w are both w; stress always on the first syllable. There is no second scheme.

What it rests on. The module does not cite a source. Sound is derived from each word's scholarly normalization in texts/runes/edition.yaml, never from the runes, because the younger futhark has 16 runes for about 30 sounds.

Where it is weakest. Runes do not mark vowel length or many consonant contrasts; the sound inherits every choice in the normalization (manifest). The Gallehus runes survive only in 18th-century drawings.

What a specialist could improve. The normalizations, with a named edition; the values of z and ʀ and of the fricatives in runic.py; a citation for the scheme as a whole.

Old English

Scheme. west-saxon, "Late West Saxon, around the year 1000". One scheme. Stress on the first syllable of the root; ge-, be- and, on verbs, a-, for-, of-, on-, to-, un-, ymb- unstressed. Palatal c and g, sc, cg, voicing of f s þ between voiced sounds, and h by position.

What it rests on. Consonant rules after Mitchell and Robinson, A Guide to Old English (module). Vowel length comes from the edition's marks (Harrison and Sharp), corrected by pipeline/weft/data/ang_quantities.yaml (values after Clark Hall, A Concise Anglo-Saxon Dictionary), then from the hand-annotated lemma, whose macrons carry over along a shared prefix. ang_prefixes.yaml lists prefixes the rules cannot see without the part of speech.

Where it is weakest. Palatal c and g are inferred from spelling; an edition that dots them would remove the guess (manifest). Forms that diverge from their lemma before the last macron are reported quantity-uncertain. Compound stress relies on the editor's hyphens.

What a specialist could improve. The quantity table (three entries at present) and the prefix list; the palatalization rules in oldenglish.py; a second scheme, since none is offered.

Old Norse

Schemes.

What it rests on. E. V. Gordon, An Introduction to Old Norse, and Haugen (module; no Haugen title is given). The normalized spelling marks length, so no quantity table is used.

Where it is weakest. The editions' ö covers both ǫ and ø and is read as ǫ by default (Eddic and Snorri manifests). For Snorri, ǫ and ø were merging about 1220 and long vowels may have begun to change. Grettis saga is read from IcePaHC's modern Icelandic spelling, so its Old Norse reading is reconstructed from modern forms; its notes date the saga to about 1310, later than the scheme.

What a specialist could improve. A per-word distinction of ǫ and ø; rules in norse.py for a later date (about 1300) as a separate scheme; old forms for Grettis saga.

Old East Slavic

Schemes.

What it rests on. The handbooks of Russian historical phonology, which rest on the spelling of dated manuscripts and birchbark letters: Shakhmatov; Kuznetsov; Shevelov, A Prehistory of Slavic (1964); Schenker, The Dawn of Slavic (1995) (module). The text is Weft's transcription of the Laurentian copy of 1377 as printed by Karsky in Полное собрание русских летописей, vol. 1 (1926), checked against the OCR of the scan with declared corrections (manifest). No treebank is used: TOROT covers the chronicle but is licensed NC (manifest).

Where it is weakest. The jers: their loss in weak position is dated to the 12th century, so at 1100 the weak ones were fading, and the scheme treats every written jer alike; the module calls this its weakest point. The value of ѣ varied by region, and г in Kiev was probably a fricative, which is not shown. The manuscript does not mark stress; it is hand-entered in n from the stress of the same words in Russian and Ukrainian and from the accent paradigms of the handbooks (module and manifest). Softening before е, ѣ and и is assumed and not marked. The years, written in letter numerals, are read as numbers; the number words are not reconstructed (manifest).

What a specialist could improve. The stress marks in texts/primary-chronicle-varangians/edition.yaml; the jer rule and a southern variant with fricative г in oldeastslavic.py; the regional value of ѣ.

Persian

Schemes.

What it rests on. The grammarians, early manuscripts, and Dari and Tajik (module). The notes cite Lazard, La langue des plus anciens monuments de la prose persane (1963), for phonology, and Whinfield's notes; the rubāʿī metre follows Elwell-Sutton, The Persian Metres (1976), and Farzaad's division of the line. The sound is computed from a hand transliteration that writes every vowel.

Where it is weakest. Stress is placed by the modern rules, "the least certain part of the reconstruction" (module); the ð rule and the vowel qualities are approximate (manifest). Short vowels are supplied from Whinfield's notes, the metre and the dictionaries and are checked against the script only for its letters. The metre's licences are a reciter's judgment.

What a specialist could improve. The stress rule and the ð rule in persian.py; the transliterations in texts/persian-rubaiyat/edition.yaml.

Middle Mongolian

Scheme.

What it rests on. The Ming transcription itself: which Chinese characters the transcribers chose for which sounds (aspirated initials for t, č, k, q and unaspirated for d, j, g, b, as Shiratori's preface also notes); the Uyghur-script spelling of later Mongolian; and comparison with modern Mongolian (module). The sound is derived from the romanization row, not from the Chinese characters read aloud. That row is Shiratori's romanization (1943) converted by rule into current conventions and hand-corrected for eighteen words, each correction marked in the word's provenance (module and manifest). Two sensors test it against the Ming spelling: the shoulder marks 舌 and 中 (every r has its mark, every mark its r or q) and vowel harmony (manifest).

Where it is weakest. The whole scheme is a reconstruction (manifest sound_confidence: low). The aspiration contrast is inferred from the Chinese characters; the hiatus written ' may already have been a long vowel by the 14th century, and the scheme does not decide where; first-syllable stress is assumed and not attested for the period (module and manifest). The romanization is a transcription of the Ming spelling, not a reconstruction of the lost Uyghur-script original. The shoulder-mark sensors report disagreements as telemetry rather than failures, because the base text sometimes omits a mark.

What a specialist could improve. The conversion rule and the eighteen n overrides in texts/secret-history-mongols/edition.yaml; the vowel values and the aspiration reading in mongolian.py; evidence for where hiatus had become length.

Japanese

Schemes.

Hyphens divide morae; pitch accent is not marked in either scheme.

What it rests on. Reconstructions that keep づ and ぢ distinct into the 1600s, and the Portuguese missionaries' spellings around 1600 (Rodrigues) for ye (module). The sound is computed from each token's reading in historical kana.

Where it is weakest. The dzu of づ, the ye of え and ゑ, and the ō of medial -afu are each uncertain by the 1680s (manifest and key). The module reads a token written only は as the particle wa; a noun spelled は alone would be misread. Pitch accent is absent.

What a specialist could improve. The three uncertain rules in japanese.py; pitch accent, if a period source supports it; the kana readings in texts/basho-furuike/edition.yaml.

Old French and Franco-Italian

Schemes.

What it rests on. Nyrop, Grammaire historique de la langue française, vol. 1 (1899), and Pope, From Latin to Modern French (1934) (module and notes). The sound comes from pipeline/weft/data/fro_lexicon.yaml, keyed by a hand normalization into central French spelling. For it1300, the spelling of Franco-Italian manuscripts; the module states that no grammarian of the period describes such a reading.

Where it is weakest. When final consonants fell silent before consonants, and whether ch and j were still affricates (manifest; the module calls the first "the weakest point of the scheme"). it1300 is a hypothesis. The speech of the writers (a Venetian and a Pisan) is not modelled.

What a specialist could improve. Lexicon entries; the final-consonant rule in oldfrench.py; evidence for or against the Italian reading.

Italian

Schemes.

What it rests on. pipeline/weft/data/it_lexicon.yaml, following the DOP (Migliorini, Tagliavini, Fiorelli, Dizionario d'ortografia e di pronunzia), checked against the Latin source vowel (module). The notes add Migliorini, Storia della lingua italiana (1960), and Izzo, Tuscan and Etruscan (1972).

Where it is weakest. The vowel lexicon is built on the modern Tuscan standard; learned words are less certain for these dates. Latinizing spellings are read as written, which may be ornament rather than speech. The gorgia is left out as not securely attested this early (module and all three manifests).

What a specialist could improve. Lexicon entries, especially those marked # learned; a separate scheme for 1513 if the evidence supports one; the metre rules for synaeresis and dialefe.

Spanish

Schemes.

What it rests on. Nebrija, Gramática de la lengua castellana (1492) and Reglas de orthographía (1517), and his Vocabulario for aspirated h; Lapesa, Historia de la lengua española; Penny, A History of the Spanish Language (module). The sound comes from pipeline/weft/data/es_lexicon.yaml, a normalized period spelling per word.

Where it is weakest. The value of h from Latin f, the b and v distinction, and the quality of the sibilants (manifest). How far ts and dz had lost their stop element by 1492 is uncertain (module). Columbus's own Genoese and Portuguese-coloured speech is not modelled.

What a specialist could improve. Lexicon entries; the b and v assignments word by word.

Dutch

Schemes.

What it rests on. The Twe-spraack vande Nederduitsche letterkunst (1584); de Heuiter, Nederduitse orthographie (1581); Schönfeld, Historische grammatica van het Nederlands, revised by van Loey (1959); de Vooys, Geschiedenis van de Nederlandse taal (1952) (module). Syllables and stress come from pipeline/weft/data/nl_lexicon.yaml.

Where it is weakest. Whether ij was a diphthong or still [iː] in Amsterdam in 1596; the n of -en; the stress of French and Latin loans (manifest). The sound is computed from modern spelling, so the sharp-long and soft-long ee and oo are not shown.

What a specialist could improve. Lexicon entries; the two long ee and oo, if a per-word source exists.

French

Schemes.

No syllable is capitalized: the module states that French has no distinctive word stress.

What it rests on. For m1580: Meigret (1542, 1550), Peletier du Mans (1550), Ramus (1562, 1572), Bèze (1584) and Henri Estienne (1578), as collected by Thurot (1881-1883); Palsgrave (1530) on final consonants (module). For fr1885 the evidence is direct, not reconstructed from rhymes or spelling: Passy, Les sons du français (1887), the transcriptions of Le Maître phonétique (from 1886), and Littré's Dictionnaire de la langue française (1863-1872), which gives a pronunciation for each word (module). The sound comes from pipeline/weft/data/fr_lexicon.yaml; where the lexicon has no fr1885 form, the word is taken from its modern form with the length rule applied.

Where it is weakest. m1580: the value of oi, the r of -er infinitives, and the strength of the e caduc (manifest). Montaigne's Gascon-coloured speech is not modelled. fr1885: vowel length is shown for the word said alone and was reduced inside a phrase; the optional liaisons of formal reading are marked by judgment, following Littré where he gives a note; the scheme models the Paris norm of the language the Act was written in, not the accents of the delegates, who were not French (manifest).

What a specialist could improve. Lexicon entries, including fr1885 forms from Littré; the liaison exceptions (nz) and formal liaisons (zf) in the editions; the length rule in french.py.

German

Schemes.

What it rests on. For northern: Viëtor, Die Aussprache des Schriftdeutschen (1885); Siebs, Deutsche Bühnenaussprache (1898); von Polenz, Deutsche Sprachgeschichte, vols. 2 and 3. For ecg1545: Luther's own spelling and the Saxon chancery usage; Ickelsamer (1527, about 1534); Frangk, Orthographia (1531); von Polenz, vol. 1; Besch, Luther und die deutsche Sprache (2014) (module). Syllables and stress come from pipeline/weft/data/de_lexicon.yaml. Kant's Latin phrases are given in the German school pronunciation of Latin (manifest).

Where it is weakest. northern: postvocalic r and long ä (manifests); no description of either author's speech; regional features not modelled. ecg1545: the diphthong values, the extent of the p/b and t/d merger, the value of w, and whether unrounding of ü and ö reached a reading of Scripture (manifest). Dauid with f is a judgment (manifest).

What a specialist could improve. Lexicon entries, including old vowel lengths where they differ from modern ones; the lenis rule and w in german.py.

Swahili

Schemes.

What it rests on. Steere's own account of his spelling in Swahili Tales (1870), preface, p. xiv ("the vowels are pronounced as in Italian, the consonants as in English, and ... there is always an accent on the last syllable but one"), and the letter values in A Handbook of the Swahili Language as Spoken at Zanzibar, third edition (1884), pp. 8-15 (module). The text is Weft's transcription from the page scans, checked against the Wikisource transcription and the OCR, with the OCR's misreadings declared token by token (manifest).

Where it is weakest. The Handbook used is the third edition, revised by Madan after Steere's death, and no recording exists (manifest). Not modelled: the aspirated p, t and k Steere hears where a nasal has been lost; the implosive b, d and g that later descriptions report; tone and intonation (module). The long consonants of Arabic loans are uncertain, since the Handbook notes a tendency to drop one of them (module and manifest).

What a specialist could improve. The syllabic-nasal rule and the a + e coalescence in swahili.py; the modern forms in texts/swahili-tales-steere/edition.yaml; aspiration where a nasal was lost.

Quenya and Sindarin (Tolkien)

No public work is in either language: Tolkien's texts are in copyright until 2043, so Namárië and the hymn to Elbereth exist only as private works built from a reader's own copy (private/README.md). The module is public because it states pronunciation rules and holds no text.

Scheme.

What it rests on. Appendix E, part I, "Pronunciation of Words and Names", in The Lord of the Rings (1955): the author's own account of how his transcription is to be read. This is the one language in the library whose author wrote down its pronunciation, so the scheme is a statement of his rules rather than a reconstruction. The twelve stress examples Tolkien gives there are the module's regression tests, and so are the stress marks he printed on Namárië in The Road Goes Ever On (1967), as recorded word by word in Eldamo (P. Strack, eldamo.org): the module reproduces every one. Lemma and gloss in the private works follow Tolkien's own word glosses through Eldamo's citations to the page and line; Eldamo's neo-Eldarin forms (fan reconstructions) are never used.

Where it is weakest. Appendix E gives letter values "approximately" and only by English keywords, so the IPA is broad. The quality of Sindarin ae and oe is not described (the respelling reads them as ai, oi, which Appendix E allows). Written th in Quenya is read θ, though Tolkien says the sound had become s in speech. Secondary stress in compounds, which Tolkien marks with a grave in 1967, is not shown. Tolkien's 1952 recordings are not used: they differ from the printed text in places and are a performance, not a rule.

What a specialist could improve. The digraph list and the syllable split in elvish.py; whether ly, ny, ry count as one consonant or two for stress (Tolkien's 1967 marks on ómaryo and Calaciryo support two, as the module has it); the Sindarin long-vowel qualities.

Open inconsistencies

Known inconsistencies between the modules, the manifests and other documentation, listed so a reviewer can weigh them. Several found while this page was written have been corrected (the Koine and Tang labels on later or earlier texts, the Akkadian and runic language names, the scanner notes, the README sample respelling, the Grettis saga scheme order, a Spanish key example).