home edit page issue tracker

This page pertains to UD version 2.

UD for Pali

Pāli is an Indo-Aryan language, widely studied as the liturgical language of Theravāda Buddhism.

Orthography

Sentence Segmentation

While we strip out punctuation in our # text rows, we do recommend using any punctuation in your source to help segment into sentences. In the Pāli Text Society and Mahāsaṅgīti editions, you should generally use both semicolons (;) and periods (.) to split sentences. However, be aware that Pāli verses may sometimes use semicolons to mark pāda “line” boundaries which are not semantic. Always segment sentences by syntax.

Tokenization and Word Segmentation

Morphology

Lemmatization

The LEMMA for all verb forms is the 3rd-person singular present indicative of the underlying verb. For example, the past participle gata “gone” gets the LEMMA gacchati.

Noun Lemmas

Lemmas are always lowercase, even for proper nouns.

For personal pronouns, collapse the cases down to the nominative. For example, amhākaṃ gets mayaṃ and te gets tvaṃ. Keep first and second, singular and plural separate.

For demonstrative pronouns, collapse the cases, gender, and number to the nominative, singular, neuter form:

Substantives and adjectives should get the stem form as their lemma. For example, itthiyo has the lemma itthi, bhikkhave gets the lemma bhikkhu, rasso the lemma rassa, rāja the lemma rājan, bhagavato bhagavant, etc.

Part of Speech Tags

Participles are tagged as VERBs and not ADJs or ADVs. This means that they cannot get the amod dependency relation when used to modify a noun, so we use acl instead to attach it to its target noun.

Because Pāli is originally an oral language, we are stripping out symbols and punctuation, meaning we do not use the SYM or PUNCT tags.

Determiners

If a pronoun such as imaṃ is used as a determiner (i.e. it gets deprel det pointing to a substantive) then it should get the UPOS DET instead of PRON.

Particles

The following particles get the PART UPOS:

Features

Some forms in Pāli are ambiguous. For example, a feminine noun of the class with a -āya suffix might be Singular Instrumental, Ablative, Genitive, Dative, or Locative! When Case can be inferred from context, feel free to mark only the semantically correct case. If multiple parses are reasonable in a given context, list all the plausible values in alphabetical order, separated by commas (e.g. if you have an -āya noun, and all but instrumental are reasonable parses, mark it Case=Abl,Dat,Gen,Loc|Gender=Fem|Number=Sing).

Nouns

Nouns in Pāli have Case, Gender, and Number.

If a genitive form is used in a dative sense, you can mark Case=Gen because we prefer syntactic over semantic readings when they conflict. Do, however, use the Case=Dat annotation for true datives. For example, in the sentence “Uyyānabhūmiṃ gacchāma subhūmidassanāyāti,” dassanāya here is morphologically dative (“for seeing”) and so should get Case=Dat.

Verbs

All verb forms in Pāli should get the VERB UPOS.

Finite Verbs

Finite verbs should have the VerbForm=Fin along with Mood, Tense, Person, Number, and Voice.

Absolutives

Absolutives (sometimes called “gerunds”) formed in Pāli with -tvā / -ya get the VerbForm=Conv feature and are attached to the main verb via advcl.

Infinitives

Infinitives in Pāli end in -tuṃ. They are noted with VerbForm=Inf and are typically xcomp.

Participles

Participles in Pāli are VERBs with the following FEATS:

Syntax

Since all the participles have VERB as their UPOS, they cannot be used as amod. They are either acl if they modify a noun or advcl if they modify a verb or a clause.

We have two subtypes of the obl relation:

Treebanks

There is one Pāli UD treebank: