home edit page issue tracker

This page pertains to UD version 2.

UD for Tamil

Preprocessing

Before UD annotation, all texts should be normalised and checked for ambiguous characters. The aim is not to correct the language, but to ensure consistent text representation.

Unicode Normalisation

Vowel modifier/issue Example Normalised form Possible non-normalised form
ொ / o கொ க + ொ க + ெ + ா
ோ / ō கோ க + ோ க + ே + ா
ௌ / au கௌ க + ௌ க + ெ + ௗ
Independent vowel au ஒள

Ambiguous Characters

Note: Ambiguous character replacement should be done carefully. In some contexts, writers may intentionally use Tamil numerals in the text. For example, Tamil numerals may be used in Tamil calendars, traditional documents, or very old texts where Tamil number symbols are expected. In such cases, these characters should not be automatically replaced with visually similar Tamil letters. Annotators should first check the context and replace characters only when it is clear that they are wrongly encoded or unintentionally used.

Ambiguous character/Tamil numeral Correct character Example correction
௨லகம் → உலகம்
௭ன்று → என்று
௮வன் → அவன்
௧லம் → கலம்

General Rules