home edit page issue tracker

This page pertains to UD version 2.

PUNCT: punctuation

Definition

Punctuation marks are non-alphabetical characters and character groups used in many languages to delimit linguistic units in printed text.

Punctuation is not taken to include logograms such as $, %, and §, which are instead tagged as SYM.

Note that there is internal punctuation (exclamation, emphasis, and question marks). Words containing internal punctuation are treated as multiword tokens, as in ինչո՞ւ “why?”, which is split into two tokens: ինչու and ՞ (for more details see the tokenization page).

Examples


PUNCT in other languages: [axm] [bej] [bg] [ca] [cs] [cy] [da] [el] [en] [es] [et] [fi] [fr] [ga] [grc] [hbo] [hy] [hyw] [it] [ja] [ka] [kk] [kpv] [ky] [myv] [naq] [no] [oge] [pal] [pt] [ru] [sl] [sv] [tr] [tt] [u] [uk] [urj] [xcl] [xmf] [yue] [zh]