Encoding diachrony: digital editions of serbian 18th-century texts

Toma Tasovac,Natalia Ermolaev

Encoding diachrony: digital editions of serbian 18th-century texts

2011

Texts in the "Digital Library of Serbian Cultural Heritage of the 18th Century" are encoded as a word-aligned corpus of TEI XML documents in two versions: one using traditional 18th-century orthography, including the graphemes which have since disappeared from Serbian, and one using modernized and standardized Serbian spelling rules that increase the legibility and searchability of these texts for modern users. The corpus also contains linguistic and semantic annotations that add modern phonetic, morphological, lexical and conceptual equivalents to the largely archaic vocabulary. By applying basic techniques of cross-lingual information retrieval to a historical dimension of one language, and making provisions for multiple indexing and annotations, our project exposes a notoriously difficult chapter in the development of the Serbian language to a wider audience, without sacrificing the edition's scholarly potential.

Keywords:

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations