Hlavní stránka/Archiv/2026, ročník 109 / číslo 2

Obálka

Počet zobrazení: 46
Rok 2026, ročník 109, číslo 2

Tiráž

Počet zobrazení: 53
Rok 2026, ročník 109, číslo 2

Obsah

Počet zobrazení: 66
Rok 2026, ročník 109, číslo 2s. 69
Rok 2026, ročník 109, číslo 2s. 71–92
Rok 2026, ročník 109, číslo 2s. 93–123
Rok 2026, ročník 109, číslo 2s. 136–142

Z jazykové poradny

Počet zobrazení: 121
Rok 2026, ročník 109, číslo 2s. 156

Knihy zaslané redakci

Počet zobrazení: 66
Rok 2026, ročník 109, číslo 2

Pokyny pro autory

Počet zobrazení: 54
Rok 2026, ročník 109, číslo 2

Obálka

Počet zobrazení: 41
Rok 2026, ročník 109, číslo 2

Niektoré problémy pri tvorbe ukrajinského webového korpusu Araneum Ucrainicum

Rok 2026, ročník 109, číslo 2

Datum publikování: 6.2026
Autor: Benko, Vladimír
Klíčová slova: Aranea web-crawled corpora, deduplication, ensemble method of lemmatization and tagging, NoSketch Engine corpus manager, tokenization, Ukrainian language, ansámblová metóda lematizácie a morfosyntaktickej anotácie, deduplikácia, korpusový manažér NoSketch Engine, tokenizácia, ukrajinský jazyk, webové korpusy Aranea
Abstrakt: The Ukrainian corpus is being developed in the framework of the Aranea project, aiming to create web-crawled corpora for several (primarily) European languages. The data were downloaded from the Internet across multiple sessions between 2014 and 2026, so the texts also cover the discourse of the pandemic period and the Russo-Ukrainian war. In this paper, we describe the main specific challenges and language-motivated decisions related to the creation of this Ukrainian corpus; particularly the filtering, deduplication, and tokenization of the texts. We present the ensemble method for lemmatization and morphosyntactic annotation using four different tools, followed by aggregation, classification of individual corpus tokens using regular expressions, and the identification of the language of individual corpus sentences. Using a sample of 170 million tokens from the latest data download session, we demonstrate searching the corpus by means of the NoSketch Engine corpus manager and the calculation of subcorpus keywords using the A. Kilgarriffʼs method.
Rubrika: Hlavní články
Rozsah stran: s. 93–123
Status recenzování: recenzovaný článek
Licence: CC BY 4.0

Citace (ISO 690)

error gettting citation: Something bad happened; please try again later.

Dostupné také z:
http://asjournals.lib.cas.cz/naserec/article/uuid:4afca445-02ef-4d09-b44c-03a9964ba4d9/detail