Datasets
Every Prompt
A million prompts generated from structured web data (How-to, FAQ, recipes and more).
MMZNO — multimodal benchmark from ZNO tests
Multimodal benchmark for evaluating models on Ukrainian external independent testing (ZNO) tasks.
Authors
Юрій Панів, Артур Кюлян, Дмитро Чаплинський, Микола Хандога, Антон Полішко, Тетяна Бас, Guillermo Gabrielli
License
MIT
Multi30k Extended translation dataset
Extended Ukrainian version of the Multi30k dataset for multimodal machine translation.
Authors
Nataliia Saichyshyna, Daniil Maksymenko, Oleksii Turuta, Andriy Yerokhin, Andrii Babii, Olena Turuta; фільтрація та контроль якості — Юрій Панів
License
Apache-2.0
ParaCrawl 3M English-Ukrainian parallel corpus
Three million filtered English-Ukrainian parallel sentence pairs based on ParaCrawl.
Authors
Дмитро Чаплинський
Links
Hugging Face
UACuisine — Ukrainian recipes dataset
A dataset of Ukrainian culinary recipes.
Authors
Юрій Панів, Артур Кюлян, Дмитро Чаплинський, Микола Хандога, Антон Полішко, Тетяна Бас, Guillermo Gabrielli
UberText-NER-Silver
Silver-standard NER dataset, automatically annotated over UberText 2.0 texts.
Recruitment datasets (Djinni)
Anonymized candidate profiles and job descriptions from Djinni, in English and Ukrainian.
Authors
Назарій Друщак, Мар'яна Романишин
License
MIT
OmniGEC corpora for grammatical error correction
Corpora for training GEC models: Wikipedia edits, Reddit and UberText data.
Authors
Роман Ковальчук, Мар'яна Романишин, Петро Іванюк
License
MIT
Ukrainian hypernymy pairs
A dataset of hypernym-hyponym pairs for Ukrainian.
Authors
Наталія Романишин
Links
Hugging Face