Datasets

Every Prompt

A million prompts generated from structured web data (How-to, FAQ, recipes and more).

Permanent link

Authors
Дмитро Чаплинський
License
MIT
MMZNO — multimodal benchmark from ZNO tests

Multimodal benchmark for evaluating models on Ukrainian external independent testing (ZNO) tasks.

Permanent link

Authors
Юрій Панів, Артур Кюлян, Дмитро Чаплинський, Микола Хандога, Антон Полішко, Тетяна Бас, Guillermo Gabrielli
License
MIT
Multi30k Extended translation dataset

Extended Ukrainian version of the Multi30k dataset for multimodal machine translation.

Permanent link

Authors
Nataliia Saichyshyna, Daniil Maksymenko, Oleksii Turuta, Andriy Yerokhin, Andrii Babii, Olena Turuta; фільтрація та контроль якості — Юрій Панів
License
Apache-2.0
ParaCrawl 3M English-Ukrainian parallel corpus

Three million filtered English-Ukrainian parallel sentence pairs based on ParaCrawl.

Permanent link

Authors
Дмитро Чаплинський
UACuisine — Ukrainian recipes dataset

A dataset of Ukrainian culinary recipes.

Permanent link

Authors
Юрій Панів, Артур Кюлян, Дмитро Чаплинський, Микола Хандога, Антон Полішко, Тетяна Бас, Guillermo Gabrielli
UberText-NER-Silver

Silver-standard NER dataset, automatically annotated over UberText 2.0 texts.

Permanent link

Authors
Владислав Радченко, Назарій Друщак
License
Apache-2.0
Recruitment datasets (Djinni)

Anonymized candidate profiles and job descriptions from Djinni, in English and Ukrainian.

Permanent link

Authors
Назарій Друщак, Мар'яна Романишин
License
MIT
OmniGEC corpora for grammatical error correction

Corpora for training GEC models: Wikipedia edits, Reddit and UberText data.

Permanent link

Authors
Роман Ковальчук, Мар'яна Романишин, Петро Іванюк
License
MIT
Ukrainian hypernymy pairs

A dataset of hypernym-hyponym pairs for Ukrainian.

Permanent link

Authors
Наталія Романишин