Introducing the Djinni Recruitment Dataset: A Corpus of Anonymized CVs and Job Postings

Nazarii Drushchak, Mariana Romanyshyn


Abstract
This paper introduces the Djinni Recruitment Dataset, a large-scale open-source corpus of candidate profiles and job descriptions. With over 150,000 jobs and 230,000 candidates, the dataset includes samples in English and Ukrainian, thereby facilitating advancements in the recruitment domain of natural language processing (NLP) for both languages. It is one of the first open-source corpora in the recruitment domain, opening up new opportunities for AI-driven recruitment technologies and related fields. Notably, the dataset is accessible under the MIT license, encouraging widespread adoption for both scientific research and commercial projects.
Anthology ID:
2024.unlp-1.2
Volume:
Proceedings of the Third Ukrainian Natural Language Processing Workshop (UNLP) @ LREC-COLING 2024
Month:
May
Year:
2024
Address:
Torino, Italia
Editors:
Mariana Romanyshyn, Nataliia Romanyshyn, Andrii Hlybovets, Oleksii Ignatenko
Venue:
UNLP
SIG:
Publisher:
ELRA and ICCL
Note:
Pages:
8–13
Language:
URL:
https://aclanthology.org/2024.unlp-1.2
DOI:
Bibkey:
Cite (ACL):
Nazarii Drushchak and Mariana Romanyshyn. 2024. Introducing the Djinni Recruitment Dataset: A Corpus of Anonymized CVs and Job Postings. In Proceedings of the Third Ukrainian Natural Language Processing Workshop (UNLP) @ LREC-COLING 2024, pages 8–13, Torino, Italia. ELRA and ICCL.
Cite (Informal):
Introducing the Djinni Recruitment Dataset: A Corpus of Anonymized CVs and Job Postings (Drushchak & Romanyshyn, UNLP 2024)
Copy Citation:
PDF:
https://preview.aclanthology.org/nschneid-patch-3/2024.unlp-1.2.pdf