Show simple item record

AuthorAlsarsour, Israa
AuthorMohamed, Esraa
AuthorSuwaileh, Reem
AuthorElsayed, Tamer
Available date2020-07-16T20:11:04Z
Publication Date2019
Publication NameLREC 2018 - 11th International Conference on Language Resources and Evaluation
ResourceScopus
URIhttps://www.scopus.com/inward/record.uri?eid=2-s2.0-85059884453&partnerID=40&md5=b3a515e144e74a7cb819868d62d1b814
URIhttp://hdl.handle.net/10576/15265
AbstractIn this paper, we present a new large manually-annotated multi-dialect dataset of Arabic tweets that is publicly available. The Dialectal ARabic Tweets (DART) dataset has about 25K tweets that are annotated via crowdsourcing and it is well-balanced over five main groups of Arabic dialects: Egyptian, Maghrebi, Levantine, Gulf, and Iraqi. The paper outlines the pipeline of constructing the dataset from crawling tweets that match a list of dialect phrases to annotating the tweets by the crowd. We also touch some challenges that we face during the process. We evaluate the quality of the dataset from two perspectives: the inter-annotator agreement and the accuracy of the final labels. Results show that both measures were substantially high for the Egyptian, Gulf, and Levantine dialect groups, but lower for the Iraqi and Maghrebi dialects, which indicates the difficulty of identifying those two dialects manually and hence automatically.
SponsorThis work was made possible by NPRP grant# NPRP 7-1313-1-245 from the Qatar National Research Fund (a member of Qatar Foundation). The statements made herein are solely the responsibility of the authors. The work was also supported by grant QUST-CENG-SPR-2017-21 from College of Engineering at Qatar University.
Languageen
PublisherEuropean Language Resources Association (ELRA)
SubjectAnnotations
Arabic
Corpus
Crowdsourcing
Multi-Dialect
Twitter
TitleDART: A large dataset of dialectal Arabic tweets
TypeConference Paper


Files in this item

FilesSizeFormatView

There are no files associated with this item.

This item appears in the following Collection(s)

Show simple item record