عرض بسيط للتسجيلة

المؤلفHassaine A.
المؤلفSafi Z.
المؤلفOtaibi J.
المؤلفJaoua A.
تاريخ الإتاحة2019-11-04T05:19:30Z
تاريخ النشر2018
اسم المنشورProceedings of IEEE/ACS International Conference on Computer Systems and Applications, AICCSA
اسم المنشور14th IEEE/ACS International Conference on Computer Systems and Applications, AICCSA 2017
المصدرScopus
الترقيم الدولي الموحد للكتاب 9781538635810
الرقم المعياري الدولي للكتاب21615322
معرّف المصادر الموحدhttp://dx.doi.org/10.1109/AICCSA.2017.102
معرّف المصادر الموحدhttp://hdl.handle.net/10576/12298
الملخصText categorization is an important research field that finds many applications nowadays. It is usually performed in two steps: feature extraction and classification. In the feature extraction step, discriminating keywords are extracted in order to distinguish between different categories of documents. In the classification step, the extracted keywords are fed to a classifier in order to detect the category of each document. In this paper, we use the hyper rectangle method which represents the corpus of documents using a binary relation in which the documents correspond to objects and words to attributes. The hyper rectangle method extracts a tree of keywords such that most discriminative keywords are at the top levels and less discriminative keywords are in the deep levels. We are particularly interested to study different proposed weighting metrics that yield different orderings of keywords. We study how these weighting metrics impact the categorization performance. For the classification step we used both a logistic regression and random forests classifiers. We tested our method on both the 20 newsgroups dataset as well as the Reuters R8 dataset. Our method achieves high performance on both datasets which compete very well with state-of-the-art methods.
راعي المشروعACKNOWLEDGMENT This contribution was made possible by NPRP grant #06-1220-1-233 from the Qatar National Research Fund (a member of Qatar Foundation). The statements made herein are solely the responsibility of the authors.
اللغةen
الناشرIEEE Computer Society
العنوانText categorization using weighted hyper rectangular keyword extraction
النوعConference Paper
الصفحات959-965
رقم المجلد2017-October


الملفات في هذه التسجيلة

الملفاتالحجمالصيغةالعرض

لا توجد ملفات لها صلة بهذه التسجيلة.

هذه التسجيلة تظهر في المجموعات التالية

عرض بسيط للتسجيلة