Document classification using term frequency-inverse document frequency and K-means clustering

Wasseem N. Ibrahem Al-Obaydy; Hala A. Hashim; Yassen AbdelKhaleq Najm; Ahmed Adeeb Jalal

doi:10.11591/ijeecs.v27.i3.pp1517-1524

Document classification using term frequency-inverse document frequency and K-means clustering

Wasseem N. Ibrahem Al-Obaydy, Hala A. Hashim, Yassen AbdelKhaleq Najm, Ahmed Adeeb Jalal

Abstract

Increased advancement in a variety of study subjects and information technologies, has increased the number of published research articles. However, researchers are facing difficulties and devote a significant time amount in locating scientific research publications relevant to their domain of expertise. In this article, an approach of document classification is presented to cluster the text documents of research articles into expressive groups that encompass a similar scientific field. The main focus and scopes of target groups were adopted in designing the proposed method, each group include several topics. The word tokens were separately extracted from topics related to a single group. The repeated appearance of word tokens in a document has an impact on the document's weight, which is computed using the term frequency-inverse document frequency (TF-IDF) numerical statistic. To perform the categorization process, the proposed approach employs the paper's title, abstract, and keywords, as well as the categories' topics. We exploited the K-means clustering algorithm for classifying and clustering the documents into primary categories. The K-means algorithm uses category weights to initialize the cluster centers (or centroids). Experimental results have shown that the suggested technique outperforms the k-nearest neighbors algorithm in terms of accuracy in retrieving information.

Keywords

Data mining; Document classification; K-means clustering; TF-IDF; Topics

Full Text:

PDF

DOI: http://doi.org/10.11591/ijeecs.v27.i3.pp1517-1524

Refbacks

There are currently no refbacks.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

Indonesian Journal of Electrical Engineering and Computer Science (IJEECS)
p-ISSN: 2502-4752, e-ISSN: 2502-4760
This journal is published by the Institute of Advanced Engineering and Science (IAES).

IJEECS visitor statistics

Username
Password
Remember me