A clustering approach for data quality results of research information systems

Date03 November 2022
Pages337-348
DOIhttps://doi.org/10.1108/IDD-07-2022-0063
Published date03 November 2022
Subject MatterLibrary & information science,Library & information services,Lending,Document delivery,Collection building & management,Stock revision,Consortia
AuthorReza Edris Abadi,Mohammad Javad Ershadi,Seyed Taghi Akhavan Niaki
A clustering approach for data quality results of
research information systems
Reza Edris Abadi
Central Tehran Branch, Islamic Azad University, Tehran, Iran
Mohammad Javad Ershadi
Information Technology Department, Iranian Research Institute for Information Science and Technology (IranDoc), Tehran, Iran, and
Seyed Taghi Akhavan Niaki
Industrial Engineering Department, Sharif University of Technology, Tehran, Iran
Abstract
Purpose The overall goal of the data mining process is to extract information from an extensive data set and make it understandable for further
use. When working with large volumes of unstructured data in research information systems, it is necessary to divide the information into logical
groupings after examining their quality before attempting to analyze it. On the other hand, data qual ity results are valuable resources for def‌ining
quality excellence programs of any information system. Hence, the purpose of this study is to discover and extract knowledge to evaluate and
improve data quality in research information systems.
Design/methodology/approach Clustering in data analysis and exploiting the outputs allows practitioners to gain an in-depth and extensive look at their
information to form some logical structures based on what they havefound. In this study, data extracted from an information system are used in the f‌irst stage.
Then, the data quality results are classif‌ied into an organized structure based on data quality dimension standards. Next, clustering algorithms (K-Means),
density-based clustering (density-based spatial clustering of applications with noise [DBSCAN]) and hierarchical clustering (balanced iterative reducing and
clustering using hierarchies [BIRCH]) are applied to compare and f‌indt hem ostapp ropriate clustering algorithms in the research information system.
Findings This paper showed that quality control results of an information system could be categorized through well-kn own data quality
dimensions, including precision, accuracy, completeness, consistency, reputation and timeliness. Furthermore, among different well-known
clustering approaches, the BIRCH algorithm of hierarchical clustering methods performs better in data clustering and gives the highest silhouette
coeff‌icient value. Next in line is the DBSCAN method, which performs better than the K-Means method.
Research limitations/implications In the data quality assessment process, the discrepancies identif‌ied and the lack of proper classif‌ication for
inconsistent data have led to unstructured reports, making the statistical analysis of qualitative metadata problems diff‌icult and thus impossible to
root out the observed errors. Therefore, in this study, the evaluation results of data quality have been categorized into various data quality
dimensions, based on which multiple analyses have been performed in the form of data mining methods.
Originality/value Although several pieces of research have been conducted to assess data quality results of research information systems ,
knowledge extraction from obtained data quality scores is a crucial work that has rarely been studied in the literature. Besides, clustering in data
quality analysis and exploiting the outputs allows practitioners to gain an in-depth and extensive look at their information to form some logical
structures based on what they have found.
Keywords Clustering, Data mining, Data quality, Information systems, Knowledge management, Machine learning
Paper type Research paper
1. Introduction
Ever since the computer was used to analyze and store data, after
about 20 years, the amount of data in the database has doubled.
This upward trend has continued, so we are dealing with the term
information explosion in todays world. Organizations and
companies have purchased or created a database to advance their
workf‌lows. This database can be essential for managers, planners
and researchers to make strategic decisions, prepare various
reports and describe their current situation.
Today, with integrated information systems, integrated
banking and e-commerce systems, the amount of data in a
database is increasing moment by moment, ultimately leading
to massive data warehouses. Therefore, the need for rapid and
accurate discovery and knowledge extraction from these
databases has become more apparent. For example, extensive
databases of studentscharacteristics include information on
family, academic and other features. Finding patterns and
knowledge contained in this information can greatly help
higher educationdecision-makers.
The current issue and full text archiveof this journal is available on Emerald
Insight at: https://www.emerald.com/insight/2398-6247.htm
Information Discovery and Delivery
51/4 (2023) 337348
© Emerald Publishing Limited [ISSN 2398-6247]
[DOI 10.1108/IDD-07-2022-0063]
This paper part of special section Information and data quality for
intelligent systems, guest edited by Junhua Ding, Haihua Chen, Lei Li
and Ismini Lourentzou.
Received 5 July 2022
Revised 3 September 2022
23 September 2022
Accepted 23 September 2022
337

Get this document and AI-powered insights with a free trial of vLex and Vincent AI

Get Started for Free

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex