Construction and evaluation of a domain-specific knowledge graph for knowledge discovery
| Date | 03 February 2023 |
| Pages | 358-370 |
| DOI | https://doi.org/10.1108/IDD-06-2022-0054 |
| Published date | 03 February 2023 |
| Subject Matter | Library & information science,Library & information services,Lending,Document delivery,Collection building & management,Stock revision,Consortia |
| Author | Huyen Nguyen,Haihua Chen,Jiangping Chen,Kate Kargozari,Junhua Ding |
Construction and evaluation of a
domain-specific knowledge graph for
knowledge discovery
Huyen Nguyen, Haihua Chen, Jiangping Chen and Kate Kargozari
Department of Information Science, University of North Texas, Denton, Texas, USA, and
Junhua Ding
University of North Texas, Denton, Texas, USA
Abstract
Purpose –This study aims to evaluate a method of building a biomedical knowledge graph (KG).
Design/methodology/approach –This research first constructs a COVID-19 KG on the COVID-19 Open Research Data Set, covering information
over six categories (i.e. disease, drug, gene, species, therapy and symptom). The construction used open-source tools to extract entities, relations
and triples. Then, the COVID-19 KG is evaluated on three data-quality dimensions: correctness, relatedness and comprehensiveness, using a
semiautomatic approach. Finally, this study assesses the application of the KG by building a question answering (Q&A) system. Five queries
regarding COVID-19 genomes, symptoms, transmissions and therapeutics were submitted to the system and the results were analyzed.
Findings –With current extraction tools, the quality of the KG is moderate and difficult to improve, unless more efforts are made to improve the
tools for entity extraction, relation extraction and others. This study finds that comprehensiveness and relatedness positively correlate with the data
size. Furthermore, the results indicate the performances of the Q&A systems built on the larger-scale KGs are better than the smaller ones for most
queries, proving the importance of relatedness and comprehensiveness to ensure the usefulness of the KG.
Originality/value –The KG construction process, data-quality-based and application-based evaluations discussed in this paper provide valuable
references for KG researchers and practitioners to build high-quality domain-specific knowledge discovery systems.
Keywords Knowledge graph, Knowledge graph evaluation, Data quality, Question answering, COVID-19, CORD-19 data set
Paper type Research paper
1. Introduction
Knowledge graph (KG), with the advantage of capturing the
context of individual entities rather than separating individual
entities in a traditional database, has been recognized as the new
direction for knowledge representation on the semantic web
(Bonatti et al., 2019) and the foundation of building knowledge
discovery systems (Ji et al., 2022). Miscellaneous real-world
knowledge discovery systems, including natural language
understanding, question answering (Q&A) and recommendation
systems (RS), have been built on top of KGs (Ji et al., 2021). The
performance of a knowledge discovery system largely depends on
the quality of the KG. The familiar “garbage in, garbage out”in a
data-driven system emphasizes that data quality is the key
determinant of the quality of the application (Stvilia et al.,2007).
KGs, known as a type of structured data, inheriting the attribu tes of
data such as being stored, managed, extended, qu ality-assured and
queried, should also conform to the data-qualityrequirements when
building applications. In summary, there are three main
construction and application stages for KG. These stages are for
“before,”“under”and “after”construction, which respectively
include data sources’acquisition and evaluation, knowledge
extraction and fusion and interesting application (Xue and Zou,
2022).
However, it is challenging to apply KGs for domain-specific
applications such as medical,legal and agriculture. Reasons are
mainly from two aspects:
1 Due to the lack of high-precision domain information
extraction tools, constructing a domain-specificandhigh-
quality KG is difficult (Mishra et al., 2017;Kejriwal, 2020).
Manually constructing a KG bydomain experts can assure the
quality, but it is time-consuming and extremely expensive.
More importantly, the manually crafted KG can hardly be
scalable.
2 Quantitatively and rigorously evaluating the quality of a KG
is expensive due to lack of gold standards; besides, KG
quality has not received enough attention from the research
community so far (Paulheim, 2017;Rula et al., 2020).
The current issue and full text archiveof this journal is available on Emerald
Insight at: https://www.emerald.com/insight/2398-6247.htm
Information Discovery and Delivery
51/4 (2023) 358–370
© Emerald Publishing Limited [ISSN 2398-6247]
[DOI 10.1108/IDD-06-2022-0054]
This paper part of special section “Information and data quality for
intelligent systems”, guest edited by Junhua Ding, Haihua Chen, Lei Li
and Ismini Lourentzou.
Received 21 June 2022
Revised 28 October 2022
26 November 2022
Accepted 11 December 2022
358
Get this document and AI-powered insights with a free trial of vLex and Vincent AI
Get Started for FreeStart Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant
-
Access comprehensive legal content with no limitations across vLex's unparalleled global legal database
-
Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength
-
Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities
-
Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting
Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant
-
Access comprehensive legal content with no limitations across vLex's unparalleled global legal database
-
Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength
-
Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities
-
Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting
Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant
-
Access comprehensive legal content with no limitations across vLex's unparalleled global legal database
-
Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength
-
Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities
-
Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting
Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant
-
Access comprehensive legal content with no limitations across vLex's unparalleled global legal database
-
Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength
-
Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities
-
Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting
Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant
-
Access comprehensive legal content with no limitations across vLex's unparalleled global legal database
-
Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength
-
Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities
-
Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting