Construction and evaluation of a domain-specific knowledge graph for knowledge discovery

Date03 February 2023
Pages358-370
DOIhttps://doi.org/10.1108/IDD-06-2022-0054
Published date03 February 2023
Subject MatterLibrary & information science,Library & information services,Lending,Document delivery,Collection building & management,Stock revision,Consortia
AuthorHuyen Nguyen,Haihua Chen,Jiangping Chen,Kate Kargozari,Junhua Ding
Construction and evaluation of a
domain-specif‌ic knowledge graph for
knowledge discovery
Huyen Nguyen, Haihua Chen, Jiangping Chen and Kate Kargozari
Department of Information Science, University of North Texas, Denton, Texas, USA, and
Junhua Ding
University of North Texas, Denton, Texas, USA
Abstract
Purpose This study aims to evaluate a method of building a biomedical knowledge graph (KG).
Design/methodology/approach This research f‌irst constructs a COVID-19 KG on the COVID-19 Open Research Data Set, covering information
over six categories (i.e. disease, drug, gene, species, therapy and symptom). The construction used open-source tools to extract entities, relations
and triples. Then, the COVID-19 KG is evaluated on three data-quality dimensions: correctness, relatedness and comprehensiveness, using a
semiautomatic approach. Finally, this study assesses the application of the KG by building a question answering (Q&A) system. Five queries
regarding COVID-19 genomes, symptoms, transmissions and therapeutics were submitted to the system and the results were analyzed.
Findings With current extraction tools, the quality of the KG is moderate and diff‌icult to improve, unless more efforts are made to improve the
tools for entity extraction, relation extraction and others. This study f‌inds that comprehensiveness and relatedness positively correlate with the data
size. Furthermore, the results indicate the performances of the Q&A systems built on the larger-scale KGs are better than the smaller ones for most
queries, proving the importance of relatedness and comprehensiveness to ensure the usefulness of the KG.
Originality/value The KG construction process, data-quality-based and application-based evaluations discussed in this paper provide valuable
references for KG researchers and practitioners to build high-quality domain-specif‌ic knowledge discovery systems.
Keywords Knowledge graph, Knowledge graph evaluation, Data quality, Question answering, COVID-19, CORD-19 data set
Paper type Research paper
1. Introduction
Knowledge graph (KG), with the advantage of capturing the
context of individual entities rather than separating individual
entities in a traditional database, has been recognized as the new
direction for knowledge representation on the semantic web
(Bonatti et al., 2019) and the foundation of building knowledge
discovery systems (Ji et al., 2022). Miscellaneous real-world
knowledge discovery systems, including natural language
understanding, question answering (Q&A) and recommendation
systems (RS), have been built on top of KGs (Ji et al., 2021). The
performance of a knowledge discovery system largely depends on
the quality of the KG. The familiar garbage in, garbage outin a
data-driven system emphasizes that data quality is the key
determinant of the quality of the application (Stvilia et al.,2007).
KGs, known as a type of structured data, inheriting the attribu tes of
data such as being stored, managed, extended, qu ality-assured and
queried, should also conform to the data-qualityrequirements when
building applications. In summary, there are three main
construction and application stages for KG. These stages are for
before,”“underand afterconstruction, which respectively
include data sourcesacquisition and evaluation, knowledge
extraction and fusion and interesting application (Xue and Zou,
2022).
However, it is challenging to apply KGs for domain-specif‌ic
applications such as medical,legal and agriculture. Reasons are
mainly from two aspects:
1 Due to the lack of high-precision domain information
extraction tools, constructing a domain-specif‌icandhigh-
quality KG is diff‌icult (Mishra et al., 2017;Kejriwal, 2020).
Manually constructing a KG bydomain experts can assure the
quality, but it is time-consuming and extremely expensive.
More importantly, the manually crafted KG can hardly be
scalable.
2 Quantitatively and rigorously evaluating the quality of a KG
is expensive due to lack of gold standards; besides, KG
quality has not received enough attention from the research
community so far (Paulheim, 2017;Rula et al., 2020).
The current issue and full text archiveof this journal is available on Emerald
Insight at: https://www.emerald.com/insight/2398-6247.htm
Information Discovery and Delivery
51/4 (2023) 358370
© Emerald Publishing Limited [ISSN 2398-6247]
[DOI 10.1108/IDD-06-2022-0054]
This paper part of special section Information and data quality for
intelligent systems, guest edited by Junhua Ding, Haihua Chen, Lei Li
and Ismini Lourentzou.
Received 21 June 2022
Revised 28 October 2022
26 November 2022
Accepted 11 December 2022
358

Get this document and AI-powered insights with a free trial of vLex and Vincent AI

Get Started for Free

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex