Knowledge graph embedding for experimental uncertainty estimation

Date08 February 2023
Pages371-383
DOIhttps://doi.org/10.1108/IDD-06-2022-0060
Published date08 February 2023
Subject MatterLibrary & information science,Library & information services,Lending,Document delivery,Collection building & management,Stock revision,Consortia
AuthorEdoardo Ramalli,Barbara Pernici
Knowledge graph embedding for
experimental uncertainty estimation
Edoardo Ramalli and Barbara Pernici
Department of Electronics, Information and Bioengineering, Politecnico di Milano, Milan, Italy
Abstract
Purpose Experiments are the backbone of the development process of data-driven predictive models for scientif‌ic applications. The quality of the
experiments directly impacts the model performance. Uncertainty inherently affects experiment measurements and is often missing in the available
data sets due to its estimation cost. For similar reasons, experiments are very few compared to other data sources. Discarding experiments based on
the missing uncertainty values would preclude the development of predictive models. Data prof‌iling techniques are fundamental t o assess data
quality, but some data quality dimensions are challenging to evaluate without knowing the uncertainty. In this context, this paper aims to pre dict
the missing uncertainty of the experiments.
Design/methodology/approach This work presents a methodology to forecast the experimentsmissing uncertainty, given a data set and its
ontological description. The approach is based on knowledge graph embeddings and leverages the task of link prediction over a knowledge graph
representation of the experiments database. The validity of the methodology is f‌irst tested in multiple conditions using synthetic data and then
applied to a large data set of experiments in the chemical kinetic domain as a case study.
Findings The analysis results of different test case scenarios suggest that knowledge graph embedding can be used to predict the missing unc ertainty of the
experiments when there is a hidden relationship between the experiment metadata and the uncertainty values. The link prediction task is also resilient to
random noise in the relationship. The knowledge graph embedding outperforms the baseline resultsif the uncertainty depends upon multiple metadata.
Originality/value The employment of knowledge graph embedding to predict the missing experimental uncertainty is a novel alternative to the
current and more costly techniques in the literature. Such contribution permits a better data quality prof‌iling of scientif‌ic repositories and improves
the development process of data-driven models based on scientif‌ic experiments.
Keywords Uncertainty prediction, Data uncertainty, Data quality, Data quality management, Data uncertainty management,
Experimental data, Experimental measurement, Uncertainty prediction
Paper type Research paper
1. Introduction
Experimental data (also experiments in the following) are
fundamental to generate chemicalphysical predictive models.
Such models predict complex systems leveraging chemical
physical equations that describe the domain phenomena.
However, some of them are still challenging to explain with
theory (i.e. chemicalphysical equations), and the experiments,
with their observations, can provide phenomenological
evidence about a domainsetting. This information can be used
to ref‌ine and validate a model. During model validation, the
models predictions are compared against the experimental
data, estimating the predictive model performance. For these
reasons, chemicalphysical predictive models often are data-
driven models (Pelucchi et al.,2019). Experiments, unlikely
other types of data such as social media, are rare and expensive
in terms of time and cost to collect. An experiment measures
physical properties in a given domain setting. They are a
particular kind of data because they record physical
measurements inherently affected by experimental uncertainty,
also known asexperimental error (Ramalliet al., 202 1b).This is
usually obtained by repeating the experiment under the same
conditions. Multiple sources of the samefact are hence used to
build a ground truth, which is compared with the experimental
data to compute the difference and thus estimate the
uncertainty (Moffat, 1985). Unfortunately, many experiments
lack to report experimental uncertainty due to the cost of
replicating them (Dai et al.,2019). Similarly, old experiments
are more likely to be imprecise, hence with biggeruncertainties,
due to the use ofold instruments to perform themeasurements.
However, it is unlikely that the community will invest in
The current issue and full text archiveof this journal is available on Emerald
Insight at: https://www.emerald.com/insight/2398-6247.htm
Information Discovery and Delivery
51/4 (2023) 371383
Emerald Publishing Limited [ISSN 2398-6247]
[DOI 10.1108/IDD-06-2022-0060]
© Edoardo Ramalli and Barbara Pernici. Published by Emerald Publishing
Limited. This article is published under the Creative Commons
Attribution (CC BY 4.0) licence. Anyone may reproduce, distribute,
translate and create derivative works of this article (for both commercial
and non-commercial purposes), subject to full attribution to the original
publication and authors. The full terms of this licence may be seen at
http://creativecommons.org/licences/by/4.0/legalcode
This paper part of special section Information and data quality for
intelligent systems, guest edited by Junhua Ding, Haihua Chen, Lei Li
and Ismini Lourentzou.
The work of ER is supported by the interdisciplinarity PhD project of
Politecnico di Milano.
Received 30 June 2022
Revised 24 October 2022
21 December 2022
Accepted 21 December 2022
371

Get this document and AI-powered insights with a free trial of vLex and Vincent AI

Get Started for Free

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex