Online digital library sampling based on query related graph

Date10 December 2018
Pages1082-1098
Published date10 December 2018
DOIhttps://doi.org/10.1108/EL-08-2017-0163
AuthorWei Liu,Jing Su
Subject MatterInformation & knowledge management,Information & communications technology,Internet
Online digital library sampling
based on query related graph
Wei Liu
Institute of Scientif‌ic and Technical Information of China, Beijing, China, and
Jing Su
Shaanxi Normal University, Xian, China
Abstract
Purpose Digital library sampling is used to obtain a collection of random literature records from the
backend database, whichis a crucial issue for a variety of importantpurposes in many online digital library
applications. Digital libraries can only be accessed through their query interfaces. The challenge is how to
ensure therandomness of the sample via the autonomous queryinterface.
Design/methodology/approach This paper presentsan iterative and incrementalapproach to obtain
samples throughthe query interface of a digital library. In theapproach, a novel graph model, query-related
graph, is proposed to transform the f‌lat literature records into a graph structure, and samples are obtained
iteratively by traveling the query-related graph. Besides query-related graph, the key components, query
generation,termination conditionand amending deviation,are also discussed in detail.
Findings The extensive experiments over two real digital libraries, ISTIC and IEEE Xplore, show the
proposed approachresults in a better performance. First, the approach is veryeffective to obtain high-quality
samples which are evaluated by the measure sample deviation.Second, the sampling process is very
eff‌icientby only submitting fewer random queries. Third,the approach is robust.
Research limitations/implications This sampling approach is limitedby the query interfaces on a
web page. In rare cases (<3 per cent), this approach cannot access query interfaces by sophisticated
techniques.
Practical implications Digital library sampling is very useful for a variety of important purposes:
subject distribution analysis,literature quality evaluation, digital library size estimation, source selection in
digital libraryintegration and content freshness evaluation.
Social implications Myriads of online digital librariescan be accessed online. Digital library sampling
is a usefulway to understand digital libraries for manyimportant applications.
Originality/value Most of the attributesof a digital library query interface have inf‌inite values, suchas
keyword attributes,which cannot be handled effectivelyby the existing sampling approaches.
Keywords Digital libraries, Data integration, Information analysis, Data sampling
Paper type Research paper
Introduction
With the rapid development of the internet, a myriad of online digital libraries (such as
university libraries, public libraries, etc.) can be accessed through their restricted query
interfaces on the Web. Figure 1 shows a typical query interface, the IEEE Xplore Digital
Library. The literature resources in libraries are limited and subject-oriented; however, a
single digital library cannot meet all the requirements of users. Digital library integration,
an effective way to bridge the gap, can provide users a unif‌ied interface to access multiple
online digital libraries. For instance, Google Scholar and Microsoft Academic Search are
typically comprehensive digital library integration systems which cover popular digital
libraries in most subjects.
EL
36,6
1082
Received30 August 2017
Revised15 January 2018
22April 2018
28June 2018
Accepted18 July 2018
TheElectronic Library
Vol.36 No. 6, 2018
pp. 1082-1098
© Emerald Publishing Limited
0264-0473
DOI 10.1108/EL-08-2017-0163
The current issue and full text archive of this journal is available on Emerald Insight at:
www.emeraldinsight.com/0264-0473.htm
The principle issue studied in this paper is digital library sampling: how can one
eff‌iciently obtain an approximate uniform random sample from the backend database of a
digital library only via the public query interface? As one of the crucial issues in digital
library integration, digital library sampling is very useful for applications which are
attempting to gather statisticalinformation and estimate some aggregate properties of a set
of records. Lots of important features can be estimated or predicted by analyzing samples.
For instance, someof them are listed as follows:
subject distribution analysis;
literature quality evaluation;
digital library size estimation;
source selection digital library integration; and
content freshness evaluation.
Digital library sampling belongs to the research area of Web data source sampling (Dasgupta
et al.,2007,2010; Lu and Li, 2013). Compared to traditional database sampling, Web data source
sampling have to obtain samples via restricted query interfaces. It offers a wide range of
applications for both data source owners and system integration designers. Though lots of
research studies have been done to address this issue via the query interface, they still suffer from
two serious limitations. First, they assume that the attributes in a query interface are independent
of each other. However, attribute correlation is fairly common. For example, considering the
attribute Document Titleand the attribute Index Termsin Figure 1, if the keyword data
integrationis submitted to Document Title,thevaluesofIndex Termsin the returned
records contain or are highly correlated with the two keywords. Second, they can only handle
f‌inite-value attributes, such as numeric attribute and category attribute (Dasgupta et al.,2007).
Keyword attributes are inf‌inite-value ones. In fact, many attributes in the digital library query
interfaces exist as keyword types, such as title, author name, abstract and publisher. As a result,
their solutions are not adapted to the digital library scenario.
This paper tries to f‌ind a solution to digital library sampling to overcome the limitations
above. It aims to randomly capture a number of literature records via the restricted query
interface. In the approach, the literature records in a digital library are modeled as a graph which
is called query-related graph (QRG). The sampling approach is implemented by traveling the
graph. To assure the uniformity of the captured sample, a measure is also raised to evaluate
sample bias and guide the sample process. Overall, the contributions of this paper are
summarized as follows:
A general solution of digital-library-sampler is put forward. This solution
incrementally harvests samples without the expression of the attributes in the query
interface.
Figure 1.
The advanced
searchquery
interfaceof IEEE
Xplore Digital
Library
Online digital
library
sampling
1083

Get this document and AI-powered insights with a free trial of vLex and Vincent AI

Get Started for Free

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex

Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant

  • Access comprehensive legal content with no limitations across vLex's unparalleled global legal database

  • Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength

  • Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities

  • Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting

vLex