A semantic model to publish open source software on the web of data
| Date | 16 February 2022 |
| Pages | 685-707 |
| DOI | https://doi.org/10.1108/AJIM-09-2021-0280 |
| Published date | 16 February 2022 |
| Subject Matter | Library & information science,Information behaviour & retrieval,Information & knowledge management,Information management & governance,Information management |
| Author | Maedeh Mosharraf |
A semantic model to publish open
source software on the web of data
Maedeh Mosharraf
Computer Science and Engineering, Shahid Beheshti University,
Tehran, Islamic Republic of Iran
Abstract
Purpose –The purpose of the paper is to propose a semantic model for describingopen source software (OSS)
in a machine–human understandable format. The model is extracted to support source code reusing and
revising as the two primary targets of OSS through a systematic review of related documents.
Design/methodology/approach –Conducting a systematic review, all the software reusing criteria are
identified and introduced to the web of data by an ontology for OSS (O4OSS). The software semantic model
introduced in this paper explores OSS through triple expressions in which the O4OSS properties are predicates.
Findings –This model improves the quality of web data by describing software in a structured machine–
human readable profile, which is linked to the related data that was previously published on the web.
Evaluating the OSS semantic model is accomplished through comparing it with previous approaches,
comparing the software structured metadata with profile index of software in some well-known repositories,
calculating the software retrieval rank and surveying domain experts.
Originality/value –Considering context-specific information and authority levels, the proposed software
model would be applicable to any open and close software. Using this model to publish software provides an
infrastructure of connected meaningful data and helps developers overcome some specific challenges. By
navigating software data, many questions which can be answered only through reading multiple documents
can be automatically responded on the web of data.
Keywords Open source software, Semantic model, Web of data, Linked data, Ontology, Software model,
Moodle
Paper type Research paper
1. Introduction
Transforming software development and worldwide distribution, OSS provides developers
with a collaborative environment to share their assets. Sharing and reusing software are
accomplished to improve software productivity, quality, maintainability and reliability while
reducing its development cost, time and complexity (Anguswamy and Frakes, 2012;Taibi,
2013). What is called reusing refers to two distinct purposes as follows:
(1) Reusing software as it is for the same target or any other goals, even in different
contexts and
(2) Revising or improving software for any purpose. If the software license allows, this
can be done by aggregating some modules to produce a new one.
To meet each purpose, users and developers need to find proper software. Nowadays, lots of
repositories provide OSS profiles. Some of these repositories store and index OSS packages and
others act as a portal. Anyway, finding suitable software in these repositories faces some
challenges.
(1) In which repositories can users find software that meets their needs? Is the
information provided by repositories about software up to date?
(2) Different software packages with similar functionalities may exist. To compare these
packages, can users get more information about them without testing and analyzing
software themselves?
OSS semantic
model
685
The current issue and full text archive of this journal is available on Emerald Insight at:
https://www.emerald.com/insight/2050-3806.htm
Received 27 September 2021
Revised 18 December 2021
25 January 2022
Accepted 27 January 2022
Aslib Journal of Information
Management
Vol. 75 No. 4, 2023
pp. 685-707
© Emerald Publishing Limited
2050-3806
DOI 10.1108/AJIM-09-2021-0280
(3) The huge amounts of data provided by various repositories can sometimes be
heterogeneous, leading to user confusion.
(4) In the software development life-cycle, a good number of artifacts are generated,
which are commonly encoded in different formats and can only be accessed through
proprietary and non-standard protocols. Considering this scenario, the adoption of
OSS faces a challenge that is inconsistent with the goal of its production.
In addition to these challenges, providers have some other problems. (1) Finding repositories
to register their software so that it has the most views; (2) for each software improvement, its
information must be updated in all the repositories; each of them requires specific information
in a specific format. Moreover, new paradigms evolved in the software area led to software
agents acting as their users or developers. Service-oriented architecture, mashup-oriented
programing, cloud programing and collaborative programing are the samples. In addition to
software tags, some repositories provide some descriptions or collective feedback about its
features. However, understanding natural language statements is not easy for software
agents. For human users, extracting needed information from these descriptions is also time
consuming. Publishing OSS on the web of data and linking it to other published datasets
eliminate the need to register software in different repositories. As a result, the need for
repeatedly keeping software information up to date is met. In addition, publishing OSS on the
web of data can enhance its accessibility. Providing structured information to end-users by
aggregating data from different software artifacts, responses to complicated queries can
possibly happen. This information that is actually a software model serves some other
important objectives: the model is used for software analysis, facilitates stakeholders’
communication, captures and organizes tool understanding, facilitates OSS reusing, helps
users to decide on choosing software without testing it and facilitates its improvements.
The model should satisfy reusing and revising requirements as the two purposes of OSS
by expressing it in a structure defined by a consistent set of rules. The main contribution of
this paper is developing this model to publish OSS on the web of data. The proposed model
also should express the numerous semantic propositions in a machine- and human-
understandable format, demonstrating the need for meaning. Meaning is added to this model
by appending the ontology entitled O4OSS into it. Utilizing a systematic review, all the
requirements of software reusing and revising which define the O4OSS properties are
extracted. Due to its free licenses considerations, we focus on OSS. However, our approach is
also applicable for close software attending its authorized data. Moodle, a well-known e-
learning software, as the test data illustrates how the possibility of retrieving software can be
increased using the proposed model.
The rest of the paper is organized as follows: Section 2 investigates the previous works
related to this research. In Section 3, we have a systematic review of software reusing criteria and
introduce an ontology that identifies all the extracted software elements as its properties.
Utilizing O4OSS, we introduce a model to publish OSS on the web of data that is adopted by the
Linked Dataprinciples. To illustrate the impact of our model,we publish Moodleon the web of
data and investigate it as a case study in Section 4. Themodel evaluation, whichis divided into
four main parts, is described in Section 5, and finally, in Section 6, the work is concluded.
2. Related works
To facilitate the development activities and avoid coding from scratch, different efforts have
been devoted according to data mining, statistical analysis and knowledge inference
techniques. Nguyen et al. (2018) suggest a framework to assist software developers in mining
OSS repositories and recommending appropriate third-party libraries. In the project entitled
OSSMETER, Ruscio et al. (2015) provide software measurements and an analysis platform to
AJIM
75,4
686
Get this document and AI-powered insights with a free trial of vLex and Vincent AI
Get Started for FreeStart Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant
-
Access comprehensive legal content with no limitations across vLex's unparalleled global legal database
-
Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength
-
Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities
-
Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting
Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant
-
Access comprehensive legal content with no limitations across vLex's unparalleled global legal database
-
Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength
-
Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities
-
Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting
Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant
-
Access comprehensive legal content with no limitations across vLex's unparalleled global legal database
-
Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength
-
Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities
-
Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting
Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant
-
Access comprehensive legal content with no limitations across vLex's unparalleled global legal database
-
Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength
-
Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities
-
Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting
Start Your Free Trial of vLex and Vincent AI, Your Precision-Engineered Legal Assistant
-
Access comprehensive legal content with no limitations across vLex's unparalleled global legal database
-
Build stronger arguments with verified citations and CERT citator that tracks case history and precedential strength
-
Transform your legal research from hours to minutes with Vincent AI's intelligent search and analysis capabilities
-
Elevate your practice by focusing your expertise where it matters most while Vincent handles the heavy lifting