VLDB 2026 Research / reviewers in the wild / expert
Bojan Bozic
dblp:24/10773
· DBLP profile ↗
13ranked-venue papers
6as first author
3since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 11 · 6 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Data Quality Assessment and Recommendation of Feature Selection Algorithms: An Ontological ApproachabstractFeature selection plays an important role in machine learning and data mining problems. Identifying the best feature selection algorithm that helps to remove irrelevant and redundant features is a complex task. This research tries to address it by recommending a feature selection algorithm based on dataset meta-features. The main contribution of the work is the use of Semantic Web principles to develop a recommendation model for the feature selection algorithm. As a result, dataset meta-features are modeled in a domain ontology, and a set of Semantic Web rule language (SWRL) predictive rules have been proposed to recommend a feature selection algorithm. The result of this research is a feature selection algorithm recommendation based on the data characteristics and quality (FSDCQ) ontology, which not only helps with recommendations but also finds the data points with data quality violations. An experiment is conducted on the classification datasets from the UCI repository to evaluate the proposed ontology. The usefulness and effectiveness of the proposed method is evaluated by comparing it with the widely used method in the literature for the recommendation. Results show that the ontology-based recommendations are equally good as the widely used recommendation model, which is k-NN, with added benefits. Aparna Nayak, Bojan Bozic, Luca Longo |
J. Web Eng. | 2 |
| 2022 | An Ontological Approach for Recommending a Feature Selection AlgorithmabstractFeature selection plays an important role in machine learning or data mining problems. Removing irrelevant features increases model accuracy and reduces the computational cost. However, selecting important features is not a simple task as one feature selection algorithm does not perform well on all the datasets that are of interest. This paper tries to address the recommendation of a feature selection algorithm based on dataset characteristics and quality. The research uses three types of dataset characteristics along with data quality metrics. The main contribution of the work is the utilization of Semantic Web techniques to develop a novel system that can aid in robust feature selection algorithm recommendations. The system’s strength lies in assisting users of machine learning algorithms by providing more relevant feature selection algorithms for the dataset using an ontology called Feature Selection algorithm recommendation based on Data Characteristics and Quality (FSDCQ). Results are generated using six different feature selection algorithms and four types of classifiers on ten datasets from UCI repository. Recommendations take the form of “Feature selection algorithm X is recommended for dataset i, as it performed better on dataset j, similar to dataset i in terms of class overlap 0.3, label noise 0.2, completeness 0.9, conciseness 0.8 units". While the domain-specific ontology FSDCQ was created to aid in the task of algorithm recommendation for feature selection, it is easily applicable to other meta-learning scenarios. Aparna Nayak, Bojan Bozic, Luca Longo |
ICWE | 2 |
| 2021 | KnowText: Auto-generated Knowledge Graphs for custom domain applicationsabstractWhile industrial Knowledge Graphs enable information extraction from massive data volumes creating the backbone of the Semantic Web, the specialised, custom designed knowledge graphs focused on enterprise specific information are an emerging trend. We present “KnowText”, an application that performs automatic generation of custom Knowledge Graphs from unstructured text and enables fast information extraction based on graph visualisation and free text query methods designed for non-specialist users. An OWL ontology automatically extracted from text is linked to the knowledge graph and used as a knowledge base. A basic ontological schema is provided including 16 Classes and Data type Properties. The extracted facts and the OWL ontology can be downloaded and further refined. KnowText is designed for applications in business (CRM, HR, banking). Custom KG can serve for locally managing existing data, often stored as “sensitive” information or proprietary accounts, which are not on open web access. KnowText deploys a custom KG from a collection of text documents and enable fast information extraction based on its graph based visualisation and text based query methods. Bojan Bozic, Jayadeep Kumar Sasikumar, Tamara Matthews |
iiWAS | 1 |
| 2019 | Towards Linked Data for Wikidata Revisions and Twitter Trending HashtagsabstractThis paper uses Twitter as a microblogging platform to link hashtags, which relate the message to a topic that is shared among users, to Wikidata, a central knowledge base of information relying on its members and machine bots to keeping its content up to date. The data is stored in a highly structured format, with the added SPARQL Protocol And RDF Query Language (SPARQL) endpoint to allow users to query its knowledge base. Paula Dooley, Bojan Bozic |
iiWAS | 2 |
| 2019 | Towards a knowledge driven framework for bridging the gap between software and data engineering
Monika Solanki, Bojan Bozic, Christian Dirschl, Rob Brennan |
J. Syst. Softw. | 2 |
| 2018 | Validation of Tagging Suggestion Models for a Hotel Ticketing CorpusabstractThis paper investigates methods for the prediction of tags on a textual corpus that describes hotel staff inputs in a ticketing system. The aim is to improve the tagging process and find the most suitable method for suggesting tags for a new text entry. The paper consists of two parts: (i) exploration of existing sample data, which includes statistical analysis and visualisation of the data to provide an overview, and (ii) evaluation of tag prediction approaches. We have included different approaches from different research fields in order to cover a broad spectrum of possible solutions. As a result, we have tested a machine learning model for multi-label classification (using gradient boosting), a statistical approach (using frequency heuristics), and two simple similarity-based classification approaches (Nearest Centroid and k-Nearest Neighbours). The experiment which compares the approaches uses recall to measure the quality of results. Finally, we provide a recommendation of the modelling approach which produces the best accuracy in terms of tag prediction on the sample data. Bojan Bozic, Andre Rios, Sarah Jane Delany |
iiWAS | 1 |
| 2016 | Building the Seshat Ontology for a Global History Databank
Rob Brennan, Kevin Feeney, Gavin Mendel-Gleason, Bojan Bozic, Peter Turchin, Harvey Whitehouse, Pieter Francois, Thomas E. Currie, Stephanie Grohmann |
ESWC | 4 |
| 2016 | Enabling Combined Software and Data Engineering at Web-Scale: The ALIGNED Suite of Ontologies
Monika Solanki, Bojan Bozic, Markus Freudenberg, Dimitris Kontokostas, Christian Dirschl, Rob Brennan |
ISWC (2) | 2 |
| 2014 | A Social Networking Platform for Semantic Time Series ProcessingabstractIn this paper, we present a social networking platform for semantic time series processing which enables expert users and time series analysts to improve their data collection and collaborate on common data in order to initiate an automated, dynamic process of assignment of right data to the right user. Our approach is the combination of the research areas of Semantic Web, Time Series Processing, and Community Building as a basis for an interactive and intelligent Web portal for expert users. The basis of our portal is a bridge ontology which enables the integration of specific domain ontologies and thus prepares the stage for the application of ontology mapping and reasoning methods. Furthermore, we present our prototype implementation and provide the validation of our concepts based on two domain ontologies from an international research project. Bojan Bozic, Werner Winiwarter, Marco DiCiano |
iiWAS | 1 |
| 2013 | Ontology Mapping and Reasoning in Semantic Time Series ProcessingabstractAfter introducing the new field of Semantic Time Series Processing, we take our findings one step further and investigate ontology mapping and reasoning in semantic time series processing. This paper shows how ontologies for time series from different areas can be mapped together and how additional meta-information can be generated by using reasoning methods. For ontology mapping we use the bridging concept and present a general time series ontology, which allows the integration of new domain ontologies. In reasoning a popular open source reasoner is used to show generation of new time series meta-information. Additionally, we discuss our results together with a description of the validation for all previously introduced methods. Bojan Bozic, Werner Winiwarter |
iiWAS | 1 |
| 2012 | Community building based on semantic time seriesabstractIn this paper, we present a new approach of time series enrichment with semantics, and the usage of this technology for building various kinds of communities interested in time series data. The paper shows the problem of assigning time series data to the right party of interest and why this problem could not be solved so far. We demonstrate a new way of processing semantic time series and the consequential ability of addressing and creating target communities. The combination of time series processing and Semantic Web technologies leads us to a new powerful method of data processing and data generation, which offers completely new opportunities to the expert user. Bojan Bozic, Werner Winiwarter |
iiWAS | 1 |
| 2012 | A Multi-domain Framework for Community Building Based on Data Tagging
Bojan Bozic |
ISWC (2) | 1 |
| 2011 | Deep web integrated systems: current achievements and open issuesabstractThe problem of extracting data that resides in the deep Web has become the center of many research efforts in the recent few years. The challenges in this research area are spanning from online databases discovery and forms extraction from query interfaces, to receiving structured queries from the user, submitting them automatically and retrieving accurate results back to the user. Therefore, the main task is to build an integrated system that connects this variety of missions. In this paper we give an overview of this area of research. We start by surveying previous deep Web systems. After that we define the basic components of a typical deep Web integrated system. Finally, we highlight the current challenges along with possible future research directions. Boutros R. El-Gamil, Werner Winiwarter, Bojan Bozic, Harald Wahl |
iiWAS | 3 |