Bela Stantic

dblp:74/1636 · DBLP profile ↗
← Back
64ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0003-0475-7951ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 36 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 31 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Systems, architecture and hardware · 2Computer networks · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
YearPublicationVenuePosition
2025 Prompts De-Biasing Augmentation to Mitigate Gender Stereotypes in Large Language Models
Jinyuan Chen, Sebastian Binnewies, Bela Stantic
ACIIDS (1)3
2025 Integration of Dynamic Window Sizing with Neural Network Architectures for Real-Time Cryptocurrency Predictions
David L. John, Sebastian Binnewies, Bela Stantic
ACIIDS (2)3
2025 A data-centric framework for combating domain shift in underwater object detection with image enhancement
abstract
Abstract Underwater object detection has numerous applications in protecting, exploring, and exploiting aquatic environments. However, underwater environments pose a unique set of challenges for object detection including variable turbidity, colour casts, and light conditions. These phenomena represent a domain shift and need to be accounted for during design and evaluation of underwater object detection models. Although methods for underwater object detection have been extensively studied, most proposed approaches do not address challenges of domain shift inherent to aquatic environments. In this work we propose a data-centric framework for combating domain shift in underwater object detection with image enhancement. We show that there is a significant gap in accuracy of popular object detectors when tested for their ability to generalize to new aquatic domains. We used our framework to compare 14 image processing and enhancement methods in their efficacy to improve underwater domain generalization using three diverse real-world aquatic datasets and two widely used object detection algorithms. Using an independent test set, our approach superseded the mean average precision performance of existing model-centric approaches by 1.7–8.0 percentage points. In summary, the proposed framework demonstrated a significant contribution of image enhancement to underwater domain generalization.
Lukas Folkman, Kylie A. Pitt, Bela Stantic
Appl. Intell.3
2024 Identifying Optimal Window Size Configurations for Big Data Time Series Forecasting
abstract
Optimal window sizing in time series forecasting emerges as a pivotal factor for enhancing predictive accuracy, particularly in the volatile cryptocurrency market. While traditional models often rely on static window sizes, resulting in compromised forecasting performance, this research explores optimal window configurations across various market volatilities. By employing a hybrid Long Short-Term Memory and Gated Recurrent Unit (LSTM-GRU) model, the study systematically identifies the most effective window sizes for high, medium, and low volatility conditions. Results demonstrate that smaller windows are preferable in highly volatile environments to capture rapid market shifts, whereas larger windows are more suitable for stable conditions to incorporate a broader historical context. By identifying the predetermined optimal window sizes for each volatility segment, this study offers valuable insights for researchers aiming to enhance the adaptability and efficacy of predictive models. These results are especially useful for exploring dynamic window sizing techniques across various domains, particularly in fields where data volatility significantly impacts model performance.
David L. John, Sebastian Binnewies, Bela Stantic
IEEE Big Data3
2024 An Iterative Graph-Based Method for Constructing Gaps in High-Voltage Bundle Conductors Using Airborne LiDAR Point Cloud Data
abstract
Transmission line safety is vital to the nation’s economy and daily lives. In recent decades, most electric utilities have relied on light detection and ranging (LiDAR) technology to inspect transmission line corridors. However, occlusion and point density gaps in scanned LiDAR data will affect power line extraction. Moreover, high-voltage transmission line (HVTL) corridors in complicated locations with multiple loops pose significant challenges. The detailed modeling of power lines is a prerequisite for detecting potential risks. Thus, this study introduces a fourfold strategy to extract and correctly reconstruct huge gaps in HVTL conductor bundles. In the first step, the power lines are retrieved as span points using the 3-D voxel grid. The second step generates bundle masks by splitting span points into several segments. A bipartite graph connects bundle segments with wide gaps utilizing bundle location information from these masks. Third, bundles are segmented again to form conductor masks for sub-conductor extraction utilizing image-based algorithms and probability. Finally, the mathematical model reconstructs the retrieved power lines as sub-conductors. The proposed technique is tested on 58 spans with different power line configurations from two low-point-density datasets with a few constant parameters. The test results demonstrate the resilience of the proposed method in effectively repairing large gaps in bundles and sub-conductors. The sub-conductors are retrieved with 90% accuracy. The suggested method robustly reconstructs sub-conductors with a fitting residual error of less than 0.07 m.
Nosheen Munir, Mohammad Awrangjeb, Bela Stantic
IEEE Trans. Geosci. Remote. Sens.3
2022 Machine Learning or Lexicon Based Sentiment Analysis Techniques on Social Media Posts
David L. John, Bela Stantic
ACIIDS (2)2
2022 3D Reconstruction of Bundle Sub-Conductors Using LiDAR Data From Forest Terrains
abstract
Utilities have recently shown a considerable interest in extracting powerlines from laser scanning data for periodic utility monitoring. However, if the powerline runs through a forest, extraction is more difficult. A robust and exact power line model is crucial for adequate clearance and detecting potential hazards. Thus, this study establishes and automated method for reconstructing high-voltage bundle conductors. The powerline bundles at various heights are first recovered and divided into segments. The fitted residuals of each segment are used to estimate the number of sub-conductors. The point distance formula and 3D line fitting are used to extract individual sub-conductors. Each part is separated according to its positive and negative distance values. Finally, the sub-conductor segments are reassembled using a random sample consensus (RANSAC) technique. On five datasets, the proposed technique extracts and reconstructs individual sub-conductors with great precision.
Nosheen Munir, Mohammad Awrangjeb, Bela Stantic
IGARSS3
2021 Empirical Study of Tweets Topic Classification Using Transformer-Based Language Models
Ranju Mandal, Susanne Becken, Bela Stantic
ACIIDS4
2019 Word Mover's Distance for Agglomerative Short Text Clustering
Nigel Franciscus, Xuguang Ren, Junhu Wang, Bela Stantic
ACIIDS (1)4
2019 Event Prediction Based on Causality Reasoning
Xuguang Ren, Nigel Franciscus, Junhu Wang, Bela Stantic
ACIIDS (1)5
2019 Top-N Hashtag Prediction via Coupling Social Influence and Homophily
Can Wang 0004, Yunwei Zhao, Chihung Chi, Willem-Jan van den Heuvel, Kwok-Yan Lam, Bela Stantic
ADMA7
2019 Mining Summary of Short Text with Centroid Similarity Distance
Nigel Franciscus, Junhu Wang, Bela Stantic
ADMA3
2019 An Unsupervised Outlier Detection Method For 3D Point Cloud Data
abstract
This paper introduces an effective method for outlier detection from the point cloud data. Although, the state-of-the-art methods offer good results in removing outliers, in most of the cases inliers are also removed erroneously. This paper focuses on this issue using the information based on a relative location from a point to its neighbours and a robust z-score based on a statistical approach. Synthetic datasets for 3D building roofs have been created to evaluate the performance. When compared with the existing methods, the proposed method exhibits better performance, i.e., 19% more recall for inliers, 6% more precision for outliers and 10% more overall accuracy. In other words, it not only preserves the inliers, but also correctly removes the outliers with a better precision rate than the current state-of-the-arts methods.
Emon Kumar Dey, Mohammad Awrangjeb, Bela Stantic
IGARSS3
2019 Handling probabilistic integrity constraints in pay-as-you-go reconciliation of data models
abstract
Data models capture the structure and characteristic properties of data entities, e.g., in terms of a database schema or an ontology. They are the backbone of diverse applications, reaching from information integration , through peer-to-peer systems and electronic commerce to social networking . Many of these applications involve models of diverse data sources. Effective utilisation and evolution of data models, therefore, calls for matching techniques that generate correspondences between their elements. Various such matching tools have been developed in the past. Yet, their results are often incomplete or erroneous, and thus need to be reconciled, i.e., validated by an expert. This paper analyses the reconciliation process in the presence of large collections of data models, where the network induced by generated correspondences shall meet consistency expectations in terms of integrity constraints. We specifically focus on how to handle data models that show some internal structure and potentially differ in terms of their assumed level of abstraction. We argue that such a setting calls for a probabilistic model of integrity constraints, for which satisfaction is preferred, but not required. In this work, we present a model for probabilistic constraints that enables reasoning on the correctness of individual correspondences within a network of data models, in order to guide an expert in the validation process. To support pay-as-you-go reconciliation, we also show how to construct a set of high-quality correspondences, even if an expert validates only a subset of all generated correspondences. We demonstrate the efficiency of our techniques for real-world datasets comprising database schemas and ontologies from various application domains.
Nguyen Quoc Viet Hung, Matthias Weidlich 0001, Thanh Tam Nguyen, Zoltán Miklós 0001, Karl Aberer, Avigdor Gal, Bela Stantic
Inf. Syst.7
2019 Multi-label classification via label correlation and first order feature dependance in a data stream
Tien Thanh Nguyen, Thi Thu Thuy Nguyen, Anh Vu Luong, Nguyen Quoc Viet Hung, Alan Wee-Chung Liew, Bela Stantic
Pattern Recognit.6
2019 From Anomaly Detection to Rumour Detection using Data Streams of Social Platforms
abstract
Social platforms became a major source of rumours. While rumours can have severe real-world implications, their detection is notoriously hard: Content on social platforms is short and lacks semantics; it spreads quickly through a dynamically evolving network; and without considering the context of content, it may be impossible to arrive at a truthful interpretation. Traditional approaches to rumour detection, however, exploit solely a single content modality, e.g., social media posts, which limits their detection accuracy. In this paper, we cope with the aforementioned challenges by means of a multi-modal approach to rumour detection that identifies anomalies in both, the entities (e.g., users, posts, and hashtags) of a social platform and their relations. Based on local anomalies, we show how to detect rumours at the network level, following a graph-based scan approach. In addition, we propose incremental methods, which enable us to detect rumours using streaming data of social platforms. We illustrate the effectiveness and efficiency of our approach with a real-world dataset of 4M tweets with more than 1000 rumours.
Thanh Tam Nguyen, Matthias Weidlich 0001, Bolong Zheng, Hongzhi Yin, Nguyen Quoc Viet Hung, Bela Stantic
Proc. VLDB Endow.6
2019 User Guidance for Efficient Fact Checking
abstract
The Web constitutes a valuable source of information. In recent years, it fostered the construction of large-scale knowledge bases, such as Freebase, YAGO, and DBpedia. The open nature of the Web, with content potentially being generated by everyone, however, leads to inaccuracies and misinformation. Construction and maintenance of a knowledge base thus has to rely on fact checking, an assessment of the credibility of facts. Due to an inherent lack of ground truth information, such fact checking cannot be done in a purely automated manner, but requires human involvement. In this paper, we propose a comprehensive framework to guide users in the validation of facts, striving for a minimisation of the invested effort. Our framework is grounded in a novel probabilistic model that combines user input with automated credibility inference. Based thereon, we show how to guide users in fact checking by identifying the facts for which validation is most beneficial. Moreover, our framework includes techniques to reduce the manual effort invested in fact checking by determining when to stop the validation and by supporting efficient batching strategies. We further show how to handle fact checking in a streaming setting. Our experiments with three real-world datasets demonstrate the efficiency and effectiveness of our framework: A knowledge base of high quality, with a precision of above 90%, is constructed with only a half of the validation effort required by baseline techniques.
Thanh Tam Nguyen, Hongzhi Yin, Matthias Weidlich 0001, Bolong Zheng, Nguyen Quoc Viet Hung, Bela Stantic
Proc. VLDB Endow.6
2019 Efficient User Guidance for Validating Participatory Sensing Data
abstract
Participatory sensing has become a new data collection paradigm that leverages the wisdom of the crowd for big data applications without spending cost to buy dedicated sensors. It collects data from human sensors by using their own devices such as cell phone accelerometers, cameras, and GPS devices. This benefit comes with a drawback: human sensors are arbitrary and inherently uncertain due to the lack of quality guarantee. Moreover, participatory sensing data are time series that exhibit not only highly irregular dependencies on time but also high variance between sensors. To overcome these limitations, we formulate the problem of validating uncertain time series collected by participatory sensors. In this article, we approach the problem by an iterative validation process on top of a probabilistic time series model. First, we generate a series of probability distributions from raw data by tailoring a state-of-the-art dynamical model, namely Generalised Auto Regressive Conditional Heteroskedasticity (GARCH), for our joint time series setting. Second, we design a feedback process that consists of an adaptive aggregation model to unify the joint probabilistic time series and an efficient user guidance model to validate aggregated data with minimal effort. Through extensive experimentation, we demonstrate the efficiency and effectiveness of our approach on both real data and synthetic data. Highlights from our experiences include the fast running time of a probabilistic model, the robustness of an aggregation model to outliers, and the significant effort saving of a guidance model.
Thanh Cong Phan, Thanh Tam Nguyen, Hongzhi Yin, Bolong Zheng, Bela Stantic, Nguyen Quoc Viet Hung
ACM Trans. Intell. Syst. Technol.5
2018 Trace Ratio Optimization With Feature Correlation Mining for Multiclass Discriminant Analysis
abstract
Fisher's linear discriminant analysis is a widely accepted dimensionality reduction method, which aims to find a transformation matrix to convert feature space to a smaller space by maximising the between-class scatter matrix while minimising the within-class scatter matrix. Although the fast and easy process of finding the transformation matrix has made this method attractive, overemphasizing the large class distances makes the criterion of this method suboptimal. In this case, the close class pairs tend to overlap in the subspace. Despite different weighting methods having been developed to overcome this problem, there is still a room to improve this issue. In this work, we study a weighted trace ratio by maximising the harmonic mean of the multiple objective reciprocals. To further improve the performance, we enforce the l2,1-norm to the developed objective function. Additionally, we propose an iterative algorithm to optimise this objective function. The proposed method avoids the domination problem of the largest objective, and guarantees that no objectives will be too small. This method can be more beneficial if the number of classes is large. The extensive experiments on different datasets show the effectiveness of our proposed method when compared with four state-of-the-art methods.
Forough Rezaei Boroujeni, Sen Wang 0001, Zhihui Li 0001, Nicholas West, Bela Stantic, Lina Yao 0001, Guodong Long
AAAI5
2018 An Ensemble System with Random Projection and Dynamic Ensemble Selection
Anh Vu Luong, Tuyet-Trinh Vu, Nguyen Quoc Viet Hung, Tien Thanh Nguyen, Bela Stantic
ACIIDS (1)6
2018 Beyond Word-Cloud: A Graph Model Derived from Beliefs
Nigel Franciscus, Xuguang Ren, Bela Stantic
ACIIDS (2)3
2018 Automatic Image Region Annotation by Genetic Algorithm-Based Joint Classifier and Feature Selection in Ensemble System
Anh Vu Luong, Tien Thanh Nguyen, Xuan Cuong Pham, Thi Thu Thuy Nguyen, Alan Wee-Chung Liew, Bela Stantic
ACIIDS (1)6
2018 Experimental Clarification of Some Issues in Subgraph Isomorphism Algorithms
Xuguang Ren, Junhu Wang, Nigel Franciscus, Bela Stantic
ACIIDS (2)4
2018 "Quality" vs. "Readability" in Document Images: Statistical Analysis of Human Perception
abstract
Based on the hypothesis that a good / poor quality document image is most probably a readable / unreadable document, document image quality and readability have interchangeably been used in the literature. These two terms, however, have different meanings implying two different perspectives of looking at document images by human being. In document images, the level of quality and the degree of readability may have a relation / correlation considering human perception. However, to the best of our knowledge there is no specific study to characterise this relation and also validate the abovementioned hypothesis. In this work, at first, we created a dataset composed of mostly camera-based document images with various distortion levels. Each document image has then been assessed with regard to two different measures, the level of quality and the degree of readability, by different individuals. A detailed Normalised Cross Correlation analysis along with different statistical analysis based on Shapiro-Wilks and Wilcoxon tests has further been provided to demonstrate how document image quality and readability are linked. Our findings indicate that the quality and readability were somewhat different in terms of the population distributions. However, the correlation between quality and readability was 0.99, which implies document quality and readability are highly correlated based on human perception.
Alireza Alaei, Romain Raveaux, Donatello Conte, Bela Stantic
DAS4
2018 A System for Spatial-Temporal Trajectory Data Integration and Representation
Douglas Alves Peixoto, Xiaofang Zhou 0001, Nguyen Quoc Viet Hung, Dan He 0009, Bela Stantic
DASFAA (2)5
2018 What-If Analysis with Conflicting Goals: Recommending Data Ranges for Exploration
abstract
What-if analysis is a data-intensive exploration to inspect how changes in a set of input parameters of a model influence some outcomes. It is motivated by a user trying to understand the sensitivity of a model to a certain parameter in order to reach a set of goals that are defined over the outcomes. To avoid an exploration of all possible combinations of parameter values, efficient what-if analysis calls for a partitioning of parameter values into data ranges and a unified representation of the obtained outcomes per range. Traditional techniques to capture data ranges, such as histograms, are limited to one outcome dimension. Yet, in practice, what-if analysis often involves conflicting goals that are defined over different dimensions of the outcome. Working on each of those goals independently cannot capture the inherent trade-off between them. In this paper, we propose techniques to recommend data ranges for what-if analysis, which capture not only data regularities, but also the trade-off between conflicting goals. Specifically, we formulate a parametric data partitioning problem and propose a method to find an optimal solution for it. Targeting scalability to large datasets, we further provide a heuristic solution to this problem. By theoretical and empirical analyses, we establish performance guarantees in terms of runtime and result quality.
Nguyen Quoc Viet Hung, Kai Zheng 0001, Matthias Weidlich 0001, Bolong Zheng, Hongzhi Yin, Thanh Tam Nguyen, Bela Stantic
ICDE7
2018 Assessing fish abundance from underwater video using deep neural networks
abstract
Uses of underwater videos to assess diversity and abundance of fish are being rapidly adopted by marine biologists. Manual processing of videos for quantification by human analysts is time and labour intensive. Automatic processing of videos can be employed to achieve the objectives in a cost and time-efficient way. The aim is to build an accurate and reliable fish detection and recognition system, which is important for an autonomous robotic platform. However, there are many challenges involved in this task (e.g. complex background, deformation, low resolution and light propagation). Recent advancement in the deep neural network has led to the development of object detection and recognition in real time scenarios. An end-to-end deep learningbased architecture is introduced which outperformed the state of the art methods and first of its kind on fish assessment task. A Region Proposal Network (RPN) introduced by an object detector termed as Faster R-CNN was combined with three classification networks for detection and recognition of fish species obtained from Remote Underwater Video Stations (RUVS). An accuracy of 82.4% (mAP) obtained from the experiments are much higher than previously proposed methods.
Ranju Mandal, Rod M. Connolly, Thomas A. Schlacher, Bela Stantic
IJCNN4
2018 AMGA: An Adaptive and Modular Genetic Algorithm for the Traveling Salesman Problem
Ryoma J. Ohira, Md. Saiful Islam 0003, Jun Jo 0001, Bela Stantic
ISDA (2)4
2018 Concept for Evaluation of Techniques for Trajectory Distance Measures
abstract
Measuring the similarity (or distance) between trajectories of moving objects is a common procedure taken by most trajectory data-driven applications. One of the biggest challenges of trajectory distances measurement is that the distance needs to be carefully defined in order to reflect the true underlying similarity. This is due to the fact that trajectories are essentially non-uniform sequential data with variable length, attached with both spatial and temporal attributes, which may or may not be considered for similarity measures. Therefore, tens of similarity measures for trajectory data have been proposed; every technique claim an advantage over the others in a different aspect. Hence, it's difficult for users to choose the best-suited technique, as well as the appropriate parameter values, since each technique has distinct performance and characteristics depending on various factors. In this paper, we develop an application that allows to evaluate several techniques in different aspects (accuracy, sensitivity to trajectory features, performance, etc.). We believe that this tool will be able to serve as a practical guideline for both researchers and developers. While researchers can use our tool to assess existing or new techniques, developers can reuse its components to reduce the development complexity.
Douglas Alves Peixoto, Han Su 0001, Nguyen Quoc Viet Hung, Bela Stantic, Bolong Zheng, Xiaofang Zhou 0001
MDM4
2018 Representing and querying now-relative relational medical data
Luca Anselma, Luca Piovesan, Bela Stantic, Paolo Terenziani
Artif. Intell. Medicine3
2018 Answering why-not questions on semantic multimedia queries
Meng Wang 0009, Weitong Chen 0001, Sen Wang 0001, Jun Liu 0002, Xue Li 0001, Bela Stantic
Multim. Tools Appl.6
2018 Coupled Clustering Ensemble by Exploring Data Interdependence
abstract
Clustering ensembles combine multiple partitions of data into a single clustering solution. It is an effective technique for improving the quality of clustering results. Current clustering ensemble algorithms are usually built on the pairwise agreements between clusterings that focus on the similarity via consensus functions, between data objects that induce similarity measures from partitions and re-cluster objects, and between clusters that collapse groups of clusters into meta-clusters. In most of those models, there is a strong assumption on IIDness (i.e., independent and identical distribution), which states that base clusterings perform independently of one another and all objects are also independent. In the real world, however, objects are generally likely related to each other through features that are either explicit or even implicit. There is also latent but definite relationship among intermediate base clusterings because they are derived from the same set of data. All these demand a further investigation of clustering ensembles that explores the interdependence characteristics of data. To solve this problem, a new coupled clustering ensemble (CCE) framework that works on the interdependence nature of objects and intermediate base clusterings is proposed in this article. The main idea is to model the coupling relationship between objects by aggregating the similarity of base clusterings, and the interactive relationship among objects by addressing their neighborhood domains. Once these interdependence relationships are discovered, they will act as critical supplements to clustering ensembles. We verified our proposed framework by using three types of consensus function: clustering-based, object-based, and cluster-based. Substantial experiments on multiple synthetic and real-life benchmark datasets indicate thatCCEcan effectively capture the implicit interdependence relationships among base clusterings and among objects with higher clustering accuracy, stability, and robustness compared to 14 state-of-the-art techniques, supported by statistical analysis. In addition, we show that the final clustering quality is dependent on the data characteristics (e.g., quality and consistency) of base clusterings in terms of sensitivity analysis. Finally, the applications in document clustering, as well as on the datasets with much larger size and dimensionality, further demonstrate the effectiveness, efficiency, and scalability of our proposed models.
Can Wang 0004, Chihung Chi, Zhong She, Longbing Cao, Bela Stantic
ACM Trans. Knowl. Discov. Data5
2018 Syntax-Preserving Belief Change Operators for Logic Programs
abstract
Recent methods have adapted the well-established AGM and belief base frameworks for belief change to cover belief revision in logic programs. In this study here, we present two new sets of belief change operators for logic programs. They focus on preserving the explicit relationships expressed in the rules of a program, a feature that is missing in purely semantic approaches that consider programs only in their entirety. In particular, operators of the latter class fail to satisfy preservation and support, two important properties for belief change in logic programs required to ensure intuitive results. We address this shortcoming of existing approaches by introducing partial meet and ensconcement constructions for logic program belief change, which allow us to define syntax-preserving operators for satisfying preservation and support. Our work is novel in that our constructions not only preserve more information from a logic program during a change operation than existing ones, but they also facilitate natural definitions of contraction operators, the first in the field to the best of our knowledge. To evaluate the rationality of our operators, we translate the revision and contraction postulates from the AGM and belief base frameworks to the logic programming setting. We show that our operators fully comply with the belief base framework and formally state the interdefinability between our operators. We further compare our approach to two state-of-the-art logic program revision methods and demonstrate that our operators address the shortcomings of one and generalise the other method.
Sebastian Binnewies, Zhiqiang Zhuang, Kewen Wang 0001, Bela Stantic
ACM Trans. Comput. Log.4
2017 Multi-View Correlated Feature Learning by Uncovering Shared Component
abstract
Learning multiple heterogeneous features from different data sources is challenging. One research topic is how to exploit and utilize the correlations among various features across multiple views with the aim of improving the performance of learning tasks, such as classification. In this paper, we propose a new multi-view feature learning algorithm that simultaneously analyzes features from different views. Compared to most of the existing subspace learning methods that only focus on exploiting a shared latent subspace, our algorithm not only learns individual information in each view but also captures feature correlations among multiple views by learning a shared component. By assuming that such a component is shared by all views, we simultaneously exploit the shared component and individual information of each view in a batch mode. Since the objective function is non-smooth and difficult to solve, we propose an efficient iterative algorithm for optimization with guaranteed convergence. Extensive experiments are conducted on several benchmark datasets. The results demonstrate that our proposed algorithm performs better than all the compared multi-view learning algorithms.
Xiaowei Xue, Feiping Nie 0001, Sen Wang 0001, Xiaojun Chang, Bela Stantic
AAAI5
2017 Answering Temporal Analytic Queries over Big Data Based on Precomputing Architecture
Nigel Franciscus, Xuguang Ren, Bela Stantic
ACIIDS (1)3
2017 Computing Influence of a Product through Uncertain Reverse Skyline
abstract
Understanding the influence of a product is crucially important for making informed business decisions. This paper introduces a new type of skyline queries, called uncertain reverse skyline, for measuring the influence of a probabilistic product in uncertain data settings. More specifically, given a dataset of probabilistic products P and a set of customers C, an uncertain reverse skyline of a probabilistic product q retrieves all customers c ∈ C which include q as one of their preferred products. We present efficient pruning ideas and techniques for processing the uncertain reverse skyline query of a probabilistic product using R-Tree data index. We also present an efficient parallel approach to compute the uncertain reverse skyline and influence score of a probabilistic product. Our approach significantly outperforms the baseline approach derived from the existing literature. The efficiency of our approach is demonstrated by conducting experiments with both real and synthetic datasets.
Md. Saiful Islam 0003, Wenny Rahayu, Chengfei Liu, Tarique Anwar, Bela Stantic
SSDBM5
2016 A Comprehensive Approach to 'Now' in Temporal Relational Databases: Semantics and Representation
abstract
Now-related temporal data play an important role in many applications. Clifford et al.'s approach is a milestone to model the semantics of `now' in temporal relational databases. Several relational representation models for now-related data have been presented; however, the semantics of such representations has not been explicitly studied. Additionally, the definition of a relational algebra to query now-related data is an open problem. We propose the first integrated approach that provides both a neat semantics for now-related data and a compact 1NF representation (data model and relational algebra) for them. Additionally, our approach also extends current approaches to consider (i) domains where it is not always possible to know when changes in the world are recorded in the database and (ii) now-related data with a bound on their persistency in the future. To do so, we explicitly model the notion of temporal indeterminacy in the future for now-related data. The properties of our approach are also analyzed both from a theoretical (semantic correctness and reducibility of the algebra) and from an experimental point of view. Experiments show that, despite the fact that our approach is a major extension to current temporal relational approaches, no significant overhead is added to deal with `now'.
Luca Anselma, Luca Piovesan, Abdul Sattar 0001, Bela Stantic, Paolo Terenziani
IEEE Trans. Knowl. Data Eng.4
2015 A General Approach to Represent and Query Now-Relative Medical Data in Relational Databases
Luca Anselma, Luca Piovesan, Abdul Sattar 0001, Bela Stantic, Paolo Terenziani
AIME4
2015 DDIG-in: detecting disease-causing genetic variations due to frameshifting indels and nonsense mutations employing sequence and structural properties at nucleotide and protein levels
abstract
Abstract Motivation: Frameshifting (FS) indels and nonsense (NS) variants disrupt the protein-coding sequence downstream of the mutation site by changing the reading frame or introducing a premature termination codon, respectively. Despite such drastic changes to the protein sequence, FS indels and NS variants have been discovered in healthy individuals. How to discriminate disease-causing from neutral FS indels and NS variants is an understudied problem. Results: We have built a machine learning method called DDIG-in (FS) based on real human genetic variations from the Human Gene Mutation Database (inherited disease-causing) and the 1000 Genomes Project (GP) (putatively neutral). The method incorporates both sequence and predicted structural features and yields a robust performance by 10-fold cross-validation and independent tests on both FS indels and NS variants. We showed that human-derived NS variants and FS indels derived from animal orthologs can be effectively employed for independent testing of our method trained on human-derived FS indels. DDIG-in (FS) achieves a Matthews correlation coefficient (MCC) of 0.59, a sensitivity of 86%, and a specificity of 72% for FS indels. Application of DDIG-in (FS) to NS variants yields essentially the same performance (MCC of 0.43) as a method that was specifically trained for NS variants. DDIG-in (FS) was shown to make a significant improvement over existing techniques. Availability and implementation: The DDIG-in web-server for predicting NS variants, FS indels, and non-frameshifting (NFS) indels is available at http://sparks-lab.org/ddig. Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online.
Lukas Folkman, Yuedong Yang, Zhixiu Li, Bela Stantic, Abdul Sattar 0001, Matthew E. Mort, David N. Cooper, Yaoqi Zhou
Bioinform.4
2014 Possible Routes on a Highway of eLearning - Promising Architecture for eLearning Systems
Mirjana Ivanovic, Zoran Putnik, Dejan Mitrovic, Bela Stantic
KES-AMSTA4
2013 Periodic Data, Burden or Convenience
Bela Stantic
ADBIS1
2013 A new operator for efficient stream-relation join processing in data streaming engines
abstract
In the last decade, Stream Processing Engines (SPEs) have emerged as a new processing paradigm that can process huge amounts of data while retaining low latency and high-throughputs. Yet, it is often necessary to join streaming data with traditional databases to provide more contextual information for the end-users and applications. The major problem that we confront is to join the fast arriving stream tuples with the static relation tuples that are on a slow database. This is what we call the Stream-Relation Join (SRJ) problem. Currently, SPEs use a naive tuple-by-tuple approach for SRJ processing where the SPE accesses the database for every incoming tuple. Some SPEs use cache to avoid accessing the database for every incoming tuple, while others do not because of the stochastic nature of streaming data. In this paper, we propose a new SRJ operator to facilitate SRJ processing regardless of the cache performance using two techniques: batching and out-of-order processing. The proposed operator provides an effective generic solution to the SRJ problem and the cost of incorporating our operator into different SPEs is minimal. Our experiments use a variety of synthetic and real data sets demonstrating that our operator outperforms the state-of-the-art tuple-by-tuple approach in terms of maximizing the throughput under ordering and memory constraints.
Roozbeh Derakhshan, Abdul Sattar 0001, Bela Stantic
CIKM3
2013 Sequence-only evolutionary and predicted structural features for the prediction of stability changes in protein mutants
abstract
BACKGROUND: Even a single amino acid substitution in a protein sequence may result in significant changes in protein stability, structure, and therefore in protein function as well. In the post-genomic era, computational methods for predicting stability changes from only the sequence of a protein are of importance. While evolutionary relationships of protein mutations can be extracted from large protein databases holding millions of protein sequences, relevant evolutionary features for the prediction of stability changes have not been proposed. Also, the use of predicted structural features in situations when a protein structure is not available has not been explored. RESULTS: We proposed a number of evolutionary and predicted structural features for the prediction of stability changes and analysed which of them capture the determinants of protein stability the best. We trained and evaluated our machine learning method on a non-redundant data set of experimentally measured stability changes. When only the direction of the stability change was predicted, we found that the best performance improvement can be achieved by the combination of the evolutionary features mutation likelihood and SIFT score in conjunction with the predicted structural feature secondary structure. The same two evolutionary features in the combination with the predicted structural feature accessible surface area achieved the lowest error when the prediction of actual values of stability changes was assessed. Compared to similar studies, our method achieved improvements in prediction performance. CONCLUSION: Although the strongest feature for the prediction of stability changes appears to be the vector of amino acid identities in the sequential neighbourhood of the mutation, the most relevant combination of evolutionary and predicted structural features further improves prediction performance. Even the predicted structural features, which did not perform well on their own, turn out to be beneficial when appropriately combined with evolutionary features. We conclude that a high prediction accuracy can be achieved knowing only the sequence of a protein when the right combination of both structural and evolutionary features is used.
Lukas Folkman, Bela Stantic, Abdul Sattar 0001
BMC Bioinform.2
2013 Querying now-relative data
Luca Anselma, Bela Stantic, Paolo Terenziani, Abdul Sattar 0001
J. Intell. Inf. Syst.2
2013 An intensional approach for periodic data in relational databases
Paolo Terenziani, Bela Stantic, Alessio Bottrighi, Abdul Sattar 0001
J. Intell. Inf. Syst.2
2013 Minimising collisions in RFID data streams using probabilistic Cluster-Based Technique
Prapassara Pupunwiwat, Bela Stantic
Wirel. Networks2
2012 Towards Real Intelligent Web Exploration
Pavel Kalinov, Abdul Sattar 0001, Bela Stantic
APWeb3
2012 Refining Genetic Algorithm twin removal for high-resolution protein structure prediction
abstract
To gain a better understanding of how proteins function a process known as protein structure prediction (PSP) is carried out. However, experimental PSP methods, such as X-ray crystallography and Nuclear Magnetic Resonance (NMR), can be time-consuming and inaccurate. This has given rise to numerous computational PSP approaches to try and elicit a protein's three-dimensional conformation. A popular PSP search strategy is Genetic Algorithms (GA). GAs allow for a generic search approach, which can provide a generic improvement to alleviate the need to redefine the search strategies for separate sequences. Though GA's working principles are remarkable, a serious problem that is inherent in the GA search process is the growth of twins or identical chromosomes. Therefore, enhanced twin removal strategies are crucial for any GA search solving hard-optimisation problems like PSP. In this paper we explain our high-resolution GA feature-based resampling PSP approach and propose a twin removal strategy to further enhance its prediction accuracy. This includes investigating the optimal chromosome correlation factor (CCF) for our approach and defining a pre-built structure library for twin removal. We have also compared our GA approach with the popular Monte Carlo (MC) method for PSP. Our results indicate that out of all the CCF values we tested a CCF value of 0.8 provided the best level of diversity within our GA population. It also generated, on average, more native-like structures than any of the other CCF values, and clearly demonstrated that twin removal is needed in PSP when using GAs to obtain more accurate results.
Trent Higgs, Bela Stantic, Tamjidul Hoque, Abdul Sattar 0001
IEEE Congress on Evolutionary Computation2
2012 An implicit approach to deal with periodically repeated medical data
Bela Stantic, Paolo Terenziani, Guido Governatori, Alessio Bottrighi, Abdul Sattar 0001
Artif. Intell. Medicine1
2012 X-CleLo: intelligent deterministic RFID data and event transformer
Peter Darcy, Bela Stantic, Abdul Sattar 0001
Pers. Ubiquitous Comput.2
2011 A Novel Integrated Classifier for Handling Data Warehouse Anomalies
Peter Darcy, Bela Stantic, Abdul Sattar 0001
ADBIS2
2011 Variable Granularity Space Filling Curve for Indexing Multidimensional Data
Justin Terry, Bela Stantic, Paolo Terenziani, Abdul Sattar 0001
ADBIS2
2011 Generic Parallel Genetic Algorithm Framework for Protein Optimisation
Lukas Folkman, Wayne J. Pullan, Bela Stantic
ICA3PP (2)3
2011 An intelligent approach to handle False-Positive Radio Frequency Identification Anomalies
abstract
Radio Frequency Identification (RFID) technology allows wireless interaction between tagged objects and readers to automatically identify large groups of items. This technology is widely accepted in a number of application domains, however, it suffers from data anomalies such as false-positive obse rvations. Existing methods, such as manual tools, user specified rules and filtering algorithms, lack the automation and intelligence to effectively remove ambiguous false-positive readings. In this paper, we propose a methodology which incorporates a highly intelligent feature set definition utilised in conjunction with various state-of-the-art classifying techniques to correctly determine if a reading flagged as a potential false-positive anomaly should be discarded. Through experimental study we have shown that our approach cleans highly ambiguous false-positive observational data effectively. We have also discovered that the Non-Monotonic Reasoning classifier obtained the highest cleaning rate when handling false-positive RFID readings.
Peter Darcy, Bela Stantic, Abdul Sattar 0001
Intell. Data Anal.2
2010 Correcting Missing Data Anomalies with Clausal Defeasible Logic
Peter Darcy, Bela Stantic, Abdul Sattar 0001
ADBIS2
2010 Indexing Temporal Data with Virtual Structure
Bela Stantic, Justin Terry, Rodney W. Topor, Abdul Sattar 0001
ADBIS1
2010 Genetic algorithm feature-based resampling for protein structure prediction
abstract
Proteins carry out the majority of functionality on a cellular level. Computational protein structure prediction (PSP) methods have been introduced to speed up the PSP process due to manual methods, like nuclear magnetic resonance (NMR) and x-ray crystallography (XC) taking numerous months even years to produce a predicted structure for a target protein. A lot of work in this area is focused on the type of search strategy to employ. Two popular methods in the literature are: Monte Carlo based algorithms and Genetic Algorithms. Genetic Algorithms (GA) have proven to be quite useful in the PSP field, as they allow for a generic search approach, which alleviates the need to redefine the search strategies for separate sequences. They also lend themselves well to feature-based resampling techniques. Feature-based resampling works by taking previously computed local minima and combining features from them to create new structures that are more uniformly low in free energy. In this work we present a feature-based resampling genetic algorithm to refine structures that are outputted by PSP software. Our results indicate that our approach performs well, and produced an average 9.5% root mean square deviation (RMSD) improvement and a 17.36% template modeling score (TM-Score) improvement.
Trent Higgs, Bela Stantic, Tamjidul Hoque, Abdul Sattar 0001
IEEE Congress on Evolutionary Computation2
2010 Multiagent Based Scheduling of Elective Surgery
Sankalp Khanna, Timothy William Cleaver, Abdul Sattar 0001, David P. Hansen, Bela Stantic
PRIMA5
2010 An Intelligent Approach to Surgery Scheduling
Sankalp Khanna, Abdul Sattar 0001, Justin R. Boyle, David P. Hansen, Bela Stantic
PRIMA5
2010 Let's Trust Users It is Their Search
abstract
The current search engine model considers users not trustworthy, so no tools are provided to let them specify what they are looking for or in what context, which severely limits what they are able to achieve. Instead, search engines try to guess that, which is currently done using "implicit feedback''. In this paper we propose a "web exploration engine'' - a model where users can use the search engine as their tool and explicitly specify the context of their search. Information about the web has been pre-classified in a large number of categories; users can explore this hierarchy by providing relevance feedback or search within a particular category. Search is truly ``local'' in the sense that keyword relevance is not global, but specific to the category. In contrast to using a search engine, users can guide the exploration engine with relevance feedback alone without entering keywords.
Pavel Kalinov, Bela Stantic, Abdul Sattar 0001
Web Intelligence2
2009 The POINT approach to represent now in bitemporal databases
Bela Stantic, Abdul Sattar 0001, Paolo Terenziani
J. Intell. Inf. Syst.1
2008 Coping efficiently with now-relative medical data
Bela Stantic, Paolo Terenziani, Abdul Sattar 0001
AMIA1
2008 Parallel Simulated Annealing for Materialized View Selection in Data Warehousing Environments
Roozbeh Derakhshan, Bela Stantic, Othmar Korn, Frank Dehne
ICA3PP2
2003 A Novel Approach to Model NOW in Temporal Databases
abstract
In bitemporal databases, current facts and transaction states are modeled using a special value to represent the current time (such as a minimum or maximum timestamp or NULL). Previous studies indicate that the choice of value for now (i.e. the current time) significantly influences the efficiency of accessing bitemporal data. This paper introduces a new approach to represent now, in which current tuples and facts are represented as points on the transaction time and valid time line respectively. This allows us to exploit the computational advantages of point-based query languages. Via an empirical study, we demonstrate that our new approach to representing now offers considerable performance benefits over existing techniques for accessing bitemporal data.
Bela Stantic, John Thornton 0001, Abdul Sattar 0001
TIME1