EDBT 2026 Demo / reviewers in the wild / expert
Bela Stantic
dblp:74/1636
· DBLP profile ↗
36ranked-venue papers in the field
3as first author
5since 2021 · last 2025
0000-0003-0475-7951ORCID · reported
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 24 (2 first)Data Mining & Knowledge Discovery · 4Knowledge Engineering, Semantic Web & Information Systems · 3 (1 first)Information Retrieval & Web Search · 2Other / Interdisciplinary · 2Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Prompts De-Biasing Augmentation to Mitigate Gender Stereotypes in Large Language Models
Jinyuan Chen, Sebastian Binnewies, Bela Stantic |
ACIIDS (1) | 3 |
| 2025 | Integration of Dynamic Window Sizing with Neural Network Architectures for Real-Time Cryptocurrency Predictions
David L. John, Sebastian Binnewies, Bela Stantic |
ACIIDS (2) | 3 |
| 2024 | Identifying Optimal Window Size Configurations for Big Data Time Series ForecastingabstractOptimal window sizing in time series forecasting emerges as a pivotal factor for enhancing predictive accuracy, particularly in the volatile cryptocurrency market. While traditional models often rely on static window sizes, resulting in compromised forecasting performance, this research explores optimal window configurations across various market volatilities. By employing a hybrid Long Short-Term Memory and Gated Recurrent Unit (LSTM-GRU) model, the study systematically identifies the most effective window sizes for high, medium, and low volatility conditions. Results demonstrate that smaller windows are preferable in highly volatile environments to capture rapid market shifts, whereas larger windows are more suitable for stable conditions to incorporate a broader historical context. By identifying the predetermined optimal window sizes for each volatility segment, this study offers valuable insights for researchers aiming to enhance the adaptability and efficacy of predictive models. These results are especially useful for exploring dynamic window sizing techniques across various domains, particularly in fields where data volatility significantly impacts model performance. David L. John, Sebastian Binnewies, Bela Stantic |
IEEE Big Data | 3 |
| 2022 | Machine Learning or Lexicon Based Sentiment Analysis Techniques on Social Media Posts
David L. John, Bela Stantic |
ACIIDS (2) | 2 |
| 2021 | Empirical Study of Tweets Topic Classification Using Transformer-Based Language Models
Ranju Mandal, Susanne Becken, Bela Stantic |
ACIIDS | 4 |
| 2019 | Word Mover's Distance for Agglomerative Short Text Clustering
Nigel Franciscus, Xuguang Ren, Junhu Wang, Bela Stantic |
ACIIDS (1) | 4 |
| 2019 | Event Prediction Based on Causality Reasoning
Xuguang Ren, Nigel Franciscus, Junhu Wang, Bela Stantic |
ACIIDS (1) | 5 |
| 2019 | Top-N Hashtag Prediction via Coupling Social Influence and Homophily
Can Wang 0004, Yunwei Zhao, Chihung Chi, Willem-Jan van den Heuvel, Kwok-Yan Lam, Bela Stantic |
ADMA | 7 |
| 2019 | Mining Summary of Short Text with Centroid Similarity Distance
Nigel Franciscus, Junhu Wang, Bela Stantic |
ADMA | 3 |
| 2019 | Handling probabilistic integrity constraints in pay-as-you-go reconciliation of data modelsabstractData models capture the structure and characteristic properties of data entities, e.g., in terms of a database schema or an ontology. They are the backbone of diverse applications, reaching from information integration , through peer-to-peer systems and electronic commerce to social networking . Many of these applications involve models of diverse data sources. Effective utilisation and evolution of data models, therefore, calls for matching techniques that generate correspondences between their elements. Various such matching tools have been developed in the past. Yet, their results are often incomplete or erroneous, and thus need to be reconciled, i.e., validated by an expert. This paper analyses the reconciliation process in the presence of large collections of data models, where the network induced by generated correspondences shall meet consistency expectations in terms of integrity constraints. We specifically focus on how to handle data models that show some internal structure and potentially differ in terms of their assumed level of abstraction. We argue that such a setting calls for a probabilistic model of integrity constraints, for which satisfaction is preferred, but not required. In this work, we present a model for probabilistic constraints that enables reasoning on the correctness of individual correspondences within a network of data models, in order to guide an expert in the validation process. To support pay-as-you-go reconciliation, we also show how to construct a set of high-quality correspondences, even if an expert validates only a subset of all generated correspondences. We demonstrate the efficiency of our techniques for real-world datasets comprising database schemas and ontologies from various application domains. Nguyen Quoc Viet Hung, Matthias Weidlich 0001, Thanh Tam Nguyen, Zoltán Miklós 0001, Karl Aberer, Avigdor Gal, Bela Stantic |
Inf. Syst. | 7 |
| 2019 | From Anomaly Detection to Rumour Detection using Data Streams of Social PlatformsabstractSocial platforms became a major source of rumours. While rumours can have severe real-world implications, their detection is notoriously hard: Content on social platforms is short and lacks semantics; it spreads quickly through a dynamically evolving network; and without considering the context of content, it may be impossible to arrive at a truthful interpretation. Traditional approaches to rumour detection, however, exploit solely a single content modality, e.g., social media posts, which limits their detection accuracy. In this paper, we cope with the aforementioned challenges by means of a multi-modal approach to rumour detection that identifies anomalies in both, the entities (e.g., users, posts, and hashtags) of a social platform and their relations. Based on local anomalies, we show how to detect rumours at the network level, following a graph-based scan approach. In addition, we propose incremental methods, which enable us to detect rumours using streaming data of social platforms. We illustrate the effectiveness and efficiency of our approach with a real-world dataset of 4M tweets with more than 1000 rumours. Thanh Tam Nguyen, Matthias Weidlich 0001, Bolong Zheng, Hongzhi Yin, Nguyen Quoc Viet Hung, Bela Stantic |
Proc. VLDB Endow. | 6 |
| 2019 | User Guidance for Efficient Fact CheckingabstractThe Web constitutes a valuable source of information. In recent years, it fostered the construction of large-scale knowledge bases, such as Freebase, YAGO, and DBpedia. The open nature of the Web, with content potentially being generated by everyone, however, leads to inaccuracies and misinformation. Construction and maintenance of a knowledge base thus has to rely on fact checking, an assessment of the credibility of facts. Due to an inherent lack of ground truth information, such fact checking cannot be done in a purely automated manner, but requires human involvement. In this paper, we propose a comprehensive framework to guide users in the validation of facts, striving for a minimisation of the invested effort. Our framework is grounded in a novel probabilistic model that combines user input with automated credibility inference. Based thereon, we show how to guide users in fact checking by identifying the facts for which validation is most beneficial. Moreover, our framework includes techniques to reduce the manual effort invested in fact checking by determining when to stop the validation and by supporting efficient batching strategies. We further show how to handle fact checking in a streaming setting. Our experiments with three real-world datasets demonstrate the efficiency and effectiveness of our framework: A knowledge base of high quality, with a precision of above 90%, is constructed with only a half of the validation effort required by baseline techniques. Thanh Tam Nguyen, Hongzhi Yin, Matthias Weidlich 0001, Bolong Zheng, Nguyen Quoc Viet Hung, Bela Stantic |
Proc. VLDB Endow. | 6 |
| 2019 | Efficient User Guidance for Validating Participatory Sensing DataabstractParticipatory sensing has become a new data collection paradigm that leverages the wisdom of the crowd for big data applications without spending cost to buy dedicated sensors. It collects data from human sensors by using their own devices such as cell phone accelerometers, cameras, and GPS devices. This benefit comes with a drawback: human sensors are arbitrary and inherently uncertain due to the lack of quality guarantee. Moreover, participatory sensing data are time series that exhibit not only highly irregular dependencies on time but also high variance between sensors. To overcome these limitations, we formulate the problem of validating uncertain time series collected by participatory sensors. In this article, we approach the problem by an iterative validation process on top of a probabilistic time series model. First, we generate a series of probability distributions from raw data by tailoring a state-of-the-art dynamical model, namely Generalised Auto Regressive Conditional Heteroskedasticity (GARCH), for our joint time series setting. Second, we design a feedback process that consists of an adaptive aggregation model to unify the joint probabilistic time series and an efficient user guidance model to validate aggregated data with minimal effort. Through extensive experimentation, we demonstrate the efficiency and effectiveness of our approach on both real data and synthetic data. Highlights from our experiences include the fast running time of a probabilistic model, the robustness of an aggregation model to outliers, and the significant effort saving of a guidance model. Thanh Cong Phan, Thanh Tam Nguyen, Hongzhi Yin, Bolong Zheng, Bela Stantic, Nguyen Quoc Viet Hung |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2018 | An Ensemble System with Random Projection and Dynamic Ensemble Selection
Anh Vu Luong, Tuyet-Trinh Vu, Nguyen Quoc Viet Hung, Tien Thanh Nguyen, Bela Stantic |
ACIIDS (1) | 6 |
| 2018 | Beyond Word-Cloud: A Graph Model Derived from Beliefs
Nigel Franciscus, Xuguang Ren, Bela Stantic |
ACIIDS (2) | 3 |
| 2018 | Automatic Image Region Annotation by Genetic Algorithm-Based Joint Classifier and Feature Selection in Ensemble System
Anh Vu Luong, Tien Thanh Nguyen, Xuan Cuong Pham, Thi Thu Thuy Nguyen, Alan Wee-Chung Liew, Bela Stantic |
ACIIDS (1) | 6 |
| 2018 | Experimental Clarification of Some Issues in Subgraph Isomorphism Algorithms
Xuguang Ren, Junhu Wang, Nigel Franciscus, Bela Stantic |
ACIIDS (2) | 4 |
| 2018 | "Quality" vs. "Readability" in Document Images: Statistical Analysis of Human PerceptionabstractBased on the hypothesis that a good / poor quality document image is most probably a readable / unreadable document, document image quality and readability have interchangeably been used in the literature. These two terms, however, have different meanings implying two different perspectives of looking at document images by human being. In document images, the level of quality and the degree of readability may have a relation / correlation considering human perception. However, to the best of our knowledge there is no specific study to characterise this relation and also validate the abovementioned hypothesis. In this work, at first, we created a dataset composed of mostly camera-based document images with various distortion levels. Each document image has then been assessed with regard to two different measures, the level of quality and the degree of readability, by different individuals. A detailed Normalised Cross Correlation analysis along with different statistical analysis based on Shapiro-Wilks and Wilcoxon tests has further been provided to demonstrate how document image quality and readability are linked. Our findings indicate that the quality and readability were somewhat different in terms of the population distributions. However, the correlation between quality and readability was 0.99, which implies document quality and readability are highly correlated based on human perception. Alireza Alaei, Romain Raveaux, Donatello Conte, Bela Stantic |
DAS | 4 |
| 2018 | A System for Spatial-Temporal Trajectory Data Integration and Representation
Douglas Alves Peixoto, Xiaofang Zhou 0001, Nguyen Quoc Viet Hung, Dan He 0009, Bela Stantic |
DASFAA (2) | 5 |
| 2018 | What-If Analysis with Conflicting Goals: Recommending Data Ranges for ExplorationabstractWhat-if analysis is a data-intensive exploration to inspect how changes in a set of input parameters of a model influence some outcomes. It is motivated by a user trying to understand the sensitivity of a model to a certain parameter in order to reach a set of goals that are defined over the outcomes. To avoid an exploration of all possible combinations of parameter values, efficient what-if analysis calls for a partitioning of parameter values into data ranges and a unified representation of the obtained outcomes per range. Traditional techniques to capture data ranges, such as histograms, are limited to one outcome dimension. Yet, in practice, what-if analysis often involves conflicting goals that are defined over different dimensions of the outcome. Working on each of those goals independently cannot capture the inherent trade-off between them. In this paper, we propose techniques to recommend data ranges for what-if analysis, which capture not only data regularities, but also the trade-off between conflicting goals. Specifically, we formulate a parametric data partitioning problem and propose a method to find an optimal solution for it. Targeting scalability to large datasets, we further provide a heuristic solution to this problem. By theoretical and empirical analyses, we establish performance guarantees in terms of runtime and result quality. Nguyen Quoc Viet Hung, Kai Zheng 0001, Matthias Weidlich 0001, Bolong Zheng, Hongzhi Yin, Thanh Tam Nguyen, Bela Stantic |
ICDE | 7 |
| 2018 | Concept for Evaluation of Techniques for Trajectory Distance MeasuresabstractMeasuring the similarity (or distance) between trajectories of moving objects is a common procedure taken by most trajectory data-driven applications. One of the biggest challenges of trajectory distances measurement is that the distance needs to be carefully defined in order to reflect the true underlying similarity. This is due to the fact that trajectories are essentially non-uniform sequential data with variable length, attached with both spatial and temporal attributes, which may or may not be considered for similarity measures. Therefore, tens of similarity measures for trajectory data have been proposed; every technique claim an advantage over the others in a different aspect. Hence, it's difficult for users to choose the best-suited technique, as well as the appropriate parameter values, since each technique has distinct performance and characteristics depending on various factors. In this paper, we develop an application that allows to evaluate several techniques in different aspects (accuracy, sensitivity to trajectory features, performance, etc.). We believe that this tool will be able to serve as a practical guideline for both researchers and developers. While researchers can use our tool to assess existing or new techniques, developers can reuse its components to reduce the development complexity. Douglas Alves Peixoto, Han Su 0001, Nguyen Quoc Viet Hung, Bela Stantic, Bolong Zheng, Xiaofang Zhou 0001 |
MDM | 4 |
| 2018 | Coupled Clustering Ensemble by Exploring Data InterdependenceabstractClustering ensembles combine multiple partitions of data into a single clustering solution. It is an effective technique for improving the quality of clustering results. Current clustering ensemble algorithms are usually built on the pairwise agreements between clusterings that focus on the similarity via consensus functions, between data objects that induce similarity measures from partitions and re-cluster objects, and between clusters that collapse groups of clusters into meta-clusters. In most of those models, there is a strong assumption on IIDness (i.e., independent and identical distribution), which states that base clusterings perform independently of one another and all objects are also independent. In the real world, however, objects are generally likely related to each other through features that are either explicit or even implicit. There is also latent but definite relationship among intermediate base clusterings because they are derived from the same set of data. All these demand a further investigation of clustering ensembles that explores the interdependence characteristics of data. To solve this problem, a new coupled clustering ensemble (CCE) framework that works on the interdependence nature of objects and intermediate base clusterings is proposed in this article. The main idea is to model the coupling relationship between objects by aggregating the similarity of base clusterings, and the interactive relationship among objects by addressing their neighborhood domains. Once these interdependence relationships are discovered, they will act as critical supplements to clustering ensembles. We verified our proposed framework by using three types of consensus function: clustering-based, object-based, and cluster-based. Substantial experiments on multiple synthetic and real-life benchmark datasets indicate thatCCEcan effectively capture the implicit interdependence relationships among base clusterings and among objects with higher clustering accuracy, stability, and robustness compared to 14 state-of-the-art techniques, supported by statistical analysis. In addition, we show that the final clustering quality is dependent on the data characteristics (e.g., quality and consistency) of base clusterings in terms of sensitivity analysis. Finally, the applications in document clustering, as well as on the datasets with much larger size and dimensionality, further demonstrate the effectiveness, efficiency, and scalability of our proposed models. Can Wang 0004, Chihung Chi, Zhong She, Longbing Cao, Bela Stantic |
ACM Trans. Knowl. Discov. Data | 5 |
| 2017 | Answering Temporal Analytic Queries over Big Data Based on Precomputing Architecture
Nigel Franciscus, Xuguang Ren, Bela Stantic |
ACIIDS (1) | 3 |
| 2017 | Computing Influence of a Product through Uncertain Reverse SkylineabstractUnderstanding the influence of a product is crucially important for making informed business decisions. This paper introduces a new type of skyline queries, called uncertain reverse skyline, for measuring the influence of a probabilistic product in uncertain data settings. More specifically, given a dataset of probabilistic products P and a set of customers C, an uncertain reverse skyline of a probabilistic product q retrieves all customers c ∈ C which include q as one of their preferred products. We present efficient pruning ideas and techniques for processing the uncertain reverse skyline query of a probabilistic product using R-Tree data index. We also present an efficient parallel approach to compute the uncertain reverse skyline and influence score of a probabilistic product. Our approach significantly outperforms the baseline approach derived from the existing literature. The efficiency of our approach is demonstrated by conducting experiments with both real and synthetic datasets. Md. Saiful Islam 0003, Wenny Rahayu, Chengfei Liu, Tarique Anwar, Bela Stantic |
SSDBM | 5 |
| 2016 | A Comprehensive Approach to 'Now' in Temporal Relational Databases: Semantics and RepresentationabstractNow-related temporal data play an important role in many applications. Clifford et al.'s approach is a milestone to model the semantics of `now' in temporal relational databases. Several relational representation models for now-related data have been presented; however, the semantics of such representations has not been explicitly studied. Additionally, the definition of a relational algebra to query now-related data is an open problem. We propose the first integrated approach that provides both a neat semantics for now-related data and a compact 1NF representation (data model and relational algebra) for them. Additionally, our approach also extends current approaches to consider (i) domains where it is not always possible to know when changes in the world are recorded in the database and (ii) now-related data with a bound on their persistency in the future. To do so, we explicitly model the notion of temporal indeterminacy in the future for now-related data. The properties of our approach are also analyzed both from a theoretical (semantic correctness and reducibility of the algebra) and from an experimental point of view. Experiments show that, despite the fact that our approach is a major extension to current temporal relational approaches, no significant overhead is added to deal with `now'. Luca Anselma, Luca Piovesan, Abdul Sattar 0001, Bela Stantic, Paolo Terenziani |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2013 | Periodic Data, Burden or Convenience
Bela Stantic |
ADBIS | 1 |
| 2013 | A new operator for efficient stream-relation join processing in data streaming enginesabstractIn the last decade, Stream Processing Engines (SPEs) have emerged as a new processing paradigm that can process huge amounts of data while retaining low latency and high-throughputs. Yet, it is often necessary to join streaming data with traditional databases to provide more contextual information for the end-users and applications. The major problem that we confront is to join the fast arriving stream tuples with the static relation tuples that are on a slow database. This is what we call the Stream-Relation Join (SRJ) problem. Currently, SPEs use a naive tuple-by-tuple approach for SRJ processing where the SPE accesses the database for every incoming tuple. Some SPEs use cache to avoid accessing the database for every incoming tuple, while others do not because of the stochastic nature of streaming data. In this paper, we propose a new SRJ operator to facilitate SRJ processing regardless of the cache performance using two techniques: batching and out-of-order processing. The proposed operator provides an effective generic solution to the SRJ problem and the cost of incorporating our operator into different SPEs is minimal. Our experiments use a variety of synthetic and real data sets demonstrating that our operator outperforms the state-of-the-art tuple-by-tuple approach in terms of maximizing the throughput under ordering and memory constraints. Roozbeh Derakhshan, Abdul Sattar 0001, Bela Stantic |
CIKM | 3 |
| 2013 | Querying now-relative data
Luca Anselma, Bela Stantic, Paolo Terenziani, Abdul Sattar 0001 |
J. Intell. Inf. Syst. | 2 |
| 2013 | An intensional approach for periodic data in relational databases
Paolo Terenziani, Bela Stantic, Alessio Bottrighi, Abdul Sattar 0001 |
J. Intell. Inf. Syst. | 2 |
| 2012 | Towards Real Intelligent Web Exploration
Pavel Kalinov, Abdul Sattar 0001, Bela Stantic |
APWeb | 3 |
| 2011 | A Novel Integrated Classifier for Handling Data Warehouse Anomalies
Peter Darcy, Bela Stantic, Abdul Sattar 0001 |
ADBIS | 2 |
| 2011 | Variable Granularity Space Filling Curve for Indexing Multidimensional Data
Justin Terry, Bela Stantic, Paolo Terenziani, Abdul Sattar 0001 |
ADBIS | 2 |
| 2010 | Correcting Missing Data Anomalies with Clausal Defeasible Logic
Peter Darcy, Bela Stantic, Abdul Sattar 0001 |
ADBIS | 2 |
| 2010 | Indexing Temporal Data with Virtual Structure
Bela Stantic, Justin Terry, Rodney W. Topor, Abdul Sattar 0001 |
ADBIS | 1 |
| 2010 | Let's Trust Users It is Their SearchabstractThe current search engine model considers users not trustworthy, so no tools are provided to let them specify what they are looking for or in what context, which severely limits what they are able to achieve. Instead, search engines try to guess that, which is currently done using "implicit feedback''. In this paper we propose a "web exploration engine'' - a model where users can use the search engine as their tool and explicitly specify the context of their search. Information about the web has been pre-classified in a large number of categories; users can explore this hierarchy by providing relevance feedback or search within a particular category. Search is truly ``local'' in the sense that keyword relevance is not global, but specific to the category. In contrast to using a search engine, users can guide the exploration engine with relevance feedback alone without entering keywords. Pavel Kalinov, Bela Stantic, Abdul Sattar 0001 |
Web Intelligence | 2 |
| 2009 | The POINT approach to represent now in bitemporal databases
Bela Stantic, Abdul Sattar 0001, Paolo Terenziani |
J. Intell. Inf. Syst. | 1 |