VLDB 2026 Research / reviewers in the wild / expert
Veselka Boeva
dblp:68/648
· DBLP profile ↗
28ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0003-3128-191XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Theory of computation · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An evidence-based neuro-symbolic framework for ambiguous image scene classificationabstractIn this study, we propose a novel neuro-symbolic approach to deal with the inherent ambiguity in image scene classification, combining the usage of pre-trained deep learning (DL) models with concepts from modal logic and evidence theory. The DL models are used to detect objects and estimate their depth in a set of labeled images. The obtained outputs are employed to form a dataset of instances characterizing the possible classes. Subsequently, a multi-valued mapping is defined between the data instances and the considered images resulting into each image being represented by the set of instances associated with it. The obtained mapping is utilized to infer necessity and possibility conditions of each class, or equivalently its upper (plausibility) and lower (belief) probabilities. Based on these interval evaluations, a rule-based and a score-based classifiers are built. The overall method is explainable and directly interpretable, robust to data scarcity and data imbalance. The presented framework is studied and evaluated on an abandoned bag detection use case. Giulia Murtas, Veselka Boeva, Elena Tsiporkova |
NeSy | 2 |
| 2025 | FedCluLearn: Federated Continual Learning Using Stream Micro-cluster Indexing Scheme
Milena Angelova, Veselka Boeva, Shahrooz Abghari, Selim Ickin, Xiaoyu Lan |
ECML/PKDD (2) | 2 |
| 2025 | Exploring Dynamic Hypergraphs for Clustering Analysis of District Heating DataabstractIn the District Heating (DH) sector, the analysis and monitoring of data from DH substations is crucial to keeping the entire DH network running efficiently. Clustering of DH substations based on multivariate data helps analyze their behavior over time. In this context, a visualization-based analysis approach can be particularly beneficial. In this paper, we explore the use of dynamic hypergraph visualization to analyze the clustering results of DH network substations over time. We present the initial results of designing and implementing a visual analytics dashboard that supports DH experts in analyzing different behaviors of DH substations. In the proposed dashboard, we adopt the Parallel Aggregated Ordered Hypergraph (PAOH) technique to visualize dynamic hypergraphs, which provides a compact visualization of multivariate data clustering over time. Moreover, we include additional views with complementary visualizations supporting the analysis and understanding of the dynamic hypergraph main view. We showcase the capability of our dashboard applied on a real DH dataset. Valeria Garro, Ilir Jusufi, Shahrooz Abghari, Jens Brage, Veselka Boeva |
VINCI | 5 |
| 2025 | Contribution prediction in federated learning via client behavior evaluationabstractFederated learning (FL), a decentralized machine learning framework that allows edge devices (i.e., clients) to train a global model while preserving data/client privacy, has become increasingly popular recently. In FL, a shared global model is built by aggregating the updated parameters in a distributed manner. To incentivize data owners to participate in FL, it is essential for service providers to fairly evaluate the contribution of each data owner to the shared model during the learning process. To the best of our knowledge, most existing solutions are resource-demanding and usually run as an additional evaluation procedure. The latter produces an expensive computational cost for large data owners. In this paper, we present simple and effective FL solutions that show how the clients’ behavior can be evaluated during the training process with respect to reliability, and this is demonstrated for two existing FL models, Cluster Analysis-based Federated Learning (CA-FL) and Group-Personalized FL (GP-FL), respectively. In the former model, CA-FL, the frequency of each client to be selected as a cluster representative and in that way to be involved in the building of the shared model is assessed. This can eventually be considered as a measure of the respective client data reliability. In the latter model, GP-FL, we calculate how many times each client changes a cluster it belongs to during FL training, which can be interpreted as a measure of the client’s unstable behavior, i.e., it can be considered as not very reliable. We validate our FL approaches on three LEAF datasets and benchmark their performance to two baseline contribution evaluation approaches. The experimental results demonstrate that by applying the two FL models we are able to get robust evaluations of clients’ behavior during the training process. These evaluations can be used for further studying, comparing, understanding, and eventually predicting clients’ contributions to the shared global model. • New federated learning solutions for quantifying the contribution of each client to the overall global model according to its behavior. • The clients’ behavior can be evaluated during the training process concerning reliability. • Without incurring any considerable communication costs, our methods have shown robustness in evaluating each client’s behavior respectively contribution to the federated model and have the potential to be used for this purpose. Ahmed Abbas Mohsin Al-Saedi, Veselka Boeva, Emiliano Casalicchio |
Future Gener. Comput. Syst. | 2 |
| 2024 | Multi-layered Clustering for Context-aware Monitoring of District Heating NetworkabstractIn this study, we propose to explore multi-layered clustering to provide a context-aware data analytics tool for monitoring the network behavior of subsystems, such as a district heating (DH) network. Multi-layer clustering, in contrast to multi-view clustering, does not assume conditional independence of layers. The main idea of our approach is based on the integration of clustering models produced by considering different perspectives that capture information about the monitored subsystems’ operational behavior or performance as well as their contextual environment. The initial clustering layer can reflect a static context, which is important for the subsystems’ performance. It will be used as a base on which clustering models produced with respect to other analyzed operational characteristics and contexts will be layered. This will facilitate analysis and comparison of the subsystems’ behavior in two comparable time periods and, eventually, identification of deviations that need attention. The proposed approach is evaluated and validated in a use case from the DH domain. The multi-layered clustering is applied and demonstrated to be robust for continuous context-aware analysis of the performance of a network of DH substations. Veselka Boeva, Shahrooz Abghari, Vishnu Manasa Devagiri, Jens Brage |
IEEE Big Data | 1 |
| 2024 | Interpretable Data-Driven Risk Assessment in Support of Predictive Maintenance of a Large Portfolio of Industrial VehiclesabstractIn this study, we propose a data-driven survival risk analysis approach in support of predictive maintenance management of a large portfolio of industrial assets. The concrete use case considered is a large portfolio of industrial vehicles (trucks). However, the approach is generic (i.e., asset-type agnostic) in nature and can be applied in different industrial contexts. It is able to employ different data sources in the risk analysis workflow, e.g., time series operation data collected via a multitude of sensor measurements combined with tabular data recording the technical specifications of the assets (vehicles). Subsequently, several different risk assessment strategies can be considered: 1) operation-related risk at each time step for any asset computed on the operation data across the whole portfolio; 2) the failure predisposition of each asset determined by its technical specification; 3) hybrid risk analysis, which innovatively combines the different data types to estimate overall risk at any time in the future for any asset. Our validation, conducted on real-world data, demonstrates that the hybrid approach provides a realistic temporal risk assessment during vehicle operation that also reflects adequately the inherent (contextual) risk predisposition of the vehicle due its technical specification. The proposed approach derives diverse survival risk estimations, which are interpretable by design and in this way facilitate both prognostic health monitoring and root cause analysis of the factors impacting vehicles’ risk of failure. Fabian Fingerhut, Elena Tsiporkova, Veselka Boeva |
IEEE Big Data | 3 |
| 2024 | Putting Sense into Incomplete Heterogeneous Data with Hypergraph Clustering Analysis
Vishnu Manasa Devagiri, Pierre Dagnely, Veselka Boeva, Elena Tsiporkova |
IDA (2) | 3 |
| 2023 | Mitigating Concept Drift in Distributed Contexts with Dynamic Repository of Federated ModelsabstractThis paper proposes a novel federated learning methodology, called FedRepo, that copes with concept drift issues in a statistically heterogeneous distributed learning environment. The proposed horizontal federated learning methodology, based on random forest (RF), can be used for collaborative training and maintenance of a dynamic repository of federated RF models, each one customized to a group of clients/devices. The clients are grouped together if their performance patterns with respect to the global RF model are similar. The performance of the customized RF global models is continuously monitored during the inference phase and the repository is accordingly adapted to mitigate the detected concept drift. The proposed methodology is studied and evaluated against an electricity consumption forecasting use case. The evaluation results demonstrate clearly that the proposed methodology is able to deal with concept drift issues in an efficient and adequate fashion without compromising the overall performance of the distributed environment. Elena Tsiporkova, Michiel De Vis, Sarah Klein, Anna Hristoskova, Veselka Boeva |
IEEE Big Data | 5 |
| 2023 | Group-Personalized Federated Learning for Human Activity Recognition Through Cluster Eccentricity Analysis
Ahmed Abbas Mohsin Al-Saedi, Veselka Boeva |
EANN | 2 |
| 2021 | Understanding Traffic Cruising Causation Via Parking Data EnhancementabstractThis study is devoted to understanding traffic cruising causation through exploring and enhancing parking data. Five recent (2017–2020) studies modeling parking congestion relied on occupancy as their only parking lot feature, then compared modeling techniques using this feature, to find the best performance. However, recently some computer scientists pointed out that it is more effective for the computer science community to focus more on data preparation for performance improvements, rather than exclusively comparing modeling techniques. This inspired us to add more parking lot features and evaluate them, to investigate how they should be composed into a congestion score, acting as a more accurate picture of reality. The score is then compared to the performance of a version where occupancy is the only parking lot feature. An experimental case study is designed in three parts. The first measures how the features should be summed into a score according to drivers’ expectations. The second analyzes how much data can be reused from the real data, and whether spatial or temporal comparisons are better for data synthesis of parking data. The third part compares the performance of the score against the occupancy-only version using k-means clustering algorithm and dynamic time warping distance. The experimental results show performance improvements in all spatial and temporal categories, and increasing improvement as the sample sizes grow. Mirza Jasarevic, Veselka Boeva, Fredrik Sjölin, Per-Olav Gramstad |
ICMLA | 2 |
| 2021 | Reducing Communication Overhead of Federated Learning through Clustering AnalysisabstractTraining of machine learning models in a Datacen-ter, with data originated from edge nodes, incurs high communication overheads and violates a user's privacy. These challenges may be tackled by employing Federated Learning (FL) machine learning technique to train a model across multiple decentralized edge devices (workers) using local data. In this paper, we explore an approach that identifies the most representative updates made by workers and those are only uploaded to the central server for reducing network communication costs. Based on this idea, we propose a FL model that can mitigate communication overheads via clustering analysis of the worker local updates. The Cluster Analysis-based Federated Learning (CA-FL) model is studied and evaluated in human activity recognition (HAR) datasets. Our evaluation results show the robustness of CA - FL in comparison with traditional FL in terms of accuracy and communication costs on both IID and non-IID cases. Ahmed Abbas Mohsin Al-Saedi, Veselka Boeva, Emiliano Casalicchio |
ISCC | 2 |
| 2020 | Multi-view Clustering Analyses for District Heating SubstationsabstractIn this study, we propose a multi-view clustering approach for mining and analysing multi-view network datasets. The proposed approach is applied and evaluated on a real-world scenario for monitoring and analysing district heating (DH) network conditions and identifying substations with sub-optimal behaviour. Initially, geographical locations of the substations are used to build an approximate graph representation of the DH network. Two different analyses can further be applied in this context: step-wise and parallel-wise multi-view clustering. The step-wise analysis is meant to sequentially consider and analyse substations with respect to a few different views. At each step, a new clustering solution is built on top of the one generated by the previously considered view, which organizes the substations in a hierarchical structure that can be used for multi-view comparisons. The parallel-wise analysis on the other hand, provides the opportunity to analyse substations with regards to two different views in parallel. Such analysis is aimed to represent and identify the relationships between substations by organizing them in a bipartite graph and analysing the substations’ distribution with respect to each view. The proposed data analysis and visualization approach arms domain experts with means for analysing DH network performance. In addition, it will facilitate the identification of substations with deviating operational behaviour based on comparative analysis with their closely located neighbours. Shahrooz Abghari, Veselka Boeva, Jens Brage, Håkan Grahn |
DATA | 2 |
| 2020 | Multi-view Data Mining Approach for Behaviour Analysis of Smart Control ValveabstractIn this study, we propose a multi-view data analysis approach that can be used for modelling and monitoring smart control valve system behaviour. The proposed approach consists of four distinctive steps: (i) multi-view interpretation of the available data attributes by separating them into several representations (views), e.g., operational parameters, contextual factors, and performance indicators; (ii) modelling different control valve system operating modes by clustering analyses of the operational data view; (iii) annotating each operating mode (cluster) by using the remaining views (i.e., contextual and system performance data); (iv) context-aware monitoring of the control valve system operating behaviour by applying the built model. In addition, the data points (daily profiles) observed during the monitoring can be annotated by comparing them with the known typical behavioural modes. This information can be further analysed and used for continuous updating and improvement of the model.The potential of the proposed approach has been evaluated and demonstrated on real-world sensor data originating from a company in the smart building domain. The obtained results show the robustness of the proposed approach in modelling, analysing, and monitoring the control valve system behaviour. Amirmohammad Eghbalian, Shahrooz Abghari, Veselka Boeva, Farhad Basiri |
ICMLA | 3 |
| 2020 | Split-Merge Evolutionary Clustering for Multi-View Streaming DataabstractIn this study, we propose a new multi-view stream clustering approach, called MV Split-Merge Clustering. The proposed approach is an extension of an existing split-merge evolutionary clustering algorithm (entitled Split-Merge Clustering) to multi-view data applications. The extended version can be used to integrate data from multiple views in a streaming manner and discover cluster structure for each data chunk. The MV Split-Merge Clustering can be applied for grouping distinct chunks of multi-view streaming data so that a global integrated clustering model is built on each data chunk. At each time window, an updated clustering solution (local model) is initially produced on each view of the current data chunk by applying the Split-Merge Clustering algorithm. Formal Concept Analysis is then used in order to integrate information from the multiple views (local clustering models) and generate a global model (formal concept lattice) that reveals the correlations among the clusters of the local models. The proposed MV Split-Merge Clustering has been initially evaluated on a publicly available data set. Our results show that the approach is able to identify a clustering structure and relationships among the different views comparable to those produced in a batch scenario. Vishnu Manasa Devagiri, Veselka Boeva, Elena Tsiporkova |
KES | 2 |
| 2019 | Higher Order Mining for Monitoring District Heating SubstationsabstractWe propose a higher order mining (HOM) approach for modelling, monitoring and analyzing district heating (DH) substations' operational behaviour and performance. HOM is concerned with mining over patterns rather than primary or raw data. The proposed approach uses a combination of different data analysis techniques such as sequential pattern mining, clustering analysis, consensus clustering and minimum spanning tree (MST). Initially, a substation's operational behaviour is modeled by extracting weekly patterns and performing clustering analysis. The substation's performance is monitored by assessing its modeled behaviour for every two consecutive weeks. In case some significant difference is observed, further analysis is performed by integrating the built models into a consensus clustering and applying an MST for identifying deviating behaviours. The results of the study show that our method is robust for detecting deviating and sub-optimal behaviours of DH substations. In addition, the proposed method can facilitate domain experts in the interpretation and understanding of the substations' behaviour and performance by providing different data analysis and visualization techniques. Shahrooz Abghari, Veselka Boeva, Jens Brage, Christian Johansson, Håkan Grahn, Niklas Lavesson |
DSAA | 2 |
| 2019 | A Split-Merge Evolutionary Clustering AlgorithmabstractIn this article we propose a bipartite correlation clustering technique that can be used to adapt the existing clustering solution to a clustering of newly collected data elements. The proposed tec ... Veselka Boeva, Milena Angelova, Elena Tsiporkova |
ICAART (2) | 1 |
| 2018 | Hoeffding Trees with Nmin AdaptationabstractMachine learning software accounts for a significant amount of energy consumed in data centers. These algorithms are usually optimized towards predictive performance, i.e. accuracy, and scalability. This is the case of data stream mining algorithms. Although these algorithms are adaptive to the incoming data, they have fixed parameters from the beginning of the execution. We have observed that having fixed parameters lead to unnecessary computations, thus making the algorithm energy inefficient. In this paper we present the nmin adaptation method for Hoeffding trees. This method adapts the value of the nmin parameter, which significantly affects the energy consumption of the algorithm. The method reduces unnecessary computations and memory accesses, thus reducing the energy, while the accuracy is only marginally affected. We experimentally compared VFDT (Very Fast Decision Tree, the first Hoeffding tree algorithm) and CVFDT (Concept-adapting VFDT) with the VFDT-nmin (VFDT with nmin adaptation). The results show that VFDT-nmin consumes up to 27% less energy than the standard VFDT, and up to 92% less energy than CVFDT, trading off a few percent of accuracy in a few datasets. Eva García Martín, Niklas Lavesson, Håkan Grahn, Emiliano Casalicchio, Veselka Boeva |
DSAA | 5 |
| 2018 | Evolutionary Clustering Techniques for Expertise Mining ScenariosabstractThe problem addressed in this article concerns the development of evolutionary clustering techniques that can be applied to adapt the existing clustering solution to a clustering of newly collected ... Veselka Boeva, Milena Angelova, Niklas Lavesson, Oliver Rosander, Elena Tsiporkova |
ICAART (2) | 1 |
| 2018 | A Minimum Spanning Tree Clustering Approach for Outlier Detection in Event SequencesabstractOutlier detection has been studied in many domains. Outliers arise due to different reasons such as mechanical issues, fraudulent behavior, and human error. In this paper, we propose an unsupervised approach for outlier detection in a sequence dataset. The proposed approach combines sequential pattern mining, cluster analysis, and a minimum spanning tree algorithm in order to identify clusters of outliers. Initially, the sequential pattern mining is used to extract frequent sequential patterns. Next, the extracted patterns are clustered into groups of similar patterns. Finally, the minimum spanning tree algorithm is used to find groups of outliers. The proposed approach has been evaluated on two different real datasets, i.e., smart meter data and video session data. The obtained results have shown that our approach can be applied to narrow down the space of events to a set of potential outliers and facilitate domain experts in further analysis and identification of system level issues. Shahrooz Abghari, Veselka Boeva, Niklas Lavesson, Håkan Grahn, Selim Ickin, Jörgen Gustafsson |
ICMLA | 2 |
| 2017 | Data-driven Techniques for Expert Finding
Veselka Boeva, Milena Angelova, Elena Tsiporkova |
ICAART (2) | 1 |
| 2014 | A method for evaluation of learning components
Niklas Lavesson, Veselka Boeva, Elena Tsiporkova, Paul Davidsson |
Autom. Softw. Eng. | 2 |
| 2014 | A Formal Concept Analysis Approach to Consensus Clustering of Multi-Experiment Expression DataabstractBACKGROUND: Presently, with the increasing number and complexity of available gene expression datasets, the combination of data from multiple microarray studies addressing a similar biological question is gaining importance. The analysis and integration of multiple datasets are expected to yield more reliable and robust results since they are based on a larger number of samples and the effects of the individual study-specific biases are diminished. This is supported by recent studies suggesting that important biological signals are often preserved or enhanced by multiple experiments. An approach to combining data from different experiments is the aggregation of their clusterings into a consensus or representative clustering solution which increases the confidence in the common features of all the datasets and reveals the important differences among them. RESULTS: We propose a novel generic consensus clustering technique that applies Formal Concept Analysis (FCA) approach for the consolidation and analysis of clustering solutions derived from several microarray datasets. These datasets are initially divided into groups of related experiments with respect to a predefined criterion. Subsequently, a consensus clustering algorithm is applied to each group resulting in a clustering solution per group.These solutions are pooled together and further analysed by employing FCA which allows extracting valuable insights from the data and generating a gene partition over all the experiments. In order to validate the FCA-enhanced approach two consensus clustering algorithms are adapted to incorporate the FCA analysis. Their performance is evaluated on gene expression data from multi-experiment study examining the global cell-cycle control of fission yeast. The FCA results derived from both methods demonstrate that, although both algorithms optimize different clustering characteristics, FCA is able to overcome and diminish these differences and preserve some relevant biological signals. CONCLUSIONS: The proposed FCA-enhanced consensus clustering technique is a general approach to the combination of clustering algorithms with FCA for deriving clustering solutions from multiple gene expression matrices. The experimental results presented herein demonstrate that it is a robust data integration technique able to produce good quality clustering solution that is representative for the whole set of expression matrices. Anna Hristoskova, Veselka Boeva, Elena Tsiporkova |
BMC Bioinform. | 2 |
| 2014 | Detecting serial residential burglaries using clustering
Anton Borg, Martin Boldt, Niklas Lavesson, Ulf Melander, Veselka Boeva |
Expert Syst. Appl. | 5 |
| 2006 | Multi-step ranking of alternatives in a multi-criteria and multi-expert decision making environment
Elena Tsiporkova, Veselka Boeva |
Inf. Sci. | 2 |
| 2004 | A transition logic for schemata conflicts
Veselka Boeva, Love Ekenberg |
Data Knowl. Eng. | 1 |
| 1999 | Dempster's rule of conditioning translated into modal logic
Elena Tsiporkova, Bernard De Baets, Veselka Boeva |
Fuzzy Sets Syst. | 3 |
| 1999 | Dempster-Shafer theory framed in modal logic
Elena Tsiporkova, Veselka Boeva, Bernard De Baets |
Int. J. Approx. Reason. | 2 |
| 1999 | Evidence Measures Induced by Kripke's Accessibility RelationsabstractModal logic interpretations of plausibility and belief measures are developed based on the observation that the accessibility relation in a model of modal logic, regarded as a multivalued mapping, induces a plausibility measure and a belief measure on the set of possible worlds. Elena Tsiporkova, Veselka Boeva, Bernard De Baets |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 2 |