EDBT 2026 Demo / reviewers in the wild / expert
Dariusz Mrozek
dblp:99/3102
· DBLP profile ↗
16ranked-venue papers in the field
6as first author
9since 2021 · last 2025
0000-0001-6764-6656ORCID · corroborated
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 6Knowledge Engineering, Semantic Web & Information Systems · 5 (4 first)Data Mining & Knowledge Discovery · 3 (1 first)Database Systems & Data Management · 2 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Genetic Reads to Information Granules: Scalable Big NGS Data Cleaning with Apache Pig
Bozena Malysiak-Mrozek, Tomasz Sitek, Vaidy S. Sunderam, Boleslaw Pochopien, Krzysztof Tokarz, Che-Lun Hung, Dariusz Mrozek |
IEEE Big Data | 7 |
| 2025 | GPU-Enabled Edge-based Federated Continual Learning with Resource Adaptation for Efficient Non-Stationary Visual Anomaly Detection
Alexandre Niyomugaba, Dariusz Mrozek |
IEEE Big Data | 2 |
| 2024 | Fuzzy Querying in the Cloud-based Environment for Data Stream-driven Predictive Maintenance in AGV-enabled Smart FactoriesabstractFuzzy data processing enables data enrichment and increases data interpretation in industrial environments. In the cloud-based IoT data ingestion pipelines, fuzzy data processing can be implemented in several locations, closer to the IoT events gateways, stream processors, or the persistence layer before the data is visualized. Since Automated Guided Vehicles (AGV)-enabled manufacturing can produce vast amounts of data, the decision on the placement of the fuzzy data processing can be important for secondary processes performed on the enriched data, like the predictive maintenance inferencing. In this paper, we analyze two locations of fuzzy data processing in the cloud-based environment built for monitoring AGVs in smart factories - by formulating fuzzy queries against data streams on stream processing units and data at rest in a database. The querying scenarios cover fuzzy filtering with simple and complex criteria, fuzzy filtering through assignment to a linguistic variable, and joining data streams by representing joining attributes as fuzzy numbers. The experimental results show that querying the data stream can be more efficient and profitable in the scalable environment of many AGVs. However, the enrichment provided for the data at rest is also beneficial when gathering data for building future predictive maintenance models. Bozena Malysiak-Mrozek, Dominik Romanów, Piotr Grzesik, Pawel Benecki, Alexandre Niyomugaba, Theodore Habimana, Daniel Kostrzewa, Krzysztof Tokarz, Che-Lun Hung, Dariusz Mrozek |
IEEE Big Data | 10 |
| 2024 | Decoding the Granular Puzzle of Macromolecules: Efficient 3D Protein Structure Alignment in the Age of Big Data with Apache SparkabstractProteins are complex biological information granules that play a crucial role in various cellular processes within living organisms. Processing 3D protein structures, which are the most informative from the biological point of view, is both intricate and time-consuming. In particular, performing 3D protein structure searches against large protein datasets involves identifying similarities and conducting structural alignments across numerous molecules (granules). This task demands advanced methods for matching identical and similar regions within protein structures and substantial computational resources to handle large collections of macromolecular data efficiently. In this paper, we present our parallel implementation of scalable 3D structural alignment on the Apache Spark big data platform. We describe a customized approach that leverages Spark data transformations within the data processing pipeline for the alignment process. Our experimental results demonstrate that this solution, tightly integrated with the Spark processing model, is both efficient and scalable, even with the increasing volume of protein structure data. Bozena Malysiak-Mrozek, Paulina Pawlowicz, Vaidy S. Sunderam, Che-Lun Hung, Andrzej Kwiecien, Dariusz Mrozek |
IEEE Big Data | 6 |
| 2023 | Effective Prediction of Energy Consumption in Automated Guided Vehicles with Recurrent and Convolutional Neural NetworksabstractDetection and prediction of failures in Automated Guided Vehicles (AGV) are essential for the uninterrupted operation of production plants. Anomaly detection is usually achieved by comparing expected measurement values with actual observations. Thus, it is crucial to predict telemetry signals properly. In this paper, we research the prediction of energy consumption using state-of-the-art Artificial Neural Networks architectures (SCINet) compared with other Recurrent Neural Network (RNN) approaches on the data streams acquired from CoBotAGV. We especially focus on the possibility of applying feature weighting. We show that it can improve prediction capabilities. We also investigate resource utilization in terms of time to fit the embedded AGV environment. Pawel Benecki, Daniel Kostrzewa, Piotr Grzesik, Bohdan Shubyn, Jia-Hao Syu, Jerry Chun-Wei Lin, Vaidy S. Sunderam, Dariusz Mrozek |
IEEE Big Data | 8 |
| 2023 | Predicting Conflict Zones on Terrestrial Routes of Automated Guided Vehicles with Fuzzy Querying on Apache KafkaabstractIn today’s world, smart factories are a coexisting element of smarticizing cities. Smart manufacturing of today relies on the automation of many component tasks of the production process. Automated guided vehicles (AGVs) that transport materials on the production lines are important elements of this automation. Appropriate management of a fleet of AGVs requires avoiding collisions. However, prediction and early detection of approaching collision points on the transportation routes not only prevent collisions but also enables adjusting the AGV operation and improving its flow. In this paper, we demonstrate the use of fuzzy sets and linguistic variables in collision prevention by processing AGV data streams with Apache Kafka. We extend the capabilities of Apache Kafka and ksqlDB towards fuzzy stream processing and use fuzzy KSQL queries to predict collisions. Our experiments prove that fuzzy querying against AGV data streams does not consume much time and computational resources, and we can successfully avoid collisions by predicting future positions of the AGV for various densities of data streams and widths of time windows. Bozena Malysiak-Mrozek, Mario Bas, Vaidy S. Sunderam, Stanislaw Kozielski, Dariusz Mrozek |
DSAA | 5 |
| 2023 | YOLO-based Object Detection in Panoramic Images of Smart BuildingsabstractCollecting and analyzing spherical images is one of the crucial fields supporting the digitization of indoor environments, like houses, apartments, offices, or factories, and creating virtual tours of smart buildings. Such images also constitute an important feed for smart indoor devices, like autonomous vacuum cleaners, intelligent production lines, or assistive droids. However, smart devices must correctly identify internal objects on high-resolution spherical images, usually heavily distorted and made with variable lighting conditions. Moreover, the recognition task must be relatively accurate and take place onboard the device, as data transmission to larger analysis centers or workstations does not allow for real-time operation. In this paper, we compare two object detectors from the YOLO family (YOLOv5 and YOLOv8) on the publicly available dataset adjusted to a newer annotation representation that allows for adapting YOLO detectors to spherical images. We verify the effectiveness of these two families of detectors for varying sizes of object detection models, image sizes, and training batch sizes. Our research proves that YOLO models can successfully detect most indoor objects, and their effectiveness is comparable to dedicated detectors having much higher computational complexity. Sebastian Pokucinski, Dariusz Mrozek |
DSAA | 2 |
| 2023 | Special issue on Recent Advances in Fuzzy Deep Learning for Uncertain Medicine Data
Weiping Ding 0001, Jun Liu 0001, Chin-Teng Lin, Dariusz Mrozek |
Inf. Sci. | 4 |
| 2022 | An Efficient and Secured Energy Management System for Automated Guided VehiclesabstractIn this paper, we propose a Secure Energy Management System (SEMS) with anomaly detection and Q-Learning decision modules for Automated Guided Vehicles (AGV). The anomaly detection module is a multi-task learning network to simultaneously classify suppliers and predict the real supply quantities. The Q-learning decision module can then determine operating reserve and subsidies to manage the energy grid. Experimental results illustrate that the proposed anomaly detection module has an excellent performance in classifying malicious suppliers, excels at shaping supply distribution, and outperforms the existing benchmark systems. Jia-Hao Syu, Jerry Chun-Wei Lin, Dariusz Mrozek |
IEEE Big Data | 3 |
| 2020 | Fall detection in older adults with mobile IoT devices and machine learning in the cloud and on the edgeabstractRemote monitoring of older adults and detecting dangers in the state of human health have become essential elements of modern telemedicine. Falls are a frequent reason for deaths or post-traumatic complications in the elderly. Therefore, the early detection of falls can be crucial for the survival of a person or for providing necessary support. However, telemedicine data centers require scalable computing and storage resources for the growing number of monitored people. Dedicated approaches that allow for minimal data transmission of strictly interesting cases are also required. In this paper, we show a scalable architecture of a system that can monitor thousands of older adults, detect falls, and notify caregivers. Scalability tests that disclose requirements to enable large scale system operations were also performed. Moreover, we validated several Machine Learning models to evaluate their suitability in the detection process. Among the tested models, Boosted Decisions Trees resulted in the best classification performance. We also experimentally tested the detection of falls inside a Cloud-based data center and on an Edge IoT device. Results of tests on the device-to-cloud data transmission confirmed that significant reduction in the size of stored and transmitted data can be achieved while performing fall detection on the Edge. Dariusz Mrozek, Anna Koczur, Bozena Malysiak-Mrozek |
Inf. Sci. | 1 |
| 2019 | High-throughput and scalable protein function identification with Hadoop and Map-only pattern of the MapReduce processing modelabstractEfficient computational solutions for identification of protein functions or finding structural homologs of proteins gain importance in the era of structural genomics and in the face of growing volumes of biological data. Structural alignments, which underlie these two processes, take a lot of time to complete, especially when performed for large collections of 3D protein structures. Fortunately, structural alignments can be carried out on well-separable and independent subsets of the whole macromolecular data repository, which perfectly fits the MapReduce processing paradigm of bringing computations to data. In this paper, we show how the protein function identification and finding structural homologs can be efficiently accelerated with the use of the MapReduce procedure executed on Hadoop cluster established in a virtualized compute environment or a private cloud. For this purpose, we propose Map-only processing pattern of the MapReduce procedure, which is formally defined in this paper. The solution that we show joins advantages of performing computations in small virtualized compute environments with large-scale computations in public clouds, thus allowing to perform structural alignments for a number of usage scenarios, including comparison of pairs of 3D protein structures during evaluation of predicted protein models, one-to-many comparisons while identifying possible functions of the given structure, or all-to-all alignments while investigating the divergence between known protein structures and classifying proteins by their fold. In this paper, we also present results of performance tests when scaling up nodes of the Hadoop cluster and increasing the degree of parallelism with the intention of improving efficiency of the computations. Dariusz Mrozek, Marek Suwala, Bozena Malysiak-Mrozek |
Knowl. Inf. Syst. | 1 |
| 2017 | Orchestrating Task Execution in Cloud4PSi for Scalable Processing of Macromolecular Data of 3D Protein Structures
Dariusz Mrozek, Artur Klapcinski, Bozena Malysiak-Mrozek |
ACIIDS (2) | 1 |
| 2017 | Scalability of a Genomic Data Analysis in the BioTest Platform
Krzysztof Psiuk-Maksymowicz, Dariusz Mrozek, Roman Jaksik, Damian Borys, Krzysztof Fujarewicz, Andrzej Swierniak |
ACIIDS (2) | 2 |
| 2017 | Life Sciences Data Analysis
Dariusz Mrozek, Pawel Kasprowski, Bozena Malysiak-Mrozek, Stanislaw Kozielski |
Inf. Sci. | 1 |
| 2016 | HDInsight4PSi: Boosting performance of 3D protein structure similarity searching with HDInsight clusters in Microsoft Azure cloud
Dariusz Mrozek, Pawel Danilowicz, Bozena Malysiak-Mrozek |
Inf. Sci. | 1 |
| 2016 | An efficient and flexible scanning of databases of protein secondary structures - with the segment index and multithreaded alignmentabstractProtein secondary structure describe protein construction in terms of regular spatial shapes, including alpha-helices, beta-strands, and loops, which protein amino acid chain can adopt in some of its regions. This information is supportive for protein classification, functional annotation, and 3D structure prediction. The relevance of this information and the scope of its practical applications cause the requirement for its effective storage and processing. Relational databases, widely-used in commercial systems in recent years, are one of the serious alternatives honed by years of experience, enriched with developed technologies, equipped with the declarative SQL query language, and accepted by the large community of programmers. Unfortunately, relational database management systems are not designed for efficient storage and processing of biological data, such as protein secondary structures. In this paper, we present a new search method implemented in the search engine of the PSS-SQL language. The PSS-SQL allows formulation of queries against a relational database in order to find proteins having secondary structures similar to the structural pattern specified by a user. In the paper, we will show how the search process can be accelerated by multiple scanning of the Segment Index and parallel implementation of the alignment procedure using multiple threads working on multiple-core CPUs. Dariusz Mrozek, Bartek Socha, Stanislaw Kozielski, Bozena Malysiak-Mrozek |
J. Intell. Inf. Syst. | 1 |