VLDB 2026 Research / reviewers in the wild / expert
Guilherme Weigert Cassales
dblp:257/5616 · also Guilherme W. Cassales
· DBLP profile ↗
12ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0003-4029-2047ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adaptive Isolation Forest
Justin Jia Liu, Guilherme Weigert Cassales, Fei Tony Liu, Bernhard Pfahringer, Albert Bifet |
DS | 2 |
| 2025 | Streaming Isolation Forest
Justin Jia Liu, Guilherme Weigert Cassales, Fei Tony Liu, Bernhard Pfahringer, Albert Bifet |
PAKDD (1) | 2 |
| 2025 | Detecting Domain Shifts in Myoelectric Activations: Challenges and Opportunities in Stream Learning
Yibin Sun, Nick Jin Sean Lim, Guilherme Weigert Cassales, Heitor Murilo Gomes, Bernhard Pfahringer, Albert Bifet, Anany Dwivedi |
PRICAI (5) | 3 |
| 2025 | RMIDDM: an unsupervised and interpretable concept drift detection method for data streams
Ruivaldo Lobão-Neto, Brenno de Mello Alencar, Heitor Murilo Gomes, Albert Bifet, João Gama 0001, Guilherme Weigert Cassales, Ricardo Araújo Rios |
Data Min. Knowl. Discov. | 6 |
| 2025 | Accelerated Weka: GPU Machine Learning with Weka WorkbenchabstractRecent advancements in Machine Learning (ML) have driven increased interest and demand in the field. However, the complexity and mathematical background required for many techniques can be challenging for newcomers and enthusiasts. For many years, Weka has been a cornerstone in introductory ML courses worldwide, offering a user-friendly interface that facilitates understanding of various methods. As an open-source software with an active community, Weka has continued to evolve with new attributes. With the growing size of datasets, users are increasingly looking for ways to accelerate computations. GPUs have become a preferred solution for efficiently building and deploying ML models. This paper presents Accelerated Weka, a software framework that integrates GPU-accelerated methods within Weka’s intuitive graphical user interface, significantly reducing execution times for large datasets while preserving Weka’s accessibility. Benchmark results indicate that Accelerated Weka can achieve 2,198x speedup on an A100 chip. The framework retains Weka’s GPL 3.0 license and offers a straightforward installation process through the Conda environment. Guilherme Weigert Cassales, Justin Jia Liu, Albert Bifet |
Neurocomputing | 1 |
| 2024 | Mini-batching with Fused Training and Testing for Data Streams Processing on the EdgeabstractEdge Computing (EC) has emerged as a solution to reduce energy demand and greenhouse gas emissions from digital technologies. EC supports low latency, mobility, and location awareness for delay-sensitive applications by bridging the gap between cloud computing services and end-users. Machine learning (ML) methods have been applied in EC for data classification and information processing. Ensemble learners have often proven to yield high predictive performance on data stream classification problems. Mini-batching is a technique proposed for improving cache reuse in multi-core architectures of bagging ensembles for the classification of online data streams, which benefits application speedup and reduces energy consumption. However, the original mini-batching presents limited benefits in terms of cache reuse and it hinders the accuracy of the ensembles (i.e., their capacity to detect behavior changes in data streams). In this paper, we improve mini-batching by fusing continuous training and test loops for the classification of data streams. We evaluated the new strategy by comparing its performance and energy efficiency with the original mini-batching for data stream classification using six ensemble algorithms and four benchmark datasets. We also compare mini-batching strategies with two hardware-based strategies supported by commodity multi-core processors commonly used in EC. Results show that mini-batching strategies can significantly reduce energy consumption in 95% of the experiments. Mini-batching improved energy efficiency by 96% on average and 169% in the best case. Likewise, our new mini-batching strategy improved energy efficiency by 136% on average and 456% in the best case. These strategies also support better control of the balance between performance, energy efficiency, and accuracy. Reginaldo Luna, Guilherme Weigert Cassales, Bernhard Pfahringer, Albert Bifet, Heitor Murilo Gomes, Hermes Senger |
CF | 2 |
| 2024 | Online Isolation ForestabstractThe anomaly detection literature is abundant with offline methods, which require repeated access to data in memory, and impose impractical assumptions when applied to a streaming context. Existing online anomaly detection methods also generally fail to address these constraints, resorting to periodic retraining to adapt to the online context. We propose Online-iForest, a novel method explicitly designed for streaming conditions that seamlessly tracks the data generating process as it evolves over time. Experimental validation on real-world datasets demonstrated that Online-iForest is on par with online alternatives and closely rivals state-of-the-art offline anomaly detection techniques that undergo periodic retraining. Notably, Online-iForest consistently outperforms all competitors in terms of efficiency, making it a promising solution in applications where fast identification of anomalies is of primary importance such as cybersecurity, fraud and fault detection. Filippo Leveni, Guilherme Weigert Cassales, Bernhard Pfahringer, Albert Bifet, Giacomo Boracchi |
ICML | 2 |
| 2024 | Time-Evolving Data Science and Artificial Intelligence for Advanced Open Environmental Science (TAIAO) Programme
Yun Sing Koh, Albert Bifet, Karin R. Bryan, Guilherme Weigert Cassales, Olivier Graffeuille, Nick Jin Sean Lim, Phil Mourot, Ding Ning, Bernhard Pfahringer, Varvara Vetrova, Heitor Murilo Gomes |
IJCAI | 4 |
| 2023 | Balancing Performance and Energy Consumption of Bagging Ensembles for the Classification of Data Streams in Edge ComputingabstractIn recent years, the Edge Computing (EC) paradigm has emerged as an enabling factor for developing technologies like the Internet of Things (IoT) and 5G networks, bridging the gap between Cloud Computing services and end-users, supporting low latency, mobility, and location awareness to delay-sensitive applications. An increasing number of solutions in EC have employed machine learning (ML) methods to perform data classification and other information processing tasks on continuous and evolving data streams. Usually, such solutions have to cope with vast amounts of data that come as data streams while balancing energy consumption, latency, and the predictive performance of the algorithms. Ensemble methods achieve remarkable predictive performance when applied to evolving data streams due to several models and the possibility of selective resets. This work investigates a strategy that introduces short intervals to defer the processing of mini-batches. Well balanced, our strategy can improve the performance (i.e., delay, throughput) and reduce the energy consumption of bagging ensembles to classify data streams. The experimental evaluation involved six state-of-art ensemble algorithms (OzaBag, OzaBag Adaptive Size Hoeffding Tree, Online Bagging ADWIN, Leveraging Bagging, Adaptive RandomForest, and Streaming Random Patches) applying five widely used machine learning benchmark datasets with varied characteristics on three computer platforms. As a result, our strategy can significantly reduce energy consumption in 96% of the experimental scenarios evaluated. Despite the trade-offs, it is possible to balance them to avoid significant loss in predictive performance. Guilherme Weigert Cassales, Heitor Murilo Gomes, Albert Bifet, Bernhard Pfahringer, Hermes Senger |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2021 | Distributed Novelty Detection at the Edge for IoT Network Security
Luís Puhl, Guilherme Weigert Cassales, Hélio Crestana Guardia, Hermes Senger |
ICCSA (3) | 2 |
| 2021 | Improving the performance of bagging ensembles for data streams through mini-batching
Guilherme Weigert Cassales, Heitor Murilo Gomes, Albert Bifet, Bernhard Pfahringer, Hermes Senger |
Inf. Sci. | 1 |
| 2019 | IDSA-IoT: An Intrusion Detection System Architecture for IoT NetworksabstractThe Internet of Things (IoT) allows large amounts and variety of devices to connect, interact and exchange data. The IoT network creates numerous opportunities for novel attacks that can compromise information and systems integrity. Intrusion detection systems have been studied over two decades, mostly employing traditional data mining and machine learning techniques that require an offline phase for model training on large amounts of data. This paper presents three data stream novelty detection techniques applied to the intrusion detection problem and proposes IDSA-IoT, a novel implementation architecture, which combines the use of resources at the edge of the network and a public cloud. After an extensive empirical evaluation, results show that it is possible to identify new attack patterns soon after their emergence and to adapt the models in an efficient way. Guilherme Weigert Cassales, Hermes Senger, Elaine Ribeiro de Faria, Albert Bifet |
ISCC | 1 |