Guilherme Weigert Cassales

dblp:257/5616 · also Guilherme W. Cassales · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0003-4029-2047ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Adaptive Isolation Forest
Justin Jia Liu, Guilherme Weigert Cassales, Fei Tony Liu, Bernhard Pfahringer, Albert Bifet
DS2
2025 Streaming Isolation Forest
Justin Jia Liu, Guilherme Weigert Cassales, Fei Tony Liu, Bernhard Pfahringer, Albert Bifet
PAKDD (1)2
2025 Detecting Domain Shifts in Myoelectric Activations: Challenges and Opportunities in Stream Learning
Yibin Sun, Nick Jin Sean Lim, Guilherme Weigert Cassales, Heitor Murilo Gomes, Bernhard Pfahringer, Albert Bifet, Anany Dwivedi
PRICAI (5)3
2025 RMIDDM: an unsupervised and interpretable concept drift detection method for data streams
Ruivaldo Lobão-Neto, Brenno de Mello Alencar, Heitor Murilo Gomes, Albert Bifet, João Gama 0001, Guilherme Weigert Cassales, Ricardo Araújo Rios
Data Min. Knowl. Discov.6
2025 Accelerated Weka: GPU Machine Learning with Weka Workbench
abstract
Recent advancements in Machine Learning (ML) have driven increased interest and demand in the field. However, the complexity and mathematical background required for many techniques can be challenging for newcomers and enthusiasts. For many years, Weka has been a cornerstone in introductory ML courses worldwide, offering a user-friendly interface that facilitates understanding of various methods. As an open-source software with an active community, Weka has continued to evolve with new attributes. With the growing size of datasets, users are increasingly looking for ways to accelerate computations. GPUs have become a preferred solution for efficiently building and deploying ML models. This paper presents Accelerated Weka, a software framework that integrates GPU-accelerated methods within Weka’s intuitive graphical user interface, significantly reducing execution times for large datasets while preserving Weka’s accessibility. Benchmark results indicate that Accelerated Weka can achieve 2,198x speedup on an A100 chip. The framework retains Weka’s GPL 3.0 license and offers a straightforward installation process through the Conda environment.
Guilherme Weigert Cassales, Justin Jia Liu, Albert Bifet
Neurocomputing1
2024 Mini-batching with Fused Training and Testing for Data Streams Processing on the Edge
abstract
Edge Computing (EC) has emerged as a solution to reduce energy demand and greenhouse gas emissions from digital technologies. EC supports low latency, mobility, and location awareness for delay-sensitive applications by bridging the gap between cloud computing services and end-users. Machine learning (ML) methods have been applied in EC for data classification and information processing. Ensemble learners have often proven to yield high predictive performance on data stream classification problems. Mini-batching is a technique proposed for improving cache reuse in multi-core architectures of bagging ensembles for the classification of online data streams, which benefits application speedup and reduces energy consumption. However, the original mini-batching presents limited benefits in terms of cache reuse and it hinders the accuracy of the ensembles (i.e., their capacity to detect behavior changes in data streams). In this paper, we improve mini-batching by fusing continuous training and test loops for the classification of data streams. We evaluated the new strategy by comparing its performance and energy efficiency with the original mini-batching for data stream classification using six ensemble algorithms and four benchmark datasets. We also compare mini-batching strategies with two hardware-based strategies supported by commodity multi-core processors commonly used in EC. Results show that mini-batching strategies can significantly reduce energy consumption in 95% of the experiments. Mini-batching improved energy efficiency by 96% on average and 169% in the best case. Likewise, our new mini-batching strategy improved energy efficiency by 136% on average and 456% in the best case. These strategies also support better control of the balance between performance, energy efficiency, and accuracy.
Reginaldo Luna, Guilherme Weigert Cassales, Bernhard Pfahringer, Albert Bifet, Heitor Murilo Gomes, Hermes Senger
CF2
2024 Online Isolation Forest
abstract
The anomaly detection literature is abundant with offline methods, which require repeated access to data in memory, and impose impractical assumptions when applied to a streaming context. Existing online anomaly detection methods also generally fail to address these constraints, resorting to periodic retraining to adapt to the online context. We propose Online-iForest, a novel method explicitly designed for streaming conditions that seamlessly tracks the data generating process as it evolves over time. Experimental validation on real-world datasets demonstrated that Online-iForest is on par with online alternatives and closely rivals state-of-the-art offline anomaly detection techniques that undergo periodic retraining. Notably, Online-iForest consistently outperforms all competitors in terms of efficiency, making it a promising solution in applications where fast identification of anomalies is of primary importance such as cybersecurity, fraud and fault detection.
Filippo Leveni, Guilherme Weigert Cassales, Bernhard Pfahringer, Albert Bifet, Giacomo Boracchi
ICML2
2024 Time-Evolving Data Science and Artificial Intelligence for Advanced Open Environmental Science (TAIAO) Programme
Yun Sing Koh, Albert Bifet, Karin R. Bryan, Guilherme Weigert Cassales, Olivier Graffeuille, Nick Jin Sean Lim, Phil Mourot, Ding Ning, Bernhard Pfahringer, Varvara Vetrova, Heitor Murilo Gomes
IJCAI4
2023 Balancing Performance and Energy Consumption of Bagging Ensembles for the Classification of Data Streams in Edge Computing
abstract
In recent years, the Edge Computing (EC) paradigm has emerged as an enabling factor for developing technologies like the Internet of Things (IoT) and 5G networks, bridging the gap between Cloud Computing services and end-users, supporting low latency, mobility, and location awareness to delay-sensitive applications. An increasing number of solutions in EC have employed machine learning (ML) methods to perform data classification and other information processing tasks on continuous and evolving data streams. Usually, such solutions have to cope with vast amounts of data that come as data streams while balancing energy consumption, latency, and the predictive performance of the algorithms. Ensemble methods achieve remarkable predictive performance when applied to evolving data streams due to several models and the possibility of selective resets. This work investigates a strategy that introduces short intervals to defer the processing of mini-batches. Well balanced, our strategy can improve the performance (i.e., delay, throughput) and reduce the energy consumption of bagging ensembles to classify data streams. The experimental evaluation involved six state-of-art ensemble algorithms (OzaBag, OzaBag Adaptive Size Hoeffding Tree, Online Bagging ADWIN, Leveraging Bagging, Adaptive RandomForest, and Streaming Random Patches) applying five widely used machine learning benchmark datasets with varied characteristics on three computer platforms. As a result, our strategy can significantly reduce energy consumption in 96% of the experimental scenarios evaluated. Despite the trade-offs, it is possible to balance them to avoid significant loss in predictive performance.
Guilherme Weigert Cassales, Heitor Murilo Gomes, Albert Bifet, Bernhard Pfahringer, Hermes Senger
IEEE Trans. Netw. Serv. Manag.1
2021 Distributed Novelty Detection at the Edge for IoT Network Security
Luís Puhl, Guilherme Weigert Cassales, Hélio Crestana Guardia, Hermes Senger
ICCSA (3)2
2021 Improving the performance of bagging ensembles for data streams through mini-batching
Guilherme Weigert Cassales, Heitor Murilo Gomes, Albert Bifet, Bernhard Pfahringer, Hermes Senger
Inf. Sci.1
2019 IDSA-IoT: An Intrusion Detection System Architecture for IoT Networks
abstract
The Internet of Things (IoT) allows large amounts and variety of devices to connect, interact and exchange data. The IoT network creates numerous opportunities for novel attacks that can compromise information and systems integrity. Intrusion detection systems have been studied over two decades, mostly employing traditional data mining and machine learning techniques that require an offline phase for model training on large amounts of data. This paper presents three data stream novelty detection techniques applied to the intrusion detection problem and proposes IDSA-IoT, a novel implementation architecture, which combines the use of resources at the edge of the network and a public cloud. After an extensive empirical evaluation, results show that it is possible to identify new attack patterns soon after their emergence and to adapt the models in an efficient way.
Guilherme Weigert Cassales, Hermes Senger, Elaine Ribeiro de Faria, Albert Bifet
ISCC1