EDBT 2026 Demo / reviewers in the wild / expert
Samuel Xavier de Souza
dblp:96/303 · also Samuel Xavier-de-Souza
· DBLP profile ↗
14ranked-venue papers
2as first author
4since 2021 · last 2026
0000-0001-8747-4580ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Accelerating high-accuracy runtime estimation in adaptive sample selection for parallel scalability analysis with reiforcement learningabstractThe efficient allocation of resources in high-performance computing (HPC) depends on accurate analysis of application behavior. However, these scalability analyses are often computationally expensive and time-consuming. Although there are tools that automate this process, such as PaScal Suite, the cost of execution is still a challenge. To mitigate these costs, prediction strategies have proven effective, notably the use of neural networks as regression models. The obstacle, however, lies in choosing the data to train these models: the process can be slow and expensive, as depending on the order of the data, it may require that nearly all of the application’s settings be configured in PaScal Suite. This work proposes a strategy based on Reinforcement Learning (RL) to optimize the choice of samples in performance modeling. The main goal is to minimize the number of executions required to train a regression model, while ensuring high accuracy. The methodology employs a Deep Q-Learning agent that iteratively selects the ideal configurations (number of cores and problem size). These selected data are then used to train a neural network capable of predicting the execution time for each configuration, always aiming to maximize information gain and minimize time spent. Experimental analyses were performed to evaluate this approach in comparison to random and heuristic methods. Results demonstrate that the strategy reduced the total exhaustive execution time by approximately 71% and outperformed the random search baseline by nearly 40%, all while maintaining an error rate of less than 10%. This validates the feasibility of using RL to drastically reduce time costs in scalability prediction without compromising accuracy. Pedro Rici, Elisa Gabriela Lucena, Samuel Xavier de Souza |
HPDC | 3 |
| 2025 | Integration framework for online thread throttling with thread and page mapping on NUMA systems
Janaina Schwarzrock, Hiago Rocha, Arthur Francisco Lorenzon, Samuel Xavier de Souza, Antonio Carlos Schneider Beck |
J. Parallel Distributed Comput. | 4 |
| 2022 | Speculative guardband: exploiting critical-delay variations across cached instructionsabstractSeveral studies have been published to discuss methods of making components, such as CPUs, GPUs, and FPGAs, more energy efficient. Well-known techniques such as dynamic voltage and frequency scaling (DVFS) and power gating are alternatives since the supply voltage is directly related to power consumption. However, to guarantee correct operation without critical-path-timing violations, the systems must impose conservative static voltage guardbands. We propose to create a predictive guardband reduction technique for CPUs by analyzing the cached instructions. The contribution is expected to be the development of a machine learning solution capable of reducing average supply voltage levels by capturing the intrisic critical path variations according to which internal circuits are used by different sets of instructions. Johannes W. Farias, Diego V. Cirilo do Nascimento, Tiago Barros, Samuel Xavier de Souza |
VLSI-SoC | 4 |
| 2021 | Low learning-cost offline strategies for EDP optimization of parallel applications
Gustavo Berned, Fábio D. Rossi, Marcelo Caggiani Luizelli, Samuel Xavier de Souza, Antonio Carlos Schneider Beck, Arthur Francisco Lorenzon |
J. Syst. Archit. | 4 |
| 2019 | Performance and Energy Efficiency Trade-Offs in Single-ISA Heterogeneous Multi-Processing for Parallel ApplicationsabstractThis work proposes a novel methodology to predict the optimal performance and energy efficiency trade-off configurations of parallel applications running on a two-cluster Heterogeneous Multi-Processing (HMP) system. we propose an analytic performance and power model that are generated offline using data measurements. These models are then used to estimate the whole configuration space to predict the application's performance and energy consumption. Then, we use these off-line predictions to choose Pareto-optimal configurations, which is the most efficient among all configurations for the given architecture and multi-threaded application. We validated our methodology on an ODROID XU3 board on several PARSEC and Phoronix Test Suite applications. Demetrios Coutinho, Kyriakos Georgiou, Kerstin Eder, José L. Núñez-Yáñez, Samuel Xavier de Souza |
VLSI-SoC | 5 |
| 2019 | A Machine Learning-Based Framework for Throughput Estimation of Time-Varying Applications in Multi-Core ServersabstractAccurate workload prediction and throughput estimation are keys in efficient proactive power and performance management of multi-core platforms. Although hardware performance counters available on modern platforms contain important information about the application behavior, employing them efficiently is not straightforward when dealing with time-varying applications even if they have iterative structures. In this work, we propose a machine learning-based framework for workload prediction and throughput estimation using hardware events. Our framework enables throughput estimation over various available system configurations, namely, number of parallel threads and operating frequency. In particular, we first employ workload clustering and classification techniques along with Markov chains to predict the next workload for each available system configuration. Then, the predicted workload is used to estimate the next expected throughput through a machine learning-based regression model. The comparison with state of the art demonstrates that our framework is able to improve Quality of Service (QoS) by 3.4x, while consuming 15% less power thanks to the more accurate throughput estimation. Arman Iranfar, Wellington Silva de Souza, Marina Zapater, Katzalin Olcoz, Samuel Xavier de Souza, David Atienza 0001 |
VLSI-SoC | 5 |
| 2019 | Exploiting guard band limits for energy gains in MPSoCsabstractThe critical path delay and, as a consequence, the maximum operating frequency of a digital system are dependent on the supply voltage, so is its power consumption. Typically, voltage guard bands are added in order to improve reliability, at the expense of power efficiency. It has been shown that error detection and correction (EDAC) techniques can be used to mitigate the effects of the reduced safety margins. This Ph.D. project aims to develop a MPSoC capable of self regulate its operating voltage by monitoring the error rate reported by a EDAC system, virtually eliminating voltage safety margins. Literature review and preliminary experiments supports the viability of this approach, motivating further investigation. Diego V. Cirilo do Nascimento, Kyriakos Georgiou, Kerstin Eder, Samuel Xavier de Souza |
VLSI-SoC | 4 |
| 2019 | A QoS and Container-Based Approach for Energy Saving and Performance Profiling in Multi-Core ServersabstractIn this work we present ContainEnergy, a new performance evaluation and profiling tool that uses software containers to perform application runtime assessment, providing energy and performance profiling data. It is focused on energy efficiency for next generation workloads and IT infrastructure. Wellington Silva de Souza, Arman Iranfar, Anderson B. N. da Silva, Marina Zapater, Samuel Xavier de Souza, Katzalin Olcoz, David Atienza 0001 |
VLSI-SoC | 5 |
| 2018 | Scalable Shared-Memory Parallelization of the Block Recursive Inversion AlgorithmabstractThe Block Recursive Inversion (BRI) algorithm calculates the inversion of large k x k block matrices with limited memory during the entire processing because it calculates one block of the inverse at a time. However, the lower memory consumption is counterbalanced by higher computational complexity. We propose a parallel BRI implementation, which also calculates one block at a time, to reduce execution time and extend its applicability by exploiting modern multi-core architectures. The proposed parallel BRI was implemented for shared memory systems in OpenMP. The results of a performance and scalability analysis for different use cases reveals opposite trends in execution time, with the proposed parallel implementation being faster for larger k. Although not weakly scalable for a fixed k, scalability tends to increase with the increase of k or, equivalently, with the reduction of memory requirements. Maria Clara Silva, Iria C. S. Cosme, Idalmis M. Sardina, Samuel Xavier de Souza |
CLUSTER | 4 |
| 2018 | Less is More: Exploiting the Standard Compiler Optimization Levels for Better Performance and Energy ConsumptionabstractThis paper presents the interesting observation that by performing fewer of the optimizations available in a standard compiler optimization level such as -02, while preserving their original ordering, significant savings can be achieved in both execution time and energy consumption. This observation has been validated on two embedded processors, namely the ARM Cortex-M0 and the ARM Cortex-M3, using two different versions of the LLVM compilation framework; v3.8 and v5.0. Experimental evaluation with 71 embedded benchmarks demonstrated performance gains for at least half of the benchmarks for both processors. An average execution time reduction of 2.4% and 5.3% was achieved across all the benchmarks for the Cortex-M0 and Cortex-M3 processors, respectively, with execution time improvements ranging from 1% up to 90% over the -02. The savings that can be achieved are in the same range as what can be achieved by the state-of-the-art compilation approaches that use iterative compilation or machine learning to select flags or to determine phase orderings that result in more efficient code. In contrast to these time consuming and expensive to apply techniques, our approach only needs to test a limited number of optimization configurations, less than 64, to obtain similar or even better savings. Furthermore, our approach can support multi-criteria optimization as it targets execution time, energy consumption and code size at the same time. Kyriakos Georgiou, Craig Blackmore, Samuel Xavier de Souza, Kerstin Eder |
SCOPES | 3 |
| 2018 | Parallel synchronous and asynchronous coupled simulated annealing
Kayo Gonçalves-e-Silva, Daniel Aloise, Samuel Xavier de Souza |
J. Supercomput. | 3 |
| 2015 | Parallel Scalability of a Fine-Grain Prestack Reverse Time Migration AlgorithmabstractSeismic imaging has evolved significantly due to the high demand from the oil/gas industry for hardware technological advancements, boosting the development of more sophisticated algorithms. In order to deliver the quality and accuracy required, the execution of these algorithms may lead to time infeasible solutions. Aiming at performance improvement, this work conducted the parallelization of the core of a reverse time migration (RTM) algorithm. Furthermore, analysis such as speedup and efficiency was performed in order to assess the scalability of the proposed method. While the many parallelization efforts so far deal with coarse-grain approaches, this letter tackles the intrashot fine-grain parallelization of prestack RTM, which increases the overall concurrency degree of the algorithm. Results using 2-D synthetic data show that the proposed approach is scalable, which means that an increase in hardware resources and/or in problem size will lead to a proportional increase in speed and/or accuracy. Desnes A. Nunes-do-Rosario, Samuel Xavier de Souza, Rosangela C. Maciel, Jesse C. Costa |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2012 | Analysis of Multilayer Perceptron networks in the multicore eraabstractIn this paper we present and analyze a modular implementation of the Multilayer Perceptron (MLP) network in the view of the recent paradigm shift called the multicore era. The implementation parallelizes the pass of an input through the network by distributing the neurons of a given layer among parallel executed modules. Each module has full connection among its local neurons and an adjustable number of remote connections with other modules. We analyzed the parallel speedup and parallel efficiency for different total number of synapses and number of remote connections per module. The fully connected case showed weak scalability for future parallel architectures with non-uniform memory access. The results for the proposed implementation showed that it is quite scalable for a small number of remote connections. Samuel Xavier de Souza, Francisco Ary Alves de Souza, Adrião Duarte Dória Neto |
IJCNN | 1 |
| 2010 | Coupled Simulated AnnealingabstractWe present a new class of methods for the global optimization of continuous variables based on simulated annealing (SA). The coupled SA (CSA) class is characterized by a set of parallel SA processes coupled by their acceptance probabilities. The coupling is performed by a term in the acceptance probability function, which is a function of the energies of the current states of all SA processes. A particular CSA instance method is distinguished by the form of its coupling term and acceptance probability. In this paper, we present three CSA instance methods and compare them with the uncoupled case, i.e., multistart SA. The primary objective of the coupling in CSA is to create cooperative behavior via information exchange. This aim helps in the decision of whether uphill moves will be accepted. In addition, coupling can provide information that can be used online to steer the overall optimization process toward the global optimum. We present an example where we use the acceptance temperature to control the variance of the acceptance probabilities with a simple control scheme. This approach leads to much better optimization efficiency, because it reduces the sensitivity of the algorithm to initialization parameters while guiding the optimization process to quasioptimal runs. We present the results of extensive experiments and show that the addition of the coupling and the variance control leads to considerable improvements with respect to the uncoupled case and a more recently proposed distributed version of SA. Samuel Xavier de Souza, Johan A. K. Suykens, Joos Vandewalle, Désiré Bollé |
IEEE Trans. Syst. Man Cybern. Part B | 1 |