VLDB 2026 Research / reviewers in the wild / expert
Alfredo Goldman
dblp:g/AlfredoGoldman · also Alfredo Goldman vel Lejbman
· DBLP profile ↗
65ranked-venue papers
9as first author
15since 2021 · last 2025
0000-0001-5746-4154ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 6 first-author · 4 since 2021Software engineering, systems software and programming languages · 17 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Computer networks · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Boosting Task Scheduling Data Locality with Low-latency, HW-accelerated Label PropagationabstractTask Scheduling is a popular technique for exploiting parallelism in modern computing systems.In particular, HW-accelerated Task Scheduling has been shown to be effective at improving the performance of fine-grained workloads by dynamically assigning tasks to cores based on their data dependencies with minimal overhead, allowing the handling of tasks with execution times in the order of thousands of cycles.However, the performance of applications assisted by accelerated Task Scheduling is limited by the fact that once a task has all its dependencies fulfilled, it is typically executed on the first available core, which might not be locality-optimal.We thus propose a novel approach to Task Scheduling that leverages HW-accelerated Label Propagation (LP), a graph clustering algorithm, to group tasks with intersecting data patterns such that they are executed on the same core.We show that our approach can significantly improve the performance of task-based applications, improving overall program execution times by up to 1.50× while simultaneously reducing average task sizes by up to 1.81×, augmenting both synthetic benchmarks and real-world applications running on a 24-core RISC-V processor mapped to the Alveo U55C FPGA.These gains rely heavily on the low-latency nature of our proposed label propagation accelerator, which will typically cluster dynamic task graphs in under 300 cycles, up to 581× faster than an equivalent software implementation.Furthermore, by ensuring that ideal placement predictions are used as a hint rather than a hard constraint, we allow the system to benefit from improved data locality for memory-intensive applications while also maintaining high core utilization in compute-bound scenarios.Our results hence demonstrate the potential of HW-accelerated label propagation to improve the performance of Task Scheduling systems with low-latency, dynamic data locality optimization. Lucas Morais, Juan Miguel De Haro Ruiz, Alfredo Goldman, Guido Araujo, Giacomo Pedretti, Jim Ignowski, Michael Frank 0008, Xavier Martorell, Daniel Jiménez-González, Carlos Álvarez 0001 |
MICRO | 3 |
| 2025 | Generative Fabrication of Medical Images for Machine Learning TrainingabstractTraining in supervised machine learning is based on the availability of datasets; however, medical datasets must comply with stringent privacy regulations. Generative Adversarial Networks (GANs) are a relevant alternative to solve the limitation of small medical datasets due to their ability to generate additional data with desired features. A significant drawback of these models is that they may produce unrealistic, blurred, or insufficiently diverse images. This paper proposes a data augmentation technique using GANs to create synthetic Magnetic Resonance Imaging (MRI) of four stages of Alzheimer's Disease (AD): non-demented, very mild demented, mild demented, and moderate demented. We designed a GAN based on the Pix2Pix model, which learns the features of each AD stage. Generated images are evaluated by multistage Convolutional Neural Network (CNN) models, greyscale histograms of the distribution of pixel intensities, and brain mass measurements on binarized images. The results indicate that AD synthetic MRI effectively captures disease patterns, demonstrating the potential of GANs to improve training and diagnosis of neurodegenerative diseases. Andres G. Calzada-Jasso, Andrei Tchernykh, Ixchel D. Avendaño-Pacheco, Jorge M. Cortés-Mendoza, Luis Bernardo Pulido-Gaytan, Mikhail G. Babenko, Alfredo Goldman, Horacio González-Vélez |
SBAC-PAD | 7 |
| 2024 | Enabling HW-Based Task Scheduling in Large Multicore ArchitecturesabstractDynamic Task Scheduling is an enticing programming model aiming to ease the development of parallel programs with intrinsically irregular or data-dependent parallelism. The performance of such solutions relies on the ability of the Task Scheduling HW/SW stack to efficiently evaluate dependencies at runtime and schedule work to available cores. Traditional SW-only systems implicate scheduling overheads of around 30K processor cycles per task, which severely limit the (core count,task granularity) combinations that they might adequately handle. Previous work on HW-accelerated Task Scheduling has shown that such systems might support high performance scheduling on processors with up to eight cores, but questions remained regarding the viability of such solutions to support the greater number of cores now frequently found in high-end SMP systems.The present work presents an FPGA-proven, tightly-integrated, Linux-capable, 30-core RISC-V system with hardware accelerated Task Scheduling. We use this implementation to show that HW Task Scheduling can still offer competitive performance at such high core count, and describe how this organization includes hardware and software optimizations that make it even more scalable than previous solutions. Finally, we outline ways in which this architecture could be augmented to overcome inter-core communication bottlenecks, mitigating the cache-degradation effects usually involved in the parallelization of highly optimized serial code. Lucas Morais, Carlos Álvarez 0001, Daniel Jiménez-González, Juan Miguel De Haro Ruiz, Guido Araujo, Michael Frank 0008, Alfredo Goldman, Xavier Martorell |
IEEE Trans. Computers | 7 |
| 2023 | Discriminant Audio Properties in Deep Learning Based Respiratory Insufficiency Detection in Brazilian Portuguese
Marcelo M. Gauy, Larissa Cristina Berti, Arnaldo Cândido Jr., Augusto Camargo Neto, Alfredo Goldman, Anna Sara Shafferman Levin, Marcus Martins, Beatriz Raposo de Medeiros, Marcelo Queiroz, Ester C. Sabino, Flaviane Romani Fernandes Svartman, Marcelo Finger |
AIME | 5 |
| 2023 | Towards the Detection of Microservice Patterns Based on MetricsabstractMicroservices is a popular architectural approach for complex systems in companies, despite its nature of decentralization. There is a comprehensive set of microservices architectural patterns that guides implementations and helps developers to overcome issues. However, the community still scarcely adopts these patterns and only has a theoretical understanding of them. In this work, in order to increase awareness of such patterns and provide aid to developers to better understand an architecture based on microservices, we propose a detection approach based on metrics for microservices patterns. We focused on structural or architectural patterns, and implemented detection for five of them. We conducted two case studies with real-world applications and evaluated the accuracy and applicability of our approach with the developers of those applications. João Francisco Lino Daniel, Eduardo Guerra 0001, Thatiane de Oliveira Rosa, Alfredo Goldman |
SEAA | 4 |
| 2023 | An Experimental Analysis of Regression-Obtained HPC Scheduling Heuristics
Lucas Rosa, Danilo Carastan-Santos, Alfredo Goldman |
JSSPP | 3 |
| 2023 | Evaluating execution time predictions on GPU kernels using an analytical model and machine learning techniquesabstractPredicting the performance of applications executed on GPUs is a great challenge and is essential for efficient job schedulers. There are different approaches to do this, namely analytical modeling and machine learning (ML) techniques. Machine learning requires large training sets and reliable features, nevertheless it can capture the interactions between architecture and software without manual intervention. In this paper, we compared a BSP-based analytical model to predict the time of execution of kernels executed over GPUs. The comparison was made using three different ML techniques. The analytical model is based on the number of computations and memory accesses of the GPU, with additional information on cache usage obtained from profiling. The ML techniques Linear Regression, Support Vector Machine, and Random Forest were evaluated over two scenarios: first, data input or features for ML techniques were the same as the analytical model and, second, using a process of feature extraction, which used correlation analysis and hierarchical clustering. Our experiments were conducted with 20 CUDA kernels, 11 of which belonged to 6 real-world applications of the Rodinia benchmark suite, and the other were classical matrix-vector applications commonly used for benchmarking. We collected data over 9 NVIDIA GPUs in different machines. We show that the analytical model performs better at predicting when applications scale regularly. For the analytical model a single parameter λ is capable of adjusting the predictions, minimizing the complex analysis in the applications. We show also that ML techniques obtained high accuracy when a process of feature extraction is implemented. Sets of 5 and 10 features were tested in two different ways, for unknown GPUs and for unknown Kernels. For ML experiments with a process of feature extractions, we got errors around 1.54% and 2.71%, for unknown GPUs and for unknown Kernels, respectively. Marcos Amaris, Raphael Y. de Camargo, Daniel Cordeiro, Alfredo Goldman, Denis Trystram |
J. Parallel Distributed Comput. | 4 |
| 2023 | CharM - Evaluating a model for characterizing service-based architectures
Thatiane de Oliveira Rosa, Eduardo Guerra 0001, Filipe Figueiredo Correia, Alfredo Goldman |
J. Syst. Softw. | 4 |
| 2022 | Are knowledge and usage of microservices patterns aligned? An exploratory study with professionalsabstractMicroservices Architecture is a trending solution for large systems, which counts with an extensive pattern language that defines its base practices and documents solutions to recurrent problems. However, there is a lack of studies investigating how these patterns are known and applied by professionals. Understanding how the patterns are used enables to comprehend the design process for this architectural style and identify opportunities for improvement. So, this work aims to collect and analyze information about how professionals know and use microservice patterns. To achieve that, we conducted a questionnaire study focused on eleven patterns that directly influence the architecture and components structure. The questionnaire was answered by 63 participants and revealed that, in general, they know the patterns, but with a significant amount declaring that it was known only as a practice. Additionally, among other results, our study also identified that the patterns are more commonly adopted at the project beginning rather than by refactoring and that they frequently are adopted more than once in the same system. João Francisco Lino Daniel, Alfredo Goldman, Eduardo Guerra 0001 |
COMPSAC | 2 |
| 2022 | Sonarlizer xplorer: a tool to mine github projects and identify technical debt items using SonarQubeabstractThe advancement of artificial intelligence and the implementation of machine learning capabilities in programming languages such as Python, along with cloud services, allow researchers to apply methods to cluster and predict behaviors and patterns in software engineering data. On the other hand, these methods need a large amount of data in order to work with high accuracy in different contexts. This paper introduces Sonarlizer Xplorer: a tool that captures a large number of technical debt items and code metrics from public GitHub projects. Sonarlizer Xplorer is composed of two sub-tools. The first is Github Xplorer, responsible for mining public Github repositories from an initial project. The second is Sonarlizer, responsible for taking projects and analyzing them using SonarQube. We used the tool over four months, collecting technical debt items and code metrics on almost 46,000 public Java projects. In addition, we mined over 57 million repositories and 4 million users. Diogo Pina, Alfredo Goldman, Carolyn B. Seaman |
TechDebt@ICSE | 2 |
| 2022 | Technical debt prioritization: a developer's perspectiveabstractBackground: The prioritization of technical debt is an essential task in managing software projects because, with current analysis tools, it is possible to find thousands of technical debt items in the software that would take months or even years to be fully paid. Aims: In this study, we aim to understand which criteria software developers use to prioritize code technical debt in real software projects. Methods: We performed a survey to collect data from open-source software projects in order to reach a large and diverse set of experiences. We analyzed the data using Straussian Grounded Theory techniques: open coding, axial coding, and selective coding. Results: We grouped the criteria into 15 categories and divided them into 2 super-categories related to paying off the technical debt and 3 related to not paying it. Conclusions: When participants decided to pay off technical debt, they wanted to do it soon. However, when they decided not to pay it, it is often because the debt occurred intentionally due to a project decision. Also, participants using similar criteria for their decisions tended to choose similar priority levels for those decisions. Finally, we observed that each software project needs to tailor the rules used to identify code technical debt to their project context. Diogo Pina, Carolyn B. Seaman, Alfredo Goldman |
TechDebt@ICSE | 3 |
| 2022 | Parallelizing Git Checkout: a Case Study of I/O ParallelismabstractVersion control systems (VCS) are tools used to track and manage the changes made to a set of files over time. Among the VCS tools available today, Git has become the most popular for software development. Being used in small personal projects of a few megabytes and massive corporate repositories with more than 300 GB and 3.5 million files, speed and scalability are among the top priorities for the tool. However, its performance sometimes falls short of what is desired on networked file systems (e.g. NFS), where input and output (I/O) operations tend to be more costly. In particular, that is the case for the checkout command, which is responsible for restoring files from specific versions of a project. Despite the optimizations implemented over the years, the sequential processing of files still carried a large time penalty for NFS, as well as being suboptimal for local file systems on SSDs. In this project, we worked to parallelize the Git checkout machinery, resulting in speedups of up to 4.5x on NFS and 3.6x on SSDs. We also studied how parallelism affects the I/O requests performed by checkout on different storage systems. The optimization was submitted upstream and made available to all Git users starting at version 2.32.0, from June 2021. Matheus Tavares Bernardino, Alfredo Goldman |
SBAC-PAD | 2 |
| 2021 | Technical Debt Prioritization: Taxonomy, Methods Results, and Practical CharacteristicsabstractTechnical debt is the metaphor for shortcuts in software development that bring short-term benefits, but long-term consequences hinder the process of maintaining and developing software. It is important to manage these technical debt items, as not all of them need to be paid. Having a list of prioritized debts is an essential step in decision-making in the management process. This work aims at finding technical debt prioritization methods, providing a classification of them. That is, methods to identify whether and when a technical debt should be paid off. We performed a systematic mapping review to find and analyze the main papers of the area, covering the main bases. We selected 112 studies, resulting in 51 unique papers. We classified the methods in a two-level taxonomy containing 10 categories according to their different possible outcomes. In addition, we have identified three methods results: boolean, category and ordered list. Finally, we have also identified practical technical characteristics and requirements for a method to prioritize technical debt items in real projects. Although several methods have been found in literature, none of them are adaptive to the context and are language-independent, nor cover several technical debt types. Moreover, there is a clear lack of tools to use them. So, in conclusion, the research on technical debt prioritization is still wide open. From this study, a combination of the techniques used in these methods can be tested and automated to assist in the decision-making process on which debts should be paid. Diogo Pina, Alfredo Goldman, Graziela Tonin |
SEAA | 2 |
| 2021 | Short-Term Ambient Temperature Forecasting for Smart HeatersabstractMaintaining Cloud data centers is a worrying challenge in terms of energy efficiency. This challenge leads to solutions such as deploying Edge nodes that operate inside buildings without massive cooling systems. Edge nodes can act as smart heaters by recycling their consumed energy to heat these buildings. We propose a novel technique to perform temperature forecasting for Edge Computing smart heater environments. Our approach uses time series algorithms to exploit historical air temperature data, smart heaters’ power consumption and temperature to create models to predict short-term ambient temperature over one hour horizon. We implemented our approach on top of Facebook's Prophet time series forecasting framework, and we used the real-time logs from Qarnot Computing as a use-case of a smart heater Edge platform. Our best trained model yields ambient temperature forecasts with less than 2.66% Mean Absolute Percentage Error. Danilo Carastan-Santos, Anderson Andrei Da Silva, Alfredo Goldman, Angan Mitra, Yanik Ngoko, Clément Mommessin, Denis Trystram |
ISCC | 3 |
| 2021 | A CPU-FPGA heterogeneous approach for biological sequence comparison using high-level synthesisabstractSummary This article presents a high‐level synthesis implementation of the longest common subsequence (LCS) algorithm combined with a weighted‐based scheduler for comparing biological sequences prioritizing energy consumption or execution time. The LCS algorithm has been thoroughly tailored using Vivado High‐Level Synthesis tool, which is able to synthesize register transfer level (RTL) from high‐level language descriptions, such as C/C++. Performance and energy consumption results were obtained with a CPU Intel Core i7‐3770 CPU and an Alpha‐Data ADM‐PCIE‐KU3 board that has a Xilinx Kintex UltraScale XCKU060 FPGA chip. We executed a batch of 20 comparisons of sequences on 10k, 20k, and 50k sizes. Our experiments showed that the energy consumption on the combined approach was significantly lower when compared to the CPU, achieving 75% energy reduction on 50k comparisons. We also used the tool proposed in this article to do a case study on Covid‐19, with real SARS‐CoV‐2 sequences, comparing their LCS scores. Carlos Antônio Campos Jorge, Alexandre Solon Nery, Alba Cristina Magalhaes Alves de Melo, Alfredo Goldman |
Concurr. Comput. Pract. Exp. | 4 |
| 2020 | PDAWL: Profile-Based Iterative Dynamic Adaptive WorkLoad Balance on Heterogeneous Architectures
Tongsheng Geng, Marcos Amaris, Stéphane Zuckerman, Alfredo Goldman, Guang R. Gao, Jean-Luc Gaudiot |
JSSPP | 4 |
| 2020 | Agile ways of working: A team maturity perspectiveabstractAbstract With the agile approach to managing software development projects comes an increased dependability on well‐functioning teams, since many of the practices are built on teamwork. The objective of this study was to investigate if, and how, team development from a group psychological perspective is related to some work practices of agile teams. Data were collected from 34 agile teams (200 individuals) from six software development organizations and one university in both Brazil and Sweden using the Group Development Questionnaire (Scale IV) and the Perceptive Agile Measurement (PAM). The result indicates a strong correlation between levels of group maturity and the two agile practices iterative development and retrospectives. We, therefore, conclude that agile teams at different group development stages adopt parts of team agility differently, thus confirming previous studies but with more data and by investigating concrete and applied agile practices. We thereby add evidence to the hypothesis that an agile implementation and management of agile projects need to be adapted to the group maturity levels of the agile teams. Lucas Gren, Alfredo Goldman, Christian Jacobsson |
J. Softw. Evol. Process. | 2 |
| 2019 | Autotuning Under Tight Budget Constraints: A Transparent Design of Experiments ApproachabstractA large amount of resources is spent writing, porting, and optimizing scientific and industrial High Performance Computing applications, which makes autotuning techniques fundamental to lower the cost of leveraging the improvements on execution time and power consumption provided by the latest software and hardware platforms. Despite the need for economy, most autotuning techniques still require large budgets of costly experimental measurements to provide good results, while rarely providing exploitable knowledge after optimization. The contribution of this paper is a user-transparent autotuning technique based on Design of Experiments that operates under tight budget constraints by significantly reducing the measurements needed to find good optimizations. Our approach enables users to make informed decisions on which optimizations to pursue and when to stop. We present an experimental evaluation of our approach and show it is capable of leveraging user decisions to find the best global configuration of a GPU Laplacian kernel using half of the measurement budget used by other common autotuning techniques. We show that our approach is also capable of finding speedups of up to 50x, compared to gcc's -O3, for some kernels from the SPAPT benchmark suite, using up to 10x fewer measurements than random sampling. Pedro Bruel, Steven Quinito Masnada, Brice Videau, Arnaud Legrand, Jean-Marc Vincent, Alfredo Goldman |
CCGRID | 6 |
| 2019 | A Students' Perspective of Native and Cross-Platform Approaches for Mobile Application Development
Paulo Meirelles, Carla S. R. Aguiar, Felipe Assis, Rodrigo Siqueira, Alfredo Goldman |
ICCSA (5) | 5 |
| 2019 | Adding Tightly-Integrated Task Scheduling Acceleration to a RISC-V Multi-core ProcessorabstractTask Parallelism is a parallel programming model that provides code annotation constructs to outline tasks and describe how their pointer parameters are accessed so that they might be executed in parallel, and asynchronously, by a runtime capable of inferring and honoring their data dependence relationships. It is supported by several parallelization frameworks, as OpenMP and StarSs. Lucas Morais, Vitor Silva, Alfredo Goldman, Carlos Álvarez 0001, Jaume Bosch, Michael Frank 0008, Guido Araujo |
MICRO | 3 |
| 2019 | A model of requirements engineering in software startupsabstractContext: Over the past 20 years, software startups have created many products that have changed human life. Since these companies are creating brand-new products or services, requirements are difficult to gather and highly volatile. Although scientific interest in software development in this context has increased, the studies on requirements engineering in software startups are still scarce and mostly focused on elicitation activities. Objective: This study overcomes this gap by answering how requirements engineering practices are performed in this context. Method: We conducted a grounded theory study based on 17 interviews with software startups practitioners. Results: We constructed a model to show that software startups do not follow a single set of practices but, instead, build a custom process, changed throughout the development of the company, combining different practices according to a set of influences (Founders, Software Development Manager, Developers, Market, Business Model and Startup Ecosystem). Conclusion: Our findings show that requirements engineering activities in software startups are similar to those in agile teams, but some steps vary as a consequence of the lack of an accessible customer. Jorge Melegati, Alfredo Goldman, Fabio Kon, Xiaofeng Wang 0001 |
Inf. Softw. Technol. | 2 |
| 2018 | Efficient Resources Utilization by Different Microservices Deployment ModelsabstractThe adoption of microservice-based architecture has become increasingly popular. Microservice containerization is a technique used by developers to facilitate the deployment process of applications based on this architecture. There are several implementation models for microservices. In this paper we study and analyze the performance of these models in terms of the use of network, CPU, memory and disk. Pros and cons related to the development process are also discussed. Among the results obtained with measurements made in a public cloud, the significant reductions in network consumption (up to 99%) are noteworthy when using one container per microservice. Fernando H. L. Buzato, Alfredo Goldman, Daniel M. Batista |
NCA | 2 |
| 2018 | Sensing Trees in Smart Cities with Open-Design HardwareabstractTree fall is a major issue in large cities as it may obstruct roads and lead to traffic jams, and even injure people. Following these kind of incidents, monitoring the environment through sensors and Internet of Things technologies emerges as an important preventive approach. Another point is that the net primary productivity has a sensitive relation with trees characteristics including plant respiration and net photosynthesis which are highly sensitive to temperature. The health of the tree can then be monitored through leaf transpiration, sap flow measurement, environment evaluation as these methods offer indicators on any chance of a fall. The solution proposed in this paper includes an open-design hardware setup for monitoring trees in Smart Cities with Internet of Things. The setup is connected to a Smart City platform aiming to facilitate the use by other researchers and botanics, specially. Therefore, we noticed that the daily sensing data results in a huge amount of information that requires a balance between Big Data and Edge Computing approaches to process the whole collected data. The evaluation of the sensing data is presented together with the proposed architecture. Antonio Deusany de Carvalho Junior, Victor Seiji Hariki, Alfredo Goldman |
NCA | 3 |
| 2017 | Is Intel high performance analytics toolkit a good alternative to Apache Spark?abstractThis paper compares the performance and stability of two Big Data processing tools: the Apache Spark and the High Performance Analytics Toolkit (HPAT). The comparison was performed using two applications: a unidimensional vector sum and the K-means clustering algorithm. The experiments were performed in distributed and shared memory environments with different numbers and configurations of virtual machines. By analyzing the results we are able to conclude that HPAT has performance improvements in relation to Apache Spark in our case studies. We independently validated the results and potential presented by the HPAT developers. We also provide an analysis of both frameworks in the presence of failures. Rafael Aquino de Carvalho, Alfredo Goldman, Gerson G. H. Cavalheiro |
NCA | 2 |
| 2017 | Software Development Practices PatternsabstractOur ultimate goal is to propose a catalog with recommendations on how to organize the work of programmers. In this research we intend to provide experiments to explore the most suitable forms to allow programmers to develop software, either alone, in pair programming or in group. We also explore other approaches like code review. Our goal is not only to reduce the software development cost, but also to improve programmers life quality. Herez Moise Kattan, Alfredo Goldman |
XP | 2 |
| 2017 | Effects of Technical Debt Awareness: A Classroom StudyabstractTechnical Debt is a metaphor that has, in recent years, helped developers to think about and to monitor software quality. The metaphor refers to flaws in software (usually caused by shortcuts to save time) that may affect future maintenance and evolution. We conducted an empirical study in an academic environment, with nine teams of graduate and undergraduate students during two offerings of a laboratory course on Extreme Programming (XP Lab). The teams had a comprehensive lecture about several alternative ways to identify and manage Technical Debt. We monitored the teams, performed interviews, did close observations and collected feedback. The results show that the awareness of Technical Debt influences team behavior. Team members report thinking and discussing more about software quality after becoming aware of Technical Debt in their projects. Graziela Tonin, Alfredo Goldman, Carolyn B. Seaman, Diogo Pina |
XP | 2 |
| 2017 | Autotuning CUDA compiler parameters for heterogeneous applications using the OpenTuner frameworkabstractSummary A Graphics Processing Unit (GPU) is a parallel computing coprocessor specialized in accelerating vector operations. The enormous heterogeneity of parallel computing platforms justifies and motivates the development of automated optimization tools and techniques. The Algorithm Selection Problem consists in finding a combination of algorithms, or a configuration of an algorithm, that optimizes the solution of a set of problem instances. An autotuner solves the Algorithm Selection Problem using search and optimization techniques. In this paper, we implement an autotuner for the Compute Unified Device Architecture compiler's parameters using the OpenTuner framework. The autotuner searches for a set of compilation parameters that optimizes the time to solve a problem. We analyze the performance speedups, in comparison with high‐level compiler optimizations, achieved in three different GPU devices, for 17 heterogeneous GPU applications, 12 of which are from the Rodinia Benchmark Suite. The autotuner often beats the compiler's high‐level optimizations, but underperformed for some problems. We achieved over 2x speedup forGaussian Eliminationand almost 2x speedup forHeart Wall, both problems from the Rodinia Benchmark, and over 4x speedup for a matrix multiplication algorithm. Copyright © 2017 John Wiley & Sons, Ltd. Pedro Bruel, Marcos Amaris, Alfredo Goldman |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | Computer architecture and high performance computingabstractComputer architecture and high performance computingThis special issue of Concurrency and Computation Practice and Experience gathers eleven selected research articles that were previously presented at the Brazilian "XVII Simpósio em Sistemas Computacionais de Alto Desempenho," WSCAD 2016, held in conjunction with 28th International Symposium on Computer Architecture and High Performance Computing, SBAC-PAD 2015, Florianópolis, SC, Brazil, from the 19th to the 21st October 2015.Since 2000, this workshop has presented important and interesting research in the fields of computer architectures, high performance computing, and distributed systems.The scope of the current special issue is broad and representative of the multidisciplinary nature of high performance and distributed computing, covering a wide range of subjects such as architecture issues, compiler optimization, analysis of HPC applications, job scheduling, and energy efficiency.The title of the first paper is "An efficient virtual system clock for the wireless Raspberry Pi computer platform," by Diego L. C. Dutra, Edilson C. Corrêa, and Claudio L. Amorim [1].In this paper, the authors present the design and experimental evaluation of an implementation of the RVEC virtual system clock in the Linux kernel for the EE (Energy-Efficient) Wireless Raspberry Pi (RasPi) platform.In the RasPi platform, the use of DVFS (Dynamic Voltage and Frequency) for reducing the energy consumption hinders the direct use of the cycle count of the ARM11 processor core for building an efficient system clock.Therefore, a distinct feature of RVEC is to obviate this obstacle, such that it can make use of the cycle count circuit for precise and accurate time measurements, concurrently with the use of DVFS by the operating system of the ARM11 processor core.In the second contribution, entitled "Portability with efficiency of the advection of BRAMS between multi-core and many-core architectures," the authors, Manoel Baptista Silva Junior, Jairo Panetta, and Stephan Stephany [2], show the feasibility of writing a single portable code embedding both interfaces (the OpenMP programming interface and OpenACC).It presents acceptable efficiency when executed on nodes with multi-core or many-core architecture.The code chosen as a case study is the advection of scalars, a part of the dynamics of the regional atmospheric model Brazilian Regional Atmospheric Modeling System (BRAMS).The dynamics of this model is hard to parallelize due to data dependencies between adjacent grid points.Single-node executions of the advections of scalars for different grid sizes using OpenMP or OpenACC yielded similar speed-ups, showing the feasibility of the proposed approach.In the third contribution, entitled "SMT-based context-bounded model checking for CUDA programs," the authors ( Alfredo Goldman, Luciana Arantes, Edward Moreno |
Concurr. Comput. Pract. Exp. | 1 |
| 2016 | A comparison of GPU execution time prediction using machine learning and analytical modelingabstractToday, most high-performance computing (HPC) platforms have heterogeneous hardware resources (CPUs, GPUs, storage, etc.) A Graphics Processing Unit (GPU) is a parallel computing coprocessor specialized in accelerating vector operations. The prediction of application execution times over these devices is a great challenge and is essential for efficient job scheduling. There are different approaches to do this, such as analytical modeling and machine learning techniques. Analytic predictive models are useful, but require manual inclusion of interactions between architecture and software, and may not capture the complex interactions in GPU architectures. Machine learning techniques can learn to capture these interactions without manual intervention, but may require large training sets. In this paper, we compare three different machine learning approaches: linear regression, support vector machines and random forests with a BSP-based analytical model, to predict the execution time of GPU applications. As input to the machine learning algorithms, we use profiling information from 9 different applications executed over 9 different GPUs. We show that machine learning approaches provide reasonable predictions for different cases. Although the predictions were inferior to the analytical model, they required no detailed knowledge of application code, hardware characteristics or explicit modeling. Consequently, whenever a database with profile information is available or can be generated, machine learning techniques can be useful for deploying automated on-line performance prediction for scheduling applications on heterogeneous architectures containing GPUs. Marcos Amaris, Raphael Y. de Camargo, Mohamed Dyab, Alfredo Goldman, Denis Trystram |
NCA | 4 |
| 2016 | Using NAS Parallel Benchmarks to evaluate HPC performance in cloudsabstractCloud computing is a reality nowadays, however there are few studies trying to understand what happens in the actual cloud infrastructures for HPC applications. The focus of this study is the evaluation of NAS Parallel Benchmarks on cloud computing environments. We analyze the execution of applications from NAS Parallel Benchmarks (LU and SP), comparing the execution behavior in different infrastructures: a public cloud, a private cloud and a NUMA multiprocessor system. Our broad goal is to estimate the performance of actual HPC applications on cloud, based on its communication characteristics. We conclude that HPC users should be careful with Virtual Machines with higher virtual CPU count, thanks to Google usage of Hyper-Threading technology and Virtual Machine instances scheduling. Thiago Kenji Okada, Alfredo Goldman, Gerson G. H. Cavalheiro |
NCA | 2 |
| 2015 | A Simple BSP-based Model to Predict Execution Time in GPU ApplicationsabstractModels are useful to represent abstractions of software and hardware processes. The Bulk Synchronous Parallel (BSP) is a bridging model for parallel computation that allows algorithmic analysis of programs on parallel computers using performance modeling. The main idea of BSP model is the treatment of communication and computation as abstractions of a parallel system. Meanwhile, the use of GPU devices are becoming more widespread and they are currently capable of performing efficient parallel computation for applications that can be decomposed on thousands of simple threads. However, few models for predicting application execution time on GPUs have been proposed. In this work we present a simple and intuitive BSP-based model for predicting the CUDA application execution times on GPUs. The model is based on the number of computations and memory accesses of the GPU, with additional information on cache usage obtained from profiling. Scalability, divergence, effect of optimizations and differences of architectures are adjusted by a single parameter. We evaluated our model using two applications and six different boards. We showed by using profile information for a single board, that the model is general enough to predict the execution time of an application with different input sizes and on different boards with the same architecture. Our model predictions were within 0.8 to 1.2 times the measured execution times, which are reasonable for such a simple model. These results indicate that the model is good enough to generalize the predictions for different problem sizes and GPU configurations. Marcos Amaris, Daniel Cordeiro, Alfredo Goldman, Raphael Y. de Camargo |
HiPC | 3 |
| 2015 | IEEE Services Visionary Track on Service Composition for the Future Internet (SCFI 2015)abstractThis document is a summary paper reporting on the IEEE Services 2015 Visionary Track on Service Composition for the Future Internet (SCFI 2015). Marco Autili, Alfredo Goldman, Massimo Tivoli |
SERVICES | 2 |
| 2015 | Fostering effective inter-team knowledge sharing in agile software development
Viviane A. Santos, Alfredo Goldman, Cleidson R. B. de Souza |
Empir. Softw. Eng. | 2 |
| 2014 | IEEE First International Workshop on Service Orchestration and Choreography for the Future Internet (OrChor 2014)abstractSummary of the IEEE First International Workshop on Service Orchestration and Choreography for the Future Internet (OrChor 2014) Marco Autili, Alfredo Goldman, Massimo Tivoli |
SERVICES | 2 |
| 2014 | A comprehensive view of Hadoop research - A systematic literature review
Ivanilton Polato, Reginaldo Ré, Alfredo Goldman, Fabio Kon |
J. Netw. Comput. Appl. | 3 |
| 2013 | A NUMA-Aware Runtime Environment for the Actor ModelabstractThe actor model is present in several mission-critical systems, such as those supporting WhatsApp and Twitter. These systems serve thousands of clients simultaneously, therefore demanding substantial computing resources usually provided by multiprocessor and multicore platforms. Non-Uniform Memory Access (NUMA) architectures account for an important share of these platforms. Yet, little or no research has been done on the suitability of the current actor runtime environments for these machines. Current runtime environments assume a flat memory space, thus not performing as well as they could. The NUMA environment presents challenges to the actor model runtime environment in fields varying from memory management to scheduling and load-balancing. In this document we analyze and characterize actor based applications to, in light of the above, propose improvements to actor runtime environments. As a proof of concept, we have applied our ideas in a real actor runtime environment, the Erlang virtual machine. This modified virtual machine uses the NUMA characteristics and the application knowledge to take better memory management, scheduling and load-balancing decisions. We have evaluated this modified runtime environment using standard benchmarks and, taking the default virtual machine as a baseline, we improved the performance of the tested applications by a factor of 2.50 on the best case while limiting our slowdown on the worst case by a factor of 1.09. Emilio Francesquini, Alfredo Goldman, Jean-François Méhaut |
ICPP | 2 |
| 2013 | A Delay-Tolerant Network Routing Algorithm Based on Column GenerationabstractDelay-Tolerant Networks (DTN) model systems that are characterized by intermittent connectivity and frequent partitioning. Routing in DTNs has drawn much research effort recently. Since very different kinds of networks fall in the DTN category, many routing approaches have been proposed. In particular, the routing layer in some DTNs have information about the schedules of contacts between nodes and about data traffic demand. Such systems can benefit from a previously proposed routing algorithm based on linear programming that minimizes the average message delay. This algorithm, however, is known to have performance issues that limit its applicability to very simple scenarios. In this work, we propose an alternative linear programming approach for routing in Delay-Tolerant Networks. We show that our formulation is equivalent to that presented in a seminal work in this area, but it contains fewer LP constraints and has a structure suitable to the application of Column Generation (CG). Simulation shows that our CG implementation arrives at an optimal solution up to three orders of magnitude faster than the original linear program in the considered DTN examples. Guilherme Amantea, Hervé Rivano, Alfredo Goldman |
NCA | 3 |
| 2012 | Malleable resource sharing algorithms for cooperative resolution of problemsabstractGiven multiple parallel heuristics solving the same problem, we are interested in combining them for taking advantage of their diversity. We propose to use the algorithm portfolio model of execution. In this model, we have multiple resources on which the candidate heuristics can be executed. An instance is solved through a concurrent execution of heuristics (each on a fraction of resources) that is stopped as soon as one of them completes its execution. The efficiency of this model depends among other things of the resource sharing adopted in a concurrent execution. In most algorithm portfolio studies, the resources fraction of a heuristic is fixed. In this paper, we consider malleable algorithm portfolio. In this portfolio model, the fraction of resources of a heuristic can be changed during its execution. We extend the computational model proposed in [1] to formalize the problem of resource sharing construction in malleable portfolio. We then propose an efficient algorithm based on the combination of two guaranteed approximation algorithms for solving it. Finally, we evaluate the proposed algorithm with multiple simulations on a database of SAT solvers. The obtained results show that even in considering that the resource allocation of a heuristic can just be changed once, malleable allocations in comparison to static ones lead to an improvement of the spent time for solving an instance in algorithm portfolio.time for solving an instance in algorithm portfolio. Alfredo Goldman, Yanik Ngoko, Denis Trystram |
IEEE Congress on Evolutionary Computation | 1 |
| 2011 | Introduction
Leonel Sousa, Frédéric Suter, Alfredo Goldman, Rizos Sakellariou, Oliver Sinnen |
Euro-Par (1) | 3 |
| 2011 | Formalization of the Necessary and Sufficient Connectivity Conditions to the Distributed Mutual Exclusion Problem in Dynamic NetworksabstractInternational audience Paulo Floriano, Alfredo Goldman, Luciana Arantes |
NCA | 2 |
| 2011 | A view towards Organizational Learning: An empirical study on Scrum implementation
Viviane A. Santos, Alfredo Goldman, Ana Carolina M. Shinoda, André L. Fischer |
SEKE | 2 |
| 2011 | From Manufacture to Software Development: A Comparative Review
Eduardo T. Katayama, Alfredo Goldman |
XP | 2 |
| 2011 | An Approach on Applying Organizational Learning in Agile Software Organizations
Viviane A. Santos, Alfredo Goldman |
XP | 2 |
| 2011 | Adaptive fault tolerance mechanisms for opportunistic environments: a mobile agent approachabstractSUMMARY The mobile agent paradigm has emerged as a promising alternative to overcome the construction challenges of opportunistic grid environments. This model can be used to implement mechanisms that enable application execution progress even in the presence of failures such as the mechanisms provided by the MAG middleware (Mobile Agents for Grids). MAG includes retrying, replication, and checkpointing as fault tolerance techniques; they operate independently from each other and they are not capable of detecting changes on resource availability. In this paper, we describe a MAG extension that is capable of migrating agents when nodes fail, which optimizes application progress by keeping only the most advanced checkpoint, and also migrates slow replicas. The proposed approach was evaluated via simulations and experiments, which showed significant improvements. Copyright © 2011 John Wiley & Sons, Ltd. Vinicius Pinheiro, Alfredo Goldman, Fabio Kon |
Concurr. Comput. Pract. Exp. | 2 |
| 2010 | Reinforcing the Learning of Agile Practices Using Coding Dojos
Mariana V. Bravo, Alfredo Goldman |
XP | 2 |
| 2010 | Open Source and Agile Methods: Two Worlds Closer than It Seems
Hugo Corbucci, Alfredo Goldman |
XP | 2 |
| 2010 | Application execution management on the InteGrade opportunistic grid middleware
Francisco José da Silva e Silva, Fabio Kon, Alfredo Goldman, Marcelo Finger, Raphael Y. de Camargo, Fernando Castor Filho, Fábio M. Costa |
J. Parallel Distributed Comput. | 3 |
| 2010 | Performance evaluation of routing protocols for MANETs with known connectivity patterns using evolving graphs
Afonso Ferreira, Alfredo Goldman, Julian Monteiro |
Wirel. Networks | 2 |
| 2009 | Combining multiple heuristics on discrete resourcesabstractIn this work we study the portfolio problem which is to find a good combination of multiple heuristics to solve given instances on parallel resources in minimum time. The resources are assumed to be discrete, it is not possible to allocate a resource to more than one heuristic. Our goal is to minimize the average completion time of the set of instances, given a set of heuristics on homogeneous discrete resources. This problem has been studied in the continuous case in [T. Sayag et al., 2006]. We first show that the problem is hard and that there is no constant ratio polynomial approximation unlessP=NPin the general case. Then, we design several approximation schemes for a restricted version of the problem where each heuristic must be used at least once. These results are obtained by using oracle with several guesses, leading to various tradeoff between the size of required information and the approximation ratio. Some additional results based on simulations are finally reported using a benchmark of instances on SAT solvers. Marin Bougeret, Pierre-François Dutot, Alfredo Goldman, Yanik Ngoko, Denis Trystram |
IPDPS | 3 |
| 2008 | A MILP Approach to Schedule Parallel Independent TasksabstractWe propose a new mixed integer linear programming approach to solve the classical problem of scheduling independent parallel tasks without preemption. We propose a formulation where the goal is to minimize the makespan.Then we show the flexibility of this approach by extending the result to the contiguous case. We validate this approach with some experiments on the execution times and comparing the optimal results with the solutions provided by list algorithms. Alfredo Goldman, Yanik Ngoko |
ISPDC | 1 |
| 2007 | Load Balancing on an Interactive Multiplayer Game Server
Daniel Cordeiro, Alfredo Goldman, Dilma Da Silva |
Euro-Par | 2 |
| 2007 | On the Evaluation of Shortest Journeys in Dynamic NetworksabstractThe assessment of routing protocols for wireless networks is a difficult task, because of the networks' highly dynamic behavior and the absence of benchmarks. However, some of these networks, such as intermittent wireless sensors networks, periodic or cyclic networks, and low earth orbit (LEO) satellites systems, have more predictable dynamics, as the temporal variations in the network topology are somehow deterministic, which may make them easier to study. The graph theoretic model - the evolving graphs - was proposed to help capture the dynamic behavior of these networks, in view of the construction of least cost routing and other algorithms. Our recent experiments showed that evolving graphs have all the potentials to be an effective and powerful tool in the development of routing protocols for dynamic networks. In this paper, we evaluated the shortest journey evolving graph algorithm when used in a routing protocol for MANETs. We use the NS2 network simulator to compare this first implementation to the four well known protocols, namely AODV, DSR, DSDV, and OLSR. In this paper we present simulation results on the energy consumption of the nodes. We also included other EG protocol, namely EGForemost, in the experiments. Afonso Ferreira, Alfredo Goldman, Julian Monteiro |
NCA | 2 |
| 2007 | Tracking the Evolution of Object-Oriented Quality Metrics on Agile Projects
Danilo T. Sato, Alfredo Goldman, Fabio Kon |
XP | 2 |
| 2006 | Performance Evaluation of Dynamic Networks using an Evolving Graph Combinatorial ModelabstractThe highly dynamic behavior of wireless networks make them very difficult to evaluate, e.g. as far as the performance of routing algorithms is concerned. However, some of these networks, such as intermittent wireless sensors networks, periodic or cyclic networks, and low Earth orbit (LEO) satellites systems have more predictable dynamics, as the temporal variations in the network topology are somehow deterministic. Recently, a graph theoretic model-the evolving graphs-was proposed to help capture the dynamic behavior of these networks, in view of the construction of least cost routing and other algorithms. The algorithms and insights obtained through this model are theoretically very efficient and intriguing. However, there is no study on the uses of these theoretical results into practical situations. Therefore, the objective of this work is to analyze the applicability of the evolving graph theory in the construction of efficient routing protocols in realistic scenarios. In this paper, we used the NS2 network simulator to first implement an evolving graph based routing protocol, and then to evaluate such protocol compared to three major ad-hoc protocols (DSDV, DSR, AODV). Interestingly, our experiments showed that evolving graphs have all the potentials to be an effective and powerful tool in the development of algorithms for dynamic networks, with predictable dynamics at least. In order to make this model widely applicable, however, some practical issues still have to be addressed and incorporated into the model, like stochastically predictable behavior. We also discuss such issues in this paper, as a result of our experience Julian Monteiro, Alfredo Goldman, Afonso Ferreira |
WiMob | 2 |
| 2006 | Checkpointing BSP parallel applications on the InteGrade Grid middlewareabstractAbstract InteGrade is a Grid middleware infrastructure that enables the use of idle computing power from user workstations. One of its goals is to support the execution of long‐running parallel applications that present a considerable amount of communication among application nodes. However, in an environment composed of shared user workstations spread across many different LANs, machines may fail, become inaccessible, or may switch from idle to busy very rapidly, compromising the execution of the parallel application in some of its nodes. Thus, to provide some mechanism for fault tolerance becomes a major requirement for such a system. In this paper, we describe the support for checkpoint‐based rollback recovery of Bulk Synchronous Parallel applications running over the InteGrade middleware. This mechanism consists of periodically saving application state to permit the application to restart its execution from an intermediate execution point in case of failure. A precompiler automatically instruments the source code of a C/C++ application, adding code for saving and recovering application state. A failure detector monitors the application execution. In case of failure, the application is restarted from the last saved global checkpoint. Copyright © 2005 John Wiley & Sons, Ltd. Raphael Y. de Camargo, Andrei Goldchleger, Fabio Kon, Alfredo Goldman |
Concurr. Comput. Pract. Exp. | 4 |
| 2006 | Exchanging messages of different sizes
Alfredo Goldman, Joseph G. Peters, Denis Trystram |
J. Parallel Distributed Comput. | 1 |
| 2005 | Scheduling Moldable BSP Tasks
Pierre-François Dutot, Marco Aurélio Stelmar Netto, Alfredo Goldman, Fabio Kon |
JSSPP | 3 |
| 2005 | Portable checkpointing and communication for BSP applications on dynamic heterogeneous Grid environmentsabstractExecuting long-running parallel applications in opportunistic grid environments composed of heterogeneous, shared user workstations, is a daunting task. Machines may fail, become inaccessible, or may switch from idle to busy unexpectedly, compromising the execution of applications. A mechanism for fault-tolerance that supports these heterogeneous architectures is an important requirement for such a system. In this paper, we describe the support for fault-tolerant execution of BSP parallel applications on heterogeneous, shared workstations, precompiler instruments application source code to save state periodically into checkpoint files. In case of failure, it is possible to recover the stored state from these files. Generated checkpoints are portable and can be recovered in a machine of different architecture, with data representation conversions being performed at recovery time. The precompiler also modifies BSP parallel applications to allow execution on a grid composed of machines with different architectures. We implemented a monitoring and recovering infrastructure in the InteGrade grid middleware. Experimental results evaluate the overhead incurred and the viability of using this approach in a grid environment. Raphael Y. de Camargo, Fabio Kon, Alfredo Goldman |
SBAC-PAD | 3 |
| 2004 | InteGrade: object-oriented Grid middleware leveraging the idle computing power of desktop machinesabstractAbstract Grid computing technology improves the computing experiences at organizations by effectively integrating distributed computing resources. However, just a small fraction of currently available Grid infrastructures focuses on reutilization of existing commodity computing resources. This paper introducesInteGrade, a novel object‐oriented middleware Grid infrastructure that focuses on leveraging the idle computing power of shared desktop machines. Its features include support for a wide range of parallel applications and mechanisms to assure that the owners of shared resources do not perceive any loss in the quality of service. A prototype implementation is under construction and the current version is available for download. Copyright © 2004 John Wiley & Sons, Ltd. Andrei Goldchleger, Fabio Kon, Alfredo Goldman, Marcelo Finger, Germano Capistrano Bezerra |
Concurr. Pract. Exp. | 3 |
| 2004 | A model for parallel job scheduling on dynamical computer GridsabstractAbstract This work presents a model that allows the execution of parallel applications in a Grid environment. Our main focus is on how to share the idle cycles of clusters and computers to execute real parallel applications. We present a new model which introduces the notions of locality and adaptability. The locality is used for job allocation, and for job migration. The adaptability provides a simple mechanism to allow clusters to join or leave a Grid. We also propose the middleware architecture to implement the model, and provide some simulation results to show the main characteristics of the model. Copyright © 2004 John Wiley & Sons, Ltd. Alfredo Goldman, Carlos Queiroz |
Concurr. Pract. Exp. | 1 |
| 2004 | An efficient parallel algorithm for solving the Knapsack problem on hypercubes
Alfredo Goldman, Denis Trystram |
J. Parallel Distributed Comput. | 1 |
| 2003 | A Parallel Algorithm for Enumerating CombinationsabstractWe propose an efficient parallel algorithm with simple static and dynamic scheduling for generating combinations. It can use any number of processors (NPlesn-m+1) in order to generate the set of all combinations of C(n,m). The main characteristic of this algorithm is to require no integer larger than n during the whole computation. The performance results show that even without a perfect load balance, this algorithm has very good performance, mainly when n is large. Besides, the dynamic algorithm presents a good performance on heterogeneous parallel platforms Martha Torres, Alfredo Goldman, Junior Barrera |
ICPP | 2 |
| 2003 | 1-optimality of static BSP computations: scheduling independent chains as a case study
Alfredo Goldman, Grégory Mounié, Denis Trystram |
Theor. Comput. Sci. | 1 |
| 2002 | Scalable Algorithms for Complete Exchange on Multi-Cluster NetworksabstractInside a dedicated parallel computer the communication times were generally modeled in the same way, independently of which processors communicate. In a network where the links among the computers are heterogeneous, or in hierarchical clusters, this might not be true anymore. Computers that have faster links, or are closer to each other should be able to exchange messages faster. These differences on communication times should be considered, not only for attributing tasks to the processors but also in global synchronization/communication. The goal of this paper is to study irregular all-to-all communications in a network of dedicated clusters. Alfredo Goldman |
CCGRID | 1 |
| 1998 | Near optimal algorithms for scheduling independent chains in BSPabstractThe aim of this work is to show that scheduling a set of independent chains on a parallel machine under the BSP model is a difficult optimization problem which can be easily approximated in practice. BSP is a machine independent computational model which is becoming more and more popular. Finding the optimal solution when the number of processors is fixed is shown to be hard. Efficient heuristics including communications are proposed and analyzed. We particularly focus on the influence of synchronization between consecutive supersteps. Simulations of a large number of instances have been carried out to complement the theoretical worst case analysis. They confirm the very good behaviour of the algorithm on average. Alfredo Goldman, Grégory Mounié, Denis Trystram |
HiPC | 1 |