EDBT 2026 Demo / reviewers in the wild / expert
Cristina Boeres
dblp:89/6061 · also Maria Cristina Silva Boeres
· DBLP profile ↗
36ranked-venue papers
14as first author
7since 2021 · last 2026
0000-0002-1679-6643ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 10 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorComputer networks · 1Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MetaCS-FL: A metaheuristic-based framework for client selection in federated learning systems
Alan L. Nunes, Cristina Boeres, Laércio Lima Pilla, Lúcia M. A. Drummond |
Future Gener. Comput. Syst. | 2 |
| 2025 | Spotting the Right Cloud Instances with Multiple AWS EC2 FleetsabstractHigh Performance Computing (HPC) is increasingly transitioning to the cloud, although cost remains a significant barrier. While On-demand instances and Committed Use Discounts provide predictable pricing, the Spot market offers an appealing opportunity for substantial cost savings – though it does not guarantee resource availability. Effectively managing multiple instances for parallel HPC applications is essential. Services like AWS EC2 Fleet or Spot Fleet help with this, but they come with limitations, notably being constrained to a single region. Furthermore, simply selecting the lowest-priced instance often leads to suboptimal performance and, surprisingly, higher overall costs. To truly economize, a more sophisticated approach is required: one that involves profiling applications and instances to understand their intricate cost-performance trade-offs. The most cost-effective execution prioritizes instances that strike a better balance between the price per hour being charged and the actual performance they offer the application, even if its current Spot price is not the lowest available. This paper addresses these challenges by conducting a thorough analysis of existing EC2 (Spot) Fleet policies and introducing Fleet-MR, a novel multi-region instance selection framework. Fleet-MR aims to improve execution times or reduce costs, and its effectiveness is validated through experimental evaluations. Daniel B. Sodré, Lucas Serrano, Miguel De Lima, Cristina Boeres, Lúcia M. A. Drummond, Vinod E. F. Rebello |
SBAC-PAD | 4 |
| 2024 | A Framework for Executing Long Simulation Jobs Cheaply in the CloudabstractThis paper presents the framework SIM@ ClOUD that optimizes cost-related resource allocation decisions for simulation jobs in cloud environments. SIM@ CLOUD offers comprehensive management of simulations throughout their execution life-cycle in the cloud, including the selection of Virtual Machine (VM) types across different regions and markets. By leveraging Spot VMs and application checkpointing, the framework transparently reduces the monetary costs associated with the execution without client intervention. Historical data analysis enables the prediction of simulation execution times, which is refined further by a dynamic predictor for adaptive VM selection. SIM@ CLOUD is being deployed in an industrial setting and employs a cachebased storage solution to improve access latency to in-house data by VMs located in geographically distinct regions. An evaluation carried out on AWS EC2, using real oil reservoir simulations, demonstrates the effectiveness of the framework. Alan L. Nunes, Daniel B. Sodré, Cristina Boeres, José Viterbo, Lúcia M. A. Drummond, Vinod E. F. Rebello, Luan Teylo, Felipe Albuquerque Portella, Paulo J. B. Estrela, Renzo Q. Malini |
IC2E | 3 |
| 2024 | Optimal Time and Energy-Aware Client Selection Algorithms for Federated Learning on Heterogeneous ResourcesabstractFederated Learning systems allow training machine learning models distributed across multiple clients, each one using private local data. Iteratively, the clients send their training contributions to a server, which performs a merge to produce an enhanced global model. Due to resource and data heterogeneity, client selection is crucial to optimize the system efficiency and improve the global model generalization. Selecting more clients is likely to increase the overall energy consumption, while a small number of clients may decline the performance of the trained model or require longer training time. We propose two time- and energy-aware client selection algorithms, MEC and ECMTC, which are proven regarding their optimality and evaluated against state-of-the-art algorithms on an extensive series of experiments in both simulation and HPC platform scenarios. The results indicate the benefits of jointly optimizing the time and energy consumption metrics using our proposals. Alan L. Nunes, Cristina Boeres, Lúcia M. A. Drummond, Laércio Lima Pilla |
SBAC-PAD | 2 |
| 2023 | Optimizing computational costs of Spark for SARS-CoV-2 sequences comparisons on a commercial cloudabstractSummary Cloud computing is currently one of the prime choices in the computing infrastructure landscape. In addition to advantages such as the pay‐per‐use bill model and resource elasticity, there are technical benefits regarding heterogeneity and large‐scale configuration. Alongside the classical need for performance, for example, time, space, and energy, there is an interest in the financial cost that might come from budget constraints. Based on scalability considerations and the pricing model of traditional public clouds, a reasonable optimization strategy output could be the most suitable configuration of virtual machines to run a specific workload. From the perspective of runtime and monetary cost optimizations, we provide the adaptation of a Hadoop applications execution cost model extracted from the literature aiming at Spark applications modeled with the MapReduce paradigm. We evaluate our optimizer model executing an improved version of the Diff Sequences Spark application to perform SARS‐CoV‐2 coronavirus pairwise sequence comparisons using the AWS EC2's virtual machine instances. The experimental results with our model outperformed 80% of the random resource selection scenarios. By only employing spot worker nodes exposed to revocation scenarios rather than on‐demand workers, we obtained an average monetary cost reduction of 35.66% with a slight runtime increase of 3.36%. Alan L. Nunes, Alba Cristina Magalhaes Alves de Melo, Claude Tadonki, Cristina Boeres, Daniel de Oliveira 0001, Lúcia M. A. Drummond |
Concurr. Comput. Pract. Exp. | 4 |
| 2021 | Comparing SARS-CoV-2 Sequences using a Commercial Cloud with a Spot Instance Based Dynamic SchedulerabstractThere has been an increasing interest in running High Performance Computing (HPC) applications in the cloud, mainly due to rapid resource provisioning and significant reduction of operational costs. Biological sequence comparison is an important HPC application that compares sequences in search of similarities. MASA-OpenMP is a highly optimized sequence comparison tool that obtains optimal results. Yet, it can take a long time, depending on the number of sequences compared and their lengths. The Covid-19 pandemic study is of particular interest nowadays, and the comparison of SARS-CoV-2 sequences is crucial to understanding this disease. In this paper, we compare SARS-CoV-2 sequences with MASA-OpenMP in the Amazon Elastic Compute Cloud (Amazon EC2), using both spot and on-demand instances. To efficiently execute a MASA-OpenMP application composed of more than 22,000 tasks on EC2 respecting a given deadline, we propose an execution modeling for MASA-OpenMP on top of the Burst-HADS framework. Burst-HADS is a spot instance-based dynamic scheduler for Bag-of-Tasks applications in the cloud, which minimizes both execution time and financial costs regarding a given deadline even in the presence of spot interruptions. Performance results reveal that, by using spots, our Burst-HADS strategy considerably reduces the monetary cost for executing 22,600 SARS-CoV-2 sequence comparisons with MASA-OpenMP when contrasted to the on-demand only approach. We also show that our strategy can meet the deadlines, even in scenarios with several spot interruptions. Luan Teylo, Alan L. Nunes, Alba Cristina Magalhaes Alves de Melo, Cristina Boeres, Lúcia M. A. Drummond, Natália Florencio Martins |
CCGRID | 4 |
| 2021 | Towards optimizing the execution of spark scientific workflows using machine learning-based parameter tuningabstractSummary In the last few years, Apache Spark has become a de facto the standard framework for big data systems on both industry and academy projects. Spark is used to execute compute‐ and data‐intensive workflows in distinct areas like biology and astronomy. Although Spark is an easy‐to‐install framework, it has more than one hundred parameters to be set, besides domain‐specific parameters of each workflow. In this way, to execute Spark‐based workflows efficiently, the user has to fine‐tune a myriad of Spark and workflow parameters (eg, partitioning strategy, the average size of a DNA sequence, etc.). This configuration task cannot be manually performed in a trial‐and‐error manner since it is tedious and error‐prone. This article proposes an approach that focuses on generating interpretable predictive machine learning models (ie, decision trees), and then extract useful rules (ie, patterns) from these models that can be applied to configure parameters of future executions of the workflow and Spark for nonexperts users. In the experiments presented in this article, the proposed parameter configuration approach led to better performance in processing Spark workflows. Finally, the approach introduced here reduced the number of parameters to be configured by identifying the most relevant domain‐specific ones related to the workflow performance in the predictive model. Douglas E. M. de Oliveira, Fábio Porto 0001, Cristina Boeres, Daniel de Oliveira 0001 |
Concurr. Comput. Pract. Exp. | 3 |
| 2020 | Static job scheduling for environments with vertical elasticityabstractSummary In virtualized environments, such as Clouds, allocating a fixed amount of resources to a job a priori, may result in underutilization of the shared host. Meanwhile, vertical elasticity can be adopted to reduce the impact by resizing virtual machines (VMs) dynamically, in conjunction with suspension and/or migration before the host had been overloaded. In order to reduce the number of such events, but still avoid overloading the host, it is important to find an effective initial VM allocation. To achieve this, the scheduler must work in unison with the local host elasticity controllers to reduce interference. This work proposes and evaluates a framework for job scheduling in elastic memory managed virtualized environments. The memory elasticity management in clouds (MEMiC) framework is a two‐tier VM scheduler for batch jobs which attempts to predict the impact caused by competition for the memory of a host in shared Cloud‐like environments for harnessing VM allocation, suspension, and migration. Evaluations show that the MEMiC achieves a reduction of approximately 19% in the total time for jobs execution in comparison to other approaches, by reducing the degree of interference. Henrique Kloh, Vinod E. F. Rebello, Cristina Boeres, Bruno Schulze, Mariza Ferro |
Concurr. Comput. Pract. Exp. | 3 |
| 2019 | New advances in high-performance computing systemsabstractThis Special Issue of Concurrency and Computation: Practice and Experience gathers seven selected research articles, which are the result of extended work previously presented at the Brazilian “XVII Simpósio em Sistemas Computacionais de Alto Desempenho,” WSCAD 2016, held in conjunction with the 28th International Symposium on Computer Architecture and High Performance Computing, SBAC-PAD 2016, Aracaju, SE, Brazil, October 5-7, 2016. Since 2000, this workshop has presented important and interesting research in the fields of computer architecture, high-performance computing, and distributed systems. The scope of the current special Issue is broad and representative of the multidisciplinary nature of high-performance and distributed computing, covering important issues of the field, such as parallel applications, GPU computing, energy consumption, and cloud computing. The papers and their main focus are the following. The paper entitled “Parallel rule-based selective sampling and on-demand learning to rank” by Mateus F. e Freitas, Daniel X. Sousa, Wellington S. Martins, Thierson C. Rosa, Rodrigo M. Silva, and Marcos A. Gonçalves uses parallelism techniques to improve the performance of the Learning to Rank (L2R) task. More specifically, the authors1 propose two methods to exploit parallelism on rule-based systems: Learning to Rank (L2R) training data sets using selective sampling and query-customized ranking models generated on the fly. The authors propose parallel algorithms and GPU implementations for these two cases showing that data set reduction takes only a few seconds with speedups of up to 148× over a serial baseline and that queries can be processed in only a few milliseconds with speedups of 1000× over a serial baseline and 29× over a parallel baseline for the best case. The extended work provides the implementations on multiple GPUs, further increasing the speedup over the baselines. Focusing on GPU computing, in the contribution2 entitled “Maximizing the GPU resource usage by reordering concurrent kernels submission,” authors Rommel A.Q. Cruz, Cristiana Bentes, Bernardo Breder, Eduardo Vasconcellos, Esteban Clua, Pablo M.C. de Carvalho, and Lúcia M.A. Drummond propose a reordering strategy to identify the best order to submit the kernels to execute on the GPU, aiming at maximizing its utilization and increasing the overall throughput. The authors model the problem as a series of knapsack problems and use a dynamic programming approach to solve them. The amount of resources is modeled as the knapsack capacity, and the strategy tries to fulfill the knapsack with kernels that take the most advantage of the available resources, favoring kernels with smaller execution time. The proposed strategy was evaluated using real-world and synthetic applications with different numbers of kernels and resource requirements. The results show that the reordering strategy provides significant gains in the average turnaround time and system throughput compared to the kernels submission implemented in modern GPUs. Tackling energy efficiency issues, in the work3 entitled “Energy efficiency and I/O performance of low-power architectures,” authors Pablo J. Pavan, Ricardo K. Lorenzoni, Vinícius R. Machado, Jean L. Bez, Edson L. Padoin, Francieli Z. Boito, Philippe O.A. Navaux, and Jean-François Méhaut present an energy efficiency and I/O performance analysis of low-power architectures with the goal of evaluating the viability of using them as storage servers. The results show that despite the fact that the power demand of the storage device accounts for a small fraction of the power demand of the whole system, significant increases in power demand are observed when accessing the storage device. The authors also investigate the access pattern impact on power demand, looking at the whole system and at the storage device by itself, finally comparing all of the tested configurations regarding energy efficiency. This study provides guidelines for the replacement of traditional storage servers by low-power alternatives. As a consequence, the choice will depend on the expected workload, estimates of power demand of the system, and factors limiting the performance. Related to cloud computing models, in the contribution4 entitled “Statistical analysis of Amazon EC2 cloud pricing models,” authors Gustavo Portella, Genaina N. Rodrigues, Eduardo Nakano, and Alba C.M.A. Melo show a statistical analysis for two Amazon cloud pricing models: on demand and spot. While the on-demand cloud instances are charged a fixed price and can only be terminated by the user, with very high availability, the spot instances are charged dynamically, which price is determined by a market-driven model and can be revoked by the provider when the spot price becomes higher than the user-defined price, having possibly low availability. The analysis for on-demand instances resulted in multiple linear regression equations that represent the influence of the characteristics of the processor and RAM memory in the composition of the price of different types of instances available on the Amazon EC2 provider. To analyze the Amazon spot pricing, the authors used time-smoothed moving averages by 12-hour periods, aiming to provide a price-availability trade-off to the user. With an extensive experimental evaluation, the authors conclude that the user's bid can be set at 30% of the on-demand price, with an availability above 90%, depending on the instance type. Also dealing with GPU computing, in the contribution5 entitled “DTM@GPU: characterizing and evaluating trace redundancy in GPU,” authors Leandro A. J. Marzulo, Alexandre C. Sena, Alexandre S. Nery, Cristiana Bentes, Igor M. Coelho, Maria Clicia S. de Castro, Saulo T. Oliveira, Tiago A. O. Alves, and Felipe M. G. França explore the concept of instruction reuse in GPUs. They propose the DTMGPU model that adapts the DTM (Dynamic Trace Memoization) technique to the NVIDIA GPU architecture. The proposed model considers the particularities of the GPU architecture in the reuse mechanism. Instruction reuse can be of two types: intra-thread and inter-threads. The authors perform a detailed investigation on the characteristics of the reused traces. This characterization shows the number and size of the reused traces, the influence of the cache size on reuse rates, and the cycles that are saved when all threads in a warp reuse instructions or traces. The results show approximately up to 35.3% of reuse, yielding an estimated speedup gain of 10.7%. Dealing with parallel and dataflow models, the paper6 entitled “DF-DTM: Dynamic Task Memoization and reuse in dataflow” and authored by Leandro Rouberte, Alexandre C. Sena, Alexandre S. Nery, Leandro A. J. Marzulo, Tiago A. O. Alves, and Felipe M. G. França also discusses instruction reuse and the DTM (Dynamic Trace Memoization) technique, but in a different environment, a Dataflow System. The authors propose the Dataflow Dynamic Task Memoization (DF-DTM) technique that allows the reuse of both nodes and subgraphs in dataflow, which are analogous to instructions and traces, respectively. The potential of DF-DTM is evaluated by a series of experiments that analyze the behavior of redundant tasks in five relevant benchmarks, where up to 99.70% of the instantiated tasks could be reused. This work also evaluates how reuse rates can be affected by limiting the subgraph size, memoization table size, task granularity, and problem size, showing that DF-DTM can yield good reuse rates in more realistic environments. Finally, considering GPU issues, the paper in this Special Issue,7 “Evaluating optimizations that reduce global memory accesses of stencil computations in GPGPUs,” authored by Thiago Nasciutti, Jairo Panetta, and Pedro Lopes, presents a performance study on different optimizations for 3D stencil computations on GPUs. The focus is on memory-oriented optimizations with varying grid and stencil sizes. The optimizations reduce global memory contention caused by the set of multiprocessors. The optimizations evaluated are grid tiling, inserting spatial and temporal loops into kernels, register reuse, and some of their combinations. A standardized experiment evaluates performance variation with grid size and stencil size for each optimization. Experimental data show that codes that use these optimizations are up to 3.3 times faster than the classical stencil formulation and that the optimization degree depends on the grid and stencil sizes. We hope that the readers can benefit from the perspectives presented in this Special Issue. The topics covered in the papers are timely and important, and the authors have done an excellent job in presenting their innovative contributions. Regarding the reviewing process, our referees (integrated by recognized researchers from the international community) made a great effort in evaluating the papers. We would like to acknowledge their effort in providing us with excellent feedback at the right time. Therefore, we wish to thank all the authors and reviewers. Finally, we would also like to express our gratitude to the Editor in Chief of this Journal, for his advice, vision, and support. Cristina Boeres, Cristiana Bentes, Edward D. Moreno |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | Exploring parallel multi-GPU local search strategies in a metaheuristic framework
Eyder Rios, Luiz Satoru Ochi, Cristina Boeres, Vitor Nazário Coelho, Igor Machado Coelho, Ricardo C. Farias |
J. Parallel Distributed Comput. | 3 |
| 2016 | MEC: The Memory Elasticity ControllerabstractThe dynamic nature of application and service requirements has lead the cloud community to invest in the study and development of elasticity features that have the ability to re-dimension resource capacities dynamically. Both online applications under unpredictable workloads or scientific workflows with different datasets require autonomic scaling in order to avoid either performance degradation due to an insufficient resource capacity or paying for additional, sub-utilized and possibly unnecessary capacity. In an effort to improve utilization, focus has recently turned to vertical elasticity where the processing, memory or storage capacity of a single virtual machine (VM) is adjusted in accordance with the application's needs. Given the increasing influence of memory availability on performance, this work presents the main features of the Memory Elasticity Controller (MEC). This VM allocation tool aims to improve the throughput of jobs or workflow tasks by continuously and judiciously calibrating the amount of host memory allocated to each VM in accordance with that VM's respective job's changing run time requirements while, at the same time, trying to avoid compromising the job's performance. This paper presents the tool's architecture and describes the functionality adopted to manage and provide vertical memory elasticity to concurrently executing VMs on a server. Results show that MEC is able to provide both resource providers and applications with an additional opportunity to improve throughput and performance. Roberto Sawamura, Cristina Boeres, Vinod E. F. Rebello |
HiPC | 2 |
| 2015 | Evaluating the Impact of Memory Allocation and Swap for Vertical Memory Elasticity in VMsabstractTypically, virtual machine (VM) allocation is based on the host server's ability to meet the VM's maximum CPU, I/O and memory requirements. However, given that the requirements of applications within the VM may vary during execution, it might be more efficient to also vary over time the amount of resources dedicated to the VM. In cloud systems, vertical elasticity is the dynamic adjustment of the amount of a physical resource, such as memory, CPU cores, etc., that is allocated to a VM. With technology pushing up core counts and speeds of modern servers, and given the growing trend towards server consolidation, making the most of the available memory is crucial for good application performance. This paper investigates the impact of memory allocation and swap usage on VM performance. Through an experimental evaluation, hyper visor independent metrics and policies are identified for consideration by tools that claim to offer vertical memory elasticity. Based on the conclusions, the paper goes on to present a framework of a tool to dynamically manage memory allocations of VMs. Preliminary results with the proposed vertical Memory Elasticity Controller, MEC, highlight some of the benefits to both resource providers and applications through improved efficiency, throughput and performance. Ongoing work will continue to expand on the current evaluation and refine the scheduling policies to further improve the tool. Roberto Sawamura, Cristina Boeres, Vinod E. F. Rebello |
SBAC-PAD | 2 |
| 2015 | Memory aware load balance strategy on a parallel branch-and-bound applicationabstractAbstract The latest trends in high performance computing systems show an increasing demand on the use of a large scale multicore system in an efficient way so that high compute‐intensive applications can be executed reasonably well. However, the exploitation of the degree of parallelism available at each multicore component can be limited by the poor utilization of the memory hierarchy. Actually, the multicore architecture introduces some distinct features that are already observed in shared memory and distributed environments. One example is that subsets of cores can share different subsets of memory. In order to achieve high performance, it is imperative that a careful allocation scheme of an application is carried out on the available cores, based on a scheduling specification that considers not only processors characteristics but also memory contention. This paper proposes a multicore cluster representation that captures relevant performance characteristics in multicores systems such as the influence of memory hierarchy and contention on application performance. Improved performance was achieved by a branch‐and‐bound application applied to the partitioning sets problem that incorporated a memory aware load balancing strategy based on the proposed multicore cluster representation. An in‐depth analysis on this application execution showed its applicability to modern systems. Copyright © 2014 John Wiley & Sons, Ltd. Juliana M. N. Silva, Cristina Boeres, Lúcia M. A. Drummond, Artur Alves Pessoa |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | Autonomic Malleability in Iterative MPI ApplicationsabstractDuring their execution, a significant number of applications often sub utilize the capacity of the resources to which they are allocated or require more. Furthermore, with the current scale up trend in server design, effective utilization can only be achieved by applications sharing such resources. Cluster management systems already support static resource partitioning at job submission time and given that application utilization more than often varies during the execution, it will become increasingly more important to permit applications to harness all available spare capacity. This paper investigates the feasibility of malleable evolving versions of applications to improve performance and system efficiency. Extending a previous classification, we show that improvements can be achieved for a real astrophysics application. Alexandre da Costa Sena, Felipe Ribeiro, Vinod E. F. Rebello, Aline de Paula Nascimento, Cristina Boeres |
SBAC-PAD | 5 |
| 2011 | An Approach to Optimise the Execution of RTM Algorithm in Multicore MachinesabstractThe new oil fields discovered in the Gulf of Mexico and in Brazil's southeast coast are located in deep water, which imposes new challenges for sub salt seismic imaging. To produce sufficient accurate imaging of such fields, the compute intensive RTM method is the currently favoured approach, despite its high computational cost. This work evaluates the RTM code in multicore machines and proposes a new version that is more than 2.4 faster than the original. Besides, how the memory hierarchy is utilised has a crucial impact on the performance of the code. This article also presents an innovative approach to calculate efficient block values for a better memory utilisation on multicore archictetures. Alexandre da Costa Sena, Aline de Paula Nascimento, Cristina Boeres, Vinod E. F. Rebello, André Bulcão |
eScience | 3 |
| 2011 | Fault Tolerance in an Industrial Seismic Processing Application for Multicore Clusters
Alexandre Domingues Gonçalves, Matheus Bersot, André Bulcão, Cristina Boeres, Lúcia M. A. Drummond, Vinod E. F. Rebello |
EuroMPI | 4 |
| 2011 | An efficient weighted bi-objective scheduling algorithm for heterogeneous systems
Cristina Boeres, Idalmis Milián Sardiña, Lúcia M. A. Drummond |
Parallel Comput. | 1 |
| 2009 | On the Feasibility of Dynamically Scheduling DAG Applications on Shared Heterogeneous Systems
Aline de Paula Nascimento, Alexandre da Costa Sena, Cristina Boeres, Vinod E. F. Rebello |
Euro-Par | 3 |
| 2008 | EasyGrid Enabling of Iterative Tightly-Coupled Parallel MPI ApplicationsabstractThis paper addresses the challenge of how to permit tightly coupled parallel applications, optimised for uniform, stable, static environments, execute equally efficiently in environments which exhibit the complete opposite characteristics. Using the N-body problem as a case study, both the traditional and proposed grid enabled MPI implementations of the popular ring algorithm are analysed. Results with respect to performance show the latter approach to be competitive on a cluster and significantly more effective in heterogeneous and dynamic environments. Alexandre da Costa Sena, Aline de Paula Nascimento, Cristina Boeres, Vinod E. F. Rebello |
ISPA | 3 |
| 2007 | On the Advantages of an Alternative MPI Execution Model for GridsabstractThe MPI message passing library is used extensively in the scientific community as a tool for parallel programming. Even though improvements have been made to existing implementations to support execution on computational grids, MPI was initially designed to deal with homogeneous, fault- free, static environments such as computing clusters. The typical programming approach is to execute a single MPI process on each resource. However, this may not be appropriate for heterogeneous, non-dedicated and dynamic environments such as grids. This paper aims to show that programmers can implement parallel MPI solutions to their problems in an architectural independent style and obtain good performance on a grid by transferring responsibility to an application management system (AMS). A comparison of program implementations under a traditional MPI execution model and a fine-grain model highlight the advantages of using the latter. Alexandre da Costa Sena, Aline de Paula Nascimento, Jacques Alves da Silva, Daniela Vianna 0001, Cristina Boeres, Vinod E. F. Rebello |
CCGRID | 5 |
| 2007 | MOS-Based Rate Adaption for VoIP SourcesabstractThis paper proposes an algorithm for the adaptive adjustment of the transmission rate of VoIP sources based on the voice quality estimated at the receiver. This adjustment is achieved through the appropriate use of differing voice codecs, as the conditions of the network change, in order to maintain an efficient utilization of the available resources. To validate our proposal realistically, we have made an effort to simulate VoIP calls using sources that follow Brady's model of human conversations. We investigate the effects of the proposed model on the aggregate network traffic and compare the results with existing related work. Simulation results show that the proposed algorithm makes better use of the available bandwidth, achieving superior performance in comparison to similar works. Nilmax Teones Moura, Bruno A. Vianna, Célio Vinicius N. de Albuquerque, Vinod E. F. Rebello, Cristina Boeres |
ICC | 5 |
| 2007 | Distributed and dynamic self-scheduling of parallel MPI Grid applicationsabstractAbstract The execution of distributed applications on the Grid is already a reality. However, as both the number of applications grow and Grids increase in scale, the efficient utilization of the available but shared heterogeneous resources will become increasingly essential to the Grid's successful maturity. Furthermore, it is unclear whether existing Grid management systems are capable of meeting this challenge. The EasyGrid middleware is a hierarchically distributed application management system (AMS) that is embedded into MPI applications to autonomously orchestrate their execution efficiently in computational Grids. The overhead of employing a distinct AMS to make each application system aware brings at least two benefits. First, the adopted policies can be tailored to the specific needs of each application, leading to improved performance. Second, distributing the management effort among executing applications makes Grid management more scalable. This article focuses on scheduling policies of an AMS for a particular class of application, describing a low intrusion implementation of a hybrid scheduling strategy designed to elicit good performance even in dynamic environments such as Grids. Using application‐specific scheduling policies, near‐optimal runtimes highlight the advantages of self‐scheduling when executing one or more system aware applications on a Grid. Copyright © 2006 John Wiley & Sons, Ltd. Aline de Paula Nascimento, Alexandre da Costa Sena, Cristina Boeres, Vinod E. F. Rebello |
Concurr. Comput. Pract. Exp. | 3 |
| 2006 | An EasyGrid portal for scheduling system-aware applications on computational GridsabstractAbstract One of the objectives of computational Grids is to offer applications the collective computational power of distributed but typically shared heterogeneous resources. Unfortunately, efficiently harnessing the performance potential of such systems (i.e. how and where applications should execute on the Grid) is a challenging endeavor due principally to the very distributed, shared and heterogeneous nature of the resources involved. A crucial step towards solving this problem is the need to identify both an appropriate scheduling model and scheduling algorithm(s). This paper presents a tool to aid the design and evaluation of scheduling policies suitable for efficient execution of system‐aware parallel applications on computational Grids. Copyright © 2005 John Wiley & Sons, Ltd. Cristina Boeres, Ariel Alves Fonseca, Helder de Amorim Mendes, Luiz Toscano Menezes, Nilmax Teones Moura, Jacques Alves da Silva, Bruno de Azevedo Vianna, Vinod E. F. Rebello |
Concurr. Comput. Pract. Exp. | 1 |
| 2005 | Hybrid evolutionary static scheduling for heterogeneous systemsabstractThe complexity of the static scheduling problem on heterogeneous resources has motivated the development of low complexity heuristics such as list scheduling. However, the greedy characteristic of such heuristics can, in many cases, generate poor results. This work proposes the integration of list scheduling heuristics with search mechanisms based on both genetic algorithms and GRASP, to efficiently schedule tasks on distributed systems. The results show that the hybrid approach is robust and can converge quickly to good quality solutions. Cristina Boeres, Eyder Rios, Luiz Satoru Ochi |
Congress on Evolutionary Computation | 1 |
| 2005 | Managing the Execution of Large Scale MPI Applications on Computational GridsabstractComputational grids aim to aggregate significant numbers of resources to provide sufficient, but low cost, computational power to an ever growing variety applications. Writing applications capable of executing efficiently in these grid environments is however extremely difficult for inexperienced users. The grid's geographically distributed resources are typically heterogeneous, non-dedicated, and are offered without any performance or availability guarantees. This work investigates an alternative approach (based on smarter system-aware applications) to solve the problem of developing and managing the execution of grid applications efficiently. Results show that these system-aware MPI applications are indeed faster than their conventional implementations and easily grid enabled. Aline de Paula Nascimento, Alexandre da Costa Sena, Jacques Alves da Silva, Daniela Vianna 0001, Cristina Boeres, Vinod E. F. Rebello |
SBAC-PAD | 5 |
| 2004 | A Cluster-based Strategy for Scheduling Task on Heterogeneous ProcessorsabstractEfficient task scheduling is fundamental for parallel applications to achieve good performance on distributed systems. While extensive work exists for scheduling tasks on homogeneous processors, fewer algorithms exist for the more common problem of scheduling in heterogeneous processor environments. In this paper, we propose coupling a replication-based clustering heuristic for homogeneous processors, with a mechanism to map the generated clusters to the heterogeneous environment. Experimental results show that this strategy compares favourably in terms of the makespan with traditional list scheduling approaches to this problem, particularly when communication costs are high. Cristina Boeres, José Viterbo, Vinod E. F. Rebello |
SBAC-PAD | 1 |
| 2004 | EasyGrid: towards a framework for the automatic Grid enabling of legacy MPI applicationsabstractAbstract One of the goals of the Grid is to aggregate collections of shared, heterogeneous, and distributed resources to provide computational ‘power’ to parallel applications. However, designing applications capable of exploiting this potential with ease remains a challenge. This paper outlines the EasyGrid methodology for the efficient and robust execution of (legacy) MPI programs across distributed computing clusters. The principal objective of this work is to identify the application‐oriented middleware necessary for, as well as to develop a framework to automatically generate, system ‐ aware applications capable of executing in dynamic, unstable, distributed environments such as computational Grids. Copyright © 2004 John Wiley & Sons, Ltd. Cristina Boeres, Vinod E. F. Rebello |
Concurr. Pract. Exp. | 1 |
| 2003 | Hybrid Task Scheduling: Integrating Static and Dynamic HeuristicsabstractResearchers are constantly looking for ways to improve the execution time of parallel applications on distributed systems. Although compile-time static scheduling heuristics employ complex mechanisms, the quality of their schedules are handicapped by estimated run-time costs. On the other hand, while dynamic schedulers use actual run-time costs, they have to be of low complexity in order to reduce the scheduling overhead. We investigate the viability of integrating these two approaches into a hybrid scheduling framework. The relationship between static schedulers, dynamic heuristics and scheduling events are examined. The results show that a hybrid scheduler can indeed improve the schedules produced by good traditional static list scheduling algorithms. Cristina Boeres, Alexandre Lima, Vinod E. F. Rebello |
SBAC-PAD | 1 |
| 2002 | Cluster-Based Static Scheduling: Theory and PracticeabstractTask scheduling is a key element in achieving high performance from multicomputer systems. To be efficient, scheduling algorithms must be based on a cost model appropriate for computing systems in use. The optimal scheduling of tasks is NP-hard, and a large number of heuristic algorithms have been proposed for a variety of scheduling conditions (graph types, granularities or cost models). This paper studies the problem of task scheduling under the LogP model and presents both theoretical and experimental results for a cluster-based, task duplication methodology. Cristina Boeres, Vinod E. F. Rebello |
SBAC-PAD | 1 |
| 2001 | On Minimising the Processor Requirements of LogP Schedules
Cristina Boeres, Gerson N. da Cunha, Vinod E. F. Rebello |
Euro-Par | 1 |
| 2001 | On the Design of Clustering-based Scheduling Algorithms for Realistic Machine ModelsabstractWhile the NP-complete problem of scheduling weighted arbitrary directed acyclic graphs under the delay model has been studied extensively, comparatively little work exists for this problem under more realistic models such as the LogP model. Recently, a number of LogP-based scheduling heuristics and related works have appeared in the literature, including a task clustering algorithm design methodology which identifies four crucial design issues (Boeres et al., 1997). Through the use of five task replication-based scheduling heuristics based on this design methodology, this paper investigates the effect of various implementations of these design issues on the schedules produced for the allocation of arbitrary task graphs to fully connected networks of processors under a LogP-type model. The quality of the schedules produced by these algorithms are also compared with good, well-known delay model-based algorithms and an existing LogP strategy. Cristina Boeres, Vinod E. F. Rebello |
IPDPS | 1 |
| 1999 | Scheduling Arbitrary Task Graphs on LogP Machines
Cristina Boeres, Aline de Paula Nascimento, Vinod E. F. Rebello |
Euro-Par | 1 |
| 1999 | A versatile cost modelling approach for multicomputer task scheduling
Cristina Boeres, Vinod E. F. Rebello |
Parallel Comput. | 1 |
| 1998 | Static Scheduling Using Task Replication for LogP and BSP Models
Cristina Boeres, Vinod E. F. Rebello, David B. Skillicorn |
Euro-Par | 1 |
| 1997 | Versatile Task Scheduling of Binary Trees for Realistic Machines
Cristina Boeres, Vinod E. F. Rebello |
Euro-Par | 1 |
| 1990 | An Occam-based evaluation of a parallel version of simulated annealing
Valmir C. Barbosa, Cristina Boeres |
Microprocessing and Microprogramming | 2 |