EDBT 2026 Demo / reviewers in the wild / expert
Raphael Y. de Camargo
dblp:31/4450 · also Raphael Yokoingawa de Camargo
· DBLP profile ↗
24ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0001-6021-747XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Efficient Computation of Attractor Fields in Coupled Boolean Networks
Luiz C. S. Rozante, Carlos Reynaldo Portocarrero Tovar, David Correa Martins Jr., Raphael Y. de Camargo, Luciana Arantes, Pierre Sens 0001 |
ICCSA (1) | 4 |
| 2023 | Applying Independent Vector Analysis on EEG-Based Motor Imagery ClassificationabstractJoint Blind Source Separation (JBSS) is an essential and versatile research topic that has attracted the attention of researchers in the last decade. Independent Vector Analysis (IVA) is an exciting approach in the context of the JBSS method since it is an extension of Independent Component Analysis (ICA) towards the exploitation of the statistical dependency between different datasets through the use of Mutual Information. In this work, we propose an original approach of IVA as a feature extraction step for Brain-Computer Interfaces, focused on the Motor Imagery (MI) paradigm. For this, we use the BCI Competition IV - Dataset 1. Since the participants of the experiment are performing the same MI tasks, we assume that the channels related to MI present correlated signals across subjects that might be explored by IVA techniques. The results show that the algorithm could classify the MI movements using a consolidated and low-cost classifier, Support Vector Machine, achieving an accuracy of 85%. Caroline P. A. Moraes, Bruno Aristimunha, Lucas Heck Dos Santos, Walter H. L. Pinaya, Raphael Y. de Camargo, Denis G. Fantinato, Aline Neves 0001 |
ICASSP | 5 |
| 2023 | Evaluating execution time predictions on GPU kernels using an analytical model and machine learning techniquesabstractPredicting the performance of applications executed on GPUs is a great challenge and is essential for efficient job schedulers. There are different approaches to do this, namely analytical modeling and machine learning (ML) techniques. Machine learning requires large training sets and reliable features, nevertheless it can capture the interactions between architecture and software without manual intervention. In this paper, we compared a BSP-based analytical model to predict the time of execution of kernels executed over GPUs. The comparison was made using three different ML techniques. The analytical model is based on the number of computations and memory accesses of the GPU, with additional information on cache usage obtained from profiling. The ML techniques Linear Regression, Support Vector Machine, and Random Forest were evaluated over two scenarios: first, data input or features for ML techniques were the same as the analytical model and, second, using a process of feature extraction, which used correlation analysis and hierarchical clustering. Our experiments were conducted with 20 CUDA kernels, 11 of which belonged to 6 real-world applications of the Rodinia benchmark suite, and the other were classical matrix-vector applications commonly used for benchmarking. We collected data over 9 NVIDIA GPUs in different machines. We show that the analytical model performs better at predicting when applications scale regularly. For the analytical model a single parameter λ is capable of adjusting the predictions, minimizing the complex analysis in the applications. We show also that ML techniques obtained high accuracy when a process of feature extraction is implemented. Sets of 5 and 10 features were tested in two different ways, for unknown GPUs and for unknown Kernels. For ML experiments with a process of feature extractions, we got errors around 1.54% and 2.71%, for unknown GPUs and for unknown Kernels, respectively. Marcos Amaris, Raphael Y. de Camargo, Daniel Cordeiro, Alfredo Goldman, Denis Trystram |
J. Parallel Distributed Comput. | 2 |
| 2022 | Predicting Dengue Outbreaks with Explainable Machine LearningabstractSeasonal infectious diseases, such as dengue, have been causing great losses in many countries around the world in terms of deaths, quality of life, and economic burden. In Brazil, this is relevant not only in large cities such as Rio de Janeiro and São Paulo but, according to the Ministry of Health, in another 500 cities throughout the country. Predicting the occurrence of diseases, such as dengue bursts, can be a valuable instrument for public health management as health officials can better prepare and redirect resources to the affected areas. In this paper, we present an explainable machine learning model to forecast the number of dengue occurrences in a large metropolis, Rio de Janeiro. We focus on explainable models, which provide health authorities with the reasons for outbreak predictions, allowing them to plan their actions accordingly. We trained a gradient boosting decision tree algorithm (CatBoost) with data from the National System of Information on Notifiable Diseases (SINAN), weather data, and socio-demographic data from The Brazilian Institute of Geoaraphy and Statistics (IBGE). Robson Aleixo, Fabio Kon, Rudi Rocha, Marcela Santos Camargo, Raphael Y. de Camargo |
CCGRID | 5 |
| 2022 | Improving the performance of batch schedulers using online job runtime classification
Salah Zrigui, Raphael Y. de Camargo, Arnaud Legrand, Denis Trystram |
J. Parallel Distributed Comput. | 2 |
| 2021 | Computer architecture and high performance computingabstractIn this special issue of Concurrency and Computation Practice and Experience, we are pleased to present eight selected papers that were previously presented at the Brazilian "XX Simpósio em Sistemas Computacionais de Alto Desempenho," WSCAD 2019. The event was held in conjunction with the 31st International Symposium on Computer Architecture and High Performance Computing, SBAC-PAD 2019, in Campo Grande, MS, Brazil, from October 15 to 18, 2019. The WSCAD workshop has been presenting important research in the fields of computer architectures, high performance computing, and distributed systems, since the beginning of the 2000s. The scope of the current special issue is broad and representative, with different forms of contributions to our discipline: methodological papers, technology papers, application papers, and system papers. The topics covered in the papers include architecture issues, compiler optimization, performance evaluation, parallel algorithms, energy efficiency, and applications. The title of the first paper is "Structural testing for communication events into loops of message-passing parallel programs," by Diaz et al.1 In this paper, the authors propose new structural testing criteria for message-passing parallel programs, focusing on defects from communication primitives into loops. A new test model is presented to support their criteria for structural testing of MPI-applications. The testing criteria are validated through experimental studies using a tool called ValiMPI. The results show that unknown defects from communication and synchronization events can be revealed in different loop iterations, increasing the quality of message-passing parallel programs. In the second contribution, entitled "Smart selection of optimizations in dynamic compilers," Rosario et al.2 present an approach that uses machine learning to select sequences of optimization for dynamic compilation that considers both code quality and compilation overhead. Their approach starts by training a model, offline, with a knowledge bank of those sequences with low overhead and high-quality code generation capability using a genetic heuristic. Then, this bank is used to guide the smart selection of optimizations sequences for the compilation of code fragments during the emulation of an application. The proposed strategy is evaluated in two LLVM-based dynamic binary translators, namely, OI-DBT and HQEMU, showing that these two translators can achieve average speedups of 1.26× and 1.15× in MiBench and Spec Cpu benchmarks, respectively. In the third contribution, entitled "Memory allocation anomalies in high-performance computing applications: A study with numerical simulations," Gomes et al.3 propose a method for identifying, locating, characterizing, and fixing allocation anomalies, and a tool for developers to apply the method. A numerical simulator that approximates the solutions to partial differential equations using a finite element method is used in the experiments. It is shown that taming allocation anomalies in the simulator reduces both its execution time and the memory footprint of its processes, irrespective of the specific heap allocator being employed with it. They conclude that the developer of HPC applications can benefit from the method and tool during the software development cycle. The fourth contribution, entitled "Investigating memory prefetcher performance over parallel applications: From real to simulated," by Girelli et al.,4 contributes to shed light on the memory prefetcher's role in the performance of parallel high-performance computing applications, considering the prefetcher algorithms offered by both the real hardware and the simulators. The authors performed a careful experimental investigation, executing the NAS parallel benchmark (NPB) on a real Skylake machine, and as well in a simulated environment with the ZSim and Sniper simulators, taking into account the prefetcher algorithms offered by both Skylake and the simulators. The experimental results show that: (i) prefetching from the L3 to L2 cache presents better performance gains, (ii) the memory contention in the parallel execution constrains the prefetcher's effect, (iii) Skylake's parallel memory contention is poorly simulated by ZSim and Sniper, and (iv) Skylake's noninclusive L3 cache hinders the accurate simulation of NPB with the Sniper's prefetchers. In the fifth contribution, entitled "Energy efficiency and portability of oil and gas simulations on multicore and graphics processing unit architectures," Serpa et al.5 propose three optimizations for an oil and gas application, reverse time migration (RTM), which reduce the floating-point operations by changing the equation derivatives. They evaluate these optimizations in different multicore and GPU architectures, investigating the impact of different APIs on the performance, energy efficiency, and portability of the code. The experimental results show that the dedicated CUDA implementation running on the NVIDIA Volta architecture has the best performance and energy efficiency for RTM on GPUs, while the OpenMP version is the best for Intel Broadwell in the multicore. Also, the OpenACC version, which has a lower programming effort and executes on both architectures, has up to 20% better performance and energy efficiency than the nonportable ones. In the sixth paper, entitled "An open computing language-based parallel Brute Force algorithm for formal concept analysis on heterogeneous architectures," Novais et al.6 propose and evaluate an Open Computing Language (OpenCL)-based Brute Force algorithm for formal concept extraction on heterogeneous architectures (CPU + GPU and CPU + FPGA). The CPU + GPU architecture presents higher performance and scalability than other architectures when the Brute Force algorithm processes high dimensional contexts with many objects and attributes. Their parallel approach shows performance results up to 18× better than a smarter sequential algorithm called Data-Peeler. Moreover, the Brute Force algorithm running on CPU + GPU architecture has greater energy efficiency, reaching at least 1.79× more operations per energy consumption than other algorithms on different architectures explored in the work. In the seventh paper, entitled "Contextual contracts for component-oriented resource abstraction in a cloud of high performance computing services," Junior et al.7 present HPC Shelf, a cloud computing services platform to build and deploy large-scale parallel computing systems. They introduce Alite, the contextual contract system of HPC Shelf, to select component implementations according to requirements of the host application, target parallel computing platform characteristics (e.g., clusters and MPPs), quality of service (QoS) properties, and cost restrictions. It is evaluated through a small-scale case study employing two complementary component-based frameworks. The first one aims to represent components that implement linear algebra computations based on the BLAS interface. In turn, the second one aims to represent parallel computing platforms on the IaaS cloud offered by Amazon EC2 Service. The last paper in this special issue, "High-performance IO for seismic processing on the cloud" authored by Guimarães et al.,8 analyzes the main file structures currently used to store seismic data and propose a new intermediate data structure to improve IO performance while still complying with established standards. They show that, throughout a common workflow in seismic data analysis, the IO performance gain greatly surpasses the overhead of translating data to the intermediate structure. The approach enables a speedup of up to 208 times in reading time when using classical standards (e.g., SEG-Y) and the intermediate structure is up to 1.8 times more efficient than modern formats (e.g., ASDF). Considering cache-friendly applications, the speedups over the direct use of SEG-Y reach 8000 times. They also performed a cost analysis on the AWS cloud showing that HDDs can be 1.25 times more cost-effective than SSDs. The research papers presented in this special issue provide insights in fields related to high performance computing, including performance evaluation, parallel algorithms, and applications in science and engineering. We believe that the main contributions presented in the research papers are timely and important, and hope that readers can benefit from the papers and contribute to these rapidly growing areas. Many individuals contributed a great deal of time and energy toward the success of this special issue. We would like to thank all the authors who provided valuable contributions to this special issue. We are also grateful to the reviewers for their many hours of dedicated efforts, with valuable feedback to the authors. Finally, we would also like to express our gratitude to the Editor-in-Chief of CCPE, for his advice, vision, and support, making this special issue possible. Raphael Y. de Camargo, Fabrizio Marozzo, Wellington Santos Martins |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | One Can Only Gain by Replacing EASY Backfilling: A Simple Scheduling Policies Case StudyabstractHigh-Performance Computing (HPC) platforms are growing in size and complexity. In order to improve the quality of service of such platforms, researchers are devoting a great amount of effort to devise algorithms and techniques to improve different aspects of performance such as energy consumption, total usage of the platform, and fairness between users. In spite of this, system administrators are always reluctant to deploy state of the art scheduling methods and most of them revert to EASY-backfilling, also known as EASY-FCFS (EASY-First-Come-First-Served). Newer methods frequently are complex and obscure and the simplicity and transparency of EASY are too important to sacrifice. In this work, we used execution logs from five HPC platforms to compare four simple scheduling policies: FCFS, Shortest estimated Processing time First (SPF), Smallest Requested Resources First (SQF), and Smallest estimated Area First (SAF). Using simulations, we performed a thorough analysis of the cumulative results for up to 180 weeks and considered three scheduling objectives: waiting time, slowdown and per-processor slowdown. We also evaluated other effects, such as the relationship between job size and slowdown, the distribution of slowdown values, and the number of backfilled jobs, for each HPC platform and scheduling policy. We conclude that one can only gain by replacing EASY-backfilling with SAF with backfilling, as it offers improvements in performance by up to 80% in the slowdown metric while maintaining the simplicity and the transparency of FCFS. Moreover, SAF reduces the number of jobs with large slowdowns and the inclusion of a simple thresholding mechanism guarantees that no starvation occurs. Finally, we propose SAF as a new benchmark for future scheduling studies. Danilo Carastan-Santos, Raphael Y. de Camargo, Denis Trystram, Salah Zrigui |
CCGRID | 2 |
| 2019 | Real-Time Scheduling Policy Selection from Queue and Machine StatesabstractTask Scheduling in large-scale HPC platforms is normally accomplished with simple heuristics combined with a backfilling algorithm. Some strategies, such as the First-Come-First-Serve (FCFS) with backfilling, provide reasonable results in a variety of scenarios, including different HPC platforms and task set characteristics. But for each scenario, a different strategy might be the most appropriate for minimizing some metric, such as the average task waiting time or turnaround time. In this work, we present a real-time scheduling policy selection algorithm, which takes as input the running queue job characteristics and machine states. We evaluated the use of logistic regression and support-vector machines to perform the mapping from queue and machine state to selected scheduling policy. The machine learning algorithms are trained and evaluated using simulations configured using HPC platform traces. When selecting among 8 (eight) scheduling policies, we obtained an accuracy above 80%, when compared to the best selection. When simulating the online real-time selection of policies for a period of one year, we obtained a reduction in the mean queue waiting time of tasks of up to 40% over using FCFS and 10% over randomly selecting policies. Moreover, the method performed close the best possible selection of policies, with a maximum of 9% increase in the mean queue waiting time. Luis Sant'Ana, Danilo Carastan-Santos, Daniel Cordeiro, Raphael Y. de Camargo |
CCGRID | 4 |
| 2019 | PLB-HAC: Dynamic Load-Balancing for Heterogeneous Accelerator Clusters
Luis Sant'Ana, Daniel Cordeiro, Raphael Y. de Camargo |
Euro-Par | 3 |
| 2019 | A hybrid CPU-GPU-MIC algorithm for minimal hitting set enumerationabstractSummary We present a hybrid exact algorithm for the Minimal Hitting Set (MHS) Enumeration Problem for highly heterogeneous CPU‐GPU‐MIC platforms. With several techniques that permit an efficient exploitation of each architecture, low communication cost, and effective load balancing, we were able to enumerate MHSs for large instances in reasonable time, achieving good performance and scalability. We obtained speedups of up to 25.32 in comparison with using two six‐core CPUs and we also enumerated MHSs for instances with tens of thousands of variables in less than 5 hours. We also evaluated our algorithm with a real‐world driven dataset, and with a large CPU‐GPU cluster, we unprecedentedly enumerated in parallel large minimal hitting sets of this dataset in less than 8 hours. These results reinforce the statement that heterogeneous clusters of CPUs, GPUs, and MICs can be used efficiently for high‐performance computing. Danilo Carastan-Santos, David Correa Martins Jr., Siang Wun Song, Luiz C. S. Rozante, Raphael Y. de Camargo |
Concurr. Comput. Pract. Exp. | 5 |
| 2017 | Obtaining dynamic scheduling policies with simulation and machine learningabstractDynamic scheduling of tasks in large-scale HPC platforms is normally accomplished using ad-hoc heuristics, based on task characteristics, combined with some backfilling strategy. Defining heuristics that work efficiently in different scenarios is a difficult task, specially when considering the large variety of task types and platform architectures. In this work, we present a methodology based on simulation and machine learning to obtain dynamic scheduling policies. Using simulations and a workload generation model, we can determine the characteristics of tasks that lead to a reduction in the mean slowdown of tasks in an execution queue. Modeling these characteristics using a nonlinear function and applying this function to select the next task to execute in a queue improved the mean task slowdown in synthetic workloads. When applied to real workload traces from highly different machines, these functions still resulted in performance improvements, attesting the generalization capability of the obtained heuristics. Danilo Carastan-Santos, Raphael Y. de Camargo |
SC | 2 |
| 2017 | Finding exact hitting set solutions for systems biology applications using heterogeneous GPU clustersabstractThe Systems Biology field presents several complex combinatorial problems that can be in part reduced to an instance of the Hitting Set Problem (HSP), which is NP-Hard. These reduced problems often come with a large amount of data that needs to be processed, such as gene expression profiles, resulting in prohibitive computational costs for finding the exact solutions. There are some proposals to obtain exact solutions for HSP, including an approach which uses GPUs. However, such an approach is not scalable for real input sizes (thousands of variables). We propose a novel algorithm for solving HSP instances with thousands of variables by using: (i) clause sorting, which enables the efficient discarding of non-solution candidates, (ii) parallel generation and evaluation of candidate solutions through the use of GPUs, and (iii) support for multiple GPUs. To permit the execution on heterogeneous clusters, we determine the minimum kernel size that does not incur extra overhead and distribute tasks among available GPUs on demand. Our experimental results show that the combination of these techniques results in a speedup of 118.5, when using eight NVIDIA Tesla K20c in comparison with a ten-core Intel Xeon E5-2690 processor. Consequently, our algorithm can enable the usage of exact algorithms for solving the Hitting Set problem and applying it to real world problems. Danilo Carastan-Santos, Raphael Y. de Camargo, David Correa Martins Jr., Siang Wun Song, Luiz C. S. Rozante |
Future Gener. Comput. Syst. | 2 |
| 2016 | A comparison of GPU execution time prediction using machine learning and analytical modelingabstractToday, most high-performance computing (HPC) platforms have heterogeneous hardware resources (CPUs, GPUs, storage, etc.) A Graphics Processing Unit (GPU) is a parallel computing coprocessor specialized in accelerating vector operations. The prediction of application execution times over these devices is a great challenge and is essential for efficient job scheduling. There are different approaches to do this, such as analytical modeling and machine learning techniques. Analytic predictive models are useful, but require manual inclusion of interactions between architecture and software, and may not capture the complex interactions in GPU architectures. Machine learning techniques can learn to capture these interactions without manual intervention, but may require large training sets. In this paper, we compare three different machine learning approaches: linear regression, support vector machines and random forests with a BSP-based analytical model, to predict the execution time of GPU applications. As input to the machine learning algorithms, we use profiling information from 9 different applications executed over 9 different GPUs. We show that machine learning approaches provide reasonable predictions for different cases. Although the predictions were inferior to the analytical model, they required no detailed knowledge of application code, hardware characteristics or explicit modeling. Consequently, whenever a database with profile information is available or can be generated, machine learning techniques can be useful for deploying automated on-line performance prediction for scheduling applications on heterogeneous architectures containing GPUs. Marcos Amaris, Raphael Y. de Camargo, Mohamed Dyab, Alfredo Goldman, Denis Trystram |
NCA | 2 |
| 2015 | A Multi-GPU Hitting Set Algorithm for GRNs InferenceabstractGene regulatory networks inference is one of the crucial problems of the Systems Biology field. It is still an open problem, mainly because of its high dimensionality (thousands of genes) with a limited number of samples (dozens), making it difficult to estimate dependencies among genes. Besides the estimation problem, another important hindrance is the inherent computational complexity of GRN inference methods. In this work, we focus on circumventing performance issues of a technique based on signal perturbations to infer gene dependencies. One of its main steps consists in solving the Hitting Set problem (HSP), which is NP-Hard. There are many proposals to obtain approximate or exact solutions to this problem. One of these proposals consists of a Graphical Processing Unit (GPU) based algorithm to obtain exact solutions to the HSP. However, such method is not scalable for real size GRNs. We propose an extension of the HSP algorithm to deal with input sets containing thousands of variables by introducing innovations in the data structures and a sorting scheme to allow efficient discarding of Hitting Set non-solution candidates. We provide an implementation for multi-core CPUs and GPU clusters. Our experimental results show that the usage of the sorting scheme brings speedups of up to 3.5 in the CPU implementation. Moreover, using a single GPU, we could obtain an additional speedup of up to 4.7, in comparison with the multithreaded CPU implementation. Finally, usage of eight GPUs from a GPU cluster brought an additional speedup of up to 6.6. Combining all techniques, speedups above 60 were obtained for the parallel part of the algorithm. Danilo Carastan-Santos, Raphael Y. de Camargo, David Correa Martins Jr., Siang Wun Song, Luiz C. S. Rozante, Fabrizio F. Borelli |
CCGRID | 2 |
| 2015 | PLB-HeC: A Profile-Based Load-Balancing Algorithm for Heterogeneous CPU-GPU ClustersabstractThe use of GPU clusters for scientific applications in areas such as physics, chemistry and bioinformatics is becoming more widespread. These clusters frequently have different types of processing devices, such as CPUs and GPUs, which can themselves be heterogeneous. To use these devices in an efficient manner, it is crucial to find the right amount of work for each processor that balances the computational load among them. This problem is not only NP-hard on its essence, but also tricky due to the variety of architectures of those devices. We present PLB-HeC, a Profile-based Load-Balancing algorithm for Heterogeneous CPU-GPU Clusters that performs an online estimation of performance curve models for each GPU and CPU processor. Its main difference to existing algorithms is the generation of a non-linear system of equations representing the models and its solution using a interior point method, improving the accuracy of block distribution among processing units. We implemented the algorithm in the StarPU framework and compared its performance with existing load-balancing algorithms using applications from linear algebra, stock markets and bioinformatics. We show that it reduces the application execution times in almost all scenarios, when using heterogeneous clusters with two or more machine configurations. Luis Sant'Ana, Daniel Cordeiro, Raphael Y. de Camargo |
CLUSTER | 3 |
| 2015 | A Simple BSP-based Model to Predict Execution Time in GPU ApplicationsabstractModels are useful to represent abstractions of software and hardware processes. The Bulk Synchronous Parallel (BSP) is a bridging model for parallel computation that allows algorithmic analysis of programs on parallel computers using performance modeling. The main idea of BSP model is the treatment of communication and computation as abstractions of a parallel system. Meanwhile, the use of GPU devices are becoming more widespread and they are currently capable of performing efficient parallel computation for applications that can be decomposed on thousands of simple threads. However, few models for predicting application execution time on GPUs have been proposed. In this work we present a simple and intuitive BSP-based model for predicting the CUDA application execution times on GPUs. The model is based on the number of computations and memory accesses of the GPU, with additional information on cache usage obtained from profiling. Scalability, divergence, effect of optimizations and differences of architectures are adjusted by a single parameter. We evaluated our model using two applications and six different boards. We showed by using profile information for a single board, that the model is general enough to predict the execution time of an application with different input sizes and on different boards with the same architecture. Our model predictions were within 0.8 to 1.2 times the measured execution times, which are reasonable for such a simple model. These results indicate that the model is good enough to generalize the predictions for different problem sizes and GPU configurations. Marcos Amaris, Daniel Cordeiro, Alfredo Goldman, Raphael Y. de Camargo |
HiPC | 4 |
| 2013 | Gene regulatory networks inference using a multi-GPU exhaustive search algorithmabstractBACKGROUND: Gene regulatory networks (GRN) inference is an important bioinformatics problem in which the gene interactions need to be deduced from gene expression data, such as microarray data. Feature selection methods can be applied to this problem. A feature selection technique is composed by two parts: a search algorithm and a criterion function. Among the search algorithms already proposed, there is the exhaustive search where the best feature subset is returned, although its computational complexity is unfeasible in almost all situations. The objective of this work is the development of a low cost parallel solution based on GPU architectures for exhaustive search with a viable cost-benefit. We use CUDA™, a general purpose parallel programming platform that allows the usage of NVIDIA® GPUs to solve complex problems in an efficient way. RESULTS: We developed a parallel algorithm for GRN inference based on multiple GPU cards and obtained encouraging speedups (order of hundreds), when assuming that each target gene has two multivariate predictors. Also, experiments using single and multiple GPUs were performed, indicating that the speedup grows almost linearly with the number of GPUs. CONCLUSION: In this work, we present a proof of principle, showing that it is possible to parallelize the exhaustive search algorithm in GPUs with encouraging results. Although our focus in this paper is on the GRN inference problem, the exhaustive search technique based on GPU developed here can be applied (with minor adaptations) to other combinatorial problems. Fabrizio F. Borelli, Raphael Y. de Camargo, David Correa Martins Jr., Luiz C. S. Rozante |
BMC Bioinform. | 2 |
| 2011 | A multi-GPU algorithm for communication in neuronal network simulationsabstractGraphical Processing Units (GPUs) are frequently used for simulations of physical and biological systems. The simulated systems are often composed of simple elements that communicate only with their neighbors. But in some systems, such as large-scale neuronal networks, each element can communicate with any other element in the simulation. In this work, we present an efficient CUDA algorithm that enables this type of communication, even when using multiple GPUs. We show that it can benefit from the large memory bandwidth and number of cores in the GPU, despite the small number of required floating point operations. We implemented and evaluated this algorithm in a GPU simulator for large-scale neuronal networks. We obtained speedups of over 10 for the communication steps for simulations with 50k neurons and 50M connections, using a single computer with 2 graphic boards with 2 GPUs each, when compared with a modern quad-core CPU. When we consider the complete neuronal network simulation, its execution was nearly 40 times faster in the GPU than in the CPU. Raphael Y. de Camargo |
HiPC | 1 |
| 2011 | A multi-GPU algorithm for large-scale neuronal networksabstractAbstract Large‐scale simulations of parts of the brain using detailed neuronal models to improve our understanding of brain functions are becoming a reality with the usage of supercomputers and large clusters. However, the high acquisition and maintenance cost of these computers, including the physical space, air conditioning, and electrical power, limits the number of simulations of this kind that scientists can perform. Modern commodity graphical cards, based on the CUDA platform, contain graphical processing units (GPUs) composed of hundreds of processors that can simultaneously execute thousands of threads and thus constitute a low‐cost solution for many high‐performance computing applications. In this work, we present a CUDA algorithm that enables the execution, on multiple GPUs, of simulations of large‐scale networks composed of biologically realistic Hodgkin–Huxley neurons. The algorithm represents each neuron as a CUDA thread, which solves the set of coupled differential equations that model each neuron. Communication among neurons located in different GPUs is coordinated by the CPU. We obtained speedups of 40 for the simulation of 200k neurons that received random external input and speedups of 9 for a network with 200k neurons and 20M neuronal connections, in a single computer with two graphic boards with two GPUs each, when compared with a modern quad‐core CPU. Copyright © 2010 John Wiley & Sons, Ltd. Raphael Y. de Camargo, Luiz C. S. Rozante, Siang Wun Song |
Concurr. Comput. Pract. Exp. | 1 |
| 2010 | Exploiting a Generic Approach to Construct Component-Based Systems Software in Linux EnvironmentsabstractComponent-based software engineering has recently emerged as a promising solution to the development of system-level software. Unfortunately, current approaches are limited to specific platforms and domains. This lack of generality is particularly problematic as it prevents knowledge sharing and generally drives development costs up. In the past, we have developed a generic approach to component-based software engineering for system-level software called OpenCom. In this paper, we present OpenComL an instantiation of OpenCom to Linux environments and show how it can be profiled to meet a range of system-level software in Linux environments. For this, we demonstrate its application to constructing a programmable router platform and a middleware for parallel environments. Jo Ueyama, Edmundo Roberto Mauro Madeira, François Taïani, Raphael Y. de Camargo, Paul Grace, Geoff Coulson |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2010 | Application execution management on the InteGrade opportunistic grid middleware
Francisco José da Silva e Silva, Fabio Kon, Alfredo Goldman, Marcelo Finger, Raphael Y. de Camargo, Fernando Castor Filho, Fábio M. Costa |
J. Parallel Distributed Comput. | 5 |
| 2007 | Design and Implementation of a Middleware for Data Storage in Opportunistic GridsabstractShared machines in opportunistic grids typically have large quantities of unused disk space. These resources could be used to store application and checkpointing data when the machines are idle, allowing those machines to share not only computational cycles, but also disk space. In this paper, we present the design and implementation of OppStore, a middleware that provides reliable distributed data storage using the free disk space from shared grid machines. The system utilizes a two-level peer-to-peer organization to connect grid machines in a scalable and fault- tolerant way. Finally, we use the concept of virtual ids to deal with resource heterogeneity, enabling heterogeneity- aware load-balancing selection of storage sites. Raphael Y. de Camargo, Fabio Kon |
CCGRID | 1 |
| 2006 | Checkpointing BSP parallel applications on the InteGrade Grid middlewareabstractAbstract InteGrade is a Grid middleware infrastructure that enables the use of idle computing power from user workstations. One of its goals is to support the execution of long‐running parallel applications that present a considerable amount of communication among application nodes. However, in an environment composed of shared user workstations spread across many different LANs, machines may fail, become inaccessible, or may switch from idle to busy very rapidly, compromising the execution of the parallel application in some of its nodes. Thus, to provide some mechanism for fault tolerance becomes a major requirement for such a system. In this paper, we describe the support for checkpoint‐based rollback recovery of Bulk Synchronous Parallel applications running over the InteGrade middleware. This mechanism consists of periodically saving application state to permit the application to restart its execution from an intermediate execution point in case of failure. A precompiler automatically instruments the source code of a C/C++ application, adding code for saving and recovering application state. A failure detector monitors the application execution. In case of failure, the application is restarted from the last saved global checkpoint. Copyright © 2005 John Wiley & Sons, Ltd. Raphael Y. de Camargo, Andrei Goldchleger, Fabio Kon, Alfredo Goldman |
Concurr. Comput. Pract. Exp. | 1 |
| 2005 | Portable checkpointing and communication for BSP applications on dynamic heterogeneous Grid environmentsabstractExecuting long-running parallel applications in opportunistic grid environments composed of heterogeneous, shared user workstations, is a daunting task. Machines may fail, become inaccessible, or may switch from idle to busy unexpectedly, compromising the execution of applications. A mechanism for fault-tolerance that supports these heterogeneous architectures is an important requirement for such a system. In this paper, we describe the support for fault-tolerant execution of BSP parallel applications on heterogeneous, shared workstations, precompiler instruments application source code to save state periodically into checkpoint files. In case of failure, it is possible to recover the stored state from these files. Generated checkpoints are portable and can be recovered in a machine of different architecture, with data representation conversions being performed at recovery time. The precompiler also modifies BSP parallel applications to allow execution on a grid composed of machines with different architectures. We implemented a monitoring and recovering infrastructure in the InteGrade grid middleware. Experimental results evaluate the overhead incurred and the viability of using this approach in a grid environment. Raphael Y. de Camargo, Fabio Kon, Alfredo Goldman |
SBAC-PAD | 1 |