George Teodoro

dblp:78/87 · also George Luiz Medeiros Teodoro · DBLP profile ↗
← Back
59ranked-venue papers
15as first author
16since 2021 · last 2025
0000-0001-6289-3914ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 24 · 8 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-authorArtificial intelligence and machine learning · 3 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
YearPublicationVenuePosition
2025 PA-Star2: Fast Optimal Multiple Sequence Alignment for Asymmetric Multicore Processors
abstract
Multiple Sequence Alignment (MSA) is an important operation in Bioinformatics, used to simultaneously compare 3 or more sequences. The MSA problem was proven NP-Hard, so strategies have been proposed to reduce the search space and solve it in parallel. Recently, asymmetric multicore processors (AMPs) have become popular, with performance and energy-efficient cores, like the P-Cores and E-cores from Intel. However, parallel MSA applications have complex access patterns and adapting them for AMPs can be challenging. In this paper, we propose PA-Star21, an asymmetric-aware strategy based on A-Star, which computes optimal MSAs taking asymmetry into account when distributing the search space among threads. Our experimental results show that the proposed optimizations can reduce considerably the average execution time of PA-Star2 achieving a speedup of up to 7.70×. We also show that the asymmetric-aware strategy can reduce the average execution time for one of the hardest sequences set from the BAliBASE benchmark, when compared to the symmetric counterpart. Finally, we show that our approach is energy-efficient.11PA-Star2 is open source and the code is publicly available at1PA-Star2 is open source and the code is publicly available at https://github.com/danielsundfeld/astar_msa
Daniel Sundfeld, George Teodoro, Alba Cristina Magalhaes Alves de Melo
PDP2
2025 STEval: A framework for evaluating spatio-temporal crime prediction models
Gabriel Amarante, Matheus Pimenta, Yan Andrade, Matheus Senna, Rainer Menezes, Antônio Hot Faria, Marcelo Vilas-Boas, Frederico Martins de Paula Neto, João Paulo da Silva, Everton Renato de Sousa, Jamicel da Silva, Wagner Meira Jr., George Teodoro, Leonardo Rocha 0001, Renato Ferreira 0001
Eng. Appl. Artif. Intell.13
2025 IMI-GPU: Inverted multi-index for billion-scale approximate nearest neighbor search with GPUs
Alan Araujo, Willian de Oliveira Barreiros Junior, Jun Kong 0002, Renato Ferreira 0001, George Teodoro
J. Parallel Distributed Comput.5
2025 The Megapixel Approach for Efficient Execution of Irregular Wavefront Algorithms on GPUs
Mathias Oliveira, Willian de Oliveira Barreiros Junior, Renato Ferreira 0001, Alba Cristina Magalhaes Alves de Melo, George Teodoro
IEEE Trans. Parallel Distributed Syst.5
2024 A Descriptive and Predictive Analysis Tool for Criminal Data: A Case Study from Brazil
Yan Andrade, Matheus Pimenta, Gabriel Amarante, Antônio Hot Faria, Marcelo Vilas-Boas, João Paulo da Silva, Felipe Rocha, Jamicel da Silva, Wagner Meira Jr., George Teodoro, Leonardo Rocha 0001, Renato Ferreira 0001
ICCSA (2)10
2024 DuMato: An efficient warp-centric subgraph enumeration system for GPU
Samuel Ferraz, Vinícius Vitor dos Santos Dias, Carlos H. C. Teixeira, Srinivasan Parthasarathy 0001, George Teodoro, Wagner Meira Jr.
J. Parallel Distributed Comput.5
2023 Effective and efficient active learning for deep learning-based tissue image analysis
abstract
MOTIVATION: Deep learning attained excellent results in digital pathology recently. A challenge with its use is that high quality, representative training datasets are required to build robust models. Data annotation in the domain is labor intensive and demands substantial time commitment from expert pathologists. Active learning (AL) is a strategy to minimize annotation. The goal is to select samples from the pool of unlabeled data for annotation that improves model accuracy. However, AL is a very compute demanding approach. The benefits for model learning may vary according to the strategy used, and it may be hard for a domain specialist to fine tune the solution without an integrated interface. RESULTS: We developed a framework that includes a friendly user interface along with run-time optimizations to reduce annotation and execution time in AL in digital pathology. Our solution implements several AL strategies along with our diversity-aware data acquisition (DADA) acquisition function, which enforces data diversity to improve the prediction performance of a model. In this work, we employed a model simplification strategy [Network Auto-Reduction (NAR)] that significantly improves AL execution time when coupled with DADA. NAR produces less compute demanding models, which replace the target models during the AL process to reduce processing demands. An evaluation with a tumor-infiltrating lymphocytes classification application shows that: (i) DADA attains superior performance compared to state-of-the-art AL strategies for different convolutional neural networks (CNNs), (ii) NAR improves the AL execution time by up to 4.3×, and (iii) target models trained with patches/data selected by the NAR reduced versions achieve similar or superior classification quality to using target CNNs for data selection. AVAILABILITY AND IMPLEMENTATION: Source code: https://github.com/alsmeirelles/DADA.
André L. S. Meirelles, Tahsin M. Kurç, Jun Kong 0002, Renato Ferreira 0001, Joel H. Saltz, George Teodoro
Bioinform.6
2023 Self-supervised semantic segmentation of retinal pigment epithelium cells in flatmount fluorescent microscopy images
abstract
MOTIVATION: Morphological analyses with flatmount fluorescent images are essential to retinal pigment epithelial (RPE) aging studies and thus require accurate RPE cell segmentation. Although rapid technology advances in deep learning semantic segmentation have achieved great success in many biomedical research, the performance of these supervised learning methods for RPE cell segmentation is still limited by inadequate training data with high-quality annotations. RESULTS: To address this problem, we develop a Self-Supervised Semantic Segmentation (S4) method that utilizes a self-supervised learning strategy to train a semantic segmentation network with an encoder-decoder architecture. We employ a reconstruction and a pairwise representation loss to make the encoder extract structural information, while we create a morphology loss to produce the segmentation map. In addition, we develop a novel image augmentation algorithm (AugCut) to produce multiple views for self-supervised learning and enhance the network training performance. To validate the efficacy of our method, we applied our developed S4 method for RPE cell segmentation to a large set of flatmount fluorescent microscopy images, we compare our developed method for RPE cell segmentation with other state-of-the-art deep learning approaches. Compared with other state-of-the-art deep learning approaches, our method demonstrates better performance in both qualitative and quantitative evaluations, suggesting its promising potential to support large-scale cell morphological analyses in RPE aging investigations. AVAILABILITY AND IMPLEMENTATION: The codes and the documentation are available at: https://github.com/jkonglab/S4_RPE.
Hanyi Yu, Fusheng Wang 0001, George Teodoro, Xiaoyuan Guo, John Nickerson, Jun Kong 0002
Bioinform.3
2023 Spatial-aware data partition for distributed memory parallelization of ANN search in multimedia retrieval
Guilherme Andrade, Renato Ferreira 0001, George Teodoro
Parallel Comput.3
2022 Efficient Strategies for Graph Pattern Mining Algorithms on GPUs
abstract
Graph Pattern Mining (GPM) is an important, rapidly evolving, and computation demanding area. GPM computation relies on subgraph enumeration, which consists in extracting subgraphs that match a given property from an input graph. Graphics Processing Units (GPUs) have been an effective platform to accelerate applications in many areas. However, the irregularity of subgraph enumeration makes it challenging for efficient execution on GPU due to typical uncoalesced memory access, divergence, and load imbalance. Unfortunately, these aspects have not been fully addressed in previous work. Thus, this work proposes novel strategies to design and implement subgraph enumeration efficiently on GPU. We support a depth-first search style search (DFS-wide) that maximizes memory performance while providing enough parallelism to be exploited by the GPU, along with a warp-centric design that minimizes execution divergence and improves utilization of the computing capabilities. We also propose a low-cost load balancing layer to avoid idleness and redistribute work among thread warps in a GPU. Our strategies have been deployed in a system named DuMato, which provides a simple programming interface to allow efficient implementation of GPM algorithms. Our evaluation has shown that DuMato is often an order of magnitude faster than state-of-the-art GPM systems and can mine larger subgraphs (up to 12 vertices).
Samuel Ferraz, Vinícius Vitor dos Santos Dias, Carlos H. C. Teixeira, George Teodoro, Wagner Meira Jr.
SBAC-PAD4
2022 A spatial attention guided deep learning system for prediction of pathological complete response using breast cancer histopathology images
abstract
MOTIVATION: Predicting pathological complete response (pCR) to neoadjuvant chemotherapy (NAC) in triple-negative breast cancer (TNBC) patients accurately is direly needed for clinical decision making. pCR is also regarded as a strong predictor of overall survival. In this work, we propose a deep learning system to predict pCR to NAC based on serial pathology images stained with hematoxylin and eosin and two immunohistochemical biomarkers (Ki67 and PHH3). To support human prior domain knowledge-based guidance and enhance interpretability of the deep learning system, we introduce a human knowledge-derived spatial attention mechanism to inform deep learning models of informative tissue areas of interest. For each patient, three serial breast tumor tissue sections from biopsy blocks were sectioned, stained in three different stains and integrated. The resulting comprehensive attention information from the image triplets is used to guide our prediction system for prognostic tissue regions. RESULTS: The experimental dataset consists of 26 419 pathology image patches of 1000×1000 pixels from 73 TNBC patients treated with NAC. Image patches from randomly selected 43 patients are used as a training dataset and images patches from the rest 30 are used as a testing dataset. By the maximum voting from patch-level results, our proposed model achieves a 93% patient-level accuracy, outperforming baselines and other state-of-the-art systems, suggesting its high potential for clinical decision making. AVAILABILITY AND IMPLEMENTATION: The codes, the documentation and example data are available on an open source at: https://github.com/jkonglab/PCR_Prediction_Serial_WSIs_biomarkers. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hongyi Duanmu, Shristi Bhattarai, Hongxiao Li, Fusheng Wang 0001, George Teodoro, Keerthi Gogineni, Preeti Subhedar, Umay Kiraz, Emiel A. M. Janssen, Ritu Aneja, Jun Kong 0002
Bioinform.6
2022 Efficient microscopy image analysis on CPU-GPU systems with cost-aware irregular data partitioning
Willian de Oliveira Barreiros Junior, Alba Cristina Magalhaes Alves de Melo, Jun Kong 0002, Renato Ferreira 0001, Tahsin M. Kurç, Joel H. Saltz, George Teodoro
J. Parallel Distributed Comput.7
2021 Spatial Attention-Based Deep Learning System for Breast Cancer Pathological Complete Response Prediction with Serial Histopathology Images in Multiple Stains
Hongyi Duanmu, Shristi Bhattarai, Hongxiao Li, Chia Cheng Cheng, Fusheng Wang 0001, George Teodoro, Emiel A. M. Janssen, Keerthi Gogineni, Preeti Subhedar, Ritu Aneja, Jun Kong 0002
MICCAI (8)6
2021 Foveal blur-boosted segmentation of nuclei in histopathology images with shape prior knowledge and probability map constraints
abstract
MOTIVATION: In most tissue-based biomedical research, the lack of sufficient pathology training images with well-annotated ground truth inevitably limits the performance of deep learning systems. In this study, we propose a convolutional neural network with foveal blur enriching datasets with multiple local nuclei regions of interest derived from original pathology images. We further propose a human-knowledge boosted deep learning system by inclusion to the convolutional neural network new loss function terms capturing shape prior knowledge and imposing smoothness constraints on the predicted probability maps. RESULTS: Our proposed system outperforms all state-of-the-art deep learning and non-deep learning methods by Jaccard coefficient, Dice coefficient, Accuracy and Panoptic Quality in three independent datasets. The high segmentation accuracy and execution speed suggest its promising potential for automating histopathology nuclei segmentation in biomedical research and clinical settings. AVAILABILITY AND IMPLEMENTATION: The codes, the documentation and example data are available on an open source at: https://github.com/HongyiDuanmu26/FovealBoosted. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hongyi Duanmu, Fusheng Wang 0001, George Teodoro, Jun Kong 0002
Bioinform.3
2021 Online multimedia retrieval on CPU-GPU platforms with adaptive work partition
Rafael Souza, André Fernandes, Thiago S. F. X. Teixeira, George Teodoro, Renato Ferreira 0001
J. Parallel Distributed Comput.4
2021 Parallel Fine-Grained Comparison of Long DNA Sequences in Homogeneous and Heterogeneous GPU Platforms With Pruning
abstract
The parallelization of Smith-Waterman (SW) sequence comparison tools for long DNA sequences has been a big challenge over the years, requesting the use of several devices and sophisticated optimizations. Pruning is one of these optimizations, which can reduce considerably the amount of computation. This article proposes MultiBP, a sequence comparison solution in multiple GPUs with block pruning. Two MultiBP strategies are proposed. In static score-sharing, workload is statically distributed to the GPUs, and the best score is sent to neighbor GPUs to simulate a global view. In the dynamic strategy, execution is divided into cycles and workload is dynamically assigned, according to the GPUs processing rate. MultiBP was integrated to MASA-CUDAlign and tested in homogeneous and heterogeneous platforms, with different NVidia GPU architectures. The best results in our homogeneous and heterogeneous platforms were mostly obtained by the static and dynamic approaches, respectively. We also show that our decision module is able to select the best strategy in most cases. Finally, the comparison of the human and chimpanzee chromosomes 1 in a cluster with 512 V100 NVidia GPUs took 11 minutes and obtained the impressive rate of 82,822 GCUPS (Billions of Cells Updated per Second) which is, to our knowledge, the best performance for SW tools in GPUs.
Marco Antonio C. de Figueiredo, João Paulo Navarro, Edans Flavius de Oliveira Sandes, George Teodoro, Alba Cristina Magalhaes Alves de Melo
IEEE Trans. Parallel Distributed Syst.4
2020 Parallel Comparison of Huge DNA Sequences in Multiple GPUs with Block Pruning
abstract
Sequence comparison is a task performed in several Bioinformatics applications daily all over the world. Algorithms that retrieve the optimal result have quadratic time complexity, requiring a huge amount of computing power when the sequences compared are long. In order to reduce the execution time, many parallel solutions have been proposed in the literature. Nevertheless, depending on the sizes of the sequences, even those parallel solutions take hours or days to complete. Pruning techniques can significantly improve the performance of the parallel solutions and a few approaches have been proposed to provide pruning capabilities for sequence comparison applications. This paper proposes and evaluates a variant of the block pruning approach that runs in multiple GPUs, in homogeneous or heterogeneous environments. Experimental results obtained with DNA sequences in two testbeds show that significant performance gains are obtained with pruning, compared to its non-pruning counterpart, achieving the impressive performance of 694.8 GCUPS (Billions of Cells Updated per Second) for four GPUs.
Marco Antonio C. de Figueiredo, Edans Flavius de Oliveira Sandes, George Teodoro, Alba Cristina Magalhaes Alves de Melo
PDP3
2020 Scalable and Efficient Spatial-Aware Parallelization Strategies for Multimedia Retrieval
abstract
Similarity search is a key operation in several multimedia applications, including online Content-Based Multimedia Retrieval (CBMR) services. These applications have to deal with very large databases and are submitted to high query rates. In this context, scalability in distributed memory system is critical to assemble the required computing power and memory space. However, we have identified that the Data Equal Split (DES) parallelization and associated data partition strategy employed by the related works on the domain have limitations in terms of efficiency and scalability. Therefore, in this paper, we developed and implemented a framework for similarity search execution on distributed memory machines and proposed a novel class of data partition strategies that takes into account the data spatial organization in its distribution. This approach leads to a reduction in communication traffic and in costs associated with processing each task in local searches carried out in the distributed machine. Our approach attained a speedup of 2.4× on top of DES in the baseline case (5 nodes) and also achieves higher scalability efficiency and is 14.5× faster when 160 nodes are used. In fact, our novel data organization led to superlinear scalability in all configurations evaluated.
Guilherme Andrade, George Teodoro, Renato Ferreira 0001
SBAC-PAD2
2020 Optimizing parameter sensitivity analysis of large-scale microscopy image analysis workflows with multilevel computation reuse
abstract
Parameter sensitivity analysis (SA) is an effective tool to gain knowledge about complex analysis applications and assess the variability in their analysis results. However, it is an expensive process as it requires the execution of the target application multiple times with a large number of different input parameter values. In this work, we propose optimizations to reduce the overall computation cost of SA in the context of analysis applications that segment high-resolution slide tissue images, ie, images with resolutions of 100k × 100k pixels. Two cost-cutting techniques are combined to efficiently execute SA: use of distributed hybrid systems for parallel execution and computation reuse at multiple levels of an analysis pipeline to reduce the amount of computation. These techniques were evaluated using a cancer image analysis workflow on a hybrid cluster with 256 nodes, each with an Intel Phi and a dual socket CPU. Our parallel execution method attained an efficiency of over 90% on 256 nodes. The hybrid execution on the CPU and Intel Phi improved the performance by 2×. Multilevel computation reuse led to performance gains of over 2.9×.
Willian de Oliveira Barreiros Junior, Jeremias Moreira, Tahsin M. Kurç, Jun Kong 0002, Alba Cristina Magalhaes Alves de Melo, Joel H. Saltz, George Teodoro
Concurr. Comput. Pract. Exp.7
2020 Using GPU to accelerate the pairwise structural RNA alignment with base pair probabilities
abstract
Summary Structural alignments of Ribonucleic acid (RNA) sequences solved by the Sankoff algorithm are computationally expensive and often require constraints to be used in practice. Modern Graphics Processing Units (GPUs) contain more than 1000 cores, which compute in parallel to speed up applications. Here, we present a GPU‐based solution to the RNA structural alignment problem that makes use of precalculated base pair probabilities on the individual sequences. We designed and developed an unconstrained version of the Sankoff algorithm, obtaining the optimal result and calculating the entire four‐dimension dynamic programming matrix (4D DP). Our approach uses a two‐level wavefront strategy to exploit parallelism. The 4D DP matrix is divided in one external matrix (EM) and several internal matrices (IM). We applied wavefront strategies on the EM and IMs in a two‐level hierarchical way. At the first level, the wavefront is applied to the EM, calculating the cells that belong to the same diagonal in parallel. In the second level, since each cell in the EM is itself an IM matrix, the cells that belong to the same IM diagonal are calculated in parallel. The results obtained with real RNA sequences show that our GPU version is capable of outperforming a multicore CPU version of the unconstrained version of the Sankoff algorithm. Compared with the CPU‐based version running on 32 cores, our approach is able to achieve a speedup of 7.81x on the NVidia Tesla P100. In this case, the execution time was reduced from 6 hours and 18 minutes (32 cores) to 48 minutes and 20 seconds (GPU).
Daniel Sundfeld, George Teodoro, Jakob Hull Havgaard, Jan Gorodkin, Alba Cristina Magalhaes Alves de Melo
Concurr. Comput. Pract. Exp.2
2020 Bitmap filter: Speeding up exact set similarity joins with bitwise operations
Edans Flavius de Oliveira Sandes, George Teodoro, Alba Cristina Magalhaes Alves de Melo
Inf. Syst.2
2019 Medical Imaging Processing Architecture on ATMOSPHERE Federated Platform
abstract
[EN] This paper describes the development of applications in the frame of the ATMOSPHERE platform. ATMOSPHERE provides means for developing container-based applications over a federated cloud offering measurin he trustworthiness of the applications. In this paper we show the design of a transcontinental application in the frame of medical imaging that keeps the data at one end and uses the processing capabilities of the resources available at the other end. The applications are described using TOSCA blueprints and the federation of IaaS resources is performed by the Fogbow middleware. Privacy guarantees are provided by means of SCONE and intensive computing resources are integrated through the use of GPUs directly mounted on the containers.
Ignacio Blanquer, Angel Alberich-Bayarri, Fabio García-Castro, George Teodoro, André Meirelles, Bruno Nascimento, Wagner Meira Jr., Antônio L. P. Ribeiro
CLOSER4
2019 MASA-OpenCL: Parallel pruned comparison of long DNA sequences with OpenCL
abstract
Summary Biological sequence comparison is often used as an auxiliary task in the analysis of genetic material. Pairwise comparison algorithms like Smith‐Waterman evaluate two strings representing sequences of proteins, DNA or RNA to obtain optimal alignment between them. Many applications have been proposed to address the sequence comparison problem, prioritizing the use of graphics cards and proprietary languages such as CUDA. In this paper, we propose and evaluate MASA‐OpenCL, an OpenCL solution for comparing long DNA sequences that is based on the MASA sequence alignment framework, with pruning capability proportional to the similarity of the sequences compared. The results of MASA‐OpenCL were compared to its CUDA counterpart (MASA‐CUDAlign) and, in most cases, MASA‐OpenCL achieved better performance. In order to better understand the behavior of MASA‐OpenCL, we performed a statistical analysis considering 11 comparisons of sequences with high, medium and low similarity in 4 GPUs. As a result, we obtained a multiple linear regression model that considers (a) the sizes of the sequences, (b) the similarity between them, (c) the computational power of the GPU, and (d) the GPU memory bandwidth. We used this model to predict the performance in two other GPUs, with low error rates.
Marco Antonio C. de Figueiredo, Edans Flavius de Oliveira Sandes, Genaína Nunes Rodrigues, George Teodoro, Alba Cristina Magalhaes Alves de Melo
Concurr. Comput. Pract. Exp.4
2019 MaReIA: a cloud MapReduce based high performance whole slide image analysis framework
Hoang Vo, Jun Kong 0002, Dejun Teng, Yanhui Liang, Ablimit Aji, George Teodoro, Fusheng Wang 0001
Distributed Parallel Databases6
2019 Large-scale parallel similarity search with Product Quantization for online multimedia services
Guilherme Andrade, André Fernandes, Jeremias M. Gomes, Renato Ferreira 0001, George Teodoro
J. Parallel Distributed Comput.5
2018 Formalization of Block Pruning: Reducing the Number of Cells Computed in Exact Biological Sequence Comparison Algorithms
abstract
This is a pre-copyedited, author-produced version of an article accepted for publication in The Computer Journal following peer review. The version of record Edans F O Sandes, George L M Teodoro, Maria Emilia M T Walter, Xavier Martorell, Eduard Ayguade, Alba C M A Melo; Formalization of Block Pruning: Reducing the Number of Cells Computed in Exact Biological Sequence Comparison Algorithms, The Computer Journal, Volume 61, Issue 5, 1 May 2018, Pages 687–713 is available online at: The Computer Journal https://academic.oup.com/comjnl/article-abstract/61/5/687/4539903 and https://doi.org/10.1093/comjnl/bxx090.
Edans Flavius de Oliveira Sandes, George Teodoro, Maria Emília M. T. Walter, Xavier Martorell, Eduard Ayguadé, Alba Cristina Magalhaes Alves de Melo
Comput. J.2
2018 Cooperative and out-of-core execution of the irregular wavefront propagation pattern on hybrid machines with Intel® Xeon Phi™
abstract
The Irregular Wavefront Propagation Pattern (IWPP) is a core computing structure in several image analysis operations. Efficient implementation of IWPP on the Intel Xeon Phi is difficult because of the irregular data access and computation characteristics. The traditional IWPP algorithm relies on atomic instructions, which are not available in the SIMD set of the Intel Phi. To overcome this limitation, we have proposed a new IWPP algorithm that can take advantage of non-atomic SIMD instructions supported on the Intel Xeon Phi. We have also developed and evaluated methods to use CPU and Intel Phi cooperatively for parallel execution of the IWPP algorithms. Our new cooperative IWPP version is also able to handle large out-of-core images that would not fit into the memory of the accelerator. The new IWPP algorithm is used to implement the Morphological Reconstruction and Fill Holes operations, which are operations commonly found in image analysis applications. The vectorization implemented with the new IWPP has attained improvements of up to about 5× on top of the original IWPP and significant gains as compared to state-of-the-art the CPU and GPU versions. The new version running on an Intel Phi is 6.21× and 3.14× faster than running on a 16-core CPU and on a GPU, respectively. Finally, the cooperative execution using two Intel Phi devices and a multi-core CPU has reached performance gains of 2.14× as compared to the execution using a single Intel Xeon Phi.
Jeremias M. Gomes, Alba Cristina Magalhaes Alves de Melo, Jun Kong 0002, Tahsin M. Kurç, Joel H. Saltz, George Teodoro
Concurr. Comput. Pract. Exp.6
2018 PA-Star: A disk-assisted parallel A-Star strategy with locality-sensitive hash for multiple sequence alignment
Daniel Sundfeld, Caina Razzolini, George Teodoro, Azzedine Boukerche, Alba Cristina Magalhaes Alves de Melo
J. Parallel Distributed Comput.3
2017 Parallel and Efficient Sensitivity Analysis of Microscopy Image Segmentation Workflows in Hybrid Systems
abstract
We investigate efficient sensitivity analysis (SA) of algorithms that segment and classify image features in a large dataset of high-resolution images. Algorithm SA is the process of evaluating variations of methods and parameter values to quantify differences in the output. A SA can be very compute demanding because it requires re-processing the input dataset several times with different parameters to assess variations in output. In this work, we introduce strategies to efficiently speed up SA via runtime optimizations targeting distributed hybrid systems and reuse of computations from runs with different parameters. We evaluate our approach using a cancer image analysis workflow on a hybrid cluster with 256 nodes, each with an Intel Phi and a dual socket CPU. The SA attained a parallel efficiency of over 90% on 256 nodes. The cooperative execution using the CPUs and the Phi available in each node with smart task assignment strategies resulted in an additional speedup of about 2×. Finally, multi-level computation reuse lead to an additional speedup of up to 2.46× on the parallel version. The level of performance attained with the proposed optimizations will allow the use of SA in large-scale studies.
Willian de Oliveira Barreiros Junior, George Teodoro, Tahsin M. Kurç, Jun Kong 0002, Alba Cristina Magalhaes Alves de Melo, Joel H. Saltz
CLUSTER2
2017 Online Multimedia Similarity Search with Response Time-Aware Parallelism and Task Granularity Auto-Tuning
abstract
This paper presents an efficient parallel implementation of the Product Quantization based approximate nearest neighbor multimedia similarity search indexing (PQANNS). The parallel PQANNS efficiently answers nearest neighbor queries by exploiting the ability of the quantization approach to reduce the data dimensionality (and memory demand) and by leveraging parallelism to speed up the search capabilities of the application. Our solution is also optimized to minimize query response times under scenarios with fluctuating query rates (load) as observed in online services. To achieve this goal, we have developed strategies to dynamically select the parallelism configuration and task granularity that minimizes the query response times during the execution. The proposed strategies (ADAPT and ADAPT+G) were thoroughly evaluated and have shown, for instance, to reduce the query response times in 6.4× as compared to the best static configuration of parallelism and task granularity.
Guilherme Andrade, George Teodoro, Renato Ferreira 0001
SBAC-PAD2
2017 Algorithm sensitivity analysis and parameter tuning for tissue image segmentation pipelines
abstract
Motivation: Sensitivity analysis and parameter tuning are important processes in large-scale image analysis. They are very costly because the image analysis workflows are required to be executed several times to systematically correlate output variations with parameter changes or to tune parameters. An integrated solution with minimum user interaction that uses effective methodologies and high performance computing is required to scale these studies to large imaging datasets and expensive analysis workflows. Results: The experiments with two segmentation workflows show that the proposed approach can (i) quickly identify and prune parameters that are non-influential; (ii) search a small fraction (about 100 points) of the parameter search space with billions to trillions of points and improve the quality of segmentation results (Dice and Jaccard metrics) by as much as 1.42× compared to the results from the default parameters; (iii) attain good scalability on a high performance cluster with several effective optimizations. Conclusions: Our work demonstrates the feasibility of performing sensitivity analyses, parameter studies and auto-tuning with large datasets. The proposed framework can enable the quantification of error estimations and output variations in image segmentation pipelines. Availability and Implementation: Source code: https://github.com/SBU-BMI/region-templates/ . Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
George Teodoro, Tahsin M. Kurç, Luis F. R. Taveira, Alba Cristina Magalhaes Alves de Melo, Yi Gao 0002
Bioinform.1
2016 MASE-BDI: agent-based simulator for environmental land change with efficient and parallel auto-tuning
Cássio Giorgio Couto Coelho, Carolina G. Abreu, Rafael Marconi Ramos, Aldo H. D. Mendes, George Teodoro, Célia Ghedini Ralha
Appl. Intell.5
2016 CUDAlign 4.0: Incremental Speculative Traceback for Exact Chromosome-Wide Alignment in GPU Clusters
abstract
This paper proposes and evaluates CUDAlign 4.0, a parallel strategy to obtain the optimal alignment of huge DNA sequences in multi-GPU platforms, using the exact Smith–Waterman (SW) algorithm. In the first phase of CUDAlign 4.0, a huge Dynamic Programming (DP) matrix is computed by multiple GPUs, which asynchronously communicate border elements to the right neighbor in order to find the optimal score. After that, the traceback phase of SW is executed. The efficient parallelization of the traceback phase is very challenging because of the high amount of data dependency, which particularly impacts the performance and limits the application scalability. In order to obtain a multi-GPU highly parallel traceback phase, we propose and evaluate a new parallel traceback algorithm called Incremental Speculative Traceback (IST), which pipelines the traceback phase, speculating incrementally over the values calculated so far, producing results in advance. With CUDAlign 4.0, we were able to calculate SW matrices with up to 60 Peta cells, obtaining the optimal local alignments of all Human and Chimpanzee homologous chromosomes, whose sizes range from 26 Millions of Base Pairs (MBP) up to 249 MBP. As far as we know, this is the first time such comparison was made with the SW exact method. We also show that the IST algorithm is able to reduce the traceback time from 2.15$\times$up to 21.03$\times$, when compared with the baseline traceback algorithm. The human$\times$chimpanzee chromosome 5 comparison (180 MBP$\times$183 MBP) attained 10,370.00 GCUPS (Billions of Cells Updated per Second) using 384 GPUs, with a speculation hit ratio of 98.2 percent.
Edans Flavius de Oliveira Sandes, Guillermo Miranda, Xavier Martorell, Eduard Ayguadé, George Teodoro, Alba Cristina Magalhaes Alves de Melo
IEEE Trans. Parallel Distributed Syst.5
2015 Parallel A-Star Multiple Sequence Alignment with Locality-Sensitive Hash Functions
abstract
In this paper, we propose and evaluate a parallel solution for the exact Multiple Sequence Alignment problem based on the A-Star algorithm. In our parallel solution, we use a multi-index data structure, templates and a locality-sensitive hash function. The results were collected in two machines (4 cores and 32 cores), with real and synthetic sequence sets ranging from 3 to 14 sequences. We show that our parallel solution executes 2.89× and 4.77× faster than a state-of-the-art parallel MSA tool, with a proportional increase in memory usage, when comparing 2 hard instances of the benchmark Bali base reference set 1.
Daniel Sundfeld, George Teodoro, Alba Cristina Magalhaes Alves de Melo
CISIS2
2015 A 3D Primary Vessel Reconstruction Framework with Serial Microscopy Images
Yanhui Liang, Fusheng Wang 0001, Darren Treanor, Derek R. Magee, George Teodoro, Jun Kong 0002
MICCAI (3)5
2015 Efficient Irregular Wavefront Propagation Algorithms on Intel(R) Xeon Phi(TM)
abstract
We investigate the execution of the Irregular Wave front Propagation Pattern (IWPP), a fundamental computing structure used in several image analysis operations, on the Intel® Xeon PhiTM co-processor. An efficient implementation of IWPP on the Xeon Phi is a challenging problem because of IWPP's irregularity and the use of atomic instructions in the original IWPP algorithm to resolve race conditions. On the Xeon Phi, the use of SIMD and vectorization instructions is critical to attain high performance. However, SIMD atomic instructions are not supported. Therefore, we propose a new IWPP algorithm that can take advantage of the supported SIMD instruction set. We also evaluate an alternate storage container (priority queue) to track active elements in the wave front in an effort to improve the parallel algorithm efficiency. The new IWPP algorithm is evaluated with Morphological Reconstruction and Imfill operations as use cases. Our results show performance improvements of up to 5.63× on top of the original IWPP due to vectorization. Moreover, the new IWPP achieves speedups of 45.7× and 1.62×, respectively, as compared to efficient CPU and GPU implementations.
Jeremias M. Gomes, George Teodoro, Alba Cristina Magalhaes Alves de Melo, Jun Kong 0002, Tahsin M. Kurç, Joel H. Saltz
SBAC-PAD2
2015 Scalable analysis of Big pathology image data cohorts using efficient methods and high-performance computing strategies
abstract
BACKGROUND: We describe a suite of tools and methods that form a core set of capabilities for researchers and clinical investigators to evaluate multiple analytical pipelines and quantify sensitivity and variability of the results while conducting large-scale studies in investigative pathology and oncology. The overarching objective of the current investigation is to address the challenges of large data sizes and high computational demands. RESULTS: The proposed tools and methods take advantage of state-of-the-art parallel machines and efficient content-based image searching strategies. The content based image retrieval (CBIR) algorithms can quickly detect and retrieve image patches similar to a query patch using a hierarchical analysis approach. The analysis component based on high performance computing can carry out consensus clustering on 500,000 data points using a large shared memory system. CONCLUSIONS: Our work demonstrates efficient CBIR algorithms and high performance computing can be leveraged for efficient analysis of large microscopy images to meet the challenges of clinically salient applications in pathology. These technologies enable researchers and clinical investigators to make more effective use of the rich informational content contained within digitized microscopy specimens.
Tahsin M. Kurç, Xin Qi 0007, Daihou Wang, Fusheng Wang 0001, George Teodoro, Lee A. D. Cooper, Michael Nalisnik, Lin Yang 0002, Joel H. Saltz, David J. Foran
BMC Bioinform.5
2014 Comparative Performance Analysis of Intel (R) Xeon Phi (TM), GPU, and CPU: A Case Study from Microscopy Image Analysis
abstract
We study and characterize the performance of operations in an important class of applications on GPUs and Many Integrated Core (MIC) architectures. Our work is motivated by applications that analyze low-dimensional spatial datasets captured by high resolution sensors, such as image datasets obtained from whole slide tissue specimens using microscopy scanners. Common operations in these applications involve the detection and extraction of objects (object segmentation), the computation of features of each extracted object (feature computation), and characterization of objects based on these features (object classification). In this work, we have identify the data access and computation patterns of operations in the object segmentation and feature computation categories. We systematically implement and evaluate the performance of these operations on modern CPUs, GPUs, and MIC systems for a microscopy image analysis application. Our results show that the performance on a MIC of operations that perform regular data access is comparable or sometimes better than that on a GPU. On the other hand, GPUs are significantly more efficient than MICs for operations that access data irregularly. This is a result of the low performance of MICs when it comes to random data access. We also have examined the coordinated use of MICs and CPUs. Our experiments show that using a performance aware task strategy for scheduling application operations improves performance about 1.29× over a first-come-first-served strategy. This allows applications to obtain high performance efficiency on CPU-MIC systems - the example application attained an efficiency of 84% on 192 nodes (3072 CPU cores and 192 MICs).
George Teodoro, Tahsin M. Kurç, Jun Kong 0002, Lee A. D. Cooper, Joel H. Saltz
IPDPS1
2014 Efficient Execution of Microscopy Image Analysis on CPU, GPU, and MIC Equipped Cluster Systems
abstract
High performance computing is experiencing a major paradigm shift with the introduction of accelerators, such as graphics processing units (GPUs) and Intel Xeon Phi (MIC). These processors have made available a tremendous computing power at low cost, and are transforming machines into hybrid systems equipped with CPUs and accelerators. Although these systems can deliver a very high peak performance, making full use of its resources in real-world applications is a complex problem. Most current applications deployed to these machines are still being executed in a single processor, leaving other devices underutilized. In this paper we explore a scenario in which applications are composed of hierarchical data flow tasks which are allocated to nodes of a distributed memory machine in coarse-grain, but each of them may be composed of several finer-grain tasks which can be allocated to different devices within the node. We propose and implement novel performance aware scheduling techniques that can be used to allocate tasks to devices. We evaluate our techniques using a pathology image analysis application used to investigate brain cancer morphology, and our experimental evaluation shows that the proposed scheduling strategies significantly outperforms other efficient scheduling techniques, such as Heterogeneous Earliest Finish Time - HEFT, in cooperative executions using CPUs, GPUs, and MICs. We also experimentally show that our strategies are less sensitive to inaccuracy in the scheduling input data and that the performance gains are maintained as the application scales.
Guilherme Andrade, Renato Ferreira 0001, George Teodoro, Leonardo Rocha 0001, Joel H. Saltz, Tahsin M. Kurç
SBAC-PAD3
2014 Large-Scale Distributed Locality-Sensitive Hashing for General Metric Data
Eliezer S. Silva, Thiago S. F. X. Teixeira, George Teodoro, Eduardo Valle
SISAP3
2014 Region templates: Data representation and management for high-throughput image analysis
abstract
We introduce a region template abstraction and framework for the efficient storage, management and processing of common data types in analysis of large datasets of high resolution images on clusters of hybrid computing nodes. The region template abstraction provides a generic container template for common data structures, such as points, arrays, regions, and object sets, within a spatial and temporal bounding box. It allows for different data management strategies and I/O implementations, while providing a homogeneous, unified interface to applications for data storage and retrieval. A region template application is represented as a hierarchical dataflow in which each computing stage may be represented as another dataflow of finer-grain tasks. The execution of the application is coordinated by a runtime system that implements optimizations for hybrid machines, including performance-aware scheduling for maximizing the utilization of computing devices and techniques to reduce the impact of data transfers between CPUs and GPUs. An experimental evaluation on a state-of-the-art hybrid cluster using a microscopy imaging application shows that the abstraction adds negligible overhead (about 3%) and achieves good scalability and high data transfer rates. Optimizations in a high speed disk based storage implementation of the abstraction to support asynchronous data transfers and computation result in an application performance gain of about 1.13×. Finally, a processing rate of 11,730 4K×4K tiles per minute was achieved for the microscopy imaging application on a cluster with 100 nodes (300 GPUs and 1,200 CPU cores). This computation rate enables studies with very large datasets.
George Teodoro, Tony Pan, Tahsin M. Kurç, Jun Kong 0002, Lee A. D. Cooper, Scott Klasky, Joel H. Saltz
Parallel Comput.1
2014 Approximate similarity search for online multimedia services on distributed CPU-GPU platforms
George Teodoro, Eduardo Valle, Nathan Mariano, Ricardo da Silva Torres, Wagner Meira Jr., Joel H. Saltz
VLDB J.1
2013 High-performance computational analysis of glioblastoma pathology images with database support identifies molecular and survival correlates
abstract
In this paper, we present a novel framework for microscopic image analysis of nuclei, data management, and high performance computation to support translational research involving nuclear morphometry features, molecular data, and clinical outcomes. Our image analysis pipeline consists of nuclei segmentation and feature computation facilitated by high performance computing with coordinated execution in multi-core CPUs and Graphical Processor Units (GPUs). All data derived from image analysis are managed in a spatial relational database supporting highly efficient scientific queries. We applied our image analysis workflow to 159 glioblastomas (GBM) from The Cancer Genome Atlas dataset. With integrative studies, we found statistics of four specific nuclear features were significantly associated with patient survival. Additionally, we correlated nuclear features with molecular data and found interesting results that support pathologic domain knowledge. We found that Proneural subtype GBMs had the smallest mean of nuclear Eccentricity and the largest mean of nuclear Extent, and MinorAxisLength. We also found gene expressions of stem cell marker MYC and cell proliferation maker MKI67 were correlated with nuclear features. To complement and inform pathologists of relevant diagnostic features, we queried the most representative nuclear instances from each patient population based on genetic and transcriptional classes. Our results demonstrate that specific nuclear features carry prognostic significance and associations with transcriptional and genetic classes, highlighting the potential of high throughput pathology image analysis as a complementary approach to human-based review and translational research.
Jun Kong 0002, Fusheng Wang 0001, George Teodoro, Lee A. D. Cooper, Carlos Sanchez Moreno, Tahsin M. Kurç, Tony Pan, Joel H. Saltz, Daniel J. Brat
BIBM3
2013 High-throughput Analysis of Large Microscopy Image Datasets on CPU-GPU Cluster Platforms
abstract
Analysis of large pathology image datasets offers significant opportunities for the investigation of disease morphology, but the resource requirements of analysis pipelines limit the scale of such studies. Motivated by a brain cancer study, we propose and evaluate a parallel image analysis application pipeline for high throughput computation of large datasets of high resolution pathology tissue images on distributed CPU-GPU platforms. To achieve efficient execution on these hybrid systems, we have built runtime support that allows us to express the cancer image analysis application as a hierarchical data processing pipeline. The application is implemented as a coarse-grain pipeline of stages, where each stage may be further partitioned into another pipeline of fine-grain operations. The fine-grain operations are efficiently managed and scheduled for computation on CPUs and GPUs using performance aware scheduling techniques along with several optimizations, including architecture aware process placement, data locality conscious task assignment, data prefetching, and asynchronous data copy. These optimizations are employed to maximize the utilization of the aggregate computing power of CPUs and GPUs and minimize data copy overheads. Our experimental evaluation shows that the cooperative use of CPUs and GPUs achieves significant improvements on top of GPU-only versions (up to 1.6×) and that the execution of the application as a set of fine-grain operations provides more opportunities for runtime optimizations and attains better performance than coarser-grain, monolithic implementations used in other works. An implementation of the cancer image analysis pipeline using the runtime support was able to process an image dataset consisting of 36,848 4Kx4K-pixel image tiles (about 1.8TB uncompressed) in less than 4 minutes (150 tiles/second) on 100 nodes of a state-of-the-art hybrid cluster system.
George Teodoro, Tony Pan, Tahsin M. Kurç, Jun Kong 0002, Lee A. D. Cooper, Norbert Podhorszki, Scott Klasky, Joel H. Saltz
IPDPS1
2013 Efficient irregular wavefront propagation algorithms on hybrid CPU-GPU machines
George Teodoro, Tony Pan, Tahsin M. Kurç, Jun Kong 0002, Lee A. D. Cooper, Joel H. Saltz
Parallel Comput.1
2012 High Performance Computing for Integrative Analysis of Large Pathology Image Datasets
Tahsin M. Kurç, Joel H. Saltz, George Teodoro, Tony Pan, Lee A. D. Cooper, Jun Kong 0002, David A. Gutman, Daniel J. Brat, Fusheng Wang 0001
AMIA3
2012 Accelerating Large Scale Image Analyses on Parallel, CPU-GPU Equipped Systems
abstract
The past decade has witnessed a major paradigm shift in high performance computing with the introduction of accelerators as general purpose processors. These computing devices make available very high parallel computing power at low cost and power consumption, transforming current high performance platforms into heterogeneous CPU-GPU equipped systems. Although the theoretical performance achieved by these hybrid systems is impressive, taking practical advantage of this computing power remains a very challenging problem. Most applications are still deployed to either GPU or CPU, leaving the other resource under- or un-utilized. In this paper, we propose, implement, and evaluate a performance aware scheduling technique along with optimizations to make efficient collaborative use of CPUs and GPUs on a parallel system. In the context of feature computations in large scale image analysis applications, our evaluations show that intelligently co-scheduling CPUs and GPUs can significantly improve performance over GPU-only or multi-core CPU-only approaches.
George Teodoro, Tahsin M. Kurç, Tony Pan, Lee A. D. Cooper, Jun Kong 0002, Patrick M. Widener, Joel H. Saltz
IPDPS1
2011 Adaptive parallel approximate similarity search for responsive multimedia retrieval
abstract
This paper introduces Hypercurves, a flexible framework for pro- viding similarity search indexing to high throughput multimedia services. Hypercurves efficiently and effectively answers k-nearest neighbor searches on multigigabyte high-dimensional databases. It supports massively parallel processing and adapts at runtime its parallelization regimens to keep answer times optimal for either low and high demands. In order to achieve its goals, Hypercurves introduces new techniques for selecting parallelism configurations and allocating threads to computation cores, including hyperthreaded cores. Its efficiency gains are throughly validated on a large database of multimedia descriptors, where it presented near linear speedups and superlinear scaleups. The adaptation reduces query response times in 43% and 74% for both platforms tested, when compared to the best static parallelism regimens.
George Teodoro, Eduardo Valle, Nathan Mariano, Ricardo da Silva Torres, Wagner Meira Jr.
CIKM1
2010 Run-time optimizations for replicated dataflows on heterogeneous environments
abstract
The increases in multi-core processor parallelism and in the flexibility of many-core accelerator processors, such as GPUs, have turned traditional SMP systems into hierarchical, heterogeneous computing environments. Fully exploiting these improvements in parallel system design remains an open problem. Moreover, most of the current tools for the development of parallel applications for hierarchical systems concentrate on the use of only a single processor type (e.g., accelerators) and do not coordinate several heterogeneous processors. Here, we show that making use of all of the heterogeneous computing resources can significantly improve application performance. Our approach, which consists of optimizing applications at run-time by efficiently coordinating application task execution on all available processing units is evaluated in the context of replicated dataflow applications. The proposed techniques were developed and implemented in an integrated run-time system targeting both intra- and inter-node parallelism. The experimental results with a real-world complex biomedical application show that our approach nearly doubles the performance of the GPU-only implementation on a distributed heterogeneous accelerator cluster.
George Teodoro, Timothy D. R. Hartley, Ümit V. Çatalyürek, Renato Ferreira 0001
HPDC1
2010 Tree Projection-Based Frequent Itemset Mining on Multicore CPUs and GPUs
abstract
Frequent itemset mining (FIM) is a core operation for several data mining applications as association rules computation, correlations, document classification, and many others, which has been extensively studied over the last decades. Moreover, databases are becoming increasingly larger, thus requiring a higher computing power to mine them in reasonable time. At the same time, the advances in high performance computing platforms are transforming them into hierarchical parallel environments equipped with multi-core processors and many-core accelerators, such as GPUs. Thus, fully exploiting these systems to perform FIM tasks poses as a challenging and critical problem that we address in this paper. We present efficient multi-core and GPU accelerated parallelizations of the Tree Projection, one of the most competitive FIM algorithms. The experimental results show that our Tree Projection implementation scales almost linearly in a CPU shared-memory environment after careful optimizations, while the GPU versions are up to 173 times faster than standard the CPU version.
George Teodoro, Nathan Mariano, Wagner Meira Jr., Renato Ferreira 0001
SBAC-PAD1
2009 Coordinating the use of GPU and CPU for improving performance of compute intensive applications
abstract
GPUs have recently evolved into very fast parallel co-processors capable of executing general purpose computations extremely efficiently. At the same time, multi-core CPUs evolution continued and today's CPUs have 4-8 cores. These two trends, however, have followed independent paths in the sense that we are aware of very few works that consider both devices cooperating to solve general computations. In this paper we investigate the coordinated use of CPU and GPU to improve efficiency of applications even further than using either device independently. We use Anthill runtime environment, a data-flow oriented framework in which applications are decomposed into a set of event-driven filters, where for each event, the runtime system can use either GPU or CPU for its processing. For evaluation, we use a histopathology application that uses image analysis techniques to classify tumor images for neuroblas-toma prognosis. Our experimental environment includes dual and octa-core machines, augmented with GPUs and we evaluate our approach's performance for standalone and distributed executions. Our experiments show that a pure GPU optimization of the application achieved a factor of 15 to 49 times improvement over the single core CPU version, depending on the versions of the CPUs and GPUs. We also show that the execution can be further reduced by a factor of about 2 by using our runtime system that effectively choreographs the execution to run cooperatively both on GPU and on a single core of CPU. We improve on that by adding more cores, all of which were previously neglected or used ineffectively. In addition, the evaluation on a distributed environment has shown near linear scalability to multiple hosts.
George Teodoro, Rafael Sachetto Oliveira, Olcay Sertel, Metin Nafi Gürcan, Wagner Meira Jr., Ümit V. Çatalyürek, Renato Ferreira 0001
CLUSTER1
2009 Profiling General Purpose GPU Applications
abstract
We are witnessing an increasing adoption of GPUs for performing general purpose computation, which is usually known as GPGPU. The main challenge in developing such applications is that they often do not fit in the model required by the graphics processing devices, limiting the scope of applications that may be benefit from the computing power provided by GPUs. Even when the application fits GPU model, obtaining optimal resource usage is a complex task. In this work we propose a profiling tool for GPGPU applications. This tool use a profiling strategy based on performance predicates and is able to quantify the major sources of performance degradation while providing hints on how to improve the applications. We used our tool in CUDA programs and were able to understand and improve their performance.
Bruno Coutinho, George Teodoro, Rafael Sachetto Oliveira, Dorgival O. Guedes, Renato Ferreira 0001
SBAC-PAD2
2009 Exploiting Computational Resources in Distributed Heterogeneous Platforms
abstract
We have been witnessing a continuous growth of both heterogeneous computational platforms (e.g., Cell blades, or the joint use of traditional CPUs and GPUs) and multi- core processor architecture; and it is still an open question how applications can fully exploit such computational potential efficiently. In this paper we introduce a run-time environment and programming framework which supports the implementation of scalable and efficient parallel applications in such heterogeneous, distributed environments. We assess these issues through well-known kernels and actual applications that behave regularly and irregularly, which are not only relevant but also demanding in terms of computation and I/O. Moreover, the irregularity of these, as well as many other applications poses a challenge to the design and implementation of efficient parallel algorithms. Our experimental environment includes dual and octa-core machines augmented with GPUs and we evaluate our framework performance for standalone and distributed executions. The evaluation on a distributed environment has shown near to linear scale-ups for two data mining applications, while the applications performance, when using CPU and GPU, has been improved into around 25%, compared to the GPU-only versions.
George Teodoro, Rafael Sachetto Oliveira, Daniel Fireman, Dorgival O. Guedes, Renato Ferreira 0001
SBAC-PAD1
2008 Achieving Multi-Level Parallelism in the Filter-Labeled Stream Programming Model
abstract
New architectural trends in chip design resulted in machines with multiple processing units as well as efficient communication networks, leading to the wide availability of systems that provide multiple levels of parallelism, both inter- and intra-machine. Developing applications that efficiently make use of such systems is a challenge, specially for application-domain programmers. In this paper we present a new version of the Anthill programming environment that efficiently exploits multi-level parallelism and experimental results that demonstrate such efficiency. Anthill is based on the filter-stream model; in this model, applications are decomposed into a set of filters communicating through streams, which has already been shown to be efficient for expressing inter-machine parallelism. We replaced the filter run-time environment, originally process-oriented, with an event-oriented version. This new version allow programmers to efficiently express opportunities for parallelism within each compute node through a higher-level programming abstraction. We evaluated our solution on dual- and quad-core machines with two data mining applications: Eclat and KNN. Both had drops in execution time nearly proportional to the number of cores on a single machine. When using a cluster of dual-core machines, speed-ups were close to linear on the number of available cores for both applications, confirming event-oriented Anthill performs well both on the inter- and intra-machine parallelism levels.
George Teodoro, Daniel Fireman, Dorgival O. Guedes, Wagner Meira Jr., Renato Ferreira 0001
ICPP1
2008 A Reconfigurable Run-Time System for Filter-Stream Applications
abstract
The development of high level abstractions for programming distributed systems is becoming a crucial effort in computer science. Several frameworks have been proposed, which expose simplified programming abstractions that are useful for a broad class of applications and can be implemented efficiently on distributed systems. One such system is Anthill, based on the filter-stream programming model, in which applications are decomposed into sets of independent filters that communicate via streams. Anthill achieves high performance by allowing filters to be transparently replicated across several compute nodes.In this paper we present a global state manager for Anthill, which exports a simple abstraction to manipulate state variables for application filters. The state is distributed transparently among the instances of that filter, and our manager is designed to allow data migration from one filter instance to another, enabling Anthill to dynamically reconfigure applications at execution time.To evaluate our system, we used two well known data mining algorithms: a priori and k-means. Our results show that the framework incurs low overhead, 1.8% on average, and that the resulting system can effectively make use of new resources as they are made available, with execution times 3.57% slower, on average, than the minimum expected time for the reconfiguration scenario.
Daniel Fireman, George Teodoro, André Cardoso, Renato Ferreira 0001
SBAC-PAD2
2007 An Efficient and Reliable Scientific Workflow System
abstract
This paper presents a fault tolerance framework for applications that process data using a distributed network of user-defined operations in a pipelined fashion. The framework saves intermediate results and messages exchanged among application components in a distributed data management system to facilitate quick recovery from failures. The experimental results show that the framework scales well and our approach introduces very little overhead to application execution.
Tulio Tavares, George Teodoro, Tahsin M. Kurç, Renato Ferreira 0001, Dorgival O. Guedes, Wagner Meira Jr., Ümit V. Çatalyürek, Shannon Hastings, Scott Oster, Stephen Langella, Joel H. Saltz
CCGRID2
2006 A Run-time System for Efficient Execution of Scientific Workflows on Distributed Environments
abstract
Scientific workflow systems have been introduced in response to the demand of researchers from several domains of science who need to process and analyze increasingly larger datasets. The design of these systems is largely based on the observation that data analysis applications can be composed as pipelines or networks of computations on data. In this paper we present a run-time support system that is designed to facilitate this type of computation in distributed computing environments. Our system is optimized for data-intensive workflows, in which efficient management and retrieval of data, coordination of data processing and data movement, and check-pointing of intermediate results are critical and challenging issues. Experimental evaluation of our system shows that linear speedups can be achieved for sophisticated applications, which are implemented as a network of multiple data processing components
George Teodoro, Tulio Tavares, Renato Ferreira 0001, Tahsin M. Kurç, Wagner Meira Jr., Dorgival O. Guedes, Tony Pan, Joel H. Saltz
SBAC-PAD1
2005 Anthill: A Scalable Run-Time Environment for Data Mining Applications
abstract
Data mining techniques are becoming increasingly more popular as a reasonable means to collect summaries from the rapidly growing datasets in many areas. However, as the size of the raw data increases, parallel data mining algorithms are becoming a necessity. In this paper, we present a run-time support system that was designed to allow the efficient implementation of data-mining algorithms on heterogeneous distributed environments. We believe that the runtime framework is suitable for a broader class of applications, beyond data mining. We also present a parallelization strategy that is supported by the run-time system. We show scalability results of three different data-mining algorithms that were parallelized using our approach and our run-time support. All applications scale almost linearly up to a large number of nodes.
Renato Ferreira 0001, Wagner Meira Jr., Dorgival O. Guedes, Lúcia M. A. Drummond, Bruno Coutinho, George Teodoro, Tulio Tavares, Renata Braga Araújo, Guilherme T. Ferreira
SBAC-PAD6
2003 Load Balancing on Stateful Clustered Web Servers
abstract
One of the main challenges to the wide use of the Internet is the scalability of the servers, that is, their ability to handle the increasing demand. Scalability in stateful servers, which comprise e-commerce and other transaction-oriented servers, is even more difficult, since it is necessary to keep transaction data across requests from the same user. One common strategy for achieving scalability is to employ clustered servers, where the load is distributed among the various servers. However, as a consequence of the workload characteristics and the need of maintaining data coherent among the servers that compose the cluster, load imbalance arise among servers, reducing the efficiency of the server as a whole. We propose and evaluate a strategy for load balancing in stateful clustered servers. Our strategy is based on control theory and allowed significant gains over configurations that do not employ the load balancing strategy, reducing the response time in up to 50% and increasing the throughput in up to 16%.
George Teodoro, Tulio Tavares, Bruno Coutinho, Wagner Meira Jr., Dorgival O. Guedes
SBAC-PAD1