EDBT 2026 Demo / reviewers in the wild / expert
Renato Ferreira 0001
dblp:f/RACFerreira · also Renato A. C. Ferreira, Renato Antônio Celso Ferreira, Renato C. Ferreira
· DBLP profile ↗
61ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-4372-8996ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 1 since 2021Databases, data management, data science and information retrieval · 5Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | STEval: A framework for evaluating spatio-temporal crime prediction models
Gabriel Amarante, Matheus Pimenta, Yan Andrade, Matheus Senna, Rainer Menezes, Antônio Hot Faria, Marcelo Vilas-Boas, Frederico Martins de Paula Neto, João Paulo da Silva, Everton Renato de Sousa, Jamicel da Silva, Wagner Meira Jr., George Teodoro, Leonardo Rocha 0001, Renato Ferreira 0001 |
Eng. Appl. Artif. Intell. | 15 |
| 2025 | IMI-GPU: Inverted multi-index for billion-scale approximate nearest neighbor search with GPUs
Alan Araujo, Willian de Oliveira Barreiros Junior, Jun Kong 0002, Renato Ferreira 0001, George Teodoro |
J. Parallel Distributed Comput. | 4 |
| 2025 | The Megapixel Approach for Efficient Execution of Irregular Wavefront Algorithms on GPUs
Mathias Oliveira, Willian de Oliveira Barreiros Junior, Renato Ferreira 0001, Alba Cristina Magalhaes Alves de Melo, George Teodoro |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2024 | A Descriptive and Predictive Analysis Tool for Criminal Data: A Case Study from Brazil
Yan Andrade, Matheus Pimenta, Gabriel Amarante, Antônio Hot Faria, Marcelo Vilas-Boas, João Paulo da Silva, Felipe Rocha, Jamicel da Silva, Wagner Meira Jr., George Teodoro, Leonardo Rocha 0001, Renato Ferreira 0001 |
ICCSA (2) | 12 |
| 2024 | Serious Games for Children with Autism Spectrum Disorder: A Systematic Literature ReviewabstractAutism Spectrum Disorder (ASD) prevalence rates are increasing and serious games have shown a valuable potential to aid the treatment of autistic individuals. Hence, a Systematic Literature Review was conducted aiming at categorizing serious games for ASD children regarding which skills they aim to develop, how their activities were operationalized, and which customization options they provide to users. Our results showed that a large number of serious games aimed at developing distinct skills in ASD children have been proposed, with their main focus being on social and socio-emotional skills. Nonetheless, for each skill we characterize the existing games according to their features, that is the platform they are developed for, I/O devices used, required users’ action, audiovisual elements, and number of players. We also identify strategies adopted in the games regarding specific features and skills that emerged from the analysis. Finally, our results highlight that offering broader customization options in the games could expand their applicability and utility to their users. This work provides a thorough examination of research on games for ASD children, contributing both to researchers interested in the topic, by identifying existing contributions and open issues for research, and to professionals interested in developing serious games for this public. Ana Paula de Carvalho, Camila Santana Braz, Sibele M. dos Santos, Renato Ferreira 0001, Raquel Oliveira Prates |
Int. J. Hum. Comput. Interact. | 4 |
| 2023 | Effective and efficient active learning for deep learning-based tissue image analysisabstractMOTIVATION: Deep learning attained excellent results in digital pathology recently. A challenge with its use is that high quality, representative training datasets are required to build robust models. Data annotation in the domain is labor intensive and demands substantial time commitment from expert pathologists. Active learning (AL) is a strategy to minimize annotation. The goal is to select samples from the pool of unlabeled data for annotation that improves model accuracy. However, AL is a very compute demanding approach. The benefits for model learning may vary according to the strategy used, and it may be hard for a domain specialist to fine tune the solution without an integrated interface. RESULTS: We developed a framework that includes a friendly user interface along with run-time optimizations to reduce annotation and execution time in AL in digital pathology. Our solution implements several AL strategies along with our diversity-aware data acquisition (DADA) acquisition function, which enforces data diversity to improve the prediction performance of a model. In this work, we employed a model simplification strategy [Network Auto-Reduction (NAR)] that significantly improves AL execution time when coupled with DADA. NAR produces less compute demanding models, which replace the target models during the AL process to reduce processing demands. An evaluation with a tumor-infiltrating lymphocytes classification application shows that: (i) DADA attains superior performance compared to state-of-the-art AL strategies for different convolutional neural networks (CNNs), (ii) NAR improves the AL execution time by up to 4.3×, and (iii) target models trained with patches/data selected by the NAR reduced versions achieve similar or superior classification quality to using target CNNs for data selection. AVAILABILITY AND IMPLEMENTATION: Source code: https://github.com/alsmeirelles/DADA. André L. S. Meirelles, Tahsin M. Kurç, Jun Kong 0002, Renato Ferreira 0001, Joel H. Saltz, George Teodoro |
Bioinform. | 4 |
| 2023 | Spatial-aware data partition for distributed memory parallelization of ANN search in multimedia retrieval
Guilherme Andrade, Renato Ferreira 0001, George Teodoro |
Parallel Comput. | 2 |
| 2022 | Efficient microscopy image analysis on CPU-GPU systems with cost-aware irregular data partitioning
Willian de Oliveira Barreiros Junior, Alba Cristina Magalhaes Alves de Melo, Jun Kong 0002, Renato Ferreira 0001, Tahsin M. Kurç, Joel H. Saltz, George Teodoro |
J. Parallel Distributed Comput. | 4 |
| 2021 | Online multimedia retrieval on CPU-GPU platforms with adaptive work partition
Rafael Souza, André Fernandes, Thiago S. F. X. Teixeira, George Teodoro, Renato Ferreira 0001 |
J. Parallel Distributed Comput. | 5 |
| 2020 | Scalable and Efficient Spatial-Aware Parallelization Strategies for Multimedia RetrievalabstractSimilarity search is a key operation in several multimedia applications, including online Content-Based Multimedia Retrieval (CBMR) services. These applications have to deal with very large databases and are submitted to high query rates. In this context, scalability in distributed memory system is critical to assemble the required computing power and memory space. However, we have identified that the Data Equal Split (DES) parallelization and associated data partition strategy employed by the related works on the domain have limitations in terms of efficiency and scalability. Therefore, in this paper, we developed and implemented a framework for similarity search execution on distributed memory machines and proposed a novel class of data partition strategies that takes into account the data spatial organization in its distribution. This approach leads to a reduction in communication traffic and in costs associated with processing each task in local searches carried out in the distributed machine. Our approach attained a speedup of 2.4× on top of DES in the baseline case (5 nodes) and also achieves higher scalability efficiency and is 14.5× faster when 160 nodes are used. In fact, our novel data organization led to superlinear scalability in all configurations evaluated. Guilherme Andrade, George Teodoro, Renato Ferreira 0001 |
SBAC-PAD | 3 |
| 2019 | Large-scale parallel similarity search with Product Quantization for online multimedia services
Guilherme Andrade, André Fernandes, Jeremias M. Gomes, Renato Ferreira 0001, George Teodoro |
J. Parallel Distributed Comput. | 4 |
| 2018 | Learning to Rank with Deep Autoencoder FeaturesabstractLearning to rank in Information Retrieval is the problem of learning the full order of a set of documents from their partially observed order. Datasets used by learning to rank algorithms are growing enormously in terms of number of features, but it remains costly and laborious to reliably label large datasets. This paper is about learning feature transformations using inexpensive unlabeled data and available labeled data, that is, building alternate features so that it becomes easier for existing learning to rank algorithms to find better ranking models from labeled datasets that are limited in size and quality. Deep autoencoders have proven powerful as nonlinear feature extractors, and thus we exploit deep autoencoder features for semi-supervised learning to rank. Typical approaches for learning autoencoder features are based on updating model parameters using either unlabeled data only, or unlabeled data first and then labeled data. We propose a novel approach which updates model parameters using unlabeled and labeled data simultaneously, enabling label propagation from labeled to unlabeled data. We present a comprehensive study on how deep autoencoder features improve the ranking performance of representative learning to rank algorithms, revealing the importance of building an effective feature set to describe the input data. Alberto Albuquerque, Tiago Amador, Renato Ferreira 0001, Adriano Veloso, Nivio Ziviani |
IJCNN | 3 |
| 2018 | From Java to FPGA: An Experience with the Intel HARP SystemabstractRecent years have seen a surge in the popularity of Field-Programmable Gate Arrays (FPGAs). Programmers can use them to develop high-performance systems that are not only efficient in time, but also in energy. Yet, programming FPGAs remains a difficult task. Even though there exist today OpenCL interfaces to synthesize such hardware, higher-level programming languages, such as Java, C# or Python remain distant from them. In this paper, we describe a compiler, and its supporting runtime environment, that reduces this distance, translating functional code written in Java to the Intel HARP platform. Thus, we bring two contributions. First, the insight that a functional-style library is a good starting point to bridge the gap between high-level programming idioms and FPGAs. Second, the implementation of this system itself, including the compiler, its intermediate representation, and all the runtime support necessary to shield developers from the task of transferring data back and forth between the host CPU and the accelerator. To demonstrate the effectiveness of our system, we have used it to implement different benchmarks, used in image processing and data-mining. For large inputs, we can observe consistent 20x speedups over the Java Virtual Machine across all our benchmarks. Depending on the target function that we compile, this speedup can achieve 280x. Pedro Caldeira, Jeronimo Costa Penha, Lucas B. da Silva, Ricardo S. Ferreira 0001, José A. M. Nacif, Renato Ferreira 0001, Fernando Magno Quintão Pereira |
SBAC-PAD | 6 |
| 2018 | A 3D modeling methodology based on a concavity-aware geometric test to create 3D textured coarse models from concept art and orthographic projections
Sergio N. Silva Junior, Felipe C. Chamone, Renato Ferreira 0001, Erickson R. Nascimento |
Comput. Graph. | 3 |
| 2017 | Online Multimedia Similarity Search with Response Time-Aware Parallelism and Task Granularity Auto-TuningabstractThis paper presents an efficient parallel implementation of the Product Quantization based approximate nearest neighbor multimedia similarity search indexing (PQANNS). The parallel PQANNS efficiently answers nearest neighbor queries by exploiting the ability of the quantization approach to reduce the data dimensionality (and memory demand) and by leveraging parallelism to speed up the search capabilities of the application. Our solution is also optimized to minimize query response times under scenarios with fluctuating query rates (load) as observed in online services. To achieve this goal, we have developed strategies to dynamically select the parallelism configuration and task granularity that minimizes the query response times during the execution. The proposed strategies (ADAPT and ADAPT+G) were thoroughly evaluated and have shown, for instance, to reduce the query response times in 6.4× as compared to the best static configuration of parallelism and task granularity. Guilherme Andrade, George Teodoro, Renato Ferreira 0001 |
SBAC-PAD | 3 |
| 2017 | Exploring Heterogeneous Mobile Architectures with a High-Level Programming ModelabstractThe development of new technologies is setting a new era characterized, among other factors, by the rise of sophisticated mobile devices containing CPUs and GPUs. This emerging scenario of heterogeneous mobile architectures brings challenging issues regarding the use of the available computing resources. Such issues are mainly related to the intrinsic complexity of coordinating these processors in order to increase application performance. In this sense, this paper presents a high-level programming model to implement parallel patterns that can be executed in a coordinate way by heterogeneous mobile architectures. A comparative analysis of performance and programming complexity is presented, contrasting code generated automatically from the proposed programming model with low-level manually-optimized implementations. Wilson de Carvalho Moreira, Guilherme Andrade, Pedro Henrique Moreira Caldeira, Renato Utsch Goncalves, Renato Ferreira 0001, Leonardo Rocha 0001, Renan de Carvalho Sousa, Millas Nasser Ramsses Avelar |
SBAC-PAD | 5 |
| 2016 | Connecting Opinions to Opinion-Leaders: A Case Study on Brazilian Political ProtestsabstractSocial media applications have assumed an important role in decision-making process of users, affecting their choices about products and services. In this context, understanding and modeling opinions, as well as opinion-leaders, have implications for several tasks, such as recommendation, advertising, brand evaluation etc. Despite the intrinsic relation between opinions and opinion-leaders, most recent works focus exclusively on either understanding the opinions, by Sentiment Analysis (SA) proposals, or identifying opinion-leaders using Influential Users Detection (IUD). This paper presents a preliminary evaluation about a combined analysis of SA and IUD. In this sense, we propose a methodology to quantify factors in real domains that may affect such analysis, as well as the potential benefits of combining SA Methods with IUD ones. Empirical assessments on a sample of tweets about the Brazilian president reveal that the collective opinion and the set of top opinion-leaders over time are inter-related. Further, we were able to identify distinct characteristics of opinion propagation, and that the collective opinion may be accurately estimated by using a few top-k opinion-leaders. These results point out the combined analysis of SA and IUD as a promising research direction to be further exploited. Leonardo Rocha 0001, Fernando Mourão, Ramon Vieira, Alan Neves, Dárlinton Barbosa Feres Carvalho, Bortik Bandyopadhyay, Srinivasan Parthasarathy 0001, Renato Ferreira 0001 |
DSAA | 8 |
| 2016 | ParallelME: A Parallel Mobile Engine to Explore Heterogeneity in Mobile Computing Architectures
Guilherme Andrade, Wilson de Carvalho, Renato Utsch, Pedro Caldeira, Alberto Albuquerque, Fabricio Ferracioli, Leonardo Rocha 0001, Michael Frank 0008, Dorgival O. Guedes, Renato Ferreira 0001 |
Euro-Par | 10 |
| 2016 | Watershed-ng: an extensible distributed stream processing frameworkabstractSummary Most high‐performance data processing (a.k.a. big data) systems allow users to express their computation using abstractions (like MapReduce), which simplify the extraction of parallelism from applications. Most frameworks, however, do not allow users to specify how communication must take place: That element is deeply embedded into the run‐time system abstractions, making changes hard to implement. In this work, we describe Wathershed‐ng, our re‐engineering of the Watershed system, a framework based on the filter–stream paradigm and originally focused on continuous stream processing. Like other big‐data environments, Watershed provided object‐oriented abstractions to express computation (filters), but the implementation of streams was a run‐time system element. By isolating stream functionality into appropriate classes, combination of communication patterns and reuse of common message handling functions (like compression and blocking) become possible. The new architecture even allows the design of new communication patterns, for example, allowing users to choose MPI, TCP, or shared memory implementations of communication channels as their problem demands. Applications designed for the new interface showed reductions in code size on the order of 50%and above in some cases. The performance results also showed significant improvements, because some implementation bottlenecks were removed in the re‐engineering process. Copyright © 2016 John Wiley & Sons, Ltd. Rodrigo Caetano Rocha, Bruno Hott, Vinícius Vitor dos Santos Dias, Renato Ferreira 0001, Wagner Meira Jr., Dorgival O. Guedes |
Concurr. Comput. Pract. Exp. | 4 |
| 2015 | SACI: Sentiment analysis by collective inspection on social media content
Leonardo Rocha 0001, Fernando Mourão, Thiago Silveira, Rodrigo Chaves, Giovanni Sá, Felipe Teixeira, Ramon Vieira, Renato Ferreira 0001 |
J. Web Semant. | 8 |
| 2014 | Scalable Feature Extraction for Visual Surveillance
Antonio C. Nazare, Renato Ferreira 0001, William Robson Schwartz |
CIARP | 2 |
| 2014 | Efficient Execution of Microscopy Image Analysis on CPU, GPU, and MIC Equipped Cluster SystemsabstractHigh performance computing is experiencing a major paradigm shift with the introduction of accelerators, such as graphics processing units (GPUs) and Intel Xeon Phi (MIC). These processors have made available a tremendous computing power at low cost, and are transforming machines into hybrid systems equipped with CPUs and accelerators. Although these systems can deliver a very high peak performance, making full use of its resources in real-world applications is a complex problem. Most current applications deployed to these machines are still being executed in a single processor, leaving other devices underutilized. In this paper we explore a scenario in which applications are composed of hierarchical data flow tasks which are allocated to nodes of a distributed memory machine in coarse-grain, but each of them may be composed of several finer-grain tasks which can be allocated to different devices within the node. We propose and implement novel performance aware scheduling techniques that can be used to allocate tasks to devices. We evaluate our techniques using a pathology image analysis application used to investigate brain cancer morphology, and our experimental evaluation shows that the proposed scheduling strategies significantly outperforms other efficient scheduling techniques, such as Heterogeneous Earliest Finish Time - HEFT, in cooperative executions using CPUs, GPUs, and MICs. We also experimentally show that our strategies are less sensitive to inaccuracy in the scheduling input data and that the performance gains are maintained as the application scales. Guilherme Andrade, Renato Ferreira 0001, George Teodoro, Leonardo Rocha 0001, Joel H. Saltz, Tahsin M. Kurç |
SBAC-PAD | 2 |
| 2014 | Economically-efficient sentiment stream analysisabstractText-based social media channels, such as Twitter, produce torrents of opinionated data about the most diverse topics and entities. The analysis of such data (aka. sentiment analysis) is quickly becoming a key feature in recommender systems and search engines. A prominent approach to sentiment analysis is based on the application of classification techniques, that is, content is classified according to the attitude of the writer. A major challenge, however, is that Twitter follows the data stream model, and thus classifiers must operate with limited resources, including labeled data and time for building classification models. Also challenging is the fact that sentiment distribution may change as the stream evolves. In this paper we address these challenges by proposing algorithms that select relevant training instances at each time step, so that training sets are kept small while providing to the classifier the capabilities to suit itself to, and to recover itself from, different types of sentiment drifts. Simultaneously providing capabilities to the classifier, however, is a conflicting-objective problem, and our proposed algorithms employ basic notions of Economics in order to balance both capabilities. We performed the analysis of events that reverberated on Twitter, and the comparison against the state-of-the-art reveals improvements both in terms of error reduction (up to 14%) and reduction of training resources (by orders of magnitude). Roberto L. de Oliveira Jr., Adriano Veloso, Adriano C. M. Pereira, Wagner Meira Jr., Renato Ferreira 0001, Srinivasan Parthasarathy 0001 |
SIGIR | 5 |
| 2014 | Smart surveillance framework: A versatile tool for video analysisabstractComputer Vision problems applied to visual surveillance have been studied for several years aiming at finding accurate and efficient solutions, required to allow the execution of surveillance systems in real environments. The main goal of such systems is to analyze the scene focusing on the detection and recognition of suspicious activities performed by humans in the scene, so that the security personnel can pay closer attention to these preselected activities. To accomplish that, several problems have to be solved first, for instance background subtraction, person detection, tracking and re-identification, face recognition, and action recognition. Even though each of these problems have been researched in the past decades, they are hardly considered in a sequence, each one is usually solved individually. However, in a real surveillance scenarios, the aforementioned problems have to be solved in sequence considering only videos as the input. Aiming at the direction of evaluating approaches in more realistic scenarios, this work proposes a framework called Smart Surveillance Framework (SSF), to allow researchers to implement their solutions to the above problems as a sequence of processing modules that communicate through a shared memory. Antonio C. Nazare, Cassio E. dos Santos, Renato Ferreira 0001, William Robson Schwartz |
WACV | 3 |
| 2014 | Thread scheduling and memory coalescing for dynamic vectorization of SPMD workloads
Teo Milanez, Caroline Collange, Fernando Magno Quintão Pereira, Wagner Meira Jr., Renato Ferreira 0001 |
Parallel Comput. | 5 |
| 2013 | GPU-NB: A Fast CUDA-Based Implementation of Naïve BayesabstractThe advent of the Web 2.0 has given rise to an interesting phenomenon: there is currently much more data than what can be effectively analyzed without relying on sophisticated automatic tools. Some of these tools, which target the organization and extraction of useful knowledge from this huge amount of data, rely on machine learning and data or text mining techniques, specifically automatic document classification algorithms. However, these algorithms are still a computational challenge because of the volume of data that needs to be processed. Some of the strategies available to address this challenge are based on the parallelization of ADC algorithms. In this work, we present GPU-NB, a parallel version of one of the most widely used document classification algorithms, the Naïve Bayes, that uses graphics processing units (GPUs). In our evaluation using 6 different document collections, we show that the GPU-NB can maintain the same classification effectiveness (in most cases) while increasing the efficiency by up to 34x faster than its sequential version using CPU. GPU-NB is also up to 11x faster than a CPU-based parallel implementation of Naive Bayes running with 4 threads. Moreover, assuming an optimistic behavior of the CPU parallelization, GPU-NB should outperform the CPU-based implementation with up to 32 cores, at a small fraction of the cost. We also show that the efficiency of the GPU-NB parallelization is impacted by features of the document collections, particularly the number of classes, although the density of the collection (average number of occurrences of terms per document) has a significant impact as well. Guilherme Andrade, Felipe Viegas, Gabriel Spada Ramos, Jussara M. Almeida, Leonardo Rocha 0001, Marcos André Gonçalves, Renato Ferreira 0001 |
SBAC-PAD | 7 |
| 2013 | Automatic parallelization of canonical loops
Leonardo Luiz Padovani da Mata, Fernando Magno Quintão Pereira, Renato Ferreira 0001 |
Sci. Comput. Program. | 3 |
| 2012 | Data and Instruction Uniformity in Minimal Multi-threadingabstractSimultaneous Multi-Threading (SMT) is a hardware model in which different threads share the same instruction fetching unit. This model is a compromise between high parallelism and low hardware cost. Minimal Multi-Threading (MMT) is a technique recently proposed to share instructions and execution between threads in a SMT machine. In this paper we propose new ways to explore redundancies in the MMT execution model. First, we propose and evaluate a new thread reconvergence heuristics that handles function calls better than previous approaches. Second, we demonstrate the existence of substantial regularity in inter-thread memory access patterns. We validate our results on the four data-parallel applications present in the PARSEC benchmark suite. The new thread reconvergence heuristics is, on the average, 82% more efficient than MMT's original reconvergence method. Furthermore, about 69% to 87% of all the memory addresses are either the same for all the threads, or are affine expressions of the thread identifier. This observation motivates the design of newly proposed hardware that benefits from regularity in inter-thread memory accesses. Teo Milanez, Caroline Collange, Fernando Magno Quintão Pereira, Wagner Meira Jr., Renato Ferreira 0001 |
SBAC-PAD | 5 |
| 2011 | Speeding Up Learning in Real-Time Search through Parallel ComputingabstractReal-time search algorithms solve the problem of path planning, regardless the size and complexity of the maps, and the massive presence of entities in the same environment. In such methods, the learning step aims to avoid local minima and improve the results for future searches, ensuring the convergence to the optimal path when the same planning task is solved repeatedly. However, performing search in a limited area due to real-time constraints makes the run to convergence a lengthy process. In this work, we present a parallelization strategy that aims to reduce the time to convergence, maintaining the real-time properties of the search. The parallelization technique consists on using auxiliary searches without the real-time restrictions present in the main search. In addition, the same learning is shared by all searches. The empirical evaluation shows that even with the additional cost required to coordinate the auxiliary searches, the reduction in time to convergence is significant, showing gains from searches occurring in environments with fewer local minima to larger searches on complex maps, where performance improvement is even better. Vinícius Marques, Luiz Chaimowicz, Renato Ferreira 0001 |
SBAC-PAD | 3 |
| 2011 | Watershed: A High Performance Distributed Stream Processing SystemabstractThe task of extracting information from datasets that become larger at a daily basis, such as those collected from the web, is an increasing challenge, but also provides more interesting insights and analysis. Current analyses went beyond content and now focus on tracking and understanding users' relationships and interactions. Such computation is intensive both in terms of the processing demand imposed by the algorithms and also the sheer amount of data that has to handled. In this paper we introduce Watershed, a distributed computing framework designed to support the analysis of very large data streams online and in real-time. Data are obtained from streams by the system's processing components, transformed, and directed to other streams, creating large flows of information. The processing components are decoupled from each other and their connections are strictly data-driven. They can be dynamically inserted and removed, providing an environment in which it is feasible that different applications share intermediate results or cooperate to a global purpose. Our experiments demonstrate the flexibility in creating a set of data analysis algorithms and their composition into a powerful stream analysis environment. Thatyene Louise Alves de Souza Ramos, Rodrigo Silva Oliveira, Ana Paula de Carvalho, Renato Ferreira 0001, Wagner Meira Jr. |
SBAC-PAD | 4 |
| 2011 | Effective sentiment stream analysis with self-augmenting training and demand-driven projectionabstractHow do we analyze sentiments over a set of opinionated Twitter messages? This issue has been widely studied in recent years, with a prominent approach being based on the application of classification techniques. Basically, messages are classified according to the implicit attitude of the writer with respect to a query term. A major concern, however, is that Twitter (and other media channels) follows the data stream model, and thus the classifier must operate with limited resources, including labeled data for training classification models. This imposes serious challenges for current classification techniques, since they need to be constantly fed with fresh training messages, in order to track sentiment drift and to provide up-to-date sentiment analysis. Ismael S. Silva, Janaína Gomide, Adriano Veloso, Wagner Meira Jr., Renato Ferreira 0001 |
SIGIR | 5 |
| 2010 | Run-time optimizations for replicated dataflows on heterogeneous environmentsabstractThe increases in multi-core processor parallelism and in the flexibility of many-core accelerator processors, such as GPUs, have turned traditional SMP systems into hierarchical, heterogeneous computing environments. Fully exploiting these improvements in parallel system design remains an open problem. Moreover, most of the current tools for the development of parallel applications for hierarchical systems concentrate on the use of only a single processor type (e.g., accelerators) and do not coordinate several heterogeneous processors. Here, we show that making use of all of the heterogeneous computing resources can significantly improve application performance. Our approach, which consists of optimizing applications at run-time by efficiently coordinating application task execution on all available processing units is evaluated in the context of replicated dataflow applications. The proposed techniques were developed and implemented in an integrated run-time system targeting both intra- and inter-node parallelism. The experimental results with a real-world complex biomedical application show that our approach nearly doubles the performance of the GPU-only implementation on a distributed heterogeneous accelerator cluster. George Teodoro, Timothy D. R. Hartley, Ümit V. Çatalyürek, Renato Ferreira 0001 |
HPDC | 4 |
| 2010 | Tree Projection-Based Frequent Itemset Mining on Multicore CPUs and GPUsabstractFrequent itemset mining (FIM) is a core operation for several data mining applications as association rules computation, correlations, document classification, and many others, which has been extensively studied over the last decades. Moreover, databases are becoming increasingly larger, thus requiring a higher computing power to mine them in reasonable time. At the same time, the advances in high performance computing platforms are transforming them into hierarchical parallel environments equipped with multi-core processors and many-core accelerators, such as GPUs. Thus, fully exploiting these systems to perform FIM tasks poses as a challenging and critical problem that we address in this paper. We present efficient multi-core and GPU accelerated parallelizations of the Tree Projection, one of the most competitive FIM algorithms. The experimental results show that our Tree Projection implementation scales almost linearly in a CPU shared-memory environment after careful optimizations, while the GPU versions are up to 173 times faster than standard the CPU version. George Teodoro, Nathan Mariano, Wagner Meira Jr., Renato Ferreira 0001 |
SBAC-PAD | 4 |
| 2009 | Coordinating the use of GPU and CPU for improving performance of compute intensive applicationsabstractGPUs have recently evolved into very fast parallel co-processors capable of executing general purpose computations extremely efficiently. At the same time, multi-core CPUs evolution continued and today's CPUs have 4-8 cores. These two trends, however, have followed independent paths in the sense that we are aware of very few works that consider both devices cooperating to solve general computations. In this paper we investigate the coordinated use of CPU and GPU to improve efficiency of applications even further than using either device independently. We use Anthill runtime environment, a data-flow oriented framework in which applications are decomposed into a set of event-driven filters, where for each event, the runtime system can use either GPU or CPU for its processing. For evaluation, we use a histopathology application that uses image analysis techniques to classify tumor images for neuroblas-toma prognosis. Our experimental environment includes dual and octa-core machines, augmented with GPUs and we evaluate our approach's performance for standalone and distributed executions. Our experiments show that a pure GPU optimization of the application achieved a factor of 15 to 49 times improvement over the single core CPU version, depending on the versions of the CPUs and GPUs. We also show that the execution can be further reduced by a factor of about 2 by using our runtime system that effectively choreographs the execution to run cooperatively both on GPU and on a single core of CPU. We improve on that by adding more cores, all of which were previously neglected or used ineffectively. In addition, the evaluation on a distributed environment has shown near linear scalability to multiple hosts. George Teodoro, Rafael Sachetto Oliveira, Olcay Sertel, Metin Nafi Gürcan, Wagner Meira Jr., Ümit V. Çatalyürek, Renato Ferreira 0001 |
CLUSTER | 7 |
| 2009 | Profiling General Purpose GPU ApplicationsabstractWe are witnessing an increasing adoption of GPUs for performing general purpose computation, which is usually known as GPGPU. The main challenge in developing such applications is that they often do not fit in the model required by the graphics processing devices, limiting the scope of applications that may be benefit from the computing power provided by GPUs. Even when the application fits GPU model, obtaining optimal resource usage is a complex task. In this work we propose a profiling tool for GPGPU applications. This tool use a profiling strategy based on performance predicates and is able to quantify the major sources of performance degradation while providing hints on how to improve the applications. We used our tool in CUDA programs and were able to understand and improve their performance. Bruno Coutinho, George Teodoro, Rafael Sachetto Oliveira, Dorgival O. Guedes, Renato Ferreira 0001 |
SBAC-PAD | 5 |
| 2009 | Exploiting Computational Resources in Distributed Heterogeneous PlatformsabstractWe have been witnessing a continuous growth of both heterogeneous computational platforms (e.g., Cell blades, or the joint use of traditional CPUs and GPUs) and multi- core processor architecture; and it is still an open question how applications can fully exploit such computational potential efficiently. In this paper we introduce a run-time environment and programming framework which supports the implementation of scalable and efficient parallel applications in such heterogeneous, distributed environments. We assess these issues through well-known kernels and actual applications that behave regularly and irregularly, which are not only relevant but also demanding in terms of computation and I/O. Moreover, the irregularity of these, as well as many other applications poses a challenge to the design and implementation of efficient parallel algorithms. Our experimental environment includes dual and octa-core machines augmented with GPUs and we evaluate our framework performance for standalone and distributed executions. The evaluation on a distributed environment has shown near to linear scale-ups for two data mining applications, while the applications performance, when using CPU and GPU, has been improved into around 25%, compared to the GPU-only versions. George Teodoro, Rafael Sachetto Oliveira, Daniel Fireman, Dorgival O. Guedes, Renato Ferreira 0001 |
SBAC-PAD | 5 |
| 2008 | Achieving Multi-Level Parallelism in the Filter-Labeled Stream Programming ModelabstractNew architectural trends in chip design resulted in machines with multiple processing units as well as efficient communication networks, leading to the wide availability of systems that provide multiple levels of parallelism, both inter- and intra-machine. Developing applications that efficiently make use of such systems is a challenge, specially for application-domain programmers. In this paper we present a new version of the Anthill programming environment that efficiently exploits multi-level parallelism and experimental results that demonstrate such efficiency. Anthill is based on the filter-stream model; in this model, applications are decomposed into a set of filters communicating through streams, which has already been shown to be efficient for expressing inter-machine parallelism. We replaced the filter run-time environment, originally process-oriented, with an event-oriented version. This new version allow programmers to efficiently express opportunities for parallelism within each compute node through a higher-level programming abstraction. We evaluated our solution on dual- and quad-core machines with two data mining applications: Eclat and KNN. Both had drops in execution time nearly proportional to the number of cores on a single machine. When using a cluster of dual-core machines, speed-ups were close to linear on the number of available cores for both applications, confirming event-oriented Anthill performs well both on the inter- and intra-machine parallelism levels. George Teodoro, Daniel Fireman, Dorgival O. Guedes, Wagner Meira Jr., Renato Ferreira 0001 |
ICPP | 5 |
| 2008 | Translational research design templates, Grid computing, and HPCabstractDesign templates that involve discovery, analysis, and integration of information resources commonly occur in many scientific research projects. In this paper we present examples of design templates from the biomedical translational research domain and discuss the requirements imposed on Grid middleware infrastructures by them. Using caGrid, which is a Grid middleware system based on the model driven architecture (MDA) and the service oriented architecture (SOA) paradigms, as a starting point, we discuss architecture directions for MDA and SOA based systems like caGrid to support common design templates. Joel H. Saltz, Scott Oster, Shannon Hastings, Stephen Langella, Renato Ferreira 0001, Justin Permar, Ashish Sharma 0001, David Ervin, Tony Pan, Ümit V. Çatalyürek, Tahsin M. Kurç |
IPDPS | 5 |
| 2008 | A Reconfigurable Run-Time System for Filter-Stream ApplicationsabstractThe development of high level abstractions for programming distributed systems is becoming a crucial effort in computer science. Several frameworks have been proposed, which expose simplified programming abstractions that are useful for a broad class of applications and can be implemented efficiently on distributed systems. One such system is Anthill, based on the filter-stream programming model, in which applications are decomposed into sets of independent filters that communicate via streams. Anthill achieves high performance by allowing filters to be transparently replicated across several compute nodes.In this paper we present a global state manager for Anthill, which exports a simple abstraction to manipulate state variables for application filters. The state is distributed transparently among the instances of that filter, and our manager is designed to allow data migration from one filter instance to another, enabling Anthill to dynamically reconfigure applications at execution time.To evaluate our system, we used two well known data mining algorithms: a priori and k-means. Our results show that the framework incurs low overhead, 1.8% on average, and that the resulting system can effectively make use of new resources as they are made available, with execution times 3.57% slower, on average, than the minimum expected time for the reconfiguration scenario. Daniel Fireman, George Teodoro, André Cardoso, Renato Ferreira 0001 |
SBAC-PAD | 4 |
| 2007 | An Efficient and Reliable Scientific Workflow SystemabstractThis paper presents a fault tolerance framework for applications that process data using a distributed network of user-defined operations in a pipelined fashion. The framework saves intermediate results and messages exchanged among application components in a distributed data management system to facilitate quick recovery from failures. The experimental results show that the framework scales well and our approach introduces very little overhead to application execution. Tulio Tavares, George Teodoro, Tahsin M. Kurç, Renato Ferreira 0001, Dorgival O. Guedes, Wagner Meira Jr., Ümit V. Çatalyürek, Shannon Hastings, Scott Oster, Stephen Langella, Joel H. Saltz |
CCGRID | 4 |
| 2007 | Fault-tolerance in filter-labeled-stream applicationsabstractFault tolerance is a desirable feature in distributed high-performance systems, since applications tend to run for long periods of time and faults become more likely as the number of nodes in the system increase. However, most distributed environments lack any fault tolerant features, since they tend to be hard to implement and use, and often hurt performance dramatically. In this paper we discuss how we successfully added fault-tolerance to the Anthill distributed programming environment by using an application-level checkpoint/rollback solution. The programming model offers an abstraction where the programmer can easily identify points during the execution where the communication pattern is well defined, forming a consistent cut where checkpoints may be saved consistently without requiring extra communication, avoiding any domino effect during recovery from faults. We present the new abstractions for fault tolerance, describe how the solution was implemented and present performance results that show the efficiency of the solution with both regular and irregular applications. Bruno Coutinho, Dorgival O. Guedes, Wagner Meira Jr., Renato Ferreira 0001 |
SBAC-PAD | 4 |
| 2007 | A Scalable Parallel Deduplication AlgorithmabstractThe identification of replicas in a database is fundamental to improve the quality of the information. Deduplication is the task of identifying replicas in a database that refer to the same real world entity. This process is not always trivial, because data may be corrupted during their gathering, storing or even manipulation. Problems such as misspelled names, data truncation, data input in a wrong format, lack of conventions (like how to abbreviate a name), missing data or even fraud may lead to the insertion of replicas in a database. The deduplication process may be very hard, if not impossible, to be performed manually, since actual databases may have hundreds of millions of records. In this paper, we present our parallel deduplication algorithm, called FER- APARDA. By using probabilistic record linkage, we were able to successfully detect replicas in synthetic datasets with more than 1 million records in about 7 minutes using a 20- computer cluster, achieving an almost linear speedup. We believe that our results do not have similar in the literature when it comes to the size of the data set and the processing time. Walter Santos, Thiago Teixeira, Carla Machado, Wagner Meira Jr., Renato Ferreira 0001, Dorgival O. Guedes, Altigran S. da Silva |
SBAC-PAD | 5 |
| 2006 | Assessing Data Virtualization for Irregularly Replicated Large DatasetsabstractLarge volumes of data are generated every day by experiments, simulations and all sorts of applications. It is common to observe situations where portions of data are irregularly replicated and distributed in different data sources. It would be desirable to be able to handle these several pieces of irregular data (replicated or not) as a unique large dataset. This is called data virtualization and is the focus of this paper. In this paper, we present a system which is capable of dealing with irregularly replicated data and is able to create a virtual view of the union of the individual irregular portions of data hosted by each data source. Our system indexes the data intervals from each data source and allows clients to submit queries against the virtual dataset created. In order to select what server will be responsible for each data interval of a query, we use and compare three algorithms, namely Random, Round-Robin and Weighted Round-Robin. The comparison is driven by simulation and the parameters for the simulation are all taken from a real data-centered application (the Virtual Microscope). Bruno Diniz, Diego L. Nogueira, André Cardoso, Renato Ferreira 0001, Dorgival O. Guedes, Wagner Meira Jr. |
CCGRID | 4 |
| 2006 | ParTriCluster: A Scalable Parallel Algorithm for Gene Expression AnalysisabstractAnalyzing gene expression patterns is becoming a highly relevant task in the bio informatics area. This analysis makes it possible to determine the behavior patterns of genes under various conditions, a fundamental information for treating diseases, among other applications. An advance in this area is the tricluster algorithm, which is the first algorithm capable of determining 3D clusters, that is, it determines clusters of sets of genes that behave similarly in a set of samples and set of time stamps. However, while biological experiments collect an increasing amount of data to be analyzed and correlated, the triclustering problem is NP-complete, and its parallelization seems to be an essential step towards obtaining feasible solutions. In this paper we propose and evaluate the implementation of a parallel version of the tricluster algorithm using the filter-labeled-stream paradigm supported by the Anthill parallel programming environment. The results show that our parallelization scales linearly with the data size. Further, the parallelization strategy is applicable to any depth-first searches Renata Braga Araújo, Guilherme Henrique Trielli Ferreira, Gustavo Henrique Orair, Wagner Meira Jr., Renato Ferreira 0001, Dorgival O. Guedes |
SBAC-PAD | 5 |
| 2006 | A Run-time System for Efficient Execution of Scientific Workflows on Distributed EnvironmentsabstractScientific workflow systems have been introduced in response to the demand of researchers from several domains of science who need to process and analyze increasingly larger datasets. The design of these systems is largely based on the observation that data analysis applications can be composed as pipelines or networks of computations on data. In this paper we present a run-time support system that is designed to facilitate this type of computation in distributed computing environments. Our system is optimized for data-intensive workflows, in which efficient management and retrieval of data, coordination of data processing and data movement, and check-pointing of intermediate results are critical and challenging issues. Experimental evaluation of our system shows that linear speedups can be achieved for sophisticated applications, which are implemented as a network of multiple data processing components George Teodoro, Tulio Tavares, Renato Ferreira 0001, Tahsin M. Kurç, Wagner Meira Jr., Dorgival O. Guedes, Tony Pan, Joel H. Saltz |
SBAC-PAD | 3 |
| 2005 | Scheduling Data Flow Applications Using Linear ProgrammingabstractGrid environments are becoming cost-effective substitutes to supercomputers. Datacutter is one of several initiatives in creating mechanisms for applications to efficiently exploit the vast computation power of such environments. In Datacutter, applications are modeled as a set of communicating filters that may run on several nodes of a computational grid. To achieve high performance, a number of transparent copies of each of the filters that comprise the application need to be appropriately placed on different nodes of the grid. Such task is carried out by a scheduler which is the focus of this work. We present LPSched, a scheduler for Datacutter applications which uses linear programming to make decisions about the number of copies of each filter as well as the placement of each of the copies across the nodes. LPSched bases its decisions upon the performance behavior of each filter as well as the resources currently available on the grid. Luiz Thomaz do Nascimento, Renato Ferreira 0001, Wagner Meira Jr., Dorgival O. Guedes |
ICPP | 2 |
| 2005 | AnthillSched: A Scheduling Strategy for Irregular and Iterative I/O-Intensive Parallel Jobs
Fabrício Góes, Pedro Henrique Calais Guerra, Bruno Coutinho, Leonardo Rocha 0001, Wagner Meira Jr., Renato Ferreira 0001, Dorgival O. Guedes, Walfredo Cirne |
JSSPP | 6 |
| 2005 | Anthill: A Scalable Run-Time Environment for Data Mining ApplicationsabstractData mining techniques are becoming increasingly more popular as a reasonable means to collect summaries from the rapidly growing datasets in many areas. However, as the size of the raw data increases, parallel data mining algorithms are becoming a necessity. In this paper, we present a run-time support system that was designed to allow the efficient implementation of data-mining algorithms on heterogeneous distributed environments. We believe that the runtime framework is suitable for a broader class of applications, beyond data mining. We also present a parallelization strategy that is supported by the run-time system. We show scalability results of three different data-mining algorithms that were parallelized using our approach and our run-time support. All applications scale almost linearly up to a large number of nodes. Renato Ferreira 0001, Wagner Meira Jr., Dorgival O. Guedes, Lúcia M. A. Drummond, Bruno Coutinho, George Teodoro, Tulio Tavares, Renata Braga Araújo, Guilherme T. Ferreira |
SBAC-PAD | 1 |
| 2004 | Asynchronous and Anticipatory Filter-Stream Based Parallel Algorithm for Frequent Itemset Mining
Adriano Veloso, Wagner Meira Jr., Renato Ferreira 0001, Dorgival O. Guedes, Srinivasan Parthasarathy 0001 |
PKDD | 3 |
| 2003 | Compiler support for efficient processing of XML datasetsabstractDeclarative, high-level, and/or application-class specific languages are often successful in easing application development. In this paper, we report our experiences in compiling a recently developed XML Query Language, XQuery for applications that process scientific datasets.Though scientific data processing applications can be conveniently represented in XQuery, compiling them to achieve efficient execution involves a number of challenges. These are, 1) analysis of recursive functions to identify reduction computations involving only associative and commutative operations, 2) replacement of recursive functions with iterative constructs, 3) parallelization of generalized reduction functions, which particularly requires the synthesis of global reduction functions, 4) application of data-centric transformations on the structure of XQuery, and 5) translation of XQuery processing to an imperative language like C/C++, which is required for using a middleware that offers low-level functionality.This paper describes our solutions towards these problems. By implementing the techniques in a compiler and generating code for a runtime system called Active Data Repository (ADR), we are able to achieve efficient processing of disk-resident datasets and parallelization on a cluster of machines. Our experimental results show that: 1) restructuring transformations, i.e. removing recursion and applying data-centric execution, result in several-folds improvement in performance, and 2) parallel versions achieve good load-balance, and incur no significant overheads besides communication. Xiaogang Li 0001, Renato Ferreira 0001, Gagan Agrawal |
ICS | 2 |
| 2003 | Compiler Support for Exploiting Coarse-Grained Pipelined ParallelismabstractThe emergence of grid and a new class of data-driven applications is making a new form of parallelism desirable, which we refer to as coarse-grained pipelined parallelism. This paper reports on a compilation system developed to exploit this form of parallelism. We use a dialect of Java that exposes both pipelined and data parallelism to the compiler. Our compiler is responsible for selecting a set of candidate filter boundaries, determining the volume of communication required if a particular boundary is chosen, performing the decomposition, and generating code. We have developed a one-pass algorithm for determining the required communication between consecutive filters. We have developed a cost model for estimating the execution time for a given decomposition, and a dynamic programming algorithm for performing the decomposition. Detailed evaluation of our current compiler using four data-driven applications demonstrate the feasibility of our approach. 1. Renato Ferreira 0001, Gagan Agrawal |
SC | 2 |
| 2003 | Optimizing Reduction Computations In a Distributed EnvironmentabstractWe investigate runtime strategies for data-intensive applications that invovle generalized reductions on large, distributed datasets.Our set of strategies includes replicated filter state, partitioned filter state, and hybrid options between these two extremes.We evaluate these strategies using emulators of three real applications, different query and output sizes, and a number of configurations.We consider execution in a homogeneous cluster and in a distributed environment where only a subset of nodes hst the data.Our results show replicating the filter state scales well and outperforms other schemes, if sufficient memory is available and sufficient computation is involved to offset the cost of global merge step.In other cases, hybrid is usually the best.Moreover, in almost all cases, the performance of the hybrid strategy is quite close to the best strategy. Thus, we believe that hybrid is an attractive approach when the relative performance of different schemes cannot be predicted. Tahsin M. Kurç, Feng Lee, Gagan Agrawal, Ümit V. Çatalyürek, Renato Ferreira 0001, Joel H. Saltz |
SC | 5 |
| 2002 | Compiler supported high-level abstractions for sparse disk-resident datasetsabstractProcessing and analyzing large volumes of data plays an increasingly important role in many domains of scientific research. The complexity and irregularity of datasets in many domains make the task of developing such processing applications tedious and error-prone.We propose use of high-level abstractions for hiding the irregularities in these datasets and enabling rapid development of correct data processing applications. We present two execution strategies and a set of compiler analysis techniques for obtaining high performance from applications written using our proposed high-level abstractions. Our execution strategies achieve high locality in disk accesses. Once a disk block is read from the disk, all iterations that access any of the elements from this disk block are performed. To support our execution strategies and improve the performance, we have developed static analysis techniques for: 1) computing the set of iterations that access a particular right-hand-side element, 2) generating a function that can be applied to the meta-data associated with each disk block, for determining if that disk block needs to be read, and 3) performing code hoisting of conditionals.We present experimental results from a prototype compiler implementing our techniques to demonstrate the effectiveness of our approach. Renato Ferreira 0001, Gagan Agrawal, Joel H. Saltz |
ICS | 1 |
| 2002 | Executing multiple pipelined data analysis operations in the gridabstractProcessing of data in many data analysis applications can be represented as an acyclic, coarse grain data flow, from data sources to the client. This paper is concerned with scheduling of multiple data analysis operations, each of which is represented as a pipelined chain of processing on data. We define the scheduling problem for effectively placing components onto Grid resources, and propose two scheduling algorithms. Experimental results are presented using a visualization application. Matthew Spencer, Renato Ferreira 0001, Michael D. Beynon, Tahsin M. Kurç, Ümit V. Çatalyürek, Alan Sussman, Joel H. Saltz |
SC | 2 |
| 2002 | Processing large-scale multi-dimensional data in parallel and distributed environments
Michael D. Beynon, Chialin Chang, Ümit V. Çatalyürek, Tahsin M. Kurç, Alan Sussman, Henrique Andrade, Renato Ferreira 0001, Joel H. Saltz |
Parallel Comput. | 7 |
| 2002 | Data parallel language and compiler support for data intensive applications
Renato Ferreira 0001, Gagan Agrawal, Joel H. Saltz |
Parallel Comput. | 1 |
| 2000 | Compiling object-oriented data intensive applicationsabstractProcessing and analyzing large volumes of data plays an increasingly important role in many domains of scientific research. High-level language and compiler support for developing applications that analyze and process such datasets has, however, been lacking so far. In this paper, we present a set of language extensions and a prototype compiler for supporting high-level object-oriented programming of data intensive reduction operations over multidimensional data. We have chosen a dialect of Java with data-parallel extensions for specifying collection of objects, a parallel for loop, and reduction variables as our source high-level language. Our compiler analyzes parallel loops and optimizes the processing of datasets through the use of an existing run-time system, called Active Data Repository (ADR). We show how loop fission followed by interprocedural static program slicing can be used by the compiler to extract required information for the run-time system. We present the design of a co... Renato Ferreira 0001, Gagan Agrawal, Joel H. Saltz |
ICS | 1 |
| 1999 | A High-Performance Database System for Managing Large Multi-resolution Medical Images
Tahsin M. Kurç, Michael D. Beynon, Chialin Chang, Renato Ferreira 0001, Benjamin B. Bederson, Joel H. Saltz, Alan Sussman |
AMIA | 4 |
| 1999 | Querying Very Large Multi-dimensional Datasets in ADRabstractApplications that make use of very large scientific datasets have become an increasingly important subset of scientific applications.In these applications, datasets are often multi-dimensional, i.e., data items are associated with points in a multi-dimensional attribute space, and access to data items is described by range queries.The basic processing involves mapping input data items to output data items, and some form of aggregation of all the input data items that project to the each output data item.We have developed an infrastructure, called the Active Data Repository (ADR), that integrates storage, retrieval and processing of multi-dimensional datasets on distributed-memory parallel architectures with multiple disks attached to each node.In this paper we address efficient execution of range queries on distributed memory parallel machines within ADR framework.We present three potential strategies, and evaluate them under different application scenarios and machine configurations.We present experimental results on the scalability and performance of the strategies on a 128-node IBM SP. Tahsin M. Kurç, Chialin Chang, Renato Ferreira 0001, Alan Sussman, Joel H. Saltz |
SC | 3 |
| 1998 | Digital dynamic telepathology-the Virtual Microscope
Asmara Afework, Michael D. Beynon, Fabián E. Bustamante, Soon Cho, Angelo Demarzo, Renato Ferreira 0001, Mark Silberman, Joel H. Saltz, Alan Sussman, Hubert Tsang |
AMIA | 6 |
| 1997 | The Virtual Microscope
Renato Ferreira 0001, Bongki Moon, Jim Humphries, Alan Sussman, Joel H. Saltz, Angelo Demarzo |
AMIA | 1 |