Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Luiz Rodolpho Monnerat

dblp:m/LuizRodolphoMonnerat · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
0since 2021 · last 2015
0000-0001-7146-0731ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 first-authorComputer networks · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 94% Parallel and multicore computing · 6%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › cache coherence › coherence protocol optimization
adaptive coherence protocol
0.011998
Efficiently Adapting to Sharing Patterns in Software DSMs · HPCA 1998
Memory systems
cache coherence
0.011998
Efficiently Adapting to Sharing Patterns in Software DSMs · HPCA 1998
Memory systems › shared memory
distributed shared memory
0.011998
Efficiently Adapting to Sharing Patterns in Software DSMs · HPCA 1998
Memory systems › memory consistency › memory consistency model › release consistency
lazy release consistency
0.011998
Efficiently Adapting to Sharing Patterns in Software DSMs · HPCA 1998
Parallel and multicore computing
parallel programming models
0.011998
Efficiently Adapting to Sharing Patterns in Software DSMs · HPCA 1998
Memory systems › shared memory › distributed shared memory
software distributed shared memory
0.011998
Efficiently Adapting to Sharing Patterns in Software DSMs · HPCA 1998

Methods — techniques the papers use, named apart from their topics

single-writer mode · 0.0multiple-writer mode · 0.0dynamic categorization · 0.0
YearPublicationVenuePosition
2015 An effective single-hop distributed hash table with high lookup performance and low traffic overhead
abstract
SUMMARY Distributed hash tables (DHTs) have been used in several applications, but most DHTs have opted to solve lookups with multiple hops, to minimize bandwidth costs while sacrificing lookup latency. This paper presents D1HT, an original DHT that has a peer‐to‐peer and self‐organizing architecture and maximizes lookup performance with reasonable maintenance traffic, and a Quarantine mechanism to reduce overheads caused by volatile peers. We implemented both D1HT and a prominent single‐hop DHT, and we performed an extensive and highly representative DHT experimental comparison, followed by complementary analytical studies. In comparison with current single‐hop DHTs, our results showed that D1HT consistently had the lowest bandwidth requirements, with typical reductions of up to one order of magnitude, and that D1HT could be used even in popular Internet applications with millions of users. In addition, we ran the first latency experiments comparing DHTs to directory servers, which revealed that D1HT can achieve latencies equivalent to or better than a directory server, and confirmed its greater scalability properties. Overall, our extensive set of results allowed us to conclude that D1HT can provide a very effective solution for a broad range of environments, from large‐scale corporate data centers to widely deployed Internet applications. Copyright © 2014 John Wiley & Sons, Ltd.
Luiz Rodolpho Monnerat, Claudio Luis de Amorim
Concurr. Comput. Pract. Exp.1
2009 Peer-to-Peer Single Hop Distributed Hash Tables
abstract
Efficiently locating information in large-scale distributed systems is a challenging problem to which Peer-to-Peer (P2P) Distributed Hash Tables (DHTs) can provide a highly scalable and cost-effective solution. However, there is very little experience on using DHTs in performance sensitive environments such as High Performance Computing (HPC) datacenters, and there is no published experimental comparison among low-latency DHTs. To fill this gap, we conducted an in-depth performance comparison of three proposed low-latency single-hop DHTs namely 1h-Calot, D1HT, and OneHop. Specifically, we compared experimentally the lookup latency and CPU use of D1HT with those of 1h-Calot by running each of them concurrently with the normal workload production for a subset of 1,800 nodes of a heavy-loaded HPC datacenter. In addition, we carried out an analytical performance comparison among the three single-hop DHTs for system sizes of up to 10 million nodes. The results showed that D1HT consistently had the smallest overhead and in most cases it required one order of magnitude less bandwidth than 1h-Calot and OneHop. Overall, the combination of our experimental and analytical results suggests that D1HT can provide a very effective solution for a broad range of environments, from large-scale HPC datacenters to widely deployed Internet P2P applications such as BitTorrent with up to one million peers. This ability to support such a wide range of environments may allow D1HT to be used as an inexpensive and scalable commodity software substrate for large-scale distributed applications.
Luiz Rodolpho Monnerat, Claudio Luis de Amorim
GLOBECOM1
2009 Accelerating Kirchhoff Migration by CPU and GPU Cooperation
abstract
We discuss the performance of Petrobras production Kirchhoff prestack seismic migration on a cluster of 64 GPUs and 256 CPU cores. Porting and optimization of the application hot spot (98.2% of a single CPU core execution time) to a single GPU reduces total execution time by a factor of 36 on a control run. We then argue against the usual practice of porting the next hot spot (1.5% of single CPU core execution time) to the GPU. Instead, we show that cooperation of CPU and GPU reduces total execution time by a factor of 59 on the same control run. Remaining GPU idle cycles are eliminated by overloading the GPU with multiple requests originated from distinct CPU cores. However, increasing the number of CPU cores in the computation reduces the gain due to the combination of enhanced parallelism in the runs without GPUs and GPU saturation on runs with GPUs. We proceed by obtaining close to perfect speed-up on the full cluster over homogeneous load obtained by replicating control run data. To cope with the heterogeneous load of real world data we show a dynamic load balancing scheme that reduces total execution time by a factor of 20 on runs that use all GPUs and half of the cluster CPU cores with respect to runs that use all CPU cores but no GPU.
Jairo Panetta, Thiago Teixeira, Paulo R. P. de Souza Filho, Carlos A. da Cunha Filho, David Sotelo, Fernando M. Roxo da Motta, Silvio Sinedino Pinheiro, Ivan Pedrosa Junior, Andre L. Romanelli Rosa, Luiz Rodolpho Monnerat, Leandro T. Carneiro, Carlos H. B. de Albrecht
SBAC-PAD10
2007 Computational Characteristics of Production Seismic Migration and its Performance on Novel Processor Architectures
abstract
We describe the computational characteristics of the Kirchhoff prestack seismic migration currently used in daily production runs at Petrobras and its port to novel architectures. Fully developed in house, this portable and fault tolerant application has high sequential and parallel efficiency, with parallel scalability tested up to 8192 processors on the IBM Blue Gene without exhausting parallelism. Production load comprises thousands of jobs per year, consuming the installed park of a few thousand x86 CPU cores, with top production runs continuously using up to 1000 dedicated processors during 20 days. A built in mechanism automatically produces, collects and stores job performance data, allowing single job performance analysis and multi job performance statistics. Its port to quad-core x86 and Sony PlayStation 3 achieved very high price/performance and performance/watt gains over single core x86 machines. Port to the PS3 is described in detail. Experimental performance data on a modest PS3 cluster is also presented.
Jairo Panetta, Paulo R. P. de Souza Filho, Carlos A. da Cunha Filho, Fernando M. Roxo da Motta, Silvio Sinedino Pinheiro, Ivan Pedrosa Junior, Andre L. Romanelli Rosa, Luiz Rodolpho Monnerat, Leandro T. Carneiro, Carlos H. B. de Albrecht
SBAC-PAD8
2006 D1HT: a distributed one hop hash table
abstract
Distributed hash tables (DHTs) have been used in a variety of applications, but most DHTs so far have opted to solve lookups with multiple hops, which sacrifices performance in order to keep little routing information and minimize maintenance traffic. In this paper, we introduce D1HT, a novel single hop DHT that is able to maximize performance with reasonable maintenance traffic overhead even for huge and dynamic peer-to-peer (P2P) systems. We formally define the algorithm we propose to detect and notify any membership change in the system, prove its correctness and performance properties, and present a quarantine-like mechanism to reduce the overhead caused by volatile peers. Our analyses show that D1HT has reasonable maintenance bandwidth requirements even for very large systems, while presenting at least twice less bandwidth overhead than previous single hop DHT
Luiz Rodolpho Monnerat, Claudio Luis de Amorim
IPDPS1
1998 Efficiently Adapting to Sharing Patterns in Software DSMs
abstract
In this paper we introduce a page-based Lazy Release Consistency protocol called ADSM that constantly and efficiently adapts to the applications' sharing patterns. Adaptation in ADSM is based on our dynamic categorization of the type of sharing experienced by each page. Pages can be categorized as falsely-shared, migratory, or producer/consumer(s). Migratory and producer/consumer(s) pages are managed in single-writer mode, while falsely-shared data are managed in multiple-writer mode. Coherence is kept with invalidations for most types of the shared data, but updates are used for lock-protected data in migratory state and barrier-protected delta in producer/consumer(s) state. We performed experiments with 6 parallel applications on an 8-node SP2 system, comparing our protocol against standard TreadMarks and a version of TreadMarks that also adapts to sharing patterns. Our results show that ADSM consistently outperforms its competitors; our protocol can improve the TreadMarks speedups by as much as 155%, while surpassing the performance of the adaptive TreadMarks implementation by as much as 67%. Our main conclusions are that our categorization and adaptation strategies are useful techniques for improving the performance of page-based software DSMs, while ADSM is a highly-efficient option for low-cost parallel computing.
Luiz Rodolpho Monnerat, Ricardo Bianchini
HPCA1