Juan Salamanca 0001

dblp:151/3212-1 · DBLP profile ↗
← Back
6ranked-venue papers
5as first author
2since 2021 · last 2024
0000-0002-0569-2806ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Parallel and multicore computing · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing › transactional memory
hardware transactional memory
0.312018
Using Hardware-Transactional-Memory Support to Implement Thread-Level Speculation · IEEE Trans. Parallel Distributed Syst. 2018
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP extensions
0.312018
Using Hardware-Transactional-Memory Support to Implement Thread-Level Speculation · IEEE Trans. Parallel Distributed Syst. 2018
Parallel and multicore computing
parallel programming models
0.312018
Using Hardware-Transactional-Memory Support to Implement Thread-Level Speculation · IEEE Trans. Parallel Distributed Syst. 2018
Parallel and multicore computing › speculative parallelization
thread-level speculation
0.312018
Using Hardware-Transactional-Memory Support to Implement Thread-Level Speculation · IEEE Trans. Parallel Distributed Syst. 2018
Parallel and multicore computing
transactional memory
0.312018
Using Hardware-Transactional-Memory Support to Implement Thread-Level Speculation · IEEE Trans. Parallel Distributed Syst. 2018

Methods — techniques the papers use, named apart from their topics

performance evaluation · 0.3loop parallelization · 0.3
YearPublicationVenuePosition
2024 Using hardware-transactional-memory support to implement speculative task execution
Juan Salamanca 0001, Alexandro Baldassin
J. Parallel Distributed Comput.1
2022 Performance Comparison of Speculative Taskloop and OpenMP-for-Loop Thread-Level Speculation on Hardware Transactional Memory
abstract
Speculative Taskloop (STL) is a loop parallelization technique that takes the best of Task-based Parallelism and Thread-Level Speculation to speed up loops with may loop-carried dependencies that were previously difficult for compilers to parallelize. Previous studies show the efficiency of STL when implemented using Hardware Transactional Memory and the advantages it offers compared to a typical DOACROSS technique such as OpenMP ordered. This paper presents a performance comparison between STL and a previously proposed technique that implements Thread-Level Speculation (TLS) in the for worksharing construct (FOR-TLS) over a set of loops from cbench and SPEC2006 benchmarks. The results show interesting insights on how each technique can be more appropriate depending on the characteristics of the evaluated loop. Experimental results reveal that by implementing both techniques on top of HTM, speed-ups of up to 2.41× can be obtained for STL and up to 2× for FOR-TLS.
Juan Salamanca 0001
ISPDC1
2018 DOACROSS Parallelization Based on Component Annotation and Loop-Carried Probability
abstract
Although modern compilers implement many loop parallelization techniques, their application is typically restricted to loops that have no loop-carried dependences (DOALL) or that contain well-known structured dependence patterns (e.g. reduction). These restrictions preclude the parallelization of many computational intensive DOACROSS loops. In such loops, either the compiler finds at least one loop-carried dependence or it cannot prove, at compile-time, that the loop is free of such dependences, even though they might never show-up at runtime. In any case, most compilers end-up not parallelizing DOACROSS loops. This paper brings three contributions to address this problem. First, it integrates three algorithms (TLS, DOAX, and BDX) into a simple openMP clause that enables the programmer to select the best algorithm for a given loop. Second, it proposes an annotation approach to separate the sequential components of a loop, thus exposing other components to parallelization. Finally, it shows that loop-carried probability is an effective metric to decide when to use TLS or other non-speculative techniques (e.g. DOAX or BDX) to parallelize DOACROSS loops. Experimental results reveal that, for certain loops, slow-downs can be transformed in 2× speed-ups by quickly selecting the appropriate algorithm.
Luis Mattos, Divino Cesar S. Lucas, Juan Salamanca 0001, João P. L. de Carvalho, Márcio Machado Pereira, Guido Araujo
SBAC-PAD3
2018 Using Hardware-Transactional-Memory Support to Implement Thread-Level Speculation
abstract
This paper presents a detailed analysis of the application of Hardware Transactional Memory (HTM) support for loop parallelization with Thread-Level Speculation (TLS) and describes a careful evaluation of the implementation of TLS on the HTM extensions available in such machines. The sample implementation of TLS over HTM described in this paper also provides evidence that the programming effort to implement TLS over HTM support is non-trivial. Thus the paper also describes an extension to OpenMP that both makes TLS more accessible to OpenMP programmers and allows for the easytuning of TLS parameters. As a result, it provides evidence to support several important claims about the performance of TLS over HTM in the Intel Core and the IBM POWER8 architectures. Experimental results reveal that by implementing TLS on top of HTM, speed-ups of up to 3.8x can be obtained for some loops.
Juan Salamanca 0001, José Nelson Amaral, Guido Araujo
IEEE Trans. Parallel Distributed Syst.1
2017 Performance Evaluation of Thread-Level Speculation in Off-the-Shelf Hardware Transactional Memories
Juan Salamanca 0001, José Nelson Amaral, Guido Araujo
Euro-Par1
2016 Evaluating and Improving Thread-Level Speculation in Hardware Transactional Memories
abstract
This paper presents a detailed analysis of the application of Hardware Transactional Memory (HTM) support for loop parallelization with Thread-Level Speculation (TLS). As a result it provides three contributions: (a) it shows that performance issues well-known to loop parallelism (e.g. false sharing) are exacerbated in the presence of HTM, and that capacity aborts can increase when one tries to overcome them, (b) it reveals that, although modern HTM extensions can provide support for TLS, they are not powerful enough to fully implement TLS, (c) it shows that simple code transformations, such as judicious strip mining and privatization techniques, can overcome such shortcomings, delivering speed-ups for programs that contain loop-carried dependencies. Experimental results reveal that, when these code transformations are used, speed-ups of up to 30% can be achieved for some loops for which previous research had reported slowdowns.
Juan Salamanca 0001, José Nelson Amaral, Guido Araujo
IPDPS1