EDBT 2026 Demo / reviewers in the wild / expert
Alexandro Baldassin
dblp:22/3545 · also Alexandro José Baldassin
· DBLP profile ↗
30ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0001-8824-3055ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FUR: Fast and Unlimited Reads on Persistent Memory TransactionsabstractDespite the recent improvements in supporting Persistent Hardware Transactions (PHTs) on emerging persistent memories (PM), they have largely overlooked the poor performance of Read-Only (RO) transactions, which suffer from two crucial bottlenecks: i) the considerable post-commit delays required to ensure consistency with concurrent update transactions; and ii) the well-known tight read capacity limits of the commercially available HTM implementations. João Barreto 0001, Daniel Castro 0004, Paolo Romano 0002, Alexandro Baldassin |
EuroSys | 4 |
| 2025 | CausalBioCF: Causal Counterfactuals for Machine Learning Interpretability
Gabriel C. Furlanetto, Alexandro Baldassin, Aleardo Manacero |
ICCSA (2) | 2 |
| 2024 | Using hardware-transactional-memory support to implement speculative task execution
Juan Salamanca 0001, Alexandro Baldassin |
J. Parallel Distributed Comput. | 2 |
| 2024 | Rank-based Hashing for Effective and Efficient Nearest Neighbor Search for Image RetrievalabstractThe large and growing amount of digital data creates a pressing need for approaches capable of indexing and retrieving multimedia content. A traditional and fundamental challenge consists of effectively and efficiently performing nearest-neighbor searches. After decades of research, several different methods are available, including trees, hashing, and graph-based approaches. Most of the current methods exploit learning to hash approaches based on deep learning. In spite of effective results and compact codes obtained, such methods often require a significant amount of labeled data for training. Unsupervised approaches also rely on expensive training procedures usually based on a huge amount of data. In this work, we propose an unsupervised data-independent approach for nearest neighbor searches, which can be used with different features, including deep features trained by transfer learning. The method uses a rank-based formulation and exploits a hashing approach for efficient ranked list computation at query time. A comprehensive experimental evaluation was conducted on seven public datasets, considering deep features based on CNNs and Transformers. Both effectiveness and efficiency aspects were evaluated. The proposed approach achieves remarkable results in comparison to traditional and state-of-the-art methods. Hence, it is an attractive and innovative solution, especially when costly training procedures need to be avoided. Vinicius Atsushi Sato Kawai, Lucas Pascotti Valem, Alexandro Baldassin, Edson Borin, Daniel C. G. Pedronette, Longin Jan Latecki |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | On the impact of mode transition on phased transactional memory performance
Catalina Munoz Morales, Bruno C. Honorio, João P. L. de Carvalho, Alexandro Baldassin, Guido Araujo |
J. Parallel Distributed Comput. | 4 |
| 2022 | Using Barrier Elision to Improve Transactional Code GenerationabstractWith chip manufacturers such as Intel, IBM, and ARM offering native support for transactional memory in their instruction set architectures, memory transactions are on the verge of being considered a genuine application tool rather than just an interesting research topic. Despite this recent increase in popularity on the hardware side of transactional memory (HTM) , software support for transactional memory (STM) is still scarce and the only compiler with transactional support currently available, the GNU Compiler Collection (GCC) , does not generate code that achieves desirable performance. For hybrid solutions of TM (HyTM) , which are frameworks that leverage the best aspects of HTM and STM, the subpar performance of the software side, caused by inefficient compiler generated code, might forbid HyTM to offer optimal results. This article extends previous work focused exclusively on STM implementations by presenting a detailed analysis of transactional code generated by GCC in the context of HybridTM implementations. In particular, it builds on previous research of transactional memory support in the Clang/LLVM compiler framework, which is decoupled from any TM runtime, and presents the following novel contributions: (a) it shows that STM’s performance overhead, due to an excessive amount of read and write barriers added by the compiler, also impacts the performance of HyTM systems; and (b) it reveals the importance of the previously proposed annotation mechanism to reduce the performance gap between HTM and STM in phased runtime systems. Furthermore, it shows that, by correctly using the annotations on just a few lines of code, it is possible to reduce the total number of instrumented barriers by 95% and to achieve speed-ups of up to 7× when compared to the original code generated by GCC and the Clang compiler. 1 Bruno C. Honorio, João P. L. de Carvalho, Catalina Munoz Morales, Alexandro Baldassin, Guido Araujo |
ACM Trans. Archit. Code Optim. | 4 |
| 2021 | Accelerating Graph Applications Using Phased Transactional Memory
Catalina Munoz Morales, Rafael Murari, João P. L. de Carvalho, Bruno C. Honorio, Alexandro Baldassin, Guido Araujo |
Euro-Par | 5 |
| 2021 | SPHT: Scalable Persistent Hardware Transactions
Daniel Castro 0004, Alexandro Baldassin, João Barreto 0001, Paolo Romano 0002 |
FAST | 2 |
| 2021 | Improving Phased Transactional Memory via Commit Throughput and Capacity EstimationabstractTransactional Memory (TM) is a programming abstraction that aims to ease parallel programming in shared-memory architectures. Both Hardware (HTM) and Software Transactional Memory (STM) implementations have been extensively studied in the literature. Modern approaches seek to combine both HTM and STM to better exploit performance. In particular, Phased TMs (PhTMs) systems execute transactions in phases, not allowing both hardware and software transactions to run concurrently to avoid coordination overheads. The main challenge in designing PhTM systems is to dynamically choose a proper execution mode. Usually, a transition mechanism is developed based on metrics such as transaction size and abort rates to guide the phase migration. However, the tuning of such metrics is not an easy task, since it may lead to over-fitting and poor performance for the general case. This paper advances state-of-the-art research on PhTM by proposing a different approach to phase selection: the use of commit throughput and cache simulation to mimic the behavior of HTM storage constraints while in STM mode. When compared to previous work, this approach leads to a simpler and more efficient mechanism to assess the state of the execution modes in run time. Experimental results using STAMP and two graph processing applications show how the Commit Throughput-based mechanism is able to outperform a state-of-the-art Phased TM runtime (PhTM*) with speedups of up to 5x. Catalina Munoz Morales, Bruno C. Honorio, Alexandro Baldassin, Guido Araujo |
SBAC-PAD | 3 |
| 2020 | NV-PhTM: An Efficient Phase-Based Transactional System for Non-volatile Memory
Alexandro Baldassin, Rafael Murari, João P. L. de Carvalho, Guido Araujo, Daniel Castro 0004, João Barreto 0001, Paolo Romano 0002 |
Euro-Par | 1 |
| 2020 | Improving Transactional Code Generation via Variable Annotation and Barrier ElisionabstractWith chip manufacturers such as Intel, IBM and ARM offering native support for transactional memory in their instruction set architectures, memory transactions are on the verge of being considered a genuine application tool rather than just an interesting research topic. Despite this recent increase in popularity on the hardware side of transactional memory (HTM), software support for transactional memory (STM) is still scarce and the only compiler with transactional support currently available, the GNU Compiler Collection (GCC), does not generate code that achieves desirable performance. This paper presents a detailed analysis of transactional code generated by GCC and by a proposed transactional memory support added to the Clang/LLVM compiler framework. Experimental results support the following contributions: (a) STM's performance overhead is due to an excessive amount of read and write barriers added by the compiler; (b) a new annotation mechanism for the Clang/LLVM compiler framework that aims to overcome the barrier over-instrumentation problem by allowing programmers to specify which variables should be free from transactional instrumentation; (c) a profiling tool that ranks the most accessed memory locations at runtime, working as a guiding tool for programmers to annotate the code. Furthermore, it is revealed that, by correctly using the annotations on just a few lines of code, it is possible to reduce the total number of instrumented barriers by 95% and to achieve speed-ups of up to 7× when compared to the original code generated by GCC and the Clang compiler. João P. L. de Carvalho, Bruno C. Honorio, Alexandro Baldassin, Guido Araujo |
IPDPS | 3 |
| 2020 | An efficient parallel implementation for training supervised optimum-path forest classifiers
Aldo Culquicondor, Alexandro Baldassin, César Castelo-Fernández, João P. L. de Carvalho, João Paulo Papa |
Neurocomputing | 2 |
| 2019 | Semi-supervised and active learning through Manifold Reciprocal kNN Graph for image retrieval
Daniel C. G. Pedronette, Ying Weng, Alexandro Baldassin, Chaohuan Hou |
Neurocomputing | 3 |
| 2019 | An optimized unsupervised manifold learning algorithm for manycore architectures
Alexandro Baldassin, Ying Weng, Daniel C. G. Pedronette, Jurandy Almeida |
Inf. Sci. | 1 |
| 2019 | The Case for Phase-Based Transactional MemoryabstractIn recent years, Hybrid TM (HyTM) has been proposed as a transactional memory approach that leverages on the advantages of both hardware (HTM) and software (STM) execution modes. HyTM assumes that concurrent transactions have very different phases and thus should run under different execution modes. Conversely, Phased Transactional Memory (PhTM) considers that concurrent transactions have similar phases, and thus all transactions could run under the same mode. In this paper we make the case for phase-based transactional systems using PhTM*, the first implementation of PhTM on modern HTM-ready processors. PhTM* novelty relies on avoiding unnecessary transitions to software mode. Experimental results using Broadwell's TSX reveal that, for the STAMP benchmark suite, PhTM* performs on average 1.68x better than PhTM, a previous phase-based TM, 2.08x better than HyTM-NOrec, a state-of-the-art HyTM, and 2.28x better than HyCO, the most recent hybrid system in the literature. We also show that STAMP applications do not exhibit hybrid behavior to justify the use of conventional hybrid systems, thus making PhTM* a better solution to those type of programs. Finally, we show for the first time that conventional hybrid systems do not perform better than phased-based system in a scenario with hybrid-behaved transactions. João P. L. de Carvalho, Guido Araujo, Alexandro Baldassin |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2018 | Pattern Analysis in Drilling Reports using Optimum-Path ForestabstractWell drilling monitoring is an essential task to prevent faults, save resources, and take care of environmental and eco-planning businesses. During drilling, it is required that staff fill out a log to keep track of the activities that are currently occurring. With such data analyzed and processed, it is possible to learn how to prevent faults and take corrective actions in realtime. However, the most important information is usually stored in a free-text format, thus complicating the task of automated text mining. In this work, we introduce the Optimum-Path Forest (OPF) for sentence classification in drilling reports and compare its results against some state-of-art results. We show that OPF combined with text-based features are a compelling source to learn patterns in drilling reports. Gustavo José de Sousa, Daniel C. G. Pedronette, Alexandro Baldassin, Pedro Ivo Monteiro Privatto, M. Gaseta, Ivan Rizzo Guilherme, Danilo Colombo, Luis C. S. Afonso, João Paulo Papa |
IJCNN | 3 |
| 2017 | Quaternionic Flower Pollination Algorithm
Gustavo H. Rosa, Luis C. S. Afonso, Alexandro Baldassin, João Paulo Papa, Xin-She Yang 0001 |
CAIP (2) | 3 |
| 2017 | Revisiting phased transactional memoryabstractIn recent years, Hybrid TM (HyTM) has been proposed as a transactional memory approach that leverages on the advantages of both hardware (HTM) and software (STM) execution modes. HyTM assumes that concurrent transactions can have very different phases and thus should run under different execution modes. Although HyTM has shown to improve performance, the overall solution can be complicated to manage, both in terms of correctness and performance. On the other hand, Phased Transactional Memory (PhTM) considers that concurrent transactions have similar phases, and thus all transactions could run under the same mode. As a result, PhTM does not require coordination between transactions on distinct modes making its implementation simpler and more flexible. In this paper we claim that PhTM is a competitive alternative to HyTM and propose PhTM*, the first implementation of PhTM on modern HTM-ready processors. PhTM* novelty relies in avoiding unnecessary transitions to software mode by: (i) taking into account the categories of hardware aborts; (ii) adding a new serialization mode. Experimental results with Haswell's TSX reveal that, for the STAMP benchmark suite, PhTM* performs on average 11% better than PhTM, a previous phase-based TM, and 15% better than HyTM-NOrec, a state-of-the-art HyTM. In addition, PhTM* showed to be even more effective running on a Power8 machine by performing over 25% and 36% better than PhTM and HyTM-NOrec, respectively. João P. L. de Carvalho, Guido Araujo, Alexandro Baldassin |
ICS | 3 |
| 2017 | FGSCM: A Fine-Grained Approach to Transactional Lock ElisionabstractSpeculative Lock Elision (SLE) is a technique that allows critical sections to be executed optimistically by eliding the lock operation and enabling multiple threads to execute concurrently. In case of inconsistencies, the hardware automatically rolls back the execution and pessimistically acquires the original lock during runtime. The decision to elide the lock in SLE is performed transparently at the microarchitecture level and, although being convenient, it may sometimes hurt performance. To avoid that case, researchers have investigated Transactional Lock Elision (TLE), in which software-controlled hardware transactions are used instead, allowing the creation of policies and heuristics to manage lock elision. Typical implementations of TLE make use of a single lock to serialize the execution in case the original lock cannot be elided, which can potentially degrade performance. In order to improve on such cases, this paper proposes the Fine-Grained Software-assisted Conflict Management (FGSCM) scheme, a TLE technique that employs multiple locks so as to avoid unnecessary serialization of the code. The main idea of FGSCM is that not all threads that conflict inside a critical section are acessing the same region of shared memory. By automatically assigning distinct locks to these threads according to the memory section they access, the level of concurrency can be increased. In this paper we formalize FGSCM and provide an in-depth performance evaluation using a microbenchmark to stress several conflict behaviors. Our initial results with a prototype implementation using Intels Restricted Transactional Memory (RTM) are encouraging. With a quadcore machine, we observed an average performance gain of 11% compared to the single-auxiliary-lock SCM and 36% compared to a standard lock scheme, both for typical read-dominated workloads. Gustavo Sousa, Alexandro Baldassin |
SBAC-PAD | 2 |
| 2015 | Performance implications of dynamic memory allocators on transactional memory systemsabstractAlthough dynamic memory management accounts for a significant part of the execution time on many modern software systems, its impact on the performance of transactional memory systems has been mostly overlooked. In order to shed some light into this subject, this paper conducts a thorough investigation of the interplay between memory allocators and software transactional memory (STM) systems. We show that allocators can interfere with the way memory addresses are mapped to versioned locks on state-of-the-art software transactional memory implementations. Moreover, we observed that key aspects of allocators such as false sharing avoidance, scalability, and locality have a drastic impact on the final performance. For instance, we have detected performance differences of up to 171% in the STAMP applications when using distinct allocators. Moreover, we show that optimizations at the STM-level (such as caching transactional objects) are not effective when a modern allocator is already in use. All in all, our study highlights the importance of reporting the allocator utilized in the performance evaluation of transactional memory systems. Alexandro Baldassin, Edson Borin, Guido Araujo |
PPoPP | 1 |
| 2013 | Transaction scheduling using conflict avoidance and Contention IntensityabstractIn the last few years, Transactional Memories (TMs) have been shown to be a parallel programming model that can effectively combine performance improvement with ease of programming. Moreover, the recent introduction of TM-based ISA extensions, by major microprocessor manufacturers, also seems to endorse TM as a programming model for today's parallel applications. One of the central issues in designing Software TM (STM) systems is to identify mechanisms/heuristics that can minimize contention arising from conflicting transactions. Although a number of mechanisms have been proposed to tackle contention, such techniques have a limited scope, as conflict is avoided by either interrupting or serializing transaction execution, thus considerably impacting performance. To deal with this limitation, we have proposed a new effective transaction scheduler, along with a conflict-avoidance heuristic, that implements a fully cooperative scheduler that switches a conflicting transaction by another with a lower conflicting probability. This paper extends such framework and introduces a new heuristic, built from the combination of our previous conflict avoidance technique with the Contention Intensity heuristic proposed by Yoo and Lee. Experimental results, obtained using the STMBench7 and STAMP benchmarks atop tinySTM, show that the proposed heuristic produces significant speedups when compared to other four solutions. Márcio Machado Pereira, Alexandro Baldassin, Guido Araujo, Luiz Eduardo Buzato |
HiPC | 2 |
| 2012 | Vectorized Algorithms for Quadtree Construction and Descent
Eraldo P. Marinho, Alexandro Baldassin |
ICA3PP (1) | 2 |
| 2012 | Energy-Performance Tradeoffs in Software Transactional MemoryabstractTransactional memory (TM) is a new synchronization mechanism devised to simplify parallel programming, thereby helping programmers to unleash the power of current multicore processors. Although software implementations of TM (STM) have been extensively analyzed in terms of runtime performance, little attention has been paid to an equally important constraint faced by nearly all computer systems: energy consumption. In this work we conduct a comprehensive study of energy and runtime tradeoff sin software transactional memory systems. We characterize the behavior of three state-of-the-art lock-based STM algorithms, along with three different conflict resolution schemes. As a result of this characterization, we propose a DVFS-based technique that can be integrated into the resolution policies so as to improve the energy-delay product (EDP). Experimental results show that our DVFS-enhanced policies are indeed beneficial for applications with high contention levels. Improvements of up to 59% in EDP can be observed in this scenario, with an average EDP reduction of 16% across the STAMP workloads. Alexandro Baldassin, João P. L. de Carvalho, Leonardo A. G. Garcia, Rodolfo Azevedo |
SBAC-PAD | 1 |
| 2012 | A transactional runtime system for the Cell/BE architecture
Alexandro Baldassin, Felipe Goldstein, Rodolfo Azevedo |
J. Parallel Distributed Comput. | 1 |
| 2011 | LUTS: A Lightweight User-Level Transaction Scheduler
Daniel Nicácio, Alexandro Baldassin, Guido Araujo |
ICA3PP (1) | 2 |
| 2010 | STM versus lock-based systems: an energy consumption perspectiveabstractThe shift towards multicore processors and the well-known drawbacks imposed by lock-based synchronization have forced researchers to devise new alternatives for building concurrent software, of which transactional memory is a promising one. This work presents a comprehensive study on the energy consumption of a state-of-the-art STM (Software Transactional Memory) implementation using STAMP, a representative set of transactional workloads, comparing it to its lock-based counterpart. Our results show that STM can be up to 22x (~3x on average) more energy-inefficient when compared to locks. This work is a novel step towards a better understanding of the energy behavior of STM systems. Felipe Klein, Alexandro Baldassin, Paulo Centoducatte, Sandro Rigo, Rodolfo Azevedo |
ISLPED | 2 |
| 2010 | Concurrent programming with revisions and isolation typesabstractBuilding applications that are responsive and can exploit parallel hardware while remaining simple to write, understand, test, and maintain, poses an important challenge for developers. In particular, it is often desirable to enable various tasks to read or modify shared data concurrently without requiring complicated locking schemes that may throttle concurrency and introduce bugs. Sebastian Burckhardt, Alexandro Baldassin, Daan Leijen |
OOPSLA | 2 |
| 2008 | A Software Transactional Memory System for an Asymmetric Processor ArchitectureabstractDue to the advent of multi-core processors and the consequent need for better concurrent programming abstractions, new synchronization paradigms have emerged. A promising one, known as software transactional memory (STM), aims to use transactions as the key synchronization mechanism to ease program development as well as increase its performance. Many (if not all) of the current STM implementations target homogeneous architectures. In this paper we describe an implementation of an STM system for an asymmetric architecture, the Cell BE. We evaluated our Transactional Software Cache (TSC) mechanism using a well-known micro-benchmark (IntSet) and the Genome application from STAMP. The results show that an STM implementation for the Cell architecture is feasible if the shared-memory programming model is adopted. When compared to a conventional lock-based implementation, the STM version of Genome obtained a performance gain of 84% and 24% with large and small input sets, respectively. Felipe Goldstein, Alexandro Baldassin, Paulo Centoducatte, Rodolfo Azevedo, Leonardo A. G. Garcia |
SBAC-PAD | 2 |
| 2008 | An open-source binary utility generatorabstractElectronic system level (ESL) modeling allows early hardware-dependent software (HDS) development. Due to broad CPU diversity and shrinking time-to-market, HDS development can neither rely on hand-retargeting binary tools, nor can it rely on pre-existent tools within standard packages. As a consequence, binary utilities which can be easily adapted to new CPU targets are of increasing interest. We present in this article a framework for automatic generation of binary utilities. It relies on two innovative ideas: platform-aware modeling and more inclusive relocation handling. Generated assemblers, linkers, disassemblers and debuggers were validated for MIPS, SPARC, PowerPC, i8051 and PIC16F84. An open-source prototype generator is available for download. Alexandro Baldassin, Paulo Centoducatte, Sandro Rigo, Daniel C. Casarotto, Luiz Cláudio Villar dos Santos, Max R. de O. Schultz, Olinto J. V. Furtado |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2005 | Extending the ArchC Language for Automatic Generation of AssemblersabstractIn this paper, we extend the ArchC language with new constructs to describe the assembly language syntax and operand encoding of an instruction set architecture. Based on the extended language we have created a tool which can automatically generate assemblers. Our tool uses the GNU Binutils framework in order to produce the assembler, generating the architecture dependent files necessary to retarget the GNU assembler and the Binutils libraries. We have generated assemblers for the MIPS-I and SPARC-V8 architectures based on ArchC models using our tool. The assemblers generated for both architectures were compared with the default gas assemblers for a set of files taken from the MiBench benchmark, and the ELF object files generated by each pair of assemblers were equivalent in both cases. Alexandro Baldassin, Paulo Centoducatte, Sandro Rigo |
SBAC-PAD | 1 |