EDBT 2026 Demo / reviewers in the wild / expert
Estanislao Mercadal
dblp:35/9592
· DBLP profile ↗
7ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0002-1835-8671ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 94% Memory systems · 6% |
Topics — the 5 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation › profiling
application profiling |
0.8 | 1 | 2024 | A Mess of Memory System Benchmarking, Simulation and Application Profiling · MICRO 2024 |
Performance modeling and evaluation › benchmarking › computer architecture benchmarking
memory system benchmarking |
0.8 | 1 | 2024 | A Mess of Memory System Benchmarking, Simulation and Application Profiling · MICRO 2024 |
Performance modeling and evaluation › simulation › architectural simulation
memory system simulation |
0.8 | 1 | 2024 | A Mess of Memory System Benchmarking, Simulation and Application Profiling · MICRO 2024 |
Performance modeling and evaluation
simulation |
0.8 | 1 | 2024 | A Mess of Memory System Benchmarking, Simulation and Application Profiling · MICRO 2024 |
Performance modeling and evaluation
workload characterization |
0.8 | 1 | 2024 | A Mess of Memory System Benchmarking, Simulation and Application Profiling · MICRO 2024 |
Methods — techniques the papers use, named apart from their topics
bandwidth-latency curve measurement · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 15+ years of joint parallel application performance analysis/tools training with Scalasca/Score-P and Paraver/Extrae toolsetsabstractThe diverse landscape of distributed heterogeneous computer systems currently available and being created to address computational challenges with the highest performance requirements presents daunting complexity for application developers. They must effectively decompose and distribute their application functionality and data, efficiently orchestrating the associated communication and synchronisation, on multi/manycore CPU processors with multiple attached acceleration devices structured within compute nodes with interconnection networks of various topologies. Sophisticated compilers, runtime systems and libraries are (loosely) matched with debugging, performance measurement and analysis tools, with proprietary versions by integrators/vendors provided exclusively for their systems complemented by portable (primarily) open-source equivalents developed and supported by the international research community over many years. The Scalasca and Paraver toolsets are two widely employed examples of the latter, installed on personal notebook computers through to the largest leadership HPC systems. Over more than fifteen years their developers have worked closely together in numerous collaborative projects culminating in the creation of a universal parallel performance assessment and optimisation methodology focused on application execution efficiency and scalability, and the associated training and coaching of application developers (often in teams) in its productive use, reviewed in this article with lessons learnt therefrom. Brian J. N. Wylie, Judit Giménez, Christian Feld, Markus Geimer, Germán Llort, Sandra Méndez, Estanislao Mercadal, Anke Visser, Marta García-Gasulla |
Future Gener. Comput. Syst. | 7 |
| 2024 | A Mess of Memory System Benchmarking, Simulation and Application ProfilingabstractThe Memory stress (Mess) framework provides a unified view of the memory system benchmarking, simulation and application profiling. The Mess benchmark provides a holistic and detailed memory system characterization. It is based on hundreds of measurements that are represented as a family of bandwidth-latency curves. The benchmark increases the coverage of all the previous tools and leads to new findings in the behavior of the actual and simulated memory systems. We deploy the Mess benchmark to characterize Intel, AMD, IBM, Fujitsu, Amazon and NVIDIA servers with DDR4, DDR5, HBM2 and HBM2E memory. The Mess memory simulator uses bandwidth-latency concept for the memory performance simulation. We integrate Mess with widely-used CPUs simulators enabling modeling of all high-end memory technologies. The Mess simulator is fast, easy to integrate and it closely matches the actual system performance. By design, it enables a quick adoption of new memory technologies in hardware simulators. Finally, the Mess application profiling positions the application in the bandwidth-latency space of the target memory system. This information can be correlated with other application runtime activities and the source code, leading to a better overall understanding of the application's behavior. The current Mess benchmark release covers all major CPU and GPU ISAs, x86, ARM, Power, RISC-V, and NVIDIA's PTX. We also release as open source the ZSim, gem5 and OpenPiton Metro-MPI integrated with the Mess simulator for DDR4, DDR5, Optane, HBM2, HBM2E and CXL memory expanders. The Mess application profiling is already integrated into a suite of production HPC performance analysis tools. Pouya Esmaili-Dokht, Francesco Sgherzi, Valéria Soldera Girelli, Isaac Boixaderas, Mariana Carmin, Alireza Monemi, Adrià Armejach, Estanislao Mercadal, Germán Llort, Petar Radojkovic, Miquel Moretó, Judit Giménez, Xavier Martorell, Eduard Ayguadé, Jesús Labarta, Emanuele Confalonieri, Rishabh Dubey, Jason Adlard |
MICRO | 8 |
| 2020 | Experiences on the characterization of parallel applications in embedded systems with Extrae/ParaverabstractCutting-edge functionalities in embedded systems require the use of parallel architectures to meet their performance requirements. This imposes the introduction of a new layer in the software stacks of embedded systems: the parallel programming model. Unfortunately, the tools used to analyze embedded systems fall short to characterize the performance of parallel applications at a parallel programming model level, and correlate this with information about non-functional requirements such as real-time, energy, memory usage, etc. HPC tools, like Extrae, are designed with that level of abstraction in mind, but their main focus is on performance evaluation. Overall, providing insightful information about the performance of parallel embedded applications at the parallel programming model level, and relate it to the non-functional requirements, is of paramount importance to fully exploit the performance capabilities of parallel embedded architectures. Adrian Munera, Sara Royuela, Germán Llort, Estanislao Mercadal, Franck Wartel, Eduardo Quiñones |
ICPP | 4 |
| 2020 | Analyzing the Efficiency of Hybrid CodesabstractHybrid parallelization may be the only path for most codes to use HPC systems on a very large scale. Even within a small scale, with an increasing number of cores per node, combining MPI with some shared memory thread-based library allows to reduce the application network requirements. Despite the benefits of a hybrid approach, it is not easy to achieve an efficient hybrid execution. This is not only because of the added complexity of combining two different programming models, but also because in many cases the code was initially designed with just one level of parallelization and later extended to a hybrid mode. This paper presents our model to diagnose the efficiency of hybrid applications, distinguishing the contribution of each parallel programming paradigm. The flexibility of the proposed methodology allows us to use it for different paradigms and scenarios, like comparing the MPI+OpenMP and MPI+CUDA versions of the same code. Judit Giménez, Estanislao Mercadal, Germán Llort, Sandra Méndez |
ISPDC | 2 |
| 2017 | Automating the Application Data Placement in Hybrid Memory SystemsabstractMulti-tiered memory systems, such as those based on Intel®Xeon Phi™ processors, are equipped with several memory tiers with different characteristics including, among others, capacity, access latency, bandwidth, energy consumption, and volatility. The proper distribution of the application data objects into the available memory layers is key to shorten the time- to-solution, but the way developers and end-users determine the most appropriate memory tier to place the application data objects has not been properly addressed to date. In this paper we present a novel methodology to build an extensible framework to automatically identify and place the application's most relevant memory objects into the Intel Xeon Phi fast on-package memory. Our proposal works on top of inproduction binaries by first exploring the application behavior and then substituting the dynamic memory allocations. This makes this proposal valuable even for end-users who do not have the possibility of modifying the application source code. We demonstrate the value of a framework based in our methodology for several relevant HPC applications using different allocation strategies to help end-users improve performance with minimal intervention. The results of our evaluation reveal that our proposal is able to identify the key objects to be promoted into fast on-package memory in order to optimize performance, leading to even surpassing hardware-based solutions. Harald Servat, Antonio J. Peña, Germán Llort, Estanislao Mercadal, Hans-Christian Hoppe, Jesús Labarta |
CLUSTER | 4 |
| 2013 | Improving the dynamism of mobile agent applications in wireless sensor networks through separate itineraries
Estanislao Mercadal, Carlos Vidueira, Cormac J. Sreenan, Joan Borrell |
Comput. Commun. | 1 |
| 2012 | Towards efficient access control in a mobile agent based wireless sensor networkabstractPublic key authorization credentials provide a flexible approach to implementing access control in open distributed systems. Wireless sensor networks, are examples of such systems; however, their low-power sensors have energy efficiency requirements that may mean it is not practical to carry out computationally intensive operations, such as public key operations. This paper describes a distributed access control system for a Wireless Sensor Network application that uses computationally efficient one-way hash-functions to implement authorization credentials. Estanislao Mercadal, Guillermo Navarro-Arribas, Simon N. Foley, Joan Borrell |
CRiSIS | 1 |