EDBT 2026 Demo / reviewers in the wild / expert
Judit Giménez
dblp:39/5410 · also Judit Giménez Lucas
· DBLP profile ↗
21ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0002-2501-2791ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 2 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Performance modeling and evaluation · 95% Memory systems · 5% |
Topics — the 5 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Performance modeling and evaluation
workload characterization |
0.9 | 2 | 2024 | A Mess of Memory System Benchmarking, Simulation and Application Profiling · MICRO 2024 On the usefulness of object tracking techniques in performance analysis · SC 2013 |
Performance modeling and evaluation › profiling
application profiling |
0.8 | 1 | 2024 | A Mess of Memory System Benchmarking, Simulation and Application Profiling · MICRO 2024 |
Performance modeling and evaluation › benchmarking › computer architecture benchmarking
memory system benchmarking |
0.8 | 1 | 2024 | A Mess of Memory System Benchmarking, Simulation and Application Profiling · MICRO 2024 |
Performance modeling and evaluation › simulation › architectural simulation
memory system simulation |
0.8 | 1 | 2024 | A Mess of Memory System Benchmarking, Simulation and Application Profiling · MICRO 2024 |
Performance modeling and evaluation
simulation |
0.8 | 1 | 2024 | A Mess of Memory System Benchmarking, Simulation and Application Profiling · MICRO 2024 |
Methods — techniques the papers use, named apart from their topics
bandwidth-latency curve measurement · 0.8object tracking · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | 15+ years of joint parallel application performance analysis/tools training with Scalasca/Score-P and Paraver/Extrae toolsetsabstractThe diverse landscape of distributed heterogeneous computer systems currently available and being created to address computational challenges with the highest performance requirements presents daunting complexity for application developers. They must effectively decompose and distribute their application functionality and data, efficiently orchestrating the associated communication and synchronisation, on multi/manycore CPU processors with multiple attached acceleration devices structured within compute nodes with interconnection networks of various topologies. Sophisticated compilers, runtime systems and libraries are (loosely) matched with debugging, performance measurement and analysis tools, with proprietary versions by integrators/vendors provided exclusively for their systems complemented by portable (primarily) open-source equivalents developed and supported by the international research community over many years. The Scalasca and Paraver toolsets are two widely employed examples of the latter, installed on personal notebook computers through to the largest leadership HPC systems. Over more than fifteen years their developers have worked closely together in numerous collaborative projects culminating in the creation of a universal parallel performance assessment and optimisation methodology focused on application execution efficiency and scalability, and the associated training and coaching of application developers (often in teams) in its productive use, reviewed in this article with lessons learnt therefrom. Brian J. N. Wylie, Judit Giménez, Christian Feld, Markus Geimer, Germán Llort, Sandra Méndez, Estanislao Mercadal, Anke Visser, Marta García-Gasulla |
Future Gener. Comput. Syst. | 2 |
| 2024 | A Mess of Memory System Benchmarking, Simulation and Application ProfilingabstractThe Memory stress (Mess) framework provides a unified view of the memory system benchmarking, simulation and application profiling. The Mess benchmark provides a holistic and detailed memory system characterization. It is based on hundreds of measurements that are represented as a family of bandwidth-latency curves. The benchmark increases the coverage of all the previous tools and leads to new findings in the behavior of the actual and simulated memory systems. We deploy the Mess benchmark to characterize Intel, AMD, IBM, Fujitsu, Amazon and NVIDIA servers with DDR4, DDR5, HBM2 and HBM2E memory. The Mess memory simulator uses bandwidth-latency concept for the memory performance simulation. We integrate Mess with widely-used CPUs simulators enabling modeling of all high-end memory technologies. The Mess simulator is fast, easy to integrate and it closely matches the actual system performance. By design, it enables a quick adoption of new memory technologies in hardware simulators. Finally, the Mess application profiling positions the application in the bandwidth-latency space of the target memory system. This information can be correlated with other application runtime activities and the source code, leading to a better overall understanding of the application's behavior. The current Mess benchmark release covers all major CPU and GPU ISAs, x86, ARM, Power, RISC-V, and NVIDIA's PTX. We also release as open source the ZSim, gem5 and OpenPiton Metro-MPI integrated with the Mess simulator for DDR4, DDR5, Optane, HBM2, HBM2E and CXL memory expanders. The Mess application profiling is already integrated into a suite of production HPC performance analysis tools. Pouya Esmaili-Dokht, Francesco Sgherzi, Valéria Soldera Girelli, Isaac Boixaderas, Mariana Carmin, Alireza Monemi, Adrià Armejach, Estanislao Mercadal, Germán Llort, Petar Radojkovic, Miquel Moretó, Judit Giménez, Xavier Martorell, Eduard Ayguadé, Jesús Labarta, Emanuele Confalonieri, Rishabh Dubey, Jason Adlard |
MICRO | 12 |
| 2022 | Feature Space Curvature Map: A Method To Homogenize Cluster DensitiesabstractThe majority of density-based clustering algorithms can not perform properly when data expose very different density through the feature space. These algorithms implicitly presume that all clusters almost have the same density, therefore, they normally use global parameters. Consequently, they are often biased towards finding dense clusters in front of sparse ones. In this paper, we propose a parametric multilinear transformation method to homogenize cluster densities while preserving the topological structure of the dataset. The transformed clusters have approximately the same density while all inter-cluster regions become globally low-density. In our method, the feature space is locally bent by dense data point concentrations the same way as stars bend the space-time dimensions in Theory of Relativity. We present a new Gravitational Self-organization Map to model the feature space curvature by plugging the concepts of gravity and fabric of space into the Self-organization Map algorithm to mathematically describe the density structure of the data. To homogenize the cluster density, we introduce a novel mapping mechanism to project the data from a non-Euclidean curved space to a new Euclidean flat space. Specifically, this mechanism transfers the basis vectors instead of the feature vectors to guarantee the continuity of the mapping function and optimize the computation cost of the algorithm. As a result, our method can efficiently and explicitly homogenize the density of any dataset globally to then apply existing clustering algorithms without modification. Our experimental results over both real-world and synthetic datasets show that our approach outperforms the current statistical-based methods. Kaveh Mahdavi, Jesús Labarta, Judit Giménez, Atefeh Mousavinia, Atiyeh Mousavinia |
IJCNN | 3 |
| 2021 | Organization Component Analysis: The method for extracting insights from the shape of clusterabstractClustering analysis is widely used to stratify data in the same cluster when they are similar according to specific metrics. The process of understanding and interpreting clusters is mostly intuitive. However, we observe each cluster has unique shape that comes out of metrics on data, which can represent the organization of categorized data mathematically. In this paper, we apply novel topological based method to study potentially complex high-dimensional categorized data by quantifying their shapes and extracting fine-grain insights about them to interpret the clustering result. We introduce our Organization Component Analysis method for the purpose of the automatic arbitrary cluster-shape study without assumption about the data distribution. Our method explores a topology-preserving map of a data cluster manifold to extract the main organization structure of a cluster by the leveraging of the self-organization map technique. To do this, we represent self-organization map as graph. We introduce organization components to geometrically describe the shape of cluster and their endogenous phenomena. Specifically, we propose an innovative way to measure the alignment between two sequences of momentum changes on geodesic path over the embedded graph to quantify the extent to which the feature is related to a given component. As a result, we can describe variability among stratified data, correlated features in terms of lower number of organization components. We illustrate the utilization of our method by applying it to two quite different types of data, in each case mathematically detecting the organization structure of categorized data which are much profounder and finer than those produced by standard methods. Kaveh Mahdavi, Jesús Labarta, Judit Giménez |
IJCNN | 3 |
| 2020 | Analyzing the Efficiency of Hybrid CodesabstractHybrid parallelization may be the only path for most codes to use HPC systems on a very large scale. Even within a small scale, with an increasing number of cores per node, combining MPI with some shared memory thread-based library allows to reduce the application network requirements. Despite the benefits of a hybrid approach, it is not easy to achieve an efficient hybrid execution. This is not only because of the added complexity of combining two different programming models, but also because in many cases the code was initially designed with just one level of parallelization and later extended to a hybrid mode. This paper presents our model to diagnose the efficiency of hybrid applications, distinguishing the contribution of each parallel programming paradigm. The flexibility of the proposed methodology allows us to use it for different paradigms and scenarios, like comparing the MPI+OpenMP and MPI+CUDA versions of the same code. Judit Giménez, Estanislao Mercadal, Germán Llort, Sandra Méndez |
ISPDC | 1 |
| 2019 | Unsupervised Feature Selection for Noisy Data
Kaveh Mahdavi, Jesús Labarta, Judit Giménez |
ADMA | 3 |
| 2018 | Understanding memory access patterns using the BSC performance tools
Harald Servat, Jesús Labarta, Hans-Christian Hoppe, Judit Giménez, Antonio J. Peña |
Parallel Comput. | 4 |
| 2016 | Bio-Inspired Call-Stack Reconstruction for Performance AnalysisabstractThe correlation of performance bottlenecks and their associated source code has become a cornerstone of performance analysis. It allows understanding why the efficiency of an application falls behind the computer's peak performance and enabling optimizations on the code ultimately. To this end, performance analysis tools collect the processor call-stack and then combine this information with measurements to allow the analyst comprehend the application behavior. Some tools modify the call-stack during run-time to diminish the collection expense but at the cost of resulting in non-portable solutions. In this paper, we present a novel portable approach to associate performance issues with their source code counterpart. To address it, we capture a reduced segment of the call-stack (up to three levels) and then process the segments using an algorithm inspired by multi-sequence alignment techniques. The results of our approach are easily mapped to detailed performance views, enabling the analyst to unveil the application behavior and its corresponding region of code. To demonstrate the usefulness of our approach, we have applied the algorithm to several first-time seen in-production applications to describe them finely, and optimize them by using tiny modifications based on the analyses. Harald Servat, Germán Llort, Juan Gonzalez, Judit Giménez, Jesús Labarta |
PDP | 4 |
| 2016 | Detailed and simultaneous power and performance analysisabstractSummary On the road to Exascale computing, both performance and power areas are meant to be tackled at different levels, from system to processor level. The processor itself is the main responsible for the serial node performance and also for the most of the energy consumed by the system. Thus, it is important to have tools to simultaneously analyze both performance and energy efficiency at processor level. Performance tools have allowed analysts to understand, and even improve, the performance of an application that runs in a system. With the advent of recent processor capabilities to measure its own power consumption, performance tools can increase their collection of metrics by adding those related to energy consumption and provide a correlation between the source code, its performance and its energy efficiency. In this paper, we present a performance tool that has been extended to gather such energy metrics. The results of this tool are passed to a mechanism called folding that produces detailed metrics and source code references by using coarse grain sampling. We have used the tool with multiple serial benchmarks as well as parallel applications to demonstrate its usefulness by locating hot spots in terms of performance and power drained. Copyright © 2013 John Wiley & Sons, Ltd. Harald Servat, Germán Llort, Judit Giménez, Jesús Labarta |
Concurr. Comput. Pract. Exp. | 3 |
| 2015 | Low-Overhead Detection of Memory Access Patterns and Their Time Evolution
Harald Servat, Germán Llort, Juan Gonzalez, Judit Giménez, Jesús Labarta |
Euro-Par | 4 |
| 2014 | Identifying Code Phases Using Piece-Wise Linear RegressionsabstractNode-level performance is one of the factors that may limit applications from reaching the supercomputers' peak performance. Studying node-level performance and attributing it to the source code results into valuable insight that can be used to improve the application efficiency, albeit performing such a study may be an intimidating task due to the complexity and size of the applications. We present in this paper a mechanism that takes advantage of combining piece-wise linear regressions, coarse-grain sampling, and minimal instrumentation to detect performance phases in the computation regions even if their granularity is very fine. This mechanism then maps the performance of each phase into the application syntactical structure displaying a correlation between performance and source code. We introduce a methodology on top of this mechanism to describe the node-level performance of parallel applications, even for first-time seen applications. Finally, we demonstrate the methodology describing optimized in-production applications and further improving their performance applying small transformations to the code based on the hints discovered. Harald Servat, Germán Llort, Juan Gonzalez, Judit Giménez, Jesús Labarta |
IPDPS | 4 |
| 2013 | On the usefulness of object tracking techniques in performance analysisabstractUnderstanding the behavior of a parallel application is crucial if we are to tune it to achieve its maximum performance. Yet the behavior the application exhibits may change over time and depend on the actual execution scenario: particular inputs and program settings, the number of processes used, or hardware-specific problems. So beyond the details of a single experiment a far more interesting question arises: how does the application behavior respond to changes in the execution conditions? Germán Llort, Harald Servat, Juan Gonzalez, Judit Giménez, Jesús Labarta |
SC | 4 |
| 2013 | Scalability analysis of Dalton, a molecular structure program
Xavier Aguilar, Michael Schliephake, Olav Vahtras, Judit Giménez, Erwin Laure |
Future Gener. Comput. Syst. | 4 |
| 2013 | Framework for a productive performance optimization
Harald Servat, Germán Llort, Kevin A. Huck, Judit Giménez, Jesús Labarta |
Parallel Comput. | 4 |
| 2011 | Scaling Dalton, A Molecular Electronic Structure ProgramabstractDalton is a molecular electronic structure program featuring common methods of computational chemistry that are based on pure quantum mechanics (QM) as well as hybrid quantum mechanics/molecular mechanics (QM/MM). It is specialized and has a leading position in calculation of molecular properties with a large world-wide user community (over 2000 licenses issued). In this paper, we present a characterization and performance optimization of Dalton that increases the scalability and parallel efficiency of the application. We also propose a solution that helps to avoid the master/worker design of Dalton to become a performance bottleneck for larger process numbers and increase the parallel efficiency. Xavier Aguilar, Michael Schliephake, Olav Vahtras, Judit Giménez, Erwin Laure |
eScience | 4 |
| 2011 | Trace Spectral Analysis toward Dynamic Levels of DetailabstractThe emergence of Petascale systems has raised new challenges to performance analysis tools. Understanding every single detail of an execution is important to bridge the gap between the theoretical peak and the actual performance achieved. Tracing tools are the best option when it comes to providing detailed information about the application behavior, but not without liabilities. The amount of information that a single execution can generate grows so fast that it easily becomes unmanageable. An effective analysis in such scenarios necessitates the intelligent selection of information. In this paper we present an on-line performance tool based on spectral analysis of signals that automatically identifies the different computing phases of the application as it runs, selects a few representative periods and decides the granularity of the information gathered for these regions. As a result, the execution is completely characterized at different levels of detail, reducing the amount of data collected while maximizing the amount of useful information presented for the analysis. Germán Llort, Marc Casas, Harald Servat, Kevin A. Huck, Judit Giménez, Jesús Labarta |
ICPADS | 5 |
| 2011 | Unveiling Internal Evolution of Parallel Application Computation PhasesabstractAs access to supercomputing resources is becoming more and more commonplace, performance analysis tools are gaining importance in order to decrease the gap between the application performance and the supercomputers' peak performance. Performance analysis tools allow the analyst to understand the idiosyncrasies of an application in order to improve it. However, these tools require monitoring regions of the application to provide information to the analysts, leaving non-monitored regions of code unknown, which may result in lack of understanding of important regions of the application. In this paper we describe an automated methodology that reports very detailed application insights and improves the analysis experience of performance tools based on traces. We apply this methodology to three production applications and provide suggestions on how to improve their performance. Our methodology uses computation burst clustering and a mechanism called folding. While clustering automatically detects application structure, folding combines instrumentation and sampling to augment the performance analysis details. Folding provides fine grain performance information from coarse grain sampling on iterative applications. Folding results closely resemble the performance data gathered from fine grain sampling with an absolute mean difference less than 5% without overhead of fine grain. Harald Servat, Germán Llort, Judit Giménez, Kevin A. Huck, Jesús Labarta |
ICPP | 3 |
| 2010 | Performance Data Extrapolation in Parallel CodesabstractMeasuring the performance of parallel codes is a compromise between lots of factors. The most important one is which data has to be analyzed. Current supercomputers are able to run applications in large number of processors as well as the analysis data that can be extracted is also large and varied. That implies a hard compromise between the potential problems one want to analyze and the information one is able to capture during the application execution. In this paper we present an extrapolation methodology to maximize the information extracted in a single application execution. It is based on a structural characterization of the applications, performed using clustering techniques, the ability to multiplex the read of performance hardware counters, plus a projection process. As a result, we obtain the approximated values of a large set of metrics for each phase of the application, with minimum error. Juan Gonzalez, Judit Giménez, Jesús Labarta |
ICPADS | 2 |
| 2010 | On-line detection of large-scale parallel application's structureabstractWith larger and larger systems being constantly deployed, trace-based performance analysis of parallel applications has become a challenging task. Even if the amount of performance data gathered per single process is small, traces rapidly become unmanageable when merging together the information collected from all processes. In general, an efficient analysis of such a large volume of data is subject to a previous filtering step that directs the analyst's attention towards what is meaningful to understand the observed application behavior. Furthermore, the iterative nature of most scientific applications usually ends up producing repetitive information. Discarding irrelevant data aims at reducing both the size of traces, and the time required to perform the analysis and deliver results. In this paper, we present an on-line analysis framework that relies on clustering techniques to intelligently select the most relevant information to understand how the application behaves, while keeping the volume of performance data at a reasonable size. Germán Llort, Juan Gonzalez, Harald Servat, Judit Giménez, Jesús Labarta |
IPDPS | 4 |
| 2009 | Automatic detection of parallel applications computation phasesabstractAnalyzing parallel programs has become increasingly difficult due to the immense amount of information collected on large systems. The use of clustering techniques has been proposed to analyze applications. However, while the objective of previous works is focused on identifying groups of processes with similar characteristics, we target a much finer granularity in the application behavior. In this paper, we present a tool that automatically characterizes the different computation regions between communication primitives in message-passing applications. This study shows how some of the clustering algorithms which may be applicable at a coarse grain are no longer adequate at this level. Density-based clustering algorithms applied to the performance counters offered by modern processors are more appropriate in this context. This tool automatically generates accurate displays of the structure of the application as well as detailed reports on a broad range of metrics for each individual region detected. Juan Gonzalez, Judit Giménez, Jesús Labarta |
IPDPS | 2 |
| 2009 | Automatic Evaluation of the Computation Structure of Parallel ApplicationsabstractMany data mining techniques have been proposed for parallel applications performance analysis, the most interesting being clustering analysis. Most cases have been used to detect processors with similar behavior. In previous work, we presented a different approach: clustering was used to detect the computation structure of the applications and how these different computation phases behave. In this paper, we present a method to evaluate the accuracy of this structure detection. This new method is based on the Single Program Multiple Data (SPMD) paradigm exhibited by real parallel programs. Assuming an SPMD structure, we expect that all tasks of a parallel application execute the same operation sequence. Using a Multiple Sequence Alignment (MSA) algorithm, we check the sequence ordering of the detected clusters to evaluate the quality of the clustering results. Juan Gonzalez, Judit Giménez, Jesús Labarta |
PDCAT | 2 |