Julita Corbalán

dblp:55/6520 · also Julita Corbalán González · DBLP profile ↗
← Back
40ranked-venue papers
10as first author
3since 2021 · last 2025
0000-0002-3926-5634ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 7 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 85% Performance modeling and evaluation · 10% Electronic design automation · 5%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Parallel and multicore computing
parallel programming models
0.112008
An adaptive cut-off for task parallelism · SC 2008
Parallel and multicore computing › parallel programming runtimes
runtime systems and scheduling
0.112008
An adaptive cut-off for task parallelism · SC 2008
Parallel and multicore computing › parallel programming models
task parallelism
0.112008
An adaptive cut-off for task parallelism · SC 2008
Parallel and multicore computing
processor allocation
0.122005
Performance-Driven Processor Allocation · IEEE Trans. Parallel Distributed Syst. 2005
Performance-Driven Processor Allocation · OSDI 2000
Parallel and multicore computing › task allocation
performance-driven allocation
0.122005
Performance-Driven Processor Allocation · IEEE Trans. Parallel Distributed Syst. 2005
Performance-Driven Processor Allocation · OSDI 2000
Parallel and multicore computing › parallel scheduling
adaptive scheduling
0.112005
Performance-Driven Processor Allocation · IEEE Trans. Parallel Distributed Syst. 2005
Performance modeling and evaluation › workload characterization
multiprogramming level
0.112005
Performance-Driven Processor Allocation · IEEE Trans. Parallel Distributed Syst. 2005
Electronic design automation › high-level synthesis
scheduling
0.012000
Performance-Driven Processor Allocation · OSDI 2000
Parallel and multicore computing › parallel programming models › task parallelism
OpenMP tasking
0.012008
An adaptive cut-off for task parallelism · SC 2008
Operating systems › resource management › process management › CPU scheduling
multiprocessor scheduling
0.012005
Performance-Driven Processor Allocation · IEEE Trans. Parallel Distributed Syst. 2005

Methods — techniques the papers use, named apart from their topics

simulation · 0.1performance measurement · 0.1runtime information collection · 0.1adaptive cut-off · 0.1
YearPublicationVenuePosition
2025 Static Powercap vs. EAR Hard-Powercap: Performance Evaluation
Julita Corbalán, Lluis Alonso, Luigi Brochard, Jordi Aneas, Oriol Vidal
JSSPP1
2021 Explicit uncore frequency scaling for energy optimisation policies with EAR in Intel architectures
abstract
EAR is an energy management framework which offers three main services: energy accounting, energy control and energy optimisation. The latter is done through the EAR runtime library (EARL). EARL is a dynamic, transparent, and lightweight runtime library that provides energy optimisation and control. It implements energy optimisation policies that selects the optimal CPU frequency based on runtime application characteristics and policy settings. Given that EARL defines a policy API and a plugin mechanism, different policies can be easily evaluated. In this paper we propose and evaluate the utilisation of explicit Uncore Frequency Scaling (explicit UFS) in Intel architectures to increase the energy savings opportunities in the cases where the hardware cannot select the optimal frequency for the Integrated Memory Controller (IMC). We extended the min_energy_to_solution policy to select the CPU and IMC frequencies and we executed and evaluated it with some kernels and six real applications. Results showed an average energy saving of 9% with an average time penalty of 3%. On some use cases, the impact of explicit UFS compared with HW UFS was up to 8% of extra energy savings.
Julita Corbalán, Oriol Vidal, Lluis Alonso, Jordi Aneas
CLUSTER1
2021 Modular Workload Format: Extending SWF for Modular Systems
Julita Corbalán, Marco D'Amico
JSSPP1
2020 Energy Optimization and Analysis with EAR
abstract
EAR is an energy management framework which offers three main services: energy accounting, energy control, and energy optimization. The latter is done through the EAR runtime library (EARL). EARL is a dynamic, transparent, and lightweight runtime library that provides energy optimisation and control. EARL optimises energy by selecting the optimal CPU frequency, based on the energy policy selected and application runtime characteristics without any application modification or user input. Currently EARL only works for MPI applications but EAR itself can still operate for non-MPI applications. It automatically (and transparently) identifies iterative regions (loops) and computes a set of metrics per iteration, application signature, and, together with the system signature, applies energy models to estimate the execution time and power for the CPU frequencies available. System signature is a set of coefficients per-node computed during EAR installation via a learning phase. Given time and power projections, EARL selects the best frequency based on policy settings. This papers shows how to optimize energy using the EAR library with min_time_to_solution energy policy and how to analyse applications through EAR framework. Evaluation includes eight applications with different sizes and application signatures. Results show how EARL computes each application signature on the fly and applies the CPU frequency selected by the min_time_to_solution policy.
Julita Corbalán, Lluis Alonso, Jordi Aneas, Luigi Brochard
CLUSTER1
2019 Holistic Slowdown Driven Scheduling and Resource Management for Malleable Jobs
abstract
In job scheduling, the concept of malleability has been explored since many years ago. Research shows that malleability improves system performance, but its utilization in HPC never became widespread. The causes are the difficulty in developing malleable applications, and the lack of support and integration of the different layers of the HPC software stack. However, in the last years, malleability in job scheduling is becoming more critical because of the increasing complexity of hardware and workloads. In this context, using nodes in an exclusive mode is not always the most efficient solution as in traditional HPC jobs, where applications were highly tuned for static allocations, but offering zero flexibility to dynamic executions. This paper proposes a new holistic, dynamic job scheduling policy, Slowdown Driven (SD-Policy), which exploits the malleability of applications as the key technology to reduce the average slowdown and response time of jobs. SD-Policy is based on backfill and node sharing. It applies malleability to running jobs to make room for jobs that will run with a reduced set of resources, only when the estimated slowdown improves over the static approach. We implemented SD-Policy in SLURM and evaluated it in a real production environment, and with a simulator using workloads of up to 198K jobs. Results show better resource utilization with the reduction of makespan, response time, slowdown, and energy consumption, up to respectively 7%, 50%, 70%, and 6%, for the evaluated workloads.
Marco D'Amico, Ana Jokanovic, Julita Corbalán
ICPP3
2017 DJSB: Dynamic Job Scheduling Benchmark
Víctor López 0003, Ana Jokanovic, Marco D'Amico, Marta Garcia-Gasulla, Raül Sirvent, Julita Corbalán
JSSPP6
2014 Evaluation and assessment of professional skills in the Final Year Project
abstract
In this paper, we present a methodology for Final Year Project (FYP) monitoring and assessment that considers the inclusion of the professional skills required in the particular engineering degree. This proper monitoring and clear evaluation framework provides the student with valuable support for the project implementation as well as for improving the quality of the projects, thereby reducing the academic drop-out rate. The proposed methodology has been implemented at the Barcelona School of Informatics at the Universität Politècnica de Catalunya - BarcelonaTech. The FYP is structured around three milestones: project definition, project monitoring and project completion. Skills are assigned to each milestone according to the tasks required in that phase, and a list of indicators is defined for each phase. The evaluation criteria for each indicator at each phase are specified in a rubric, and are made public both to students and teachers. Thus, the FYP includes an exhaustive evaluation method distributed throughout the whole project implementation, thereby facilitating project organization for the student as well as providing a clear and homogeneous assessment framework. The methodology for the FYP organization, assessment and evaluation was launched and piloted over two semesters. We believe the experience to be general in the sense that it has been conducted as part of an ICT engineering degree, but may easily be extended to any other engineering degree.
Fermín Sánchez, Joan Climent, Julita Corbalán, Pau Fonseca i Casas, Jordi Garcia 0001, José R. Herrero 0001, Xavier Llinas, Horacio Rodríguez, Maria-Ribera Sancho, Marc Alier Forment, Jose Cabré, David López 0001
FIE3
2014 Hints to improve automatic load balancing with LeWI for hybrid applications
Marta Garcia-Gasulla, Jesús Labarta, Julita Corbalán
J. Parallel Distributed Comput.3
2014 Scheduling parallel jobs on multicore clusters using CPU oversubscription
Gladys Utrera, Julita Corbalán, Jesús Labarta
J. Supercomput.2
2012 A Job Scheduling Approach for Multi-core Clusters Based on Virtual Malleability
Gladys Utrera, Siham Tabik, Julita Corbalán, Jesús Labarta
Euro-Par3
2012 Understanding the future of energy-performance trade-off via DVFS in HPC environments
Maja Etinski, Julita Corbalán, Jesús Labarta, Mateo Valero
J. Parallel Distributed Comput.2
2012 Parallel job scheduling for power constrained HPC systems
Maja Etinski, Julita Corbalán, Jesús Labarta, Mateo Valero
Parallel Comput.2
2010 Enabling GPU and Many-Core Systems in Heterogeneous HPC Environments Using Memory Considerations
abstract
Increasing the utilization of many-core systems has been one of the forefront topics these last years. Although many-cores architectures were merely theoretical models few years ago, they have become an important part of the high performance computing market. The semiconductor industry has developed Graphical Processing Units (GPU) systems that provide access to many cores (i.e: Larrabee, Fermi or Tesla) that can be used for General Purpose (GP) computing. In this paper, we propose and evaluate a scheduling strategy for GPU and many-core architectures for HPC environments. Specifically, our strategy is a variant of the backfilling scheduling policy with resource sharing considerations. We propose a scheduling strategy that considers the differences between GP processors and GPU computing elements in terms of computational capacity and memory bandwidth. To do this, our approach uses a resource model that predicts how shared resources are used in both GP and GPU/many-core elements. Furthermore, it considers the differences between these elements in terms of performance. First, it models their differences in terms of computational power and how they share the access to the node's memory bandwidth. Second, it characterizes how the processes are allocated to the GPU. Using this resource model, we design the Power Aware resource selection policy, which we combine with the LessConsume scheduling policy. Our strategy tries to allocate jobs aiming at reducing the memory contention and the energy consumption. Results show that the scheduling strategies proposed in this work are able to save over 40% of energy and improve the system performance up to 30% with respect to traditional backfilling strategies.
Francesc Guim 0001, Ivan Rodero, Julita Corbalán, Manish Parashar
HPCC3
2010 Grid broker selection strategies using aggregated resource information
Ivan Rodero, Francesc Guim 0001, Julita Corbalán, Liana L. Fong, Seyed Masoud Sadjadi
Future Gener. Comput. Syst.3
2009 Evaluation of Coordinated Grid Scheduling Strategies
abstract
Grid computing has emerged as a way to share geographically and organizationally distributed resources that may belong to different institutions or administrative domains. In this context, the scheduling and resource management is usually performed by a grid resource broker. The scheduling task consists of distributing the jobs among the different centers resources and the need to coordinate the grid with the underlying scheduling levels which have already been identified. However, there is still a lack of policies for this approach. In this paper we describe and evaluate our coordinated grid scheduling strategy. We take as a reference the FCFS job scheduling policy and the matchmaking approach for the resource selection. We also present a new job scheduling policy based on backfilling (JR-backfilling) that aims to improve the workloads execution performance, avoiding starvation and the SLOW-coordinated resource selection policy that considers the average bounded slowdown of the resources as the main parameter to perform the resource selection. From our evaluation, based on trace-driven simulations of real grid systems, we state that our proposed coordinated strategy can substantially improve the workloads execution performance as well as the resource utilization.
Ivan Rodero, Francesc Guim 0001, Julita Corbalán
HPCC3
2009 LeWI: A Runtime Balancing Algorithm for Nested Parallelism
abstract
We present LeWI: a novel load balancing algorithm, that can balance applications with very different patterns of imbalance. Our algorithm can balance fine grain imbalances, non iterative applications and applications with irregular imbalance. To achieve this LeWI reassigns the computational resources of blocked processes to other processes more loaded. We have implemented LeWI within DLB a Dynamic Load Balancing Library developed by us. DLB helps parallel programming models to make the most of the computational power available with the minimum effort. It solves the imbalance among processes in applications with two levels of parallelism using the malleability of the inner level. The performance evaluation shows that LeWI, the novel balancing algorithm we are presenting in this paper, together with DLB is able to improve the performance of a different range of unbalanced applications and when applied to well balanced applications it does not introduce significant overhead. Therefore we present a mechanism that can be used with any hybrid application without needing a programmer to analyze the application nor modify it.
Marta Garcia-Gasulla, Julita Corbalán, Jesús Labarta
ICPP2
2009 Broker Selection Strategies in Interoperable Grid Systems
abstract
The increasing demand for resources of the high performance computing systems has led to new forms of collaboration of distributed systems such as interoperable grid systems that contain and manage their own resources. While with a single grid domain one of the most important tasks is the selection of the most appropriate set of resources to dispatch a job, in an interoperable grid environment this problem shifts to selecting the most appropriate domain containing the requiring resources for the job. In this paper, we present and evaluate broker selection strategies for interoperable grid systems. They use aggregated resource information as well as dynamic performance information of the underlying scheduling layers. From our evaluations performed with simulation tools, we conclude that aggregation techniques do not penalize performance significantly, and that delegating part of the scheduling responsibilities to the underlying scheduling layers is a good way to balance the load among the different grid systems.
Ivan Rodero, Francesc Guim 0001, Julita Corbalán, Liana L. Fong, Seyed Masoud Sadjadi
ICPP3
2009 Power-aware load balancing of large scale MPI applications
abstract
Power consumption is a very important issue for HPC community, both at the level of one application or at the level of whole workload. Load imbalance of a MPI application can be exploited to save CPU energy without penalizing the execution time. An application is load imbalanced when some nodes are assigned more computation than others. The nodes with less computation can be run at lower frequency since otherwise they have to wait for the nodes with more computation blocked in MPI calls. A technique that can be used to reduce the speed is Dynamic Voltage Frequency Scaling (DVFS). Dynamic power dissipation is proportional to the product of the frequency and the square of the supply voltage, while static power is proportional to the supply voltage. Thus decreasing voltage and/or frequency results in power reduction. Furthermore, over-clocking can be applied in some CPUs to reduce overall execution time. This paper investigates the impact of using different gear sets, over-clocking, and application and platform properties to reduce CPU power. A new algorithm applying DVFS and CPU over-clocking is proposed that reduces execution time while achieving power savings comparable to prior work. The results show that it is possible to save up to 60% of CPU energy in applications with high load imbalance. Our results show that six gear sets achieve, on average, results close to the continuous frequency set that has been used as a baseline.
Maja Etinski, Julita Corbalán, Jesús Labarta, Mateo Valero, Alexander V. Veidenbaum
IPDPS2
2009 The Resource Usage Aware Backfilling
Francesc Guim 0001, Ivan Rodero, Julita Corbalán
JSSPP3
2008 Coordinated Co-allocation Scheduling on Heterogeneous Clusters of SMPs
abstract
Job scheduling research for parallel systems has been widely exploited in recent years, especially in centers with high performance computing facilities. In the recent past we presented the eNANOS execution environment which is based on a coordinated architecture, from the CPU allocation to the grid scheduling, providing a good low level support to perform an efficient high level scheduling. In this paper we present and evaluate the multi-node scheduling configuration of eNANOS that is implemented through the eNANOS Scheduler. Moreover, we introduce our scheduling strategy based on co-allocation and the coordination with dynamic processor allocation techniques. Finally, through experimental evaluation we state that our architecture and scheduling strategy can improve the applications execution and the system performance on heterogeneous clusters composed of SMP architectures.
Ivan Rodero, Julita Corbalán
eScience2
2008 Modeling and Evaluating Interoperable Grid Systems
abstract
Grid resource management tools have evolved from manual discovery and job submission to sophisticated brokering solutions. User requirements have created certain properties that resource managers have learned to support. This development is still continuing, and users already find it difficult to distinguish brokers and to migrate their applications when they move to a different grid. Moreover, new architectures are continuously being proposed, such as multi-site and interoperable grid systems. This paper presents the Alvio simulation framework which is designed to evaluate job scheduling strategies in complex HPC infrastructures. The main contribution of this simulator is that it allows modeling from local systems to interoperable grid scenarios. We also present an evaluation of multi-site and grid interoperable systems which shows the effect of job forwarding between different brokers.
Ivan Rodero, Francesc Guim 0001, Julita Corbalán
eScience3
2008 Balancing HPC applications through smart allocation of resources in MT processors
abstract
Many studies have shown that load imbalancing causes significant performance degradation in high performance computing (HPC) applications. Nowadays, multi-threaded (MT1) processors are widely used in HPC for their good performance/energy consumption and performance/cost ratios achieved sharing internal resources, like the instruction window or the physical register. Some of these processors provide the software hardware mechanisms for controlling the allocation of processor's internal resources. In this paper, we show, for the first time, that by appropriately using these mechanisms, we are able to control the tasks speed, reducing the imbalance in parallel applications transparently to the user and, hence, reducing the total execution time. Our results show that our proposal leads to a performance improvement up to 18% for one of the NAS benchmark. For a real HPC application (much more dynamic than the benchmark) the performance improvement is 8.1%. Our results also show that, if resource allocation is not used properly, the imbalance of applications is worsened causing performance loss.
Carlos Boneti, Roberto Gioiosa, Francisco J. Cazorla, Julita Corbalán, Jesús Labarta, Mateo Valero
IPDPS4
2008 An adaptive cut-off for task parallelism
abstract
In task parallel languages, an important factor for achieving a good performance is the use of a cut-off technique to reduce the number of tasks created. Using a cut-off to avoid an excessive number of tasks helps the runtime system to reduce the total overhead associated with task creation, particularlt if the tasks are fine grain. Unfortunately, the best cut-off technique its usually dependent on the application structure or even the input data of the application. We propose a new cut-off technique that, using information from the application collected at runtime, decides which tasks should be pruned to improve the performance of the application. This technique does not rely on the programmer to determine the cut-off technique that is best suited for the application. We have implemented this cut-off in the context of the new OpenMP tasking model. Our evaluation, with a variety of applications, shows that our adaptive cut-off is able to make good decisions and most of the time matches the optimal cut-off that could be set by hand by a programmer.
Alejandro Duran, Julita Corbalán, Eduard Ayguadé
SC2
2007 A Job Self-scheduling Policy for HPC Infrastructures
Francesc Guim 0001, Julita Corbalán
JSSPP2
2007 Modeling the Impact of Resource Sharing in Backfilling Policies using the Alvio Simulator
abstract
Job scheduling policies for HPC centers have been extensively studied during these last years, specially backfilling based policies. Almost all of these studies have been done using simulation tools. These tools evaluate the performance of scheduling policies using the workloads and the resource definition as an input. To the best of our knowledge, all the existent simulators use the runtime (either requested or real) provided in the workload as a basis of their simulations. However, the runtime of a job, even executed with a fixed number of processors, depends on runtime issues such as the specific resource selection policy used for allocate the jobs or the resource jobs requirements. This paper is the first part of a more complex research project that analyzes the impact in the system performance of considering the resource sharing of running jobs. With this purpose we have included in our job scheduler simulator (the Alvio simulator) a performance model that estimates the penalty introduced in the application runtime when sharing the memory bandwidth. Experiments have been conducted with two resource selection policies and we present both the impact from the point of view of global performance metrics, such as average slowdown, and per job impact such as percentage of penalized runtime.
Francesc Guim 0001, Julita Corbalán, Jesús Labarta
MASCOTS2
2007 Prediction f Based Models for Evaluating Backfilling Scheduling Policies
abstract
The research on the usage of prediction techniques in HPC scheduling policies rather than user estimates has increased it relevance these recent years. In the coming scheduling architectures, like grids and very heterogeneous computational resources, such techniques are having a crucial relevance due to users in most of the cases will not have enough information or enough skills for specify for how long will their jobs run. Many studies have analyzed the impact of the user runtime estimates accuracy in the performance of the scheduling policies. Using user runtime estimation models, such as the f-model, researchers have evaluated how the accuracy of the runtime estimates provided by the user at the job submission can affect the performance of the backfilling policies and its variants. However, these traditional estimation models can not applied to backfilling scheduling policies that use runtime predictions rather than user estimates. Clearly, predictions can not be characterized with these models. For instance because the underestimation of the runtime is not considered by them and obviously it can occurs. In this paper we describe and evaluate a set of f-model based prediction models that characterize the behavior that prediction techniques have shown in HPC centers. They have been designed for evaluate scheduling policies that use predictions rather than user estimates.
Francesc Guim 0001, Julita Corbalán, Jesús Labarta
PDCAT2
2006 Uniform Job Monitoring using the HPC-Europa Single Point of Access
Francesc Guim 0001, Ivan Rodero, Julita Corbalán, Jesús Labarta, Ariel Oleksiak, Tomasz Kuczynski, Dawid Szejnfeld, Jarek Nabrzyski
CCGRID3
2006 How the JSDL can Exploit the Parallelism?
abstract
The description of the jobs is a very important issue for the scheduling and management of grid jobs. Since there are a lot of different languages for describing grid jobs, the GGF have presented the Job Submission Description Language (JSDL) to standardize the job submission language. We believe that the JSDL is a good solution but it has some deficiencies regarding the parallelism issues. In this paper, we propose an extension of the JSDL to specify the parallelism details of grid jobs. This extension is proposed in general terms for supporting current multilevel parallel applications and incoming approaches in parallel programming models. We also discus the suitability of the multilevel parallel programming models for grids, in particular the MPI+OpenMP since our project, the eNANOS project, is based on this hybrid programming model.
Ivan Rodero, Francesc Guim 0001, Julita Corbalán, Jesús Labarta
CCGRID3
2005 Automatic thread distribution for nested parallelism in OpenMP
abstract
OpenMP is becoming the standard programming model for shared-memory parallel architectures. One of its most interesting features in the language is the support for nested parallelism. Previous research and parallelization experiences have shown the benefits of using nested parallelism as an alternative to combining several programming models such as MPI and OpenMP. However, all these works rely on the manual definition of an appropriate distribution of all the available thread across the different levels of parallelism. Some proposals have been made to extend the OpenMP language to allow the programmers to specify the thread distribution.This paper proposes a mechanism to dynamically compute the most appropriate thread distribution strategy. The mechanism is based on gathering information at runtime to derive the structure of the nested parallelism. This information is used to determine how the overall computation is distributed between the parallel branches in the outermost level of parallelism, which is constant in this work. According to this, threads in the innermost level of parallelism are distributed.The proposed mechanism is evaluated in two different environments: a research environment, the Nanos OpenMP research platform, and a commercial environment, the IBM XL runtime library. The performance numbers obtained validate the mechanism in both environments and they show the importance of selecting the proper amount of parallelism in the outer level.
Alejandro Duran, Marc González 0001, Julita Corbalán
ICS3
2005 Another approach to backfilled jobs: applying virtual malleability to expired windows
abstract
An efficient job scheduling must ensure high throughput and good performance. Moreover in highly parallel systems where processors are a critical resource, high machine utilization becomes an essential aspect.Backfilling consists on moving jobs ahead in the queue, given that they do not delay certain previously submitted jobs. When the execution time of a backfilled job was underestimated, some action has to be taken with it: abort, suspend/resume, checkpoint/restart, remain executing.In this paper we propose an alternative choice for that situation which consists on apply Virtual Malleability to the backfilled job. This means that its processors partition will be reduced, and as MPI jobs aren't really malleable, we make the job contend with itself for the use of processors by applying Co-scheduling. In this way resources are freed and the job at the head of the queue have a chance to start executing. In addition to this, as MPI parallel jobs can be Moldable, we add this possibility to the scheme.We obtained better performance than traditional backfilling in about 25 %, especially in high machine utilization. We claim also for the portability of our technique which does not requires special support from the operating system as checkpointing does.
Gladys Utrera, Julita Corbalán, Jesús Labarta
ICS2
2005 Performance-Driven Processor Allocation
abstract
In current multiprogrammed multiprocessor systems, to take into account the performance of parallel applications is critical to decide an efficient processor allocation. In this paper, we present the performance-driven processor allocation policy (PDPA). PDPA is a new scheduling policy that implements a processor allocation policy and a multiprogramming-level policy, in a coordinated way, based on the measured application performance. With regard to the processor allocation, PDPA is a dynamic policy that allocates to applications the maximum number of processors to reach a given target efficiency. With regard to the multiprogramming level, PDPA allows the execution of a new application when free processors are available and the allocation of all the running applications is stable, or if some applications show bad performance. Results demonstrate that PDPA automatically adjusts the processor allocation of parallel applications to reach the specified target efficiency, and that it adjusts the multiprogramming level to the workload characteristics. PDPA is able to adjust the processor allocation and the multiprogramming level without human intervention, which is a desirable property for self-configurable systems, resulting in a better individual application response time.
Julita Corbalán, Xavier Martorell, Jesús Labarta
IEEE Trans. Parallel Distributed Syst.1
2004 Scheduling of MPI Applications: Self-co-scheduling
Gladys Utrera, Julita Corbalán, Jesús Labarta
Euro-Par2
2004 Dynamic Load Balancing of MPI+OpenMP Applications
abstract
The hybrid programming model MPI+OpenMP are useful to solve the problems of load balancing of parallel applications independently of the architecture. Typical approaches to balance parallel applications using two levels of parallelism or only MPI consist of including complex codes that dynamically detect which data domains are more computational intensive and either manually redistribute the allocated processors or manually redistribute data. This approach has two drawbacks: it is time consuming and it requires an expert in application analysis. In this paper we present an automatic and dynamic approach for load balancing MPI+OpenMP applications. The system calculates the percentage of load imbalance and decides a processor distribution for the MPI processes that eliminates the computational load imbalance. Results show that this method can balance effectively applications without analyzing nor modifying them and that in the cases that the application was well balanced does not incur in a great overhead for the dynamic instrumentation and analysis realized.
Julita Corbalán, Alejandro Duran, Jesús Labarta
ICPP1
2003 Evaluation of the memory page migration influence in the system performance: the case of the SGI O2000
abstract
Current shared-memory multiprocessor CC-NUMA architectures provide a global address space to applications by hardware. However, even though the memory is virtually shared, it is actually physically distributed. Since memory nodes are distributed across the system, the cost of the memory accesses depends on the distance between the node that accesses the data and the node that physically contains the data. To reduce the impact of a bad initial memory placement, some operating systems offer a dynamic memory migration mechanism.In this paper, we want to demonstrate that memory migration mechanisms are a useful approach, but that their performance depends more on related issues, such as the processor scheduling, than on the mechanism itself. To show that, we evaluate the case of the automatic memory migration mechanism provided by IRIX, in Origin systems.We have evaluated several workloads of OpenMP applications under different system conditions such as the processor scheduling policy or the system load. In particular, we have focused on the effects of the page migration mechanism on the CPU time consumed by each application, the processor allocation received, and the speedup, when applying performance-driven scheduling policies.Results show that, if the scheduler is memory conscious, that is, it maintains as much as possible the system stable, the automatic memory page migration mechanism provided by IRIX will improve the execution time of OpenMPapplications. Experiments also show that the combination of performance-driven policies and the memory migration mechanism results in a system that can be automatically self-evaluated and self-configured.
Julita Corbalán, Xavier Martorell, Jesús Labarta
ICS1
2001 Improving Gang Scheduling through job performance analysis and malleability
abstract
The OpenMP programming model provides parallel applications a very important feature: job malleability. Job malleability is the capacity of an application to dynamically adapt its parallelism to the number of processors allocated to it. We believe that job malleability provides to applications the flexibility that a system needs to achieve its maximum performance. We also defend that a system has to take its decisions not only based on user requirements but also based on run-time performance measurements to ensure the efficient use of resources. Job malleability is the application characteristic that makes possible the run-time performance analysis. Without malleability applications would not be able to adapt their parallelism to the system decisions. To support these ideas, we present two new approaches to attack the two main problems of Gang Scheduling: the excessive number of time slots and the fragmentation. Our first proposal is to apply a scheduling policy inside each time slot of Gang Scheduling to distribute processors among applications considering their efficiency, calculated based on run-time measurements. We call this policy Performance-Driven Gang Scheduling. Our second approach is a new re-packing algorithm, Compress&Join, that exploits the job malleability. This algorithm modifies the processor allocation of running applications to adapt it to the system necessities and minimize the fragmentation and number of time slots. These proposals have been implemented in a SGI Origin 2000 with 64 processors. Results show the validity and convenience of both, to consider the job performance analysis calculated at run-time to decide the processor allocation, and to use a flexible programming model that adapts applications to system decisions.
Julita Corbalán, Xavier Martorell, Jesús Labarta
ICS1
2001 Improving Processor Allocation through Run-Time Measured Efficiency
abstract
In a multiprocessor architecture it is very important to allocate processors to applications in a proportional way to the performance that applications are achieving, not considering this performance can result in an under-utilization of the multiprocessor and also it can slowdown the execution time of parallel applications. However the performance of parallel applications is not known before their execution. In this work, we propose to use dynamically measured application efficiency of OpenMP applications to improve the performance of two scheduling policies proposed so far, the equipartition and the equal efficiency. The modified scheduling policies will request parallel applications to achieve a target efficiency to receive more processors. We refer to the modified equipartition and equal efficiency as equip++ and equal eff++. We also propose to use a dynamic multiprogramming level to avoid the under-utilization of the machine introduced by these new scheduling policies when using a static multiprogramming level. We have evaluated this work by executing several workloads in an SGI Origin2000 with 64 processors. Results show that the combination of (target efficiency+dynamic multiprogramming level) achieves, in the worst case, the same performance as the equipartition and the equal efficiency, and in the best case it achieves a speedup of up to 1.3 in individual applications and in specific workloads a speedup of up to 2.5, with respect to the original algorithms.
Julita Corbalán, Jesús Labarta
IPDPS1
2001 A Dynamic Periodicity Detector: Application to Speedup Computation
abstract
We propose a dynamic periodicity detector (DPD) for the estimation of periodicities in data series obtained from the execution of applications. We analyze the algorithm used by the periodicity detector and its performance on a number of data streams. It is shown how the periodicity detector is used for the segmentation and prediction of data streams. In an application case we describe how the periodicity detector is applied to the dynamic detection of iterations in parallel applications, where the detected segments are evaluated by a speedup computation tool. We test the performance of the periodicity detector on a number of parallelized benchmarks. The periodicity detector correctly identifies the iterations of parallel structures also in the case where the application has nested parallelism. In our implementation we measure only a negligible overhead produced by the periodicity detector. We find the DPD to be useful and suitable for the incorporation in dynamic optimization tools.
Felix Freitag, Julita Corbalán, Jesús Labarta
IPDPS2
2000 A Tool to Schedule Parallel Applications on Multiprocessors: The NANOS CPU MANAGER
Xavier Martorell, Julita Corbalán, Dimitrios S. Nikolopoulos, Nacho Navarro, Eleftherios D. Polychronopoulos, Theodore S. Papatheodorou, Jesús Labarta
JSSPP2
2000 Performance-Driven Processor Allocation
Julita Corbalán, Xavier Martorell, Jesús Labarta
OSDI1
1999 Thread fork/join techniques for multi-level parallelism exploitation in NUMA multiprocessors
abstract
This paper presents some techniques for efficient thread forking and joining in parallel execution environments, taking into consideration the physical structure of NUMA machines and the support for multi-level parallelization and processor grouping.Two work generation schemes and one join mechanism are designed, implemented, evaluated and compared with the ones used in the IFUX MP library, an efficient implementation which supports a single level of parallelism.Supporting multiple levels of parallelism is a current research goal, both in shared and distributed memory machines.Our proposals include a first work generation scheme (GWD, or global work descriptor) which supports multiple levels of parallelism, but not processor grouping.The second work generation scheme (LWD, or local work descriptor) has been designed to support multiple levels of parallelism and processor grouping.Processor grouping is needed to distribute processors among different parts of the computation and maintain the working set of each processor across different parallel constructs.The mechanisms are evaluated using synthetic benchmarks, two SPEC95fp applications and one NAS application.The performance evaluation concludes that: i) the overhead of the proposed mechanisms is similar to the overhead of the existing ones when exploiting a single level of parallelism, and ii) a remarkable improvement in performance is obtained for applications that have multiple levels of parallelism.The comparison with the traditional single-level parallelism exploitation gives an improvement in the range of 3065% for these applications.
Xavier Martorell, Eduard Ayguadé, Nacho Navarro, Julita Corbalán, Marc González 0001, Jesús Labarta
International Conference on Supercomputing4