Ivan Rodero

dblp:85/1328 · DBLP profile ↗
← Back
43ranked-venue papers
11as first author
3since 2021 · last 2026
0000-0002-0675-6728ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 9 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-authorSoftware engineering, systems software and programming languages · 5 · 2 first-authorArtificial intelligence and machine learning · 3Databases, data management, data science and information retrieval · 2Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
YearPublicationVenuePosition
2026 interTwin: Advancing Scientific Digital Twins through AI, Federated Computing and Data
abstract
Data will be made available on request.
Andrea Manzi, Raul Bardaji, Ivan Rodero, Germán Moltó, Sandro Fiore, Isabel Campos Plasencia, Donatello Elia, Francesco Sarandrea, A. Paul Millar, Daniele Spiga, Matteo Bunino, Gabriele Accarino, Lorenzo Asprea, Samuel Bernardo, Miguel Caballer, Charis Chatzikyriakou, Diego Ciangottini, Michele Claus, Andrea Cristofori, Davide Donno, Emanuele Donno, Iacopo Ferrario, Massimiliano Fronza, Alexander W. Jacob, Javad Komijani, Marina Krstic Marinkovic, Federica Legger, Ivan Palomo, Estíbaliz Parcero, Rakesh Sarma, Gaurav Sinha Ray, Sara Vallero, Juraj Zvolensky
Future Gener. Comput. Syst.3
2021 Facilitating Data Discovery for Large-scale Science Facilities using Knowledge Networks
abstract
Large-scale multiuser scientific facilities, such as geographically distributed observatories, remote instruments, and experimental platforms, represent some of the largest national investments and can enable dramatic advances across many areas of science. Recent examples of such advances include the detection of gravitational waves and the imaging of a black hole's event horizon. However, as the number of such facilities and their users grow, along with the complexity, diversity, and volumes of their data products, finding and accessing relevant data is becoming increasingly challenging, limiting the potential impact of facilities. These challenges are further amplified as scientists and application workflows increasingly try to integrate facilities' data from diverse domains. In this paper, we leverage concepts underlying recommender systems, which are extremely effective in e-commerce, to address these data-discovery and data-access challenges for large-scale distributed scientific facilities. We first analyze data from facilities and identify and model user-query patterns in terms of facility location and spatial localities, domain-specific data models, and user associations. We then use this analysis to generate a knowledge graph and develop the collaborative knowledge-aware graph attention network (CKAT) recommendation model, which leverages graph neural networks (GNNs) to explicitly encode the collaborative signals through propagation and combine them with knowledge associations. Moreover, we integrate a knowledge-aware neural attention mechanism to enable the CKAT to pay more attention to key information while reducing irrelevant noise, thereby increasing the accuracy of the recommendations. We apply the proposed model on two real-world facility datasets and empirically demonstrate that the CKAT can effectively facilitate data discovery, significantly outperforming several compelling state-of-the-art baseline models.
Yubo Qin, Ivan Rodero, Manish Parashar
IPDPS2
2021 Leveraging user access patterns and advanced cyberinfrastructure to accelerate data delivery from shared-use scientific observatories
Yubo Qin, Ivan Rodero, Anthony Simonet, Charles Meertens, Daniel Reiner, James Riley, Manish Parashar
Future Gener. Comput. Syst.2
2020 A Distributed Multi-Sensor Machine Learning Approach to Earthquake Early Warning
abstract
Our research aims to improve the accuracy of Earthquake Early Warning (EEW) systems by means of machine learning. EEW systems are designed to detect and characterize medium and large earthquakes before their damaging effects reach a certain location. Traditional EEW methods based on seismometers fail to accurately identify large earthquakes due to their sensitivity to the ground motion velocity. The recently introduced high-precision GPS stations, on the other hand, are ineffective to identify medium earthquakes due to its propensity to produce noisy data. In addition, GPS stations and seismometers may be deployed in large numbers across different locations and may produce a significant volume of data consequently, affecting the response time and the robustness of EEW systems.In practice, EEW can be seen as a typical classification problem in the machine learning field: multi-sensor data are given in input, and earthquake severity is the classification result. In this paper, we introduce the Distributed Multi-Sensor Earthquake Early Warning (DMSEEW) system, a novel machine learning-based approach that combines data from both types of sensors (GPS stations and seismometers) to detect medium and large earthquakes. DMSEEW is based on a new stacking ensemble method which has been evaluated on a real-world dataset validated with geoscientists. The system builds on a geographically distributed infrastructure, ensuring an efficient computation in terms of response time and robustness to partial infrastructure failures. Our experiments show that DMSEEW is more accurate than the traditional seismometer-only approach and the combined-sensors (GPS and seismometers) approach that adopts the rule of relative strength.
Kevin Fauvel, Daniel Balouek-Thomert, Diego Melgar, Pedro Silva 0007, Anthony Simonet, Gabriel Antoniu, Alexandru Costan, Véronique Masson, Manish Parashar, Ivan Rodero, Alexandre Termier
AAAI10
2020 Submarine: A subscription-based data streaming framework for integrating large facilities and advanced cyberinfrastructure
abstract
Summary Large scientific facilities provide researchers with instrumentation, data, and data products that can accelerate scientific discovery. However, increasing data volumes coupled with limited local computational power prevents researchers from taking full advantage of what these facilities can offer. Many researchers looked into using commercial and academic cyberinfrastructure (CI) to process these data. Nevertheless, there remains a disconnect between large facilities and CI that requires researchers to be actively part of the data processing cycle. The increasing complexity of CI and data scale necessitates new data delivery models, those that can autonomously integrate large‐scale scientific facilities and CI to deliver real‐time data and insights. In this paper, we present our initial efforts using the Ocean Observatories Initiative project as a use case. In particular, we present a subscription‐based data streaming service for data delivery that leverages the Apache Kafka data streaming platform. We also show how our solution can automatically integrate large‐scale facilities with CI services for automated data processing.
Ali Reza Zamani, Moustafa AbdelBaky, Daniel Balouek-Thomert, Juan J. Villalobos, Ivan Rodero, Manish Parashar
Concurr. Comput. Pract. Exp.5
2020 An edge-aware autonomic runtime for data streaming and in-transit processing
Ali Reza Zamani, Daniel Balouek-Thomert, Juan J. Villalobos, Ivan Rodero, Manish Parashar
Future Gener. Comput. Syst.4
2019 Optimizing Performance and Computing Resource Management of In-memory Big Data Analytics with Disaggregated Persistent Memory
abstract
The performance of modern Big Data frameworks, e.g. Spark, depends greatly on high-speed storage and shuffling, which impose a significant memory burden on production data centers. In many production situations, the persistence and shuffling intensive applications can suffer a major performance loss due to lack of memory. Thus, the common practice is usually to over-allocate the memory assigned to the data workers for production applications, which in turn reduces overall resource utilization. One efficient way to address the dilemma between the performance and cost efficiency of Big Data applications is through data center computing resource disaggregation. This paper proposes and implements a system that incorporates the Spark Big Data framework with a novel in-memory distributed file system to achieve memory disaggregation for data persistence and shuffling. We address the challenge of optimizing performance at affordable cost by co-designing the proposed in-memory distributed file system with large-volume DIMM-based persistent memory (PMEM) and RDMA technology. The disaggregation design allows each part of the system to be scaled independently, which is particularly suitable for cloud deployments. The proposed system is evaluated in a production-level cluster using real enterprise-level Spark production applications. The results of an empirical evaluation show that the system can achieve up to a 3.5- fold performance improvement for shuffle-intensive applications with the same amount of memory, compared to the default Spark setup. Moreover, by leveraging PMEM, we demonstrate that our system can effectively increase the memory capacity of the computing cluster with affordable cost, with a reasonable execution time overhead with respect to using local DRAM only.
Shouwei Chen, Xueyang Wu 0002, Zhen Fan 0009, Kunwu Huang, Peiyu Zhuang, Ivan Rodero, Manish Parashar, Dennis Z. Weng
CCGRID8
2019 Toward a Dynamic Network-Centric Distributed Cloud Platform for Scientific Workflows: A Case Study for Adaptive Weather Sensing
abstract
Computational science today depends on complex, data-intensive applications operating on datasets from a variety of scientific instruments. A major challenge is the integration of data into the scientist's workflow. Recent advances in dynamic, networked cloud resources provide the building blocks to construct reconfigurable, end-to-end infrastructure that can increase scientific productivity. However, applications have not adequately taken advantage of these advanced capabilities. In this work, we have developed a novel network-centric platform that enables high-performance, adaptive data flows and coordinated access to distributed cloud resources and data repositories for atmospheric scientists. We demonstrate the effectiveness of our approach by evaluating time-critical, adaptive weather sensing workflows, which utilize advanced networked infrastructure to ingest live weather data from radars and compute data products used for timely response to weather events. The workflows are orchestrated by the Pegasus workflow management system and were chosen because of their diverse resource requirements. We show that our approach results in timely processing of Nowcast workflows under different infrastructure configurations and network conditions. We also show how workflow task clustering choices affect throughput of an ensemble of Nowcast workflows with improved turnaround times. Additionally, we find that using our network-centric platform powered by advanced layer2 networking techniques results in faster, more reliable data throughput, makes cloud resources easier to provision, and the workflows easier to configure for operational use and automation.
Eric Lyons 0001, Anirban Mandal, George Papadimitriou 0002, Cong Wang 0014, Komal Thareja, Paul Ruth, Juan J. Villalobos, Ivan Rodero, Ewa Deelman, Michael Zink
eScience8
2018 Exploring the Potential of Next Generation Software-Defined in Memory Frameworks
abstract
As in-memory data analytics become increasingly important in a wide range of domains, the ability to develop large-scale and sustainable platforms faces significant challenges related to storage latency and memory size constraints. These challenges can be resolved by adopting new and effective formulations and novel architectures such as software-defined infrastructure. This paper investigates the key issue of data persistency for in-memory processing systems by evaluating persistence methods using different storage and memory devices for Apache Spark and the use of Alluxio. It also proposes and evaluates via simulation a Spark execution model for using disaggregated off-rack memory and non-volatile memory targeting next-generation software-defined infrastructure. Experimental results provide better understanding of behaviors and requirements for improving data persistence in current in-memory systems and provide data points to better understand requirements and design choices for next-generation software-defined infrastructure. The findings indicate that in-memory processing systems can benefit from ongoing software-defined infrastructure implementations; however current frameworks need to be enhanced appropriately to run efficiently at scale.
Shouwei Chen, Ivan Rodero
SBAC-PAD2
2018 Exploring Power Budget Scheduling Opportunities and Tradeoffs for AMR-Based Applications
abstract
Computational demand has brought major changes to Advanced Cyber-Infrastructure (ACI) architectures. It is now possible to run scientific simulations faster and obtain more accurate results. However, power and energy have become critical concerns. Also, the current roadmap toward the new generation of ACI includes power budget as one of the main constraints. Current research efforts have studied power and performance tradeoffs and how to balance these (e.g., using Dynamic Voltage and Frequency Scaling (DVFS) and power capping for meeting power constraints, which can impact performance). However, applications may not tolerate degradation in performance, and other tradeoffs need to be explored to meet power budgets (e.g., involving the application in making energy-performance-quality tradeoff decisions). This paper proposes using the properties of AMR-based algorithms (e.g., dynamically adjusting the resolution of a simulation in combination with power capping techniques) to schedule or re-distribute the power budget. It specifically explores the opportunities to realize such an approach using checkpointing as a proof-of-concept use case and provides a characterization of a representative set of applications that use Adaptive Mesh Refinement (AMR) methods, including a Low-Mach-Number Combustion (LMC) application. It also explores the potential of utilizing power capping to understand power-quality tradeoffs via simulation.
Yubo Qin, Ivan Rodero, Pradeep Subedi, Manish Parashar, Sandro Rigo
SBAC-PAD2
2018 Runtime Management of Data Quality for Scientific Observatories Using Edge and In-Transit Resources
abstract
Modern Cyberinfrastructures (CIs) operate to bring content produced from remote data sources such as sensors and scientific instruments and deliver it to end users and workflow applications. Maintaining data quality/resolution and on-time data delivery while considering an increasing number of computing, storage and network resources requires a reactive system, able to adapt to changing demands. In this paper, we propose a modelization of such system by expressing the dynamic stage of resources in the context of edge and in-transit computing. By considering resource utilization, approximation techniques and users' constraints, our proposed engine is generating mappings of workflow stages on heterogeneous geo-distributed resources. We specifically propose a runtime management layer that adapts the data resolution being delivered to the users by implementing feedback loops over the resources involved in the delivery and processing of the data streams. We implement our model into a subscription-based data streaming framework which enables integration of large facilities and advanced CIs. Experimental results show that dynamically adapting data resolution can overcome bandwidth limitation in wide area streaming analytics.
Ali Reza Zamani, Daniel Balouek-Thomert, Juan J. Villalobos, Ivan Rodero, Manish Parashar
SBAC-PAD4
2018 End-to-end energy models for Edge Cloud-based IoT platforms: Application to data stream analysis in IoT
Yunbo Li, Anne-Cécile Orgerie, Ivan Rodero, Betsegaw Lemma Amersho, Manish Parashar, Jean-Marc Menaud
Future Gener. Comput. Syst.3
2017 Understanding Behavior Trends of Big Data Frameworks in Ongoing Software-Defined Cyber-Infrastructure
abstract
As data analytics applications become increasingly important in a wide range of domains, the ability to develop large-scale and sustainable platforms and software infrastructure to support these applications has significant potential to drive research and innovation in both science and business domains. This paper characterizes performance and power-related behavior trends and tradeoffs of the two predominant frameworks for Big Data analytics (i.e., Apache Hadoop and Spark) for a range of representative applications. It also evaluates system design knobs, such as storage and network technologies and power capping techniques. Experimental results from empirical executions provide meaningful data points for exploring the potential of software-defined infrastructure for Big Data processing systems through simulation. The results provide better understanding of the design space to build multi-criteria application-centric models as well as show significant advantages of software-defined infrastructure in terms of execution time, energy and cost. It motivates further research focused on in-memory processing formulations regarding systems with deeper memory hierarchies and software-defined infrastructure.
Shouwei Chen, Ivan Rodero
BDCAT2
2017 An Unsupervised Approach for Online Detection and Mitigation of High-Rate DDoS Attacks Based on an In-Memory Distributed Graph Using Streaming Data and Analytics
abstract
A Distributed Denial of Service (DDoS) attack is an attempt to make an online service, a network, or even an entire organization, unavailable by saturating it with traffic from multiple sources. DDoS attacks are among the most common and most devastating threats that network defenders have to watch out for. DDoS attacks are becoming bigger, more frequent, and more sophisticated. Volumetric attacks are the most common types of DDoS attacks. A DDoS attack is considered volumetric, or high-rate, when within a short period of time it generates a large amount of packets or a high volume of traffic. High-rate attacks are well-known and have received much attention in the past decade; however, despite several detection and mitigation strategies have been designed and implemented, high-rate attacks are still halting the normal operation of information technology infrastructures across the Internet when the protection mechanisms are not able to cope with the aggregated capacity that the perpetrators have put together. With this in mind, the present paper aims to propose and test a distributed and collaborative architecture for online high-rate DDoS attack detection and mitigation based on an in-memory distributed graph data structure and unsupervised machine learning algorithms that leverage real-time streaming data and analytics. We have successfully tested our proposed mechanism using a real-world DDoS attack dataset at its original rate in pursuance of reproducing the conditions of an actual large scale attack.
Juan J. Villalobos, Ivan Rodero, Manish Parashar
BDCAT2
2017 Enabling Distributed Software-Defined Environments Using Dynamic Infrastructure Service Composition
abstract
Service-based access models coupled with emerging application deployment technologies are enabling opportunities for realizing highly customized software-defined environments, which can support dynamic and data-driven applications. However, this requires rethinking traditional resource federation models to support dynamic resource compositions, which can adapt to evolving application needs and the dynamic state of underlying resources. In this paper, we present a programmable approach that leverages software-defined techniques to create a dynamic space-time infrastructure service composition. We propose the use of Constraint Programming as a formal language to allow users, applications, and service providers to define the desired state of the execution environment. The resulting distributed software-defined environment continually adapts to meet objectives/constraints set by the users, applications, and/or resource providers. We present the design and prototype implementation of such distributed software-defined environment. We use a cancer informatics workflow to demonstrate the operation of our framework using resources from five different cloud providers, which are aggregated on-demand based on dynamic user and resource provider constraints.
Moustafa AbdelBaky, Javier Diaz Montes, Merve Unuvar, Melissa Romanus, Ivan Rodero, Malgorzata Steinder, Manish Parashar
CCGrid5
2017 Leveraging Renewable Energy in Edge Clouds for Data Stream Analysis in IoT
abstract
The emergence of Internet of Things (IoT) is participating to the increase of data-and energy-hungry applications. As connected devices do not yet offer enough capabilities for sustaining these applications, users perform computation offloading to the cloud. To avoid network bottlenecks and reduce the costs associated to data movement, edge cloud solutions have started being deployed, thus improving the Quality of Service. In this paper, we advocate for leveraging on-site renewable energy production in the different edge cloud nodes to green IoT systems while offering improved QoS compared to core cloud solutions. We propose an analytic model to decide whether to offload computation from the objects to the edge or to the core Cloud, depending on the renewable energy availability and the desired application QoS. This model is validated on our application use-case that deals with video stream analysis from vehicle cameras.
Yunbo Li, Anne-Cécile Orgerie, Ivan Rodero, Manish Parashar, Jean-Marc Menaud
CCGrid3
2017 Supporting Data-Driven Workflows Enabled by Large Scale Observatories
abstract
Large scale observatories are shared-use resources that provide open access to data from geographically distributed sensors and instruments. This data has the potential to accelerate scientific discovery. However, seamlessly integrating the data into scientific workflows remains a challenge. In this paper, we summarize our ongoing work in supporting data-driven and data-intensive workflows and outline our vision for how these observatories can improve large-scale science. Specifically, we present programming abstractions and runtime management services to enable the automatic integration of data in scientific workflows. Further, we show how approximation techniques can be used to address network and processing variations by studying constraint limitations and their associated latencies. We use the Ocean Observatories Initiative (OOI) as a driving use case for this work.
Ali Reza Zamani, Moustafa AbdelBaky, Daniel Balouek-Thomert, Ivan Rodero, Manish Parashar
eScience4
2017 WA-Dataspaces: Exploring the Data Staging Abstractions for Wide-Area Distributed Scientific Workflows
abstract
Data staging has been shown to be very effective for supporting data intensive in-situ workflows and coupling of applications. Experimental sciences are increasingly becoming collaborative among geographically distributed teams, and include experimental instruments and HPC facilities. This new way of doing science poses new challenges due to data sizes, complexity of computation, and the use of wide area networks between couplings. In this paper, we explore how the staging abstraction can be extended to support such workflows. Specifically, we develop a NUMA-like abstraction that orchestrates multiple distributed local-area staging abstractions, and provides asynchronous data put/get semantics to enable data sharing across them. To mask data movement overhead and provide in-time data access, we propose the use of predictive prefetching approaches that leverage the iterative nature of the coupling. We evaluate our prototype implementation using a fusion workflow and show that our design can effectively and transparently support widearea coupled workflows. Additionally, results show that the use of prefetching techniques leads to significant gains in data access times of data that needs to be moved over the wide area network.
Mehmet Fatih Aktas, Javier Diaz Montes, Ivan Rodero, Manish Parashar
ICPP3
2017 Modelling and Implementing Social Community Clouds
abstract
As the number of people who interact on social networks increases, and coupled with the greater capability made available within our computational devices, there is the potential to establish “Social Clouds”-a resource sharing infrastructure that enable people who have trust relationships to come together to share computational/ data services within a community. Social clouds can also provide the means to enhance multi-user collaboration and greatly stimulate the exchange of resources among participants. Recent research in the establishment and use of Social Clouds has raised significant interest by proposing an environment where users are able to trade resources mediated by a social networking mechanism. In such a cloud environment the incentives for sharing can represent a solution for improving resource utilisation and for making available additional capacity to friends and collaborators. In this paper we demonstrate how revenue can be earned within a social cloud community, by executing internal (intra community) and external (inter community) tasks. A number of different scenarios are first investigated through simulation, using the PeerSim simulator, in order to validate our approach. We use two key metrics: revenue and reputation, to evaluate how the system dynamics change as new tasks are added to one or more communities for execution, along with additional behaviours, such as nodes migrating from one community to another, or selectively reporting on the outcome of task execution. Subsequently, we develop a practical deployment using a federated cloud scenario using the CometCloud system-deployed over three sites: Cardiff (UK), Rutgers and Indiana. We show how approaches that have been simulated in PeerSim can be implemented in practice.
Ioan Petri, Javier Diaz Montes, Omer F. Rana, Magdalena Punceva, Ivan Rodero, Manish Parashar
IEEE Trans. Serv. Comput.5
2016 Evaluation of In-Situ Analysis Strategies at Scale for Power Efficiency and Scalability
abstract
The increasing gap between available compute power and I/O capabilities is resulting in simulation pipelines running on leadership computing facilities being reformulated. In particular, in-situ processing is complementing conventional post-process analysis, however, it can be performed by using the same compute resources as the simulation or using secondary dedicated resources. In this paper, we focus on three different in-situ analysis strategies, which use the same compute resources as the ongoing simulation but different data movement strategies. We evaluate the costs incurred by these strategies in terms of run time, scalability and power/energy consumption. Furthermore, we extrapolate power behavior to peta-scale and investigate different design choices through projections. Experimental evaluation at full machine scale on Titan supports that using fewer cores per node for in-situ analysis is the optimum choice in terms of scalability. Hence, further research effort should be devoted towards developing in-situ analysis techniques following this strategy in future high-end systems.
Ivan Rodero, Manish Parashar, Aaditya G. Landge, Sidharth Kumar, Valerio Pascucci, Peer-Timo Bremer
CCGrid1
2015 Incentivising resource sharing in social clouds
abstract
Summary Social Clouds provide the capability to share resources among participants within a social network—leveraging on the trust relationships already existing between such participants. In such a system, users are able to trade resources between each other rather than make use of capability offered at a (centralized) data center. Although such an environment has significant potential for improving resource utilization and making available additional capacity that remains dormant, incentives for sharing remain an important hurdle limiting its effective. In this paper, we utilize the socioeconomic model proposed by Silvio Gesell to demonstrate how a ‘virtual currency’ can be used to incentivise sharing of resources within a ‘community’. We subsequently demonstrate, through simulations, the benefit provided to participants within such a community using a variety of economic (such as overall credits gained) and technical (number of successfully completed transactions) metrics. Further, we describe our implementation of such a Social Cloud using CometCloud. CometCloud is an autonomic computing engine for cloud and grid environments. It supports highly heterogeneous and dynamic federated cloud/Grid infrastructures, integration of public/private clouds and autonomic cloudbursts. We demonstrate the implementation of two designs on the basis of the master/worker approach: (i) one tuple space per cluster and (ii) one coordination tuple space and multiple transient spaces—one per each cluster. Finally, we discuss an extended version of our Social Cloud model where intermediary relay nodes take on more active roles as traders in a transaction. Copyright © 2013 John Wiley & Sons, Ltd.
Magdalena Punceva, Ivan Rodero, Manish Parashar, Omer F. Rana, Ioan Petri
Concurr. Comput. Pract. Exp.2
2015 Uncertainty-Aware Autonomic Resource Provisioning for Mobile Cloud Computing
abstract
Mobile platforms are becoming the predominant medium of access to Internet services due to the tremendous increase in their computation and communication capabilities. However, enabling applications that require real-time in-the-field data collection and processing using mobile platforms is still challenging due to i) the insufficient computing capabilities and unavailability of complete data on individual mobile devices and ii) the prohibitive communication cost and response time involved in offloading data to remote computing resources such as cloud datacenters for centralized computation. A novel resource provisioning framework for organizing the heterogeneous sensing, computing, and communication capabilities of static and mobile devices in the vicinity in order to form an elastic resource pool—a hybrid static/mobile computing grid (also called a loosely-coupled mobile device cloud)—is presented. This local computing grid can be harnessed to enable innovative data- and compute-intensive mobile applications such as ubiquitous context-aware health and wellness monitoring of the elderly, distributed rainfall and flood-risk estimation, distributed object recognition and tracking, and content-based distributed multimedia search and sharing. In order to address challenges such as the inherent uncertainty in the hybrid grid (in terms of network connectivity and device availability), the proposed role-based resource provisioning framework is imparted with autonomic capabilities, namely, self-organization, self-optimization, and self-healing. A thorough experimental analysis aimed at verifying and demonstrating the benefits brought by autonomic capabilities of the framework is also presented in detail.
Hariharasudhan Viswanathan, Ivan Rodero, Dario Pompili
IEEE Trans. Parallel Distributed Syst.3
2014 Content-based histopathology image retrieval using CometCloud
abstract
BACKGROUND: The development of digital imaging technology is creating extraordinary levels of accuracy that provide support for improved reliability in different aspects of the image analysis, such as content-based image retrieval, image segmentation, and classification. This has dramatically increased the volume and rate at which data are generated. Together these facts make querying and sharing non-trivial and render centralized solutions unfeasible. Moreover, in many cases this data is often distributed and must be shared across multiple institutions requiring decentralized solutions. In this context, a new generation of data/information driven applications must be developed to take advantage of the national advanced cyber-infrastructure (ACI) which enable investigators to seamlessly and securely interact with information/data which is distributed across geographically disparate resources. This paper presents the development and evaluation of a novel content-based image retrieval (CBIR) framework. The methods were tested extensively using both peripheral blood smears and renal glomeruli specimens. The datasets and performance were evaluated by two pathologists to determine the concordance. RESULTS: The CBIR algorithms that were developed can reliably retrieve the candidate image patches exhibiting intensity and morphological characteristics that are most similar to a given query image. The methods described in this paper are able to reliably discriminate among subtle staining differences and spatial pattern distributions. By integrating a newly developed dual-similarity relevance feedback module into the CBIR framework, the CBIR results were improved substantially. By aggregating the computational power of high performance computing (HPC) and cloud resources, we demonstrated that the method can be successfully executed in minutes on the Cloud compared to weeks using standard computers. CONCLUSIONS: In this paper, we present a set of newly developed CBIR algorithms and validate them using two different pathology applications, which are regularly evaluated in the practice of pathology. Comparative experimental results demonstrate excellent performance throughout the course of a set of systematic studies. Additionally, we present and evaluate a framework to enable the execution of these algorithms across distributed resources. We show how parallel searching of content-wise similar images in the dataset significantly reduces the overall computational time to ensure the practical utility of the proposed CBIR algorithms.
Xin Qi 0007, Daihou Wang, Ivan Rodero, Javier Diaz Montes, Rebekah H. Gensure, Fuyong Xing, Lauri A. Goodell, Manish Parashar, David J. Foran, Lin Yang 0002
BMC Bioinform.3
2013 Exploring energy and performance behaviors of data-intensive scientific workflows on systems with deep memory hierarchies
abstract
The increasing gap between the rate at which large scale scientific simulations generate data and the corresponding storage speeds and capacities is leading to more complex system architectures with deep memory hierarchies. Advances in non-volatile memory (NVRAM) technology have made it an attractive candidate as intermediate storage in this memory hierarchy to address the latency and performance gap between main memory and disk storage. As a result, it is important to understand and model its energy/performance behavior from an application perspective as well as how it can be effectively used for staging data within an application workflow. In this paper, we target a NVRAM-based deep memory hierarchy and explore its potential for supporting in-situ/in-transit data analytics pipelines that are part of application workflows patterns. Specifically, we model the memory hierarchy and experimentally explore energy/performance behaviors of different data management strategies and data exchange patterns, as well as the tradeoffs associated with data placement, data movement and data processing.
Marc Gamell, Ivan Rodero, Manish Parashar, Stephen W. Poole
HiPC2
2013 Exploring power behaviors and trade-offs of in-situ data analytics
abstract
As scientific applications target exascale, challenges related to data and energy are becoming dominating concerns. For example, coupled simulation workflows are increasingly adopting in-situ data processing and analysis techniques to address costs and overheads due to data movement and I/O. However it is also critical to understand these overheads and associated trade-offs from an energy perspective. The goal of this paper is exploring data-related energy/performance trade-offs for end-to-end simulation workflows running at scale on current high-end computing systems. Specifically, this paper presents: (1) an analysis of the data-related behaviors of a combustion simulation workflow with an in-situ data analytics pipeline, running on the Titan system at ORNL; (2) a power model based on system power and data exchange patterns, which is empirically validated; and (3) the use of the model to characterize the energy behavior of the workflow and to explore energy/performance trade-offs on current as well as emerging systems.
Marc Gamell, Ivan Rodero, Manish Parashar, Janine Bennett, Hemanth Kolla, Jacqueline Chen, Peer-Timo Bremer, Aaditya G. Landge, Attila Gyulassy, Patrick S. McCormick, Scott Pakin, Valerio Pascucci, Scott Klasky
SC2
2013 Enabling Interoperability among Grid Meta-Schedulers
Ivan Rodero, David Villegas, Norman Bobroff, Liana L. Fong, Seyed Masoud Sadjadi
J. Grid Comput.1
2012 Accelerating MapReduce Analytics Using CometCloud
abstract
MapReduce-Hadoop has emerged as an effective framework for large-scale data analytics, providing support for executing jobs and storing data in a parallel and distributed manner. MapReduce has been shown to perform very well on large datacenters running applications where the data can be effectively divided into homogeneous chunks running across homogeneous hardware. However, the performance of MapReduceHadoop is far from ideal when either or both hardware and datasets are heterogeneous. Such heterogeneity is unavoidable in many academic computing environments that use multiple generations of hardware, and share resources among users. Heterogeneity is also unavoidable in scientific applications that process a varying number of datasets of different sizes. In these cases, the performance of MapReduce-Hadoop can be a concern. In this paper, we implement MapReduce on top of CometCloud to address the issue of heterogeneity and support applications classes that involve irregular datasets (e.g. large number of small data files or datasets of varying sizes). Furthermore, we develop an autonomic manager that can schedule MapReduce tasks based on user objective, provision resources accordingly, and support on-demand scale up and cloudbursts. These resources can be selected from a hybrid infrastructure such as local clusters, data centers, and public clouds. The performance of the developed solution is verified using a protein data mining application operating on data from the Protein Data Bank. The application is deployed, based on deadline and budget constraints, on a cluster at Rutgers and/or Amazon EC2 resources. The experimental results show that the MapReduce-CometCloud framework can effectively support applications operating on large numbers of small data files on a heterogeneous and distributed environment, and satisfy user objective autonomically using cloudbursts.
Moustafa AbdelBaky, Hyunjoo Kim, Ivan Rodero, Manish Parashar
IEEE CLOUD3
2012 Exploring cross-layer power management for PGAS applications on the SCC platform
abstract
High-performance parallel computing architectures are increasingly based on multi-core processors. While current commercially available processors are at 8 and 16 cores, technological and power constraints are limiting the performance growth of the cores and are resulting in architectures with much higher core counts, such as the experimental many-core Intel Single-chip Cloud Computer (SCC) platform. These trends are presenting new sets of challenges to HPC applications including programming complexity and the need for extreme energy efficiency.
Marc Gamell, Ivan Rodero, Manish Parashar, Rajeev Muralidhar
HPDC2
2012 Energy-Efficient Thermal-Aware Autonomic Management of Virtualized HPC Cloud Infrastructure
Ivan Rodero, Hariharasudhan Viswanathan, Marc Gamell, Dario Pompili, Manish Parashar
J. Grid Comput.1
2012 Cloud federation in a layered service model
David Villegas, Norman Bobroff, Ivan Rodero, Javier Delgado, Aditya Devarakonda, Liana L. Fong, Seyed Masoud Sadjadi, Manish Parashar
J. Comput. Syst. Sci.3
2011 Adaptive memory power management techniques for HPC workloads
abstract
The memory subsystem is responsible for a large fraction of the energy consumed by compute nodes in High Performance Computing (HPC) systems. The rapid increase in the number of cores has been accompanied by a corresponding increase in the DRAM capacity and bandwidth, and as a result, the memory system consumes a significant amount of the power budget available to a compute node. Consequently, there is a broad research effort focused on power management techniques using DRAM low-power modes. However, memory power management continues to present many challenges. In this paper, we study the potential of Dynamic Voltage and Frequency Scaling (DVFS) of the memory subsystems, and consider the ability to select different frequencies for different memory channels. Our approach is based on tuning voltage and frequency dynamically to maximize the energy savings while maintaining performance degradation within tolerable limits. We assume that HPC applications do not demand maximum bandwidth throughout the entire period of execution. We can use these low memory demand intervals to tune down the frequency and, as a result, applications can tolerate a reduction in bandwidth to save energy. In this paper, we study application channel access patterns, and use these patterns to determine potential additional energy savings that can be achieved by accordingly controlling the channels independently. We then evaluate the proposed DVFS algorithm using a novel hybrid evaluation methodology that includes simulation as well as executions on real hardware. Our results demonstrate the large potential of adaptive memory power management techniques based on DVFS for HPC workloads.
Karthik Elangovan, Ivan Rodero, Manish Parashar, Francesc Guim 0001, Isaac Hernandez
HiPC2
2010 Investigating the potential of application-centric aggressive power management for HPC workloads
abstract
Energy efficiency of large-scale data centers is becoming a major concern not only for reasons of energy conservation, failures, and cost reduction, but also because such sys tems are soon reaching the limits of power available to them. Like High Performance Computing (HPC) systems, large-scale clu ster-based data centers can consume power in megawatts, and of all the power consumed by such a system, only a fraction is used for actual computations. In this paper, we study the potential of application-centric aggressive power management of data center's resources for HPC workloads. Specifically, we consider power management mechanisms and controls (currently or soon to be) available at different levels and for different subsystems, and leverage several innovative approaches that have been taken to tackle this problem in the last few years, can be effectively used in a application-aware manner for HPC workloads. To do this, we first profile sta ndard HPC benchmarks with respect to behaviors, resource usage and power impact on individual computing nodes. Based on a power and latency model and the workload profiles, we develop an algorithm that can improve energy efficiency with little or no performance loss. We then evaluate our proposed algorithm through simulations using empirical power characterization and quantification. Finally, we validate the simulation results with actual executions on real hardware. The obtained results show that by using application aware power management, we can re-du ce the average energy consumption without significant penalty in performance. This motivates us to investigate autonomic approaches for application-aware aggressive power management and cross layer and cross function predictive subsystem level power management for large-scale data centers.
Ivan Rodero, Sharat Chandra, Manish Parashar, Rajeev Muralidhar, Harinarayanan Seshadri, Stephen W. Poole
HiPC1
2010 Enabling GPU and Many-Core Systems in Heterogeneous HPC Environments Using Memory Considerations
abstract
Increasing the utilization of many-core systems has been one of the forefront topics these last years. Although many-cores architectures were merely theoretical models few years ago, they have become an important part of the high performance computing market. The semiconductor industry has developed Graphical Processing Units (GPU) systems that provide access to many cores (i.e: Larrabee, Fermi or Tesla) that can be used for General Purpose (GP) computing. In this paper, we propose and evaluate a scheduling strategy for GPU and many-core architectures for HPC environments. Specifically, our strategy is a variant of the backfilling scheduling policy with resource sharing considerations. We propose a scheduling strategy that considers the differences between GP processors and GPU computing elements in terms of computational capacity and memory bandwidth. To do this, our approach uses a resource model that predicts how shared resources are used in both GP and GPU/many-core elements. Furthermore, it considers the differences between these elements in terms of performance. First, it models their differences in terms of computational power and how they share the access to the node's memory bandwidth. Second, it characterizes how the processes are allocated to the GPU. Using this resource model, we design the Power Aware resource selection policy, which we combine with the LessConsume scheduling policy. Our strategy tries to allocate jobs aiming at reducing the memory contention and the energy consumption. Results show that the scheduling strategies proposed in this work are able to save over 40% of energy and improve the system performance up to 30% with respect to traditional backfilling strategies.
Francesc Guim 0001, Ivan Rodero, Julita Corbalán, Manish Parashar
HPCC2
2010 Towards energy-aware autonomic provisioning for virtualized environments
abstract
As energy efficiency and associated costs become key concerns, consolidated and virtualized data centers and clouds are attractive computing platforms for data- and compute-intensive applications. Recently, these platforms are also being considered for more traditional high-performance computing (HPC) applications. However, maximizing energy efficiency, cost-effectiveness, and utilization for these applications while ensuring performance and other Quality of Service (QoS) guarantees, requires leveraging important and extremely challenging tradeoffs. These include, for example, the tradeoff between the need to efficiently create and provision Virtual Machines (VMs) on data center resources and the need to accommodate the heterogeneous resource demands and runtimes of the applications that run on them. In this paper we propose an energy-aware online provisioning approach for HPC applications on consolidated and virtualized computing platforms. Energy efficiency is achieved using a workload-aware, just-right dynamic provisioning mechanism and the ability to power down subsystems of a host system that are not required by the VMs mapped to it. Our preliminary evaluations show that our approach can improve energy efficiency with an acceptable QoS penalty.
Ivan Rodero, Juan Jaramillo, Andres Quiroz, Manish Parashar, Francesc Guim 0001
HPDC1
2010 Grid broker selection strategies using aggregated resource information
Ivan Rodero, Francesc Guim 0001, Julita Corbalán, Liana L. Fong, Seyed Masoud Sadjadi
Future Gener. Comput. Syst.1
2009 Evaluation of Coordinated Grid Scheduling Strategies
abstract
Grid computing has emerged as a way to share geographically and organizationally distributed resources that may belong to different institutions or administrative domains. In this context, the scheduling and resource management is usually performed by a grid resource broker. The scheduling task consists of distributing the jobs among the different centers resources and the need to coordinate the grid with the underlying scheduling levels which have already been identified. However, there is still a lack of policies for this approach. In this paper we describe and evaluate our coordinated grid scheduling strategy. We take as a reference the FCFS job scheduling policy and the matchmaking approach for the resource selection. We also present a new job scheduling policy based on backfilling (JR-backfilling) that aims to improve the workloads execution performance, avoiding starvation and the SLOW-coordinated resource selection policy that considers the average bounded slowdown of the resources as the main parameter to perform the resource selection. From our evaluation, based on trace-driven simulations of real grid systems, we state that our proposed coordinated strategy can substantially improve the workloads execution performance as well as the resource utilization.
Ivan Rodero, Francesc Guim 0001, Julita Corbalán
HPCC1
2009 Broker Selection Strategies in Interoperable Grid Systems
abstract
The increasing demand for resources of the high performance computing systems has led to new forms of collaboration of distributed systems such as interoperable grid systems that contain and manage their own resources. While with a single grid domain one of the most important tasks is the selection of the most appropriate set of resources to dispatch a job, in an interoperable grid environment this problem shifts to selecting the most appropriate domain containing the requiring resources for the job. In this paper, we present and evaluate broker selection strategies for interoperable grid systems. They use aggregated resource information as well as dynamic performance information of the underlying scheduling layers. From our evaluations performed with simulation tools, we conclude that aggregation techniques do not penalize performance significantly, and that delegating part of the scheduling responsibilities to the underlying scheduling layers is a good way to balance the load among the different grid systems.
Ivan Rodero, Francesc Guim 0001, Julita Corbalán, Liana L. Fong, Seyed Masoud Sadjadi
ICPP1
2009 The Resource Usage Aware Backfilling
Francesc Guim 0001, Ivan Rodero, Julita Corbalán
JSSPP2
2008 Enabling Interoperability among Meta-Schedulers
abstract
Grid computing supports shared access to computing resources from cooperating organizations or institutes in the form of virtual organizations. Resource brokering middleware, commonly known as a meta-scheduler or a resource broker, matches jobs to distributed resources. Recent advances in meta- scheduling capabilities are extended to enable resource matching across multiple virtual organizations. Several architectures have been proposed for interoperating meta-scheduling systems. This paper presents a hybrid approach, combining hierarchical and peer-to-peer architectures for flexibility and extensibility of these systems. A set of protocols are introduced to allow different meta-scheduler instances to communicate over Web Services. Interoperability between three heterogeneous and distributed organizations (namely, BSC, FIU, and IBM), each using different meta-scheduling technologies, is demonstrated under these protocols and resource models.
Norman Bobroff, Liana L. Fong, Selim Kalayci, Ivan Rodero, Seyed Masoud Sadjadi, David Villegas
CCGRID6
2008 Coordinated Co-allocation Scheduling on Heterogeneous Clusters of SMPs
abstract
Job scheduling research for parallel systems has been widely exploited in recent years, especially in centers with high performance computing facilities. In the recent past we presented the eNANOS execution environment which is based on a coordinated architecture, from the CPU allocation to the grid scheduling, providing a good low level support to perform an efficient high level scheduling. In this paper we present and evaluate the multi-node scheduling configuration of eNANOS that is implemented through the eNANOS Scheduler. Moreover, we introduce our scheduling strategy based on co-allocation and the coordination with dynamic processor allocation techniques. Finally, through experimental evaluation we state that our architecture and scheduling strategy can improve the applications execution and the system performance on heterogeneous clusters composed of SMP architectures.
Ivan Rodero, Julita Corbalán
eScience1
2008 Modeling and Evaluating Interoperable Grid Systems
abstract
Grid resource management tools have evolved from manual discovery and job submission to sophisticated brokering solutions. User requirements have created certain properties that resource managers have learned to support. This development is still continuing, and users already find it difficult to distinguish brokers and to migrate their applications when they move to a different grid. Moreover, new architectures are continuously being proposed, such as multi-site and interoperable grid systems. This paper presents the Alvio simulation framework which is designed to evaluate job scheduling strategies in complex HPC infrastructures. The main contribution of this simulator is that it allows modeling from local systems to interoperable grid scenarios. We also present an evaluation of multi-site and grid interoperable systems which shows the effect of job forwarding between different brokers.
Ivan Rodero, Francesc Guim 0001, Julita Corbalán
eScience1
2006 Uniform Job Monitoring using the HPC-Europa Single Point of Access
Francesc Guim 0001, Ivan Rodero, Julita Corbalán, Jesús Labarta, Ariel Oleksiak, Tomasz Kuczynski, Dawid Szejnfeld, Jarek Nabrzyski
CCGRID2
2006 How the JSDL can Exploit the Parallelism?
abstract
The description of the jobs is a very important issue for the scheduling and management of grid jobs. Since there are a lot of different languages for describing grid jobs, the GGF have presented the Job Submission Description Language (JSDL) to standardize the job submission language. We believe that the JSDL is a good solution but it has some deficiencies regarding the parallelism issues. In this paper, we propose an extension of the JSDL to specify the parallelism details of grid jobs. This extension is proposed in general terms for supporting current multilevel parallel applications and incoming approaches in parallel programming models. We also discus the suitability of the multilevel parallel programming models for grids, in particular the MPI+OpenMP since our project, the eNANOS project, is based on this hybrid programming model.
Ivan Rodero, Francesc Guim 0001, Julita Corbalán, Jesús Labarta
CCGRID1