EDBT 2026 Demo / reviewers in the wild / expert
Claris Castillo
dblp:25/442
· DBLP profile ↗
12ranked-venue papers
6as first author
0since 2021 · last 2020
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 first-authorArtificial intelligence and machine learning · 2Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Parallel and multicore computing · 56% Cloud and datacenter computing · 44% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Parallel and multicore computing › task scheduling
online scheduling |
0.1 | 1 | 2009 | Resource co-allocation for large-scale distributed environments · HPDC 2009 |
Cloud and datacenter computing › resource allocation › multi-resource allocation
resource co-allocation |
0.1 | 1 | 2009 | Resource co-allocation for large-scale distributed environments · HPDC 2009 |
Parallel and multicore computing
scheduling algorithms |
0.1 | 1 | 2009 | Resource co-allocation for large-scale distributed environments · HPDC 2009 |
Cloud and datacenter computing › resource allocation › resource allocation policy
advance reservation |
0.0 | 1 | 2009 | Resource co-allocation for large-scale distributed environments · HPDC 2009 |
Methods — techniques the papers use, named apart from their topics
temporal availability data structures · 0.1simulation · 0.1range searches · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | A Cloud-Agnostic Framework to Enable Cost-Aware Scheduling of Applications in a Multi-Cloud EnvironmentabstractWe have witnessed a surge in both the big data applications being hosted by an assortment of cloud vendors, and in the astronomical amount of data they produce and consume on a daily basis. Traditional cluster computing frameworks can hardly cope with the unprecedented data volume and the geo-distributed, cross-cloud data distribution due to their limited scalability and adaptability across the heterogeneous clouds. Moreover, running data-intensive applications across clouds at will is extremely cost-inefficient and likely to incur outrageous expenses. Hence, we introduce our cloud-agnostic system PIVOT with the novel cost-aware scheduling algorithm, which enables data-intensive applications to run and scale across clouds instantly in a cost-efficient manner. We evaluate our system and scheduling algorithm extensively with the Alibaba production cluster trace, as well as real-world big data applications on a 100-node deployment across 11 regions (31 availability zones) on AWS and GCP. The experimental results show that PIVOT achieves up to 90.8% saving in expense for VM subscription and 99.2% for egress network traffic compared to the state-of-the-art baselines. Notably, the cost-aware scheduling also achieves over 4x speedup in data transfers for data-intensive applications. Kyle Ferriter, Claris Castillo |
NOMS | 3 |
| 2018 | Cachalot: A network-aware, cooperative cache network for geo-distributed, data-intensive applicationsabstractCollaborative and data-intensive applications are hosted on geo-distributed infrastructures to exploit computing resources at scale. However, these applications typically incur massive data transfers over bandwidth-constrained wide- area networks (WANs) which impose significant performance overhead. Conventional distributed computing platforms (e.g., Spark) leverage caching to avoid duplicate executions of common computations and thus reduce network traffic. However, these techniques were developed for data center environments and therefore lack advanced network-aware mechanisms to support high-performance, data-intensive applications over the WAN in geo-distributed environments. Hence, we develop Cachalot - a novel network-aware, cooperative cache network for caching datasets generated by common computations shared among geo- distributed, data-intensive applications. We perform a simulation- based deep evaluation using both synthetic and real traces. The experimental results indicate Cachalot speeds up data-intensive applications by over 50%, reducing network traffic by up to 60%; and, outperforms state-of-the-art baselines by over 20% in geo-distributed environments for various common user-driven performance metrics. Claris Castillo, Stanley C. Ahalt |
NOMS | 2 |
| 2016 | RADU: Bridging the divide between data and infrastructure management to support data-driven collaborationsabstractWe have witnessed a dramatic increase in national cyberinfrastructure resources to support data-driven research. Orchestrating these resources to enable the creation of collaborative infrastructure capable of supporting data intensive activities is challenging. In this work we present RADII, a novel architecture and system that enables the provisioning and configuration of collaborative infrastructure by orchestrating data and infrastructure management in an integrated manner. We also introduce a cross-layer data annotation mechanism that together with Software-Defined Networking (SDN) support allows the embedding of user-defined performance policies into file metadata which are translated into executable optimal network plans. We have deployed RADII on a worldwide production testbed and demonstrated through experimentation that RADII can improve network throughput of data transfers by 28.2% as compared to conventional approaches. Claris Castillo, Charles Schmitt |
IEEE BigData | 2 |
| 2014 | Enabling genomic analysis on federated cloudsabstractGenomic research involves a considerable amount of intensive computational and data management challenges including varying demand for computing resources, large data staging, management of complex workflows, and managing data and metadata across thousands of experiments and datasets. To address these challenges, researchers are typically forced to acquire and maintain experienced IT staff and informatics infrastructure. Increasingly researchers have explored cloud technologies, yet these lack key capabilities for data and workflow management. We introduce work to address these issues through use of federated cloud infrastructure coupled with data and workflow management technology. We present preliminary work toward the integration of three major technologies: ExoGENI, integrated Rule Oriented Data System (iRODS), and Pegasus/HTCondor, to develop a software infrastructure that better supports data-and workflow-centric genomic analysis. Michael Shoffner, Claris Castillo, Charles Schmitt |
IEEE BigData | 3 |
| 2012 | Enabling Efficient Placement of Virtual Infrastructures in the Cloud
Ioana Giurgiu, Claris Castillo, Asser N. Tantawi, Malgorzata Steinder |
Middleware | 2 |
| 2012 | Cost-aware replication for dataflowsabstractIn this work we are concerned with the cost associated with replicating intermediate data for dataflows in Cloud environments. This cost is attributed to the extra resources required to create and maintain the additional replicas for a given data set. Existing data-analytic platforms such as Hadoop provide for fault-tolerance guarantee by relying on aggressive replication of intermediate data. We argue that the decision to replicate along with the number of replicas should be a function of the resource usage and utility of the data in order to minimize the cost of reliability. Furthermore, the utility of the data is determined by the structure of the dataflow and the reliability of the system. We propose a replication technique, which takes into account resource usage, system reliability and the characteristic of the dataflow to decide what data to replicate and when to replicate. The replication decision is obtained by solving a constrained integer programming problem given information about the dataflow up to a decision point. In addition, we built a working prototype, CARDIO of our technique which shows through experimental evaluation using a real testbed that finds an optimal solution. Claris Castillo, Asser N. Tantawi, Diana Arroyo, Malgorzata Steinder |
NOMS | 1 |
| 2011 | Towards efficient resource management for data-analytic platformsabstractWe present architectural and experimental work exploring the role of intermediate data handling in the performance of MapReduce workloads. Our findings show that: (a) certain jobs are more sensitive to disk cache size than others and (b) this sensitivity is mostly due to the local file I/O for the intermediate data. We also show that a small amount of memory is sufficient for the normal needs of map workers to hold their intermediate data until it is read. We introduce Hannibal, which exploits the modesty of that need in a simple and direct way — holding the intermediate data in application-level memory for precisely the needed time — to improve performance when the disk cache is stressed. We have implemented Hannibal and show through experimental evaluation that Hannibal can make MapReduce jobs run faster than Hadoop when little memory is available to the disk cache. This provides better performance insulation between concurrent jobs. Claris Castillo, Mike Spreitzer, Malgorzata Steinder |
Integrated Network Management | 1 |
| 2011 | Resource-Aware Adaptive Scheduling for MapReduce Clusters
Jorda Polo, Claris Castillo, David Carrera 0001, Yolanda Becerra 0001, Ian Whalley, Malgorzata Steinder, Jordi Torres, Eduard Ayguadé |
Middleware | 2 |
| 2011 | Online algorithms for advance resource reservations
Claris Castillo, George N. Rouskas, Khaled Harfoush |
J. Parallel Distributed Comput. | 1 |
| 2009 | Resource co-allocation for large-scale distributed environmentsabstractAdvances in the development of large scale distributed computing systems such as Grids and Computing Clouds have intensified the need for developing scheduling algorithms capable of allocating multiple resources simultaneously. In principle, the required resources may be allocated by sequentially scheduling each resource individually. However, such a solution can be computationally expensive, hence inappropriate for time-sensitive applications, and may lead to deadlocks. In this work we present an efficient online algorithm for co-allocating resources that also provides support for advance reservations. The algorithm utilizes data structures specifically designed to organize the temporal availability of resources, and implements co-allocation through efficient range searches that identify all available resources simultaneously. We use simulations driven by real workloads to show that the co-allocation algorithm scales to systems with large numbers of users and resources, and we perform an in-depth comparative analysis against existing batch scheduling mechanisms. Our findings indicate that the online scheduling algorithms may achieve higher utilization while providing smaller delays and better QoS guarantees without adding much complexity. Claris Castillo, George N. Rouskas, Khaled Harfoush |
HPDC | 1 |
| 2008 | Efficient resource management using advance reservations for heterogeneous GridsabstractSupport for advance reservations of resources plays a key role in Grid resource management as it enables the system to meet user expectations with respect to time requirements and temporal dependence of applications, increases predictability of the system and enables co- allocation of resources. Despite these attractive features, adoption of advance reservations is limited mainly due to the fact that related algorithms are typically complex and fail to scale to large and loaded systems. In this work we consider two aspects of advance reservations. First, we investigate the impact of heterogeneity on Grid resource management when advance reservations are supported. Second, we employ techniques from computational geometry to develop an efficient heterogeneity-aware scheduling algorithm. Our main finding is that Grids may benefit from high levels of resource heterogeneity, independently of the total system capacity. Our results show that our algorithm performs well across several user and system performance and overcome the lack of scalability and adaptability of existing mechanisms. Claris Castillo, George N. Rouskas, Khaled Harfoush |
IPDPS | 1 |
| 2007 | On the Design of Online Scheduling Algorithms for Advance Reservations and QoS in GridsabstractWe consider the problem of providing QoS guarantees to Grid users through advance reservation of resources. Advance reservation mechanisms provide the ability to allocate resources to users based on agreed-upon QoS requirements and increase the predictability of a Grid system, yet incorporating such mechanisms into current Grid environments has proven to be a challenging task due to the resulting resource fragmentation. We use concepts from computational geometry to present a framework for tackling the resource fragmentation, and for formulating a suite of scheduling strategies. We also develop efficient implementations of the scheduling algorithms that scale to large Grids. We conduct a comprehensive performance evaluation study using simulation, and we present numerical results to demonstrate that our strategies perform well across several metrics that reflect both user-and system-specific goals. Our main contribution is a timely, practical, and efficient solution to the problem of scheduling resources in emerging on-demand computing environments. Claris Castillo, George N. Rouskas, Khaled Harfoush |
IPDPS | 1 |