VLDB 2026 Research / reviewers in the wild / expert
Trilce Estrada
dblp:20/6272
· DBLP profile ↗
37ranked-venue papers
8as first author
9since 2021 · last 2025
0000-0001-7743-8754ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 2 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 7 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling Laws for the Workload Throughput of Emerging Heterogeneous ClustersabstractNext-generation HPC clusters are evolving into highly heterogeneous systems that integrate traditional computing resources with emerging accelerator technologies such as quantum processors, neuromorphic units, dataflow architectures, and specialized AI accelerators within a unified infrastructure. These advanced systems enable workloads to dynamically utilize different accelerators during various computation phases, creating complex execution patterns. The performance of the workloads can therefore be impacted by many factors, including how the accelerators are shared, their utilization, and their placement within the system. Moreover, effects such as the system and network state due to the overall system load can significantly impact the job completion rate. Understanding, identifying, and quantifying the impact of the most critical factors (e.g., the number of allocated accelerators) will help decide the investment decisions for accelerator acquisition and deployment that can improve the overall system throughput. This paper extensively studies these complex interactions among advanced accelerators within an HPC cluster and various workloads. We introduce a novel analytical model which predicts the speedup of a workload given an accelerator/system configuration. This model can be used to quantify the effect of augmenting additional accelerators on job performance running on an HPC cluster. We validate the model using both simulated and real environments. Akhil Alasandagutti, Joshua Suetterlein, Jesun Sahariar Firoz, Stephen J. Young, Joseph B. Manzano, Jason R. Stewart, Patrick G. Bridges, Trilce Estrada, Kevin J. Barker |
CCGrid | 8 |
| 2025 | Grey-Box Machine Learning Prediction of Parallel Application ScalingabstractAccurate prediction of parallel application performance in HPC systems is essential for efficient resource allocation and system design. Classical performance models estimate of speedup based on theoretical assumptions, but their applicability is limited by parameter estimation, data acquisition, and real-world system issues such as latency and network congestion. This paper describes performance prediction using classical performance models boosted by a trainable machine learning framework. Domain-informed machine-learning models estimate the overhead of an application for a given problem size and resource configuration as a coefficient of the estimated speedup provided by performance laws. We evaluate this approach on two HPC mini-applications and two full applications with varying patterns of computation and communication and also evaluate the prediction accuracy on runs with varying processors-per-node configurations. Our results show that this method significantly improves the accuracy of performance predictions over standard analytical models and black-box regressors, while remaining robust even with limited training data. Akhil Alasandagutti, Patrick G. Bridges, Trilce Estrada |
HiPC | 3 |
| 2023 | DRUM: A Real Time Detector for Regime Shifts in Data Streams via an Unsupervised, Multivariate Framework
Adnan Bashir, Trilce Estrada |
DaWaK | 2 |
| 2023 | Online Boosted Gaussian Learners for In-Situ Detection and Characterization of Protein Folding States in Molecular Dynamics SimulationsabstractMolecular Dynamics (MD) simulations are a crucial tool for understanding how proteins fold. In its easiest form, MD simulations can be scaled through data parallelism, this means that multiple folding trajectories can be spawned and executed in parallel, facilitating a more efficient exploration of the protein folding space. However, due to data dependencies, the analysis of MD simulations remains largely as a centralized process. In this work, we propose a data parallel, lightweight technique to learn the characteristics of protein folding states in MD simulations. Contrary to other methods, ours can differentiate relevant states in a single protein folding trajectory without requiring centralized global knowledge of the protein dynamics. As its processing and memory overheads are negligible (in the order of milliseconds per window of frames, and kilo bytes respectively) this technique can be coupled with the simulation for in-situ analysis. Harshita Sahni, Hector Carrillo-Cabada, Ekaterina D. Kots, Silvina Caíno-Lores, Jack D. Marquez, Ewa Deelman, Michel A. Cuendet, Harel Weinstein, Michela Taufer, Trilce Estrada |
e-Science | 10 |
| 2023 | NSF/IEEE-TCPP Curriculum on Parallel and Distributed Computing for Undergraduates - Version II - Big Data, Energy, and Distributed ComputingabstractThis special session will report on the updated NSF/IEEE-TCPP Curriculum on Parallel and Distributed Computing released in Nov 2020 by the Center for Parallel and Distributed Computing Curriculum Development and Educational Resources (CDER). The purpose of the special session is to obtain SIGCSE community feedback on this curriculum in a highly interactive manner employing the hybrid modality and supported by a full-time CDER booth for the duration of SIGCSE. In this era of big data, cloud, and multi- and many-core systems, it is essential that the computer science (CS) and computer engineering (CE) graduates have basic skills in parallel and distributed computing (PDC). The topics are primarily organized into the areas of architecture, programming, and algorithms topics. A set of pervasive concepts that percolate across area boundaries are also identified. Version 1 of this curriculum was released in December 2012. That curriculum guideline has over 140 early adopter institutions worldwide and has been incorporated into the 2013 ACM/IEEE Computer Science curricula. This Version-II represents a major revision. The updates have focused on enhancing coverage related to the topical aspects of Big Data, Energy, and Distributed Computing. Sushil K. Prasad, Charles C. Weems, Alan Sussman, Trilce Estrada, Ramachandran Vaidyanathan, Sheikh K. Ghafoor, Krishna Kant 0001, Craig B. Stunkel |
SIGCSE (2) | 5 |
| 2021 | PASCAL-G: a Probabilistic Stream Clustering Analysis on GraphsabstractGraph clustering is a fundamental problem widely used in many applications, most of them require the efficient processing of very large input graphs in real time or require to process data that may not reside in a centralized location. Several algorithms to solve the graph clustering problem have been proposed and published, but traditional methods are not sufficient anymore. This paper proposes a novel graph clustering algorithm that is designed to effectively find community structure in large graphs. PASCAL-G uncovers the network structure on the go, is incremental in nature and works without needing to store the entire graph before-hand. All of this makes PASCAL-G capable of handling data in an online or streaming fashion. Nidia Yadira Vaquera Chavez, Trilce Estrada |
IEEE BigData | 2 |
| 2021 | A Roadmap to Robust Science for High-throughput Applications: The Developers' PerspectiveabstractScientists using the high-throughput computing (HTC) paradigm for scientific discovery rely on complex software systems and heterogeneous architectures that must deliver robust science (i.e., ensuring performance scalability in space and time; trust in technology, people, and infrastructures; and reproducible or confirmable research). Developers must overcome a variety of obstacles to pursue workflow interoperability, identify tools and libraries for robust science, port codes across different architectures, and establish trust in non-deterministic results. This poster presents recommendations to build a roadmap to overcome these challenges and enable robust science for HTC applications and workflows. The findings were collected from an international community of software developers during a Virtual World Cafe in May 2021. Michela Taufer, Ewa Deelman, Rafael Ferreira da Silva, Trilce Estrada, Mary W. Hall, Miron Livny |
CLUSTER | 4 |
| 2021 | A Roadmap to Robust Science for High-throughput Applications: The Scientists' PerspectiveabstractThis poster presents our first steps to define a roadmap to robust science for high-throughput applications used in scientific discovery. These applications combine multiple components into increasingly complex multi-modal workflows that are often executed in concert on heterogeneous systems. The increasing complexity hinders the ability of scientists to generate robust science (i.e., ensuring performance scalability in space and time; trust in technology, people, and infrastructures; and reproducible or confirmable research). Scientists must withstand and overcome adverse conditions such as heterogeneous and unreliable architectures at all scales (including extreme scale), rigorous testing under uncertainties, unexplainable algorithms in machine learning, and black-box methods. This poster presents findings and recommendations to build a roadmap to overcome these challenges and enable robust science. The data was collected from an international community of scientists during a virtual world café in February 2021. Michela Taufer, Ewa Deelman, Rafael Ferreira da Silva, Trilce Estrada, Mary W. Hall |
e-Science | 4 |
| 2021 | A Graphic Encoding Method for Quantitative Classification of Protein Structure and Representation of Conformational ChangesabstractIn order to successfully predict a proteins function throughout its trajectory, in addition to uncovering changes in its conformational state, it is necessary to employ techniques that maintain its 3D information while performing at scale. We extend a protein representation that encodes secondary and tertiary structure into fix-sized, color images, and a neural network architecture (called GEM-net) that leverages our encoded representation. We show the applicability of our method in two ways: (1) performing protein function prediction, hitting accuracy between 78 and 83 percent, and (2) visualizing and detecting conformational changes in protein trajectories during molecular dynamics simulations. Hector Carrillo-Cabada, Jeremy Benson, Asghar M. Razavi, Brianna Mulligan, Michel A. Cuendet, Harel Weinstein, Michela Taufer, Trilce Estrada |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2019 | DeepManner: Automatically Determining Manner of DeathabstractMedical examiners are trained medical professional who are responsible for determining cause and manner of death (natural, accident, suicide, homicide, and undetermined). Their jurisdiction includes unexpec ted, violent, and unexplained deaths. The final conclusions determined by the medical examiners can have significant consequences, therefore it is important every case is handled in a consistent and unbiased fashion. We will show how modern natural language processing techniques can be used to automatically determine the manner of death from the textual descriptions of injuries, demographic info, scene description, and the microscopy results found in autopsy reports. Esteban Guillen, Trilce Estrada, Matthew Cain |
IEEE BigData | 2 |
| 2019 | Characterizing In Situ and In Transit Analytics of Molecular Dynamics Simulations for Next-Generation SupercomputersabstractMolecular Dynamics (MD) simulations executed on state-of-the-art supercomputers are producing data at rates faster than it can be written out to disk. In situ and in transit analysis of data generated by MD simulations reduce the original volume of information by several orders of magnitude, thereby alleviating the negative impact of I/O bottlenecks. This work focuses on characterizing the impact of in situ and in transit analytics on the overall MD workflow performance, and the capability for capturing rapid, rare events in the simulated molecular system. The MD simulation and analysis processes share data via remote direct memory access (RDMA) using DataSpaces. Our metrics of interest are time spent waiting in I/O by the MD simulation, lost frames of the MD simulation, and idle time of the analysis. We measure these metrics for a diverse set of molecular systems and characterize their trends for in situ and in transit configurations. We then model which frames are dropped and which ones are analyzed for a real use case. The insights gained from this study are generally applicable for in situ and in transit workflows that require optimization of parameters to minimize loss in workflow performance and analytic accuracy. Michela Taufer, Ewa Deelman, Michael R. Wyatt II, Tu Mai Anh Do, Loïc Pottier, Rafael Ferreira da Silva, Harel Weinstein, Michel A. Cuendet, Trilce Estrada |
eScience | 10 |
| 2018 | In situ TensorView: In situ Visualization of Convolutional Neural NetworksabstractConvolutional Neural Networks(CNNs) are complex systems trained to recognize images, texts and more. However, once trained, they are regarded as black-boxes that are not easy to analyze and understand. Visualizing the dynamics within such deep artificial neural networks can provide a better understanding of how they are learning and making predictions. In the field of scientific simulations, visualization tools like Paraview have long been utilized to provide insights. We present in situ TensorView to visualize the training and functioning of CNNs as if they are systems of scientific simulations. In situ TensorView is a loosely coupled in situ visualization open framework that provides multiple viewers with the ability to visualize and understand their networks. It leverages the capability of co-processing from Paraview to provide real-time visualization during training and predicting phases, and avoids heavy I/O overhead. Tensorview is easily coupled with Tensorflow, as it only requires the insertion of a few lines of code into a TensorFlow framework. In this work, we showcase visualizing LeNet-5 and VGG16 using in situ TensorView. With the insight provided by Tensorview, users can adjust network architectures, or compress pre-trained networks guided by visualization results. Qiang Guan, Li-Ta Lo, Simon Su, Zhengyong Ren, James P. Ahrens, Trilce Estrada |
IEEE BigData | 7 |
| 2018 | KeyBin2: Distributed Clustering for Scalable and In-Situ AnalysisabstractWe present KeyBin2, a key-based clustering method that is able to learn from distributed data in parallel. KeyBin2 uses random projections and discrete optimizations to efficiently clustering very high dimensional data. Because it is based on keys computed independently per dimension and per data point, KeyBin2 scales linearly. We perform accuracy and scalability tests to evaluate our algorithm's performance using synthetic and real datasets. The experiments show that KeyBin2 outperforms other parallel clustering methods for problems with increased complexity. Finally, we present an application of KeyBin2 for in-situ clustering of protein folding trajectories. Jeremy Benson, Matt Peterson, Michela Taufer, Trilce Estrada |
ICPP | 5 |
| 2017 | keybin: Key-Based Binning for Distributed ClusteringabstractTraditional machine learning algorithms often require computations on centralized data, but modern datasets are collected and stored in a distributed way. In addition to the cost of moving data to centralized locations, increasing concerns about privacy and security warrant distributed approaches. We propose keybin, a distributed key-based binning clustering algorithm for high-dimensional spaces. keybin locally generates a spatial key for each data point across all dimensions without needing knowledge of other data. Then, it performs a conceptual Map- Reduce procedure in the index space to form a global clustering assignment. We present an implementation and a case study on the capabilities and limitations of this approach, showing that this algorithm can learn a global clustering structure with limited communication and can scale with the dimensionality and size of data sets. Jeremy Benson, Trilce Estrada |
CLUSTER | 3 |
| 2017 | Index Clustering: A Map-reduce Clustering Approach using Numba
Trilce Estrada |
DATA | 2 |
| 2017 | Table Interpretation and Extraction of Semantic Relationships to Synthesize Digital Documents
Martha O. Perez-Arriaga, Trilce Estrada, Soraya Abad-Mota |
DATA | 2 |
| 2017 | Enabling scalable and accurate clustering of distributed ligand geometries on supercomputers
Boyu Zhang 0002, Trilce Estrada, Pietro Cicotti, Pavan Balaji, Michela Taufer |
Parallel Comput. | 2 |
| 2016 | Crowdlearning: A framework for collaborative and personalized learningabstractThis work presents a new technological learning paradigm that we call crowdlearning, where students are upgraded from mere passive content consumers to primary content creators and curators. Crowdlearning is all about allowing learners to share their particular vision of the world, understanding how different users learn, and putting in place mechanisms to safely, scalable, and accurately share and consume learning materials online. By providing an interactive learning framework, we expect that attention of underrepresented groups in STEM, particularly women, will be drawn towards computer science and other technology-related fields. In this paper we introduce all the necessary components of a crowdlearning framework, we provide a middleware implementation along with a proof of concept application. Finally, we discuss possible implications, limitations, and prospects of this approach, and how it could be shaped into a mainstream collaborative way of enhancing cyber learning. If successful in the long run, this paradigm has the potential of advancing technology aided education tools while promoting teaching, training, and collaborative learning. Thilak R. Balasubramanian, Trilce Estrada |
FIE | 2 |
| 2016 | Scheduling Matters: Area-Oriented Heuristic for Resource ManagementabstractParallel and distributed systems that provide compute resources on demand are convenient, cost-effective, and becoming increasingly common. Boosting workload performance in such environments through scheduling has been of great interest, as users and providers aim to increase parallelism and reduce execution times. For modern data centers, leaving a smaller carbon footprint while maintaining high performance and low cost is becoming the next big challenge. With this in mind, we analyze the relative impacts on resource utilization of three well-motivated platform-oblivious scheduling heuristics. We simulate over 50,000 DAG workflow executions and measure performance, cost, and resource utilization under the three scheduling heuristics. Our results provide insights to better enable high-performance execution of workflows and advanced capacity planning while increasing resource utilization and reducing costs. Jeremy Benson, Trilce Estrada, Arnold L. Rosenberg, Michela Taufer |
SBAC-PAD | 2 |
| 2015 | Accurate Scoring of Drug Conformations at the Extreme ScaleabstractWe present a scalable method to extensively search for and accurately select pharmaceutical drug candidates in large spaces of drug conformations computationally generated and stored across the nodes of a large distributed system. For each legend conformation in the dataset, our method first extracts relevant geometrical properties and transforms the properties into a single metadata point in the three-dimensional space. Then, it performs an ochre-based clustering on the metadata to search for predominant clusters. Our method avoids the need to move legend conformations among nodes because it extracts relevant data properties locally and concurrently. By doing so, we can perform accurate and scalable distributed clustering analysis on large distributed datasets. We scale the analysis of our pharmaceutical datasets a factor of 400X higher in performance and 500X larger in size than ever before. We also show that our clustering achieves higher accuracy compared with that of traditional clustering methods and conformational scoring based on minimum energy. Boyu Zhang 0002, Trilce Estrada, Pietro Cicotti, Pavan Balaji, Michela Taufer |
CCGRID | 2 |
| 2015 | A Genetic Programming Approach to Design Resource Allocation Policies for Heterogeneous Workflows in the CloudabstractWhen dealing with very large applications in the cloud, higher costs do not always result in better turnaround times, particularly for complex workflows with multiple task dependencies. Thus, resource allocation policies are needed that can determine when using expensive but faster resources is best and when it is not. Manually developing such heuristics is time consuming and limited by the subjective beliefs of the developer. To overcome such impediments, we present an automatic method that designs and evaluates a large set of policies using a genetic programming approach. Our method finds a robust set of policies that adapt to changes in workload while using resources efficiently. Our results show that our genetic programming designed policies perform better than greedy and other human designed policies do. Trilce Estrada, Michael R. Wyatt II, Michela Taufer |
ICPADS | 1 |
| 2015 | A Testing Engine for High-Performance and Cost-Effective Workflow Execution in the CloudabstractWhile pursuing high performance and cost effectiveness for directed acyclic graph (DAG)-structured scientific workflow executions in the cloud, it is critical to identify appropriate resource instances and their quantity. This paper presents a testing engine that employs a resource-selection heuristic, which statically analyzes the DAG structure to guide the selection of resource instances, how many and which ones. The testing engine combines the heuristic with two platform-independent DAG-scheduling policies, the Area-oriented DAG-scheduling heuristic (AO) and the Locally-Optimal heuristic (L-OPT), to perform extensive validation assessments. The testing engine ensures the realism of these assessments by modeling the performance variability of the cloud platform using real traces. The testing engine also enables cost-effectiveness analysis that guides users to select a small set of instance candidates that provide performance-cost trade off. Our empirical results show that the pairing of the resource-selection heuristic with AO scheduling policy is a powerful method for cost-effective DAG-structured workflow execution in the cloud. Vivek K. Pallipuram, Trilce Estrada, Michela Taufer |
ICPP | 2 |
| 2014 | Gender and volunteer computing: A survey studyabstractVolunteer computing is a form of citizen science that has a significant gender imbalance. Far fewer women than men participate; women are typically less than ten percent of a project's participants. To better understand the experience of women in volunteer computing and seek clues as to methods for using volunteer computing experience as a recruiting tool, we analyze participant survey data from a project that tries to engage new communities with an interactive infrastructure. Our results showed very few gender differences among the responding men and women in volunteer computing. Our findings add to evidence that men and women engaged in computing activities are overwhelmingly similar. The challenge for gender balance seems to be informing and engaging larger networks of diverse women to promote volunteer computing and contribute to achieving its goals. Feng Raoking, Joanne McGrath Cohoon, Kathryn Cooke, Michela Taufer, Trilce Estrada |
FIE | 5 |
| 2014 | Time Series Join on Subsequence CorrelationabstractWe consider the problem of joining two long time series based on their most correlated segments. Two time series can be joined at any locations and for arbitrary length. Such join locations and length provide useful knowledge about the synchrony of the two time series and have applications in many domains including environmental monitoring, patient monitoring and power monitoring. However, join on correlation is a computationally expensive task, specially when the time series are large. The naive algorithm requires O (n4) computation where n is the length of the time series. We propose an algorithm, named Jocor, that uses two algorithmic techniques to tackle the complexity. First, the algorithm reuses the computation by caching sufficient statistics and second, the algorithm prunes unnecessary correlation computation by admissible heuristics. The algorithm runs orders of magnitude faster than the naive algorithm and enables us to join long time series as well as many small time series. We propose a variant of Jocor for fast approximation and an extension to a GPU-based parallel method to bring down the running-time to interactive level for analytics applications. We show three independent uses of time series join on correlation which are made possible by our algorithm. Abdullah Mueen, Hossein Hamooni, Trilce Estrada |
ICDM | 3 |
| 2014 | Enabling In-Situ Data Analysis for Large Protein-Folding Trajectory DatasetsabstractThis paper presents a one-pass, distributed method that enables in-situ data analysis for large protein folding trajectory datasets by executing sufficiently fast, avoiding moving trajectory data, and limiting the memory usage. First, the method extracts the geometric shape features of each protein conformation in parallel. Then, it classifies sets of consecutive conformations into meta-stable and transition stages using a probabilistic hierarchical clustering method. Lastly, it rebuilds the global knowledge necessary for the intraand inter-trajectory analysis through a reduction operation. The comparison of our method with a traditional approach for a villin headpiece sub domain shows that our method generates significant improvements in execution time, memory usage, and data movement. Specifically, to analyze the same trajectory consisting of 20,000 protein conformations, our method runs in 41.5 seconds while the traditional approach takes approximately 3 hours, uses 6.9MB memory per core while the traditional method uses 16GB on one single node where the analysis is performed, and communicates only 4.4KB while the traditional method moves the entire dataset of 539MB. The overall results in this paper support our claim that our method is suitable for in-situ data analysis of folding trajectories. Boyu Zhang 0002, Trilce Estrada, Pietro Cicotti, Michela Taufer |
IPDPS | 2 |
| 2013 | Benchmarking Gender Differences in Volunteer Computing ProjectsabstractVolunteer Computing (VC) uses the computational resources of volunteers with Internet-connected personal computers to address fundamental problems in science. Docking Home (D@H) is a VC project targeting drug discovery through high throughput docking simulations i.e., by docking small molecules (ligands) into target proteins associated to diseases. Currently there are more than 27,000 volunteers (and 70,000 computers) worldwide supporting D@H. Similar to national trends in STEM fields, in general, the huge majority of volunteers engaged in VC projects, and in D@H in particular, are Caucasian males. This paper aims to characterize the current VC community supporting D@H and uses the information to define strategies that can help attract and retain female and ethnic minority volunteers. Trilce Estrada, Kathleen L. Pusecker, Manuel R. Torres, Joanne McGrath Cohoon, Michela Taufer |
e-Science | 1 |
| 2013 | On the powerful use of simulations in the Quake-Catcher Network to efficiently position low-cost earthquake sensors
Kyle E. Benson, Samuel Schlachter, Trilce Estrada, Michela Taufer, Jesse Lawrence, Elizabeth Cochran |
Future Gener. Comput. Syst. | 3 |
| 2012 | ExSciTecH: Expanding volunteer computing to Explore Science, Technology, and HealthabstractThis paper presents ExSciTecH, an NSF-funded project deploying volunteer computing (VC) systems to Explore Science, Tecenology, and Health. ExSciTecH aims at radically transforming VC systems and the volunteer's experience. To pursue this goal, ExSciTecH integrates and uses gameplay environments into BOINC, a well-known VC middleware, to involve the volunteers not only for simply donating idle cycles but also for actively participating in scientific discovery, i.e., generating new simulations side by side with the scientists. More specifically, ExSciTecH plugs into the BOINC framework extending it with two main gaming components, i.e., a learning component that includes a suite of games for training users on relevant biochemical concepts, and an engaging component that includes a suite of games to engage volunteers in drug design and scientific discovery. We assessed the impact of a first implementation of the learning game on a group of students at the University of Delaware. Our tests clearly show how ExSciTecH can generate higher levels of enthusiasm than more traditional learning tools in our students. Michael Matheny, Samuel Schlachter, L. M. Crouse, E. T. Kimmel, Trilce Estrada, Marcel Schumann, Roger S. Armen, Gary M. Zoppetti, Michela Taufer |
eScience | 5 |
| 2012 | On the effectiveness of application-aware self-management for scientific discovery in volunteer computing systemsabstractAn important challenge faced by high-throughput, multiscale applications is that human intervention has a central role in driving their success. However, manual intervention is inefficient, error-prone and promotes resource wasting. This paper presents an application-aware modular framework that provides self-management for computational multiscale applications in volunteer computing (VC). Our framework consists of a learning engine and three modules that can be easily adapted to different distributed systems. The learning engine of this framework is based on our novel tree-like structure called KOTree. KOTree is a fully automatic method that organizes statistical information in a multidimensional structure that can be efficiently searched and updated at runtime. Our empirical evaluation shows that our framework can effectively provide application-aware self-management in VC systems. Additionally, we observed that our KOTree algorithm is able to predict accurately the expected length of new jobs, resulting in an average of 85% increased throughput with respect to other algorithms. Trilce Estrada, Michela Taufer |
SC | 1 |
| 2011 | On the Powerful Use of Simulations in the Quake-Catcher Network to Efficiently Position Low-cost Earthquake SensorsabstractThe Quake-Catcher Network (QCN) uses low-cost sensors connected to volunteer computers across the world to monitor seismic events. The location and density of these sensors' placement can impact the accuracy of the event detection. Because testing different special arrangements of new sensors could disrupt the currently active project, this would best be accomplished in a simulated environment. This paper presents an accurate and efficient framework for simulating the low cost QCN sensors and identifying their most effective locations and densities. Results presented show how our simulations are reliable tools to study diverse scenarios under different geographical and infrastructural constraints. Kyle E. Benson, Trilce Estrada, Michela Taufer, Jesse Lawrence, Elizabeth Cochran |
eScience | 2 |
| 2011 | Providing Quality of Science in Volunteer ComputingabstractWe propose an application-centric approach for tuning, at runtime, parameters in volunteer computing (VC). Our approach defines Quality of Science (QSc) as the ability of a system to meet application-specific goals within time constraints. We first estimate the impact of different application parameters on both QSc metrics and resource usage from empirical information gathered at runtime and then reconfigure the parameters accordingly. Our approach results in more accurate solutions in shorter times and lower resource consumption than other, more traditional approaches. Trilce Estrada, Michela Taufer |
HPCC | 1 |
| 2009 | Modeling Job Lifespan Delays in Volunteer Computing ProjectsabstractVolunteer computing (VC) projects harness the power of computers owned by volunteers across the Internet to perform hundreds of thousands of independent jobs. In VC projects, the path leading from the generation of jobs to the validation of the job results is characterized by delays hidden in the job lifespan, i.e., distribution delay, in-progress delay, and validation delay. These delays are difficult to estimate because of the dynamic behavior and heterogeneity of VC resources. A wrong estimation of these delays can cause the loss of project throughput and job latency in VC projects. In this paper, we evaluate the accuracy of several probabilistic methods to model the upper time bounds of these delays. We show how our selected models predict up-and-down trends in traces from existing VC projects. The use of our models provides valuable insights on selecting project deadlines and taking scheduling decisions. By accurately predicting job lifespan delays, our models lead to more efficient resource use, higher project throughput, and lower job latency in VC projects. Trilce Estrada, Michela Taufer, Kevin Reed |
CCGRID | 1 |
| 2009 | EmBOINC: An emulator for performance analysis of BOINC projectsabstractBOINC is a platform for volunteer computing. The server component of BOINC embodies a number of scheduling policies and parameters that have a large impact on the projects throughput and other performance metrics. We have developed a system, EmBOINC, for studying these policies and parameters. EmBOINC uses a hybrid approach: it simulates a population of volunteered clients (including heterogeneity, churn, availability, reliability) and it emulates the server component; that is, it uses the actual server software and its associated database. This paper describes the design of EmBOINC and validates its results based on trace data from an existing BOINC project. Trilce Estrada, Michela Taufer, Kevin Reed, David P. Anderson |
IPDPS | 1 |
| 2009 | Performance Prediction and Analysis of BOINC Projects: An Empirical Study with EmBOINCabstractMiddleware systems for volunteer computing convert a set of computers that is large and diverse (in terms of hardware, software, availability, reliability, and trustworthiness) into a unified computing resource. This involves a number of scheduling policies and parameters, which have a large impact on the throughput and other performance metrics. How can we study and refine these policies? Experimentation in the context of a working project is problematic, and it is difficult to accurately model complex middleware in a conventional simulator. Instead, we use an approach in which the policies being studied are “emulated”, using parts of the actual middleware. In this paper we describe EmBOINC, an emulator based on the BOINC middleware system. EmBOINC simulates a population of volunteered clients (including heterogeneity, churn, availability, and reliability) and emulates the BOINC server components. After describing the design of EmBOINC and its validation, we present three case studies in which the impact of different scheduling policies are quantified in terms of throughput, latency, and starvation metrics. Trilce Estrada, Michela Taufer, David P. Anderson |
J. Grid Comput. | 1 |
| 2007 | Moving Volunteer Computing towards Knowledge-Constructed, Dynamically-Adaptive Modeling and SchedulingabstractVolunteer computing projects supported by BOINC have been exploring new research directions. For example, mature projects like Folding@home are moving towards the use of a broader range of architectures and computers. Other projects such as Docking@Home are exploring multi-scale, resource-driven and application-driven adaptations of the volunteer system. This paper presents results that enforce the need for knowledge-constructed capabilities in volunteer computing projects, i.e., the capability to drive simulations based on application-results and resource-status. The Docking@Home project, which uses volunteer resources to study putative drugs by computationally simulating the behavior of small molecules (ligands) when docking to a protein, serves as a case study to positively assess two key hypotheses. The first hypothesis claims that the adaptive selection of computational models for docking simulations based on the features of the protein and ligand can positively affect the final accuracy of the prediction. The second hypothesis claims that the adaptive selection of volunteer resources can ultimately improve project throughput. Michela Taufer, Andre Kerstens, Trilce Estrada, David A. Flores, Richard Zamudio, Patricia J. Teller, Roger S. Armen, Charles L. Brooks III |
IPDPS | 3 |
| 2006 | The Effectiveness of Threshold-Based Scheduling Policies in BOINC ProjectsabstractSeveral scientific projects use BOINC (Berkeley Open Infrastructure for Network Computing) to perform largescale simulations using volunteers' computers (workers) across the Internet. In general, the scheduling of tasks in BOINC uses a First-Come-First-Serve policy and no attention is paid to workers' past performance, such as whether or not they have tended to perform tasks promptly and correctly. In this paper we use SimBA, a discrete-event Simulator of BOINC Applications, to study new threshold-based scheduling strategies for BOINC projects that use availability and reliability metrics to classify workers and distribute tasks according to this classification. We show that if availability and reliability thresholds are selected properly, then the workers' throughput of valid results increases significantly in BOINC projects. Trilce Estrada, David A. Flores, Michela Taufer, Patricia J. Teller, Andre Kerstens, David P. Anderson |
e-Science | 1 |
| 2006 | Poster reception - SimBA: a discrete event simulator for performance prediction of volunteer computing projectsabstractSimBA (Simulator of BOINC Applications) is a discrete event simulator that accurately models the main functions of BOINC, a master-worker runtime framework. SimBA generates, distributes, and monitors tasks executed in a highly volatile, heterogeneous, and distributed environment such as Volunteer Computing (VC). In addition, it collects and validates results of executed tasks.To understand the strengths and weaknesses of BOINC under distinct scenarios, project designers must study and quantify its performance without affecting the VC community. Although this is not possible on production systems, it is possible using SimBA, which is capable of testing a wide range of hypotheses in a short period of time.Our experience to date indicates that SimBA is a reliable tool for performance prediction of VC projects. Preliminary results show that SimBA's predictions of [email protected] performance are within approximately 5% of the performance reported by this BOINC project. David A. Flores, Trilce Estrada, Michela Taufer, Patricia J. Teller, Andre Kerstens |
SC | 2 |