Thu D. Nguyen

dblp:n/ThuDNguyen · also Thu Duc Nguyen · DBLP profile ↗
← Back
71ranked-venue papers
5as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 3 first-author · 2 since 2021Software engineering, systems software and programming languages · 12 · 1 first-authorComputer networks · 7Security and privacy · 7Human-computer interaction and ubiquitous computing · 6Databases, data management, data science and information retrieval · 5Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
28 papers
Energy-efficient computing · 38% Cloud and datacenter computing · 26% Storage systems · 11%
Databases, data mining, and information retrieval
4 papers
Information retrieval · 66% Query processing and optimization · 23% Database system architecture and tuning · 10%

Topics — the 30 heaviest of 76, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Energy-efficient computing › thermal management
cooling optimization
0.912025
SibylOpt: Managing Green Data Centers Using Off-Online Deep Reinforcement Learning · HPDC 2025
Cloud and datacenter computing
datacenter operations
0.912025
SibylOpt: Managing Green Data Centers Using Off-Online Deep Reinforcement Learning · HPDC 2025
Energy-efficient computing
green datacenter
0.912025
SibylOpt: Managing Green Data Centers Using Off-Online Deep Reinforcement Learning · HPDC 2025
Energy-efficient computing
renewable energy
0.642025
SibylOpt: Managing Green Data Centers Using Off-Online Deep Reinforcement Learning · HPDC 2025
Parasol and GreenSwitch: managing datacenters powered by renewable energy · ASPLOS 2013
GreenHadoop: leveraging green energy in data-processing frameworks · EuroSys 2012
Storage systems › storage reliability
disk reliability
0.522016
Environmental Conditions and Disk Reliability in Free-cooled Datacenters · USENIX ATC 2016
Environmental Conditions and Disk Reliability in Free-cooled Datacenters · FAST 2016
Storage systems
storage reliability
0.522016
Environmental Conditions and Disk Reliability in Free-cooled Datacenters · USENIX ATC 2016
Environmental Conditions and Disk Reliability in Free-cooled Datacenters · FAST 2016
Energy-efficient computing
datacenter power management
0.322013
Parasol and GreenSwitch: managing datacenters powered by renewable energy · ASPLOS 2013
GreenHadoop: leveraging green energy in data-processing frameworks · EuroSys 2012
Embedded and real-time systems › real-time scheduling › multicore scheduling
heterogeneous multicore scheduling
0.312017
Exploiting heterogeneity for tail latency and energy efficiency · MICRO 2017
Electronic design automation › high-level synthesis
scheduling
0.312017
Exploiting heterogeneity for tail latency and energy efficiency · MICRO 2017
Distributed systems
fault tolerance
0.252010
Barricade: defending systems against operator mistakes · EuroSys 2010
Quantifying and Improving the Availability of High-Performance Cluster-Based Internet Services · SC 2003
Lazy Garbage Collection of Recovery State for Fault-Tolerant Distributed Shared Memory · IEEE Trans. Parallel Distributed Syst. 2002
Cloud and datacenter computing › resource management
datacenter resource management
0.222011
GreenSlot: scheduling energy consumption in green datacenters · SC 2011
Managing the cost, energy consumption, and carbon footprint of internet services · SIGMETRICS 2010
Emerging computing paradigms
approximate computing
0.212015
ApproxHadoop: Bringing Approximations to MapReduce Frameworks · ASPLOS 2015
Cloud and datacenter computing › cluster computing framework
mapreduce framework
0.212015
ApproxHadoop: Bringing Approximations to MapReduce Frameworks · ASPLOS 2015
Hardware reliability and fault tolerance › reliability analysis
thermal reliability
0.212015
CoolAir: Temperature- and Variation-Aware Management for Free-Cooled Datacenters · ASPLOS 2015
Energy-efficient computing
power management
0.222013
Parasol and GreenSwitch: managing datacenters powered by renewable energy · ASPLOS 2013
Managing the cost, energy consumption, and carbon footprint of internet services · SIGMETRICS 2010
Cloud and datacenter computing
cluster resource management and scheduling
0.222012
GreenHadoop: leveraging green energy in data-processing frameworks · EuroSys 2012
Improving cluster availability using workstation validation · SIGMETRICS 2002
Query processing and optimization › flexible queries
fuzzy query processing
0.112012
Efficient Multidimensional Fuzzy Search for Personal Information Management Systems · IEEE Trans. Knowl. Data Eng. 2012
Information retrieval
personal information management
0.112012
Efficient Multidimensional Fuzzy Search for Personal Information Management Systems · IEEE Trans. Knowl. Data Eng. 2012
Energy-efficient computing › datacenter power management
electricity cost minimization
0.112011
Reducing electricity cost through virtual machine placement in high performance computing clouds · SC 2011
Cloud and datacenter computing › datacenter architecture
geo-distributed datacenters
0.112011
Reducing electricity cost through virtual machine placement in high performance computing clouds · SC 2011
Energy-efficient computing › energy-aware scheduling
green energy-aware scheduling
0.112011
GreenSlot: scheduling energy consumption in green datacenters · SC 2011
Cloud and datacenter computing
job scheduling
0.112011
GreenSlot: scheduling energy consumption in green datacenters · SC 2011
Cloud and datacenter computing › virtualization › virtual machine management
virtual machine placement
0.112011
Reducing electricity cost through virtual machine placement in high performance computing clouds · SC 2011
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.112010
Barricade: defending systems against operator mistakes · EuroSys 2010
Memory systems › shared memory
distributed shared memory
0.132002
Lazy Garbage Collection of Recovery State for Fault-Tolerant Distributed Shared Memory · IEEE Trans. Parallel Distributed Syst. 2002
Lazy Garbage Collection of Recovery State for Fault-Tolerant Distributed Shared Memory · IEEE Trans. Parallel Distributed Syst. 2002
Scalable Fault-Tolerant Distributed Shared Memory · SC 2000
Memory systems › memory consistency › memory consistency model › release consistency
lazy release consistency
0.132002
Lazy Garbage Collection of Recovery State for Fault-Tolerant Distributed Shared Memory · IEEE Trans. Parallel Distributed Syst. 2002
Lazy Garbage Collection of Recovery State for Fault-Tolerant Distributed Shared Memory · IEEE Trans. Parallel Distributed Syst. 2002
Scalable Fault-Tolerant Distributed Shared Memory · SC 2000
Processor architecture and microarchitecture › multicore design › heterogeneous multicore
asymmetric multicore
0.112017
Exploiting heterogeneity for tail latency and energy efficiency · MICRO 2017
Energy-efficient computing › power management
dynamic voltage and frequency scaling
0.112017
Exploiting heterogeneity for tail latency and energy efficiency · MICRO 2017
Information retrieval › similarity search
fuzzy search
0.112008
Fuzzy Multi-Dimensional Search in the Wayfinder File System · ICDE 2008
Information retrieval › document retrieval › structured document retrieval
semi-structured retrieval
0.112008
Fuzzy Multi-Dimensional Search in the Wayfinder File System · ICDE 2008

Methods — techniques the papers use, named apart from their topics

offline reinforcement learning · 0.9deep reinforcement learning · 0.9advantage-weighted actor-critic · 0.9PPO · 0.9threshold-based scheduling · 0.3control theory · 0.3simulation · 0.3environmental condition analysis · 0.2multidimensional scoring · 0.2statistical error bounds · 0.2input data sampling · 0.2indexing · 0.1fuzzy search · 0.1mistake detection · 0.1mistake confinement · 0.1fuzzy querying · 0.1runtime monitors · 0.1heuristic dependency detection · 0.1
YearPublicationVenuePosition
2025 SibylOpt: Managing Green Data Centers Using Off-Online Deep Reinforcement Learning
abstract
We introduce SibylOpt, a system that applies deep reinforcement learning (DRL) to optimize the operation of a green data center (DC). SibylOpt uses an offline reinforcement learning algorithm, Advantage-Weighted Actor-Critic (AWAC), that requires historical data for training but does not depend on a DC simulator, which is effort intensive to build and keep updated as the DC evolves. SibylOpt then augments the offline training with online learning for refinement and adaptation to changes. We apply SibylOpt to the management of a small green DC that has onsite solar energy generation and a hybrid cooling system that includes "free-cooling." Evaluation results (using simulation) show that the offline trained SibylOpt achieves higher rewards trading off job wait time, cooling, and grid electricity consumption compared to two baseline policies. It is also competitive with PPO, a DRL approach that requires an accurate DC simulator for training.
Ning Gu 0004, Thu D. Nguyen, Peijian Wang, Tania Lorido-Botran
HPDC3
2022 GreenDRL: managing green datacenters using deep reinforcement learning
abstract
Managing datacenters to maximize efficiency and sustain-ability is a complex and challenging problem. In this work, we explore the use of deep reinforcement learning (RL) to manage "green" datacenters, bringing a robust approach for designing efficient management systems that account for specific workload, datacenter, and environmental characteristics. We design and evaluate GreenDRL, a system that combines a deep RL agent with simple heuristics to manage workload, energy consumption, and cooling in the presence of onsite generation of renewable energy to minimize brown energy consumption and cost. Our design addresses several important challenges, including adaptability, robustness, and effective learning in an environment comprising an enormous state/action space and multiple stochastic processes. Evaluation results (using simulation) show that GreenDRL is able to learn important principles such as delaying deferrable jobs to leverage variable generation of renewable (solar) energy, and avoiding the use of power-intensive cooling settings even at the expense of leaving some renewable energy unused. In an environment where a fraction of the workload is deferrable by up to 12 hours, GreenDRL can reduce grid electricity consumption for days with different solar energy generation and temperature characteristics by 32--54% compared to a FIFO baseline approach. GreenDRL also matches or outperforms a management approach that uses linear programming together with oracular future knowledge to manage workload and server energy consumption, but leaves the management of the cooling system to a separate (and independent) controller. Overall, our work shows that deep RL is a promising technique for building efficient management systems for green datacenters.
Peijian Wang, Ning Gu 0004, Thu D. Nguyen
SoCC4
2022 Gender Diversity in Computer Science at a Large Public R1 Research University: Reporting on a Self-study
abstract
With the number of jobs in computer occupations on the rise, there is a greater need for computer science (CS) graduates than ever. At the same time, most CS departments across the country are only seeing 25–30% of women students in their classes, meaning that we are failing to draw interest from a large portion of the population. In this work, we explore the gender gap in CS at Rutgers University–New Brunswick, a large public R1 research university, using three data sets that span thousands of students across six academic years. Specifically, we combine these data sets to study the gender gaps in four core CS courses and explore the correlation of several factors with retention and the impact of these factors on changes to the gender gap as students proceed through the CS courses toward completing the CS major. For example, we find that a significant percentage of women students taking the introductory CS1 course for majors do not intend to major in CS, which may be a contributing factor to a large increase in the gender gap immediately after CS1. This finding implies that part of the retention task is attracting these women students to further explore the major. Results from our study include both novel findings and findings that are consistent with known challenges for increasing gender diversity in CS. In both cases, we provide extensive quantitative data in support of the findings.
Monica Babes-Vroman, Thuytien N. Nguyen, Thu D. Nguyen
ACM Trans. Comput. Educ.3
2021 CSF: Formative Feedback in Autograding
abstract
Autograding systems are being increasingly deployed to meet the challenges of teaching programming at scale. Studies show that formative feedback can greatly help novices learn programming. This work extends an autograder, enabling it to provide formative feedback on programming assignment submissions. Our methodology starts with the design of a knowledge map, which is the set of concepts and skills that are necessary to complete an assignment, followed by the design of the assignment and that of a comprehensive test suite for identifying logical errors in the submitted code. Test cases are used to test the student submissions and learn classes of common errors. For each assignment, we train a classifier that automatically categorizes errors in a submission based on the outcome of the test suite. The instructor maps the errors to corresponding concepts and skills and writes hints to help students find their misconceptions and mistakes. We apply this methodology to two assignments in our Introduction to Computer Science course and find that the automatic error categorization has a 90% average accuracy. We report and compare data from two semesters, one semester when hints are given for the two assignments and one when hints are not given. Results show that the percentage of students who successfully complete the assignments after an initial erroneous submission is three times greater when hints are given compared to when hints are not given. However, on average, even when hints are provided, almost half of the students fail to correct their code so that it passes all the test cases. The initial implementation of the framework focuses on the functional correctness of the programs as reflected by the outcome of the test cases. In our future work, we will explore other kinds of feedback and approaches to automatically generate feedback to better serve the educational needs of the students.
Georgiana Haldeman, Monica Babes-Vroman, Andrew Tjang, Thu D. Nguyen
ACM Trans. Comput. Educ.4
2020 Exploring Novice Programmers' Homework Practices: Initial Observations of Information Seeking Behaviors
abstract
There are many factors that contribute to the success of students learning to code. For students in introductory programming classes, one source of complexity is the availability of a wide variety of information sources. In this paper, we report observations of students seeking information when working on programming homework assignments. Our data was collected from a think-aloud protocol embedded into semi-structured, individual interviews with students enrolled in a CS1 course. We analyze our data through the lens of information seeking behavior. We observed students using multiple sources of information, including referring back to course materials and searching for information online, and discussing how they sought help from friends, classmates, and family members. Herein, we discuss implications for teaching and future research based on our initial observations. For example, instructors could consider designing early homework assignments that would prompt students to seek information and follow up this assignment with an in-class discussion about homework strategies. Future research could investigate the mechanisms by which students progress from haphazard to more strategic information seeking behaviors.
Silvia Muller, Monica Babes-Vroman, Mary Emenike, Thu D. Nguyen
SIGCSE4
2019 Catalyst: A Cloud-based Media Processing Framework
abstract
Massive increases in production and consumption of digital media are motivating cloud "video encoding as a service." In this paper, we explore efficient platforms for supporting such services. Specifically, we explore transcoding performance on a software encoder and two hardware-accelerated encoders representing widely different points in the hardware acceleration design space, as well as interference effects when heterogenous encoders are run concurrently. Using results from our exploratory study, we next propose a framework for providing video encoding as a service on a single server, and implement a prototype of the framework in a system called Catalyst. Catalyst accepts transcoding tasks and intelligently uses all available resources, including software and hardware-accelerated encoders, to maximize throughput while respecting task deadlines, and provides a foundational building block for a cluster-based cloud video encoding service. We evaluate Catalyst using a synthetic transcoding workload designed to emulate an IP-TV/Cloud-DVR workload. Evaluation results show that Catalyst can significantly increase throughput while meeting task deadlines compared to naive use of hardware-accelerated encoding.
William A. Katsak, Kiran Nagaraja, Joacim Halén, Nimish Radia, Thu D. Nguyen
ICDCS6
2019 Approximation with Error Bounds in Spark
abstract
Many decision-making queries are based on aggregating massive amounts of data, where sampling is an important approximation technique for reducing execution times. It is important to estimate error bounds when sampling to help users balance between precision and performance. However, error bound estimation is challenging because data processing pipelines often transform the input dataset in complex ways before computing the final aggregated values. In this paper, we introduce a sampling framework to support approximate computing with estimated error bounds in Spark. Our framework allows sampling to be performed at multiple arbitrary points within a sequence of transformations preceding an aggregation operation. The framework constructs a data provenance tree to maintain information about how transformations are clustering output data items to be aggregated. It then uses the tree and multi-stage sampling theories to compute the approximate aggregate values and corresponding error bounds. When information about output keys are available early, the framework can also use adaptive stratified reservoir sampling to avoid (or reduce) key losses in the final output and to achieve more consistent error bounds across popular and rare keys. Finally, the framework includes an algorithm to dynamically choose sampling rates to meet user-specified constraints on the CDF of error bounds in the outputs. We have implemented a prototype of our framework called ApproxSpark and used it to implement five approximate applications from different domains. Evaluation results show that ApproxSpark can (a) significantly reduce execution time if users can tolerate small amounts of uncertainties and, in many cases, loss of rare keys, and (b) automatically find sampling rates to meet user-specified constraints on error bounds. We also explore and discuss extensively tradeoffs between sampling rates, execution time, accuracy and key loss.
Guangyan Hu, Sandro Rigo, Desheng Zhang 0002, Thu D. Nguyen
MASCOTS4
2019 Dynamic Recitation: A Student-Focused, Goal-Oriented Recitation Management Platform
abstract
Computer science universities and colleges around the nation are experiencing large growth in enrollments. To maintain in-person interaction with students, large courses typically include multiple recitations, each led by a Teaching Assistant (TA). Students in each group struggle with various course content, and these weaknesses are best evaluated by the TAs working closely with the students. TAs, however, have only a superficial understanding of education theory, and instructors must closely monitor and evaluate the content of recitations. Existing instruction management systems can be used to organize course-wide content, but, to our knowledge, none of them operate at the granularity of recitations. This poster presents Dynamic Recitation, an open-source platform through which TAs and instructors can create and share practice problems and lesson plans tagged with the learning objectives that they cover. The poster illustrates the interface offered to the TAs for creating problems, designing lesson plans for individual sections, and submitting feedback about student progress. It shows examples of lesson plans created in the system, as well as reports that can be used to identify elements that promote desired learning objectives and refine recitations.
Joseph A. Boyle, Georgiana Haldeman, Andrew Tjang, Monica Babes-Vroman, Ana Paula Centeno, Thu D. Nguyen
SIGCSE6
2019 Living-Learning Community for Women in Computer Science at Rutgers
abstract
We describe our experience developing and running a Computer Science Living-Learning Community (LLC) for first-year women at Rutgers University, now in its third year. Each year, around 20 first-year undergraduate women who intend to major in computer science (CS) apply and are selected to participate. LLC participants live in a common residence hall and are provided with an educational, mentoring, and community-building program that supports their progress as students and CS majors. Participants take a "house course," Great Ideas and Insights in Computer Science, as a group, and also take a course on Knowledge and Power: Issues in Women's Leadership. Program activities include study sessions and industry interactions, as well as opportunities to participate in K-12 outreach programs, hackathons, and computing research. To evaluate the program, participants and a similar comparison group are surveyed at the beginning and end of the academic year and a focus group is conducted with program participants. Program participants find the program valuable and would recommend it to others, but both program participants and the comparison group report some lack of confidence in their potential success as computer scientists.
Rebecca N. Wright, Sally J. Nadler, Thu D. Nguyen, Cynthia Sanchez Gomez, Heather M. Wright
SIGCSE3
2018 Uncertainty Propagation in Data Processing Systems
abstract
We are seeing an explosion of uncertain data---i.e., data that is more properly represented by probability distributions or estimated values with error bounds rather than exact values---from sensors in IoT, sampling-based approximate computations and machine learning algorithms. In many cases, performing computations on uncertain data as if it were exact leads to incorrect results. Unfortunately, developing applications for processing uncertain data is a major challenge from both the mathematical and performance perspectives. This paper proposes and evaluates an approach for tackling this challenge in DAG-based data processing systems. We present a framework for uncertainty propagation (UP) that allows developers to modify precise implementations of DAG nodes to process uncertain inputs with modest effort. We implement this framework in a system called UP-MapReduce, and use it to modify ten applications, including AI/ML, image processing and trend analysis applications to process uncertain data. Our evaluation shows that UP-MapReduce propagates uncertainties with high accuracy and, in many cases, low performance overheads. For example, a social network trend analysis application that combines data sampling with UP can reduce execution time by 2.3x when the user can tolerate a maximum relative error of 5% in the final answer. These results demonstrate that our UP framework presents a compelling approach for handling uncertain data in DAG processing.
Ioannis Manousakis, Íñigo Goiri, Ricardo Bianchini, Sandro Rigo, Thu D. Nguyen
SoCC5
2018 Providing Meaningful Feedback for Autograding of Programming Assignments
abstract
Autograding systems are increasingly being deployed to meet the challenge of teaching programming at scale. We propose a methodology for extending autograders to provide meaningful feedback for incorrect programs. Our methodology starts with the instructor identifying the concepts and skills important to each programming assignment, designing the assignment, and designing a comprehensive test suite. Tests are then applied to code submissions to learn classes of common errors and produce classifiers to automatically categorize errors in future submissions. The instructor maps the errors to concepts and skills and writes hints to help students find their misconceptions and mistakes. We have applied the methodology to two assignments from our Introduction to Computer Science course. We used submissions from one semester of the class to build classifiers and write hints for observed common errors. We manually validated the automatic error categorization and potential usefulness of the hints using submissions from a second semester. We found that the hints given for erroneous submissions should be helpful for 96% or more of the cases. Based on these promising results, we have deployed our hints and are currently collecting submissions and feedback from students and instructors.
Georgiana Haldeman, Andrew Tjang, Monica Babes-Vroman, Stephen Bartos, Jay Shah, Danielle Yucht, Thu D. Nguyen
SIGCSE7
2018 Computer Science Living-Learning Community for Women at Rutgers: Initial Experiences and Outcomes (Abstract Only)
abstract
We have developed the Douglass-SAS-DIMACS Computer Science Living-Learning Community (LLC) for first-year women at Rutgers, now in its second year. Each year, around 20 first-year women undergraduates at Rutgers who intend to major in computer science are selected for the LLC. LLC participants live in a common dorm and are provided with an educational, mentoring, and community-building program that supports their progress as Rutgers students and as computer science majors. To our knowledge, this is the first undergraduate living-learning community for women in computer science at any university. A focus group conducted with women from the inaugural cohort revealed that faculty support contributed to feelings of belonging, both in the program and in the CS department, among the participants; participants valued the academic support they received as part of the program and felt communication structures within the program were effective; and participants expressed a desire for advanced undergraduate peer mentors. A quasi-experimental study of this cohort indicated that LLC participants showed a decrease in satisfaction with the CS department at Rutgers; a decrease in computing-related self-efficacy; and an increase in the belief that computing ability is inborn. Follow up interviews suggested that the efficacy of the LLC might be dependent on two factors: participants' commitment to a CS major coming into the program and participants' level of involvement with the LLC group. In response to these results, we have made some changes to the program and continue to carefully study the program in order to maximize its effectiveness.
Rebecca N. Wright, Jane Stout, Geraldine Cochran, Thu D. Nguyen, Cynthia Sanchez Gomez
SIGCSE4
2017 Exploiting heterogeneity for tail latency and energy efficiency
abstract
Interactive service providers have strict requirements on high-percentile (tail) latency to meet user expectations. If providers meet tail latency targets with less energy, they increase profits, because energy is a significant operating expense. Unfortunately, optimizing tail latency and energy are typically conflicting goals. Our work resolves this conflict by exploiting servers with per-core Dynamic Voltage and Frequency Scaling (DVFS) and Asymmetric Multicore Processors (AMPs). We introduce the Adaptive Slow-to-Fast scheduling framework, which matches the heterogeneity of the workload --- a mix of short and long requests --- to the heterogeneity of the hardware --- cores running at different speeds. The scheduler prioritizes long requests to faster cores by exploiting the insight that long requests reveal themselves. We use control theory to design threshold-based scheduling policies that use individual request progress, load, competition, and latency targets to optimize performance and energy. We configure our framework to optimize Energy Efficiency for a given Tail Latency (EETL) for both DVFS and AMP. In this framework, each request self-schedules, starting on a slow core and then migrating itself to faster cores. At high load, when a desired AMP core speed s is not available for a request but a faster core is, the longest request on an s core type migrates early to make room for the other request. Compared to per-core DVFS systems, EETL for AMPs delivers the same tail latency, reduces energy by 18% to 50%, and improves capacity (throughput) by 32% to 82%. We demonstrate that our framework effectively exploits dynamic DVFS and static AMP heterogeneity to reduce provisioning and operational costs for interactive services.
Md. Enamul Haque, Yuxiong He, Sameh Elnikety, Thu D. Nguyen, Ricardo Bianchini, Kathryn S. McKinley
MICRO4
2017 Exploring Gender Diversity in CS at a Large Public R1 Research University
abstract
With the number of Computer Science (CS) jobs on the rise, there is a greater need for Computer Science graduates than ever. At the same time, most CS departments across the country are only seeing 25-30% of female students in their classes, meaning that we are failing to draw interest from a large portion of the population. In this work, we explore the gender gap in CS at Rutgers University using three data sets that span thousands of students across 3.5 academic years. By combining these data sets, we can explore interesting issues such as retention, as students progress through the CS major. For example, we find that a large percentage of women taking the Introductory CS1 course for majors do not intend to major in CS, which contributes to a large increase in the gender gap immediately after CS1. This finding implies that a large part of the retention task is attracting these women to further explore the major. We correlate our findings with initiatives that some CS programs across the country have taken to significantly improve their gender diversity, and identify initiatives that we can start with in our effort to increase the diversity in our program. These findings may also be applicable to the computing programs at other large public research universities.
Monica Babes-Vroman, Isabel Juniewicz, Bruno Lucarelli, Nicole Fox, Thu D. Nguyen, Andrew Tjang, Georgiana Haldeman, Ashni Mehta, Risham Chokshi
SIGCSE5
2016 Environmental Conditions and Disk Reliability in Free-cooled Datacenters
Ioannis Manousakis, Sriram Sankar, Gregg McKnight, Thu D. Nguyen, Ricardo Bianchini
FAST4
2016 Environmental Conditions and Disk Reliability in Free-cooled Datacenters
Ioannis Manousakis, Sriram Sankar, Gregg McKnight, Thu D. Nguyen, Ricardo Bianchini
USENIX ATC4
2015 ApproxHadoop: Bringing Approximations to MapReduce Frameworks
abstract
We propose and evaluate a framework for creating and running approximation-enabled MapReduce programs. Specifically, we propose approximation mechanisms that fit naturally into the MapReduce paradigm, including input data sampling, task dropping, and accepting and running a precise and a user-defined approximate version of the MapReduce code. We then show how to leverage statistical theories to compute error bounds for popular classes of MapReduce programs when approximating with input data sampling and/or task dropping. We implement the proposed mechanisms and error bound estimations in a prototype system called ApproxHadoop. Our evaluation uses MapReduce applications from different domains, including data analytics, scientific computing, video encoding, and machine learning. Our results show that ApproxHadoop can significantly reduce application execution time and/or energy consumption when the user is willing to tolerate small errors. For example, ApproxHadoop can reduce runtimes by up to 32x when the user can tolerate an error of 1% with 95% confidence. We conclude that our framework and system can make approximation easily accessible to many application domains using the MapReduce model.
Íñigo Goiri, Ricardo Bianchini, Santosh Nagarakatte, Thu D. Nguyen
ASPLOS4
2015 CoolAir: Temperature- and Variation-Aware Management for Free-Cooled Datacenters
abstract
Despite its benefits, free cooling may expose servers to high absolute temperatures, wide temperature variations, and high humidity when datacenters are sited at certain locations. Prior research (in non-free-cooled datacenters) has shown that high temperatures and/or wide temporal temperature variations can harm hardware reliability. In this paper, we identify the runtime management strategies required to limit absolute temperatures, temperature variations, humidity, and cooling energy in free-cooled datacenters. As the basis for our study, we propose CoolAir, a system that embodies these strategies. Using CoolAir and a real free-cooled datacenter prototype, we show that effective management requires cooling infrastructures that can act smoothly. In addition, we show that CoolAir can tightly manage temperature and significantly reduce temperature variation, often at a lower cooling cost than existing free-cooled datacenters. Perhaps most importantly, based on our results, we derive several principles and lessons that should guide the design of management systems for free-cooled datacenters of any size.
Íñigo Goiri, Thu D. Nguyen, Ricardo Bianchini
ASPLOS2
2015 CoolProvision: underprovisioning datacenter cooling
abstract
Cloud providers have made significant strides in reducing the cooling capital and operational costs of their datacenters, for example, by leveraging outside air ("free") cooling where possible. Despite these advances, cooling costs still represent a significant expense mainly because cloud providers typically provision their cooling infrastructure for the worst-case scenario (i.e., very high load and outside temperature at the same time). Thus, in this paper, we propose to reduce cooling costs by underprovisioning the cooling infrastructure. When the cooling is underprovisioned, there might be (rare) periods when the cooling infrastructure cannot cool down the IT equipment enough. During these periods, we can either (1) reduce the processing capacity and potentially degrade the quality of service, or (2) let the IT equipment temperature increase in exchange for a controlled degradation in reliability. To determine the ideal amount of underprovisioning, we introduce CoolProvision, an optimization and simulation framework for selecting the cheapest provisioning within performance constraints defined by the provider. CoolProvision leverages an abstract trace of the expected workload, as well as cooling, performance, power, reliability, and cost models to explore the space of potential provisionings. Using data from a real small free-cooled datacenter, our results suggest that CoolProvision can reduce the cost of cooling by up to 55%. We extrapolate our experience and results to larger cloud datacenters as well.
Ioannis Manousakis, Íñigo Goiri, Sriram Sankar, Thu D. Nguyen, Ricardo Bianchini
SoCC4
2015 GreenPar: Scheduling Parallel High Performance Applications in Green Datacenters
abstract
We propose GreenPar, a scheduler for parallel high-perormance applications in datacenters partially powered by on-site generation of renewable ("green'') energy. GreenPar schedules the workload to maximize the green energy consumption and minimize the grid ("brown'') energy consumption, while respecting a performance service-level agreement (SLA). When green energy is available, GreenPar increases the resource allocations of active jobs to reduce runtimes. When using brown energy, GreenPar reduces resource allocations within the constraints imposed by the performance SLA to conserve energy. GreenPar makes its decisions based on the speedup profile of each job. We have implemented GreenPar in a real solar-powered datacenter. Our results show that GreenPar can increase the green energy consumption and reduce both the average job runtime and the brown energy consumption, compared to schedulers that are oblivious to on-site green energy.
Md. Enamul Haque, Íñigo Goiri, Ricardo Bianchini, Thu D. Nguyen
ICS4
2015 EdgeBuffer: Caching and prefetching content at the edge in the MobilityFirst future Internet architecture
abstract
The prevalence of mobile devices especially smartphones has attracted research on mobile content delivery techniques. In this paper, we propose to take advantage of the storage available at wireless access points to bring content closer to mobile devices, hence improving the downloading performance. Specifically, we propose to have a separate popularity based cache and a prefetch buffer at the network edge to capture both long-term and short-term content access patterns. Further, we point out that it is insufficient to rely on a device's past history to predict when and where to prefetch, especially in urban settings; instead, we propose to derive a prediction model based on the aggregated network-level statistics. We discuss the proposed mobile content caching/prefetching method in the context of the MobilityFirst future Internet architecture. In MobilityFirst, when mobile clients move between network attachment points (e.g., Wi-Fi access points), their network association records are logged by the network, which then naturally facilitates the network-level mobility prediction. Through detailed simulations with real taxi mobility traces, we show that such a strategy is more effective than earlier schemes in satisfying content requests at the edge (higher cache hit ratios), leading to shorter content download latencies. Specifically, the fraction of requests satisfied at the edge increases by a factor of 2.9 compared to a caching only approach, and by 45% compared to individual user-based prediction and prefetching.
Feixiong Zhang, Chenren Xu, Yanyong Zhang, K. K. Ramakrishnan, Shreyasee Mukherjee, Roy D. Yates, Thu D. Nguyen
WOWMOM7
2015 Matching renewable energy supply and demand in green datacenters
Íñigo Goiri, Md. Enamul Haque, Kien Le, Ryan Beauchea, Thu D. Nguyen, Jordi Guitart, Jordi Torres, Ricardo Bianchini
Ad Hoc Networks5
2014 Building Green Cloud Services at Low Cost
abstract
Interest in powering data enters at least partially using on-site renewable sources, e.g. solar or wind, has been growing. In fact, researchers have studied distributed services comprising networks of such "green" data centers, and load distribution approaches that "follow the renewables" to maximize their use. However, prior works have not considered where to site such a network for efficient production of renewable energy, while minimizing both data center and renewable plant building costs. Moreover, researchers have not built real load management systems for follow-the-renewables services. Thus, in this paper, we propose a framework, optimization problem, and solution approach for sitting and provisioning green data centers for a follow-the-renewables HPC cloud service. We illustrate the location selection tradeoffs by quantifying the minimum cost of achieving different amounts of renewable energy. Finally, we design and implement a system capable of migrating virtual machines across the green data centers to follow the renewables. Among other interesting results, we demonstrate that one can build green HPC cloud services at a relatively low additional cost compared to existing services.
Josep Lluís Berral, Íñigo Goiri, Thu D. Nguyen, Ricard Gavaldà, Jordi Torres, Ricardo Bianchini
ICDCS3
2013 Parasol and GreenSwitch: managing datacenters powered by renewable energy
abstract
Several companies have recently announced plans to build "green" datacenters, i.e. datacenters partially or completely powered by renewable energy. These datacenters will either generate their own renewable energy or draw it directly from an existing nearby plant. Besides reducing carbon footprints, renewable energy can potentially reduce energy costs, reduce peak power costs, or both. However, certain renewable fuels are intermittent, which requires approaches for tackling the energy supply variability. One approach is to use batteries and/or the electrical grid as a backup for the renewable energy. It may also be possible to adapt the workload to match the renewable energy supply. For highest benefits, green datacenter operators must intelligently manage their workloads and the sources of energy at their disposal.
Íñigo Goiri, William A. Katsak, Kien Le, Thu D. Nguyen, Ricardo Bianchini
ASPLOS4
2013 Quantifying and improving I/O predictability in virtualized systems
abstract
Virtualization enables the consolidation of virtual machines (VMs) to increase the utilization of physical servers in Infrastructure-as-a-Service (IaaS) cloud providers. However, our experience shows that storage I/O performance varies wildly in the face of consolidation. Since many users may desire consistent performance, we argue that IaaS providers should offer a class of predictable-performance service in addition to existing (predictability-oblivious) services. Thus, we propose VirtualFence, a storage system that provides predictable VM performance. VirtualFence uses three main techniques: (1) non-work-conserving time-division I/O scheduling, (2) a small solid-state (SSD) cache in front of a much larger hard disk drive (HDD), and (3) space-partitioning of both the SSD cache and the HDD. Our evaluation shows that VirtualFence improves predictability significantly, while allowing cloud providers to reach any desired compromise between predictability and performance.
Íñigo Goiri, Abhishek Bhattacharjee, Ricardo Bianchini, Thu D. Nguyen
IWQoS5
2012 GreenHadoop: leveraging green energy in data-processing frameworks
abstract
Interest has been growing in powering datacenters (at least partially) with renewable or "green" sources of energy, such as solar or wind. However, it is challenging to use these sources because, unlike the "brown" (carbon-intensive) energy drawn from the electrical grid, they are not always available. This means that energy demand and supply must be matched, if we are to take full advantage of the green energy to minimize brown energy consumption. In this paper, we investigate how to manage a datacenter's computational workload to match the green energy supply. In particular, we consider data-processing frameworks, in which many background computations can be delayed by a bounded amount of time. We propose GreenHadoop, a MapReduce framework for a datacenter powered by a photovoltaic solar array and the electrical grid (as a backup). GreenHadoop predicts the amount of solar energy that will be available in the near future, and schedules the MapReduce jobs to maximize the green energy consumption within the jobs' time bounds. If brown energy must be used to avoid time bound violations, GreenHadoop selects times when brown energy is cheap, while also managing the cost of peak brown power consumption. Our experimental results demonstrate that GreenHadoop can significantly increase green energy consumption and decrease electricity cost, compared to Hadoop.
Íñigo Goiri, Kien Le, Thu D. Nguyen, Jordi Guitart, Jordi Torres, Ricardo Bianchini
EuroSys3
2012 DMap: A Shared Hosting Scheme for Dynamic Identifier to Locator Mappings in the Global Internet
abstract
This paper presents the design and evaluation of a novel distributed shared hosting approach, DMap, for managing dynamic identifier to locator mappings in the global Internet. DMap is the foundation for a fast global name resolution service necessary to enable emerging Internet services such as seamless mobility support, content delivery and cloud computing. Our approach distributes identifier to locator mappings among Autonomous Systems (ASs) by directly applying K>1 consistent hash functions on the identifier to produce network addresses of the AS gateway routers at which the mapping will be stored. This direct mapping technique leverages the reach ability information of the underlying routing mechanism that is already available at the network layer, and achieves low lookup latencies through a single overlay hop without additional maintenance overheads. The proposed DMap technique is described in detail and specific design problems such as address space fragmentation, reducing latency through replication, taking advantage of spatial locality, as well as coping with inconsistent entries are addressed. Evaluation results are presented from a large-scale discrete event simulation of the Internet with ~26,000 ASs using real-world traffic traces from the DIMES repository. The results show that the proposed method evenly balances storage load across the global network while achieving lookup latencies with a mean value of ~50 ms and 95th percentile value of ~100 ms, considered adequate for support of dynamic mobility across the global Internet.
Tam Vu 0001, Akash Baid, Yanyong Zhang, Thu D. Nguyen, Junichiro Fukuyama, Richard P. Martin, Dipankar Raychaudhuri
ICDCS4
2012 Efficient Multidimensional Fuzzy Search for Personal Information Management Systems
abstract
With the explosion in the amount of semistructured data users access and store in personal information management systems, there is a critical need for powerful search tools to retrieve often very heterogeneous data in a simple and efficient way. Existing tools typically support some IR-style ranking on the textual part of the query, but only consider structure (e.g., file directory) and metadata (e.g., date, file type) as filtering conditions. We propose a novel multidimensional search approach that allows users to perform fuzzy searches for structure and metadata conditions in addition to keyword conditions. Our techniques individually score each dimension and integrate the three dimension scores into a meaningful unified score. We also design indexes and algorithms to efficiently identify the most relevant files that match multidimensional queries. We perform a thorough experimental evaluation of our approach and show that our relaxation and scoring framework for fuzzy query conditions in noncontent dimensions can significantly improve ranking accuracy. We also show that our query processing strategies perform and scale well, making our fuzzy search approach practical for every day usage.
Wei Wang 0014, Christopher Peery, Amélie Marian, Thu D. Nguyen
IEEE Trans. Knowl. Data Eng.4
2011 Unified structure and content search for personal information management systems
abstract
User data stored in personal information systems is growing massively. Simultaneously, this data is increasingly distributed across multiple organizational domains such as email, music databases, and photo albums, some of which are structured automatically by applications. Powerful search tools are needed to help users locate data in these expanding yet fragmented data sets. In this paper, we present a novel fuzzy search approach that considers approximate matches to structure and content query conditions. Our framework uses unified data and query processing models so that structure conditions can be approximately matched by content and vice versa. Our models also unify external structure (e.g., directories) with internal structure (e.g., XML structure), supporting integrated queries matched to a single data domain. We propose indexes and algorithms for efficient query processing. We evaluate our approach using a real data set, showing that it can leverage structure information to significantly improve search accuracy, yet is robust to mistakes in query conditions.
Wei Wang 0014, Amélie Marian, Thu D. Nguyen
EDBT3
2011 GreenSlot: scheduling energy consumption in green datacenters
abstract
In this paper, we propose GreenSlot, a parallel batch job scheduler for a datacenter powered by a photovoltaic solar array and the electrical grid (as a backup). GreenSlot predicts the amount of solar energy that will be available in the near future, and schedules the workload to maximize the green energy consumption while meeting the jobs' deadlines. If grid energy must be used to avoid deadline violations, the scheduler selects times when it is cheap. Our results for production scientific workloads demonstrate that Green-Slot can increase green energy consumption by up to 117% and decrease energy cost by up to 39%, compared to a conventional scheduler. Based on these positive results, we conclude that green datacenters and green-energy-aware scheduling can have a significant role in building a more sustainable IT ecosystem.
Íñigo Goiri, Ryan Beauchea, Kien Le, Thu D. Nguyen, Md. Enamul Haque, Jordi Guitart, Jordi Torres, Ricardo Bianchini
SC4
2011 Reducing electricity cost through virtual machine placement in high performance computing clouds
abstract
In this paper, we first study the impact of load placement policies on cooling and maximum data center temperatures in cloud service providers that operate multiple geographically distributed data centers. Based on this study, we then propose dynamic load distribution policies that consider all electricity-related costs as well as transient cooling effects. Our evaluation studies the ability of different cooling strategies to handle load spikes, compares the behaviors of our dynamic cost-aware policies to cost-unaware and static policies, and explores the effects of many parameter settings. Among other interesting results, we demonstrate that (1) our policies can provide large cost savings, (2) load migration enables savings in many scenarios, and (3) all electricity-related costs must be considered at the same time for higher and consistent cost savings.
Kien Le, Ricardo Bianchini, Yogesh Jaluria, Jiandong Meng, Thu D. Nguyen
SC6
2011 MassConf: automatic configuration tuning by leveraging user community information
abstract
Configuring modern enterprise software can be extremely difficult because their behaviors often depend on large numbers of configuration parameters. Software vendors can simplify the configuration process for new users by collecting and using configuration information from existing users. In particular, we observe that (1) a 'good' configuration may work well for many different users, and (2) multiple configurations may work well for each user. We leverage these observations to design MassConf, a system that collects and uses existing configurations to automatically configure new software installations. Our evaluations with a case study confirm our observations and show that MassConf successfully reaches the targets of many more new installations than an existing efficient optimization algorithm.
Ricardo Bianchini, Thu D. Nguyen
ICPE3
2010 Barricade: defending systems against operator mistakes
abstract
In this paper, we propose a management framework for protecting large computer systems against operator mistakes. By detecting and confining mistakes to isolated portions of the managed system, our framework facilitates correct operation even by inexperienced operators. We built a prototype management system called Barricade based on our framework. We evaluate Barricade by deploying it for two different systems, a prototype Internet service and an enterprise computer infrastructure, and conducting experiments with 20 volunteer operators. Our results are very promising. For example, we show that Barricade can detect and contain 39 out of the 43 mistakes that we observed in 49 live operator experiments performed with our Internet service.
Fábio Oliveira, Andrew Tjang, Ricardo Bianchini, Richard P. Martin, Thu D. Nguyen
EuroSys5
2010 Managing the cost, energy consumption, and carbon footprint of internet services
abstract
The large amount of energy consumed by Internet services represents significant and fast-growing financial and environmental costs. This paper introduces a general, optimization-based framework and several request distribution policies that enable multi-data-center services to manage their brown energy consumption and leverage green energy, while respecting their service-level agreements (SLAs) and minimizing energy cost. Our policies can be used to abide by caps on brown energy consumption that might arise from various scenarios such as government imposed Kyoto-style carbon limits. Extensive simulations and real experiments show that our policies allow a service to trade off consumption and cost. For example, using our policies, a service can reduce brown energy consumption by 24% for only a 10% increase in cost, while still abiding by SLAs.
Kien Le, Ozlem Bilgir, Ricardo Bianchini, Margaret Martonosi, Thu D. Nguyen
SIGMETRICS5
2009 Model-Based Validation for Internet Services
abstract
Operator mistakes are a significant source of unavailability in Internet services. In our previous work, we proposed operator action validation as an approach for detecting mistakes while hiding them from the service and its users. Previous validation strategies have limitations, however, including the need for instances of correct behavior for comparison. In this paper, we propose a novel model-based validation strategy that addresses these limitations and complements our previous techniques. Model-based validation calls for service engineers to define models of Internet services that can be used to differentiate between correct and incorrect configurations and behaviors. These models are then used to guide the specification of validation assertions that check the correctness of operator actions before they are exposed. We have implemented a prototype model-based validation system for two services, the Web crawler of a commercial search engine (Ask.com) and an academic yet realistic online auction service. Experimentation with model-based validation demonstrates that it is highly effective at detecting and hiding both activated and latent mistakes.
Andrew Tjang, Fábio Oliveira, Ricardo Bianchini, Richard P. Martin, Thu D. Nguyen
SRDS5
2008 Multi-dimensional search for personal information management systems
abstract
With the explosion in the amount of semi-structured data users access and store in personal information management systems, there is a need for complex search tools to retrieve often very heterogeneous data in a simple and efficient way. Existing tools usually index text content, allowing for some IR-style ranking on the textual part of the query, but only consider structure (e.g., file directory) and metadata (e.g., date, file type) as filtering conditions. We propose a novel multi-dimensional approach to semi-structured data searches in personal information management systems by allowing users to provide fuzzy structure and metadata conditions in addition to keyword conditions. Our techniques provide a complex query interface that is more comprehensive than content-only searches as it considers three query dimensions (content, structure, metadata) in the search. We propose techniques to individually score each dimension, as well as a framework to integrate the three dimension scores into a meaningful unified score. Our work is integrated in Wayfinder, an existing fully-functioning file system. We perform a thorough experimental evaluation of our techniques to show the effect of approximating individual dimensions on the overall scores and ranks of files, as well as on query performance. Our experiments show that our scoring strategy adequately takes into account the approximation in each dimension to efficiently evaluate fuzzy multi-dimensional queries. In addition, fuzzy query conditions in non-content dimensions can significantly improve scoring (and thus ranking) accuracy.
Christopher Peery, Wei Wang 0014, Amélie Marian, Thu D. Nguyen
EDBT4
2008 Fuzzy Multi-Dimensional Search in the Wayfinder File System
abstract
With the explosion in the amount of semi-structured data users access and store, there is a need for complex search tools to retrieve often very heterogeneous data in a simple and efficient way. Existing tools usually index text content, allowing for some IR-style ranking on the textual part of the query, but only consider structure (e.g., file directory) and metadata (e.g., date, file type) as filtering conditions. We propose a novel multidimensional querying approach to semi-structured data searches in personal information systems by allowing users to provide fuzzy structure and metadata conditions in addition to traditional keyword conditions. The provided query interface is more comprehensive than content-only searches as it considers three query dimensions (content, structure, metadata) in the search. We have implemented our proposed approach in the Wayfinder file system. In this demo, we will use this implementation to both present an overview of the unified scoring framework underlying the fuzzy multi-dimensional querying approach and demonstrate its potential in improving search results.
Christopher Peery, Wei Wang 0014, Amélie Marian, Thu D. Nguyen
ICDE4
2007 Automatic configuration of internet services
abstract
Recent research has found that operators frequently misconfigure Internet services, causing various availability and performance problems. In this paper, we propose a software infrastructure that eliminates several types of misconfiguration by automating the generation of configuration files in Internet services, even as the services evolve. The infrastructure comprises a custom scripting language, configuration file templates, communicating runtime monitors, and heuristic algorithms to detect dependencies between configuration parameters and select ideal configurations. To demonstrate our infrastructure experimentally, we apply it to a realistic online auction service. Our results show that the infrastructure can simplify operation significantly while eliminating 58% of the misconfigurations found in a previous study of the same service. Furthermore, our results show that the infrastructure can efficiently determine the configuration parameters that lead to high performance as the service evolves through a hardware upgrade and the scheduled maintenance of a few nodes.
Ricardo Bianchini, Thu D. Nguyen
EuroSys3
2007 A Cost-Effective Distributed File Service with QoS Guarantees
Kien Le, Ricardo Bianchini, Thu D. Nguyen
Middleware3
2006 A: an assertion language for distributed systems
abstract
Operator mistakes have been identified as a significant source of unavailability in Internet services. In this paper, we propose a new language, A, for service engineers to write assertions about expected behaviors, proper configurations, and proper structural characteristics. This formalized specification of correct behavior can be used to bolster system understanding, as well as help to flag operator mistakes in a distributed system. Operator mistakes can be caused by anything from static misconfiguration to physical placement of wires and machines. This language, along with its associated runtime system, seeks to be flexible and robust enough to deal with the wide array of operator mistakes while maintaining a simple interface for designers or programmers.
Andrew Tjang, Fábio Oliveira, Richard P. Martin, Thu D. Nguyen
PLOS4
2006 Reducing the Availability Management Overheads of Federated Content Sharing Systems
abstract
We consider the problem of ensuring high data availability in federated content sharing systems. Ideally, such a system would provide high data availability in a device transparent manner so that users are not faced with the time-consuming and error-prone task of managing data replicas across the constituent devices of the system. We propose a novel unified availability model and a decentralized replication algorithm to approximate this ideal. Our availability model addresses three different concerns: availability during connected operation (online), availability during disconnected operation (offline), and availability after permanent disconnection from the federated system (ownership). Our replication algorithm centers around the intuition that devices should selfishly use their local storage to ensure offline and ownership availability for their individual owners. Excess storage, however, is used communally to ensure high online availability for all shared content. Evaluation of an implementation shows that our algorithm rapidly reaches stable and communally desirable configurations when there is sufficient space. Consistent with the fact that devices in a federated system are owned by different users, however, as space becomes highly constrained, the system approaches a non-cooperative configuration where devices only hoard content to serve their individual owners' needs
Christopher Peery, Thu D. Nguyen, Francisco Matias Cuenca-Acuna
SRDS2
2006 Understanding and Validating Database System Administration
Fábio Oliveira, Kiran Nagaraja, Rekha Bachwani, Ricardo Bianchini, Richard P. Martin, Thu D. Nguyen
USENIX ATC, General Track6
2005 Human-Aware Computer System Design
Ricardo Bianchini, Richard P. Martin, Kiran Nagaraja, Thu D. Nguyen, Fábio Oliveira
HotOS4
2005 Model-based validation for dealing with operator mistakes
abstract
Online services are rapidly becoming the supporting infrastructure for numerous users' work and leisure, placing higher demands on their availability and correct functioning. Increasingly, these services are comprised of complex conglomerates of distributed hardware and software components. Added to this complexity, these services evolve quite frequently accumulating considerable heterogeneity within them, while allowing little time for their in-depth understanding by service personnel. Thus, it is not surprising that mistakes by service operators are common, and have been deemed to be the primary cause of service downtime.
Kiran Nagaraja, Andrew Tjang, Fábio Oliveira, Ricardo Bianchini, Richard P. Martin, Thu D. Nguyen
SOSP6
2005 Enforcing Enterprise-wide Policies Over Standard Client-Server Interactions
abstract
We propose and evaluate a novel framework for enforcing global coordination and control policies over interacting software components in enterprise computing environments. This framework combines a per-node reference monitor with two existing coordination and control systems to enforce policies that, among other properties, are stateful and communal. Each reference monitor filters messages exchanged between the interacting software components similar to a firewall, passing only messages that are allowed by the policies in effect. This filtering approach decouples coordination and control from application implementation, allowing the coordination and control mechanism and application implementations to evolve independently of each other. We demonstrate the power of our framework by using it to specify and enforce an RBAC policy with delegation, revocation, and separation-of-duty over accesses to a cluster of NFS and SMB file servers without changing any client or server implementations. Measurements show that our framework imposes acceptable overheads when enforcing this policy.
Zhijun He 0002, Tuan Phan, Thu D. Nguyen
SRDS3
2005 Quantifying the Performability of Cluster-Based Services
abstract
In this paper, we propose a two-phase methodology for systematically evaluating the performability (performance and availability) of cluster-based Internet services. In the first phase, evaluators use a fault-injection infrastructure to characterize the service's behavior in the presence of faults. In the second phase, evaluators use an analytical model to combine an expected fault load with measurements from the first phase to assess the service's performability. Using this model, evaluators can study the service's sensitivity to different design decisions, fault rates, and other environmental factors. To demonstrate our methodology, we study the performability of a multitier Internet service. In particular, we evaluate the performance and availability of three soft state maintenance strategies for an online bookstore service in the presence of seven classes of faults. Among other interesting results, we clearly isolate the effect of different faults, showing that the tier of Web servers is responsible for an often dominant fraction of the service unavailability. Our results also demonstrate that storing the soft state in a database achieves better performability than storing it in main memory (even when the state is efficiently replicated) when we weight performance and availability equally. Based on our results, we conclude that service designers may want an unbalanced system in which they heavily load highly available components and leave more spare capacity for components that are likely to fail more often.
Kiran Nagaraja, Gustavo Machado Campagnani Gama, Ricardo Bianchini, Richard P. Martin, Wagner Meira Jr., Thu D. Nguyen
IEEE Trans. Parallel Distributed Syst.6
2004 Enforcement of Communal Policies for P2P Systems
Mihail F. Ionescu, Naftaly H. Minsky, Thu D. Nguyen
COORDINATION3
2004 Understanding and Dealing with Operator Mistakes in Internet Services
Kiran Nagaraja, Fábio Oliveira, Ricardo Bianchini, Richard P. Martin, Thu D. Nguyen
OSDI5
2004 Using adaptive range control to maximize 1-hop broadcast coverage in dense wireless networks
abstract
We present a distributed algorithm for maximizing 1-hop broadcast coverage in dense wireless sensor networks. Our strategy is built upon an analytic model that predicts the optimal range for maximizing 1-hop broadcast coverage given information like network density and node sending rate. The algorithm allows each node to set the maximizing radio range using only the locally observed sending rate and node density. The algorithm is thus critically dependent on the empirical determination of these parameters. Our algorithm can observe the parameters using only message eavesdropping and thus does not require extra protocol messages. Using simulation, we show that in spite of many simplifications in the model and incomplete density information in a live network, our algorithm converges fairly quickly and provides good coverage for both uniform and non-uniform networks across a wide range of conditions. We also demonstrate the utility of our algorithm for higher layer protocols by showing that it significantly improves the reception rate for a flooding application as well as the performance of a localization protocol.
Thu D. Nguyen, Richard P. Martin
SECON2
2004 Self-Managing Federated Services
abstract
We consider the problem of deploying and managing federated services that run on federated systems spanning multiple collaborative organizations. In particular, we present a peer-to-peer framework targeted to the construction of self-managing services that automatically adjust the number of service components and their placements in response to changes in the system or client loads. Our framework is completely decentralized, depending only on a modest amount of loosely synchronized global state. More specifically, our framework is comprised of a set of per-node monitoring agents and per-service-component management agents that periodically exchange information about the state of the system and of the service with each other using a gossiping protocol. Each management agent then periodically searches for configurations that are better than the current one according to an application model and explicit performance and availability targets. On finding a better configuration, an agent will enact the new configuration after a random delay to avoid possible collisions. We evaluate our framework by studying a prototype UDDI service. We show that while agents act autonomously, the service rapidly reaches a stable and appropriate configuration in response to system dynamics.
Francisco Matias Cuenca-Acuna, Thu D. Nguyen
SRDS2
2004 State Maintenance and its Impact on the Performability of Multi-tiered Internet Services
abstract
In this paper, we evaluate the performance, availability, and combined performability of four soft state maintenance strategies in two multitier Internet services, an online book store and an auction service. To take soft state and service latency into account, we propose an extension of our previous quantification methodology, and novel availability and performability metrics. Our results demonstrate that storing the soft state in a database can achieve better performability than storing it in main memory, even when the state is efficiently replicated. Strategies that offload the handling of soft state from the database increase the load on other tiers and, consequently, increase the impact of faults in these tiers on service availability. Based on these results, we conclude that service designers need to provision the cluster and balance the load with availability and cost, as well as performance, in mind.
Gustavo Machado Campagnani Gama, Kiran Nagaraja, Ricardo Bianchini, Richard P. Martin, Wagner Meira Jr., Thu D. Nguyen
SRDS6
2003 Compiler-Directed Program-Fault Coverage for Highly Available Java Internet Services
abstract
We present a new approach that uses compiler-directed fault-injection for coverage testing of recovery code in Internet services, to evaluate their robustness to operating system and I/O hardware faults. We define a set of program-fault coverage metrics that enable quantification of Java catch blocks exercised during fault-injection experiments. We use compiler analyses to instrument application code in two ways: to direct fault injection to occur at appropriate points during execution, and to measure the resulting coverage. As a proof of concept for these ideas, we have applied our techniques manually to Muffin, a proxy server; we obtained a high degree of coverage of catch blocks, with on average 85% of the expected faults per catch being experienced as caught exceptions.
Richard P. Martin, Kiran Nagaraja, Thu D. Nguyen, Barbara G. Ryder, David G. Wonnacott
DSN4
2003 Evaluating the Impact of Communication Architecture on the Performability of Cluster-Based Services
abstract
We consider the impact of different communication architectures on the performability (performance plus availability) of cluster-based servers. In particular, we use a combination of fault-injection experiments and analytic modeling to evaluate the performability of two popular communication protocols, TCP and VIA, as the intra-cluster communication substrate of a sophisticated Web server. Our analysis leads to several interesting conclusions, the most surprising of which is, under the same fault load, VIA-based servers deliver greater availability than TCP-based servers. If we assume higher fault rates for VIA-based servers because the underlying technology is more immature and programming model more complex, we find that packet errors or application faults would have to occur at approximately 4 times the rate in TCP-based servers before their performabilities equalize. We use our results from the study to suggest that high-performance and robust communication layers for highly available cluster-based servers should preserve message boundaries, as opposed to using byte streams, use single-copy transfers, pre-allocate channel resources, and report errors in manner consistent with the network fabric's fault model.
Kiran Nagaraja, Neeraj Krishnan, Ricardo Bianchini, Richard P. Martin, Thu D. Nguyen
HPCA5
2003 PlanetP: Using Gossiping to Build Content Addressable Peer-to-Peer Information Sharing Communities
abstract
We introduce PlanetP, content addressable publish/subscribe service for unstructured peer-to-peer (P2P) communities. PlanetP supports content addressing by providing: (1) a gossiping layer used to globally replicate a membership directory and an extremely compact content index; and (2) a completely distributed content search and ranking algorithm that help users find the most relevant information. PlanetP is a simple, yet powerful system for sharing information. PlanetP is simple because each peer must only perform a periodic, randomized, point-to-point message exchange with other peers. PlanetP is powerful because it maintains a globally content-ranked view of the shared data. Using simulation and a prototype implementation, we show that PlanetP achieves ranking accuracy that is comparable to a centralized solution and scales easily to several thousand peers while remaining resilient to rapid membership changes.
Francisco Matias Cuenca-Acuna, Christopher Peery, Richard P. Martin, Thu D. Nguyen
HPDC4
2003 Quantifying and Improving the Availability of High-Performance Cluster-Based Internet Services
abstract
Cluster-based servers can substantially increase performance when nodes cooperate to globally manage resources. However, in this paper we show that cooperation results in a substantial availability loss, in the absence of high-availability mechanisms. Specifically, we show that a sophisticated cluster-based Web server, which gains a factor of 3 in performance through cooperation, increases service unavailability by a factor of 10 over a non-cooperative version. We then show how to augment this Web server with software components embodying a small set of high-availability techniques to regain the lost availability. Among other interesting observations, we show that the application of multiple high-availability techniques, each implemented independently in its own subsystem, can lead to inconsistent recovery actions. We also show that a novel technique called Fault Model Enforcement can be used to resolve such inconsistencies. Augmenting the server with these techniques led to a final expected availability of close to 99.99%.
Kiran Nagaraja, Neeraj Krishnan, Ricardo Bianchini, Richard P. Martin, Thu D. Nguyen
SC5
2003 Using adaptive range control to optimize 1-hop broadcast coverage in dense wireless networks
abstract
No abstract available.
Thu D. Nguyen, Richard P. Martin
SenSys2
2003 Autonomous Replication for High Availability in Unstructured P2P Systems
abstract
We consider the problem of increasing the availability of shared data in peer-to-peer systems. In particular, we conservatively estimate the amount of excess storage required to achieve a practical availability of 99.9% by studying a decentralized algorithm that only depends on a modest amount of loosely synchronized global state. Our algorithm uses randomized decisions extensively together with a novel application of an erasure code to tolerate autonomous peer actions as well as staleness in the loosely synchronized global state. We study the behavior of this algorithm in three distinct environments modeled on previously reported measurements. We show that while peers act autonomously, the community as a whole will reach a stable configuration. We also show that space is used fairly and efficiently, delivering three times availability at a cost of six times the storage footprint of the data collection when the average peer availability is only 24%.
Francisco Matias Cuenca-Acuna, Richard P. Martin, Thu D. Nguyen
SRDS3
2002 Improving cluster availability using workstation validation
abstract
We demonstrate a framework for improving the availability of cluster based Internet services. Our approach models Internet services as a collection of interconnected components, each possessing well defined interfaces and failure semantics. Such a decomposition allows designers to engineer high availability based on an understanding of the interconnections and isolated fault behavior of each component, as opposed to ad-hoc methods. In this work, we focus on using the entire commodity workstation as a component because it possesses natural, fault-isolated interfaces. We define a failure event as a reboot because not only is a workstation unavailable during a reboot, but also because reboots are symptomatic of a larger class of failures, such as configuration and operator errors. Our observations of 3 distinct clusters show that the time between reboots is best modeled by a Weibull distribution with shape parameters of less than 1, implying that a workstation becomes more reliable the longer it has been operating. Leveraging this observed property, we design an allocation strategy which withholds recently rebooted workstations from active service, validating their stability before allowing them to return to service. We show via simulation that this policy leads to a 70-30 rule-of-thumb: For a constant utilization, approximately 70% of the workstation failures can be masked from end clients with 30% extra capacity added to the cluster, provided reboots are not strongly correlated. We also found our technique is most sensitive to the burstiness of reboots as opposed to absolute lengths of workstation uptimes.
Taliver Heath, Richard P. Martin, Thu D. Nguyen
SIGMETRICS3
2002 Lazy Garbage Collection of Recovery State for Fault-Tolerant Distributed Shared Memory
abstract
In this paper, we address the problem of garbage collection in a single-failure fault-tolerant home-based lazy release consistency (HLRC) distributed shared-memory (DSM) system based on independent checkpointing and logging. Our solution uses laziness in garbage collection and exploits consistency constraints of the HLRC memory model for low overhead and scalability. We prove safe bounds on the state that must be retained in the system to guarantee correct recovery after a failure. We devise two algorithms for garbage collection of checkpoints and logs, checkpoint garbage collection (CGC), and lazy log trimming (LLT). The proposed approach targets large-scale distributed shared-memory computing on local-area clusters of computers. In such systems, using global synchronization or extra communication for garbage collection is inefficient or simply impractical due to system scale and temporary disconnections in communication. The challenge lies in controlling the size of the logs and the number of checkpoints without global synchronization while tolerating transient disruptions in communication. Our garbage collection scheme is completely distributed, does not force processes to synchronize, does not add extra messages to the base DSM protocol, and uses only the available DSM protocol information. Evaluation results for real applications show that it effectively bounds the number of past checkpoints to be retained and the size of the logs in stable storage.
Florin Sultan, Thu D. Nguyen, Liviu Iftode
IEEE Trans. Parallel Distributed Syst.2
2002 Lazy Garbage Collection of Recovery State for Fault-Tolerant Distributed Shared Memory
abstract
We address the problem of garbage collection in a single-failure fault-tolerant home-based lazy release consistency (HLRC) distributed shared-memory (DSM) system based on independent checkpointing and logging. Our solution uses laziness in garbage collection and exploits consistency constraints of the HLRC memory model for low overhead and scalability. We prove safe bounds on the state that must be retained in the system to guarantee correct recovery after a failure. We devise two algorithms for garbage collection of checkpoints and logs, checkpoint garbage collection (CGC), and lazy log trimming (LLT). The proposed approach targets large-scale distributed shared-memory computing on local-area clusters of computers. The challenge lies in controlling the size of the logs and the number of checkpoints without global synchronization while tolerating transient disruptions in communication. Evaluation results for real applications show that it effectively bounds the number of past checkpoints to be retained and the size of the logs in stable storage.
Florin Sultan, Thu D. Nguyen, Liviu Iftode
IEEE Trans. Parallel Distributed Syst.2
2001 Quantifying the Impact of Architectural Scaling on Communication
abstract
This work quantifies how persistent increases in processor speed compared to I/O speed reduce the performance gap between specialized, high performance messaging layers and general purpose protocols such as TCP/IP and UDP/IP. The comparison is important because specialized layers sacrifice considerable system connectivity and robustness to obtain increased performance. We first quantify the scaling effects on small messages by measuring the LogP performance of two Active Message II layers, one running over a specialized VIA layer and the other over stock UDP as we scale the CPU and I/O components. We then predict future LogP performance by mapping the LogP model's network parameters, particularly overhead into architectural components. Our projections show that the performance benefit afforded by specialized messaging for small messages will erode to a factor of 2 in the next 5 years. Our models further show that the performance differential between the two approaches will continue to erode without a radical restructuring of the I/O system. For long messages, we quantify the variable per-page instruction budget that a zero-copy messaging approach has for page table manipulations if it is to outperform a single-copy approach. Finally we conclude with an examination of future I/O advances that would result in substantial improvements to messaging performance.
Taliver Heath, Samian Kaur, Richard P. Martin, Thu D. Nguyen
HPCA4
2001 Cooperative Caching Middleware for Cluster-Based Servers
abstract
Considers the use of cooperative caching to manage the memories of cluster-based servers. Over the last several years, a number of researchers have proposed content-aware servers that implement locality-conscious request distribution to address this memory management problem. During this development, it has become conventional wisdom that cooperative caching cannot match the performance of these servers. Unfortunately, while content-aware servers provide very high performance, their request distribution algorithms are typically bound to specific applications. The advantage of building distributed servers on top of a block-based cooperative caching layer is the generality of such a layer; it can be used as a building block for diverse services, ranging from file systems to web servers. In this paper, we reexamine the question of whether a server built on top of a generic block-based cooperative caching algorithm can perform competitively with content-aware servers. Specifically, we compare the performance of a cooperative caching-based Web server against L2S, a highly optimized locality- and load-conscious server. Our results show that, by modifying the replacement policy of traditional cooperative caching algorithms, we can achieve much of the performance provided by locality-conscious servers. Our modification increases network communication to reduce disk accesses, a reasonable trade-off considering the current trend of relative performance between LANs and disks.
Francisco Matias Cuenca-Acuna, Thu D. Nguyen
HPDC2
2001 DDDDRRaW: A Prototype Toolkit for Distributed Real-Time Rendering on Commodity Clusters
abstract
We describe DDDDRRaW, a prototype toolkit for distributed real-time rendering on commodity clusters. In constrast to most work on cluster computing, DDDDRRaW supports a repeated, low-latency computation, the drawing of frames which must take place on a time scale of 30-100 ms. DDDDRRaW employs image layer decomposition, a rendering-specific work partitioning algorithm described and evaluated using simulation. In this paper we address implementation issues. In particular, one important issue we explore is how to exploit the potential parallelism afforded by the multiple hardware resources of each node: the CPU, the network adapter and the video card. We evaluate DDDDRRaW's live performance on two small workstation clusters representing different points in the technology spectrum. Our results show that DDDDRRaW effectively exploits cluster resources to improve real-time rendering performance and should scale well to moderately sized clusters.
Thu D. Nguyen, Christopher Peery, John Zahorjan
IPDPS1
2000 Law-Governed Internet Communities
Xuhui Ao, Naftaly H. Minsky, Thu D. Nguyen, Victoria Ungureanu
COORDINATION3
2000 Image Layer Decomposition for Distributed Real-Time Rendering on Clusters
abstract
We propose a novel work partitioning technique, image layer decomposition (ILD), designed specifically to support distributed real-time rendering on commodity clusters. ILD has several advantages over previous partitioning algorithms for our targeted environment, including its compatibility with the use of hardware graphics accelerators, decoupling of communication bandwidth requirement from scene complexity, and reduced communication bandwidth growth as the system size increases. Furthermore, ILD tries to optimize the rendering of a sequence of frames (of an interactive application) instead of only individual frames. We simulate ILD using traces taken from a VRML viewer Our results show that ILD can be expected to work well up to moderately sized clusters and to outperform sort-last, a common partitioning approach, because of its smaller communication bandwidth requirement.
Thu D. Nguyen, John Zahorjan
IPDPS1
2000 Scalable Fault-Tolerant Distributed Shared Memory
abstract
This paper shows how a state-of-the-art software distributed shared-memory (DSM) protocol can be efficiently extended to tolerate single-node failures. In particular, we extend a home-based lazy release consistency (HLRC) DSM system with independent check- pointing and logging to volatile memory, targeting shared-memory computing on very large LAN-based clusters. In these environments, where global coordination may be expensive, independent checkpointing becomes critical to scalability. However, independent checkpointing is only practical if we can control the size of the log and checkpoints in the absence of global coordination. In this paper we describe the design of our fault-tolerant DSM system and present our solutions to the problems of checkpoint and log management. We also present experimental results showing that our fault tolerance support is light-weight, adding only low messaging, logging and checkpointing overheads, and that our management algorithms can be expected to effectively bound the size of the checkpoints and logs or real applications.
Florin Sultan, Thu D. Nguyen, Liviu Iftode
SC2
1998 Scheduling Policies to Support Distributed 3D Multimedia Applications
abstract
We consider the problem of scheduling the rendering component of 3D multimedia applications on a cluster of workstations connected via a local area network. Our goal is to meet a periodic real-time constraint.In abstract terms, the problem we address is how best to schedule tasks with unpredictable service times on distinct processing nodes so as to meet a real-time deadline, given that all communication among nodes entails some (possibly large) overhead. We consider two distinct classes of schemes, static, in which task reallocations are scheduled to occur at specific times, and dynamic, in which reallocations are triggered by some processor going idle. For both classes we further examine both global reassignments, in which all nodes are rescheduled at a rescheduling moment, and local reassignments, in which only a subset of the nodes engage in rescheduling at any one time.We show that global dynamic policies work best over a range of parameterizations appropriate to such systems. We introduce a new policy, Dynamic with Shadowing, that places a small number of tasks in the schedules of multiple workstations to reduce the amount of communication required to complete the schedule. This policy is shown to dominate the other alternatives considered over most of the parameter space.
Thu D. Nguyen, John Zahorjan
SIGMETRICS1
1996 Using Runtime Measured Workload Characteristics in Parallel Processor Scheduling
Thu D. Nguyen, Raj Vaswani, John Zahorjan
JSSPP1
1996 Parallel Application Characteristics for Multiprocessor Scheduling Policy Design
Thu D. Nguyen, Raj Vaswani, John Zahorjan
JSSPP1
1993 Implementing Network Protocols at User Level
abstract
Traditionally, network software has been structured in a monolithic fashion with all protocol stacks executing either within the kernel or in a single trusted user-level server. This organization is motivated by performance and security concerns. However, considerations of code maintenance, ease of debugging, customization, and the simultaneous existence of multiple protocols argue for separating the implementations into more manageable user-level libraries of protocols. This paper describes the design and implementation of transport protocols as user-level libraries.We begin by motivating the need for protocol implementations as user-level libraries and placing our approach in the context of previous work. We then describe our alternative to monolithic protocol organization, which has been implemented on Mach workstations connected not only to traditional Ethernet, but also to a more modern network, the DEC SRC ANI. Based on our experience, we discuss the implications for host-network interface design and for overall system structure to support efficient user-level implementations of network protocols.
Chandramohan A. Thekkath, Thu D. Nguyen, Evelyn Moy, Edward D. Lazowska
SIGCOMM2
1993 Implementing network protocols at user level
abstract
Traditionally, network software has been structured in a monolithic fashion with all protocol stacks executing either within the kernel or in a single trusted user-level server. This organization is motivated by performance and security concerns. However, considerations of code maintenance, ease of debugging, customization, and the simultaneous existence of multiple protocols argue for separating the implementations into more manageable user-level libraries of protocols. The present paper describes the design and implementation of transport protocols as user-level libraries. The authors begin by motivating the need for protocol implementations as user-level libraries and placing their approach in the context of previous work. They then describe their alternative to monolithic protocol organization, which has been implemented on Mach workstations connected not only to traditional Ethernet, but also to a more modern network, the DEC SRC AN1. Based on the authors' experience, they discuss the implications for host-network interface design and for overall system structure to support efficient user-level implementations of network protocols.>
Chandramohan A. Thekkath, Thu D. Nguyen, Evelyn Moy, Edward D. Lazowska
IEEE/ACM Trans. Netw.2