Sarah Bird

dblp:53/8277 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
0since 2021 · last 2020
0000-0002-5469-5149ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5Databases, data management, data science and information retrieval · 3 · 2 first-authorArtificial intelligence and machine learning · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 78% Efficient and distributed learning · 22%
Computer architecture, parallel and distributed computing, and storage systems
4 papers
Cloud and datacenter computing · 23% Memory systems · 23% Performance modeling and evaluation · 16%
Network and information security
1 paper
Privacy and data protection · 100%
Computer networks
1 paper
Network measurement and analytics · 100%
Software engineering, system software, and programming languages
2 papers
Operating systems · 100%
Databases, data mining, and information retrieval
2 papers
Machine learning and data management · 50% Data mining · 50%

Topics — the 17 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
0.822019
Fairness-Aware Machine Learning: Practical Challenges and Lessons Learned · WSDM 2019
Fairness-Aware Machine Learning: Practical Challenges and Lessons Learned · KDD 2019
Network measurement and analytics
web measurement
0.412020
The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing · WWW 2020
Privacy and data protection
online tracking
0.412020
The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing · WWW 2020
Machine learning › Trustworthy machine learning › fairness
algorithmic fairness
0.412019
Fairness-Aware Machine Learning: Practical Challenges and Lessons Learned · KDD 2019
Machine learning › Efficient and distributed learning
distributed training
0.312018
Applied Machine Learning at Facebook: A Datacenter Infrastructure Perspective · HPCA 2018
Cloud and datacenter computing
datacenter infrastructure
0.312018
Applied Machine Learning at Facebook: A Datacenter Infrastructure Perspective · HPCA 2018
Operating systems › resource management
resource containers
0.212013
Tessellation: refactoring the OS around explicit resource containers with continuous adaptation · DAC 2013
Operating systems
resource management
0.212013
Tessellation: refactoring the OS around explicit resource containers with continuous adaptation · DAC 2013
Memory systems
cache
0.212013
A hardware evaluation of cache partitioning to improve utilization and energy-efficiency while preserving responsiveness · ISCA 2013
Memory systems › cache management
cache partitioning
0.212013
A hardware evaluation of cache partitioning to improve utilization and energy-efficiency while preserving responsiveness · ISCA 2013
Privacy and data protection › web tracking
browser fingerprinting
0.112020
The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing · WWW 2020
Reconfigurable computing and FPGAs
FPGA-accelerated simulation
0.112010
A case for FAME: FPGA architecture model execution · ISCA 2010
Performance modeling and evaluation › simulation › processor simulation
multicore simulation
0.112010
A case for FAME: FPGA architecture model execution · ISCA 2010
Performance modeling and evaluation
simulation
0.112010
A case for FAME: FPGA architecture model execution · ISCA 2010
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.112018
Applied Machine Learning at Facebook: A Datacenter Infrastructure Perspective · HPCA 2018
Operating systems › resource management › process management
CPU scheduling
0.012013
A hardware evaluation of cache partitioning to improve utilization and energy-efficiency while preserving responsiveness · ISCA 2013
Processor architecture and microarchitecture
many-core architecture
0.012013
Tessellation: refactoring the OS around explicit resource containers with continuous adaptation · DAC 2013

Methods — techniques the papers use, named apart from their topics

fairness-aware machine learning · 1.5measurement study · 0.9comparative analysis · 0.9GPU training · 0.7CPU inference · 0.7hardware cache partitioning · 0.3dynamic partition sizing · 0.3
YearPublicationVenuePosition
2020 The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing
abstract
Large-scale Web crawls have emerged as the state of the art for studying characteristics of the Web. In particular, they are a core tool for online tracking research. Web crawling is an attractive approach to data collection, as crawls can be run at relatively low infrastructure cost and don’t require handling sensitive user data such as browsing histories. However, the biases introduced by using crawls as a proxy for human browsing data have not been well studied. Crawls may fail to capture the diversity of user environments, and the snapshot view of the Web presented by one-time crawls does not reflect its constantly evolving nature, which hinders reproducibility of crawl-based studies. In this paper, we quantify the repeatability and representativeness of Web crawls in terms of common tracking and fingerprinting metrics, considering both variation across crawls and divergence from human browser usage. We quantify baseline variation of simultaneous crawls, then isolate the effects of time, cloud IP address vs. residential, and operating system. This provides a foundation to assess the agreement between crawls visiting a standard list of high-traffic websites and actual browsing behaviour measured from an opt-in sample of over 50,000 users of the Firefox Web browser. Our analysis reveals differences between the treatment of stateless crawling infrastructure and generally stateful human browsing, showing, for example, that crawlers tend to experience higher rates of third-party activity than human browser users on loading pages from the same domains.
David Zeber, Sarah Bird, Camila Oliveira, Walter Rudametkin, Ilana Segall, Fredrik Wollsén, Martin Lopatka
WWW2
2019 Fairness-Aware Machine Learning: Practical Challenges and Lessons Learned
abstract
Researchers and practitioners from different disciplines have highlighted the ethical and legal challenges posed by the use of machine learned models and data-driven systems, and the potential for such systems to discriminate against certain population groups, due to biases in algorithmic decision-making systems. This tutorial aims to present an overview of algorithmic bias / discrimination issues observed over the last few years and the lessons learned, key regulations and laws, and evolution of techniques for achieving fairness in machine learning systems. We will motivate the need for adopting a "fairness-first" approach (as opposed to viewing algorithmic bias / fairness considerations as an afterthought), when developing machine learning based models and systems for different consumer and enterprise applications. Then, we will focus on the application of fairness-aware machine learning techniques in practice, by highlighting industry best practices and case studies from different technology companies. Based on our experiences in industry, we will identify open problems and research challenges for the data mining / machine learning community.
Sarah Bird, Ben Hutchinson, Krishnaram Kenthapadi, Emre Kiciman, Margaret Mitchell
KDD1
2019 Fairness-Aware Machine Learning: Practical Challenges and Lessons Learned
abstract
Researchers and practitioners from different disciplines have highlighted the ethical and legal challenges posed by the use of machine learned models and data-driven systems, and the potential for such systems to discriminate against certain population groups, due to biases in algorithmic decision-making systems. This tutorial aims to present an overview of algorithmic bias / discrimination issues observed over the last few years and the lessons learned, key regulations and laws, and evolution of techniques for achieving fairness in machine learning systems. We will motivate the need for adopting a "fairness-first" approach (as opposed to viewing algorithmic bias / fairness considerations as an afterthought), when developing machine learning based models and systems for different consumer and enterprise applications. Then, we will focus on the application of fairness-aware machine learning techniques in practice, by presenting case studies from different technology companies. Based on our experiences in industry, we will identify open problems and research challenges for the data mining / machine learning community.
Sarah Bird, Krishnaram Kenthapadi, Emre Kiciman, Margaret Mitchell
WSDM1
2018 Applied Machine Learning at Facebook: A Datacenter Infrastructure Perspective
abstract
Machine learning sits at the core of many essential products and services at Facebook. This paper describes the hardware and software infrastructure that supports machine learning at global scale. Facebook's machine learning workloads are extremely diverse: services require many different types of models in practice. This diversity has implications at all layers in the system stack. In addition, a sizable fraction of all data stored at Facebook flows through machine learning pipelines, presenting significant challenges in delivering data to high-performance distributed training flows. Computational requirements are also intense, leveraging both GPU and CPU platforms for training and abundant CPU capacity for real-time inference. Addressing these and other emerging challenges continues to require diverse efforts that span machine learning algorithms, software, and hardware design.
Kim M. Hazelwood, Sarah Bird, David Brooks 0001, Soumith Chintala, Utku Diril, Dmytro Dzhulgakov, Mohamed Fawzy, Bill Jia, Yangqing Jia, Aditya Kalro, James Law, Jason Lu, Pieter Noordhuis, Mikhail Smelyanskiy, Liang Xiong, Xiaodong Wang 0020
HPCA2
2013 Tessellation: refactoring the OS around explicit resource containers with continuous adaptation
abstract
Adaptive Resource-Centric Computing (ARCC) enables a simultaneous mix of high-throughput parallel, real-time, and interactive applications through automatic discovery of the correct mix of resource assignments necessary to achieve application requirements. This approach, embodied in the Tessellation manycore operating system, distributes resources to QoS domains called cells. Tessellation separates global decisions about the allocation of resources to cells from application-specific scheduling of resources within cells. We examine the implementation of ARCC in the Tessellation OS, highlight Tessellation's ability to provide predictable performance, and investigate the performance of Tessellation services within cells.
Juan A. Colmenares, Gage Eads, Steven Hofmeyr, Sarah Bird, Miquel Moretó, Brian Gluzman, Eric Roman, Davide B. Bartolini, Nitesh Mor, Krste Asanovic, John Kubiatowicz
DAC4
2013 A hardware evaluation of cache partitioning to improve utilization and energy-efficiency while preserving responsiveness
abstract
Computing workloads often contain a mix of interactive, latency-sensitive foreground applications and recurring background computations. To guarantee responsiveness, interactive and batch applications are often run on disjoint sets of resources, but this incurs additional energy, power, and capital costs. In this paper, we evaluate the potential of hardware cache partitioning mechanisms and policies to improve efficiency by allowing background applications to run simultaneously with interactive foreground applications, while avoiding degradation in interactive responsiveness. We evaluate these tradeoffs using commercial x86 multicore hardware that supports cache partitioning, and find that real hardware measurements with full applications provide different observations than past simulation-based evaluations. Co-scheduling applications without LLC partitioning leads to a 10% energy improvement and average throughput improvement of 54% compared to running tasks separately, but can result in foreground performance degradation of up to 34% with an average of 6%. With optimal static LLC partitioning, the average energy improvement increases to 12% and the average throughput improvement to 60%, while the worst case slowdown is reduced noticeably to 7% with an average slowdown of only 2%. We also evaluate a practical low-overhead dynamic algorithm to control partition sizes, and are able to realize the potential performance guarantees of the optimal static approach, while increasing background throughput by an additional 19%.
Henry Cook, Miquel Moretó, Sarah Bird, Khanh Dao, David A. Patterson 0001, Krste Asanovic
ISCA3
2011 Tessellation operating system: Building a real-time, responsive, high-throughput client OS for many-core architectures
Juan A. Colmenares, Sarah Bird, Gage Eads, Steven Hofmeyr, Albert Kim, Rohit Poddar, Hilfi Alkaff, Krste Asanovic, John Kubiatowicz
Hot Chips Symposium2
2010 A case for FAME: FPGA architecture model execution
abstract
Given the multicore microprocessor revolution, we argue that the architecture research community needs a dramatic increase in simulation capacity. We believe FPGA Architecture Model Execution (FAME) simulators can increase the number of useful architecture research experiments per day by two orders of magnitude over Software Architecture Model Execution (SAME) simulators. To clear up misconceptions about FPGA-based simulation methodologies, we propose a FAME taxonomy to distinguish the costperformance of variations on these ideas. We demonstrate our simulation speedup claim with a case study wherein we employ a prototype FAME simulator, RAMP Gold, to research the interaction between hardware partitioning mechanisms and operating system scheduling policy. The study demonstrates FAME's capabilities: we run a modern parallel benchmark suite on a research operating system, simulate 64-core target architectures with multi-level memory hierarchy timing models, and add experimental hardware mechanisms to the target machine. The simulation speedup achieved by our adoption of FAME-250×-enables experiments with more realistic time scales and data set sizes thanare possible with SAME.
Zhangxi Tan, Andrew Waterman, Henry Cook, Sarah Bird, Krste Asanovic, David A. Patterson 0001
ISCA4