VLDB 2026 Research / reviewers in the wild / expert
Sarah Bird
dblp:53/8277
· DBLP profile ↗
8ranked-venue papers
2as first author
0since 2021 · last 2020
0000-0002-5469-5149ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5Databases, data management, data science and information retrieval · 3 · 2 first-authorArtificial intelligence and machine learning · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Trustworthy machine learning · 78% Efficient and distributed learning · 22% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Cloud and datacenter computing · 23% Memory systems · 23% Performance modeling and evaluation · 16% | |
| Network and information security
1 paper |
Privacy and data protection · 100% | |
| Computer networks
1 paper |
Network measurement and analytics · 100% | |
| Software engineering, system software, and programming languages
2 papers |
Operating systems · 100% | |
| Databases, data mining, and information retrieval
2 papers |
Machine learning and data management · 50% Data mining · 50% |
Topics — the 17 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
0.8 | 2 | 2019 | Fairness-Aware Machine Learning: Practical Challenges and Lessons Learned · WSDM 2019 Fairness-Aware Machine Learning: Practical Challenges and Lessons Learned · KDD 2019 |
Network measurement and analytics
web measurement |
0.4 | 1 | 2020 | The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing · WWW 2020 |
Privacy and data protection
online tracking |
0.4 | 1 | 2020 | The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing · WWW 2020 |
Machine learning › Trustworthy machine learning › fairness
algorithmic fairness |
0.4 | 1 | 2019 | Fairness-Aware Machine Learning: Practical Challenges and Lessons Learned · KDD 2019 |
Machine learning › Efficient and distributed learning
distributed training |
0.3 | 1 | 2018 | Applied Machine Learning at Facebook: A Datacenter Infrastructure Perspective · HPCA 2018 |
Cloud and datacenter computing
datacenter infrastructure |
0.3 | 1 | 2018 | Applied Machine Learning at Facebook: A Datacenter Infrastructure Perspective · HPCA 2018 |
Operating systems › resource management
resource containers |
0.2 | 1 | 2013 | Tessellation: refactoring the OS around explicit resource containers with continuous adaptation · DAC 2013 |
Operating systems
resource management |
0.2 | 1 | 2013 | Tessellation: refactoring the OS around explicit resource containers with continuous adaptation · DAC 2013 |
Memory systems
cache |
0.2 | 1 | 2013 | A hardware evaluation of cache partitioning to improve utilization and energy-efficiency while preserving responsiveness · ISCA 2013 |
Memory systems › cache management
cache partitioning |
0.2 | 1 | 2013 | A hardware evaluation of cache partitioning to improve utilization and energy-efficiency while preserving responsiveness · ISCA 2013 |
Privacy and data protection › web tracking
browser fingerprinting |
0.1 | 1 | 2020 | The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing · WWW 2020 |
Reconfigurable computing and FPGAs
FPGA-accelerated simulation |
0.1 | 1 | 2010 | A case for FAME: FPGA architecture model execution · ISCA 2010 |
Performance modeling and evaluation › simulation › processor simulation
multicore simulation |
0.1 | 1 | 2010 | A case for FAME: FPGA architecture model execution · ISCA 2010 |
Performance modeling and evaluation
simulation |
0.1 | 1 | 2010 | A case for FAME: FPGA architecture model execution · ISCA 2010 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.1 | 1 | 2018 | Applied Machine Learning at Facebook: A Datacenter Infrastructure Perspective · HPCA 2018 |
Operating systems › resource management › process management
CPU scheduling |
0.0 | 1 | 2013 | A hardware evaluation of cache partitioning to improve utilization and energy-efficiency while preserving responsiveness · ISCA 2013 |
Processor architecture and microarchitecture
many-core architecture |
0.0 | 1 | 2013 | Tessellation: refactoring the OS around explicit resource containers with continuous adaptation · DAC 2013 |
Methods — techniques the papers use, named apart from their topics
fairness-aware machine learning · 1.5measurement study · 0.9comparative analysis · 0.9GPU training · 0.7CPU inference · 0.7hardware cache partitioning · 0.3dynamic partition sizing · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | The Representativeness of Automated Web Crawls as a Surrogate for Human BrowsingabstractLarge-scale Web crawls have emerged as the state of the art for studying characteristics of the Web. In particular, they are a core tool for online tracking research. Web crawling is an attractive approach to data collection, as crawls can be run at relatively low infrastructure cost and don’t require handling sensitive user data such as browsing histories. However, the biases introduced by using crawls as a proxy for human browsing data have not been well studied. Crawls may fail to capture the diversity of user environments, and the snapshot view of the Web presented by one-time crawls does not reflect its constantly evolving nature, which hinders reproducibility of crawl-based studies. In this paper, we quantify the repeatability and representativeness of Web crawls in terms of common tracking and fingerprinting metrics, considering both variation across crawls and divergence from human browser usage. We quantify baseline variation of simultaneous crawls, then isolate the effects of time, cloud IP address vs. residential, and operating system. This provides a foundation to assess the agreement between crawls visiting a standard list of high-traffic websites and actual browsing behaviour measured from an opt-in sample of over 50,000 users of the Firefox Web browser. Our analysis reveals differences between the treatment of stateless crawling infrastructure and generally stateful human browsing, showing, for example, that crawlers tend to experience higher rates of third-party activity than human browser users on loading pages from the same domains. David Zeber, Sarah Bird, Camila Oliveira, Walter Rudametkin, Ilana Segall, Fredrik Wollsén, Martin Lopatka |
WWW | 2 |
| 2019 | Fairness-Aware Machine Learning: Practical Challenges and Lessons LearnedabstractResearchers and practitioners from different disciplines have highlighted the ethical and legal challenges posed by the use of machine learned models and data-driven systems, and the potential for such systems to discriminate against certain population groups, due to biases in algorithmic decision-making systems. This tutorial aims to present an overview of algorithmic bias / discrimination issues observed over the last few years and the lessons learned, key regulations and laws, and evolution of techniques for achieving fairness in machine learning systems. We will motivate the need for adopting a "fairness-first" approach (as opposed to viewing algorithmic bias / fairness considerations as an afterthought), when developing machine learning based models and systems for different consumer and enterprise applications. Then, we will focus on the application of fairness-aware machine learning techniques in practice, by highlighting industry best practices and case studies from different technology companies. Based on our experiences in industry, we will identify open problems and research challenges for the data mining / machine learning community. Sarah Bird, Ben Hutchinson, Krishnaram Kenthapadi, Emre Kiciman, Margaret Mitchell |
KDD | 1 |
| 2019 | Fairness-Aware Machine Learning: Practical Challenges and Lessons LearnedabstractResearchers and practitioners from different disciplines have highlighted the ethical and legal challenges posed by the use of machine learned models and data-driven systems, and the potential for such systems to discriminate against certain population groups, due to biases in algorithmic decision-making systems. This tutorial aims to present an overview of algorithmic bias / discrimination issues observed over the last few years and the lessons learned, key regulations and laws, and evolution of techniques for achieving fairness in machine learning systems. We will motivate the need for adopting a "fairness-first" approach (as opposed to viewing algorithmic bias / fairness considerations as an afterthought), when developing machine learning based models and systems for different consumer and enterprise applications. Then, we will focus on the application of fairness-aware machine learning techniques in practice, by presenting case studies from different technology companies. Based on our experiences in industry, we will identify open problems and research challenges for the data mining / machine learning community. Sarah Bird, Krishnaram Kenthapadi, Emre Kiciman, Margaret Mitchell |
WSDM | 1 |
| 2018 | Applied Machine Learning at Facebook: A Datacenter Infrastructure PerspectiveabstractMachine learning sits at the core of many essential products and services at Facebook. This paper describes the hardware and software infrastructure that supports machine learning at global scale. Facebook's machine learning workloads are extremely diverse: services require many different types of models in practice. This diversity has implications at all layers in the system stack. In addition, a sizable fraction of all data stored at Facebook flows through machine learning pipelines, presenting significant challenges in delivering data to high-performance distributed training flows. Computational requirements are also intense, leveraging both GPU and CPU platforms for training and abundant CPU capacity for real-time inference. Addressing these and other emerging challenges continues to require diverse efforts that span machine learning algorithms, software, and hardware design. Kim M. Hazelwood, Sarah Bird, David Brooks 0001, Soumith Chintala, Utku Diril, Dmytro Dzhulgakov, Mohamed Fawzy, Bill Jia, Yangqing Jia, Aditya Kalro, James Law, Jason Lu, Pieter Noordhuis, Mikhail Smelyanskiy, Liang Xiong, Xiaodong Wang 0020 |
HPCA | 2 |
| 2013 | Tessellation: refactoring the OS around explicit resource containers with continuous adaptationabstractAdaptive Resource-Centric Computing (ARCC) enables a simultaneous mix of high-throughput parallel, real-time, and interactive applications through automatic discovery of the correct mix of resource assignments necessary to achieve application requirements. This approach, embodied in the Tessellation manycore operating system, distributes resources to QoS domains called cells. Tessellation separates global decisions about the allocation of resources to cells from application-specific scheduling of resources within cells. We examine the implementation of ARCC in the Tessellation OS, highlight Tessellation's ability to provide predictable performance, and investigate the performance of Tessellation services within cells. Juan A. Colmenares, Gage Eads, Steven Hofmeyr, Sarah Bird, Miquel Moretó, Brian Gluzman, Eric Roman, Davide B. Bartolini, Nitesh Mor, Krste Asanovic, John Kubiatowicz |
DAC | 4 |
| 2013 | A hardware evaluation of cache partitioning to improve utilization and energy-efficiency while preserving responsivenessabstractComputing workloads often contain a mix of interactive, latency-sensitive foreground applications and recurring background computations. To guarantee responsiveness, interactive and batch applications are often run on disjoint sets of resources, but this incurs additional energy, power, and capital costs. In this paper, we evaluate the potential of hardware cache partitioning mechanisms and policies to improve efficiency by allowing background applications to run simultaneously with interactive foreground applications, while avoiding degradation in interactive responsiveness. We evaluate these tradeoffs using commercial x86 multicore hardware that supports cache partitioning, and find that real hardware measurements with full applications provide different observations than past simulation-based evaluations. Co-scheduling applications without LLC partitioning leads to a 10% energy improvement and average throughput improvement of 54% compared to running tasks separately, but can result in foreground performance degradation of up to 34% with an average of 6%. With optimal static LLC partitioning, the average energy improvement increases to 12% and the average throughput improvement to 60%, while the worst case slowdown is reduced noticeably to 7% with an average slowdown of only 2%. We also evaluate a practical low-overhead dynamic algorithm to control partition sizes, and are able to realize the potential performance guarantees of the optimal static approach, while increasing background throughput by an additional 19%. Henry Cook, Miquel Moretó, Sarah Bird, Khanh Dao, David A. Patterson 0001, Krste Asanovic |
ISCA | 3 |
| 2011 | Tessellation operating system: Building a real-time, responsive, high-throughput client OS for many-core architectures
Juan A. Colmenares, Sarah Bird, Gage Eads, Steven Hofmeyr, Albert Kim, Rohit Poddar, Hilfi Alkaff, Krste Asanovic, John Kubiatowicz |
Hot Chips Symposium | 2 |
| 2010 | A case for FAME: FPGA architecture model executionabstractGiven the multicore microprocessor revolution, we argue that the architecture research community needs a dramatic increase in simulation capacity. We believe FPGA Architecture Model Execution (FAME) simulators can increase the number of useful architecture research experiments per day by two orders of magnitude over Software Architecture Model Execution (SAME) simulators. To clear up misconceptions about FPGA-based simulation methodologies, we propose a FAME taxonomy to distinguish the costperformance of variations on these ideas. We demonstrate our simulation speedup claim with a case study wherein we employ a prototype FAME simulator, RAMP Gold, to research the interaction between hardware partitioning mechanisms and operating system scheduling policy. The study demonstrates FAME's capabilities: we run a modern parallel benchmark suite on a research operating system, simulate 64-core target architectures with multi-level memory hierarchy timing models, and add experimental hardware mechanisms to the target machine. The simulation speedup achieved by our adoption of FAME-250×-enables experiments with more realistic time scales and data set sizes thanare possible with SAME. Zhangxi Tan, Andrew Waterman, Henry Cook, Sarah Bird, Krste Asanovic, David A. Patterson 0001 |
ISCA | 4 |