Yee Jiun Song

dblp:93/2816 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 8 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2Computer networks · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
Distributed systems · 41% Cloud and datacenter computing · 29% Storage systems · 17%
Computer networks
1 paper
Content delivery and video streaming · 54% Network optimization and economics · 23% Routing and switching · 23%

Topics — the 27 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
datacenter storage
0.522021
Log-structured Protocols in Delos · SOSP 2021
TAO: Facebook's Distributed Data Store for the Social Graph · USENIX ATC 2013
Storage systems › distributed storage
shared log
0.512021
Log-structured Protocols in Delos · SOSP 2021
Distributed systems › replication
state machine replication
0.512021
Log-structured Protocols in Delos · SOSP 2021
Distributed systems
consensus
0.412020
Virtual Consensus in Delos · OSDI 2020
Cloud and datacenter computing › datacenter operations
datacenter reliability
0.312018
Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently · OSDI 2018
Distributed systems
fault tolerance
0.312018
Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently · OSDI 2018
Distributed systems
replication
0.322015
Existential consistency: measuring and understanding consistency at Facebook · SOSP 2015
Wormhole: Reliable Pub-Sub to Support Geo-replicated Internet Services · NSDI 2015
Cloud and datacenter computing
cluster resource management and scheduling
0.212016
Kraken: Leveraging Live Traffic Tests to Identify and Resolve Resource Utilization Bottlenecks in Large Scale Web Services · OSDI 2016
Cloud and datacenter computing
datacenter operations
0.212016
Kraken: Leveraging Live Traffic Tests to Identify and Resolve Resource Utilization Bottlenecks in Large Scale Web Services · OSDI 2016
Energy-efficient computing
datacenter power management
0.212016
Dynamo: Facebook's Data Center-Wide Power Management System · ISCA 2016
Energy-efficient computing › power management
power capping
0.212016
Dynamo: Facebook's Data Center-Wide Power Management System · ISCA 2016
Distributed systems
consistency models
0.212015
Existential consistency: measuring and understanding consistency at Facebook · SOSP 2015
Distributed systems › replication
geo-replication
0.212015
Wormhole: Reliable Pub-Sub to Support Geo-replicated Internet Services · NSDI 2015
Distributed systems
publish/subscribe systems
0.212015
Wormhole: Reliable Pub-Sub to Support Geo-replicated Internet Services · NSDI 2015
Storage systems › data redundancy
replicated storage
0.212015
Existential consistency: measuring and understanding consistency at Facebook · SOSP 2015
Storage systems
distributed storage
0.212013
TAO: Facebook's Distributed Data Store for the Social Graph · USENIX ATC 2013
Storage systems
key-value storage
0.112021
Log-structured Protocols in Delos · SOSP 2021
Distributed systems › operating system support › interprocess communication
client-server communication
0.112009
RPC Chains: Efficient Client-Server Communication in Geodistributed Systems · NSDI 2009
Distributed systems › distributed system architecture
geo-distributed systems
0.112009
RPC Chains: Efficient Client-Server Communication in Geodistributed Systems · NSDI 2009
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.112016
Dynamo: Facebook's Data Center-Wide Power Management System · ISCA 2016
Distributed systems › fault tolerance
reliable communication
0.112015
Wormhole: Reliable Pub-Sub to Support Geo-replicated Internet Services · NSDI 2015
Routing and switching › traffic engineering
bandwidth cost optimization
0.112005
CobWeb: a proactive analysis-driven approach to content distribution · SOSP 2005
Content delivery and video streaming
content delivery network
0.112005
CobWeb: a proactive analysis-driven approach to content distribution · SOSP 2005
Content delivery and video streaming
content placement
0.112005
CobWeb: a proactive analysis-driven approach to content distribution · SOSP 2005
Network optimization and economics
resource allocation
0.112005
CobWeb: a proactive analysis-driven approach to content distribution · SOSP 2005
Distributed systems
remote procedure call
0.012009
RPC Chains: Efficient Client-Server Communication in Geodistributed Systems · NSDI 2009
Content delivery and video streaming › caching
cache management
0.012005
CobWeb: a proactive analysis-driven approach to content distribution · SOSP 2005

Methods — techniques the papers use, named apart from their topics

state machine replication · 0.5leasing · 0.5batching · 0.5power monitoring · 0.2coordinated control · 0.2measurement study · 0.2consistency analysis · 0.2optimization · 0.1analytical modeling · 0.1
YearPublicationVenuePosition
2021 Log-structured Protocols in Delos
abstract
Developers have access to a wide range of storage APIs and functionality in large-scale systems, such as relational databases, key-value stores, and namespaces. However, this diversity comes at a cost: each API is implemented by a complex distributed system that is difficult to develop and operate. Delos amortizes this cost by enabling different APIs on a shared codebase and operational platform. The primary innovation in Delos is a log-structured protocol: a fine-grained replicated state machine executing above a shared log that can be layered into reusable protocol stacks under different databases. We built and deployed two production databases using Delos at Facebook, creating nine different log-structured protocols in the process. We show via experiments and production data that log-structured protocols impose low overhead, while allowing optimizations that can improve latency by up to 100X (e.g., via leasing) and throughput by up to 2X (e.g., via batching).
Mahesh Balakrishnan 0001, Ahmed Jafri, Suyog Mapara, David Geraghty, Jason Flinn, Vidhya Venkat, Ivailo Nedelchev, Santosh Ghosh, Mihir Dharamshi, Jingming Liu, Filip Gruszczynski, Rounak Tibrewal, Ali Zaveri, Rajeev Nagar, Ahmed Yossef, Francois Richard, Yee Jiun Song
SOSP19
2020 Virtual Consensus in Delos
Mahesh Balakrishnan 0001, Jason Flinn, Mihir Dharamshi, Ahmed Jafri, Santosh Ghosh, Hazem Hassan, Aaryaman Sagar, Rhed Shi, Jingming Liu, Filip Gruszczynski, Xianan Zhang, Huy Hoang, Ahmed Yossef, Francois Richard, Yee Jiun Song
OSDI17
2018 Maelstrom: Mitigating Datacenter-level Disasters by Draining Interdependent Traffic Safely and Efficiently
Kaushik Veeraraghavan, Justin Meza, Scott Michelson, Sankaralingam Panneerselvam, Alex Gyori, Sonia Margulis, Daniel Obenshain, Shruti Padmanabha, Ashish Shah, Yee Jiun Song, Tianyin Xu
OSDI11
2017 Canopy: An End-to-End Performance Tracing And Analysis System
abstract
This paper presents Canopy, Facebook's end-to-end performance tracing infrastructure. Canopy records causally related performance data across the end-to-end execution path of requests, including from browsers, mobile applications, and backend services. Canopy processes traces in near real-time, derives user-specified features, and outputs to performance datasets that aggregate across billions of requests. Using Canopy, Facebook engineers can query and analyze performance data in real-time. Canopy addresses three challenges we have encountered in scaling performance analysis: supporting the range of execution and performance models used by different components of the Facebook stack; supporting interactive ad-hoc analysis of performance data; and enabling deep customization by users, from sampling traces to extracting and visualizing features. Canopy currently records and processes over 1 billion traces per day. We discuss how Canopy has evolved to apply to a wide range of scenarios, and present case studies of its use in solving various performance challenges.
Jonathan Kaldor, Jonathan Mace, Michal Bejda, Edison Gao, Wiktor Kuropatwa, Joe O'Neill, Kian Win Ong, Bill Schaller, Pingjia Shan, Brendan Viscomi, Vinod Venkataraman, Kaushik Veeraraghavan, Yee Jiun Song
SOSP13
2016 Dynamo: Facebook's Data Center-Wide Power Management System
abstract
Data center power is a scarce resource that often goes underutilized due to conservative planning. This is because the penalty for overloading the data center power delivery hierarchy and tripping a circuit breaker is very high, potentially causing long service outages. Recently, dynamic server power capping, which limits the amount of power consumed by a server, has been proposed and studied as a way to reduce this penalty, enabling more aggressive utilization of provisioned data center power. However, no real at-scale solution for data center-wide power monitoring and control has been presented in the literature. In this paper, we describe Dynamo -- a data center-wide power management system that monitors the entire power hierarchy and makes coordinated control decisions to safely and efficiently use provisioned data center power. Dynamo has been developed and deployed across all of Facebook's data centers for the past three years. Our key insight is that in real-world data centers, different power and performance constraints at different levels in the power hierarchy necessitate coordinated data center-wide power management. We make three main contributions. First, to understand the design space of Dynamo, we provide a characterization of power variation in data centers running a diverse set of modern workloads. This characterization uses fine-grained power samples from tens of thousands of servers and spanning a period of over six months. Second, we present the detailed design of Dynamo. Our design addresses several key issues not addressed by previous simulation-based studies. Third, the proposed techniques and design have been deployed and evaluated in large scale data centers serving billions of users. We present production results showing that Dynamo has prevented 18 potential power outages in the past 6 months due to unexpected power surges, that Dynamo enables optimizations leading to a 13% performance boost for a production Hadoop cluster and a nearly 40% performance increase for a search cluster, and that Dynamo has already enabled an 8% increase in the power capacity utilization of one of our data centers with more aggressive power subscription measures underway.
Qingyuan Deng, Lakshmi Ganesh, Chang-Hong Hsu, Yun Jin, Justin Meza, Yee Jiun Song
ISCA9
2016 Kraken: Leveraging Live Traffic Tests to Identify and Resolve Resource Utilization Bottlenecks in Large Scale Web Services
Kaushik Veeraraghavan, Justin Meza, Wonho Kim, Sonia Margulis, Scott Michelson, Rajesh Nishtala, Daniel Obenshain, Dmitri Perelman, Yee Jiun Song
OSDI10
2015 Wormhole: Reliable Pub-Sub to Support Geo-replicated Internet Services
Yogeshwer Sharma, Philippe Ajoux, Petchean Ang, David Callies, Abhishek Choudhary, Laurent Demailly, Thomas Fersch, Liat Atsmon Guz, Andrzej Kotulski, Sachin Kulkarni, Harry C. Li, Evgeniy Makeev, Kowshik Prakasam, Robbert van Renesse, Sabyasachi Roy, Pratyush Seth, Yee Jiun Song, Benjamin Wester, Kaushik Veeraraghavan, Peter Xie
NSDI19
2015 Existential consistency: measuring and understanding consistency at Facebook
abstract
Replicated storage for large Web services faces a trade-off between stronger forms of consistency and higher performance properties. Stronger consistency prevents anomalies, i.e., unexpected behavior visible to users, and reduces programming complexity. There is much recent work on improving the performance properties of systems with stronger consistency, yet the flip-side of this trade-off remains elusively hard to quantify. To the best of our knowledge, no prior work does so for a large, production Web service.
Haonan Lu, Kaushik Veeraraghavan, Philippe Ajoux, Jim Hunt, Yee Jiun Song, Wendy Tobagus, Wyatt Lloyd
SOSP5
2013 TAO: Facebook's Distributed Data Store for the Social Graph
Nathan Bronson, Zach Amsden, George Cabrera, Prasad Chakka, Peter Dimov, Jack Ferris, Anthony Giardullo, Sachin Kulkarni, Harry C. Li, Mark Marchukov, Dmitri Petrov, Lovro Puzar, Yee Jiun Song, Venkateshwaran Venkataramani
USENIX ATC14
2009 RPC Chains: Efficient Client-Server Communication in Geodistributed Systems
Yee Jiun Song, Marcos K. Aguilera, Ramakrishna Kotla, Dahlia Malkhi
NSDI1
2008 Bosco: One-Step Byzantine Asynchronous Consensus
Yee Jiun Song, Robbert van Renesse
DISC1
2005 CobWeb: a proactive analysis-driven approach to content distribution
abstract
CobWeb is an open-access content distribution network (CDN) that provides low latency lookups, resilience to flash crowds, and optimal utilization of network resources. Unlike traditional Web caches and CDNs, which rely on ad hoc heuristics for replica placement and cache management, CobWeb achieves superior performance through a unique analysis driven approach. CobWeb derives the optimal replica placement strategy by posing the fundamental performance-overhead tradeoff as a resource constraint problem. Analytically modeling the costs and performance benefits of replicas enables CobWeb to convert the systems problem to an optimization problem. The optimization problem can then be solved to provide, for instance, minimal lookup latency while achieving a targeted bandwidth cost, or to achieve a targeted lookup performance while minimizing bandwidth consumption. CobWeb is currently deployed on Planet-Lab and is available for open access through an intuitive easy-to-use interface.
Yee Jiun Song, Venugopalan Ramasubramanian, Emin Gün Sirer
SOSP1