VLDB 2026 Research / reviewers in the wild / expert
Yutian Sun
dblp:98/10492
· DBLP profile ↗
13ranked-venue papers
6as first author
4since 2021 · last 2024
0000-0002-0848-9029ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 4 first-authorDatabases, data management, data science and information retrieval · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Data Caching for Enterprise-Grade Petabyte-Scale OLAP
Chunxu Tang, Beinan Wang, Ziyue Qiu, Lu Qiu, Shouzhuo Sun, Saiguang Che, Jiaming Mai, Shouwei Chen, Jianjian Xie, Yutian Sun, Mingmin Chen |
USENIX ATC | 16 |
| 2023 | Shared Foundations: Modernizing Meta's Data Lakehouse
Biswapesh Chattopadhyay, Pedro Pedreira, Yutian Sun, Suketu Vakharia, Sundaram Narayanan |
CIDR | 4 |
| 2023 | Presto: A Decade of SQL Analytics at MetaabstractPresto is an open-source distributed SQL query engine that supports analytics workloads involving multiple exabyte-scale data sources. Presto is used for low-latency interactive use cases as well as long-running ETL jobs at Meta. It was originally launched at Meta in 2013 and donated to the Linux Foundation in 2019. Over the last ten years, upholding query latency and scalability with the hyper growth of data volume at Meta as well as new SQL analytics requirements have raised impressive challenges for Presto. A top priority has been ensuring query reliability does not regress with the shift towards smaller, more elastic container allocation, which requires queries to run with substantially smaller memory headroom and can be preempted at any time. Additionally, new demands from machine learning, privacy, and graph analytics have driven Presto maintainers to think beyond traditional data analytics. In this paper, we discuss several successful evolutions in recent years that have improved Presto latency as well as scalability by several orders of magnitude in production at Meta. Some of the notable ones are hierarchical caching, native vectorized execution engines, materialized views, and Presto on Spark. With these new capabilities, we have deprecated or are in the process of deprecating various legacy query engines so that Presto becomes the single piece to serve interactive, ad-hoc, ETL, and graph processing workloads for the entire data warehouse. Yutian Sun, Tim Meehan, Rebecca Schlussel, Wenlei Xie, Masha Basmanova, Orri Erling, Andrii Rosa, Shixuan Fan, Rongrong Zhong, Arun Thirupathi, Nikhil Collooru, Dionysios Logothetis, Kostas Xirogiannopoulos, Varun Gajjala, Rohit Jain, Ajay Palakuzhy, Prithvi Pandian, Sergey Pershin, Abhisek Saikia, Pranjal Shankhdhar, Neerad Somanchi, Swapnil Tailor, Jialiang Tan, Sreeni Viswanadha, Zac Wen, Biswapesh Chattopadhyay, Deepak Majeti, Aditi Pandit |
Proc. ACM Manag. Data | 1 |
| 2022 | From Batch Processing to Real Time Analytics: Running Presto® at ScaleabstractPresto is an open source distributed query engine used widely at Facebook, Uber, Twitter, Pinterest, and many other internet companies. Since open sourced in 2013, the Presto community has made several rounds of design and implementations, to support a variety of use cases, including interactive analytics, real time reporting and dashboard, ETL workloads, A/B testing, monitoring and alerts, etc. In this paper, we'd like to introduce some of the most important features and performance improvements the open source Presto community made in recent years, which enables companies running Presto at scale, supporting millions of queries per day, with hundreds of thousands of machines. Specifically, how Presto provides unified SQL on heterogeneous storage systems without data copy; how Presto deals with complex data, including nested columnar data and schema evolution; How Presto supports geospatial queries efficiently, and how file list cache works in Presto. We also talk about cluster federation, and Presto on cloud. Experimental results and our production experience could help others running interactive SQL systems at scale. Zhenxiao Luo, Lu Niu, Venki Korukanti, Yutian Sun, Masha Basmanova, Beinan Wang, Devesh Agrawal, Chunxu Tang, Girish Baliga, Maosong Fu |
ICDE | 4 |
| 2019 | Presto: SQL on EverythingabstractPresto is an open source distributed query engine that supports much of the SQL analytics workload at Facebook. Presto is designed to be adaptive, flexible, and extensible. It supports a wide variety of use cases with diverse characteristics. These range from user-facing reporting applications with sub-second latency requirements to multi-hour ETL jobs that aggregate or join terabytes of data. Presto's Connector API allows plugins to provide a high performance I/O interface to dozens of data sources, including Hadoop data warehouses, RDBMSs, NoSQL systems, and stream processing systems. In this paper, we outline a selection of use cases that Presto supports at Facebook. We then describe its architecture and implementation, and call out features and performance optimizations that enable it to support these use cases. Finally, we present performance results that demonstrate the impact of our main design decisions. Raghav Sethi, Martin Traverso, Dain Sundstrom, David Phillips, Wenlei Xie, Yutian Sun, Nezih Yegitbasi, Haozhun Jin, Eric Hwang, Nileema Shingte, Christopher Berner |
ICDE | 6 |
| 2014 | Separating Execution and Data Management: A Key to Business-Process-as-a-Service (BPaaS)
Yutian Sun, Jianwen Su, Jian Yang 0001 |
BPM | 1 |
| 2014 | Modeling data for business processesabstractAn important omission in current development practice for business process (or workflow) management systems is modeling of data & access for a business process, including relationship of the process data and the persistent data in the underlying enterprise database(s). This paper develops and studies a new approach to modeling data for business processes: representing data used by a process as a hierarchically structured business entity with (i) keys, local keys, and update constraints, and (ii) a set of data mapping rules defining exact correspondence between entity data values and values in the enterprise database. This paper makes the following technical contributions: (1) A data mapping language is formulated based on path expressions, and shown to coincide with a subclass of the schema mapping language Clio. (2) Two new notions are formulated: Updatability allows each update on a business entity (or database) to be translated to updates on the database (or resp. business entity), a fundamental requirement for process implementation. Isolation reflects that updates by one process execution do not alter data used by another running process. The property provides an important clue in process design. (3) Decision algorithms for updatability and isolation are presented, and they can be easily adapted for data mappings expressed in the subclass of Clio. Yutian Sun, Jianwen Su, Budan Wu, Jian Yang 0001 |
ICDE | 1 |
| 2014 | Conformance for DecSerFlow Constraints
Yutian Sun, Jianwen Su |
ICSOC | 1 |
| 2014 | Splitting GSM schemas: A framework for outsourcing of declarative artifact systems
Rik Eshuis, Richard Hull 0001, Yutian Sun, Roman Vaculín |
Inf. Syst. | 3 |
| 2013 | Splitting GSM Schemas: A Framework for Outsourcing of Declarative Artifact Systems
Rik Eshuis, Richard Hull 0001, Yutian Sun, Roman Vaculín |
BPM | 3 |
| 2013 | Barcelona: A Design and Runtime Environment for Declarative Artifact-Centric BPM
Terry Heath, David Boaz, Manmohan Gupta, Roman Vaculín, Yutian Sun, Richard Hull 0001, Lior Limonad |
ICSOC | 5 |
| 2012 | Declarative Choreographies for Artifacts
Yutian Sun, Jianwen Su |
ICSOC | 1 |
| 2011 | Computing Degree of Parallelism for BPMN Processes
Yutian Sun, Jianwen Su |
ICSOC | 1 |