Tao He 0013

dblp:94/5035-13 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
0009-0007-7687-7342ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Distributed systems · 44% Storage systems · 25% High-performance computing · 19%
Databases, data mining, and information retrieval
2 papers
Distributed and cloud data management · 64% Graph data management · 36%
Artificial intelligence
1 paper
Graph learning · 100%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed systems › resource sharing
data sharing
0.712023
Vineyard: Optimizing Data Sharing in Data-Intensive Analytics · Proc. ACM Manag. Data 2023
Storage systems
object storage
0.712023
Vineyard: Optimizing Data Sharing in Data-Intensive Analytics · Proc. ACM Manag. Data 2023
Machine learning › Graph learning
graph neural network
0.512021
GraphScope: A Unified Engine For Big Graph Processing · Proc. VLDB Endow. 2021
Distributed systems › distributed graph processing
distributed graph processing engine
0.512021
GraphScope: A Unified Engine For Big Graph Processing · Proc. VLDB Endow. 2021
High-performance computing
large-scale graph processing
0.512021
GraphScope: A Unified Engine For Big Graph Processing · Proc. VLDB Endow. 2021
Compilers and program optimization
intermediate representation
0.312025
Moko: Marrying Python with Big Data Systems · EuroSys 2025
Parallel and multicore computing › parallel algorithms › parallel primitives
data-parallel primitives
0.112021
GraphScope: A Unified Engine For Big Graph Processing · Proc. VLDB Endow. 2021
Parallel and multicore computing
parallel programming models
0.112021
GraphScope: A Unified Engine For Big Graph Processing · Proc. VLDB Endow. 2021

Methods — techniques the papers use, named apart from their topics

IR-based code generation · 1.7graph engine optimization · 1.5declarative data-parallel operators · 1.5method sharing · 0.7memory mapping · 0.7VCDL · 0.7
YearPublicationVenuePosition
2025 Moko: Marrying Python with Big Data Systems
abstract
Python stands as the preferred language for data science, thanks to its user-friendly syntax and a robust ecosystem that effortlessly accommodates a variety of data types and workloads, such as relational/tabular data, tensors, and graphs. While Python thrives in smaller data settings, it struggles to scale in distributed big data environments. MOKO is an IR-based execution framework designed to extend Python's reach into the distributed big data domain by generating code that can utilize existing systems such as Spark, Dask, Torch, and GRAPE. Moko preserves Python's key features---interoperability, ease of use, and support for multi-model data types and workloads---while enabling efficient execution in a distributed setting. Our evaluation indicates that MOKO can accelerate Python applications by up to 11× across diverse systems, diminish data alignment overhead by 28×, and outperform hand-optimized solutions by 2.5×.
Tao He 0013, Sijie Shen, Lei Wang 0004, Wenyuan Yu, Jingren Zhou 0001
EuroSys2
2023 Vineyard: Optimizing Data Sharing in Data-Intensive Analytics
abstract
Modern data analytics and AI jobs become increasingly complex and involve multiple tasks performed on specialized systems. Sharing of intermediate data between different systems is often a significant bottleneck in such jobs. When the intermediate data is large, it is mostly exchanged through files in standard formats (e.g., CSV and ORC), causing high I/O and (de)serialization overheads. To solve these problems, we develop Vineyard, a high-performance, extensible, and cloud-native object store, trying to provide an intuitive experience for users to share data across systems in complex real-life workflows. Since different systems usually work on data structures (e.g., dataframes, graphs, hashmaps) with similar interfaces, and their computation logic is often loosely-coupled with how such interfaces are implemented over specific memory layouts, it enables Vineyard to conduct data sharing efficiently at a high level via memory mapping and method sharing. Vineyard provides an IDL named VCDL to facilitate users to register their own intermediate data types into Vineyard such that objects of the registered types can then be efficiently shared across systems in a polyglot workflow. As a cloud-native system, Vineyard is designed to work closely with Kubernetes, as well as achieve fault-tolerance and high performance in production environments. Evaluations on real-life datasets and data analytics jobs show that the above optimizations of Vineyard can significantly improve the end-to-end performance of data analytics jobs, by reducing their data-sharing time up to 68.4x.
Wenyuan Yu, Tao He 0013, Lei Wang 0004, Ye Cao 0004, Diwen Zhu, Sanhong Li, Jingren Zhou 0001
Proc. ACM Manag. Data2
2021 GraphScope: A Unified Engine For Big Graph Processing
abstract
GraphScope is a system and a set of language extensions that enable a new programming interface for large-scale distributed graph computing. It generalizes previous graph processing frameworks (e.g. , Pregel, GraphX) and distributed graph databases ( e.g ., Janus-Graph, Neptune) in two important ways: by exposing a unified programming interface to a wide variety of graph computations such as graph traversal, pattern matching, iterative algorithms and graph neural networks within a high-level programming language; and by supporting the seamless integration of a highly optimized graph engine in a general purpose data-parallel computing system. A GraphScope program is a sequential program composed of declarative data-parallel operators, and can be written using standard Python development tools. The system automatically handles the parallelization and distributed execution of programs on a cluster of machines. It outperforms current state-of-the-art systems by enabling a separate optimization (or family of optimizations) for each graph operation in one carefully designed coherent framework. We describe the design and implementation of GraphScope and evaluate system performance using several real-world applications.
Wenfei Fan, Tao He 0013, Longbin Lai, Xue Li 0024, Yong Li 0020, Zhao Li 0007, Zhengping Qian, Chao Tian 0001, Lei Wang 0004, Jingbo Xu 0001, Youyang Yao, Qiang Yin 0002, Wenyuan Yu, Kai Zeng 0002, Jingren Zhou 0001, Diwen Zhu
Proc. VLDB Endow.2