Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Bowen Yu 0003

dblp:95/10266-3 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
5since 2021 · last 2023
0000-0001-5537-8244ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
8 papers
High-performance computing · 30% Distributed systems · 22% Storage systems · 22%
Software engineering, system software, and programming languages
3 papers
Compilers and program optimization · 64% Operating systems · 26% Programming languages and type systems · 10%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 100%

Topics — the 24 heaviest of 25, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems › cache management › storage caching
block cache
1.222023
TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory Programs · ACM Trans. Storage 2023
TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory Programs · OSDI 2022
Storage systems
out-of-core computation
1.222023
TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory Programs · ACM Trans. Storage 2023
TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory Programs · OSDI 2022
Compilers and program optimization › parallelization
automatic parallelization
0.712023
Mat2Stencil: A Modular Matrix-Based DSL for Explicit and Implicit Matrix-Free PDE Solvers on Structured Grid · Proc. ACM Program. Lang. 2023
Compilers and program optimization
domain-specific compilation
0.712023
Mat2Stencil: A Modular Matrix-Based DSL for Explicit and Implicit Matrix-Free PDE Solvers on Structured Grid · Proc. ACM Program. Lang. 2023
Storage systems
file systems
0.712023
TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory Programs · ACM Trans. Storage 2023
High-performance computing › scientific computing systems
partial differential equation solver
0.712023
Mat2Stencil: A Modular Matrix-Based DSL for Explicit and Implicit Matrix-Free PDE Solvers on Structured Grid · Proc. ACM Program. Lang. 2023
High-performance computing
stencil computation
0.712023
Mat2Stencil: A Modular Matrix-Based DSL for Explicit and Implicit Matrix-Free PDE Solvers on Structured Grid · Proc. ACM Program. Lang. 2023
Distributed systems › fault tolerance
checkpointing
0.622018
An Efficient In-Memory Checkpoint Method and its Practice on Fault-Tolerant HPL · IEEE Trans. Parallel Distributed Syst. 2018
Self-Checkpoint: An In-Memory Checkpoint Method Using Less Space and Its Practice on Fault-Tolerant HPL · PPoPP 2017
Distributed systems › fault tolerance › checkpointing
diskless checkpointing
0.622018
An Efficient In-Memory Checkpoint Method and its Practice on Fault-Tolerant HPL · IEEE Trans. Parallel Distributed Syst. 2018
Self-Checkpoint: An In-Memory Checkpoint Method Using Less Space and Its Practice on Fault-Tolerant HPL · PPoPP 2017
Distributed systems
fault tolerance
0.622018
An Efficient In-Memory Checkpoint Method and its Practice on Fault-Tolerant HPL · IEEE Trans. Parallel Distributed Syst. 2018
Self-Checkpoint: An In-Memory Checkpoint Method Using Less Space and Its Practice on Fault-Tolerant HPL · PPoPP 2017
Query processing and optimization › query compilation
operator fusion
0.512021
Chukonu: A Fully-Featured Big Data Processing System by Efficiently Integrating a Native Compute Engine into Spark · Proc. VLDB Endow. 2021
Query processing and optimization
query execution
0.512021
Chukonu: A Fully-Featured Big Data Processing System by Efficiently Integrating a Native Compute Engine into Spark · Proc. VLDB Endow. 2021
Operating systems › resource management
memory management
0.312018
Spindle: Informed Memory Access Monitoring · USENIX ATC 2018
High-performance computing
large-scale graph processing
0.312018
ShenTu: processing multi-trillion edge graphs on millions of cores in seconds · SC 2018
Memory systems › memory access
memory access monitoring
0.312018
Spindle: Informed Memory Access Monitoring · USENIX ATC 2018
Parallel and multicore computing
parallel graph algorithms
0.312018
ShenTu: processing multi-trillion edge graphs on millions of cores in seconds · SC 2018
High-performance computing
scientific computing systems
0.312018
ShenTu: processing multi-trillion edge graphs on millions of cores in seconds · SC 2018
Programming languages and type systems
domain-specific languages
0.212023
Mat2Stencil: A Modular Matrix-Based DSL for Explicit and Implicit Matrix-Free PDE Solvers on Structured Grid · Proc. ACM Program. Lang. 2023
Operating systems › resource management › memory management
page cache
0.212023
TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory Programs · ACM Trans. Storage 2023
High-performance computing › numerical linear algebra
linpack
0.222018
An Efficient In-Memory Checkpoint Method and its Practice on Fault-Tolerant HPL · IEEE Trans. Parallel Distributed Syst. 2018
Self-Checkpoint: An In-Memory Checkpoint Method Using Less Space and Its Practice on Fault-Tolerant HPL · PPoPP 2017
Memory systems
cache management
0.212022
TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory Programs · OSDI 2022
Cloud and datacenter computing
cluster computing framework
0.112021
Chukonu: A Fully-Featured Big Data Processing System by Efficiently Integrating a Native Compute Engine into Spark · Proc. VLDB Endow. 2021
Performance modeling and evaluation
benchmarking
0.112018
An Efficient In-Memory Checkpoint Method and its Practice on Fault-Tolerant HPL · IEEE Trans. Parallel Distributed Syst. 2018
Distributed systems
distributed graph processing
0.112018
ShenTu: processing multi-trillion edge graphs on millions of cores in seconds · SC 2018

Methods — techniques the papers use, named apart from their topics

task partitioning · 1.3multi-stage programming · 1.3multi-level block cache · 1.3fine-grained synchronization · 1.3address translation · 1.3vectorization · 1.0compaction · 1.0DAG splitting · 1.0memory access monitoring · 0.7in-memory checkpointing · 0.6
YearPublicationVenuePosition
2023 Mat2Stencil: A Modular Matrix-Based DSL for Explicit and Implicit Matrix-Free PDE Solvers on Structured Grid
abstract
Partial differential equation (PDE) solvers are extensively utilized across numerous scientific and engineering fields. However, achieving high performance and scalability often necessitates intricate and low-level programming, particularly when leveraging deterministic sparsity patterns in structured grids. In this paper, we propose an innovative domain-specific language (DSL), Mat2Stencil, with its compiler, for PDE solvers on structured grids. Mat2Stencil introduces a structured sparse matrix abstraction, facilitating modular, flexible, and easy-to-use expression of solvers across a broad spectrum, encompassing components such as Jacobi or Gauss-Seidel preconditioners, incomplete LU or Cholesky decompositions, and multigrid methods built upon them. Our DSL compiler subsequently generates matrix-free code consisting of generalized stencils through multi-stage programming. The code allows spatial loop-carried dependence in the form of quasi-affine loops, in addition to the Jacobi-style stencil’s embarrassingly parallel on spatial dimensions. We further propose a novel automatic parallelization technique for the spatially dependent loops, which offers a compile-time deterministic task partitioning for threading, calculates necessary inter-thread synchronization automatically, and generates an efficient multi-threaded implementation with fine-grained synchronization. Implementing 4 benchmarking programs, 3 of them being the pseudo-applications in NAS Parallel Benchmarks with 6.3% lines of code and 1 being matrix-free High Performance Conjugate Gradients with 16.4% lines of code, we achieve up to 1.67× and on average 1.03× performance compared to manual implementations.
Huanqi Cao, Shizhi Tang, Qianchao Zhu, Bowen Yu 0003
Proc. ACM Program. Lang.4
2023 TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory Programs
abstract
Out-of-core systems rely on high-performance cache sub-systems to reduce the number of I/O operations. Although the page cache in modern operating systems enables transparent access to memory and storage devices, it suffers from efficiency and scalability issues on cache misses, forcing out-of-core systems to design and implement their own cache components, which is a non-trivial task. This study proposes TriCache, a cache mechanism that enables in-memory programs to efficiently process out-of-core datasets without requiring any code rewrite. It provides a virtual memory interface on top of the conventional block interface to simultaneously achieve user transparency and sufficient out-of-core performance. A multi-level block cache design is proposed to address the challenge of per-access address translations required by a memory interface. It can exploit spatial and temporal localities in memory or storage accesses to render storage-to-memory address translation and page-level concurrency control adequately efficient for the virtual memory interface. Our evaluation shows that in-memory systems operating on top of TriCache can outperform Linux OS page cache by more than one order of magnitude, and can deliver performance comparable to or even better than that of corresponding counterparts designed specifically for out-of-core scenarios.
Guanyu Feng, Huanqi Cao, Xiaowei Zhu 0001, Bowen Yu 0003, Yuanwei Wang, Zixuan Ma, Shengqi Chen 0001
ACM Trans. Storage4
2022 TriCache: A User-Transparent Block Cache Enabling High-Performance Out-of-Core Processing with In-Memory Programs
Guanyu Feng, Huanqi Cao, Xiaowei Zhu 0001, Bowen Yu 0003, Yuanwei Wang, Zixuan Ma, Shengqi Chen 0001
OSDI4
2021 Sparker: Efficient Reduction for More Scalable Machine Learning with Spark
abstract
Machine learning applications on Spark suffers from poor scalability. In this paper, we reveal that the key reasons is the non-scalable reduction, which is restricted by the non-splittable object programming interface in Spark. This insight guides us to propose Sparker, Spark with Efficient Reduction. By providing a split aggregation interface, Sparker is able to perform split aggregation with scalable reduction while being backward compatible with existing applications. We implemented Sparker in 2,534 lines of code. Sparker can improve the aggregation performance by up to 6.47 × and can improve the end-to-end performance of MLlib model training by up to 3.69 × with a geometric mean of 1.81 × .
Bowen Yu 0003, Huanqi Cao, Tianyi Shan, Haojie Wang 0004, Xiongchao Tang
ICPP1
2021 Chukonu: A Fully-Featured Big Data Processing System by Efficiently Integrating a Native Compute Engine into Spark
abstract
Apache Spark is a widely deployed big data analytics framework that offers such attractive features as resiliency, load-balancing, and a rich ecosystem. However, there is still plenty of room for improvement in its performance. Although a data-parallel system in a native programming language significantly improves performance, it may require re-implementing many functionalities of Spark to become a full-featured system. It is desirable for native big data systems to just write a compute engine in native languages to ensure high efficiency, and reuse other mature features provided by Spark rather than re-implement everything. But the interaction between the JVM and the native world risks becoming a bottleneck. This paper proposes Chukonu, a native big data framework that re-uses critical big data features provided by Spark. Owing to our novel DAG-splitting approach, the potential Spark integration overhead is alleviated, and its even outperforms existing pure native big data frameworks. Chukonu splits DAG programs into run-time parts and compile-time parts: The run-time parts are delegated to Spark to offload the complexities due to feature implementations. The compile-time parts are natively compiled. We propose a series of optimization techniques to be applied to the compile-time parts, such as operator fusion, vectorization, and compaction, to significantly reduce the Spark integration overhead. The results of evaluation show that Chukonu has a speedup of up to 71.58X (geometric mean 6.09X) over Apache Spark, and up to 7.20X (geometric mean 2.30X) over pure-native frameworks on six commonly-used big data applications. By translating the physical plan produced by SparkSQL into Chukonu programs, Chukonu accelerates Spark-SQL's TPC-DS performance by 2.29X.
Bowen Yu 0003, Guanyu Feng, Huanqi Cao, Zhenbo Sun, Haojie Wang 0004, Xiaowei Zhu 0001
Proc. VLDB Endow.1
2018 ShenTu: processing multi-trillion edge graphs on millions of cores in seconds
Heng Lin, Xiaowei Zhu 0001, Bowen Yu 0003, Xiongchao Tang, Wei Xue 0003, Lufei Zhang, Torsten Hoefler, Xiaosong Ma, Xin Liu 0081, Jingfang Xu
SC3
2018 Spindle: Informed Memory Access Monitoring
Haojie Wang 0004, Jidong Zhai, Xiongchao Tang, Bowen Yu 0003, Xiaosong Ma
USENIX ATC4
2018 An Efficient In-Memory Checkpoint Method and its Practice on Fault-Tolerant HPL
abstract
Fault tolerance is increasingly important in high-performance computing due to the substantial growth of system scale and decreasing system reliability. In-memory/diskless checkpoint has gained extensive attention as a solution to avoid the IO bottleneck of traditional disk-based checkpoint methods. However, applications using previous in-memory checkpoint suffer from little available memory space. To provide high reliability, previous in-memory checkpoint methods either need to keep two copies of checkpoints to tolerate failures while updating old checkpoints or trade performance for space by flushing in-memory checkpoints into disk. In this paper, we propose a novel in-memory checkpoint method, called self-checkpoint, which can not only achieve the same reliability of previous in-memory checkpoint methods, but also increase the available memory space for applications by almost 50 percent. To validate our method, we apply self-checkpoint method to an important problem: High-Performance Linpack (HPL) with fault tolerance. We implement a scalable and fault tolerant HPL based on this new method, called SKT-HPL, and validate it on two large-scale systems. Experimental results with 24,576 processes show that SKT-HPL achieves over 95 percent of the performance of the original HPL. Compared to the state-of-the-art in-memory checkpoint method, it improves the available memory size by 47 percent and the performance by 5 percent.
Xiongchao Tang, Jidong Zhai, Bowen Yu 0003, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.3
2017 Scalable Graph Traversal on Sunway TaihuLight with Ten Million Cores
abstract
Interest has recently grown in efficiently analyzing unstructured data such as social network graphs and protein structures. A fundamental graph algorithm for doing such task is the Breadth-First Search (BFS) algorithm, the foundation for many other important graph algorithms such as calculating the shortest path or finding the maximum flow in graphs. In this paper, we share our experience of designing and implementing the BFS algorithm on Sunway TaihuLight, a newly released machine with 40,960 nodes and 10.6 million accelerator cores. It tops the Top500 list of June 2016 with a 93.01 petaflops Linpack performance [1]. Designed for extremely large-scale computation and power efficiency, processors on Sunway TaihuLight employ a unique heterogeneous many-core architecture and memory hierarchy. With its extremely large size, the machine provides both opportunities and challenges for implementing high-performance irregular algorithms, such as BFS. We propose several techniques, including pipelined module mapping, contention-free data shuffling, and group-based message batching, to address the challenges of efficiently utilizing the features of this large scale heterogeneous machine. We ultimately achieved 23755.7 giga-traversed edges per second (GTEPS), which is the best among heterogeneous machines and the second overall in the Graph500s June 2016 list [2].
Heng Lin, Xiongchao Tang, Bowen Yu 0003, Youwei Zhuo, Jidong Zhai, Wanwang Yin
IPDPS3
2017 Self-Checkpoint: An In-Memory Checkpoint Method Using Less Space and Its Practice on Fault-Tolerant HPL
abstract
Fault tolerance is increasingly important in high performance computing due to the substantial growth of system scale and decreasing system reliability. In-memory/diskless checkpoint has gained extensive attention as a solution to avoid the IO bottleneck of traditional disk-based checkpoint methods. However, applications using previous in-memory checkpoint suffer from little available memory space. To provide high reliability, previous in-memory checkpoint methods either need to keep two copies of checkpoints to tolerate failures while updating old checkpoints or trade performance for space by flushing in-memory checkpoints into disk.
Xiongchao Tang, Jidong Zhai, Bowen Yu 0003
PPoPP3