Yuchen Zhong

dblp:164/5547 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Efficient and distributed learning · 58% Graph learning · 30% Planning, search and constraint satisfaction · 11%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Distributed systems · 76% Parallel and multicore computing · 18% GPUs and heterogeneous computing · 7%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
2.332025
Heta: Distributed Training of Heterogeneous Graph Neural Networks · Proc. VLDB Endow. 2025
Swift: Expedited Failure Recovery for Large-Scale DNN Training · IEEE Trans. Parallel Distributed Syst. 2024
Swift: Expedited Failure Recovery for Large-Scale DNN Training · PPoPP 2023
Distributed systems › fault tolerance
failure recovery
1.422024
Swift: Expedited Failure Recovery for Large-Scale DNN Training · IEEE Trans. Parallel Distributed Syst. 2024
Swift: Expedited Failure Recovery for Large-Scale DNN Training · PPoPP 2023
Distributed systems
fault tolerance
1.422024
Swift: Expedited Failure Recovery for Large-Scale DNN Training · IEEE Trans. Parallel Distributed Syst. 2024
Swift: Expedited Failure Recovery for Large-Scale DNN Training · PPoPP 2023
Machine learning › Efficient and distributed learning › distributed training
communication-efficient training
0.912025
Heta: Distributed Training of Heterogeneous Graph Neural Networks · Proc. VLDB Endow. 2025
Machine learning › Graph learning
graph neural network
0.912025
Heta: Distributed Training of Heterogeneous Graph Neural Networks · Proc. VLDB Endow. 2025
Machine learning › Graph learning › graph neural network
heterogeneous graph neural network
0.912025
Heta: Distributed Training of Heterogeneous Graph Neural Networks · Proc. VLDB Endow. 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › plan execution
failure recovery
0.712023
Swift: Expedited Failure Recovery for Large-Scale DNN Training · PPoPP 2023
Parallel and multicore computing
distributed deep learning training
0.712023
Swift: Expedited Failure Recovery for Large-Scale DNN Training · PPoPP 2023
GPUs and heterogeneous computing › GPU memory management
GPU cache management
0.312025
Heta: Distributed Training of Heterogeneous Graph Neural Networks · Proc. VLDB Endow. 2025
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training
0.212024
Swift: Expedited Failure Recovery for Large-Scale DNN Training · IEEE Trans. Parallel Distributed Syst. 2024
Image and video processing › image segmentation
contour detection
0.212015
Thin Structure Estimation with Curvature Regularization · ICCV 2015
Image and video processing › image restoration › inverse problem › inverse problem regularization › image regularization
curvature regularization
0.212015
Thin Structure Estimation with Curvature Regularization · ICCV 2015

Methods — techniques the papers use, named apart from their topics

data parallelism · 2.8relation-aggregation-first paradigm · 1.7meta-partitioning · 1.7heterogeneity-aware caching · 1.7logging-based recovery · 1.5computation replay · 1.5logging and replay · 1.3checkpointing · 1.3lower bound optimization · 0.4joint optimization · 0.4curvature regularization · 0.4
YearPublicationVenuePosition
2025 Heta: Distributed Training of Heterogeneous Graph Neural Networks
abstract
Heterogeneous Graphs (HetGs) that capture relationships among different types of nodes are ubiquitous in real-world applications such as academic networks and e-commerce. Although Heterogeneous Graph Neural Networks (HGNNs) have demonstrated superior performance in learning from these complex structures, distributed training of HGNNs on large-scale graphs with billions of edges faces substantial communication overhead. This challenge is exacerbated by heterogeneous characteristics such as varying feature dimensions across node types and featureless nodes requiring learnable parameters. Existing systems and communication reduction techniques designed for homogeneous graphs become suboptimal or even inapplicable for HetGs and HGNNs by overlooking both these heterogeneous characteristics and the inherent computational structure of HGNNs. We present Heta , a framework designed to address the communication bottleneck in distributed HGNN training. Heta leverages the key insight that HGNN aggregation is order-invariant and decomposable into relation-specific computations. Built on this insight, we introduce three key innovations: (1) a Relation-Aggregation-First (RAF) paradigm that conducts relation-specific aggregations within partitions and exchanges only partial aggregations across machines, proven to reduce communication complexity; (2) a meta-partitioning strategy that divides a HetG based on its graph schema and HGNN computation dependency while minimizing cross-partition communication and maintaining computation and storage balance; and (3) a heterogeneity-aware GPU cache system that accounts for varying miss-penalty ratios across node types. Through extensive evaluation of billion-edge heterogeneous graphs, we demonstrate that Heta achieves up to 5.3X and 4.4X speedup over state-of-the-art systems DGL and GraphLearn while maintaining model accuracy.
Yuchen Zhong, Junwei Su, Chuan Wu 0001
Proc. VLDB Endow.1
2024 Contrast-based unsupervised hashing method with margin limit
Hai Su, Zhenyu Ke, Songsen Yu, Jianwei Fang, Yuchen Zhong
Multim. Tools Appl.5
2024 Swift: Expedited Failure Recovery for Large-Scale DNN Training
abstract
As the size of deep learning models gets larger and larger, training takes longer time and more resources, making fault tolerance more and more critical. Existing state-of-the-art methods like CheckFreq and Elastic Horovod need to back up a copy of the model state (i.e., parameters and optimizer states) in memory, which is costly for large models and leads to non-trivial overhead. This article presentsSwift, a novel recovery design for distributed deep neural network training that significantly reduces the failure recovery overhead without affecting training throughput and model accuracy. Instead of making an additional copy of the model state,Swiftresolves the inconsistencies of the model state caused by the failure and exploits the replicas of the model state in data parallelism for failure recovery. We propose a logging-based approach when replicas are unavailable, which records intermediate data and replays the computation to recover the lost state upon a failure. The re-computation is distributed across multiple machines to accelerate failure recovery further. We also log intermediate data selectively, exploring the trade-off between recovery time and intermediate data storage overhead. Evaluations show thatSwiftsignificantly reduces the failure recovery time and achieves similar or better training throughput during failure-free execution compared to state-of-the-art methods without degrading final model accuracy.Swiftcan also achieve up to 1.16x speedup in total training time compared to state-of-the-art methods.
Yuchen Zhong, Guangming Sheng, Jinhui Yuan, Chuan Wu 0001
IEEE Trans. Parallel Distributed Syst.1
2023 Swift: Expedited Failure Recovery for Large-Scale DNN Training
abstract
As the size of deep learning models gets larger and larger, training takes longer time and more resources, making fault tolerance critical. Existing state-of-the-art methods like Check-Freq and Elastic Horovod need to back up a copy of the model state in memory, which is costly for large models and leads to non-trivial overhead. This paper presents Swift, a novel failure recovery design for distributed deep neural network training that significantly reduces the failure recovery overhead without affecting training throughput and model accuracy. Instead of making an additional copy of the model state, Swift resolves the inconsistencies of the model state caused by the failure and exploits replicas of the model state in data parallelism for failure recovery. We propose a logging-based approach when replicas are unavailable, which records intermediate data and replays the computation to recover the lost state upon a failure. Evaluations show that Swift significantly reduces the failure recovery time and achieves similar or better training throughput during failure-free execution compared to state-of-the-art methods without degrading final model accuracy.
Yuchen Zhong, Guangming Sheng, Jinhui Yuan, Chuan Wu 0001
PPoPP1
2015 Thin Structure Estimation with Curvature Regularization
abstract
Many applications in vision require estimation of thin structures such as boundary edges, surfaces, roads, blood vessels, neurons, etc. Unlike most previous approaches, we simultaneously detect and delineate thin structures with sub-pixel localization and real-valued orientation estimation. This is an ill-posed problem that requires regularization. We propose an objective function combining detection likelihoods with a prior minimizing curvature of the center-lines or surfaces. Unlike simple block-coordinate descent, we develop a novel algorithm that is able to perform joint optimization of location and detection variables more effectively. Our lower bound optimization algorithm applies to quadratic or absolute curvature. The proposed early vision framework is sufficiently general and it can be used in many higher-level applications. We illustrate the advantage of our approach on a range of 2D and 3D examples.
Dmitrii Marin, Yuchen Zhong, Maria Drangova, Yuri Boykov
ICCV2