VLDB 2026 Research / reviewers in the wild / expert
Weijian Chen 0002
dblp:164/0613-2
· DBLP profile ↗
11ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0001-8296-2673ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 4 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Frenzy: A Memory-Aware Serverless LLM Training System for Heterogeneous GPU Clusters
Zihan Chang, Sheng Xiao, Shuibing He, Xuechen Zhang 0001, Siling Yang, Zhenxin Li, Weijian Chen 0002 |
Euro-Par (2) | 8 |
| 2025 | LeapGNN: Accelerating Distributed GNN Training Leveraging Feature-Centric Model Migration
Weijian Chen 0002, Shuibing He, Haoyang Qu, Xuechen Zhang 0001 |
FAST | 1 |
| 2025 | IMPRESS: An Importance-Informed Multi-Tier Prefix KV Storage System for Large Language Model Inference
Weijian Chen 0002, Shuibing He, Haoyang Qu, Siling Yang, Baoxing Huai, Gang Chen 0001 |
FAST | 1 |
| 2025 | GoPIM: GCN-Oriented Pipeline Optimization for PIM AcceleratorsabstractGraph convolutional networks (GCNs) are popular for a variety of graph learning tasks. ReRAM-based processing-in-memory (PIM) accelerators are promising to expedite GCN training owing to their in-situ computing capability. However, existing accelerators can be severely underutilized even with pipelines, due to the oversight of the skewed execution times of various GCN stages and the ignorance of skewed degrees of graph vertices. In this work, we propose GOPIM, a GCN-oriented pipeline optimization for PIM accelerators to expedite GCN training. First, GOPIM proposes an ML-based scheme that allocates crossbar resources to the most needed stages to streamline the overall pipeline. Second, GOPIM utilizes a selective vertex updating technique that evenly distributes vertices on crossbars by interleaved mapping. These techniques collectively reduce the overall execution time without losing much accuracy. We also provide a practical architecture design for GOPIM. Our experimental results show that, GoPIM achieves up to 191 × speedup and 16.1 × energy saving, compared to the state-of-the-art work. Siling Yang, Shuibing He, Wenjiong Wang, Yanlong Yin, Weijian Chen 0002, Xuechen Zhang 0001, Xian-He Sun |
HPCA | 6 |
| 2025 | ImPACT: Importance-Informed Prefetching and Caching for I/O-Bound DNN TrainingabstractFetching large amounts of DNN training data from storage systems causes high I/O latency and GPU stalls. Importance sampling can reduce data processing on GPUs while maintaining model accuracy, but current frameworks lack a prefetching and caching layer to optimize data fetches and cache management based on sample importance. This leads to unnecessary fetches, poor cache hit ratios, and random I/Os. We present ImPACT, an importance-informed prefetching and caching system, to accelerate I/O-bound DNN training. First, we propose an importance-informed prefetching technique to reduce the prefetching of unimportant data. Then, we introduce an importance-aware caching layer, partitioned into two regions: H-cache and L-cache, which store samples of high importance and low importance respectively. Rather than using recency or frequency, we manage data items in H-cache according to their corresponding sample importance. When there is a cache miss in L-cache, we use sample substitutability and dynamic packaging to improve the cache hit ratio and reduce the number of random I/Os. Our experimental results show that ImPACT has a negligible impact on training accuracy while speeding up DNN training by up to 3.5× compared to state-of-the-art prefetching and caching systems. Weijian Chen 0002, Shuibing He, Xuechen Zhang 0001, Siling Yang, Haoyang Qu, Xuan Zhan |
IEEE Trans. Computers | 1 |
| 2024 | AUTOHET: An Automated Heterogeneous ReRAM-Based Accelerator for DNN InferenceabstractReRAM-based accelerators have become prevalent in accelerating deep neural network inference owing to their in-situ computing capability of ReRAM crossbars. However, most existing ReRAM-based accelerators are designed with homogeneous crossbars, leading to either low resource utilization or sub-optimal energy efficiency. In this paper, we propose AutoHet, an automated heterogeneous ReRAM-based accelerator with varied-size crossbars for different DNN layers. To achieve both high crossbar utilization and energy efficiency, AutoHet uses a reinforcement learning algorithm to automatically determine the proper crossbar configuration for each DNN layer. Additionally, AutoHet introduces rectangle crossbars and a tile-shared crossbar allocation scheme to reduce crossbar wastage and energy consumption. Experiment results show that AutoHet effectively improves crossbar utilization by up to 3.1 × and reduces energy consumption by up to 94.6%, compared to approaches with homogeneous ReRAM crossbars. Shuibing He, Weijian Chen 0002, Siling Yang, Yanlong Yin, Xuechen Zhang 0001, Xian-He Sun, Gang Chen 0001 |
ICPP | 4 |
| 2024 | IOWA: An I/O-Aware Adaptive Sampling Framework for Deep LearningabstractTraining deep DNN models is time-consuming, especially when using large datasets. In the standard model training process, data instances are sampled uniformly and fed into the neural networks. However, not all instances contribute equally to the resulting model, and even the same data instance may affect the model differently in different training iterations. In addition to computational costs, I/O overhead can significantly impact the training speed, particularly for I/O-intensive processes. Given these observations, we propose an I/O-aware sampling metric in this paper. Building on this, we introduce an I/O-Aware Adaptive Sampling Framework (IOWA), which includes data profiling, adaptive data sampling, and redundant data instance replacement to accelerate the training process. Extensive exper-iments demonstrate that, compared to traditional DNN training processes, our approach can achieve up to a 3 x speedup without compromising the resulting model. Weijian Chen 0002, Yanlong Yin, Shuibing He |
NAS | 2 |
| 2023 | iCache: An Importance-Sampling-Informed Cache for Accelerating I/O-Bound DNN Model TrainingabstractFetching a large amount of DNN training data from storage systems incurs long I/O latency and fetch stalls of GPUs. Importance sampling in DNN training can reduce the amount of data computing on GPUs while maintaining a similar model accuracy. However, existing DNN training frameworks do not have a cache layer that reduces the number of data fetches and manages cached items according to sample importance, resulting in unnecessary data fetches, poor cache hit ratios, and random I/Os when importance sampling is used.In this paper, we design a new importance-sampling-informed cache, namely, iCache, to accelerate I/O bound DNN training jobs. iCache only fetches parts of samples instead of all samples in the dataset. The cache is partitioned into two regions: H-cache and L-cache, which store samples of high importance and low importance respectively. Rather than using recency or frequency, we manage data items in H-cache according to their corresponding sample importance. When there is a cache miss in L-cache, we use sample substitutability and dynamic packaging to improve the cache hit ratio and reduce the number of random I/Os. When multiple concurrent jobs access the same datasets in H-cache, we design a model to assign the relative importance values to cached samples to avoid cache thrashing, which may happen when there is no coordination among the concurrent training jobs. Our experimental results show that iCache has a negligible impact on training accuracy and speeds up the DNN training time by up to 2.0× compared to the state-of-the-art caching systems. Weijian Chen 0002, Shuibing He, Yaowen Xu, Xuechen Zhang 0001, Siling Yang, Xian-He Sun, Gang Chen 0001 |
HPCA | 1 |
| 2023 | HOME: A Holistic GPU Memory Management Framework for Deep LearningabstractWe propose HOlistic MEmory management (HOME), a new framework for performing tensor placements in large DNN training when GPU memory space is not enough. HOME combines tensor swapping with tensor recomputation to reduce GPU memory footprint. Different from existing work that only considers partial DNN model information, HOME takes the holistic DNN model information into account in tensor placement decisions. More specifically, HOME uses a custom-designed particle swarm optimization algorithm to achieve the globally optimized placement for each tensor of the DNN model with a greatly reduced searching space. This holistic awareness of the whole model information enables HOME to obtain high performance under the given GPU memory constraint. We implement HOME in PyTorch and conduct our experiments using six popular DNN models. Experimental results show that HOME can outperform vDNN and Capuchin by up to 5.7× and 1.3× in throughput. Furthermore, HOME can improve the maximum batch size by up to 2.8× than the original PyTorch and up to 1.3× than Capuchin. Shuibing He, Shuaiben Chen, Zheng Li 0006, Siling Yang, Weijian Chen 0002, Lidan Shou |
IEEE Trans. Computers | 6 |
| 2023 | APQ: Automated DNN Pruning and Quantization for ReRAM-Based AcceleratorsabstractEmerging ReRAM-based accelerators support in-memory computation to accelerate deep neural network (DNN) inference. Weight matrix pruning is a widely used technique to reduce the size of DNN models, thereby reducing the resource and energy consumption of ReRAM-based accelerators. However, existing pruning works for ReRAM-based accelerators have three major issues. First, they use heuristics or rules from domain experts to prune the weights, leading to sub-optimal pruning policies. Second, they use row or column-level coarse-granularity methods to prune weights, resulting in poor compression rates with model accuracy constraints. Third, they only apply the weight pruning technique individually, losing the compression opportunity of both pruning and quantization. In this article, we propose an Automated DNN Pruning and Quantization framework, namedAPQ, for ReRAM-based accelerators. First,APQadopts reinforcement learning (RL) to automatically determine the pruning policy for DNN layers for a global optimum. Second, it prunes and maps weight matrices to a ReRAM-based accelerator in a finer granularity of column-vector, which improves the compression rates with the accuracy constraints. To address the dislocation problem, it uses a new data path in ReRAM-based accelerators to correctly index and feed input to matrix-vector computation. Third, to further reduce resource consumption,APQalso leverages reinforcement learning to automatically determine the quantization bitwidth of each layer of the pruned DNN model. Experimental results show that,APQachieves up to 4.52X compression rate, 4.11X area efficiency, and 4.51X energy efficiency with similar or even higher model accuracy, compared to the state-of-the-art work. Siling Yang, Shuibing He, Hexiao Duan, Weijian Chen 0002, Xuechen Zhang 0001, Yanlong Yin |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | AUTO-PRUNE: automated DNN pruning and mapping for ReRAM-based acceleratorabstractEmergent ReRAM-based accelerators support in-memory computation to accelerate deep neural network (DNN) inference. Weight matrix pruning of DNNs is a widely used technique to reduce the size of DNN models, thereby reducing the resource and energy consumption of ReRAM-based accelerators. However, conventional works on weight matrix pruning for ReRAM-based accelerators have three major issues. First, they use heuristics or rules from domain experts to prune the weights, leading to suboptimal pruning policies. Second, they mostly focus on improving compression ratio, thus may not meet accuracy constraints. Third, they ignore direct feedback of hardware. In this paper, we introduce an automated DNN pruning and mapping framework, named AUTO-PRUNE. It leverages reinforcement learning (RL) to automatically determine the pruning policy considering the constraint of accuracy loss. The reward function of RL agents is designed using hardware’s direct feedback (i.e., accuracy and compression rate of occupied crossbars). The function directs the search of the pruning ratio of each layer for a global optimum considering the characteristics of individual layers of DNN models. Then AUTO-PRUNE maps the pruned weight matrices to crossbars to store only nontrivial elements. Finally, to avoid the dislocation problem, we design a new data-path in ReRAM-based accelerators to correctly index and feed input to matrix-vector computation leveraging the mechanism of operation units. Experimental results show that, compared to the state-of-the-art work, AUTO-PRUNE achieves up to 3.3X compression rate, 3.1X area efficiency, and 3.3X energy efficiency with a similar or even higher accuracy. Siling Yang, Weijian Chen 0002, Xuechen Zhang 0001, Shuibing He, Yanlong Yin, Xian-He Sun |
ICS | 2 |