Jia Wei 0002

dblp:354/4904-2 · DBLP profile ↗
← Back
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-2234-0378ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MAFNet: Multi-scale active fusion network for long-term time series forecasting
Qianyang Li, Xingjun Zhang, Shaoxun Wang, Jia Wei 0002
Neurocomputing4
2026 Dual-Pronged Deep Learning Preprocessing on Heterogeneous Platforms With CPU, Accelerator and CSD
abstract
For image-related deep learning tasks, the first step often involves reading data from external storage and performing preprocessing on the CPU. As accelerator speed increases and the number of single compute node accelerators increases, the computing and data transfer capabilities gap between accelerators and CPUs gradually increases. Data reading and preprocessing become progressively the bottleneck of these tasks. Our work, DDLP, addresses the data computing and transfer bottleneck of deep learning preprocessing using Computable Storage Devices (CSDs). DDLP allows the CPU and CSD to efficiently parallelize preprocessing from both ends of the datasets, respectively. To this end, we propose two adaptive dynamic selection strategies to make DDLP control the accelerator to automatically read data from different sources. The two strategies trade-off between consistency and efficiency. DDLP achieves sufficient computational overlap between CSD data preprocessing and CPU preprocessing, accelerator computation, and accelerator data reading. In addition, DDLP leverages direct storage technology to enable efficient SSD-to-accelerator data transfer. In addition, DDLP reduces the use of expensive CPU and DRAM resources with more energy-efficient CSDs, alleviating preprocessing bottlenecks while significantly reducing power consumption. Extensive experimental results show that DDLP can improve learning speed by up to 23.5% on ImageNet Dataset while reducing energy consumption by 19.7% and CPU and DRAM usage by 37.6%. DDLP also improves the learning speed by up to 27.6% on the Cifar-10 dataset.
Jia Wei 0002, Xingjun Zhang, Witold Pedrycz, Jie Zhao 0002
IEEE Trans. Computers1
2025 Dynamic Fuzzy Sampler for Graph Neural Networks
abstract
Graph neural networks (GNNs) have been widely used in many fields. Inductive learning has replaced transductive learning as the current mainstream paradigm for GNN training due to its higher memory efficiency, computing speed, and stronger generalization. Neighbor node sampling as a key step in GNN inductive learning is crucial to the model performance. However, existing samplers only focus on how to sample nodes from the adjacency matrix, ignoring the fact that different neighbors have different impacts on the target node at different moments. They usually aggregate the neighbor information in a simple way such as averaging or summing, which limits the information representation, robustness, and generalization. In order to address the shortcomings of existing graph inductive learning samplers, this article proposes a dynamic fuzzy sampler (DFS) based on a Gaussian fuzzy system. The DFS fully takes into account the diversity of nodes in the graph-structured data, and efficiently models and handles the uncertainties and fuzziness of the mutual information of various nodes at different moments. Specifically, DFS first innovatively constructs a learnable Gaussian fuzzy set system for determining the membership degree of different neighbors to the target node at different moments. Subsequently, DFS aggregates the target node embeddings and membership-weighted neighbor embeddings to update the target node's features, which makes the target node utilize the sampling information more effectively. The aggregated target node effectively captures the graph structure information and neighbor node information, which can facilitate the subsequent GNN-based graph representation model with stronger representation and generalization capabilities. Our supervised and self-supervised experimental results on graph datasets of different sizes show that DFS has consistently excellent performance, significantly outperforming other state-of-the-art sampling schemes. DFS achieves up to 1.90% and 9.52% F1-score improvement compared to the state-of-the-art schemes on small- and large-scale graphs, respectively.
Jia Wei 0002, Xingjun Zhang, Witold Pedrycz, Weiping Ding 0001
IEEE Trans. Fuzzy Syst.1
2024 Revisit and Benchmarking of Automated Quantization Toward Fair Comparison
abstract
Automated quantization has emerged as an entirely new design paradigm to automate the optimal configuration of bitwidth for deep neural networks (DNNs), making the DNN more memory-efficient and faster to execute on hardware with limited resources. Reinforcement learning (RL) and differentiable neural architecture search (DNAS) are two main solution paths that have shown their superiority. Yet, there are countless methods with various implementations within each path. It has been hard to comprehend their differences and make a relatively fair comparison due to the lack of a benchmark framework and a clear analysis of which aspects are common, respectively distinct, between different implementations. To this end, we introduce BenQ to pave the way towards fair comparisons in two separate race tracks, i.e., intra-comparison of the RL-based and the DNAS-based methods, respectively. We provide a systematic approach, which helps to reveal relatively vital aspects of different implementations. Finally, we conduct comprehensive experi-ments on VGG, AlexNet, ResNet, GoogleNet, MobileNet-V2, and Vision Transformer (ViT), and the new observations shed light on potential future directions for automated quantization to move forward.
Xingjun Zhang, Zeyu Ji, Jia Wei 0002
IEEE Trans. Computers5
2023 Leader population learning rate schedule
Jia Wei 0002, Xingjun Zhang, Zhimin Zhuo, Zeyu Ji, Qianyang Li
Inf. Sci.1
2023 Fastensor: Optimise the Tensor I/O Path from SSD to GPU for Deep Learning Training
abstract
In recent years, benefiting from the increase in model size and complexity, deep learning has achieved tremendous success in computer vision (CV) and (NLP). Training deep learning models using accelerators such as GPUs often requires much iterative data to be transferred from NVMe SSD to GPU memory. Much recent work has focused on data transfer during the pre-processing phase and has introduced techniques such as multiprocessing and GPU Direct Storage (GDS) to accelerate it. However, tensor data during training (such as Checkpoints, logs, and intermediate feature maps), which is also time-consuming, is often transferred using traditional serial, long-I/O-path transfer methods. In this article, based on GDS technology, we built Fastensor, an efficient tool for tensor data transfer between the NVMe SSDs and GPUs. To achieve higher tensor data I/O throughput, we optimized the traditional data I/O process. We also proposed a data and runtime context-aware tensor I/O algorithm. Fastensor can select the most suitable data transfer tool for the current tensor from a candidate set of tools during model training. The optimal tool is derived from a dictionary generated by our adaptive exploration algorithm in the first few training iterations. We used Fastensor’s unified interface to test the read/write bandwidth and energy consumption of different transfer tools for different sizes of tensor blocks. We found that the execution efficiency of different tensor transfer tools is related to both the tensor block size and the runtime context. We then deployed Fastensor in the widely applicable Pytorch deep learning framework. We showed that Fastensor could perform superior in typical scenarios of model parameter saving and intermediate feature map transfer with the same hardware configuration. Fastensor achieves a 5.37x read performance improvement compared to torch.save () when used for model parameter saving. When used for intermediate feature map transfer, Fastensor can increase the supported training batch size by 20x, while the total read and write speed is increased by 2.96x compared to the torch I/O API.
Jia Wei 0002, Xingjun Zhang
ACM Trans. Archit. Code Optim.1
2022 BenQ: Benchmarking Automated Quantization on Deep Neural Network Accelerators
abstract
Hardware-aware automated quantization promises to unlock an entirely new algorithm-hardware co-design paradigm for efficiently accelerating deep neural network (DNN) inference by incorporating the hardware cost into the reinforcement learning (RL) -based quantization strategy search process. Existing works usually design an automated quantization algorithm targeting one hardware accelerator with a device-specific performance model or pre-collected data. However, determining the hardware cost is non-trivial for algorithm experts due to their lack of cross-disciplinary knowledge in computer architecture, compiler, and physical chip design. Such a barrier limits reproducibility and fair comparison. Moreover, it is notoriously challenging to interpret the results due to the lack of quantitative metrics. To this end, we first propose BenQ, which includes various RL-based automated quantization algorithms with aligned settings and encapsulates two off-the-shelf performance predictors with standard OpenAI Gym API. Then, we leverage cosine similarity and manhattan distance to interpret the similarity between the searched policies. The experiments show that different automated quantization algorithms can achieve near equivalent optimal trade-offs because of the high similarity between the searched policies, which provides insights for revisiting the innovations in automated quantization algorithms.
Xingjun Zhang, Zeyu Ji, Jia Wei 0002
DATE5
2022 Status, challenges and trends of data-intensive supercomputing
Jia Wei 0002, Pei Ren, Yujia Lei, Yuqi Qu, Qiyu Jiang, Xiaoshe Dong, Weiguo Wu, Qiang Wang 0062, Xingjun Zhang
CCF Trans. High Perform. Comput.1
2022 GARLSched: Generative adversarial deep reinforcement learning task scheduling optimization for large-scale high performance computing systems
Xingjun Zhang, Jia Wei 0002, Zeyu Ji
Future Gener. Comput. Syst.3
2022 DPLRS: Distributed Population Learning Rate Schedule
Jia Wei 0002, Xingjun Zhang, Zeyu Ji
Future Gener. Comput. Syst.1
2022 EP4DDL: addressing straggler problem in heterogeneous distributed deep learning
Zeyu Ji, Xingjun Zhang, Jia Wei 0002
J. Supercomput.4
2021 Energy-aware task scheduling optimization with deep reinforcement learning for large-scale heterogeneous systems
Xingjun Zhang, Jia Wei 0002, Zeyu Ji
CCF Trans. High Perform. Comput.4
2021 A tile-fusion method for accelerating Winograd convolutions
Zeyu Ji, Xingjun Zhang, Jia Wei 0002
Neurocomputing5