Wenbin Jiang 0001

dblp:96/5583-1 · DBLP profile ↗
← Back
47ranked-venue papers
18as first author
15since 2021 · last 2026
0000-0001-5628-8806ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 8 first-author · 8 since 2021Computer networks · 7 · 3 first-authorDatabases, data management, data science and information retrieval · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 CPU-Oblivious Offloading of Failure-Atomic Transactions for Disaggregated Memory
abstract
Memory disaggregation introduces new challenges for application reliability, as compute server or interconnection failures can interrupt execution and lead to data inconsistency in the memory server. This paper presents Fanmem, a novel failure-atomic transaction system designed specifically for disaggregated memory architectures. Fanmem ensures data consistency in the presence of failures, drawing inspiration from persistent memory transactions while tailored for memory disaggregation. The key innovations of Fanmem include an asynchronous transaction model and the integration of a processing unit within the switch, enabling the offloading of time-consuming log persistency operations to the switch processing unit and significantly reducing the overhead on the compute servers. Evaluation confirms the effectiveness of Fanmem on two representative memory-disaggregated architectures. Compared to the state-of-the-art persistent memory transaction system, Fanmem achieves an average performance improvement of 1.2X and 1.7X on the respective architectures.
Chencheng Ye 0001, Yuanchao Xu 0001, Xipeng Shen, Xiaofei Liao, Hai Jin 0001, Wenbin Jiang 0001, Yan Solihin
ASPLOS (2)7
2026 DTMiner: A Data-Centric System for Efficient Temporal Motif Mining
abstract
Mining temporal motifs in temporal graphs is essential for many critical applications. Although several solutions have been proposed to handle temporal motif mining, they still suffer from substantial inefficiencies due to significant redundant graph traversals and fragmented memory access, both caused by irregular search tree expansions across different motif matching tasks. In this work, we observe that data accesses issued by these tasks exhibit strong spatial similarity and temporal monotonicity. Based on these observations, this paper proposes an efficient data-centric temporal motif mining system DTMiner, which introduces a novel Load-Explore-Synchronize (LES) execution model to efficiently regularize data accesses to the common temporal graph data among different tasks. Specifically, DTMiner enables the temporal graph chunks to be sequentially loaded into the cache in temporal order and then triggers all relevant tasks to explore only these loaded data for search tree expansions in a fine-grained synchronization mechanism. In this way, different tasks can share the graph traversal corresponding to the same chunks, while fragmented memory accesses are restricted to the graph data residing in the cache, significantly reducing data access overhead. Experimental results demonstrate that DTMiner achieves 1.14×-11.98× performance improvement in comparison with the state-of-the-art temporal motif mining solutions.
Yinbo Hou, Hao Qi 0004, Ligang He, Jin Zhao 0003, Yu Zhang 0027, Longlong Lin, Lin Gu 0002, Wenbin Jiang 0001, Xiaofei Liao, Hai Jin 0001
PPoPP9
2025 Dataflow-Guided Neuro-Symbolic Language Models for Type Inference
abstract
Language Models (LMs) are increasingly used for type inference, aiding in error detection and software development. Some real-world deployments of LMs require the model to run on local machines to safeguard the intellectual property of the source code. This setting often limits the size of the LMs that can be used. We present Nester, the first neuro-symbolic approach that enhances LMs for type inference by integrating symbolic learning without increasing model size. Nester breaks type inference into sub-tasks based on the data and control flow of the input code, encoding them as a modular high-level program. This program executes multi-step actions, such as evaluating expressions and analyzing conditional branches of the target code, combining static typing with LMs to infer potential types. Evaluated on the ManyTypes4Py dataset in Python, Nester outperforms two state-of-the-art type inference methods (HiTyper and TypeGen), achieving 70.7\% Top-1 Exact Match, which is 18.3\% and 3.6\% higher than HiTyper and TypeGen, respectively. For complex type annotations like typing.Optional and typing.Union, Nester achieves 51.0\% and 16.7\%, surpassing TypeGen by 28.3\% and 5.8\%.
Ge Li 0001, Yao Wan 0001, Hongyu Zhang 0002, Zhou Zhao 0001, Wenbin Jiang 0001, Xuanhua Shi, Hai Jin 0001, Zheng Wang 0001
ICML5
2025 BRP-SpMM: Block-Row Partition Based Sparse Matrix Multiplication with Tensor and CUDA Cores
abstract
Sparse-Dense Matrix Multiplication (SpMM) is a fundamental computational operation in various domains, and leveraging Tensor cores or CUDA cores on GPU to accelerate SpMM has become common practice. While Tensor cores offer notable advantages in dense matrix multiplication, their efficiency significantly decreases when handling sparse matrices. Besides, the computational power of CUDA cores should not be overlooked. Nevertheless, the differences in supported data formats present challenges in effectively leveraging both types of cores to accelerate SpMM. To this end, we propose BRP-SpMM, a block and row partition approach designed for efficient SpMM on GPU. BRP-Sp partitions the sparse matrix into two parts: TC Block and Residual Row part, which are computed on Tensor cores and CUDA cores separately. Meanwhile, a customized storage format is proposed to manage these two distinct parts. Our two GPU kernels incorporate several advanced techniques including load balance, register remapping, and 1-D tiling. BRP-SpMM can achieve higher memory access efficiency and more rational computing resource utilization of GPU. Extensive experiments on modern NVIDIA A800 GPU show that BRPSpMM outperforms SOTA libraries by up to$2.9 \times$(on average$2.1 \times)$. Furthermore, BRP-SpMM accelerates end-to-end GNN training by up to$1.9 \times$compared to popular frameworks.
Yukang Dong, Wenbin Jiang 0001, Xinhai Shen, Haihong Guo, Zhiyuan Shao, Hai Jin 0001
IPDPS2
2025 LaTCoder: Converting Webpage Design to Code with Layout-as-Thought
abstract
Converting webpage designs into code (design-to-code) plays a vital role in User Interface (UI) development for front-end developers, bridging the gap between visual design and functional implementation. While recent Multimodal Large Language Models (MLLMs) have shown significant potential in design-to-code tasks, they often fail to accurately preserve the layout during code generation. To this end, we draw inspiration from the Chain-of-Thought (CoT) reasoning in human cognition and propose LaTCoder, a novel approach that enhances layout preservation in webpage design during code generation with Layout-as-Thought (LaT). Specifically, we first introduce a simple yet efficient algorithm to divide the webpage design into image blocks. Next, we prompt MLLMs using a CoT-based approach to generate code for each block. Finally, we apply two assembly strategies-absolute positioning and an MLLM-based method-followed by dynamic selection to determine the optimal output. We evaluate the effectiveness of LaTCoder using multiple backbone MLLMs (i.e., DeepSeek-VL2, Gemini, and GPT-4o) on both a public benchmark and a newly introduced, more challenging benchmark (CC-HARD) that features complex layouts. The experimental results on automatic metrics demonstrate significant improvements. Specifically, TreeBLEU scores increased by 66.67% and MAE decreased by 38% when using DeepSeek-VL2, compared to direct prompting. Moreover, the human preference evaluation results indicate that annotators favor the webpages generated by LaTCoder in over 60% of cases, providing strong evidence of the effectiveness of our method.
Yi Gui, Zhen Li 0050, Guohao Wang, Tianpeng Lv, Gaoyang Jiang, Yi Liu 0069, Dongping Chen, Yao Wan 0001, Hongyu Zhang 0002, Wenbin Jiang 0001, Xuanhua Shi, Hai Jin 0001
KDD (2)11
2025 Bridging the Gap between Unstructured SpMM and Structured Sparse Tensor Cores
abstract
The acceleration of Sparse-dense Matrix Multiplication (SpMM) using Tensor Cores (TCs) in GPUs has recently garnered significant attention. TCs are designed for block-wise matrix multiplication, however, block partitioning of general unstructured sparse matrices often results in low-level density, causing a substantial waste of computational resources. Sparse Tensor Cores (SpTCs) can mitigate this issue by skipping 50% of zero values, however, SpTCs are limited to strict 2:4 or 1:2 structured sparsity. To bridge this gap, we propose MP-SpMM, a novel Matching and Padding approach that transforms general sparse matrices into structured sparsity, drawing inspiration from the maximum matching problem in graph theory. Moreover, we introduce a novel storage format and a highly optimized GPU kernel that fully exploits the capabilities of SpTCs. Extensive experiments on modern GPUs demonstrate that MP-SpMM outperforms state-of-the-art SpMM libraries, DTC-SpMM and RoDe, with an average speedup of 2.42 × (up to 7.65 ×) and 1.92 × (up to 8.60 ×).
Yukang Dong, Ziyuan Shen, Wenbin Jiang 0001, Zhenghang Liu, Bingyi He, Hai Jin 0001
SC3
2025 UICopilot: Automating UI Synthesis via Hierarchical Code Generation from Webpage Designs
abstract
Automating the synthesis of User Interfaces (UIs) plays a crucial role in enhancing productivity and accelerating the development lifecycle, reducing both development time and manual effort. Recently, the rapid development of Multimodal Large Language Models (MLLMs) has made it possible to generate front-end Hypertext Markup Language (HTML) code directly from webpage designs. However, real-world webpages encompass not only a diverse array of HTML tags but also complex stylesheets, resulting in significantly lengthy code. The lengthy code poses challenges for the performance and efficiency of MLLMs, especially in capturing the structural information of UI designs. To address these challenges, this paper proposes UICopilot, a novel approach to automating UI synthesis via hierarchical code generation from webpage designs. To validate the effectiveness of UICopilot, we conduct experiments on a real-world dataset, i.e., WebCode2M. Experimental results demonstrate that UICopilot significantly outperforms existing baselines in both automatic evaluation metrics and human evaluations. Specifically, statistical analysis reveals that the majority of human annotators prefer the webpages generated by UICopilot over those produced by GPT-4V.
Yi Gui, Yao Wan 0001, Zhen Li 0050, Dongping Chen, Hongyu Zhang 0002, Bohua Chen, Wenbin Jiang 0001, Xiangliang Zhang 0001
WWW10
2025 WebCode2M: A Real-World Dataset for Code Generation from Webpage Designs
abstract
Automatically generating webpage code from webpage designs can significantly reduce the workload of front-end developers, and recent Multimodal Large Language Models (MLLMs) have shown promising potential in this area. However, our investigation reveals that most existing MLLMs are constrained by the absence of high-quality, large-scale, real-world datasets, resulting in inadequate performance in automated webpage code generation. To fill this gap, this paper introduces WebCode2M, a new dataset comprising 2.56 million instances, each containing a design image along with the corresponding webpage code and layout details. Sourced from real-world web resources, WebCode2M offers a rich and valuable dataset for webpage code generation across a variety of applications. The dataset quality is ensured by a scoring model that filters out instances with aesthetic deficiencies or other incomplete elements. To validate the effectiveness of WebCode2M, we introduce a baseline model based on the Vision Transformer (ViT), named WebCoder, and establish a benchmark for fair comparison. Additionally, we introduce a new metric, TreeBLEU, to measure the structural hierarchy recall. The benchmarking results demonstrate that our dataset significantly improves the ability of MLLMs to generate code from webpage designs, confirming its effectiveness and usability for future applications in front-end design tools. Finally, we highlight several practical challenges introduced by our dataset, calling for further research. The code and dataset are publicly available at our project homepage: https://webcode2m.github.io.
Yi Gui, Zhen Li 0050, Yao Wan 0001, Yemin Shi 0001, Hongyu Zhang 0002, Bohua Chen, Dongping Chen, Siyuan Wu 0001, Wenbin Jiang 0001, Hai Jin 0001, Xiangliang Zhang 0001
WWW11
2025 RT-GNN: Accelerating Sparse Graph Neural Networks by Tensor-CUDA Kernel Fusion
abstract
Graph Neural Networks (GNNs) have achieved remarkable successes in various graph-based learning tasks, thanks to their ability to leverage advanced GPUs. However, GNNs currently face challenges arising from the concurrent use of advanced Tensor Cores (TCs) and CUDA Cores (CDs) in GPUs. These challenges are further exacerbated due to repeated, inefficient, and redundant aggregations in GNN that result from the high sparsity and irregular non-zero distribution of real-world graphs. We propose RT-GNN, a GNN framework based on the fusion of advanced TC and CD units, to eliminate the aforementioned redundancies by exploiting the properties of an adjacency matrix. First, a novel GNN representation technique, hierarchical embedding graph (HEG) is proposed to manage the intermediate aggregation results hierarchically, which can further avoid redundancy in intermediate aggregations elegantly. Next, to address the inherent sparsity of graphs, RT-GNN places the blocks (a.k.a. tiles) in HEG onto TCs and CDs according to their sparsity by a new block-based row-wise multiplication approach, which assembles TCs and CDs to work concurrently. Experimental results demonstrate that HEG outperforms HAG by an average speedup of 19.3× for redundancy elimination performance, especially up to 72× speedup on the dataset of ARXIV. Moreover, for overall performance, RT-GNN outperforms state-of-the-art GNN frameworks (including DGL, HAG, GNNAdvisor, and TC-GNN) by an average factor of 3.1× while maintaining or even improving the task accuracy.
Jianrong Yan, Wenbin Jiang 0001, Dongao He, Suyang Wen, Hai Jin 0001, Zhiyuan Shao
ACM Trans. Archit. Code Optim.2
2025 Comprehensive Architecture Search for Deep Graph Neural Networks
abstract
In recent years, Neural Architecture Search (NAS) has emerged as a promising approach for automatically discovering superior model architectures for deep Graph Neural Networks (GNNs). Different methods have paid attention to different types of search spaces. However, due to the time-consuming nature of training deep GNNs, existing NAS methods often fail to explore diverse search spaces sufficiently, which constrains their effectiveness. To crack this hard nut, we propose CAS-DGNN, a novelcomprehensivearchitecturesearch method fordeepGNNs. It encompasses four kinds of search spaces that are the composition of aggregate and update operators, different types of aggregate operators, residual connections, and hyper-parameters. To meet the needs of such a complex situation, a phased and hybrid search strategy is proposed to accommodate the diverse characteristics of different search spaces. Specifically, we divide the search process into four phases, utilizing evolutionary algorithms and Bayesian optimization. Meanwhile, we design two distinct search methods for residual connections (All-connected search and Initial Residual search) to streamline the search space, which enhances the scalability of CAS-DGNN. The experimental results show that CAS-DGNN achieves higher accuracy with competitive search costs across ten public datasets compared to existing methods.
Yukang Dong, Fanxing Pan, Yi Gui, Wenbin Jiang 0001, Yao Wan 0001, Hai Jin 0001
IEEE Trans. Big Data4
2024 Accelerating Large-Scale GNN by Combining k-Hop Neighbors Based Feature Caching and Hierarchical GPU-Centric Data Access
Wenbin Jiang 0001, Fanxing Pan, Xinhai Shen, Dongao He, Hai Jin 0001
GPC1
2023 PDAS: Improving network pruning based on Progressive Differentiable Architecture Search for DNNs
Wenbin Jiang 0001, Suyang Wen, Long Zheng 0003, Hai Jin 0001
Future Gener. Comput. Syst.1
2023 Accelerating Graph Convolutional Networks Through a PIM-Accelerated Approach
abstract
Graph convolutional networks(GCNs) are promising to enable machine learning on graph data. GCNs show potential vertex-level and intra-vertex parallelism for GPU acceleration, but their irregular memory accesses arising in aggregation operations and the inherent sparsity for vertex features of graphs cause inefficiencies on the GPU. In this paper, we present gPIM, which aims to accelerate GCNs inference through aprocessing-in-memory(PIM) enabled architecture. gPIM is expected to perform compute-intensive combination on the GPU while aggregation and memory-bound combination are offloaded to the PIM-featuredhybrid memory cubes(HMCs). To maximize the efficiency of such GPU-HMC architecture, gPIM is novel with two key designs: 1) A GCN-induced graph partitioning that minimizes communication overheads between cubes, 2) A programmer-transparent performance estimation mechanism that predicts the performance bound of operations accurately for workload offloading. Experimental results show that gPIM significantly outperforms Intel Xeon E5-2680v3 CPU (8,979.52×), NVIDIA Tesla V100 GPU (96.01×), and a state-of-the-art GCN accelerator AWB-GCN (4.18×).
Hai Jin 0001, Dan Chen 0006, Long Zheng 0003, Yu Huang 0013, Pengcheng Yao, Jin Zhao 0003, Xiaofei Liao, Wenbin Jiang 0001
IEEE Trans. Computers8
2021 GraSU: A Fast Graph Update Library for FPGA-based Dynamic Graph Processing
abstract
Existing FPGA-based graph accelerators, typically designed for static graphs, rarely handle dynamic graphs that often involve substantial graph updates (e.g., edge/node insertion and deletion) over time. In this paper, we aim to fill this gap. The key innovation of this work is to build an FPGA-based dynamic graph accelerator easily from any off-the-shelf static graph accelerator with minimal hardware engineering efforts (rather than from scratch). We observe \em spatial similarity of dynamic graph updates in the sense that most of graph updates get involved with only a small fraction of vertices. We therefore propose an FPGA library, called GraSU, to exploit spatial similarity for fast graph updates. GraSU uses a differential data management, which retains the high-value data (that will be frequently accessed) in the specialized on-chip UltraRAM while the overwhelming majority of low-value ones reside in the off-chip memory. Thus, GraSU can transform most of off-chip communications arising in dynamic graph updates into fast on-chip memory accesses. Our experiences show that GraSU can be easily integrated into existing state-of-the-art static graph accelerators with only 11 lines of code modifications. Our implementation atop AccuGraph using a Xilinx Alveo#8482; \ U250 board outperforms two state-of-the-art CPU-based dynamic graph systems, Stinger and Aspen, by an average of 34.24× and 4.42× in terms of update throughput, improving further overall efficiency by 9.80× and 3.07× on average.
Qinggang Wang, Long Zheng 0003, Yu Huang 0013, Pengcheng Yao, Chuangyi Gui, Xiaofei Liao, Hai Jin 0001, Wenbin Jiang 0001, Fubing Mao
FPGA8
2021 FDGLib: A Communication Library for Efficient Large-Scale Graph Processing in FPGA-Accelerated Data Centers
Qinggang Wang, Long Zheng 0003, Xiaofei Liao, Hai Jin 0001, Wenbin Jiang 0001, Kan Hu
J. Comput. Sci. Technol.6
2020 An Efficient Data Prefetch Strategy for Deep Learning Based on Non-volatile Memory
Wenbin Jiang 0001, Pai Liu, Hai Jin 0001
GPC1
2020 GradSA: Gradient Sparsification and Accumulation for Communication-Efficient Distributed Deep Learning
Bo Liu 0057, Wenbin Jiang 0001, Shaofeng Zhao, Hai Jin 0001, Bingsheng He
GPC2
2020 Exploiting potential of deep neural networks by layer-wise fine-grained parallelism
Wenbin Jiang 0001, Yangsong Zhang 0001, Pai Liu, Laurence T. Yang, Geyan Ye, Hai Jin 0001
Future Gener. Comput. Syst.1
2020 Layup: Layer-adaptive and Multi-type Intermediate-oriented Memory Optimization for GPU-based CNNs
abstract
Although GPUs have emerged as the mainstream for the acceleration of convolutional neural network (CNN) training processes, they usually have limited physical memory, meaning that it is hard to train large-scale CNN models. Many methods for memory optimization have been proposed to decrease the memory consumption of CNNs and to mitigate the increasing scale of these networks; however, this optimization comes at the cost of an obvious drop in time performance. We propose a new memory optimization strategy named Layup that realizes both better memory efficiency and better time performance. First, a fast layer-type-specific method for memory optimization is presented, based on the new finding that a single memory optimization often shows dramatic differences in time performance for different types of layers. Second, a new memory reuse method is presented in which greater attention is paid to multi-type intermediate data such as convolutional workspaces and cuDNN handle data. Experiments show that Layup can significantly increase the scale of extra-deep network models on a single GPU with lower performance loss. It even can train ResNet with 2,504 layers using 12GB memory, outperforming the state-of-the-art work of SuperNeurons with 1,920 layers (batch size = 16).
Wenbin Jiang 0001, Bo Liu 0057, Haikun Liu, Bing Bing Zhou, Song Wu 0001, Hai Jin 0001
ACM Trans. Archit. Code Optim.1
2019 A Novel Stochastic Gradient Descent Algorithm Based on Grouping over Heterogeneous Cluster Systems for Distributed Deep Learning
abstract
On heterogeneous cluster systems, the convergence performances of neural network models are greatly troubled by the different performances of machines. In this paper, we propose a novel distributed Stochastic Gradient Descent (SGD) algorithm named Grouping-SGD for distributed deep learning, which converges faster than Sync-SGD, Async-SGD, and Stale-SGD. In Grouping-SGD, machines are partitioned into multiple groups, ensuring that machines in the same group have similar performances. Machines in the same group update the models synchronously, while different groups update the models asynchronously. To improve the performance of Grouping-SGD further, the parameter servers are arranged from fast to slow, and they are responsible for updating the model parameters from the lower layer to the higher layer respectively. The experimental results indicate that Grouping-SGD can achieve 1.2~3.7 times speedups using popular image classification benchmarks: MNIST, Cifar10, Cifar100, and ImageNet, compared to Sync-SGD, Async-SGD, and Stale-SGD.
Wenbin Jiang 0001, Geyan Ye, Laurence T. Yang, Hai Jin 0001
CCGRID1
2019 HiNUMA: NUMA-Aware Data Placement and Migration in Hybrid Memory Systems
abstract
Non-uniform memory access (NUMA) architectures feature asymmetrical memory access latencies on different CPU nodes. Hybrid memory systems composed of non-volatile memory (NVM) and DRAM further diversify memory access latencies due to the relatively large performance gap between NVM and DRAM. Traditional NUMA memory management policies fail to manage hybrid memories effectively and may even hurt application performance. In this paper, we present HiNUMA, a new NUMA abstraction for memory allocation and migration in hybrid memory systems. HiNUMA advocates NUMA topologyaware hybrid memory allocation policies for the initial data placement. HiNUMA also proposes a new NUMA balancing mechanism called HANB for memory migration at runtime. HANB considers both data access frequency and memory bandwidth utilization to reduce the cost of memory accesses in hybrid memory systems. We evaluate the performance of HiNUMA with several typical workloads. Experimental results show that HiNUMA can effectively utilize hybrid memories, and deliver much higher application performance than conventional NUMA memory management policies and other state-of-the-art work.
Zhuohui Duan, Haikun Liu, Xiaofei Liao, Hai Jin 0001, Wenbin Jiang 0001, Yu Zhang 0027
ICCD5
2018 FiLayer: A Novel Fine-Grained Layer-Wise Parallelism Strategy for Deep Neural Networks
Wenbin Jiang 0001, Yangsong Zhang 0001, Pai Liu, Geyan Ye, Hai Jin 0001
ICANN (3)1
2018 Layrub: layer-centric GPU memory reuse and data migration in extreme-scale deep learning systems
abstract
Growing accuracy and robustness of Deep Neural Networks (DNN) models are accompanied by growing model capacity (going deeper or wider). However, high memory requirements of those models make it difficult to execute the training process in one GPU. To address it, we first identify the memory usage characteristics for deep and wide convolutional networks, and demonstrate the opportunities of memory reuse on both intra-layer and inter-layer levels. We then present Layrub, a runtime data placement strategy that orchestrates the execution of training process. It achieves layer-centric reuse to reduce memory consumption for extreme-scale deep learning that cannot be run on one single GPU.
Bo Liu 0057, Wenbin Jiang 0001, Hai Jin 0001, Xuanhua Shi
PPoPP2
2018 FIPIP: A novel fine-grained parallel partition based intra-frame prediction on heterogeneous many-core systems
Wenbin Jiang 0001, Laurence T. Yang, Xiaobai Liu, Hai Jin 0001, Alan L. Yuille, Ye Chi
Future Gener. Comput. Syst.1
2018 Layer-Centric Memory Reuse and Data Migration for Extreme-Scale Deep Learning on Many-Core Architectures
abstract
Due to the popularity of Deep Neural Network (DNN) models, we have witnessed extreme-scale DNN models with the continued increase of the scale in terms of depth and width. However, the extremely high memory requirements for them make it difficult to run the training processes on single many-core architectures such as a Graphic Processing Unit (GPU), which compels researchers to use model parallelism over multiple GPUs to make it work. However, model parallelism always brings very heavy additional overhead. Therefore, running an extreme-scale model in a single GPU is urgently required. There still exist several challenges to reduce the memory footprint for extreme-scale deep learning. To address this tough problem, we first identify the memory usage characteristics for deep and wide convolutional networks, and demonstrate the opportunities for memory reuse at both the intra-layer and inter-layer levels. We then present Layrub, a runtime data placement strategy that orchestrates the execution of the training process. It achieves layer-centric reuse to reduce memory consumption for extreme-scale deep learning that could not previously be run on a single GPU. Experiments show that, compared to the original Caffe, Layrub can cut down the memory usage rate by an average of 58.2% and by up to 98.9%, at the moderate cost of 24.1% higher training execution time on average. Results also show that Layrub outperforms some popular deep learning systems such as GeePS, vDNN, MXNet, and Tensorflow. More importantly, Layrub can tackle extreme-scale deep learning tasks. For example, it makes an extra-deep ResNet with 1,517 layers that can be trained successfully in one GPU with 12GB memory, while other existing deep learning systems cannot.
Hai Jin 0001, Bo Liu 0057, Wenbin Jiang 0001, Xuanhua Shi, Bingsheng He, Shaofeng Zhao
ACM Trans. Archit. Code Optim.3
2017 Detail enhancement of image super-resolution based on detail synthesis
Jinsheng Xiao, Enyu Liu, Yuan-Fang Wang, Wenbin Jiang 0001
Signal Process. Image Commun.5
2016 A Fine-Grained Parallel Intra Prediction for HEVC Based on GPU
abstract
Intra prediction in HEVC is much more complex compared to the one in H.264 because of the more diversifications of the block sizes and prediction modes. The state-of-the-art researches for its parallelization only focus on block-level methods, which only take very limited advantage of GPUs. It is still a big challenge to implement fine-grained parallelism on GPU in consideration of the HEVC branch instructions and the different prediction formulae. We present a novel pixel-level parallelism method for the intra prediction of HEVC based on GPU combined with mode-level parallelism. By unifying not only the prediction formulae between angular mode and planar mode but also a predictor array, an algorithm based on look-up table is proposed to greatly reduce branches and improve prediction efficiency. With the help of look-up table algorithm, each pixel in a block can obtain the offset of corresponding reference pixels and find the value in the unifying predictor array at the same time which makes it possible to predict all pixels in parallel regardless of their relative positions in the block. The experimental results show that the proposed algorithm outperforms previous work and can reduce encoding time effectively.
Wenbin Jiang 0001, Ye Chi, Hai Jin 0001, Xiaofei Liao, Yangsong Zhang 0001, Geyan Ye
ICPADS1
2016 A novel parallelized motion estimation algorithm for GPU based video encoding
abstract
Cloud-based video encoding has become more and more popular in Internet, especially for mobile clients, considering their limited resources. Recently, GPUs (Graphics Processor Units) make the cloud-based video encoding more economic and efficient. However, the motion estimation in inter prediction, which usually occupies about 70% encoding time in H.264/AVC, is still a big headache because of its complexity. In this paper, a novel motion estimation algorithm is proposed, which is customized for GPU-based cloud encoding, considering motion tendency. A mean subsampling template is presented for a pre-search approach to get motion tendency, which can reduce computation cost obviously with less quality loss. To improve the efficiency of the CUDA (Compute Unified Device Architecture) thread organization for motion estimation, a section-division method is presented. Experimental results show that the proposed algorithm can reduce nearly 22% computation time with less video quality loss, compared with the state-of-the-art work.
Wenbin Jiang 0001, Hai Jin 0001
WoWMoM1
2016 A novel parallel deblocking filtering strategy for HEVC/H.265 based on GPU
abstract
Summary The deblocking filter inhigh‐efficiency video coding(HEVC) has huge computational complexity because of its high content‐adaptive coding structure as well as high‐definition. Parallelization for it based on massively parallel architectures such asgraphics processing unitbecomes an urgent demand. However, a large number of conditional branches and data dependencies severely hinder its efficient parallelization. In this paper, a novel parallel optimization strategy based on graphics processing unit is presented for concurrent deblocking in HEVC/H.265 standard to improve the parallel performance. First, by reducing various conditional branches, a normalization mechanism for instruction stream based on feature vector is proposed, which improves the efficiency of boundary strength computation dramatically. The idea can also be applied to edge discrimination. Second, a parallel mechanism based on an adaptive post‐correction is presented to process vertical and horizontal edges filtering concurrently, which improves the processing speed obviously, while producing negligible quality loss. Experimental results show that the strategy presented outperforms the existing state‐of‐the‐art method with accelerating factor up to 32. Copyright © 2016 John Wiley & Sons, Ltd.
Wenbin Jiang 0001, Hongyan Mei, Feng Lu 0003, Hai Jin 0001, Laurence T. Yang, Bin Luo 0006, Ye Chi
Concurr. Comput. Pract. Exp.1
2015 Parallel Top-k Query Processing on Uncertain Strings Using MapReduce
Xiaofeng Ding 0001, Hai Jin 0001, Wenbin Jiang 0001
DASFAA (2)4
2015 A New Approach for Vehicle Recognition and Tracking in Multi-camera Traffic System
Wenbin Jiang 0001, Hai Jin 0001, Ye Chi
ICA3PP (3)1
2015 Optimization strategies for inter-thread synchronization overhead on NUMA machine
abstract
Overhead caused by data consistence issue in inter-thread synchronization probably degrades the performance of parallel applications. Non-Uniform Memory Access (NUMA), as the mainstream architecture in today's multicore processor, further exacerbates this issue due to the significant overhead incurred by Remote Memory Reference (RMR). Therefore, to reduce synchronization overhead, it is important to solve the data consistence issue. In this paper, we classify the overhead into two kinds: (1) overhead incurred by algorithms themselves, and (2) overhead incurred by critical sections. To reduce two kinds of overhead on NUMA machine, we present two optimization strategies called search and backtrace (SAB) and reorder critical section and non-critical section (RCAN), respectively. In SAB, a server thread tries to search a thread coming from master NUMA node, and designates it as the new server thread. In this way, most of the time, shared data resides in the cache of master NUMA node, resulting in lower overhead caused by data consistence issue in critical section. In RCAN, each thread consecutively posts synchronization requests, followed by consecutively executing non-critical section. In this way, server threads could serve enough requests, resulting in better data locality. We design an algorithm named R-Synch based on SAB, while designing an algorithm named H-STA based on RCAN. Our evaluation with representative synchronization algorithms demonstrates the effectiveness of R-Synch and H-STA.
Song Wu 0001, Yaqiong Peng, Hai Jin 0001, Wenbin Jiang 0001
IPCCC5
2014 Fine-grained CUDA-based Parallel Intra Prediction for H.264/AVC
abstract
Recently, the power of the Graphics Processing Unit (GPU) has largely increased, whereas previous works of intra prediction on the GPU could not efficiently exploit the massive parallel opportunity. The related work only achieves frame-level, slice-level or block-level parallelism. It is a challenge to implement fine-grained parallelism on the Compute Unified Device Architecture (CUDA), such as pixel-level and mode-level, because the irregular formulas of intra prediction and the constraints posed by H.264/AVC cause significant branch instructions and the CUDA architecture is inherently not good at handling branches. In this paper, a CUDA-based approach that adopts fine-grained parallelism is presented. By transforming the various prediction formulas to the same form and introducing the predictor unit, an algorithm based on a lookup table is proposed to efficiently eliminate the branches. In addition, the combinatorial frame technique and the optimized encoding order are adopted to maximize the parallelism. Experimental results show that significant encoding time reduction can be achieved and the proposed algorithm outperforms previous works.
Wenbin Jiang 0001, Hai Jin 0001
NOSSDAV1
2014 Memshepherd: comprehensive memory bug fault-tolerance system
abstract
Abstract Among all software vulnerabilities, memory bugs are most common and dangerous. Programs written in unsafe languages such as C and C++ are vulnerable to stack‐based buffer overflow, heap buffer overflow, dangling pointer, and double free. Although there are a number of proposed solutions to tolerate heap related bugs, most of the existing solutions terminates the vulnerable program after a stack‐based buffer overflow attempt. There is no comprehensive solution to actively tolerate all of the four kinds of bugs mentioned previously currently. This paper presents Memshepherd, a system that can probabilistically prevent software from both stack and heap memory bugs and guarantee soundness of the software execution. It dynamically reallocates stack‐based buffers in the heap space during software execution, thus transforms a stack memory problem into a heap memory problem. By adaptively sizing buffers to be M times of their defined size and randomly placing them, Memshepherd keeps the buffers far from each other. When a buffer is to be deallocated, Memshepherd checks invalid and double frees. A Linux prototype is implemented and tested against four kinds of memory bugs. The experiment results prove that Memshepherd is effective in eliminating crashes, erroneous execution, as well as security vulnerability. Copyright © 2013 John Wiley & Sons, Ltd.
Deqing Zou, Weide Zheng, Wenbin Jiang 0001, Hai Jin 0001
Secur. Commun. Networks3
2013 Capacity and delay of heterogeneous wireless networks with correlated mobility
abstract
Although the capacity of wireless ad hoc networks has been extensively studied under different mobility models and network settings, few work has been done on the effect of heterogeneous mobile nodes and correlated mobility. In this paper, we consider the heterogeneous wireless networks consisting of two types of nodes, called user nodes and master nodes. Specifically, user nodes are combined into groups and each group is equipped with a more powerful master node serving as relay for packet transmissions among groups. By proposing a simple, asymptotically optimal scheduling and routing scheme, we present the maximum per-node throughput and the end-to-end delay in order sense, respectively. We also explore the trade-offs between capacity and delay by adjusting network settings.
Yanzhi Tao, Xiaoliang Wang 0001, Sanglu Lu, Wenbin Jiang 0001
WCNC5
2012 A Novel Task Management System for Modelica-Based Multi-discipline Virtual Experiment Platform
abstract
Currently, there is few uniform modelling standards for virtual experiment (VE) systems of different disciplines. The scalability and compatibility of existing systems are relatively poor. The idea of Modelica provides a good opportunity for the unification of the modelling of multi-discipline VEs (MDVE). However, Modelica is an original multi-domain modelling method for scientific research, instead of for VE education. There are some obvious gaps to bring it into a MDVE platform (MDVEP), especially, if it is designed to support massive users and parallel modelling and resolving. This paper presents a new virtual experiment distributed task management system (VETMS) for MDVEP, which can improve the efficiency, stability and availability of the platform. It uses hierarchical design method to decouple the different modules and also provides a set of Application Programming Interfaces (APIs) for external calls. Besides, the system performance is also considered. A Modelica-oriented mechanism is proposed to tackle service failures. A parallelized Twisted framework is presented to overcome the problem of limitation of concurrent requests. Meanwhile, a NAT (Network Address Translator) traversal module of TCP based STUNT protocol is added to reduce the amount of data through master node. Experiment results show that the system can serve as a task management service for MDVEP with good performance.
Wenbin Jiang 0001, Shuguang Wang, Hai Jin 0001
APSCC1
2011 Special issue on information dissemination and new services in P2P systems
Min Song 0002, Sachin Shetty, Wenbin Jiang 0001, E. K. Park
Peer-to-Peer Netw. Appl.3
2011 Integrated buffering schemes for P2P VoD services
Linchen Yu, Xiaofei Liao, Hai Jin 0001, Wenbin Jiang 0001
Peer-to-Peer Netw. Appl.4
2011 Adaptive Object Tracking by Learning Hybrid Template Online
abstract
This paper presents an adaptive tracking algorithm by learning hybrid object templates online in video. The templates consist of multiple types of features, each of which describes one specific appearance structure, such as flatness, texture, or edge/corner. Our proposed solution consists of three aspects. First, in order to make the features of different types comparable with each other, a unified statistical measure is defined to select the most informative features to construct the hybrid template. Second, we propose a simple yet powerful generative model for representing objects. This model is characterized by its simplicity since it could be efficiently learnt from the currently observed frames. Last, we present an iterative procedure to learn the object template from the currently observed frames, and to locate every feature of the object template within the observed frames. The former step is referred to as feature pursuit, and the latter step is referred to as feature alignment, both of which are performed over a batch of observations. We fuse the results of feature alignment to locate objects within frames. The proposed solution to object tracking is in essence robust against various challenges, including background clutters, low-resolution, scale changes, and severe occlusions. Extensive experiments are conducted over several publicly available databases and the results with comparisons show that our tracking algorithm clearly outperforms the state-of-the-art methods.
Xiaobai Liu, Liang Lin 0004, Shuicheng Yan, Hai Jin 0001, Wenbin Jiang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2009 Load Balancing Routing Algorithm for Ad Hoc Networks
abstract
In ad hoc networks, when the load is heavy, the performance of on-demand routing protocols, such as DSR protocol, will suffer large degradation. For instance, the absence of balanced distribution of network flow among nodes in the whole network is highly prone to producing heavy-load key nodes. This results in fast energy consumption of these key nodes, and finally decreasing the survival time of the network caused by their premature failures. To address this problem, a novel algorithm called cross-layer based load balancing algorithm for DSR (CLB-DSR) is proposed. In CLB-DSR, the cross-layer strategy is adopted to exchange the information between data link layer and network layer. Simulation results show that the improved routing protocol using CLB-DSR can effectively balance network load, improve the network throughput and reduce the end-to-end delay, the rate of data dropped.
Wenbin Jiang 0001, Zhaojing Li, Chunqiang Zeng, Hai Jin 0001
MSN1
2009 VRFPS: A Novel Virtual Machine-Based Real-time File Protection System
abstract
With the development of virtualization technology, file protection in virtual machine, especially in guest OS, becomes more and more important. Traditional host-based file protection system resides the critical modules in monitored system, which is easily explored and destroyed by malwares. Moreover, in order to protect the multiple operation systems running on the same platform, it is necessary to install independent file protection system for each of them, which greatly wastes computing resources and brings serious performance overhead. In this paper, a novel VM-based real-time file protection system, named VRFPS, is proposed to solve these problems. First, virtual machine monitor introspects all file operations of guest OS. Then, semantic gap between disk block and logic files is narrowed by blktap. Finally, a virtual sandbox is implemented in privileged domain to prevent protected files in guest domain from modifying illegally. Our approach is highly isolated, transparent and without modification on virtual machine monitor and guest OS. The experimental results show that the presented system is validate and of low performance overhead.
Feng Zhao 0003, Guofu Xiang, Hai Jin 0001, Wenbin Jiang 0001
SERA5
2008 A New Proxy Scheme for Large-Scale P2P VoD System
abstract
Large-scale stream media delivery over Internet is one hot and hard problem. Three main solutions have been developed for this purpose. Content delivery network (CDN) can provide high quality streaming service with high cost. Server-based proxies are cost-effective but not scalable due to the limited proxy capacity and its centralized control. Peer-to-peer (P2P) network is scalable but does not guarantee high quality streaming service due to its self-restrained characteristic. This paper proposes a new chunk-based scalable and reliable proxy scheme for P2P network. Clients are self-organized in unstructured P2P VoD system GridCast and requests for media contents that are not found in P2P overlay from proxy servers. This scheme can address the limitations of server-based proxy model and client-based P2P model. The proposed scheme is evaluated by trace-driven simulations from logs collected in GridCast. The results show that the approach proposed significantly improves the quality of media streaming and the system scalability.
Wenbin Jiang 0001, Hai Jin 0001, Xiaofei Liao
EUC (1)1
2008 ER-TCP: an efficient TCP fault-tolerance scheme for cluster computing
Zhiyuan Shao, Hai Jin 0001, Bin Cheng 0001, Wenbin Jiang 0001
J. Supercomput.4
2006 FreeSpeech: A Novel Wireless Approach for Conference Projecting and Cooperating
Wenbin Jiang 0001, Hai Jin 0001, Zhiyuan Shao, Qiwei Ye
UIC1
2005 ER-TCP: An Efficient Fault-Tolerance Scheme for TCP Connections
Zhiyuan Shao, Hai Jin 0001, Bin Cheng 0001, Wenbin Jiang 0001
ISPA4
2005 TCP-ABC: From Multiple TCP Connections to Atomic Broadcasting
Zhiyuan Shao, Hai Jin 0001, Wenbin Jiang 0001, Bin Cheng 0001
NPC3
2004 A fast BMA based on combining search candidate subsampling and APDS
abstract
A new faster block-matching algorithm (BMA) named SSC-APDS is presented by subsampling search candidates in adjustable partial distortion search (APDS). Firstly APDS is modified to visit about half points of all search candidates by taking subsampling on them, using a spiral-scanning path with one skip. Two selected candidates that have minimal and second minimal block distortion measures are obtained Then a fine-tune step is taken around them to find the best one, while at most 8 more candidates were required. Experimental results show that the SSC-APDS can maintain its MSE performance very close to that of the APDS with higher speedup ratio. Moreover, the wider the search window is, the better SSC-APDS performs
Wenbin Jiang 0001, Manli Zhou
ICME1