Rongchun Li

dblp:23/11277 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
9since 2021 · last 2025
0000-0001-5922-7961ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 1 since 2021Systems, architecture and hardware · 8 · 3 first-author · 3 since 2021Computer networks · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Highly Parallelized Reinforcement Learning Training with Relaxed Assignment Dependencies
abstract
As the demands for superior agents grow, the training complexity of Deep Reinforcement Learning (DRL) becomes higher. Thus, accelerating training of DRL has become a major research focus. Dividing the DRL training process into sub-tasks and using parallel computation can effectively reduce training costs. However, current DRL training systems lack sufficient parallelization due to data assignment between sub-task components. This assignment issue has been ignored, but addressing it can further boost training efficiency. Therefore, we propose a high-throughput distributed RL training system called TianJi. It relaxes assignment dependencies between sub-task components and enables event-driven asynchronous communication. Meanwhile, TianJi maintains clear boundaries between sub-task components. To address convergence uncertainty from relaxed assignment dependencies, TianJi proposes a distributed strategy based on the balance of sample production and consumption. The strategy controls the staleness of samples to correct their quality, ensuring convergence. We conducted extensive experiments. TianJi achieves a convergence time acceleration ratio of up to 4.37 compared to related comparison frameworks. When scaled to eight computational nodes, TianJi shows a convergence time speedup of 1.6 and a throughput speedup of 7.13 relative to XingTian, emonstrating its capability to accelerate training and scalability. In data transmission efficiency experiments, TianJi significantly outperforms other frameworks, approaching hardware limits. TianJi also shows effectiveness in on-policy algorithms, achieving convergence time acceleration ratios of 4.36 and 2.95 compared to RLlib and XingTian.
Zhouyu He, Peng Qiao, Rongchun Li, Yong Dou, Yusong Tan
AAAI3
2025 Beyond Synthetic Data: Leveraging Natural Image Pretraining and Finetuning for Mixed Exposure Correction in Capsule Endoscopy
abstract
Wireless capsule endoscopy (WCE) has become standard in gastrointestinal diagnostics, but its frames often exhibit mixed exposure problem with both over- and underexposed regions. Existing methods mostly use pixel-level supervision in RGB space and introduce unnatural color shifts because exposure, color, and texture are tightly entangled. Another challenge is the lack of reliable paired WCE training data. Available synthetic datasets are small, low in quality, and lack diversity, while real paired examples are hard to obtain. Fully supervised models trained on such data tend to overfit and generalize poorly. We first extract exposure priors from a model pretrained on large natural-image datasets and then fine-tune it on WCE data. Leveraging hue stability and pixel dispersion nature in HSV space, our separated exposure and color correction method comprises three significant components: the training-free Single-Image Exposure Fusion (SIEF) module for precise exposure partitioning, the Sequential Interaction Correction (SIC) module for efficient information exchange between the exposure and color correction branches, and the Phase-Shifting Coder (PSC) to resolve discontinuities in the circular H channel and ensure smooth, stable hue prediction. Extensive experiments demonstrate that our method achieves SOTA results on natural-image exposure correction and transfers effectively to WCE mixed-exposure cases. It outperforms prior supervised approaches and shows robust recovery across multiple WCE datasets.
Tongrui Hu, Peng Qiao, Keran Ding, Yong Dou, Rongchun Li
BIBM5
2025 Only One Stage: A Chemical-Aware Model for Accurate Combustion Chemical Kinetics Prediction
abstract
The combustion chemical kinetics simulation, which focuses on the change in species mass fractions during reactions, is vital for clean energy development. Due to the sample complexities introduced by chemical kinetics, current multi-stage methods employ data preprocessing stage to separate it into subspaces, aiming to ease training. However, the current approaches to sample space separation are not effective, which not only affects the training accuracy of the network but also introduces a complex preprocessing procedure. To solve this, we propose a one-stage, end-to-end model with chemical-aware capabilities, using an auto dividing mechanism and spatio-temporal convolution for feature extraction. In hydrogen simulation experiments, our model achieved an L1 error of 10−6level, which is nearly identical to the standard results of numerical calculations, outperforming the usually used Multilayer Perceptron (MLP) method by 41.9 times.
Zhenglun Sun, Peng Qiao, Yong Dou, Rongchun Li, Sidun Liu
ICME4
2024 End-To-End High-Quality Transformer Object Detection Model Applied to Human Head Detection
Rongchun Li, Peng Qiao, Jingfei Jiang
PRCV (12)1
2024 High performance dilated convolutions on multi-core DSPs
Xiangdong Pei, Songzhu Mei, Rongchun Li, Jie Liu 0002
CCF Trans. High Perform. Comput.5
2022 MLPs: Efficient Training of MiniGo on Large-scale Heterogeneous Computing System
abstract
Deep Reinforcement Learning has been successfully applied in various applications and achieved impressive performance compared with previous traditional methods but suffers from high computation cost and long training time. MLPerf takes deep reinforcement learning as one of the benchmark tracks and provides a single node training version of MiniGo as a reference. A key challenge is to achieve efficient MiniGo training on a large-scale computing system. According to the training computation pattern in MiniGo and the characteristics of our large-scale heterogeneous computing system, we propose a MultiLevel Parallel strategy, MLPs, including task-level parallelism between nodes, CPU-DSP heterogeneous parallelism, and DSP multi-core parallelism. The proposed method reduces the overall execution time from 43 hours to 16 hours while scaling the node size from 1067 to 4139. The scaling efficiency is 69.1%. According to our fitting method, the scaling efficiency is 46.5% when scaling to 8235 nodes. The experimental results show that the proposed method achieves the efficient training of MiniGo on the largescale heterogeneous computing system.
Peng Qiao, Zhouyu He, Rongchun Li, Jingfei Jiang, Yong Dou, Dongsheng Li 0001
ICPADS3
2022 A pipelining strategy for accelerating convolution neural networks on ARM CPUs
abstract
Abstract Convolution is a primary operation in convolution neural networks. The speed of inference is mainly decided by the speed of the convolutional layer. Improving the performance of embedded processors makes it possible to process the inference on embedded devices. In this article, a pipelining strategy of single instruction and multiple data (SIMD) instructions is proposed to finely optimize the process of the 3 × 3 convolution on ARM‐based CPUs. We implement the SIMD group to improve the efficiency of the SIMD pipeline. A tiling method is exploited to increase data reuse during the process. An evaluation model is proposed to guide the design of the tiling method and register allocation. The speed of our implementation is 5.18 times of the GNU compiler collection compiled unoptimized version on RK3288. The effect of our optimizing method is measured by a performance profiling tool, the performance information suggests that the pipelining strategy has a significant effect for both normal and depthwise separable convolution. By implementing multithread processing, the speedup achieves 18.3 compared with the single thread unoptimized version.
Yong Dou, Rongchun Li, Peng Zhang 0035, Yuntao Liu 0004
Concurr. Comput. Pract. Exp.3
2022 Hierarchical learning with backtracking algorithm based on the Visual Confusion Label Tree for large-scale image classification
Yuntao Liu 0004, Yong Dou, Ruochun Jin, Rongchun Li, Peng Qiao
Vis. Comput.4
2021 Ddper: Decentralized Distributed Prioritized Experience Replay
abstract
In off-policy reinforcement learning, prioritized experience replay plays an important role. However, the centralized prioritized experience replay becomes the bottleneck for efficient training. We propose to approximate the centralized prioritized experience replay in a distributed and decentralized way under certain mild assumptions. To be specific, each actor stores samples in its local replay in the same way as prioritized experience replay, the learner fetches a batch of samples from these replays following a certain strategy. We implement a Deep Q-Learning off-policy algorithm upon the proposed framework. The comparison experiments are performed on a commonly used subset of the Atari-57 learning environment. The experimental results show that the proposed framework speeds up training as the number of actors increases. With the same algorithm and hyper-parameter settings, the proposed framework with 16 actors achieves superior performance that Ape-X with 32 and even more actors does.
Sidun Liu, Peng Qiao, Yong Dou, Rongchun Li
ICME4
2020 End-to-end Spatial Attention Network with Feature Mimicking for Head Detection
abstract
Human head detection is a widely used task and suitable for identifying persons in practical applications. Although existing methods have achieved significant progress, the problems of false alarm and miss detection are still challenging, which arise from weak classification power of detector in the face of variability in occlusion, illumination, etc. In this paper, we present an effective end-to-end head detector called Spatial Attention Network with feature Mimicking(SANM) that can obtain better feature and enhanced classification power, through attention mechanism and a feature mimic method. The spatial-wise attention is extracted from several levels of feature and supervised by the bounding-box annotated heat map. The attention improves the quality of the features in the head and opposite area. To further improve the classification ability, we utilize the feature mimicking method to drive network learning the feature refined by a deep cascading classifier. Compared with the baseline model, our method achieves better performance and produces leading results on head detection benchmarks.
Yuntao Liu 0004, Rongchun Li, Yong Dou
FG3
2020 Towards Precise End-to-end Semi-Supervised Human Head Detection Network
abstract
Head detection, as a fundamental task in practice for many head-related problems, requires an enormous number of annotated boxes to maintain the performance. To alleviate the time and cost of labeling each image in the dataset, we propose an end-to-end semi-supervised head detection frame-work, which shows competitive results with only a small set of data. Specifically, under the setting of semi-supervised, we introduce a weak boxes generate branch and a weak boxes refine branch to produce pseudo ground truth label for unlabeled images with the guidance of annotated images. The weak boxes generate branch is embedded in the detection framework taking the proposals as input and outputting the initial weak boxes that coarsely locate the place of the head. Then, the weak boxes refine branch adjusts the weak boxes more accurate gradually by training a transferred sub-network with the established relation between proposals, weak boxes and labeled boxes. In the training process, we jointly train the two branches in an end-to-end manner, which can generate better pseudo bounding boxes with a small dataset online to avoid over-fitting and obtain a more precise head detector. The results on the public head detection benchmark Brainwash and SCUT-HEAD show the effectiveness of our method.
Rongchun Li, Yuntao Liu 0004, Yong Dou
IJCNN1
2020 A High-Throughput LDPC Decoder Based on GPUs for 5G New Radio
abstract
In this paper, we propose a GPU-based QC-LDPC decoder for 5G New Radio(NR). Different from existing LDPC decoders based on GPUs, our decoder achieves high throughput when decoding LDPC codes with high code rates. Moreover, we implement the shortening and puncturing techniques which are exploited by 5G NR. The decoding algorithm Min-Sum approximation algorithm(MSA) is optimized to implement efficient parallel decoding on the GPU. In order to save the on-chip and the off-chip bandwidth, we propose the two-level quantization scheme and implement data packing on the GPU. We also analyse the optimum thread assignment for different code rates based on our implementation. By using the optimum settings on the GPU, the decoding throughput achieves 1.38 Gbps in the case of (2080, 1760), r=5/6 on Nvidia RTX 2080Ti.
Rongchun Li, Hengyue Pan, Huayou Su, Yong Dou
ISCC1
2020 Annealed gradient descent for deep learning
Hengyue Pan, Xin Niu 0002, Rongchun Li, Yong Dou
Neurocomputing3
2020 DropFilterR: A Novel Regularization Method for Learning Convolutional Neural Networks
Hengyue Pan, Xin Niu 0002, Rongchun Li, Yong Dou
Neural Process. Lett.3
2019 Accelerated Inference Framework of Sparse Neural Network Based on Nested Bitmask Structure
abstract
In order to satisfy the ever-growing demand for high-performance processors for neural networks, the state-of-the-art processing units tend to use application-oriented circuits to replace Processing Engine (PE) on the GPU under circumstances where low-power solutions are required. The application-oriented PE is fully optimized in terms of the circuit architecture and eliminates incorrect data dependency and instructional redundancy. In this paper, we propose a novel encoding approach on a sparse neural network after pruning. We partition the weight matrix into numerous blocks and use a low-rank binary map to represent the validation of these blocks. Furthermore, the elements in each nonzero block are also encoded into two submatrices: one is the binary stream discriminating the zero/nonzero position, while the other is the pure nonzero elements stored in the FIFO. In the experimental part, we implement a well pre-trained sparse neural network on the Xilinx FPGA VC707. Experimental results show that our algorithm outperforms the other benchmarks. Our approach has successfully optimized the throughput and the energy efficiency to deal with a single frame. Accordingly, we contend that Nested Bitmask Neural Network (NBNN), is an efficient neural network structure with only minor accuracy loss on the SoC system.
Yipeng Zhang 0001, Bo Du 0001, Lefei Zhang, Rongchun Li, Yong Dou
IJCAI4
2018 Deep Image Clustering Using Convolutional Autoencoder Embedding with Inception-Like Block
abstract
Image clustering is one of the challenging tasks in machine learning, and has been extensively used in various applications. Recently, various deep clustering methods has been proposed. These methods take a two-stage approach, feature learning and clustering, sequentially or jointly. We observe that these works usually focus on the combination of reconstruction loss and clustering loss, relatively little work has focused on improving the learning representation of the neural network for clustering. In this paper, we propose a deep convolutional embedded clustering algorithm with inception-like block (DCECI). Specifically, an inception-like block with different type of convolution filters are introduced in the symmetric deep convolutional network to preserve the local structure of convolution layers. We simultaneously minimize the reconstruction loss of the convolutional autoencoders with inception-like block and the clustering loss. Experimental results on multiple image datasets exhibit the promising performance of our proposed algorithm compared with other competitive methods.
Qiang Wang 0006, Rongchun Li, Peng Qiao, Ke Yang 0004, Shijie Li 0002, Yong Dou
ICIP3
2018 Temporal Pyramid Relation Network for Video-Based Gesture Recognition
abstract
Gesture recognition in video is an important application of computer vision. However, there are few works talked about the temporal order or relation of the frames in video, which is important for model gestures. In this paper, we propose Temporal Pyramid Relation Network (TPRN) which can model the temporal relation of video frames effectively and efficiently. First, we use Temporal Pyramid Pooling (TPP) layer to get temporal feature sequences of multiple scale pyramids. Then, a Temporal Relation Network (TRN) is stacked on the feature sequence of each scale respectively to model the temporal relations of video frames at multiple scales. At last, representations of all scales are aggregated to get the final prediction. TPRN can take video clips of various length as input and is scalable for video length. We evaluate TPRN on a recently released very large video-based gesture recognition dataset - 20BN-Jester dataset v1, and TPRN achieves competitive performance.
Ke Yang 0004, Rongchun Li, Peng Qiao, Qiang Wang 0006, Dongsheng Li 0001, Yong Dou
ICIP2
2018 Visual Confusion Label Tree for Image Classification
abstract
Convolution neural network models are widely used in image classification tasks. However, the running time of such models is so long that it is not the conforming to the strict real-time requirement of mobile devices. In order to optimize models and meet the requirement mentioned above, we propose a method that replaces the fully-connected layers of convolution neural network models with a tree classifier. Specifically, we construct a Visual Confusion Label Tree based on the output of the convolution neural network models, and use a multi-kernel SVM plus classifier with hierarchical constraints to train the tree classifier. Focusing on those confusion subsets instead of the entire set of categories makes the tree classifier more discriminative and the replacement of the fully-connected layers reduces the original running time. Experiments show that our tree classifier obtains a significant improvement over the state-of-the-art tree classifier by 4.3% and 2.4% in terms of top-l accuracy on CIFAR-100 and ImageNet datasets respectively. Additionally, our method achieves 124× and 115× speedup ratio compared with fully-connected layers on AlexNet and VGG16 without accuracy decline.
Yuntao Liu 0004, Yong Dou, Ruochun Jin, Rongchun Li
ICME4
2018 An efficient CPU-GPU hybrid parallel implementation for DVB-RCS2 receiver
abstract
Summary The second‐generation digital video broadcasting return channel via satellite (DVB‐RCS2) is a promising real‐time wireless protocol that has been widely used in many applications, such as video conferences, video feeds, and video multicasting. However, the receiver end of DVB‐RCS2 is time consuming and should be accelerated by high‐performance processing systems. Today, graphic processing units (GPUs) have been applied in communication systems due to high parallel capability and processing throughput. In this study, we design a novel pipeline of the receiver on the CPU‐GPU platform. Moreover, we propose a CPU‐GPU hybrid strategy to fully utilize resources and reduce communication latency. Compared with the parallel turbo decoder proposed in other work on the same platform, our parallel implementation achieves higher throughput. For the entire DVB‐RCS2 receiver, compared with the non‐pipelined serial and non‐pipelined parallel algorithms, our proposed pipeline obtains 20 times and 6 times speedup, respectively. In addition, the latency of our implementation is lower than that of non‐pipelined CPU‐GPU implementation, which is equal to 1.06 ms.
Yueqing Wang, Fang Wang 0004, Rongchun Li, Yong Dou
Concurr. Comput. Pract. Exp.3
2017 Multiple Kernel Clustering Framework with Improved Kernels
abstract
Multiple kernel clustering (MKC) algorithms have been successfully applied into various applications. However, these successes are largely dependent on the quality of pre-defined base kernels, which cannot be guaranteed in practical applications. This may adversely affect the clustering performance. To address this issue, we propose a simple while effective framework to adaptively improve the quality of these base kernels. Under our framework, we instantiate three MKC algorithms based on the widely used multiple kernel $k$-means clustering (MKKM), MKKM with matrix-induced regularization (MKKM-MR) and co-regularized multi-view spectral clustering (CRSC). After that, we design the corresponding algorithms with proved convergence to solve the resultant optimization problems. To the best of our knowledge, our framework fills the gap between kernel adaption and clustering procedure for the first time in the literature and is readily extendable. Extensive experimental research has been conducted on 7 MKC benchmarks. As is shown, our algorithms consistently and significantly improve the performance of the base MKC algorithms, indicating the effectiveness of the proposed framework. Meanwhile, our framework shows better performance than compared ones with imperfect kernels.
Yueqing Wang, Xinwang Liu 0002, Yong Dou, Rongchun Li
IJCAI4
2017 Approximate Large-scale Multiple Kernel k-means Using Deep Neural Network
abstract
Multiple kernel clustering (MKC) algorithms have been extensively studied and applied to various applications. Although they demonstrate great success in both the theoretical aspects and applications, existing MKC algorithms cannot be applied to large-scale clustering tasks due to: i) the heavy computational cost to calculate the base kernels; and ii) insufficient memory to load the kernel matrices. In this paper, we propose an approximate algorithm to overcome these issues, and to make it be applicable to large-scale applications. Specifically, our algorithm trains a deep neural network to regress the indicating matrix generated by MKC algorithms on a small subset, and then obtains the approximate indicating matrix of the whole data set using the trained network, and finally performs the $k$-means on the output of our network. By mapping features into indicating matrix directly, our algorithm avoids computing the full kernel matrices, which dramatically decreases the memory requirement. Extensive experiments show that our algorithm consumes less time than most comparatively similar algorithms, while it achieves comparable performance with MKC algorithms.
Yueqing Wang, Xinwang Liu 0002, Yong Dou, Rongchun Li
IJCAI4
2017 Learning Non-local Image Diffusion for Image Denoising
abstract
Image diffusion plays a fundamental role for the task of image denoising. The recently proposed trainable nonlinear reaction diffusion (TNRD) model defines a simple but very effective framework for image denoising. However, as the TNRD model is a local model, whose diffusion behavior is purely controlled by information of local patches, it is prone to create artifacts in the homogenous regions and over-smooth highly textured regions, especially in the case of strong noise levels. Meanwhile, it is widely known that the non-local self-similarity (NSS) prior stands as an effective image prior for image denoising, which has been widely exploited in many non-local methods. In this work, we are highly motivated to embed the NSS prior into the TNRD model to tackle its weaknesses. In order to preserve the expected property that end-to-end training remains available, we exploit the NSS prior by defining a set of non-local filters, and derive our proposed trainable non-local reaction diffusion (TNLRD) model for image denoising. Together with the local filters and influence functions, the non-local filters are learned by employing loss-specific training. The experimental results show that the trained TNLRD model produces visually plausible recovered images with more textures and less artifacts, compared to its local versions. Moreover, the trained TNLRD model can achieve strongly competitive performance to recent state-of-the-art image denoising methods in terms of peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM).
Peng Qiao, Yong Dou, WenSen Feng, Rongchun Li, Yunjin Chen
ACM Multimedia4
2015 Efficient graphics processing unit based layered decoders for quasicyclic low-density parity-check codes
abstract
SUMMARY Because layered low‐density parity‐check (LDPC) decoding algorithm was proposed, one can exploit the diversity gain to achieve performance comparable to the traditional two‐phase message passing (TPMP) decoding but with about twice faster decoding convergence compared to TPMP. In order to reduce the decoding time of layered LDPC decoder, a graphics processing unit (GPU) is exploited as the modem processor so that the decoding procedure can be processed in parallel using numerous threads in the GPU. In this paper, we present the parallel algorithms and efficient implementations on the GPU for two different layered message passing schemes, the row‐layered and column‐layered decoding. In the experiments, the quasicyclic LDPC codes for WiFi (802.11n) and WiMAX (802.16e) are decoded by the proposed layered LDPC decoders. The experimental results show that our decoder has good bit error ratio (BER) performance comparable to TPMP decoder. The peak throughput is 712 Mbps, which is about two orders of magnitude faster than that of CPU implementation and comparable to the dedicated hardware solutions. Compared to the existing fastest GPU‐based implementation, the presented decoder can achieve a performance improvement of 2.3 times. Copyright © 2013 John Wiley & Sons, Ltd.
Rongchun Li, Yong Dou, Dan Zou
Concurr. Comput. Pract. Exp.1
2014 Efficient parallel implementation of three-point viterbi decoding algorithm on CPU, GPU, and FPGA
abstract
SUMMARY In wireless communication, Viterbi decoding algorithm (VDA) is the one of most popular channel decoding algorithms, which is widely used in WLAN, WiMAX, or 3G communications. However, the throughput of Viterbi decoder is constrained by the convolutional characteristic. Recently, the three‐point VDA (TVDA) was proposed to solve this problem. In TVDA, the whole procedure can be divided into three phases, the forward, trace‐back, and decoding phases. In this paper, we analyze the parallelism of TVDA and propose parallel TVDA on the multi‐core CPU, graphics processing unit (GPU), and field programmable gate array (FPGA). We demonstrate approaches that fully exploit its performance potential on CPU, GPU, and FPGA computing platforms. For CPU platforms, we perform two optimization methods, single instruction multiple data and multithreading to gain over 145 × speedup over the naive CPU version on a quad‐core CPU platform. For GPU platforms, we propose the combination of cached memory optimization, coalesced global memory accesses, codeword packing scheme, and asynchronous data transition, achieving the throughput of 404.65 Mbps and 12 × speedup over initial GPU versions on an NVIDIA GeForce GTX580 card and 7 × speedup over Intel quad‐core CPU i5‐2300, under the same manufacturing year and both with fully optimized schemes. In addition, for FPGA platforms, we customize a radix‐4 pipelined architecture for the TVDA in a 45‐nm FPGA chip from Xilinx (XC6VLX760). Under 209.15‐MHz clock rate, it achieves a throughput of 418.30 Mbps. Finally, we also discuss the performance evaluation and efficiency comparison of different flexible architectures for real‐time Viterbi decoding in terms of the decoding throughput, power consumption, optimization schemes, programming costs, and price costs.Copyright © 2013 John Wiley & Sons, Ltd.
Rongchun Li, Yong Dou, Dan Zou
Concurr. Comput. Pract. Exp.1
2014 Supernodal sparse Cholesky factorization on graphics processing units
abstract
SUMMARY Sparse Cholesky factorization is the most computationally intensive component in solving large sparse linear systems and is the core algorithm of numerous scientific computing applications. A large number of sparse Cholesky factorization algorithms have previously emerged, exploiting architectural features for various computing platforms. The recent use of graphics processing units (GPUs) to accelerate structured parallel applications shows the potential to achieve significant acceleration relative to desktop performance. However, sparse Cholesky factorization has not been explored sufficiently because of the complexity involved in its efficient implementation and the concerns of low GPU utilization. In this paper, we present a new approach for sparse Cholesky factorization on GPUs. We present the organization of the sparse matrix supernode data structure for GPU and propose a queue‐based approach for the generation and scheduling of GPU tasks with dense linear algebraic operations. We also design a subtree‐based parallel method for multi‐GPU system. These approaches increase GPU utilization, thus resulting in substantial computational time reduction. Comparisons are made with the existing parallel solvers by using problems arising from practical applications. The experiment results show that the proposed approaches can substantially improve sparse Cholesky factorization performance on GPUs. Relative to a highly optimized parallel algorithm on a 12‐core node, we were able to obtain speedups in the range 1.59× to 2.31× by using one GPU and 1.80× to 3.21× by using two GPUs. Relative to a state‐of‐the‐art solver based on supernodal method for CPU‐GPU heterogeneous platform, we were able to obtain speedups in the range 1.52× to 2.30× by using one GPU and 2.15× to 2.76× by using two GPUs. Concurrency and Computation: Practice and Experience, 2013. Copyright © 2013 John Wiley & Sons, Ltd.
Dan Zou, Yong Dou, Song Guo 0003, Rongchun Li
Concurr. Comput. Pract. Exp.4
2014 CuSora: Real-time software radio using multi-core graphics processing unit
Rongchun Li, Yong Dou, Jie Zhou 0007
J. Syst. Archit.1
2013 A fully parallel truncated Viterbi decoder for Software Defined Radio on GPUs
abstract
Software-Defined Radio (SDR) on the Graphics Processing Unit (GPU) platform. We exploit a map-reduce strategy based on the three-point Viterbi decoding algorithm (TVDA) due to the high parallelization potential. The trellis of Viterbi decoding algorithm can be divided into sub-trellises in truncation, which can perform independent forward metrics computing and trace-back procedure in parallel. The parallel Viterbi decoding algorithm is mapped on a GPU named NVIDIA GTX580. The experiment shows that our method shows low BER and 36.0x speedup over a C implementation on a CPU with the frequency of 2.0GHz. At the meantime, our method achieves a performance improvement of 1.2x-3.6x times that of the existing GPU-based implementation.
Rongchun Li, Yong Dou
WCNC1