Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Falei Luo

dblp:159/3838 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
2since 2021 · last 2022
0000-0003-3263-9549ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 2 since 2021Systems, architecture and hardware · 4 · 3 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
3 papers
Image and video coding · 80% Image and video processing · 20%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
GPUs and heterogeneous computing · 64% Parallel and multicore computing · 36%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video coding
video compression
1.022022
Joint Local and Nonlocal Progressive Prediction for Versatile Video Coding · IEEE Trans. Image Process. 2022
GPU-Based Hierarchical Motion Estimation for High Efficiency Video Coding · IEEE Trans. Multim. 2019
Image and video coding › video compression
intra prediction
0.612022
Joint Local and Nonlocal Progressive Prediction for Versatile Video Coding · IEEE Trans. Image Process. 2022
Image and video coding › video compression
in-loop filtering
0.512021
Fast Non-Local Adaptive In-Loop Filter Optimization on GPU · IEEE Trans. Multim. 2021
GPUs and heterogeneous computing
GPU performance optimization
0.512021
Fast Non-Local Adaptive In-Loop Filter Optimization on GPU · IEEE Trans. Multim. 2021
Parallel and multicore computing
parallel computing
0.512021
Fast Non-Local Adaptive In-Loop Filter Optimization on GPU · IEEE Trans. Multim. 2021
Image and video processing
motion estimation
0.412019
GPU-Based Hierarchical Motion Estimation for High Efficiency Video Coding · IEEE Trans. Multim. 2019
GPUs and heterogeneous computing
GPU computing
0.412019
GPU-Based Hierarchical Motion Estimation for High Efficiency Video Coding · IEEE Trans. Multim. 2019
Image and video processing › image restoration
compression artifact removal
0.112021
Fast Non-Local Adaptive In-Loop Filter Optimization on GPU · IEEE Trans. Multim. 2021
Image and video coding › video compression › video codec
HEVC
0.112019
GPU-Based Hierarchical Motion Estimation for High Efficiency Video Coding · IEEE Trans. Multim. 2019

Methods — techniques the papers use, named apart from their topics

thread allocation optimization · 1.0patch matching · 1.0GPU parallelization · 1.0quadtree coding · 0.8parallel motion estimation · 0.8template matching · 0.6progressive prediction · 0.6
YearPublicationVenuePosition
2022 Joint Local and Nonlocal Progressive Prediction for Versatile Video Coding
abstract
In the latest video coding standard, namely Versatile Video Coding (VVC), more directional intra modes and reference lines have been utilized to improve prediction efficiency. However, complex content still cannot be predicted well with only the adjacent reference samples. Although nonlocal prediction has been proposed to further improve the prediction efficiency in existing algorithms, explicit signalling or matching error potentially limits the coding efficiency. To address these issues, we propose a joint local and nonlocal progressive prediction scheme, targeting at improving nonlocal prediction accuracy without additional signalling. Specifically, template matching based prediction (TMP) is conducted firstly to derive an initial nonlocal predictor. Based on the first prediction and previously decoded reconstruction information, a local template, including inner textures and neighboring reconstruction, is carefully designed. With the local template involved in nonlocal matching process, a more accurate nonlocal predictor can be found progressively in the second prediction. Finally, the coefficients from the two predictions are fused and transmitted in bitstreams. In this way, more accurate nonlocal predictor can be derived implicitly with local information instead of being explicitly signalled. Experimental results on the reference software VTM-9.0 of VVC show that the method achieves 1.02% BD-Rate reduction for natural sequences and 2.31% BD-Rate reduction for screen content videos under all intra (AI) configuration.
Meng Lei, Falei Luo, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001
IEEE Trans. Image Process.2
2021 Fast Non-Local Adaptive In-Loop Filter Optimization on GPU
abstract
The non-local adaptive in-loop filter (NALF) for video coding has achieved significant coding gain by exploiting image non-local self-similarity (NSS) to efficiently reduce the compression artifacts. However, the intensive computation of NALF hinders its practical deployment in video standardizations. In this paper, we propose a fast NALF optimization algorithm in parallel-computing framework by leveraging the massive parallel execution resources of GPU. First, the computational complexity of original NALF is analyzed in depth, then the pipelines of computational-intensive modules are re-designed to adapt to the general-purpose GPU with more parallel-friendly consideration. Specifically, we speed up the NALF by optimizing thread allocation to maximize the parallelism degree and elaborately designing the GPU block dimension to avoid access conflict. The group-level and pixel-level parallelization for collaboratively filtering and patch matching modules are designed respectively. To reduce the cost in data transmission, the whole filtering process is implemented on GPU by taking the advantage of low data dependency in NALF. Extensive experimental results show that the proposed fast NALF optimization using GPU architecture achieves high-speeed processing while maintaining the significant coding performance of original NALF, which shows the potential of NALF in the future video coding standard.
Chuanmin Jia, Falei Luo, Xinfeng Zhang 0001, Shiqi Wang 0001, Shanshe Wang, Siwei Ma 0001
IEEE Trans. Multim.2
2020 Two-Step Progressive Intra Prediction For Versatile Video Coding
abstract
In traditional intra prediction, nearest reference samples are utilized to generate the prediction block. Although more directional intra modes and reference lines have been utilized, encoders could not predict complex content with only the 10-cal reference samples efficiently. To address this issue, a twostep progressive prediction method combining local and nonlocal information is proposed. The non-local information can be obtained through template matching based prediction, and the local information can be derived by the high frequency coefficients from the first prediction step. Experimental results show that the proposed method can achieve 0.87% BD-rate reduction in VTM-7.0. In particular, the method is of significant advantages over prediction schemes using only non-local information.
Meng Lei, Falei Luo, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001
ICIP2
2019 Efficient GPU-Based Inter Prediction for Video Decoder
abstract
Interpolation is a very important module in inter prediction for any video decoder, e.g. AVS2 [1] and HEVC [2], which occupies most of the time in the whole decoding process . Thus, the real-time decoder is largely limited by the speed of inter prediction. To solve this problem, we propose an efficient GPU-based interpolation framework for inter prediction. Through optimizing shared memory allocation and thread scheduling on the GPU side, GPU are utilized efficiently and inter prediction is accelerated effectively. The experimental results on AVS2 show that for all Ultra HD 4K, WQXGA and full HD video sequences tested, the inter prediction acceleration ratio is over 6 times, and the average processing time is up to 1.25ms, 0.75ms and 0.45ms, respectively, with the NVIDIA GeForce GTX 1080TI GPU.
Falei Luo, Shanshe Wang, Siwei Ma 0001
ICIP2
2019 Look-Ahead Prediction Based Coding Unit Size Pruning for VVC Intra Coding
abstract
In the emerging video coding standard, Versatile Video Coding (VVC), a quadtree with nested multi-type tree (MTT) using binary and ternary tree structure was proposed. MTT brings significant coding efficiency but increases the encoding complexity. In this paper, a look-ahead prediction based coding unit size pruning algorithm is proposed to cut down redundant MTT partitions. The proposed scheme aims to identify the unnecessary partition direction in advance and consists of two steps, i.e. SATD-based mode decision (SMD) for possible blocks and refined cost derivation based on rate-distortion optimization. Experimental results show that the proposed method can save 41% encoder time with only 0.84% increase in bit rate on average.
Meng Lei, Falei Luo, Xiang Zhang 0004, Shanshe Wang, Siwei Ma 0001
ICIP2
2019 GPU-Based Hierarchical Motion Estimation for High Efficiency Video Coding
abstract
Motion estimation (ME) plays a crucial role in removing the temporal redundancy for video compression. However, during the encoding process a substantial computational burden is imposed by ME due to the exhaustive evaluations of possible candidates within the searching window. In view of the increasing computing capacity of GPU, we propose a GPU-based low delay parallel ME scheme for high efficiency video coding (HEVC). In particular, considering the quadtree coding structure of HEVC, we achieve the parallelization in a hierarchical way by optimizing the ME process in a coding tree unit (CTU), prediction unit (PU), and motion vector (MV) layers. Specifically, in the CTU layer, a novel motion vector predictor determination scheme is proposed to alleviate the side effects of inaccurate MV prediction due to the removal of the CTU-level dependency. In the PU layer, a novel indexing table is particularly designed to realize an efficient cost derivation strategy. As such, the cost of each PU can be computed in a convenient and efficient manner. In an MV layer, we propose a compact descriptor to represent MV and its corresponding cost as a whole, such that the redundant branches can be further avoided in the searching process. With such an optimization strategy, the proposed scheme can completely save the encoding time for ME on CPU. Experimental results demonstrate that the proposed scheme can achieve 41% encoding time savings with the ME acceleration up to 12.7 times, and the incurred BD-BR loss is only 0.52% on average. Moreover, further experimental results show that the proposed GPU-based ME can achieve up to 200 times acceleration compared to the full search ME on CPU.
Falei Luo, Shanshe Wang, Shiqi Wang 0001, Xinfeng Zhang 0001, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Multim.1
2017 An adaptive and low-complexity all-zero block detection for HEVC encoder
abstract
To improve the detection accuracy of SAD or SATD based threshold and save the time cost of RDO determination, we proposed an All Zero Block (AZB) detection method by adaptively searching the maximum transform coefficient amplitude in low frequency of TU after conventional SATD detection was failed. The experimental results show that our algorithm can achieve around 39% transform and quantization time-saving with only 0.1% on average RD performance reduction. The detection accuracy of larger TU size, i.e. 16×16 and 32 × 32, can reach up to about 95% on average.
Ruiqin Xiong, Falei Luo, Shanshe Wang, Siwei Ma 0001
ISCAS3
2017 Fast intra coding unit size decision for HEVC with GPU based keypoint detection
abstract
In this paper, a fast intra Coding Unit (CU) size decision framework based on keypoint detection on Graphic Processing Unit (GPU) is proposed. In this framework, firstly the original frames are sent to GPU and then keypoint detection is conducted with numerous threads, which is able to avoid bringing in additional computational complexity even in realtime systems. Then, based on the keypoint distribution, whether to split the CU to the next coding depth is efficiently predicted. Experiments show that the proposed algorithm can achieve over 25% time saving under all intra (AI) configuration with ignorable performance loss.
Falei Luo, Shanshe Wang, Siwei Ma 0001, Nan Zhang 0015, Wen Gao 0001
ISCAS1
2016 GPU based sample adaptive offset parameter decision and perceptual optimization for HEVC
abstract
In this paper, a graphics processing unit (GPU) based sample adaptive offset (SAO) parameters decision scheme is proposed for High Efficiency Video Coding (HEVC). Then, in order to further improve the performance of SAO, a perceptual based optimization scheme is provided according to the adjustment of Lagrange multiplier aiming to improve the subjective performance of SAO. Experimental results demonstrate that the proposed GPU based SAO parameter decision scheme can achieve average 0.76% and 0.78% BD-rate gain in terms of PSNR (Peak Signal to Noise Ratio) and SSIM (Structure Similarity) respectively. Combined with the perceptual optimization scheme, the maximum BD-rate gain in terms of PSNR and SSIM can be up to 1.77% and 3.3% with the average as 1.23% and 1.37%. Moreover, much computation complexity of SAO can be distributed to GPU.
Falei Luo, Shanshe Wang, Nan Zhang 0015, Siwei Ma 0001, Wen Gao 0001
ISCAS1
2016 A novel mode decision for depth map coding in 3D-AVS
abstract
In this paper, a new mode decision scheme is proposed for depth map coding in 3D-AVS. The novelty of the paper mainly contains the following two points. Firstly, an improved distortion estimation model of synthesized views is proposed. Secondly, for the mode decision of depth map coding, the distortion is represented to be the weighted sum of depth distortion and estimated distortion of the synthesized view. We proposed a new scheme to derive the weighting factors adaptively based on the disparity. Then the distortion is utilized to calculate the rate distortion cost for mode decision. Experimental results demonstrate that the proposed scheme achieves remarkable performance improvement in 3D-AVS. The average BD-rate gain is about 12%.
Falei Luo, Shanshe Wang, Shiqi Wang 0001, Siwei Ma 0001
VCIP2
2016 Low complexity encoder optimization for HEVC
Shanshe Wang, Falei Luo, Siwei Ma 0001, Xiang Zhang 0004, Shiqi Wang 0001, Debin Zhao, Wen Gao 0001
J. Vis. Commun. Image Represent.2
2015 Multiple layer parallel motion estimation on GPU for High Efficiency Video Coding (HEVC)
abstract
This paper provides a multiple-layer parallel motion estimation (ME) scheme implemented on GPU for High Efficiency Video Coding (HEVC). The scheme is hierarchically structured, including four layers: coding tree unit (CTU), prediction unit (PU), motion vector (MV) selection and instruction optimization. In PU-layer, costs of various PU sizes were obtained through a SAD (sum of absolute differences) look-up table instead of progressive cost merging. And during MV selection, GPU's comparison instruction was used to avoid branches. At the same time, concurrent CTUs processing and SIMD (Single Instruction, Multiple Data) optimization also improve the performance significantly. Experimental results show that the proposed scheme can take full advantage of GPU and achieves over 90 times speedup compared with the HM10.0 using fast ME.
Falei Luo, Siwei Ma 0001, Juncheng Ma, Honggang Qi, Li Su 0003, Wen Gao 0001
ISCAS1
2015 Parallel intra coding for HEVC on CPU plus GPU platform
abstract
In High Efficiency Video Coding (HEVC), the intra coding performance is significantly improved due to the recursive splitting structure and up to 35 intra prediction modes. However, the computational complexity of intra coding increases largely as well. In this paper, a fast intra coding scheme is proposed based on CPU and GPU cooperation. Firstly, the intra prediction of variable blocks is performed in parallel on multi-cores GPU. Secondly, the intra prediction mode with minimum Sum of Absolute Difference (SAD) cost is selected and transmitted to the host CPU. Instead of exhaustively searching all the intra modes in Rough Mode Decision (RMD) process, the mode returned by the GPU is directly selected. Lastly, the texture gradient of each coding unit (CU) is assessed during parallel intra prediction, then used by the CPU for fast CU size decision. Experiment results show that the proposed parallel intra coding method achieves up to 62% complexity reduction with acceptable coding performance loss.
Juncheng Ma, Falei Luo, Shanshe Wang, Nan Zhang 0015, Siwei Ma 0001
VCIP2
2015 Adaptive motion vector resolution prediction in block-based video coding
abstract
In the classical block-based video coding, motion vector is derived for each coding block to remove the inter-frame redundancy. However, the motion vector resolution is usually restricted to be identical, typically 1/4-pixel resolution, regardless of the different video contents. In this paper, we propose an algorithm that can adaptively select the optimal motion vector resolution at frame level according to the characteristics of the video contents. We first derived a residual energy model, and the major factors that may impact the motion vector resolution are considered, including the texture complexity, motion scale, inter-frame noise and quantization parameter. Experimental results have shown that the proposed scheme can achieve 1.8% BD-rate gain on average without complexity increment.
Zhao Wang 0004, Juncheng Ma, Falei Luo, Siwei Ma 0001
VCIP3
2014 Flexible CTU-level parallel motion estimation by CPU and GPU pipeline for HEVC
abstract
In the high efficiency video coding (HEVC) encoder, motion estimation (ME) takes up more than 50% encoding time. To reduce the complexity of the ME module in HEVC, this paper proposes a flexible coding tree unit (CTU)-level parallel ME method through CPU and GPU pipeline collaboration. Firstly a highly scalable CTU-level parallel motion search scheme on GPU is provided, in which, the parallel CTU group can be configured to be any size to adapt to the variable sequence resolution and hardware configurations. Then, the motion search range can be adaptively adjusted based on the motion intensity. Therefore, the unnecessary GPU time wasting can be further avoided for slow-moving scenes, while high performance kept for fast-moving scenes. Moreover, the ME information returned from GPU can be used by CPU for fast mode decision. Experimental results show that the proposed method achieves up to 73% complexity reduction than HM10.0 anchor using CPU only with acceptable coding performance loss, providing higher performance than the state-of-the-art scheme.
Juncheng Ma, Falei Luo, Shanshe Wang, Siwei Ma 0001
VCIP2