EDBT 2026 Demo / reviewers in the wild / expert
Linwei Zhu
dblp:131/2809
· DBLP profile ↗
40ranked-venue papers
14as first author
24since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 14 first-author · 19 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 4 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | CDV-PCQA: Content-distortion-guided dynamic viewpoint quality assessment for 3D point clouds
Qihao Liang, Li Li 0014, Ting Luo 0001, Gangyi Jiang, Wujie Zhou, Linwei Zhu, Zhouyan He |
Expert Syst. Appl. | 6 |
| 2026 | ARMLF: Anomalous region representation learning for multi-exposure fused light field image quality assessment
Guanglong Liao, Gangyi Jiang, Linwei Zhu, Yeyao Chen, Yueli Cui, Ting Luo 0001, Haiyong Xu |
Expert Syst. Appl. | 3 |
| 2026 | Deep Feature Prior-Guided Conditional Diffusion Model for Underwater Image EnhancementabstractUnderwater imaging always suffers from color distortion and reduced visibility due to light absorption and scattering, severely hindering visual perception and analysis. In this letter, we propose an underwater image enhancement framework based on diffusion model augmented with two lightweight guidance modules. The first module is a conditional branch that extracts structural features from a coarsely enhanced version to guide the denoising process toward more faithful restoration. While the second module retrieves high-quality features from a pre-constructed feature dictionary as priors, effectively restoring colors and fine details in degraded regions. Extensive experiments on public underwater image datasets demonstrate that our proposed method outperforms the state-of-the-art approaches both quantitatively and visually. It also generalizes well across various underwater environments, highlighting the effectiveness of incorporating structural and feature-level guidance into the diffusion process. The source code and pre-trained model are available at https://github.com/Juneit/PGUIE. Linwei Zhu, Tao Tian, Wenhui Wu 0001, Jingchao Cao |
IEEE Signal Process. Lett. | 2 |
| 2026 | Deep-JGAC: End-to-End Deep Joint Geometry and Attribute Compression for Dense Colored Point CloudsabstractColored point cloud becomes a fundamental representation in the realm of 3D vision. Effective Point Cloud Compression (PCC) is urgently needed due to the huge amount of data. In this paper, we propose an end-to-end Deep Joint Geometry and Attribute Compression (Deep-JGAC) method for dense colored point clouds. First, we propose a flexible Deep-JGAC framework, where the geometry and attribute encoders are compatible with either learning or non-learning encoders. Second, we propose an end-to-end deep residual self-attention-based geometry encoder to improve geometry coding efficiency, where a Hybrid Residual Self-attention Module (HRSM) is proposed to enhance geometry representation by considering its geometrical importance. Third, to solve the mismatch between the point cloud geometry and attribute caused by the geometry compression distortion, we propose an optimized re-colorization module to attach attribute to the geometrically distorted point cloud for attribute coding, which lowers the computational complexity. Extensive experimental results demonstrate that, in terms of the geometry quality metric D1-PSNR, the proposed Deep-JGAC achieves average Bjøntegaard Delta Bit Rate (BDBR) of -82.96%, -44.63%, -36.46%, -41.72%, and -31.16% compared to the G-PCC (Octree), G-PCC (Trisoup), V-PCC, GRASP, and PCGCv2, respectively. For the perceptual joint quality metric MS-GraphSIM, Deep-JGAC achieves an average BDBR of - 48.72%, -57.14%, -14.67% and -13.37% against G-PCC(Octree), IT-DL-PCC, V-PCC, and DeepPCC, respectively. In addition, the costs of encoding/decoding time are reduced by 32.8%/30.8%, 80.1%/81.8%, 97.2%/35.7%, 98.4%/92.3%, and 96.4%/99.6% on average compared to G-PCC (Octree), G-PCC (Trisoup), V-PCC, IT-DL-PCC and DeepPCC. The code and pre-trained models are available at https://github.com/SYSU-Video/Deep-JGAC. Yun Zhang 0002, Zixi Guo, Linwei Zhu, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | C-CTX: Cubic-Checkerboard Context Entropy Model for Learned Image CompressionabstractLearned Image Compression (LIC) has achieved superior performance in recent years, of which the context entropy model is an important component. However, in the context entropy model, there is no deterministic correlation between neighboring channels, and it is difficult to capture inter-channel correlation as well as spatial correlation for further improving the performance. To address this issue, a Cubic-Checkerboard conTeXt entropy model (C-CTX) for LIC is proposed in this work, which is able to refer uniformly across the channel domain and maintain the correlations in the spatial domain. To make neighboring channels have more similar distribution, Cubic Checkerboard Mask (CCM) with channel- wise mask convolution is utilized to achieve uniform distribution in different domains and Channel Wise Re-Arrangement (CWRA) is performed in terms of entropy. Based on CCM and CWRA, two Feature Disentangle Modules (FDMs) are designed in C-CTX to project the context information within sub-spaces for catching spatial correlation and channel correlation separately. Extensive experimental evaluations show that our method outperforms the state-of-the-art works on six datasets, i.e., Kodak, Tecnick, CLIC'20, CLIC'21, CLIC'22, and JPEG-AI. Shiyu Feng, Linwei Zhu, Yun Zhang 0002, Na Li 0015, Shiqi Wang 0001 |
IEEE Trans. Multim. | 2 |
| 2026 | Temporal Consistency-Aware Dynamic Point Clouds Color Attribute EnhancementabstractDynamic point clouds, widely used in virtual reality and autonomous driving systems, often suffer from distortions due to quantization in the process of compression. These distortions significantly degrade the visual quality of dynamic point clouds, especially temporal inconsistency. To address this issue, a temporal consistency-aware dynamic point clouds color attribute enhancement method is proposed in this work. Specifically, a 3D Spatial-Temporal Search (STS) module is designed to adaptively search point cloud patches in the temporal domain for feature alignment. These matched patches are then individually fed into Single Frame Feature Extraction (SFFE) module that comprises of multi-head attention and graph convolution to exploit latent features of point cloud color attribute. In addition, to further capture both the spatial and temporal dependencies, a Convolutional Point cloud Long Short-Term Memory (Conv-PointLSTM) network is applied, which integrates convolution and max pooling with LSTM mechanism to facilitate the color attribute correspondents across the spatial-temporal latent features. Experimental results demonstrate that the proposed method can achieve 0.44 dB gains on average in terms of Peak Signal-to-Noise Ratio (PSNR) and 1.50%/5.31%bit rate reductions at the low/high bit rate, which outperforms the state-of-the-art works. The source code and trained models are available athttps://github.com/xu-coder-666/DPC. Linwei Zhu, Ruxu Liang, Yun Zhang 0002, Hui Yuan 0001, Sam Kwong |
IEEE Trans. Multim. | 1 |
| 2026 | Enhanced Quality-Aware Scalable Underwater Image CompressionabstractUnderwater imaging plays a pivotal role in marine exploration and ecological monitoring. However, it faces significant challenges of limited transmission bandwidth and severe distortion in the aquatic environment. In this work, to achieve the target of both underwater image compression and enhancement simultaneously, an enhanced quality-aware scalable underwater image compression framework is presented, which comprises a Base Layer (BL) and an Enhancement Layer (EL). In the BL, the underwater image is represented by a controllable number of non-zero sparse coefficients for coding bits saving. Furthermore, the underwater image enhancement dictionary is derived with shared sparse coefficients to make reconstruction close to the enhanced version. In the EL, a dual-branch filter comprising rough filtering and detail refinement branches is designed to produce a pseudo-enhanced version for residual redundancy removal and to improve the quality of final reconstruction. Extensive experimental results demonstrate that the proposed scheme outperforms the state-of-the-art works under five large-scale underwater image datasets in terms of Underwater Image Quality Measure (UIQM). Linwei Zhu, Xu Zhang 0044, Huan Zhang 0008, Ye Li 0002, Runmin Cong, Sam Kwong |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Mamba-Based Blind Stitched Wide Field of View Light Field Image Quality Assessment via Dual-Viewport SamplingabstractDue to the limitations of commercial light field camera hardware, the field of view (FOV) of light field images (LFIs) is relatively narrow. To expand the FOV, various LFI stitching algorithms have been developed. However, these algorithms inevitably introduce localized distortions and angular consistency disruptions, which conventional LFI quality assessment metrics struggle to evaluate effectively. To address this issue, a novel Mamba-based blind quality assessment metric for stitched wide field of view light field images (WLFIs) using dual-viewport sampling is proposed. Firstly, sub-aperture images from horizontal and vertical directions are stacked to characterize angular information, and a dual-viewport sampling pattern is designed to enhance data augmentation and capture spatial details. After that, a multi-scale state space block is proposed to improve distortion feature extraction, complemented by an auxiliary distortion discrimination task. Finally, experimental results demonstrate that the proposed metric outperforms state-of-the-art metrics on the benchmark WLFI dataset. Gangyi Jiang, Linwei Zhu, Yeyao Chen, Yueli Cui, Ting Luo 0001, Haiyong Xu |
ICME | 3 |
| 2025 | Learning adaptive distractor-aware-suppression appearance model for visual tracking
Huanlong Zhang, Linwei Zhu, Yanchun Zhao, De-Shuang Huang |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Enhancing 3D video watching experiences: Tackling compression and 3D warping distortions in synthesized view with perceptual guidance
Huan Zhang 0008, Xu Zhang 0044, Linwei Zhu, Yun Zhang 0002, Jiang-Zhong Cao, Bingo Wing-Kuen Ling |
Expert Syst. Appl. | 3 |
| 2025 | Geometry-Aware RWKV for Heterogeneous Light Field Spatial Super-ResolutionabstractHeterogeneous Light Field (LF) spatial Super-Resolution (SR) aims to significantly enhance the spatial resolution of LF imaging by integrating an extra 2D digital camera. Inspired by the Receptance Weighted Key Value (RWKV), a simple yet effective heterogeneous LF spatial SR method is proposed. Specifically, a texture transfer module with channel correlation is designed, which leverages a feature distillation strategy to transfer texture information from the high-resolution 2D image to the low-resolution LF image. Meanwhile, a spatial-angular rectification module is constructed to restore the spatial-angular coherence damaged in texture transfer. It employs geometry-aware RWKV to capture the intrinsic geometric structure of LFs. Experimental results show that the proposed method outperforms the state-of-the-art methods in both quantitative and qualitative comparisons, while achieving higher efficiency in terms of inference time and memory usage. Zean Chen, Yeyao Chen, Linwei Zhu, Haiyong Xu, Gangyi Jiang |
IEEE Signal Process. Lett. | 3 |
| 2025 | Blind Light Field Image Quality Assessment via Frequency Domain Analysis and Auxiliary LearningabstractDue to the distortions occurring at various stages from acquisition to visualization, light field image quality assessment (LFIQA) is crucial for guiding the processing of light field images (LFIs). In this letter, we propose a new blind LFIQA metric via frequency domain analysis and auxiliary learning, termed as FABLFQA. First, spatial-angular patches are extracted from LFIs and further processed through discrete cosine transform to obtain light field frequency maps. Subsequently, a concise and efficient frequency-aware deep learning network is designed to extract frequency features, including the frequency descriptor, 3D ConvBlock, and frequency transformer. Finally, a distortion type discrimination auxiliary task is employed to facilitate the learning of the main quality assessment task. Experimental results on three representative LFI datasets show that the proposed metric outperforms the state-of-the-art metrics. Gangyi Jiang, Linwei Zhu, Yueli Cui, Ting Luo 0001 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Geometry-Guided Latent Diffusion Model for Static Point Cloud Color Attribute Denoising
Linwei Zhu, Ruxu Liang, Yun Zhang 0002, Gangyi Jiang, Yo-Sung Ho |
IEEE Signal Process. Lett. | 1 |
| 2025 | DA-Net: A Double Alignment Multimodal Learning Network for Point Cloud Quality AssessmentabstractExisting multimodal point cloud quality assessment (PCQA) methods usually integrate 3D and 2D information to simulate human visual perception of distortions. However, due to the lack of consideration of spatial correspondence, they have difficulty to learn consistent distortion representations from different modalities in the same region of the PC. In addition, they also ignore the heterogeneity of modalities and rely on complex fusion mechanisms (e.g., attention) to integrate multimodal features. Both lead to limited performance and increased computational complexity. To address these limitations, we propose a novel double alignment multimodal learning network (DA-Net), which introduces two key alignment strategies. Specifically, the first is spatial pre-alignment strategy, which generates informative 2D patch for each 3D patch via an adaptive patch projection module (APPM), ensuring accurate spatial correspondence of different modalities prior to feature extraction. The second is a uniform feature alignment strategy, which includes feature disentanglement module (FDM) and feature mapping module (FMM) to relieve heterogeneity of modalities and guide the optimization of 2D and 3D encoder. Finally, multimodal features are simply integrated and regressed to obtain the quality score. Experimental results demonstrate that the DA-Net exhibits outstanding performance and generalization ability. It also achieves lower computational complexity compared with other multimodal PCQA methods. The source codes of DA-Net will be available at https://github.com/Rphone/DA-Net. Xinqiang Wu, Zhouyan He, Ting Luo 0001, Gangyi Jiang, Wujie Zhou, Linwei Zhu, Weisi Lin |
IEEE Trans. Image Process. | 6 |
| 2024 | Neural Network Based Multi-Level In-Loop Filtering for Versatile Video CodingabstractTo further improve the performance of Versatile Video Coding (VVC), a neural network based multi-level in-loop filtering framework for luma and chroma is presented in this letter, which includes Reference pixel Level (RL), Coding tree unit Level (CL), and Frame Level (FL). The neural network based filters in these levels can be flexibly enabled. In RL, the coding performance upper bound is analyzed and asymmetric convolution is designed. In CL, the pixels located at the bottom and rightmost have been assigned greater weights for loss calculation during training. In addition, the co-located luma is adopted in CL and FL chroma filtering for guiding chroma enhancement due to the high correlation between them. For the architecture of neural network, two input channel fusion schemes are combined to enjoy both of their benefits. Extensive experimental results show that the proposed multi-level in-loop filtering method can achieve 6.87%, 32.8%, and 36.9% bit rate reductions on average for Y, U, and V components under all intra configuration, which outperforms the state-of-the-art works. Linwei Zhu, Yun Zhang 0002, Na Li 0015, Wenhui Wu 0001, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Learning to Predict Object-Wise Just Recognizable Distortion for Image and Video CompressionabstractJust Recognizable Distortion (JRD) refers to the minimum distortion that notably affects the recognition performance of a machine vision model. If a distortion added to images or videos falls within this JRD threshold, the degradation of the recognition performance will be unnoticeable. Based on this JRD property, it will be useful to Video Coding for Machine (VCM) to minimize the bit rate while maintaining the recognition performance of compressed images. In this study, we propose a deep learning-based JRD prediction model for image and video compression. We first construct a large image dataset of Object-Wise JRD (OW-JRD) containing 29,218 original images with 80 object categories, and each image was compressed into 64 distorted versions using Versatile Video Coding (VVC). Secondly, we analyze of the distribution of the OW-JRD, formulate JRD prediction as binary classification problems and propose a deep learning-based OW-JRD prediction framework. Thirdly, we propose a deep learning based binary OW-JRD predictor to predict whether an image object is still detectable or not under different compression levels. Also, we propose an error-tolerance strategy that corrects misclassifications from the binary classifier. Finally, extensive experiments on large JRD image datasets demonstrate that the Mean Absolute Errors (MAEs) of the predicted OW-JRD are 4.90 and 5.92 on different numbers of the classes, which is significantly better than the state-of-the-art JRD prediction model. Moreover, ablation studies on deep network structures, object sizes, features, data padding strategies and image/video coding schemes are presented to validate the effectiveness of the proposed JRD model. Yun Zhang 0002, Haoqin Lin, Jing Sun 0010, Linwei Zhu, Sam Kwong |
IEEE Trans. Multim. | 4 |
| 2023 | Adaptive distractor-aware for siamese tracking via enhancement confidence evaluator
Huanlong Zhang, Linwei Zhu, Huaiguang Wu, Yanchun Zhao, Yingzi Lin, Jianwei Zhang 0014 |
Appl. Intell. | 2 |
| 2023 | Atmospheric Scattering Model Induced Statistical Characteristics Estimation for Underwater Image RestorationabstractUnderwater images often suffer from color deviation and low contrast due to selective absorption and light scattering, whose degradation is generally described by an Atmospheric Scattering Model (ASM). However, it is challenging to design hand-craft priors to estimate the transmission map and global light within ASM. To avoid the estimation on these two variables, in this paper, we establish a statistical characteristics relationship between underwater and recovered images based on ASM. With this relationship, a novel lightweight model is proposed for efficient Underwater Image Restoration (UIR). Within our proposed model, the UIR problem is disentangled into global restoration and local compensation, for which two modules are developed. Extensive experimental results demonstrate that our proposed method can effectively improve color deviation and low contrast while preserving details, and outperform state-of-the-art methods. Shuaibo Gao, Wenhui Wu 0001, Hua Li 0012, Linwei Zhu, Xu Wang 0006 |
IEEE Signal Process. Lett. | 4 |
| 2023 | Deep Learning-Based Intra Mode Derivation for Versatile Video CodingabstractIn intra coding, Rate Distortion Optimization (RDO) is performed to achieve the optimal intra mode from a pre-defined candidate list. The optimal intra mode is also required to be encoded and transmitted to the decoder side besides the residual signal, where lots of coding bits are consumed. To further improve the performance of intra coding in Versatile Video Coding (VVC) , an intelligent intra mode derivation method is proposed in this paper, termed as Deep Learning based Intra Mode Derivation (DLIMD) . In specific, the process of intra mode derivation is formulated as a multi-class classification task, which aims to skip the module of intra mode signaling for coding bits reduction. The architecture of DLIMD is developed to adapt to different quantization parameter settings and variable coding blocks including non-square ones, where only one single trained model is required. Different from the existing deep learning based classification problems, the hand-crafted features are also fed into intra mode derivation network besides the learned features from feature learning network. To compete with traditional methods, one additional binary flag is utilized in the video codec to indicate the selected scheme with RDO. Extensive experimental results reveal that the proposed method can achieve 2.28%, 1.74%, and 2.18% bit rate reduction on average for Y, U, and V components on the platform of VVC test model, which outperforms the state-of-the-art works. Linwei Zhu, Yun Zhang 0002, Na Li 0015, Gangyi Jiang, Sam Kwong |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | Texture-Aware Spherical Rotation for High Efficiency Omnidirectional Intra Video CodingabstractTo adapt to the existing video coding standards, omnidirectional videos are usually projected from Three-Dimensional (3D) sphere to Two-Dimensional (2D) plane. However, this projection will cause geometrical stretching distortion and boundary discontinuity, which may degrade coding efficiency. In this paper, we present a Spherical Rotation based Omnidirectional Video Coding (SROVC) method, which exploits the textural properties of omnidirectional videos with spherical rotation. Firstly, SROVC framework is presented and Full-traversal Spherical Rotation (FSR) is developed to derive the optimal rotation angle with frame-level Rate Distortion Optimization (RDO). Secondly, to achieve comparable coding gains and lower computational complexity when compared with FSR, a Texture-aware Spherical Rotation (TSR) method is proposed to predict the rotation angle. Finally, to further reduce complexity and maintain coding efficiency, a Group-oriented TSR (G-TSR) approach is presented, in which the group length is statistically determined. Extensive experiments demonstrate that the proposed TSR and G-TSR schemes can achieve bit rate reductions up to 4.38%, 0.94% and 0.89% on average for CubeMap Projection (CMP) based high efficiency omnidirectional video coding. Additionally, the TSR scheme achieves bit rate saving from 0.91% to 1.19% on average under three more CMP-based projection formats, and 1.89% for joint rotation of X, Y, and Z axes. Jinyong Pi, Yun Zhang 0002, Linwei Zhu, Jinzhi Lin, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Deep Learning-Based Perceptual Video Quality Enhancement for 3D Synthesized ViewabstractDue to occlusion among views and temporal inconsistency in depth video, spatio-temporal distortion occurs in 3D synthesized video with depth image-based rendering. In this paper, we propose a deep Convolutional Neural Network (CNN)-based synthesized video denoising algorithm to reduce temporal flicker distortion and improve perceptual quality of 3D synthesized video. First, we analyze the spatio-temporal distortion, and model eliminating spatio-temporal distortion as a perceptual video denoising problem. Then, a deep learning-based synthesized video denoising network is proposed, in which a CNN-friendly spatio-temporal loss function is derived from a synthesized video quality metric and integrated with a single image denoising network architecture. Finally, specific schemes, i.e., specific Synthesized Video Denoising Networks (SynVD-Nets), and a general scheme, i.e., General SynVD-Net (GSynVD-Net), based on existing CNN-based denoising models, are developed to handle synthesized video with different distortion levels more effectively. Experimental results show that the proposed SynVD-Net and GSynVD-Net can outperform deep learning-based counterparts and conventional denoising methods, and significantly enhance perceptual quality of 3D synthesized video. Huan Zhang 0008, Yun Zhang 0002, Linwei Zhu, Weisi Lin |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Circular intra prediction for 360 degree video coding
Linwei Zhu, Yun Zhang 0002, Na Li 0015, Jinyong Pi, Shiqi Wang 0001 |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | Deep Learning-Based Chroma Prediction for Intra Versatile Video CodingabstractColor images always exhibit a high correlation between luma and chroma components. Cross component linear model (CCLM) has been introduced to exploit such correlation for removing redundancy in the on-going video coding standard, i.e., versatile video coding (VVC). To further improve the coding performance, this paper presents a deep learning based intra chroma prediction method, termed as convolutional neural network based chroma prediction (CNNCP). More specifically, the process of chroma prediction is formulated to produce the colorful version from available information input. CNNCP includes two sub-networks for luma down-sampling and chroma prediction, which are jointly optimized to fully exploit spatial and cross component information. In addition, the outputs of CCLM are adopted as chroma initialization for performance enhancement, and the coding distortion level characterized by quantization parameter is fed into the network to release the negative affect from compression artifacts. To further improve the coding performance, the competition is performed between the conventional chroma prediction and CNNCP in terms of rate-distortion cost with a binary flag signalled. The learned CNNCP is incorporated into both video encoder and decoder. Extensive experimental results demonstrate that the proposed scheme can achieve 4.283%, 3.343%, and 4.634% bit rate savings for luma and two chroma components, compared with the VVC test model version 4.0 (VTM 4.0). Linwei Zhu, Yun Zhang 0002, Shiqi Wang 0001, Sam Kwong, Xin Jin 0002, Yu Qiao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Highly Efficient Multiview Depth Coding Based on Histogram Projection and Allowable Depth DistortionabstractMismatches between the precisions of representing the disparity, depth value and rendering position in 3D video systems cause redundancies in depth map representations. In this paper, we propose a highly efficient multiview depth coding scheme based on Depth Histogram Projection (DHP) and Allowable Depth Distortion (ADD) in view synthesis. Firstly, DHP exploits the sparse representation of depth maps generated from stereo matching to reduce the residual error from INTER and INTRA predictions in depth coding. We provide a mathematical foundation for DHP-based lossless depth coding by theoretically analyzing its rate-distortion cost. Then, due to the mismatch between depth value and rendering position, there is a many-to-one mapping relationship between them in view synthesis, which induces the ADD model. Based on this ADD model and DHP, depth coding with lossless view synthesis quality is proposed to further improve the compression performance of depth coding while maintaining the same synthesized video quality. Experimental results reveal that the proposed DHP based depth coding can achieve an average bit rate saving of 20.66% to 19.52% for lossless coding on Multiview High Efficiency Video Coding (MV-HEVC) with different groups of pictures. In addition, our depth coding based on DHP and ADD achieves an average depth bit rate reduction of 46.69%, 34.12% and 28.68% for lossless view synthesis quality when the rendering precision varies from integer, half to quarter pixels, respectively. We obtain similar gains for lossless depth coding on the 3D-HEVC, HEVC Intra coding and JPEG2000 platforms. Yun Zhang 0002, Linwei Zhu, Raouf Hamzaoui, Sam Kwong, Yo-Sung Ho |
IEEE Trans. Image Process. | 2 |
| 2020 | Towards Modality Transferable Visual Information Representation with Optimal Model CompressionabstractCompactly representing the visual signals is of fundamental importance in various image/video-centered applications. Although numerous approaches were developed for improving the image and video coding performance by removing the redundancies within visual signals, much less work has been dedicated to the transformation of the visual signals to another well-established modality for better representation capability. In this paper, we propose a new scheme for visual signal representation that leverages the philosophy of transferable modality. In particular, the deep learning model, which characterizes and absorbs the statistics of the input scene with online training, could be efficiently represented in the sense of rate-utility optimization to serve as the enhancement layer in the bitstream. As such, the overall performance can be further guaranteed by optimizing the new modality incorporated. The proposed framework is implemented on the state-of-the-art video coding standard (i.e., versatile video coding), and significantly better representation capability has been observed based on extensive evaluations. Rongqun Lin, Linwei Zhu, Shiqi Wang 0001, Sam Kwong |
ACM Multimedia | 2 |
| 2020 | Content-aware Hybrid Equi-angular Cubemap Projection for Omnidirectional Video CodingabstractOmnidirectional video is required to be projected from the Three-Dimensional (3D) sphere to a Two-Dimensional (2D) plane before compression due to its spherical characteristics. Therefore, various projection formats have been proposed in recent years. However, these existing projection methods have problems of either oversampling or discontinuous boundary, which penalize the coding performance. Among them, Hybrid Equiangular Cubemap (HEC) projection has achieved significant coding gains by keeping boundary continuity when compared with Equi-Angular Cubemap (EAC) projection. However, the parameters of its mapping function are fixed and cannot adapt to the video contents, which results in non-uniform sampling in certain regions. To address this limitation, a projection method named Content-aware HEC (CHEC) is presented in this paper. In particular, these parameters of mapping function are adaptively achieved by minimizing the projection conversion distortion. Additionally, an omnidirectional video coding framework with adaptive parameters of mapping function is proposed to effectively improve the coding performance. Experimental results show that the proposed scheme achieves 8.57% and 0.11% bit rate reduction on average in terms of End-to-End Weighted to Spherically uniform Peak Signal to Noise Ratio (E2E WS-PSNR) when compared with Equi-Rectangular Projection (ERP) and HEC projections, respectively. Jinyong Pi, Yun Zhang 0002, Linwei Zhu, Xinju Wu, Xuemei Zhou |
VCIP | 3 |
| 2020 | Sparse Representation-Based Intra Prediction for Lossless/Near Lossless Video CodingabstractIn this paper, a novel intra prediction method is presented for lossless/near lossless High Efficiency Video Coding (HEVC), termed as Sparse Representation based Intra Prediction (SRIP). In specific, the existing Angular Intra Prediction (AIP) modes in HEVC are organized as a mode dictionary, which is utilized to sparsely represent the visual signal by minimizing the difference with respect to the ground truth. For the match of encoding and decoding, the sparse coefficients are also required to be encoded and transmitted to the decoder side. To further improve the coding performance, an additional binary flag is included in the video codec to indicate which strategy is finally adopted with the rate distortion optimization, i.e., SRIP or traditional AIP. Extensive experimental results reveal that the proposed method can achieve 0.36% bit rate saving on average in case of lossless scenario. Linwei Zhu, Yun Zhang 0002, Na Li 0015, Jinyong Pi, Xinju Wu |
VCIP | 1 |
| 2020 | Asynchronous Acoustic Localization and Tracking for Mobile TargetsabstractRecently, acoustic-based indoor localization has attracted much attention due to its affordable infrastructure costs and high localization accuracy. However, previous work is infeasible in mobile target tracking for its long latency in obtaining sufficient beacon messages. In addition, the performance can further deteriorate due to device diversity, varying channel gains, and background noises. To this end, we propose an asynchronous acoustic localization and tracking system (AALTS), which utilizes distributed acoustic anchor nodes to locate passive off-the-shelf mobile devices. In AALTS, we propose an orthogonal chirp spread spectrum (OCSS) modulation technique, which doubles the data rate and thus mitigates the latency. We design a more robust method to capture acoustic signals which embody timestamps for localization, accounting for device diversity, varying channel gains, and the multipath effect. Finally, we incorporate an acoustic Doppler speed estimation module with a path-based particle filter framework to accurately track the moving targets. We have evaluated AALTS in an indoor testbed of size 8×12 m2with commodity mobile phones and customized acoustic anchors. Our evaluation results demonstrate remarkable performance: AALTS achieves 90-percentile tracking errors of 0.49 m for mobile targets and a median of 0.12 m for stationary ones with only four anchor nodes. Chao Cai 0001, Rong Zheng 0001, Jun Li 0067, Linwei Zhu, Henglin Pu, Menglan Hu |
IEEE Internet Things J. | 4 |
| 2020 | Generative Adversarial Network-Based Intra Prediction for Video CodingabstractIn this paper, a novel intra prediction method is proposed to improve the video coding performance, in which the generative adversarial network (GAN) is adopted to intelligently remove the spatial redundancy with the inference process. The proposed GAN-based method improves the prediction by exploiting more information and generating more flexible prediction patterns. In particular, the intra prediction is modeled as an inpainting task, which is accomplished with the GAN model to fill in the missing part by conditioning on the available reconstructed pixels. As such, the learned GAN model is incorporated into both video encoder and decoder, and the rate-distortion optimization is performed for the competition between GAN-based intra prediction and traditional angular-based intra prediction to achieve better coding performance. The proposed scheme is implemented into the high-efficiency video coding test model (HM 16.17) and the versatile video coding test model (VTM 1.1). The experimental results show that the proposed algorithm can achieve 6.6%, 7.5%, and 7.5% under HM 16.17 and 6.75%, 7.63%, and 7.65% under VTM 1.1 bit rate savings on average for luma and chroma components in the intra coding scenario. Linwei Zhu, Sam Kwong, Yun Zhang 0002, Shiqi Wang 0001, Xu Wang 0006 |
IEEE Trans. Multim. | 1 |
| 2020 | HackMan: hacking commodity millimeter-wave hardware for a measurement study
Chao Cai 0001, Jun Luo 0001, Linwei Zhu, Menglan Hu |
Wirel. Networks | 4 |
| 2019 | Fast Coding Unit Decision for Intra Screen Content Coding Based on Ensemble LearningabstractThe Screen Content Coding (SCC) is an extension of High Efficiency Video Coding (HEVC), and it achieves significant improvement on compression ratio. However, the obtained coding efficiency is at the cost of high computational complexity. In this paper, to reduce the computation complexity, we propose to use an ensemble classifier for predicting the coding unit (CU) in intra-coding. Firstly, the L1-loss based linear support vector machine (SVM) is employed as basic classifier for its simplicity. Then, a bagging scheme is applied to train the linear classifiers and boost the prediction accuracy by ensemble learning. Compared with the reference software SCM-5.0, the proposed scheme can achieve 30% complexity reduction on average with only 1.64% bit rates increase. Yali Xue, Xu Wang 0006, Linwei Zhu, Zhaoqing Pan, Sam Kwong |
ICASSP | 3 |
| 2019 | Reinforcement learning based coding unit early termination algorithm for high efficiency video coding
Na Li 0015, Yun Zhang 0002, Linwei Zhu, Wenhan Luo, Sam Kwong |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Convolutional Neural Network-Based Synthesized View Quality Enhancement for 3D Video CodingabstractThe quality of synthesized view plays an important role in the three dimensional (3D) video system. In this paper, to further improve the coding efficiency, a convolutional neural network (CNN) based synthesized view quality enhancement method for 3D High Efficiency Video Coding (HEVC) is proposed. Firstly, the distortion elimination in synthesized view is formulated as an image restoration task with the aim to reconstruct the latent distortion free synthesized image. Secondly, the learned CNN models are incorporated into 3D HEVC codec to improve the view synthesis performance for both view synthesis optimization (VSO) and the final synthesized view, where the geometric and compression distortions are considered according to the specific characteristics of synthesized view. Thirdly, a new Lagrange multiplier in the rate-distortion (RD) cost function is derived to adapt the CNN based VSO process to embrace a better 3D video coding performance. Extensive experimental results show that the proposed scheme can efficiently eliminate the artifacts in the synthesized image, and reduce 25.9% and 11.7% bit rate in terms of peak-signal-to-noise ratio (PSNR) and structural similarity (SSIM) index, which significantly outperforms the state-of-theart methods. Linwei Zhu, Yun Zhang 0002, Shiqi Wang 0001, Hui Yuan 0001, Sam Kwong, Horace Ho-Shing Ip |
IEEE Trans. Image Process. | 1 |
| 2017 | Multi-class ranking based most probable prediction unit selection for HEVC encodingabstractIn this paper, an incremental learning based multi-class Prediction Units (PUs) ranking approach is presented for High Efficiency Video Coding (HEVC) Rate-Distortion-Complexity (RDC) optimization. In particular, the process of PUs selection is formulated as a binary classification plus multi-class ranking task, and incremental learning is applied for classifier training to better exploit the information in the emerging training data. Furthermore, the proposed most probable PUs selection scheme is incorporated into a joint RDC optimization framework, where the complexity can be flexibly allocated targeting at minimizing computational cost under a constrained RD performance degradation. Experimental results demonstrate that the proposed approach can reduce 53.7% and 50.4% computational complexity on average under low delay P and random access configurations with ignorable RD performance degradation, which outperforms the state-of-the-art approaches in terms of RDC performance. Linwei Zhu, Sam Kwong, Yun Zhang 0002, Xu Wang 0006, Shiqi Wang 0001 |
VCIP | 1 |
| 2017 | Allowable depth distortion based fast mode decision and reference frame selection for 3D depth coding
Yun Zhang 0002, Zhaoqing Pan, Yang Zhou 0052, Linwei Zhu |
Multim. Tools Appl. | 4 |
| 2016 | Allowable depth distortion based depth filtering for 3D high efficiency video codingabstractDepth videos shall be efficiently compressed and transmitted to the client for view synthesis in Three-Dimensional (3D) video system. Since depth video may contain noise that reduce the coding efficiency, we propose a depth filtering algorithm for 3D depth coding, which exploits the Allowable Depth Distortion (ADD) in view synthesis and is able to improve the coding performance of the depth encoder. Firstly, the depth values has the same rendering position based on the ADD model are clustered. Then, the clustered depth are filtered and set to the optimal depth value for each group by minimizing the view synthesis error. The filtered depth videos are smoother and can be more effectively compressed by the existing 3D High Efficiency Video Coding (HEVC) depth encoder. Experimental results show that the proposed depth filtering method can assist the depth encoder achieve 5.87% bit rate reduction in terms of Bjonteggard Delta Bit Rate (BDBR) and 0.25dB quality gain in terms of Bjonteggard Delta Peak-Signal-to-Noise Ratio (BDPSNR) on average as compared with that of coding the original depth maps. Yun Zhang 0002, Linwei Zhu, Xiangkai Liu, Gangyi Jiang |
ISCAS | 2 |
| 2016 | Machine learning based fast H.264/AVC to HEVC transcoding exploiting block partition similarity
Linwei Zhu, Yun Zhang 0002, Na Li 0015, Gangyi Jiang, Sam Kwong |
J. Vis. Commun. Image Represent. | 1 |
| 2015 | View synthesis distortion elimination filter for depth video coding in 3D video broadcasting
Linwei Zhu, Yun Zhang 0002, Xu Wang 0006, Sam Kwong |
Multim. Tools Appl. | 1 |
| 2013 | Analysis of Emotion and Engagement in a STEM Alternate Reality Game
Yu-Han Chang, Rajiv T. Maheswaran, Jihie Kim, Linwei Zhu |
AIED | 4 |
| 2013 | View-spatial-temporal post-refinement for view synthesis in 3D video systems
Linwei Zhu, Yun Zhang 0002, Mei Yu 0001, Gangyi Jiang, Sam Kwong |
Signal Process. Image Commun. | 1 |