EDBT 2026 Demo / reviewers in the wild / expert
Lu Yu 0003
dblp:04/1781-3
· DBLP profile ↗
180ranked-venue papers
2as first author
79since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 119 · 1 first-author · 50 since 2021Systems, architecture and hardware · 25 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 16 · 13 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Computer networks · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MPEG Explorations Toward 3D Gaussian Splat Coding and Standardizationabstract3D Gaussian splats (3DGS) have rapidly gained traction as a 3D scene representation technique that enables efficient real-time rendering and highfidelity novel view synthesis. This paper reports on ongoing MPEG GSC efforts conducted jointly by the MPEG Video Coding group (WG 4) and the Coding of 3D Graphics and Haptics group (WG 7) to define a practical and interoperable compression framework for 3DGS. MPEG GSC is planning GSC standardization with short-term and long-term timelines to address market requirements. The short-term objective is to standardize coding tools that build on proven MPEG ecosystems while introducing only the minimal set of extensions, syntax, and processing required for INRIA-3DGS format (referred to in MPEG as I-3DGS). In the long-term, MPEG is also investigating broader alternatives for 3DGS representation and compression, including approaches that integrate training during compression, collectively referred to as Alternative-3DGS (A-3DGS). More specifically, this paper focuses on the I-3DGS and explores both geometrybased and video-based coding frameworks within MPEG GSC. Gun Bang, Yiyi Liao, Alexandre Zaghetto, Marius Preda, Lu Yu 0003 |
DCC | 5 |
| 2026 | Accelerating Entropy Coding via Probability State Caching Mechanism (PSCM)
Lu Yu 0003 |
ISCAS | 3 |
| 2026 | Enhanced RDOQ for Dependent Quantization in VVC
Lu Yu 0003 |
ISCAS | 3 |
| 2026 | Training-Free Adaptation from Visible to Thermal Domain for Learned Image Compression
Jiawang Liu, Hualong Yu, Lu Yu 0003 |
ISCAS | 3 |
| 2026 | Dual-Hilbert Scan for Efficient Video-Based Gaussian Splatting Compression
Yiyi Liao, Lu Yu 0003 |
ISCAS | 3 |
| 2026 | MLANet: Multilevel aggregation network for binocular eye-fixation prediction
Wujie Zhou, Jiabao Ma, Yulai Zhang, Lu Yu 0003, Weijia Gao, Ting Luo 0001 |
Signal Process. Image Commun. | 4 |
| 2026 | GSCodec Studio: A Modular Framework for Gaussian Splat Compressionabstract3D Gaussian Splatting and its extension to 4D dynamic scenes enable photorealistic, real-time rendering from real-world captures, positioning Gaussian Splats (GS) as a promising format for next-generation immersive media. However, their high storage requirements pose significant challenges for practical use in sharing, transmission, and storage. Despite various studies exploring GS compression from different perspectives, these efforts remain scattered across separate repositories, complicating benchmarking and the integration of best practices. To address this gap, we present GSCodec Studio, a unified and modular framework for GS reconstruction, compression, and rendering. The framework incorporates a diverse set of 3D/4D GS reconstruction methods and GS compression techniques as modular components, facilitating flexible combinations and comprehensive comparisons. By integrating best practices from community research and our own explorations, GSCodec Studio supports the development of compact representation and compression solutions for static and dynamic Gaussian Splats. Specifically, we present Static and Dynamic GSCodec: Static GSCodec achieves competitive 3D Gaussian Splat rate-distortion performance with low decoding complexity, while Dynamic GSCodec delivers advanced 4D Gaussian Splat compression performance. The code for our framework is publicly available at https://github.com/JasonLSC/GSCodec_Studio, to advance the research on Gaussian Splats compression. Sicheng Li 0003, Chengzhen Wu, Hao Li 0069, Yiyi Liao, Lu Yu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Low-Rank Approximation for Efficient Compression of Gaussian Splatting Spherical Harmonicsabstract3D Gaussian Splatting (3DGS) enables real-time, high-fidelity rendering but suffers from large model sizes, mainly due to the Spherical Harmonics (SH) coefficients used for view-dependent appearance modeling, whose parameter count scales quadratically(O(L2))with the degreeL. This paper presents the first systematic study that reveals and leverages the intrinsic low-rank structure of SH coefficients in 3DGS. Unlike prior approaches that truncate spectral energy, the proposed low-rank paradigm compactly preserves spectral information, achieving high visual quality with substantially reduced storage. Two complementary approaches are introduced. (1) SHAC-PCA (Principal Component Analysis) is a plug-and-play post-hoc compressor that retains principal spectral variance for high-fidelity compression. (2) SHAC-LST (Learned Subset Transformation) is a training-integrated approach that decomposes Alternating Current (AC) of SH coefficients into low-dimensional subset coefficients and a shared transformation matrix, offering superior compression and even quality improvements through regularization. Both methods effectively reduce SH coefficients storage complexity toO(L). Extensive experiments demonstrate that our approaches significantly reduce memory usage while maintaining or even improving rendering quality. The proposed techniques are highly versatile: SHAC-PCA can be applied to any pre-trained 3DGS model, while SHAC-LST supports end-to-end training or fine-tuning. Both methods are compatible with existing 3DGS compression pipelines, providing a practical and general solution for efficient compression of SH coefficients in 3DGS. Sicheng Li 0003, Yiyi Liao, Lu Yu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | GIFStream: 4D Gaussian-based Immersive Video with Feature StreamabstractImmersive video offers a 6-Dof-Free viewing experience, potentially playing a key role in future video technology. Recently, 4D Gaussian Splatting has gained attention as an effective approach for immersive video due to its high rendering efficiency and quality, though maintaining quality with manageable storage remains challenging. To address this, we introduce GIFStream, a novel 4D Gaussian representation using a canonical space and a deformation field enhanced with time-dependent feature streams. These feature streams enable complex motion modeling and allow efficient compression by leveraging their motion-awareness and temporal correspondence. Additionally, we incorporate both temporal and spatial compression networks for end-to-end compression. Experimental results show that GIFStream delivers high-quality immersive video at 30 Mbps, with real-time rendering and fast decoding on an RTX 4090. Hao Li 0069, Sicheng Li 0003, Abudouaihati Batuer, Lu Yu 0003, Yiyi Liao |
CVPR | 5 |
| 2025 | Assessing the Reusability of Cloud-Received Feature Streams on Advanced NetworksabstractThe goal of this paper is to raise awareness of challenges and opportunities in the Collaborative Intelligence (CI) field and promote research on related standards. We begin by identifying a key challenge in CI applications, i.e., is it still possible for feature streams received in the cloud to be reused in the future by more advanced multitasking networks to achieve effective task accuracy? We then propose a framework to explore the generalization ability of cloud-received feature streams on more advanced networks from a coarse-grained to a fine-grained manner. We design a series of adapters of varying complexity to further explore the potential of feature streams for task network adaptation. Experiments show that sharing feature streams across multiple task networks could achieve an average of nearly 80% bitrate saving compared to Versatile Video Coding (VVC), which demonstrates the reuse potential of cloud-received feature streams. In addition, we make theoretical inferences about the adaptation range of shared feature streams, especially for those networks with high precision. Jiawang Liu, Hualong Yu, Heming Sun, Lu Yu 0003 |
ISCAS | 5 |
| 2025 | Learning-based Image Coding for Machine Intelligence with Variable-RateabstractImage Coding for Machines (ICM) has yielded significant developments recently. Variable-rate support is necessary for image coding, while performance gap still exists, in learning-based image coding, between the single-model and multiple-fixed-models methods. This paper proposes a Machine Intelligence Variable-Rate Codec (MIVRCodec) with single-model method. We introduce a method to generate, compress, and utilize image semantic feature information, enabling the codec to adaptively process different semantic content of the image. Additionally, current studies employ fixed methods to remove redundant information between luminance and chrominance components, neglecting the dynamic characteristics of this redundancy and leading to its inappropriate utilization. We further propose a Color Dynamic Fusion Module (CDFM), which adaptively fuses image color component features based on various conditions (e.g., bitrate and image content) to utilize the redundancy among image color components as appropriately as possible. Lastly, we propose a Progressive Training Strategy (PTS) for training MIVRCodec. These proposed methods not only reduce performance loss in variable-rate ICM but also improve baseline performance. Experimental results demonstrate that our proposed MIVRCodec works well in the bitrate range corresponding to meaningful accuracy intervals in machine intelligence tasks using a single model, achieving coding efficiency on par with multiple fixed-rate models and surpassing existing state-of-the-art codecs. Hualong Yu, Jiawang Liu, Qiqi He, Lu Yu 0003 |
ISCAS | 5 |
| 2025 | Eliminating Geometric Representation Redundancy for 3D Gaussian Splat Codingabstract3D Gaussian Splatting (3DGS) enables photorealistic, real-time rendering, yet its native representation imposes significant storage and transmission overhead, hindering widespread deployment. Most existing approaches focus on reducing data volume or improving the efficiency of lossy coding. However, they overlook the high data entropy caused by inherent representation ambiguity, where multiple geometric attribute values can define the same geometry. To address this, we introduce a lightweight, plug-and-play preprocessing method that lowers raw data entropy by canonicalizing scale and quaternion attributes. Specifically, our method first resolves geometric representation ambiguity via a deterministic regularization rule that enforces a unique representation; reduces dimensionality by converting 4D quaternions to minimal 3D Rodrigues parameters; and addresses numerical redundancy by clamping perceptually insignificant scale values. Our method is a generic, plug-and-play preprocessing module, fully orthogonal to existing 3DGS compression schemes, and effectively boosts their coding efficiency. When combined with a baseline compression pipeline, it yields an average BD-Rate reduction of 15.81% compared to the same pipeline without our preprocessing. Shanchuan Liu, Sicheng Li 0003, Yiyi Liao, Lu Yu 0003 |
VCIP | 5 |
| 2025 | A Subjective and Objective Study on HDR Video Quality Assessment for Learned CompressionabstractHigh Dynamic Range (HDR) video provides a more realistic and immersive visual experience compared to Standard Dynamic Range (SDR) video, which is typically compressed by codecs for transmission and storage. Recently, deep learning-based codecs have demonstrated superior compression performance over traditional codecs. However, they also introduce distinctive artifacts, posing new challenges for video quality assessment (VQA) metrics. Existing VQA benchmarks are primarily designed for traditional codecs and thus fail to assess the capability of VQA models in detecting artifacts specific to learning-based codecs. To address this gap, we conducted a pioneering subjective quality assessment study of HDR videos compressed using both conventional and learning-based approaches. Eight HDR 10-bit video sequences were encoded with six different codecs, yielding 178 videos under various settings and over 4,800 subjective quality ratings. Based on subjective scores, we evaluate five full-reference objective video quality assessment metrics (i.e., PSNR, VMAF, SSIM, ColorVideoVDP, and HDRMAX+VMAF). Our findings reveal that: 1) existing metrics present inadequate generalization capabilities across codecs, with notably lower performance on learning-based codecs compared to traditional ones; 2) current HDR metrics have not demonstrated significant advantages over SDR metrics, highlighting the need for further research to advance HDR metric development. The resources are available at: https://github.com/blindwang/ZJUHDR. Lu Yu 0003 |
VCIP | 3 |
| 2025 | Neural mesh refinementabstractSubdivision is a widely used technique for mesh refinement. Classic methods rely on fixed manually defined weighting rules and struggle to generate a finer mesh with appropriate details, while advanced neural subdivision methods achieve data-driven nonlinear subdivision but lack robustness, suffering from limited subdivision levels and artifacts on novel shapes. To address these issues, this paper introduces a neural mesh refinement (NMR) method that uses the geometric structural priors learned from fine meshes to adaptively refine coarse meshes through subdivision, demonstrating robust generalization. Our key insight is that it is necessary to disentangle the network from non-structural information such as scale, rotation, and translation, enabling the network to focus on learning and applying the structural priors of local patches for adaptive refinement. For this purpose, we introduce an intrinsic structure descriptor and a locally adaptive neural filter. The intrinsic structure descriptor excludes the non-structural information to align local patches, thereby stabilizing the input feature space and enabling the network to robustly extract structural priors. The proposed neural filter, using a graph attention mechanism, extracts local structural features and adapts learned priors to local patches. Additionally, we observe that Charbonnier loss can alleviate over-smoothing compared to L2 loss. By combining these design choices, our method gains robust geometric learning and locally adaptive capabilities, enhancing generalization to various situations such as unseen shapes and arbitrary refinement levels. We evaluate our method on a diverse set of complex three-dimensional (3D) shapes, and experimental results show that it outperforms existing subdivision methods in terms of geometry quality. See https://zhuzhiwei99.github.io/NeuralMeshRefinement for the project page. Lu Yu 0003, Yiyi Liao |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2025 | AESeg: Affinity-enhanced segmenter using feature class mapping knowledge distillation for efficient RGB-D semantic segmentation of indoor scenes
Wujie Zhou, Yuxiang Xiao, Fangfang Qiang, Xiena Dong, Caie Xu, Lu Yu 0003 |
Neural Networks | 6 |
| 2025 | Q-LIC: Quantizing Learned Image Compression With Channel SplittingabstractLearned image compression (LIC) has reached a comparable coding gain with traditional hand-crafted methods such as VVC intra. However, the large network complexity prohibits the usage of LIC on resource-limited embedded systems. Network quantization is an efficient way to reduce the network burden. This paper presents a quantized LIC (QLIC) by channel splitting. First, we explore that the influence of quantization error to the reconstruction error is different for various channels. Second, we split the channels whose quantization has larger influence to the reconstruction error. After the splitting, the dynamic range of channels is reduced so that the quantization error can be reduced. Finally, we prune several channels to keep the number of overall channels as origin. By using the proposal, in the case of 8-bit quantization for weight and activation of both main and hyper path, we can reduce the BD-rate by 0.61%-4.74% compared with the previous QLIC. Besides, we can reach better coding gain compared with the state-of-the-art network quantization method when quantizing MS-SSIM models. Moreover, our proposal can be combined with other network quantization methods to further improve the coding gain. The moderate coding loss caused by the quantization validates the feasibility of the hardware implementation for QLIC in the future. Heming Sun, Lu Yu 0003, Jiro Katto |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | IRFR-Net: Interactive Recursive Feature-Reshaping Network for Detecting Salient Objects in RGB-D ImagesabstractUsing attention mechanisms in saliency detection networks enables effective feature extraction, and using linear methods can promote proper feature fusion, as verified in numerous existing models. Current networks usually combine depth maps with red-green-blue (RGB) images for salient object detection (SOD). However, fully leveraging depth information complementary to RGB information by accurately highlighting salient objects deserves further study. We combine a gated attention mechanism and a linear fusion method to construct a dual-stream interactive recursive feature-reshaping network (IRFR-Net). The streams for RGB and depth data communicate through a backbone encoder to thoroughly extract complementary information. First, we design a context extraction module (CEM) to obtain low-level depth foreground information. Subsequently, the gated attention fusion module (GAFM) is applied to the RGB depth (RGB-D) information to obtain advantageous structural and spatial fusion features. Then, adjacent depth information is globally integrated to obtain complementary context features. We also introduce a weighted atrous spatial pyramid pooling (WASPP) module to extract the multiscale local information of depth features. Finally, global and local features are fused in a bottom-up scheme to effectively highlight salient objects. Comprehensive experiments on eight representative datasets demonstrate that the proposed IRFR-Net outperforms 11 state-of-the-art (SOTA) RGB-D approaches in various evaluation indicators. Wujie Zhou, Qinling Guo, Jingsheng Lei, Lu Yu 0003, Jenq-Neng Hwang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | NeRFCodec: Neural Feature Compression Meets Neural Radiance Fields for Memory-Efficient Scene Representation
Sicheng Li 0003, Hao Li 0069, Yiyi Liao, Lu Yu 0003 |
CVPR | 4 |
| 2024 | Perceptual Image Compression with Text-Guided Multi-level Fusion
Jiaqi Hu 0005, Jiedong Zhuang, Lu Yu 0003, Haoji Hu |
PRCV (5) | 5 |
| 2024 | PET-NeRV: Bridging Generalized Video Codec and Content-Specific Neural RepresentationabstractGeneralized models dominate neural video compression methods to compress arbitrary video with the same model. However, obtaining a universal neural codec with high compression efficiency on all videos is challenging. Existing work attempts to adapt the decoder-side model per video content through full parameter tuning, but this requires a large Group of Pictures (GOP) to compensate for the cost of transmitting updated parameters. To tackle this challenge, we propose to tune generalized video codecs per video content in a parameter-efficient manner. The per-content tuned parameters are further compressed with entropy coding using adaptive distribution estimations. This allows for enhancing the compression efficiency while maintaining a normal GOP size for random access capabilities. To validate the generality and validity of our approach, we apply it to two representative methods: CNN-based, DCVC-HEM, and Transformer-based, VCT. Our results demonstrate that introducing content-specific representation leads to a notable improvement in compression efficiency compared to the original methods. Hao Li 0069, Lu Yu 0003, Yiyi Liao |
VCIP | 2 |
| 2024 | Distance-based feature repack algorithm for video coding for machines
Yuan Zhang 0023, Xiaoli Gong, Hualong Yu, Lu Yu 0003 |
J. Vis. Commun. Image Represent. | 5 |
| 2024 | PGGNet: Pyramid gradual-guidance network for RGB-D indoor scene semantic segmentation
Wujie Zhou, Gao Xu, Meixin Fang, Shanshan Mao, Rongwang Yang, Lu Yu 0003 |
Signal Process. Image Commun. | 6 |
| 2024 | CMPFFNet: Cross-Modal and Progressive Feature Fusion Network for RGB-D Indoor Scene Semantic SegmentationabstractDepth information can contribute to the semantic segmentation of scenes from red–green–blue (RGB) images. Therefore, the amount of information that can be obtained from RGB and RGB-depth (RGB-D) images is significantly greater for this task. However, RGB and RGB-D modalities are different in terms of object representation. Features that are extracted from these modalities and fused effectively are key to scene semantic segmentation. In addition, complete segmentation requires the fusion of multiscale features to unify global information. However, existing approaches primarily use multiscale features for sequential integration. This study introduces a cross-modal and progressive feature fusion network (CMPFFNet) for semantic segmentation of indoor scenes in RGB-D images. First, a multimodal adaptive alignment fusion (MAAF) module based on an attention mechanism is introduced. This module aligns the two modal channels by additive attention and then computes the spatial similarity between the two modalities based on the dot product to incorporate the complementary information of the depth modality into the RGB modality. In addition, a reverse attention augmentation (RAA) module is introduced to augment the more abstract high-level features for two adjacent multilevel features using the concrete semantic information of the lower-level features in them. After augmenting the extracted multilevel features, a multilevel feature progressive fusion (MFPF) module is deployed; this module sequentially fuses the neighboring two features progressively with emphasis on the spatial semantics. The network uses the Segformer network with high performance as a backbone in multiple computer vision tasks to enhance the segmentation capability. Experimental results obtained from two publicly available datasets of indoor scenes reveal that the proposed CMPFFNet outperforms existing models in semantic segmentation of indoor scenes of RGB-D images.Note to Practitioners—This study introduces a cross-modal and progressive feature fusion network (CMPFFNet) for indoor scene semantic segmentation in RGB-D images. The complementary information of the depth modality is incorporated into the RGB modality in both channel and spatial forms to form a discriminative representation for easy segmentation. A multilevel feature aggregation decoder is proposed to predict the results of semantic segmentation of scenes. The network uses the Segformer network with high performance as a backbone in multiple computer vision tasks to enhance the segmentation capability. Wujie Zhou, Yuxiang Xiao, Weiqing Yan, Lu Yu 0003 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2024 | Guest Editorial Special Section on Recent Standardization Efforts for Learning-Based Visual Data CodingabstractVisual data coding is an enabling technology for various applications and is now ubiquitously adopted in modern image processing, communications, and computer vision systems. To enable interoperability between devices manufactured and services provided by different enterprises, a series of standards targeting visual data coding have been crafted in the past three decades. Several standardization organizations, such as ISO/IEC JTC 1/SC 29 consisting of Joint Picture Experts Group (JPEG) and Moving Picture Experts Group (MPEG),1ITU-T SG 16 Video Coding Experts Group (VCEG),2IEEE Data Compression Standards Committee Audio Video Coding Working Group (1857 WG),3MPAI Community,4have been creating these standards from many contributions of academia and industry. While most of these visual coding standards have been successfully deployed in many applications, there are more challenges nowadays, especially to accommodate the large volume of visual data in limited storage and limited bandwidth transmission links. Compression efficiency improvements are still needed, especially considering emerging data representation formats ranging from 8K/HDR image/video to rich plenoptic data. Dong Liu 0002, Shan Liu 0001, João Ascenso, Dong Tian, Lu Yu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | MC3Net: Multimodality Cross-Guided Compensation Coordination Network for RGB-T Crowd CountingabstractOwing to the expansion in processing of industrial information through advances in machine learning, the demand for accurate crowd counting in various applications is increasing. We propose a multimodality cross-guided compensation coordination network (MC$^{3}$Net) for accurate red–green–blue and thermal (RGB-T) crowd counting. The network includes modules of intricate interactive fusion, feature difference compensation, and complementary attention enhancement. We use ConvNext as the backbone and process the three streams from RGB, thermal, and spliced RGB-T inputs. The multimodality data are sequentially guided and fused hierarchically, fully combining features extracted from the RGB and thermal images. Thereafter, difference compensation is applied to compress fusion and splicing features. Redundant information is removed. Then, feature mismatch is mitigated to enhance complementary information, reduce the loss of details, and finally obtain crowd statistics. Results from extensive experiments on the RGBT-CC dataset indicate the robustness and effectiveness of MC$^{3}$Net, which also achieves high performance on the DroneRGBT dataset and ShanghaiTechRGBD dataset, outperforming existing crowd counting methods. The code and models are available at: https://github.com/WBangG/MC3Net. Wujie Zhou, Jingsheng Lei, Weiqing Yan, Lu Yu 0003 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | UTLNet: Uncertainty-Aware Transformer Localization Network for RGB-Depth Mirror SegmentationabstractMirror segmentation, an emerging discipline in the field of computer vision, involves the identification and marking of mirrors in an image. Current mirror segmentation methods rely on fixed mirror elements as features for object segmentation. However, these methods do not account for the varied quality of feature images obtained under complex real-world conditions, leading to inaccurate segmentation results. To address these limitations, we propose a novel uncertainty-aware transformer localization network (UTLNet) for RGB-D mirror segmentation. Our approach draws inspiration from biomimicry, specifically the behavior pattern of human observation. We aim to explore features from different angles and focus on complex features that are challenging to determine during the coding stage. Additionally, we employ graph convolution to construct complementary dual-modal fusion features. Furthermore, we design a multiscale interaction transformer module using the shifted-window self-attention mechanism to acquire precise position information. In our experiments, the proposed UTLNet surpasses the current state-of-the-art mirror segmentation method as well as alternative task-specific methods. It achieves superior performance across various evaluation scenarios. Wujie Zhou, Yuqi Cai, Weiqing Yan, Lu Yu 0003 |
IEEE Trans. Multim. | 5 |
| 2024 | DHFNet: dual-decoding hierarchical fusion network for RGB-thermal semantic segmentation
Yuqi Cai, Wujie Zhou, Lu Yu 0003, Ting Luo 0001 |
Vis. Comput. | 4 |
| 2023 | SteerNeRF: Accelerating NeRF Rendering via Smooth Viewpoint TrajectoryabstractNeural Radiance Fields (NeRF) have demonstrated superior novel view synthesis performance but are slow at rendering. To speed up the volume rendering process, many acceleration methods have been proposed at the cost of large memory consumption. To push the frontier of the efficiency-memory trade-off, we explore a new perspective to accelerate NeRF rendering, leveraging a key fact that the view-point change is usually smooth and continuous in interactive viewpoint control. This allows us to leverage the information of preceding viewpoints to reduce the number of rendered pixels as well as the number of sampled points along the ray of the remaining pixels. In our pipeline, a low-resolution feature map is rendered first by volume rendering, then a lightweight 2D neural renderer is applied to generate the output image at target resolution leveraging the features of preceding and current frames. We show that the proposed method can achieve competitive rendering quality while reducing the rendering time with little memory overhead, enabling 30FPS at 1080P image resolution with a low memory footprint. Sicheng Li 0003, Hao Li 0069, Yue Wang 0020, Yiyi Liao, Lu Yu 0003 |
CVPR | 5 |
| 2023 | Video Surveillance on Mobile Edge Networks: Exploiting Multi-Exit NetworkabstractVideo surveillance systems are playing increasingly important roles in our everyday lives. To get meaningful surveillance information in a timely and accurate manner, it is vital to optimally allocate computation and communication resources for image classification tasks. In this paper, taking face recognition as an example, we propose a novel end-to-edge collaborative computing system based on a multi-exit network to dynamically allocate computation at the front end (the camera sensor) and back end (the mobile edge computing server). With the ∊-greedy algorithm for reinforcement learning, the decision module decides whether to obtain recognition results from earlier exits at the front end or transmit the feature maps to the back end to obtain more accurate results. The module balances recognition accuracy and time overhead under different channel conditions. Experimental results show that the proposed system can significantly save inference time and maintain competitive accuracy in various communication channel conditions. Yuchen Cao 0005, Siming Fu, Xiaoxuan He, Haoji Hu, Hangguan Shan, Lu Yu 0003 |
ICC | 6 |
| 2023 | VeRi3D: Generative Vertex-based Radiance Fields for 3D Controllable Human Image SynthesisabstractUnsupervised learning of 3D-aware generative adversarial networks has lately made much progress. Some recent work demonstrates promising results of learning human generative models using neural articulated radiance fields, yet their generalization ability and controllability lag behind parametric human models, i.e., they do not perform well when generalizing to novel pose/shape and are not part controllable. To solve these problems, we propose VeRi3D, a generative human vertex-based radiance field parameterized by vertices of the parametric human template, SMPL. We map each 3D point to the local coordinate system defined on its neighboring vertices, and use the corresponding vertex feature and local coordinates for mapping it to color and density values. We demonstrate that our simple approach allows for generating photorealistic human images with free control over camera pose, human pose, shape, as well as enabling part-level editing. Xinya Chen, Jiaxin Huang 0012, Yanrui Bin, Lu Yu 0003, Yiyi Liao |
ICCV | 4 |
| 2023 | Bird's-Eye-View-Based LiDAR Point Cloud Coding For MachinesabstractRecently, there has been a growing interest in image/video coding tailored for intelligent analysis tasks, such as Image/Video Coding for Machines (ICM/VCM). These approaches have shown remarkable results when compared to human perception-based coding methods. However, point cloud coding methods for machine intelligence have not been extensively studied. Inspired by current LiDAR point cloud intelligent analysis methods which convert point cloud into bird’s-eye-view (BEV) perspective, we propose an end-to-end learnt point cloud coding framework for 3D machine intelligent tasks with BEV representation, named PC4M. Specifically, the PC4M system consists of a LiDAR encoder, a learnt BEV feature codec and a BEV region proposal network with a task-specific head. To achieve better rate-distortion performance for analysis tasks, we propose an efficient Res-NeXt fusion block with powerful multi-scale modeling ability to compress sparse BEV features, and design a long-distance adaptive attention module by using VanAtten block. Experimental results demonstrate that our method outperforms the state-of-the-art MPEG standard Geometry-based Point Cloud Coding (G-PCC) on the object detection and BEV map segmentation by 83.12% and 85.32% of BD-rate gain on nuScenes, respectively. To the best of our knowledge, this is the first end-to-end learnt task-oriented point cloud codec. Lu Yu 0003 |
VCIP | 3 |
| 2023 | Evaluation on the generalization of coded features across neural networks of different tasksabstractRecent advances in deep neural networks (DNNs) for computer vision tasks have made intelligent analysis on edge devices more prevalent and practical. To better distribute computational load between edge devices and the cloud, a novel deep learning deployment strategy called Collaborative Intelligence (CI) has been proposed. In this strategy, features extracted from edge devices are first compressed and then transmitted to the cloud. However, it is unclear whether these compressed features have enough information to perform diverse downstream tasks. This paper focuses on the generalization of compressed features from one neural network among other object detection and instance segmentation task networks. We first propose a scheme to evaluate the generalization of features and further perform experiments on feature compression. Our experiments show that the extracted features contain enough information for other task networks and feature compression scheme for multi-task networks offers a 82.04% average bitrate saving compared to VVC. Jiawang Liu, Ke Jia, Hualong Yu, Lu Yu 0003 |
VCIP | 5 |
| 2023 | Learned Image Compression With Variable Neural Network ArchitectureabstractWith the development of neural networks, the coding efficiency of learned image compression methods gradually exceeds that of traditional image codecs that are carefully designed and optimized by experts. However, the deep image compression model with an invariable network structure and fixed network parameters can only cover a limited range of bit rates. Using the same network structure for different bit rate ranges leads to sub-optimal coding efficiency. Meanwhile, the more complex network structure and the increase of network parameters cause higher model complexity, which makes it difficult for network training and practical application. In this paper, a new perspective is proposed to solve the above challenges. The idea of variable network architecture is introduced to the field of learned image compression. We devise two strategies with variable network architecture, which improve the coding performance of the learned image compression model from the coding efficiency and complexity respectively. Compared to the image compression method of fixed network structure, our first approach attains an average BDPSNR of 0.771dB and an average BDBR of −14.876%. The second attains an average BDPSNR of 0.320dB and an average BDBR of −5.471%. These results demonstrate the effectiveness of variable network structure in improving coding efficiency and controlling model complexity. Lu Yu 0003 |
VCIP | 2 |
| 2023 | Global contextually guided lightweight network for RGB-thermal urban scene understanding
Tingting Gong, Wujie Zhou, Xiaohong Qian, Jingsheng Lei, Lu Yu 0003 |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | CGINet: Cross-modality grade interaction network for RGB-T crowd counting
Wujie Zhou, Xiaohong Qian, Shanshan Mao, Rongwang Yang, Lu Yu 0003 |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | DRNet: Dual-stage refinement network with boundary inference for RGB-D semantic segmentation of indoor scenes
Enquan Yang, Wujie Zhou, Xiaohong Qian, Jingsheng Lei, Lu Yu 0003 |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | MENet: Lightweight multimodality enhancement network for detecting salient objects in RGB-thermal images
Wujie Zhou, Xiaohong Qian, Jingsheng Lei, Lu Yu 0003, Ting Luo 0001 |
Neurocomputing | 5 |
| 2023 | CCFNet: Cross-Complementary fusion network for RGB-D scene parsing of clothing images
Gao Xu, Wujie Zhou, Xiaohong Qian, Lv Ye, Jingsheng Lei, Lu Yu 0003 |
J. Vis. Commun. Image Represent. | 6 |
| 2023 | AMCFNet: Asymmetric multiscale and crossmodal fusion network for RGB-D semantic segmentation in indoor service robots
Wujie Zhou, Yuchun Yue, Meixin Fang, Shanshan Mao, Rongwang Yang, Lu Yu 0003 |
J. Vis. Commun. Image Represent. | 6 |
| 2023 | Edge Detection Guide Network for Semantic Segmentation of Remote-Sensing ImagesabstractThe acquisition of high-resolution satellite and airborne remote sensing images has been significantly simplified due to the rapid development of sensor technology. Several practical applications of high-resolution remote sensing images (HRRSIs) are based on semantic segmentation. However, single-modal HRRSIs are difficult to classify accurately in the complex situation of some scene objects; therefore, the semantic segmentation of multi-source information fusion is gaining popularity. The inherent difference between multimodal features and the semantic gap between multi-level features typically affect the performance of existing multi-mode fusion methods. We propose a multimodal fusion network based on edge detection to address these issues. This method aids multimodal information fusion by utilizing spatial information contained in the boundary. An edge detection guide module is included in the feature extraction stage to realize the boundary information through the fusion of details and semantics between high-level and low-level features. The boundary information is extended into the well-designed multimodal adaptive fusion block (MAFB) to obtain the multimodal fusion features. Furthermore, a residual adaptive fusion block (RAFB) and a spatial position module (SPM) in the feature decoding stage have been designed to fuse multi-level features from the standpoint of local and global dependence. We compared our method to several state-of-the-art (SOTA) methods using the International Society for Photogrammetry and Remote Sensing’s (ISPRS) Vaihingen and Potsdam datasets. The final results demonstrate that our method achieves excellent performance. Jianhui Jin, Wujie Zhou, Rongwang Yang, Lv Ye, Lu Yu 0003 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2023 | Adjacent Bi-Hierarchical Network for Scene Parsing of Remote Sensing ImagesabstractDriven by the rapid development and application of earth observation sensors, the scene parsing of remote sensing images (RSIs) has attracted extensive research attention in recent years. Restricted by the limited local receptive field of successive convolution layers, traditional models of scene parsing cannot effectively and interactively utilize the local-global information and digital surface model (DSM) of RSIs. Comparatively, accurate scene parsing faces more challenges because of unbalanced categories, small targets, and more complex scenes. To address these challenges, herein, we propose a novel adjacent bi-hierarchical network (ABHNet). Specifically, we introduce a DSM-enhanced (DSE) module to excavate characteristic DSM information from DSM images and enhance the red, green, and blue (RGB) features by exploiting informative cues between RGB and DSM modalities. In addition, an adjacent context exploration (ACE) module is proposed, which contains current and adjacent branches. The branches first exploit multiscale complementary characteristics of multilevel features and then integrate these features by applying adjacent exploration. Our model includes five ACE modules—three are deployed to activate detailed features and two obtain deep-guided features. The mutual collaboration of deep and detailed features is more beneficial to the segmentation of small objects. Extensive experiments on two remote sensing benchmark datasets (ISPRS Potsdam and Vaihingen) showed that the proposed ABHNet qualitatively and quantitatively outperformed other methods. Jiabao Ma, Wujie Zhou, Jingsheng Lei, Lu Yu 0003 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2023 | AuxBranch: Binarization residual-aware network design via auxiliary branch search
Siming Fu, Huanpeng Chu, Lu Yu 0003, Zheyang Li, Wenming Tan, Haoji Hu |
Pattern Recognit. | 3 |
| 2023 | LSNet: Lightweight Spatial Boosting Network for Detecting Salient Objects in RGB-Thermal ImagesabstractMost recent methods for RGB (red-green-blue)-thermal salient object detection (SOD) involve several floating-point operations and have numerous parameters, resulting in slow inference, especially on common processors, and impeding their deployment on mobile devices for practical applications. To address these problems, we propose a lightweight spatial boosting network (LSNet) for efficient RGB-thermal SOD with a lightweight MobileNetV2 backbone to replace a conventional backbone (e.g., VGG, ResNet). To improve feature extraction using a lightweight backbone, we propose a boundary boosting algorithm that optimizes the predicted saliency maps and reduces information collapse in low-dimensional features. The algorithm generates boundary maps based on predicted saliency maps without incurring additional calculations or complexity. As multimodality processing is essential for high-performance SOD, we adopt attentive feature distillation and selection and propose semantic and geometric transfer learning to enhance the backbone without increasing the complexity during testing. Experimental results demonstrate that the proposed LSNet achieves state-of-the-art performance compared with 14 RGB-thermal SOD methods on three datasets while improving the numbers of floating-point operations (1.025G) and parameters (5.39M), model size (22.1 MB), and inference speed (9.95 fps for PyTorch, batch size of 1, and Intel i5-7500 processor; 93.53 fps for PyTorch, batch size of 1, and NVIDIA TITAN V graphics processor; 936.68 fps for PyTorch, batch size of 20, and graphics processor; 538.01 fps for TensorRT and batch size of 1; and 903.01 fps for TensorRT/FP16 and batch size of 1). The code and results can be found from the link of https://github.com/zyrant/LSNet. Wujie Zhou, Yun Zhu 0011, Jingsheng Lei, Rongwang Yang, Lu Yu 0003 |
IEEE Trans. Image Process. | 5 |
| 2023 | Embedded Control Gate Fusion and Attention Residual Learning for RGB-Thermal Urban Scene ParsingabstractThe semantic segmentation of road scenes is an important task in autonomous driving. Deep learning has enabled the development of a variety of semantic segmentation networks using RGB and depth data. However, poor lighting conditions and long-distance sensing limit the applicability of RGB and depth cameras. Nevertheless, many existing methods still rely on precise depth maps for scene segmentation. Unlike depth information, thermal imaging provides a visual heat representation that remains accurate under a variety of lighting conditions and over longer distances. For robust and accurate segmentation of scenes collected during autonomous driving, we used the advanced MobileNetV2 network for feature extraction and a fusion strategy with an embedded control gate. In addition, we adopted an encoder–decoder scheme for semantic segmentation and developed an attention residual learning strategy to restore the resolution of the feature map. Finally, semantic and boundary supervision is introduced to optimize parameters of the proposed network. Experimental results show that the proposed network outperforms existing networks on segmentation of urban scenes, and our network can be generalized to depth data. Wujie Zhou, Jingsheng Lei, Lu Yu 0003 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | PGDENet: Progressive Guided Fusion and Depth Enhancement Network for RGB-D Indoor Scene ParsingabstractScene parsing is a fundamental task in computer vision. Various RGB-D (color and depth) scene parsing methods based on fully convolutional networks have achieved excellent performance. However, color and depth information are different in nature and existing methods cannot optimize the cooperation of high-level and low-level information when aggregating modal information, which introduces noise or loss of key information in the aggregated features and generates inaccurate segmentation maps. The features extracted from the depth branch are weak because of the low quality of the depth map, which results in unsatisfactory feature representation. To address these drawbacks, we propose a progressive guided fusion and depth enhancement network (PGDENet) for RGB-D indoor scene parsing. First, high-quality RGB images are used to improve depth data through a depth enhancement module, in which the depth maps are strengthened in terms of channel and spatial correlations. Then, we integrate information from the RGB and enhance depth modalities using a progressive complementary fusion module, in which we start with high-level semantic information and move down layerwise to guide the fusion of adjacent layers while reducing hierarchy-based differences. Extensive experiments are conducted on two public indoor scene datasets, and the results show that the proposed PGDENet outperforms state-of-the-art methods in RGB-D scene parsing. Wujie Zhou, Enquan Yang, Jingsheng Lei, Jian Wan 0001, Lu Yu 0003 |
IEEE Trans. Multim. | 5 |
| 2023 | DBCNet: Dynamic Bilateral Cross-Fusion Network for RGB-T Urban Scene Understanding in Intelligent VehiclesabstractUnderstanding urban scenes is a fundamental capability required of intelligent vehicles. Depth cues provide useful geometric information for semantic segmentation, thus complementing RGB (color) data. Although single-modal RGB images are improved by depth information, semantic segmentation may be degraded in poor-visibility conditions. Thermal imaging can address some limitations of depth data. Therefore, we leverage the multimodal information in RGB-and-thermal (RGB-T) images by introducing a dynamic bilateral cross-fusion network (DBCNet) for RGB-T urban scene understanding. First, RGB-T features extracted by a given backbone are regrouped as high- or low-level features. Second, multimodal high-level features are sent to a dynamic bilateral cross-fusion module for further refinement. Third, a bounded high-level semantic-feature integration module is added to provide feature guidance, and a multitask supervision mechanism is used for fine-tuning. Extensive experiments on two RGB-T urban scene-understanding datasets indicate that DBCNet aggregates multilevel deep features effectively and outperforms state-of-the-art deep-learning scene-understanding methods. Wujie Zhou, Tingting Gong, Jingsheng Lei, Lu Yu 0003 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2022 | Project-Based Learning: Bridging the Gap Between Algorithm and Architecture in Neural Network CourseabstractNeural network has shown its powerful ability in many research fields in the recent years. By using different network structures, many new algorithms are developed to enhance the accuracy. Along with the algorithm development, corresponding architectures are also proposed for the acceleration. However, pure algorithm may not be hardware friendly. As a result, we need to find an optimal trade-off between algorithmic accuracy and architectural efficiency. To help students build the gap between algorithm and architecture, this paper introduces a project-based learning. The project is called learned image compression, which is composed of three phases: algorithm design, architecture mapping and algorithm-architecture co-optimization. Through the project, the students are expected to develop a neural network with high image compression ratio and hardware performance. Furthermore, these kind of knowledge can be extended to any neural network applications. Heming Sun, Lu Yu 0003 |
ISCAS | 2 |
| 2022 | Mutual Learning Inspired Prediction Network for Video Anomaly Detection
Yuan Zhang 0023, Fan Li 0003, Lu Yu 0003 |
PRCV (3) | 4 |
| 2022 | Improving Latent Quantization of Learned Image Compression with Gradient ScalingabstractLearned image compression (LIC) has shown its superior compression ability. Quantization is an inevitable stage to generate quantized latent for the entropy coding. To solve the non-differentiable problem of quantization in the training phase, many differentiable approximated quantization methods have been proposed. However, the derivative of quantized latent to non-quantized latent are set as one in most of the previous methods. As a result, the quantization error between non-quantized and quantized latent is not taken into consideration in the gradient descent. To address this issue, we exploit the gradient scaling method to scale the gradient of non-quantized latent in the back-propagation. The experimental results show that we can outperform the recent LIC quantization methods. Heming Sun, Lu Yu 0003, Jiro Katto |
VCIP | 2 |
| 2022 | Real-time Learned Image Codec on FPGAabstractThis demo paper gives a real-time learned image codec on FPGA. By using Xilinx VCU128, the proposed system reaches 720P@30fps codec, which is 7.76x faster than prior work. Heming Sun, Qingyang Yi, Fangzheng Lin, Lu Yu 0003, Jiro Katto |
VCIP | 4 |
| 2022 | RLLNet: a lightweight remaking learning network for saliency redetection on RGB-D images
Wujie Zhou, Jingsheng Lei, Lu Yu 0003 |
Sci. China Inf. Sci. | 4 |
| 2022 | GCNet: Grid-like context-aware network for RGB-thermal semantic segmentation
Wujie Zhou, Yueli Cui, Lu Yu 0003, Ting Luo 0001 |
Neurocomputing | 4 |
| 2022 | HFNet: Hierarchical feedback network with multilevel atrous spatial pyramid pooling for RGB-D saliency detection
Wujie Zhou, Jingsheng Lei, Lu Yu 0003, Ting Luo 0001 |
Neurocomputing | 4 |
| 2022 | Deep image compression based on multi-scale deformable convolution
Daowen Li, Yingming Li, Heming Sun, Lu Yu 0003 |
J. Vis. Commun. Image Represent. | 4 |
| 2022 | A no-reference perceptual image quality assessment database for learned image codecs
Zhigao Fang, Lu Yu 0003 |
J. Vis. Commun. Image Represent. | 3 |
| 2022 | GEBNet: Graph-Enhancement Branch Network for RGB-T Scene ParsingabstractRGB-T (red–green–blue and thermal) scene parsing has recently drawn considerable research attention. Although existing methods efficiently conduct RGB-T scene parsing, their performance remains limited by a small receptive field. Unlike methods that capture the global context by fusing multiscale features or using an attention mechanism, we propose a graph-enhancement branch network (GEBNet), which uses long-range dependencies obtained from the branch to refine a coarse semantic map generated by the decoder. Semantic and detail modules embedded in the graph-enhancement branch fuse high- and low-level features. Furthermore, inspired by the ability of graph neural networks to capture the global context, we integrate a novel graph-enhancement module into the network branch to obtain global information from both high-level semantic information and low-level details. Results from extensive experiments on the MFNet and PST900 datasets demonstrate the high performance of the proposed GEBNet and the contributions of its main components to the parsing performance. Shaohua Dong, Wujie Zhou, Xiaohong Qian, Lu Yu 0003 |
IEEE Signal Process. Lett. | 4 |
| 2022 | Hierarchical Decoding Network Based on Swin Transformer for Detecting Salient Objects in RGB-T ImagesabstractAlthough conventional deep convolutional neural networks are effective for contextual semantic segmentation of objects, recent vision transformers can capture global information of an image and are better at capturing semantic associations over longer ranges. In addition, some existing saliency detection methods disregard the guidance of high-level semantic information for low-level features during decoding, and only use layer-by-layer transmission for encoding. Therefore, we propose a hierarchical decoding network based on a swin transformer to perform red–green–blue and thermal (RGB-T) salient object detection (SOD). First, a sine–cosine fusion module performs multimodality intersections and exploits complementarity. As a second fusion stage, an advanced semantic information guidance module adjusts high-level semantic information and low-level detailed characteristics. Finally, a global saliency perception module fuses cross-layer information in a top-down path. Comprehensive experiments demonstrate that the proposed network outperforms 12 state-of-the-art methods on three RGB-T SOD datasets. Wujie Zhou, Lv Ye, Lu Yu 0003 |
IEEE Signal Process. Lett. | 4 |
| 2022 | Depth Repeated-Enhancement RGB Network for Rail Surface Defect InspectionabstractSurface defect inspection of railways is important to ensure safe transportation. However, challenging conditions, such as uneven illumination and similar foreground and background, hinder defect inspection. With the development of deep learning and the wide application of the computer vision, defect inspection has made great progress. Accordingly, we propose a depth repeated-enhancement RGB (red–green–blue) network (DRERNet) for rail surface defect inspection. DRERNet fully uses depth and RGB information to better inspect defects on rail surfaces using an encoder–decoder architecture. In the encoder, a novel cross modality enhancement fusion module uses details from RGB maps and location information from depth maps to perform cross-modality fusion. In the decoder, the details and location information in a multimodality complementation module are repeatedly used to progressively refine the DRERNet prediction. We performed extensive experiments, and compared the proposed DRERNet with 10 state-of-the-art methods on the industrial NEU RSDDS-AUG RGB-depth dataset. The comparison results demonstrate that DRERNet consistently performs better than other methods in the all evaluation measures. Wujie Zhou, Weiwei Qiu, Lu Yu 0003 |
IEEE Signal Process. Lett. | 4 |
| 2022 | MGCNet: Multilevel Gated Collaborative Network for RGB-D Semantic Segmentation of Indoor SceneabstractRGB-D semantic segmentation of indoor scenes has long been an enduring research topic. However, because of the intrinsic differences in modal information and large gaps in multi-level feature cues, adopting the traditional U-Net framework provides suboptimal indoor scene segmentation. In this paper, we consider an effective feature exploration approach to achieve accurate segmentation. Specifically, it consists of three steps. First, in the encoder, we design a difference-exploration fusion module, which extracts the difference weights of the two modalities to guide them for fusion, so as to achieve intrinsically consistent feature fusion. The gated decoder module relates to the remaining two steps. Second, we use a gating unit for each level of fusion information to reduce the difference between layers, which also increases the unique distinction of a specific layer while avoiding the exclusion between layers of information. Finally, we use a serial-parallel alternation strategy to increase the ability to capture contextual knowledge. Considering the above three steps, we construct the multilevel gated collaborative network (MGCNet). Extensive experiments indicate the performance of the proposed MGCNet can compete favorably against state-of-the-art models under three standard metrics. Enquan Yang, Wujie Zhou, Xionghong Qian, Lu Yu 0003 |
IEEE Signal Process. Lett. | 4 |
| 2022 | RTLNet: Recursive Triple-Path Learning Network for Scene Parsing of RGB-D ImagesabstractScene parsing approaches have attracted extensive attention in recent years; although several methods have been developed for scene parsing, most include complex modules for both cross-modality fusion between RGB and depth images in the encoder and image scale level recovery in the decoder under label supervision for high inference accuracy. Cross-modality information in the encoder may be diluted when processed through the decoder, and the supervision results may not be reused effectively, which adversely affects scene parsing. To address these problems, we propose a recursive triple-path learning network (RTLNet) for cross-modality interactions in the decoder using global context and cross-modality fusion modules. The proposed modules fully use cross-modality information to reduce information loss. To enhance the robustness of RTLNet, we add a path to reuse the initial predictions from the decoder and introduce a ladder-shaped feature consistency module to further leverage multiscale features. Experiments are conducted with the proposed RTLNet and nine recent RGB-D indoor scene parsing methods on the NYUv2 and SUN-RGBD indoor scene datasets; the results show that the RTLNet outperforms the other methods. Yuchun Yue, Wujie Zhou, Jingsheng Lei, Lu Yu 0003 |
IEEE Signal Process. Lett. | 4 |
| 2022 | Perception-Based Pseudo-Motion Response for 360-Degree Video StreamingabstractStreaming high-quality 360-degree video over constrained networks with low latency is very challenging due to high bandwidth requirement. Tile-based viewport adaptive streaming that proactively delivers predicted visible fields with higher quality is bandwidth-friendly, but limited prediction accuracy of head movement results in degraded viewport quality. In this letter, we propose a perception-based pseudo-motion response strategy to mitigate the damage to viewport quality, benefiting from human perception thresholds for head rotation losses and gains in virtual environment. It employs imperceptible virtual rotation losses when the imminent physical viewports may exceed the high-quality region, and immediate losses compensation once the prediction performs well. Experiments results show that our proposed strategy achieves an additional average 1.43% coding gain compared to traditional tile-based video streaming. Most notably, the proposed method is compatible with any tile-based video streaming. Lu Yu 0003, Hualong Yu |
IEEE Signal Process. Lett. | 2 |
| 2022 | ECFFNet: Effective and Consistent Feature Fusion Network for RGB-T Salient Object DetectionabstractUnder ideal environmental conditions, RGB-based deep convolutional neural networks can achieve high performance for salient object detection (SOD). In scenes with cluttered backgrounds and many objects, depth maps have been combined with RGB images to better distinguish spatial positions and structures during SOD, achieving high accuracy. However, under low-light and uneven lighting conditions, RGB and depth information may be insufficient for detection. Thermal images are insensitive to lighting and weather conditions, being able to capture important objects even during nighttime. By combining thermal images and RGB images, we propose an effective and consistent feature fusion network (ECFFNet) for RGB-T SOD. In ECFFNet, an effective cross-modality fusion module fully fuses features of corresponding sizes from the RGB and thermal modalities. Then, a bilateral reversal fusion module performs bilateral fusion of foreground and background information, enabling the full extraction of salient object boundaries. Finally, a multilevel consistent fusion module combines features across different levels to obtain complementary information. Comprehensive experiments on three RGB-T SOD datasets show that the proposed ECFFNet outperforms 12 state-of-the-art methods under different evaluation indicators. Wujie Zhou, Qinling Guo, Jingsheng Lei, Lu Yu 0003, Jenq-Neng Hwang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Hilbert Space Filling Curve Based Scan-Order for Point Cloud Attribute CompressionabstractPoint cloud is a set of three-dimensional points in arbitrary order, which is a popular representation of 3D scene in autonomous navigation and immersive applications in recent years. Compression becomes an inevitable issue due to the huge data volume of point cloud. In order to effectively compress attributes of those points, proper reordering is important. The existing voxel-based point cloud attributes compression scheme uses a naive scan for points reordering. In this paper, we theoretically analyzed 3C properties of point cloud, i.e., Compactness, Clustering and Correlation, of different scan-orders defined by different space filling curves and disclosed that the Hilbert curve can provide the best spatial correlation preservation compared with Z-order and Gray-coded curves. It is also statistically verified that the Hilbert curve always has the best ability of attributes correlation preservation for point clouds with different sparsity. We also proposed a fast and iterative Hilbert address code generation method to implement points reordering. The Hilbert scan-order could be combined with various point cloud attribute coding methods. Experiments show that the correlation preservation feature of the proposed scan-order can bring us 6.1% and 6.5% coding gain for prediction and transform coding, respectively. Jiafeng Chen, Lu Yu 0003 |
IEEE Trans. Image Process. | 2 |
| 2022 | DEFNet: Dual-Branch Enhanced Feature Fusion Network for RGB-T Crowd CountingabstractMost existing crowd counting approaches use limited information of RGB (red–green–blue) images and fail to suitably extract potential pedestrians in unconstrained scenarios. Moreover, complementary depth maps do not provide information of locations where people are more likely to be present. However, by incorporating optical and thermal information, the recognition of pedestrians may be enhanced considerably. In fact, thermal imaging information is robust to weather and lighting scenarios, and information from targets can be extracted even at nighttime. By combining RGB and thermal imaging information, we propose a dual-branch enhanced feature fusion network (DEFNet) for RGB-T (RGB and thermal) crowd counting. In DEFNet, an intensive data-enhancement module fuses complementary features of the same sizes from the RGB and thermal modalities, thus combining various rich receptive fields and generating powerful fused RGB-T features. These features describe both spatial structures and appearance details, highlighting information of crowd location. Then, an efficient dilation fusion module applies convolutions to the RGB -T features to obtain flexible and specific features, effectively eliminating the influence of background on the crowd information for density map prediction. Finally, high- and low-level features are used to efficiently obtain density maps through a fusion decoding module. Experimental results on an RGB-T crowd counting dataset indicate that the proposed DEFNet outperforms existing approaches. Furthermore, DEFNet can be generalized to handle RGB and depth data. Wujie Zhou, Jingsheng Lei, Lv Ye, Lu Yu 0003 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Design and Analysis of MEC- and Proactive Caching-Based 360° Mobile VR Video StreamingabstractRecently, 360-degree mobile virtual reality video (MVRV) has become increasingly popular because it can provide users with an immersive experience. However, MVRV is usually recorded in a high resolution and is sensitive to latency, which indicates that broadband, ultra-reliable, and low-latency communication is necessary to guarantee the users’ quality of experience. In this paper, we propose a mobile edge computing (MEC)-based 360-degree MVRV streaming scheme with field-of-view (FoV) prediction, which jointly considers video coding, proactive caching, computation offloading, and data transmission. To meet the requirement of stringent end-to-end (E2E) latency, the user’s viewpoint prediction is utilized to cache video data proactively, and computing tasks are partially offloaded to the MEC server. In addition, we propose an analytical model based on diffusion process to study the packet transmission process of 360-degree MVRV in multihop wired/wireless networks and analyze the performance of the MEC-enabled scheme. The simulation results verify the accuracy of the analysis and the effectiveness of the proposed MVRV streaming scheme in reducing the E2E delay. Furthermore, the analytical framework sheds some light on the impacts of system parameters, e.g., FoV prediction accuracy and transmission rate, on the balance between computation delay and communication delay. Qi Cheng 0006, Hangguan Shan, Weihua Zhuang, Lu Yu 0003, Zhaoyang Zhang 0001, Tony Q. S. Quek |
IEEE Trans. Multim. | 4 |
| 2022 | MFFENet: Multiscale Feature Fusion and Enhancement Network For RGB-Thermal Urban Road Scene ParsingabstractCompared with traditional handcrafted features, deep learning has greatly improved the performance of scene parsing. However, it remains challenging under various environmental conditions caused by imaging limitations. Thermal imaging cameras have several advantages over cameras for the visible spectrum, such as operation in total darkness, robustness to shadow effects, insensitivity to illumination variations, and strong ability to penetrate smog and haze. These advantages of thermal imaging cameras make them ideal for the scene parsing of semantic objects in daytime and nighttime. In this paper, we propose a novel multiscale feature fusion and enhancement network (MFFENet) for accurate parsing of RGB–thermal urban road scenes even when the quality of the available RGB data is compromised. The proposed MFFENet consists of two encoders, a feature fusion layer, and a multi-label supervision layer. We concatenate the multi-scale features with the features that contain global semantic information. Furthermore, we explore the cross-modal fusion of RGB and thermal features at multiple stages, rather than fusing them once at the low or high stage. Then, we propose a spatial attention mechanism module that provides a higher weight to (focuses more on) the foreground area, allowing MFFENet to emphasize foreground objects. Finally, multi-label supervision is introduced to optimize parameters of the proposed MFFENet. Experimental results confirm that the proposed MFFENet outperforms similar high-performing methods. Wujie Zhou, Xinyang Lin, Jingsheng Lei, Lu Yu 0003, Jenq-Neng Hwang |
IEEE Trans. Multim. | 4 |
| 2022 | CCAFNet: Crossflow and Cross-Scale Adaptive Fusion Network for Detecting Salient Objects in RGB-D ImagesabstractOwing to the widespread adoption of depth sensors, salient object detection (SOD) supported by depth maps for reliable complementary information is being increasingly investigated. Existing SOD models mainly exploit the relation between an RGB image and its corresponding depth information across three fusion domains: input RGB-D images, extracted feature maps, and output salient object. However, these models do not leverage the crossflows between high- and low-level information well. Moreover, the decoder in these models uses conventional convolution that involves several calculations. To further improve RGB-D SOD, we propose a crossflow and cross-scale adaptive fusion network (CCAFNet) to detect salient objects in RGB-D images. First, a channel fusion module allows for effective fusing depth and high-level RGB features. This module extracts accurate semantic information features from high-level RGB features. Meanwhile, a spatial fusion module combines low-level RGB and depth features with accurate boundaries and subsequently extracts detailed spatial information from low-level depth features. Finally, a purification loss is proposed to precisely learn the boundaries of salient objects and obtain additional details of the objects. The results of comprehensive experiments on seven common RGB-D SOD datasets indicate that the performance of the proposed CCAFNet is comparable to those of state-of-the-art RGB-D SOD models. Wujie Zhou, Yun Zhu 0011, Jingsheng Lei, Jian Wan 0001, Lu Yu 0003 |
IEEE Trans. Multim. | 5 |
| 2021 | Learned Image Compression with Fixed-point ArithmeticabstractLearned image compression (LIC) has achieved superior coding performance than traditional image compression standards such as HEVC intra in terms of both PSNR and MS-SSIM. However, most LIC frameworks are based on floating-point arithmetic which has two potential problems. First is that using traditional 32-bit floating-point will consume huge memory and computational cost. Second is that the decoding might fail because of the floating-point error coming from different encoding/decoding platforms. To solve the above two problems. 1) We linearly quantize the weight in the main path to 8-bit fixed-point arithmetic, and propose a fine tuning scheme to reduce the coding loss caused by the quantization. Analysis transform and synthesis transform are fine tuned layer by layer. 2) We exploit look-up-table (LUT) for the cumulative distribution function (CDF) to avoid the floating-point error. When the latent node follows non-zero mean Gaussian distribution, to share the CDF LUT for different mean values, we restrict the range of latent node to be within a certain range around mean. As a result, 8-bit weight quantization can achieve negligible coding gain loss compared with 32-bit floating-point anchor. In addition, proposed CDF LUT can ensure the correct coding at various CPU and GPU hardware platforms. Heming Sun, Lu Yu 0003, Jiro Katto |
PCS | 2 |
| 2021 | MPEG Immersive Video Coding StandardabstractThis article introduces the ISO/IEC MPEG Immersive Video (MIV) standard, MPEG-I Part 12, which is undergoing standardization. The draft MIV standard provides support for viewing immersive volumetric content captured by multiple cameras with six degrees of freedom (6DoF) within a viewing space that is determined by the camera arrangement in the capture rig. The bitstream format and decoding processes of the draft specification along with aspects of the Test Model for Immersive Video (TMIV) reference software encoder, decoder, and renderer are described. The use cases, test conditions, quality assessment methods, and experimental results are provided. In the TMIV, multiple texture and geometry views are coded as atlases of patches using a legacy 2-D video codec, while optimizing for bitrate, pixel rate, and quality. The design of the bitstream format and decoder is based on the visual volumetric video-based coding (V3C) and video-based point cloud compression (V-PCC) standard, MPEG-I Part 5. Jill M. Boyce, Renaud Doré, Adrian Dziembowski, Julien Fleureau, Joël Jung, Bart Kroon, Basel Salahieh, Vinod Kumar Malamal Vadakital, Lu Yu 0003 |
Proc. IEEE | 9 |
| 2021 | Special issue on Open Media Compression: Overview, Design Criteria, and Outlook on Emerging StandardsabstractUniversal access to and provisioning of multimedia content is now a reality. It is easy to generate, distribute, share, and consume any multimedia content, anywhere, anytime, or any device. Open media standards took a crucial role toward enabling all these use cases leading to a plethora of applications and services that have now become a commodity in our daily life. Interestingly, most of these services adopt a streaming paradigm, are typically deployed over the open, unmanaged Internet, and account for most of today’s Internet traffic. Currently, the global video traffic is greater than 60% of all Internet traffic[1], and it is expected that this share will grow to more than 80% in the near future[2]. In addition, Nielsen’s law of Internet bandwidth states that the users’ bandwidth grows by 50% per year, which roughly fits data from 1983 to 2019[3]. Thus, the users’ bandwidth can be expected to reach approximately 1 Gb/s by 2022. At the same time, network applications will grow and utilize the bandwidth provided, just like programs and their data expand to fill the memory available in a computer system. Most of the available bandwidth today is consumed by video applications, and the amount of data is further increasing due to already established and emerging applications, e.g., ultrahigh definition, high dynamic range, or virtual, augmented, mixed realities, or immersive media applications in general. Christian Timmerer, Mathias Wien, Lu Yu 0003, Amy R. Reibman |
Proc. IEEE | 3 |
| 2021 | Multiscale multilevel context and multimodal fusion for RGB-D salient object detection
Junwei Wu 0001, Wujie Zhou, Ting Luo 0001, Lu Yu 0003, Jingsheng Lei |
Signal Process. | 4 |
| 2021 | Multi-layer fusion network for blind stereoscopic 3D visual quality prediction
Wujie Zhou, Xinyang Lin, Jingsheng Lei, Lu Yu 0003, Ting Luo 0001 |
Signal Process. Image Commun. | 5 |
| 2021 | TSFNet: Two-Stage Fusion Network for RGB-T Salient Object DetectionabstractSalient object detection (SOD) based on convolutional neural networks has achieved remarkable success. However, further improving the detection performance on challenging scenes (e.g., low-light scenes) requires additional investigation. Thermal infrared imaging captures thermal radiation from the surface of objects. Thus, it is insensitive to lighting conditions and can provide uniform imaging of objects. Accordingly, we propose a two-stage fusion network (TSFNet) integrating RGB and thermal information for RGB-T SOD. For the first fusion stage, we propose a feature-wise fusion module that captures and aggregates united information and intersecting information in each local region of the RGB and thermal images, and then independent decoding is applied to the RGB and thermal features. For the second fusion stage, we propose a bilateral auxiliary fusion module that extracts auxiliary spatial features from the foreground and background of the thermal and RGB modalities. Finally, we use multiple supervision to further improve the SOD performance. Comprehensive experiments demonstrate that TSFNet outperforms 11 state-of-the-art models under various indicators on three RGB-T SOD datasets. Qinling Guo, Wujie Zhou, Jingsheng Lei, Lu Yu 0003 |
IEEE Signal Process. Lett. | 4 |
| 2021 | Two-Stage Cascaded Decoder for Semantic Segmentation of RGB-D ImagesabstractExploiting RGB and depth information can boost the performance of semantic segmentation. However, owing to the differences between RGB images and the corresponding depth maps, such multimodal information should be effectively used and combined. Most existing methods use the same fusion strategy to explore multilevel complementary information at various levels, likely ignoring different feature contributions at various levels for segmentation. To address this problem, we propose a network using a two-stage cascaded decoder (TCD), embedding a detail polishing module, to effectively integrate high- and low-level features and suppress noise from low-level details. Additionally, we introduce a depth filter and fusion module to extract informative regions from depth cues with the guidance of RGB images. The proposed TCD network achieves comparable performance to state-of-the-art RGB-D semantic segmentation methods on the benchmark NYUDv2 and SUN RGB-D datasets. Yuchun Yue, Wujie Zhou, Jingsheng Lei, Lu Yu 0003 |
IEEE Signal Process. Lett. | 4 |
| 2021 | MRINet: Multilevel Reverse-Context Interactive-Fusion Network for Detecting Salient Objects in RGB-D ImagesabstractThe use of RGB-D information for salient object detection (SOD) is being increasingly explored. Traditional multilevel models handle both low- and high-level features similarly, as they use the same number of features for blending. Unlike these models, in this paper, we propose multilevel reverse-context interactive-fusion (MRI) network (MRINet) for RGB-D SOD. Specifically, first, we extract and reuse different numbers of features depending on their level; the deeper the information, the more times do we perform the extraction. Deeper information contains more semantic cues, which are important for locating salient regions. Thereafter, we use an RGB MRI block (MRIB) to merge RGB information at different levels; furthermore, we use depth features as auxiliary information and an RGB-D MRIB for full merging with RGB information. RGB and RGB-D MRIBs can reconstruct the high-level feature map in high resolution and integrate the low-level feature map to enhance boundary details. Extensive experiments demonstrate the effectiveness of the proposed MRINet and its state-of-the-art performance in RGB-D SOD. Wujie Zhou, Sijia Pan, Jingsheng Lei, Lu Yu 0003 |
IEEE Signal Process. Lett. | 4 |
| 2021 | Parallax-Estimation-Enhanced Network With Interweave Consistency Feature Fusion for Binocular Salient Object DetectionabstractSalient object detection (SOD) has received extensive attention in recent years, and many models have been developed. However, most SOD models only consider monocular images and not binocular images, which resemble the human vision and can better reflect human perception for distinguishing salient objects. To leverage the information in binocular images, we propose herein a first-of-its-kind parallax-estimation-enhanced network (PEENet) for binocular SOD. More specifically, we use a weighted binocular fusion module and a parallax correlation fusion module to explore the complementary and different information in binocular images. In addition, a parallax enhancing module and interweave consistency fusion use complementary saliency information and parallax information to enhance saliency and parallax representations. Finally, a transformation module avoids global and local information loss during decoding. Experiments were performed to validate the effectiveness and robustness of the proposed PEENet, which outperforms 10-RGB/RGB-D SOD methods on two binocular SOD datasets. Yun Zhu 0011, Wujie Zhou, Lu Yu 0003 |
IEEE Signal Process. Lett. | 4 |
| 2021 | GMNet: Graded-Feature Multilabel-Learning Network for RGB-Thermal Urban Scene Semantic SegmentationabstractSemantic segmentation is a fundamental task in computer vision, and it has various applications in fields such as robotic sensing, video surveillance, and autonomous driving. A major research topic in urban road semantic segmentation is the proper integration and use of cross-modal information for fusion. Here, we attempt to leverage inherent multimodal information and acquire graded features to develop a novel multilabel-learning network for RGB-thermal urban scene semantic segmentation. Specifically, we propose a strategy for graded-feature extraction to split multilevel features into junior, intermediate, and senior levels. Then, we integrate RGB and thermal modalities with two distinct fusion modules, namely a shallow feature fusion module and deep feature fusion module for junior and senior features. Finally, we use multilabel supervision to optimize the network in terms of semantic, binary, and boundary characteristics. Experimental results confirm that the proposed architecture, the graded-feature multilabel-learning network, outperforms state-of-the-art methods for urban scene semantic segmentation, and it can be generalized to depth data. Wujie Zhou, Jingsheng Lei, Lu Yu 0003, Jenq-Neng Hwang |
IEEE Trans. Image Process. | 4 |
| 2021 | Salient Object Detection in Stereoscopic 3D Images Using a Deep Convolutional Residual AutoencoderabstractIn recent years, the detection of distinctive objects in stereoscopic 3D images has drawn increasing attention. Unlike 2D salient object detection, salient object detection in stereoscopic 3D images is highly challenging. Hence, we propose a novel Deep Convolutional Residual Autoencoder (DCRA) for end-to-end salient object detection in stereoscopic 3D images. The core trainable architecture of the salient object detection model employs raw stereoscopic 3D images as the inputs and their corresponding ground truth saliency masks as the labels. A convolutional residual module is applied to both the encoder and the decoder as a basic building block in the DCRA, and long-range skip connections are employed to bypass the equal-sized feature maps between the encoder and the decoder. To explore the complex relationships and exploit the complementarity between RGB (photometric) and depth (geometric) information, multiple feature map fusion modules are constructed. These modules integrate texture and structure information between the RGB and depth branches of the encoder and fuse their features over several multiscale layers. Finally, to efficiently optimize DCRA parameters, a supervision pyramid based on boundary loss and background prior loss is adopted, which employs supervised learning over the multiscale layers in the decoder to prevent vanishing gradients and accelerate the training at the fusion stage. We compare the proposed DCRA with state-of-the-art methods on two challenging benchmark datasets. The results of these experiments demonstrate that our proposed DCRA performs favorably against the comparison models. Wujie Zhou, Junwei Wu 0001, Jingsheng Lei, Jenq-Neng Hwang, Lu Yu 0003 |
IEEE Trans. Multim. | 5 |
| 2021 | Global and Local-Contrast Guides Content-Aware Fusion for RGB-D Saliency PredictionabstractMany RGB-D visual attention models have been proposed with diverse fusion models; thus, the main challenge lies in the differences in the results between the different models. To address this challenge, we propose a local-global fusion model for fixation prediction on an RGB-D image; this method combines global and local information through a content-aware fusion module (CAFM) structure. First, it comprises a channel-based upsampling block for exploiting global contextual information and scaling up this information to the same resolution as the input. Second, our Deconv block contains a contrast feature module to utilize multilevel local features stage-by-stage for superior local feature representation. The experimental results demonstrate that the proposed model exhibits competitive performance on two databases. Wujie Zhou, Jingsheng Lei, Lu Yu 0003 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2020 | Triplet Distillation For Deep Face RecognitionabstractConvolutional neural networks (CNNs) have achieved great successes in face recognition, which unfortunately comes at the cost of massive computation and storage consumption. Many compact face recognition networks are thus proposed to resolve this problem, and triplet loss is effective to further improve the performance of these compact models. However, it normally employs a fixed margin to all the samples, which neglects the informative similarity structures between different identities. In this paper, we borrow the idea of knowledge distillation and define the informative similarity as the transferred knowledge. Then, we propose an enhanced version of triplet loss, named triplet distillation, which exploits the capability of a teacher model to transfer the similarity information to a student model by adaptively varying the margin between positive and negative pairs. Experiments on the LFW, AgeDB and CPLFW datasets show the merits of our method compared to the original triplet loss. Yushu Feng, Huan Wang 0014, Haoji Hu, Lu Yu 0003, Wei Wang 0118, Shiyan Wang |
ICIP | 4 |
| 2020 | Affine Deformation Model Based Intra Block Copy for Intra Frame CodingabstractIn intra frame coding, many methods are proposed to reduce spatial redundancy, including angular intra prediction and intra block copy. However, these two methods cannot deal with complex structure redundancy, such as scaling and rotation relationship where spatial structures are the same but sizes and orientations are different, leading to limited coding efficiency. In this paper, we first mathematically prove that 4-parameter affine deformation model can describe this complex relationship. Then we propose an affine deformation model based intra block copy (ADMIBC) scheme. In the scheme, we focus on several problems, including a fast displacement vector validity judgement algorithm to reduce encoding time, a padding method for unavailable reference pixels during sub-pixel interpolation to improve prediction accuracy and candidate list building methods of control point displacement vector predictor to reduce coding bits. Experimental results show that ADMIBC achieves up to -2.16% BD-rate reduction and -0.99% on average in all intra configuration, compared to the latest video coding standard Versatile Video Coding (VVC). Daowen Li, Kaitian Qiu, Yaqing Pan, Yingming Li, Haoji Hu, Lu Yu 0003 |
ISCAS | 7 |
| 2020 | Fully Neural Network Mode Based Intra Prediction of Variable Block SizeabstractIntra prediction is an essential component in the image coding. This paper gives an intra prediction framework completely based on neural network modes (NM). Each NM can be regarded as a regression from the neighboring reference blocks to the current coding block. (1) For variable block size, we utilize different network structures. For small blocks 4×4 and 8×8, fully connected networks are used, while for large blocks 16×16 and 32×32, convolutional neural networks are exploited. (2) For each prediction mode, we develop a specific pre-trained network to boost the regression accuracy. When integrating into HEVC test model, we can save 3.55%, 3.03% and 3.27% BD-rate for Y, U, V components compared with the anchor. As far as we know, this is the first work to explore a fully NM based framework for intra prediction, and we reach a better coding gain with a lower complexity compared with the previous work. Heming Sun, Lu Yu 0003, Jiro Katto |
VCIP | 2 |
| 2020 | Compressing Facial Makeup Transfer Networks by Collaborative Distillation and Kernel DecompositionabstractAlthough the facial makeup transfer network has achieved high-quality performance in generating perceptually pleasing makeup images, its capability is still restricted by the massive computation and storage of the network architecture. We address this issue by compressing facial makeup transfer networks with collaborative distillation and kernel decomposition. The main idea of collaborative distillation is underpinned by a finding that the encoder-decoder pairs construct an exclusive collaborative relationship, which is regarded as a new kind of knowledge for low-level vision tasks. For kernel decomposition, we apply the depth-wise separation of convolutional kernels to build a light-weighted Convolutional Neural Network (CNN) from the original network. Extensive experiments show the effectiveness of the compression method when applied to the state-of-the-art facial makeup transfer network - BeautyGAN [1]. Bianjiang Yang, Zi Hui, Haoji Hu, Lu Yu 0003 |
VCIP | 5 |
| 2020 | NOMA based VR Video Transmissions Exploiting User Behavioral CoherenceabstractIn this work, we study the cooperative and non-cooperative transmission schemes design for live VR video broadcast scenarios by utilizing non-orthogonal multiple access (NOMA), considering that users’ viewports partly overlap due to behavioral coherence. To characterize the performance of the proposed cooperative and non-cooperative transmission schemes, the exact and asymptotic expressions of outage probability, as well as the average outage capacity under imperfect successive interference cancellation (SIC), are derived, respectively. Based on the asymptotic outage probability results, we optimize the power allocation to maximize the average outage capacity of the proposed schemes. Finally, simulation results demonstrate that both of the proposed schemes can achieve a considerable performance gain over the traditional orthogonal multiple access (OMA) scheme in average outage capacity, and each of the proposed schemes has its advantages and applicable scenarios. Ping Xiang, Hangguan Shan, Zhaoyang Zhang 0001, Lu Yu 0003, Tony Q. S. Quek |
WCNC | 4 |
| 2020 | Video Surveillance on Mobile Edge Networks - A Reinforcement-Learning-Based ApproachabstractVideo surveillance systems or Internet of Multimedia Things are playing a more and more important role in our daily life. To obtain useful surveillance information timely and accurately, not only image recognition algorithms but also computing and communication resources can be bottlenecks of the whole system. In this article, taking face recognition application as an example, we study how to build video surveillance systems by utilizing mobile edge computing (MEC), one of the 5G's key technologies. Specifically, to achieve high recognition accuracy and low recognition time, we design image recognition algorithms for both the camera sensor and MEC server, and utilize the action-value methods to train actions of the system by jointly optimizing offloading decision and image compression parameters. The experimental results show the advantages of the proposed system for enabling communication environment-adaptive, efficient, and intelligent video surveillance. Haoji Hu, Hangguan Shan, Chuankun Wang, Tengxu Sun, Xiaojian Zhen, Kunpeng Yang, Lu Yu 0003, Zhaoyang Zhang 0001, Tony Q. S. Quek |
IEEE Internet Things J. | 7 |
| 2020 | Crowdsourcing Based Cross Random Access Point Referencing for Video CodingabstractIn video coding, Random Access Points (RAPs) are inserted in a bitstream to support flexible tune-in but divide it into multiple independent Random Access Segments (RASs) that may have similar contents. To reduce redundancy between RASs, this letter proposes a novel Cross Random-access-point Referencing (CRR) structure to provide inter prediction for RAP pictures by using multiple External Reference Pictures (ERPs) across RAPs that are selected from preceding or following RASs other than the current RAS. With ERPs shared by multiple RASs, a crowdsourcing method is proposed to optimize the joint rate distortion costs of RASs and ERPs to generate an optimal set of ERPs. Content preparation and bitstream splicing processes supported by system environments are also designed to ensure random access functionality of CRR coded RASs. Simulation results show that CRR achieves significant coding gain compared to Versatile Video Coding (VVC), i.e., 12.00% on sequences in common test condition and 25.48% on long drama sequences. Hualong Yu, Xiaoding Gao, Lu Yu 0003 |
IEEE Signal Process. Lett. | 3 |
| 2020 | GFNet: Gate Fusion Network With Res2Net for Detecting Salient Objects in RGB-D ImagesabstractThe performance of recent RGB-D salient object detectors has significantly improved owing to their integration of convolutional neural networks (CNNs). However, most existing salient object detection (SOD) methods represent features using a VGGNet backbone, which lacks the ability to retain complete RGB and depth modals and must compensate by applying several skip connections. In this letter, we propose a gate fusion network (GFNet) with Res2Net architecture to solve this problem. GFNet consists of two interacting Res2Net block encoder streams and four gate fusion block (GFB) decoders to interconnect the streams and fuse features. Res2Net blocks have a robust feature retention mechanism to ensure that the decoders can learn complete information, while the GFB formulates the interdependences of the encoders and eliminates noise via a gate mechanism. We evaluated GFNet using two popular RGB-D salient detection benchmark datasets (NJU2000 and NLPR) and achieved state-of-the art performance. Wujie Zhou, Lu Yu 0003 |
IEEE Signal Process. Lett. | 4 |
| 2020 | Guest Editorial Introduction to Special Section on Learning-Based Image and Video CompressionabstractVideo is being watched more than ever before. It is estimated that in 2020, 82% of global IP traffic and 79% of global Internet traffic will come from video; globally 3 trillion minutes (5 million years) of video content will cross the Internet each month. According to the Cisco 2020 Forecast, that is one million minutes of video streamed or downloaded every second[1]. The rapidly increasing consumption of storage capacity and transmission bandwidth from video, especially HD and UHD video content, has made video compression a critical stage to guarantee the quality of delivery and playback. Shan Liu 0001, Wen-Hsiao Peng, Lu Yu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Convolutional Neural Network Based Bi-Prediction Utilizing Spatial and Temporal Information in Video CodingabstractWith the growing popularity of high-resolution videos, the demand for higher coding efficiency is increasing to cope with multimedia transmission challenges on the communication network. Since conventional linearly weighted bi-prediction does not handle inhomogeneous motion activities inside one block well, in recent works, Convolutional Neural Network (CNN) is explored to tackle inhomogeneous motion by utilizing patch-level information to predict each individual pixel. However, as only two reference blocks are used as input information, those works ignore the variation of pixel values between reference blocks and current block, and ignore the differences between extrapolation and interpolation. This work utilizes both spatial neighboring pixels and temporal display orders as extra inputs for CNN models to further improve the prediction accuracy of a bi-predictor. The extra input information has the following advantages. First, variations among spatial neighboring pixels of both reference blocks and the current block reflect variations between current block and reference blocks. Together with temporal distance, spatial neighboring pixels are able to address extrapolation and interpolation uniformly. Second, spatial neighboring pixels of the current block have a high correlation with current predicted signals, which helps to reduce prediction residuals around the block boundary and alleviate block artifacts. Last, temporal distances help to improve the accuracy of prediction signals based on its ability of reflecting the correlation of video frames. Experimental results show that our proposed network achieves 2.92% and 5.09% bit-rate savings on average compared with HEVC, under Low-Delay B (LDB) and Random-Access (RA) configurations, respectively. As temporal information is used in our network, the LDB and RA configurations share the same networks in this work. Jue Mao, Lu Yu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | A YCbCr Color Depth Packing Method and its Extension for 3D Video Broadcasting ServicesabstractTo deliver three-dimension (3D) video services through the current two-dimension (2D) broadcasting systems, the frame compatible packing formats of views with depth maps will be an effective approach. With texture and depth information, the 3D videos with a depth image-based rendering engine can easily support all glasses and glasses-free 3D displays. In this paper, we propose a new YCbCr color depth packing method based on the centralized texture-depth packing (CTDP) formats to deliver effective 3D video services. Simulations show that the CTDP formats with YCbCr color depth packing method can achieve better objective and subjective texture and depth quality than the 2D-plus-depth packing (2DDP) formats. The benefits from using CTDP include light complexity and maturity and wide application of existing video codec. The packing scheme is compatible with any video codecs. Before the 3D-HEVC-based video broadcasting system, the proposed CTDP formats with YCbCr color depth packing method could help to deliver 3D videos in the current 2D broadcasting systems simply and efficiently. Jar-Ferr Yang, Kuan-Ting Lee, Guan-Cheng Chen, Wei-Jong Yang, Lu Yu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | End-To-End Convolutional Network for Video Rain Streaks RemovalabstractExisting video rain streaks removal methods utilize various manual models to represent the appearance of rain streaks, and only use convolutional neural network (CNN) as a post-processing part to compensate the artifacts like misalignment caused by traditional de-raining operations. However, these manual models only work for some particular scenes because the distribution of rain streaks is complex and random. Moreover, since CNN network and previous traditional de-raining operations cannot be trained jointly, the output of CNN network may still contain artifacts. To address these problems, we propose an end-to-end video rain streaks removal CNN network called EEVRSR net. Experimental results of both synthetic and real data demonstrate that the proposed EEVRSR net achieves better performance in both speed and effectiveness over state-of-the-art methods. Xiaoding Gao, Jue Mao, Hualong Yu, Lu Yu 0003 |
ICIP | 4 |
| 2019 | Three-Dimensional Convolutional Neural Network Pruning with Regularization-Based MethodabstractDespite enjoying extensive applications in video analysis, three-dimensional convolutional neural networks (3D CNNs) are restricted by their massive computation and storage consumption. To solve this problem, we propose a three-dimensional regularization-based neural network pruning method to assign different regularization parameters to different weight groups based on their importance to the network. Further we analyze the redundancy and computation cost for each layer to determine the different pruning ratios. Experiments show that pruning based on our method can lead to 2× theoretical speedup with only 0.41% accuracy loss for 3D-ResNet18 and 3.28% accuracy loss for C3D. The proposed method performs favorably against other popular methods for model compression and acceleration. Huan Wang 0014, Lu Yu 0003, Haoji Hu, Hangguan Shan, Tony Q. S. Quek |
ICIP | 4 |
| 2019 | Structured Pruning for Efficient ConvNets via Incremental RegularizationabstractParameter pruning is a promising approach for CNN compression and acceleration by eliminating redundant model parameters with tolerable performance degrade. Despite its effectiveness, existing regularization-based parameter pruning methods usually drive weights towards zero with large and constant regularization factors, which neglects the fragility of the expressiveness of CNNs, and thus calls for a more gentle regularization scheme so that the networks can adapt during pruning. To achieve this, we propose a new and novel regularization-based pruning method, named IncReg, to incrementally assign different regularization factors to different weights based on their relative importance. Empirical analysis on CIFAR-10 dataset verifies the merits of IncReg. Further extensive experiments with popular CNNs on CIFAR-10 and ImageNet datasets show that IncReg achieves comparable to even better results compared with state-of-the-arts. Our source codes and trained models are available here: https://github.com/mingsun-tse/caffe_increg. Huan Wang 0014, Qiming Zhang 0001, Yuehai Wang, Lu Yu 0003, Haoji Hu |
IJCNN | 4 |
| 2019 | Parallel-to-Axis Uniform Cubemap Projection for Omnidirectional VideoabstractOmnidirectional video can provide an immersive visual experience by presenting a 360° video content. For omnidirectional video, pixels on the sphere need to be mapped onto a two-dimensional plane to adapt to the existing coding standards. Equirectangular Projection (ERP) and Cubemap Projection (CMP) have excessive sampling density on some sampling regions, which results in unnecessary sampling pixels. In this paper, Parallel-to-Axis Uniform Cubemap Projection (PAU) is proposed. After projecting the sphere to a cube like CMP, an equal-angle mapping position adjustment is used to sample the sphere more uniformly. After adjustment, one chosen axis is equal-angle sampled and each line parallel to the other axis is sampled equiangularly in the proposed format. Experimental results show that our proposed projection format can achieve 10.68% bitrate gain compared to ERP. Xuchang Huangfu, Yule Sun, Ruidi Zheng, Lu Yu 0003 |
ISCAS | 5 |
| 2019 | An In-Loop Filter Based on Low-Complexity CNN using Residuals in Intra Video CodingabstractNeural network-based filters have shown their potential in removing video compression artifacts. However, previously studied neural networks have achieved boosted filtering performance by continuously increasing network complexity, causing heavy burden on memory cost and computation speed. In this paper, we firstly analyze properties of original residuals which are the difference between original and predicted pixel values. Then an in-loop filter based on low-complexity CNN using residuals(CNNF-R), which are generated after compression and reconstruction from original residuals, is proposed for intra video coding. Insights of designing the network are also demonstrated. Compared with the state-of-the-art video coding standard HEVC, CNNF-R achieves up to 6.8% BD-rate reduction and 4.8% on average under all intra configuration, and 2.3% on average under random access configuration. Meanwhile, CNNF-R outperforms the previous network VRCNN in terms of nearly 70% decrease in computation complexity, considerable decrease in memory consumption and 1.2% increase in BD-rate reduction. Daowen Li, Lu Yu 0003 |
ISCAS | 2 |
| 2019 | CNN-Based Bi-Prediction Utilizing Spatial Information for Video CodingabstractIn video coding, slice-level and block-level weighted bi-prediction are used for scenes with temporal brightness variation. However, there are still structured residuals when applying weighted bi-prediction in slice and block level. Recently, CNN-based bi-prediction has achieved remarkable success on reducing significant structured residuals, in which bi-predictor is generated by CNN model using two reference blocks as inputs. Inspired by high spatial correlation of pixels, this paper uses spatial neighboring pixels of both current block and two reference blocks as the additional information of the proposed CNN model to further reduce residual and generate a more accurate bi-predictor. Moreover, by comparing AMVP and merge/skip mode, this paper illustrates that CNN-based bi-prediction is more efficient for merge/skip mode than for AMVP mode. Experimental results show that proposed method reaches 3.46% BD-rate saving for random access configuration on average compared to HM 16.15. Jue Mao, Hualong Yu, Xiaoding Gao, Lu Yu 0003 |
ISCAS | 4 |
| 2019 | Standard Designs for Cross Random Access Point Reference in Video CodingabstractIn videos like movies and TV shows, similar scenes usually appear alternately. The period between these similar scenes is so long that it even exceeds the length of a random access (RA) segment. This means that the temporal correlation among these similar scenes crosses random access point (RAP). Besides, in videos like surveillance videos, a similar scene usually lasts for a long time which is longer than a RA segment. In other words, the temporal correlation among this scene also crosses RAP. To exploit the cross-RAP temporal correlation for further improving coding performance and keep the random access functionality at the same time, library picture based cross random access point reference is proposed and adopted in AVS3, which is going to be introduced in this paper. Firstly, a library picture is used as a reference picture for similar RA segments to exploit the temporal correlation among these RA segments. Secondly, to support random access, decoder and system layer are both designed to ensure that the decoder can get the corresponding library picture when random access occurs to correctly decode RA segments. Experimental results show that this method can save 6.4% and 31.5% BD-rate on the AVS3 general sequences and special sequences. Xiaoding Gao, Hualong Yu, Qichao Yuan, Xiangyu Lin, Lu Yu 0003 |
PCS | 5 |
| 2019 | Complementary Motion Vector for Motion Prediction in Video Coding with Long-Term ReferenceabstractIn HEVC, there are two types of motion vector (MV) when long-term reference is enabled: short-term MV (SMV) pointing to short-term reference frame and long-term MV (LMV) pointing to long-term reference frame. And cross-class prediction between SMV and LMV is not allowed because of their low correlation. Therefore, MV predictor candidates of current block would be inadequate when neighboring MVs are in different types. This paper proposes a complementary MV to enrich MV predictor candidates for current block. There would be two types of MV for each neighboring inter block. In addition to MV used in motion compensation, a complementary MV of the other type is derived by reconstructed pixels. This paper also proposed a reliability-based MV predictor candidate list construction method to improve the prediction efficiency of complementary MVs. Experimental results show that the proposed method can achieve 1.18% coding performance improvement on average. Jue Mao, Hualong Yu, Xiaoding Gao, Lu Yu 0003 |
PCS | 4 |
| 2019 | Adaptive QP with Tile Partition and Padding to Remove Boundary Artifacts for 360 Video CodingabstractTo adapt to existing high-efficiency video coding standards or technologies, 360 video is represented by different projection formats with one or multiple planes. However, continuous spherical content will have discontinuous boundaries on representation planes, which results in obvious artifacts on viewing viewports rendered from coded projection formats. To solve this subjective issue, a boundary based adaptive QP method together with tile partition and padding is proposed in this paper. Tile partition and padding can effectively reduce obvious seam artifacts. Boundary areas are coded with better quality by using lower QP to further reduce visible artifacts. Experiments are conducted based on hybrid equi-angular cubemap projection (HEC). Experimental results show that the proposed method can significantly improve subjective quality by removing boundary artifacts for 360 video coding. Yule Sun, Lu Yu 0003 |
PCS | 3 |
| 2019 | Low Pixel Rate 3DoF+ Video Compression Via Unpredictable Region CroppingabstractEnhanced three degrees of freedom (3DoF+) video system enables both rotational and translational movements for viewers within a limited scene. It introduces interactive motion parallax that provides viewers with a more natural and immersive visual experience on head-mounted displays (HMDs). A large set of views are required to support 3DoF+ video system, hence huge data costs a tremendous amount of computation and bandwidth. In this paper, a method based on unpredictable region cropping is proposed to reduce the size of data. The proposed method can be separated into two steps: basic view selection and sub-image cropping. One or several views in the source view set are selected as basic views adaptively. The goal of the basic views is to predict the remaining views, and unpredictable regions are cropped to multiple rectangular sub-images. The basic views and the cropped sub-images, called to-be-coded views, are coded by using High Efficiency Video Coding (HEVC) standard. Experimental results show the proposed method can achieve up to 75.0% pixel-rate saving and improve the quality of rendered views by 1.5dB. The method of basic view selection was adopted by the moving picture experts group (MPEG) video subgroup as a tool for the first version of Test Model of Immersive Video (TMIV). Yule Sun, Lu Yu 0003 |
PCS | 3 |
| 2019 | Perceptual Tolerance to Motion-To-Photon Latency with Head Movement in Virtual RealityabstractSince Motion-To-Photon (MTP) latency is inevitable and can be perceived in virtual reality, quantifying perception of MTP latency becomes necessary. In this paper, we investigate perceptual tolerance to MTP latency, including perception threshold of the latency and user acceptance of delays above the threshold. It is affected by different head motion events, such as Motion-Static-Alternate (MSA, i.e., an acceleration or deceleration in one direction) and Motion-To-Reverse (MTR, i.e., a movement reverses direction). In each motion event, rotation angle and angular velocity also influence perception of MTP latency. Experimental results show that subjects are more intolerant of MTP latency in MTR than MSA. The latency perception threshold is about 23 ms when subjects turn their heads at the maximum speed of human limits. When the angular velocity decreases, the perception threshold increases. The maximum threshold is ~41 ms at 20 °/s in this study. Inversely proportional models are established to describe the relationship between threshold and angular velocity. Besides, MTP latency over the threshold is harder to be accepted with the rotation angle decreasing or the angular velocity increasing. Minxia Yang, Lu Yu 0003 |
PCS | 3 |
| 2019 | Novel Coding Tools Based on Characteristics for Short VideosabstractCurrently, short videos occupy large proportion of live streaming media and short video coding attracts intense attention. Different from traditional online videos, short videos feature frequent scene switching and allow various complex special effects to be added. For the unique characteristics of these short videos, we proposed an advanced video compression platform, which includes Library Picture based Cross Random Access Point Reference (LP-CRAPR), Affine-based Intra Block Copy (AIBC) and Local Illumination Compensation (LIC), to meet the challenges of short video coding. The experimental results show that our proposed platform has performed 5.84% and 2.74% gain under Random Access (RA) and All Intra (AI) configuration compared to VTM 4.0, with approximately two times encoding complexity increasing and no additional decoding time. For our proposed short video coding, high coding efficiency and improved subjective results are achieved. Xiangyu Lin, Daowen Li, Jue Mao, Yaqing Pan, Lu Yu 0003 |
PCS | 7 |
| 2019 | Deep blind quality evaluator for multiply distorted images based on monogenic binary coding
Wujie Zhou, Lu Yu 0003, Yaguan Qian, Weiwei Qiu, Yang Zhou 0011, Ting Luo 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Deep Road Scene UnderstandingabstractRoad scene understanding is a difficult task in autonomous driving. In this letter, we propose a novel deep encoder-decoder architecture for road scene understanding in an end-to-end manner. This core trainable understanding engine includes an encoder network, a decoder network with two streams, and a pixel-level fusion network with classification layer. The encoder network is composed of the front-end model of the classical convolution neural network, VGGNet. The decoder network with two streams includes multi-scale skip connection modules to reduce the down-scaling effect. Finally, a fusion network fuses the two-level information from the two streams of the decoder network for precise pixel-level classification. Additionally, the convolution layer is added to each skip connection module to increase the depth of the architecture. Our architecture achieves outstanding performance on the publicly available CamVid dataset and significantly outperforms previous architectures. This deep architecture is ideal for road scene understanding. Wujie Zhou, Sijia Lv, Qiuping Jiang, Lu Yu 0003 |
IEEE Signal Process. Lett. | 4 |
| 2018 | Improved DASH for Cross-Random-Access Prediction Structure in Video CodingabstractIn video coding, temporal correlations between pictures are utilized by short-term and long-term reference, which are limited inside a random access segment and cannot cross random access points (RAPs). To improve coding efficiency, correlations across RAPs are exploited by some novel video coding schemes. The cross-random-access prediction structure results in alterable dependency between pictures, which is affected by random access and different from the fixed dependency in conventional video coding standard. However, current media streaming scheme, such as Dynamic Adaptive Streaming over HTTP (DASH), cannot describe dependency across RAPs. This paper introduces the cross-random-access dependency description in DASH syntax. With dependency information, client can download segments in proper temporal order but may re-download segments. This paper further improves downloading management in client. Experiments show that improved DASH system transmitting stream with cross-random-access prediction structure can save 22.1% on maximum and 14.3% on average transmission bits in contrast to conventional DASH system. Hualong Yu, Lu Yu 0003 |
ISCAS | 2 |
| 2018 | Adaptive Weighted Averaged Template Matching Prediction for Intra CodingabstractTemplate matching (TM) method has been used in video coding to improve the performance of intra prediction by exploiting non-local correlations. One or more blocks, which are selected by matching the block templates, are used to create prediction samples for the target block. Each selected block is a degraded target block superimposed with the noise caused by compression and inner dissimilarities between the blocks. In this paper, an adaptive weighted averaged template matching (AWTM) method is proposed to minimize the noise variance in prediction samples. The prediction samples are generated by weighted average of several selected blocks. The weights are determined according to the correlations adaptively. Furthermore, a novel approach is proposed to integrate the AWTM method into HEVC test model, including a new derivation algorithm of most probable mode (MPM) list and modifications in coding related syntax elements. Our method can achieve 1.05% coding gain on average and up to 4.24% for luminance component under all-intra configuration compared to HEVC test model HM16.0. Lu Yu 0003 |
ISCAS | 4 |
| 2018 | Projection Deformation Based Motion Compensation for Panoramic VideoabstractFor these multi-face projection formats of panoramic video, such as cube map projection format, projection on different faces of a rigid object may change in size and shape when it moves across planes. If a reference block of the current block is out of the range of the current face, projection deformation based motion compensation is proposed. Different with the conventional block-matching motion compensation, the matching block in the reference frame is allowed to be deformed for better motion compensation. Taking parallelizability into consideration, the proposed motion compensation method is applied to each sub-block of the current block. Deformation of reference block is reflected in the relative location of reference sub-blocks while each sub-block still sticks to the translational motion model. Projection deformation based motion compensation is embedded into HM16.15 and achieves BD rate reductions of 0.9% for sequences with relatively rough motions. Ruidi Zheng, Xuchang Huangfu, Yule Sun, Lu Yu 0003 |
ISCAS | 4 |
| 2018 | Adaptive Weighted Bi-Prediction based on Template Similarity in Video CodingabstractIn bi-prediction of merge/skip and advanced motion vector prediction(AMVP) modes, current block is predicted by averaging two uni-reference blocks. In this paper, statistical experiments show that averaged bi-prediction is not a good choice for merge/skip mode, compared with AMVP mode. And we propose an adaptive weighted bi-prediction method for merge/skip mode according to the similarities between current block and its two uni-predictors. The uni-predictor, which is more similar to current block, would be assigned with larger weighting factor. As pixel values of current block are unknown in decoder before conducting motion compensation, the similarity between current block and uni-reference block is estimated by similarity between their corresponding spatial neighboring pixels. The results show that the proposed method achieves 0.55% BD reduction on average over the reference software JEM 6.0. Jue Mao, Yin Zhao, Lu Yu 0003 |
VCIP | 4 |
| 2018 | Long-term prediction for hierarchical-B-picture-based coding of video with repeated shotsabstractThe latest video coding standard High Efficiency Video Coding (HEVC) can achieve much higher coding efficiency than previous video coding standards. Particularly, by exploiting the hierarchical B-picture prediction structure, temporal redundancy among neighbor frames is eliminated remarkably well. In practice, videos available to consumers usually contain many repeated shots, such as TV series, movies, and talk shows. According to our observations, when these videos are encoded by HEVC with the hierarchical B-picture structure, the temporal correlation in each shot is well exploited. However, the long-term correlation between repeated shots has not been used. We propose a long-term prediction (LTP) scheme to use the long-term temporal correlation between correlated shots in a video. The long-term reference (LTR) frames of a source video are chosen by clustering similar shots and extracting the representative frames, and a modified hierarchical B-picture coding structure based on an LTR frame is introduced to support long-term temporal prediction. An adaptive quantization method is further designed for LTR frames to improve the overall video coding efficiency. Experimental results show that up to 22.86% coding gain can be achieved using the new coding scheme. Xuguang Zuo, Lu Yu 0003 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2018 | Possibility distribution based lossless coding and its optimization
Zhichu He, Lu Yu 0003 |
Signal Process. | 2 |
| 2018 | A Spatio-Temporal Perceptual Quality Index Measuring Compression Distortions of Three-Dimensional VideoabstractObjective video quality assessment (VQA) of stereoscopic three-dimensional (3-D) video plays a vital role in the application of video compression and transmission. This letter presents an efficient full-reference metric, called 3-D perceptual quality index (3-D-PQI), to measure video compression distortion of stereoscopic videos. In the metric, local video compression distortions in left and right view are measured in both spatial and temporal channels by considering contrast and motion masking effect. Then, a stereo saliency based pooling strategy is used to accumulate local spatial and temporal distortions. Finally, 3-D-PQI is derived from texture energy based fusion of distortion measurement of left and right view. A verification test of the proposed metric is conducted on publicly available stereoscopic VQA databases, which shows that 3-D-PQI is superior to other VQA metrics with respect to correlation against subjective quality judgment. Lu Yu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2018 | Local and Global Feature Learning for Blind Quality Evaluation of Screen Content and Natural Scene ImagesabstractThe blind quality evaluation of screen content images (SCIs) and natural scene images (NSIs) has become an important, yet very challenging issue. In this paper, we present an effective blind quality evaluation technique for SCIs and NSIs based on a dictionary of learned local and global quality features. First, a local dictionary is constructed using local normalized image patches and conventional -means clustering. With this local dictionary, the learned local quality features can be obtained using a locality-constrained linear coding with max pooling. To extract the learned global quality features, the histogram representations of binary patterns are concatenated to form a global dictionary. The collaborative representation algorithm is used to efficiently code the learned global quality features of the distorted images using this dictionary. Finally, kernel-based support vector regression is used to integrate these features into an overall quality score. Extensive experiments involving the proposed evaluation technique demonstrate that in comparison with most relevant metrics, the proposed blind metric yields significantly higher consistency in line with subjective fidelity ratings. Wujie Zhou, Lu Yu 0003, Yang Zhou 0011, Weiwei Qiu, Mingwei Wu 0001, Ting Luo 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Geometric derived motion vector for motion prediction in block-based video codingabstractMotion compensation plays an important role in inter picture prediction but the representation of motion vector (MV) is highly redundant. The existing motion model only selects the MVs of neighbor blocks as candidates even if they are very different from the MV of current block, which results in additional RD cost. In this paper, a geometric motion model is proposed to derive a geometric derived motion vector (GDMV) when an object has a relative movement that is close to or away from the camera. The proposed GDMV is integrated to HM 16.0 to further improve the motion prediction in HEVC. In order to ensure the accuracy of the motion vector predictor, a refining process is adopted before integrating the model into HEVC motion vector prediction. Experimental results show that compared to HM 16.0, the proposed GDMV can bring 0.14%~1.13% BD bitrate saving in low delay P main test. Lu Yu 0003 |
ICIP | 2 |
| 2017 | A higher order transform domain filter exploiting non-local spatial correlation for video codingabstractIn-loop filters are widely employed in video coding standards to improve coding efficiency by reducing quantization noise. However, only local spatial correlation of video frames is exploited in existing in-loop filters, which limits the improvement of video coding efficiency. In this paper, a higher order transform domain filter (HOTDF) is proposed to further reduce quantization noise. The filter fully exploits non-local spatial correlation by grouping similar patches and sparsely representing them in a higher order transform domain. Then quantization noise can be well removed by preserving large transform coefficients conveying majority of original signals and discarding the remaining smaller ones mainly caused by quantization. Experimental results show that compared with HM16.0, the proposed algorithm achieves 1.9%, 3.0% and 2.5% BD-rate reduction for luminance component on average for all intra, low delay P and random access configurations, respectively. Lu Yu 0003 |
ISCAS | 2 |
| 2017 | Coding optimization based on weighted-to-spherically-uniform quality metric for 360 videoabstractTo optimize coding efficiency, an encoder minimizes total distortion under a constraint rate. Distortion criteria for conventional video uniformly measures coding error around picture. For non-equal area projection formats of 360 video, e.g., equirectangular projection (ERP), cube map projection (CMP), uniformly measured distortion cannot correctly represent coding loss on a watching sphere of panoramic visual content. In this paper, coding optimization based on weighted-to-spherically-uniform quality criteria is proposed to improve coding performance of 360 video, where adaptive quantization is applied in CTU level according to WS-PSNR weights. Since quantization parameter (QP) offset can be derived according to CTU position of 360 video at decoder side, no additional parameters need to be transmitted. Experimental results show that compared to coding with fixed QP, the proposed adaptive quantization method achieves 4.50%, 8.51% and 3.21% BD-rate reduction for ERP, rotated ERP and CMP on average, respectively. Yule Sun, Lu Yu 0003 |
VCIP | 2 |
| 2017 | Blind 3D image quality assessment based on self-similarity of binocular features
Wujie Zhou, Shuangshuang Zhang, Lu Yu 0003, Weiwei Qiu, Yang Zhou 0011, Ting Luo 0001 |
Neurocomputing | 4 |
| 2017 | Local gradient patterns (LGP): An effective local-statistical-feature extraction scheme for no-reference image quality assessment
Wujie Zhou, Lu Yu 0003, Weiwei Qiu, Yang Zhou 0011, Mingwei Wu 0001 |
Inf. Sci. | 2 |
| 2017 | Blind quality estimator for 3D images based on binocular combination and extreme learning machine
Wujie Zhou, Lu Yu 0003, Yang Zhou 0011, Weiwei Qiu, Mingwei Wu 0001, Ting Luo 0001 |
Pattern Recognit. | 2 |
| 2017 | Weighted-to-Spherically-Uniform Quality Evaluation for Omnidirectional VideoabstractOmnidirectional video records a scene in all directions around one central position. It allows users to select viewing content freely in all directions. Assuming that viewing directions are uniformly distributed, the isotropic observation space can be regarded as a sphere. Omnidirectional video is commonly represented by different projection formats with one or multiple planes. To measure objective quality of omnidirectional video in observation space more accurately, a weighted-to-spherically-uniform quality evaluation method is proposed in this letter. The error of each pixel on projection planes is multiplied by a weight to ensure the equivalent spherical area in observation space, in which pixels with equal mapped spherical area have the same influence on distortion measurement. Our method makes the quality evaluation results more accurate and reliable since it avoids error propagation caused by the conversion from resampling representation space to observation space. As an example of such quality evaluation method, weighted-to-spherically-uniform peak signal-to-noise ratio is described and its performance is experimentally analyzed. Yule Sun, Ang Lu, Lu Yu 0003 |
IEEE Signal Process. Lett. | 3 |
| 2016 | General Synthesized View Distortion Estimation for Depth Map Compression of FTVabstractThe state-of-the-art depth map coding distortion measurement in view synthesis optimization (VSO) of 3D-HEVC reference software applies only for 1D parallel camera arrangement and requires explicit specification of virtual camera position(s) to be optimized. This paper presents a general joint distortion metric (GJDM) of synthesized views for depth map coding of FTV. The proposed metric estimates the expected distortion of all potential synthesized views generated by the depth map regardless of the arrangement of virtual cameras, thus supports arbitrary camera setups and various applications, e.g., free navigation. Furthermore, the proposed distortion metric is rendering-free and prone to parallelization. In a word, the proposed metric supports more general camera arrangement and achieves 21% coding gain and 6% encoding time reduction in average in free navigation test sequences. Ang Lu, Lu Yu 0003 |
DCC | 3 |
| 2016 | Full-reference perceptual quality assessment for stereoscopic images based on primary visual processing mechanismabstractWith the development of 3D technology, there is an urgent demand for accurate and efficient perceptual quality evaluation of stereoscopic images. In this paper, we propose a primary-visual-processing-mechanism-based metric (PVPM) for stereoscopic image quality assessment (SIQA). Considering binocular and monocular cells in primary visual cortex, we classify image regions into binocular and monocular ones. Inter-view difference of dominance appears in perception of binocular cells under asymmetric quality, so local quality of left and right binocular regions are weighted by stimuli strength indexes to obtain joint binocular local quality. The overall quality is generated through pooling of binocular and monocular local quality, in which visual sensitivity and region sizes are taken into account. Experimental results demonstrate that PVPM can achieve higher accuracy as well as faster speed compared to state-of-the-art SIQA metrics. Lu Yu 0003 |
ICME | 3 |
| 2016 | Effective HEVC intra coding unit size decision based on online progressive Bayesian classificationabstractIn High Efficiency Video Coding (HEVC), optimal coding unit (CU) size is decided based on recursive rate-distortion cost comparison, which consumes high computational resources. In this paper, an effective HEVC intra CU size decision algorithm is proposed to speed up the encoding process. The algorithm is based on progressive Bayesian classification, which is composed of two cascade classifiers: a three-class classifier and a binary classifier, at every coding depth. The three-class classifier firstly partitions the feature space into three regions: split, indistinct and non-split regions. Then, the binary classifier further partitions the indistinct region into split and non-split regions by utilizing additional complicated features. The thresholds of the two classifiers are well designed based on Bayesian risk to balance coding efficiency and complexity. All classifiers are online trained and experimental results show that the proposed algorithm can save approximately 51%~63% of the total encoding time of HM 15.0 with negligible loss on rate-distortion performance. Lu Yu 0003 |
ICME | 2 |
| 2016 | Binocular rivalry detection in natural image pairsabstractWhen different images are presented to two eyes, they compete for perceptual dominance, such that a region of one image is visible while corresponding region of the other is suppressed. This visual phenomenon is called binocular rivalry. Binocular rivalry may be introduced in stereoscopic images of natural scene, leading to strong visual discomfort and visual fatigue. When binocular differences exceed a threshold, binocular rivalry occurs. Binocular difference threshold of corresponding regions increases with background complexity of the region. Based on these features of human visual system, we proposed a detection model of binocular rivalry for natural image pairs. A binocular rivalry database was established and used to evaluate our detection model with the database. Experimental results demonstrate that the proposed model can achieve higher hit rate and lower false alarm rate compared with existing method. Yapeng Xue, Lu Yu 0003 |
ICME | 4 |
| 2016 | Spatial quality index based rate perceptual-distortion optimization for video coding
Xingguo Zhu, Lu Yu 0003, Yin Zhao |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Utilizing binocular vision to facilitate completely blind 3D image quality measurement
Wujie Zhou, Lu Yu 0003, Weiwei Qiu, Ting Luo 0001, Zhongpeng Wang, Mingwei Wu 0001 |
Signal Process. | 2 |
| 2016 | Binocular Responses for No-Reference 3D Image Quality AssessmentabstractPerceptual quality assessment of distorted three-dimensional (3D) images has become a fundamental yet challenging issue in the field of 3D imaging. In this paper, we propose a general-purpose blind/no-reference (NR) 3D image quality assessment (IQA) metric that utilizes the complementary local patterns (the local magnitude pattern and the proposed generalized local directional pattern) of binocular energy response (BER) and binocular rivalry response (BRR). The main technical contribution of this research is that binocular visual perception and local structural distribution are considered for NR 3D-IQA. More specifically, the metric simulates the binocular visual perception using BER and BRR. Subsequently, the local patterns of the binocular responses' encoding maps are used to form various binocular quality-predictive features, which will change in the presence of distortions. After feature extraction, we use k-nearest neighbors-based machine learning to drive the overall quality score. We tested our proposed metric against two publicly available 3D databases; these tests confirm that the proposed metric's results consistently align with human subjective judgments. Wujie Zhou, Lu Yu 0003 |
IEEE Trans. Multim. | 2 |
| 2015 | On the Efficiency of View Synthesis Prediction for 3D Video CodingabstractWe study the efficiency of view synthesis prediction (VSP). The proposed spectral domain analysis relates the power spectral density of the VSP error to the probability density function of the warping error. The analysis takes into account the warping error induced by (i) depth coding and (ii) disparity rounding at integer-pel, half-pel and quarter-pel warping accuracy. The interaction between VSP efficiency and interpolation filter is also studied. We validate our proposed model with empirical data. Using the proposed model, we discuss the interaction between prediction efficiency, depth image distortion, warping accuracy and interpolation filter. The proposed model provides theoretical insights of VSP in 3D-HEVC. It could be used to guide optimization of VSP designs. Ngai-Man Cheung, Lu Yu 0003 |
DCC | 3 |
| 2015 | The efficiency of view synthesis prediction for 3D video coding: A spectral domain analysisabstractWe study the coding efficiency of view synthesis prediction (VSP) in 3D video coding. Our spectral domain analysis relates the power spectral density (PSD) of the VSP prediction error to the probability density function (pdf) of the warping error. Our analysis takes into account the warping error induced by (i) depth coding and (ii) rounding error at integer-pel, half-pel and quarter-pel warping accuracy. We also study the interaction between depth coding error and warping accuracy. Our model suggests that the coding gain with using higher warping accuracy diminishes as the depth coding error increases. Our analysis results are validated with empirical data. Ngai-Man Cheung, Lu Yu 0003 |
ICASSP | 3 |
| 2015 | A novel interpolation-free scheme for fractional pixel motion estimationabstractFractional pixel motion estimation (ME), including fractional pixel interpolation and fractional motion vector (MV) prediction, has been implemented within many modern video codecs to optimize the rate-distortion performance. However, it brings about significant computational complexity increase. This paper proposes an interpolation-free fractional pixel ME scheme. Based on the analysis of the property of the error surface established with the block distortion measurement (BDM) of the integer pixel points, a parabolic curve passing the minimum point of the error surface is derived. Then the optimal fractional MV position is determined as the minimum point of the curve. Experimental results show that the proposed algorithm significantly reduces computational complexity of fractional pixel ME and outperforms the state-of-the-art interpolation-free fractional pixel ME algorithms. Xuguang Zuo, Lu Yu 0003 |
PCS | 2 |
| 2015 | Library based coding for videos with repeated scenesabstractIn video coding, the coding efficiency of intra frames (I frames) is limited since they are coded independently. Besides, in videos with repeated scenes, high correlations exist between video contents far apart, which can be used to improve the coding efficiency. In this article, a library based video coding scheme is proposed. In the proposed scheme, a scene library containing the common scene frames (CSFs) of a video is firstly built, and then original I frames are coded as inter frames by referencing CSFs in the library. With the aid of the library, the correlations between video contents far part can be well exploited. Also by only referencing the library, the coding bits of traditional I frames are reduced while the functionalities of original I frames are still preserved. Experimental results show that the coding efficiency of videos with repeated scenes is significantly improved by the proposed scheme. Xuguang Zuo, Lu Yu 0003 |
PCS | 2 |
| 2015 | A reversibility-gain model for integer Karhunen-Loève transform design in video codingabstractKarhunen-Loève transform (KLT) is the optimal transform that minimizes distortion at a given bit allocation for Gaussian source. As a KLT matrix usually contains non-integers, integer-KLT design is a classical problem. In this paper, a joint reversibility-gain (R-G) model is proposed for integer-KLT design in video coding. Specifically, the ‘reversibility’ is modeled according to distortion analysis in using forward and inverse integer transform without quantization. It not only measures how invertible a transform is, but also bounds the distortion introduced by the non-orthonormal integer transform process. The ‘gain’ means transform coding gain (TCG), which is a widely used criterion for transform design in video coding. Since KLT maximizes the TCG under some assumptions, here we define the TCG loss ratio (LR) to measure how much coding gain an integer-KLT loses when compared with the original KLT. Thus, the R-G model can be explained as follows: subject to a certain TCG LR, an integer- KLT with the best reversibility is the optimal integer transform for a given non-integer-KLT. Experimental results show that the R-G model can guide the design of integer-KLTs with good performance. Xingguo Zhu, Lu Yu 0003 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2014 | A reconfiguration system for video decoderabstractThis demonstration system shows a kind of video decoder's implementation in Reconfigurable Video Coding (RVC) framework on Open RVC-CAL Compiler (Orcc) platform. Differently from tradition video decoder, the reconfigurable video decoder is not a decoder conforming a special video coding standard, but dynamically built according to actual bitsteams, which may not conform any standard. The reconfigurable video decoder receives not only the compressed video bitstream but also the decoder description. As an example, in this demo, we reconfigure AVS and H.264/AVC decoders using Just-In-Time Adaptive Decoder Engine (Jade). Honggang Qi, Dandan Ding, Lu Yu 0003 |
VCIP | 4 |
| 2014 | A hardware-oriented IME algorithm and its implementation for HEVCabstractThe flexible coding structure in High Efficiency Video Coding (HEVC) introduces many challenges to real-time implementation of the integer-pel motion estimation (IME). In this paper, a hardware-oriented IME algorithm naming parallel clustering tree search (PCTS) is proposed, where various prediction units (PU) are processed simultaneously with a parallel scheme. The PCTS consists of four hierarchical search steps. After each search step, PUs with the same MV candidate are clustered to one group. And the next search step is shared by PUs in the same group. Owing to the top-down tree-structure search strategy of the PCTS, search processes are highly shared among different PUs and system throughput is thus significantly increased. As a result, the hardware implementation based on the proposed algorithm can support real-time video applications of QFHD (3840×2160) at 30fps. Dandan Ding, Lu Yu 0003 |
VCIP | 3 |
| 2014 | A cost-efficient hardware architecture of deblocking filter in HEVCabstractThis paper presents a hardware architecture of deblocking filter (DBF) for High Efficiency Video Coding (HEVC) by jointly considering system throughput and hardware cost. A hybrid pipeline with two processing levels is adopted to improve system performance. With the hybrid pipeline, only one 1-D filter and single-port on-chip SRAM are used. According to the data dependence between neighbouring edges, a shifted 16×16 basic processing unit as well as corresponding filtering order is proposed. It reduces memory cost and makes the DBF friendlier to work in a coding/decoding system. The proposed hardware architecture is synthesized under 0.13um standard CMOS technology and result shows that it consumes 17.6k gates at an operating frequency of 250MHz. Consequently, the design can support real-time processing of QFHD (3840×2160) video applications at 60 fps. Dandan Ding, Lu Yu 0003 |
VCIP | 3 |
| 2014 | Fast mode decision method for all intra spatial scalability in SHVCabstractScalable high efficiency video coding (SHVC) is now being developed by the Joint Collaborative Team on Video Coding (JCT-VC). In SHVC, the enhancement layer (EL) employs the same tree structured coding unit (CU) and 35 intra prediction modes as the base layer (BL), which results in heavy computation load. To speed up the mode decision process in the EL, the correlations of the CU depth and intra prediction modes between the BL and the EL are exploited in this paper. Based on the correlations an EL CU depth early skip algorithm and a fast intra prediction mode decision algorithm are proposed for all intra spatial scalability. Experimental results show that 45.3% and 42.3% coding time of the EL can be saved in AH Intra 1.5× spatial scalability and 2× spatial scalability respectively. In the meantime, the R-D performance degraded less than 0.05% compared with SHVC Test Model (SHM) 5.0. Xuguang Zuo, Lu Yu 0003 |
VCIP | 2 |
| 2014 | Multipath Routing of Multiple Description Coded Images in Wireless Networks
Yuanyuan Xu 0001, Ce Zhu, Lu Yu 0003 |
J. Comput. Sci. Technol. | 3 |
| 2014 | Pixel-Based Inter Prediction in Coded Texture Assisted Depth CodingabstractThis letter presents a pixel-based motion estimation scheme assisted with the coded texture video for depth inter-prediction, in view of motion similarity between depth and texture video. The proposed scheme can achieve higher inter-prediction gain without transmitting any motion vector in the pixel-based motion estimation. Coupled with depth-texture structure similarity, the inter prediction method is further extended to an integrated prediction approach by making use of both intra and inter information. Experimental results show that our proposed method achieves superior rate-distortion performance. Shuai Li 0005, Jianjun Lei 0001, Ce Zhu, Lu Yu 0003, Chunping Hou |
IEEE Signal Process. Lett. | 4 |
| 2014 | Block-Based In-Loop View Synthesis for 3-D Video CodingabstractView synthesis prediction (VSP) employs a synthesized picture as a reference picture for current-view texture coding, which is an advanced disparity-compensated prediction. However, the picture-based view synthesis demands huge complexity, especially for decoders. Therefore, we propose a block-based in-loop view synthesis scheme which generates VSP samples only for blocks using VSP modes (called target blocks). For a target block, a window in reference view is estimated. Then, pixels within the window are warped to the current view, producing VSP samples for the target block. The proposed method turns the picture-level VSP sample generation into macroblock-level process, and significantly reduces complexity of the VSP module while maintaining coding efficiency. Yin Zhao, Lu Yu 0003 |
IEEE Signal Process. Lett. | 3 |
| 2013 | Synthesis distortion estimation in 3D video using frequency and spatial analysisabstractWe propose an analytical model to estimate the synthesized view quality in 3D video. Specifically, we estimate the depth-error induced distortion using an approach that combines frequency and spatial domain analysis. We also propose to decompose the spatial-variant video signals into gradient-based representations to capture the interaction between image gradients, depth errors and synthesis distortion. Experiment results with video sequences and coding/rendering tools used in MPEG 3DV activities show that our analytical model can accurately estimate the synthesis noise power. Lu Fang 0001, Ngai-Man Cheung, Dong Tian, Anthony Vetro, Huifang Sun, Lu Yu 0003 |
ICIP | 6 |
| 2013 | Framework of AVS2-video codingabstractThe first generation of Audio-video coding standard (AVS1) will become an international standard of IEEE (IEEE.1857). The second generation of Audio-video coding standard (AVS2) is a successor to IEEE.1857 which targets to higher coding efficiency, especially for high resolution videos, with reasonable complexity increase. Currently, AVS2 working draft is developed, in which symbolic techniques of IEEE.1857, such as logarithmic arithmetic coding, are still used. It also introduces new coding tools, e.g. flexible partition structure which is based on macro-block structure and high efficient prediction methods utilizing more texture information and temporal redundancies, etc. The performance of AVS2 is improved by 37% compared to the IEEE.1857 Jiaqiang Profile and 16.4% compared to H.264 in average. This paper provides an overview of the framework of AVS2-video coding standard. Zhichu He, Lu Yu 0003, Xiaozhen Zheng, Siwei Ma 0001 |
ICIP | 2 |
| 2013 | Synthesized disparity vectors for 3D video codingabstractIn 3D video coding, some dependent-view coding tools may utilize derived disparity vectors (DV) to locate inter-view correspondences, such as the Backward block-based View Synthesis Prediction (BVSP) and Depth-based Motion Vector Prediction (DMVP) in ATM. A typical way to derive a DV for a target block, as performed in ATM, is to convert reconstructed depth values associated with the block to a DV. However, this approach only works in depth-first coding order (DFCO) and becomes inapplicable when the dependent-view texture is coded prior to the depth, i.e., in texture-first coding order (TFCO). In this paper, a sparse DV field is synthesized from the depth map of a coded reference view to provide the DVs required by dependent-view coding tools in TFCO. The synthesized sparse DV field is accurate to support disparity-aided coding tools in different coding orders, while only introducing 2% decoding time increase. With the proposed method, BVSP and DMVP can be applied in both TFCO and DFCO, which mitigates the large coding performance gap (around 20% BD rate) between TFCO and DFCO of ATM. Yin Zhao, Lu Yu 0003 |
ICIP | 4 |
| 2013 | Subjective study of binocular rivalry in stereoscopic images with transmission and compression artifactsabstractBinocular rivalry is a visual phenomenon that occurs when two eyes are presented with different patterns. Instead of being fused to form a unitary visual impression, the two patterns may be perceived alternately or evoke a peculiar shimmering effect. In stereoscopic images, corresponding regions in two views may be dissimilar due to artifacts from lossy video processing, e.g., transmission and compression. In this case, binocular rivalry may be introduced, which impairs 3D visual quality. In this paper, we established a database of binocular rivalry artifacts in stereoscopic images with transmission and compression distortions. Ten subjects were engaged to mark binocular rivalry artifacts in ten stereoscopic images. The performance of each subject was analyzed using the hit rate and false alarm rate of the subject compared with the average marking results of the other subjects. The subjective data were then combined into maps which indicate the locations and strength of the rivalry artifacts in the stereo pairs. The database is aimed to facilitate the validation of binocular rivalry artifact detection algorithms, and provides cues for stereoscopic 3D quality assessment and binocular rivalry artifact removal. Yin Zhao, Lu Yu 0003 |
ICIP | 3 |
| 2013 | A hardware CABAC encoder for HEVCabstractThis paper presents a hardware design of context-based adaptive binary arithmetic coding (CABAC) for the emerging High efficiency video coding (HEVC) standard. While aiming at higher compression efficiency, the CABAC in HEVC also invests a lot of effort in the pursuit of parallelism and reducing hardware cost. Simulation results show that our design processes 1.18 bins per cycle on average. It can work at 357 MHz with 48.940K gates targeting 0.13 μm CMOS process. This processing rate can support real-time encoding for all sequences under common test conditions of HEVC standard conforming to the main profile level 6.1 of main tier or main profile level 5.1 of high tier. Dandan Ding, Xingguo Zhu, Lu Yu 0003 |
ISCAS | 4 |
| 2013 | Rate-Distortion Optimization for depth map coding with distortion estimation of synthesized viewabstractIn 3D video systems, depth maps supply geometry information which is used to generate virtual views. The coding of the depth maps can be improved by considering distortion of synthesized views instead of depth map distortion. Therefore, this paper proposes a novel metric for depth coding, which quantifies the impact of the depth distortion on the fidelity of synthesized views, without banding with any complete rendering process. The metric is integrated into Rate-Distortion Optimization (RDO) process for mode decision in depth map coding, replacing the metric in HTM software. Experimental results show that the proposed metric significantly improves the coding efficiency of 3D video by about 10% bitrate saving compared with conventional RDO. Moreover, the RDO with the proposed metric achieves stable coding efficiency for different decoder-side renderers with low computation complexity. Lu Yu 0003 |
ISCAS | 2 |
| 2013 | Fast HEVC intra coding unit size decision based on an improved Bayesian classification frameworkabstractHEVC achieves much higher coding efficiency compared with previous coding standards at cost of significant computation complexity. In this paper, an improved Bayesian classification framework is constructed to accelerate the coding unit size decision for HEVC intra coding. Under this framework, different decision mechanisms are applied for different regions of feature space. Specifically, the feature space is partitioned into two regions: high-discrimination region and low-discrimination region. In high-discrimination region, Bayesian classifier is employed to assist mode decision, while in low-discrimination region, Rate-Distortion-cost based decision is utilized. An optimal method of feature space partition is theoretically analyzed by modeling the problem as a constrained optimization problem. Moreover, the feature used in the improved Bayesian classification framework is carefully selected. Experimental results show that the proposed algorithm can save on average 45.3% of the total encoding time of HM 10.0 with only 0.5% increases in BD-rate. Xuefei Fang, Xingguo Zhu, Lu Yu 0003, Xiao-lin Shen |
PCS | 3 |
| 2013 | On hardware architecture and processing order of HEVC intra prediction moduleabstractThis article presents a parallel and memory optimized hardware architecture for intra prediction of the High Efficiency Video Coding (HEVC) standard. The architecture consists of 64 parallel reconfigurable Processing Elements as datapaths and supports all 35 intra prediction modes and all prediction sizes from 4×4 to 64×64. In order to avoid implementing large area memory-datapaths interconnections and save memory usage, the maximum number of reference registers is reduced from 129 to 72 by reclassifying 35 prediction modes into 3 general categories. In addition, a 3 stage hierarchical processing order including an S-shaped scan order of blocks and a Bidirectional Ring Register File is proposed to avoid bandwidth bottleneck and increase system throughput. This architecture is synthesized using TSMC 130nm technology and the working frequency is up to 400MHz with 324K gates area. Running at 300MHz, it supports real time 1080p@60fps full modes and full sizes HEVC intra prediction. Dandan Ding, Lu Yu 0003 |
PCS | 3 |
| 2012 | Fast coding unit size selection for HEVC based on Bayesian decision ruleabstractHigh Efficiency Video Coding (HEVC) is being developed by Joint Collaborative Team on Video Coding (JCT-VC). The new standard aims at a significant improvement on Rate Distortion (RD) performance over state-of-the-art video coding standards. Some new efficient coding tools are adopted, such as flexible data structure representation: Coding Unit (CU), Prediction Unit (PU), and Transform Unit (TU). Meanwhile, these new tools impose lots of complexity. We propose a CU size decision algorithm that greatly reduces the complexity of HEVC. The proposed algorithm collects relevant and computational-friendly features to assist decisions on CU splitting. To minimize the RD loss of the fast algorithm, a Bayesian decision rule is defined, based on which the encoder can save 41.4% encoding time on average (up to 66%) whilst suffers from negligible loss (1.88%) on RD performance. Xiao-lin Shen, Lu Yu 0003 |
PCS | 2 |
| 2012 | Encoder-embedded temporal-spatial wiener filter for video encodingabstractNoise not only degrades the visual quality of video contents, but also significantly affects the coding efficiency. Video filters can achieve a better compression performance by reducing the noise. However, they will blur video details. So a subjective quality considered filter with motion detection, working in the DCT domain, is proposed. This is a motion compensated 3D filter with low complexity. The picture will be divided into two areas. The static area will be filtered much more than the motion area. The experimental results show that noise is well eliminated and coding efficiency is improved. Huiming Tang, Minjun Fan, Lu Yu 0003 |
PCS | 3 |
| 2012 | Binarization and context model selection of CABAC based on the distribution of syntax elementabstractContext-based Adaptive Binary Arithmetic Coding (CABAC) is an efficient entropy coding method in H.264/AVC, which consists of three processes: binarization, context model selection (CS) and binary arithmetic encoding (BAE). This paper proposes a new binarization and CS method to encode mb_type of depth videos. The proposed method includes (1) remapping of mb_type values based on the distribution of the mb_type; (2) classification and binarization inspired from Configurable Variable Length Code (CVLC) and (3) simple CS without using neighboring context information. Experimental results show that up to 12.55% bitrate reduction can be achieved at hierarchical B structure compared to Multi-view Video Coding (MVC) on depth videos. Xingguo Zhu, Lu Yu 0003 |
PCS | 2 |
| 2011 | A parallel and area-efficient architecture for deblocking filter and Adaptive Loop FilterabstractAdaptive Loop Filter (ALF) has been developed lately to improve the video coding performance. It is inserted between deblocking and inter-prediction, which makes deblocking and ALF very time-critical because they are conducted sequentially. In this paper, we propose an efficient architecture integrating deblocking and ALF for the decoder. The architecture not only implements deblocking and ALF in parallel but also reduces area cost as much as possible. These are achieved by shared hybrid organized memory architecture and one-block-two-edge parallel strategy using a novel filter order. The proposed architecture is implemented in verilog HDL and can achieve real-time decoding for 1080p @ 30 fps applications by working at 211MHz in a Xilinx Virtex-5 FPGA. Lu Yu 0003 |
ISCAS | 2 |
| 2011 | Motion estimation with Second Order PredictionabstractTo exploit both spatial and temporal correlation of video signal, a Second Order Prediction (SOP) scheme has been presented to eliminate the spatial correlation existing in motion compensated residue. Although SOP achieves better RD performance by adopting residue prediction in Mode Decision (MD) stage, it can be further improved by combining residue prediction into Motion Estimation (ME). Our analysis demonstrates that a more accurate motion vector could be obtained if both ME and MD take residue prediction into account. Implementation of the proposed algorithm into the H.264/AVC reference software demonstrates that the enhanced SOP can achieve up to 18.46% of bit-rate saving at equivalent objective quality. Shangwen Li, Lu Yu 0003 |
ISCAS | 2 |
| 2011 | Stereoscopic video coding in AVSabstractThis paper is an overview for AVS stereoscopic video coding technology, including two channels based inter-view prediction coding and stereo packing mode coding. The first one utilizes inter-view prediction to efficiently exploit the redundancy between the two channels of stereoscopic video. The superior coding performance of the inter-view prediction scheme benefits from an enhanced block prediction algorithm, which includes an improvement of direct mode for B-picture and motion vector prediction for P-picture. In addition, stereo packing mode, including side by side and top bottom, is adopted in AVS stereoscopic video coding to support the stereoscopic video service deployments based on the frame-compatible approach. Furthermore, an enhanced stereo packing mode is also developed to allow the prediction between signals coming from different channels within one packed frame. The simulation results demonstrate that the adopted techniques in AVS stereoscopic video are able to improve the compression efficiency of stereoscopic videos compared to simulcast one. Xiangyang Ji, Yongbing Zhang 0002, Lu Yu 0003, Gwo Giun Lee |
VCIP | 3 |
| 2011 | Cross-view post-filtering for fidelity enhancement on asymmetric coding of 3D videoabstract3D video employing depth-image-based rendering (DIBR) typically contains a stereo pair plus two associated depth maps. The stereo pair may be compressed asymmetrically (e.g., using mixed resolution coding) to effectively reduce the bit rate while maintaining the stereoscopic visual quality at the same level as that from symmetric coding. With depth information, it is straightforward to locate high-fidelity (HF) inter-view correspondences of pixels in the low-fidelity (LF) view. However, we find that LF-view fidelity cannot be consistently improved by simply substituting LF pixels with their HF counterparts, due to inter-view color/geometric differences and depth errors. In this paper, we propose an effective post-filter which first checks the coherence between local LF and HF waveforms to distinguish unreliable correspondences. Then, the LF view is adaptively rectified by the reliable HF pixels to improve its fidelity. Experimental results show that the proposed cross-view fidelity enhancement scheme can promote Peak Signal-to-Noise Ratio (PSNR) of LF views by up to 1.1 dB, and can effectively suppress ringing and blocking artifacts in flat regions of LF views. Yin Zhao, Lu Yu 0003 |
VCIP | 2 |
| 2011 | Distributed video coding with adaptive selection of hash functionsabstractWe address the compression efficiency of feedback-free and hash-check distributed video coding, which generates and transmits a hash code of a source information sequence. The hash code helps the decoder perform a motion search. A hash collision is a special case in which the hash codes of wrongly reconstructed information sequences occasionally match the hash code of the source information sequence. This deteriorates the quality of the decoded image greatly. In this paper, the statistics of hash collision are analyzed to help the codec select the optimal trade-off between the probability of hash collision and the length of the hash code, according to the principle of rate-distortion optimization. Furthermore, two novel algorithms are proposed: (1) the nonzero prefix of coefficients (NPC), which indicates the count of nonzero coefficients of each block for the second algorithm, and also saves 8.4% bitrate independently; (2) the adaptive selection of hash functions (AHF), which is based on the NPC and saves a further 2%–6% bitrate on average. The detailed optimization of the parameters of AHF is also presented. Xin-Hao Chen, Lu Yu 0003 |
J. Zhejiang Univ. Sci. C | 2 |
| 2011 | Hash signature saving in distributed video codingabstractIn transform-domain distributed video coding (DVC), the correlation noises (denoted as N ) between the source block and its temporal predictor can be modeled as Laplacian random variables. In this paper we propose that the noises (denoted as N ′) between the source block and its co-located block in a reference frame can also be modeled as Laplacian random variables. Furthermore, it is possible to exploit the relationship between N and N ′ to improve the performance of the DVC system. A practical scheme based on theoretical insights, the hash signature saving scheme, is proposed. Experimental results show that the proposed scheme saves on average 83.2% of hash signatures, 13.3% of bit-rate, and 3.9% of encoding time. Xin-Hao Chen, Xingguo Zhu, Xiao-lin Shen, Lu Yu 0003 |
J. Zhejiang Univ. Sci. C | 4 |
| 2011 | An efficient hardware design for HDTV H.264/AVC encoderabstractThis paper presents a hardware efficient high definition television (HDTV) encoder for H.264/AVC. We use a two-level mode decision (MD) mechanism to reduce the complexity and maintain the performance, and design a sharable architecture for normal mode fractional motion estimation (NFME), special mode fractional motion estimation (SFME), and luma motion compensation (LMC), to decrease the hardware cost. Based on these technologies, we adopt a four-stage macro-block pipeline scheme using an efficient memory management strategy for the system, which greatly reduces on-chip memory and bandwidth requirements. The proposed encoder uses about 1126k gates with an average Bjontegaard-Delta peak signal-to-noise ratio (BD-PSNR) decrease of 0.5 dB, compared with JM15.0. It can fully satisfy the real-time video encoding for 1080p@30 frames/s of H.264/AVC high profile. Dandan Ding, Bin-bin Yu, Lu Yu 0003 |
J. Zhejiang Univ. Sci. C | 5 |
| 2011 | Binocular Just-Noticeable-Difference Model for Stereoscopic ImagesabstractConventional 2-D Just-Noticeable-Difference (JND) models measure the perceptible distortion of visual signal based on monocular vision properties by presenting a single image for both eyes. However, they are not applicable for stereoscopic displays in which a pair of stereoscopic images is presented to a viewer's left and right eyes, respectively. Some unique binocular vision properties, e.g., binocular combination and rivalry, need to be considered in the development of a JND model for stereoscopic images. In this letter, we propose a binocular JND (BJND) model based on psychophysical experiments which are conducted to model the basic binocular vision properties in response to asymmetric noises in a pair of stereoscopic images. The first experiment exploits the joint visibility thresholds according to the luminance masking effect and the binocular combination of noises. The second experiment examines the reduction of visual sensitivity in binocular vision due to the contrast masking effect. Based on these experiments, the developed BJND model measures the perceptible distortion of binocular vision for stereoscopic images. Subjective evaluations on stereoscopic images validate of the proposed BJND model. Yin Zhao, Ce Zhu, Yap-Peng Tan, Lu Yu 0003 |
IEEE Signal Process. Lett. | 5 |
| 2011 | Video Quality Assessment Based on Measuring Perceptual Noise From Spatial and Temporal PerspectivesabstractVideo quality assessment (VQA) exploits important properties of the sophisticated human visual system (HVS). In this paper, we study a series of fundamental HVS characteristics for subjective video quality assessment, and incorporate them into a systematic framework to simulate subjective evaluation on impaired videos. Based on this framework, we develop a novel full-reference metric, namely, perceptual quality index (PQI). Specifically, the proposed PQI metric comprises four major modules: 1) visual performance equation for the foveal and extra-foveal vision based on the cortical magnification theory; 2) perceptible noise detection using a spatial-temporal just noticeable difference model, and its quantification in both spatial and temporal channels, considering the varying error sensitivity due to the contrast and motion masking effects; 3) instantaneous error summation with inhibition of weak local distortions, and quality degradation accumulation over time that models the visual persistence and recency effect; and 4) fusion of the spatial and temporal noise intensities into a perceptual quality index. Compared with some state-of-the-art VQA models, the PQI metric, which exploits multiple visual properties, measures video quality more accurately and reliably on two VQA databases. Yin Zhao, Lu Yu 0003, Ce Zhu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Depth No-Synthesis-Error Model for View Synthesis in 3-D VideoabstractCurrently, 3-D Video targets at the application of disparity-adjustable stereoscopic video, where view synthesis based on depth-image-based rendering (DIBR) is employed to generate virtual views. Distortions in depth information may introduce geometry changes or occlusion variations in the synthesized views. In practice, depth information is stored in 8-bit grayscale format, whereas the disparity range for a visually comfortable stereo pair is usually much less than 256 levels. Thus, several depth levels may correspond to the same integer (or sub-pixel) disparity value in the DIBR-based view synthesis such that some depth distortions may not result in geometry changes in the synthesized view. From this observation, we develop a depth no-synthesis-error (D-NOSE) model to examine the allowable depth distortions in rendering a virtual view without introducing any geometry changes. We further show that the depth distortions prescribed by the proposed D-NOSE profile also do not compromise the occlusion order in view synthesis. Therefore, a virtual view can be synthesized losslessly if depth distortions follow the D-NOSE specified thresholds. Our simulations validate the proposed D-NOSE model in lossless view synthesis and demonstrate the gain with the model in depth coding. Yin Zhao, Ce Zhu, Lu Yu 0003 |
IEEE Trans. Image Process. | 4 |
| 2010 | Evaluating video quality with temporal noiseabstractHuman can only perceive video distortions beyond certain intensity. Temporal variations of suprathreshold noise in a video induce annoying temporal noise (e.g., flickering artifacts). Prior studies on full-reference video quality assessment (VQA) focused mainly on spatial quality of impaired videos, and ignored or underestimated the impact of temporal noise on visual quality degradation. According to some characteristics of human visual system, we consider temporal noise as the most salient distortion in impaired videos and it can greatly influence the perceived video quality. Thus, we propose a novel and simple full-reference metric that evaluates quality of impaired videos by measuring the energy of temporal noise in them. This metric shows competitive performance with some state-of-the-art objective models on the LIVE VQA database. Yin Zhao, Lu Yu 0003 |
ICME | 2 |
| 2010 | Temporal consistency enhancement on depth sequencesabstractCurrently, depth sequences generated by automatic depth estimation suffer from the temporal inconsistency problem. Estimated depth values of some objects vary in adjacent frames, whereas the objects actually remain on the same depth planes. These temporal depth errors significantly impair the visual quality of the synthesized virtual view as well as the coding efficiency of the depth sequences. Since depth sequences correspond to texture sequences, some erroneous temporal depth variations can be detected by analyzing temporal variations of the texture sequences. Utilizing this property, we propose a novel solution to enhance the temporal consistency of depth sequences by applying adaptive temporal filtering on them. Experiments demonstrate that the proposed depth filtering algorithm can effectively suppress transient depth errors and generate more stable depth sequences, resulting in notable temporal quality improvement of the synthesized views and higher coding efficiency on the depth sequences. Deliang Fu, Yin Zhao, Lu Yu 0003 |
PCS | 3 |
| 2010 | Block-based second order prediction on AVS-Part 2abstractAVS-Part 2 is a mainstream video coding standard with high compression efficiency similar to H.264/AVC. A technique named Second Order Prediction (SOP) has been presented based on H.264/AVC to decrease the signal correlation after motion-compensated prediction. To achieve better coding performance, this paper presents a method named Block-based Second Order Prediction (BSOP) to ameliorate SOP to adapt to the features of the motion-compensation in AVS-P2 with analysis and demonstration in detail. Experimental results show that the proposed BSOP can outperform AVS-P2 P-picture coding by 3.99% bit-rate saving (0.126dB BD-PSNR gain) on average, and performs better than SOP implemented on AVS by 1.81% bit-rate saving. Bin-bin Yu, Shangwen Li, Lu Yu 0003 |
PCS | 3 |
| 2010 | Suppressing texture-depth misalignment for boundary noise removal in view synthesisabstractDuring view synthesis based on depth maps, also known as Depth-Image-Based Rendering (DIBR), annoying artifacts are often generated around foreground objects, yielding the visual effects that slim silhouettes of foreground objects are scattered into the background. The artifacts are referred as the boundary noises. We investigate the cause of boundary noises, and find out that they result from the misalignment between texture and depth information along object boundaries. Accordingly, we propose a novel solution to remove such boundary noises by applying restrictions during forward warping on the pixels within the texture-depth misalignment regions. Experiments show this algorithm can effectively eliminate most boundary noises and it is also robust for view synthesis with compressed depth and texture information. Yin Zhao, Dong Tian, Ce Zhu, Lu Yu 0003 |
PCS | 5 |
| 2010 | Motion-compensated filtering of reference picture for video codingabstractNoisy video occurs frequently in many visual communication applications, especially in video surveillance and video phone. The noise affects compression a lot by decreasing the accuracy of prediction and increasing the residual. In this paper, we propose a novel inter-frame filtering method based on motion compensation for reference picture in video coding. It constructs the low-pass temporal filter by modulating the constructed picture with previous reference picture. As a result, it greatly reduces the noise in reference picture, which results in the improvement of compression efficiency. As the replacement of traditional reference picture, this proposed dynamical motion compensation based Synthesized Reference (SR) picture is computation efficient and easy for implementation. It can be applied to any block-based hybrid coding structure, such as H.264/AVC, AVS (Audio Video Coding Standard), MPEG-4. Simulation results show that more than 20% bit rate is reduced when applying SR picture to the encoding of typical video surveillance sequences. At the same time, the visual quality is greatly improved. Huiming Tang, Shenghui Lin, Lu Yu 0003, Yunhai Liu, Libo Yang |
VCIP | 5 |
| 2010 | A perceptual metric for evaluating quality of synthesized sequences in 3DV systemabstract3D Video system based on Depth-Image-Based Rendering relies on high quality depth data. Errors distributed randomly in depth map sequences induce annoying temporal noise, such as flickering and object shifting. Prior studies on video quality assessment focused mainly on spatial quality of the tested sequence and often ignored its temporal performance. In synthesized sequences, a large number of tiny geometric distortions and illumination differences are temporally constant and perceptually invisible. The dynamic noise impairs subjective quality of the sequences more greatly than the static spatial noise. Temporal quality plays a dominant role in overall quality assessment on the synthesized sequences with temporal instability problem. We propose a simple full-reference metric, Peak Signal to Perceptible Temporal Noise Ratio, to evaluate quality of synthesized sequences by measuring the perceptible temporal noise in them. Yin Zhao, Lu Yu 0003 |
VCIP | 2 |
| 2010 | Review of the current and future technologies for video compressionabstractMany important developments in video compression technologies have occurred during the past two decades. The block-based discrete cosine transform with motion compensation hybrid coding scheme has been widely employed by most available video coding standards, notably the ITU-T H.26 x and ISO/IEC MPEG- x families and video part of China audio video coding standard (AVS). The objective of this paper is to provide a review of the developments of the four basic building blocks of hybrid coding scheme, namely predictive coding, transform coding, quantization and entropy coding, and give theoretical analyses and summaries of the technological advancements. We further analyze the development trends and perspectives of video compression, highlighting problems and research directions. Lu Yu 0003 |
J. Zhejiang Univ. Sci. C | 1 |
| 2009 | A Proposed AVS Decoder Configuration in the Reconfigurable Video Coding FrameworkabstractThis demonstration shows an AVS intra decoder configuration in the RVC framework. It explains how to use the dataflow mechanism offered by the RVC framework to support AVS decoder configuration. It also shows the flexibility and convenience to reconfigure decoders in the RVC framework. In this work, the AVS VTL is established containing FUs from AVS. The proposed AVS decoder configuration is implemented by connecting some FUs from AVS VTL and reusing some FUs from MPEG VTL. The demonstration shows that the decoder can decode AVS conformance bitstreams correctly in RVC simulator. Dandan Ding, Honggang Qi, Lu Yu 0003, Tiejun Huang 0001, Wen Gao 0001 |
ISCAS | 3 |
| 2009 | A Hybrid Decoder Configuration of MPEG-4 and AVS in Reconfigurable Video Coding FrameworkabstractThis demonstration shows that the RVC framework offers a great flexibility in selecting coding tools for decoder reconfigurations to satisfy a wide variety of different applications by showing a hybrid decoder reconfiguration using coding tools from MPEG and AVS Video Tool Library (VTL). Compared with the original MPEG-4 Simple Profile (SP), complexity of the reconfigured decoder is reduced whereas the performance is improved gradually as the bitrate increases. Dandan Ding, Lu Yu 0003, Christophe Lucarz, Marco Mattavelli |
ISCAS | 2 |
| 2009 | Second Order Prediction on H.264/AVCabstractBased on a relatively simple motion assumption, the technique of motion-compensated prediction will result in large residue when dealing with video contents containing complex motions. To achieve higher coding efficiency, this paper presents a method named second order prediction (SOP) to exploit remaining signal correlation after motion-compensated prediction. Demonstration and analysis on the SOP concept are provided along with its technical details. Experimental results show that the proposed SOP can outperform H.264/AVC P-picture coding 10.28% bit-rate saving under the same PSNR. Shangwen Li, Lu Yu 0003 |
PCS | 4 |
| 2009 | Reconfigurable video coding framework and decoder reconfiguration instantiation of AVS
Dandan Ding, Honggang Qi, Lu Yu 0003, Tiejun Huang 0001, Wen Gao 0001 |
Signal Process. Image Commun. | 3 |
| 2009 | Special issue on AVS and its applications: Guest editorial
Wen Gao 0001, King Ngi Ngan, Lu Yu 0003 |
Signal Process. Image Commun. | 3 |
| 2009 | Overview of AVS-video coding standards
Lu Yu 0003 |
Signal Process. Image Commun. | 1 |
| 2008 | L-shaped segmentations in motion-compensated prediction of H.264abstractPlaying the role of removing temporal correlation in video signals, motion-compensated prediction aims at the ideal prediction and eliminates the compensation. However, the block-based video coding infrastructure at present was not competent in exactly present the various shapes of moving objects in actual video. Contrast to that, an exact segmentation of coding area according to the actual moving objects in video is also not easy to achieve and will introduce significant computational complexity into the video coding scheme. Concerned with this conflict, this paper proposes a new set of segmentations of macroblocks, in which one macroblock splits into one L-shaped and one square segment. Experiments of this named "L-shaped segmentations" technique show that it can outperform the H.264 P-picture coding by 0.11 dB. Qichao Sun, Xiaoyang Wu 0001, Lu Yu 0003 |
ISCAS | 4 |
| 2008 | The Technique of Prescaled Integer Transform: Concept, Design and ApplicationsabstractInteger cosine transform (ICT) is adopted by H.264/AVC for its bit-exact implementation and significant complexity reduction compared to the discrete cosine transform (DCT) with an impact in peak signal-to-noise ratio (PSNR) of less than 0.02 dB. In this paper, a new technique, named prescaled integer transform (PIT), is proposed. With PIT, while all the merits of ICT are kept, the implementation complexity of decoder is further reduced compared to corresponding conventional ICT, which is especially important and beneficial for implementation on low-end processors. Since not all PIT kernels are good in respect of coding efficiency, design rules that lead to good PIT kernels are considered in this paper. Different types of PIT and their target applications are examined. Both fixed block-size transform and adaptive block-size transform (ABT) schemes of PIT are also studied. Experimental results show that no penalty in performance is observed with PIT when the PIT kernels employed are derived from the design rules. Up to 0.2 dB of improvement in PSNR for all intra frame coding compared to H.264/AVC can be achieved and the subjective quality is also slightly improved when PIT scheme is carefully designed. Using the same concept, a variation of PIT, post-scaled integer transform, can also be potentially designed to simplify the encoder in some special applications. PIT has been adopted in audio video coding standard (AVS), Chinese National Coding standard. Cixun Zhang, Lu Yu 0003, Wai-kuen Cham, Jie Dong 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Modeling Natural Image for Estimating DCT Coefficient Properties of Intra PredictionabstractAn elliptical symmetric model in spatial domain of natural image is proposed in this paper. Based on the model, the statistical properties of intra prediction error are studied. Furthermore, using the linear characteristic of Discrete Cosine Transform (DCT) and the proposed model, we offer a mathematical analysis of the properties of DCT coefficient for intra prediction error. The horizontal intra prediction mode is selected as the example. DCT coefficient properties of the rest prediction modes can be derived by the same method. Xiaoyang Wu 0001, Qichao Sun, Lu Yu 0003 |
ICME | 4 |
| 2007 | Block-Interleaved Error-Resilient Entropy CodingabstractThe variable-length coding (VLC) is widely used in video coding to improve compression efficiency. However, suffering from the loss of synchronization, VLC bit stream is much more sensitive to random errors than fixed-length coding (FLC) bit stream. The EREC is a valid tool combating random errors in VLC bit stream. Due to its intrinsic property of error propagation, when the EREC is applied to video bit stream, those blocks placed later become much more likely to be lost. This paper proposes a simple method to further improve the error robustness of video bit stream by interleaving transform coefficients of blocks so that low-frequency information is always placed ahead of high-frequency information. Thus, low-frequency information of greater significance is less likely to be lost. Experimental results prove the superiority of the proposed method. In addition, block interleaving can also be used in data-partitioned video bit stream with ease. Yong Fang 0001, Lu Yu 0003 |
ISCAS | 2 |
| 2007 | A Content-adaptive Fast Multiple Reference Frames Motion Estimation in H.264abstractThe H.264/AVC video coding standard adopts various advanced coding techniques such as variable block sizes and multiple reference frames motion compensation (MRF-MC). The computational burden of motion estimation increases as the number of reference frames searched. In this paper, we present a content-adaptive fast motion estimation algorithm. This algorithm fully uses the temporal and spatial content information at frame, MB and block levels to speed up the search process for multiple reference frames. Simulation results show that the proposed algorithm can be 3.6 times faster than the multiple reference frames UMHexagonS scheme, with negligible degradation of video quality. Qichao Sun, Xin-Hao Chen, Xiaoyang Wu 0001, Lu Yu 0003 |
ISCAS | 4 |
| 2007 | An Area-efficient VLSI Implementation of CA-2D-VLC Decoder for AVSabstractContext-based adaptive 2D-VLC (CA-2D-VLC) is adopted by AVS. In this paper, we present an area-efficient VLSI implementation of CA-2D-VLC decoder. Data Compression Storage (DCS) method is proposed in memory optimization for VLC tables and a reduction of 30% in on-chip memory cost is achieved. Furthermore, an Exp-Golomb decoder is developed with CodeWord Segmentation Decoding (CSD) method, which saves 70% hardware cost compared with the prior work. Synthesized with 0.18μCMOS standard-cell library, the overall hardware cost of the proposed CA-2D-VLC decoder is 1540 gates at the clock frequency constraint of 180MHz. With an average throughput of one symbol per cycle, the proposed design is suitable for cost-aware and high-resolution AVS video decoding applications. Though designed for AVS originally, the proposed architecture can be adapted to other coding standards easily. Xiaoyang Wu 0001, Lu Yu 0003 |
ISCAS | 3 |
| 2007 | Video transmission using advanced partial backward decodable bit stream (APBDBS)
Yong Fang 0001, Chengke Wu 0001, Lu Yu 0003 |
J. Vis. Commun. Image Represent. | 3 |
| 2006 | Low-Complexity Tools in AVS Part 7
Feng Yi, Qichao Sun, Jie Dong 0001, Lu Yu 0003 |
J. Comput. Sci. Technol. | 4 |