Shuyuan Zhu

dblp:84/6329 · DBLP profile ↗
← Back
113ranked-venue papers
28as first author
47since 2021 · last 2026
0000-0003-4450-3868ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 96 · 25 first-author · 36 since 2021Systems, architecture and hardware · 11 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 10 · 10 since 2021Computer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RIR-Agent: An interactive framework for effective and adaptive restoration of remote sensing imagery
Junyu Liu, Tianyu Li 0003, Lanyue Liang, Gang Fu 0003, Guoqing Wang 0001, Quan Rui, Xiongxin Tang, Shuyuan Zhu, Yang Yang 0002
Expert Syst. Appl.8
2026 Inter predictive coding for point cloud attributes with online coordinate alignment and multi-scale latent prediction
Yu Liu 0091, Shuyuan Zhu, Zeliang Li, Jeff Siu-Kei Au-Yeung, Fan Zhang 0017, Bing Zeng 0001
Neurocomputing2
2026 A fast video coding scheme based on perceptual rate distortion optimized preprocessing
Luheng Jia, Yifan Zang 0002, Jiyong Yu, Shuyuan Zhu, Li Song 0001, Kebin Jia
Multim. Syst.5
2026 FCVSR: A Frequency-Aware Method for Compressed Video Super-Resolution
abstract
Compressed video super-resolution (SR) aims to generate high-resolution (HR) videos from the corresponding low-resolution (LR) compressed videos. Recently, some compressed video SR methods attempt to exploit the spatio-temporal information in the frequency domain, showing great promise in super-resolution performance. However, these methods do not differentiate various frequency subbands spatially or capture the temporal frequency dynamics, potentially leading to suboptimal results. In this paper, we propose a deep frequency-based compressed video SR model (FCVSR) consisting of a motion-guided adaptive alignment (MGAA) network and a multi-frequency feature refinement (MFFR) module. Additionally, a frequency-aware contrastive loss is proposed for training FCVSR, in order to reconstruct finer spatial details. The proposed model has been evaluated on three public compressed video super-resolution datasets, with results demonstrating its effectiveness when compared to existing works in terms of super-resolution performance and complexity.
Fan Zhang 0017, Feiyu Chen 0001, Shuyuan Zhu, David Bull 0001, Bing Zeng 0001
IEEE Trans. Multim.4
2025 Plug-and-Play Tri-Branch Invertible Block for Image Rescaling
abstract
High-resolution (HR) images are commonly downscaled to low-resolution (LR) to reduce bandwidth, followed by upscaling to restore their original details. Recent advancements in image rescaling algorithms have employed invertible neural networks (INNs) to create a unified framework for downscaling and upscaling, ensuring a one-to-one mapping between LR and HR images. Traditional methods, utilizing dual-branch based vanilla invertible blocks, process high-frequency and low-frequency information separately, often relying on specific distributions to model high-frequency components. However, processing the low-frequency component directly in the RGB domain introduces channel redundancy, limiting the efficiency of image reconstruction. To address these challenges, we propose a plug-and-play tri-branch invertible block (T-InvBlocks) that decomposes the low- frequency branch into luminance (Y) and chrominance (CbCr) components, reducing redundancy and enhancing feature processing. Additionally, we adopt an all-zero mapping strategy for high-frequency components during upscaling, focusing essential rescaling information within the LR image. Our T-InvBlocks can be seamlessly integrated into existing rescaling models, improving performance in both general rescaling tasks and scenarios involving lossy compression. Extensive experiments confirm that our method advances the state of the art in HR image reconstruction.
Jingwei Bao, Jinhua Hao, Ming Sun 0008, Chao Zhou 0003, Shuyuan Zhu
AAAI6
2025 Blind Video Super-Resolution Based on Implicit Kernels
abstract
Blind video super-resolution (BVSR) is a low-level vision task which aims to generate high-resolution videos from low-resolution counterparts in unknown degradation scenarios. Existing approaches typically predict blur kernels that are spatially invariant in each video frame or even the entire video. These methods do not consider potential spatio-temporal varying degradations in videos, resulting in suboptimal BVSR performance. In this context, we propose a novel BVSR model based on Implicit Kernels, BVSR-IK, which constructs a multi-scale kernel dictionary parameterized by implicit neural representations. It also employs a newly designed recurrent Transformer to predict the coefficient weights for accurate filtering in both frame correction and feature alignment. Experimental results have demonstrated the effectiveness of the proposed BVSR-IK, when compared with four state-of-the-art BVSR models on three commonly used datasets, with BVSR-IK outperforming the second best approach, FMA-Net, by up to 0.59 dB in PSNR. Source code will be available at https://github.com/QZ1-boy/BVSR-IK.
Yuxuan Jiang 0015, Shuyuan Zhu, Fan Zhang 0017, David Bull 0001, Bing Zeng 0001
ICCV3
2025 Cross-Space Alignment-based Attribute Artifact Removal for V-PCC
abstract
In this paper, we propose a cross-space alignment-based attribute artifact removal method for V-PCC. Firstly, we construct a cross-space frame alignment module to align the adjacent projected 2D frames based on the temporal continuity between the 3D point clouds. Then, we design a multi-frame fusion network that is developed based on U-Net to fuse the current frame with the aligned frames to remove the artifact that occurred in the attribute. Experimental results demonstrate that our proposed method achieves superior performance for the enhancement of point clouds.
Yu Liu 0091, Mingfei Hu, Zeliang Li, Jeff Siu-Kei Au-Yeung, Shuyuan Zhu, Lihuo He, Fan Zhang 0017
ISCAS6
2025 Blind Image Super-Resolution with Local and Global Dual-Guidance
abstract
Blind image super-resolution (BISR) aims to recover the high-resolution image from its degraded low-resolution version with unknown degradation. Recent research on BISR has demonstrated impressive results by using convolutional neural network (CNN) based techniques. However, these methods suffer from limited receptive fields brought by CNN. In addition, they have limited adaptivity to different frequency components of natural images. To address those drawbacks, we propose a deep local and global dual-guidance degradation-adaptive BISR network that exploits global information by involving Fourier coefficients in the degradation representation and image reconstruction process. Additionally, we propose a novel frequency component enhancement module that explicitly decomposes images into multiple frequency bands and assign different weights for each band so as to construct high-quality image. Experimental results demonstrate the superior performance of our method.
Yajun Qiu, Shuyuan Zhu, Lantao Yu, Bing Zeng 0001
MMSP2
2025 Task-Aware Optimized Color Image Demosaicing
abstract
In this paper, we propose a task-aware deep demosaicing network that is designed to produce images for object detection and image compression, targeting high performance for both tasks. The proposed network consists of a color restoration module and a semantic enhancement module. Specifically, the color restoration module converts the Bayer-pattern raw images into full-color images. The semantic enhancement module integrates the semantic information via an adaptive feature fusion to enhance the task-relevant features while suppressing the task-irrelevant content to save bitrate. Experimental results demonstrate that using the color images produced by our demosaicing network can achieve a better trade-off between detection accuracy and compression efficiency.
Feiyu Chen 0001, Shuyuan Zhu, Bing Zeng 0001
MMSP5
2025 High Dynamic Range Imaging for Dynamic Scenes Based on Multi-Level Spike Camera
abstract
Spike camera is a retina-inspired neuromorphic camera which can capture dynamic scenes of high-speed motion by firing a continuous stream of spikes at an extremely high temporal resolution. The limitation in the current design is that each spike only represents the arrival of a fixed amount of photons. It can not deal with strong light areas in which the amount of accumulated photons reaches the pre-specified threshold multiple times within a single readout interval. In this paper, we propose a new spike camera model of high-speed imaging for high dynamic range scenarios. In this scheme, each pixel accumulates the incoming photons persistently and generates a new type of spike stream in which each spike symbol may be associated with different levels, indicating the arrival of different amounts of photons since the last readout. This enables the camera to support dynamic scenes with wider dynamic range. To achieve this, we propose a two-level buffer mechanism, one for photon accumulation and one for spike-firing encoding. We use a register to hold the number of spike-firings which has not been read out yet. At each readout time, the major part in the counter is read out via a carefully designed exponential encoding and the counter is updated. Such encoding and readout strategy enables a very efficient expansion of the dynamic range using a small number of encoding bits. Furthermore, we propose an image reconstruction scheme for the proposed camera, utilizing both spike intervals and spike levels to recover the light intensity. We incorporate Mamba and propose a temporal-spatial selective scan mechanism to extract temporal-spatial correlation within spike streams. We employ a pyramid adaptive filtering and alignment module to achieve coarse-to-fine feature alignment. Experimental results show that the proposed scheme can achieve better imaging quality and outperform the existing spike camera in high dynamic range scenarios.
Zhenkun Zhu 0001, Ruiqin Xiong, Jing Zhao 0011, Rui Zhao 0010, Xiaopeng Fan 0001, Shuyuan Zhu, Tiejun Huang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Associate Everything Detected: Facilitating Tracking-by-Detection to the Unknown
abstract
Multi-object tracking (MOT) emerges as a pivotal and highly promising branch in the field of computer vision. Classical closed-vocabulary MOT (CV-MOT) methods aim to track objects of predefined categories. Recently, some open-vocabulary MOT (OV-MOT) methods have successfully addressed the problem of tracking unknown categories. However, we found that the CV-MOT and OV-MOT methods each struggle to excel in the tasks of the other. In this paper, we present a unified framework, Associate Everything Detected (AED), that simultaneously tackles CV-MOT and OV-MOT by integrating with any off-the-shelf detector and supports unknown categories. Different from existing tracking-by-detection MOT methods, AED gets rid of prior knowledge (e.g., motion cues) and relies solely on highly robust feature learning to handle complex trajectories in OV-MOT tasks while keeping excellent performance in CV-MOT tasks. Specifically, we model the association task as a similarity decoding problem and propose a sim-decoder with an association-centric learning mechanism. The sim-decoder calculates similarities in three aspects: spatial, temporal, and cross-clip. Subsequently, association-centric learning leverages these threefold similarities to ensure that the extracted features are appropriate for continuous tracking and robust enough to generalize to unknown categories. Compared with existing powerful OV-MOT and CV-MOT methods, AED achieves superior performance on TAO, SportsMOT, and DanceTrack without any prior knowledge. Our code is available at https://github.com/balabooooo/AED.
Zimeng Fang, Shuyuan Zhu, Xi Li 0001
IEEE Trans. Image Process.4
2025 Enhancing Few-Shot 3D Point Cloud Classification With Soft Interaction and Self-Attention
abstract
Few-shot learning is a crucial aspect of modern machine learning that enables models to recognize and classify objects efficiently with limited training data. The shortage of labeled 3D point cloud data calls for innovative solutions, particularly when novel classes emerge more frequently. In this paper, we propose a novel few-shot learning method for recognizing 3D point clouds. More specifically, this paper addresses the challenges of applying few-shot learning to 3D point cloud data, which poses unique difficulties due to the unordered and irregular nature of these data. We propose two new modules for few-shot based 3D point cloud classification, i.e., the Soft Interaction Module (SIM) and Self-Attention Residual Feedforward (SARF) Module. These modules balance and enhance the feature representation by enabling more relevant feature interactions and capturing long-range dependencies between query and support features. To validate the effectiveness of the proposed method, extensive experiments are conducted on benchmark datasets, including ModelNet40, ShapeNetCore, and ScanObjectNN. Our approach demonstrates superior performance in handling abrupt feature changes occurring during the meta-learning process. The results of the experiments indicate the superiority of our proposed method by demonstrating its robust generalization ability and better classification performance for 3D point cloud data with limited training samples.
Abdullah Aman Khan, Jie Shao 0001, Sidra Shafiq, Shuyuan Zhu, Heng Tao Shen
IEEE Trans. Multim.4
2025 Projection Difference-Guided Geometry Quality Enhancement for Video-Based Point Cloud Compression
Yu Liu 0091, Jingwei Bao, Zeliang Li, Shuyuan Zhu, Jeff Siu-Kei Au-Yeung, Bing Zeng 0001
IEEE Trans. Multim.5
2025 DVSRNet: Deep Video Super-Resolution Based on Progressive Deformable Alignment and Temporal-Sparse Enhancement
abstract
Video super-resolution (VSR) is used to compose high-resolution (HR) video from low-resolution video. Recently, the deformable alignment-based VSR methods are becoming increasingly popular. In these methods, the features extracted from video are aligned to eliminate the motion error targeting high super-resolution (SR) quality. However, these methods often suffer from misalignment and the lack of enough temporal information to compose HR frames, which accordingly induce artifacts in the SR result. In this article, we design a deep VSR network (DVSRNet) based on the proposed progressive deformable alignment (PDA) module and temporal-sparse enhancement (TSE) module. Specifically, the PDA module is designed to accurately align features and to eliminate artifacts via the bidirectional information propagation. The TSE module is constructed to further eliminate artifacts and to generate clear details for the HR frame. In addition, we construct a lightweight deep optical flow network (OFNet) to obtain the bidirectional optical flows for the implementation of the PDA module. Moreover, two new loss functions are designed for our proposed method. The first one is adopted in OFNet and the second one is constructed to guarantee the generation of sharp and clear details for the HR frames. The experimental results demonstrate that our method performs better than the state-of-the-art methods.
Feiyu Chen 0001, Shuyuan Zhu, Yu Liu 0091, Ruiqin Xiong, Bing Zeng 0001
IEEE Trans. Neural Networks Learn. Syst.3
2024 Joint Demosaicing and Denoising for Spike Camera
abstract
As a neuromorphic camera with high temporal resolution, spike camera can capture dynamic scenes with high-speed motion. Recently, spike camera with a color filter array (CFA) has been developed for color imaging. There are some methods for spike camera demosaicing to reconstruct color images from Bayer-pattern spike streams. However, the demosaicing results are bothered by severe noise in spike streams, to which previous works pay less attention. In this paper, we propose an iterative joint demosaicing and denoising network (SJDD-Net) for spike cameras based on the observation model. Firstly, we design a color spike representation (CSR) to learn latent representation from Bayer-pattern spike streams. In CSR, we propose an offset-sharing deformable convolution module to align temporal features of color channels. Then we develop a spike noise estimator (SNE) to obtain features of the noise distribution. Finally, a color correlation prior (CCP) module is proposed to utilize the color correlation for better details. For training and evaluation, we designed a spike camera simulator to generate Bayer-pattern spike streams with synthesized noise. Besides, we captured some Bayer-pattern spike streams, building the first real-world captured dataset to our knowledge. Experimental results show that our method can restore clean images from Bayer-pattern spike streams. The source codes and dataset are available at https://github.com/csycdong/SJDD-Net.
Yanchen Dong 0001, Ruiqin Xiong, Jing Zhao 0011, Jian Zhang 0018, Xiaopeng Fan 0001, Shuyuan Zhu, Tiejun Huang 0001
AAAI6
2024 Super-Resolution Reconstruction from Bayer-Pattern Spike Streams
abstract
Spike camera is a neuromorphic vision sensor that can capture highly dynamic scenes by generating a continuous stream of binary spikes to represent the arrival of photons at very high temporal resolution. Equipped with Bayer color filter array (CFA), color spike camera (CSC) has been invented to capture color information. Although spike camera has already demonstrated great potential for high-speed imaging, its spatial resolution is limited compared with conventional digital cameras. This paper proposes a Color Spike Camera Super-Resolution (CSCSR) network to super-resolve higher-resolution color images from spike camera streams with Bayer CFA. To be specific, we first propose a representation for Bayer-pattern spike streams, exploring local temporal information with global perception to represent the binary data. Then we exploit the CFA layout and sub-pixel level motion to collect temporal pixels for the spatial super-resolution of each color channel. In particular, a residual-based module for feature refinement is developed to reduce the impact of motion estimation errors. Considering color correlation, we jointly utilize the multi-stage temporal-pixel features of color channels to reconstruct the high-resolution color image. Experimental results demonstrate that the proposed scheme can reconstruct satisfactory color images with both high temporal and spatial resolution from low-resolution Bayerpattern spike streams. The source codes are available at https://github.com/csycdong/CSCSR.
Yanchen Dong 0001, Ruiqin Xiong, Jian Zhang 0018, Zhaofei Yu, Xiaopeng Fan 0001, Shuyuan Zhu, Tiejun Huang 0001
CVPR6
2024 CPGA: Coding Priors-Guided Aggregation Network for Compressed Video Quality Enhancement
abstract
Recently, numerous approaches have achieved notable success in compressed video quality enhancement (VQE). However, these methods usually ignore the utilization of valuable coding priors inherently embedded in compressed videos, such as motion vectors and residual frames, which carry abundant temporal and spatial information. To remedy this problem, we propose the Coding Priors-Guided Aggregation (CPGA) network to utilize temporal and spatial information from coding priors. The CPGA mainly consists of an inter-frame temporal aggregation (ITA) module and a multi-scale non-local aggregation (MNA) module. Specifically, the ITA module aggregates temporal information from consecutive frames and coding priors, while the MNA module globally captures spatial information guided by residual frames. In addition, to facilitate research in VQE task, we newly construct the Video Coding Priors (VCP) dataset, comprising 300 videos with various coding priors extracted from corresponding bitstreams. It remedies the shortage of previous datasets on the lack of coding information. Experimental results demonstrate the superiority of our method compared to existing state-of-the-art methods. The code and dataset will be released at https://github.com/VQE-CPGA/CPGA.
Jinhua Hao, Yukang Ding, Yu Liu 0091, Qiao Mo, Ming Sun 0008, Chao Zhou 0003, Shuyuan Zhu
CVPR8
2024 OAPT: Offset-Aware Partition Transformer for Double JPEG Artifacts Removal
Qiao Mo, Yukang Ding, Jinhua Hao, Ming Sun 0008, Chao Zhou 0003, Feiyu Chen 0001, Shuyuan Zhu
ECCV (22)8
2024 Filamentary Convolution for Spoken Language Identification: A Brain-Inspired Approach
abstract
Spoken language identification (SLI) by human beings relies on the hierarchical understanding of one or a few words within the voice signal, encapsulated within the corresponding time windows. Concurrently, frequency-domain features play a crucial role in enhancing identification. The short-time Fourier transform (STFT) has conventionally served as a pivotal component in the forefront of most SLI systems, including deep-learning networks (DLNs). Nevertheless, the use of rectangle-shaped masks in STFT introduces spectral component mixing across different time windows, potentially resulting in an aliasing effect. To address this limitation, we propose a novel filamentary convolution framework to replace the conventional rectangle-shaped convolutions. This framework not only reduces complexity but also enhances feature learning within each frame. Leveraging filamentary convolution, we formulate an encoding module with a non-overlapping strategy and a multi-level information extraction (MIE) module featuring unbalanced dual-route convolution (UDRC) blocks. The frequency features learned from filamentary convolutions are seamlessly integrated through a long-short term memory (LSTM) structure. In summary, our decision-making process employs the filamentary convolution kernel-based hierarchical neural network (FCK-NN), comprising an encoding module, MIE module, and LSTM module. We conduct experiments on a novel dataset encompassing 44 languages, curated by ourselves, and the results validate that our FCK-NN yields a significant improvement in performance.
Shuyuan Zhu, Tong Xie, Xibang Yang, Bing Zeng 0001
ICASSP2
2024 Region Motion-based Adaptive Composite Long-Term Reference Coding for VVC
abstract
The adoption of composite long-term reference (CLTR) in versatile video coding (VVC) has demonstrated good performance, especially for the coding of video containing a large amount of stationary background areas. However, if the long-term reference (LTR) picture is constructed by using the frames that contain lots of foreground contents, the composed LTR picture cannot offer enough background information for the coding of target frame. To effectively apply CLTR to video coding, especially to VVC, we propose a region motion-based determination method to adaptively choose long-term and short-term reference pictures for inter-picture coding. The experimental results demonstrate that our proposed method can achieve promising performance improvement compared with the results generated from VVC common test condition and the traditional LTR-based coding scheme.
Xiaozhen Zheng, Yu Liu 0091, Jianglin Wang, Zihao Ren 0002, Shuyuan Zhu, Qingmin Liao
ISCAS5
2024 Color Enhancement for V-PCC Compressed Point Cloud via 2D Attribute Map Optimization
abstract
Video-based point cloud compression (V-PCC) converts the dynamic point cloud data into video sequences using traditional video codecs for efficient encoding. However, this lossy compression scheme introduces artifacts that degrade the color attributes of the data. This paper introduces a framework designed to enhance the color quality in the V-PCC compressed point clouds. We propose the lightweight de-compression Unet (LDC-Unet), a 2D neural network, to optimize the projection maps generated during V-PCC encoding. The optimized 2D maps will then be back-projected to the 3D space to enhance the corresponding point cloud attributes. Additionally, we introduce a transfer learning strategy and develop a customized natural image dataset for the initial training. The model was then fine-tuned using the projection maps of the compressed point clouds. The whole strategy effectively addresses the scarcity of point cloud training data. Our experiments, conducted on the public 8i voxelized full bodies long sequences (8iVSLF) dataset, demonstrate the effectiveness of our proposed method in improving the color quality.
Jingwei Bao, Yu Liu 0091, Zeliang Li, Shuyuan Zhu, Jeff Siu-Kei Au-Yeung
VCIP4
2024 Compressed Video Quality Enhancement With Temporal Group Alignment and Fusion
abstract
In this paper, we propose a temporal group alignment and fusion network to enhance the quality of compressed videos by using the long-short term correlations between frames. The proposed model consists of the intra-group feature alignment (IntraGFA) module, the inter-group feature fusion (InterGFF) module, and the feature enhancement (FE) module. We form the group of pictures (GoP) by selecting frames from the video according to their temporal distances to the target enhanced frame. With this grouping, the composed GoP can contain either long- or short-term correlated information of neighboring frames. We design the IntraGFA module to align the features of frames of each GoP to eliminate the motion existing between frames. We construct the InterGFF module to fuse features belonging to different GoPs and finally enhance the fused features with the FE module to generate high-quality video frames. The experimental results show that our proposed method achieves up to 0.05 dB gain and lower complexity compared to the state-of-the-art method.
Yajun Qiu, Yu Liu 0091, Shuyuan Zhu, Bing Zeng 0001
IEEE Signal Process. Lett.4
2024 Dual Circle Contrastive Learning-Based Blind Image Super-Resolution
abstract
Blind image super-resolution (BISR) aims to construct high-resolution image from low-resolution (LR) image that contains unknown degradation. Although the previous methods demonstrated impressive performance by introducing the degradation representation in BISR task, there still exist two problems in most of them. First, they ignore the degradation characteristics of different image regions when generating degradation representation. Second, they lack effective supervision on the generation of both degradation representation and super-resolution result. To solve these problems, we propose the dual circle contrastive learning (DCCL) with the high-efficiency modules to implement BISR. In our proposed method, we design the degradation extraction network to obtain the degradation representations from different texture regions of LR image. Meanwhile, we propose DCCL coupled with the degrading network to guarantee the obtained degradation representation to contain the degradation of LR image as much as possible. The application of DCCL also makes the SR results contain degradation as little as possible. Additionally, we develop an information distillation module for our proposed BISR model to guarantee the SR images with high quality. The experimental results demonstrate that our proposed method achieves the state-of-the-art BISR performance.
Yajun Qiu, Shuyuan Zhu, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 Spike Camera Image Reconstruction Using Deep Spiking Neural Networks
abstract
Spike camera is a bio-inspired sensor with ultra-high temporal resolution and low energy consumption. It captures visual signals using an “integrate-and-fire" mechanism and outputs a continuous stream of binary spikes. Reconstructing image sequence from spikes streams is critical for spike camera. Several reconstruction methods have been proposed in recent years. However, the computational cost of these methods is relatively high. Inspired by the fact that spiking neural networks (SNNs) are energy efficient and support time-series signal processing inherently, we propose a lightweight SNN for spike camera image reconstruction (abbreviated to SSIR). Experimental results show that SSIR achieves comparable performance with the state-of-the-art (SOTA) methods at much lower computation and energy cost.
Rui Zhao 0010, Ruiqin Xiong, Jian Zhang 0018, Zhaofei Yu, Shuyuan Zhu, Lei Ma 0008, Tiejun Huang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 Learning a Deep Demosaicing Network for Spike Camera With Color Filter Array
abstract
For capturing dynamic scenes with ultra-fast motion, neuromorphic cameras with extremely high temporal resolution have demonstrated their great capability and potential. Different from the event cameras that only record relative changes in light intensity, spike camera fires a stream of spikes according to a full-time accumulation of photons so that it can recover the texture details for both static areas and dynamic areas. Recently, color spike camera has been invented to record color information of dynamic scenes using a color filter array (CFA). However, demosaicing for color spike cameras is an open and challenging problem. In this paper, we develop a demosaicing network, called CSpkNet, to reconstruct dynamic color visual signals from the spike stream captured by the color spike camera. Firstly, we develop a light inference module to convert binary spike streams to intensity estimates. In particular, a feature-based channel attention module is proposed to reduce the noises caused by quantization errors. Secondly, considering both the Bayer configuration and object motion, we propose a motion-guided filtering module to estimate the missing pixels of each color channel, without undesired motion blur. Finally, we design a refinement module to improve the intensity and details, utilizing the color correlation. Experimental results demonstrate that CSpkNet can reconstruct color images from the Bayer-pattern spike stream with promising visual quality.
Yanchen Dong 0001, Ruiqin Xiong, Jing Zhao 0011, Jian Zhang 0018, Xiaopeng Fan 0001, Shuyuan Zhu, Tiejun Huang 0001
IEEE Trans. Image Process.6
2024 Depth-Guided Deep Video Inpainting
abstract
Video inpainting aims to fill in missing regions of a video after any undesired contents are removed from it. This technique can be applied to repair the broken video or edit the video content. In this paper, we propose a depth-guided deep video inpainting network (DGDVI) and demonstrate its effectiveness in processing challenging broken areas crossing multiple depth layers. To achieve our goal, we divide the inpainting into depth completion, content reconstruction, and content enhancement. Three corresponding modules are designed to implement a process-flow. Firstly, we develop a depth completion module based upon the spatio-temporal Transformer which is used to obtain the completed depth information for each video frame. Secondly, we design a content reconstruction module to generate initially inpainted video. With this module, the contents of the missing regions are composed via the depth-guided feature propagation. Thirdly, we construct a content enhancement module to enhance the temporal coherence and texture quality for the inpainted video. All of proposed modules are jointly optimized to guarantee the high inpainting efficiency. The experimental results demonstrate that our proposed method provides better inpainting results, both qualitatively and quantitatively, compared with the previous state-of-the-art.
Shuyuan Zhu, Yao Ge 0002, Bing Zeng 0001, Muhammad Ali Imran 0001, Qammer H. Abbasi, Jonathan M. Cooper
IEEE Trans. Multim.2
2023 Multivariate, Multi-Frequency and Multimodal: Rethinking Graph Neural Networks for Emotion Recognition in Conversation
abstract
Complex relationships of high arity across modality and context dimensions is a critical challenge in the Emotion Recognition in Conversation (ERC) task. Yet, previous works tend to encode multimodal and contextual relationships in a loosely-coupled manner, which may harm relationship modelling. Recently, Graph Neural Networks (GNN) which show advantages in capturing data relations, offer a new solution for ERC. However, existing GNN-based ERC models fail to address some general limits of GNNs, including assuming pairwise formulation and erasing high-frequency signals, which may be trivial for many applications but crucial for the ERC task. In this paper, we propose a GNN-based model that explores multivariate relationships and captures the varying importance of emotion discrepancy and commonality by valuing multi-frequency signals. We empower GNNs to better capture the inherent relationships among utterances and deliver more sufficient multimodal and contextual modelling. Experimental results show that our proposed method outperforms previous state-of-the-art works on two popular multimodal ERC datasets.
Feiyu Chen 0001, Jie Shao 0001, Shuyuan Zhu, Heng Tao Shen
CVPR3
2023 Learned Image Compression with Large Capacity and Low Redundancy of Latent Representation
abstract
Learned image compression has attracted a lot of attention in recent years. Currently, popular learned image compression methods usually exploit hyperprior and autoregressive models to facilitate probability estimation and reduce the redundancy of latent representation. These models ignore different image contents, and it is difficult to eliminate the spatial redundancy of image, resulting in the performance saturation. In this work, we propose a learned image compression method with large capacity and low redundancy of latent representation. We design two enhancement modules, i.e., the network capacity expansion module (NCEM) and the high-entropy content guided reconstruction module (HCGR), to construct network architectures with better rate-distortion performance than the existing hyperprior and autoregressive models. Experimental results show that our method can produce superior results compared to the state-of-the-art methods.
Xiandong Meng, Shuyuan Zhu, Siwei Ma 0001, Bing Zeng 0001
ICIP2
2023 Optimization-Inspired Deep Network for Image Restoration from Partial Random Samples
abstract
Image Restoration from Partial Random Samples (RRS) has been studied in many image restoration works. There are also some attempts to use convolutional neural networks (CNNs) to handle it. However, most existing neural network-based methods perform poorly in generalization and we need to train a specific model for each degradation situation. Besides, the sampling mask which represents the positions of the sampled pixels is not used effectively in these methods. To address the problems, we propose an optimization-inspired network called RRSNet based on our derivation of the iterative optimization formulas for RRS. In our method, we design a CNN with two encoders and one decoder for training, setting up a flexible and effective prior. To make the most of the sampling information, we concatenate the degraded image with the mask and input them into one encoder for better generalization. Then we split the pixels into two groups according to the mask and extract their features as the input of another encoder. Experiments demonstrate that our RRSNet with the mask input can handle various sampling ratios using only one trained model and achieve the best restoration performance among all comparison methods.
Yanchen Dong 0001, Rui Zhao 0010, Ruiqin Xiong, Shuyuan Zhu, Xiaopeng Fan 0001, Tiejun Huang 0001
ISCAS4
2023 WiFi sensing of Human Activity Recognition using Continuous AoA-ToF Maps
abstract
Joint communication and sensing technique has been adopted for smart home design and other applications recently. WiFi sensing, which utilizes mutually orthogonal channel response to monitor the changes in the medium, is regarded as one of key techniques in this field. Human activity recognition using wireless communication systems is a key function of future internet of things systems. The effective and inexpensive WiFi sensing system can help people with device-free controlling, and healthcare monitoring without concern of image information leakage that uses a camera system. In this article, we proposed a continuous angle of arrival and time of flight (AoA-ToF) maps based method that adopts multiple signals classification analysis on commercial and off-the-shelf WiFi devices to detect human activities. Our experimental results ensure the effectiveness of the proposed system for the human activity recognition (HAR) task with 8 activities among 5 users in three directions. The performance of our system achieves 85.6% accuracy on average. Meanwhile, we evaluate the performance of our system under different conditions, including direction and user identity. The results show the system’s robustness for human activity recognition under such conditions.
Yao Ge 0002, Liyuan Qi, Shuyuan Zhu, Jonathan M. Cooper, Muhammad Ali Imran 0001, Qammer H. Abbasi
WCNC5
2023 Quality-Constrained Encoding Optimization for Omnidirectional Video Streaming
abstract
Omnidirectional video streaming is usually implemented based on the representations of tiles, where the tiles are obtained by splitting the video frame into several rectangular areas and each tile is converted into multiple representations with different resolutions and encoded at different bitrates. One key issue in omnidirectional video streaming is how to choose the optimal representations for each tile at the server to save the overall transmission bitrate to all users while offering them satisfactory quality. This is different from the adaptive bitrate-based method that optimizes the downloading procedure of individual users, where the given video representations are stored on the server. In this work, we focus on optimization for the encoding of omnidirectional video streaming by using the optimal combination of tile representations. To achieve our goal, we formulate the selection of the representations into an optimization problem in which the transmission bitrate of all the representations is minimized with a quality constraint. By using this constraint, we can improve the average quality of omnidirectional videos for users. More specifically, we first construct the tile-level rate-distortion (R-D) model and determine the available tile bandwidth based on the previous viewers’ statistics. Then, we formulate the representation selection problem based on the obtained R-D model and tile bandwidth. Finally, we solve this problem to obtain the optimal combination of tile representations so that we can transmit the omnidirectional video to users with satisfactory quality but low bitrate. The experimental results demonstrate the effectiveness of our proposed approach when it is applied to omnidirectional video streaming.
Chaofan He, Roberto Gerson De Albuquerque Azevedo, Shuyuan Zhu, Bing Zeng 0001, Pascal Frossard
IEEE Trans. Circuits Syst. Video Technol.4
2023 NOMA-Based Uncoded Video Transmission With Optimization of Joint Resource Allocation
abstract
The non-orthogonal multiple access (NOMA) technique has demonstrated potential for the multicast of multiple videos. However, it simply multiplexes limited number of signals in each single channel and cannot satisfy different video quality requirements of multiple users. To resolve this problem, we construct the NOMA-based uncoded multi-user video transmission (NOMA-UMVT) system in which the allocation of power and channel resources to all the users is jointly optimized to guarantee high video quality. Specifically, we first implement the power allocation of multiuser by converting it into the inter-channel and intra-channel allocation sub-problems. After solving these sub-problems for power allocation, we assign channels with the proposed two-staged channel assignment algorithm. The simulation results demonstrate the superior performance of our proposed NOMA-UMVT system when it is applied to transmit videos to multiple users.
Chaofan He, Shuyuan Zhu, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2023 DFCE: Decoder-Friendly Chrominance Enhancement for HEVC Intra Coding
abstract
We propose a decoder-friendly chrominance enhancement method for the compressed images. Our proposed method is developed based on the luminance-guided chrominance enhancement network (LGCEN) and online learning. With LGCEN, the textures of the compressed chrominance components are enhanced by the guidance of luminance component. Moreover, LGCEN is constructed with the recursive design and the light-weight channel attention mechanism to achieve high performance as well as low complexity. It is arranged at both encoder and decoder sides. Given the input image, we train LGCEN at encoder side by using online learning. With online learning, we partially update network parameters and transmit them to decoder to update LGCEN arranged there. The adoption of online learning effectively reduces the workload of decoder and guarantee high robustness. Compared with the state-of-the-art methods, our proposed approach achieves superior performance.
Renwei Yang, Hewei Liu, Shuyuan Zhu, Xiaozhen Zheng, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Local and Global Fusion Network For Learned Image Compression
abstract
Convolution autoencoder is a widely utilized framework for learned image compression. However, it is facing performance bottleneck because the convolution with limited receptive field mainly extracts image local information. In this paper, we propose a novel neural network based image compression framework by removing both local and global redundancy. Herein, convolution autoencoder and Generative flow (Glow) are utilized to extract image local and global information respectively. Glow is a lossless invertible neural network and can facilitate global information extraction. Furthermore, we design a DenseNet module to fuse the local and global information extracted from convolution autoencoder and Glow. Extensive experimental results show that the proposed framework outperform the intra-frame coding of Versatile Video Coding (VVC) and state-of-the-art neural network based image compression methods.
Gai Zhang, Xinfeng Zhang 0001, Shuyuan Zhu
ICIP3
2022 Hierarchical Coding for Talking-Head Video
abstract
Talking-head video is very popular in video conference and social media, where the camera captures the movement of user’s head and the change of facial expression. In this paper, we propose a hierarchical coding scheme for the compression of talking-head video. In our proposed method, three data layers, including one base layer, one enhancement layer and one feature layer, are formed as the input of encoder. More specifically, the base layer is generated by spatially sub-sampling the source video. The enhancement layer is composed by the specific key frames and the feature layer is produced based on the extracted facial landmarks. These layers are separately compressed but fused together to reconstruct the video signal in the decoder side. To achieve a high-quality reconstruction, we design the multi-feature fusion network in which the feature layer is used to guide the fusion of base layer and enhancement layer. The experiment results demonstrate the good performance of our proposed method for the coding of talking-head video.
Yu Liu 0091, Shuyuan Zhu, Jeff Siu-Kei Au-Yeung, Bing Zeng 0001
ISCAS3
2022 Luminance-Guided Chrominance Image Enhancement for HEVC Intra Coding
abstract
In this paper, we propose a luminance-guided chrominance image enhancement convolutional neural network for HEVC intra coding. Specifically, we firstly develop a gated recursive asymmetric-convolution block to restore each degraded chrominance image, which generates an intermediate output. Then, guided by the luminance image, the quality of this intermediate output is further improved, which finally produces the high-quality chrominance image. When our proposed method is adopted in the compression of color images with HEVC intra coding, it achieves 28.96% and 16.74% BD-rate gains over HEVC for the U and V images, respectively, which accordingly demonstrate its superiority. The code is available online: https://github.com/Nickyang4900/Luminance-Guided-Chrominance-Enhancement-for-HEVC-Intra-Coding.
Hewei Liu, Renwei Yang, Shuyuan Zhu, Bing Zeng 0001
ISCAS3
2022 Deep Feature Compression with Collaborative Coding of Image Texture
abstract
In this paper, we propose a coding scheme for the deep intermediate feature and it is implemented with the collaborative compression of image texture. More specifically, we separately compress the feature and texture of the image to form two data layers. The first one is the intermediate feature layer and the second one is the texture layer. The texture layer can provide an image for users and the feature layer can be used to implement the computer vision (CV) task. With our proposed deep reconstruction network (RecNet), the texture and features cooperate to achieve a high-quality visual output as well as a high-efficiency CV task. The experimental results demonstrate the excellent performance by using our proposed method to compress the deep features.
Hewei Liu, Shuyuan Zhu, Xiaozhen Zheng, Ruiqin Xiong, Bing Zeng 0001
ISCAS3
2022 Deep Video Super-Resolution with Flow-Guided Deformable Alignment and Sparsity-based Temporal-Spatial Enhancement
abstract
Video super resolution (VSR) aims to construct the high resolution (HR) frames from the low resolution ones. In this paper, we propose a new deep VSR network based on the flow-guided deformable alignment (FGDA) module and the sparsity-based temporal-spatial enhancement (STSE) module. More specifically, the FGDA module is designed to generate temporally-aligned features with the bidirectional propagation. Meanwhile, the STSE module is constructed to eliminate the alignment error for the features and strengthen the sparsity of them to enhance the spatial details to construct the high-quality HR result. In addition, we design a sparsity-based loss function to guarantee the germinated HR frames with sharp details. The experimental results demonstrate that our proposed method achieves superior performance compared with the existing popular methods.
Shuyuan Zhu, Guanghui Liu 0001, Bing Zeng 0001, Xiaozhen Zheng
MMSP3
2022 Infrared Small Target Detection Using Local Feature-Based Density Peaks Searching
abstract
In this letter, we propose a new local feature-based method for the detection of infrared small targets with the density feature map that can effectively suppress noise in the feature domain. First, we combine the local tetra pattern (LTrP) and the second-order LTrP to generate the density feature map. Second, we apply density peaks searching to the feature map to obtain candidate targets. Third, we generate two local features, i.e., the entropy and third-order moment, for each image patch whose center is a candidate target and then fuse them to employ the fused feature to find the real target. The experimental results demonstrate that our proposed method achieves better performance compared with the state-of-the-art approaches.
Shuyuan Zhu, Guanghui Liu 0001, Zhenming Peng
IEEE Geosci. Remote. Sens. Lett.2
2022 Inter-Frame Dependency-Based Rate Control for VVC Low-Delay Coding
abstract
In this letter, we propose two solutions for the rate control of the VVC low-delay coding. Both solutions are developed by determining the bit allocation factors for video frames based on their dependency. Specifically, we design the first solution according to the distortion correlation between the key-frame and its subsequent frames. With this solution, the bit allocation factors are determined by applying multi-pass coding on the video to build up the cross-frame distortion model. This model offers us a straightforward way to achieve better rate control performance but results in a rather high complexity. To solve the complexity problem, we propose the second solution based on the difference between frames. In this solution, we construct the bit allocation model and apply it to frames so that we can adaptively determine the allocation factors with a low complexity. The experimental results demonstrate that our proposed two solutions can offer better rate-distortion performances than the state-of-the-art method.
Hewei Liu, Shuyuan Zhu, Bing Zeng 0001
IEEE Signal Process. Lett.2
2022 DeepOIS: Gyroscope-Guided Deep Optical Image Stabilizer Compensation
abstract
Mobile captured images can be aligned using their gyroscope sensors. Optical image stabilizer (OIS) terminates this possibility by adjusting the images during the capturing. In this work, we propose a deep network that compensates for the motions caused by the OIS, such that the gyroscopes can be used for image alignment on the OIS cameras. To achieve this, we first record both videos and gyroscope readings with an OIS camera as training data. Then, we convert gyroscope readings into motion fields. Second, we propose an Essential Mixtures motion model for rolling shutter cameras, where an array of rotations within a frame are extracted as the ground-truth guidance. Third, we train a convolutional neural network with gyroscope motions as input to compensate for the OIS motion. Once finished, the compensation network can be applied for other scenes, where the image alignment is purely based on gyroscopes with no need for images contents, delivering strong robustness. Experiments show that our results are comparable with that of non-OIS cameras, and outperform image-based alignment results with a relatively large margin. Code and dataset is available at:https://github.com/lhaippp/DeepOIS.
Shuaicheng Liu, Haipeng Li 0001, Zhengning Wang, Jue Wang 0001, Shuyuan Zhu, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 Quadratic Terms Based Point-to-Surface 3D Representation for Deep Learning of Point Cloud
abstract
In this paper, we introduce a novel point-to-surface representation for 3D point cloud learning. Unlike the previous methods that mainly adopt voxel, mesh, or point coordinates, we propose to tackle this problem from a new perspective: learn a set of quadratic terms based static and global reference surfaces to describe 3D shapes, such that the coordinates of a 3D point (x, y, z) can be extended to quadratic terms (xy, xz, yz,$\ldots $) and transformed to the relationship between the local point and the global reference surfaces. Then, the static surfaces are changed into dynamic surfaces by adaptive contribution weighting to improve the descriptive capability. Towards this end, we propose our point-to-surface representation, a new representation for 3D point cloud learning that has not been attempted before, which can assemble local and global geometric information effectively by building connections between the point cloud and the learned reference surfaces. Given 3D points, we show how the reference surfaces are constructed, and how they are inserted into the 3D learning pipeline for different tasks. The experimental results confirm the effectiveness of our new representation, which has outperformed the state-of-the-art methods on the tasks of 3D classification and segmentation.
Tiecheng Sun, Guanghui Liu 0001, Ru Li 0002, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 Rethinking the Competition Between Detection and ReID in Multiobject Tracking
abstract
Due to balanced accuracy and speed, one-shot models which jointly learn detection and identification embeddings, have drawn great attention in multi-object tracking (MOT). However, the inherent differences and relations between detection and re-identification (ReID) are unconsciously overlooked because of treating them as two isolated tasks in the one-shot tracking paradigm. This leads to inferior performance compared with existing two-stage methods. In this paper, we first dissect the reasoning process for these two tasks, which reveals that the competition between them inevitably would destroy task-dependent representations learning. To tackle this problem, we propose a novel reciprocal network (REN) with a self-relation and cross-relation design so that to impel each branch to better learn task-dependent representations. The proposed model aims to alleviate the deleterious tasks competition, meanwhile improve the cooperation between detection and ReID. Furthermore, we introduce a scale-aware attention network (SAAN) that prevents semantic level misalignment to improve the association capability of ID embeddings. By integrating the two delicately designed networks into a one-shot online MOT system, we construct a strong MOT tracker, namely CSTrack. Our tracker achieves the state-of-the-art performance on MOT16, MOT17 and MOT20 datasets, without other bells and whistles. Moreover, CSTrack is efficient and runs at 16.4 FPS on a single modern GPU, and its lightweight version even runs at 34.6 FPS. The complete code has been released at https://github.com/JudasDie/SOTS.
Bing Li 0001, Shuyuan Zhu, Weiming Hu 0004
IEEE Trans. Image Process.5
2021 Recover The Residual Of Residual: Recurrent Residual Refinement Network For Image Super-Resolution
abstract
Benefiting from learning the residual between low resolution (LR) image and high resolution (HR) image, image super-resolution (SR) networks demonstrate superior reconstruction performance in recent studies. However, for the images with rich texture information, the residuals are complex and difficult for networks to learn. To address this problem, we propose a recurrent residual refinement network (RRRN) to gradually refine the residual with a recurrent structure. Instead of directly reconstructing the residual between LR image and HR image, each sub-network in our framework reconstructs the residual between SR image from previous stage and HR image, i.e. recovers the residual of residual (RoR). Considering the domain gap between the image feature and the RoR feature, we introduce a residual projection block to explicitly transform the feature from image domain to RoR domain. The RoR feature is further optimized in an iterative up- and down-sampling manner with a residual learning block. We construct the structure of each block based on the optimization methods of conventional SR and improve our network with dense connections. Experimental results prove that our method improves the quality of super-resolution images on different datasets with variable scenes.
Tianxiao Gao, Ruiqin Xiong, Rui Zhao 0010, Jian Zhang 0018, Shuyuan Zhu, Tiejun Huang 0001
ICIP5
2021 Cross-Block Difference Guided Fast CU Partition for VVC Intra Coding
abstract
In this paper, we propose a new fast CU partition method for VVC intra coding based on the cross-block difference. This difference is measured by the gradient and the content of sub-blocks obtained from partition and is employed to guide the skipping of unnecessary horizontal and vertical partition modes. With this guidance, a fast determination of block partitions is accordingly achieved. Compared with VVC, our proposed method can save 41.64% (on average) encoding time with only 0.97% (on average) increase of BD-rate.
Hewei Liu, Shuyuan Zhu, Ruiqin Xiong, Guanghui Liu 0001, Bing Zeng 0001
VCIP2
2021 A Robust Quality Enhancement Method Based on Joint Spatial-Temporal Priors for Video Coding
abstract
Quality enhancement of HEVC compressed videos has attracted a lot of attentions in recent years. In this article, we propose a robust multi-frame guided attention network (MGANet) to reconstruct high-quality frames based on HEVC compressed videos. In our network, we first use an advanced motion flow algorithm to estimate the motion information of input frames so as to guide the warping of adjacent frames. After performing the alignment, we find that large residuals still appear in the edge area of moving objects of the warped frames. Then, we design a temporal encoder based on a bi-directional convolutional long short term memory (ConvLSTM) with residual structure to further discover the variations between the current frame and its adjacent warped frames. Finally, we feed the extracted temporal information and a partitioned average image (PAI) to a multi-scale guided encoder-decoder subnet to reconstruct high-quality frames. Here, each PAI is generated according to the transform unit (TU) partitioning map that can be extracted directly from the coded bit-streams, thus enabling our network to focus on the TU boundaries while optimizing the global content. We present extensive experimental results to demonstrate the robustness of our method, especially for the high bit-rate coding case and large motion scenes. Due to the lightweight design structure, our proposed MGANet also has a very competitive inference time.
Xiandong Meng, Shuyuan Zhu, Xinfeng Zhang 0001, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 NTSDCN: New Three-Stage Deep Convolutional Image Demosaicking Network
abstract
In this letter, we compose a new three-stage deep convolutional neural network (NTSDCN) for image demosaicking, and it consists of our proposed Laplacian energy-constrained local residual unit (LC-LRU) and a feature-guided prior fusion unit (FG-PFU). Specifically, the LC-LRU is used to refine the learning target of the specific residual blocks in the network and enhance the dominant information of the residual features. The FG-PFU is designed to guide the feature extraction of the red (R) and blue (B) channels by utilizing prior information from the reconstructed green (G) channel. In our proposed NTSDCN, we recover the G channel image in the first stage with the CFA image and reconstruct the R and B images in the second stage. Finally, we fine-tune the resulting R, G and B images in the third stage to compose a full-color RGB image. The experimental results show that our proposed method achieves better performance than the state-of-the-art methods. The code is available at https://github.com/wyannn/NTSDCN.
Shiying Yin, Shuyuan Zhu, Zhan Ma 0001, Ruiqin Xiong, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Flow-Guided Temporal-Spatial Network for HEVC Compressed Video Quality Enhancement
abstract
In this paper, a flow-guided temporal-spatial network (FGTSN) is proposed to enhance the quality of HEVC compressed video. Specifically, we first employ a robust motion estimation subnet via trainable optical flow module to estimate the motion flow between the target frame and its adjacent frames, and these adjacent frames are pre-warped guided by the predicted motion flow. Then, a temporal encoder is proposed to fuse the related information between the target frame and its pre-warped frames. Finally, a quality enhancement subnet with multi-scale encoder-decoder structure is designed to generate high quality frame by training the network in a multi-supervised fashion. Experimental results show the superior performance of our proposed FGTSN method for the reconstruction quality of HEVC compressed frames, much better than the state-of-the-art quality enhancement methods. In addition, our FGTSN method can also effectively mitigate the quality fluctuation of adjacent frames.
Xiandong Meng, Shuyuan Zhu, Shuaicheng Liu, Bing Zeng 0001
DCC3
2020 Salient Object Detection Based On Image Bit-Map
abstract
In this paper, we propose a novel salient object detection framework, which makes full use of the essential image compression. More specifically, we first compose an intuitive measure of compressibility from JPEG compression, namely bit-map. Then, depending on the relationship between bitmap and salient object, we generate the salient object window directly from bit-map without utilizing any features from the compressed image. Finally, the saliency map is calculated according to the salient object window and with a ranking algorithm. The proposed method achieves good performance as well as low complexity. The experimental results demonstrate the effectiveness of our proposed method compared with other existing approaches.
Bangqi Cao, Xiandong Meng, Shuyuan Zhu, Bing Zeng 0001
ICASSP3
2020 Optical Flow Estimation Between Images of Different Resolutions via Variational Method
abstract
Traditional optical flow estimation methods mostly focus on images of the same resolution. However, there are some situations requiring optical flow between images of different resolutions, where the traditional approaches suffer from the inequality of spectrum aliasing level. In this paper, we propose a method estimating the flow fields between a clear image and a highly undersampled one. The proposed method simultaneously describes the motion and integral relationship between the images via an integral form image under the assumption of brightness and gradient consistency as well as motion smoothness. We also derive the numerical solution briefly, through which we can solve the equations easily via linearizations. Experimental results on Middlebury and MPI-Sintel datasets demonstrate that our proposed method outperforms traditional methods preprocessing images of different resolutions to be the same size, offering more accurate results.
Rui Zhao 0010, Ruiqin Xiong, Shuyuan Zhu, Bing Zeng 0001, Tiejun Huang 0001, Wen Gao 0001
VCIP3
2020 A New Polyphase Down-Sampling-Based Multiple Description Image Coding
abstract
Multiple description coding (MDC) is an efficient source coding technique for error-prone transmission over multiple channels. In this paper, we focus on the design of a new polyphase down-sampling based MDC (NPDS-MDC) for image signals. The encoding of our proposed NPDS-MDC consists of three steps. First, we perform down-sampling on each N×N image block according to the quincunx down-sampling pattern. Second, we propose a new transform and apply it to the down-sampled pixels to produce the side descriptions. Third, we develop an error compensation algorithm to reduce the compression distortion occurring on the down-sampled pixels. In our scheme, the side decoding is performed posterior to image interpolation with reference to the down-sampled compressed pixels. Moreover, the central decoding is achieved by interlacing the side descriptions. We also propose a compression-constrained central deblocking algorithm to further improve the efficiency of the central decoding. The experimental results indicate that our proposed MDC scheme offers clearly superior performance, especially at high bit rates, as compared to the state-of-the-art methods for various types of images.
Shuyuan Zhu, Zhiying He, Xiandong Meng, Jiantao Zhou 0001, Yuanfang Guo, Bing Zeng 0001
IEEE Trans. Image Process.1
2019 Enhancing Quality for VVC Compressed Videos by Jointly Exploiting Spatial Details and Temporal Structure
abstract
In this paper, we propose a quality enhancement network of versatile video coding (VVC) compressed videos by jointly exploiting spatial details and temporal structure (SDTS). The proposed network consists of a temporal structure fusion subnet and a spatial detail enhancement subnet. The former subnet is used to estimate and compensate the temporal motion across frames, and the latter subnet is used to reduce the compression artifacts and enhance the reconstruction quality of compressed video. Experimental results demonstrate the effectiveness of our SDTS-based method. The code of our proposed method is available at https://github.com/mengab/SDTS.
Xiandong Meng, Shuyuan Zhu, Bing Zeng 0001
ICIP3
2019 Multiple Description Image Coding Based on Compression-Guided Optimization
abstract
In this paper, we design a new multiple description coding scheme for image signals based on our proposed compression-guided optimization. Firstly, we propose a compression-constrained adaptive filtering method to produce two descriptions for the source image, where the proposed filtering algorithm works not only to guarantee a high-quality side decoding but also make a high-efficient central decoding. Secondly, we design a compression-dependent deblocking algorithm based on the transform coefficients which are decoded from both descriptions to improve the performance for the cental decoding. Experimental results demonstrate that our proposed method achieves impressive performance gains when it is applied to image signals.
Shuyuan Zhu, Zhiying He, Xiandong Meng, Guanghui Liu 0001, Bing Zeng 0001
PCS1
2019 Color Image Compression with Transform Domain Down-Sampling and Deep Convolutional Reconstruction
abstract
In this paper, we build up a new block-based color image compression scheme based on our proposed transform domain down-sampling method and deep convolutional reconstruction algorithm. Specifically, our proposed down-sampling scheme aims to down-sample each N × N transform block into the N/2 × N/2 block for the saving of bit-cost. On the other hand, the proposed deep convolutional reconstruction algorithm is employed to reconstruct the down-sampled block for a full- resolution reconstruction. We apply our proposed methods to both the chrominance components to compress color images. Experimental results show that our proposed method achieves excellent results when used in practice.
Shuyuan Zhu, Xiandong Meng, Bing Zeng 0001, Yuanfang Guo, Ruiqin Xiong
VCIP2
2019 Efficient Chroma Sub-Sampling and Luma Modification for Color Image Compression
abstract
In color image compression, the chroma components are often sub-sampled before compression and up-sampled after compression. Although sub-sampling the chroma components saves the bit-cost for compression, it often induces extra color distortions in the compressed images. In this paper, we propose two approaches to tackle this problem. First, we propose a sub-sampling method in the transform domain and apply it to both chroma components. Then, based on this sub-sampling, we propose a novel method to modify the luma component. In our proposed luma modification algorithm, the distortions that occurred in the two chroma components can be coupled together and utilized to modify the luma component. With our proposed chroma sub-sampling and luma modification algorithms, we can achieve a low RGB distortion in practical image coding. The experimental results demonstrate that our proposed methods offer more significant coding gains compared with the state-of-the-art methods for the compression of color images.
Shuyuan Zhu, Chang Cui, Ruiqin Xiong, Yuanfang Guo, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 High-Quality Color Image Compression by Quantization Crossing Color Spaces
abstract
Coding of a color image usually happens in the YCbCr space so that the rate-distortion optimization is conducted in this space. Due to the use of a non-unitary matrix in the RGB-to-YCbCr conversion, an optimal coding performance achieved in the YCbCr space does not guarantee an optimal quality in the RGB space, which would impact most display devices that need RGB signals as the inputs. In this paper, we first study the relationship between the coding distortions of the compressed RGB signals and the quantization errors occurred in the coded YCbCr signals. Then, we design a new quantization scheme crossing the RGB and YCbCr spaces to achieve a high-quality color image compression with the YCbCr 4:4:4 format. Although our proposed quantization takes place in the YCbCr space, it aims at reducing the coding distortion in the RGB space as much as possible. Experimental results demonstrate that our proposed method offers a significant quality gain over the existing block-based coding methods for various images.
Shuyuan Zhu, Zhiying He, Chen Chen 0015, Shuaicheng Liu, Jiantao Zhou 0001, Yuanfang Guo, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 A New HEVC In-Loop Filter Based on Multi-channel Long-Short-Term Dependency Residual Networks
abstract
In this paper, we propose a new HEVC in-loop filter based on a multi-channel long-short-term dependency residual network (MLSDRN). Inspired by the information storage and information update function of human memory cell, our MLSDRN introduces an update cell to adaptively store and select the long-term and short-term dependency information through an adaptive learning process. In addition, we leverage the block boundary information that recorded in the bit-streams to improve the filter performance, which also makes our MLSDRN to unequally treat the video content. Meanwhile, the multi-channel is introduced to solve the illumination discrepancy problem. We integrate the novel in-loop filter into HM reference software, and applying it to luma and chroma components, simulation results demonstrate that the proposed in-loop filter can save BD-rate reduction up to 15.9% with ALF off. For luma component, the novel in-loop filter achieves 6.0%, 8.1%, 7.4% BD-rate saving for all intra, low delay and random access configurations, respectively.
Xiandong Meng, Chen Chen 0015, Shuyuan Zhu, Bing Zeng 0001
DCC3
2018 Coding Trajectory: Enable Video Coding for Video Denoising
abstract
We introduce a novel video denoising approach which can produce a clean video by utilizing redundant image patches existed in the video frames. Previous multi-frame video denosing approaches either require image registration or employ Patch Match algorithms for the discovery of the patch redundancy. However, these computations are time-consuming and prone to errors. On the other hand, nearly all captured videos have been compressed. Such a compression can produce a rich set of block-based motion vectors that can be utilized for the redundant patch extraction, leading to the efficient video denosing. To be specific, the motion vectors and frame references can be obtained from the video coding. Given a noised frame block, we follow its motion vectors from the coding to form a trajectory and gather a set of block candidates along the routes from its nearby frames. The trajectory is referred to as Coding Trajectory. Then, the corresponding denoised block is generated by weighted fusing the block candidates with outlier rejections. A denoised frame is consisted of all the denoised blocks. We compare our method with several state-of-the-art approaches, such as VBM3D and VB-M4D, in terms of PSNR and SSIM. The experiments show that our method can achieve high quality results while runs much faster then the other approaches.
Zhihang Ren, Peng Dai 0003, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001
ICIP4
2018 A 3D Descriptor based on Local Height Image
abstract
This paper proposes a novel 3D local descriptor, which seeks a good balance between the efficiency and the accuracy. We use the Local Reference Frame (LRF) to estimate a robust coordinate system to describe the local 3D shape. A novel Local Height Image (LHI) is defined by projecting the 3D points in the support region onto the tangent plane of the basis point. The Local Height Image Descriptor (LHID) is then defined by calculating the averaged projection distances. We further smooth the LHID to resist various kinds of interferences. We setup several experiments to assess the performance of our descriptor by comparison with the state-of-the-art algorithms. The experimental results demonstrate the effectiveness of the proposed method, which not only achieves the high accuracy as well as the robustness, but also possesses low complexity for the efficiency.
Tiecheng Sun, Shuaicheng Liu, Guanghui Liu 0001, Shuyuan Zhu, Zhipeng Zhu
ISCAS4
2018 Block-based Image Coding by Compression-Constrained Transform Domain Down-Scaling
abstract
Transform domain down-scaling (TDDS) is traditionally implemented by dropping most of high-frequency components of the transformed block. Applying it to image compression can improve the compression efficiency by saving considerable bit-cost. Due to losing some necessary high-frequency information, the resulted image compressed by using the traditional TDDS-based coding often suffers a serious quality degradation. In this paper, we propose a compression-constrained TDDS and perform it on each N × N block to produce an N/2 × N/2 coefficient block for the compression. Our proposed TDDS not only guarantees a high reconstruction quality but also makes a low bit-cost for compression. We integrate it in practical image coding to build up our proposed compression scheme. Experimental results show that our proposed method demonstrates excellent coding performance when used to compress image signals.
Chang Cui, Shuyuan Zhu, Xiandong Meng, Shuaicheng Liu, Bing Zeng 0001
VCIP2
2018 Multi-exposure Fusion With JPEG Compression Guidance
abstract
Construct a High Dynamic Range (HDR) image is the primary method to solve the information loss caused by insufficient dynamic range of cameras. We propose a technique for fusing a bracketed low dynamic range (LDR) image sequence of varying exposures into an HDR image, skipping the physically-based HDR assembly step. Traditionally approaches often rely on complicated algorithms to select good regions from the input LDR images for the fusion. However, we found that the selection strategy can purely base on JPEG compression bits, bypassing the calculations of image low-level features, such as image gradients, local saturations, over/under exposure evaluations, as long as the input LDR image is compressed by the JPEG formats. In this way, lots of computations can be saved. In particular, we extract the coding bits from the intermediate product of the JPEG. The coding bits of blocks can be modified as the weights for the exposure fusion. Well-exposed regions often require higher bits for the compression while overexposure or saturated regions often correspond to lower bits. The objective and subjective evaluations demonstrate the effectiveness of our method.
Xingdi Zhang, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001
VCIP3
2018 DC Coefficient Estimation of Intra-Predicted Residuals in HEVC
abstract
This paper presents a DC coefficient estimation algorithm for intra-predicted residual blocks in the High-Efficiency Video Coding (HEVC) standard. Discarding the DC coefficient directly in each transform block leads to substantial bit-saving, but at the same time produces strong discontinuities between neighboring blocks. To overcome this problem, we propose an estimation algorithm for the DC coefficient, which solves an optimal offset in a closed-form to recover the corresponding block edges. Then, we embed this algorithm into HEVC in its rate-distortion optimized quantization and sign bit hiding steps. Furthermore, a flag is signaled to decide whether the DC estimation strategy is used for each transform block. Test results under the common test condition show that our algorithm achieves 1.5% and 1.6% BD-rate reduction on average for luma and chroma, respectively, under all intra configuration. In the meantime, our simulation results show that both encoding time and decoding time increase only slightly (about 10%, without any special optimization on programming our proposed algorithm). When testing the proposed DC estimation algorithm on inter coding configurations, including low delay with P pictures, low delay with B pictures, and random access, we can also achieve 0.5%-1.1% bit-rate savings on average, while nearly no extra encoding and decoding time is needed.
Chen Chen 0015, Zexiang Miao, Xiandong Meng, Shuyuan Zhu, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.4
2018 Cross-Space Distortion Directed Color Image Compression
abstract
Traditional color image compression is usually conducted in the YCbCr space but many color displayers only accept RGB signals as inputs. Due to the use of a non-unitary matrix in the YCbCr-RGB conversion, low distortion achieved in the YCbCr space cannot guarantee low distortion for the RGB signals. To solve this problem, we propose a novel compression scheme for color images through defining a cross-space distortion so as to reduce as much as possible the distortion in the RGB space. To this end, we first derive the relationship between the distortions in the YCbCr space and RGB space. Then, we develop two solutions to implement color image compression for the most popular 4:2:0 chroma format. The first solution focuses on the design of a new spatial downsampling method to generate the 4:2:0 YCbCr image for a high-efficiency compression. The second one provides a novel way to reduce the distortion of the compressed color image by controlling the quantization error of the 4:2:0 YCbCr image, especially the one generated by using the traditional spatial downsampling. Experimental results show that both proposed solutions offer a remarkable quality gain over some state-of-the-art approaches when tested on various textured color images.
Shuyuan Zhu, Chen Chen 0015, Shuaicheng Liu, Bing Zeng 0001
IEEE Trans. Multim.1
2018 Robust Privacy-Preserving Image Sharing over Online Social Networks (OSNs)
abstract
Sharing images online has become extremely easy and popular due to the ever-increasing adoption of mobile devices and online social networks (OSNs). The privacy issues arising from image sharing over OSNs have received significant attention in recent years. In this article, we consider the problem of designing a secure, robust, high-fidelity, storage-efficient image-sharing scheme over Facebook, a representative OSN that is widely accessed. To accomplish this goal, we first conduct an in-depth investigation on the manipulations that Facebook performs to the uploaded images. Assisted by such knowledge, we propose a DCT-domain image encryption/decryption framework that is robust against these lossy operations. As verified theoretically and experimentally, superior performance in terms of data privacy, quality of the reconstructed images, and storage cost can be achieved.
Weiwei Sun 0009, Jiantao Zhou 0001, Shuyuan Zhu, Yuan Yan Tang
ACM Trans. Multim. Comput. Commun. Appl.3
2017 Shape Recovery of Endoscopic Videos by Shape from Shading Using Mesh Regularization
Zhihang Ren, Lingbing Peng, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001
ICIG (3)5
2017 Endoscopic video deblurring via synthesis
abstract
Endoscopic videos have been widely used for stomach diagnoses. However, endoscopic devices often capture videos with motion blurs, due to the dimly-lit environment and the camera shakiness during the capturing, which severely disturbs the diagnoses. In this paper, we present a framework that can restore blurry frames by synthesizing image details from the nearby sharp frames. Specifically, the blurry frame and their corresponding nearby sharp frames are identified according to the image gradient sharpness. To restore one blurry frame, a non-parametric mesh-based motion model is proposed to align the sharp frame to the blurry frame. The motion model leverages motions from image feature matches and optical flows, which yields high quality alignments to overcome challenges such as noisy, blurry, reflective and textureless interferences. After the alignment, the deblurred frame is synthesized by matching patches locally between the blurry frame and the aligned sharp frame. Without the estimation of blur kernels, we show that it is possible to directly compare a blurry patch against the sharp patches for the nearest neighbor matches in endoscopic images. The experiments demonstrate the effectiveness of our algorithm.
Lingbing Peng, Shuaicheng Liu, Dehua Xie, Shuyuan Zhu, Bing Zeng 0001
VCIP4
2017 MMSE-Directed Linear Image Interpolation Based on Nonlocal Geometric Similarity
abstract
In this letter, we propose a minimum mean square error (MMSE) directed linear interpolation to compose the high-resolution image from a single low-resolution image. We build up our interpolation model by using some similar image patches selected according to the nonlocal geometric similarity. First, we use a two-stage search scheme to collect the matched patches inside the whole image. Second, a similarity scaling factor is used in the second search to refine the collected patches so as to help find a robust solution to the MMSE-directed interpolation. Third, our MMSE-directed interpolation is regularized by the involved reference patches to make the solved interpolation coefficients more reliable. Experimental results show that our proposed method outperforms the state-of-the-art MMSE-directed linear interpolation schemes and works competitively with the state-of-the-art learning-based ones.
Shuyuan Zhu, Zhiying He, Shuaicheng Liu, Bing Zeng 0001
IEEE Signal Process. Lett.1
2017 A New Block-Based Method for HEVC Intra Coding
abstract
This paper presents a new block-based method for the High Efficiency Video Coding (HEVC) intra coding. First, we have found through analysis and test that the prediction errors on some pixels in each prediction block (PB) that are neighboring to the reference pixels would be no bigger than the corresponding coding errors. Based on this observation, the pixels in each PB are divided into two parts: half pixels are coded via a novel padding technique together with a constrained quantization algorithm (leading to around 3 dB gain under the same bit rate), whereas the other half are reconstructed by linear interpolations along a prediction direction by utilizing the neighboring reference pixels and the first half coded pixels. In the final implementation, a competition mechanism is employed between this new method and the original HEVC intra coding in order to choose the best mode for each PB. Experimental results show that about 2% BD-rate reduction has been achieved both for luma and chroma with respect to the original HEVC intra coding, whereas the encoder complexity increases by 130%, but the decoding time remains nearly unchanged.
Chen Chen 0015, Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj
IEEE Trans. Circuits Syst. Video Technol.2
2017 A Hybrid Approach for Near-Range Video Stabilization
abstract
Near-range videos contain objects that are close to the camera. These videos often contain discontinuous depth variation (DDV), which is the main challenge to the existing video stabilization methods. Traditionally, 2D methods are robust to various camera motions (e.g., quick rotation and zooming) under scenes with continuous depth variation (CDV). However, in the presence of DDV, they often generate wobbled results due to the limited ability of their 2D motion models. Alternatively, 3D methods are more robust in handling near-range videos. We show that, by compensating rotational motions and ignoring translational motions, near-range videos can be successfully stabilized by 3D methods without sacrificing the stability too much. However, it is time-consuming to reconstruct the 3D structures for the entire video and sometimes even impossible due to rapid camera motions. In this paper, we combine the advantages of 2D and 3D methods, yielding a hybrid approach that is robust to various camera motions and can handle the near-range scenarios well. To this end, we automatically partition the input video into CDV and DDV segments. Then, the 2D and 3D approaches are adopted for CDV and DDV clips, respectively. Finally, these segments are stitched seamlessly via a constrained optimization. We validate our method on a large variety of consumer videos.
Shuaicheng Liu, Binhan Xu, Chuang Deng, Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj
IEEE Trans. Circuits Syst. Video Technol.4
2017 CodingFlow: Enable Video Coding for Video Stabilization
abstract
Video coding focuses on reducing the data size of videos. Video stabilization targets at removing shaky camera motions. In this paper, we enable video coding for video stabilization by constructing the camera motions based on the motion vectors employed in the video coding. The existing stabilization methods rely heavily on image features for the recovery of camera motions. However, feature tracking is time-consuming and prone to errors. On the other hand, nearly all captured videos have been compressed before any further processing and such a compression has produced a rich set of block-based motion vectors that can be utilized for estimating the camera motion. More specifically, video stabilization requires camera motions between two adjacent frames. However, motion vectors extracted from video coding may refer to non-adjacent frames. We first show that these non-adjacent motions can be transformed into adjacent motions such that each coding block within a frame contains a motion vector referring to its adjacent previous frame. Then, we regularize these motion vectors to yield a spatially-smoothed motion field at each frame, named as CodingFlow, which is optimized for a spatially-variant motion compensation. Based on CodingFlow, we finally design a grid-based 2D method to accomplish the video stabilization. Our method is evaluated in terms of efficiency and stabilization quality, both quantitatively and qualitatively, which shows that our method can achieve high-quality results compared with the state-of-the-art methods (feature-based).
Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001
IEEE Trans. Image Process.3
2016 Joint bundled camera paths for stereoscopic video stabilization
abstract
This paper presents a method to stabilize shaky stereoscopic videos captured by hand-held devices. Directly applying traditional monocular video stabilization techniques to two views independently is problematic as it often brings undesirable vertical disparities and produces inaccurate horizontal disparities, which violate original stereoscopic disparity constraints, leading to erroneous depth perception. In this paper, we show that monocular video stabilization methods, such as the bundled camera paths stabilization, can be extended for stereoscopic videos by taking additional disparity constraints during the stabilization. In particular, we first estimate disparities between two views. Then, we compute camera motions as meshes of bundled paths for each view. Next, we smooth paths of two views separately and iteratively. During each iteration, we adjust the meshes of one view by our proposed `Joint Disparity and Stability mesh Warp (JDSW)'. The final result is generated after several iterations of paths smoothing and meshes adjusting, in which temporal stability and correct depth perception are achieved simultaneously. We evaluate our method by various challenging stereoscopic videos with different camera motions and scene types. The experiments demonstrate the effectiveness of our method.
Heng Guo 0003, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001
ICIP3
2016 A framework of single-image deraining method based on analysis of rain characteristics
abstract
In this paper, we propose an algorithm to remove rain streaks from single color image. Firstly, the guided filter, cooperated with rain pixels detection are used to separate a color image into low-frequency and high-frequency parts so that most rain components exist in the high-frequency part. Then, we focus on the high-frequency part to extract the non-rain details according to the characteristics of the rain in which a dictionary learning method is used. Meanwhile, to enhance the quality of the rain-removed image, the proposed principal direction of an image patch (PDIP) and the sensitivity of variance of color channels (SVCC) are employed in our work to help extract more non-rain details. Compared with the state-of-the-art works, our proposed method can remove the rain (especially heavy rain) from color images more efficiently.
Yinglong Wang 0002, Chen Chen 0015, Shuyuan Zhu, Bing Zeng 0001
ICIP3
2016 Intrinsic decomposition for stereoscopic images
abstract
Intrinsic image decomposition is an important technique that decomposes an image into reflectance and shading components. In this paper, we enable intrinsic decomposition for stereoscopic images. Traditional approaches cannot be directly applied to decompose stereoscopic images, yielding inconsistent reflectance and 3D artifacts after recoloring. To solve this problem, we propose a straight yet effective method for stereoscopic intrinsic decomposition, which consists of classical retinex constraint as well as disparity constraint. The former encodes the shading smoothness prior while the latter controls the reflectance similarity between two views. To further reduce ambiguity, we employ local and non-local texture cues by using superpixels within and across two views. The experiments show that our method can effectively decompose stereoscopic images with high quality and offer a comfortable 3D viewing experience.
Dehua Xie, Shuaicheng Liu, Kaimo Lin, Shuyuan Zhu, Bing Zeng 0001
ICIP4
2016 Constrained quantization based transform domain down-conversion for image compression
abstract
The image down-conversion may be used in the block-based image compression because it can help save lots of bit-counts for each individual block. A straightforward way to implement the transform domain down-conversion is to truncate some high-frequency components to get a down-sized coefficient block. However, directly using this down-sized coefficient block to reconstruct a completed image block will lead to a serious quality degradation. In this paper, we propose a constrained quantization based transform domain down-conversion (CQTDD) to help compress each 16×16 macro-block and it makes the coding quality of 1/4 selected pixels (according to a regular pattern) in each macro-block much higher than that can be achieved by using the traditional truncation based approach. Meanwhile, the other 3/4 pixels will be interpolated by using those 1/4 well-reconstructed pixels. Furthermore, these 1/4 pixels are optimized before the compression to help get a more efficient interpolation. Finally, the proposed CQTDD works with the JPEG baseline coding together as two candidate coding modes in our proposed compression scheme. Experimental results demonstrate that our proposed method may offer a remarkable quality gain, both objectively and subjectively, compared with some existing methods.
Shuyuan Zhu, Liaoyuan Zeng, Bing Zeng 0001, Jiantao Zhou 0001
ISCAS1
2016 Processing-Aware Privacy-Preserving Photo Sharing over Online Social Networks
abstract
With the ever-increasing popularity of mobile devices and online social networks (OSNs), sharing photos online has become extremely easy and popular. The privacy issues of shared photos and the associated protection schemes have received significant attention in recent years. In this work, we address the problem of designing privacy-preserving, high-fidelity, storage-efficient photo sharing solution over Facebook. We first conduct an in-depth study on the manipulations that Facebook performs to the uploaded images. With the awareness of such information, we suggest a DCT-domain image encryption scheme that is robust against these lossy operations. As validated by our experimental results, superior performance in terms of security, quality of the reconstructed images, and storage cost can be achieved.
Weiwei Sun 0009, Jiantao Zhou 0001, Ran Lyu, Shuyuan Zhu
ACM Multimedia4
2016 DC coefficient estimation of intra-predicted residuals in high efficiency video coding
abstract
This paper proposes a DC coefficient estimation algorithm for intra-predicted residual blocks in the High Efficiency Video Coding (HEVC) standard. Discarding the DC coefficient in the current coding block leads to a substantial bit-saving but produces at the same time strong discontinuities between this block and its neighboring reconstructed blocks. To overcome this problem, we propose an estimation algorithm for the DC coefficient, which solves an optimal offset in a closed-form in the pixel domain to recover the corresponding block edges. Test results show that our algorithm achieves 1.0% and 1.4% BD-rate reduction on average for luma and chroma as compared with HM-16.6, respectively, when the sign-bit-hiding (SBH) technique is disabled. When SDH is set on, namely under the common test condition (CTC), the BD-rate reduction drops slightly to 0.7% and 1.1% for luma and chroma, respectively. In the meantime, the test results show that both encoding time and decoding time increase only slightly (about 10%, without any special optimization on programming our proposed algorithm).
Chen Chen 0015, Zexiang Miao, Xiandong Meng, Shuyuan Zhu, Bing Zeng 0001
VCIP4
2016 Mode-dependent transforms based on elliptical model for high efficiency video coding
abstract
High efficiency video coding (HEVC) defines 35 prediction modes in its intra prediction stage to signal the direction information of residual blocks. Traditionally, separable two-dimension (2-D) transforms (integer DCT and DST) are utilized in a similar manner as in the previous H.264/AVC standards. However, such 2-D transforms cannot yield the best energy compaction for a 2-D directional source where the dominating directional information is other than the horizontal or vertical one. In order to overcome this drawback, we build an elliptical model with directionality and design some non-separable transforms based on the Karhunen-Loeve transform in this paper. Specifically, we derive a non-separable transform in closed-form for each intra-prediction mode and replace the default transform in HEVC. Simulation results reveal that 1.7% and 2.0% on average and up to 7.7% and 8.1% BD-rate reduction can be achieved for luma and chroma component, respectively. In the meantime, the test results show that both the encoding time and decoding time increase only about 5%.
Kaiyuan Jia, Chen Chen 0015, Xiandong Meng, Shuyuan Zhu, Bing Zeng 0001
VCIP4
2016 Interpolation-directed transform domain downward conversion for block-based image compression
abstract
In this paper, we design an interpolation-directed transform domain downward conversion (ITDDC) to build up a new block-based image compression scheme. This ITDDC is derived from our proposed 2-D padding and performed on each 16×16 macro-block of pixels to convert it into an 8×8 coefficient block, leading to a downward image conversion in the transform domain. More interestingly, the further compression is just performed on the down-sized coefficient block and the reconstruction for an entire macro-block is achieved via the interpolation by using the decoded pixels only locating in some specific positions of it. To make the interpolation more efficient, the pixels participating in the interpolation will be optimized before the compression. The ITDDC-based coding is used competitively with the JPEG baseline coding to compress each macro-block in our proposed compression scheme according to a simple but efficient rate-distortion optimization based criterion. Experimental results demonstrate that our proposed method gets a remarkable quality gain over the existing approaches.
Shuyuan Zhu, Jinglin Yu, Chen Chen 0015, Liaoyuan Zeng, Bing Zeng 0001
VCIP1
2016 Joint Video Stitching and Stabilization From Moving Cameras
abstract
In this paper, we extend image stitching to video stitching for videos that are captured for the same scene simultaneously by multiple moving cameras. In practice, videos captured under this circumstance often appear shaky. Directly applying image stitching methods for shaking videos often suffers from strong spatial and temporal artifacts. To solve this problem, we propose a unified framework in which video stitching and stabilization are performed jointly. Specifically, our system takes several overlapping videos as inputs. We estimate both inter motions (between different videos) and intra motions (between neighboring frames within a video). Then, we solve an optimal virtual 2D camera path from all original paths. An enlarged field of view along the virtual path is finally obtained by a space-temporal optimization that takes both inter and intra motions into consideration. Two important components of this optimization are that: 1) a grid-based tracking method is designed for an improved robustness, which produces features that are distributed evenly within and across multiple views and 2) a mesh-based motion model is adopted for the handling of the scene parallax. Some experimental results are provided to demonstrate the effectiveness of our approach on various consumer-level videos and a Plugin, named "Video Stitcher" is developed at Adobe After Effects CC2015 to show the processed videos.
Heng Guo 0003, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj
IEEE Trans. Image Process.4
2016 Image Interpolation Based on Non-local Geometric Similarities and Directional Gradients
abstract
Image interpolation offers an efficient way to compose a high-resolution (HR) image from the observed low-resolution (LR) image. Advanced interpolation techniques design the interpolation weighting coefficients by solving a minimum mean-square-error (MMSE) problem in which the local geometric similarity is often considered. However, using local geometric similarities cannot usually make the MMSE-based interpolation as reliable as expected. To solve this problem, we propose a robust interpolation scheme by using the nonlocal geometric similarities to construct the HR image. In our proposed method, the MMSE-based interpolation weighting coefficients are generated by solving a regularized least squares problem that is built upon a number of dual-reference patches drawn from the given LR image and regularized by the directional gradients of these patches. Experimental results demonstrate that our proposed method offers a remarkable quality improvement as compared to some state-of-the-art methods, both objectively and subjectively.
Shuyuan Zhu, Bing Zeng 0001, Liaoyuan Zeng, Moncef Gabbouj
IEEE Trans. Multim.1
2015 Image interpolation based on non-local geometric similarities
abstract
Image interpolation refers to constructing a high-resolution (HR) image from a low-resolution (LR) image. Traditionally, an HR image can be produced from an observed LR image via the polynomial-based interpolation (bi-linear or bi-cubic interpolations, involving a small number of neighbors around each interpolated position). The advanced interpolation makes use of the so-called “geometric similarity” to design a set of optimal interpolation weighting coefficients. However, better geometric similarities can perhaps be found from a non-local area within the LR source image or even from other but similar images (possibly with higher resolutions). Based on this fact, we propose in this paper a non-local geometric similarity based interpolation scheme to construct HR images. In our proposed method, optimal weighting coefficients are determined by solving a regularized least squares problem which is built upon a number of dual reference patches drawn from the observed LR image and regularized by the variation of directional gradients of the image patch. Experimental results demonstrate that our proposed method offers a remarkable quality improvement, both objectively and subjectively.
Shuyuan Zhu, Bing Zeng 0001, Guanghui Liu 0001, Liaoyuan Zeng, Moncef Gabbouj
ICME1
2015 No reference image quality assessment metric via multi-domain structural information and piecewise regression
Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan, Shuyuan Zhu
J. Vis. Commun. Image Represent.5
2015 Adaptive sampling for compressed sensing based image compression
Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj
J. Vis. Commun. Image Represent.1
2015 Constrained Directed Graph Clustering and Segmentation Propagation for Multiple Foregrounds Cosegmentation
abstract
This paper proposes a new constrained directed graph clustering (DGC) method and segmentation propagation method for the multiple foreground cosegmentation. We solve the multiple object cosegmentation with the perspective of classification and propagation, where the classification is used to obtain the object prior of each class and the propagation is used to propagate the prior to all images. In our method, the DGC method is designed for the classification step, which adds clustering constraints in cosegmentation to prevent the clustering of the noise data. A new clustering criterion such as the strongly connected component search on the graph is introduced. Moreover, a linear time strongly connected component search algorithm is proposed for the fast clustering performance. Then, we extract the object priors from the clusters, and propagate these priors to all the images to obtain the foreground maps, which are used to achieve the final multiple objects extraction. We verify our method on both the cosegmentation and clustering tasks. The experimental results show that the proposed method can achieve larger accuracy compared with both the existing cosegmentation methods and clustering methods.
Fanman Meng, Hongliang Li 0001, Shuyuan Zhu, Bing Luo 0003, Chao Huang 0003, Bing Zeng 0001, Moncef Gabbouj
IEEE Trans. Circuits Syst. Video Technol.3
2014 Robust interactive image segmentation via iterative refinement
abstract
Image segmentation with user inputs gets more and more popular in recent years and always performs better compared with automatic methods. However, existing interactive image segmentation methods still might fail if the image contains messy textures, or the user inputs are sparse or at inappropriate locations. In this paper, we propose a novel iterative refinement framework which leads to robust segmentation performance even with sparse and improper input strokes. Specifically, a geodesic distance based energy is introduced and combined with convex active contour model, and an iterative seeds refinement technique is put forward to handle the sparse input problem. Extensive experiments using real world images, and segmentation benchmark dataset show that our proposed method has superior performance compared with representative state-of-the-art methods.
Juyong Zhang, Yancheng Yuan, Shuyuan Zhu
ICIP4
2014 Fast and efficient inter CU decision for high efficiency video coding
abstract
In this paper, a graph cut based fast Coding Unit (CU) decision algorithm is proposed for HEVC inter frames. Firstly, a feature called pyramid variance of the absolute difference (PVAD) is designed for the CU selection. Secondly, the CU decision is modeled as a Markov Random Field (MRF) inference problem, which can be optimized by the graph cut algorithm. Thirdly, a maximum a posteriori (MAP) approach based on the R-D cost is conducted to evaluate whether the unsplit CUs should be further split or not. Experimental results show the effectiveness of the proposed method.
Jian Xiong 0005, Hongliang Li 0001, Fanman Meng, Bing Zeng 0001, Shuyuan Zhu, Qingbo Wu 0001
ICIP5
2014 Downward spatially-scalable image reconstruction based on compressed sensing
abstract
According to the compressed sensing (CS) theory, we can sample a sparse signal at a rate that is (much) lower than the required Nyquist rate, while still enabling a nearly exact reconstruction. Image signals are sparse when represented in a certain domain, and because of this, a large number of CS-based image sampling and reconstruction techniques have been developed recently. In this paper, we focus on the design of the downward spatially-scalable image reconstruction from the CS-sampled data. Traditional methods usually reconstruct an image whose size is the same as the original source image and then achieve the downward scalability through sub-sampling. In our proposed method, we unify these two steps into a single one and promise to deliver a much improved quality.
Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj
ICIP1
2014 Adaptive sampling for compressed sensing based image compression
abstract
The compressed sensing (CS) theory shows that a sparse signal can be recovered at a sampling rate that is (much) lower than the required Nyquist rate. In practice, many image signals are sparse in a certain domain, and because of this, the CS theory has been successfully applied to the image compression in the past few years. The most popular CS-based image compression scheme is the block-based CS (BCS). In this paper, we focus on the design of an adaptive sampling mechanism for the BCS through a deep analysis of the statistical information of each image block. Specifically, this analysis will be carried out at the encoder side (which needs a few overhead bits) and the decoder side (which requires a feedback to the encoder side), respectively. Two corresponding solutions will be compared carefully in our work. We also present experimental results to show that our proposed adaptive method offers a remarkable quality improvement compared with the traditional BCS schemes.
Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj
ICME1
2014 Adaptive reweighted compressed sensing for image compression
abstract
According to the compressed sensing (CS) theory, a signal that is sparse in a certain domain can be nearly exactly recovered from a few measurements where the sampling rate is lower than the Nyquist rate. This theory has been successfully applied to the image compression in the past few years as most image signals are highly sparse. In this paper, we apply an adaptive sampling mechanism to the reweighted block-based CS (BCS). The proposed adaptive sampling allocates the measurements to each image block according to the statistical information of the block so as to sample and recover the image more efficiently. Experimental results demonstrate that our adaptive reweighted method offers a very significant quality improvement compared with the traditional BCS schemes, including the non-reweighted and reweighted ones.
Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj
ISCAS1
2014 Bird breed classification and annotation using saliency based graphical model
Chao Huang 0003, Fanman Meng, Shuyuan Zhu
J. Vis. Commun. Image Represent.4
2014 Perceptual Encryption of H.264 Videos: Embedding Sign-Flips Into the Integer-Based Transforms
abstract
An alternative-transforms-based scheme has recently been proposed to achieve perceptual encryption of video signals in which multiple transforms are designed by using different rotation angles at the final stage of the discrete cosine transforms (DCTs) butterfly flow-graph structure. More recently, it is found that a set of more efficient alternative transforms can be derived by introducing sign-flips at the same stage, which is equivalent to an extra rotation angle of π. In this paper, we generalize this sign-flipping technique by randomly embedding sign-flips into all stages of the DCTs butterfly structure so that the encryption space becomes much larger to yield a higher security. We pursue this study for H.264-compatible videos, assuming that the integer DCT of size 4 × 4 is used. First, we follow the separable implementation of the 4 × 4 2-D DCT in which different sign-flipping strategies will be employed along its horizontal and vertical dimensions. Second, we convert the 4 × 4 2-D DCT into a 16-point 1-D butterfly structure so that more sign-flips can be embedded at its various stages. Third, we choose different schemes to pair the node-variables in the 16-point 1-D butterfly structure, thus further enlarging the encryption space. Extensive experiments are conducted to show the performance of these improved encryption schemes and some security analyzes are also presented to confirm their persistence to various attacking strategies.
Bing Zeng 0001, Jeff Siu-Kei Au-Yeung, Shuyuan Zhu, Moncef Gabbouj
IEEE Trans. Inf. Forensics Secur.3
2014 Repairing Bad Co-Segmentation Using Its Quality Evaluation and Segment Propagation
abstract
In this paper, we improve co-segmentation performance by repairing bad segments based on their quality evaluation and segment propagation. Starting from co-segmentation results of the existing co-segmentation method, we first perform co-segmentation quality evaluation to score each segment. Good segments can be filter out based on the scores. Then, a propagation method is designed to transfer good segments to the rest bad ones so as to repair the bad segmentation. In our method, the quality evaluation is implemented by the measurements of foreground consistency and segment completeness. Two propagation methods such as global propagation and local region propagation are then defined to achieve the more accurate propagation. We verify the proposed method using four state-of-the-arts co-segmentation methods and two public datasets such as ICoseg dataset and MSRC dataset. The experimental results demonstrate the effectiveness of the proposed quality evaluation method. Furthermore, the proposed method can significantly improve the performance of existing methods with larger intersection-over-union score values.
Hongliang Li 0001, Fanman Meng, Bing Luo 0003, Shuyuan Zhu
IEEE Trans. Image Process.4
2014 MRF-Based Fast HEVC Inter CU Decision With the Variance of Absolute Differences
abstract
The newly developed High Efficiency Video Coding (HEVC) Standard has improved video coding performance significantly in comparison to its predecessors. However, more intensive computation complexity is introduced by implementing a number of new coding tools. In this paper, a fast coding unit (CU) decision based on Markov random field (MRF) is proposed for HEVC inter frames. First, it is observed that the variance of the absolute difference (VAD) is proportional with the rate-distortion (R-D) cost. The VAD based feature is designed for the CU selection. Second, the decision of CU splittings is modeled as an MRF inference problem, which can be optimized by the Graphcut algorithm. Third, a maximum a posteriori (MAP) approach based on the R-D cost is conducted to evaluate whether the unsplit CUs should be further split or not. Experimental results show that the proposed algorithm can achieve about 53% reduction of the coding time with negligible coding performance degradation, which outperforms the state-of-the-art algorithms significantly.
Jian Xiong 0005, Hongliang Li 0001, Fanman Meng, Shuyuan Zhu, Qingbo Wu 0001, Bing Zeng 0001
IEEE Trans. Multim.4
2013 A novel enhancement for hierarchical image coding
Shuyuan Zhu, Bing Zeng 0001
J. Vis. Commun. Image Represent.1
2012 An enhanced block-based hierarchical image coding scheme
abstract
A block-based, two-layer hierarchical image coding scheme codes a down-sampled version of the original image first, and then the residue between the original one and a reconstructed version that is interpolated from the upper layer. When the residual layer is coded at a lower quality as compared to its upper layer, it often happens that all pixels in the upper layer will be deteriorated if the corresponding coded residuals are added into them. In this paper, we demonstrate that there exists a critical region for each image such that the deterioration indeed happens and this region includes nearly all typical bit-rates used in practice. To avoid this problem, we first propose a “naive” solution and then apply an advanced quantization technique to handle the residual layer. To verify the effectiveness, we conduct extensive tests to show that the gap between the hierarchical coding scheme and its single-level counterpart (which is typically around 2~3 dB in the two-layer scenario) has been filled up by a rather big percentage (about 30~50%).
Shuyuan Zhu, Bing Zeng 0001
ICIP1
2012 Image Super-Resolution via Low-Pass Filter Based Multi-scale Image Decomposition
abstract
This paper presents a spatial-varying minimum mean square error (MMSE)-based approach to construct super-resolution images from single source image of a lower resolution. The unique feature of this approach is that it works on a set of sub-images (also called multi-scale images) that are generated via decomposing the original source image. To do the decomposition, we design a number of low-pass filters with overlapped pass-bands so that sub-images are correlated with each other. Then, an MMSE-based estimation, involving all sub-images, is solved (after making use of the geometric-duality principle) to construct each missing pixel in the super-resolution image. Experimental results show that our new method offers a clearly-noticeable improvement over the existing MMSE-based methods (without decomposition). We believe that this is mainly attributing to the fact that both intra-scale and inter-scale correlations among the sub-images have been utilized in our approach.
Shuyuan Zhu, Bing Zeng 0001, Shuicheng Yan
ICME1
2012 Total-variation based picture reconstruction in multiple description image and video coding
Shuyuan Zhu, Jeff Siu-Kei Au-Yeung, Bing Zeng 0001, Jiying Wu
Signal Process. Image Commun.1
2011 Perceptual video encryption using multiple 8×8 transforms in H.264 and MPEG-4
abstract
It has been demonstrated in our earlier works [1, 2] that perceptual video encryption can be effectively achieved by using multiple transforms where the block size 4×4 has been considered. In this paper, we study the extension to the transforms of size 8×8. In this case, a more complex flow-graph structure is resulted, thus leading to a larger room for encryption. In addition, special technique on controlling the encrypted video quality is presented by carefully selecting the number of rotations in the flow-graph structure of an 8×8 transform. The proposed scheme is first evaluated using the high profile of H.264. It is then further tested for the MPEG-4 standard that completely relies on 8×8 transform. Both cases show that promising results can be achieved with our proposed scheme.
Jeff Siu-Kei Au-Yeung, Shuyuan Zhu, Bing Zeng 0001
ICASSP2
2011 SEAMLESS P2P-MDVC with well-balanced descriptions
abstract
Multiple-description coding (MDC) provides a promising solution to support the error-prone transmission over multiple channels. One extremely important application is to design efficient multiple description video coding (MDVC) systems for the peer to peer (P2P) scenario. To this end, one needs to solve the mismatching problem between the reference frames used at the encoder and decoder sides (for motion compensation) in most non-scalable MDVC schemes or get rid of the inter-dependency within various enhancement layers in the scalable MDVC schemes. In the meantime, it is highly preferable to enforce that all descriptions transmitted over the network are well-balanced so as to have an equal payload to each peer. In this paper, we propose an MDVC scheme with a number of well-balanced descriptions. These descriptions are generated from some newly-developed unitary transforms and they can solve all of the problems mentioned above.
Shuyuan Zhu, Jeff Siu-Kei Au-Yeung, Bing Zeng 0001
ICASSP1
2011 Design of non-separable transforms for directional 2-D sources
abstract
Traditionally, a 2-D block-based transform is always implemented through two separate 1-D transforms along each block's vertical and horizontal dimensions. Such a framework is however not highly suitable for a directional 2-D source in which the dominant directional information is neither horizontal nor vertical. On the other hand, the R-D performance upper bound for all block-based transform coding schemes applied on such 2-D directional sources can be obtained by the non-separable Karhunen-Loève transform (KLT) — which is unfortunately very expensive computationally. In this paper, we present a new framework for designing some non-separable transforms that offer an R-D performance closer to that of the KLT, but can be implemented with nearly the same complexity as that of the discrete cosine transform (DCT).
Haoming Chen, Shuyuan Zhu, Bing Zeng 0001
ICIP2
2011 A comparative study of image correlation models for directional two-dimensional sources
abstract
The non-separable Karhunen-Loève transform (KLT) has been proven to be optimal for coding a directional 2-D source in which the dominant directional information is neither horizontal nor vertical. However, the KLT depends on the image data, and it is difficult to apply it in a practical image/video coding application. In order to solve this problem, it is necessary to build an image correlation model, and this model needs to adapt to the directional information so as to facilitate the design of 2-D non-separable transforms. In this paper, we compare two models that have been used commonly in practice: the absolute-distance model and the Euclidean-distance model. To this end, theoretical analysis and experimental study are carried out based on these two models, and the results show that the Euclidean-distance model consistently performs better than the absolute-distance model.
Shuyuan Zhu, Bing Zeng 0001
MMSP1
2011 Design of New Unitary Transforms for Perceptual Video Encryption
abstract
In our earlier work , we proposed for the first time that the perceptual video encryption be performed at the transformation stage by selecting one out of multiple unitary transforms according to the encryption key. In this letter, we aim to design some more efficient transforms to be used in this framework. Two criteria are followed for designing such transforms: 1) they are significantly different from discrete cosine transform (DCT) or discrete sine transform (DST), and 2) the resulted coding efficiency is exactly the same to what can be achieved by using DCT or just falls very slightly. As a result, we find that these transforms are actually derived by the sign-flipping on some node-variables in the flow-graph structure of DCT - a special case of the plane-based rotations (by π or 180°). Extensive simulations based on the H.264 codec are performed to demonstrate their effectiveness through both objective and subjective assessments. Finally, we present the security analysis to show the resistance of our algorithms to different types of attacks.
Jeff Siu-Kei Au-Yeung, Shuyuan Zhu, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.2
2010 Partial video encryption based on alternative integer transforms
abstract
It has been demonstrated in our earlier work [1] that partial video encryption can be effectively achieved by using alternative transforms. In this paper, we modify those floating-point based transforms to integer-based counterparts so as to be compatible with the H.264 standard. To this end, we present the design details of these integer-based transforms and their efficient implementations. To follow the H.264 standard that combines the transform and quantization processes, we present some special handling of the necessary scaling factors and the associated quantization. Finally, we demonstrate that the derived integer transforms can achieve partial video encryption with nearly the same performance as what has been achieved with floating-point transforms.
Jeff Siu-Kei Au-Yeung, Shuyuan Zhu, Bing Zeng 0001
ISCAS2
2010 Composing better pictures in MDC: A multi-target total variational approach
abstract
In any multiple description coding (MDC) system, how to compose pictures of the “best” quality upon receiving more than one description is always important. To this goal, a DCT-domain translation and overlapping technique has been developed in for the popular multiple description scalar quantization (MDSQ) based video coding. Although certain quality gain (in terms of PSNR) has been achieved, the improvement is rather limited. In this paper, we formulate this problem into a traditional total variation (TV) regularized optimization in which all received descriptions are regarded as multiple fidelity terms. We demonstrate that, as compared to the overlapping technique, such a TV-based enhancement yields not only a higher objective quality gain but also a much more significant visual quality gain in the composed pictures.
Shuyuan Zhu, Jiying Wu, Bing Zeng 0001
ISCAS1
2010 A total variation-based approach for composing better pictures in multiple description coding
abstract
One of the important issues in multiple description coding (MDC) for image and video signals is to compose pictures of the "best" quality when more than one description is received at the decoder side. To this goal, a transform-domain overlapping technique combined with some necessary translations in the transform domain is proposed in the popular MDC schemes with staggered quantizers. However, the achievable gain is rather limited. In this paper, we formulate this enhancement problem as a total variation (TV) regularized optimization constrained by the knowledge of the quantization intervals of the DCT coefficients of each composed picture. In this TV-based approach, the transformdomain overlapping technique is used to find the more accurate quantization intervals in which the true DCT coefficients will fall while receiving more than one description. Simulation results demonstrate that such a TV-based enhancement yields a higher quality gain in both objective (e.g. PSNR-based) and subjective (i.e. visual perception) evaluations than using the transform-domain overlapping technique.
Shuyuan Zhu, Bing Zeng 0001
VCIP1
2010 In Search of "Better-than-DCT" Unitary Transforms for Encoding of Residual Signals
abstract
It is well known that the discrete cosine transform (DCT) closely approximates the optimal Karhunen-Loève transform (KLT) under the first-order stationary Markov condition with a strong inter-pixel correlation. However, if the inter-pixel correlation is weak or becomes negative, transforms other than the DCT would possibly become better. In this letter, we present a design framework for finding new unitary transforms according to two principles: 1) they are indeed more efficient than the DCT in encoding of signals with a weak or negative inter-pixel correlation (e.g., the residual signals after motion-compensation or intra- prediction) and 2) each has a fixed transform matrix (i.e., signal-independent) and can be implemented as efficiently as the DCT. Some of the new transforms will be used to demonstrate an improved R-D performance in the video coding scenario.
Shuyuan Zhu, Jeff Siu-Kei Au-Yeung, Bing Zeng 0001
IEEE Signal Process. Lett.1
2010 Constrained Quantization in the Transform Domain With Applications in Arbitrarily-Shaped Object Coding
abstract
In any block-based transform coding of image/video signals, it is well-known that the mean square error (MSE) distortion measured in the pixel domain is exactly equal to the MSE distortion resulted from quantization in the transform domain if the involved transform matrix is unitary. However, such a property no longer exists if the pixel-domain distortion is measured only on a selected part of pixels within one image block. This provides us an opportunity of dynamically shaping the quantization errors so as to make the selected pixels (much) better than the unselected ones. In this paper, we first develop a reversed iterative algorithm to guide us to perform a highly constrained quantization so that the coding quality of the selected pixels in each image block is significantly higher than what can be achieved by using the normal quantization. Then, we apply this intelligent quantization in one practical scenario-coding of arbitrarily-shaped image blocks in MPEG-4, showing remarkable improvements in comparison with the original MPEG-4.
Shuyuan Zhu, Bing Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2009 A comparative study of multiple description video coding in P2P: normal MDSQ versus flexible dead-zones
abstract
Multiple-description coding (MDC) encodes a source image or video into several distinct descriptions. Each description can be decoded independently and more than one description can also be decoded jointly so as to deliver a higher quality. One typical design of an MDC encoder is based on the so-called multiple description scalar quantization (MDSQ). A simple MDSQ system is to choose different step-sizes used in all involved quantizers. Another implementation is to apply flexible dead-zones in all quantizers (whereas keeping the same step-size across them). Two issues turn to be particularly important when MDC is used in video streaming applications in the P2P scenario. First, we need to consider more than two descriptions to fit the P2P reality. Second, we have to eliminate any possible drifting errors that are resulted from the use of different reference frames (in different descriptions) at the motion-compensation stage. In this paper, we introduce a DCT-domain translation to solve the mismatching problem in different reference frames. Then, we present some comparative results of the two approaches mentioned above, with detailed pros and cons for each one as well as some practical considerations.
Shuyuan Zhu, Jeff Siu-Kei Au-Yeung, Bing Zeng 0001
ICME1
2009 Partial Video Encryption Based on Alternating Transforms
abstract
In this letter, we propose a novel video encryption technique that is used to achievepartialencryption where an annoying video can still be reconstructed even without the security key. In contrast to the existing methods where the encryption usually takes place at the entropy-coding stage or the bit-stream level, our proposed scheme embeds the encryption at thetransformstage during the encoding process. To this end, we develop a number of new unitary transforms that are demonstrated to be equally efficient as the well-known DCT and thus used as alternates to DCT during the encoding process. Partial encryption is achieved through alternately applying these transforms to individual blocks according to a pre-designed secret key. Analysis on the security level of this partial encryption scheme is carried out against various common attacks and some experimental results based on H.264/AVC are presented.
Jeff Siu-Kei Au-Yeung, Shuyuan Zhu, Bing Zeng 0001
IEEE Signal Process. Lett.2
2009 R-D Performance Upper Bound of Transform Coding for 2-D Directional Sources
abstract
Traditionally, any 2D transform (such as 2D DCT) is implemented through two separable 1D transforms along the vertical and horizontal dimensions. Such a framework is however not most suitable for a 2D directional source in which the dominant directional information is neither horizontal nor vertical. In this letter, we attempt to determine the R-D performance upper bound for block-based transform coding schemes applied on such 2D directional sources. It is not a surprise that the Karhunen-Loeve transform (KLT) plays a critical role here. Specifically, we show that a nonseparable KLT can be determined directly from the given 2D directional source model to yield the R-D performance upper bound. We also show that there exists a significant gap between this upper bound and the R-D performance that can be achieved by using the traditional 2D DCT.
Shuyuan Zhu, Jeff Siu-Kei Au-Yeung, Bing Zeng 0001
IEEE Signal Process. Lett.1
2008 Coding of arbitrarily-shaped image blocks based on a constrained quantization
abstract
Coding of arbitrarily-shaped image/video segments is important to achieving the object-based coding - which is a key technique in many multimedia applications. MPEG-4 suggested a low-pass-extrapolation (LPE) padding method to handle each boundary block of an arbitrary shape. In this paper, we develop a constrained quantization technique and apply it to each LPE-padded boundary block. This ‘smart’ quantization is performed in such a way that the pixel-domain MSE distortion due to the DCT domain quantization is shaped as much as possible onto all padded pixels — as these pixels will be thrown away in the end. In the meantime, all pixels in each original boundary block would be coded with an improved quality. Compared to the normal rounding-based quantization used in the original MPEG-4’s LPE scheme, our new quantization needs some extra computations at the quantization stage in the encoding procedure but keeps the decoding complexity unchanged, while offering a remarkable coding gain — as demonstrated by some simulation results.
Shuyuan Zhu, Bing Zeng 0001
ICME1
2008 Seamless MDVC in P2P: A transform-domain approach
abstract
Multiple-description coding (MDC) provides an effective way to mitigate the effects of packet errors/loses by making use of multiple channels. Perhaps, the most attractive application of MDC is in the peer-to-peer (P2P) scenario to support simultaneous video streaming to a large population of clients. To this end, a number of multiple-description video coding (MDVC) schemes (both non-scalable and scalable) were proposed in the past few years. However, almost all non-scalable schemes would suffer from the prediction mismatch between the references used at the encoder and decoder sides; whereas all scalable schemes (involving a base-layer and some enhancement layers) would suffer from the inter-dependency within the enhancement-layer information. In this paper, we propose a transform-domain MDVC method that can solve these problems and at the same time offer some other interesting features.
Shuyuan Zhu, Bing Zeng 0001
MMSP1
2007 Multiple Description Transform Coding: A New Design Approach
abstract
Multiple-description coding (MDC) is an effective technique to support error-prone transmission in the scenario of multiple channels. Typically, an MDC is designed in the DCT domain where a pair-wise correlating transform (PCT) or some generalized version is further applied to some DCT coefficients of each image block. Although these correlating transforms offer a freedom of compromising the overall coding quality and the redundancy, none of them can be made optimal after taking into consideration the quantization characteristics. In this paper, we present a new design approach in which we will do a novel grouping of all DCT coefficients (not necessarily pair-wise) such that the distortion measured on some specified pixels within each image block is minimized. An interesting property of this new MDC design is that, when all descriptions are received, it provides a coding quality that becomes higher than the target quality set in the corresponding single-description coding (SDC) design.
Bing Zeng 0001, Shuyuan Zhu
ICME2