VLDB 2026 Research / reviewers in the wild / expert
Bing Zeng 0001
dblp:26/3636-1
· DBLP profile ↗
208ranked-venue papers
9as first author
76since 2021 · last 2026
0000-0002-4491-7967ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 159 · 7 first-author · 57 since 2021Artificial intelligence and machine learning · 40 · 31 since 2021Systems, architecture and hardware · 17 · 3 since 2021Computer networks · 13 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RAW-Flow: Advancing RGB-to-RAW Image Reconstruction with Deterministic Latent Flow MatchingabstractRGB-to-RAW reconstruction, or the reverse modeling of a camera Image Signal Processing (ISP) pipeline, aims to recover high-fidelity RAW data from RGB images. Despite notable progress, existing learning-based methods typically treat this task as a direct regression objective and still struggle with detail inconsistency and color deviation, due to the ill-posed nature of inverse ISP and the inherent information loss in quantized RGB images. To address these limitations, we pioneer a generative perspective by reformulating RGB-to-RAW reconstruction as a deterministic latent transport problem and introduce a novel framework named RAW-Flow, which leverages flow matching to learn a deterministic vector field in latent space, to effectively bridge the gap between RGB and RAW representations and enable accurate reconstruction of structural details and color information. To further enhance latent transport, we introduce a cross-scale context guidance module that injects hierarchical RGB features into the flow estimation process. Moreover, we design a Dual-domain Latent Autoencoder (DLAE) with a feature alignment constraint to support the proposed latent transport framework, which jointly encodes RGB and RAW inputs while promoting stable training and high-fidelity reconstruction. Extensive experiments demonstrate that RAW-Flow outperforms state-of-the-art approaches both quantitatively and visually. Zhen Liu 0022, Diedong Feng, Hai Jiang 0006, Liaoyuan Zeng, Hao Wang 0073, Chaoyu Feng, Bing Zeng 0001, Shuaicheng Liu |
AAAI | 8 |
| 2026 | Robust Partial-to-Partial Point Cloud Registration with Overlapping Mask Learning
Hao Xu 0018, Guanghui Liu 0001, Bing Zeng 0001, Shuaicheng Liu |
Int. J. Comput. Vis. | 3 |
| 2026 | Inter predictive coding for point cloud attributes with online coordinate alignment and multi-scale latent prediction
Yu Liu 0091, Shuyuan Zhu, Zeliang Li, Jeff Siu-Kei Au-Yeung, Fan Zhang 0017, Bing Zeng 0001 |
Neurocomputing | 6 |
| 2026 | Supervised Small-Baseline and Large-Baseline Homography Learning With Diffusion-Based Data GenerationabstractIn this paper, we propose an iterative framework, which consists of two phases: a generation phase and a training phase, to generate realistic training data for supervised small-baseline and large-baseline homography learning and yield a state-of-the-art homography estimation network. In the generation phase, given an unlabeled image pair, we utilize the pre-estimated dominant plane masks and homography of the pair, along with another sampled homography that serves as ground truth to generate a new labeled training pair with realistic motion. In the training phase, the generated data is used to train the supervised homography network, in which the training data is refined via a content refinement diffusion model. Once an iteration is finished, the trained network is used in the next data generation phase to update the pre-estimated homography. Through such an iterative strategy, the quality of the dataset and the performance of the network can be gradually and simultaneously improved. Experimental results show that our method outperforms existing competitors and previous supervised methods can also be improved based on the generated dataset. Hai Jiang 0006, Haipeng Li 0001, Songchen Han, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | SS-NeRF: Physically Based Sparse Spectral Rendering With Neural Radiance FieldabstractIn this paper, we propose SS-NeRF, the end-to-end Neural Radiance Field (NeRF)-based architectures for high-quality physically based rendering with sparse inputs. We modify the classical spectral rendering into two main steps, 1) the generation of a series of spectrum maps spanning different wavelengths, 2) the combination of these spectrum maps for the RGB output. The proposed architecture follows these two steps through the proposed multi-layer perceptron (MLP)-based architecture (SpectralMLP) and spectrum attention UNet (SAUNet). Given the ray origin and the ray direction, the SpectralMLP constructs the spectral radiance field to obtain spectrum maps of novel views, which are then sent to the SAUNet to produce RGB images of white-light illumination. Applying NeRF to build up the spectral rendering is a more physically-based way from the perspective of ray-tracing. Further, the spectral radiance fields decompose difficult scenes and improve the performance of NeRF-based methods. Previous baseline, such as SpectralNeRF, outperforms recent methods in synthesizing novel views but requires relatively dense viewpoints for accurate scene reconstruction. To tackle this, we propose SS-NeRF to enhance the detail of scene representation with sparse inputs. In SS-NeRF, we first design the depth-aware continuity to optimize the reconstruction based on single-view depth predictions. Then, the geometric-projected consistency is introduced to optimize the multi-view geometry alignment. Additionally, we introduce a superpixel-aligned consistency to ensure that the average color within each superpixel region remains consistent. Comprehensive experimental results demonstrate that the proposed method is superior to recent state-of-the-art methods when synthesizing new views on both synthetic and real-world datasets. Ru Li 0002, Guanghui Liu 0001, Shengping Zhang, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Learning Efficient Meshflow and Optical Flow From Event CamerasabstractIn this paper, we explore the problem of event-based meshflow estimation, a novel task that involves predicting a spatially smooth sparse motion field from event cameras. To start, we review the state-of-the-art in event-based flow estimation, highlighting two key areas for further research: i) the lack of meshflow-specific event datasets and methods, and ii) the underexplored challenge of event data density. First, we generate a large-scale High-Resolution Event Meshflow (HREM) dataset, which showcases its superiority by encompassing the merits of high resolution at 1280 × 720, handling dynamic objects and complex motion patterns, and offering both optical flow and meshflow labels. These aspects have not been fully explored in previous works. Besides, we propose Efficient Event-based MeshFlow (EEMFlow) network, a lightweight model featuring a specially crafted encoder-decoder architecture to facilitate swift and accurate meshflow estimation. Furthermore, we upgrade EEMFlow network to support dense event optical flow, in which a Confidence-induced Detail Completion (CDC) module is proposed to preserve sharp motion boundaries. We conduct comprehensive experiments to show the exceptional performance and runtime efficiency (30×faster) of our EEMFlow model compared to the recent state-of-the-art flow method. As an extension, we expand HREM into HREM+, a multi-density event dataset contributing to a thorough study of the robustness of existing methods across data with varying densities, and propose an Adaptive Density Module (ADM) to adjust the density of input event data to a more optimal range, enhancing the model's generalization ability. We empirically demonstrate that ADM helps to significantly improve the performance of EEMFlow and EEMFlow+ by 8% and 10%, respectively. Xinglong Luo, Ao Luo, Kunming Luo, Zhengning Wang, Ping Tan 0002, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Solving ILL-posed Regions in High Dynamic Range Reconstruction With Uncertainty-Aware Diffusion ModelsabstractLearning-based approaches have achieved promising progress in High Dynamic Range (HDR) image reconstruction, particularly in ghost removal. However, they often struggle in ill-posed regions, such as areas with occlusion or saturation, where insufficient or unreliable information leads to persistent residual ghosting artifacts and structural distortions. In this paper, we present UA-Diff, an uncertainty-aware diffusion framework designed to generate visually coherent, ghost-free HDR images. Specifically, our approach introduces an Uncertainty Generation Module (UGM) that estimates pixel-wise reconstruction confidence via a probabilistic Laplacian loss, producing an uncertainty map that explicitly highlights challenging ill-posed regions. To address these regions effectively, we develop an Uncertainty-Aware Diffusion Module (UADM) that operates selectively on the average-coefficient component of a 2D discrete wavelet transform, where dominant artifacts tend to concentrate. This enables reduced computational overhead while preserving high-quality details. Moreover, we propose an Uncertainty-Guided Sampling (UGS) strategy that leverages the uncertainty map to guide the denoising process, ensuring faithful reconstruction in reliable regions and targeted refinement in uncertain areas. Extensive experiments on three public HDR benchmarks demonstrate that UA-Diff surpasses state-of-the-art methods both quantitatively and perceptually, especially in challenging ill-posed scenarios. Zhen Liu 0022, Hai Jiang 0006, Haipeng Li 0001, Shuaicheng Liu, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | FCVSR: A Frequency-Aware Method for Compressed Video Super-ResolutionabstractCompressed video super-resolution (SR) aims to generate high-resolution (HR) videos from the corresponding low-resolution (LR) compressed videos. Recently, some compressed video SR methods attempt to exploit the spatio-temporal information in the frequency domain, showing great promise in super-resolution performance. However, these methods do not differentiate various frequency subbands spatially or capture the temporal frequency dynamics, potentially leading to suboptimal results. In this paper, we propose a deep frequency-based compressed video SR model (FCVSR) consisting of a motion-guided adaptive alignment (MGAA) network and a multi-frequency feature refinement (MFFR) module. Additionally, a frequency-aware contrastive loss is proposed for training FCVSR, in order to reconstruct finer spatial details. The proposed model has been evaluated on three public compressed video super-resolution datasets, with results demonstrating its effectiveness when compared to existing works in terms of super-resolution performance and complexity. Fan Zhang 0017, Feiyu Chen 0001, Shuyuan Zhu, David Bull 0001, Bing Zeng 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | FlowPolicy: Enabling Fast and Robust 3D Flow-Based Policy via Consistency Flow Matching for Robot ManipulationabstractRobots can acquire complex manipulation skills by learning policies from expert demonstrations, which is often known as vision-based imitation learning. Generating policies based on diffusion and flow matching models has been shown to be effective, particularly in robotic manipulation tasks. However, recursion-based approaches are inference inefficient in working from noise distributions to policy distributions, posing a challenging trade-off between efficiency and quality. This motivates us to propose FlowPolicy, a novel framework for fast policy generation based on consistency flow matching and 3D vision. Our approach refines the flow dynamics by normalizing the self-consistency of the velocity field, enabling the model to derive task execution policies in a single inference step. Specifically, FlowPolicy conditions on the observed 3D point cloud, where consistency flow matching directly defines straight-line flows from different time states to the same action space, while simultaneously constraining their velocity values, that is, we approximate the trajectories from noise to robot actions by normalizing the self-consistency of the velocity field within the action space, thus improving the inference efficiency. We validate the effectiveness of FlowPolicy in Adroit and Metaworld, demonstrating a 7× increase in inference speed while maintaining competitive average success rates compared to state-of-the-art methods. Qinglun Zhang, Zhen Liu 0022, Haoqiang Fan, Guanghui Liu 0001, Bing Zeng 0001, Shuaicheng Liu |
AAAI | 5 |
| 2025 | Estimating 2D Camera Motion with Hybrid Motion Basis
Haipeng Li 0001, Tianhao Zhou, Zhanglei Yang, Yan Chen 0007, Zijing Mao, Shen Cheng, Bing Zeng 0001, Shuaicheng Liu |
ICCV | 8 |
| 2025 | Blind Video Super-Resolution Based on Implicit KernelsabstractBlind video super-resolution (BVSR) is a low-level vision task which aims to generate high-resolution videos from low-resolution counterparts in unknown degradation scenarios. Existing approaches typically predict blur kernels that are spatially invariant in each video frame or even the entire video. These methods do not consider potential spatio-temporal varying degradations in videos, resulting in suboptimal BVSR performance. In this context, we propose a novel BVSR model based on Implicit Kernels, BVSR-IK, which constructs a multi-scale kernel dictionary parameterized by implicit neural representations. It also employs a newly designed recurrent Transformer to predict the coefficient weights for accurate filtering in both frame correction and feature alignment. Experimental results have demonstrated the effectiveness of the proposed BVSR-IK, when compared with four state-of-the-art BVSR models on three commonly used datasets, with BVSR-IK outperforming the second best approach, FMA-Net, by up to 0.59 dB in PSNR. Source code will be available at https://github.com/QZ1-boy/BVSR-IK. Yuxuan Jiang 0015, Shuyuan Zhu, Fan Zhang 0017, David Bull 0001, Bing Zeng 0001 |
ICCV | 6 |
| 2025 | Coding-Prior Guided Diffusion Network for Video DeblurringabstractWhile recent video deblurring methods have advanced significantly, they often overlook two valuable prior information: (1) motion vectors (MVs) and coding residuals (CRs) from video codecs, which provide efficient inter-frame alignment cues, and (2) the rich real-world knowledge embedded in pre-trained diffusion generative models. We present CPGD-Net, a novel two-stage framework that effectively leverages both coding priors and generative diffusion priors for high-quality deblurring. First, our coding-prior feature propagation (CPFP) module utilizes MVs for efficient frame alignment and CRs to generate attention masks, addressing motion inaccuracies and texture variations. Second, a coding-prior controlled generation (CPC) module network integrates coding priors into a pre-trained diffusion model, guiding it to enhance critical regions and synthesize realistic details. Experiments demonstrate our method achieves state-of-the-art perceptual quality with up to 30% improvement in IQA metrics. The code and the coding-prior-augmented dataset are available at: https://github.com/liuyike422/CPGD-Net. Haipeng Li 0001, Shuaicheng Liu, Bing Zeng 0001 |
ACM Multimedia | 5 |
| 2025 | Blind Image Super-Resolution with Local and Global Dual-GuidanceabstractBlind image super-resolution (BISR) aims to recover the high-resolution image from its degraded low-resolution version with unknown degradation. Recent research on BISR has demonstrated impressive results by using convolutional neural network (CNN) based techniques. However, these methods suffer from limited receptive fields brought by CNN. In addition, they have limited adaptivity to different frequency components of natural images. To address those drawbacks, we propose a deep local and global dual-guidance degradation-adaptive BISR network that exploits global information by involving Fourier coefficients in the degradation representation and image reconstruction process. Additionally, we propose a novel frequency component enhancement module that explicitly decomposes images into multiple frequency bands and assign different weights for each band so as to construct high-quality image. Experimental results demonstrate the superior performance of our method. Yajun Qiu, Shuyuan Zhu, Lantao Yu, Bing Zeng 0001 |
MMSP | 4 |
| 2025 | Task-Aware Optimized Color Image DemosaicingabstractIn this paper, we propose a task-aware deep demosaicing network that is designed to produce images for object detection and image compression, targeting high performance for both tasks. The proposed network consists of a color restoration module and a semantic enhancement module. Specifically, the color restoration module converts the Bayer-pattern raw images into full-color images. The semantic enhancement module integrates the semantic information via an adaptive feature fusion to enhance the task-relevant features while suppressing the task-irrelevant content to save bitrate. Experimental results demonstrate that using the color images produced by our demosaicing network can achieve a better trade-off between detection accuracy and compression efficiency. Feiyu Chen 0001, Shuyuan Zhu, Bing Zeng 0001 |
MMSP | 6 |
| 2025 | Unsupervised Global and Local Homography Estimation With Coplanarity-Aware GANabstractUnsupervised methods have received increasing attention in homography learning due to their promising performance and label-free training. However, existing methods do not explicitly consider the plane-induced parallax, making the prediction compromised on multiple planes. In this work, we propose a novel method HomoGAN to guide unsupervised homography estimation to focus on the dominant plane. First, a multi-scale transformer is designed to predict homography from the feature pyramids of input images in a coarse-to-fine fashion. Moreover, we propose an unsupervised GAN to impose coplanarity constraint on the predicted homography, which is realized by using a generator to predict a mask of aligned regions, and then a discriminator to check if two masked feature maps are induced by a single homography. Based on the global homography framework, we extend it to the local mesh-grid homography estimation, namely, MeshHomoGAN, where plane constraints can be enforced on each mesh cell to go beyond a single dominant plane, such that scenes with multiple depth planes can be better aligned. To validate the effectiveness of our method and its components, we conduct extensive experiments on large-scale datasets. Results show that our matching error is 22% lower than previous SOTA methods. Code is available at https://github.com/megvii-research/HomoGAN. Shuaicheng Liu, Mingbo Hong, Nianjin Ye, Chunyu Lin, Bing Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Minimum Latency Deep Online Video Stabilization and Its ExtensionsabstractWe present a novel deep camera path optimization framework for minimum latency online video stabilization. Typically, a stabilization pipeline consists of three steps: motion estimation, path smoothing, and novel view synthesis. Most previous methods concentrate on motion estimation while path optimization receives less attention, particularly in the crucial online setting where future frames are inaccessible. In this work, we adopt off-the-shelf high-quality deep motion models for motion estimation and focus only on the path optimization. Specifically, our camera path smoothing network takes a short 2D camera path in a sliding window as input and outputs the stabilizing warp field of the last frame, which warps the coming frame to its stabilized position. We explore three motion densities: a global single camera path, local mesh-based bundled paths, and dense flow paths. A hybrid loss and an efficient motion smoothing attention (EMSA) module are proposed for spatially and temporally consistent path smoothing. Moreover, we build a motion dataset that contains stable and unstable motion pairs for training. Extensive experiments demonstrate that our method surpasses state-of-the-art online stabilization methods and rivals the performance of offline methods, offering compelling advancements in the field of video stabilization. Shuaicheng Liu, Zhuofan Zhang, Zhen Liu 0022, Ping Tan 0002, Bing Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Kernel Reformulation With Deep Constrained Least Squares for Blind Image Super-ResolutionabstractThis work proposes to learn blind image super-resolution (SR) using deep constrained least squares deconvolution with low-resolution (LR) space kernels. Our method recovers the high-resolution (HR) image with a kernel estimation step and a kernel-based image restoration process. Specifically, we first reformulate the classical degradation model to transfer the deblurring kernel estimation into the LR space. We show that the LR space kernel has a closed-form solution given a pair of LR-HR images, which can be learned without ground truth kernels. Next, we introduce a dynamic deep linear filter module, which can generate deblurring kernel weights adaptively. Subsequently, the estimated kernel is integrated with a deep constrained least square filtering module to produce clean features. For reconstruction, we adopt a dual-path structured SR network that inputs both the deblurred feature and the original feature to suppress deconvolution artifacts. Finally, we learn discriminative features for deblurring and then restore the HR image in a single branch, producing a lighter weight network that can achieve comparable performance while only using 56% parameters and 60% inference time. Extensive experiments on both synthetic and real-world datasets demonstrate that our method achieves better accuracy and visual improvements against state-of-the-art approaches. Ziwei Luo 0002, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Multi-Frame Rolling Shutter Correction With Diffusion Models
Zhanglei Yang, Haipeng Li 0001, Shen Cheng, Mingbo Hong, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Projection Difference-Guided Geometry Quality Enhancement for Video-Based Point Cloud Compression
Yu Liu 0091, Jingwei Bao, Zeliang Li, Shuyuan Zhu, Jeff Siu-Kei Au-Yeung, Bing Zeng 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | DVSRNet: Deep Video Super-Resolution Based on Progressive Deformable Alignment and Temporal-Sparse EnhancementabstractVideo super-resolution (VSR) is used to compose high-resolution (HR) video from low-resolution video. Recently, the deformable alignment-based VSR methods are becoming increasingly popular. In these methods, the features extracted from video are aligned to eliminate the motion error targeting high super-resolution (SR) quality. However, these methods often suffer from misalignment and the lack of enough temporal information to compose HR frames, which accordingly induce artifacts in the SR result. In this article, we design a deep VSR network (DVSRNet) based on the proposed progressive deformable alignment (PDA) module and temporal-sparse enhancement (TSE) module. Specifically, the PDA module is designed to accurately align features and to eliminate artifacts via the bidirectional information propagation. The TSE module is constructed to further eliminate artifacts and to generate clear details for the HR frame. In addition, we construct a lightweight deep optical flow network (OFNet) to obtain the bidirectional optical flows for the implementation of the PDA module. Moreover, two new loss functions are designed for our proposed method. The first one is adopted in OFNet and the second one is constructed to guarantee the generation of sharp and clear details for the HR frames. The experimental results demonstrate that our method performs better than the state-of-the-art methods. Feiyu Chen 0001, Shuyuan Zhu, Yu Liu 0091, Ruiqin Xiong, Bing Zeng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | SpectralNeRF: Physically Based Spectral Rendering with Neural Radiance FieldabstractIn this paper, we propose SpectralNeRF, an end-to-end Neural Radiance Field (NeRF)-based architecture for high-quality physically based rendering from a novel spectral perspective. We modify the classical spectral rendering into two main steps, 1) the generation of a series of spectrum maps spanning different wavelengths, 2) the combination of these spectrum maps for the RGB output. Our SpectralNeRF follows these two steps through the proposed multi-layer perceptron (MLP)-based architecture (SpectralMLP) and Spectrum Attention UNet (SAUNet). Given the ray origin and the ray direction, the SpectralMLP constructs the spectral radiance field to obtain spectrum maps of novel views, which are then sent to the SAUNet to produce RGB images of white-light illumination. Applying NeRF to build up the spectral rendering is a more physically-based way from the perspective of ray-tracing. Further, the spectral radiance fields decompose difficult scenes and improve the performance of NeRF-based methods. Comprehensive experimental results demonstrate the proposed SpectralNeRF is superior to recent NeRF-based methods when synthesizing new views on synthetic and real datasets. The codes and datasets are available at https://github.com/liru0126/SpectralNeRF. Ru Li 0002, Guanghui Liu 0001, Shengping Zhang, Bing Zeng 0001, Shuaicheng Liu |
AAAI | 5 |
| 2024 | Efficient Meshflow and Optical Flow Estimation from Event CamerasabstractIn this paper, we explore the problem of event-based meshflow estimation, a novel task that involves predicting a spatially smooth sparse motion field from event cameras. To start, we generate a large-scale High-Resolution Event Meshflow (HREM) dataset, which showcases its superiority by encompassing the merits of high resolution at 1280×720, handling dynamic objects and complex motion patterns, and offering both optical flow and meshflow labels. These aspects have not been fully explored in previous works. Besides, we propose Efficient Event-based MeshFlow (EEMFlow) network, a lightweight model featuring a specially crafted encoder-decoder architecture to facilitate swift and accurate meshflow estimation. Furthermore, we upgrade EEMFlow network to support dense event optical flow, in which a Confidence-induced Detail Completion (CDC) module is proposed to preserve sharp motion boundaries. We conduct comprehensive experiments to show the exceptional performance and runtime efficiency (39× faster) of our EEMFlow model compared to recent state-of-the-art flow methods. Our code is available at https://github.com/boomluo02/EEMFlow. Xinglong Luo, Ao Luo, Zhengning Wang, Chunyu Lin, Bing Zeng 0001, Shuaicheng Liu |
CVPR | 5 |
| 2024 | RecDiffusion: Rectangling for Image Stitching with Diffusion ModelsabstractImage stitching from different captures often results in non-rectangular boundaries, which is often considered un-appealing. To solve non-rectangular boundaries, current solutions involve cropping, which discards image content, inpainting, which can introduce unrelated content, or warping, which can distort non-linear features and introduce artifacts. To overcome these issues, we introduce a novel diffusion-based learning framework, RecDiffusion, for image stitching rectangling. This framework combines Motion Diffusion Models (MDM) to generate motion fields, ef-fectively transitioning from the stitched image's irregular borders to a geometrically corrected intermediary. Fol-lowed by Content Diffusion Models (CDM) for image de-tail refinement. Notably, our sampling process utilizes a weighted map to identify regions needing correction during each iteration of CDM. Our RecDiffusion ensures geomet-ric accuracy and overall visual appeal, surpassing all pre-vious methods in both quantitative and qualitative measures when evaluated on public benchmarks. Code is released at https://github.com/haippp/RecDiffusion. Tianhao Zhou, Haipeng Li 0001, Ao Luo, Chen-Lin Zhang, Bing Zeng 0001, Shuaicheng Liu |
CVPR | 7 |
| 2024 | Filamentary Convolution for Spoken Language Identification: A Brain-Inspired ApproachabstractSpoken language identification (SLI) by human beings relies on the hierarchical understanding of one or a few words within the voice signal, encapsulated within the corresponding time windows. Concurrently, frequency-domain features play a crucial role in enhancing identification. The short-time Fourier transform (STFT) has conventionally served as a pivotal component in the forefront of most SLI systems, including deep-learning networks (DLNs). Nevertheless, the use of rectangle-shaped masks in STFT introduces spectral component mixing across different time windows, potentially resulting in an aliasing effect. To address this limitation, we propose a novel filamentary convolution framework to replace the conventional rectangle-shaped convolutions. This framework not only reduces complexity but also enhances feature learning within each frame. Leveraging filamentary convolution, we formulate an encoding module with a non-overlapping strategy and a multi-level information extraction (MIE) module featuring unbalanced dual-route convolution (UDRC) blocks. The frequency features learned from filamentary convolutions are seamlessly integrated through a long-short term memory (LSTM) structure. In summary, our decision-making process employs the filamentary convolution kernel-based hierarchical neural network (FCK-NN), comprising an encoding module, MIE module, and LSTM module. We conduct experiments on a novel dataset encompassing 44 languages, curated by ourselves, and the results validate that our FCK-NN yields a significant improvement in performance. Shuyuan Zhu, Tong Xie, Xibang Yang, Bing Zeng 0001 |
ICASSP | 6 |
| 2024 | GyroFlow+: Gyroscope-Guided Unsupervised Deep Homography and Optical Flow Learning
Haipeng Li 0001, Kunming Luo, Bing Zeng 0001, Shuaicheng Liu |
Int. J. Comput. Vis. | 3 |
| 2024 | GLOCAL: A self-supervised learning framework for global and local motion estimation
Yihao Zheng 0002, Kunming Luo, Shuaicheng Liu, Zun Li 0001, Ye Xiang, Lifang Wu, Bing Zeng 0001, Chang Wen Chen |
Pattern Recognit. Lett. | 7 |
| 2024 | Compressed Video Quality Enhancement With Temporal Group Alignment and FusionabstractIn this paper, we propose a temporal group alignment and fusion network to enhance the quality of compressed videos by using the long-short term correlations between frames. The proposed model consists of the intra-group feature alignment (IntraGFA) module, the inter-group feature fusion (InterGFF) module, and the feature enhancement (FE) module. We form the group of pictures (GoP) by selecting frames from the video according to their temporal distances to the target enhanced frame. With this grouping, the composed GoP can contain either long- or short-term correlated information of neighboring frames. We design the IntraGFA module to align the features of frames of each GoP to eliminate the motion existing between frames. We construct the InterGFF module to fuse features belonging to different GoPs and finally enhance the fused features with the FE module to generate high-quality video frames. The experimental results show that our proposed method achieves up to 0.05 dB gain and lower complexity compared to the state-of-the-art method. Yajun Qiu, Yu Liu 0091, Shuyuan Zhu, Bing Zeng 0001 |
IEEE Signal Process. Lett. | 5 |
| 2024 | PBR-GAN: Imitating Physically-Based Rendering With Generative Adversarial NetworksabstractWe propose a Generative Adversarial Network (GAN)-based architecture for achieving high-quality physically based rendering (PBR). Conventional PBR relies heavily on ray tracing, which is computationally expensive in complicated environments. Some recent deep learning-based methods can improve efficiency but cannot deal with illumination variation well. In this paper, we propose PBR-GAN, an end-to-end GAN-based network that solves these problems while generating natural photo-realistic images. Two encoders (the shading encoder and albedo encoder) and two decoders (the image decoder and light decoder) are introduced to achieve our target. The two encoders and the image decoder constitute the generator that learns the mapping between the generated domain and the real domain. The light decoder produces light maps that pay more attention to the highlight and shadow regions. The discriminator aims to optimize the generator by distinguishing target images from the generated ones. Three novel loss items, concentrating on domain translation, overall shading preservation, and light map estimation, are proposed to optimize the photo-realistic outputs. Furthermore, a real dataset is collected to provide realistic information for training GAN architecture. Extensive experiments indicate that PBR-GAN can preserve the illumination variation and improve the image perceptual quality. Ru Li 0002, Peng Dai 0003, Guanghui Liu 0001, Shengping Zhang, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | CodingHomo: Bootstrapping Deep Homography With Video CodingabstractHomography estimation is a fundamental task in computer vision with applications in diverse fields. Recent advances in deep learning have improved homography estimation, particularly with unsupervised learning approaches, offering increased robustness and generalizability. However, accurately predicting homography, especially in complex motions, remains a challenge. In response, this work introduces a novel method leveraging video coding, particularly by harnessing inherent motion vectors (MVs) present in videos. We present CodingHomo, an unsupervised framework for homography estimation. Our framework features a Mask-Guided Fusion (MGF) module that identifies and utilizes beneficial features among the MVs, thereby enhancing the accuracy of homography prediction. Additionally, the Mask-Guided Homography Estimation (MGHE) module is presented for eliminating undesired features in the coarse-to-fine homography refinement process. CodingHomo outperforms existing state-of-the-art unsupervised methods, delivering good robustness and generalizability. The code and dataset are available at:https://github.com/liuyike422/CodingHomo. Haipeng Li 0001, Shuaicheng Liu, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Dual Circle Contrastive Learning-Based Blind Image Super-ResolutionabstractBlind image super-resolution (BISR) aims to construct high-resolution image from low-resolution (LR) image that contains unknown degradation. Although the previous methods demonstrated impressive performance by introducing the degradation representation in BISR task, there still exist two problems in most of them. First, they ignore the degradation characteristics of different image regions when generating degradation representation. Second, they lack effective supervision on the generation of both degradation representation and super-resolution result. To solve these problems, we propose the dual circle contrastive learning (DCCL) with the high-efficiency modules to implement BISR. In our proposed method, we design the degradation extraction network to obtain the degradation representations from different texture regions of LR image. Meanwhile, we propose DCCL coupled with the degrading network to guarantee the obtained degradation representation to contain the degradation of LR image as much as possible. The application of DCCL also makes the SR results contain degradation as little as possible. Additionally, we develop an information distillation module for our proposed BISR model to guarantee the SR images with high quality. The experimental results demonstrate that our proposed method achieves the state-of-the-art BISR performance. Yajun Qiu, Shuyuan Zhu, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Single-Image-Based Deep Learning for Segmentation of Early Esophageal Cancer LesionsabstractAccurate segmentation of lesions is crucial for diagnosis and treatment of early esophageal cancer (EEC). However, neither traditional nor deep learning-based methods up to today can meet the clinical requirements, with the mean Dice score - the most important metric in medical image analysis - hardly exceeding 0.75. In this paper, we present a novel deep learning approach for segmenting EEC lesions. Our method stands out for its uniqueness, as it relies solely on a single input image from a patient, forming the so-called "You-Only-Have-One" (YOHO) framework. On one hand, this "one-image-one-network" learning ensures complete patient privacy as it does not use any images from other patients as the training data. On the other hand, it avoids nearly all generalization-related problems since each trained network is applied only to the same input image itself. In particular, we can push the training to "over-fitting" as much as possible to increase the segmentation accuracy. Our technical details include an interaction with clinical doctors to utilize their expertise, a geometry-based data augmentation over a single lesion image to generate the training dataset (the biggest novelty), and an edge-enhanced UNet. We have evaluated YOHO over an EEC dataset collected by ourselves and achieved a mean Dice score of 0.888, which is much higher as compared to the existing deep-learning methods, thus representing a significant advance toward clinical applications. The code and dataset are available at: https://github.com/lhaippp/YOHO. Haipeng Li 0001, Dingrui Liu, Shuaicheng Liu, Tao Gan, Nini Rao, Jinlin Yang, Bing Zeng 0001 |
IEEE Trans. Image Process. | 8 |
| 2024 | Depth-Guided Deep Video InpaintingabstractVideo inpainting aims to fill in missing regions of a video after any undesired contents are removed from it. This technique can be applied to repair the broken video or edit the video content. In this paper, we propose a depth-guided deep video inpainting network (DGDVI) and demonstrate its effectiveness in processing challenging broken areas crossing multiple depth layers. To achieve our goal, we divide the inpainting into depth completion, content reconstruction, and content enhancement. Three corresponding modules are designed to implement a process-flow. Firstly, we develop a depth completion module based upon the spatio-temporal Transformer which is used to obtain the completed depth information for each video frame. Secondly, we design a content reconstruction module to generate initially inpainted video. With this module, the contents of the missing regions are composed via the depth-guided feature propagation. Thirdly, we construct a content enhancement module to enhance the temporal coherence and texture quality for the inpainted video. All of proposed modules are jointly optimized to guarantee the high inpainting efficiency. The experimental results demonstrate that our proposed method provides better inpainting results, both qualitatively and quantitatively, compared with the previous state-of-the-art. Shuyuan Zhu, Yao Ge 0002, Bing Zeng 0001, Muhammad Ali Imran 0001, Qammer H. Abbasi, Jonathan M. Cooper |
IEEE Trans. Multim. | 4 |
| 2024 | DMHomo: Learning Homography with Diffusion ModelsabstractSupervised homography estimation methods face a challenge due to the lack of adequate labeled training data. To address this issue, we propose DMHomo , a diffusion model-based framework for supervised homography learning. This framework generates image pairs with accurate labels, realistic image content, and realistic interval motion, ensuring that they satisfy adequate pairs. We utilize unlabeled image pairs with pseudo labels such as homography and dominant plane masks, computed from existing methods, to train a diffusion model that generates a supervised training dataset. To further enhance performance, we introduce a new probabilistic mask loss, which identifies outlier regions through supervised training, and an iterative mechanism to optimize the generative and homography models successively. Our experimental results demonstrate that DMHomo effectively overcomes the scarcity of qualified datasets in supervised homography learning and improves generalization to real-world scenes. The code and dataset are available at GitHub ( https://github.com/lhaippp/DMHomo ). Haipeng Li 0001, Hai Jiang 0006, Ao Luo, Ping Tan 0002, Haoqiang Fan, Bing Zeng 0001, Shuaicheng Liu |
ACM Trans. Graph. | 6 |
| 2023 | Supervised Homography Learning with Realistic Dataset GenerationabstractIn this paper, we propose an iterative framework, which consists of two phases: a generation phase and a training phase, to generate realistic training data and yield a supervised homography network. In the generation phase, given an unlabeled image pair, we utilize the pre-estimated dominant plane masks and homography of the pair, along with another sampled homography that serves as ground truth to generate a new labeled training pair with realistic motion. In the training phase, the generated data is used to train the supervised homography network, in which the training data is refined via a content consistency module and a quality assessment module. Once an iteration is finished, the trained network is used in the next data generation phase to update the pre-estimated homography. Through such an iterative strategy, the quality of the dataset and the performance of the network can be gradually and simultaneously improved. Experimental results show that our method achieves state-of-the-art performance and existing supervised methods can be also improved based on the generated dataset. Code and dataset are available at https://github.com/JianghaiSCU/RealSH. Hai Jiang 0006, Haipeng Li 0001, Songchen Han, Haoqiang Fan, Bing Zeng 0001, Shuaicheng Liu |
ICCV | 5 |
| 2023 | Deep Homography Mixture for Single Image Rolling Shutter CorrectionabstractWe present a deep homography mixture motion model for single image rolling shutter correction. Rolling shutter (RS) effects are often caused by row-wise exposure delay in the widely adopted CMOS sensor. Previous methods often require more than one frame for the correction, leading to data quality requirements. Few approaches address the more challenging task of single image RS correction, which often adopt designs like trajectory estimation or long rectangular kernels, to learn the camera motion parameters of an RS image, to restore the global shutter (GS) image. In this work, we adopt a more straightforward method to learn deep homography mixture motion between an RS image and its corresponding GS image, without large solution space or strict restrictions on image features. We show that dividing an image into blocks with a Gaussian weight of block scanlines fits well for the RS setting. Moreover, instead of directly learning the motion mapping, we learn coefficients that assemble several motion bases to produce the correction motion, where these bases are learned from the consecutive frames of natural videos beforehand. Experiments show that our method outperforms existing single RS methods statistically and visually, in both synthesized and real RS images. Our code and dataset are available at https://github.com/DavidYan2001/Deep_HM. Weilong Yan, Robby T. Tan, Bing Zeng 0001, Shuaicheng Liu |
ICCV | 3 |
| 2023 | Minimum Latency Deep Online Video StabilizationabstractWe present a novel camera path optimization framework for the task of online video stabilization. Typically, a stabilization pipeline consists of three steps: motion estimating, path smoothing, and novel view rendering. Most previous methods concentrate on motion estimation, proposing various global or local motion models. In contrast, path optimization receives relatively less attention, especially in the important online setting, where no future frames are available. In this work, we adopt recent off-the-shelf high-quality deep motion models for motion estimation to recover the camera trajectory and focus on the latter two steps. Our network takes a short 2D camera path in a sliding window as input and outputs the stabilizing warp field of the last frame in the window, which warps the coming frame to its stabilized position. A hybrid loss is well-defined to constrain the spatial and temporal consistency. In addition, we build a motion dataset that contains stable and unstable motion pairs for the training. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art online methods both qualitatively and quantitatively and achieves comparable performance to offline methods. Our code and dataset are available at https://github.com/liuzhen03/NNDVS. Zhuofan Zhang, Zhen Liu 0022, Ping Tan 0002, Bing Zeng 0001, Shuaicheng Liu |
ICCV | 4 |
| 2023 | Learned Image Compression with Large Capacity and Low Redundancy of Latent RepresentationabstractLearned image compression has attracted a lot of attention in recent years. Currently, popular learned image compression methods usually exploit hyperprior and autoregressive models to facilitate probability estimation and reduce the redundancy of latent representation. These models ignore different image contents, and it is difficult to eliminate the spatial redundancy of image, resulting in the performance saturation. In this work, we propose a learned image compression method with large capacity and low redundancy of latent representation. We design two enhancement modules, i.e., the network capacity expansion module (NCEM) and the high-entropy content guided reconstruction module (HCGR), to construct network architectures with better rate-distortion performance than the existing hyperprior and autoregressive models. Experimental results show that our method can produce superior results compared to the state-of-the-art methods. Xiandong Meng, Shuyuan Zhu, Siwei Ma 0001, Bing Zeng 0001 |
ICIP | 4 |
| 2023 | Dynamic updating self-training for semi-weakly supervised object detection
Shuaicheng Liu, Bing Zeng 0001 |
Neurocomputing | 3 |
| 2023 | Unsupervised Global and Local Homography Estimation With Motion Basis LearningabstractIn this paper, we introduce a new framework for unsupervised deep homography estimation. Our contributions are 3 folds. First, unlike previous methods that regress 4 offsets for a homography, we propose a homography flow representation, which can be estimated by a weighted sum of 8 pre-defined homography flow bases. Second, considering a homography contains 8 Degree-of-Freedoms (DOFs) that is much less than the rank of the network features, we propose a Low Rank Representation (LRR) block that reduces the feature rank, so that features corresponding to the dominant motions are retained while others are rejected. Last, we propose a Feature Identity Loss (FIL) to enforce the learned image feature warp-equivariant, meaning that the result should be identical if the order of warp operation and feature extraction is swapped. With this constraint, the unsupervised optimization can be more effective and the learned features are more stable. With global-to-local homography flow refinement, we also naturally generalize the proposed method to local mesh-grid homography estimation, which can go beyond the constraint of a single homography. Extensive experiments are conducted to demonstrate the effectiveness of all the newly proposed components, and results show that our approach outperforms the state-of-the-art on the homography benchmark dataset both qualitatively and quantitatively. Code is available at https://github.com/megvii-research/BasesHomo. Shuaicheng Liu, Hai Jiang 0006, Nianjin Ye, Chuan Wang 0001, Bing Zeng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Quality-Constrained Encoding Optimization for Omnidirectional Video StreamingabstractOmnidirectional video streaming is usually implemented based on the representations of tiles, where the tiles are obtained by splitting the video frame into several rectangular areas and each tile is converted into multiple representations with different resolutions and encoded at different bitrates. One key issue in omnidirectional video streaming is how to choose the optimal representations for each tile at the server to save the overall transmission bitrate to all users while offering them satisfactory quality. This is different from the adaptive bitrate-based method that optimizes the downloading procedure of individual users, where the given video representations are stored on the server. In this work, we focus on optimization for the encoding of omnidirectional video streaming by using the optimal combination of tile representations. To achieve our goal, we formulate the selection of the representations into an optimization problem in which the transmission bitrate of all the representations is minimized with a quality constraint. By using this constraint, we can improve the average quality of omnidirectional videos for users. More specifically, we first construct the tile-level rate-distortion (R-D) model and determine the available tile bandwidth based on the previous viewers’ statistics. Then, we formulate the representation selection problem based on the obtained R-D model and tile bandwidth. Finally, we solve this problem to obtain the optimal combination of tile representations so that we can transmit the omnidirectional video to users with satisfactory quality but low bitrate. The experimental results demonstrate the effectiveness of our proposed approach when it is applied to omnidirectional video streaming. Chaofan He, Roberto Gerson De Albuquerque Azevedo, Shuyuan Zhu, Bing Zeng 0001, Pascal Frossard |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | NOMA-Based Uncoded Video Transmission With Optimization of Joint Resource AllocationabstractThe non-orthogonal multiple access (NOMA) technique has demonstrated potential for the multicast of multiple videos. However, it simply multiplexes limited number of signals in each single channel and cannot satisfy different video quality requirements of multiple users. To resolve this problem, we construct the NOMA-based uncoded multi-user video transmission (NOMA-UMVT) system in which the allocation of power and channel resources to all the users is jointly optimized to guarantee high video quality. Specifically, we first implement the power allocation of multiuser by converting it into the inter-channel and intra-channel allocation sub-problems. After solving these sub-problems for power allocation, we assign channels with the proposed two-staged channel assignment algorithm. The simulation results demonstrate the superior performance of our proposed NOMA-UMVT system when it is applied to transmit videos to multiple users. Chaofan He, Shuyuan Zhu, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | DFCE: Decoder-Friendly Chrominance Enhancement for HEVC Intra CodingabstractWe propose a decoder-friendly chrominance enhancement method for the compressed images. Our proposed method is developed based on the luminance-guided chrominance enhancement network (LGCEN) and online learning. With LGCEN, the textures of the compressed chrominance components are enhanced by the guidance of luminance component. Moreover, LGCEN is constructed with the recursive design and the light-weight channel attention mechanism to achieve high performance as well as low complexity. It is arranged at both encoder and decoder sides. Given the input image, we train LGCEN at encoder side by using online learning. With online learning, we partially update network parameters and transmit them to decoder to update LGCEN arranged there. The adoption of online learning effectively reduces the workload of decoder and guarantee high robustness. Compared with the state-of-the-art methods, our proposed approach achieves superior performance. Renwei Yang, Hewei Liu, Shuyuan Zhu, Xiaozhen Zheng, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Exploring Fine Polarimetric Decomposition Technique for Built-Up Area MonitoringabstractHighly variable polarimetric signatures caused by complex structures in built-up areas make interpretation of these scattering behaviors intractable for PolSAR remote sensing. This paper proposes a fine polarimetric decomposition method and derives several products to finely simulate the scattering mechanisms of urban buildings, thus fulfilling its use for effective surveillance. First, through theoretically establishing the roll-invariant condition for a completely general scatterer, a roll-invariant cross polarization (RICP) scattering model is constructed, which characterizes the cross polarization scattering in the manner of planar structure distribution. Second, by designing a root-discriminant-based parameter inversion strategy, a fine seven-component decomposition is proposed, which achieves the complete physical interpretation of matrix elements and reasonable inversion of model parameters. Third, by analyzing the external and internal scattering difference, the derivative products, i.e., scattering contribution synthesizers are derived for built-up area monitoring. Experimental results derived from real PolSAR data confirm the superiority and effectiveness of the constructed descriptors on the one hand. On the other hand, the extensibility of fine polarimetric decomposition in specific remote sensing is also explicitly demonstrated. Sinong Quan, Tao Zhang 0027, Wei Wang 0099, Gangyao Kuang, Xuesong Wang 0003, Bing Zeng 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Stereo RGB and Deeper LIDAR-Based Network for 3D Object Detection in Autonomous Drivingabstract3D object detection has become an emerging task in autonomous driving scenarios. Most of previous works process 3D point clouds using either projection-based or voxel-based models. However, both approaches contain some drawbacks. The voxel-based methods lack semantic information, while the projection-based methods suffer from numerous spatial information loss when projected to different views. In this paper, we propose the Stereo RGB and Deeper LIDAR (SRDL) framework which can utilize semantic and spatial information simultaneously such that the performance of network for 3D object detection can be improved naturally. Specifically, the network generates candidate boxes from stereo pairs and combines different region-wise features using a deep fusion scheme. The stereo strategy offers more information for prediction compared with prior works. Then, several local and global feature extractors are stacked in the segmentation module to capture richer deep semantic geometric features from point clouds. After aligning the interior points with fused features, the proposed network refines the prediction in a more accurate manner and encodes the whole box in a novel compact method. The decent experimental results on the challenging KITTI detection benchmark demonstrate the effectiveness of utilizing both stereo images and point clouds for 3D object detection. Qingdong He, Zhengning Wang, Yijun Liu 0012, Shuaicheng Liu, Bing Zeng 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2023 | Learning Fashion Compatibility With Context Conditioning EmbeddingabstractFashion compatibility predictions have obtained a lot of attention recently. Mining the compatibility between fashion items in an outfit is different from learning the visual similarity, since this relationship is more delicate. Decomposing the outfit compatibility into pairwise item matching is a popular way to treat the problem. However, in most existing methods, the items are matched without considering the context, i.e, the remaining items in the outfit. Recent efforts have been made to learn the underlying high order relationships among items by treating the outfit as a whole. These models could be sensitive to the properties of different datasets, and the item representations in these models are not as compact as those in the pairwise models. In this paper, we propose a context conditioning embedding approach to learn compact representations that preserve the shared information among items under the existence of contextual items. We use two different spaces, the general and the contextual spaces, to embed items, where the representation in the contextual space contains information from the context. We employ mutual information maximization for model learning, which is shown to be more appropriate for the problem. With extensive experiments, we show that our model achieves superior performance than other state-of-the-art methods. Yang Hu 0006, Cong Yu 0011, Yan Chen 0007, Bing Zeng 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Personalized Fashion Recommendation With Discrete Content-Based Tensor FactorizationabstractFashion outfit recommendation has attracted lots of attention recently. The problem becomes even more interesting and challenging when considering users’ personalized fashion preferences. Although existing works have successfully improved the recommendation accuracy, the efficiency issue of computation and storage is still under-investigated and often ignored. In this paper, we propose a discrete content-based tensor factorization model that maps items and user to binary codes for efficient fashion recommendation. We introduce a probabilistic perspective for learning to hash, where the binary codes are sampled from a set of underlying Bernoulli variables. To demonstrate the effectiveness of our model, we collect a large-scale outfit dataset together with user label information from a fashion-focused social website. Extensive experiments on our dataset show that the proposed model outperforms other state-of-the-art methods. Yang Hu 0006, Cong Yu 0011, Yunchao Jiang, Yan Chen 0007, Bing Zeng 0001 |
IEEE Trans. Multim. | 6 |
| 2022 | FINet: Dual Branches Feature Interaction for Partial-to-Partial Point Cloud RegistrationabstractData association is important in the point cloud registration. In this work, we propose to solve the partial-to-partial registration from a new perspective, by introducing multi-level feature interactions between the source and the reference clouds at the feature extraction stage, such that the registration can be realized without the attentions or explicit mask estimation for the overlapping detection as adopted previously. Specifically, we present FINet, a feature interactionbased structure with the capability to enable and strengthen the information associating between the inputs at multiple stages. To achieve this, we first split the features into two components, one for rotation and one for translation, based on the fact that they belong to different solution spaces, yielding a dual branches structure. Second, we insert several interaction modules at the feature extractor for the data association. Third, we propose a transformation sensitivity loss to obtain rotation-attentive and translation-attentive features. Experiments demonstrate that our method performs higher precision and robustness compared to the state-of-the-art traditional and learning-based methods. Code is available at https://github.com/megvii-research/FINet. Hao Xu 0018, Nianjin Ye, Guanghui Liu 0001, Bing Zeng 0001, Shuaicheng Liu |
AAAI | 4 |
| 2022 | Learning to Zoom Inside Camera Imaging PipelineabstractExisting single image super-resolution methods are either designed for synthetic data, or for real data but in the RGB-to-RGB or the RAW-to-RGB domain. This paper proposes to zoom an image from RAW to RAW inside the camera imaging pipeline. The RAW-to-RAW domain closes the gap between the ideal and the real degradation models. It also excludes the image signal processing pipeline, which refocuses the model learning onto the super-resolution. To these ends, we design a method that receives a low-resolution RAW as the input and estimates the desired higher-resolution RAW jointly with the degradation model. In our method, two convolutional neural networks are learned to constrain the high-resolution image and the degradation model in lower-dimensional subspaces. This subspace constraint converts the ill-posed SISR problem to a well-posed one. To demonstrate the superiority of the proposed method and the RAW-to-RAW domain, we conduct evaluations on the RealSR and the SR-RAW datasets. The results show that our method performs superiorly over the state-of-the-arts both qualitatively and quantitatively, and it also generalizes well and enables zero-shot transfer across different sensors. Chengzhou Tang, Yuqiang Yang, Bing Zeng 0001, Ping Tan 0002, Shuaicheng Liu |
CVPR | 3 |
| 2022 | Ghost-free High Dynamic Range Imaging with Context-Aware Transformer
Zhen Liu 0022, Yinglong Wang 0002, Bing Zeng 0001, Shuaicheng Liu |
ECCV (19) | 3 |
| 2022 | Hierarchical Coding for Talking-Head VideoabstractTalking-head video is very popular in video conference and social media, where the camera captures the movement of user’s head and the change of facial expression. In this paper, we propose a hierarchical coding scheme for the compression of talking-head video. In our proposed method, three data layers, including one base layer, one enhancement layer and one feature layer, are formed as the input of encoder. More specifically, the base layer is generated by spatially sub-sampling the source video. The enhancement layer is composed by the specific key frames and the feature layer is produced based on the extracted facial landmarks. These layers are separately compressed but fused together to reconstruct the video signal in the decoder side. To achieve a high-quality reconstruction, we design the multi-feature fusion network in which the feature layer is used to guide the fusion of base layer and enhancement layer. The experiment results demonstrate the good performance of our proposed method for the coding of talking-head video. Yu Liu 0091, Shuyuan Zhu, Jeff Siu-Kei Au-Yeung, Bing Zeng 0001 |
ISCAS | 6 |
| 2022 | Luminance-Guided Chrominance Image Enhancement for HEVC Intra CodingabstractIn this paper, we propose a luminance-guided chrominance image enhancement convolutional neural network for HEVC intra coding. Specifically, we firstly develop a gated recursive asymmetric-convolution block to restore each degraded chrominance image, which generates an intermediate output. Then, guided by the luminance image, the quality of this intermediate output is further improved, which finally produces the high-quality chrominance image. When our proposed method is adopted in the compression of color images with HEVC intra coding, it achieves 28.96% and 16.74% BD-rate gains over HEVC for the U and V images, respectively, which accordingly demonstrate its superiority. The code is available online: https://github.com/Nickyang4900/Luminance-Guided-Chrominance-Enhancement-for-HEVC-Intra-Coding. Hewei Liu, Renwei Yang, Shuyuan Zhu, Bing Zeng 0001 |
ISCAS | 5 |
| 2022 | Deep Feature Compression with Collaborative Coding of Image TextureabstractIn this paper, we propose a coding scheme for the deep intermediate feature and it is implemented with the collaborative compression of image texture. More specifically, we separately compress the feature and texture of the image to form two data layers. The first one is the intermediate feature layer and the second one is the texture layer. The texture layer can provide an image for users and the feature layer can be used to implement the computer vision (CV) task. With our proposed deep reconstruction network (RecNet), the texture and features cooperate to achieve a high-quality visual output as well as a high-efficiency CV task. The experimental results demonstrate the excellent performance by using our proposed method to compress the deep features. Hewei Liu, Shuyuan Zhu, Xiaozhen Zheng, Ruiqin Xiong, Bing Zeng 0001 |
ISCAS | 6 |
| 2022 | Deep Video Super-Resolution with Flow-Guided Deformable Alignment and Sparsity-based Temporal-Spatial EnhancementabstractVideo super resolution (VSR) aims to construct the high resolution (HR) frames from the low resolution ones. In this paper, we propose a new deep VSR network based on the flow-guided deformable alignment (FGDA) module and the sparsity-based temporal-spatial enhancement (STSE) module. More specifically, the FGDA module is designed to generate temporally-aligned features with the bidirectional propagation. Meanwhile, the STSE module is constructed to eliminate the alignment error for the features and strengthen the sparsity of them to enhance the spatial details to construct the high-quality HR result. In addition, we design a sparsity-based loss function to guarantee the germinated HR frames with sharp details. The experimental results demonstrate that our proposed method achieves superior performance compared with the existing popular methods. Shuyuan Zhu, Guanghui Liu 0001, Bing Zeng 0001, Xiaozhen Zheng |
MMSP | 5 |
| 2022 | Inter-Frame Dependency-Based Rate Control for VVC Low-Delay CodingabstractIn this letter, we propose two solutions for the rate control of the VVC low-delay coding. Both solutions are developed by determining the bit allocation factors for video frames based on their dependency. Specifically, we design the first solution according to the distortion correlation between the key-frame and its subsequent frames. With this solution, the bit allocation factors are determined by applying multi-pass coding on the video to build up the cross-frame distortion model. This model offers us a straightforward way to achieve better rate control performance but results in a rather high complexity. To solve the complexity problem, we propose the second solution based on the difference between frames. In this solution, we construct the bit allocation model and apply it to frames so that we can adaptively determine the allocation factors with a low complexity. The experimental results demonstrate that our proposed two solutions can offer better rate-distortion performances than the state-of-the-art method. Hewei Liu, Shuyuan Zhu, Bing Zeng 0001 |
IEEE Signal Process. Lett. | 3 |
| 2022 | A Fast CABAC Hardware Design for Accelerating the Rate Estimation in HEVCabstractThe latest High Efficiency Video Coding standard achieves twice the coding efficiency of the H264 standard through a complex rate-distortion optimization (RDO). The coded bit-streams are produced with context adaptive binary arithmetic coding (CABAC). CABAC itself is a very time-consuming process that includes binarization, context modeling, interval subdivision, renormalization, outstanding bit handling, and context updating. The aim of this research is to speed up the CABAC process through several simplifications. First, we approximate three parts of the CABAC, i.e., interval subdivision, renormalization, and outstanding bit handling, with a piecewise-linear function that is very friendly to hardware implementation. In order to achieve better hardware parallelism, we also improve the coding process at the sub-block level. The context of syntax elements in a sub-block is redistributed to skip the complex calculation of context indexing. We perform context updating at the granularity of sub-blocks so that the data dependency of the context updating is removed completely, and the original serial encoding process is changed to a parallel encoding process. At the same time, we make another simplification for the context modeling ofcu_skip_flag. Based on these simplifications, we build a parallel hardware architecture for the rate estimation of the RDO process. This architecture completes the bit estimation of a$32\times 32$coding tree unit (CTU) in 220.8 nano-seconds, whereas the Bjøntegaard Delta rate increases by only 2.225%. We believe that the proposed architecture can meet the requirements of 8K@120 fps ultra-high-definition videos. This is the first study to simplify the hardware design of rate estimation by changing the context allocation and updating rules. Yujie Cai, Yibo Fan, Leilei Huang, Xiaoyang Zeng, Haibing Yin, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | UPHDR-GAN: Generative Adversarial Network for High Dynamic Range Imaging With Unpaired DataabstractThe paper proposes a method to effectively fuse multi-exposure inputs and generate high-quality high dynamic range (HDR) images with unpaired datasets. Deep learning-based HDR image generation methods rely heavily on paired datasets. The ground truth images play a leading role in generating reasonable HDR images. Datasets without ground truth are hard to be applied to train deep neural networks. Recently, Generative Adversarial Networks (GAN) have demonstrated their potentials of translating images from source domain$X$to target domain$Y$in the absence of paired examples. In this paper, we propose a GAN-based network for solving such problems while generating enjoyable HDR results, named UPHDR-GAN. The proposed method relaxes the constraint of the paired dataset and learns the mapping from the LDR domain to the HDR domain. Although the pair data are missing, UPHDR-GAN can properly handle the ghosting artifacts caused by moving objects or misalignments with the help of the modified GAN loss, the improved discriminator network and the useful initialization phase. The proposed method preserves the details of important regions and improves the total image perceptual quality. Qualitative and quantitative comparisons against the representative methods demonstrate the superiority of the proposed UPHDR-GAN. Ru Li 0002, Chuan Wang 0001, Jue Wang 0001, Guanghui Liu 0001, Heng-Yu Zhang, Bing Zeng 0001, Shuaicheng Liu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | ASFlow: Unsupervised Optical Flow Learning With Adaptive Pyramid SamplingabstractWe present an unsupervised optical flow estimation method by proposing an adaptive pyramid sampling in the deep pyramid network. Specifically, in the pyramid downsampling, we propose a Content-Aware Pooling (CAP) module, which promotes local feature gathering by avoiding cross region pooling, so that the learned features become more representative. In the pyramid upsampling, we propose an Adaptive Flow Upsampling (AFU) module, where cross edge interpolation can be avoided, producing sharp motion boundaries. Equipped with these two modules, our method achieves the best performance for unsupervised optical flow estimation on multiple leading benchmarks, including MPI-Sintel, KITTI 2012 and KITTI 2015. Particularly, we achieve EPE=1.5 on KITTI 2012 and F1=9.67% KITTI 2015, which outperform the previous state-of-the-art methods by 16.7% and 13.1%, respectively. Shuaicheng Liu, Kunming Luo, Ao Luo, Chuan Wang 0001, Fanman Meng, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | DeepOIS: Gyroscope-Guided Deep Optical Image Stabilizer CompensationabstractMobile captured images can be aligned using their gyroscope sensors. Optical image stabilizer (OIS) terminates this possibility by adjusting the images during the capturing. In this work, we propose a deep network that compensates for the motions caused by the OIS, such that the gyroscopes can be used for image alignment on the OIS cameras. To achieve this, we first record both videos and gyroscope readings with an OIS camera as training data. Then, we convert gyroscope readings into motion fields. Second, we propose an Essential Mixtures motion model for rolling shutter cameras, where an array of rotations within a frame are extracted as the ground-truth guidance. Third, we train a convolutional neural network with gyroscope motions as input to compensate for the OIS motion. Once finished, the compensation network can be applied for other scenes, where the image alignment is purely based on gyroscopes with no need for images contents, delivering strong robustness. Experiments show that our results are comparable with that of non-OIS cameras, and outperform image-based alignment results with a relatively large margin. Code and dataset is available at:https://github.com/lhaippp/DeepOIS. Shuaicheng Liu, Haipeng Li 0001, Zhengning Wang, Jue Wang 0001, Shuyuan Zhu, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Quadratic Terms Based Point-to-Surface 3D Representation for Deep Learning of Point CloudabstractIn this paper, we introduce a novel point-to-surface representation for 3D point cloud learning. Unlike the previous methods that mainly adopt voxel, mesh, or point coordinates, we propose to tackle this problem from a new perspective: learn a set of quadratic terms based static and global reference surfaces to describe 3D shapes, such that the coordinates of a 3D point (x, y, z) can be extended to quadratic terms (xy, xz, yz,$\ldots $) and transformed to the relationship between the local point and the global reference surfaces. Then, the static surfaces are changed into dynamic surfaces by adaptive contribution weighting to improve the descriptive capability. Towards this end, we propose our point-to-surface representation, a new representation for 3D point cloud learning that has not been attempted before, which can assemble local and global geometric information effectively by building connections between the point cloud and the learned reference surfaces. Given 3D points, we show how the reference surfaces are constructed, and how they are inserted into the 3D learning pipeline for different tasks. The experimental results confirm the effectiveness of our new representation, which has outperformed the state-of-the-art methods on the tasks of 3D classification and segmentation. Tiecheng Sun, Guanghui Liu 0001, Ru Li 0002, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | JigsawGAN: Auxiliary Learning for Solving Jigsaw Puzzles With Generative Adversarial NetworksabstractThe paper proposes a solution based on Generative Adversarial Network (GAN) for solving jigsaw puzzles. The problem assumes that an image is divided into equal square pieces, and asks to recover the image according to information provided by the pieces. Conventional jigsaw puzzle solvers often determine the relationships based on the boundaries of pieces, which ignore the important semantic information. In this paper, we propose JigsawGAN, a GAN-based auxiliary learning method for solving jigsaw puzzles with unpaired images (with no prior knowledge of the initial images). We design a multi-task pipeline that includes, (1) a classification branch to classify jigsaw permutations, and (2) a GAN branch to recover features to images in correct orders. The classification branch is constrained by the pseudo-labels generated according to the shuffled pieces. The GAN branch concentrates on the image semantic information, where the generator produces the natural images to fool the discriminator, while the discriminator distinguishes whether a given image belongs to the synthesized or the real target domain. These two branches are connected by a flow-based warp module that is applied to warp features to correct the order according to the classification results. The proposed method can solve jigsaw puzzles more efficiently by utilizing both semantic information and boundary information simultaneously. Qualitative and quantitative comparisons against several representative jigsaw puzzle solvers demonstrate the superiority of our method. Ru Li 0002, Shuaicheng Liu, Guangfu Wang, Guanghui Liu 0001, Bing Zeng 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Multi-Decoding Deraining Network and Quasi-Sparsity Based TrainingabstractExisting deep deraining models are mainly learned via directly minimizing the statistical differences between rainy images and rain-free ground truths. They emphasize learning a mapping from rainy images to rain-free images with supervision. Despite the demonstrated success, these methods do not perform well on restoring the fine-grained local details or removing blurry rainy traces. In this work, we aim to exploit the intrinsic priors of rainy images and develop intrinsic loss functions to facilitate training deraining networks, which decompose a rainy image into a rain-free background layer and a rainy layer containing intact rain streaks. To this end, we introduce the quasi-sparsity prior to train network so as to generate two sparse layers with intact textures of different objects. Then we explore the low-value prior to compensate sparsity, forcing all rain streaks to enter into one layer while non-rain con-tents into another layer to restore image details. We introduce a multi-decoding structure to specially supervise the generation of multi-type deraining features. This helps to learn the most contributory features to deraining in respective spaces. Moreover, our model stabilizes the feature values from multi-spaces via information sharing to alleviate potential artifacts, which also accelerates the running speed. Extensive experiments show that the proposed de-raining method outperforms the state-of-the-art approaches in terms of effectiveness and efficiency. Yinglong Wang 0002, Chao Ma 0004, Bing Zeng 0001 |
CVPR | 3 |
| 2021 | Personalized Outfit Recommendation With Learnable AnchorsabstractThe multimedia community has recently seen a tremendous surge of interest in the fashion recommendation problem. A lot of efforts have been made to model the compatibility between fashion items. Some have also studied users’ personal preferences for the outfits. There is, however, another difficulty in the task that hasn’t been dealt with carefully by previous work. Users that are new to the system usually only have several (less than 5) outfits available for learning. With such a limited number of training examples, it is challenging to model the user’s preferences reliably. In this work, we propose a new solution for personalized outfit recommendation that is capable of handling this case. We use a stacked self-attention mechanism to model the high-order interactions among the items. We then embed the items in an outfit into a single compact representation within the outfit space. To accommodate the variety of users’ preferences, we characterize each user with a set of anchors, i.e. a group of learnable latent vectors in the outfit space that are the representatives of the outfits the user likes. We also learn a set of general anchors to model the general preference shared by all users. Based on this representation of the outfits and the users, we propose a simple but effective strategy for the new user profiling tasks. Extensive experiments on large scale real-world datasets demonstrate the performance of our proposed method. Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001 |
CVPR | 4 |
| 2021 | Holistic 3D Scene Understanding From a Single Image With Implicit RepresentationabstractWe present a new pipeline for holistic 3D scene understanding from a single image, which could predict object shapes, object poses, and scene layout. As it is a highly ill-posed problem, existing methods usually suffer from inaccurate estimation of both shapes and layout especially for the cluttered scene due to the heavy occlusion between objects. We propose to utilize the latest deep implicit representation to solve this challenge. We not only propose an image-based local structured implicit network to improve the object shape estimation, but also refine the 3D object pose and scene layout via a novel implicit scene graph neural network that exploits the implicit local object features. A novel physical violation loss is also proposed to avoid incorrect context between objects. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods in terms of object shape, scene layout estimation, and 3D object detection. Zhaopeng Cui, Yinda Zhang 0001, Bing Zeng 0001, Marc Pollefeys, Shuaicheng Liu |
CVPR | 4 |
| 2021 | OMNet: Learning Overlapping Mask for Partial-to-Partial Point Cloud RegistrationabstractPoint cloud registration is a key task in many computational fields. Previous correspondence matching based methods require the inputs to have distinctive geometric structures to fit a 3D rigid transformation according to point-wise sparse feature matches. However, the accuracy of transformation heavily relies on the quality of extracted features, which are prone to errors with respect to partiality and noise. In addition, they can not utilize the geometric knowledge of all the overlapping regions. On the other hand, previous global feature based approaches can utilize the entire point cloud for the registration, however they ignore the negative effect of non-overlapping points when aggregating global features. In this paper, we present OM-Net, a global feature based iterative network for partial-to-partial point cloud registration. We learn overlapping masks to reject non-overlapping regions, which converts the partial-to-partial registration to the registration of the same shape. Moreover, the previously used data is sampled only once from the CAD models for each object, resulting in the same point clouds for the source and reference. We propose a more practical manner of data generation where a CAD model is sampled twice for the source and reference, avoiding the previously prevalent over-fitting issue. Experimental results show that our method achieves state-of-the-art performance compared to traditional and deep learning based methods. Code is available at https://github.com/megvii-research/OMNet. Hao Xu 0018, Shuaicheng Liu, Guangfu Wang, Guanghui Liu 0001, Bing Zeng 0001 |
ICCV | 5 |
| 2021 | DeepPanoContext: Panoramic 3D Scene Understanding with Holistic Scene Context Graph and Relation-based OptimizationabstractPanorama images have a much larger field-of-view thus naturally encode enriched scene context information compared to standard perspective images, which however is not well exploited in the previous scene understanding methods. In this paper, we propose a novel method for panoramic 3D scene understanding which recovers the 3D room layout and the shape, pose, position, and semantic category for each object from a single full-view panorama image. In order to fully utilize the rich context information, we design a novel graph neural network based context model to predict the relationship among objects and room layout, and a differentiable relationship-based optimization module to optimize object arrangement with well-designed objective functions on-the-fly. Realizing the existing data are either with incomplete ground truth or overly-simplified scene, we present a new synthetic dataset with good diversity in room layout and furniture placement, and realistic image quality for total panoramic 3D scene understanding. Experiments demonstrate that our method outperforms existing methods on panoramic scene understanding in terms of both geometry accuracy and object arrangement. Code is available at https://chengzhag.github.io/publication/dpc. Zhaopeng Cui, Cai Chen 0002, Shuaicheng Liu, Bing Zeng 0001, Hujun Bao, Yinda Zhang 0001 |
ICCV | 5 |
| 2021 | Hierarchical Region Proposal Refinement Network for Weakly Supervised Object DetectionabstractWeakly supervised object detection (WSOD) has attracted more attention because it only requires image-level annotations to indicate whether a certain class exists. Most WSOD methods utilize multiple instance learning (MIL) to train an object detector where an image is treated as a bag of candidate proposals. Unlike fully supervised object detection (FSOD) that uses the object-aware region proposal network (RPN) to generate effective candidate proposals, WSOD only utilizes region proposal methods (e.g., selective search or edge boxes) due to the lack of instance-level annotations (i.e., bounding boxes). However, the quality of proposals can influence the training of the detector. To solve this problem, we propose a hierarchical region proposal refinement network (HRPRN) to refine these proposals gradually. Specifically, our network contains multiple weakly supervised detectors that are trained stage by stage. In addition, we propose an instance regression refinement model to generate object-aware coordinate offsets to refine proposals at each stage. In order to demonstrate the effectiveness of our method, we conduct experiments on PASCAL VOC 2007 dataset that is the widely used benchmark. Compared with our baseline method, online instance classifier refinement (OICR), our method achieves 9% and 5.6% improvements in terms of mAP and CorLoc, respectively. Shuaicheng Liu, Bing Zeng 0001 |
ICIP | 3 |
| 2021 | Simplified Power-Based Detectors for Ship Detection of PolSAR ImageryabstractShip detection of polarimetric SAR (PolSAR) imagery has attracted lots of attentions in recent years. Also, it is known that among the polarimetric channels$HH, HV$, and$VV, VV$is the most sensitive to sea clutter. Following this guidance, in this paper, a novel ship detector SVVS is first proposed via subtracting the term$C_{33}$from the total power detector SPAN. And then, the complect polarimetric covariance difference matrix [$CP$] is utilized to calculate SVVS, leading to the construction of another novel ship detector$\text{SVVS}_{CP}$. Finally, we investigate the statistical distribution of sea clutter with$\text{SVVS}_{CP}$and further develop an adaptive$\text{SVVS}_{CP}$-based C-FAR detector for ship detection. The experiment carried out on one real PolSAR imagery shows that, compared to SPAN, both SVVS and$\text{SVVS}_{CP}$hold better ship detection performances. Tao Zhang 0027, Hongping Gan, Zhen Yang 0012, Bing Zeng 0001, Jian Yang 0011 |
IGARSS | 4 |
| 2021 | GLM-Net: Global and Local Motion Estimation via Task-Oriented Encoder-Decoder StructureabstractIn this work, we study the problem of separating the global camera motion and the local dynamic motion from an optical flow. Previous methods either estimate global motions by a parametric model, such as a homography, or estimate both of them by an optical flow field. However, none of these methods can directly estimate global and local motions through an end-to-end manner. In addition, separating the two motions accurately from a hybrid flow field is challenging. Because one motion can easily confuse the estimate of the other one when they are compounded together. To this end, we propose an end-to-end global and local motion estimation network GLM-Net. We design two encoder-decoder structures for the motion separation in the optical flow based on different task orientations. One structure adopts a mask autoencoder to extract the global motion, while the other one uses attention U-net for the local motion refinement. We further designed two effective training methods to overcome the problem of lacking supervisions. We apply our method on the action recognition datasets NCAA and UCF-101 to verify the accuracy of the local motion, and the homography estimation dataset DHE for the accuracy of the global motion. Experimental results show that our method can achieve competitive performance in both tasks at the same time, validating the effectiveness of the motion separation. Ye Xiang, Shuaicheng Liu, Lifang Wu, Boxuan Zhao, Bing Zeng 0001 |
ACM Multimedia | 6 |
| 2021 | Cross-Block Difference Guided Fast CU Partition for VVC Intra CodingabstractIn this paper, we propose a new fast CU partition method for VVC intra coding based on the cross-block difference. This difference is measured by the gradient and the content of sub-blocks obtained from partition and is employed to guide the skipping of unnecessary horizontal and vertical partition modes. With this guidance, a fast determination of block partitions is accordingly achieved. Compared with VVC, our proposed method can save 41.64% (on average) encoding time with only 0.97% (on average) increase of BD-rate. Hewei Liu, Shuyuan Zhu, Ruiqin Xiong, Guanghui Liu 0001, Bing Zeng 0001 |
VCIP | 5 |
| 2021 | OAENet: Oriented attention ensemble for accurate facial expression recognition
Zhengning Wang, Fanwei Zeng, Shuaicheng Liu, Bing Zeng 0001 |
Pattern Recognit. | 4 |
| 2021 | A Robust Quality Enhancement Method Based on Joint Spatial-Temporal Priors for Video CodingabstractQuality enhancement of HEVC compressed videos has attracted a lot of attentions in recent years. In this article, we propose a robust multi-frame guided attention network (MGANet) to reconstruct high-quality frames based on HEVC compressed videos. In our network, we first use an advanced motion flow algorithm to estimate the motion information of input frames so as to guide the warping of adjacent frames. After performing the alignment, we find that large residuals still appear in the edge area of moving objects of the warped frames. Then, we design a temporal encoder based on a bi-directional convolutional long short term memory (ConvLSTM) with residual structure to further discover the variations between the current frame and its adjacent warped frames. Finally, we feed the extracted temporal information and a partitioned average image (PAI) to a multi-scale guided encoder-decoder subnet to reconstruct high-quality frames. Here, each PAI is generated according to the transform unit (TU) partitioning map that can be extracted directly from the coded bit-streams, thus enabling our network to focus on the TU boundaries while optimizing the global content. We present extensive experimental results to demonstrate the robustness of our method, especially for the high bit-rate coding case and large motion scenes. Due to the lightweight design structure, our proposed MGANet also has a very competitive inference time. Xiandong Meng, Shuyuan Zhu, Xinfeng Zhang 0001, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | NTSDCN: New Three-Stage Deep Convolutional Image Demosaicking NetworkabstractIn this letter, we compose a new three-stage deep convolutional neural network (NTSDCN) for image demosaicking, and it consists of our proposed Laplacian energy-constrained local residual unit (LC-LRU) and a feature-guided prior fusion unit (FG-PFU). Specifically, the LC-LRU is used to refine the learning target of the specific residual blocks in the network and enhance the dominant information of the residual features. The FG-PFU is designed to guide the feature extraction of the red (R) and blue (B) channels by utilizing prior information from the reconstructed green (G) channel. In our proposed NTSDCN, we recover the G channel image in the first stage with the CFA image and reconstruct the R and B images in the second stage. Finally, we fine-tune the resulting R, G and B images in the third stage to compose a full-color RGB image. The experimental results show that our proposed method achieves better performance than the state-of-the-art methods. The code is available at https://github.com/wyannn/NTSDCN. Shiying Yin, Shuyuan Zhu, Zhan Ma 0001, Ruiqin Xiong, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | SDP-GAN: Saliency Detail Preservation Generative Adversarial Networks for High Perceptual Quality Style TransferabstractThe paper proposes a solution to effectively handle salient regions for style transfer between unpaired datasets. Recently, Generative Adversarial Networks (GAN) have demonstrated their potentials of translating images from source domain X to target domain Y in the absence of paired examples. However, such a translation cannot guarantee to generate high perceptual quality results. Existing style transfer methods work well with relatively uniform content, they often fail to capture geometric or structural patterns that always belong to salient regions. Detail losses in structured regions and undesired artifacts in smooth regions are unavoidable even if each individual region is correctly transferred into the target style. In this paper, we propose SDP-GAN, a GAN-based network for solving such problems while generating enjoyable style transfer results. We introduce a saliency network, which is trained with the generator simultaneously. The saliency network has two functions: (1) providing constraints for content loss to increase punishment for salient regions, and (2) supplying saliency features to generator to produce coherent results. Moreover, two novel losses are proposed to optimize the generator and saliency networks. The proposed method preserves the details on important salient regions and improves the total image perceptual quality. Qualitative and quantitative comparisons against several leading prior methods demonstrates the superiority of our method. Ru Li 0002, Chihao Wu 0001, Shuaicheng Liu, Jue Wang 0001, Guangfu Wang, Guanghui Liu 0001, Bing Zeng 0001 |
IEEE Trans. Image Process. | 7 |
| 2021 | OIFlow: Occlusion-Inpainting Optical Flow Estimation by Unsupervised LearningabstractOcclusion is an inevitable and critical problem in unsupervised optical flow learning. Existing methods either treat occlusions equally as non-occluded regions or simply remove them to avoid incorrectness. However, the occlusion regions can provide effective information for optical flow learning. In this paper, we present OIFlow, an occlusion-inpainting framework to make full use of occlusion regions. Specifically, a new appearance-flow network is proposed to inpaint occluded flows based on the image content. Moreover, a boundary dilated warp is proposed to deal with occlusions caused by displacement beyond the image border. We conduct experiments on multiple leading flow benchmark datasets such as Flying Chairs, KITTI and MPI-Sintel, which demonstrate that the performance is significantly improved by our proposed occlusion handling framework. Shuaicheng Liu, Kunming Luo, Nianjin Ye, Chuan Wang 0001, Jue Wang 0001, Bing Zeng 0001 |
IEEE Trans. Image Process. | 6 |
| 2021 | SlimConv: Reducing Channel Redundancy in Convolutional Neural Networks by Features RecombiningabstractThe channel redundancy of convolutional neural networks (CNNs) results in the large consumption of memories and computational resources. In this work, we design a novel Slim Convolution (SlimConv) module to boost the performance of CNNs by reducing channel redundancies. Our SlimConv consists of three main steps: Reconstruct, Transform, and Fuse. It aims to reorganize and fuse the learned features more efficiently, such that the method can compress the model effectively. Our SlimConv is a plug-and-play architectural unit that can be used to replace convolutional layers in CNNs directly. We validate the effectiveness of SlimConv by conducting comprehensive experiments on various leading benchmarks, such as ImageNet, MS COCO2014, Pascal VOC2012 segmentation, and Pascal VOC2007 detection datasets. The experiments show that SlimConv-equipped models can achieve better performances consistently, less consumption of memory and computation resources than non-equipped counterparts. For example, the ResNet-101 fitted with SlimConv achieves 77.84% top-1 classification accuracy with 4.87 GFLOPs and 27.96M parameters on ImageNet, which shows almost 0.5% better performance with about 3 GFLOPs and 38% parameters reduced. Jiaxiong Qiu, Cai Chen 0002, Shuaicheng Liu, Heng-Yu Zhang, Bing Zeng 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Deep Single Image Deraining via Modeling Haze-Like EffectabstractRemoving rain from images is of a great importance to various applications such as autonomous driving, drone piloting, and photo editing. Conventional methods rely on some heuristics to handcraft various priors to remove or separate rain from images. Recently, deep learning models are proposed to learn various end-to-end methods to complete this task. However, these methods might fail in obtaining satisfactory results in some real-world scenarios, especially when the captured images suffer from heavy rain that brings not only rain streaks but also a haze-like effect (caused by the accumulation of tiny raindrops). Different from most of the existing deep learning deraining methods that focus on handling rain streaks, we add a new variable to model the haze-like effect in a general model for rain, based on which a deep neural network is designed accordingly. Specifically, in our method, two branches are designed to handle rain streaks and the haze-like effect, respectively. The output of such branch structure is fed to an additional module to further enhance the performance. Three modules are trained jointly, leading to an end-to-end network, which supports a adjustment to the strength of removing the haze-like effect. Extensive experiments on several datasets show that our method outperforms several state-of-the-art methods in both objective assessment and visual quality. Yinglong Wang 0002, Dong Gong, Jie Yang 0002, Qinfeng Shi, Anton van den Hengel, Dehua Xie, Bing Zeng 0001 |
IEEE Trans. Multim. | 7 |
| 2020 | Neural Point Cloud Rendering via Multi-Plane ProjectionabstractWe present a new deep point cloud rendering pipeline through multi-plane projections. The input to the network is the raw point cloud of a scene and the output are image or image sequences from a novel view or along a novel camera trajectory. Unlike previous approaches that directly project features from 3D points onto 2D image domain, we propose to project these features into a layered volume of camera frustum. In this way, the visibility of 3D points can be automatically learnt by the network, such that ghosting effects due to false visibility check as well as occlusions caused by noise interferences are both avoided successfully. Next, the 3D feature volume is fed into a 3D CNN to produce multiple planes of images w.r.t. the space division in the depth directions. The multi-plane images are then blended based on learned weights to produce the final rendering results. Experiments show that our network produces more stable renderings compared to previous methods, especially near the object boundaries. Moreover, our pipeline is robust to noisy and relatively sparse point cloud for a variety of challenging scenes. Peng Dai 0003, Yinda Zhang 0001, Zhuwen Li, Shuaicheng Liu, Bing Zeng 0001 |
CVPR | 5 |
| 2020 | Flow-Guided Temporal-Spatial Network for HEVC Compressed Video Quality EnhancementabstractIn this paper, a flow-guided temporal-spatial network (FGTSN) is proposed to enhance the quality of HEVC compressed video. Specifically, we first employ a robust motion estimation subnet via trainable optical flow module to estimate the motion flow between the target frame and its adjacent frames, and these adjacent frames are pre-warped guided by the predicted motion flow. Then, a temporal encoder is proposed to fuse the related information between the target frame and its pre-warped frames. Finally, a quality enhancement subnet with multi-scale encoder-decoder structure is designed to generate high quality frame by training the network in a multi-supervised fashion. Experimental results show the superior performance of our proposed FGTSN method for the reconstruction quality of HEVC compressed frames, much better than the state-of-the-art quality enhancement methods. In addition, our FGTSN method can also effectively mitigate the quality fluctuation of adjacent frames. Xiandong Meng, Shuyuan Zhu, Shuaicheng Liu, Bing Zeng 0001 |
DCC | 5 |
| 2020 | Rethinking Image Deraining via Rain Streaks and Vapors
Yinglong Wang 0002, Yibing Song, Chao Ma 0004, Bing Zeng 0001 |
ECCV (17) | 4 |
| 2020 | Salient Object Detection Based On Image Bit-MapabstractIn this paper, we propose a novel salient object detection framework, which makes full use of the essential image compression. More specifically, we first compose an intuitive measure of compressibility from JPEG compression, namely bit-map. Then, depending on the relationship between bitmap and salient object, we generate the salient object window directly from bit-map without utilizing any features from the compressed image. Finally, the saliency map is calculated according to the salient object window and with a ranking algorithm. The proposed method achieves good performance as well as low complexity. The experimental results demonstrate the effectiveness of our proposed method compared with other existing approaches. Bangqi Cao, Xiandong Meng, Shuyuan Zhu, Bing Zeng 0001 |
ICASSP | 4 |
| 2020 | Optical Flow Estimation Between Images of Different Resolutions via Variational MethodabstractTraditional optical flow estimation methods mostly focus on images of the same resolution. However, there are some situations requiring optical flow between images of different resolutions, where the traditional approaches suffer from the inequality of spectrum aliasing level. In this paper, we propose a method estimating the flow fields between a clear image and a highly undersampled one. The proposed method simultaneously describes the motion and integral relationship between the images via an integral form image under the assumption of brightness and gradient consistency as well as motion smoothness. We also derive the numerical solution briefly, through which we can solve the equations easily via linearizations. Experimental results on Middlebury and MPI-Sintel datasets demonstrate that our proposed method outperforms traditional methods preprocessing images of different resolutions to be the same size, offering more accurate results. Rui Zhao 0010, Ruiqin Xiong, Shuyuan Zhu, Bing Zeng 0001, Tiejun Huang 0001, Wen Gao 0001 |
VCIP | 4 |
| 2020 | WiFi Vision: Sensing, Recognition, and Detection With Commodity MIMO-OFDM WiFiabstractIndoor human sensing, recognition, and detection, as key enablers of building smart environments, such as smart home, smart retail, and smart museum, have gained tremendous attention in recent years. Compared with traditional vision-based and wearable sensor-based solutions, radio-frequency (RF)-based approaches are more desirable with the contactless and nonline-of-sight nature. Among all RF-based approaches, WiFi-based approaches have been the focus of many researchers because of the ubiquitous availability and cost efficiency. In this article, we present a survey of recent advances in WiFi vision problems, i.e., sensing, recognition, and detection by utilizing the channel state information (CSI) of the commodity WiFi devices. We focus on nine key applications of smart environments, including WiFi imaging, vital sign monitoring, human identification, gesture recognition, gait recognition, daily activity recognition, fall detection, human detection, and indoor positioning. Such a survey can help readers have an overall understanding of sensing, recognition, and detection with commodity WiFi, and thus expedite the development of smart environments. Ying He 0013, Yan Chen 0007, Yang Hu 0006, Bing Zeng 0001 |
IEEE Internet Things J. | 4 |
| 2020 | Predominant Instrument Recognition Based on Deep Neural Network With Auxiliary ClassificationabstractInstrument recognition plays very important roles in music information retrieval, sound source separation and automatic music transcription. However, due to different playing styles and audio qualities, this task cannot be accomplished easily. Simultaneous existence of multiple instruments in polyphonic music increases the challenge to a greater extent. This article mainly focus on the identification of the predominant instruments in polyphonic music. We propose to construct a network with an auxiliary classification designed based on the onset groups and instrument families. The principal classification and the auxiliary classification enable the network to learn the instrument categories and groups jointly in a pattern of multitask learning. The IRMAS datasetis adopted in the experiment to extract the mel-spectrogram and six other types of features. The micro and macro average of precisions, recalls and F1 measures are used to evaluate the classification results. The effect of multitask learning, batch normalization and center loss in the predominant instrument recognition are demonstrated by various experiments. By selecting the loss ratios through a development set, the micro and macro F1 measures of our proposed network can reach 0.685 and 0.597, which are 10.7% and 16.4% higher than those obtained by the baseline, the ConvNet presented in [1]. Dongyan Yu, Huiping Duan, Jun Fang 0001, Bing Zeng 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | MUcast: Linear Uncoded Multiuser Video Streaming With Channel Assignment and Power Allocation OptimizationabstractMultiuser video transmission, where the server transmits videos to multiple users that require different contents at the same time, becomes more and more popular with the development of wireless communication technology. One key problem in multiuser video transmission is how to optimally allocate system resources such as transmission power and channels to multiple users to achieve the best system performance. To resolve the problem, in this paper, we propose an uncoded multiuser video streaming system, which exploits diversities of video contents and channel conditions of multiple users. We first solve the channel assignment problem with known power allocation by taking into account the intra-block energy diffusion and inter-block energy aliasing. Then, with the obtained channel assignment, we derive a closed-form solution to the multiuser power allocation optimization problem. Finally, we conduct simulations to evaluate the proposed uncoded multiuser video streaming system by comparing with three other approaches, and the simulation results show that the proposed method can achieve the best system performance. Chaofan He, Yang Hu 0006, Yan Chen 0007, Xiaopeng Fan 0001, Houqiang Li, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | PBR-Net: Imitating Physically Based Rendering Using Deep Neural NetworkabstractPhysically based rendering has been widely used to generate photo-realistic images, which greatly impacts industry by providing appealing rendering, such as for entertainment and augmented reality, and academia by serving large scale high-fidelity synthetic training data for data hungry methods like deep learning. However, physically based rendering heavily relies on ray-tracing, which can be computational expensive in complicated environment and hard to parallelize. In this paper, we propose an end-to-end deep learning based approach to generate physically based rendering efficiently. Our system consists of two stacked neural networks, which effectively simulates the physical behavior of the rendering process and produces photo-realistic images. The first network, namely shading network, is designed to predict the optimal shading image from surface normal, depth and illumination; the second network, namely composition network, learns to combine the predicted shading image with the reflectance to generate the final result. Our approach is inspired by intrinsic image decomposition, and thus it is more physically reasonable to have shading as intermediate supervision. Extensive experiments show that our approach is robust to noise thanks to a modified perceptual loss and even outperforms the physically based rendering systems in complex scenes given a reasonable time budget. Peng Dai 0003, Zhuwen Li, Yinda Zhang 0001, Shuaicheng Liu, Bing Zeng 0001 |
IEEE Trans. Image Process. | 5 |
| 2020 | A New Polyphase Down-Sampling-Based Multiple Description Image CodingabstractMultiple description coding (MDC) is an efficient source coding technique for error-prone transmission over multiple channels. In this paper, we focus on the design of a new polyphase down-sampling based MDC (NPDS-MDC) for image signals. The encoding of our proposed NPDS-MDC consists of three steps. First, we perform down-sampling on each N×N image block according to the quincunx down-sampling pattern. Second, we propose a new transform and apply it to the down-sampled pixels to produce the side descriptions. Third, we develop an error compensation algorithm to reduce the compression distortion occurring on the down-sampled pixels. In our scheme, the side decoding is performed posterior to image interpolation with reference to the down-sampled compressed pixels. Moreover, the central decoding is achieved by interlacing the side descriptions. We also propose a compression-constrained central deblocking algorithm to further improve the efficiency of the central decoding. The experimental results indicate that our proposed MDC scheme offers clearly superior performance, especially at high bit rates, as compared to the state-of-the-art methods for various types of images. Shuyuan Zhu, Zhiying He, Xiandong Meng, Jiantao Zhou 0001, Yuanfang Guo, Bing Zeng 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Residual Carrier Frequency Offset Estimation and Compensation for Commodity WiFiabstractThe various offsets existed in the commodity WiFi devices greatly limit the use of ubiquitous WiFi signals for indoor applications. In this paper, we focus on the estimation and compensation of the residual carrier frequency offset (CFO) for the commodity WiFi devices. Specifically, we consider a distorted channel state information (CSI) model by taking into consideration various CSI errors such as packet detection delay (PDD) and CFO. We propose a multiscale sparse recovery algorithm to get rid of the effect of PDD and extract the carrier frequency component out of CSI. Then, we formulate the residual CFO estimation as a spectrum estimation problem and utilize the MUSIC algorithm to estimate the residual CFO. Real experiments and numerical simulations are conducted to evaluate the performance of the proposed method. The experimental results and simulation results show that the residual CFO is time-varying, and compared with existing methods, the proposed method can better estimate and compensate the residual CFO, and thus achieve better results. Yan Chen 0007, Yang Hu 0006, Bing Zeng 0001 |
IEEE Trans. Mob. Comput. | 4 |
| 2019 | Learning Binary Code for Personalized Fashion RecommendationabstractWith the rapid growth of fashion-focused social networks and online shopping, intelligent fashion recommendation is now in great needs. Recommending fashion outfits, each of which is composed of multiple interacted clothing and accessories, is relatively new to the field. The problem becomes even more interesting and challenging when considering users' personalized fashion style. Another challenge in a large-scale fashion outfit recommendation system is the efficiency issue of item/outfit search and storage. In this paper, we propose to learn binary code for efficient personalized fashion outfits recommendation. Our system consists of three components, a feature network for content extraction, a set of type-dependent hashing modules to learn binary codes, and a matching block that conducts pairwise matching. The whole framework is trained in an end-to-end manner. We collect outfit data together with user label information from a fashion-focused social website for the personalized recommendation task. Extensive experiments on our datasets show that the proposed framework outperforms the state-of-the-art methods significantly even with a simple backbone. Yang Hu 0006, Yunchao Jiang, Yan Chen 0007, Bing Zeng 0001 |
CVPR | 5 |
| 2019 | DeepLiDAR: Deep Surface Normal Guided Depth Prediction for Outdoor Scene From Sparse LiDAR Data and Single Color ImageabstractIn this paper, we propose a deep learning architecture that produces accurate dense depth for the outdoor scene from a single color image and a sparse depth. Inspired by the indoor depth completion, our network estimates surface normals as the intermediate representation to produce dense depth, and can be trained end-to-end. With a modified encoder-decoder structure, our network effectively fuses the dense color image and the sparse LiDAR depth. To address outdoor specific challenges, our network predicts a confidence mask to handle mixed LiDAR signals near foreground boundaries due to occlusion, and combines estimates from the color image and surface normals with learned attention maps to improve the depth accuracy especially for distant areas. Extensive experiments demonstrate that our model improves upon the state-of-the-art performance on KITTI depth completion benchmark. Ablation study shows the positive impact of each model components to the final performance, and comprehensive analysis shows that our model generalizes well to the input with higher sparsity or from indoor scenes. Jiaxiong Qiu, Zhaopeng Cui, Yinda Zhang 0001, Xingdi Zhang, Shuaicheng Liu, Bing Zeng 0001, Marc Pollefeys |
CVPR | 6 |
| 2019 | Exploiting Channel Assignment and Power Allocation for Linear Uncoded Multiuser Video StreamingabstractMultiuser video transmission, where the server transmits videos to multiple users that request different contents at the same time, becomes more and more popular with the development of wireless communication technology. One key problem in multiuser video transmission is how to optimally allocate the system resources such as transmission power and channels to multiple users to achieve the best system performance. To resolve the problem, in this paper, we propose an uncoded multiuser video streaming system, which exploits diversities of video contents and channel conditions of multiple users. We first solve the channel assignment problem with known power allocation by taking into account the intra-block energy diffusion and inter-block energy aliasing. Then, with the obtained channel assignment, we derive a closed-form solution to the multiuser power allocation optimization problem. Finally, we conduct simulations to evaluate the proposed uncoded multiuser video streaming system by comparing with three other approaches, and simulation results show that the proposed method can achieve the best system performance. Chaofan He, Yang Hu 0006, Yan Chen 0007, Xiaopeng Fan 0001, Houqiang Li, Bing Zeng 0001 |
ICC | 6 |
| 2019 | Estimating and Compensating Residual Carrier Frequency Offset for Commodity WiFiabstractThe various offsets existed on the commodity WiFi devices greatly limit the use of ubiquitous WiFi signals for indoor applications. In this paper, we focus on the estimation and compensation of the residual carrier frequency offset (CFO) for the commodity WiFi devices. Specifically, we introduce a distorted channel state information (CSI) model by taking into consideration various CSI errors such as packet detection delay (PDD) and CFO. We propose a multiscale sparse recovery algorithm to get rid of the effect of PDD and extract the carrier frequency component out of CSI. Then, we formulate the residual CFO estimation as a spectrum estimation problem and propose to utilize the MUSIC algorithm to estimate the residual CFO. Real experiments are conducted to evaluate the performance of the proposed method. The experimental results show that the residual CFO is time-varying, and compared with existing methods, the proposed method can better estimate and compensate the residual CFO, and thus achieve better results. Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001 |
ICC | 4 |
| 2019 | Personalized Fashion DesignabstractFashion recommendation is the task of suggesting a fashion item that fits well with a given item. In this work, we propose to automatically synthesis new items for recommendation. We jointly consider the two key issues for the task, i.e., compatibility and personalization. We propose a personalized fashion design framework with the help of generative adversarial training. A convolutional network is first used to map the query image into a latent vector representation. This latent representation, together with another vector which characterizes user's style preference, are taken as the input to the generator network to generate the target item image. Two discriminator networks are built to guide the generation process. One is the classic real/fake discriminator. The other is a matching network which simultaneously models the compatibility between fashion items and learns users' preference representations. The performance of the proposed method is evaluated on thousands of outfits composited by online users. The experiments show that the items generated by our model are quite realistic. They have better visual quality and higher matching degree than those generated by alternative methods. Cong Yu 0011, Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001 |
ICCV | 4 |
| 2019 | Hybrid Synthesis for Exposure Fusion from Hand-Held Camera InputsabstractThe paper proposes a hybrid synthesis method for multi-exposure image fusion taken by hand-held cameras. Motions either due to the shaky cameras or caused by dynamic scenes should be compensated before any content fusion. The misalignment will cause blurring/ghosting artifacts in the fused result. The proposed method can deal with such motions and maintain the exposure information of each input effectively. In particular, the proposed method first applies optical flow for a coarse registration, which performs well with complex non-rigid motion but produces deformations at regions with missing correspondences. To correct such error registration, we segment images into superpixels and identify problematic alignments based on each superpixel, which is further aligned by PatchMatch. After that, the proposed method obtains a fully aligned image stack which facilitates a high-quality fusion that is free from blurring/ghosting artifacts. We compare our method with existing fusion algorithms on various challenging examples, including the static/dynamic, the indoor/outdoor and the daytime/nighttime scenes. Experiment results demonstrate the effectiveness and robustness. Ru Li 0002, Shuaicheng Liu, Guanghui Liu 0001, Bing Zeng 0001 |
ICIP | 4 |
| 2019 | Enhancing Quality for VVC Compressed Videos by Jointly Exploiting Spatial Details and Temporal StructureabstractIn this paper, we propose a quality enhancement network of versatile video coding (VVC) compressed videos by jointly exploiting spatial details and temporal structure (SDTS). The proposed network consists of a temporal structure fusion subnet and a spatial detail enhancement subnet. The former subnet is used to estimate and compensate the temporal motion across frames, and the latter subnet is used to reduce the compression artifacts and enhance the reconstruction quality of compressed video. Experimental results demonstrate the effectiveness of our SDTS-based method. The code of our proposed method is available at https://github.com/mengab/SDTS. Xiandong Meng, Shuyuan Zhu, Bing Zeng 0001 |
ICIP | 4 |
| 2019 | Multiple Description Image Coding Based on Compression-Guided OptimizationabstractIn this paper, we design a new multiple description coding scheme for image signals based on our proposed compression-guided optimization. Firstly, we propose a compression-constrained adaptive filtering method to produce two descriptions for the source image, where the proposed filtering algorithm works not only to guarantee a high-quality side decoding but also make a high-efficient central decoding. Secondly, we design a compression-dependent deblocking algorithm based on the transform coefficients which are decoded from both descriptions to improve the performance for the cental decoding. Experimental results demonstrate that our proposed method achieves impressive performance gains when it is applied to image signals. Shuyuan Zhu, Zhiying He, Xiandong Meng, Guanghui Liu 0001, Bing Zeng 0001 |
PCS | 5 |
| 2019 | Color Image Compression with Transform Domain Down-Sampling and Deep Convolutional ReconstructionabstractIn this paper, we build up a new block-based color image compression scheme based on our proposed transform domain down-sampling method and deep convolutional reconstruction algorithm. Specifically, our proposed down-sampling scheme aims to down-sample each N × N transform block into the N/2 × N/2 block for the saving of bit-cost. On the other hand, the proposed deep convolutional reconstruction algorithm is employed to reconstruct the down-sampled block for a full- resolution reconstruction. We apply our proposed methods to both the chrominance components to compress color images. Experimental results show that our proposed method achieves excellent results when used in practice. Shuyuan Zhu, Xiandong Meng, Bing Zeng 0001, Yuanfang Guo, Ruiqin Xiong |
VCIP | 4 |
| 2019 | Deep Deterministic Policy Gradient (DDPG)-Based Energy Harvesting Wireless CommunicationsabstractTo overcome the difficulties of charging the wireless sensors in the wild with conventional energy supply, more and more researchers have focused on the sensor networks with renewable generations. Considering the uncertainty of the renewable generations, an effective energy management strategy is necessary for the sensors. In this paper, we propose a novel energy management algorithm based on the reinforcement learning. By utilizing deep deterministic policy gradient (DDPG), the proposed algorithm is applicable for the continuous states and realizes the continuous energy management. We also propose a state normalization algorithm to help the neural network initialize and learn. With only one day's real solar data and the simulative channel data for training, the proposed algorithm shows excellent performance in the validation with about 800 days length of real solar data. Compared with the state-of-the-art algorithms, the proposed algorithm achieves better performance in terms of long-term average net bit rate. Chengrun Qiu, Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001 |
IEEE Internet Things J. | 4 |
| 2019 | BreathTrack: Tracking Indoor Human Breath Status via Commodity WiFiabstractIn this paper, we propose a contact-free breath tracking system, BreathTrack, to track the status of breath using the off-the-shelf WiFi devices. BreathTrack exploits the phase variation of the channel state information (CSI) to track human breath. To resolve the phase distortions introduced by the hardware imperfection of the commodity WiFi chips, BreathTrack utilizes both the hardware and software correction methods. The time-invariant PLL phase offset is calibrated by the hardware correction using cables and splitters, while the time-varying carrier frequency offset, sampling frequency offset and packet detection delay are removed by the software corrections using the phase difference between the CSI at the receiver antennas and that at the reference antenna connected from the transmitter. Moreover, BreathTrack utilizes the sparse recovery method to find the dominant path in the multipath indoor environment and derive the corresponding complex attenuation coefficient. Then, the phase variation of the complex attenuation coefficient is utilized to extract the detailed breath status and the breath rate. Extensive experiments are conducted to show that BreathTrack could estimate the breath rate with the median accuracy of over 99% in most scenarios, and could track the detailed status of breath directly using the raw phase variation. Dongheng Zhang, Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001 |
IEEE Internet Things J. | 4 |
| 2019 | Joint Power Allocation and Channel Assignment for NOMA With Deep Reinforcement LearningabstractNon-orthogonal multiple access (NOMA) has been considered as a significant candidate technique for the next generation wireless communication to support high throughput and massive connectivity. It allows different users to be multiplexed on one channel through applying superposition coding at the transmitter and successive interference cancellation (SIC) at the receiver. To fully utilize the benefit of the NOMA technique, the key problem is how to optimally allocate resources, such as power and channels, to users to maximize the system performance. There have been some existing works on the power allocation for the single-carrier NOMA system. However, how to optimally assign channels in the multi-carrier NOMA system is still unclear. In this paper, we propose a deep reinforcement learning framework to allocate resources to users in a near optimal way. Specifically, we exploit an attention-based neural network (ANN) to perform the channel assignment. Simulation results show that the proposed framework can achieve better system performance, compared with the state-of-the-art approaches. Chaofan He, Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001 |
IEEE J. Sel. Areas Commun. | 4 |
| 2019 | Simultaneous color-depth super-resolution with conditional generative adversarial networks
Lijun Zhao 0002, Huihui Bai 0001, Jie Liang 0001, Bing Zeng 0001, Anhong Wang, Yao Zhao 0001 |
Pattern Recognit. | 4 |
| 2019 | Local activity-driven structural-preserving filtering for noise removal and image smoothing
Lijun Zhao 0002, Huihui Bai 0001, Jie Liang 0001, Anhong Wang, Bing Zeng 0001, Yao Zhao 0001 |
Signal Process. | 5 |
| 2019 | Lyapunov Optimized Resource Management for Multiuser Mobile Video StreamingabstractBuffering techniques have been commonly used in mobile video streaming systems to handle the bandwidth fluctuation and to mitigate the impact of the stochastic characteristic of wireless channels on a mobile user's quality of experience. However, it has been shown by measurement study that users tend to abort when watching videos with mobile devices, which results in a significant wastage of the video data in the buffer. Therefore, one important problem in mobile video streaming is how to manage the buffer at each mobile user. On the other hand, mobile users generally share the wireless media to download the video data, i.e., mobile users compete with each other for the bandwidth to download the video data. Thus, another important problem in mobile video streaming is how to allocate bandwidth among mobile users. In this paper, we propose to optimize the resource management, i.e., to design buffer management strategy at each mobile user and bandwidth allocation strategy among mobile users, for the multiuser mobile video streaming systems. Specifically, we optimize the long-term average total cost of data wastage and quality of experience of mobile users with certain constraints. By introducing virtual queues and employing the Lyapunov optimization theory, we transform the original optimization problem into the drift-plus-penalty minimization problem. Then, we adopt the primal decomposition to decouple the relationship among different mobile users, which decomposes the problem into a master problem with multiple subproblems. A one-dimension full search algorithm is applied to find the global optimal solution to each subproblem, and the subgradient descent algorithm is utilized to update the solution to the master problem. Finally, simulations are conducted to show that the proposed algorithm is effective for the buffer management and bandwidth allocation in a multiuser mobile video streaming system. Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Efficient Chroma Sub-Sampling and Luma Modification for Color Image CompressionabstractIn color image compression, the chroma components are often sub-sampled before compression and up-sampled after compression. Although sub-sampling the chroma components saves the bit-cost for compression, it often induces extra color distortions in the compressed images. In this paper, we propose two approaches to tackle this problem. First, we propose a sub-sampling method in the transform domain and apply it to both chroma components. Then, based on this sub-sampling, we propose a novel method to modify the luma component. In our proposed luma modification algorithm, the distortions that occurred in the two chroma components can be coupled together and utilized to modify the luma component. With our proposed chroma sub-sampling and luma modification algorithms, we can achieve a low RGB distortion in practical image coding. The experimental results demonstrate that our proposed methods offer more significant coding gains compared with the state-of-the-art methods for the compression of color images. Shuyuan Zhu, Chang Cui, Ruiqin Xiong, Yuanfang Guo, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2019 | High-Quality Color Image Compression by Quantization Crossing Color SpacesabstractCoding of a color image usually happens in the YCbCr space so that the rate-distortion optimization is conducted in this space. Due to the use of a non-unitary matrix in the RGB-to-YCbCr conversion, an optimal coding performance achieved in the YCbCr space does not guarantee an optimal quality in the RGB space, which would impact most display devices that need RGB signals as the inputs. In this paper, we first study the relationship between the coding distortions of the compressed RGB signals and the quantization errors occurred in the coded YCbCr signals. Then, we design a new quantization scheme crossing the RGB and YCbCr spaces to achieve a high-quality color image compression with the YCbCr 4:4:4 format. Although our proposed quantization takes place in the YCbCr space, it aims at reducing the coding distortion in the RGB space as much as possible. Experimental results demonstrate that our proposed method offers a significant quality gain over the existing block-based coding methods for various images. Shuyuan Zhu, Zhiying He, Chen Chen 0015, Shuaicheng Liu, Jiantao Zhou 0001, Yuanfang Guo, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2018 | A New HEVC In-Loop Filter Based on Multi-channel Long-Short-Term Dependency Residual NetworksabstractIn this paper, we propose a new HEVC in-loop filter based on a multi-channel long-short-term dependency residual network (MLSDRN). Inspired by the information storage and information update function of human memory cell, our MLSDRN introduces an update cell to adaptively store and select the long-term and short-term dependency information through an adaptive learning process. In addition, we leverage the block boundary information that recorded in the bit-streams to improve the filter performance, which also makes our MLSDRN to unequally treat the video content. Meanwhile, the multi-channel is introduced to solve the illumination discrepancy problem. We integrate the novel in-loop filter into HM reference software, and applying it to luma and chroma components, simulation results demonstrate that the proposed in-loop filter can save BD-rate reduction up to 15.9% with ALF off. For luma component, the novel in-loop filter achieves 6.0%, 8.1%, 7.4% BD-rate saving for all intra, low delay and random access configurations, respectively. Xiandong Meng, Chen Chen 0015, Shuyuan Zhu, Bing Zeng 0001 |
DCC | 4 |
| 2018 | Photomontage for Robust HDR Imaging with Hand-Held CamerasabstractThis paper studies the image fusion from multiple images taken by hand-held cameras with different exposures. The existing methods often generate unsatisfactory results, such as the blurring/ghosting artifacts due to the problematic handling of camera motions, dynamic contents, and inappropriate fusion of local regions (e.g., over or under exposed). They often require high quality image registration before fusion. However, the accurate alignment is hard to obtain in many scenarios, such as scenes with large depth variations and dynamic textures. Besides, high quality alignment is also time consuming. In this paper, we only enable a rough registration by a single homography and combine the inputs seamlessly to hide any possible misalignment. Specifically, we propose to use a Markov Random Filed (MRF) function for the labelling of all pixels, which assigns different labels to different aligned input images. During the labelling, we choose well-exposured regions and skip moving objects simultaneously. Then, we combine a Laplace image according to the labels and construct the fusion result by solving the Poisson equation. We present various challenging examples to demonstrate the effectiveness and practicability of our approach. Ru Li 0002, Xiaowu He, Shuaicheng Liu, Guanghui Liu 0001, Bing Zeng 0001 |
ICIP | 5 |
| 2018 | Coding Trajectory: Enable Video Coding for Video DenoisingabstractWe introduce a novel video denoising approach which can produce a clean video by utilizing redundant image patches existed in the video frames. Previous multi-frame video denosing approaches either require image registration or employ Patch Match algorithms for the discovery of the patch redundancy. However, these computations are time-consuming and prone to errors. On the other hand, nearly all captured videos have been compressed. Such a compression can produce a rich set of block-based motion vectors that can be utilized for the redundant patch extraction, leading to the efficient video denosing. To be specific, the motion vectors and frame references can be obtained from the video coding. Given a noised frame block, we follow its motion vectors from the coding to form a trajectory and gather a set of block candidates along the routes from its nearby frames. The trajectory is referred to as Coding Trajectory. Then, the corresponding denoised block is generated by weighted fusing the block candidates with outlier rejections. A denoised frame is consisted of all the denoised blocks. We compare our method with several state-of-the-art approaches, such as VBM3D and VB-M4D, in terms of PSNR and SSIM. The experiments show that our method can achieve high quality results while runs much faster then the other approaches. Zhihang Ren, Peng Dai 0003, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001 |
ICIP | 5 |
| 2018 | Block-based Image Coding by Compression-Constrained Transform Domain Down-ScalingabstractTransform domain down-scaling (TDDS) is traditionally implemented by dropping most of high-frequency components of the transformed block. Applying it to image compression can improve the compression efficiency by saving considerable bit-cost. Due to losing some necessary high-frequency information, the resulted image compressed by using the traditional TDDS-based coding often suffers a serious quality degradation. In this paper, we propose a compression-constrained TDDS and perform it on each N × N block to produce an N/2 × N/2 coefficient block for the compression. Our proposed TDDS not only guarantees a high reconstruction quality but also makes a low bit-cost for compression. We integrate it in practical image coding to build up our proposed compression scheme. Experimental results show that our proposed method demonstrates excellent coding performance when used to compress image signals. Chang Cui, Shuyuan Zhu, Xiandong Meng, Shuaicheng Liu, Bing Zeng 0001 |
VCIP | 5 |
| 2018 | Multi-exposure Fusion With JPEG Compression GuidanceabstractConstruct a High Dynamic Range (HDR) image is the primary method to solve the information loss caused by insufficient dynamic range of cameras. We propose a technique for fusing a bracketed low dynamic range (LDR) image sequence of varying exposures into an HDR image, skipping the physically-based HDR assembly step. Traditionally approaches often rely on complicated algorithms to select good regions from the input LDR images for the fusion. However, we found that the selection strategy can purely base on JPEG compression bits, bypassing the calculations of image low-level features, such as image gradients, local saturations, over/under exposure evaluations, as long as the input LDR image is compressed by the JPEG formats. In this way, lots of computations can be saved. In particular, we extract the coding bits from the intermediate product of the JPEG. The coding bits of blocks can be modified as the weights for the exposure fusion. Well-exposed regions often require higher bits for the compression while overexposure or saturated regions often correspond to lower bits. The objective and subjective evaluations demonstrate the effectiveness of our method. Xingdi Zhang, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001 |
VCIP | 4 |
| 2018 | Lyapunov Optimization for Energy Harvesting Wireless Sensor CommunicationsabstractWith the development and popularity of the renewable energy harvesting devices, the energy harvesting wireless sensor communications that can make use of the energy harvested from the nearby environments have gained more and more attentions. One key problem in the energy harvesting wireless sensor communications is the transmission strategy management, i.e., how to manage the transmission strategy at each time slot to optimize the transmission performance. In this paper, we propose to use Lyapunov optimization theory to maximize the expected good bits per packet transmission for the source node in an energy harvesting wireless communication system. Considering the channel and battery states, we adapt the transmission power and modulation type to achieve such a goal. The problem is formulated as an optimization where the objective function is the long-term average good bits per packet transmission and the constraints are the bounded long-term average battery level and bit error rate. To solve the optimization, we introduce virtual queues and employ the Lyapunov optimization theory to transform the optimization with long-term average format into optimizing the drift-plus-penalty problem. The drift-plus-penalty is further upper bounded with variables only related to current time slot, which greatly simplifies the optimization problem. Theoretic analysis is also conducted to show that the optimal solution is limited by an upper bound that is independent of the operation time index. Finally, simulation results with real solar irradiance data show that the proposed algorithm can achieve much better performance than existing approaches based on Markov decision process and water-filling. Chengrun Qiu, Yang Hu 0006, Yan Chen 0007, Bing Zeng 0001 |
IEEE Internet Things J. | 4 |
| 2018 | DC Coefficient Estimation of Intra-Predicted Residuals in HEVCabstractThis paper presents a DC coefficient estimation algorithm for intra-predicted residual blocks in the High-Efficiency Video Coding (HEVC) standard. Discarding the DC coefficient directly in each transform block leads to substantial bit-saving, but at the same time produces strong discontinuities between neighboring blocks. To overcome this problem, we propose an estimation algorithm for the DC coefficient, which solves an optimal offset in a closed-form to recover the corresponding block edges. Then, we embed this algorithm into HEVC in its rate-distortion optimized quantization and sign bit hiding steps. Furthermore, a flag is signaled to decide whether the DC estimation strategy is used for each transform block. Test results under the common test condition show that our algorithm achieves 1.5% and 1.6% BD-rate reduction on average for luma and chroma, respectively, under all intra configuration. In the meantime, our simulation results show that both encoding time and decoding time increase only slightly (about 10%, without any special optimization on programming our proposed algorithm). When testing the proposed DC estimation algorithm on inter coding configurations, including low delay with P pictures, low delay with B pictures, and random access, we can also achieve 0.5%-1.1% bit-rate savings on average, while nearly no extra encoding and decoding time is needed. Chen Chen 0015, Zexiang Miao, Xiandong Meng, Shuyuan Zhu, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | MCast: High-Quality Linear Video Transmission With Time and Frequency DiversitiesabstractUncoded linear video transmission has recently attracted people's attention due to its capacity to provide robust and scalable transmission. However, in reality, with the fluctuation of the wireless channels, the received quality may not be good enough. In such a case, the data may need to be transmitted multiple times to exploit both the time and frequency diversities to improve the received quality. Such a problem has never been investigated in the literature of uncoded video transmission. To resolve the problem, in this paper, we propose a framework, named MCast, to utilize the time and frequency diversities to achieve high-quality linear video transmission. We study how to optimally allocate the power and assign the channels at each time slot to the source data such that the overall performance is maximized. Specifically, we first derive a closed-form optimal power allocation solution for any given channel assignment. With the optimal power allocation, we then propose a suboptimal channel assignment scheme, where we sort the channels with their gains and assign the channels one-by-one to the corresponding block that can reduce the most reconstruction error. Finally, we compare MCast system with four other systems that are based on Softcast and Parcast, and simulation results show that MCast system can achieve better performance in terms of both PSNR performance and visual quality. Chaofan He, Huiying Wang, Yang Hu 0006, Yan Chen 0007, Xiaopeng Fan 0001, Houqiang Li, Bing Zeng 0001 |
IEEE Trans. Image Process. | 7 |
| 2018 | Cross-Space Distortion Directed Color Image CompressionabstractTraditional color image compression is usually conducted in the YCbCr space but many color displayers only accept RGB signals as inputs. Due to the use of a non-unitary matrix in the YCbCr-RGB conversion, low distortion achieved in the YCbCr space cannot guarantee low distortion for the RGB signals. To solve this problem, we propose a novel compression scheme for color images through defining a cross-space distortion so as to reduce as much as possible the distortion in the RGB space. To this end, we first derive the relationship between the distortions in the YCbCr space and RGB space. Then, we develop two solutions to implement color image compression for the most popular 4:2:0 chroma format. The first solution focuses on the design of a new spatial downsampling method to generate the 4:2:0 YCbCr image for a high-efficiency compression. The second one provides a novel way to reduce the distortion of the compressed color image by controlling the quantization error of the 4:2:0 YCbCr image, especially the one generated by using the traditional spatial downsampling. Experimental results show that both proposed solutions offer a remarkable quality gain over some state-of-the-art approaches when tested on various textured color images. Shuyuan Zhu, Chen Chen 0015, Shuaicheng Liu, Bing Zeng 0001 |
IEEE Trans. Multim. | 5 |
| 2017 | Shape Recovery of Endoscopic Videos by Shape from Shading Using Mesh Regularization
Zhihang Ren, Lingbing Peng, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001 |
ICIG (3) | 6 |
| 2017 | Uncertain Region Identification for Stereoscopic Foreground Cutout
Taotao Yang, Shuaicheng Liu, Zhengning Wang, Bing Zeng 0001 |
ICIG (3) | 5 |
| 2017 | Meshflow video denoisingabstractWe propose an efficient video denoising approach that produces clean videos by utilizing the recently proposed meshflow motion model for the camera motion compensation. The meshflow is a spatially-smooth sparse motion field with motion vectors located at the mesh vertexes. The model is very effective and efficient for the purpose of the multiframes denoising due to its internal characteristics such as the lightweight, the nonparametric form, and the spatially-variant motion representation. Specifically, the meshflows are estimated between adjacent frames, which are used to align frames within a sliding time window. A denoised frame is generated by fusing of several registered frames in a spatial and temporal manner with outlier rejections. Various challenging examples demonstrate the effectiveness and practicability of the proposed approach. Zhihang Ren, Shuaicheng Liu, Bing Zeng 0001 |
ICIP | 4 |
| 2017 | Sampling for Approximate Maximum Search in Factorized TensorabstractFactorization models have been extensively used for recovering the missing entries of a matrix or tensor. However, directly computing all of the entries using the learned factorization models is prohibitive when the size of the matrix/tensor is large. On the other hand, in many applications, such as collaborative filtering, we are only interested in a few entries that are the largest among them. In this work, we propose a sampling-based approach for finding the top entries of a tensor which is decomposed by the CANDECOMP/PARAFAC model. We develop an algorithm to sample the entries with probabilities proportional to their values. We further extend it to make the sampling proportional to the $k$-th power of the values, amplifying the focus on the top ones. We provide theoretical analysis of the sampling algorithm and evaluate its performance on several real-world data sets. Experimental results indicate that the proposed approach is orders of magnitude faster than exhaustive computing. When applied to the special case of searching in a matrix, it also requires fewer samples than the other state-of-the-art method. Yang Hu 0006, Bing Zeng 0001 |
IJCAI | 3 |
| 2017 | Endoscopic video deblurring via synthesisabstractEndoscopic videos have been widely used for stomach diagnoses. However, endoscopic devices often capture videos with motion blurs, due to the dimly-lit environment and the camera shakiness during the capturing, which severely disturbs the diagnoses. In this paper, we present a framework that can restore blurry frames by synthesizing image details from the nearby sharp frames. Specifically, the blurry frame and their corresponding nearby sharp frames are identified according to the image gradient sharpness. To restore one blurry frame, a non-parametric mesh-based motion model is proposed to align the sharp frame to the blurry frame. The motion model leverages motions from image feature matches and optical flows, which yields high quality alignments to overcome challenges such as noisy, blurry, reflective and textureless interferences. After the alignment, the deblurred frame is synthesized by matching patches locally between the blurry frame and the aligned sharp frame. Without the estimation of blur kernels, we show that it is possible to directly compare a blurry patch against the sharp patches for the nearest neighbor matches in endoscopic images. The experiments demonstrate the effectiveness of our algorithm. Lingbing Peng, Shuaicheng Liu, Dehua Xie, Shuyuan Zhu, Bing Zeng 0001 |
VCIP | 5 |
| 2017 | Two-stage filtering of compressed depth images with Markov Random Field
Lijun Zhao 0002, Huihui Bai 0001, Anhong Wang, Yao Zhao 0001, Bing Zeng 0001 |
Signal Process. Image Commun. | 5 |
| 2017 | MMSE-Directed Linear Image Interpolation Based on Nonlocal Geometric SimilarityabstractIn this letter, we propose a minimum mean square error (MMSE) directed linear interpolation to compose the high-resolution image from a single low-resolution image. We build up our interpolation model by using some similar image patches selected according to the nonlocal geometric similarity. First, we use a two-stage search scheme to collect the matched patches inside the whole image. Second, a similarity scaling factor is used in the second search to refine the collected patches so as to help find a robust solution to the MMSE-directed interpolation. Third, our MMSE-directed interpolation is regularized by the involved reference patches to make the solved interpolation coefficients more reliable. Experimental results show that our proposed method outperforms the state-of-the-art MMSE-directed linear interpolation schemes and works competitively with the state-of-the-art learning-based ones. Shuyuan Zhu, Zhiying He, Shuaicheng Liu, Bing Zeng 0001 |
IEEE Signal Process. Lett. | 4 |
| 2017 | A New Block-Based Method for HEVC Intra CodingabstractThis paper presents a new block-based method for the High Efficiency Video Coding (HEVC) intra coding. First, we have found through analysis and test that the prediction errors on some pixels in each prediction block (PB) that are neighboring to the reference pixels would be no bigger than the corresponding coding errors. Based on this observation, the pixels in each PB are divided into two parts: half pixels are coded via a novel padding technique together with a constrained quantization algorithm (leading to around 3 dB gain under the same bit rate), whereas the other half are reconstructed by linear interpolations along a prediction direction by utilizing the neighboring reference pixels and the first half coded pixels. In the final implementation, a competition mechanism is employed between this new method and the original HEVC intra coding in order to choose the best mode for each PB. Experimental results show that about 2% BD-rate reduction has been achieved both for luma and chroma with respect to the original HEVC intra coding, whereas the encoder complexity increases by 130%, but the decoding time remains nearly unchanged. Chen Chen 0015, Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | A Hybrid Approach for Near-Range Video StabilizationabstractNear-range videos contain objects that are close to the camera. These videos often contain discontinuous depth variation (DDV), which is the main challenge to the existing video stabilization methods. Traditionally, 2D methods are robust to various camera motions (e.g., quick rotation and zooming) under scenes with continuous depth variation (CDV). However, in the presence of DDV, they often generate wobbled results due to the limited ability of their 2D motion models. Alternatively, 3D methods are more robust in handling near-range videos. We show that, by compensating rotational motions and ignoring translational motions, near-range videos can be successfully stabilized by 3D methods without sacrificing the stability too much. However, it is time-consuming to reconstruct the 3D structures for the entire video and sometimes even impossible due to rapid camera motions. In this paper, we combine the advantages of 2D and 3D methods, yielding a hybrid approach that is robust to various camera motions and can handle the near-range scenarios well. To this end, we automatically partition the input video into CDV and DDV segments. Then, the 2D and 3D approaches are adopted for CDV and DDV clips, respectively. Finally, these segments are stitched seamlessly via a constrained optimization. We validate our method on a large variety of consumer videos. Shuaicheng Liu, Binhan Xu, Chuang Deng, Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | CodingFlow: Enable Video Coding for Video StabilizationabstractVideo coding focuses on reducing the data size of videos. Video stabilization targets at removing shaky camera motions. In this paper, we enable video coding for video stabilization by constructing the camera motions based on the motion vectors employed in the video coding. The existing stabilization methods rely heavily on image features for the recovery of camera motions. However, feature tracking is time-consuming and prone to errors. On the other hand, nearly all captured videos have been compressed before any further processing and such a compression has produced a rich set of block-based motion vectors that can be utilized for estimating the camera motion. More specifically, video stabilization requires camera motions between two adjacent frames. However, motion vectors extracted from video coding may refer to non-adjacent frames. We first show that these non-adjacent motions can be transformed into adjacent motions such that each coding block within a frame contains a motion vector referring to its adjacent previous frame. Then, we regularize these motion vectors to yield a spatially-smoothed motion field at each frame, named as CodingFlow, which is optimized for a spatially-variant motion compensation. Based on CodingFlow, we finally design a grid-based 2D method to accomplish the video stabilization. Our method is evaluated in terms of efficiency and stabilization quality, both quantitatively and qualitatively, which shows that our method can achieve high-quality results compared with the state-of-the-art methods (feature-based). Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001 |
IEEE Trans. Image Process. | 4 |
| 2017 | A Hierarchical Approach for Rain or Snow Removing in a Single Color ImageabstractIn this paper, we propose an efficient algorithm to remove rain or snow from a single color image. Our algorithm takes advantage of two popular techniques employed in image processing, namely, image decomposition and dictionary learning. At first, a combination of rain/snow detection and a guided filter is used to decompose the input image into a complementary pair: 1) the low-frequency part that is free of rain or snow almost completely and 2) the high-frequency part that contains not only the rain/snow component but also some or even many details of the image. Then, we focus on the extraction of image's details from the high-frequency part. To this end, we design a 3-layer hierarchical scheme. In the first layer, an overcomplete dictionary is trained and three classifications are carried out to classify the high-frequency part into rain/snow and non-rain/snow components in which some common characteristics of rain/snow have been utilized. In the second layer, another combination of rain/snow detection and guided filtering is performed on the rain/snow component obtained in the first layer. In the third layer, the sensitivity of variance across color channels is computed to enhance the visual quality of rain/snow-removed image. The effectiveness of our algorithm is verified through both subjective (the visual quality) and objective (through rendering rain/snow on some ground-truth images) approaches, which shows a superiority over several state-of-the-art works. Yinglong Wang 0002, Shuaicheng Liu, Chen Chen 0015, Bing Zeng 0001 |
IEEE Trans. Image Process. | 4 |
| 2016 | MeshFlow: Minimum Latency Online Video Stabilization
Shuaicheng Liu, Ping Tan 0002, Lu Yuan 0001, Jian Sun 0001, Bing Zeng 0001 |
ECCV (6) | 5 |
| 2016 | Low bit-rate intra coding scheme based on constrained quantization and median-type filterabstractThis paper presents a new intra coding scheme for low bit-rate video compression. We first propose an improved codec architecture based on HEVC encoder. Then we divide pixels in each prediction block into two parts: three quarters of pixels are coded via a smart padding technique together with a constrained quantization algorithm (leading to a significantly improved quality); whereas the other quarter are reconstructed according to median-type filtering by utilizing the 8-neighboring reference samples after all blocks have been encoded and reconstructed. Experimental results show that about 3% BD-rate reduction has been achieved both for luma and chroma components without apparent increase of encoding complexity with respect to the original HEVC intra coding. Chen Chen 0015, Bing Zeng 0001 |
ICASSP | 2 |
| 2016 | Joint bundled camera paths for stereoscopic video stabilizationabstractThis paper presents a method to stabilize shaky stereoscopic videos captured by hand-held devices. Directly applying traditional monocular video stabilization techniques to two views independently is problematic as it often brings undesirable vertical disparities and produces inaccurate horizontal disparities, which violate original stereoscopic disparity constraints, leading to erroneous depth perception. In this paper, we show that monocular video stabilization methods, such as the bundled camera paths stabilization, can be extended for stereoscopic videos by taking additional disparity constraints during the stabilization. In particular, we first estimate disparities between two views. Then, we compute camera motions as meshes of bundled paths for each view. Next, we smooth paths of two views separately and iteratively. During each iteration, we adjust the meshes of one view by our proposed `Joint Disparity and Stability mesh Warp (JDSW)'. The final result is generated after several iterations of paths smoothing and meshes adjusting, in which temporal stability and correct depth perception are achieved simultaneously. We evaluate our method by various challenging stereoscopic videos with different camera motions and scene types. The experiments demonstrate the effectiveness of our method. Heng Guo 0003, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001 |
ICIP | 4 |
| 2016 | A framework of single-image deraining method based on analysis of rain characteristicsabstractIn this paper, we propose an algorithm to remove rain streaks from single color image. Firstly, the guided filter, cooperated with rain pixels detection are used to separate a color image into low-frequency and high-frequency parts so that most rain components exist in the high-frequency part. Then, we focus on the high-frequency part to extract the non-rain details according to the characteristics of the rain in which a dictionary learning method is used. Meanwhile, to enhance the quality of the rain-removed image, the proposed principal direction of an image patch (PDIP) and the sensitivity of variance of color channels (SVCC) are employed in our work to help extract more non-rain details. Compared with the state-of-the-art works, our proposed method can remove the rain (especially heavy rain) from color images more efficiently. Yinglong Wang 0002, Chen Chen 0015, Shuyuan Zhu, Bing Zeng 0001 |
ICIP | 4 |
| 2016 | Intrinsic decomposition for stereoscopic imagesabstractIntrinsic image decomposition is an important technique that decomposes an image into reflectance and shading components. In this paper, we enable intrinsic decomposition for stereoscopic images. Traditional approaches cannot be directly applied to decompose stereoscopic images, yielding inconsistent reflectance and 3D artifacts after recoloring. To solve this problem, we propose a straight yet effective method for stereoscopic intrinsic decomposition, which consists of classical retinex constraint as well as disparity constraint. The former encodes the shading smoothness prior while the latter controls the reflectance similarity between two views. To further reduce ambiguity, we employ local and non-local texture cues by using superpixels within and across two views. The experiments show that our method can effectively decompose stereoscopic images with high quality and offer a comfortable 3D viewing experience. Dehua Xie, Shuaicheng Liu, Kaimo Lin, Shuyuan Zhu, Bing Zeng 0001 |
ICIP | 5 |
| 2016 | Constrained quantization based transform domain down-conversion for image compressionabstractThe image down-conversion may be used in the block-based image compression because it can help save lots of bit-counts for each individual block. A straightforward way to implement the transform domain down-conversion is to truncate some high-frequency components to get a down-sized coefficient block. However, directly using this down-sized coefficient block to reconstruct a completed image block will lead to a serious quality degradation. In this paper, we propose a constrained quantization based transform domain down-conversion (CQTDD) to help compress each 16×16 macro-block and it makes the coding quality of 1/4 selected pixels (according to a regular pattern) in each macro-block much higher than that can be achieved by using the traditional truncation based approach. Meanwhile, the other 3/4 pixels will be interpolated by using those 1/4 well-reconstructed pixels. Furthermore, these 1/4 pixels are optimized before the compression to help get a more efficient interpolation. Finally, the proposed CQTDD works with the JPEG baseline coding together as two candidate coding modes in our proposed compression scheme. Experimental results demonstrate that our proposed method may offer a remarkable quality gain, both objectively and subjectively, compared with some existing methods. Shuyuan Zhu, Liaoyuan Zeng, Bing Zeng 0001, Jiantao Zhou 0001 |
ISCAS | 3 |
| 2016 | Automatic Reflection Removal using Gradient Intensity and Motion CuesabstractWe present a method to separate the background image and reflection from two photos that are taken in front of a transparent glass under slightly different viewpoints. In our method, the SIFT-flow between two images is first calculated and a motion hierarchy is constructed from the SIFT-flow at multiple levels of spatial smoothness. To distinguish background edges and reflection edges, we calculate a motion score for each edge pixel by its variance along the motion hierarchy. Alternatively, we make use of the so-called superpixels to group edge pixels into edge segments and calculate the motion scores by averaging over each segment. In the meantime, we also calculate an intensity score for each edge pixel by its gradient magnitude. We combine both motion and intensity scores to get a combination score. A binary labelling (for separation) can be obtained by thresholding the combination scores. The background image is finally reconstructed from the separated gradients. Compared to the existing approaches that require a sequence of images or a video clip for the separation, we only need two images, which largely improves its feasibility. Various challenging examples are tested to validate the effectiveness of our method. Shuaicheng Liu, Taotao Yang, Bing Zeng 0001, Zhengning Wang, Guanghui Liu 0001 |
ACM Multimedia | 4 |
| 2016 | DC coefficient estimation of intra-predicted residuals in high efficiency video codingabstractThis paper proposes a DC coefficient estimation algorithm for intra-predicted residual blocks in the High Efficiency Video Coding (HEVC) standard. Discarding the DC coefficient in the current coding block leads to a substantial bit-saving but produces at the same time strong discontinuities between this block and its neighboring reconstructed blocks. To overcome this problem, we propose an estimation algorithm for the DC coefficient, which solves an optimal offset in a closed-form in the pixel domain to recover the corresponding block edges. Test results show that our algorithm achieves 1.0% and 1.4% BD-rate reduction on average for luma and chroma as compared with HM-16.6, respectively, when the sign-bit-hiding (SBH) technique is disabled. When SDH is set on, namely under the common test condition (CTC), the BD-rate reduction drops slightly to 0.7% and 1.1% for luma and chroma, respectively. In the meantime, the test results show that both encoding time and decoding time increase only slightly (about 10%, without any special optimization on programming our proposed algorithm). Chen Chen 0015, Zexiang Miao, Xiandong Meng, Shuyuan Zhu, Bing Zeng 0001 |
VCIP | 5 |
| 2016 | Mode-dependent transforms based on elliptical model for high efficiency video codingabstractHigh efficiency video coding (HEVC) defines 35 prediction modes in its intra prediction stage to signal the direction information of residual blocks. Traditionally, separable two-dimension (2-D) transforms (integer DCT and DST) are utilized in a similar manner as in the previous H.264/AVC standards. However, such 2-D transforms cannot yield the best energy compaction for a 2-D directional source where the dominating directional information is other than the horizontal or vertical one. In order to overcome this drawback, we build an elliptical model with directionality and design some non-separable transforms based on the Karhunen-Loeve transform in this paper. Specifically, we derive a non-separable transform in closed-form for each intra-prediction mode and replace the default transform in HEVC. Simulation results reveal that 1.7% and 2.0% on average and up to 7.7% and 8.1% BD-rate reduction can be achieved for luma and chroma component, respectively. In the meantime, the test results show that both the encoding time and decoding time increase only about 5%. Kaiyuan Jia, Chen Chen 0015, Xiandong Meng, Shuyuan Zhu, Bing Zeng 0001 |
VCIP | 5 |
| 2016 | Geometry-based PSF estimation and deblurring of defocused images with depth informationabstractWe propose in this paper an algorithm to recover the blurred details of an image caused by defocusing during the photo-taking. Our algorithm takes one RGB image as well as its depth map as the input. We build up a model in which each captured image pixel is regarded as a light-emitting source that goes through a synthetic camera system. Thanks to the depth map, we have the geometrical information of the scene so that the point spread function (PSF) can be derived for each pixel more accurately as compared to conventional approaches where only RGB images are involved. Then, we make use of the derived PSFs to solve an optimization so as to reconstruct an all-in-focus image. The reconstructed results are evaluated by comparison with the original all-in-focus images. Compared to other methods for the deblurring of defocused images, our method shows a better recovery of image details. Yiqun Wu 0003, Bing Zeng 0001, Dehua Xie, Shuaicheng Liu |
VCIP | 3 |
| 2016 | Interpolation-directed transform domain downward conversion for block-based image compressionabstractIn this paper, we design an interpolation-directed transform domain downward conversion (ITDDC) to build up a new block-based image compression scheme. This ITDDC is derived from our proposed 2-D padding and performed on each 16×16 macro-block of pixels to convert it into an 8×8 coefficient block, leading to a downward image conversion in the transform domain. More interestingly, the further compression is just performed on the down-sized coefficient block and the reconstruction for an entire macro-block is achieved via the interpolation by using the decoded pixels only locating in some specific positions of it. To make the interpolation more efficient, the pixels participating in the interpolation will be optimized before the compression. The ITDDC-based coding is used competitively with the JPEG baseline coding to compress each macro-block in our proposed compression scheme according to a simple but efficient rate-distortion optimization based criterion. Experimental results demonstrate that our proposed method gets a remarkable quality gain over the existing approaches. Shuyuan Zhu, Jinglin Yu, Chen Chen 0015, Liaoyuan Zeng, Bing Zeng 0001 |
VCIP | 6 |
| 2016 | Seamless Video Stitching from Hand-held Camera InputsabstractAbstract Images/videos captured by portable devices (e.g., cellphones, DV cameras) often have limited fields of view. Image stitching, also referred to as mosaics or panorama, can produce a wide angle image by compositing several photographs together. Although various methods have been developed for image stitching in recent years, few works address the video stitching problem. In this paper, we present the first system to stitch videos captured by hand‐held cameras. We first recover the 3D camera paths and a sparse set of 3D scene points using CoSLAM system, and densely reconstruct the 3D scene in the overlapping regions. Then, we generate a smooth virtual camera path, which stays in the middle of the original paths. Finally, the stitched video is synthesized along the virtual path as if it was taken from this new trajectory. The warping required for the stitching is obtained by optimizing over both temporal stability and alignment quality, while leveraging on 3D information at our disposal. The experiments show that our method can produce high quality stitching results for various challenging scenarios. Kaimo Lin, Shuaicheng Liu, Loong Fah Cheong, Bing Zeng 0001 |
Comput. Graph. Forum | 4 |
| 2016 | Localized Low-Rank Promoting for Recovery of Block-Sparse Signals With Intrablock CorrelationabstractWe consider the problem of recovering block-sparse signals with intrablock correlated entries. The block partition of the sparse signal is assumed unknown a priori. To exploit the block-sparse structure as well as the local smoothness of the sparse signal, consecutive coefficients of the sparse signal are organized into a number of 2×2 matrices, and the log-determinant function is used to promote the low rankness of these 2×2 matrices. We show that such a log-determinant function has the ability to promote the block-sparsity and local smoothness simultaneously. An iterative reweighted method is developed by iteratively minimizing a surrogate function of the original objective function. Simulation results show that our proposed method offers competitive performance for recovering block-sparse signals with intrablock correlated entries. Linxiao Yang, Jun Fang 0001, Hongbin Li 0001, Bing Zeng 0001 |
IEEE Signal Process. Lett. | 4 |
| 2016 | New R-D Optimization Criterion for Fast Mode Decision Algorithms in Video Coding and TransratingabstractMode decision has a significant effect on the quality and complexity of video coding. It is even more challenging when generating multiple bitstreams with different bitrates (BRs) in, for example, dynamic adaptive streaming for HTTP or transrating systems. Full search and simplified fast search mode decision (MD) methods either suffer from a high computational complexity or have a negative impact on quality. Furthermore, mode selection in conventional approaches strongly depends on the quantization parameter (QP). Hence, modes that have been selected for high BR compression may not be suitable for low BR when transrating a bitstream. In this paper, we propose a rate-distortion (R-D)-optimized criterion for fast MD algorithms. The proposed cost function, when adopted in different fast MD algorithms, not only improves the R-D performance by up to 6.6% in terms of Bjøntegaard delta rate, but also reduces the execution time of the encoder by up to 6.8%. We also show that modes selected by the proposed criterion are less sensitive to changes in BR or QP. As a result, the same modes in an encoded bitstream may be used even after transrating using requantization, resulting in a significant R-D performance improvement of up to 33.3%. Alireza Aminlou, Mahmoud Reza Hashemi, Moncef Gabbouj, Bing Zeng 0001, Omid Fatemi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Blind Image Quality Assessment Based on Multichannel Feature Fusion and Label TransferabstractIn this paper, we propose an efficient blind image quality assessment (BIQA) algorithm, which is characterized by a new feature fusion scheme and a k-nearest-neighbor (KNN)-based quality prediction model. Our goal is to predict the perceptual quality of an image without any prior information of its reference image and distortion type. Since the reference image is inaccessible in many applications, the BIQA is quite desirable in this context. In our method, a new feature fusion scheme is first introduced by combining an image's statistical information from multiple domains (i.e., discrete cosine transform, wavelet, and spatial domains) and multiple color channels (i.e., Y, Cb, and Cr). Then, the predicted image quality is generated from a nonparametric model, which is referred to as the label transfer (LT). Based on the assumption that similar images share similar perceptual qualities, we implement the LT with an image retrieval procedure, where a query image's KNNs are searched for from some annotated images. The weighted average of the KNN labels (e.g., difference mean opinion score or mean opinion score) is used as the predicted quality score. The proposed method is straightforward and computationally appealing. Experimental results on three publicly available databases (i.e., LIVE II, TID2008, and CSIQ) show that the proposed method is highly consistent with human perception and outperforms many representative BIQA metrics. Qingbo Wu 0001, Hongliang Li 0001, Fanman Meng, King Ngi Ngan, Bing Luo 0003, Chao Huang 0003, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2016 | Joint Video Stitching and Stabilization From Moving CamerasabstractIn this paper, we extend image stitching to video stitching for videos that are captured for the same scene simultaneously by multiple moving cameras. In practice, videos captured under this circumstance often appear shaky. Directly applying image stitching methods for shaking videos often suffers from strong spatial and temporal artifacts. To solve this problem, we propose a unified framework in which video stitching and stabilization are performed jointly. Specifically, our system takes several overlapping videos as inputs. We estimate both inter motions (between different videos) and intra motions (between neighboring frames within a video). Then, we solve an optimal virtual 2D camera path from all original paths. An enlarged field of view along the virtual path is finally obtained by a space-temporal optimization that takes both inter and intra motions into consideration. Two important components of this optimization are that: 1) a grid-based tracking method is designed for an improved robustness, which produces features that are distributed evenly within and across multiple views and 2) a mesh-based motion model is adopted for the handling of the scene parallax. Some experimental results are provided to demonstrate the effectiveness of our approach on various consumer-level videos and a Plugin, named "Video Stitcher" is developed at Adobe After Effects CC2015 to show the processed videos. Heng Guo 0003, Shuaicheng Liu, Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
IEEE Trans. Image Process. | 5 |
| 2016 | Region-Aware 3-D Warping for DIBRabstractIn 3-D video (3DV) applications, depth-image-based rendering (DIBR) has been widely employed to synthesize virtual views. However, this approach is performed in a frame-based way, meaning each whole frame is dealt with and the characteristics of different regions in the frame are ignored. As a result, redundant pixels in some regions are abused during the subsequent warping and blending stage. This paper proposes a region-aware 3-D warping approach for DIBR in which warped frames are reasonably divided beforehand so that only the indispensable regions are used. With the proposed scheme, it is possible to avoid noneffective and repeated pixels during the warping stage. In addition, the blending process is also saved. The experimental results show that compared to the state-of-the-art VSRS3.5 and VSRS-1D-fast algorithms, our approach can achieve significant computation savings without sacrificing synthesis quality. Anhong Wang, Yao Zhao 0001, Chunyu Lin, Bing Zeng 0001 |
IEEE Trans. Multim. | 5 |
| 2016 | Image Interpolation Based on Non-local Geometric Similarities and Directional GradientsabstractImage interpolation offers an efficient way to compose a high-resolution (HR) image from the observed low-resolution (LR) image. Advanced interpolation techniques design the interpolation weighting coefficients by solving a minimum mean-square-error (MMSE) problem in which the local geometric similarity is often considered. However, using local geometric similarities cannot usually make the MMSE-based interpolation as reliable as expected. To solve this problem, we propose a robust interpolation scheme by using the nonlocal geometric similarities to construct the HR image. In our proposed method, the MMSE-based interpolation weighting coefficients are generated by solving a regularized least squares problem that is built upon a number of dual-reference patches drawn from the given LR image and regularized by the directional gradients of these patches. Experimental results demonstrate that our proposed method offers a remarkable quality improvement as compared to some state-of-the-art methods, both objectively and subjectively. Shuyuan Zhu, Bing Zeng 0001, Liaoyuan Zeng, Moncef Gabbouj |
IEEE Trans. Multim. | 2 |
| 2015 | Synthesis of light-field raw data from RGB-D imagesabstractImage raw data captured by light-field cameras such as Lytro and Raytrix have been strictly protected by the manufacturers, which results in a serious hurdle to most academic researchers working in this area. In this paper, we make an attempt of developing an algorithm that can generate light-field raw data from a single input picture captured by a conventional camera together with its depth map. Our algorithm follows closely the imaging procedure of the light-field camera with micro lens, with certain simplification. Specifically, an input image (of size W × H) will be projected to a light-field data of size (M × P) × (N × P), assuming that the micro-lens array has size M × N micro lens and each micro lens is supported by P × P sensors. In our work, the main lens is partitioned into K × K cells, each pixel of the input picture is assumed to locate at certain distance from the main lens and regarded as a lighting source. The directional light emitted from each pixel goes through the main lens and the corresponding micro-lens to be recorded at a particular sensor in the CMOS array. We will derive the geometry of this imaging process as well as the intensity recorded at each sensor. We choose to run the refocusing algorithm to verify whether the light-field data generated by our algorithm is meaningful or not. Comparing the refocused images against the original input picture shows that our synthesis algorithm seems quite successful. Yiqun Wu 0003, Bing Zeng 0001 |
ICIP | 3 |
| 2015 | Image interpolation based on non-local geometric similaritiesabstractImage interpolation refers to constructing a high-resolution (HR) image from a low-resolution (LR) image. Traditionally, an HR image can be produced from an observed LR image via the polynomial-based interpolation (bi-linear or bi-cubic interpolations, involving a small number of neighbors around each interpolated position). The advanced interpolation makes use of the so-called “geometric similarity” to design a set of optimal interpolation weighting coefficients. However, better geometric similarities can perhaps be found from a non-local area within the LR source image or even from other but similar images (possibly with higher resolutions). Based on this fact, we propose in this paper a non-local geometric similarity based interpolation scheme to construct HR images. In our proposed method, optimal weighting coefficients are determined by solving a regularized least squares problem which is built upon a number of dual reference patches drawn from the observed LR image and regularized by the variation of directional gradients of the image patch. Experimental results demonstrate that our proposed method offers a remarkable quality improvement, both objectively and subjectively. Shuyuan Zhu, Bing Zeng 0001, Guanghui Liu 0001, Liaoyuan Zeng, Moncef Gabbouj |
ICME | 2 |
| 2015 | Adaptive guided image filter for improved in-loop filtering in video codingabstractThis paper proposes a new adaptive sharpening filter based on guided image filter and improves HEVC's in-loop filter architecture by embedding sharpening filter between deblocking filter and SAO. The proposed algorithm classifies pixels of a frame into several groups according to uniform quantization of each pixel's Sum-Modified-Laplacian value and assigns identical optimal filtering parameters to the pixels belonging to the same group based on rate-distortion optimization. Simulation results show that our proposed algorithm achieves 0.7% on average and up to 8% BD-rate reduction with respect to the original HEVC in-loop filtering method. Encoding time increases slightly by about 15% and decoding time increases by 70% on average without special optimization of C++ program integrated in HM-16.5. Chen Chen 0015, Zexiang Miao, Bing Zeng 0001 |
MMSP | 3 |
| 2015 | Demo: Achieving Simultaneous Screen-Human Viewing and Hidden Screen-Camera CommunicationabstractWe present and demonstrate INFRAME++, a novel system that enables concurrent, dual-mode, full-frame communication for both users and devices. It achieves unobtrusive screen-camera data communication without affecting the primary video-viewing experience for human users. It leverages the capability discrepancy and distinctive features of the human vision system and devices (modern display and camera). We have implemented INFRAME++ as a PC-phone application. Both communication will be realized through multiplexed videos frames which will be displayed on a modern monitor (120FPS) and captured by a smartphone camera for data decoding. In this demonstration (Figure 1), we will show that INFRAME++ can yield normal video-viewing experience for humans, and high-rate data communication for devices (up to 300 kbps). User participation will be welcome in this live demo. Anran Wang 0002, Gan Fang, Chunyi Peng 0001, Guobin Shen, Bing Zeng 0001 |
MobiSys | 6 |
| 2015 | InFrame++: Achieve Simultaneous Screen-Human Viewing and Hidden Screen-Camera CommunicationabstractRecent efforts in visible light communication over screen-camera links have exploited the display for data communication. Such practices, albeit convenient, have led to contention between space allocated for users and content reserved for devices, in addition to their visual anti-aesthetics and distractedness. In this paper, we propose INFRAME++, a system that enables concurrent, dual-mode, full-frame communication for both users and devices. INFRAME++ leverages the spatial-temporal flicker-fusion property of human vision system and the fast frame rate of modern display. It multiplexes data onto full-frame video contents through novel complementary frame composition, hierarchical frame structure, and CDMA-like modulation. It thus ensures opportunistic and unobtrusive screen-camera data communication without affecting the primary video-viewing experience for human users. Our prototype and experiments have confirmed its effectiveness of delivering data to devices in its visual communication with imperceptible video artifacts for viewers. INFRAME++ is able to achieve 150-240 kbps at 120FPS over a 24? LCD monitor with one data frame per 12 display frames. It supports up to 360kbps while data:video is 1:6. Anran Wang 0002, Chunyi Peng 0001, Guobin Shen, Gan Fang, Bing Zeng 0001 |
MobiSys | 6 |
| 2015 | Adaptive sampling for compressed sensing based image compression
Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Constrained Directed Graph Clustering and Segmentation Propagation for Multiple Foregrounds CosegmentationabstractThis paper proposes a new constrained directed graph clustering (DGC) method and segmentation propagation method for the multiple foreground cosegmentation. We solve the multiple object cosegmentation with the perspective of classification and propagation, where the classification is used to obtain the object prior of each class and the propagation is used to propagate the prior to all images. In our method, the DGC method is designed for the classification step, which adds clustering constraints in cosegmentation to prevent the clustering of the noise data. A new clustering criterion such as the strongly connected component search on the graph is introduced. Moreover, a linear time strongly connected component search algorithm is proposed for the fast clustering performance. Then, we extract the object priors from the clusters, and propagate these priors to all the images to obtain the foreground maps, which are used to achieve the final multiple objects extraction. We verify our method on both the cosegmentation and clustering tasks. The experimental results show that the proposed method can achieve larger accuracy compared with both the existing cosegmentation methods and clustering methods. Fanman Meng, Hongliang Li 0001, Shuyuan Zhu, Bing Luo 0003, Chao Huang 0003, Bing Zeng 0001, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2014 | InFrame: Multiflexing Full-Frame Visible Communication Channel for Humans and DevicesabstractRecent efforts in visible light communication over screen-camera links have exploited the display for data communications. Such practices, albeit convenient, have led to contention between space allocated for users and content reserved for devices, in addition to their aesthetic issues and distractive nature. In this paper, we propose InFrame--a system that enables dual-mode full-frame communication for both humans and devices simultaneously. InFrame leverages the temporal flick-fusion property of human vision system and the fast frame rate of modern display. It multiplexes data onto full-frame video contents through a novel complementary frame design and several other techniques. It thus ensures screen-camera data communication without affecting the primary video-viewing experience for human users. Our preliminary experiments have confirmed that InFrame can achieve about 12.8kbps data rate with imperceptible video artifacts when being played back at 120FPS. Anran Wang 0002, Chunyi Peng 0001, Ouyang Zhang, Guobin Shen, Bing Zeng 0001 |
HotNets | 5 |
| 2014 | Fast and efficient inter CU decision for high efficiency video codingabstractIn this paper, a graph cut based fast Coding Unit (CU) decision algorithm is proposed for HEVC inter frames. Firstly, a feature called pyramid variance of the absolute difference (PVAD) is designed for the CU selection. Secondly, the CU decision is modeled as a Markov Random Field (MRF) inference problem, which can be optimized by the graph cut algorithm. Thirdly, a maximum a posteriori (MAP) approach based on the R-D cost is conducted to evaluate whether the unsplit CUs should be further split or not. Experimental results show the effectiveness of the proposed method. Jian Xiong 0005, Hongliang Li 0001, Fanman Meng, Bing Zeng 0001, Shuyuan Zhu, Qingbo Wu 0001 |
ICIP | 4 |
| 2014 | Downward spatially-scalable image reconstruction based on compressed sensingabstractAccording to the compressed sensing (CS) theory, we can sample a sparse signal at a rate that is (much) lower than the required Nyquist rate, while still enabling a nearly exact reconstruction. Image signals are sparse when represented in a certain domain, and because of this, a large number of CS-based image sampling and reconstruction techniques have been developed recently. In this paper, we focus on the design of the downward spatially-scalable image reconstruction from the CS-sampled data. Traditional methods usually reconstruct an image whose size is the same as the original source image and then achieve the downward scalability through sub-sampling. In our proposed method, we unify these two steps into a single one and promise to deliver a much improved quality. Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
ICIP | 2 |
| 2014 | Using mid-high level cues to detect salient objectabstractThis paper proposes a novel saliency object detection method by using the mid-level and high-level visual cues. In the mid-level objectness evaluation, we generate three complementary saliency maps, such as the multi-scale segmentation cue, the background cue and the spatial color distribution cue. The first cue is used to highlight the objects via the local region segment. The second cue uses the background priors to detect the saliency information. The third cue is to capture the spatial color distribution. For the high-level visual cue, we propose an objectness evaluation model to distinguish the object and the background. All the saliency cues are finally combined to achieve the saliency detection. The experimental results show that the proposed method outperforms the state-of-the-art saliency object detection methods. Hongliang Li 0001, Yurui Xie, Bing Luo 0003, Liangzhi Tang, Bing Zeng 0001, King Ngi Ngan, Fanman Meng |
ICME | 5 |
| 2014 | Adaptive sampling for compressed sensing based image compressionabstractThe compressed sensing (CS) theory shows that a sparse signal can be recovered at a sampling rate that is (much) lower than the required Nyquist rate. In practice, many image signals are sparse in a certain domain, and because of this, the CS theory has been successfully applied to the image compression in the past few years. The most popular CS-based image compression scheme is the block-based CS (BCS). In this paper, we focus on the design of an adaptive sampling mechanism for the BCS through a deep analysis of the statistical information of each image block. Specifically, this analysis will be carried out at the encoder side (which needs a few overhead bits) and the decoder side (which requires a feedback to the encoder side), respectively. Two corresponding solutions will be compared carefully in our work. We also present experimental results to show that our proposed adaptive method offers a remarkable quality improvement compared with the traditional BCS schemes. Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
ICME | 2 |
| 2014 | Favorite object extraction using web imagesabstractIn this paper, we propose a framework to discover and segment favorite object from the natural images. The main idea is to first generate the shape based common template of the favorite object using the images collected from the web. Then, the common template is used to extract the favorite object from the original images. In the common template generation, co-segmentation is used to provide the initial segments. The median graph theory is employed to construct the common template. We also propose a new shape descriptor namely directional shape representation to handle shape variations. We test our method on the images collected from image datasets and web. Experimental results demonstrate the effectiveness of the proposed method. Fanman Meng, Bing Luo 0003, Chao Huang 0003, Liangzhi Tang, Bing Zeng 0001, Nini Rao |
ISCAS | 5 |
| 2014 | Cosegmentation from similar backgroundsabstractRecently, the common objects are often required to be extracted from a group of images in many applications, such as video coding and model training. Co-segmentation is a new and efficient method for this requirement. In realistic applications, we observe that the images usually contain similar backgrounds (namely similar scene co-segmentation), such as the city landmark images collected from the web or the key frames sampled from a video. Meanwhile, the existing co-segmentation has not paid so much attention on the similar scene co-segmentation, and the insufficiently accurate segments may be provided by the existing methods. In this paper, we propose an active contours based co-segmentation model to provide foregrounds from the similar backgrounds. We combine the background consistency constraint with the foreground consistency constraint to form the energy function, and use the method of level-set and the calculus of variations to minimize the model. We also speed up the model by the hierarchical structure and the superpixel technique. We test the method on both the image and video dataset. The results show that the proposed model can obtain larger IOU values than the state-of-the-art co-segmentation methods. Fanman Meng, Hongliang Li 0001, King Ngi Ngan, Bing Zeng 0001, Nini Rao |
ISCAS | 4 |
| 2014 | Texture classification using joint statistical representation in space-frequency domain with local quantized patternsabstractDespite its success in texture analysis, Local Binary Pattern (LBP) is operated in the original image space, and it fails to capture deeper pixel interactions to provide a more discriminative description. In this paper, we propose to explore the joint statistical representation in the space-frequency domain with local quantized patterns for texture classification. The proposed method consists of two channels. In each channel, the multi-resolution spatial filters are employed to generate multi-scale spatial maps and the local Fourier transform is subsequently applied to extract local frequency features (spectral maps). The global thresholding is adopted to quantize the spatial and spectral maps into different levels, which are then jointly encoded to built a space-frequency co-occurrence histogram. Finally, the two-channel feature histograms are combined to represent the texture. Experiments on the Outex texture database demonstrate the robustness of our method to image rotation and illumination changes, and our method outperforms the state of the art in terms of the classification accuracy. Tiecheng Song, Hongliang Li 0001, Bing Zeng 0001, Moncef Gabbouj |
ISCAS | 3 |
| 2014 | No reference image quality metric via distortion identification and multi-channel label transferabstractIn this paper, we propose a no reference image quality assessment (NR-IQA) algorithm based on distortion identification (DI) and multi-channel label transfer (LT). First, the distortion type classification is used to obtain the query image's probabilities of belonging to each distortion type. Then, the distortion specific label transfer is implemented in multiple distortion category channels. Based on the hypothesis that the similar images share the similar subjective qualities, the label transfer predicts the subjective quality of the query image by pooling the labels of its k-nearest neighbors (KNN) retrieved from the annotated samples. A weighting average of the multi-channel label transfer's outputs is computed to obtain the final perceptual quality score. The weight is the query image's probability that belongs to the corresponding distortion type. The experimental results show that the proposed method outperforms representative NR-IQA approaches and some full-reference metrics. Qingbo Wu 0001, Hongliang Li 0001, King Ngi Ngan, Bing Zeng 0001, Moncef Gabbouj |
ISCAS | 4 |
| 2014 | Adaptive reweighted compressed sensing for image compressionabstractAccording to the compressed sensing (CS) theory, a signal that is sparse in a certain domain can be nearly exactly recovered from a few measurements where the sampling rate is lower than the Nyquist rate. This theory has been successfully applied to the image compression in the past few years as most image signals are highly sparse. In this paper, we apply an adaptive sampling mechanism to the reweighted block-based CS (BCS). The proposed adaptive sampling allocates the measurements to each image block according to the statistical information of the block so as to sample and recover the image more efficiently. Experimental results demonstrate that our adaptive reweighted method offers a very significant quality improvement compared with the traditional BCS schemes, including the non-reweighted and reweighted ones. Shuyuan Zhu, Bing Zeng 0001, Moncef Gabbouj |
ISCAS | 2 |
| 2014 | Wireless multicasting of video signals based on distributed compressed sensing
Anhong Wang, Bing Zeng 0001 |
Signal Process. Image Commun. | 2 |
| 2014 | Noise-Robust Texture Description Using Local Contrast Patterns via Global MeasuresabstractThis letter presents a noise-robust descriptor by exploring a set of local contrast patterns (LCPs) via global measures for texture classification. To handle image noise, the directed and undirected difference masks are designed to calculate three types of local intensity contrasts: directed, undirected, and maximum difference responses. To describe pixel-wise features, these responses are separately quantized and encoded into specific patterns based on different global measures. These resulting patterns (i.e., LCPs) are jointly encoded to form our final texture representation. Experiments are conducted on the well-known Outex and CUReT databases in the presence of high levels of noise. Compared to many state-of-the-art methods, the proposed descriptor achieves superior texture classification performance while enjoying a compact feature representation. Tiecheng Song, Hongliang Li 0001, Fanman Meng, Qingbo Wu 0001, Bing Luo 0003, Bing Zeng 0001, Moncef Gabbouj |
IEEE Signal Process. Lett. | 6 |
| 2014 | Perceptual Encryption of H.264 Videos: Embedding Sign-Flips Into the Integer-Based TransformsabstractAn alternative-transforms-based scheme has recently been proposed to achieve perceptual encryption of video signals in which multiple transforms are designed by using different rotation angles at the final stage of the discrete cosine transforms (DCTs) butterfly flow-graph structure. More recently, it is found that a set of more efficient alternative transforms can be derived by introducing sign-flips at the same stage, which is equivalent to an extra rotation angle of π. In this paper, we generalize this sign-flipping technique by randomly embedding sign-flips into all stages of the DCTs butterfly structure so that the encryption space becomes much larger to yield a higher security. We pursue this study for H.264-compatible videos, assuming that the integer DCT of size 4 × 4 is used. First, we follow the separable implementation of the 4 × 4 2-D DCT in which different sign-flipping strategies will be employed along its horizontal and vertical dimensions. Second, we convert the 4 × 4 2-D DCT into a 16-point 1-D butterfly structure so that more sign-flips can be embedded at its various stages. Third, we choose different schemes to pair the node-variables in the 16-point 1-D butterfly structure, thus further enlarging the encryption space. Extensive experiments are conducted to show the performance of these improved encryption schemes and some security analyzes are also presented to confirm their persistence to various attacking strategies. Bing Zeng 0001, Jeff Siu-Kei Au-Yeung, Shuyuan Zhu, Moncef Gabbouj |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2014 | MRF-Based Fast HEVC Inter CU Decision With the Variance of Absolute DifferencesabstractThe newly developed High Efficiency Video Coding (HEVC) Standard has improved video coding performance significantly in comparison to its predecessors. However, more intensive computation complexity is introduced by implementing a number of new coding tools. In this paper, a fast coding unit (CU) decision based on Markov random field (MRF) is proposed for HEVC inter frames. First, it is observed that the variance of the absolute difference (VAD) is proportional with the rate-distortion (R-D) cost. The VAD based feature is designed for the CU selection. Second, the decision of CU splittings is modeled as an MRF inference problem, which can be optimized by the Graphcut algorithm. Third, a maximum a posteriori (MAP) approach based on the R-D cost is conducted to evaluate whether the unsplit CUs should be further split or not. Experimental results show that the proposed algorithm can achieve about 53% reduction of the coding time with negligible coding performance degradation, which outperforms the state-of-the-art algorithms significantly. Jian Xiong 0005, Hongliang Li 0001, Fanman Meng, Shuyuan Zhu, Qingbo Wu 0001, Bing Zeng 0001 |
IEEE Trans. Multim. | 6 |
| 2013 | A novel enhancement for hierarchical image coding
Shuyuan Zhu, Bing Zeng 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2012 | A new design of multiple transforms for perceptual video encryptionabstractPerceptual (or partial) video encryption has recently been achieved via a random-key controlled selection among multiple transforms, where these transforms are designed by modifying the rotation angles at the last stage of the DCT's butterfly structure. Later on, it was found that more efficient transforms can be obtained by applying random sign-flips at the same stage, which is equivalent to an extra rotation angle of 180°. In this paper, we generalize this technique by embedding random sign-flips into other stages in the DCT's implementation structure. To this end, we first convert the separable implementation of the H.264 4×4 2-D DCT into a 1-D 16-point structure so that more sign-flips can be embedded at its various stages. Then, we introduce different pairing schemes (i.e., forming N/2 pairs for a group of N node-variables) to further increase the encryption space. Extensive experiments are carried out to show the performance of this improved video encryption scheme; whereas security analyses are conducted to show the newly designed transforms out-perform the previous ones in terms of search space and encryption efficiency. Jeff Siu-Kei Au-Yeung, Bing Zeng 0001 |
ICIP | 2 |
| 2012 | An enhanced block-based hierarchical image coding schemeabstractA block-based, two-layer hierarchical image coding scheme codes a down-sampled version of the original image first, and then the residue between the original one and a reconstructed version that is interpolated from the upper layer. When the residual layer is coded at a lower quality as compared to its upper layer, it often happens that all pixels in the upper layer will be deteriorated if the corresponding coded residuals are added into them. In this paper, we demonstrate that there exists a critical region for each image such that the deterioration indeed happens and this region includes nearly all typical bit-rates used in practice. To avoid this problem, we first propose a “naive” solution and then apply an advanced quantization technique to handle the residual layer. To verify the effectiveness, we conduct extensive tests to show that the gap between the hierarchical coding scheme and its single-level counterpart (which is typically around 2~3 dB in the two-layer scenario) has been filled up by a rather big percentage (about 30~50%). Shuyuan Zhu, Bing Zeng 0001 |
ICIP | 2 |
| 2012 | Image Super-Resolution via Low-Pass Filter Based Multi-scale Image DecompositionabstractThis paper presents a spatial-varying minimum mean square error (MMSE)-based approach to construct super-resolution images from single source image of a lower resolution. The unique feature of this approach is that it works on a set of sub-images (also called multi-scale images) that are generated via decomposing the original source image. To do the decomposition, we design a number of low-pass filters with overlapped pass-bands so that sub-images are correlated with each other. Then, an MMSE-based estimation, involving all sub-images, is solved (after making use of the geometric-duality principle) to construct each missing pixel in the super-resolution image. Experimental results show that our new method offers a clearly-noticeable improvement over the existing MMSE-based methods (without decomposition). We believe that this is mainly attributing to the fact that both intra-scale and inter-scale correlations among the sub-images have been utilized in our approach. Shuyuan Zhu, Bing Zeng 0001, Shuicheng Yan |
ICME | 2 |
| 2012 | Design of low-complexity, non-separable 2-D transforms based on butterfly structuresabstractThe transform used in most image and video coding standards is the separable 2-D discrete cosine transform (DCT), which has been proven to be a robust approximation of the optimal Karhunen-Loève transform (KLT) for the 1st-order Markov sources with a large correlation coefficient. However, such separable 2-D DCT surely is not the best choice when it is applied on some residual or directional signals. Based on the butterfly architecture for DCT's fast implementation, we present in this paper a novel design of non-separable 2-D transforms that get much closer to the KLT but at the implementation cost no bigger than that of the DCT. The critical issue in our design is how to pair all node-variables in various stages of the butterfly structure. We propose a near-optimal pairing strategy to solve this problem and present some examples to demonstrate its effectiveness. Haoming Chen, Bing Zeng 0001 |
ISCAS | 2 |
| 2012 | Total-variation based picture reconstruction in multiple description image and video coding
Shuyuan Zhu, Jeff Siu-Kei Au-Yeung, Bing Zeng 0001, Jiying Wu |
Signal Process. Image Commun. | 3 |
| 2012 | New Transforms Tightly Bounded by DCT and KLTabstractIt is well known that the discrete cosine transform (DCT) and Karhunen–Loève transform (KLT) are two good representatives in image and video coding: the first can be implemented very efficiently while the second offers the best R-D coding performance. In this work, we attempt to design some new transforms with two goals: i) approaching to the KLT's R-D performance and ii) maintaining the implementation cost no bigger than that of DCT. To this end, we follow a cascade structure of multiple butterflies to develop an iterative algorithm: two out of N nodes are selected at each stage to form a Givens rotation (which is equivalent to a butterfly); and the best rotation angle is then determined by maximizing the resulted coding gain. We give the closed-form solutions for the node-selection as well as the angle-determination, together with some design examples to demonstrate their superiority. Haoming Chen, Bing Zeng 0001 |
IEEE Signal Process. Lett. | 2 |
| 2011 | Perceptual video encryption using multiple 8×8 transforms in H.264 and MPEG-4abstractIt has been demonstrated in our earlier works [1, 2] that perceptual video encryption can be effectively achieved by using multiple transforms where the block size 4×4 has been considered. In this paper, we study the extension to the transforms of size 8×8. In this case, a more complex flow-graph structure is resulted, thus leading to a larger room for encryption. In addition, special technique on controlling the encrypted video quality is presented by carefully selecting the number of rotations in the flow-graph structure of an 8×8 transform. The proposed scheme is first evaluated using the high profile of H.264. It is then further tested for the MPEG-4 standard that completely relies on 8×8 transform. Both cases show that promising results can be achieved with our proposed scheme. Jeff Siu-Kei Au-Yeung, Shuyuan Zhu, Bing Zeng 0001 |
ICASSP | 3 |
| 2011 | SEAMLESS P2P-MDVC with well-balanced descriptionsabstractMultiple-description coding (MDC) provides a promising solution to support the error-prone transmission over multiple channels. One extremely important application is to design efficient multiple description video coding (MDVC) systems for the peer to peer (P2P) scenario. To this end, one needs to solve the mismatching problem between the reference frames used at the encoder and decoder sides (for motion compensation) in most non-scalable MDVC schemes or get rid of the inter-dependency within various enhancement layers in the scalable MDVC schemes. In the meantime, it is highly preferable to enforce that all descriptions transmitted over the network are well-balanced so as to have an equal payload to each peer. In this paper, we propose an MDVC scheme with a number of well-balanced descriptions. These descriptions are generated from some newly-developed unitary transforms and they can solve all of the problems mentioned above. Shuyuan Zhu, Jeff Siu-Kei Au-Yeung, Bing Zeng 0001 |
ICASSP | 3 |
| 2011 | Design of non-separable transforms for directional 2-D sourcesabstractTraditionally, a 2-D block-based transform is always implemented through two separate 1-D transforms along each block's vertical and horizontal dimensions. Such a framework is however not highly suitable for a directional 2-D source in which the dominant directional information is neither horizontal nor vertical. On the other hand, the R-D performance upper bound for all block-based transform coding schemes applied on such 2-D directional sources can be obtained by the non-separable Karhunen-Loève transform (KLT) — which is unfortunately very expensive computationally. In this paper, we present a new framework for designing some non-separable transforms that offer an R-D performance closer to that of the KLT, but can be implemented with nearly the same complexity as that of the discrete cosine transform (DCT). Haoming Chen, Shuyuan Zhu, Bing Zeng 0001 |
ICIP | 3 |
| 2011 | A comparative study of image correlation models for directional two-dimensional sourcesabstractThe non-separable Karhunen-Loève transform (KLT) has been proven to be optimal for coding a directional 2-D source in which the dominant directional information is neither horizontal nor vertical. However, the KLT depends on the image data, and it is difficult to apply it in a practical image/video coding application. In order to solve this problem, it is necessary to build an image correlation model, and this model needs to adapt to the directional information so as to facilitate the design of 2-D non-separable transforms. In this paper, we compare two models that have been used commonly in practice: the absolute-distance model and the Euclidean-distance model. To this end, theoretical analysis and experimental study are carried out based on these two models, and the results show that the Euclidean-distance model consistently performs better than the absolute-distance model. Shuyuan Zhu, Bing Zeng 0001 |
MMSP | 2 |
| 2011 | Design of New Unitary Transforms for Perceptual Video EncryptionabstractIn our earlier work , we proposed for the first time that the perceptual video encryption be performed at the transformation stage by selecting one out of multiple unitary transforms according to the encryption key. In this letter, we aim to design some more efficient transforms to be used in this framework. Two criteria are followed for designing such transforms: 1) they are significantly different from discrete cosine transform (DCT) or discrete sine transform (DST), and 2) the resulted coding efficiency is exactly the same to what can be achieved by using DCT or just falls very slightly. As a result, we find that these transforms are actually derived by the sign-flipping on some node-variables in the flow-graph structure of DCT - a special case of the plane-based rotations (by π or 180°). Extensive simulations based on the H.264 codec are performed to demonstrate their effectiveness through both objective and subjective assessments. Finally, we present the security analysis to show the resistance of our algorithms to different types of attacks. Jeff Siu-Kei Au-Yeung, Shuyuan Zhu, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2010 | An overview of directional transforms in image codingabstractTransform-based image coding has been the mainstream for many years, as witnessed in from the early effort in JPEG to the recent advances in HD Photo. Traditionally, a 2-D transform used in image coding is always implemented separately along the vertical and horizontal directions, respectively. However, it is usually true that many image blocks contain oriented structures (e.g., edges) and/or textures that do not follow either the vertical or horizontal direction. The traditional 2-D transform thus may not be the most appropriate one to these image blocks. This well-known fact has recently triggered several attempts towards the development of directional transforms so as to better preserve the directional information in an image block. Some of these directional transforms have been applied in image coding, demonstrating a significant coding gain. This paper presents an overview of these directional transforms as well as a discussion of some existing problems and their potential solutions. Jizheng Xu, Bing Zeng 0001, Feng Wu 0001 |
ISCAS | 2 |
| 2010 | Partial video encryption based on alternative integer transformsabstractIt has been demonstrated in our earlier work [1] that partial video encryption can be effectively achieved by using alternative transforms. In this paper, we modify those floating-point based transforms to integer-based counterparts so as to be compatible with the H.264 standard. To this end, we present the design details of these integer-based transforms and their efficient implementations. To follow the H.264 standard that combines the transform and quantization processes, we present some special handling of the necessary scaling factors and the associated quantization. Finally, we demonstrate that the derived integer transforms can achieve partial video encryption with nearly the same performance as what has been achieved with floating-point transforms. Jeff Siu-Kei Au-Yeung, Shuyuan Zhu, Bing Zeng 0001 |
ISCAS | 3 |
| 2010 | Composing better pictures in MDC: A multi-target total variational approachabstractIn any multiple description coding (MDC) system, how to compose pictures of the “best” quality upon receiving more than one description is always important. To this goal, a DCT-domain translation and overlapping technique has been developed in for the popular multiple description scalar quantization (MDSQ) based video coding. Although certain quality gain (in terms of PSNR) has been achieved, the improvement is rather limited. In this paper, we formulate this problem into a traditional total variation (TV) regularized optimization in which all received descriptions are regarded as multiple fidelity terms. We demonstrate that, as compared to the overlapping technique, such a TV-based enhancement yields not only a higher objective quality gain but also a much more significant visual quality gain in the composed pictures. Shuyuan Zhu, Jiying Wu, Bing Zeng 0001 |
ISCAS | 3 |
| 2010 | A total variation-based approach for composing better pictures in multiple description codingabstractOne of the important issues in multiple description coding (MDC) for image and video signals is to compose pictures of the "best" quality when more than one description is received at the decoder side. To this goal, a transform-domain overlapping technique combined with some necessary translations in the transform domain is proposed in the popular MDC schemes with staggered quantizers. However, the achievable gain is rather limited. In this paper, we formulate this enhancement problem as a total variation (TV) regularized optimization constrained by the knowledge of the quantization intervals of the DCT coefficients of each composed picture. In this TV-based approach, the transformdomain overlapping technique is used to find the more accurate quantization intervals in which the true DCT coefficients will fall while receiving more than one description. Simulation results demonstrate that such a TV-based enhancement yields a higher quality gain in both objective (e.g. PSNR-based) and subjective (i.e. visual perception) evaluations than using the transform-domain overlapping technique. Shuyuan Zhu, Bing Zeng 0001 |
VCIP | 2 |
| 2010 | In Search of "Better-than-DCT" Unitary Transforms for Encoding of Residual SignalsabstractIt is well known that the discrete cosine transform (DCT) closely approximates the optimal Karhunen-Loève transform (KLT) under the first-order stationary Markov condition with a strong inter-pixel correlation. However, if the inter-pixel correlation is weak or becomes negative, transforms other than the DCT would possibly become better. In this letter, we present a design framework for finding new unitary transforms according to two principles: 1) they are indeed more efficient than the DCT in encoding of signals with a weak or negative inter-pixel correlation (e.g., the residual signals after motion-compensation or intra- prediction) and 2) each has a fixed transform matrix (i.e., signal-independent) and can be implemented as efficiently as the DCT. Some of the new transforms will be used to demonstrate an improved R-D performance in the video coding scenario. Shuyuan Zhu, Jeff Siu-Kei Au-Yeung, Bing Zeng 0001 |
IEEE Signal Process. Lett. | 3 |
| 2010 | Constrained Quantization in the Transform Domain With Applications in Arbitrarily-Shaped Object CodingabstractIn any block-based transform coding of image/video signals, it is well-known that the mean square error (MSE) distortion measured in the pixel domain is exactly equal to the MSE distortion resulted from quantization in the transform domain if the involved transform matrix is unitary. However, such a property no longer exists if the pixel-domain distortion is measured only on a selected part of pixels within one image block. This provides us an opportunity of dynamically shaping the quantization errors so as to make the selected pixels (much) better than the unselected ones. In this paper, we first develop a reversed iterative algorithm to guide us to perform a highly constrained quantization so that the coding quality of the selected pixels in each image block is significantly higher than what can be achieved by using the normal quantization. Then, we apply this intelligent quantization in one practical scenario-coding of arbitrarily-shaped image blocks in MPEG-4, showing remarkable improvements in comparison with the original MPEG-4. Shuyuan Zhu, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | A comparative study of multiple description video coding in P2P: normal MDSQ versus flexible dead-zonesabstractMultiple-description coding (MDC) encodes a source image or video into several distinct descriptions. Each description can be decoded independently and more than one description can also be decoded jointly so as to deliver a higher quality. One typical design of an MDC encoder is based on the so-called multiple description scalar quantization (MDSQ). A simple MDSQ system is to choose different step-sizes used in all involved quantizers. Another implementation is to apply flexible dead-zones in all quantizers (whereas keeping the same step-size across them). Two issues turn to be particularly important when MDC is used in video streaming applications in the P2P scenario. First, we need to consider more than two descriptions to fit the P2P reality. Second, we have to eliminate any possible drifting errors that are resulted from the use of different reference frames (in different descriptions) at the motion-compensation stage. In this paper, we introduce a DCT-domain translation to solve the mismatching problem in different reference frames. Then, we present some comparative results of the two approaches mentioned above, with detailed pros and cons for each one as well as some practical considerations. Shuyuan Zhu, Jeff Siu-Kei Au-Yeung, Bing Zeng 0001 |
ICME | 3 |
| 2009 | Multiple Description Coding in the Quincunx Sub-sampling Lattice with Diamond-shape DCTabstractPixel-domain sub-sampling based on the quincunx lattice is one of the traditional approaches to carry out multiple description coding (MDC). Unfortunately, the re-alignment (either horizontal or vertical) would deteriorate the sample correlation within each description, thus leading to a lower coding efficiency. In this paper, we propose to apply a diamond-shape DCT on the quincunx sub-sampling lattice to create an MDC representation. A theoretical analysis is first presented assuming the covariance matrix of the Toeplitz type. The results show that such DS-DCT can achieve a better coding gain as well as energy packing efficiency as compared to the normal DCT with re-alignment. Then, the proposed DS-DCT is tested on various images using the standard JPEG compression scheme and the results show that it out-performs the traditional re-alignment approach by a similar margin as obtained in our theoretical analysis. Jeff Siu-Kei Au-Yeung, Bing Zeng 0001 |
ISCAS | 2 |
| 2009 | Partial Video Encryption Based on Alternating TransformsabstractIn this letter, we propose a novel video encryption technique that is used to achievepartialencryption where an annoying video can still be reconstructed even without the security key. In contrast to the existing methods where the encryption usually takes place at the entropy-coding stage or the bit-stream level, our proposed scheme embeds the encryption at thetransformstage during the encoding process. To this end, we develop a number of new unitary transforms that are demonstrated to be equally efficient as the well-known DCT and thus used as alternates to DCT during the encoding process. Partial encryption is achieved through alternately applying these transforms to individual blocks according to a pre-designed secret key. Analysis on the security level of this partial encryption scheme is carried out against various common attacks and some experimental results based on H.264/AVC are presented. Jeff Siu-Kei Au-Yeung, Shuyuan Zhu, Bing Zeng 0001 |
IEEE Signal Process. Lett. | 3 |
| 2009 | R-D Performance Upper Bound of Transform Coding for 2-D Directional SourcesabstractTraditionally, any 2D transform (such as 2D DCT) is implemented through two separable 1D transforms along the vertical and horizontal dimensions. Such a framework is however not most suitable for a 2D directional source in which the dominant directional information is neither horizontal nor vertical. In this letter, we attempt to determine the R-D performance upper bound for block-based transform coding schemes applied on such 2D directional sources. It is not a surprise that the Karhunen-Loeve transform (KLT) plays a critical role here. Specifically, we show that a nonseparable KLT can be determined directly from the given 2D directional source model to yield the R-D performance upper bound. We also show that there exists a significant gap between this upper bound and the R-D performance that can be achieved by using the traditional 2D DCT. Shuyuan Zhu, Jeff Siu-Kei Au-Yeung, Bing Zeng 0001 |
IEEE Signal Process. Lett. | 3 |
| 2008 | Coding of arbitrarily-shaped image blocks based on a constrained quantizationabstractCoding of arbitrarily-shaped image/video segments is important to achieving the object-based coding - which is a key technique in many multimedia applications. MPEG-4 suggested a low-pass-extrapolation (LPE) padding method to handle each boundary block of an arbitrary shape. In this paper, we develop a constrained quantization technique and apply it to each LPE-padded boundary block. This ‘smart’ quantization is performed in such a way that the pixel-domain MSE distortion due to the DCT domain quantization is shaped as much as possible onto all padded pixels — as these pixels will be thrown away in the end. In the meantime, all pixels in each original boundary block would be coded with an improved quality. Compared to the normal rounding-based quantization used in the original MPEG-4’s LPE scheme, our new quantization needs some extra computations at the quantization stage in the encoding procedure but keeps the decoding complexity unchanged, while offering a remarkable coding gain — as demonstrated by some simulation results. Shuyuan Zhu, Bing Zeng 0001 |
ICME | 2 |
| 2008 | Seamless MDVC in P2P: A transform-domain approachabstractMultiple-description coding (MDC) provides an effective way to mitigate the effects of packet errors/loses by making use of multiple channels. Perhaps, the most attractive application of MDC is in the peer-to-peer (P2P) scenario to support simultaneous video streaming to a large population of clients. To this end, a number of multiple-description video coding (MDVC) schemes (both non-scalable and scalable) were proposed in the past few years. However, almost all non-scalable schemes would suffer from the prediction mismatch between the references used at the encoder and decoder sides; whereas all scalable schemes (involving a base-layer and some enhancement layers) would suffer from the inter-dependency within the enhancement-layer information. In this paper, we propose a transform-domain MDVC method that can solve these problems and at the same time offer some other interesting features. Shuyuan Zhu, Bing Zeng 0001 |
MMSP | 2 |
| 2007 | Multiple Description Transform Coding: A New Design ApproachabstractMultiple-description coding (MDC) is an effective technique to support error-prone transmission in the scenario of multiple channels. Typically, an MDC is designed in the DCT domain where a pair-wise correlating transform (PCT) or some generalized version is further applied to some DCT coefficients of each image block. Although these correlating transforms offer a freedom of compromising the overall coding quality and the redundancy, none of them can be made optimal after taking into consideration the quantization characteristics. In this paper, we present a new design approach in which we will do a novel grouping of all DCT coefficients (not necessarily pair-wise) such that the distortion measured on some specified pixels within each image block is minimized. An interesting property of this new MDC design is that, when all descriptions are received, it provides a coding quality that becomes higher than the target quality set in the corresponding single-description coding (SDC) design. Bing Zeng 0001, Shuyuan Zhu |
ICME | 1 |
| 2004 | Error-resilient unequal protection of fine granularity scalable video bitstreamsabstractThis paper deals with the optimal packet loss protection issue for streaming the fine granularity scalable (FGS) video bitstreams over IP networks. Unlike many other existing protection schemes, we develop an error-resilient unequal protection (ER-UEP) method that adds redundant information optimally for loss protection and, at the same time, cancels completely the dependency among bitstream after loss recovery. In our ER-UEP method, the FGS enhancement-layer bitstream is first packetized into a group of independent data packets, while each packet can be truncated to represent the original video signal at any fidelity (i.e., scalability). Parity packets are then created with intrinsic UEP capabilities that can easily adapt to the current channel conditions. Unlike conventional UEP schemes that suffer from bitstream contamination due to the dependency among packets, our method guarantees the successful decoding of all received bits, thus leading to a better error resilience as well as higher robustness (under varying and/or unclean channel conditions). Hua Cai, Bing Zeng 0001, Guobin Shen, Shipeng Li 0001 |
ICC | 2 |
| 2004 | Arbitrarily-shaped video coding: smart padding versus MPEG-4 LPE/zero paddingabstractAn effective padding scheme, called the smart padding (SmartPad), has been developed recently for the DCT coding of arbitrarily-shaped image/video objects; whereas its superior performance over the MPEG-4 LPE padding has been confirmed solidly. In the present paper, we propose to extend the use of SmartPad to all INTER frames (of arbitrary shapes), i.e., to use SmartPad to replace the zero padding scheme (as recommended in MPEG-4). Our simulation results show that a very substantial performance gain (3-7 dB) has been achieved, as compared to the MPEG-4 LPE/zero padding scheme. A. C. Yu, Guobin Shen, Bing Zeng 0001, Oscar C. Au |
ICME | 3 |
| 2003 | Adaptive vector quantization with codebook updating based on locality and historyabstractIn this paper, we propose two techniques that are applicable to any adaptive vector quantization (AVQ) systems. The first one is called the locality-based codebook updating: when performing a codebook updating, we update the operational codebook using not only the current input vector but also the codewords at all positions within a selected neighboring area (called the locality), while the operational codebook is organized in a "cache" manner. This technique is rationalized by the high correlation cross neighboring vectors that facilitates a more efficient coding of the indices of the codewords chosen from the codebook. The second technique is called the history aid, which makes use of the information of previously coded vectors to quantize the current input vector if it is used to update the operational codebook. A more effective AVQ system is obtained by combining together the history aid and the locality-based updating. Extensive simulations are carried out to demonstrate the improved results achieved by our AVQ systems. Particularly, when the operational codebook size is relatively small, the improvement over a benchmark AVQ system--the generalized threshold replenishment (GTR)--is drastic. For example, when the size is 32, testing on a nonstationary signal (containing frames from different video sequences, ordered in the concatenating or interleaving format) shows that the combination of history aid and locality-based updating offers more than 4 dB gain over GTR at 0.5 bpp. Guobin Shen, Bing Zeng 0001, Ming Lei Liou |
IEEE Trans. Image Process. | 2 |
| 2002 | Optimal rate allocation for macroblock-based progressive fine granularity scalable video codingabstractThis paper addresses the problem of optimal rate allocation for the macroblock-based progressive fine granularity scalable (PFGS) video coding. To solve this complicated problem, the error propagation pattern in the macroblock-based PFGS is first investigated. An effective drifting model is established subsequently for estimating the drifting for each enhancement bit stream segment encoded by the macroblock-based PFGS. The distortion reduction for the current frame and the estimated drifting suppression for the subsequent frames form the actual contribution of the enhancement layer bitstream. The equal-slope argument is then applied to select the best bit stream segments for the given bandwidth. Experiments show that our optimal rate allocation outperforms the uniform rate allocation by 0.3-1.4 dB. Hua Cai, Guobin Shen, Shipeng Li 0001, Bing Zeng 0001 |
ICIP (3) | 4 |
| 2002 | Error concealment for fine granularity scalable video transmissionabstractIn this paper we present an efficient error concealment (EC) method for the fine granularity scalable (FGS) video transmission. The proposed EC method exploits both the temporal and spatial correlations in an FGS encoded bitstream. In our scheme, the temporal redundancy is used to improve the quality of contaminated regions, and the intensity of the temporal correlation in contaminated regions is estimated by exploiting the spatial correlation in the surrounding high-quality regions. To maximally utilize the spatial correlation for estimation, we also propose two interleaving patterns that can avoid packetizing the neighboring regions into the same packet. Experiments show that our EC method achieves very good performance and is robust to different bandwidths and different sequences. Hua Cai, Guobin Shen, Feng Wu 0001, Shipeng Li 0001, Bing Zeng 0001 |
ICME (1) | 5 |
| 2002 | A novel motion estimation algorithm for arbitrarily shaped video codingabstractIn this paper, we present a fast motion estimation algorithm for arbitrarily shaped video coding in MPEG-4. This novel algorithm takes advantage of our new discovery about the close relation between the best matching block and its shape information-alpha plane. Without any additional padding computation, this new motion estimation algorithm computes the sum of absolute difference (SAD) between two alpha planes rather than the pixels' intensities. Compared with the full search block-matching algorithm (recommended in the MPEG-4 standard), the proposed algorithm achieves an impressive speed-up ratio with very minor quality degradation and little bit-count increase. Extensive simulations are provided in this paper to demonstrate this fact. Andy C.-W. Yu, Bing Zeng 0001, Oscar C. Au |
ICME (1) | 2 |
| 2001 | Arbitrarily shaped transform coding based on a new padding techniqueabstractCoding of arbitrarily shaped image segments is an important tool to achieve object-based coding, which is becoming more and more popular in today's multimedia applications. We introduce a new padding technique based on which the arbitrarily shaped DCT can be implemented using a normal N/spl times/N DCT. The new padding is carried out for each arbitrarily shaped block in such a way that there are as many transformed coefficients of high frequencies as possible that could be set to zero. In the best case, it does not expand the data set in the DCT-domain. Arbitrarily shaped DCT coding based on this padding technique is developed, and then analyzed and compared against some of the existing algorithms in terms of the rate-distortion performance, computational complexity, and implementation cost. Guobin Shen, Bing Zeng 0001, Ming Lei Liou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2000 | Achieving optimal rate-distortion performance in arbitrarily-shaped transform codingabstractIn this paper, we present a simple but effective method to enhance the coding performance of shape-adaptive DCT (SA-DCT). By choosing the first processing direction (between horizontal and vertical) within each boundary block for doing the 1D transform, the proposed method guarantees to achieve optimal rate-distortion results and actually outperforms the existing SA-DCT algorithms significantly. The proposed method is also applicable to arbitrarily-shaped coding based on some padding techniques. Guobin Shen, Bing Zeng 0001, Ming Lei Liou |
ISCAS | 2 |
| 2000 | An efficient hybrid arbitrarily-shaped object coding techniqueabstractIn this paper, we present a very efficient shape-adaptive coding method which is a hybrid between the standard SA-DCT and the padding technique we proposed in our early work [1999]. This hybrid has led to a significantly lower computation burden while the coding performance is improved. This conclusion is proved by thorough complexity analysis and extensive simulation. This method also has the shape preserving property and exhibits asymmetric complexities between the encoder and the decoder. Guobin Shen, Bing Zeng 0001, Ming Lei Liou |
ISCAS | 2 |
| 1999 | A New Padding Technique for Coding of Arbitrarily-Shaped Iamge/Video SegmentsabstractObject-based coding is becoming more and more important in today's multimedia applications. Shape-adaptive DCT (SA-DCT) provides a useful tool for coding of arbitrarily-shaped image/video segments which is indispensable to achieve object-based coding. In this paper, we introduce a new padding technique based on which the arbitrarily-shaped DCT can be implemented using normal N×N DCT. The new padding is carried out for each arbitrarily-shaped block in such a way that the number of non-zero coefficients after DCT is guaranteed to be no more than that of the original image data, thereby never expanding the data set in the DCT-domain. Arbitrarily-shaped DCT coding based on this padding technique is developed, and then analyzed and compared against some of the existing algorithms in terms of rate-distortion performance, computational complexity, and implementation cost. Guobin Shen, Bing Zeng 0001, Ming Lei Liou |
ICIP (2) | 2 |
| 1997 | Optimization of fast block motion estimation algorithmsabstractThere are basically three approaches for carrying out fast block motion estimation: (1) fast search by a reduction of motion vector candidates; (2) fast block-matching distortion (BMD) computation; and (3) motion field subsampling. The first approach has been studied more extensively since different ways of reducing motion vector candidates may result in significantly different performance; while the second and third approaches can in general be integrated into the first one so as to further accelerate the estimation process. In this paper, we first formulate the design of good fast estimation algorithms based on motion vector candidate reduction into an optimization problem that involves the checking point pattern (CPP) design via minimizing the distance from the true motion vector to the closest checking point (DCCP). Then, we demonstrate through extensive studies on the statistical behavior of real-world motion vectors that the DCCP minimization can result in fast search algorithms that are very efficient as well as highly robust. To further utilize the spatiotemporal correlation of motion vectors, we develop an adaptive search scheme and a hybrid search idea that involves a fixed CPP and a variable CPP. Simulations are performed to confirm their advantages over conventional fast search algorithms. Bing Zeng 0001, Renxiang Li, Ming Lei Liou |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1995 | Visual Pattern BTC with Two Principle Colors for Color ImagesabstractThis paper presents an extension of the "block truncation coding with visual patterns (BTC-VP)" for color image compression. Our approach BTC-VP-2PC attributes its name to the fact that it, regardless of the variety of colors, encodes each 4/spl times/4 image block into two principle colors (PC) and one visual pattern (VP) (which are both visually significant to human observers). In BTC-VP-2PC, each VP is generated from a bitmap. This bitmap, obtained from a weighted trichromatic transform, quantizes all color components (i.e. it's actually BTC). Experimental results show that only two PCs and one VP for each block are enough to represent edge regions with high subjective quality. Also a substantial compression ratio has been achieved (for NTSC color TV images). Tak Po Chan, Bing Zeng 0001, Ming Lei Liou |
ISCAS | 2 |
| 1994 | Design and Implemenation of a Programmable Stack FilterabstractProposes two simple and efficient architectures for the realization of stack filters in hardware: one for pipelined implementation and the other for non-pipelined implementation. Both architectures are interchangeable, with very minor modifications. A programmable stack filter has been layed out based on the proposed pipelined architecture in the AMS 1 micron technology using Mentor Graphics Generator Design Tools for five words of twelve bits each and simulations have been carried out successfully using the Mentor Graphics Lsim simulator. Results are encouraging that the simulations are error-free at a frequency as high as 166 MHz.> Prasad V. Lakamsani, Ruikang Yang, Bing Zeng 0001, Ming Lei Liou |
ICIP (3) | 3 |
| 1994 | A new three-step search algorithm for block motion estimationabstractThe three-step search (TSS) algorithm has been widely used as the motion estimation technique in some low bit-rate video compression applications, owing to its simplicity and effectiveness. However, TSS uses a uniformly allocated checking point pattern in its first step, which becomes inefficient for the estimation of small motions. A new three-step search (NTSS) algorithm is proposed in the paper. The features of NTSS are that it employs a center-biased checking point pattern in the first step, which is derived by making the search adaptive to the motion vector distribution, and a halfway-stop technique to reduce the computation cost. Simulation results show that, as compared to TSS, NTSS is much more robust, produces smaller motion compensation errors, and has a very compatible computational complexity.> Renxiang Li, Bing Zeng 0001, Ming Lei Liou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1994 | Optimal parallel stack filtering under the mean absolute error criterionabstractThe authors extend the configuration of stack filtering to develop a new class of stack-type filters called parallel stack filters (PSFs). As a basis for the parallel stack filtering, the block threshold decomposition (BTD) is introduced, and its properties are investigated. The design of optimal PSHs under the mean absolute error (MAE) criterion is shown to be similar to the minimum MAE stack filtering theory. The only difference is that one needs now to design more than one stack filter that together construct an optimal PSF. As a result, while reviewing briefly the optimal stack filtering theory, they will put more efforts to demonstrate, via several examples, the improvement by switching from stack filtering to parallel stack filtering for the task of image noise removal. Bing Zeng 0001, Yrjö Neuvo |
IEEE Trans. Image Process. | 1 |
| 1993 | Interpolative BTC image coding with vector quantizationabstractThe authors suggest two interpolative block truncation coding (BTC) image coding schemes with vector quantization and median filters as the interpolator. The first scheme is based on quincunx subsampling and the second one on every-other-row-and-every-other-column subsampling. It is shown that the schemes yield a significant reduction in bit rate at only a small performance degradation and, in general, better channel error resisting capabilities, as compared to the absolute moment BTC. The methods are further demonstrated to outperform the corresponding BTC schemes with pure vector quantization at the same bit rate and require minimal computations for the interpolation.> Bing Zeng 0001, Yrjö Neuvo |
IEEE Trans. Commun. | 1 |
| 1992 | Interpolative BTC image codingabstractTwo interpolative block truncation coding (BTC) schemes with vector quantization (VQ) and median filters as the interpolator are proposed. The first scheme is based on quincunx subsampling and the second one on every-other-row-and-every-other-column subsampling. Compared to the standard BTC, the new schemes yield a significant reduction in bit rate, only a small degradation in quality, and better channel error resisting capabilities. Moreover, the methods are shown to outperform the BTC schemes with pure VQ at the same bit rate and require minimal computations for the interpolation.> Bing Zeng 0001, Yrjö Neuvo, Anastasios N. Venetsanopoulos |
ICASSP | 1 |
| 1991 | Optimal stack filtering and classical Bayes decisionabstractOptimal stack filtering under the mean absolute error (MAE) criterion is studied. It is first shown that this problem is equivalent to the classical a priori Bayes minimum-cost decision. Generally, a linear program (LP) with O(b2/sup b/) variables and constraints (b is the window width) is required for finding the best filter. Instead, the authors develop a suboptimal routine which renders the use of the LP obsolete, but yields reasonably good filters. Sufficient conditions under which the proposed routine results in optimal solutions are provided and shown to hold in most practical cases. Several design examples are given.> Bing Zeng 0001, Moncef Gabbouj, Yrjö Neuvo |
ICASSP | 1 |
| 1991 | Error feedback for floating-point cascade-form digital filtersabstractThe authors propose the use of error feedback for roundoff noise reduction in cascade-form recursive digital filters employing floating-point arithmetic. The floating-point error feedback has a similar mechanism to the fixed-point error feedback. However, the optimal error feedback solution (in the sense of minimum mean-square error) for the floating-point case is different from the corresponding fixed-point case. The formulated optimal feedback coefficients are shown to be independent of the section of ordering and zero-pole pairing, which does not hold in the fixed-point case. Finally, a suboptimal error feedback is proposed and shown to provide a constant output SNR (signal-to-noise ratio) independent of the specific characteristics of the filter except the filter's order. Several examples are given to verify these results.> Bing Zeng 0001, Timo I. Laakso, Iiro Hartimo, Yrjö Neuvo |
ICASSP | 1 |
| 1991 | Synthesis of optimal detail-restoring stack filters for image processingabstractA two-step method is presented to synthesize optimal stack filters under the mean absolute error (MAE) criterion. First, the probabilities needed in the optimal filter design are estimated based on images. Second, the linear program (LP) required for finding the best filter is avoided by a 'reasonably good' suboptimal routine which only involves data comparisons. A sufficient condition under which the suboptimal routine results in optimal solutions is given and shown to hold in most practical cases. The proposed method is then applied to synthesize a family of optimal stack filters for the task of restoring an image in impulsive noise. Testing results show that the synthesized filters lead to a greatly improved image-detail restoration compared to standard median filters, although median filters remove noise better.> Bing Zeng 0001, Hongbing Zhou, Yrjö Neuvo |
ICASSP | 1 |