VLDB 2026 Research / reviewers in the wild / expert
Ruiqin Xiong
dblp:12/6908
· DBLP profile ↗
188ranked-venue papers
25as first author
60since 2021 · last 2026
0000-0001-9796-0478ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 158 · 22 first-author · 47 since 2021Artificial intelligence and machine learning · 29 · 27 since 2021Systems, architecture and hardware · 16 · 3 first-author · 2 since 2021Databases, data management, data science and information retrieval · 14 · 3 first-authorApplied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Computer networks · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spatio-Temporal Distortion Aware Omnidirectional Video Super-ResolutionabstractOmnidirectional videos (ODVs) provide an immersive visual experience by capturing the 360° scene. With the rapid advancements in virtual/augmented reality, metaverse, and generative artificial intelligence, the demand for high-quality ODVs is surging. However, ODVs often suffer from low resolution due to their wide field of view and limitations in capturing devices and transmission bandwidth. Although video super-resolution (SR) is a capable video quality enhancement technique, the performance ceiling and practical generalization of existing methods are limited when applied to ODVs due to their unique attributes. To alleviate spatial projection distortions and temporal flickering of ODVs, we propose a Spatio-Temporal Distortion Aware Network (STDAN) with joint spatio-temporal alignment and reconstruction. Specifically, we incorporate a spatio-temporal continuous alignment (STCA) to mitigate discrete geometric artifacts in parallel with temporal alignment. Subsequently, we introduce an interlaced multi-frame reconstruction (IMFR) to enhance temporal consistency. Furthermore, we employ latitude-saliency adaptive (LSA) weights to focus on regions with higher texture complexity and human-watching interest. By exploring a spatio-temporal jointly framework and real-world viewing strategies, STDAN effectively reinforces spatio-temporal coherence on a novel ODV-SR dataset and ensures affordable computational costs. Extensive experimental results demonstrate that STDAN outperforms state-of-the-art methods in improving visual fidelity and dynamic smoothness of ODVs. Hongyu An, Xinfeng Zhang 0001, Shijie Zhao 0001, Li Zhang 0006, Ruiqin Xiong |
AAAI | 5 |
| 2026 | Spike Stream Memory Transfer for Dynamic Scene ReconstructionabstractAs a retina-inspired sensor with ultra-high temporal resolution, spike camera can continuously capture dynamic scenes with high-speed motion. It is a key task to restore clear images from spike streams. The quantization effects in spike readout bring degradation to the visual quality of restored images. To tackle the degradation without introducing motion blur, existing methods often employ a short-term temporal window to infer the light intensity at a certain time point. However, these methods only focus on the spike signals within the current window, which limits their performance. Motivated by the human-like memory mechanism for visual signals from the retina, we explore Spike Stream Memory Transfer (SSMT) to restore the dynamic scenes, considering spike signals beyond the window. Specifically, we design a framework that leverages temporal memory by transferring previously inferred light intensity and motion to enhance current reconstruction. The framework enables a long-term temporal perception of spike streams to handle the spike quantization effects. Besides, we utilize the estimated motion to suppress the potential blur from inter-stream clips, considering the underlying motion of spike streams. We also develop a spike interval-guided alignment module to tackle the blur from intra-stream clips. Experimental results on both synthetic and real-captured data demonstrate that our method can restore high-quality images from spike streams. Yanchen Dong 0001, Ruiqin Xiong, Rui Zhao 0010, Xinfeng Zhang 0001, Tiejun Huang 0001 |
AAAI | 2 |
| 2026 | SpikeCV: open a continuous computer vision era
Yajing Zheng, Jiyuan Zhang 0005, Rui Zhao 0010, Jianhao Ding, Shiyan Chen, Weijian Wu, Ruiqin Xiong, Zhaofei Yu, Tiejun Huang 0001 |
Sci. China Inf. Sci. | 7 |
| 2026 | DermClinical: Clinical-oriented dataset and evaluation for computer-aided dermatological diagnosis
Zihao Liu 0009, Ruiqin Xiong, Shaoting Zhang 0001, Tingting Jiang 0001 |
Neurocomputing | 3 |
| 2026 | Spike Camera Optical Flow Estimation Based on Continuous Spike StreamsabstractSpike camera is an emerging bio-inspired vision sensor with ultra-high temporal resolution. It records scenes by accumulating photons and outputting binary spike streams. Optical flow estimation aims to estimate pixel-level correspondences between different moments, describing motion information along time, which is a key task of spike camera. High-quality optical flow is important since motion information is a foundation for analyzing spikes. However, extracting stable light-intensity information from spikes is difficult due to the randomness of binary spikes. Besides, the continuity of spikes can offer contextual information for optical flow. In this paper, we propose a network Spike2Flow++ to estimate optical flow for spike camera. In Spike2Flow++, we propose a differential of spike firing time (DSFT) to represent information in binary spikes. Moreover, we propose a dual DSFT representation and a dual correlation construction to extract stable light-intensity information for reliable correlations. To use the continuity of spikes as motion contextual information, we propose a joint correlation decoding (JCD) that jointly estimates a series of flow fields. To adaptively fuse different motions in JCD, we propose a global motion bank aggregation to construct an information bank for all motions and adaptively extract contexts from the bank for each iteration during recurrent decoding of each motion. To train and evaluate our network, we construct a real scene with spikes and flow++ (RSSF++) based on real-world scenes. Experiments demonstrate that our Spike2Flow++ achieves state-of-the-art performance on RSSF++, photo-realistic high-speed motion (PHM), and real-captured data. Rui Zhao 0010, Ruiqin Xiong, Dongkai Wang, Shiyu Xuan, Jian Zhang 0018, Xiaopeng Fan 0001, Tiejun Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Semantic-Decoupled and Knowledge-Shared Probabilistic Mapping Network for Multi-Grained Cross-Modal RetrievalabstractCross-modal retrieval is essential for exploring semantic correlations between multimodal data. However, existing approaches face challenges in resolving semantic ambiguity and transferring knowledge with sparse sample generalization. To address these challenges, we propose a new Semantic-Decoupled and Knowledge-Shared Probabilistic Mapping Network (SKPMN). Specifically, the Semantic Decoupling and Distinction (SDD) module decomposes complex word-region relationships into relevance-driven representations. The Deep Probability Mapping (DPM) module introduces a paradigm shift by mapping multimodal features into probabilistic distributions, capturing the semantic similarities and the potential uncertainties that define sparse or ambiguous relationships. By combining the Attention Probabilistic Mapping (APM) module, the model can effectively transfer knowledge across similar samples while emphasizing critical distinctions, significantly enhancing generalization to sparse and ambiguous samples. Finally, the multi-grained alignment strategy establishes a novel integration of fine-grained patch-to-word alignment and coarse-grained global alignment. Experimental results show that SKPMN achieves superior retrieval accuracy across major benchmark datasets. Furthermore, we implement a channel resource allocation technique that allocates more transmission resources to semantically significant information. In resource-constrained environments, our approach leverages Joint Source-Channel Coding (JSCC) to enhance the efficiency of visual feature transmission. Wenrui Li 0001, Yeyu Chai, Liang-Jian Deng, Ruiqin Xiong, Xiaopeng Fan 0001, Yonghong Tian 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Self-Supervised Learning for Color Spike Camera ReconstructionabstractSpike camera is a kind of neuromorphic camera with ultra-high temporal resolution, which can capture dynamic scenes by continuously firing spike signals. To capture color information, a color filter array (CFA) is employed on the sensor of the spike camera, resulting in Bayer-pattern spike streams. How to restore high-quality color images from the binary spike signals remains challenging. In this paper, we propose a motion-guided reconstruction method for spike cameras with CFA, utilizing color layout and estimated motion information. Specifically, we develop a joint motion estimation pipeline for the Bayer-pattern spike stream, exploiting the motion consistency of channels. We propose to estimate the missing pixels of each color channel according to temporally neighboring pixels of the corresponding color along the motion trajectory. As the spike signals are read out at discrete time points, there is quantization noise that impacts the image quality. Thus, we analyze the correlation of the noise in spatial and temporal domains and propose a self-supervised network utilizing a masked spike encoder to handle the noise. Experiments on real-world captured Bayer-pattern spike streams show that our method can restore color images with better visual quality, compared with state-of-the-art methods. The source codes are available at https://github.com/csycdong/SSL-CSC. Yanchen Dong 0001, Ruiqin Xiong, Xiaopeng Fan 0001, Zhaofei Yu, Yonghong Tian 0001, Tiejun Huang 0001 |
CVPR | 2 |
| 2025 | Spk2SRImgNet: Super-Resolve Dynamic Scene from Spike Stream via Motion Aligned Collaborative FilteringabstractSpike camera is a kind of neuromorphic camera that records dynamic scenes by firing a stream of binary spikes with extremely high temporal resolution. It demonstrates great potential for vision tasks in high-speed scenarios. One limitation in its current implementation is the relatively low spatial resolution. This paper develops a network called Spk2SRImgNet to super-resolve high resolution images from low resolution spike stream. However, fluctuations in spike stream hinder the performance of spike camera super resolution. To address this issue, we propose a motion aligned collaborative filtering (MACF) module, which is motivated by key ideas in classic image restoration schemes to mitigate fluctuations in spike data. MACF leverages the temporal similarity of spike stream to acquire similar features from neighboring moments via motion alignment. To separate disturbances from features, MACF filters these similar features jointly in transform domain to exploit representation sparsity, and generates refinement features that will be used to update initial fluctuated features. Specifically, MACF designs an inverse motion alignment operation to map these refinement features back to their original positions. The initial features are aggregated with the repositioned refinement features to enhance reliability. Experimental results demonstrate that the proposed method achieves state-of-the-art performance compared with existing methods. Yuanlin Wang, Ruiqin Xiong, Jing Zhao 0011, Jian Zhang 0018, Xiaopeng Fan 0001, Tiejun Huang 0001 |
CVPR | 3 |
| 2025 | ISP2HRNet: Learning to Reconstruct High Resolution Image from Irregularly Sampled Pixels via Hierarchical Gradient Learning
Yuanlin Wang, Ruiqin Xiong, Rui Zhao 0010, Jin Wang 0023, Xiaopeng Fan 0001, Tiejun Huang 0001 |
ICCV | 2 |
| 2025 | SAMPLE: Semantic Alignment through Temporal-Adaptive Multimodal Prompt Learning for Event-Based Open-Vocabulary Action Recognition
Rui Zhao 0010, Ruiqin Xiong, Xiaopeng Fan 0001, Tiejun Huang 0001 |
ICCV | 3 |
| 2025 | High Dynamic Range Imaging with Time-Encoding Spike CameraabstractAs a bio-inspired vision sensor, spike camera records light intensity by accumulating photons and firing a spike once a preset threshold is reached. For high-light regions, the accumulated photons may reach the threshold multiple times within a readout interval, while only one spike can be stored and read out, resulting in incorrect intensity representation and a limited dynamic range. Multi-level (ML) spike camera enhances the dynamic range by introducing a spike-firing counter (SFC) to count spikes within each readout interval for each pixel, and uses different spike symbols to represent the arrival of different amounts of photons. However, when the light intensity becomes even higher, each pixel requires an SFC with a higher bit depth, causing great cost to the manufacturing process. To address these issues, we propose time-encoding (TE) spike camera, which transforms the counting of spikes to recording of the time at which a specific number of spikes (i.e., an overflow) is reached. To encode time information with as few bits as possible, instead of directly utilising a timer, we leverage a periodic timing signal with a higher frequency than the readout signal. Then the recording of overflow moment can be transformed into recording the number of accumulated timing signal cycles until the overflow occurs. Additionally, we propose an image reconstruction scheme for TE spike camera, which leverages the multi-scale gradient features of spike data. This scheme includes a similarity-based pyramid alignment module to align spike streams across the temporal domain and a light intensity-based refinement module, which utilises the guidance of light intensity to fuse spatial features of the spike data. Experimental results demonstrate that TE spike camera effectively improves the dynamic range of spike camera. Zhenkun Zhu 0001, Ruiqin Xiong, Jiyu Xie, Yuanlin Wang, Xinfeng Zhang 0001, Tiejun Huang 0001 |
NeurIPS | 2 |
| 2025 | Towards Ultra High-Speed Hyperspectral Imaging by Integrating Compressive and Neuromorphic Sampling
Mengyue Geng, Lizhi Wang 0001, Lin Zhu 0012, Wei Zhang 0161, Ruiqin Xiong, Yonghong Tian 0001 |
Int. J. Comput. Vis. | 5 |
| 2025 | Event-Enhanced Snapshot Mosaic Hyperspectral Frame DeblurringabstractSnapshot Mosaic Hyperspectral Cameras (SMHCs) are popular hyperspectral imaging devices for acquiring both color and motion details of scenes. However, the narrow-band spectral filters in SMHCs may negatively impact their motion perception ability, resulting in blurry SMHC frames. In this paper, we propose a hardware-software collaborative approach to address the blurring issue of SMHCs. Our approach involves integrating SMHCs with neuromorphic event cameras for efficient event-enhanced SMHC frame deblurring. To achieve spectral information recovery guided by event signals, we formulate a spectral-aware Event-based Double Integral (sEDI) model that links SMHC frames and events from a spectral perspective, providing principled model design insights. Then, we develop a Diffusion-guided Noise Awareness (DNA) training framework that utilizes diffusion models to learn noise-aware features and promote model robustness towards camera noise. Furthermore, we design an Event-enhanced Hyperspectral frame Deblurring Network (EvHDNet) based on sEDI, which is trained with DNA and features improved spatial-spectral learning and modality interaction for reliable SMHC frame deblurring. Experiments on both synthetic data and real data show that the proposed DNA + EvHDNet outperforms state-of-the-art methods on both spatial and spectral fidelity. The code and dataset will be made publicly available. Mengyue Geng, Lizhi Wang 0001, Lin Zhu 0012, Wei Zhang 0161, Ruiqin Xiong, Yonghong Tian 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Attentive Large Kernel Network With Mixture of Experts for Video DeblurringabstractVideo deblurring is a fundamental problem in low-level vision, and many methods have employed designs based on CNNs and transformers. Traditional CNNs often require deeper architectures to achieve a larger receptive field, which may not be optimal for spatially non-uniform blurs and intense motion blurs. While transformers offer a large receptive field, their quadratic complexity due to attention designs typically imposes a significant computational burden. In addressing these issues, we present an Attentive Large Kernel Network with Mixture of Experts (ALK-MoE). In ALK-MoE, an attentive large kernel backbone network is proposed. On one hand, it inherently extends the network’s receptive field through its large kernel design. On the other hand, it addresses the quadratic complexity of attention by employing a sophisticated attention design, thus maintaining its ability to capture long-range dependencies. Furthermore, to achieve more precise and robust alignment of inter-frame features using optical flow for better utilization of clear frames, a mixture of experts model is proposed. It involves integrating optical flow updates between different experts in a residual manner. Our ablation experiments and experiments on multiple datasets indicate that ALK-MoE achieves comparable or superior performance compared to Transformer-based methods, with lower complexity. Chaopeng Zhang, Ruiqin Xiong, Xiaopeng Fan 0001, Debin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | High Dynamic Range Imaging for Dynamic Scenes Based on Multi-Level Spike CameraabstractSpike camera is a retina-inspired neuromorphic camera which can capture dynamic scenes of high-speed motion by firing a continuous stream of spikes at an extremely high temporal resolution. The limitation in the current design is that each spike only represents the arrival of a fixed amount of photons. It can not deal with strong light areas in which the amount of accumulated photons reaches the pre-specified threshold multiple times within a single readout interval. In this paper, we propose a new spike camera model of high-speed imaging for high dynamic range scenarios. In this scheme, each pixel accumulates the incoming photons persistently and generates a new type of spike stream in which each spike symbol may be associated with different levels, indicating the arrival of different amounts of photons since the last readout. This enables the camera to support dynamic scenes with wider dynamic range. To achieve this, we propose a two-level buffer mechanism, one for photon accumulation and one for spike-firing encoding. We use a register to hold the number of spike-firings which has not been read out yet. At each readout time, the major part in the counter is read out via a carefully designed exponential encoding and the counter is updated. Such encoding and readout strategy enables a very efficient expansion of the dynamic range using a small number of encoding bits. Furthermore, we propose an image reconstruction scheme for the proposed camera, utilizing both spike intervals and spike levels to recover the light intensity. We incorporate Mamba and propose a temporal-spatial selective scan mechanism to extract temporal-spatial correlation within spike streams. We employ a pyramid adaptive filtering and alignment module to achieve coarse-to-fine feature alignment. Experimental results show that the proposed scheme can achieve better imaging quality and outperform the existing spike camera in high dynamic range scenarios. Zhenkun Zhu 0001, Ruiqin Xiong, Jing Zhao 0011, Rui Zhao 0010, Xiaopeng Fan 0001, Shuyuan Zhu, Tiejun Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Color Spike Camera Reconstruction via Long Short-Term Temporal Aggregation of Spike SignalsabstractWith the prevalence of emerging computer vision applications, the demand for capturing dynamic scenes with high-speed motion has increased. A kind of neuromorphic sensor called spike camera shows great potential in this aspect since it generates a stream of binary spikes to describe the dynamic light intensity with a very high temporal resolution. Color spike camera (CSC) was recently invented to capture the color information of dynamic scenes via a color filter array (CFA) on the sensor. This paper proposes a long short-term temporal aggregation strategy of spike signals. First, we utilize short-term temporal correlation to adaptively extract temporal features of each time point. Then we align the features and aggregate them to exploit long-term temporal correlation, suppressing undesired motion blur. To implement the strategy, we design a CSC reconstruction network. Based on adaptive short-term temporal aggregation, we propose a spike representation module to extract temporal features of each color channel, leveraging multiple temporal scales. Considering the long-term temporal correlation, we develop an alignment module to align the temporal features. In particular, we perform motion alignment of red and blue channels with the guidance of the higher-sampling-rate green channel, leveraging motion consistency among color channels. Besides, we propose a module to aggregate the aligned temporal features for the restored color image, which exploits color channel correlation. We have also developed a CSC simulator for data generation. Experimental results demonstrate that our method can restore color images with fine texture details, achieving state-of-the-art CSC reconstruction performance. Yanchen Dong 0001, Ruiqin Xiong, Jing Zhao 0011, Xiaopeng Fan 0001, Xinfeng Zhang 0001, Tiejun Huang 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Spiking Variational Graph Representation Inference for Video SummarizationabstractWith the rise of short video content, efficient video summarization techniques for extracting key information have become crucial. However, existing methods struggle to capture the global temporal dependencies and maintain the semantic coherence of video content. Additionally, these methods are also influenced by noise during multi-channel feature fusion. We propose a Spiking Variational Graph (SpiVG) Network, which enhances information density and reduces computational complexity. First, we design a keyframe extractor based on Spiking Neural Networks (SNN), leveraging the event-driven computation mechanism of SNNs to learn keyframe features autonomously. To enable fine-grained and adaptable reasoning across video frames, we introduce a Dynamic Aggregation Graph Reasoner, which decouples contextual object consistency from semantic perspective coherence. We present a Variational Inference Reconstruction Module to address uncertainty and noise arising during multi-channel feature fusion. In this module, we employ Evidence Lower Bound Optimization (ELBO) to capture the latent structure of multi-channel feature distributions, using posterior distribution regularization to reduce overfitting. Experimental results show that SpiVG surpasses existing methods across multiple datasets such as SumMe, TVSum, VideoXum, and QFVS. Our codes and pre-trained models are available at https://github.com/liwrui/SpiVG. Wenrui Li 0001, Wei Han 0002, Liang-Jian Deng, Ruiqin Xiong, Xiaopeng Fan 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Super-Resolving Dynamic Scenes With Spike Camera via Multi-Frame Sequential Alignment With Motion PropagationabstractSpike camera is a neuromorphic sensor that can capture high-speed dynamic scenes by firing a continuous stream of binary spikes with extremely high temporal resolution, essentially forming a dense sampling in the temporal dimension. Due to the relative motion between camera and scene, each pixel is actually sampling at a large number of different spatial positions on the object in a short period. Converting this dense sampling from temporal dimension to spatial domain, high resolution images can be reconstructed from the spike stream. However, spike fluctuations and large motion in high-speed scenes pose great challenges for this task, especially for intensity information extraction and temporal alignment. In this paper, we propose a spike camera super resolution network to address these issues. Considering the local temporal correlation of spike stream and correlation consistency within a local region, we introduce a representation module that performs region-adaptive temporal filtering on spikes to mitigate fluctuations and extract stable intensity information from binary data. Additionally, we develop a module for multi-frame feature alignment, leveraging the long-term temporal information of spike stream. To handle large motions, we propagate the motion information from neighboring moment to current feature alignment module, which provides a prior that helps to narrow the search range for current motion offset, improving the accuracy of temporal alignment. Experimental results demonstrate that the proposed network achieves state-of-the-art performance on synthetic and real-captured spike data. Yuanlin Wang, Ruiqin Xiong, Jian Zhang 0018, Xinfeng Zhang 0001, Tiejun Huang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | GCN-Based Multi-Modality Fusion Network for Action RecognitionabstractThanks to the remarkably expressive power for depicting structural data, Graph Convolutional Network (GCN) has been extensively adopted for skeleton-based action recognition in recent years. However, GCN is designed to operate on irregular graphs of skeletons, making it difficult to deal with other modalities represented on regular grids directly. Thus, although existing works have demonstrated the necessity of multi-modality fusion, few methods in the literature explore the fusion of skeleton and other modalities within a GCN architecture. In this paper, we present a novel GCN-based framework, termed GCN-based Multi-modality Fusion Network (GMFNet), to efficiently utilize complementary information in RGB and skeleton data. GMFNet is constructed by connecting a main stream with a GCN-based multi-modality fusion module (GMFM), whose goal is to gradually combine finer and coarse action-related information extracted from skeletons and RGB videos, respectively. Specifically, a cross-modality data mapping method is designed to transform an RGB video into a$\mathit{skeleton-like}$(SL) sequence, which is then integrated with the skeleton sequence under a gradual fusion scheme in GMFM. The fusion results are fed into the following main stream to extract more discriminative features and produce the final prediction. In addition, a spatio-temporal joint attention mechanism is introduced for more accurate action recognition. Compared to the multi-stream approaches, GMFNet can be implemented within an end-to-end training pipeline and thereby reduces the training complexity. Experimental results show the proposed GMFNet achieves impressive performance on two large-scale data sets of NTU RGB+D 60 and 120. Shaocan Liu, Ruiqin Xiong, Xiaopeng Fan 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | DVSRNet: Deep Video Super-Resolution Based on Progressive Deformable Alignment and Temporal-Sparse EnhancementabstractVideo super-resolution (VSR) is used to compose high-resolution (HR) video from low-resolution video. Recently, the deformable alignment-based VSR methods are becoming increasingly popular. In these methods, the features extracted from video are aligned to eliminate the motion error targeting high super-resolution (SR) quality. However, these methods often suffer from misalignment and the lack of enough temporal information to compose HR frames, which accordingly induce artifacts in the SR result. In this article, we design a deep VSR network (DVSRNet) based on the proposed progressive deformable alignment (PDA) module and temporal-sparse enhancement (TSE) module. Specifically, the PDA module is designed to accurately align features and to eliminate artifacts via the bidirectional information propagation. The TSE module is constructed to further eliminate artifacts and to generate clear details for the HR frame. In addition, we construct a lightweight deep optical flow network (OFNet) to obtain the bidirectional optical flows for the implementation of the PDA module. Moreover, two new loss functions are designed for our proposed method. The first one is adopted in OFNet and the second one is constructed to guarantee the generation of sharp and clear details for the HR frames. The experimental results demonstrate that our method performs better than the state-of-the-art methods. Feiyu Chen 0001, Shuyuan Zhu, Yu Liu 0091, Ruiqin Xiong, Bing Zeng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Joint Demosaicing and Denoising for Spike CameraabstractAs a neuromorphic camera with high temporal resolution, spike camera can capture dynamic scenes with high-speed motion. Recently, spike camera with a color filter array (CFA) has been developed for color imaging. There are some methods for spike camera demosaicing to reconstruct color images from Bayer-pattern spike streams. However, the demosaicing results are bothered by severe noise in spike streams, to which previous works pay less attention. In this paper, we propose an iterative joint demosaicing and denoising network (SJDD-Net) for spike cameras based on the observation model. Firstly, we design a color spike representation (CSR) to learn latent representation from Bayer-pattern spike streams. In CSR, we propose an offset-sharing deformable convolution module to align temporal features of color channels. Then we develop a spike noise estimator (SNE) to obtain features of the noise distribution. Finally, a color correlation prior (CCP) module is proposed to utilize the color correlation for better details. For training and evaluation, we designed a spike camera simulator to generate Bayer-pattern spike streams with synthesized noise. Besides, we captured some Bayer-pattern spike streams, building the first real-world captured dataset to our knowledge. Experimental results show that our method can restore clean images from Bayer-pattern spike streams. The source codes and dataset are available at https://github.com/csycdong/SJDD-Net. Yanchen Dong 0001, Ruiqin Xiong, Jing Zhao 0011, Jian Zhang 0018, Xiaopeng Fan 0001, Shuyuan Zhu, Tiejun Huang 0001 |
AAAI | 2 |
| 2024 | Optical Flow for Spike Camera with Hierarchical Spatial-Temporal Spike FusionabstractAs an emerging neuromorphic camera with an asynchronous working mechanism, spike camera shows good potential for high-speed vision tasks. Each pixel in spike camera accumulates photons persistently and fires a spike whenever the accumulation exceeds a threshold. Such high-frequency fine-granularity photon recording facilitates the analysis and recovery of dynamic scenes with high-speed motion. This paper considers the optical flow estimation problem for spike cameras. Due to the Poisson nature of incoming photons, the occurrence of spikes is random and fluctuating, making conventional image matching inefficient. We propose a Hierarchical Spatial-Temporal (HiST) fusion module for spike representation to pursue reliable feature matching and develop a robust optical flow network, dubbed as HiST-SFlow. The HiST extracts features at multiple moments and hierarchically fuses the spatial-temporal information. We also propose an intra-moment filtering module to further extract the feature and suppress the influence of randomness in spikes. A scene loss is proposed to ensure that this hierarchical representation recovers the essential visual information in the scene. Experimental results demonstrate that the proposed method achieves state-of-the-art performance compared with the existing methods. The source codes are available at https://github.com/ruizhao26/HiST-SFlow. Rui Zhao 0010, Ruiqin Xiong, Jian Zhang 0018, Xinfeng Zhang 0001, Zhaofei Yu, Tiejun Huang 0001 |
AAAI | 2 |
| 2024 | Super-Resolution Reconstruction from Bayer-Pattern Spike StreamsabstractSpike camera is a neuromorphic vision sensor that can capture highly dynamic scenes by generating a continuous stream of binary spikes to represent the arrival of photons at very high temporal resolution. Equipped with Bayer color filter array (CFA), color spike camera (CSC) has been invented to capture color information. Although spike camera has already demonstrated great potential for high-speed imaging, its spatial resolution is limited compared with conventional digital cameras. This paper proposes a Color Spike Camera Super-Resolution (CSCSR) network to super-resolve higher-resolution color images from spike camera streams with Bayer CFA. To be specific, we first propose a representation for Bayer-pattern spike streams, exploring local temporal information with global perception to represent the binary data. Then we exploit the CFA layout and sub-pixel level motion to collect temporal pixels for the spatial super-resolution of each color channel. In particular, a residual-based module for feature refinement is developed to reduce the impact of motion estimation errors. Considering color correlation, we jointly utilize the multi-stage temporal-pixel features of color channels to reconstruct the high-resolution color image. Experimental results demonstrate that the proposed scheme can reconstruct satisfactory color images with both high temporal and spatial resolution from low-resolution Bayerpattern spike streams. The source codes are available at https://github.com/csycdong/CSCSR. Yanchen Dong 0001, Ruiqin Xiong, Jian Zhang 0018, Zhaofei Yu, Xiaopeng Fan 0001, Shuyuan Zhu, Tiejun Huang 0001 |
CVPR | 2 |
| 2024 | Boosting Spike Camera Image Reconstruction from a Perspective of Dealing with Spike FluctuationsabstractAs a bio-inspired vision sensor with ultra-high speed, spike cameras exhibit great potential in recording dynamic scenes with high-speed motion or drastic light changes. Different from traditional cameras, each pixel in spike cam-eras records the arrival of photons continuously by firing binary spikes at an ultra-fine temporal granularity. In this process, multiple factors impact the imaging, including the photons' Poisson arrival, thermal noises from circuits, and quantization effects in spike readout. These factors intro-duce fluctuations to spikes, making the recorded spike in-tervals unstable and unable to reflect accurate light intensi-ties. In this paper, we present an approach to deal with spike fluctuations and boost spike camera image reconstruction. We first analyze the quantization effects and reveal the unbi-ased estimation attribute of the reciprocal of differential of spike firing time (DSFT). Based on this, we propose a spike representation module to use DSFT with multiple orders for fluctuation suppression, where DSFT with higher or-ders indicates spike integration duration between multiple spikes. We also propose a module for inter-moment feature alignment at multiple granularities. The coarser alignment is based on patch-level cross-attention with a local search strategy, and the finer alignment is based on deformable convolution at the pixel level. Experimental results demon-strate the effectiveness of our method on both synthetic and real-captured data. The source code and dataset are avail-able at https://github.com/ruizhao26/BSF. Rui Zhao 0010, Ruiqin Xiong, Jing Zhao 0011, Jian Zhang 0018, Xiaopeng Fan 0001, Zhaofei Yu, Tiejun Huang 0001 |
CVPR | 2 |
| 2024 | Event-Based Visible and Infrared Fusion via Multi-Task CollaborationabstractVisible and Infrared image Fusion (VIF) offers a comprehensive scene description by combining thermal infrared images with the rich textures from visible cameras. However, conventional VIF systems may capture over/under exposure or blurry images in extreme lighting and high dynamic motion scenarios, leading to degraded fusion results. To address these problems, we propose a novel Event-based Visible and Infrared Fusion (EVIF) system that employs a visible event camera as an alternative to traditional frame-based cameras for the VIF task. With extremely low latency and high dynamic range, event cameras can effectively address blurriness and are robust against diverse luminous ranges. To produce high-quality fused images, we develop a multitask collaborative framework that simultaneously performs event-based visible texture reconstruction, event-guided infrared image deblurring, and visible-infrared fusion. Rather than independently learning these tasks, our framework capitalizes on their synergy, leveraging cross-task event enhancement for efficient deblurring and bi-level min-max mutual information optimization to achieve higher fusion quality. Experiments on both synthetic and real data show that EVIF achieves remarkable performance in dealing with extreme lighting conditions and high-dynamic scenes, ensuring high-quality fused images across a broad range of practical scenarios. Mengyue Geng, Lin Zhu 0012, Lizhi Wang 0001, Wei Zhang 0161, Ruiqin Xiong, Yonghong Tian 0001 |
CVPR | 5 |
| 2024 | Reconstruct Dynamic Scene for Spike Camera Based on 3D Space Time SimilarityabstractSpike camera is a neuromorphic camera that recurrently accumulates photons and fires spikes to record the incident light intensity at very high temporal resolution, making it particularly suitable for recording high dynamic scenes. This paper addresses the problem of image reconstruction for spike camera. Due to the Poisson effect of photon arrival and the quantization effect of spike readout, the spike interval calculated from a single spike cycle cannot reflect the light intensity accurately. Firstly, this paper analyzes the error of spike interval estimation under static light intensity. Then, this paper focuses on the temporal correlation of continuous spikes under dynamic light intensity. Specifically, it considers 3D space time similarity to weighted average multiple continuous spike intervals for the intensity estimation at a certain pixel. Experimental results demonstrate that the proposed method achieves better performance in both objective and subjective aspects compared with previous reconstruction methods. Yuanlin Wang, Ruiqin Xiong, Jing Zhao 0011, Tiejun Huang 0001 |
ICIP | 2 |
| 2024 | Mesh Denoising Using Filtering Coefficients Jointly Aware of Noise and GeometryabstractMesh denoising is a fundamental task in geometry processing, and recent studies have demonstrated the remarkable superiority of deep learning-based methods in this field. However, existing works commonly rely on neural networks without explicit designs for noise and geometry which are actually fundamental factors in mesh denoising. In this paper, by jointly considering noise intensity and geometric characteristics, a novel Filtering Coefficient Learner (FCL for short) for mesh denoising is developed, which delicately generates coefficients to filter face normals. Specifically, FCL produces filtering coefficients consisting of a noise-aware component and a geometry-aware component. The first component is inversely proportional to the noise intensity of each face, resulting in smaller coefficients for faces with stronger noise. For the effective assessment of the noise intensity, a noise intensity estimation module is designed, which predicts the angle between paired noisy-clean normals based on a mean filtering angle. The second component is derived based on two types of geometric features, namely the category feature and face-wise features. The category feature provides a global description of the input patch, while the face-wise features complement the perception of local textures. Extensive experiments have validated the superior performance of FCL over SOTA works in both noise removal and feature preservation. Xianqi Zhang, Wenxue Cui, Ruiqin Xiong, Xiaopeng Fan 0001, Debin Zhao |
ACM Multimedia | 4 |
| 2024 | Multi-Layer Probabilistic Association Reasoning Network for Image-Text RetrievalabstractWith the advancement of deep learning, the task of image-text retrieval has received widespread attention for addressing the semantic heterogeneity in multimodal data. However, many existing methods ignore the uncertainty present in manually annotated datasets. It is crucial for models to learn the potential corresponding relationships between regions in images and words in sentences. To tackle these challenges, we introduce the Multi-layer Probabilistic Association Reasoning Network (MPARN). In MPARN, the region-word association reasoning module is developed to treat each visual and textual fragment as unique probability distributions. This allows our model to imagine and capture the intricate one-to-many and many-to-many relationships between visual and textual objects. To effectively integrate the association distributions between visual and textual modalities, we propose the cross-modal association probability composer. This composer not only combines these distributions effectively but also preserves the intrinsic hierarchical structure of the elements involved. Furthermore, we introduce the semantic relationship reasoning module, which is designed to analyze the contextual semantic information within each modality. The multi-layer adaptive aggregate composer is employed to progressively explore semantic correlations within each modality and to dynamically synthesize outputs based on their relevance. Our extensive experiments on the Flickr30K and MSCOCO datasets demonstrate the MPARN’s state-of-the-art retrieval performance when compared to other baselines. The qualitative results further validate the effectiveness of the probabilistic association distributions. Wenrui Li 0001, Ruiqin Xiong, Xiaopeng Fan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Progressive Content-Aware Coded Hyperspectral Snapshot Compressive ImagingabstractHyperspectral imaging plays a pivotal role across diverse applications, like remote sensing, medicine, and cytology. The utilization of 2D sensors to acquire 3D hyperspectral images (HSIs) via a coded aperture snapshot spectral imaging (CASSI) system has proven successful, owing to its hardware-friendly implementation and fast sampling speed. Nevertheless, for less spectrally sparse scenes, the use of a single snapshot and unreasonable coded aperture design limits the efficacy of CASSI systems and renders HSI reconstruction more ill-posed, leading to compromised spatial and spectral fidelity. This paper proposes a novel Progressive Content-Aware CASSI (PCA-CASSI) framework, which progressively captures HSIs using multiple optimized content-aware coded apertures and fuses all snapshot measurements for reconstruction. By unlocking the full potential of CASSI systems and elevating their performance ceilings, this framework offers researchers new avenues for improving imaging quality. Furthermore, we develop the RndHRNet, a Range-Null space Decomposition (RND)-inspired deep unfolding network with multiple iterative phases for HSI recovery. Each unfolded recovery phase efficiently exploits the physical information within the coded apertures via explicit RND and adaptively explores the spatial-spectral correlation by dual transformer blocks. Through comprehensive experiments, our approach demonstrates superior performance compared to existing state-of-the-art methods in both the multiple- and single-shot compressive HSI imaging tasks with substantial improvements. Code is available athttps://github.com/xuanyuzhang21/PCA-CASSI. Xuanyu Zhang 0003, Bin Chen 0006, Wenzhen Zou, Yongbing Zhang 0002, Ruiqin Xiong, Jian Zhang 0018 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Spike Camera Image Reconstruction Using Deep Spiking Neural NetworksabstractSpike camera is a bio-inspired sensor with ultra-high temporal resolution and low energy consumption. It captures visual signals using an “integrate-and-fire" mechanism and outputs a continuous stream of binary spikes. Reconstructing image sequence from spikes streams is critical for spike camera. Several reconstruction methods have been proposed in recent years. However, the computational cost of these methods is relatively high. Inspired by the fact that spiking neural networks (SNNs) are energy efficient and support time-series signal processing inherently, we propose a lightweight SNN for spike camera image reconstruction (abbreviated to SSIR). Experimental results show that SSIR achieves comparable performance with the state-of-the-art (SOTA) methods at much lower computation and energy cost. Rui Zhao 0010, Ruiqin Xiong, Jian Zhang 0018, Zhaofei Yu, Shuyuan Zhu, Lei Ma 0008, Tiejun Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Learning a Deep Demosaicing Network for Spike Camera With Color Filter ArrayabstractFor capturing dynamic scenes with ultra-fast motion, neuromorphic cameras with extremely high temporal resolution have demonstrated their great capability and potential. Different from the event cameras that only record relative changes in light intensity, spike camera fires a stream of spikes according to a full-time accumulation of photons so that it can recover the texture details for both static areas and dynamic areas. Recently, color spike camera has been invented to record color information of dynamic scenes using a color filter array (CFA). However, demosaicing for color spike cameras is an open and challenging problem. In this paper, we develop a demosaicing network, called CSpkNet, to reconstruct dynamic color visual signals from the spike stream captured by the color spike camera. Firstly, we develop a light inference module to convert binary spike streams to intensity estimates. In particular, a feature-based channel attention module is proposed to reduce the noises caused by quantization errors. Secondly, considering both the Bayer configuration and object motion, we propose a motion-guided filtering module to estimate the missing pixels of each color channel, without undesired motion blur. Finally, we design a refinement module to improve the intensity and details, utilizing the color correlation. Experimental results demonstrate that CSpkNet can reconstruct color images from the Bayer-pattern spike stream with promising visual quality. Yanchen Dong 0001, Ruiqin Xiong, Jing Zhao 0011, Jian Zhang 0018, Xiaopeng Fan 0001, Shuyuan Zhu, Tiejun Huang 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | Spiking Tucker Fusion Transformer for Audio-Visual Zero-Shot LearningabstractThe spiking neural networks (SNNs) that efficiently encode temporal sequences have shown great potential in extracting audio-visual joint feature representations. However, coupling SNNs (binary spike sequences) with transformers (float-point sequences) to jointly explore the temporal-semantic information still facing challenges. In this paper, we introduce a novel Spiking Tucker Fusion Transformer (STFT) for audio-visual zero-shot learning (ZSL). The STFT leverage the temporal and semantic information from different time steps to generate robust representations. The time-step factor (TSF) is introduced to dynamically synthesis the subsequent inference information. To guide the formation of input membrane potentials and reduce the spike noise, we propose a global-local pooling (GLP) which combines the max and average pooling operations. Furthermore, the thresholds of the spiking neurons are dynamically adjusted based on semantic and temporal cues. Integrating the temporal and semantic information extracted by SNNs and Transformers are difficult due to the increased number of parameters in a straightforward bilinear model. To address this, we introduce a temporal-semantic Tucker fusion module, which achieves multi-scale fusion of SNN and Transformer outputs while maintaining full second-order interactions. Our experimental results demonstrate the effectiveness of the proposed approach in achieving state-of-the-art performance in three benchmark datasets. The harmonic mean (HM) improvement of VGGSound, UCF101 and ActivityNet are around 15.4%, 3.9%, and 14.9%, respectively. Wenrui Li 0001, Penghong Wang, Ruiqin Xiong, Xiaopeng Fan 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | A Universal Optimization Framework for Learning-based Image CodecabstractRecently, machine learning-based image compression has attracted increasing interests and is approaching the state-of-the-art compression ratio. But unlike traditional codec, it lacks a universal optimization method to seek efficient representation for different images. In this paper, we develop a plug-and-play optimization framework for seeking higher compression ratio, which can be flexibly applied to existing and potential future compression networks. To make the latent representation more efficient, we propose a novel latent optimization algorithm to adaptively remove the redundancy for each image. Additionally, inspired by the potential of side information for traditional codecs, we introduce side information into our framework, and integrate side information optimization with latent optimization to further enhance the compression ratio. In particular, with the joint side information and latent optimization, we can achieve fine rate control using only single model instead of training different models for different rate-distortion trade-offs, which significantly reduces the training and storage cost to support multiple bit rates. Experimental results demonstrate that our proposed framework can remarkably boost the machine learning-based compression ratio, achieving more than 10% additional bit rate saving on three different representative network structures. With the proposed optimization framework, we can achieve 7.6% bit rate saving against the latest traditional coding standard VVC on Kodak dataset, yielding the state-of-the-art compression ratio. Jing Zhao 0011, Bin Li 0012, Jiahao Li 0001, Ruiqin Xiong, Yan Lu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | SVFI: Spiking-Based Video Frame Interpolation for High-Speed MotionabstractOcclusion and motion blur make it challenging to interpolate video frame, since estimating complex motions between two frames is hard and unreliable, especially in highly dynamic scenes. This paper aims to address these issues by exploiting spike stream as auxiliary visual information between frames to synthesize target frames. Instead of estimating motions by optical flow from RGB frames, we present a new dual-modal pipeline adopting both RGB frames and the corresponding spike stream as inputs (SVFI). It extracts the scene structure and objects' outline feature maps of the target frames from spike stream. Those feature maps are fused with the color and texture feature maps extracted from RGB frames to synthesize target frames. Benefited by the spike stream that contains consecutive information between two frames, SVFI can directly extract the information in occlusion and motion blur areas of target frames from spike stream, thus it is more robust than previous optical flow-based methods. Experiments show SVFI outperforms the SOTA methods on wide variety of datasets. For instance, in 7 and 15 frame skip evaluations, it shows up to 5.58 dB and 6.56 dB improvements in terms of PSNR over the corresponding second best methods BMBC and DAIN. SVFI also shows visually impressive performance in real-world scenes. Lujie Xia, Jing Zhao 0011, Ruiqin Xiong, Tiejun Huang 0001 |
AAAI | 3 |
| 2023 | Learning to Super-resolve Dynamic Scenes for Neuromorphic Spike CameraabstractSpike camera is a kind of neuromorphic sensor that uses a novel ``integrate-and-fire'' mechanism to generate a continuous spike stream to record the dynamic light intensity at extremely high temporal resolution. However, as a trade-off for high temporal resolution, its spatial resolution is limited, resulting in inferior reconstruction details. To address this issue, this paper develops a network (SpikeSR-Net) to super-resolve a high-resolution image sequence from the low-resolution binary spike streams. SpikeSR-Net is designed based on the observation model of spike camera and exploits both the merits of model-based and learning-based methods. To deal with the limited representation capacity of binary data, a pixel-adaptive spike encoder is proposed to convert spikes to latent representation to infer clues on intensity and motion. Then, a motion-aligned super resolver is employed to exploit long-term correlation, so that the dense sampling in temporal domain can be exploited to enhance the spatial resolution without introducing motion blur. Experimental results show that SpikeSR-Net is promising in super-resolving higher-quality images for spike camera. Jing Zhao 0011, Ruiqin Xiong, Jian Zhang 0018, Rui Zhao 0010, Hangfan Liu, Tiejun Huang 0001 |
AAAI | 2 |
| 2023 | Optimization-Inspired Deep Network for Image Restoration from Partial Random SamplesabstractImage Restoration from Partial Random Samples (RRS) has been studied in many image restoration works. There are also some attempts to use convolutional neural networks (CNNs) to handle it. However, most existing neural network-based methods perform poorly in generalization and we need to train a specific model for each degradation situation. Besides, the sampling mask which represents the positions of the sampled pixels is not used effectively in these methods. To address the problems, we propose an optimization-inspired network called RRSNet based on our derivation of the iterative optimization formulas for RRS. In our method, we design a CNN with two encoders and one decoder for training, setting up a flexible and effective prior. To make the most of the sampling information, we concatenate the degraded image with the mask and input them into one encoder for better generalization. Then we split the pixels into two groups according to the mask and extract their features as the input of another encoder. Experiments demonstrate that our RRSNet with the mask input can handle various sampling ratios using only one trained model and achieve the best restoration performance among all comparison methods. Yanchen Dong 0001, Rui Zhao 0010, Ruiqin Xiong, Shuyuan Zhu, Xiaopeng Fan 0001, Tiejun Huang 0001 |
ISCAS | 3 |
| 2023 | Unsupervised Optical Flow Estimation with Dynamic Timing Representation for Spike CameraabstractEfficiently selecting an appropriate spike stream data length to extract precise information is the key to the spike vision tasks. To address this issue, we propose a dynamic timing representation for spike streams. Based on multi-layers architecture, it applies dilated convolutions on temporal dimension to extract features on multi-temporal scales with few parameters. And we design layer attention to dynamically fuse these features. Moreover, we propose an unsupervised learning method for optical flow estimation in a spike-based manner to break the dependence on labeled data. In addition, to verify the robustness, we also build a spike-based synthetic validation dataset for extreme scenarios in autonomous driving, denoted as SSES dataset. It consists of various corner cases. Experiments show that our method can predict optical flow from spike streams in different high-speed scenes, including real scenes. For instance, our method achieves $15\%$ and $19\%$ error reduction on PHM dataset compared to the best spike-based work, SCFlow, in $\Delta t=10$ and $\Delta t=20$ respectively, using the same settings as in previous works. The source code and dataset are available at \href{https://github.com/Bosserhead/USFlow}{https://github.com/Bosserhead/USFlow}. Lujie Xia, Ziluo Ding, Rui Zhao 0010, Jiyuan Zhang 0005, Lei Ma 0008, Zhaofei Yu, Tiejun Huang 0001, Ruiqin Xiong |
NeurIPS | 8 |
| 2023 | FCNet: Learning Noise-Free Features for Point Cloud DenoisingabstractThe acquisition of point clouds is usually accompanied by noise due to imperfect laser scanning or image-based reconstruction techniques. Deep learning-based methods have achieved impressive performance in point cloud denoising. However, the features captured by a denoising network from noisy point clouds are usually contaminated by noise during training. The feature noise will lead to the oscillation of back-propagated gradients, which interferes with parameter optimization and reduces the denoising performance. In this paper, we propose to explicitly clean up feature noise for point cloud denoising from two aspects: feature noise cleaning and network training. From the first aspect, we propose the feature clean network (FCNet for short) to explicitly clean up the feature noise. From the second aspect, we train FCNet by a teacher-student learning model to learn the noise-free features under the guidance of feature domain losses. Specifically, FCNet is designed with emphasis on two modules: non-local self-similarity (NSS) and weighted average pooling (WAP). NSS module smooths features through a non-local filter based on the inherent non-local self-similarity of point clouds. WAP module applies original weights calculated by the statistical outlier removal algorithm to suppress the feature noise induced by outliers. In the teacher-student learning model, we introduce a clean input using the noisy point and its clean neighbors. The teacher network accepts the clean input to capture noise-free features. The student network is trained to imitate the teacher network to learn noise-free features by minimizing the feature loss. The experiments on synthetic and real scanned point clouds show that FCNet outperforms state-of-the-art point cloud denoising methods. Wenxue Cui, Ruiqin Xiong, Xiaopeng Fan 0001, Debin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Tree-Structured Data Clustering-Driven Neural Network for Intra Prediction in Video CodingabstractIntra prediction is a crucial part of video compression, which utilizes local information in images to eliminate spatial redundancy. As the state-of-the-art video coding standard, Versatile Video Coding (H.266/VVC) employs multiple directional prediction modes in intra prediction to find the texture trend of local areas. Then the prediction is made based on reference samples in the selected direction. Recently, neural network-based intra prediction has achieved great success. Deep network models are trained and applied to assist the HEVC and VVC intra modes. In this paper, we propose a novel tree-structured data clustering-driven neural network (dubbed TreeNet) for intra prediction, which builds the networks and clusters the training data in a tree-structured manner. Specifically, in each network split and training process of TreeNet, every parent network on a leaf node is split into two child networks by adding or subtracting Gaussian random noise. Then data clustering-driven training is applied to train the two derived child networks using the clustered training data of their parent. On the one hand, the networks at the same level in TreeNet are trained with non-overlapping clustered datasets, and thus they can learn different prediction abilities. On the other hand, the networks at different levels are trained with hierarchically clustered datasets, and thus they will have different generalization abilities. TreeNet is integrated into VVC to assist or replace intra prediction modes to test its performance. In addition, a fast termination strategy is proposed to accelerate the search of TreeNet. The experimental results demonstrate that when TreeNet is used to assist the VVC Intra modes, TreeNet with depth = 3 can bring an average of 3.78% bitrate saving (up to 8.12%) over VTM-17.0. If TreeNet with the same depth replaces all VVC intra modes, an average of 1.59% bitrate saving can be reached. Hengyu Man, Xiaopeng Fan 0001, Ruiqin Xiong, Debin Zhao |
IEEE Trans. Image Process. | 3 |
| 2023 | CI-Net: Clinical-Inspired Network for Automated Skin Lesion RecognitionabstractThe lesion recognition of dermoscopy images is significant for automated skin cancer diagnosis. Most of the existing methods ignore the medical perspective, which is crucial since this task requires a large amount of medical knowledge. A few methods are designed according to medical knowledge, but they ignore to be fully in line with doctors' entire learning and diagnosis process, since certain strategies and steps of those are conducted in practice for doctors. Thus, we put forward Clinical-Inspired Network (CI-Net) to involve the learning strategy and diagnosis process of doctors, as for a better analysis. The diagnostic process contains three main steps: the zoom step, the observe step and the compare step. To simulate these, we introduce three corresponding modules: a lesion area attention module, a feature extraction module and a lesion feature attention module. To simulate the distinguish strategy, which is commonly used by doctors, we introduce a distinguish module. We evaluate our proposed CI-Net on six challenging datasets, including ISIC 2016, ISIC 2017, ISIC 2018, ISIC 2019, ISIC 2020 and PH2 datasets, and the results indicate that CI-Net outperforms existing work. The code is publicly available at https://github.com/lzh19961031/Dermoscopy_classification. Zihao Liu 0009, Ruiqin Xiong, Tingting Jiang 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Spatio-Temporal Recurrent Networks for Event-Based Optical Flow EstimationabstractEvent camera has offered promising alternative for visual perception, especially in high speed and high dynamic range scenes. Recently, many deep learning methods have shown great success in providing model-free solutions to many event-based problems, such as optical flow estimation. However, existing deep learning methods did not address the importance of temporal information well from the perspective of architecture design and cannot effectively extract spatio-temporal features. Another line of research that utilizes Spiking Neural Network suffers from training issues for deeper architecture. To address these points, a novel input representation is proposed that captures the events temporal distribution for signal enhancement. Moreover, we introduce a spatio-temporal recurrent encoding-decoding neural network architecture for event-based optical flow estimation, which utilizes Convolutional Gated Recurrent Units to extract feature maps from a series of event images. Besides, our architecture allows some traditional frame-based core modules, such as correlation layer and iterative residual refine scheme, to be incorporated. The network is end-to-end trained with self-supervised learning on the Multi-Vehicle Stereo Event Camera dataset. We have shown that it outperforms all the existing state-of-the-art methods by a large margin. Ziluo Ding, Rui Zhao 0010, Jiyuan Zhang 0005, Tianxiao Gao, Ruiqin Xiong, Zhaofei Yu, Tiejun Huang 0001 |
AAAI | 5 |
| 2022 | Optical Flow Estimation for Spiking CameraabstractAs a bio-inspired sensor with high temporal resolution, the spiking camera has an enormous potential in real applications, especially for motion estimation in high-speed scenes. However, frame-based and event-based methods are not well suited to spike streams from the spiking camera due to the different data modalities. To this end, we present, SCFlow, a tailored deep learning pipeline to estimate optical flow in high-speed scenes from spike streams. Importantly, a novel input representation is introduced which can adaptively remove the motion blur in spike streams according to the prior motion. Further, for training SCFlow, we synthesize two sets of optical flow data for the spiking camera, SPIkingly Flying Things and Photo-realistic Highspeed Motion, denoted as SPIFT and PHM respectively, corresponding to random high-speed and well-designed scenes. Experimental results show that the SCFlow can predict optical flow from spike streams in different high-speed scenes. Moreover, SCFlow shows promising generalization on real spike streams. Codes and datasets refer to https://github.com/Acnext/Optical-Flow-For-Spiking-Camera. Liwen Hu 0002, Rui Zhao 0010, Ziluo Ding, Lei Ma 0008, Boxin Shi, Ruiqin Xiong, Tiejun Huang 0001 |
CVPR | 6 |
| 2022 | HerosNet: Hyperspectral Explicable Reconstruction and Optimal Sampling Deep Network for Snapshot Compressive ImagingabstractHyperspectral imaging is an essential imaging modality for a wide range of applications, especially in remote sensing, agriculture, and medicine. Inspired by existing hyperspectral cameras that are either slow, expensive, or bulky, reconstructing hyperspectral images (HSIs) from a low-budget snapshot measurement has drawn wide attention. By mapping a truncated numerical optimization algorithm into a network with a fixed number of phases, recent deep unfolding networks (DUNs) for spectral snapshot compressive sensing (SCI) have achieved remarkable success. However, DUNs are far from reaching the scope of industrial applications limited by the lack of cross-phase feature interaction and adaptive parameter adjustment. In this paper, we propose a novel Hyperspectral Explicable Reconstruction and Optimal Sampling deep Network for SCI, dubbed HerosNet, which includes several phases under the ISTA-unfolding framework. Each phase can flexibly simulate the sensing matrix and contextually adjust the step size in the gradient descent step, and hierarchically fuse and interact the hidden states of previous phases to effectively recover current HSI frames in the proximal mapping step. Simultaneously, a hardware-friendly optimal binary mask is learned end-to-end to further improve the reconstruction performance. Finally, our HerosNet is validated to outperform the state-of-the-art methods on both simulation and real datasets by large margins. The source code is available at https://github.com/jianzhangcs/HerosNet. Xuanyu Zhang 0003, Yongbing Zhang 0002, Ruiqin Xiong, Qilin Sun 0001, Jian Zhang 0018 |
CVPR | 3 |
| 2022 | 3D Residual Interpolation for Spike Camera DemosaicingabstractThe recently invented spike camera can capture high-speed motion in dynamic scenes by accumulating incoming photons continuously and firing spikes at very high temporal resolution. This paper addresses the demosaicing problem in spike camera color imaging. Specifically, we propose the 3D residual interpolation (3DRI) method to convert raw spike frames to color image frames. Due to the Poisson effect of photon arrivals and the quantization effect of spike readout, the instantaneous intensity recovered from the spike stream may suffer from undesired noise. To handle the noise, we estimate the missing color pixels along motion trajectories to exploit the temporal correlation among neighboring frames. In addition, by utilizing the color channels correlation, we design a residual-based demosaicing pipeline that uses the green pixels to guide the estimation of the red or blue missing pixels. Experimental results demonstrate our proposed 3DRI can produce color images from spike streams, achieving a good objective and perceptual quality for high-motion scenes. Yanchen Dong 0001, Jing Zhao 0011, Ruiqin Xiong, Tiejun Huang 0001 |
ICIP | 3 |
| 2022 | Self-Supervised Mutual Learning for Dynamic Scene Reconstruction of Spiking CameraabstractMimicking the sampling mechanism of the primate fovea, a retina-inspired vision sensor named spiking camera has been developed, which has shown great potential for capturing high-speed dynamic scenes with a sampling rate of 40,000 Hz. Unlike conventional digital cameras, the spiking camera continuously captures photons and outputs asynchronous binary spikes with various inter-spike intervals to record dynamic scenes. However, how to reconstruct dynamic scenes from asynchronous spike streams remains challenging. In this work, we propose a novel pretext task to build a self-supervised reconstruction framework for spiking cameras. Specifically, we utilize the blind-spot network commonly used in self-supervised denoising tasks as our backbone, and perform self-supervised learning by constructing proper pseudo-labels. In addition, in view of the poor scalability and insufficient information utilization of the blind-spot network, we present a mutual learning framework to improve the overall performance of the network through mutual distillation between a non-blind-spot network and a blind-spot network. This also enables the network to bypass constraints of the blind-spot network, allowing state-of-the-art modules to be used to further improve performance. The experimental results demonstrate that our methods evidently outperform previous unsupervised spiking camera reconstruction methods and achieve desirable results compared with supervised methods. Shiyan Chen, Chaoteng Duan, Zhaofei Yu, Ruiqin Xiong, Tiejun Huang 0001 |
IJCAI | 4 |
| 2022 | Deep Feature Compression with Collaborative Coding of Image TextureabstractIn this paper, we propose a coding scheme for the deep intermediate feature and it is implemented with the collaborative compression of image texture. More specifically, we separately compress the feature and texture of the image to form two data layers. The first one is the intermediate feature layer and the second one is the texture layer. The texture layer can provide an image for users and the feature layer can be used to implement the computer vision (CV) task. With our proposed deep reconstruction network (RecNet), the texture and features cooperate to achieve a high-quality visual output as well as a high-efficiency CV task. The experimental results demonstrate the excellent performance by using our proposed method to compress the deep features. Hewei Liu, Shuyuan Zhu, Xiaozhen Zheng, Ruiqin Xiong, Bing Zeng 0001 |
ISCAS | 5 |
| 2022 | Learning Optical Flow from Continuous Spike StreamsabstractSpike camera is an emerging bio-inspired vision sensor with ultra-high temporal resolution. It records scenes by accumulating photons and outputting continuous binary spike streams. Optical flow is a key task for spike cameras and their applications. A previous attempt has been made for spike-based optical flow. However, the previous work only focuses on motion between two moments, and it uses graphics-based data for training, whose generalization is limited. In this paper, we propose a tailored network, Spike2Flow that extracts information from binary spikes with temporal-spatial representation based on the differential of spike firing time and spatial information aggregation. The network utilizes continuous motion clues through joint correlation decoding. Besides, a new dataset with real-world scenes is proposed for better generalization. Experimental results show that our approach achieves state-of-the-art performance on existing synthetic datasets and real data captured by spike cameras. The source code and dataset are available at \url{https://github.com/ruizhao26/Spike2Flow}. Rui Zhao 0010, Ruiqin Xiong, Jing Zhao 0011, Zhaofei Yu, Xiaopeng Fan 0001, Tiejun Huang 0001 |
NeurIPS | 2 |
| 2022 | High-Speed Scene Reconstruction from Low-Light Spike StreamsabstractBenefiting from the high temporal resolution, the spike camera shows promising potential in capturing high-speed scenes via accumulating luminance intensity and firing spikes. Although the spike camera compared to the high-speed camera is quite cost-effective, its performance in capturing low-light scenes is poor. Specifically, it takes more time for the spike camera to accumulate enough light signal for firing a spike in low-light scenes, while the scenes may have already changed because of the high-speed motion. There may be no effective spikes for a long time due to the insufficient incident light, and the signal-to-noise ratio of spike streams in low-light scenes is unsatisfactory. Thus, it's easy to introduce noise and motion blur while reconstructing, especially in rapidly changing scenes. To address this issue, we propose a low-light scene reconstruction method for the spike camera. In particular, we first develop a Brightness-Adaptive Light Inference (BALI) method to preliminarily reconstruct the low-light scene according to the brightness, which utilizes both the spike interval and the spike number. Considering the motion, we then estimate optical flow and filter the preliminary restored frames iteratively to handle the noise via temporal correlation. After that, there is still some noise, and we further handle it by a spatial filter according to the brightness. As a result, we restore a clear image based on both temporal and spatial correlation. The experimental results demonstrate that our method achieves good visual quality in low-light scene reconstruction. Yanchen Dong 0001, Jing Zhao 0011, Ruiqin Xiong, Tiejun Huang 0001 |
VCIP | 3 |
| 2022 | Spike Signal Reconstruction Based on Inter-Spike SimilarityabstractSpike camera is a kind of bio-inspired camera which is particularly proposed for capturing dynamic scenes with high speed motion. Spike camera works in a way simulating the retina that it receives incoming photons continuously and fires a spike whenever the accumulated photons reach a threshold. The spike stream can be recorded at an extremely high temporal resolution so that the dynamic process of light-intensity changes may be recovered. This paper addresses the problem of recovering the original visual signal from spikes. To reduce the fluctuation in spike intervals caused by the Poisson effect of photon arrivals and the quantization effect in spike reading, we estimate the real interval from a sequence of temporally neighboring spikes. To avoid mixing the spikes generated from significantly different light intensities, we propose a temporal and spatial weighting method based on the inter-spike similarity. Experimental results demonstrate that the proposed method outperforms the previous light intensity inference methods and achieves better performance under different motion conditions. Ruiqin Xiong, Tiejun Huang 0001 |
VCIP | 2 |
| 2022 | Neural Network-Based Enhancement to Inter Prediction for Video CodingabstractInter prediction is a crucial part of hybrid video coding frameworks, utilized to exploit the temporal redundancy in video sequences and improve the coding performance. During inter prediction, a predicted block is typically derived from reference pictures using motion estimation and motion compensation. To improve the coding performance of inter prediction, a neural network based enhancement to inter prediction (NNIP) is proposed in this paper. NNIP is composed of three networks, namely residue estimation network, combination network, and deep refinement network. Specifically, first, a residue estimation network is designed to estimate the residue between current block and its predicted block using their available spatial neighbors. Second, the feature maps of the estimated residue and the predicted block are extracted and concatenated in a combination network. Finally, the concatenated feature maps are fed into a deep refinement network to generate a refined residue, which is added back to the predicted block to derive a more accurate predicted block. NNIP is integrated in HEVC to evaluate its efficiency. The experimental results demonstrate that NNIP can achieve 4.6%, 3.0%, and 2.7% BD-rate reduction on average under LDP, LDB, and RA configurations compared to HEVC. Yang Wang 0048, Xiaopeng Fan 0001, Ruiqin Xiong, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Depth Map Super-Resolution Based on Dual Normal-Depth Regularization and Graph Laplacian PriorabstractThe edge information plays a key role in the restoration of a depth map. Most conventional methods assume that the color image and depth map are consistent in edge areas. However, complex texture regions in the color image do not match exactly with edges in the depth map. In this paper, firstly, we point out that in most cases the consistency between normal map and depth map is much higher than that between RGB-D pairs. Then we propose a dual normal-depth regularization term to guide the restoration of depth map, which constrains the edge consistency between normal map and depth map back and forth. Moreover, considering the bimodal characteristic of weight distribution that exists in depth discontinuous areas, a reweighted graph Laplacian regularizer is proposed to promote this bimodal characteristic. And this regularization is incorporated into a unified optimization framework to effectively protect the piece-wise smoothness(PWS) characteristics of depth map. By treating depth image as graph signal, the weight between two nodes is adapted according to its content. The proposed method is tested for both noise-free and noisy cases, and is compared against the state-of-the-art methods on both synthesis and real captured datasets. Extensive experimental results demonstrate the superior performance of our method compared with most state-of-the-art works in terms of both objective and subjective quality evaluations. Specifically, our method is more effective on edge areas and more robust to noises. Jin Wang 0023, Longhua Sun, Ruiqin Xiong, Yunhui Shi, Qing Zhu 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | MRDFlow: Unsupervised Optical Flow Estimation Network With Multi-Scale Recurrent DecoderabstractOptical flow estimation is a fundamental task in computer vision and image processing. Due to the difficulty in obtaining the ground truth of flow field, unsupervised learning approaches attract more and more research interests in recent years. However, despite of their good generalization capability, unsupervised optical flow methods suffer in the scenarios with large displacement, small objects, and occlusions. In this work, we propose a novel optical flow network based on decoder with multi-scale kernels. Different from previous U-Net like or pyramidal methods, we design our network based on RAFT architecture that with a 4D correlation layer and recurrent decoder. More importantly, we incorporate three novel ideas with regard to the input, information processing and output of the update units improve the performance. Firstly, we utilize various motion-related information as input to the update units. Secondly, we propose a module of multi-scale update unit. Thirdly, for the final flow up-sampling procedure, we propose an image-guided up-sampling loss to guide the learning of up-sampling masks. Our model is trained by the occlusion-aware photometric loss, edge-aware smoothness loss, self-supervised loss, and image-guided up-sampling loss. Experimental results demonstrate that our model achieves the state-of-the-art performance on both Sintel and KITTI and outperforms other unsupervised optical flow methods remarkably. Rui Zhao 0010, Ruiqin Xiong, Ziluo Ding, Xiaopeng Fan 0001, Jian Zhang 0018, Tiejun Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Contrastive and Selective Hidden Embeddings for Medical Image SegmentationabstractMedical image segmentation is fundamental and essential for the analysis of medical images. Although prevalent success has been achieved by convolutional neural networks (CNN), challenges are encountered in the domain of medical image analysis by two aspects: 1) lack of discriminative features to handle similar textures of distinct structures and 2) lack of selective features for potential blurred boundaries in medical images. In this paper, we extend the concept of contrastive learning (CL) to the segmentation task to learn more discriminative representation. Specifically, we propose a novel patch-dragsaw contrastive regularization (PDCR) to perform patch-level tugging and repulsing. In addition, a new structure, namely uncertainty-aware feature re- weighting block (UAFR), is designed to address the potential high uncertainty regions in the feature maps and serves as a better feature re- weighting. Our proposed method achieves state-of-the-art results across 8 public datasets from 6 domains. Besides, the method also demonstrates robustness in the limited-data scenario. The code is publicly available at https://github.com/lzh19961031/PDCR_UAFR-MIShttps://github.com/lzh19961031/PDCR_UAFR-MIS. Zihao Liu 0009, Zhuowei Li 0002, Qing Xia 0002, Ruiqin Xiong, Shaoting Zhang 0001, Tingting Jiang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Spk2ImgNet: Learning To Reconstruct Dynamic Scene From Continuous Spike StreamabstractThe recently invented retina-inspired spike camera has shown great potential for capturing dynamic scenes. Different from the conventional digital cameras that compact the photoelectric information within the exposure interval into a single snapshot, the spike camera produces a continuous spike stream to record the dynamic light intensity variation process. For spike cameras, image reconstruction remains an important and challenging issue. To this end, this paper develops a spike-to-image neural network (Spk2ImgNet) to reconstruct the dynamic scene from the continuous spike stream. In particular, to handle the challenges brought by both noise and high-speed motion, we propose a hierarchical architecture to exploit the temporal correlation of the spike stream progressively. Firstly, a spatially adaptive light inference subnet is proposed to exploit the local temporal correlation, producing basic light intensity estimates of different moments. Then, a pyramid deformable alignment is utilized to align the intermediate features such that the feature fusion module can exploit the long-term temporal correlation, while avoiding undesired motion blur. In addition, to train the network, we simulate the working mechanism of spike camera to generate a large-scale spike dataset composed of spike streams and corresponding ground truth images. Experimental results demonstrate that the proposed network evidently outperforms the state-of-the-art spike camera reconstruction methods. Jing Zhao 0011, Ruiqin Xiong, Hangfan Liu, Jian Zhang 0018, Tiejun Huang 0001 |
CVPR | 2 |
| 2021 | Super Resolve Dynamic Scene from Continuous Spike StreamsabstractRecently, a novel retina-inspired camera, namely spike camera, has shown great potential for recording high-speed dynamic scenes. Unlike conventional digital cameras that compact the visual information within an exposure interval into a single snapshot, the spike camera continuously outputs binary spike streams to record the dynamic scenes, yielding a very high temporal resolution. Most of the existing reconstruction methods for spike camera focus on reconstructing images with the same resolution as spike camera. However, as a trade-off of high temporal resolution, the spatial resolution of spike camera is limited, resulting in inferior details of the reconstruction. To address this issue, we develop a spike camera super-resolution framework, aiming to super resolve high-resolution intensity images from the low-resolution binary spike streams. Due to the relative motion between the camera and the objects to capture, the spikes fired by the same sensor pixel no longer describes the same points in the external scene. In this paper, we exploit the relative motion and derive the relationship between light intensity and each spike, so as to recover the external scene with both high temporal and high spatial resolution. Experimental results demonstrate that the proposed method can reconstruct pleasant high-resolution images from low- resolution spike streams. Jing Zhao 0011, Jiyu Xie, Ruiqin Xiong, Jian Zhang 0018, Zhaofei Yu, Tiejun Huang 0001 |
ICCV | 3 |
| 2021 | Recover The Residual Of Residual: Recurrent Residual Refinement Network For Image Super-ResolutionabstractBenefiting from learning the residual between low resolution (LR) image and high resolution (HR) image, image super-resolution (SR) networks demonstrate superior reconstruction performance in recent studies. However, for the images with rich texture information, the residuals are complex and difficult for networks to learn. To address this problem, we propose a recurrent residual refinement network (RRRN) to gradually refine the residual with a recurrent structure. Instead of directly reconstructing the residual between LR image and HR image, each sub-network in our framework reconstructs the residual between SR image from previous stage and HR image, i.e. recovers the residual of residual (RoR). Considering the domain gap between the image feature and the RoR feature, we introduce a residual projection block to explicitly transform the feature from image domain to RoR domain. The RoR feature is further optimized in an iterative up- and down-sampling manner with a residual learning block. We construct the structure of each block based on the optimization methods of conventional SR and improve our network with dense connections. Experimental results prove that our method improves the quality of super-resolution images on different datasets with variable scenes. Tianxiao Gao, Ruiqin Xiong, Rui Zhao 0010, Jian Zhang 0018, Shuyuan Zhu, Tiejun Huang 0001 |
ICIP | 2 |
| 2021 | Dual Regularization Based Depth Map Super-Resolution with Graph Laplacian PriorabstractThe edge information plays key role in the restoration of depth map. Most conventional methods assume that the RGB-D pairs are consistent in edge areas. In this paper, firstly, we point out that in most cases the consistency between normal map and depth map(N-D pairs) are much higher than that be-tween RGB-D pairs. Then we propose a dual regularization term to guide the restoration of depth map, which constrains the consistency between N-D pairs back and forth. Moreover, a reweighted graph Laplacian prior is incorporated into a unified optimization framework to effectively protect piece-wise smoothness(PWS) characteristics of depth map. By treating depth maps as graph signals, the weight between two nodes is adapted according to its content. Extensive experimental results demonstrate the superior performance of our method compared with other state-of-the-art works in terms of objective and subjective quality evaluations. Longhua Sun, Jin Wang 0023, Ruiqin Xiong, Yunhui Shi, Qing Zhu 0004 |
ICME | 3 |
| 2021 | Multi-level Relationship Capture Network for Automated Skin Lesion Recognition
Zihao Liu 0009, Ruiqin Xiong, Tingting Jiang 0001 |
MICCAI (7) | 2 |
| 2021 | Cross-Block Difference Guided Fast CU Partition for VVC Intra CodingabstractIn this paper, we propose a new fast CU partition method for VVC intra coding based on the cross-block difference. This difference is measured by the gradient and the content of sub-blocks obtained from partition and is employed to guide the skipping of unnecessary horizontal and vertical partition modes. With this guidance, a fast determination of block partitions is accordingly achieved. Compared with VVC, our proposed method can save 41.64% (on average) encoding time with only 0.97% (on average) increase of BD-rate. Hewei Liu, Shuyuan Zhu, Ruiqin Xiong, Guanghui Liu 0001, Bing Zeng 0001 |
VCIP | 3 |
| 2021 | NTSDCN: New Three-Stage Deep Convolutional Image Demosaicking NetworkabstractIn this letter, we compose a new three-stage deep convolutional neural network (NTSDCN) for image demosaicking, and it consists of our proposed Laplacian energy-constrained local residual unit (LC-LRU) and a feature-guided prior fusion unit (FG-PFU). Specifically, the LC-LRU is used to refine the learning target of the specific residual blocks in the network and enhance the dominant information of the residual features. The FG-PFU is designed to guide the feature extraction of the red (R) and blue (B) channels by utilizing prior information from the reconstructed green (G) channel. In our proposed NTSDCN, we recover the G channel image in the first stage with the CFA image and reconstruct the R and B images in the second stage. Finally, we fine-tune the resulting R, G and B images in the third stage to compose a full-color RGB image. The experimental results show that our proposed method achieves better performance than the state-of-the-art methods. The code is available at https://github.com/wyannn/NTSDCN. Shiying Yin, Shuyuan Zhu, Zhan Ma 0001, Ruiqin Xiong, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Light Field Image Compression Using Multi-branch Spatial Transformer Networks Based View SynthesisabstractThe recent years have witnessed the widespread of light field imaging in interactive and immersive visual applications. To record the directional information of the light rays, larger storage space is required by light field images compared with conventional 2D images. Hence, the efficient compression of light field image is highly desired for further applications. In this paper, we propose a novel light field image compression scheme using multi-branch spatial transformer networks based view synthesis. Firstly, a sparse subset of views are selected and are rearranged into a pseudo sequence to be encoded by a video codec at encoder. Then the other unselected views are synthesized based on the similarity between neighboring views with our proposed method at decoder. To better characterize the non-linear relationship between the sub-views, a multi-branch spatial transformer networks (MSTN) is designed to adaptively learn the affine transformations between the neighboring views, which are used to warp the input views to generate accurate approximation of the target views. Moreover, to better obtain the final view by the generated approximation views, the Wasserstein generative adversarial networks(WGAN) is applied with the improved training. Experimental results show the superior compression performance of our scheme compared with the state-of-the-art methods. Jin Wang 0023, Ruiqin Xiong, Qing Zhu 0004 |
DCC | 3 |
| 2020 | High-Speed Motion Scene Reconstruction for Spike Camera via Motion Aligned FilteringabstractA new retina-inspired bionic spike camera has recently shown great potential for capturing high speed movements. Unlike conventional cameras with a fixed low sampling rate, retina-inspired spike camera can well record fast-moving scenes by continuously accumulating luminance intensity and firing spikes. To restore the captured high-speed motion scenes from spike data, several reconstruction methods have been proposed. A typical method utilizes two neighbouring spikes to infer the instantaneous luminance intensity. Although high temporal resolution imaging can be achieved, the signal to noise ratio (SNR) of reconstructions is generally unsatisfactory. For improving the SNR, some methods propose to average the spikes in big time window. However, the reconstructions may suffer from undesired motion blur, especially when there are objects moving very fast in scenes. To address this issue, we develop a new image reconstruction approach for potential retina-inspired spike camera to recover high-speed motion scenes. Specially, we take the motion of objects into consideration and exploit optical flow to align the scenes of different moments. After motion alignment, a filtering along motion trajectory can be employed to the signals to take the advantage of temporal correlations while not introducing undesired motion blur. Experimental results demonstrate that our proposed method achieves better visual quality than previous reconstruction schemes. Jing Zhao 0011, Ruiqin Xiong, Tiejun Huang 0001 |
ISCAS | 2 |
| 2020 | Multi-class Skin Lesion Segmentation for Cutaneous T-cell Lymphomas on High-Resolution Clinical Images
Zihao Liu 0009, Haihao Pan, Zejia Fan, Yujie Wen, Tingting Jiang 0001, Ruiqin Xiong, Yang Wang 0048 |
MICCAI (6) | 7 |
| 2020 | Clinical-Inspired Network for Skin Lesion Recognition
Zihao Liu 0009, Ruiqin Xiong, Tingting Jiang 0001 |
MICCAI (6) | 2 |
| 2020 | HDR Image Compression with Convolutional AutoencoderabstractAs one of the next-generation multimedia technology, high dynamic range (HDR) imaging technology has been widely applied. Due to its wider color range, HDR image brings greater compression and storage burden compared with traditional LDR image. To solve this problem, in this paper, a two-layer HDR image compression framework based on convolutional neural networks is proposed. The framework is composed of a base layer which provides backward compatibility with the standard JPEG, and an extension layer based on a convolutional variational autoencoder neural networks and a post-processing module. The autoencoder mainly includes a nonlinear transform encoder, a binarized quantizer and a nonlinear transform decoder. Compared with traditional codecs, the proposed CNN autoencoder is more flexible and can retain more image semantic information, which will improve the quality of decoded HDR image. Moreover, to reduce the compression artifacts and noise of reconstructed HDR image, a post-processing method based on group convolutional neural networks is designed. Experimental results show that our method outperforms JPEG XT profile A, B, C and other methods in terms of HDR-VDP-2 evaluation metric. Meanwhile, our scheme also provides backward compatibility with the standard JPEG. Jin Wang 0023, Ruiqin Xiong, Qing Zhu 0004 |
VCIP | 3 |
| 2020 | Motion Estimation for Spike Camera Data Sequence via Spike Interval AnalysisabstractWith the development of emerging computer vision applications, there is an increasing demand for capturing the scenes with high-speed motion. Recently, a novel retina-inspired spike camera has shown great potential for recording the dynamic scenes at high temporal resolution. Different from the conventional digital cameras that capture the visual scene by a single snapshot, the spike camera monitors the incoming light persistently, with each pixel producing a continuous stream of spikes. Recovering the motion process from the spike data sequence is an important problem to study for the spike camera, as it is the foundation of many other tasks, such as image reconstruction, object tracking and object detection. In this paper, we carefully analyze the characteristics of spike data and develop a motion estimation algorithm to recover the continuous high-speed motion process from the spike camera data sequence. Based on the assumption that the spike intervals passed by the same motion trajectories usually have the similar spike densities, we establish a data term constraint to model the temporal consistency of spike intervals. In addition, we integrate a local smoothness constraint with the proposed data term constraint to further improve the estimation accuracy. Experimental results demonstrate that our proposed algorithm can recover high-speed motion process from the captured spike data, and the recovered motion information is beneficial for i mage reconstruction. Jing Zhao 0011, Ruiqin Xiong, Rui Zhao 0010, Jin Wang 0023, Siwei Ma 0001, Tiejun Huang 0001 |
VCIP | 2 |
| 2020 | Optical Flow Estimation Between Images of Different Resolutions via Variational MethodabstractTraditional optical flow estimation methods mostly focus on images of the same resolution. However, there are some situations requiring optical flow between images of different resolutions, where the traditional approaches suffer from the inequality of spectrum aliasing level. In this paper, we propose a method estimating the flow fields between a clear image and a highly undersampled one. The proposed method simultaneously describes the motion and integral relationship between the images via an integral form image under the assumption of brightness and gradient consistency as well as motion smoothness. We also derive the numerical solution briefly, through which we can solve the equations easily via linearizations. Experimental results on Middlebury and MPI-Sintel datasets demonstrate that our proposed method outperforms traditional methods preprocessing images of different resolutions to be the same size, offering more accurate results. Rui Zhao 0010, Ruiqin Xiong, Shuyuan Zhu, Bing Zeng 0001, Tiejun Huang 0001, Wen Gao 0001 |
VCIP | 2 |
| 2020 | Graph-Based Non-Convex Low-Rank Regularization for Image Compression Artifact ReductionabstractBlock transform coded images usually suffer from annoying artifacts at low bit-rates, because of the independent quantization of DCT coefficients. Image prior models play an important role in compressed image reconstruction. Natural image patches in a small neighborhood of the high-dimensional image space usually exhibit an underlying sub-manifold structure. To model the distribution of signal, we extract sub-manifold structure as prior knowledge. We utilize graph Laplacian regularization to characterize the sub-manifold structure at patch level. And similar patches are exploited as samples to estimate distribution of a particular patch. Instead of using Euclidean distance as similarity metric, we propose to use graph-domain distance to measure the patch similarity. Then we perform low-rank regularization on the similar-patch group, and incorporate a non-convex lp penalty to surrogate matrix rank. Finally, an alternatively minimizing strategy is employed to solve the non-convex problem. Experimental results show that our proposed method is capable of achieving more accurate reconstruction than the state-of-the-art methods in both objective and perceptual qualities. Ruiqin Xiong, Xiaopeng Fan 0001, Dong Liu 0002, Feng Wu 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Compressed Image Restoration via Artifacts-Free PCA Basis Learning and Adaptive Sparse ModelingabstractVisually unpleasant compression artifacts frequently appear in block-based transform coding, especially at low bit rates. This paper presents a new artifact reduction scheme based on Bayesian sparse modeling and artifacts-free PCA basis learning. To avoid the effect of blocking artifacts, we propose to learn artifacts-free PCA basis from clean images. We concatenate the clean patches and their compressed counterparts to learn paired distribution prior via the Gaussian Mixture Model (GMM). By this way, the GMM characterizes the mapping between the clean image and its compressed version. To restore a compressed patch, the best matched GMM component is assigned using the patch in the compressed image subspace. The artifacts-free PCA basis is obtained according to the mapping learned by the paired GMM. In practice, the statistical distributions of different sparse coefficients in different patches may dramatically vary with image contents. Instead of using a global zero-mean distribution for all coefficients, we propose to adaptively model the prior of each band in a Bayesian framework. The expectation and variance of each band are adaptively learned from the similar patches within the image. Thus, different transform bands are regularized unequally according to the learned priors. Experimental results show that the proposed scheme outperforms most of the compared schemes in terms of both objective quality and perceptual quality. Ruiqin Xiong, Xiaopeng Fan 0001, Dong Liu 0002, Feng Wu 0001, Tiejun Huang 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Visual Relationship Embedding Network for Image Paragraph GenerationabstractImage paragraph generation aims to produce a complete description of a given image. This task is more challenging than image captioning, which only generates one sentence to describe the entire image. Traditional paragraph generation methods usually produce paragraph descriptions based on individual regions that are detected by a Region Proposal Network (RPN). However, relationships among visual objects are either ignored or utilized in an implicit manner in previous work. In this paper, we attempt to explore more visual information through a novel paragraph generation network that explicitly incorporates visual relationship semantics when producing descriptions. First, a novel Relation Pair Generative Adversarial Network (RP-GAN) is designed to locate regions that may cover subjective or objective elements. Then, their relationships are inferred through an attention-based network. Finally, the visual features and relationship semantics of valid relation pairs are taken as inputs by a Long Short-Term Memory (LSTM) network for generating sentences. The experimental results show that by explicitly utilizing the predicted relationship information, our proposed method obtains more accurate and informative paragraph descriptions than previous methods. Wenbin Che, Xiaopeng Fan 0001, Ruiqin Xiong, Debin Zhao |
IEEE Trans. Multim. | 3 |
| 2020 | iWave: CNN-Based Wavelet-Like Transform for Image CompressionabstractWavelet transform is a powerful tool for multiresolution time-frequency analysis. It has been widely adopted in many image processing tasks, such as denoising, enhancement, fusion, and especially compression. Wavelets lead to the successful image coding standard JPEG-2000. Traditionally, wavelets were designed from the signal processing theory with certain assumption on the signal, but natural images are not as ideal as assumed by the theory. How to design content-adaptive wavelets for natural images remains a difficulty. Inspired by the recent progress of convolutional neural network (CNN), we propose iWave as a framework for deriving wavelet-like transform that is more suitable for natural image compression. iWave adopts an update-first lifting scheme, where the prediction filter is a trained CNN, to achieve wavelet-like transform. The CNN can be embedded into a deep network that is analogous to an auto-encoder, which is trained end-to-end. The trained wavelet-like transform still possesses the lifting structure, which ensures perfect reconstruction, supports multiresolution analysis, and is more interpretable than the deep networks trained as “black boxes.” We perform experiments to verify the generality as well as the speciality of iWave in comparison with JPEG-2000. When trained with a generic set of natural images and tested on the Kodak dataset, iWave achieves on average 4.4% and up to 14% BD-rate reductions. When trained and tested with a specific kind of textures, iWave provides as high as 27% BD-rate reduction. Haichuan Ma, Dong Liu 0002, Ruiqin Xiong, Feng Wu 0001 |
IEEE Trans. Multim. | 3 |
| 2019 | Perceptual Video Coding Based on Visual Saliency Modulated Just Noticeable DistortionabstractTo reduce the perceptual redundancy in the video coding process, human visual system (HVS)-based visual attention and visual sensitivity can be utilized due to their intrinsic natures. Just Noticeable Distortion (JND) is one of widely used models to simulate human visual sensitivity, while visual saliency map has been popular for years in image processing to describe the visual attention feature, which has been proved by the ability to enhance the visual sensitivity effect. In this paper, we proposed a perceptual video coding (PVC) scheme with visual saliency modulated JND model to suppress the DCT coefficient without resulting in noteworthy subjective quality degradation. The experimental results show that the PVC scheme with the proposed VS-JND model can save bit rates up to 35.58% in high bit rates case with the similar subjective quality compared with that of HEVC software reference code HM 16.12. Ruiqin Xiong, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001 |
DCC | 2 |
| 2019 | Referenceless Quality Assessment for Contrast Distorted Image Using Hybrid FeaturesabstractContrast distortion has a significant influence on the perceptual quality of an image, which may be generated in various image processing procedures. In this paper, we propose a no-reference image quality assessment (IQA) algorithm for contrast-distorted images based on hybrid features from information and appearance attributes. From information attribute aspect, we utilize the basic information feature to quantify the information of the visible part and the extended information feature further containing the invisible part of an image. From the appearance aspect, we propose an efficient perceptual contrast and colorfulness index to capture the direct visual changes. With the hybrid information attribute and appearance attribute features, support vector regression (SVR) is utilized to learn an IQA model to predict the quality of contrast distorted images. Extensive experimental results on CCID2014 and TID2013 databases further demonstrate the superior performance and robustness of the proposed method. Xinfeng Zhang 0001, Shanshe Wang, Xiaofei Pan, Siwei Ma 0001, Ruiqin Xiong |
ICIP | 6 |
| 2019 | A CNN-Based Image Compression Scheme Compatible with JPEG-2000abstractWe propose a convolutional neural network (CNN) based image compression scheme that is compatible with JPEG-2000. Specifically, our scheme reuses the existing JPEG-2000 encoders to achieve bitstream, and features two components in addition to JPEG-2000: bitstream re-compression and decoder-side post-processing. First, we propose an advanced arithmetic codec that adopts CNN-based probability estimation to exploit the correlation between wavelet coefficients within and across subbands. Second, we propose a CNN-based post-processing method to improve the quality of reconstructed images. Experimental results show that the proposed two CNN-based components both help improve the compression efficiency by a significant margin. Haichuan Ma, Dong Liu 0002, Ruiqin Xiong, Feng Wu 0001 |
ICIP | 3 |
| 2019 | Learning a Deep Convolutional Network for Subband Image DenoisingabstractDue to the fast inference and excellent learning capability, deep learning has become an effective means for image denoising and attracted considerable attention recently. However, for the images with rich textures and structures, the performance of deep learning approaches is still unsatisfactory. To address this issue, we develop a new convolutional neural network (CNN) for subband image denoising and name it SDCNN. In the proposed approach, we first decompose images into transform domain and denoise the coefficients of various subbands. By incorporating frequency information with spatial context, SDCNN is more effective in recovering image details. In particular, the introduced procedure of subband transform also plays the role of downsampling and enlarges the receptive field without increasing depth or sacrificing efficiency of network. Experimental results show that the SDCNN achieves promising results in terms of both objective and subjective performance. Jing Zhao 0011, Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Tiejun Huang 0001 |
ICME | 2 |
| 2019 | Color Image Compression with Transform Domain Down-Sampling and Deep Convolutional ReconstructionabstractIn this paper, we build up a new block-based color image compression scheme based on our proposed transform domain down-sampling method and deep convolutional reconstruction algorithm. Specifically, our proposed down-sampling scheme aims to down-sample each N × N transform block into the N/2 × N/2 block for the saving of bit-cost. On the other hand, the proposed deep convolutional reconstruction algorithm is employed to reconstruct the down-sampled block for a full- resolution reconstruction. We apply our proposed methods to both the chrominance components to compress color images. Experimental results show that our proposed method achieves excellent results when used in practice. Shuyuan Zhu, Xiandong Meng, Bing Zeng 0001, Yuanfang Guo, Ruiqin Xiong |
VCIP | 6 |
| 2019 | CG-Cast: Scalable Wireless Image SoftCast Using Compressive GradientabstractG-Cast is a wireless visual communication scheme that conveys visual information via image gradient. It is inspired by the characteristics of human vision systems and can provide improved perceptual quality. G-Cast is power efficient but bandwidth demanding, because gradient data have double the size of the original image. This paper presents a scheme named CG-Cast for scalable image transmission in bandwidth-limited wireless scenarios. It employs a compressive-gradient-based image representation to describe perceptually sensitive image details and reduce the bandwidth requirement at the same time, combining gradient-based visual representation with compressive sensing techniques. The compressive gradient data are transmitted in a pseudo-analog way so that it achieves elegant quality transition in a wide channel signal-to-noise ratio (CSNR) range. CG-Cast also sends a small set of low-frequency data in digital a way to provide the global and local luminance of the image. We developed an effective optimization algorithm for the decoder to reconstruct the original image from the received noisy compressive gradient and the low-frequency part of the image. Experimental results demonstrate that the proposed scheme improves the quality of received images remarkably under different CSNR and channel bandwidth conditions. Hangfan Liu, Ruiqin Xiong, Xiaopeng Fan 0001, Debin Zhao, Yongbing Zhang 0002, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Efficient Chroma Sub-Sampling and Luma Modification for Color Image CompressionabstractIn color image compression, the chroma components are often sub-sampled before compression and up-sampled after compression. Although sub-sampling the chroma components saves the bit-cost for compression, it often induces extra color distortions in the compressed images. In this paper, we propose two approaches to tackle this problem. First, we propose a sub-sampling method in the transform domain and apply it to both chroma components. Then, based on this sub-sampling, we propose a novel method to modify the luma component. In our proposed luma modification algorithm, the distortions that occurred in the two chroma components can be coupled together and utilized to modify the luma component. With our proposed chroma sub-sampling and luma modification algorithms, we can achieve a low RGB distortion in practical image coding. The experimental results demonstrate that our proposed methods offer more significant coding gains compared with the state-of-the-art methods for the compression of color images. Shuyuan Zhu, Chang Cui, Ruiqin Xiong, Yuanfang Guo, Bing Zeng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Intra Block Copy for Screen Content in the Emerging AV1 Video CodecabstractScreen content coding plays an important role in many applications. To meet the growing demands of screen content coding, the emerging AV1 video codec incorporates several coding tools, which are specially designed for screen content utilizing its distinctive characteristics. Among these tools, the intra block copy utilizes the characteristic that repeating patterns frequently occur in screen content. This paper presents the technology of intra block copy in AV1. In particular, to efficiently search the predictor in the reconstructed regions of the current picture, AV1 uses the hash matching method at the encoder side. For the generation of hash table, a bottom-to-up manner is adopted to reduce the redundant computation and then decrease the encoding time. In addition, several constraints are involved to facilitate hardware design. Experimental results demonstrate that the intra block copy in AV1 can bring 27.1% bitrate saving for screen content. When compared with the non hash-based intra block copy, the hash-based method achieves 12.2% bitrate saving. Jiahao Li 0001, Hui Su, Alex Converse, Bin Li 0012, Roger Zhou, Bruce Lin, Jizheng Xu, Yan Lu 0001, Ruiqin Xiong |
DCC | 9 |
| 2018 | Compressed Image Restoration via External-Image Assisted Band Adaptive PCA Model LearningabstractVisually annoying compression artifacts frequently appear in block-based transform coding at low bit rates, due to coarse and independent quantization of transform coefficients in coding blocks. This paper presents a subband adaptive modeling framework for reducing quantization artifacts. In this framework, each patch is jointly regularized by bandwise distribution priors adaptively learned in its PCA transform domain together with a quantization constraint prior in the DCT domain. Since the compression artifacts influence the covariance statistics of coded image patches remarkably, external images are utilized to provide more robust PCA domains for patch sparse modeling. Instead of using a global distribution model for all patches, the distribution prior of each patch is adaptively learned from similar patches within the compressed image itself to address the non-stationarity of image signals. The coefficients in different PCA bands are regularized unequally according to the learned priors. Experimental results show that the proposed scheme outperforms existing schemes in terms of both the objective and the perceptual qualities. Ruiqin Xiong, Xiaopeng Fan 0001, Xianming Liu 0005, Tiejun Huang 0001, Wen Gao 0001 |
DCC | 2 |
| 2018 | Fast and Robust Image Upsampling by Local Adaptive Gradient Field Sharpening TransformabstractThis paper proposes an image upsampling scheme by introducing a new gradient field sharpening transform that converts the blurry gradient field of upsampled low-resolution (LR) image to a much sharper gradient field of original high-resolution (HR) image. Different from the existing methods that need to figure out the whole gradient profile structure and locate the edge points, we derive a new approach that sharpens the gradient field adaptively only based on the pixels in a small neighborhood. To maintain image contrast, image gradient is adaptively scaled to keep the integral of gradient field stable. Finally the HR image is reconstructed by fusing the LR image with the sharpened HR gradient field. Experimental results demonstrate that the proposed algorithm can generate more accurate gradient field and produce super-resolved images with better objective and visual qualities. Another advantage is that the proposed gradient sharpening transform is very fast and suitable for low-complexity applications. Ruiqin Xiong, Dong Liu 0002, Zhiwei Xiong, Feng Wu 0001, Wen Gao 0001 |
DCC | 2 |
| 2018 | Residual Signals Modeling for Layered Image/Video Softcast with Hybrid Digital-Analog TransmissionabstractRecently, the SoftCast scheme has shown great potential for robust image/video communication in wireless scenarios, where the channel quality may fluctuate drastically and unpredictably. However, the analog-like transmission in Soft-Cast is not always efficient in terms of power usage, compared with digital approaches. In this paper, we propose a layered image/video SoftCast scheme, in which a coarse version of the image is transmitted by a base layer in digital way, while the residual details are delivered by an enhancement layer in pseudo-analog way. In order to achieve optimal overall transmission performance, this paper studies the problem of optimal bit rate selection for the base layer, by building a rate-residual model based on the relationship between the base layer bit rate and the residual signal spectrum. Experimental results show that the proposed scheme can improve the performance of original scheme remarkably, while still preserving the smooth quality degradation characteristic of SoftCast. Jing Zhao 0011, Jiyu Xie, Ruiqin Xiong |
ICIP | 3 |
| 2018 | Paragraph Generation Network with Visual Relationship DetectionabstractParagraph generation of images is a new concept, aiming to produce multiple sentences to describe a given image. In this paper, we propose a paragraph generation network with introducing visual relationship detection. We first detect regions which may contain important visual objects and then predict their relationships. Paragraphs are produced based on object regions which have valid relationship with others. Compared with previous works which generate sentences based on region features, we explicitly explore and utilize visual relationships in order to improve final captions. The experimental results show that such strategy could improve paragraph generating performance from two aspects: more details about object relations are detected and more accurate sentences are obtained. Furthermore, our model is more robust to region detection fluctuation. Wenbin Che, Xiaopeng Fan 0001, Ruiqin Xiong, Debin Zhao |
ACM Multimedia | 3 |
| 2018 | Frequency-Domain Dynamic Pruning for Convolutional Neural NetworksabstractDeep convolutional neural networks have demonstrated their powerfulness in a variety of applications. However, the storage and computational requirements have largely restricted their further extensions on mobile devices. Recently, pruning of unimportant parameters has been used for both network compression and acceleration. Considering that there are spatial redundancy within most filters in a CNN, we propose a frequency-domain dynamic pruning scheme to exploit the spatial correlations. The frequency-domain coefficients are pruned dynamically in each iteration and different frequency bands are pruned discriminatively, given their different importance on accuracy. Experimental results demonstrate that the proposed scheme can outperform previous spatial-domain counterparts by a large margin. Specifically, it can achieve a compression ratio of 8.4x and a theoretical inference speed-up of 9.2x for ResNet-110, while the accuracy is even better than the reference model on CIFAR-110. Zhenhua Liu 0003, Jizheng Xu, Xiulian Peng, Ruiqin Xiong |
NeurIPS | 4 |
| 2018 | Efficient Multiple-Line-Based Intra Prediction for HEVCabstractTraditional intra prediction usually utilizes the nearest reference line to generate the predicted block when considering strong spatial correlation. However, this kind of single-line-based method does not always work well due to at least two issues. One is the incoherence caused by the signal noise or the texture of other objects, where this texture deviates from the inherent texture of the current block. The other reason is that the nearest reference line usually has worse reconstruction quality in block-based video coding. Due to these two issues, this paper proposes an efficient multiple-line-based intra-prediction scheme to improve coding efficiency. Besides the nearest reference line, further reference lines are also utilized. The further reference lines with a relatively higher quality can provide potentially better prediction. At the same time, the residue compensation is introduced to calibrate the prediction of boundary regions in a block when we utilize further reference lines. To speed up the encoding process, this paper designs several fast algorithms. The experimental results show that compared with HM-16.9, the proposed fast search method achieves a 2.0% bit saving on average and up to 3.7% by increasing the encoding time by 112%. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Diversity-Based Reference Picture Management for Low Delay Screen Content CodingabstractScreen content coding plays an important role in many applications. Conventional reference picture management (RPM) strategies developed for natural content may not work well for screen content. This is because many regions in screen content remain static for a long time, causing a lot of repetitive contents to stay in the decoded picture buffer. The repetitive contents are not conducive to inter prediction, but still occupy valuable memory. This paper proposes a diversity-based RPM scheme for screen content coding. The concept of diversity is introduced for the reference picture set (RPS) to help formulate the RPM problem. By maximizing the diversity of RPS, more potentially better predictions are provided. Better compression performance can then be achieved. Meanwhile, the proposed scheme is nonnormative and compatible with existing video coding standards, such as High Efficiency Video Coding. The experimental results show that, for low delay screen content coding, the bit saving of the proposed scheme is 4.9% on average and up to 13.7%, without increasing encoding time. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Superimposed Modulation for Soft Video Delivery With Hidden ResourcesabstractAnalog-transmission-based soft video delivery suffers from the leveling-off effect when the allocated channel bandwidth is severely insufficient. Fortunately, with superimposed modulation, it is possible for analog traffic to share bandwidth with digital traffic. In this paper, we design and analyze such a hybrid digital-analog superimposed modulation (HDA-SIM) scheme for soft video delivery. Unlike previous work, we treat the bandwidth of competing digital traffic as hidden resources for the video delivery system. The key problem in this scheme is how to allocate the bandwidth and power resources among various modulation symbols so that we can improve the performance of video delivery without sacrificing the throughput of existing digital traffic. The resource allocation problem is formulated and the optimal solution under any given channel signal-to-noise ratio is derived. Based on the results, the sufficient and necessary condition for the video delivery system to achieve performance gain is given. In addition, we implement the proposed scheme for two state-of-the-art soft video delivery systems known as SoftCast and SharpCast. Both simulations and testbed evaluations show that the HDA-SIM version can achieve significant gains in the received video quality over their original designs. Chong Luo 0001, Ruiqin Xiong, Wenjun Zeng 0001, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Image Denoising via Low Rank Regularization Exploiting Intra and Inter Patch CorrelationabstractIn image restoration tasks, image priors generally utilize correlation within image contents to predict the latent image signal. In this paper, we propose to jointly exploit both intra- and inter-patch correlation of the input image, so as to further reduce the uncertainty of the unknown signal, and thus improve the prediction of the latent image. The proposed scheme evolves from the low-rank regularization for non-local highly-correlated image contents. Since the underlying cost function to pursue minimal rank is hard to solve, we use non-convex smooth surrogates for the rank penalty. Two such surrogates are utilized in order to incorporate both intra- and inter-patch correlation. To tackle the optimization problem, we use iterative alternating direction technique to divide the problem into two subproblems, each of which is solved via an empirical Bayesian procedure built upon variational approximation. Experimental results on image denoising show that the proposed approach outperforms several state-of-the-art methods in terms of peak signal-to-noise ratio, structural similarity, and perceptual quality. Hangfan Liu, Ruiqin Xiong, Dong Liu 0002, Siwei Ma 0001, Feng Wu 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Hybrid All Zero Soft Quantized Block Detection for HEVCabstractTransform and quantization account for a considerable amount of computation time in video encoding process. However, there are a large number of discrete cosine transform coefficients which are finally quantized into zeros. In essence, blocks with all zero quantized coefficients do not transmit any information, but still occupy substantial unnecessary computational resources. As such, detecting all-zero block (AZB) before transform and quantization has been recognized to be an efficient approach to speed up the encoding process. Instead of considering the hard-decision quantization (HDQ) only, in this paper, we incorporate the properties of soft-decision quantization into the AZB detection. In particular, we categorize the AZB blocks into genuine AZBs (G-AZB) and pseudo AZBs (P-AZBs) to distinguish their origins. For G-AZBs directly generated from HDQ, the sum of absolute transformed difference-based approach is adopted for early termination. Regarding the classification of P-AZBs which are generated in the sense of rate-distortion optimization, the rate-distortion models established based on transform coefficients together with the adaptive searching of the maximum transform coefficient are jointly employed for the discrimination. Experimental results show that our algorithm can achieve up to 24.16% transform and quantization time-savings with less than 0.06% RD performance loss. The total encoder time saving is about 5.18% on average with the maximum value up to 9.12%. Moreover, the detection accuracy of larger TU sizes, such as and can reach to 95% on average. Ruiqin Xiong, Xinfeng Zhang 0001, Shiqi Wang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Fully Connected Network-Based Intra Prediction for Image CodingabstractThis paper proposes a deep learning method for intra prediction. Different from traditional methods utilizing some fixed rules, we propose using a fully connected network to learn an end-to-end mapping from neighboring reconstructed pixels to the current block. In the proposed method, the network is fed by multiple reference lines. Compared with traditional single line-based methods, more contextual information of the current block is utilized. For this reason, the proposed network has the potential to generate better prediction. In addition, the proposed network has good generalization ability on different bitrate settings. The model trained from a specified bitrate setting also works well on other bitrate settings. Experimental results demonstrate the effectiveness of the proposed method. When compared with high efficiency video coding reference software HM-16.9, our network can achieve an average of 3.4% bitrate saving. In particular, the average result of 4K sequences is 4.5% bitrate saving, where the maximum one is 7.4%. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong, Wen Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Fast Image Super-Resolution via Local Adaptive Gradient Field Sharpening TransformabstractThis paper proposes a single-image super-resolution scheme by introducing a gradient field sharpening transform that converts the blurry gradient field of upsampled low-resolution (LR) image to a much sharper gradient field of original high-resolution (HR) image. Different from the existing methods that need to figure out the whole gradient profile structure and locate the edge points, we derive a new approach that sharpens the gradient field adaptively only based on the pixels in a small neighborhood. To maintain image contrast, image gradient is adaptively scaled to keep the integral of gradient field stable. Finally, the HR image is reconstructed by fusing the LR image with the sharpened HR gradient field. Experimental results demonstrate that the proposed algorithm can generate more accurate gradient field and produce super-resolved images with better objective and visual qualities. Another advantage is that the proposed gradient sharpening transform is very fast and suitable for low-complexity applications. Ruiqin Xiong, Dong Liu 0002, Zhiwei Xiong, Feng Wu 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2018 | Residual Highway Convolutional Neural Networks for in-loop Filtering in HEVCabstractHigh efficiency video coding (HEVC) standard achieves half bit-rate reduction while keeping the same quality compared with AVC. However, it still cannot satisfy the demand of higher quality in real applications, especially at low bit rates. To further improve the quality of reconstructed frame while reducing the bitrates, a residual highway convolutional neural network (RHCNN) is proposed in this paper for in-loop filtering in HEVC. The RHCNN is composed of several residual highway units and convolutional layers. In the highway units, there are some paths that could allow unimpeded information across several layers. Moreover, there also exists one identity skip connection (shortcut) from the beginning to the end, which is followed by one small convolutional layer. Without conflicting with deblocking filter (DF) and sample adaptive offset (SAO) filter in HEVC, RHCNN is employed as a high-dimension filter following DF and SAO to enhance the quality of reconstructed frames. To facilitate the real application, we apply the proposed method to I frame, P frame, and B frame, respectively. For obtaining better performance, the entire quantization parameter (QP) range is divided into several QP bands, where a dedicated RHCNN is trained for each QP band. Furthermore, we adopt a progressive training scheme for the RHCNN where the QP band with lower value is used for early training and their weights are used as initial weights for QP band of higher values in a progressive manner. Experimental results demonstrate that the proposed method is able to not only raise the PSNR of reconstructed frame but also prominently reduce the bit-rate compared with HEVC reference software. Yongbing Zhang 0002, Xiangyang Ji, Yun Zhang 0002, Ruiqin Xiong, Qionghai Dai |
IEEE Trans. Image Process. | 5 |
| 2018 | Hybrid Digital-Analog Video Delivery With Shannon-Kotel'nikov MappingabstractHybrid digital-analog (HDA) transmission is becoming an attractive solution for mobile video delivery because it not only has graceful degradation with channel variations but also yields high power efficiency. However, the heavy bandwidth demand of analog transmission is still an unsolved problem, limiting the received video quality when the bandwidth is not sufficient. To address this problem, we propose adopting Shannon-Kotel'nikov (SK) mapping for HDA video transmission and design an HDA scheme called SK-Cast. SK-Cast consists of a digital and an analog branch. In the digital branch, SK-Cast compresses the video sequence using an high efficiency video coding digital encoder to produce a base layer. The base layer is transmitted through digital methods with strong protection. The residual signals are then decorrelated using three-dimensional discrete cosine transform transform. The SK mapping is exploited to transmit these coefficients, as they can achieve efficient bandwidth compression. We address the resource allocation problems in SK-Cast, including the allocation between digital and analog branches and the allocation among analog symbols. The simulation results show that the SK-Cast outperforms the state-of-the-art HDA systems, including WSVC and SharpCast, and a digital scalable video coding system. Chong Luo 0001, Ruiqin Xiong, Wenjun Zeng 0001, Feng Wu 0001 |
IEEE Trans. Multim. | 3 |
| 2018 | Hybrid Intraprediction Based on Local and Nonlocal CorrelationsabstractIn the latest video coding standard, namely, high-efficiency video coding (HEVC), intra coding efficiency is significantly improved by a quadtree partition structure and more intra prediction modes. For intra coding, 35 intra modes are employed including 33 angular intra prediction (AIP) modes that are effective for blocks with strong directions, and a planar mode and a dc mode that are used to predict smooth regions. However, intra prediction in HEVC still cannot handle complicated blocks well. To deal with this problem, this paper proposes a hybrid intra prediction method to improve the intra prediction efficiency. The proposed hybrid intra prediction method consists of three parts: adaptive template matching prediction (ATMP) by exploring nonlocal correlation, combined local and nonlocal prediction by exploring both local correlation for AIP and non-local correlation for ATMP, and combined neighboring modes prediction that can generate smooth prediction by exploring more local correlations. Experimental results suggest that the proposed hybrid intra prediction method achieves 2.8% BD-rate reduction on average for luma compared to HEVC reference software HM-14.0 under all intra main configurations. The gain can be up to 6.9%. When integrated into joint exploration model -1.0, the proposed method still can achieve 1.2% BD-rate reduction for luma. Tao Zhang 0013, Xiaopeng Fan 0001, Debin Zhao, Ruiqin Xiong, Wen Gao 0001 |
IEEE Trans. Multim. | 4 |
| 2017 | Intra Prediction Using Multiple Reference Lines for Video CodingabstractTraditional intra prediction schemes usually only use the nearest adjacent reference line to generate the prediction. Although the nearest reference line generally has the strongest statistical correlation with current block, the farther non-adjacent reference lines can still provide potential better prediction in some cases. Thus, in this paper, not only the nearest reference line but also the farther reference lines are utilized to help intra prediction. When using the farther reference lines, an additional residue compensation procedure is introduced to further refine the prediction. In particular, this paper designs three solutions to meet different complexity requirements. They are multiple line-based intra prediction (MLIP), fast search for multiple line-based intra prediction (FS-MLIP), and dual line-based intra prediction (DLIP). Experimental results verify the effectiveness of the proposed methods. When compared with HM-16.9, the proposed MLIP achieves 2.4% bit saving on average with the encoding time increasing about 362%. The FS-MLIP achieves 2.0% bit saving on average with the encoding time increasing about 114%. The DLIP achieves 0.9% bit saving on average with the encoding time increasing about only 15%. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
DCC | 4 |
| 2017 | Wireless Image SoftCast Using Compressive GradientabstractSummary form only given: Based on observations that the visual quality has strong correlation with image gradients, gradient based image SoftCast (G-Cast) [1] advocates to convey visual information by delivering image gradients. In G-Cast, both horizontal gradients and vertical gradients needs to be transmitted, even if the channel bandwidth is insufficient. This paper propose to send out the random projection measurements of the gradients instead of delivering gradients directly, so that data size can be reduced to an arbitrary ratio and channel bandwidth occupation can be lowered. We name this scheme as compressive gradient based SoftCast (CG-Cast). At CG-Cast sender, after generated by gradient transform, the gradients are down-sampled by random projection, sample rate of which is set according to the channel bandwidth condition. Then the produced measurements are sent out for raw OFDM transmission. A few lowfrequency components are also transmitted to tell the global luminance. At CG-Cast receiver, the received noisy measurements are used for compressive gradient based reconstruction procedure, which utilizes sparsity in gradient domain and non-local similarity in spatial domain [2]. The proposed method is compared with SoftCast [3] and compressive sensing (CS) in bandwidth limited and power constrained scenarios. To make fair comparison, these three schemes are tested under the same channel signal-to-noise ratio (CSNR) conditions to transmit equal amount of data for reconstruction, using equivalent power and bandwidth. CG-Cast outperforms SoftCast and CS in terms of SSIM and gradient signal-to-noise ratio (GSNR) at different bandwidth ratios. Comparing with SoftCast in different channel conditions, the average SSIM gain of all the tested images varies from 0.04 to 0.13, and the average GSNR gain ranges from 1.5dB to 2.9dB. CS is rather unstable in noisy conditions. SoftCast performs better than CS because of its power allocation. Hangfan Liu, Ruiqin Xiong, Xiaopeng Fan 0001, Siwei Ma 0001, Wen Gao 0001 |
DCC | 2 |
| 2017 | Band-Wise Adaptive Sparsity Regularization for Quantized Compressed Sensing Exploiting Nonlocal SimilarityabstractThe theory of compressive sensing (CS) has attracted considerable research interests from signal and image processing communities. And in practice, because of the considerations of data storage and transmission, scalar quantization is necessary to be implemented on the CS measurements. In this paper, we propose an adaptive bandwise sparsity regularization to handle the recovery problem of quantized compressive sensing. The sparsity regularization constraints every patch by using bandwise distribution model in transform domain. In addition, we bring in the quantization cost function to quantify the influence of measurement quantization. Experimental results demonstrate that our CS recovery strategy achieves significant performance improvements over the current state-of-the-art schemes with both unquantized measurements and quantized measurements. Ruiqin Xiong, Xinfeng Zhang 0001, Siwei Ma 0001 |
DCC | 2 |
| 2017 | Intra prediction using fully connected network for video codingabstractTraditional intra prediction methods exploit some fixed rules to generate prediction, which might not be adaptive enough to handle complicated contents. In this paper, we investigate applying deep neural network to improve the state-of-the-art intra prediction. Considering the characteristics of block-based video coding framework, we propose a fully connected network for intra prediction where all layers except non-linear ones are fully connected. In the proposed network, the inputs are multiple reference lines of the current block and the output is the prediction for the block. When compared with the traditional intra prediction method, the richer context of current block is exploited. For this reason, the proposed network is capable of providing more accurate prediction. Experimental results demonstrate the effectiveness of proposed network. When integrated into the HEVC reference software, the proposed method can achieve up to 3.3% bitrate saving and an average of 1.6% bitrate saving for 4K sequences. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
ICIP | 4 |
| 2017 | Image restoration via multi-scale non-local total variation regularizationabstractTotal-variation (TV) regularization is widely adopted in image restoration problems to exploit the local smoothness of image. However, traditional TV regularization only models the sparsity of image gradient at the original scale. This paper introduces a multi-scale TV regularization method which models the image gradient sparsity at different scales, and constrains the gradient magnitude of different scales jointly. As different scales extract different frequency of image, our proposed multi-scale regularization method provides constraints for different frequency components. And for each scale, we adaptively estimate the gradient distribution at a particular pixel from a group of nonlocally searched similar patches. Finally, experimental results demonstrate that the proposed method outperforms the conventional TV regularization methods for image restoration. Ruiqin Xiong, Xiaopeng Fan 0001, Siwei Ma 0001 |
ICME | 2 |
| 2017 | An adaptive and low-complexity all-zero block detection for HEVC encoderabstractTo improve the detection accuracy of SAD or SATD based threshold and save the time cost of RDO determination, we proposed an All Zero Block (AZB) detection method by adaptively searching the maximum transform coefficient amplitude in low frequency of TU after conventional SATD detection was failed. The experimental results show that our algorithm can achieve around 39% transform and quantization time-saving with only 0.1% on average RD performance reduction. The detection accuracy of larger TU size, i.e. 16×16 and 32 × 32, can reach up to about 95% on average. Ruiqin Xiong, Falei Luo, Shanshe Wang, Siwei Ma 0001 |
ISCAS | 2 |
| 2017 | Compressive gradient based scalable image SoftCastabstractIn wireless visual communication systems, it is crucial to effectively utilize channel power and bandwidth in the pursue of optimal performance, and it is worthwhile to adapt the transmission scheme to human vision system (HVS) so as to achieve perceptually appealing results. Inspired by observations that visual quality of an image is closely related to the gradient data, this paper proposes to convey visual information by random projection measurements of image gradients in an analog framework. Since HVS is more sensitive to luminance variations of image contents, which are contained in the gradient data, the proposed scheme achieves better perceptual quality than conventional analog uncoded schemes like SoftCast. Besides, the gradient transform removes the low and medium frequency components of the image hence substantially reduces the power of the signal transmitted in the analog channel, thus evidently improves the power-distortion performance of the system. Furthermore, by applying random projection to the gradients, the number of transmitted data can be adjusted according to bandwidth conditions. Another contribution of this paper is to develop an effective optimization scheme for the compressive gradient based reconstruction problem. Experimental results validate the effectiveness of the proposed transmission and reconstruction scheme under different channel signal-to-noise ratio and bandwidth conditions. Hangfan Liu, Ruiqin Xiong, Xiaopeng Fan 0001, Chong Luo 0001, Wen Gao 0001 |
VCIP | 2 |
| 2017 | Low rank regularization exploiting intra and inter patch correlation for image denoisingabstractBased on the observation that a matrix X consisted of non-local highly-correlated patches is of low rank, many image restoration methods use low-rank regularization to exploit correlation between image contents, so that the uncertainty of the unknown image signal can be reduced. To tackle the problem that the underlying cost function to pursue minimal rank is hard to solve, an effective way is to employ smooth non-convex surrogate log |XXT| for the rank penalty. Essentially, such technique only considers to utilize correlation within image patches. In this paper, we propose to jointly exploit both intra- and inter-patch correlation of the input image, so as to further reduce the uncertainty of the signal, and thus improve the prediction of the latent image. The corresponding two surrogates are integrated to incorporate both intra- and inter-patch correlation. To solve the optimization problem, we use iterative alternating direction technique to divide the problem into two subproblems, each of which is solved via an empirical Bayesian procedure built upon variational approximation. Experimental results show that the proposed approach outperforms several state-of-the-art methods in terms of PSNR and perceptual quality. Hangfan Liu, Ruiqin Xiong, Dong Liu 0002, Feng Wu 0001, Wen Gao 0001 |
VCIP | 2 |
| 2017 | Image super-resolution based on adaptive joint distribution modelingabstractThis paper combines an adaptive reconstruction based approach and a learning based technique into an effective scheme for single image super-resolution. Unlike conventional schemes that adopt pre-trained dictionaries to tell the relationship between high-resolution (HR) image and the low-resolution (LR) observation, the proposed method attempts to learn the joint distribution of highly-correlated patch couples from the input image itself instead of an external dataset, so that the learnt models are specially tailored for the current patches and thus can better fit the image data of interest. To be specific, we first apply spatially adaptive gradient sparsity regularization in the reconstruction of the HR image using the contour information, and then utilize the generated HR output to guide the joint distribution learning to infer the relationship between the highly-correlated HR and LR patches. In this way, we simultaneously exploit the inter-scale correlation as well as the local and non-local correlation of the image contents. Empirical results show that the performance of the proposed method is highly competitive with state-of-the-art schemes in terms of peak signal-to-noise ratio (PSNR) and perceptual quality. Hangfan Liu, Ruiqin Xiong, Feng Wu 0001, Wen Gao 0001 |
VCIP | 2 |
| 2017 | Wireless image and video soft transmission via perception-inspired power distortion optimizationabstractRecently, a scheme called SoftCast has shown great potential for wireless image/video communication in the scenarios where the channel quality may fluctuate drastically and unpredictably. The transmission is lossy in nature, with its transmission power allocated among coefficients unequally to minimize the distortion. One problem is that its performance is optimized using mean square errors (MSE) as the quality metric, which is known for not matching the perception of human eyes. Inspired by the researches in image quality assessment, this paper proposes a power allocation and optimization scheme that minimizes the perceptual distortion of reconstruction image. In particular, we establish a perception model to evaluate the perceptual importance of different transform coefficients, based on the structure similarity (SSIM) image quality metric. Experimental results show that the proposed scheme can improve the perceptual performance of the original SoftCast scheme. Jing Zhao 0011, Ruiqin Xiong, Chong Luo 0001, Feng Wu 0001, Wen Gao 0001 |
VCIP | 2 |
| 2017 | Nonlocal Gradient Sparsity Regularization for Image RestorationabstractTotal variation (TV) regularization is widely used in image restoration to exploit the local smoothness of image content. Essentially, the TV model assumes a zero-mean Laplacian distribution for the gradient at all pixels. However, real-world images are nonstationary in general, and the zero-mean assumption of pixel gradient might be invalid, especially for regions with strong edges or rich textures. This paper introduces a nonlocal (NL) extension of TV regularization, which models the sparsity of the image gradient with pixelwise content-adaptive distributions, reflecting the nonstationary nature of image statistics. Taking advantage of the NL similarity of natural images, the proposed approach estimates the image gradient statistics at a particular pixel from a group of nonlocally searched patches, which are similar to the patch located at the current pixel. The gradient data in these NL similar patches are regarded as the samples of the gradient distribution to be learned. In this way, more accurate estimation of gradient is achieved. Experimental results demonstrate that the proposed method outperforms the conventional TV and several other anchors remarkably and produces better objective and subjective image qualities. Hangfan Liu, Ruiqin Xiong, Xinfeng Zhang 0001, Yongbing Zhang 0002, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Low-Rank-Based Nonlocal Adaptive Loop Filter for High-Efficiency Video CompressionabstractIn video coding, the in-loop filtering has emerged as a key module due to its significant improvement on compression performance since H.264/Advanced Video Coding. Existing incorporated in-loop filters in video coding standards mainly take advantage of the local smoothness prior model used for images. In this paper, we propose a novel adaptive loop filter utilizing image nonlocal prior knowledge by imposing the low-rank constraint on similar image patches for compression noise reduction. In the filtering process, the reconstructed frame is first divided into image patch groups according to image patch similarity. The proposed in-loop filtering is formulated as an optimization problem with low-rank constraint for every group of image patches independently. It can be efficiently solved by soft-thresholding singular values of the matrix composed of image patches in the same group. To adapt the properties of the input sequences and bit budget, an adaptive threshold derivation model is established for every group of image patches according to the characteristics of compressed image patches, quantization parameters, and coding modes. Moreover, frame-level and largest coding unit-level control flags are signaled to further improve the adaptability from the sense of rate-distortion optimization. The performance of the proposed in-loop filter is analyzed when it collaborates with the existing in-loop filters in High Efficiency Video Coding. Extensive experimental results show that our proposed in-loop filter can further improve the performance of state-of-the-art video coding standard significantly, with up to 16% bit-rate savings. Xinfeng Zhang 0001, Ruiqin Xiong, Weisi Lin, Jian Zhang 0018, Shiqi Wang 0001, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Video Compressive Sensing Reconstruction via Reweighted Residual SparsityabstractThe compressive sensing (CS) theory indicates that robust reconstruction of signals can be obtained from far fewer measurements than those required by the Nyquist-Shannon theorem. Thus, CS has great potential in video acquisition and processing, considering that it makes the subsequent complex data compression unnecessary. In this paper, we propose a novel algorithm for effectively reconstructing videos from CS measurements. The algorithm comprises double phases, of which the first phase exploits intra-frame correlation and provides good initial recovery for each frame, and the second phase iteratively enhances reconstruction quality by alternating interframe multihypothesis (MH) prediction and sparsity modeling of residuals in a weighted manner. The weights of residual coefficients are updated in each iteration using a statistical method based on the MH predictions. These procedures are performed in the unit of overlapped patches such that potential blocking artifacts can be effectively suppressed through averaging. In addition, we devise an effective scheme based on the split Bregman iteration algorithm to solve the formulated weighted ℓ1minimization problem. The experimental results demonstrate that the proposed algorithm outperforms the state-of-the-art methods in both objective and subjective reconstruction quality. Chen Zhao 0002, Siwei Ma 0001, Jian Zhang 0018, Ruiqin Xiong, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Power Distortion Optimization for Uncoded Linear Transformed Transmission of Images and VideosabstractRecently, there is a resurgence of interest in uncoded transmission for wireless visual communication. While conventional coded systems suffer from cliff effect as the channel condition varies dynamically, uncoded linear-transformed transmission (ULT) provides elegant quality degradation for wide channel SNR range. ULT skips non-linear operations, such as quantization and entropy coding. Instead, it utilizes linear decorrelation transform and linear scaling power allocation to achieve optimized transmission. This paper presents a theoretical analysis for power-distortion optimization of ULT. In addition to the observation in our previous work that a decorrelation transform can bring significant performance gain, this paper reveals that exploiting the energy diversity in transformed signal is the key to achieve the full potential of decorrelation transform. In particular, we investigated the efficiency of ULT with exact or inexact signal statistics, highlighting the impact of signal energy modeling accuracy. Based on that, we further proposed two practical energy modeling schemes for ULT of visual signals. Experimental results show that the proposed schemes improve the quality of reconstructed images by 3~5 dB, while reducing the signal modeling overhead from hundreds or thousands of meta data to only a few meta data. The perceptual quality of reconstruction is significantly improved. Ruiqin Xiong, Jian Zhang 0018, Feng Wu 0001, Jizheng Xu, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Content-adaptive low rank regularization for image denoisingabstractPrior knowledge plays an important role in image denoising tasks. This paper utilizes the data of the input image to adaptively model the prior distribution. The proposed scheme is based on the observation that, for a natural image, a matrix consisted of its vectorized non-local similar patches is of low rank. We use a non-convex smooth surrogate for the low-rank regularization, and view the optimization problem from the empirical Bayesian perspective. In such framework, a parameter-free distribution prior is derived from the grouped non-local similar image contents. Experimental results show that the proposed approach is highly competitive with several state-of-art denoising methods in PSNR and visual quality. Hangfan Liu, Xinfeng Zhang 0001, Ruiqin Xiong |
ICIP | 3 |
| 2016 | Adaptive multi-dimension sparsity based coefficient estimation for compression artifact reductionabstractSparsity has shown promising results in various image restoration applications. Recent advances have suggested that structured or group sparsity often leads to more powerful results in compression artifact reduction studies. In this paper, we introduce nonlocal multi-dimension sparsity in an adaptive space-transform domain, which performs multi-scale wavelet transform on DCT coefficients of similar patches. The new transform efficiently reduces image redundancies between inner block and inter block simultaneously, thus it can substantially achieve sparse representation for images. Furthermore, a band-based filter is proposed to reduce compression artifacts by shrinking transform coefficients adaptively. Because of the overlapped processing, adaptive aggregation is used to combine different estimates for each block. The proposed algorithm achieves improvement over some methods in terms of both objective and subjective qualities. Xinfeng Zhang 0001, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
ICME | 3 |
| 2016 | Compressive Sensing Based Soft Video Broadcast Using Spatial and Temporal Sparsity
Xiaopeng Fan 0001, Yunhui Shi, Ruiqin Xiong, Debin Zhao |
Mob. Networks Appl. | 4 |
| 2016 | Image Denoising via Bandwise Adaptive Modeling and Regularization Exploiting Nonlocal SimilarityabstractThis paper proposes a new image denoising algorithm based on adaptive signal modeling and regularization. It improves the quality of images by regularizing each image patch using bandwise distribution modeling in transform domain. Instead of using a global model for all the patches in an image, it employs content-dependent adaptive models to address the non-stationarity of image signals and also the diversity among different transform bands. The distribution model is adaptively estimated for each patch individually. It varies from one patch location to another and also varies for different bands. In particular, we consider the estimated distribution to have non-zero expectation. To estimate the expectation and variance parameters for every band of a particular patch, we exploit the nonlocal correlation in image to collect a set of highly similar patches as the data samples to form the distribution. Irrelevant patches are excluded so that such adaptively learned model is more accurate than a global one. The image is ultimately restored via bandwise adaptive soft-thresholding, based on a Laplacian approximation of the distribution of similar-patch group transform coefficients. Experimental results demonstrate that the proposed scheme outperforms several state-of-the-art denoising methods in both the objective and the perceptual qualities. Ruiqin Xiong, Hangfan Liu, Xinfeng Zhang 0001, Jian Zhang 0018, Siwei Ma 0001, Feng Wu 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Analysis of Decorrelation Transform Gain for Uncoded Wireless Image and Video CommunicationabstractAn uncoded transmission scheme called SoftCast has recently shown great potential for wireless video transmission. Unlike conventional approaches, SoftCast processes input images only by a series of transformations and modulates the coefficients directly to a dense constellation for transmission. The transmission is uncoded and lossy in nature, with its noise level commensurate with the channel condition. This paper presents a theoretical analysis for an uncoded visual communication, focusing on developing a quantitative measurements for the efficiency of decorrelation transform in a generalized uncoded transmission framework. Our analysis reveals that the energy distribution among signal elements is critical for the efficiency of uncoded transmission. A decorrelation transform can potentially bring a significant performance gain by boosting the energy diversity in signal representation. Numerical results on Markov random process and real image and video signals are reported to evaluate the performance gain of using different transforms in uncoded transmission. The analysis presented in this paper is verified by simulated SoftCast transmissions. This provide guidelines for designing efficient uncoded video transmission schemes. Ruiqin Xiong, Feng Wu 0001, Jizheng Xu, Xiaopeng Fan 0001, Chong Luo 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Low-Rank Decomposition-Based Restoration of Compressed Images via Adaptive Noise EstimationabstractImages coded at low bit rates in real-world applications usually suffer from significant compression noise, which significantly degrades the visual quality. Traditional denoising methods are not suitable for the content-dependent compression noise, which usually assume that noise is independent and with identical distribution. In this paper, we propose a unified framework of content-adaptive estimation and reduction for compression noise via low-rank decomposition of similar image patches. We first formulate the framework of compression noise reduction based upon low-rank decomposition. Compression noises are removed by soft thresholding the singular values in singular value decomposition of every group of similar image patches. For each group of similar patches, the thresholds are adaptively determined according to compression noise levels and singular values. We analyze the relationship of image statistical characteristics in spatial and transform domains, and estimate compression noise level for every group of similar patches from statistics in both domains jointly with quantization steps. Finally, quantization constraint is applied to estimated images to avoid over-smoothing. Extensive experimental results show that the proposed method not only improves the quality of compressed images obviously for post-processing, but are also helpful for computer vision tasks as a pre-processing method. Xinfeng Zhang 0001, Weisi Lin, Ruiqin Xiong, Xianming Liu 0005, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | CONCOLOR: Constrained Non-Convex Low-Rank Model for Image DeblockingabstractDue to independent and coarse quantization of transform coefficients in each block, block-based transform coding usually introduces visually annoying blocking artifacts at low bitrates, which greatly prevents further bit reduction. To alleviate the conflict between bit reduction and quality preservation, deblocking as a post-processing strategy is an attractive and promising solution without changing existing codec. In this paper, in order to reduce blocking artifacts and obtain high-quality image, image deblocking is formulated as an optimization problem within maximum a posteriori framework, and a novel algorithm for image deblocking using constrained non-convex low-rank model is proposed. The ℓ(p) (0 < p < 1) penalty function is extended on singular values of a matrix to characterize low-rank prior model rather than the nuclear norm, while the quantization constraint is explicitly transformed into the feasible solution space to constrain the non-convex low-rank optimization. Moreover, a new quantization noise model is developed, and an alternatively minimizing strategy with adaptive parameter adjustment is developed to solve the proposed optimization problem. This parameter-free advantage enables the whole algorithm more attractive and practical. Experiments demonstrate that the proposed image deblocking algorithm outperforms the current state-of-the-art methods in both the objective quality and the perceptual quality. Jian Zhang 0018, Ruiqin Xiong, Chen Zhao 0002, Yongbing Zhang 0002, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2015 | Image denoising via adaptive soft-thresholding based on non-local samplesabstractThis paper proposes a new image denoising approach using adaptive signal modeling and adaptive soft-thresholding. It improves the image quality by regularizing all the patches in image based on distribution modeling in transform domain. Instead of using a global model for all patches, it employs content adaptive models to address the non-stationarity of image signals. The distribution model of each patch is estimated individually and can vary for different transform bands and for different patch locations. In particular, we allow the distribution model for each individual patch to have non-zero expectation. To estimate the expectation and variance parameters for the transform bands of a particular patch, we exploit the non-local correlation of image and collect a set of similar patches as data samples to form the distribution. Irrelevant patches are excluded so that this non-local based modeling is more accurate than global modeling. Adaptive soft-thresholding is employed since we observed that the distribution of non-local samples can be approximated by Laplacian distribution. Experimental results show that the proposed scheme outperforms the state-of-the-art denoising methods such as BM3D and CSR in both the PSNR and the perceptual quality. Hangfan Liu, Ruiqin Xiong, Jian Zhang 0018, Wen Gao 0001 |
CVPR | 2 |
| 2015 | Compression artifact reduction for low bit-rate images based on non-local similarity and across-resolution coherenceabstractThis paper proposes a method to estimate coefficients for blocking artifact reduction at low bit rate. Across-resolution coherence that low and high resolution image are similar is introduced to preserve signal continuity. Non-local similarity is used to provide samples for estimation by searching similar blocks of reference block. We have two sources of estimation. One source is exploiting non-local similarity to estimate coefficients of low resolution of decoded image, and interpolating the low resolution image to high resolution. We obtain the coefficients estimation for high resolution image based on the coherence across different resolutions. The other source of estimation is the quantization coefficients. These estimations are fused by their reliability respectively. Experimental results demonstrate that the proposed algorithm outperforms some recently presented methods in terms of both objective and subjective qualities of the reconstruction images. Ruiqin Xiong, Xiaopeng Fan 0001, Siwei Ma 0001 |
ISCAS | 2 |
| 2015 | Image compressive sensing using overlapped block projection and reconstructionabstractCompressive sensing allows a signal to be sampled at sub-Nyquist rate and still get recovered exactly, if the signal is sparse in some domain. Block compressive sensing (BCS) is advocated for practical image compressive sensing, since it processes image at block level and significantly reduces the memory requirement for storing projection matrix. However, existing BCS methods process blocks separately, which breaks the continuity between blocks and usually produces blocking artifacts. This paper proposes a new image compressive sensing scheme using overlapped-block projection and reconstruction (OBPR), in which the sampling is performed on overlapped blocks. During reconstruction, the sparsity constraint in transform domain is also enforced on the overlapped blocks. An augmented Lagrangian method is used to solve the optimization problem efficiently. Experimental results show that the proposed OBPR scheme achieves significantly better results than the existing BCS schemes in reconstruction quality. Sheng Shi, Ruiqin Xiong, Siwei Ma 0001, Xiaopeng Fan 0001, Wen Gao 0001 |
ISCAS | 2 |
| 2015 | High accuracy sub-pixel image registration under noisy conditionabstractImage registration plays an important role in many image processing applications. A key problem is that the accuracy of registration can be severely affected by noise. This paper presents a sub-pixel registration method for noisy images. In particular, we investigate how to estimate transitional shift to a high precision using the noisy phase data in frequency domain. Based on theoretical analysis, we find that the noise-caused phase change for every frequency component of high signal-to-noise ratio (SNR) can be well approximated by a Gaussian distribution. Furthermore, we show that the reliability of phase data can also be measured by the SNR of corresponding frequency component. A noise-robust registration framework is proposed to utilize high-SNR frequency components adaptively, while masking out the components of SNR lower than a threshold. Experiments demonstrate that the proposed method is superior to existing image registration methods in the presence of noise. Ruiqin Xiong, Siwei Ma 0001, Xiaopeng Fan 0001, Wen Gao 0001 |
ISCAS | 2 |
| 2015 | Towards accurate visual information estimation with Entropy of PrimitiveabstractRecently, a novel concept referred to as Entropy of Primitive (EoP) has been proposed for evaluating the visual information of natural images. The idea originates from the sparse representation, which has been successfully applied in a wide variety of signal processing and analysis tasks. This is because of the high efficiency of sparse representation in dealing with rich, varied and directional information contained in the natural scene. In this paper, we further explore the EoP to bridge the sparse representation and visual perception. Sparse primitives are divided into three categories depending on their visual importance. Accordingly the visual signal is decomposed into structural and non-structural layers. It is found that the image sparse representation is highly relevant with the hierarchical visual information construction process in representing the natural scene. We evaluate the efficiency and robustness of the EoP in real applications, including surveillance video and shot boundary detection. Xiang Zhang 0004, Shiqi Wang 0001, Siwei Ma 0001, Ruiqin Xiong, Wen Gao 0001 |
ISCAS | 4 |
| 2015 | Adaptive boundary dependent transform optimization for HEVCabstractHigh Efficiency Video Coding (HEVC) adopts hybrid transform coding scheme to improve the transform performance. However, it does not consider the influence of the prediction unit boundary. This paper proposes an adaptive boundary dependent transform optimization scheme for HEVC. Based on the transform unit boundary type identification by checking whether it lies at a prediction unit boundary or not, some additional transform kernels, including DCT-IV, are incorporated for transform. Moreover, the additional adopted transform kernels are generated by re-using the existing kernels in HEVC. Therefore, the butterfly structure for transform implementation can be preserved to facilitate parallel computing, and computational complexity increasing can be avoided. Experimental results show that the coding performance improvements can be up to 0.96% and 0.64% for low delay P and random access testing configurations respectively. Juanting Fan, Jicheng An, Shanshe Wang, Nan Zhang 0015, Ruiqin Xiong, Siwei Ma 0001, Shawmin Lei |
PCS | 5 |
| 2015 | Study on subjective quality assessment of Screen Content ImagesabstractWith the coming age of big data, the cloud technology, referred to as the computations or applications through the Internet, is dramatically developed. The screen content has become one of the most common data form due to the extraordinary advance of network communication technology, and the JCT-VC has started to develop new standard focusing on improving the efficiency of screen content based on High Efficiency Video Coding (HEVC). Nevertheless, the research on the quality assessment of screen content is still quite limited at the current stage. In this paper, we present a study on subjective quality assessment of the Screen Content Images (SCIs) and investigate whether the existing objective Image Quality Assessment (IQA) methods can effectively evaluate the quality of distorted SCIs. We construct a new Screen Content Database (SCD) including 24 source SCIs and 492 compressed ones with two codecs including HEVC as well as the HEVC extension. The Single Comparison (SC) method is employed for the subjective viewing to guarantee the reliability of the results. In our experiment, the correlations of eight popular IQA methods with the obtained Mean Opinion Score (MOS) values are evaluated. The result indicates that visual information fidelity method can achieve highest consistency with human visual perception. Sheng Shi, Xiang Zhang 0004, Shiqi Wang 0001, Ruiqin Xiong, Siwei Ma 0001 |
PCS | 4 |
| 2015 | An adaptive hierarchical QP setting for screen content codingabstractScreen content refers to computer generated content like text, graphics, and animations. In such video, many regions may remain static for a long period after a sudden change. Traditional hierarchical Quantization Parameter (QP) setting may not be able to handle these regions efficiently because the encoder probably needs to refine the quality of these static regions multiple times. It will cost more bits while the quality of the static regions may reach the expected degree which the flat QP setting is able to achieve. This paper proposes using different QP settings for different regions in a picture. Region classification algorithms are developed to determine whether a flat or hierarchical QP setting is used. Experimental results demonstrate that the proposed scheme can achieve an average bitrate reduction of 3.1%, and up to 8.1% bitrate reduction for IBBB coding. The proposed method improves coding efficiency without increasing encoding complexity. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
VCIP | 4 |
| 2015 | Quad-tree based inter-view motion predictionabstractAs a 3D video extension of Audio Video Coding Standard (AVS), 3D-AVS is being developed to improve the coding efficiency of multi-view video. Since multi-view video is composed of projections of the same scenery from different viewpoints at the same time instant, it contains a large amount of inter-view redundancies. To exploit the inter-view correlation, this paper presents a method to derive the motion parameters for a coding unit (CU) in the dependent view from the already coded inter-view picture. The algorithm is based on quad-tree partitioning, each CU could be recursively split into four sub-CUs of the same size and whether each sub-CU is further split is determined by comparing the derived motion parameters. Experimental results show that the proposed method provides 8.7% BD-rate saving for both video 1 and video 2 in low delay configuration and the BD-rate saving is up to 14.5% and 13.7% on video 1 and video 2 for Balloons. This method has been proposed and adopted into the 3D-AVS standard. Na Zhang 0003, Xiaopeng Fan 0001, Ruiqin Xiong, Debin Zhao |
VCIP | 4 |
| 2015 | Registration-reliability based strategy to enhance multi-frame super-resolution algorithmsabstractImage registration plays an important role in most of multi-frame super-resolution methods. As far as we know, the accuracy of most registration algorithms is not enough for superresolution, which will lead to annoying artifacts. This paper proposes a simple but effective strategy that aims to enhance the performance of existing super-resolution methods. The idea is to measure the reliability of the estimated shifts and only choose the reliable frames to reconstruct a coarse high-resolution (HR) image. This coarse HR image helps to refine all the shifts to finally get a refined HR image. An iterative contour smoothing filter is proposed to improve the accuracy of this refining process. Experimental results demonstrate that the proposed algorithm can help to improve the performance of the existing superresolution methods with fewer artifacts. Ruiqin Xiong, Xiaopeng Fan 0001, Siwei Ma 0001, Wen Gao 0001 |
VCIP | 2 |
| 2015 | Registration of under-sampled images via higher resolution spectrum restorationabstractHigh accuracy image registration is critical for the success of multi-frame super-resolution. Conventionally, the shift between images are estimated directly based on the under-sampled low-resolution (LR) image data. However, the high-frequency of LR data is unreliable due to the aliasing effect of sub-sampling, which will deteriorate the accuracy of registration. This paper proposes to resolve the aliasing by converting the LR images to high-resolution (HR) domain and then perform registration on the restored HR spectrum. To recover the HR spectrum for each corresponding LR image, we fuse the LR images into one HR image and project the estimation difference back to the reconstructed HR spectrum iteratively. To address the unequal reliabilities of different restored frequencies, weighted least square is employed to improve the precision of registration. Experimental results show that the proposed method can outperform other existing methods and improve the quality of super-resolution image. Ruiqin Xiong, Xinfeng Zhang 0001, Siwei Ma 0001, Wen Gao 0001 |
VCIP | 2 |
| 2015 | Enhanced inter prediction with localized weighted prediction in HEVCabstractInter prediction plays an important role in most video encoding systems since it could significantly improve coding performance. The more accurate the prediction of the current block, the smaller the residual, and the higher coding efficiency could be achieved accordingly. In this paper, a localized weighted prediction method is proposed to improve inter prediction accuracy. The linear regression improvement model is employed to modify the prediction pixel values. The weighting parameters are estimated in both encoder and decoder, no additional bits are required to be transmitted. The proposed method shows better coding performance than previous methods, including the explicit weighted prediction method in High Efficiency Video Coding (HEVC). Experimental results show that the BD bit rate saving of the proposed method is up to 7.4% compared to HM12.0, while the decoding complexity is almost the same. Na Zhang 0003, Yiran Lu, Xiaopeng Fan 0001, Ruiqin Xiong, Debin Zhao, Wen Gao 0001 |
VCIP | 4 |
| 2015 | A dual structured-sparsity model for compressive-sensed video reconstructionabstractThe compressive sensing theory indicates that robust reconstruction of signals can be obtained from far fewer measurements than those required by the Nyquist theorem. Thus, it has great potential in video acquisition and processing in that it can tremendously save the complex compression required by traditional video coding standards. In this paper, we consider reconstruction of compressive-sensed videos and propose a novel structured-sparsity model with a dual prediction strategy. This structured-sparsity model goes beyond simple sparsity and characterizes the intrinsic structure within the transform coefficients. Also, it exploits the sparsity of the residual between the current patch and its prediction. The prediction process is comprised of a dual strategy, which integrates the advantages of the ambient pixel domain and the measurement domain. In addition, an effective optimization method is designed for solving the formulated problem derived from the model. Experiments demonstrate that the proposed algorithm outperforms the state-of-the art methods for compressive-sensed video reconstruction in both subjective and objective quality. Chen Zhao 0002, Jian Zhang 0018, Siwei Ma 0001, Ruiqin Xiong, Wen Gao 0001 |
VCIP | 4 |
| 2015 | Video super-resolution with registration-reliability regulation and adaptive total variation
Xinfeng Zhang 0001, Ruiqin Xiong, Siwei Ma 0001, Ge Li 0002, Wen Gao 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Layered Soft Video Broadcast for Heterogeneous ReceiversabstractWireless video broadcast poses a challenge to the conventional visual communication in providing simultaneously each receiver the best video quality under its channel condition. Soft video broadcast, as a newly emerged wireless video broadcast scheme, is able to accommodate multiple receivers of different channel SNRs. However, the current soft video broadcast frameworks such as SoftCast require the bandwidth of the wireless channel to match the number of video coefficients per second. When the channel bandwidth is larger, the existing frameworks become not very efficient in bandwidth expansion. More importantly, it is possible that the users in broadcast applications have different bandwidths. However, none of the existing soft video broadcast frameworks considers bandwidth heterogeneity. In this paper, we propose a soft video broadcast framework, called LayerCast, which can simultaneously accommodate heterogeneous users with diverse SNRs and diverse bandwidths. The bandwidth expansion problem is solved by applying layered coset coding. More importantly, we derive a globally optimal power allocation between layers and, within each layer, between each DCT chunk. In simulations, the proposed framework outperforms SoftCast of up to 4 dB in video PSNR, and outperforms H.264-based framework up to 8 dB in broadcast. Xiaopeng Fan 0001, Ruiqin Xiong, Debin Zhao, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Video Compression Artifact Reduction via Spatio-Temporal Multi-Hypothesis PredictionabstractAnnoying compression artifacts exist in most of lossy coded videos at low bit rates, which are caused by coarse quantization of transform coefficients or motion compensation from distorted frames. In this paper, we propose a compression artifact reduction approach that utilizes both the spatial and the temporal correlation to form multi-hypothesis predictions from spatio-temporal similar blocks. For each transform block, three predictions with their reliabilities are estimated, respectively. The first prediction is constructed by inversely quantizing transform coefficients directly, and its reliability is determined by the variance of quantization noise. The second prediction is derived by representing each transform block with a temporal auto-regressive (TAR) model along its motion trajectory, and its corresponding reliability is estimated from local prediction errors of the TAR model. The last prediction infers the original coefficients from similar blocks in non-local regions, and its reliability is estimated based on the distribution of coefficients in these similar blocks. Finally, all the predictions are adaptively fused according to their reliabilities to restore high-quality videos. The experimental results show that the proposed method can efficiently reduce most of the compression artifacts and improve both subjective and objective quality of block transform coded videos. Xinfeng Zhang 0001, Ruiqin Xiong, Weisi Lin, Siwei Ma 0001, Jiaying Liu 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | G-CAST: Gradient Based Image SoftCast for Perception-Friendly Wireless Visual CommunicationabstractConventional image and video communication systems are usually designed with the objective being to maximize the fidelity of reconstructed images measured by mean square errors (MSE). It is well known that the fidelity metric MSE may not reflect the visual quality perceived by human eyes. Recent advancements in image quality assessment tell us that the structural similarity (SSIM), especially the gradient similarity, reveals the perceptual fidelity of images more reliably. Inspired by this observation, this paper proposes a new image communication approach, which conveys the visual information in an image by transmitting the image gradients and recovers the image from the received gradient data at decoder side using statistical image prior knowledge. In particular, we designed a gradient-based image SoftCast scheme for wireless scenarios. Experimental results show that the proposed scheme can produce reconstruction images with much better perceptual quality. The advantage in perceptual quality is verified by the quality improvement measured by the metrics SSIM and gradient signal-to-noise ratio (GSNR). Ruiqin Xiong, Hangfan Liu, Siwei Ma 0001, Xiaopeng Fan 0001, Feng Wu 0001, Wen Gao 0001 |
DCC | 1 |
| 2014 | Hybridcast: A wireless image/video SoftCast scheme using layered representation and hybrid digital-analog modulationabstractThe recently proposed SoftCast scheme employs analog-like transmission for wireless visual communication, providing graceful reconstruction quality degradation for drastically changing channel conditions. However, the transmission in SoftCast is not always efficient in terms of power usage. In this paper, we propose a wireless image/video SoftCast scheme which employs layered representation with hybrid digital-analog modulation. In this scheme, a coarse approximation of the image is coded in a base layer in digital ways, while the residual image details are delivered in an enhancement layer in analog-like way. The outputs from the two layers are superimposed for transmission, using a hybrid digital-analog modulation scheme. Since a major part of the signal is handled by the base layer, the power efficiency of the SoftCast layer is significantly improved. Experimental results show that the proposed scheme outperforms the original SoftCast remarkably, while still preserving the smooth quality degradation characteristic of the SoftCast scheme. Zhihai Song, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
ICIP | 2 |
| 2014 | High quality image reconstruction via non-local collaborative estimation for wireless image/video softcastabstractFor wireless scenarios where the channel condition fluctuates unpredictably, a novel image/video communication scheme, named SoftCast, was recently proposed to provide graceful quality degradation and competitive performance simultaneously. Unlike conventional approaches, SoftCast decorrelates input images by a transform and modulates the coefficients directly to a dense constellation for transmission, leaving out the conventional quantization, entropy coding and channel coding. The transmission is lossy in nature, with its noise level commensurate with the channel condition. To reconstruct images from the received noisy data, SoftCast employs a linear least-square estimator (LLSE), but it tends to produce annoying reconstruction artifacts. This paper proposes a high-quality image reconstruction algorithm for SoftCast, employing a collaborative estimator to utilize both the local correlation and non-local similarity within images. Experimental results show that the proposed method outperforms the existing SoftCast scheme, achieving remarkable improvement in the objective and subjective qualities of the reconstruction images. Ruiqin Xiong, Jian Zhang 0018, Feng Wu 0001, Wen Gao 0001 |
ICIP | 1 |
| 2014 | Artifact reduction of compressed video via three-dimensional adaptive estimation of transform coefficientsabstractBlock transform compressed videos usually suffer from annoying artifacts at low bit rates, caused by the coarse quantization of transform coefficients. The inter prediction utilized in video coding also induces block boundary artifacts when the neighboring blocks using different motion vectors from previous decoded frames. In this paper, a three-dimensional adaptive estimation method for video transform coefficients is proposed to reduce the artifact in compressed video. In the proposed method, transform coefficients of each block are estimated by adaptively fusing three prediction sources based on their reliabilities. One prediction source is the transform coefficients directly acquired from the decoded video, whose reliability is determined by the distribution of the quantization noise. The other prediction source is derived by a temporal autoregressive model of transform-blocks along motion trajectory among neighboring reference frames. Its reliability is estimated from prediction variance of local blocks. The last prediction source is derived from the nonlocal transform-blocks, whose reliability is estimated based on the distribution of the nonlocal coefficients and their similarity with the estimated block. Experimental results for compressed video sequences by HEVC show that, the proposed method can reduce the compression artifacts and improve both the objective and subjective quality. Xinfeng Zhang 0001, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
ICIP | 2 |
| 2014 | Gradient based image transmission and reconstruction using non-local gradient sparsity regularizationabstractMost existing image coding and communication systems aim to minimize the mean square error (MSE) of the pixels reconstructed at receivers. However, the quality metric MSE has long been criticized for not being consistent with the perception of human vision systems. This paper considers a gradient-based image SoftCast (G-Cast) scheme, based on the recent advancements in image quality assessment which indicate that gradient similarity is highly correlated with perceptual image quality. To reconstruct the image from the received noisy gradient data, we exploit the statistical characteristics of image gradients. Instead of using the very simple Laplacian distribution for image gradient as in the total variation (TV) model, we further exploit the non-local similarity of image patches. A non-local gradient sparsity regularization (NLGSR) method is developed and solved using augmented Lagrangian method. Experimental results show that the proposed scheme provides promising perceptual image quality, and the NLGSR reconstruction scheme outperforms the existing schemes remarkably. Hangfan Liu, Ruiqin Xiong, Siwei Ma 0001, Xiaopeng Fan 0001, Wen Gao 0001 |
ICME | 2 |
| 2014 | Layered image/video softcast with hybrid digital-analog transmission for robust wireless visual communicationabstractDue to the mobility of transceivers and the interference from other signals, the condition of a wireless channel may vary drastically and unpredictably. In such scenario, conventional communication systems usually suffer from cliff effect due to the nature of entropy coding and channel coding. The recently proposed SoftCast scheme, on the contrary, achieves smooth quality degradation by employing analog-like transmission, together with efficient decorrelation and power allocation. However, the analog-like transmission in SoftCast is not always efficient in terms of power usage, when compared with digital approaches. In this paper, we propose a layered image/video SoftCast scheme, in which a coarse approximation of the image is coded in a base layer and transmitted in digital way while the remained image details are represented in an enhancement layer and sent out using the SoftCast way. Since a major part of the signal energy is transmitted in the base layer using digital approach, the efficiency of power usage in the enhancement layer is improved remarkably. Experimental results show that the proposed scheme can outperform the original SoftCast scheme remarkably, while still preserving the smooth quality degradation characteristic of SoftCast. Zhihai Song, Ruiqin Xiong, Siwei Ma 0001, Xiaopeng Fan 0001, Wen Gao 0001 |
ICME | 2 |
| 2014 | Non-local extension of total variation regularization for image restorationabstractTotal-variation (TV) regularization is widely adopted in image restoration problems to exploit the feature that natural images are smooth with small gradient values at most regions. Basic TV method assumes identical zero-mean Laplacian distribution for the gradients at all pixels. However, for real-world images, the statistics of gradients may not be stationary, and the zero-mean assumption of gradients may not be valid either for a specific pixel. This paper presents a non-local extension of TV regularization for image restoration, called Non-Local Gradient Sparsity Regularization (NGSR). The NGSR model employs a separate gradient value distribution for each pixel. To figure out the distribution parameters, the NGSR method exploits a set of patches which are similar to the patch centered at current pixel and estimates the distribution parameter adaptively. Experimental results demonstrate that the proposed NGSR outperforms traditional TV remarkably for image restoration. Hangfan Liu, Ruiqin Xiong, Siwei Ma 0001, Xiaopeng Fan 0001, Wen Gao 0001 |
ISCAS | 2 |
| 2014 | Transform domain energy modeling of natural images for wireless SoftCast optimizationabstractThe SoftCast scheme recently proposed for wireless visual communication avoids the threshold effect that traditional communication systems usually suffer from. It provides graceful quality transition by sending images in a sequence of whitened transform coefficients using dense-constellation modulation and analog-like transmission. A key point in SoftCast is that it allocates transmission power among coefficients unequally, according to the expected energy of each coefficient. Importantly, the energy diversity utilized by power allocation should be shared with the receiver for correct decoding. Signaling the energy for each coefficient individually is prohibitive since it requires a large set of meta data. Grouping coefficients into a few chunks and signaling the energy at chunk level, on the other hand, may compromise the efficiency of SoftCast remarkably. In this paper, we investigate the energy distribution of natural images in transform domain and propose a model to approximate this distribution. We apply the model to guide the power allocation in SoftCast. Experimental results show that the proposed method outperforms the equal-chunk approach in the original SoftCast by 2~5 dB, and reduces the number of required meta data significantly at the same time. Zhihai Song, Ruiqin Xiong, Xiaopeng Fan 0001, Siwei Ma 0001, Wen Gao 0001 |
ISCAS | 2 |
| 2014 | Gradient based image/video softcast with grouped-patch collaborative reconstructionabstractInspired by the recent image quality assessment (IQA) studies which indicate that the image gradient data reflects the visual information more reliably than the image pixels, gradient based transmission scheme was recently proposed to pursue better perceptual quality for wireless visual communication. This paper develops an effective method to reconstruct high quality image from the received noisy gradient data. The proposed method utilizes both local correlation and non-local similarity within the image signal to regularize the reconstruction image. Principle component analysis (PCA) is employed to learn signal-adaptive two-dimensional (2D) transform basis, and 3D transform is performed on grouped similar patches to further decorrelate the coefficients. In this way, distortions can be effectively suppressed via adaptive collaborative shrinkage on the transform coefficients. Experimental results demonstrate that the proposed method improves the reconstruction performance remarkably compared with the existing schemes. Hangfan Liu, Ruiqin Xiong, Siwei Ma 0001, Xiaopeng Fan 0001, Wen Gao 0001 |
VCIP | 2 |
| 2014 | Image Restoration Using Joint Statistical Modeling in a Space-Transform DomainabstractThis paper presents a novel strategy for high-fidelity image restoration by characterizing both local smoothness and nonlocal self-similarity of natural images in a unified statistical manner. The main contributions are three-fold. First, from the perspective of image statistics, a joint statistical modeling (JSM) in an adaptive hybrid space-transform domain is established, which offers a powerful mechanism of combining local smoothness and nonlocal self-similarity simultaneously to ensure a more reliable and robust estimation. Second, a new form of minimization functional for solving the image inverse problem is formulated using JSM under a regularization-based framework. Finally, in order to make JSM tractable and robust, a new Split Bregman-based algorithm is developed to efficiently solve the above severely underdetermined inverse problem associated with theoretical proof of convergence. Extensive experiments on image inpainting, image deblurring, and mixed Gaussian plus salt-and-pepper noise removal applications verify the effectiveness of the proposed algorithm. Jian Zhang 0018, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Performance analysis of transform in uncoded wireless visual communicationabstractIn wireless scenarios where the channel condition may vary drastically, visual communication systems using source and channel coding generally suffer from threshold effect. An uncoded transmission scheme called SoftCast [1]-[3], however, was recently shown to provide both graceful quality transition and competitive performance. In SoftCast, image signal is directly modulated to a dense constellation using proper power for transmission, solely after employing a transform for energy compaction, leaving out conventional quantization, entropy coding and channel coding. The received signal is lossy in nature, with its noise level commensurate with the channel condition. This paper presents a theoretical analysis for uncoded visual communication, focusing on the role of transform and the quantitative measurement of transform gain in a generalized uncoded transmission framework with optimal power allocation. Our analysis reveal that the energy distribution among signal elements plays an important role in the power-distortion performance. Further analysis show that the energy compaction capability of decorrelation transform can bring significant gain by boosting the energy diversity in signal representation. Numerical analysis results are reported for Markov random signals and natural images, respectively. The performance of typical transforms, e.g. KLT, DCT and DWT, and the effect of different transform sizes or levels are evaluated. These analysis results are verified by simulations. Ruiqin Xiong, Feng Wu 0001, Jizheng Xu, Wen Gao 0001 |
ISCAS | 1 |
| 2013 | Improved total variation based image compressive sensing recovery by nonlocal regularizationabstractRecently, total variation (TV) based minimization algorithms have achieved great success in compressive sensing (CS) recovery for natural images due to its virtue of preserving edges. However, the use of TV is not able to recover the fine details and textures, and often suffers from undesirable staircase artifact. To reduce these effects, this paper presents an improved TV based image CS recovery algorithm by introducing a new nonlocal regularization constraint into CS optimization problem. The nonlocal regularization is built on the well known nonlocal means (NLM) filtering and takes advantage of self-similarity in images, which helps to suppress the staircase effect and restore the fine details. Furthermore, an efficient augmented Lagrangian based algorithm is developed to solve the above combined TV and nonlocal regularization constrained problem. Experimental results demonstrate that the proposed algorithm achieves significant performance improvements over the state-of-the-art TV based algorithm in both PSNR and visual perception. Jian Zhang 0018, Shaohui Liu, Ruiqin Xiong, Siwei Ma 0001, Debin Zhao |
ISCAS | 3 |
| 2013 | Cactus: a hybrid digital-analog wireless video communication systemabstractThis paper challenges the conventional wisdom that video redundancy should be removed as much as possible for efficient communications. We discover that, by keeping spatial redundancy at the sender and properly utilizing it at the receiver, we can build a more robust and even more efficient wireless video communication system than existing ones. Hao Cui 0001, Zhihai Song, Chong Luo 0001, Ruiqin Xiong, Feng Wu 0001 |
MSWiM | 5 |
| 2013 | Power-distortion optimization for wireless image/video SoftCast by transform coefficients energy modeling with adaptive chunk divisionabstractTraditional communication systems usually suffer from the threshold effect when channel signal-to-noise ratio (CSNR) fluctuates unpredictably in wireless and mobile scenarios. The SoftCast scheme, however, provides graceful quality transition in wide CSNR range. In SoftCast, input image is decorrelated by a transform and modulated directly to a dense constellation for transmission, leaving out the conventional quantization, entropy coding and channel coding. A key point of SoftCast is that the transmission power needs to be allocated among the transform coefficients unequally, according to the energy of coefficients. Importantly, the energy diversity used to guide power allocation should be shared between the sender and the receiver for correct decoding. This paper addresses the power distortion optimization problem, introducing a new adaptive chunk division scheme to describe the energy diversity among coefficients. A concrete algorithm is developed to determine the chunk boundaries that achieve optimal transmission power usage. Experimental results show that the proposed scheme can improve the performance of the original SoftCast by 4~8dB using a smaller number of chunks. Ruiqin Xiong, Feng Wu 0001, Xiaopeng Fan 0001, Chong Luo 0001, Siwei Ma 0001, Wen Gao 0001 |
VCIP | 1 |
| 2013 | Distributed soft video broadcast with variable block size motion estimationabstractIn recent years, video broadcast has become a popular application, but the traditional hierarchical design requires the source to pick a bitrate and video resolution for encoding before transmission, it cannot be efficient to accommodate users with different channel quality. The proposed DCAST scheme can solve this problem. DCAST uses the ME and MC technology to generate a predicted frame which helps current frame to do coset coding. However, in the ME process, DCAST uses fixed block size. In this paper, we use the variable block size motion estimation to replace the fixed block size motion estimation, it can effectively reduce the block effect and improve the quality of the reconstructed frame generated by MC and predicted frame. The DCAST with variable block size motion estimation is 0.5dB better than DCAST with fixed block size motion estimation. Xiaopeng Fan 0001, Ruiqin Xiong, Debin Zhao |
VCIP | 3 |
| 2013 | Correlation estimation for distributed wireless video communicationabstractOne important problem in distributed video coding is to estimate the variance of the correlation noise between the video signal and its decoder side information. This variance is hard to estimate due to the lack of the motion vectors at the encoder side. In this paper, we first propose a linear model to estimate this variance by referring the zero motion prediction at the encoder based on a Markov field assumption. Furthermore, not only the prediction noise from the video signal itself but also the additional noise due to wireless transmission is considered in this paper. We applied our correlation estimation method in our recent distributed wireless visual communication framework called DCAST. The experimental results show that the proposed method improves the video PSNR by 0.5-1.5dB while avoiding motion estimation at encoder. Na Zhang 0003, Xiaopeng Fan 0001, Ruiqin Xiong, Debin Zhao |
VCIP | 4 |
| 2013 | PET Protection Optimization for Streaming Scalable Videos With Multiple TransmissionsabstractThis paper investigates priority encoding transmission (PET) protection for streaming scalably compressed video streams over erasure channels, for the scenarios where a small number of retransmissions are allowed. In principle, the optimal protection depends not only on the importance of each stream element, but also on the expected channel behavior. By formulating a collection of hypotheses concerning its own behavior in future transmissions, limited-retransmission PET (LR-PET) effectively constructs channel codes spanning multiple transmission slots and thus offers better protection efficiency than the original PET. As the number of transmission opportunities increases, the optimization for LR-PET becomes very challenging because the number of hypothetical retransmission paths increases exponentially. As a key contribution, this paper develops a method to derive the effective recovery-probability versus redundancy-rate characteristic for the LR-PET procedure with any number of transmission opportunities. This significantly accelerates the protection assignment procedure in the original LR-PET with only two transmissions, and also makes a quick and optimal protection assignment feasible for scenarios where more transmissions are possible. This paper also gives a concrete proof to the redundancy embedding property of the channel codes formed by LR-PET, which allows for a decoupled optimization for sequentially dependent source elements with convex utility-length characteristic. This essentially justifies the source-independent construction of the protection convex hull for LR-PET. Ruiqin Xiong, David S. Taubman, Vijay Sivaraman |
IEEE Trans. Image Process. | 1 |
| 2013 | Compression Artifact Reduction by Overlapped-Block Transform Coefficient Estimation With Block SimilarityabstractBlock transform coded images usually suffer from annoying artifacts at low bit rates, caused by the coarse quantization of transform coefficients. In this paper, we propose a new method to reduce compression artifacts by the overlapped-block transform coefficient estimation from non-local blocks. In the proposed method, the discrete cosine transform coefficients of each block are estimated by adaptively fusing two prediction values based on their reliabilities. One prediction is the quantized values of coefficients decoded from the compressed bitstream, whose reliability is determined by quantization steps. The other prediction is the weighted average of the coefficients in nonlocal blocks, whose reliability depends on the variance of the coefficients in these blocks. The weights are used to distinguish the effectiveness of the coefficients in nonlocal blocks to predict original coefficients and are determined by block similarity in transform domain. To solve the optimization problem, the overlapped blocks are divided into several subsets. Each subset contains nonoverlapped blocks covering the whole image and is optimized independently. Therefore, the overall optimization is reduced to a set of sub-optimization problems, which can be easily solved. Finally, we provide a strategy for parameter selection based on the compression levels. Experimental results show that the proposed method can remarkably reduce compression artifacts and significantly improve both the subjective and objective qualities of block transform coded images. Xinfeng Zhang 0001, Ruiqin Xiong, Xiaopeng Fan 0001, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | Multi-scale Spatial Error Concealment via Hybrid Bayesian RegressionabstractIn this paper, we propose a novel multi-scale spatial error concealment algorithm to combine the modeling strengthes of the parametric and nonparametric Bayesian regression. We progressively recover missing blocks in the scale space from coarse to fine so that the sharp edges and texture in the finest scale can be eventually recovered. On one hand, in each scale, the nonparametric part of the methodology is used to exploit the intra-scale correlation, which relies on the data itself to dictate the structure of the model. In this procedure, the non-local self-similarity property is utilized as a fruitful resource for abstracting a priori knowledge of images. On the other hand, the parametric part is used to explicitly model the inter-scale correlation, in which the local structure regularity is thoroughly explored to recover the sharp edges and major texture features of images. It is not respected if only the nonparametric modeling is considering. We achieve the best of both worlds within a multi-scale framework. Experimental results on benchmark test images demonstrate that the proposed method achieves very competitive performance with the state-of-the-art error concealment algorithms. Xianming Liu 0005, Deming Zhai, Guangtao Zhai, Debin Zhao, Ruiqin Xiong, Wen Gao 0001 |
DCC | 5 |
| 2012 | Compressed Sensing Recovery via Collaborative SparsityabstractCompressed Sensing (CS) has drawn quite an amount of attention as a joint sampling and compression approach. Its theory shows that a signal can be decoded from many fewer measurements than suggested by the Nyquist sampling theory, when the signal is sparse in some domain. So one of the most significant challenges in CS is to seek a domain where a signal can exhibit a high degree of sparsity and hence be recovered faithfully. Most of conventional CS recovery approaches, however, exploited a set of fixed bases (e.g. DCT, wavelet and gradient domain) for the entirety of a signal, which are irrespective of the nonstationarity of natural signals and cannot achieve high enough degree of sparsity, thus resulting in poor rate-distortion performance. In this paper, we propose a new framework for compressed sensing recovery via collaborative sparsity (RCoS), which enforces local two-dimensional sparsity and nonlocal three-dimensional sparsity simultaneously in an adaptive hybrid space-transform domain, thus substantially utilizing intrinsic sparsities of natural images and greatly confining the CS solution space. In addition, an efficient augmented Lagrangian based technique is developed to solve the above optimization problem. Experimental results on a wide range of natural images are presented to demonstrate the efficacy of the new CS recovery strategy. Jian Zhang 0018, Debin Zhao, Chen Zhao 0002, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
DCC | 4 |
| 2012 | Reducing Blocking Artifacts in Compressed Images via Transform-Domain Non-local Coefficients EstimationabstractBlock transform coding using discrete cosine transform is the most popular approach for image compression. However, many annoying blocking artifacts are generated due to coarse quantization on transform coefficients independently. This paper proposes an effective blocking artifacts reduction method by estimating the transform coefficients from their quantized version. In the proposed scheme, we estimate the transform coefficients based on an image statistic model and non-local similarity among blocks in transform domain. The parameters used in our proposed scheme are discussed and adaptively selected. Extensive experimental results show that our proposed method significantly reduces blocking artifacts and improves the subjective and the objective quality of block transform coded images. Xinfeng Zhang 0001, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
ICME | 2 |
| 2012 | Exploiting Image Local and Nonlocal Consistency for Mixed Gaussian-Impulse Noise RemovalabstractMost existing image denoising algorithms can only deal with a single type of noise, which violates the fact that the noisy observed images in practice are often suffered from more than one type of noise during the process of acquisition and transmission. In this paper, we propose a new variational algorithm for mixed Gaussian-impulse noise removal by exploiting image local consistency and nonlocal consistency simultaneously. Specifically, the local consistency is measured by a hyper-Lap lace prior, enforcing the local smoothness of images, while the nonlocal consistency is measured by three-dimensional sparsity of similar blocks, enforcing the nonlocal self-similarity of natural images. Moreover, a Split-Bregman based technique is developed to solve the above optimization problem efficiently. Extensive experiments for mixed Gaussian plus impulse noise show that significant performance improvements over the current state-of-the-art schemes have been achieved, which substantiates the effectiveness of the proposed algorithm. Jian Zhang 0018, Ruiqin Xiong, Chen Zhao 0002, Siwei Ma 0001, Debin Zhao |
ICME | 2 |
| 2012 | Image super-resolution via dual-dictionary learning and sparse representationabstractLearning-based image super-resolution aims to reconstruct high-frequency (HF) details from the prior model trained by a set of high- and low-resolution image patches. In this paper, HF to be estimated is considered as a combination of two components: main high-frequency (MHF) and residual high-frequency (RHF), and we propose a novel image super-resolution method via dual-dictionary learning and sparse representation, which consists of the main dictionary learning and the residual dictionary learning, to recover MHF and RHF respectively. Extensive experimental results on test images validate that by employing the proposed two-layer progressive scheme, more image details can be recovered and much better results can be achieved than the state-of-the-art algorithms in terms of both PSNR and visual perception. Jian Zhang 0018, Chen Zhao 0002, Ruiqin Xiong, Siwei Ma 0001, Debin Zhao |
ISCAS | 3 |
| 2012 | Adaptive loop filter with temporal predictionabstractIn this paper, we propose a method to improve adaptive loop filter (ALF) efficiency with temporal prediction. For one frame, two sets of adaptive loop filter parameters are adaptively selected by rate distortion optimization. The first set of ALF parameters is estimated by minimizing the mean square error between the original frame and the current reconstructed frame. The second set of filter parameters is the one that is used in the latest prior frame. The proposed algorithm is implemented in HM3.0 software. Compared with the HM3.0 anchor, the proposed method achieves 0.4%, 0.3% and 0.3% BD bitrate reduction in average for high efficiency low delay B, high efficiency low delay P and high efficiency random access configuration, respectively. The encoding and decoding time increase by 1% and 2% on average, respectively. Xinfeng Zhang 0001, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
PCS | 2 |
| 2012 | WaveCast: Wavelet based wireless video broadcast using lossy transmissionabstractWireless video broadcasting is a popular application of mobile network. However, the traditional approaches have limited supports to the accommodation of users with diverse channel conditions. The newly emerged Softcast approach provides smooth multicast performance but is not very efficient in inter frame compression. In this work, we propose a new video multicast approach: WaveCast. Different from softcast, WaveCast utilizes motion compensated temporal filter (MCTF) to exploit inter frame redundancy, and utilizes conventional framework to transmit motion information such that the MVs can be reconstructed losslessly. Meanwhile, WaveCast transmits the transform coefficients in lossy mode and performs gracefully in multicast. In experiments, Wave-Cast outperforms softcast 2dB in video PSNR at low channel SNR, and outperforms H.264 based framework up to 8dB in broadcast. Xiaopeng Fan 0001, Ruiqin Xiong, Feng Wu 0001, Debin Zhao |
VCIP | 2 |
| 2012 | Compound image compression by multi-stage predictionabstractComputer generated compound images contain not only photographic images but also text and graphics images, which makes it difficult to be efficiently compressed using exiting image and video coding standards. This paper presents a new compound image coding scheme based on the intra coding framework of the upcoming high efficiency video coding (HEVC) standard. The proposed scheme first decomposes the compound image into color components and structure components. Then two-stage prediction scheme is employed to exploit the correlation among structure components, where the first prediction is generated by directional prediction and second prediction is generated by template matching. Experimental results show that the proposed scheme achieves up to 10.2 dB coding gain on compound images compared with HEVC. And on average it achieves 42.2 percent bitrate saving compared with HEVC. Weijia Zhu, Wenpeng Ding, Ruiqin Xiong, Yunhui Shi |
VCIP | 3 |
| 2012 | Multiple Hypotheses Bayesian Frame Rate Up-Conversion by Adaptive Fusion of Motion-Compensated InterpolationsabstractFrame rate up-conversion (FRUC) improves the viewing experience of a video because the motion in a FRUC-constructed high frame-rate video looks more smooth and continuous. This paper proposes a multiple hypotheses Bayesian FRUC scheme for estimating the intermediate frame with maximum a posteriori probability, in which both temporal motion model and spatial image model are incorporated into the optimization criterion. The image model describes the spatial structure of neighboring pixels while the motion model describes the temporal correlation of pixels along motion trajectories. Instead of employing a single uniquely optimal motion, multiple “optimal” motion trajectories are utilized to form a group of motion hypotheses. To obtain accurate estimation for the pixels in missing intermediate frames, the motion-compensated interpolations generated by all these motion hypotheses are adaptively fused according to the reliability of each hypothesis. We revealed by numerical analysis that this reliability (i.e., the variance of interpolation errors along the hypothesized motion trajectory) can be measured by the variation of reference pixels along the motion trajectory. To obtain the multiple motion fields, a set of block-matching sizes is used and the motion fields are estimated by progressively reducing the size of matching block. Experimental results show that the proposed method can significantly improve both the objective and the subjective quality of the constructed high frame rate video. Hongbin Liu 0004, Ruiqin Xiong, Debin Zhao, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Transductive Regression with Local and Global Consistency for Image Super-ResolutionabstractIn this paper, we propose a novel image super-resolution algorithm, referred to as interpolation based on transductive regression with local and global consistency (TRLGC). Our algorithm first constructs a set of local interpolation models which can predict the intensity labels of all image samples, and a loss term will be minimized to keep the predicted labels of available low-resolution (LR) samples sufficiently close to the original ones. Then, all of the losses evaluated in local neighborhoods are accumulated together to measure the global consistency on all samples. Furthermore, a graph-Laplacian based manifold regularization term is incorporated to penalize the global smoothness of intensity labels, such smoothing can alleviate the insufficient training of the local models and make them more robust. Finally, we construct a unified objective function to combine together the accumulated loss of the locally linear regression, square error of prediction bias on the available LR samples and the manifold regularization term, which could be solved with a closed-form solution as a convex optimization problem. In this way, a transductive regression algorithm with local and global consistency is developed. Experimental results on benchmark test images demonstrate that the proposed image super-resolution method achieves very competitive performance with the state-of-the-art algorithms. Xianming Liu 0005, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001, Huifang Sun |
DCC | 3 |
| 2011 | Adaptive Image Deblurring via Tanner Graph Representation and Belief PropagationabstractSummary form only given. In this paper, we propose a deblurring framework based on a factor graph representation of the image and the image formation process. Each pixel is described by a variable node, while the statistical relation among pixels is formulated by two sets of check nodes, describing the local image structures and the image formation process, respectively. Belief propagation is employed to solve the pixel values and it is reduced to mechanisms to generate and fuse predictions for each pixel iteratively. A key work is that we analyzed the origin of ringing artifacts and found that it is due to the propagation of estimation error in previous iterations. We propose a method to estimate the uncertainty in each pixel of previous estimation, which is then used to adapt the generation and fusion of prediction in the next iteration. Experimental results show that the proposed solution can significantly eliminate ringing artifacts without employing any image priors. Ruiqin Xiong |
DCC | 1 |
| 2011 | Bayesian frame interpolation by fusing multiple motion-compensated prediction framesabstractFor video frame rate up-conversion, new frames need to be interpolated along motion trajectories and inserted between the original adjacent frames. This paper proposes a Bayesian frame interpolation strategy, which considers both the spatial (i.e. intra-frame) correlation model and the temporal (i.e. inter-frame) correlation model after motion compensation. Different from the conventional schemes that perform interpolation using only one estimated motion vector field (MVF), the proposed strategy estimates multiple MVFs. To cope with different scales of frame contents, we generate the multiple MVFs through variable block size motion estimation. A criterion is adopted to estimate reliability of the motion compensated prediction frames from different motion fields. Experimental results demonstrate that the proposed strategy improves the interpolation performance remarkably. Hongbin Liu 0004, Ruiqin Xiong, Siwei Ma 0001, Debin Zhao, Wen Gao 0001 |
ICIP | 2 |
| 2011 | A new image deblurring algorithm with less ringing artifacts via error variance estimation and soft decisionabstractImage deblurring is an ill-posed linear inverse problem. Most traditional algorithms suffer from severe ringing artifacts. Recent approaches handle this issue by regularization techniques based on assumed image prior models. This paper presents a new method to reduce the ringing artifacts, without introducing any image prior models. For this purpose, we revisit the deblurring problem, using a probabilistic graph to model the image formation process. We establish the link between iterative back-projection and belief propagation and show that the ringing artifacts are caused by error propagation. Based on these analysis, we introduce a method to measure the variance of an estimation image and further propose an error-variance aware deblurring algorithm. Experimental results demonstrate that the proposed algorithm is very effective in suppressing the ringing artifacts. Ruiqin Xiong, Wenpeng Ding, Siwei Ma 0001, Wen Gao 0001 |
ICIP | 1 |
| 2011 | Side information extrapolation with temporal and spatial consistencyabstractIn this paper, we present an efficient side information extrapolation scheme with temporal and spatial consistency for low delay Wyner-Ziv video coding. Our method is based on the regularized local linear regression (RLLR) model, in which each pixel in SI is approximated as a linear weighted combination of samples within a local temporal neighborhood. The optimal model parameters are estimated by projecting the transformation function onto the temporal training samples to exploit motion-related dependency. During this procedure, moving weights are incorporated into the objective function to express the relative importance of training samples in estimating parameters of the model. Furthermore, spatial correlation is explored by imposing an additional local smoothness penalty, which does good to estimate the occluded regions and complex motion regions. The learned function is smooth and locally linear, and can be obtained with a closed-form solution by solving a convex optimization problem. Experimental results demonstrate that the RLLR method achieves very competitive SI extrapolation performance compared with the state-of-the-art methods. Xianming Liu 0005, Deming Zhai, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
ISCAS | 4 |
| 2011 | GPU based fast algorithm for tanner graph based image interpolationabstractIn image/video processing software and hardware products, low complexity interpolation algorithms, such as cubic and splines methods, are commonly used. However, these methods tend to blur textures and produce jaggy effect compared with other adaptive methods such as NEDI, SAI. Tanner graph based image interpolation algorithm has better effect in dealing with edge and texture, but with high computation complexity. Thanks to the high performance parallel processing capability of today's GPU, use of complex algorithms for real time application is becoming possible. In this paper, we present a fast algorithm for tanner graph based image interpolation and it's implementation on GPU. In our algorithm, the image model training process of tanner graph based image interpolation is greatly simplified. Experimental results show that the GPU implementation can be more than 47 times as fast as the CPU implementation. Wei Lei, Ruiqin Xiong, Siwei Ma 0001, Luhong Liang |
MMSP | 2 |
| 2011 | New image coding scheme with hierarchical representation and adaptive interpolationabstractIn this paper, we propose a new image coding scheme which combines the advantage of hierarchical (i.e. multi- resolution) representation, adaptive interpolation and rate- distortion optimization capability of block-based coding. We divide the coding process into two layers by partitioning the image pixels into four groups with chessboard structure. In the first layer, pixels in one of these partitions are encoded and reconstructed using the state-of-the-art coding scheme. In the second layer, directional adaptive interpolation is employed to predict and code the pixels in the remaining three partitions. Joint mode decision is performed for the co-located macroblocks of these three partitions. Skip mode is introduced to save coding bits for blocks where the interpolation process works very well. Experimental results demonstrate that the proposed coding scheme achieves very promising performance. Xinfeng Zhang 0001, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
VCIP | 2 |
| 2011 | High-quality image restoration from partial random samples in spatial domainabstractIn this paper, a novel algorithm for high-quality image restoration is proposed. The contributions of this work are two-fold. First, a new form of minimization function for solving image inverse problems is formulated via combining local total variation model and nonlocal adaptive 3-D sparse representation model as regularizers under the regularization- based framework. Second, a new Split-Bregman based iterative algorithm is developed to solve the above optimization problem efficiently associated with proved theoretical convergence property. Experimental results on image restoration from partial random samples have shown that the proposed algorithm achieves significant performance improvements over the current state-of-the-art schemes and exhibits nice convergence property. Jian Zhang 0018, Ruiqin Xiong, Siwei Ma 0001, Debin Zhao |
VCIP | 2 |
| 2011 | Fast mode dependent directional transform via butterfly-style transform and integer lifting steps
Wenpeng Ding, Ruiqin Xiong, Yunhui Shi, Dehui Kong |
J. Vis. Commun. Image Represent. | 2 |
| 2011 | Image Interpolation Via Regularized Local Linear RegressionabstractThe linear regression model is a very attractive tool to design effective image interpolation schemes. Some regression-based image interpolation algorithms have been proposed in the literature, in which the objective functions are optimized by ordinary least squares (OLS). However, it is shown that interpolation with OLS may have some undesirable properties from a robustness point of view: even small amounts of outliers can dramatically affect the estimates. To address these issues, in this paper we propose a novel image interpolation algorithm based on regularized local linear regression (RLLR). Starting with the linear regression model where we replace the OLS error norm with the moving least squares (MLS) error norm leads to a robust estimator of local image structure. To keep the solution stable and avoid overfitting, we incorporate the l(2)-norm as the estimator complexity penalty. Moreover, motivated by recent progress on manifold-based semi-supervised learning, we explicitly consider the intrinsic manifold structure by making use of both measured and unmeasured data points. Specifically, our framework incorporates the geometric structure of the marginal probability distribution induced by unmeasured samples as an additional local smoothness preserving constraint. The optimal model parameters can be obtained with a closed-form solution by solving a convex optimization problem. Experimental results on benchmark test images demonstrate that the proposed method achieves very competitive performance with the state-of-the-art interpolation algorithms, especially in image edge structure preservation. Xianming Liu 0005, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001, Huifang Sun |
IEEE Trans. Image Process. | 3 |
| 2011 | Interpolation-Dependent Image DownsamplingabstractTraditional methods for image downsampling commit to remove the aliasing artifacts. However, the influences on the quality of the image interpolated from the downsampled one are usually neglected. To tackle this problem, in this paper, we propose an interpolation-dependent image downsampling (IDID), where interpolation is hinged to downsampling. Given an interpolation method, the goal of IDID is to obtain a downsampled image that minimizes the sum of square errors between the input image and the one interpolated from the corresponding downsampled image. Utilizing a least squares algorithm, the solution of IDID is derived as the inverse operator of upsampling. We also devise a content-dependent IDID for the interpolation methods with varying interpolation coefficients. Numerous experimental results demonstrate the viability and efficiency of the proposed IDID. Yongbing Zhang 0002, Debin Zhao, Jian Zhang 0018, Ruiqin Xiong, Wen Gao 0001 |
IEEE Trans. Image Process. | 4 |
| 2010 | Tanner Graph Based Image InterpolationabstractThis paper interprets image interpolation as a channel decoding problem and proposes a tanner graph based interpolation framework, which regards each pixel in an image as a variable node and the local image structure around each pixel as a check node. The pixels available from low-resolution image are "received" whereas other missing pixels of highresolution image are "erased", through an imaginary channel. Local image structures exhibited by the low-resolution image provide information on the joint distribution of pixels in a small neighborhood, and thus play the same role as parity symbols in the classic channel coding scenarios. We develop an efficient solution for the sum-product algorithm of belief propagation in this framework, based on a gaussian auto-regressive image model. Initial experiments show up to 3dB gain over other methods with the same image model. The proposed framework is flexible in message processing at each node and provides much room for incorporating more sophisticated image modelling techniques. Ruiqin Xiong, Wen Gao 0001 |
DCC | 1 |
| 2010 | A practical algorithm for tanner graph based image interpolationabstractThis paper interprets image interpolation as a decoding problem on tanner graph and proposes a practical belief propagation algorithm based on a gaussian autoregressive image model. This algorithm regards belief propagation as a way to generate and fuse predictions from various check nodes. A low complexity implementation of this algorithm measures and distributes the departure of current interpolation result from the image model. Convergence speed of the proposed algorithm is discussed. Experimental results show that good interpolation results can be obtained by a very small number of iterations. Ruiqin Xiong, Wenpeng Ding, Siwei Ma 0001, Wen Gao 0001 |
ICIP | 1 |
| 2010 | Image interpolation via regularized local linear regressionabstractIn this paper, we present an efficient image interpolation scheme by using regularized local linear regression (RLLR). On one hand, we introduce a robust estimator of local image structure based on moving least squares, which can efficiently handle the statistical outliers compared with ordinary least squares based methods. On the other hand, motivated by recent progress on manifold based semi-supervise learning, the intrinsic manifold structure is explicitly considered by making use of both measured and unmeasured data points. In particular, the geometric structure of the marginal probability distribution induced by unmeasured samples is incorporated as an additional locality preserving constraint. The optimal model parameters can be obtained with a closed-form solution by solving a convex optimization problem. Experimental results demonstrate that our method outperform the existing methods in both objective and subjective visual quality over a wide range of test images. Xianming Liu 0005, Debin Zhao, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
PCS | 3 |
| 2010 | Improved autoregressive image model estimation for directional image interpolationabstractFor image interpolation algorithms employing autoregressive models, a mechanism is required to estimate the model parameters piecewisely and accurately so that local structures of image can be exploited efficiently. This paper proposes a new strategy for better estimating the model. Different from conventional schemes which build the model solely upon the co-variance matrix of low-resolution image, the proposed strategy utilizes the covariance matrix of high-resolution image itself, with missing pixels properly initialized. To make the estimation robust, we adopt a general solution which exploits the covariance matrices of both scales. Experimental results demonstrate that the proposed strategy improves model estimation and the interpolation performance remarkably. Ruiqin Xiong, Wenpeng Ding, Siwei Ma 0001, Wen Gao 0001 |
PCS | 1 |
| 2010 | A robust video super-resolution algorithmabstractIn this paper, we propose a robust video super-resolution reconstruction method based on spatial-temporal orientation-adaptive kernel regression. First, we propose a robust registration efficiency model to reflect the temporal information reliability. Second, we propose a spatial-temporal steering kernel considering motions between frames and structures in each low resolution frame. Simulation results demonstrate that our new super-resolution method substantially improves both the subjective quality and objective quality than other resolution enhancement methods. Xinfeng Zhang 0001, Ruiqin Xiong, Siwei Ma 0001, Wen Gao 0001 |
PCS | 2 |
| 2010 | Low bit-rate image coding via interpolation oriented adaptive down-samplingabstractAn interpolation oriented adaptive down-sampling algorithm is proposed for low bit-rate image coding in this paper. Given an image, the proposed algorithm is able to obtain a low resolution image, from which a high quality image with the same resolution as the input image can be interpolated. Different from the traditional down-sampling algorithms, which are independent from the interpolation process, the proposed down-sampling algorithm hinges the down-sampling to the interpolation process. Consequently, the proposed down-sampling algorithm is able to maintain the original information of the input image to the largest extent. The down-sampled image is then fed into JPEG. A total variation (TV) based post processing is then applied to the decompressed low resolution image. Ultimately, the processed image is interpolated to maintain the original resolution of the input image. Experimental results verify that utilizing the downsampled image by the proposed algorithm, an interpolated image with much higher quality can be achieved. Besides, the proposed algorithm is able to achieve superior performance than JPEG for low bit rate image coding. Yongbing Zhang 0002, Jian Zhang 0018, Ruiqin Xiong, Debin Zhao, Siwei Ma 0001 |
VCIP | 3 |
| 2010 | Robust video super-resolution with registration efficiency adaptationabstractSuper-Resolution (SR) is a technique to construct a high-resolution (HR) frame by fusing a group of low-resolution (LR) frames describing the same scene. The effectiveness of the conventional super-resolution techniques, when applied on video sequences, strongly relies on the efficiency of motion alignment achieved by image registration. Unfortunately, such efficiency is limited by the motion complexity in the video and the capability of adopted motion model. In image regions with severe registration errors, annoying artifacts usually appear in the produced super-resolution video. This paper proposes a robust video super-resolution technique that adapts itself to the spatially-varying registration efficiency. The reliability of each reference pixel is measured by the corresponding registration error and incorporated into the optimization objective function of SR reconstruction. This makes the SR reconstruction highly immune to the registration errors, as outliers with higher registration errors are assigned lower weights in the objective function. In particular, we carefully design a mechanism to assign weights according to registration errors. The proposed superresolution scheme has been tested with various video sequences and experimental results clearly demonstrate the effectiveness of the proposed method. Xinfeng Zhang 0001, Ruiqin Xiong, Siwei Ma 0001, Li Zhang 0006, Wen Gao 0001 |
VCIP | 2 |
| 2010 | Optimal PET Protection for Streaming Scalably Compressed Video Streams With Limited Retransmission Based on Incomplete FeedbackabstractFor streaming scalably compressed video streams over unreliable networks, Limited-Retransmission Priority Encoding Transmission (LR-PET) outperforms PET remarkably since the opportunity to retransmit is fully exploited by hypothesizing the possible future retransmission behavior before the retransmission really occurs. For the retransmission to be efficient in such a scheme, it is critical to get adequate acknowledgment from a previous transmission before deciding what data to retransmit. However, in many scenarios, the presence of a stochastic packet delay process results in frequent late acknowledgements, while imperfect feedback channels can impair the server's knowledge of what the client has received. This paper proposes an extended LR-PET scheme, which optimizes PET-protection of transmitted bitstreams, recognizing that the received feedback information is likely to be incomplete. Similar to the original LR-PET, the behavior of future retransmissions is hypothesized in the optimization objective of each transmission opportunity. As the key contribution, we develop a method to efficiently derive the effective recovery probability versus redundancy rate characteristic for the extended LR-PET communication process. This significantly simplifies the ultimate protection assignment procedure. This paper also demonstrates the advantage of the proposed strategy over several alternative strategies. Ruiqin Xiong, David S. Taubman, Vijay Sivaraman |
IEEE Trans. Image Process. | 1 |
| 2008 | Optimal LR-PET protection for scalable video streams over lossy channels with random delayabstractThis paper investigates the optimal PET protection for streaming scalably compressed streams over networks where the delivery time constraints allow limited retransmissions (LR) and the communication channels exhibit both random losses and delays. A key property must be considered in this scenario is the possibility that a packet successfully arrives at the receiver in time, even if its acknowledgment is not received by the sender at certain deadlines. This paper proposes an extended LRPET scheme, namely random-delay LR-PET, in which additional streams may be sent to provide supplemental protection for the packets whose acknowledgments are still missing at a specified time after the transmission. To determine the optimal protection in each transmission opportunity, hypotheses concerning the number of acknowledged packets and the effect of future retransmission are considered. As the key contribution of this paper, we develop a method to derive the effective overall recovery probability versus redundancy characteristic, which significantly simplifies the actual protection assignment procedure. This paper also demonstrates the benefits of the optimization strategy proposed for this random-delay LR-PET scheme and the cruciality of time selection for scheduling retransmission. Ruiqin Xiong, David S. Taubman |
MMSP | 1 |
| 2008 | In-Scale Motion Compensation for Spatially Scalable Video CodingabstractIn existing pyramid-based spatially scalable coding schemes, such as H.264/MPEG-4 SVC (scalable video coding), video frame at a certain high-resolution layer is mainly predicted either from the same frame at the next lower resolution layer, or from the temporal neighboring frames within the same resolution layer. But these schemes fail to exploit both kinds of correlation simultaneously and therefore cannot remove the redundancies among resolution layers efficiently. This paper extends the idea of spatiotemporal subband transform and proposes a general in-scale motion compensation technique for pyramid-based spatially scalable video coding. Video frame at each high-resolution layer is partitioned into two parts in frequency. Prediction for the lowpass part is derived from the next lower resolution layer, whereas prediction for the highpass part is obtained from neighboring frames within the same resolution layer, to further utilize temporal correlation. In this way, both kinds of correlation are exploited simultaneously and the cross-resolution-layer redundancy can be highly removed. Furthermore, this paper also proposes a macroblock-based adaptive in-scale technique for hybrid spatial and SNR scalability. Experimental results show that the proposed techniques can significantly improve the spatial scalability performance of H.264/MPEG-4 SVC, especially when the bit-rate ratio of lower resolution bit stream to higher resolution bit stream is considerable. Ruiqin Xiong, Jizheng Xu, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Macroblock-Based Adaptive In-Scale Prediction for Scalable Video CodingabstractThis paper extends the so called generalized in-scale coding technique developed in our previous work to a macroblock-based adaptive in-scale motion compensation framework for H.264/MPEG-4 SVC (scalable video coding). In this framework, the in-scale technique is employed to generate predictions for video frames at high resolution layers, which can utilize both temporal correlation and cross-resolution correlation simultaneously. The lowpass content of a high resolution frame is predicted from its lower resolution layer and the highpass content is predicted from neighboring frames at the same resolution layer. For quality and resolution combined scalability, the prediction from lower resolution layer is not always better than temporal prediction. The in-scale technique is thereby employed adaptively at macroblock level according to the local characteristics of video signal. Techniques for motion estimation and mode decision are also investigated. Experimental results demonstrate that the scheme proposed in this paper can outperform H.264/MPEG-4 SVC significantly. The coding performance is quite promising especially for high-fidelity video coding Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001 |
ISCAS | 1 |
| 2007 | Generalized in-scale motion compensation framework for spatial scalable video codingabstractIn existing video coding schemes with spatial scalability based on pyramid frame representation, such as the ongoing H.264/MPEG-4 SVC (scalable video coding) standard, video frame at a high resolution is mainly predicted either from the lower-resolution image of the same frame or from the temporal neighboring frames at the same resolution. Most of these prediction techniques fail to exploit the two correlations simultaneously and efficiently. This paper extends the in-scale prediction technique developed for wavelet video coding to a generalized in-scale motion compensation framework for H.264/MPEG-4 SVC. In this framework, for a video frame at a high resolution layer, the lowpass content is predicted from the information already coded in lower resolution layer, but the highpass content is predicted by exploiting the neighboring frames at current resolution. In this way, both the cross-resolution correlation and temporal correlation are exploited simultaneously, which leads to much higher efficiency in prediction. Preliminary experimental results demonstrate that the proposed framework improves the spatial scalability performance of current H.264/MPEG-4 SVC. The improvement is significant especially for high-fidelity video coding. In addition, another advantage over wavelet-based in-scale scheme is achieved that the proposed framework can support arbitrary down-sampling and up-sampling filters. Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001 |
VCIP | 1 |
| 2007 | Barbell-Lifting Based 3-D Wavelet Coding SchemeabstractThis paper provides an overview of the Barbell lifting coding scheme that has been adopted as common software by the MPEG ad hoc group on further exploration of wavelet video coding. The core techniques used in this scheme, such as Barbell lifting, layered motion coding, 3D entropy coding and base layer embedding, are discussed. The paper also analyzes and compares the proposed scheme with the oncoming scalable video coding (SVC) standard because the hierarchical temporal prediction technique used in SVC has a close relationship with motion compensated temporal lifting (MCTF) in wavelet coding. The commonalities and differences between these two schemes are exhibited for readers to better understand modern scalable video coding technologies. Several challenges that still exist in scalable video coding, e.g., performance of spatial scalable coding and accurate MC lifting, are also discussed. Two new techniques are presented in this paper although they are not yet integrated into the common software. Finally, experimental results demonstrate the performance of the Barbell-lifting coding scheme and compare it with SVC and another well-known 3D wavelet coding scheme, MC embedded zero block coding (MC-EZBC). Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2007 | Subband Coupling Aware Rate Allocation for Spatial Scalability in 3-D Wavelet Video CodingabstractThe motion compensated temporal filtering (MCTF) technique, which is extensively used in 3-D wavelet video coding schemes nowadays, leads to signal coupling among various spatial subbands because motion alignment is introduced in the temporal filtering. Using all spatial subbands as a reference enables MCTF to fully take advantage of temporal correlation across frames but inevitably brings drifting problem in supporting spatial scalability. This paper first analyzes the signal coupling phenomenon and then proposes a quantitative model to describe signal propagation across spatial subbands during the MCTF process. The signal propagation is modeled for a single MC step based on the shifting effect of wavelet synthesis filters and then it is extended to multilevel MCTF. This model is called subband coupling aware signal propagation (SCASP) model in this paper. Based on the model, we further propose a subband coupling aware rate allocation scheme as one possible solution to the above dilemma in supporting spatial scalability. To find the optimal rate allocation among all subbands for a specified reconstruction resolution, the SCASP model is used to approximate the reconstruction process and derive the synthesis gain of each subband with regard to that reconstruction. Experimental results have fully demonstrated the advantages of our proposed rate allocation scheme in improving both objective and subjective qualities of reconstructed low-resolution video, especially at middle bit rates and high bit rates. Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001, Ya-Qin Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2006 | A Lifting-Based Wavelet Transform Supporting Non-Dyadic Spatial ScalabilityabstractIn subband image and video compression, the issue of dyadic spatial scalability with a downsizing ratio of 2:1 has already been widely investigated. In some application scenarios, however, other downsizing ratios can be desirable. This paper proposes a wavelet coding scheme to support non-dyadic spatial scalability. By extending lifting-based transform structure, we show how wavelet transform can directly support non-dyadic spatial scalability. Specially, based on the extended lifting framework, a wavelet transform is designed to decompose an image with size 3W×3H to a low-pass subband with size 2W×2H and three high-pass subbands to support 3:2-ratio spatial scalability. The characteristic of proposed wavelet low-pass filter is investigated and experimental results show that the low resolution image produced by our scheme has good quality while the coding performance just has marginal loss compared to Daubechies 9/7 filter. Ruiqin Xiong, Jizheng Xu, Feng Wu 0001 |
ICIP | 1 |
| 2006 | Adaptive MCTF based on Correlation Noise Model for SNR Scalable Video CodingabstractThis paper proposes a subband adaptive motion compensated temporal filtering (MCTF) technique for scalable video coding and introduces a revised synthesis gain model for the quantization in this adaptive MCTF scheme. In scalable video coding, hierarchical MCTF is extensively adopted to exploit the temporal correlation across video frames. In this hierarchical MCTF structure, the strength of temporal correlation varies with the level of temporal transform and varies with the various spatial frequency components in a frame. The reconstruction noises also have diverse strength at various subbands. According to the correlation and noise characteristics of various subbands, we can adjust the strength of motion compensated prediction step in MCTF to maximally take the advantage of temporal correlation but restrict the propagation of reconstruction noise. The quantization step of each subband is also adjusted according to synthesis gain determined by the MCTF structure. In this way an adaptive MCTF scheme is formed and the proposed technique improves the coding performance of scalable video coding Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001 |
ICME | 1 |
| 2006 | In-scale motion aligned temporal filteringabstractTo handle the mismatch problems of spatial-domain motion aligned temporal filtering (MATF) in providing spatial scalability, this paper presents a novel in-scale motion aligned temporal filtering (ISMATF) for highly scalable video coding. For each lifting step of ISMATF, the predict or update is carried out at various resolutions in a scale-by-scale way and the operation at higher resolution is designed to embed the operations at lower resolutions to prevent redundancy in frame representation. The proposed ISMATF has several main features, which as a whole differentiates it from other schemes (e.g. spatial-domain MATF, in-band MATF, straightforward multi-resolution MATF) and makes it effective in video coding with spatial scalabilities. These features include 1) Multi-resolution lifting with perfect reconstruction capability; 2) non-redundancy in frame representation, which favors highly scalable video coding; 3) temporal filtering performed in the image-domain of corresponding resolution, which is different from in-band schemes and preserves high efficiency of motion compensation; 4) no mismatch between the encoder and decoder, no matter which resolution is decoded. The proposed ISMATF scheme solves the mismatch problems of SDMATF at low resolution while maintains a good performance at high resolution. It can outperform redundant multi-resolution MATF schemes by up to 1.0 dB at high resolution. Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001 |
ISCAS | 1 |
| 2004 | Layered motion estimation and coding for fully scalable 3d wavelet video codingabstractThis paper proposes a framework of scalable motion estimation and coding with the structure of multilayers for 3D wavelet video coding. The motion representation consists of multiple layers. The encoder uses motion of all layers to perform analysis, while the decoder may receive only part of motion for synthesis. Different from other schemes, each layer of motion is a point optimized at a certain range of bit-rate. We observe that the distortion introduced by motion mismatch is highly independent with the rate for texture in a wide range. Therefore, to make the best trade-off between motion and texture under the constraint of a given bit rate, a motion layer decision algorithm is used to find the appropriate number of motion layers to be included into the bit-stream. The proposed framework also supports the spatial and temporal scalabilities of motion. Experimental results show significant improvement at low bit-rates and nearly no loss at high bit-rates with layered motion coding and optimal motion decision. The performance is approaching to the convex hull of those with multiple sets of nonscalable motion. Ruiqin Xiong, Jizheng Xu, Feng Wu 0001, Shipeng Li 0001, Ya-Qin Zhang |
ICIP | 1 |
| 2004 | Exploiting temporal correlation with adaptive block-size motion alignment for 3D wavelet codingabstractThis paper proposes an adaptive block-size motion alignment technique in 3D wavelet coding to further exploit temporal correlations across pictures. Similar to B picture in traditional video coding, each macroblock can motion align from forward and/or backward for temporal wavelet de-composition. In each direction, a macroblock may select its partition from one of seven modes - 16x16, 8x16, 16x8, 8x8, 8x4, 4x8 and 4x4 - to allow accurate motion alignment. Furthermore, the rate-distortion optimization criterions are proposed to select motion mode, motion vectors and partition mode. Although the proposed technique greatly improves the accuracy of motion alignment, it does not directly bring the coding efficiency gain because of smaller block size and more block boundaries. Therefore, an overlapped block motion alignment is further proposed to cope with block boundaries and to suppress spatial high-frequency components. The experimental results show the proposed adaptive block-size motion alignment with the overlapped block motion alignment can achieve up to 1.0 dB gain in 3D wavelet video coding. Our 3D wavelet coder outperforms the MC-EZBC for most sequences by 1~2dB and we are doing up to 1.5 dB better than H.264. Ruiqin Xiong, Feng Wu 0001, Shipeng Li 0001, Zixiang Xiong, Ya-Qin Zhang |
VCIP | 1 |