VLDB 2026 Research / reviewers in the wild / expert
Xiaoyun Zhang 0001
dblp:40/1945-1
· DBLP profile ↗
61ranked-venue papers
0as first author
27since 2021 · last 2025
0000-0001-7680-4062ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 21 since 2021Artificial intelligence and machine learning · 22 · 15 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VRVVC: Variable-Rate NeRF-Based Volumetric Video CompressionabstractNeural Radiance Field (NeRF)-based volumetric video has revolutionized visual media by delivering photorealistic Free-Viewpoint Video (FVV) experiences that provide audiences with unprecedented immersion and interactivity. However, the substantial data volumes pose significant challenges for storage and transmission. Existing solutions typically optimize NeRF representation and compression independently or focus on a single fixed rate-distortion (RD) tradeoff. In this paper, we propose VRVVC, a novel end-to-end joint optimization variable-rate framework for volumetric video compression that achieves variable bitrates using a single model while maintaining superior RD performance. Specifically, VRVVC introduces a compact tri-plane implicit residual representation for inter-frame modeling of long-duration dynamic scenes, effectively reducing temporal redundancy. We further propose a variable-rate residual representation compression scheme that leverages a learnable quantization and a tiny MLP-based entropy model. This approach enables variable bitrates through the utilization of predefined Lagrange multipliers to manage the quantization error of all latent representations. Finally, we present an end-to-end progressive training strategy combined with a multi-rate-distortion loss function to optimize the entire framework. Extensive experiments demonstrate that VRVVC achieves a wide range of variable bitrates within a single model and surpasses the RD performance of existing methods across various datasets. Qiang Hu 0003, Houqiang Zhong, Zihan Zheng, Xiaoyun Zhang 0001, Zhengxue Cheng, Li Song 0001, Guangtao Zhai, Yanfeng Wang 0001 |
AAAI | 4 |
| 2025 | FineVQ: Fine-Grained User Generated Content Video Quality AssessmentabstractThe rapid growth of user-generated content (UGC) videos has produced an urgent need for effective video quality assessment (VQA) algorithms to monitor video quality and guide optimization and recommendation procedures. However, current VQA models generally only give an overall rating for a UGC video, which lacks fine-grained labels for serving video processing and recommendation applications. To address the challenges and promote the development of UGC videos, we establish the first large-scale Fine-grained Video quality assessment Database, termed FineVD, which comprises 6104 UGC videos with fine-grained quality scores and descriptions across multiple dimensions. Based on this database, we propose a Fine-grained Video Quality assessment (FineVQ) model to learn the fine-grained quality of UGC videos, with the capabilities of quality rating, quality scoring, and quality attribution. Extensive experimental results demonstrate that our proposed FineVQ can produce fine-grained video-quality results and achieve state-of-the-art performance on FineVD and other commonly used UGC-VQA datasets. Both FineVD and FineVQ are publicly available at: https://github.com/IntMeGroup/FineVQ. Huiyu Duan, Qiang Hu 0003, Zitong Xu, Lu Liu 0005, Xiongkuo Min, Chunlei Cai, Tianxiao Ye, Xiaoyun Zhang 0001, Guangtao Zhai |
CVPR | 10 |
| 2025 | 4DGC: Rate-Aware 4D Gaussian Compression for Efficient Streamable Free-Viewpoint Videoabstract3D Gaussian Splatting (3DGS) has substantial potential for enabling photorealistic Free-Viewpoint Video (FVV) experiences. However, the vast number of Gaussians and their associated attributes poses significant challenges for storage and transmission. Existing methods typically handle dynamic 3DGS representation and compression separately, neglecting motion information and the rate-distortion (RD) trade-off during training, leading to performance degradation and increased model redundancy. To address this gap, we propose 4DGC, a novel rate-aware 4D Gaussian compression framework that significantly reduces storage size while maintaining superior RD performance for FVV. Specifically, 4DGC introduces a motion-aware dynamic Gaussian representation that utilizes a compact motion grid combined with sparse compensated Gaussians to exploit inter-frame similarities. This representation effectively handles large motions, preserving quality and reducing temporal redundancy. Furthermore, we present an end-to-end compression scheme that employs differentiable quantization and a tiny implicit entropy model to compress the motion grid and compensated Gaussians efficiently. The entire framework is jointly optimized using a rate-distortion trade-off. Extensive experiments demonstrate that 4DGC supports variable bitrates and consistently outperforms existing methods in RD performance across multiple datasets. Qiang Hu 0003, Zihan Zheng, Houqiang Zhong, Sihua Fu, Li Song 0001, Xiaoyun Zhang 0001, Guangtao Zhai, Yanfeng Wang 0001 |
CVPR | 6 |
| 2025 | F-Bench: Rethinking Human Preference Evaluation Metrics for Benchmarking Face Generation, Customization, and RestorationabstractArtificial intelligence generative models exhibit remarkable capabilities in content creation, particularly in face image generation, customization, and restoration. However, current AI-generated faces (AIGFs) often fall short of human preferences due to unique distortions, unrealistic details, and unexpected identity shifts, underscoring the need for a comprehensive quality evaluation framework for AIGFs. To address this need, we introduce FaceQ, a large-scale, comprehensive database of AI-generated Face images with fine-grained Quality annotations reflecting human preferences. The FaceQ database comprises 12,255 images generated by 29 models across three tasks: (1) face generation, (2) face customization, and (3) face restoration. It includes 32,742 mean opinion scores (MOSs) from 180 annotators, assessed across multiple dimensions: quality, authenticity, identity (ID) fidelity, and text-image correspondence. Using the FaceQ database, we establish F-Bench, a benchmark for comparing and evaluating face generation, customization, and restoration models, highlighting strengths and weaknesses across various prompts and evaluation dimensions. Additionally, we assess the performance of existing image quality assessment (IQA), face quality assessment (FQA), AI-generated content image quality assessment (AIGCIQA), and preference evaluation metrics, manifesting that these standard metrics are relatively ineffective in evaluating authenticity, ID fidelity, and text-image correspondence. The FaceQ database will be publicly available upon publication. Lu Liu 0005, Huiyu Duan, Qiang Hu 0003, Chunlei Cai, Tianxiao Ye, Huayu Liu, Xiaoyun Zhang 0001, Guangtao Zhai |
ICCV | 8 |
| 2025 | TD-BFR: Truncated Diffusion Model for Efficient Blind Face RestorationabstractDiffusion-based methodologies have shown significant potential in blind face restoration (BFR), leveraging their robust generative capabilities. However, they are often criticized for two significant problems: 1) slow training and inference speed, and 2) inadequate recovery of fine-grained facial details. To address these problems, we propose a novel Truncated Diffusion model for efficient Blind Face Restoration (TD-BFR), a three-stage paradigm tailored for the progressive resolution of degraded images. Specifically, TD-BFR utilizes an innovative truncated sampling method, starting from low-quality (LQ) images at low resolution to enhance sampling speed, and then introduces an adaptive degradation removal module to handle unknown degradations and connect the generation processes across different resolutions. Additionally, we further adapt the priors of pre-trained diffusion models to recover rich facial details. Our method efficiently restores high-quality images in a coarse-to-fine manner and experimental results demonstrate that TD-BFR is, on average, 4.75× faster than current state-of-the-art diffusion-based BFR methods while maintaining competitive quality. Ziying Zhang, Zhixin Wang, Qiang Hu 0003, Xiaoyun Zhang 0001 |
ICME | 5 |
| 2025 | Serial Low-rank Adaptation of Vision TransformerabstractFine-tuning large pre-trained vision foundation models in a parameter-efficient manner is critical for downstream vision tasks, considering the practical constraints of computational and storage costs. Low-rank adaptation (LoRA) is a well-established technique in this domain, achieving impressive efficiency by reducing the parameter space to a low-rank form. However, developing more advanced low-rank adaptation methods to reduce parameters and memory requirements remains a significant challenge in resource-constrained application scenarios. In this study, we consider on top of the commonly used vision transformer and propose Serial LoRA, a novel LoRA variant that introduces a shared low-rank matrix serially composite with the attention mechanism. Such a design extracts the underlying commonality of parameters in adaptation, significantly reducing redundancy. Notably, Serial LoRA uses only ${\color {Magenta}{\text{1/4}}}$ parameters of LoRA but achieves comparable performance in most cases. We conduct extensive experiments on a range of vision foundation models with the transformer structure, and the results confirm consistent superiority of our method. Houqiang Zhong, Shaocheng Shen, Ke Cai, Zhenlong Wu, Jiangchao Yao, Xiaoyun Zhang 0001, Li Song 0001, Qiang Hu 0003 |
ICME | 8 |
| 2025 | 4DGCPro: Efficient Hierarchical 4D Gaussian Compression for Progressive Volumetric Video StreamingabstractAchieving seamless viewing of high-fidelity volumetric video, comparable to 2D video experiences, remains an open challenge. Existing volumetric video compression methods either lack the flexibility to adjust quality and bitrate within a single model for efficient streaming across diverse networks and devices, or struggle with real-time decoding and rendering on lightweight mobile platforms. To address these challenges, we introduce 4DGCPro, a novel hierarchical 4D Gaussian compression framework that facilitates real-time mobile decoding and high-quality rendering via progressive volumetric video streaming in a single bitstream. Specifically, we propose a perceptually-weighted and compression-friendly hierarchical 4D Gaussian representation with motion-aware adaptive grouping to reduce temporal redundancy, preserve coherence, and enable scalable multi-level detail streaming. Furthermore, we present an end-to-end entropy-optimized training scheme, which incorporates layer-wise rate-distortion (RD) supervision and attribute-specific entropy modeling for efficient bitstream generation. Extensive experiments show that 4DGCPro enables flexible quality and variable bitrate within a single model, achieving real-time decoding and rendering on mobile devices while outperforming existing methods in RD performance across multiple datasets. Zihan Zheng, Zhenlong Wu, Houqiang Zhong, Yuan Tian 0017, Lan Xu 0003, Jiangchao Yao, Xiaoyun Zhang 0001, Qiang Hu 0003, Wenjun Zhang 0001 |
NeurIPS | 8 |
| 2025 | Robust ID-Specific Face Restoration via Alignment Learning
Yushun Fang, Lu Liu 0005, Qiang Hu 0003, Jianghe Cui, Gang Chen 0040, Xiaoyun Zhang 0001 |
PRCV (9) | 8 |
| 2025 | PrismGS: Physically-Grounded Anti-Aliasing for High-Fidelity Large-Scale 3D Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has recently enabled real-time photorealistic rendering in compact scenes, but scaling to large urban environments introduces severe aliasing artifacts and optimization instability, especially under high-resolution (e.g., 4K) rendering. These artifacts, manifesting as flickering textures and jagged edges, arise from the mismatch between Gaussian primitives and the multi-scale nature of urban geometry. While existing "divide-and-conquer" pipelines address scalability, they fail to resolve this fidelity gap. In this paper, we propose PrismGS, a physically-grounded regularization framework that improves the intrinsic rendering behavior of 3D Gaussians. PrismGS integrates two synergistic regularizers. The first is pyramidal multi-scale supervision, which enforces consistency by supervising the rendering against a pre-filtered image pyramid. This compels the model to learn an inherently anti-aliased representation that remains coherent across different viewing scales, directly mitigating flickering textures. This is complemented by an explicit size regularization that imposes a physically-grounded lower bound on the dimensions of the 3D Gaussians. This prevents the formation of degenerate, view-dependent primitives, leading to more stable and plausible geometric surfaces and reducing jagged edges. Our method is plug-and-play and compatible with existing pipelines. Extensive experiments on MatrixCity, Mill-19, and UrbanScene3D demonstrate that PrismGS achieves state-of-the-art performance, yielding significant PSNR gains around 1.5 dB against CityGaussian, while maintaining its superior quality and robustness under demanding 4K rendering. Houqiang Zhong, Zhenglong Wu, Sihua Fu, Zihan Zheng, Xin Jin 0014, Xiaoyun Zhang 0001, Li Song 0001, Qiang Hu 0003 |
VCIP | 6 |
| 2025 | MegaFusion: Extend Diffusion Models towards Higher-resolution Image Generation without Further TuningabstractDiffusion models have emerged as frontrunners in text-to-image generation, but their fixed image resolution during training often leads to challenges in high-resolution image generation, such as semantic deviations and object replication. This paper introduces MegaFusion, a novel approach that extends existing diffusion-based text-to-image models towards efficient higher-resolution generation without additional fine-tuning or adaptation. Specifically, we employ an innovative truncate and relay strategy to bridge the denoising processes across different resolutions, allowing for high-resolution image generation in a coarse-to-fine manner. Moreover, by integrating dilated convolutions and noise re-scheduling, we further adapt the model's priors for higher resolution. The versatility and efficacy of MegaFusion make it universally applicable to both latent-space and pixel-space diffusion models, along with other derivative models. Extensive experiments confirm that MegaFusion significantly boosts the capability of existing models to pro-duce images of megapixels and various aspect ratios, while only requiring about 40% of the original computational cost. Code is available at https://haoningwu3639.github.io/MegaFusion/. Haoning Wu 0002, Shaocheng Shen, Qiang Hu 0003, Xiaoyun Zhang 0001, Ya Zhang 0002, Yanfeng Wang 0001 |
WACV | 4 |
| 2025 | VARFVV: View-Adaptive Real-Time Interactive Free-View Video Streaming With Edge ComputingabstractFree-view video (FVV) allows users to explore immersive video content from multiple views. However, delivering FVV poses significant challenges due to the uncertainty in view switching, combined with the substantial bandwidth and computational resources required to transmit and decode multiple video streams, which may result in frequent playback interruptions. Existing approaches, either client-based or cloud-based, struggle to meet high Quality of Experience (QoE) requirements under limited bandwidth and computational resources. To address these issues, we propose VARFVV, a bandwidth- and computationally-efficient system that enables real-time interactive FVV streaming with high QoE and low switching delay. Specifically, VARFVV introduces a low-complexity FVV generation scheme that reassembles multiview video frames at the edge server based on user-selected view tracks, eliminating the need for transcoding and significantly reducing computational overhead. This design makes it well-suited for large-scale, mobile-based UHD FVV experiences. Furthermore, we present a popularity-adaptive bit allocation method, leveraging a graph neural network, that predicts view popularity and dynamically adjusts bit allocation to maximize QoE within bandwidth constraints. We also construct an FVV dataset comprising 330 videos from 10 scenes, including basketball, opera, etc. Extensive experiments show that VARFVV surpasses existing methods in video quality, switching latency, computational efficiency, and bandwidth usage, supporting over 500 users on a single edge server with a switching delay of 71.5ms. Our code and dataset are available at https://github.com/qianghu-huber/VARFVV. Qiang Hu 0003, Qihan He, Houqiang Zhong, Guo Lu, Xiaoyun Zhang 0001, Guangtao Zhai, Yanfeng Wang 0001 |
IEEE J. Sel. Areas Commun. | 5 |
| 2024 | Intelligent Grimm - Open-ended Visual Storytelling via Latent Diffusion ModelsabstractGenerative models have recently exhibited exceptional capabilities in text-to-image generation, but still struggle to generate image sequences coherently. In this work, we focus on a novel, yet challenging task of generating a co-herent image sequence based on a given storyline, denoted as open-ended visual storytelling. We make the following three contributions: (i) to fulfill the task of visual sto-rytelling, we propose a learning-based auto-regressive im-age generation model, termed as Story Gen, with a novel vision-language context module, that enables to generate the current frame by conditioning on the corresponding text prompt and preceding image-caption pairs; (ii) to ad-dress the data shortage of visual storytelling, we collect paired image-text sequences by sourcing from online videos and open-source E-books, establishing processing pipeline for constructing a large-scale dataset with diverse characters, storylines, and artistic styles, named StorySalon; (iii) Quantitative experiments and human evaluations have vali-dated the superiority of our StoryGen, where we show it can generalize to unseen characters without any optimization, and generate image sequences with coherent content and consistent character. Code, dataset, and models are avail-able at https://haoningwu3639.github.io/StoryGen_Webpage/. “Mirror mirror on the wall, who's the fairest of them all?” -Grimms' Fairy Tales Chang Liu 0079, Haoning Wu 0002, Xiaoyun Zhang 0001, Yanfeng Wang 0001, Weidi Xie |
CVPR | 4 |
| 2024 | JOINTRF: End-To-End Joint Optimization for Dynamic Neural Radiance Field Representation and CompressionabstractNeural Radiance Field (NeRF) excels in photo-realistically static scenes, inspiring numerous efforts to facilitate volumetric videos. However, rendering dynamic and long-sequence radiance fields remains challenging due to the significant data required to represent volumetric videos. In this paper, we propose a novel end-to-end joint optimization scheme of dynamic NeRF representation and compression, called JointRF, thus achieving significantly improved quality and compression efficiency against the previous methods. Specifically, JointRF employs a compact residual feature grid and a coefficient feature grid to represent the dynamic NeRF. This representation handles large motions without compromising quality while concurrently diminishing temporal redundancy. We also introduce a sequential feature compression subnetwork to further reduce spatial-temporal redundancy. Finally, the representation and compression subnetworks are end-to-end trained combined within the JointRF. Extensive experiments demonstrate that JointRF can achieve superior compression performance across various datasets. Zihan Zheng, Houqiang Zhong, Qiang Hu 0003, Xiaoyun Zhang 0001, Li Song 0001, Ya Zhang 0002, Yanfeng Wang 0001 |
ICIP | 4 |
| 2024 | HPC: Hierarchical Progressive Coding Framework for Volumetric VideoabstractVolumetric video based on Neural Radiance Field (NeRF) holds vast potential for various 3D applications, but its substantial data volume poses significant challenges for compression and transmission. Current NeRF compression lacks the flexibility to adjust video quality and bitrate within a single model for various network and device capacities. To address these issues, we propose HPC, a novel hierarchical progressive volumetric video coding framework achieving variable bitrate using a single model. Specifically, HPC introduces a hierarchical representation with a multi-resolution residual radiance field to reduce temporal redundancy in long-duration sequences while simultaneously generating various levels of detail. Then, we propose an end-to-end progressive learning approach with a multi-rate-distortion loss function to jointly optimize both hierarchical representation and compression. Our HPC trained only once can realize multiple compression levels, while the current methods need to train multiple fixed-bitrate models for different rate-distortion (RD) tradeoffs. Extensive experiments demonstrate that HPC achieves flexible quality levels with variable bitrate by a single model and exhibits competitive RD performance, even outperforming fixed-bitrate models across various datasets. Zihan Zheng, Houqiang Zhong, Qiang Hu 0003, Xiaoyun Zhang 0001, Li Song 0001, Ya Zhang 0002, Yanfeng Wang 0001 |
ACM Multimedia | 4 |
| 2023 | Boost Video Frame Interpolation via Motion Adaptation
Haoning Wu 0002, Xiaoyun Zhang 0001, Weidi Xie, Ya Zhang 0002, Yanfeng Wang 0001 |
BMVC | 2 |
| 2023 | DR2: Diffusion-Based Robust Degradation Remover for Blind Face RestorationabstractBlind face restoration usually synthesizes degraded low-quality data with a pre-defined degradation model for training, while more complex cases could happen in the real world. This gap between the assumed and actual degradation hurts the restoration performance where artifacts are often observed in the output. However, it is expensive and infeasible to include every type of degradation to cover real-world cases in the training data. To tackle this robustness issue, we propose Diffusion-based Robust Degradation Remover (DR2) to first transform the degraded image to a coarse but degradation-invariant prediction, then employ an enhancement module to restore the coarse prediction to a high-quality image. By leveraging a well-performing denoising diffusion probabilistic model, our DR2 diffuses input images to a noisy status where various types of degradation give way to Gaussian noise, and then captures semantic information through iterative denoising steps. As a result, DR2 is robust against common degradation (e.g. blur, resize, noise and compression) and compatible with different designs of enhancement modules. Experiments in various settings show that our framework outperforms state-of-the-art methods on heavily degraded synthetic and real-world datasets. Zhixin Wang, Ziying Zhang, Xiaoyun Zhang 0001, Huangjie Zheng, Mingyuan Zhou, Ya Zhang 0002, Yanfeng Wang 0001 |
CVPR | 3 |
| 2023 | Open-vocabulary Object Segmentation with Diffusion ModelsabstractThe goal of this paper is to extract the visual-language correspondence from a pre-trained text-to-image diffusion model, in the form of segmentation map, i.e., simultaneously generating images and segmentation masks for the corresponding visual entities described in the text prompt. We make the following contributions: (i) we pair the existing Stable Diffusion model with a novel grounding module, that can be trained to align the visual and textual embedding space of the diffusion model with only a small number of object categories; (ii) we establish an automatic pipeline for constructing a dataset, that consists of {image, segmentation mask, text prompt} triplets, to train the proposed grounding module; (iii) we evaluate the performance of open-vocabulary grounding on images generated from the text-to-image diffusion model and show that the module can well segment the objects of categories beyond seen ones at training time, as shown in Fig. 1; (iv) we adopt the augmented diffusion model to build a synthetic semantic segmentation dataset, and show that, training a standard segmentation model on such dataset demonstrates competitive performance on the zero-shot segmentation (ZS3) benchmark, which opens up new opportunities for adopting the powerful diffusion model for discriminative tasks. Ziyi Li 0004, Qinye Zhou, Xiaoyun Zhang 0001, Ya Zhang 0002, Yanfeng Wang 0001, Weidi Xie |
ICCV | 3 |
| 2023 | Adaptive Mutual Supervision for Weakly-Supervised Temporal Action LocalizationabstractWeakly-supervised temporal action localization aims to localize actions from untrimmed long videos with only video-level category labels. Most previous methods ignore the incompleteness issue of Class Activation Sequences (CAS), suffering from trivial detection results. To tackle this issue, we propose a novel Adaptive Mutual Supervision (AMS) framework with two branches, where the base branch detects the most discriminative action regions, while the supplementary branch localizes the less discriminative action regions through an adaptive sampler. The sampler dynamically updates the inputs for the supplementary branch using a sampling weight sequence negatively correlated with the CAS from the base branch, thus encouraging the supplementary branch to localize the action regions underestimated by the base branch. To promote mutual enhancement between two branches, we further construct mutual location supervision. Each branch adopts the location pseudo-labels generated from the other branch as the localization supervision. By alternately optimizing two branches for multiple iterations, we progressively complete action regions. Extensive experiments on THUMOS14 and ActivityNet1.2 demonstrate that the proposed AMS method significantly outperforms state-of-the-art methods. Chen Ju, Peisen Zhao, Siheng Chen, Ya Zhang 0002, Xiaoyun Zhang 0001, Yanfeng Wang 0001, Qi Tian 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | A Simple Plugin for Transforming Images to Arbitrary Scales
Qinye Zhou, Ziyi Li 0004, Weidi Xie, Xiaoyun Zhang 0001, Yanfeng Wang 0001, Ya Zhang 0002 |
BMVC | 4 |
| 2022 | LAR-SR: A Local Autoregressive Model for Image Super-ResolutionabstractPrevious super-resolution (SR) approaches often formulate SR as a regression problem and pixel wise restoration, which leads to a blurry and unreal SR output. Recent works combine adversarial loss with pixel-wise loss to train a GAN-based model or introduce normalizing flows into SR problems to generate more realistic images. As another powerful generative approach, autoregressive (AR) model has not been noticed in low level tasks due to its limitation. Based on the fact that given the structural in-formation, the textural details in the natural images are locally related without long term dependency, in this paper we propose a novel autoregressive model-based SR approach, namely LAR-SR, which can efficiently generate realistic SR images using a novel local autoregressive (LAR) module. The proposed LAR module can sample all the patches of textural components in parallel, which greatly reduces the time consumption. In addition to high time efficiency, it is also able to leverage contextual information of pixels and can be optimized with a consistent loss. Experimental results on the widely-used datasets show that the proposed LAR-SR approach achieves superior performance on the vi-sual quality and quantitative metrics compared with other generative models such as GAN, Flow, and is competitive with the mixture generative model. Baisong Guo, Xiaoyun Zhang 0001, Haoning Wu 0002, Yu Wang 0027, Ya Zhang 0002, Yanfeng Wang 0001 |
CVPR | 2 |
| 2022 | Task Decoupled Framework for Reference-based Super-ResolutionabstractReference-based super-resolution(RefSR) has achieved impressive progress on the recovery of high-frequency details thanks to an additional reference high-resolution(HR) image input. Although the superiority compared with Single-Image Super-Resolution(SISR), existing RefSR methods easily result in the reference-underuse issue and the reference-misuse as shown in Fig. I. In this work, we deeply investigate the cause of the two issues and further propose a novel framework to mitigate them. Our studies find that the issues are mostly due to the improper coupled framework design of current methods. Those methods conduct the super-resolution task of the input low-resolution(LR) image and the texture transfer task from the reference image together in one module, easily introducing the interference between LR and reference features. Inspired by this finding, we propose a novel framework, which decouples the two tasks of RefSR, eliminating the interference between the LR image and the reference image. The super-resolution task upsamples the LR image leveraging only the LR image itself. The texture transfer task extracts and transfers abundant textures from the reference image to the coarsely upsampled result of the super-resolution task. Extensive experiments demonstrate clear improvements in both quantitative and qualitative evaluations over state-of-the-art methods. Xiaoyun Zhang 0001, Siheng Chen, Ya Zhang 0002, Yanfeng Wang 0001, Dazhi He |
CVPR | 2 |
| 2021 | CaT: Weakly Supervised Object Detection with Category TransferabstractA large gap exists between fully-supervised object detection and weakly-supervised object detection. To narrow this gap, some methods consider knowledge transfer from additional fully-supervised dataset. But these methods do not fully exploit discriminative category information in the fully-supervised dataset, thus causing low mAP. To solve this issue, we propose a novel category transfer framework for weakly supervised object detection. The intuition is to fully leverage both visually-discriminative and semantically-correlated category information in the fully-supervised dataset to enhance the object-classification ability of a weakly-supervised detector. To handle overlapping category transfer, we propose a double-supervision mean teacher to gather common category information and bridge the domain gap between two datasets. To handle non-overlapping category transfer, we propose a semantic graph convolutional network to promote the aggregation of semantic features between correlated categories. Experiments are conducted with Pascal VOC 2007 as the target weakly-supervised dataset and COCO as the source fully-supervised dataset. Our category transfer framework achieves 63.5% mAP and 80.3% CorLoc with 5 overlapping categories between two datasets, which outperforms the state-of-the-art methods. Codes are avaliable at https://github.com/MediaBrain-SJTU/CaT. Lianyu Du, Xiaoyun Zhang 0001, Siheng Chen, Ya Zhang 0002, Yanfeng Wang 0001 |
ICCV | 3 |
| 2021 | Unsupervised Segmentation Framework with Active Contour Models for Cine Cardiac MRIabstractDeep learning methods have made remarkable progress in medical image segmentation tasks, but these methods require enough labeled data, which tends to be difficult for medical tasks. To tack this issue, we propose an unsupervised segmentation framework by combining deep learning networks with the active contour model. We design an iterative loop process that the network can be trained with outputs from the active contour model and the active contour model can be initialized by coarse predictions from the network. In this way, our approach can train the segmentation networks iterative with no annotations but only one initialization for the active contour model at the beginning. We evaluate our approach in the task of cine cardiac MRI segmentation and get very competitive results. Lianyu Du, Xiaoyun Zhang 0001, Yu-Min Zhong, Ya Zhang 0002, Yanfeng Wang 0001 |
ICIP | 3 |
| 2021 | MEMC-Net: Motion Estimation and Motion Compensation Driven Neural Network for Video Interpolation and EnhancementabstractMotion estimation (ME) and motion compensation (MC) have been widely used for classical video frame interpolation systems over the past decades. Recently, a number of data-driven frame interpolation methods based on convolutional neural networks have been proposed. However, existing learning based methods typically estimate either flow or compensation kernels, thereby limiting performance on both computational efficiency and interpolation accuracy. In this work, we propose a motion estimation and compensation driven neural network for video frame interpolation. A novel adaptive warping layer is developed to integrate both optical flow and interpolation kernels to synthesize target frame pixels. This layer is fully differentiable such that both the flow and kernel estimation networks can be optimized jointly. The proposed model benefits from the advantages of motion estimation and compensation methods without using hand-crafted features. Compared to existing methods, our approach is computationally efficient and able to generate more visually appealing results. Furthermore, the proposed MEMC-Net architecture can be seamlessly adapted to several video enhancement tasks, e.g., super-resolution, denoising, and deblocking. Extensive quantitative and qualitative evaluations demonstrate that the proposed method performs favorably against the state-of-the-art video frame interpolation and enhancement algorithms on a wide range of datasets. Wenbo Bao, Wei-Sheng Lai, Xiaoyun Zhang 0001, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | An End-to-End Learning Framework for Video CompressionabstractTraditional video compression approaches build upon the hybrid coding framework with motion-compensated prediction and residual transform coding. In this paper, we propose the first end-to-end deep video compression framework to take advantage of both the classical compression architecture and the powerful non-linear representation ability of neural networks. Our framework employs pixel-wise motion information, which is learned from an optical flow network and further compressed by an auto-encoder network to save bits. The other compression components are also implemented by the well-designed networks for high efficiency. All the modules are jointly optimized by using the rate-distortion trade-off and can collaborate with each other. More importantly, the proposed deep video compression framework is very flexible and can be easily extended by using lightweight or advanced networks for higher speed or better efficiency. We also propose to introduce the adaptive quantization layer to reduce the number of parameters for variable bitrate coding. Comprehensive experimental results demonstrate the effectiveness of the proposed framework on the benchmark datasets. Guo Lu, Xiaoyun Zhang 0001, Wanli Ouyang, Li Chen 0021, Dong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Two-Stream Compare and Contrast Network for Vertebral Compression Fracture DiagnosisabstractDifferentiating Vertebral Compression Fractures (VCFs) associated with trauma and osteoporosis (benign VCFs) or those caused by metastatic cancer (malignant VCFs) is critically important for treatment decisions. So far, automatic VCFs diagnosis is solved in a two-step manner, i.e., first identify VCFs and then classify them into benign or malignant. In this paper, we explore to model VCFs diagnosis as a three-class classification problem, i.e., normal vertebrae, benign VCFs, and malignant VCFs. However, VCFs recognition and classification require very different features, and both tasks are characterized by high intra-class variation and high inter-class similarity. Moreover, the dataset is extremely class-imbalanced. To address the above challenges, we propose a novel Two-Stream Compare and Contrast Network (TSCCN) for VCFs diagnosis. This network consists of two streams, a recognition stream which learns to identify VCFs through comparing and contrasting between adjacent vertebrae, and a classification stream which compares and contrasts between intra-class and inter-class to learn features for fine-grained classification. The two streams are integrated via a learnable weight control module which adaptively sets their contribution. TSCCN is evaluated on a dataset consisting of 239 VCFs patients and achieves the average sensitivity and specificity of 92.56% and 96.29%, respectively. Shixiang Feng, Ya Zhang 0002, Xiaoyun Zhang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Boundary-Aware Supervoxel-Level Iteratively Refined Interactive 3D Image Segmentation With Multi-Agent Reinforcement LearningabstractInteractive segmentation has recently been explored to effectively and efficiently harvest high-quality segmentation masks by iteratively incorporating user hints. While iterative in nature, most existing interactive segmentation methods tend to ignore the dynamics of successive interactions and take each interaction independently. We here propose to model iterative interactive image segmentation with a Markov decision process (MDP) and solve it with reinforcement learning (RL) where each voxel is treated as an agent. Considering the large exploration space for voxel-wise prediction and the dependence among neighboring voxels for the segmentation tasks, multi-agent reinforcement learning is adopted, where the voxel-level policy is shared among agents. Considering that boundary voxels are more important for segmentation, we further introduce a boundary-aware reward, which consists of a global reward in the form of relative cross-entropy gain, to update the policy in a constrained direction, and a boundary reward in the form of relative weight, to emphasize the correctness of boundary predictions. To combine the advantages of different types of interactions, i. e., simple and efficient for point-clicking, and stable and robust for scribbles, we propose a supervoxel-clicking based interaction design. Experimental results on four benchmark datasets have shown that the proposed method significantly outperforms the state-of-the-arts, with the advantage of fewer interactions, higher accuracy, and enhanced robustness. Chaofan Ma, Qisen Xu, Xiangfeng Wang 0001, Bo Jin 0003, Xiaoyun Zhang 0001, Yanfeng Wang 0001, Ya Zhang 0002 |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Iteratively-Refined Interactive 3D Medical Image Segmentation With Multi-Agent Reinforcement LearningabstractExisting automatic 3D image segmentation methods usually fail to meet the clinic use. Many studies have explored an interactive strategy to improve the image segmentation performance by iteratively incorporating user hints. However, the dynamic process for successive interactions is largely ignored. We here propose to model the dynamic process of iterative interactive image segmentation as a Markov decision process (MDP) and solve it with reinforcement learning (RL). Unfortunately, it is intractable to use single-agent RL for voxel-wise prediction due to the large exploration space. To reduce the exploration space to a tractable size, we treat each voxel as an agent with a shared voxel-level behavior strategy so that it can be solved with multi-agent reinforcement learning. An additional advantage of this multi-agent model is to capture the dependency among voxels for segmentation task. Meanwhile, to enrich the information of previous segmentations, we reserve the prediction uncertainty in the state space of MDP and derive an adjustment action space leading to a more precise and finer segmentation. In addition, to improve the efficiency of exploration, we design a relative cross-entropy gain-based reward to update the policy in a constrained direction. Experimental results on various medical datasets have shown that our method significantly outperforms existing state-of-the-art methods, with the advantage of less interactions and a faster convergence. Xuan Liao, Wenhao Li 0001, Qisen Xu, Xiangfeng Wang 0001, Bo Jin 0003, Xiaoyun Zhang 0001, Yanfeng Wang 0001, Ya Zhang 0002 |
CVPR | 6 |
| 2020 | Content Adaptive and Error Propagation Aware Deep Video Compression
Guo Lu, Chunlei Cai, Xiaoyun Zhang 0001, Li Chen 0021, Wanli Ouyang, Dong Xu 0001 |
ECCV (2) | 3 |
| 2020 | Adversarial Text Image Super-Resolution using Sinkhorn DistanceabstractConvolutional neural network-based methods have demonstrated promising results for single image super-resolution. However, existing methods usually approach the problem on natural scenes rather than texts, whereas the latter can provide more informative messages to viewers. In this paper, instead of using the Lp-norm as the supervision metric, we propose a novel one for better preserving semantic information in text images. Our new metric combines optimal transport in a primal form with Sinkhorn distance defined in an adversarially learned feature space. Since the Sinkhorn distance measures the similarity between two features in terms of both feature components and spatial locations, our metric can maintain the spatial structure of texts during network optimization. Experimental results on text datasets show that our method performs favorably against state-of-the-art approaches in both quantitative and qualitative evaluations. We will publish the code, datasets, and models upon acceptance. Cong Geng, Li Chen 0021, Xiaoyun Zhang 0001 |
ICASSP | 3 |
| 2020 | Dual-Task Self-supervision for Cross-modality Domain Adaptation
Yingying Xue, Shixiang Feng, Ya Zhang 0002, Xiaoyun Zhang 0001, Yanfeng Wang 0001 |
MICCAI (1) | 4 |
| 2020 | Viewport-adaptive 360-degree video coding
Qiang Hu 0003, Jun Zhou 0007, Xiaoyun Zhang 0001, Zhiru Shi |
Multim. Tools Appl. | 3 |
| 2020 | End-to-End Optimized ROI Image CompressionabstractCompressing an image with more bits automatically allocated to the region of interest (ROI) than to the background can both protect key information and reduce substantial redundancy. This paper models ROI image compression as an optimization problem of minimizing a weighted sum of the rate of the image and distortion of the ROI. The traditional framework solves this problem by cascading ROI prediction and ROI coding, through which achieving the optimized solution is impossible. To improve coding performance, we propose a novel deep-learning-based unified framework that can achieve rate distortion optimization for ROI compression. Specifically, the proposed framework includes a pair of ROI encoder and decoder convolutional neural networks and a learned entropy codec. The encoder network simultaneously generates multiscale representations that support efficient rate allocation and an implicit ROI mask that guides rate allocation. The proposed framework can automatically complete ROI image compression, and it can be optimized from data in an end-to-end manner. To effectively train the framework by back propagation, we develop a soft-to-hard ROI prediction scheme to make the entire framework differential. To improve visual quality, we propose a hierarchical distortion loss function to protect both pixel-level fidelity for ROI and structural similarity for the entire image. The proposed framework is implemented in two scenarios: salient-target and face-target ROI compression. Comparative experiments demonstrate the advantages of the proposed framework over the traditional framework, including considerably better subjective visual quality, significantly higher objective ROI compression performance and execution efficiency. Chunlei Cai, Li Chen 0021, Xiaoyun Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Deep Non-Local Kalman Network for Video Compression Artifact ReductionabstractVideo compression algorithms are widely used to reduce the huge size of video data, but they also introduce unpleasant visual artifacts due to the lossy compression. In order to improve the quality of the compressed videos, we proposed a deep non-local Kalman network for compression artifact reduction. Specifically, the video restoration is modeled as a Kalman filtering procedure and the decoded frames can be restored from the proposed deep Kalman model. Instead of using the noisy previous decoded frames as temporal information, the less noisy previous restored frame is employed in a recursive way, which provides the potential to generate high quality restored frames. In the proposed framework, several deep neural networks are utilized to estimate the corresponding states in the Kalman filter and integrated together in the deep Kalman filtering network. More importantly, we also exploit the non-local prior information by incorporating the spatial and temporal non-local networks for better restoration. Our approach takes the advantages of both the model-based methods and learning-based methods, by combining the recursive nature of the Kalman model and powerful representation ability of neural networks. Extensive experimental results on the Vimeo-90k and HEVC benchmark datasets demonstrate the effectiveness of our proposed method. Guo Lu, Xiaoyun Zhang 0001, Wanli Ouyang, Dong Xu 0001, Li Chen 0021 |
IEEE Trans. Image Process. | 2 |
| 2019 | Depth-Aware Video Frame InterpolationabstractVideo frame interpolation aims to synthesize nonexistent frames in-between the original frames. While significant advances have been made from the recent deep convolutional neural networks, the quality of interpolation is often reduced due to large object motion or occlusion. In this work, we propose a video frame interpolation method which explicitly detects the occlusion by exploring the depth information. Specifically, we develop a depth-aware flow projection layer to synthesize intermediate flows that preferably sample closer objects than farther ones. In addition, we learn hierarchical features to gather contextual information from neighboring pixels. The proposed model then warps the input frames, depth maps, and contextual features based on the optical flow and local interpolation kernels for synthesizing the output frame. Our model is compact, efficient, and fully differentiable. Quantitative and qualitative results demonstrate that the proposed model performs favorably against state-of-the-art frame interpolation methods on a wide variety of datasets. The source code and pre-trained model are available at https://github.com/baowenbo/DAIN. Wenbo Bao, Wei-Sheng Lai, Chao Ma 0004, Xiaoyun Zhang 0001, Ming-Hsuan Yang 0001 |
CVPR | 4 |
| 2019 | DVC: An End-To-End Deep Video Compression FrameworkabstractConventional video compression approaches use the predictive coding architecture and encode the corresponding motion information and residual information. In this paper, taking advantage of both classical architecture in the conventional video compression method and the powerful non-linear representation ability of neural networks, we propose the first end-to-end video compression deep model that jointly optimizes all the components for video compression. Specifically, learning based optical flow estimation is utilized to obtain the motion information and reconstruct the current frames. Then we employ two auto-encoder style neural networks to compress the corresponding motion and residual information. All the modules are jointly learned through a single loss function, in which they collaborate with each other by considering the trade-off between reducing the number of compression bits and improving quality of the decoded video. Experimental results show that the proposed approach can outperform the widely used video coding standard H.264 in terms of PSNR and be even on par with the latest standard H.265 in terms of MS-SSIM. Code is released at https://github.com/GuoLusjtu/DVC. Guo Lu, Wanli Ouyang, Dong Xu 0001, Xiaoyun Zhang 0001, Chunlei Cai |
CVPR | 4 |
| 2019 | A Novel Deep Progressive Image Compression FrameworkabstractIn Internet applications, compressing the image without perceptually distinguishable distortions and loading the images without notable delays in the client end can significantly improve the user experience. Compressing the image at high bit rates can maintain the high quality of the decoded image but in cost of long transmitting and decoding time, resulting in bad user experience. The progressive coding scheme can resolve the conflict between the high quality requirement and the large loading delay. This paper proposes a novel efficient progressive image coding framework based on deep convolutional neural networks. The proposed framework is composed of a uniform encoder network and two progressive decoder networks. The encoder network decomposes the input image into two scales of representations, that can be transmitted and reconstructed progressively into a basic quality preview image and a high-quality image by two individual decoder networks respectively. All the networks are jointly learned when achieving the rate distortion optimization of both scales. Experiments results show that the proposed method has much better coding performance than the commercial codecs WebP and JPEG, which are commonly used in Internet applications. Meanwhile, the proposed codec consumes much less time to load the image compared with WebP. Chunlei Cai, Li Chen 0021, Xiaoyun Zhang 0001, Guo Lu |
PCS | 3 |
| 2019 | FPGA Based Video Transcoding System with 2K-4K Super-Resolution ConversionabstractWe present a FPGA-based system supporting video stream transcoding with 2k full high-definition (FHD) video to 4k ultra high-definition (UHD) video super- resolution(SR) conversion. Our system focuses on building a functional pipeline with convolutional neural network (CNN) accelerator and real-time video codec unit for converting H.264 video stream to H.265/HEVC video stream. The overall video processing system can be used as an important plug-in module in the video streaming network to improve the video stream service quality. Yuzhuo Wei, Li Chen 0021, Rong Xie 0004, Li Song 0001, Xiaoyun Zhang 0001 |
VCIP | 5 |
| 2019 | Efficient Variable Rate Image Compression With Multi-Scale Decomposition NetworkabstractWhile deep learning image compression methods have shown an impressive coding performance, most of them output a single-optimized-compression rate using a trained-specific network. However, in practice, it is essential to support the variable rate compression or meet a target rate with a high-coding performance. This paper proposes a novel image compression method, making it possible for a single convolutional neural network (CNN) model to generate the variable rate efficiently with an optimized rate-distortion (RD) performance. The method consists of CNN-based multi-scale decomposition transform and content adaptive rate allocation. Specifically, the transform network is learned to decompose the input image into several scales of representations while optimizing the RD performance for all scales. Rate allocation algorithms for two typical scenarios are provided to determine the optimal scale of each image block for a given target rate or quality factor. For a target rate, the allocation is adaptive based on content complexity. In addition, for a target quality factor which indicates a tradeoff between the rate and the quality, the optimal scale is determined by minimizing the RD cost. The experimental results have shown that our method has outperformed the JPEG2000 and BPG standards with high efficiency and the state-of-the-art RD performance as measured by the multi-scale structural similarity index metric. Moreover, our method can strictly control the rate to generate the target compression result. Chunlei Cai, Li Chen 0021, Xiaoyun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | KalmanFlow 2.0: Efficient Video Optical Flow Estimation via Context-Aware Kalman FilteringabstractRecent studies on optical flow typically focus on the estimation of the single flow field in between a pair of images but pay little attention to the multiple consecutive flow fields in a longer video sequence. In this paper, we propose an efficient video optical flow estimation method by exploiting the temporal coherence and context dynamics under a Kalman filtering system. In this system, pixel's motion flow is first formulated as a second-order time-variant state vector and then optimally estimated according to the measurement and system noise levels within the system by maximum a posteriori criteria. Specifically, we evaluate the measurement noise according to the flow's temporal derivative, spatial gradient, and warping error. We determine the system noise based on the similarity of contextual information, which is represented by the compact features learned by pre-trained convolutional neural networks. The context-aware Kalman filtering helps improve the robustness of our method against abrupt change of light and occlusion/dis-occlusion in complicated scenes. The experimental results and analyses on the MPI Sintel, Monkaa, and Driving video datasets demonstrate that the proposed method performs favorably against the state-of-the-art approaches. Wenbo Bao, Xiaoyun Zhang 0001, Li Chen 0021 |
IEEE Trans. Image Process. | 2 |
| 2018 | Deep Kalman Filtering Network for Video Compression Artifact Reduction
Guo Lu, Wanli Ouyang, Dong Xu 0001, Xiaoyun Zhang 0001, Ming-Ting Sun |
ECCV (14) | 4 |
| 2018 | Rcdfnn: Robust Change Detection Based on Convolutional Fusion Neural NetworkabstractVideo change detection, which plays an important role in computer vision, is far from being well resolved due to the complexity of diverse scenes in real world. Most of the current methods are designed based on hand-crafted features and perform well in some certain scenes but may fail on others. This paper puts up forward a deep learning based method to automatically fuse multiple basic detections into an optimal one. Specifically, a convolutional fusion neural network is designed to obtain an adaptive fusion strategy based on features extracted from video content. Limited by the amount of available labeled dataset for change detection, this paper leverages an extractor that well trained on external dataset to improve generalization. Experiments show that the proposed method generates state-of-the-art result compared with nine recent outstanding algorithms and it performs well for diverse scenarios such as dynamic background, camera jitter and night videos. Chunlei Cai, Li Chen 0021, Xiaoyun Zhang 0001 |
ICASSP | 4 |
| 2018 | KalmanFlow: Efficient Kalman Filtering for Video Optical FlowabstractThis paper proposes an efficient optical flow filtering method for video sequences. Motivated by the observation that motions in videos have strong temporal coherence, we use Kalman filtering to exploit this characteristic for more accurate flow fields. In the proposed system, pixel's motion flow is formulated as a time-variant state vector and optimally estimated by Kalman filter according to the noise level, which is evaluated using flow's temporal derivative, spatial gradient and matching error. Experiments on MPI Sintel video dataset demonstrate that the temporal coherence employed during Kalman filtering has the advantage of more consistent results, and can contribute to the state-of-the-art methods. Wenbo Bao, Xiaoyun Zhang 0001, Li Chen 0021 |
ICIP | 2 |
| 2018 | Enhancing HEVC Compressed Videos with a Partition-Masked Convolutional Neural NetworkabstractIn this paper, we propose a partition-masked Convolution Neural Network (CNN) to achieve compressed-video enhancement for the state-of-the-art coding standard, High Efficiency Video Coding (HECV). More precisely, our method utilizes the partition information produced by the encoder to guide the quality enhancement process. In contrast to existing CNN-based approaches, which only take the decoded frame as the input to the CNN, the proposed approach considers the coding unit (CU) size information and combines it with the distorted decoded frame such that the degradation introduced by HEVC is reduced more efficiently. Experimental results show that our approach leads to over 9.76% BD-rate saving on benchmark sequences, which achieves the state-of-the-art performance. Xiaoyi He, Qiang Hu 0003, Xiaoyun Zhang 0001, Weiyao Lin, Xintong Han |
ICIP | 3 |
| 2018 | A Wavelet-based Learning for Face Hallucination with Loop ArchitectureabstractFace hallucination is a specific super-resolution problem which aims to generate high-resolution(HR) faces from low-resolution(LR) input. Recently, deep learning methods have been widely applied in single-image super resolution. Considering face images have great similarities in both pixel value and global structure, we propose a wavelet-based deep learning method with loop architecture for face hallucination. In contrast to existing wavelet-based methods that generate wavelet coefficients independently without considering relationships between them, we propose a three-stage method with loop architecture. This alternately updated loop structure explores the statistical relationships among wavelet coefficients and has a maximum use of information flow with a small number of parameters. Because of multi-resolution property of wavelet transform, we adopt a mixed input strategy to train images with different sizes to realize multi-scale face hallucination without retraining and adding extra sub-networks. Experiments demonstrate that our method can get a robust performance with multi-scale face hallucination. Cong Geng, Li Chen 0021, Xiaoyun Zhang 0001 |
VCIP | 3 |
| 2018 | High-Order Model and Dynamic Filtering for Frame Rate Up-ConversionabstractThis paper proposes a novel frame rate up-conversion method through high-order model and dynamic filtering (HOMDF) for video pixels. Unlike the constant brightness and linear motion assumptions in traditional methods, the intensity and position of the video pixels are both modeled with high-order polynomials in terms of time. Then, the key problem of our method is to estimate the polynomial coefficients that represent the pixel's intensity variation, velocity, and acceleration. We propose to solve it with two energy objectives: one minimizes the auto-regressive prediction error of intensity variation by its past samples, and the other minimizes video frame's reconstruction error along the motion trajectory. To efficiently address the optimization problem for these coefficients, we propose the dynamic filtering solution inspired by video's temporal coherence. The optimal estimation of these coefficients is reformulated into a dynamic fusion of the prior estimate from pixel's temporal predecessor and the maximum likelihood estimate from current new observation. Finally, frame rate up-conversion is implemented using motion-compensated interpolation by pixel-wise intensity variation and motion trajectory. Benefited from the advanced model and dynamic filtering, the interpolated frame has much better visual quality. Extensive experiments on the natural and synthesized videos demonstrate the superiority of HOMDF over the state-of-the-art methods in both subjective and objective comparisons. Wenbo Bao, Xiaoyun Zhang 0001, Li Chen 0021, Lianghui Ding |
IEEE Trans. Image Process. | 2 |
| 2018 | Novel Integration of Frame Rate Up Conversion and HEVC Coding Based on Rate-Distortion OptimizationabstractFrame rate up conversion (FRUC) can improve the visual quality by interpolating new intermediate frames. However, high frame rate videos by FRUC are confronted with more bitrate consumption or annoying artifacts of interpolated frames. In this paper, a novel integration framework of FRUC and high efficiency video coding (HEVC) is proposed based on rate-distortion optimization, and the interpolated frames can be reconstructed at encoder side with low bitrate cost and high visual quality. First, joint motion estimation (JME) algorithm is proposed to obtain robust motion vectors, which are shared between FRUC and video coding. What's more, JME is embedded into the coding loop and employs the original motion search strategy in HEVC coding. Then, the frame interpolation is formulated as a rate-distortion optimization problem, where both the coding bitrate consumption and visual quality are taken into account. Due to the absence of original frames, the distortion model for interpolated frames is established according to the motion vector reliability and coding quantization error. Experimental results demonstrate that the proposed framework can achieve 21% ~ 42% reduction in BDBR, when compared with the traditional methods of FRUC cascaded with coding. Guo Lu, Xiaoyun Zhang 0001, Li Chen 0021 |
IEEE Trans. Image Process. | 2 |
| 2017 | Iterative convolutional neural network for noisy image super-resolutionabstractImages captured by camera tend to be noisy and their qualities are often deteriorated in super-resolution. In this paper, we propose an end-to-end convolutional neural network to generate denoised, high-resolution image directly from its noisy, low-resolution counterpart. To preserve textures and eliminate noises simultaneously, the network is organized into an iterative structure for the recovery of high-quality image step by step. Each step of the structure is aimed to learn a better result with reference of its predecessor's output. Experiments show that our method is able to produce more desirable highresolution images in both objective and subjective evaluations comparing to conventional ones as well as non-iterative network based one. Wenbo Bao, Xiaoyun Zhang 0001, Shangpeng Yan |
ICIP | 2 |
| 2017 | Spatiotemporal salient object detection based on distance transform and energy optimization
Bing Yang 0003, Xiaoyun Zhang 0001, Li Chen 0021 |
Neurocomputing | 2 |
| 2017 | Edge guided salient object detection
Bing Yang 0003, Xiaoyun Zhang 0001, Li Chen 0021, Hua Yang 0001 |
Neurocomputing | 2 |
| 2016 | Principal components analysis-based visual saliency detectionabstractIn this paper, a novel patch-wise saliency detection algorithm is proposed based on Principal Component Analysis (PCA). As a powerful statistical procedure in data analysis, PCA are fully exploited to convert color space and produce compact patch representation. Specifically, images are first converted to linearly uncorrelated channels and divided into non-overlapped patches. Then the patches are represented by the coefficients of principal components using PCA analysis. Based on the compact representation of patches, two types of distinctiveness are introduced: center-surround contrast and global rarity. Experimental results demonstrate that the PCA-based color space conversion and patch representation can improve the accuracy of human fixations prediction, and the proposed algorithm outperforms the mainstream algorithms on predicting human fixations. Bing Yang 0003, Xiaoyun Zhang 0001, Jing Liu 0002, Li Chen 0021 |
ICASSP | 2 |
| 2016 | Frame rate up-conversion based on motion-region segmentationabstractThe key problem of frame rate up-conversion (FRUC) is to obtain true motion vectors (MV), especially for the motion boundaries. In this paper, we propose a novel FRUC algorithm based on motion-region segmentation. According to region's temporal consistency, motion-regions are determined by a categorization of detected feature points' true MVs. Then, constrained by MV's spatial smoothness within a region, true motions are propagated to the entire frame. This motion-region segmentation based method achieves truthful motion vector field and preferable interpolated frames. Experiments show that comparing to the state-of-art methods, the proposed algorithm produces videos with better quality in terms of objective and subjective evaluation. Wenbo Bao, Xiaoyun Zhang 0001, Li Chen 0021 |
VCIP | 2 |
| 2016 | Neyman-Pearson-Based Early Mode Decision for HEVC EncodingabstractThe high efficiency video coding (HEVC) standard has highly improved the coding efficiency by adopting hierarchical structures of coding unit (CU), prediction unit (PU), and transform unit (TU). However, enormous computational complexity is introduced due to the recursive rate-distortion optimization (RDO) process on all CUs, PUs and TUs. In this paper, we propose a fast and efficient mode decision algorithm based on the Neyman-Pearson rule, which consists of early SKIP mode decision and fast CU size decision. First, the early mode decision is modeled as a binary classification problem of SKIP/non-SKIP or split/unsplit. The Neyman-Pearson-based rule is employed to balance the rate-distortion (RD) performance loss and the complexity reduction by minimizing the missed detection with a constrained incorrect decision rate. A nonparametric density estimation scheme is also developed to calculate the likelihood function of the statistical parameters. Furthermore, an online training scheme is employed to periodically update the probability density distributions for different quantization parameters (QPs) and CU depth levels. The experimental results show that the proposed overall algorithm can save 65% and 58% computational complexity on average with a 1.29% and 1.08% Bjontegaard Delta bitrate (BDBR) increase for various test sequences under random access and low delay P conditions, respectively. The proposed overall scheme also has the advantage that it can make the trade-off between the RD performance and time saving by setting different values for the incorrect decision rate. Qiang Hu 0003, Xiaoyun Zhang 0001, Zhiru Shi |
IEEE Trans. Multim. | 2 |
| 2015 | Early SKIP mode decision based on Bayesian model for HEVCabstractIn High Efficiency Video Coding (HEVC), SKIP mode is an efficient inter prediction tool with high coding performance but low complexity. In this paper, an early SKIP mode decision algorithm is proposed to accelerate the encoding process. The rate-distortion cost (RD-cost) of SKIP mode is used as the decision criterion. A nonparametric density estimation scheme is employed to partition the SKIP RD-cost distribution space into high distinction region (HDR) and low distinction region (LDR). For a given Coding Unit (CU), if the SKIP RD-cost falls into the HDR, the SKIP mode is directly selected as the optimal mode. If the RD-cost maps into LDR, a Bayes risk minimization rule is adopted to ensure RD performance. The statistical parameters are updated according to different QPs and CU depths. Experimental results show that the proposed algorithm can reduce 47% of encoding time with only 0.34% of BD-bitrate increase on average. Qiang Hu 0003, Zhiru Shi, Xiaoyun Zhang 0001 |
VCIP | 3 |
| 2015 | Efficient SAO coding algorithm for x265 encoderabstractx265 is an open-source HEVC encoder which utilizes many optimization techniques and thus has fast coding speed as well as excellent coding performance. As a key technique in HEVC, sample adaptive offset (SAO) is also adopted by x265 to reduce the distortion between reconstructed frames and original ones. Although x265 is implemented in parallel with many speed up techniques, less efforts have been focused on the reduction of SAO related computation, which makes SAO become the speed bottleneck due to its great computational complexity. In this paper, we first thoroughly investigate and analyze the implementation and complexity of SAO in x265. Then, we propose a fast algorithm to early terminate SAO process based on inter prediction mode, CTU spatial-domain correlations, and relations between luma and chroma. Experimental results show that the proposed algorithm can save 72.2% SAO processing time with only 0.52% BDBR increase or 0.014 dB BDPSNR loss which is negligible. Shibo Yin, Xiaoyun Zhang 0001 |
VCIP | 2 |
| 2014 | Analysis and optimization of x265 encoderabstractx265 is an open-source encoder project which aims to deliver the world's fastest and most computationally efficient HEVC encoder. Although x265 has been developed efficiently with many optimization techniques, it is still not able to encode HD videos in real time even at its faster setting. In this paper, we deeply investigate the encoding framework and computational complexity of x265, and find that RDO process is the most time consuming part. Then, an efficient prediction scheme is proposed which includes decreasing the number of RDO times, early skip detection and fast intra mode decision. Experimental results show that the proposed method improves the speed of x265 from 19.86fps to 37.76fps for HD test sequences, i.e., 47.44% complexity reduction, with only 1.37% BDBR coding performance loss. Qiang Hu 0003, Xiaoyun Zhang 0001, Jun Sun 0005 |
VCIP | 2 |
| 2014 | Effective H.264/AVC to HEVC transcoder based on prediction homogeneityabstractThe new video coding standard, High Efficiency Video Coding (HEVC), has been established to succeed the widely used H.264/AVC standard. However, an enormous amount of legacy content is encoded with H.264/AVC. This makes high performance AVC to HEVC transcoding in great need. This paper presents a fast transcoding algorithm based on residual and motion information extracted from H.264 decoder. By exploiting these side information, regions' homogeneity characteristic are analysed. An efficient coding unit (CU) and prediction unit (PU) mode decision strategy is proposed combing regions' prediction homogeneity and current encoding information. The experimental results show that the proposed transcoding scheme can save up to 55% of encoding time with negligible loss of coding efficiency, when compared to that of the full decoding and full encoding transcoder. Feiyang Zheng, Zhiru Shi, Xiaoyun Zhang 0001 |
VCIP | 3 |
| 2013 | Occlusion handling frame rate up-conversionabstractMotion-compensated frame interpolation (MCFI) is a technique used extensively to enhance the temporal resolution of video sequences. In order to obtain a high quality interpolation, the motion vector field (MVF) between frames must be well-estimated. However, many current techniques for determining the MVF are prone to errors in occlusion regions. In this work, we propose an improved algorithm for improving the quality of MCFI by restoring the unreliable MVF and pixels in occlusion regions. We first utilize a dual motion estimation (DME) scheme which performs better in occlusion regions. Occlusion regions are determined by the ratio of two directional matching errors. Then, MVs in occlusion regions are refined using an orientation-based refinement (OBR) method, which promotes occluded MVs with its orthogonal neighboring MVF. Finally, regional blending (RB) is proposed to restore the unreliable pixels in occlusion regions for further error concealment. Experimental results demonstrate that the proposed algorithm provides a better quality than previous benchmark frame rate up-conversion (FRUC) methods both objectively and subjectively. Li Chen 0021, Xiaoyun Zhang 0001 |
ICASSP | 4 |
| 2013 | Effective early termination using adaptive search order for frame rate up-conversionabstractMotion Estimation (ME) palys a crucial part in frame rate up-conversion (FRUC) and video compression. One popular chip estimator, 3-D Recursive Search (3DRS), has achieved success in nowadays high definition televisions(HDTV). However, for future ultra high definition television (UHDTV) applications, the amount of computation increases exponentially. Therefore, there is an urgent need for further computational reduction. In this paper, we propose an effective early termination (ET) technique to reduce the calculation of absolute difference (AD), which is the most computation-intensive process in ME. First, motion vector (MV) candidates are examined in an adaptive search order (ASO) of their reliability to be true estimator. Then, the ET method dynamically changes the threshold for different motions to terminate accurately among reliable candidates. From the simulation, we can achieve an average computation reduction by 42% (up to 65%) with a slight PSNR degradation of 0.16dB on average. Li Chen 0021, Xiaoyun Zhang 0001 |
ISCAS | 4 |
| 2013 | A new Local-Main-Gradient-Orientation HOG and contour differences based algorithm for object classificationabstractThis paper presents a new algorithm to better classify objects in videos. In our case, the objects are cars, vans, and people on the roads. First, in order to extract the moving objects more precisely, we have proposed a method for foreground extraction based on the contour differences between the video frame and the background image. Second, after we got the integrated moving object, we have proposed a new algorithm to extract better features from the object. The new algorithm is based on two extended Histogram of Oriented Gradient (HOG) descriptor. We have improved HOG in two aspects: (a) selecting the gradient information from the moving objects and discarding the background gradient; (b) weighting every bin of gradient orientation histogram according to their significance within predefined area, in order to emphasize the important gradient information. We obtained Contour-Difference HOG (CD-HOG) from the first extension and Local-Main-Gradient-Orientation HOG (LMGO-HOG) from the second extended HOG. These extensions can cope with the cluttered background and make the features more distinguishable. Each of the extended HOG descriptors can produce a satisfying performance separately and an even better one if they are applied in cascade. From extensive evaluations, we showed the wonderful performance of our algorithm, and the accuracy rate of 94.04% can be achieved in some cases. Xiaoqiong Su, Weiyao Lin, Xiaozhen Zheng, Xintong Han, Hang Chu, Xiaoyun Zhang 0001 |
ISCAS | 6 |
| 2012 | Principal Components Analysis-Based Edge-Directed Image InterpolationabstractThis paper presents an edge-directed, noniterative image interpolation algorithm. In the proposed algorithm, the gradient directions are explicitly estimated with a statistical-based approach. The local dominant gradient directions are obtained by using principal components analysis (PCA) on the four nearest gradients. The angles of the whole gradient plane are divided into four parts, and each gradient direction falls into one part. Then we implement the interpolation with one-dimention (1-D) cubic convolution interpolation perpendicular to the gradient direction. Compared to the state of-the-art interpolation methods, simulation results show that the proposed PCA-based edge-directed interpolation method preserves edges well while maintaining a high PSNR value. Bing Yang 0003, Xiaoyun Zhang 0001 |
ICME | 3 |