VLDB 2026 Research / reviewers in the wild / expert
Rong Xie 0004
dblp:89/4026-4
· DBLP profile ↗
90ranked-venue papers
0as first author
54since 2021 · last 2026
0000-0002-8261-5337ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 79 · 48 since 2021Artificial intelligence and machine learning · 15 · 13 since 2021Computer networks · 6 · 5 since 2021Systems, architecture and hardware · 5 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Volumetric Video on Demand System Based on Scalable Spacetime Gaussian Splatting
Jun Xu 0040, Bingcong Lu, Rong Xie 0004, Li Song 0001 |
ISCAS | 5 |
| 2026 | Unified Multimodal Retrieval Framework for Multimodal RAG
Tianyi Feng, Ruiyan Wang, Fei Huang 0002, Zhengxue Cheng, Rong Xie 0004, Li Song 0001 |
PAKDD (4) | 7 |
| 2026 | OmniScaleSR: Unleashing Scale-Controlled Diffusion Prior for Faithful and Realistic Arbitrary-Scale Image Super-ResolutionabstractArbitrary-scale super-resolution (ASSR) overcomes the limitation of traditional super-resolution (SR) that works only at a fixed scale (e.g., ×4), enabling a single model to achieve arbitrary-scale SR. Most ASSR methods explicitly incorporate implicit neural representation (INR) to achieve ASSR, but INR’s inherently regression-driven feature extraction and aggregation nature restricts their capacity to synthesize meticulous details, leading to low realism. Recently, diffusion-based realistic image super-resolution (Real-ISR) methods leverage the pre-trained diffusion prior and have shown promising results at ×4 scale. We find that they could also achieve ASSR because the powerful pre-trained diffusion prior implicitly employs SR scale adaptation by encouraging the model to always generate high-realism images. However, due to the lack of explicit SR scale controls, the model fails to effectively manage the diffusion behavior according to different SR scales, causing either excessive hallucination or blurry results, especially for ultra-high magnification. To address these limitations, we proposeOmniScaleSR, a novel diffusion-based realistic arbitrary-scale super-resolution (Real-ASSR) method to achieve both high fidelity and high-realism ASSR. We introduce explicit diffusion-native SR scale controls, which could be elegantly coupled with the implicit scale adaptation, unleashing scale-controlled diffusion prior to dynamically managing the diffusion behavior in a content- and scale-aware manner. Furthermore, we incorporate multi-domain fidelity enhancement designs to achieve more faithful reconstruction. Extensive experiments on both bicubic degradation benchmarks and real-world datasets demonstrate that OmniScaleSR consistently outperforms state-of-the-art methods in terms of both fidelity and perceptual realism, with especially strong performance under high-magnification scenarios. Codes will be at https://github.com/chaixinning/OmniScaleSR. Xinning Chai, Zhengxue Cheng, Hengsheng Zhang, Yingsheng Qin, Yucai Yang, Rong Xie 0004, Li Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Diff-Restorer: Unleashing Visual Prompts for Diffusion-Based Universal Image RestorationabstractImage restoration aims to recover high-quality images from degraded observations, yet real-world degradations are complex, coupled, and difficult to model. Existing task-specific methods struggle to generalize beyond predefined degradation types, while recent all-in-one or prompt-based methods still face three key challenges: (1) they rely on task-specific training or fixed prompt pools, limiting adaptability to real-world and mixed degradations; (2) human-instruction or implicit-prompt mechanisms make them difficult to use in practice; and (3) they often fail to balance structural fidelity and perceptual realism. To address these issues, we propose Diff-Restorer, a diffusion-based universal image restoration framework that unifies diverse degradation handling within a single model. Diff-Restorer adaptively extracts decoupled visual prompts from a visual-language model (CLIP), including clear semantic and degradation embeddings. The clear semantic embeddings serve as content prompts to guide the diffusion model for generation, improving perceptual quality. The degradation embeddings as the task identifier modulate the Image-guided Control Module to generate structure control, ensuring faithfulness. Furthermore, we design a Task-aware Decoder to perform structural correction and convert the latent code to the pixel domain. Extensive experiments on various single, real-world, and mixed degradation tasks show that Diff-Restorer outperforms state-of-the-art methods in terms of generality, realism, and fidelity. Hengsheng Zhang, Xinning Chai, Zhengxue Cheng, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | PoseTalk: Exploring Text- and Audio-Based Pose Control for One-Shot Talking Face GenerationabstractAlthough audio-driven talking face generation has witnessed significant advances in recent years, two problems remain to be solved. First, the existing models cannot control the long-term head actions as humans expect, because the audio can only provide short-term cues, such as the rhythms and sentiments, for head movements. Second, generating long-term head poses and ensuring accurate lip motions remain challenging due to the difficulty in harnessing the optimization process for large-scale head movements and small-scale mouth motions. In this study, we propose a novel method to address these issues. First, to alleviate the limitations of audio conditions, we propose a Pose Latent Diffusion (PLD) model to generate head motions from two kinds of input modalities: the input audio and user-controlled text prompts. The audio provides short-term rhythm correspondence with the head movements, while the text prompts describe the long-term semantics of head motions. Second, we propose a refinement-based learning strategy to synthesize head movements and accurate lip motions using two cascaded networks, namely CoarseNet and RefineNet. The CoarseNet estimates coarse global motions to produce animated images with changed poses, and the RefineNet progressively estimates finer lip motions from low to high resolutions, yielding improved lip-synchronization performance. Experiments demonstrate that our method can achieve better pose diversity and realness compared to audio-based pose generation baselines, and our video generator model outperforms state-of-the-art methods in synthesizing natural head motions. Projects and demos are available at https://junleen.github.io/projects/posetalk . Jun Ling, Rong Xie 0004, Li Song 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2026 | A Hybrid Scheme for Face Video CompressionabstractWith the rapid development of social media, the amount of face video data has grown rapidly, making face video compression a hot research topic. Traditional video coding techniques do not discriminate video content and compress all videos in the same way, while talking head video compression should have more potential. Existing generative compression methods mostly adopt static reference frames, resulting in a decrease in fidelity caused by dynamic background or large pose change. In this article, we propose a hybrid compression scheme for face videos which combines traditional coding with generative compression. On the one hand, we sample and encode key frames with traditional codecs to provide dynamic reference frames which contain real-time background and motion information. On the other hand, we devise a deep video generation model to synthesize smooth video frames according to the extracted sparse keypoints. Combining the pixel-level recovery capability of traditional coding with the detail generation capability of deep generative models, our proposed hybrid scheme is able to implement high-fidelity face video compression at low bitrate in real time. Additionally, we also devise a Portrait Recovery module to recover the low-quality key frames, improving the reconstruction quality in low-bitrate scenarios. Extensive experiments show that our method has advantages over traditional codecs and existing generative compression methods in terms of both rate-distortion performance and coding complexity. Anni Tang, Zhiyu Zhang 0010, Jun Ling, Rong Xie 0004, Li Song 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Semantic and Temporal Integration in Latent Diffusion Space for High-Fidelity Video Super-ResolutionabstractRecent advancements in video super-resolution (VSR) models have demonstrated impressive results in enhancing low-resolution videos. However, due to limitations in adequately controlling the generation process, achieving high fidelity alignment with the low-resolution input while maintaining temporal consistency across frames remains a significant challenge. In this work, we propose Semantic and Temporal Guided Video Super-Resolution (SeTe-VSR), a novel approach that incorporates both semantic and temporal-spatio guidance in the latent diffusion space to address these challenges. By incorporating high-level semantic information and integrating spatial and temporal information, our approach achieves a seamless balance between recovering intricate details and ensuring temporal coherence. Our method not only preserves high-reality visual content but also significantly enhances fidelity. Extensive experiments demonstrate that SeTe-VSR outperforms existing methods in terms of detail recovery and perceptual quality, highlighting its effectiveness for complex video super-resolution tasks. Xinning Chai, Zhengxue Cheng, Rong Xie 0004, Li Song 0001 |
ICME | 6 |
| 2025 | SemanticGarment: Semantic-Controlled Generation and Editing of 3D Gaussian Garmentsabstract3D digital garment generation and editing play a pivotal role in fashion design, virtual try-on, and gaming. Traditional methods struggle to meet the growing demand due to technical complexity and high resource costs. Learning-based approaches offer faster, more diverse garment synthesis based on specific requirements and reduce human efforts and time costs. However, they still face challenges such as inconsistent multi-view geometry or textures and heavy reliance on detailed garment topology and manual rigging. We propose SemanticGarment, a 3D Gaussian-based method that realizes high-fidelity 3D garment generation from text or image prompts and supports semantic-based interactive editing for flexible user customization. To ensure multi-view consistency and garment fitting, we propose to leverage structural human priors for the generative model by introducing a 3D semantic clothing model, which initializes the geometry structure and lays the groundwork for view-consistent garment generation and editing. Without the need to regenerate or rely on existing mesh templates, our approach allows for rapid and diverse modifications to existing Gaussians, either globally or within a local region. To address the artifacts caused by self-occlusion for garment reconstruction based on single image, we develop a self-occlusion optimization strategy to mitigate holes and artifacts that arise when directly animating self-occluded garments. Extensive experiments are conducted to demonstrate our superior performance in 3D garment generation and editing. Ruiyan Wang, Zhengxue Cheng, Zonghao Lin, Jun Ling, Yanru An, Rong Xie 0004, Li Song 0001 |
ACM Multimedia | 7 |
| 2025 | PA-HOI: A Physics-Aware Human and Object Interaction Dataset
Ruiyan Wang, Lin Zuo, Zonghao Lin, Qiang Wang 0061, Zhengxue Cheng, Rong Xie 0004, Jun Ling, Li Song 0001 |
ACM Multimedia | 6 |
| 2025 | FreeInsert: Personalized Object Insertion with Geometric and Style ControlabstractText-to-image diffusion models have made significant progress in image generation, allowing for effortless customized generation. However, existing image editing methods still face certain limitations when dealing with personalized image composition tasks. First, there is the issue of lack of geometric control over the inserted objects. Current methods are confined to 2D space and typically rely on textual instructions, making it challenging to maintain precise geometric control over the objects. Second, there is the challenge of style consistency. Existing methods often overlook the style consistency between the inserted object and the background, resulting in a lack of realism. In addition, the challenge of inserting objects into images without extensive training remains significant. To address these issues, we propose FreeInsert, a novel training-free framework that customizes object insertion into arbitrary scenes by leveraging 3D geometric information. Benefiting from the advances in existing 3D generation models, we first convert the 2D object into 3D, perform interactive editing at the 3D level, and then re-render it into a 2D image from a specified view. This process introduces geometric controls such as shape or view. The rendered image, serving as geometric control, is combined with style and content control achieved through diffusion adapters, ultimately producing geometrically controlled, style-consistent edited images via the diffusion model. Rong Xie 0004, Li Song 0001 |
ACM Multimedia | 4 |
| 2025 | H3D-DGS: Exploring Heterogeneous 3D Motion Representation for Deformable 3D Gaussian SplattingabstractDynamic scene reconstruction poses a persistent challenge in 3D vision. Deformable 3D Gaussian Splatting has emerged as an effective method for this task, offering real-time rendering and high visual fidelity.
This approach decomposes a dynamic scene into a static representation in a canonical space and time-varying scene motion.
Scene motion is defined as the collective movement of all Gaussian points, and for compactness, existing approaches commonly adopt implicit neural fields or sparse control points.
However, these methods predominantly rely on gradient-based optimization for all motion information. Due to the high degree of freedom, they struggle to converge on real-world datasets exhibiting complex motion.
To preserve the compactness of motion representation and address convergence challenges, this paper proposes heterogeneous 3D control points, termed \textbf{H3D control points}, whose attributes are obtained using a hybrid strategy combining optical flow back-projection and gradient-based methods.
This design decouples directly observable motion components from those that are geometrically occluded.
Specifically, components of 3D motion that project onto the image plane are directly acquired via optical flow back projection, while unobservable portions are refined through gradient-based optimization.
Experiments on the Neu3DV and CMU-Panoptic datasets demonstrate that our method achieves superior performance over state-of-the-art deformable 3D Gaussian splatting techniques. Remarkably, our method converges within just 100 iterations and achieves a per-frame processing speed of 2 seconds on a single NVIDIA RTX 4070 GPU. Yunuo Chen 0002, Guo Lu, Cheems Wang, Qunshan Gu, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
NeurIPS | 6 |
| 2025 | CPIG: Controlling the Portrait Image Generation by Distilling 3D GAN's Latent DirectionsabstractABSTRACT Synthesis of 3D‐aware facial images from latent spaces has garnered significant attention in multimedia content generation due to its ability to model images with rich semantics and diverse appearances. However, existing methods often rely on labeled data or suffer from incomplete attribute control and ambiguous latent space semantics. This paper proposes an efficient semantic distillation method that learns attribute directions of pre‐trained 3D GAN models, without the supervised semantic labels. We consider the latent space of GAN models as the mixture of two featured subspaces, namely the geometry‐aware space and appearance‐aware space. Following this hypothesis, we define two sets of learnable latent bases and use linear composition to represent controllable geometry and appearance feature space, respectively. To learn semantic‐wise latent bases for attribute‐controllable image generation, we design a framework and propose a three‐staged training strategy, which optimizes the appearance‐aware and the geometry‐aware latent bases. With the two sets of latent bases, we obtain the combined latent vectors using different weights for those bases and synthesize images with specified attributes. Compared to existing methods, our approach eliminates the need for labeled data and enables more controllable attribute disentanglement while ensuring identity consistency, which can be directly applied to real‐world scenarios such as virtual avatars and augmented reality applications. Experiments demonstrate the effectiveness and insight of our approach in aiding a better understanding of the latent space of 3D GANs. Ruiyan Wang, Jun Ling, Rong Xie 0004, Li Song 0001 |
IET Image Process. | 3 |
| 2025 | SSP-IR: Semantic and Structure Priors for Diffusion-Based Realistic Image RestorationabstractRealistic image restoration is a crucial task in computer vision, and diffusion-based models for image restoration have garnered significant attention due to their ability to produce realistic results. Restoration can be seen as a controllable generation conditioning on priors. However, due to the severity of image degradation, existing diffusion-based restoration methods cannot fully exploit priors from low-quality images and still have many challenges in perceptual quality, semantic fidelity, and structure accuracy. Based on the challenges, we introduce a novel image restoration method, SSP-IR. Our approach aims to fully exploit semantic and structure priors from low-quality images to guide the diffusion model in generating semantically faithful and structurally accurate natural restoration results. Specifically, we integrate the visual comprehension capabilities of Multimodal Large Language Models (explicit) and the visual representations of the original image (implicit) to acquire accurate semantic prior. To extract degradation-independent structure prior, we introduce a Processor with RGB and FFT constraints to extract structure prior from the low-quality images, guiding the diffusion model and preventing the generation of unreasonable artifacts. Lastly, we employ a multi-level attention mechanism to integrate the acquired semantic and structure priors. The qualitative and quantitative results demonstrate that our method outperforms other state-of-the-art methods overall on both synthetic and real-world datasets. Our project page ishttps://zyhrainbow.github.io/projects/SSP-IR. Hengsheng Zhang, Zhengxue Cheng, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Implicit-Explicit Integrated Representations for Multi-View Video CompressionabstractWith the increasing consumption of 3D displays and virtual reality, multi-view video has become a promising format. However, its high resolution and multi-camera shooting result in a substantial increase in data volume, making storage and transmission a challenging task. To tackle these difficulties, we propose an implicit-explicit integrated representation for multi-view video compression. Specifically, we first use the explicit representation-based 2D video codec to encode one of the source views. Subsequently, we propose employing the implicit neural representation (INR)-based codec to encode the remaining views. The implicit codec takes the time and view index of multi-view video as coordinate input and generates the corresponding implicit reconstruction frames. To enhance the compressibility, we introduce a multi-level feature grid embedding and a fully convolutional architecture into the implicit codec. These components facilitate coordinate-feature and feature-RGB mapping, respectively. To further enhance the reconstruction quality from the INR codec, we leverage the high-quality reconstructed frames from the explicit codec to achieve inter-view compensation. Finally, the compensated results are fused with the implicit reconstructions from the INR to obtain the final reconstructed frames. Our proposed framework combines the strengths of both implicit neural representation and explicit 2D codec. Extensive experiments conducted on public datasets demonstrate that the proposed framework can achieve comparable or even superior performance to the latest multi-view video compression standard MIV and other INR-based schemes in terms of view compression and scene modeling. The source code can be found at https://github.com/zc-lynen/MV-IERV. Guo Lu, Rong Xie 0004, Li Song 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | Depth-Guided Robust and Fast Point Cloud Fusion NeRF for Sparse Input ViewsabstractNovel-view synthesis with sparse input views is important for real-world applications like AR/VR and autonomous driving. Recent methods have integrated depth information into NeRFs for sparse input synthesis, leveraging depth prior for geometric and spatial understanding. However, most existing works tend to overlook inaccuracies within depth maps and have low time efficiency. To address these issues, we propose a depth-guided robust and fast point cloud fusion NeRF for sparse inputs. We perceive radiance fields as an explicit voxel grid of features. A point cloud is constructed for each input view, characterized within the voxel grid using matrices and vectors. We accumulate the point cloud of each input view to construct the fused point cloud of the entire scene. Each voxel determines its density and appearance by referring to the point cloud of the entire scene. Through point cloud fusion and voxel grid fine-tuning, inaccuracies in depth values are refined or substituted by those from other views. Moreover, our method can achieve faster reconstruction and greater compactness through effective vector-matrix decomposition. Experimental results underline the superior performance and time efficiency of our approach compared to state-of-the-art baselines. Shuai Guo 0002, Qiuwen Wang, Yijie Gao, Rong Xie 0004, Li Song 0001 |
AAAI | 4 |
| 2024 | Disentangled Clothed Avatar Generation from Text Descriptions
Jionghao Wang, Yuan Liu 0025, Zhiyang Dou, Zhengming Yu, Yongqing Liang 0001, Cheng Lin 0001, Rong Xie 0004, Li Song 0001, Xin Li 0003, Wenping Wang 0001 |
ECCV (52) | 7 |
| 2024 | Hdrtvformer: Efficient Sdrtv-to-Hdrtv via Affine Transformation and Spatial-Aware TransformerabstractRecent works on reconstructing HDR videos in display format (HDRTV) suffer from high computational and memory requirements because they learn the SDRTV-to-HDRTV mapping directly in 4K resolution. This paper proposes an efficient SDRTV-to-HDRTV model (HDRTVFormer) that decomposes the HDRTV restoration into SDRTV-to-HDRTV Domain Mapping and HDRTV Refinement. SDRTV-to-HDRTV Domain Mapping is an affine transformation-based model that learns SDRTV-to-HDRTV affine coefficients in low-resolution space, achieving rapid processing times. To enhance the accuracy of the predicted affine coefficients, the model introduces global information-modulated feature extraction blocks and a detail guidance upsampling module. For HDRTV Refinement, we propose a spatial-aware Transformer to refine the luminance and color details. We modify the self-attention and feed-forward network of Transformer blocks to improve efficiency and feature representations. Experimental results have demonstrated that our method outperforms other state-of-the-art works in performance and efficiency. Hengsheng Zhang, Xinning Chai, Rong Xie 0004, Li Song 0001 |
ICASSP | 4 |
| 2024 | A New People-Object Interaction Dataset and NVS BenchmarksabstractRecently, NVS in human-object interaction scenes has received increasing attention. Existing human-object interaction datasets mainly consist of static data with limited views, offering only RGB images or videos, mostly containing interactions between a single person and objects. Moreover, these datasets exhibit complexities in lighting environments, poor synchronization, and low resolution, hindering high-quality human-object interaction studies. In this paper, we introduce a new people-object interaction dataset that comprises 38 series of 30-view multi-person or single-person RGB-D video sequences, accompanied by camera parameters, foreground masks, SMPL models, some point clouds, and mesh files. Video sequences are captured by 30 Kinect Azures, uniformly surrounding the scene, each in 4 K resolution 25 FPS, and lasting for 1~19 seconds. Meanwhile, we evaluate some SOTA NVS models on our dataset to establish the NVS benchmarks. We hope our work can inspire further research in human-object interaction. Shuai Guo 0002, Houqiang Zhong, Qiuwen Wang, Yijie Gao, Jiajing Yuan, Rong Xie 0004, Li Song 0001 |
ICIP | 8 |
| 2024 | Pioneer: Offline Reinforcement Learning based Bandwidth Estimation for Real-Time CommunicationabstractFor Real-time Communication (RTC), Bandwidth Estimation (BWE) is crucial for enhancing user Quality of Experience (QoE) by ensuring efficient bandwidth utilization and low latency. Recent advancements have shifted towards machine learning based algorithms, particularly online reinforcement leanring (RL), to dynamically infer future bandwidth using statistical data. However, challenges such as dependency on training settings, the necessity for extensive trial and error, and instability in complex state spaces hinder their efficacy. To address these limitations, we propose Pioneer, a novel offline RL framework for BWE in RTC systems. Unlike its predecessors, Pioneer eliminates the need for real-time environment interaction during training and achieves good performance through lightweight training. Our framework consists of a Trajectory Sampler for state information preprocessing and a Bandwidth Estimator based on offline RL model. Our test results on offline datasets show that Pioneer can achieve better performance than expert algorithms. We also tested Pioneer on online simulation platforms, and Pioneer can improve QoE by 9% compared to other offline algorithm, demonstrating good robustness. Bingcong Lu, Jun Xu 0040, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
MMSys | 4 |
| 2024 | Efficient Bitrate Ladder Construction for Per-Shot Adaptive EncodingabstractHTTP adaptive streaming (HAS) constructs bitrate ladders to deliver videos with the best possible quality under varying network conditions. Though per-shot content adaptive encoding (CAE) largely improves the compression efficiency by constructing the optimal bitrate ladder for each video shot, it suffers from excessive encoding complexity as all the points in the operating space (typically resolution × bitrate) need to be encoded and compared. To address this issue, this paper proposes an efficient bitrate ladder construction method that encodes only a subset of operating points, then uses curve fitting and inter-curve prediction to estimate other points’ RD performance. The proposed method enables low-complexity ladder construction even for high-dimension operating spaces that incorporate dimensions like encoding presets. Experiments show that this method can achieve RD performance comparable to the original per-shot CAE with only 42% encoding points. Even when minimizing the encoding points to 3.6% of the original CAE, it achieves 15% BD-Rate improvements compared to using the fixed bitrate ladder. Yan Zhao 0041, Zhengxue Cheng, Guo Lu, Rong Xie 0004, Li Song 0001 |
VCIP | 4 |
| 2024 | Depth-Guided Robust Point Cloud Fusion NeRF for Sparse Input ViewsabstractNovel-view synthesis with sparse input views is important for practical applications such as AR/VR and autonomous driving. Many works in this field have already integrated depth information into NeRF, utilizing depth priors for assistance in geometric and spatial understanding. However, most existing work tends to either overlook the inaccuracies in depth maps or only handle them roughly, limiting the effectiveness of the synthesis. To address this issue, we propose a depth-guided robust point cloud fusion NeRF for sparse input synthesis. We first construct a point cloud for each input view, with a novel point cloud representation based on learnable matrices and vectors. Then, through an additional lightweight scene fusion network, we fuse the point clouds from each input view to build a point cloud of the entire scene. By optimizing the point cloud representation and scene fusion network, inaccuracies in the depth map can be adjusted and refined, thereby achieving a more precise perception of the overall scene. Each voxel in the scene is determined by referencing the fused point cloud to establish its density and appearance. Experimental results demonstrate that our method outperforms state-of-the-art baselines. Shuai Guo 0002, Qiuwen Wang, Yijie Gao, Rong Xie 0004, Lin Li 0062, Li Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | A Character Position-Aware Compression Framework for Screen Text ImageabstractText patterns typically exhibit distinct boundaries and sparse color histograms. However, in current hybrid codec frameworks, the positions of coding units are often misaligned with the text patterns, resulting in prediction and color mapping tools consuming a large number of bits to indicate these patterns. Nowadays, some text detection and recognition methods have been proposed to accurately locate and analyze the text regions in screen images. Combined with these techniques, we propose a character position-aware compression framework for screen text image. On the encoder side, a low-complexity detection method is adopted to locate the text characters. Then it copies the detected characters to the position aligned with the coding unit (CU) grid to form a text layer. This text-layer representation can further increase the efficiency of existing screen content coding tools such as Intra Block Copy (IBC). Moreover, we design several compression tools based on this representation. We extend the two Motion Vector (MV) prediction modes: Adaptive Motion Vector Prediction (AMVP) and Merge. We modify the MV encoding syntax according to the layout characteristics of the text layer. We present a Gradient-guided In-loop Filter (GIF) to sharpen the text lines using a convolutional network. Experiments conducted on VVC reference software VTM all_intra configuration show that the proposed framework can achieve an average bitrate savings of 4.6% and 3.6% under the w/ GIF and w/o GIF versions, with a corresponding increase in CPU encoding complexity of 72% and 10%. Guo Lu, Huanbang Chen, Donghui Feng 0003, Shen Wang 0013, Yan Zhao 0041, Rong Xie 0004, Li Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Real-Time Free Viewpoint Video Synthesis System Based on DIBR and a Depth Estimation NetworkabstractDepth image-based rendering (DIBR) view synthesis is the most widely employed method in real-time FVV research. Despite recent progress, most DIBR-based FVV synthesis approaches are not sufficiently simple and effective in filling holes and artifacts. Additionally, they use RGB-D cameras, which are difficult to widely adopt or take considerable time to estimate high-quality depth images. This paper introduces a real-time FVV synthesis system based on DIBR and a depth estimation network. This system includes a 12-view synchronous camera system, a new multistage depth estimation network, a new GPU-accelerated DIBR algorithm, and a virtual view parameter generation method. This system provides the first real-time FVV solution for background-fixed fields based on DIBR and a depth estimation network. It can infer depth images for all camera views and synthesize any virtual view along the horizontal circular arc of the camera rig in real time. To our knowledge, we are the first to introduce background models and foreground masks and a refined multistage structure to address real-time high-quality depth estimation and DIBR FVV synthesis. We also build a high-quality multiview RGB-D synchronous dataset that has promising DIBR FVV synthesis performance to train and evaluate our system. The experimental results demonstrate the real-time and better performance of the proposed system. Shuai Guo 0002, Jingchuan Hu, Kai Zhou 0016, Jionghao Wang, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | ViCoFace: Learning Disentangled Latent Motion Representations for Visual-Consistent Face ReenactmentabstractUnsupervised face reenactment aims to animate a source image to imitate the motions of a target image while retaining the source portrait’s attributes like facial geometry, identity, hair texture, and background. While prior methods can extract the motion from the target image via compact representations (e.g., keypoints or latent motion bases [ 50 ]), they are not robust in predicting motions that are disentangled with portrait attributes, thus failing to preserve portrait attributes in the cross-subject reenactment. In this work, we propose an effective and cost-efficient face reenactment approach to address this issue. Our approach is highlighted by two major strengths. First, based on the theory of latent motion bases, we disentangle the full-head motion into two parts: the transferable motion and preservable motion and then compose the full motion representation using latent motions from the source image and the target image. Second, to optimize and learn disentangled motions, we introduce an efficient training framework, which features two training strategies: (1) a mixture training strategy that encompasses self-reenactment training and cross-subject training for better motion disentanglement and (2) a multi-path training strategy that improves the visual consistency of portrait attributes. Extensive experiments on widely used benchmarks demonstrate that our method exhibits a remarkable generalization ability compared to state-of-the-art baselines. Project and demos are available at https://junleen.github.io/projects/vicoface . Jun Ling, Anni Tang, Rong Xie 0004, Li Song 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Boosting Video Object Segmentation via Space-Time Correspondence LearningabstractCurrent top-leading solutions for video object segmentation (VOS) typically follow a matching-based regime: for each query frame, the segmentation mask is inferred according to its correspondence to previously processed and the first annotated frames. They simply exploit the supervisory signals from the groundtruth masks for learning mask prediction only, without posing any constraint on the space-time correspondence matching, which, however, is the fundamental building block of such regime. To alleviate this crucial yet commonly ignored issue, we devise a correspondence-aware training framework, which boosts matching-based VOS solutions by explicitly encouraging robust correspondence matching during network learning. Through comprehensively exploring the intrinsic coherence in videos on pixel and object levels, our algorithm reinforces the standard, fully supervised training of mask segmentation with label-free, contrastive correspondence learning. Without neither requiring extra annotation cost during training, nor causing speed delay during deployment, nor incurring architectural modification, our algorithm provides solid performance gains on four widely used benchmarks, i.e., DAVIS2016&2017, and YouTube-VOS2018&2019, on the top of famous matching-based VOS solutions. Liulei Li, Wenguan Wang, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
CVPR | 4 |
| 2023 | Dual-Head Fusion Network for Image EnhancementabstractImage enhancement algorithms have made great progress recently. However, most existing methods tend to construct a uniform enhancer for the color transformation of all pixels and ignore the local context information which is significant for photographs, causing unsatisfactory results. To solve these issues, we propose a novel dual-head fusion network for image enhancement, which synthetically considers both global scenario and local content information. Our network consists of four lightweight modules. We first develop a dual-head feature extraction module to extract the global condition vector and spatial context map. After that, we propose a context-aware retouching module and a global color rendering module to generate latent results. Finally, we employ the spatial attention based fusion module to adaptively aggregate the latent results. Experiments on public datasets show that our method consistently achieves the best results compared with SOTA methods both quantitatively and qualitatively. Hengsheng Zhang, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
ICASSP | 4 |
| 2023 | Divide and Conquer: a Two-Step Method for High Quality Face De-identification with Model ExplainabilityabstractFace de-identification involves concealing the true identity of a face while retaining other facial characteristics. Current target-generic methods typically disentangle identity features in the latent space, using adversarial training to balance privacy and utility. However, this pattern often leads to a trade-off between privacy and utility, and the latent space remains difficult to explain. To address these issues, we propose IDeudemon, which employs a "divide and conquer" strategy to protect identity and preserve utility step by step while maintaining good explainability. In Step I, we obfuscate the 3D disentangled ID code calculated by a parametric NeRF model to protect identity. In Step II, we incorporate visual similarity assistance and train a GAN with adjusted losses to preserve image utility. Thanks to the powerful 3D prior and delicate generative designs, our approach could protect the identity naturally, produce high quality details and is robust to different poses and expressions. Extensive experiments demonstrate that the proposed IDeudemon outperforms previous state-of-the-art methods. Yunqian Wen, Bo Liu 0001, Jingyi Cao, Rong Xie 0004, Li Song 0001 |
ICCV | 4 |
| 2023 | PACC: Perception Aware Congestion Control for Real-time CommunicationabstractDue to the network fluctuations, congestion control is indispensable to guarantee the quality of experience (QoE) for Real-Time Communication (RTC) users. This component adjusts the sending rate of media data, which determines the video encoding bitrate. However, existing control schemes either only focus on network numerical indicators or fail to adapt to various network environments. Logically, we propose PACC (Perception Aware Congestion Control) for RTC in this paper. Leveraging the convolutional neural network (CNN), we develop a quality sensor to infer the video quality increasing rate. Assisted with the variation trend analysis for user perception, PACC tunes the bitrate towards the direction of better QoE. Extensive tracedriven experiments demonstrate the effectiveness of PACC, which outperforms the existing landmark schemes by 8.2% to 32.4% and 6.8% to 18.0% in terms of transport and application layer QoE metrics, respectively. Bingcong Lu, Li Song 0001, Rong Xie 0004, Yanmei Liu, Ying Chen 0011 |
ICME | 4 |
| 2023 | 360-Degree Panorama Generation from Few Unregistered NFoV Imagesabstract360° panoramas are extensively utilized as environmental light sources in computer graphics. However, capturing a 360° × 180° panorama poses challenges due to the necessity of specialized and costly equipment, and additional human resources. Prior studies develop various learning-based generative methods to synthesize panoramas from a single Narrow Field-of-View (NFoV) image, but they are limited in alterable input patterns, generation quality, and controllability. To address these issues, we propose a novel pipeline called PanoDiff, which efficiently generates complete 360° panoramas using one or more unregistered NFoV images captured from arbitrary angles. Our approach has two primary components to overcome the limitations. Firstly, a two-stage angle prediction module to handle various numbers of NFoV inputs. Secondly, a novel latent diffusion-based panorama generation model uses incomplete panorama and text prompts as control signals and utilizes several geometric augmentation schemes to ensure geometric properties in generated panoramas. Experiments show that PanoDiff achieves state-of-the-art panoramic generation quality and high controllability, making it suitable for applications such as content editing. Jionghao Wang, Jun Ling, Rong Xie 0004, Li Song 0001 |
ACM Multimedia | 4 |
| 2023 | Achieving Privacy-Preserving Multi-View Consistency with Advanced 3D-Aware Face De-identificationabstractThe widespread application of face recognition technology has exacerbated privacy threats. Face de-identification is an effective means of protecting visual privacy by concealing identity information. While deep learning-based methods have greatly improved de-identification results, most existing algorithms rely on 2D generative models that struggle to produce identity-consistent results for multiple views. In this paper, we focus on identity disentanglement within the latest 3D-aware face generation model, and propose an advanced face de-identification framework that can be applied to various scenarios. Our proposed framework disentangles identity from other facial features, modifies only the former and generates the de-identified face using a 3D generator. This approach results in high-quality, identity-consistent de-identification that preserves other facial features. We demonstrate our approach on StyleNeRF, one of the most widely-used style-based neural radiation field models. Through extensive experiments, we demonstrate the effectiveness of our approach in achieving face de-identification both for a single image and group images with the same identity. Our work is a significant step forward in the field of face de-identification, opening up new possibilities for practical applications. Jingyi Cao, Bo Liu 0001, Yunqian Wen, Rong Xie 0004, Li Song 0001 |
MMAsia | 4 |
| 2023 | NeRF-SDP: Efficient Generalizable Neural Radiance Field with Scene Depth PerceptionabstractIn recent years, neural radiance fields have exhibited impressive performance in novel view synthesis. However, exploiting complex network structures to achieve generalizable NeRF usually results in inefficient rendering. Existing methods for accelerating rendering directly employ simpler inference networks or fewer sampling points, leading to unsatisfactory synthesis quality. To address the challenge of balancing rendering speed and quality in generalizable NeRF, we propose a novel framework, NeRF-SDP, which achieves both efficiency and high fidelity by introducing scene depth perception. We incorporate more scene information into the radiance field by using our proposed geometry feature extraction and depth-encoded ray transformer to improve the model’s inference capabilities with sparse points. With the aid of scene depth perception, NeRF-SDP can better understand the scene’s structure, thus better reconstructing the objects’ edges with significantly fewer artifacts. Experimental results demonstrate that NeRF-SDP achieves comparable synthesis quality to state-of-the-art methods while significantly improving rendering efficiency. Furthermore, ablation studies confirm that the depth-encoded ray transformer enhances the model’s robustness to varying numbers of sampling points. Qiuwen Wang, Shuai Guo 0002, Haoning Wu 0002, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
MMAsia | 4 |
| 2023 | High-Fidelity Free-View Talking Head Synthesis for Low-Bandwidth Video ConferenceabstractAs video conferencing becomes an indispensable part of human’s daliy life, how to achieve a high-fidelity calling experience under low bandwidth has been a popular and challenging issue. Deep generative models have great potential in low-bandwidth facial video compression due to the excellent generation capability based on abridged information. Nevertheless, exsiting deep generation-based compression methods tend to handle motion information in pure 2D or pseudo 3D space, causing facial distortion when large head poses are encountered. In this paper, we propose a 3D-aware high-fidelity facial video conferencing system based on a parameterized NeRF-based face model. Through the compression of the parameterized face model and the transmisstion of extracted facial parameters, we implement high-fidelity talking head synthesis for video conferencing at an ultra-low bitrate. Additionally, the 3D perception capability of the system allows for viewpoint control over the head, achieving higher interactivity and practicability. Extensive experiments verify the effectiveness of the proposed 3D-aware high-fidelity free-view facial video conferencing system. Zhiyu Zhang 0010, Anni Tang, Guo Lu, Rong Xie 0004, Li Song 0001 |
VCIP | 5 |
| 2023 | Deep Online Video Stabilization Using IMU SensorsabstractIn this paper, we propose a deep learning based sensor-driven method for online video stabilization. This method utilizes the Euler angles and acceleration values estimated from the gyroscope and accelerator to assist stable video reconstruction. We introduce two simple sub-networks for trajectory optimization. The first network exploits real unstable trajectories and camera acceleration values to detect shooting scenarios. This network also generates an attention mask to adaptively choose scenario-specific features. Then the second network predicts smooth camera paths based on real unstable trajectories using long short-term memory (LSTM) under the supervision of the above mask. The output of the trajectory optimization network is filtered with a two-step modification process to guarantee smoothness. The real and smoothed camera paths are then utilized as guidance to generate stable frames in a projective manner. We also capture videos with sensor data covering seven typical shooting scenarios and design a ground truth generation method to construct pseud-labels. Moreover, the trajectory smoothing network allows the use of 3- or 10-frame buffers as future information to construct a lookahead filter. Experimental results show that our online method could outperform other state-of-the-art offline methods in several shaky video clips with fewer buffer frames for both general and low-quality videos. Furthermore, our method could effectively reduce running times without performing image content analysis, and the stabilization efficiency reaches 25 fps on 1080p videos. Chen Li 0021, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Local Bidirection Recurrent Network for Efficient Video Deblurring with the Fused Temporal Merge ModuleabstractVideo deblurring methods exploit the correlation between consecutive blurry inputs to generate sharp frames. However, designing an effective and efficient method is a challenging problem for video deblurring. To guarantee the effectiveness and further improve the deblurring performance, we adopt the recurrent-based method as the baseline and reconsider the recurrent mechanism as well as the temporal feature alignment in the state-of-the-art methods. For the recurrent mechanism, we add the local backward connection to the global forward recurrent backbone to effectively exploit accurate future information. For the temporal alignment, we adopt a fused temporal merge module that exploits the superiority of flow-based and kernel-based methods with progressive correlation volumes estimation. In addition, we evaluate our method with both synthetic datasets (GoPro, DVD) and a realistic dataset (BSD). The experimental results demonstrate that our method achieves significant performance improvement with a slight computational cost increase against the state-of-the-art video deblurring methods. The extended ablation studies verify the effectiveness of our model. Chen Li 0021, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | High-Fidelity Face Reenactment Via Identity-Matched Correspondence LearningabstractFace reenactment aims to generate an animation of a source face using the poses and expressions from a target face. Although recent methods have made remarkable progress by exploiting generative adversarial networks, they are limited in generating high-fidelity and identity-preserving results due to the inappropriate driving information and insufficiently effective animating strategies. In this work, we propose a novel face reenactment framework that achieves both high-fidelity generation and identity preservation. Instead of sparse face representations (e.g., facial landmarks and keypoints), we utilize the Projected Normalized Coordinate Code (PNCC) to better preserve facial details. We propose to reconstruct the PNCC with the source identity parameters and the target pose and expression parameters estimated by 3D face reconstruction to factor out the target identity. By adopting the reconstructed representation as the driving information, we address the problem of identity mismatch. To effectively utilize the driving information, we establish the correspondence between the reconstructed representation and the source representation based on the features extracted by an encoder network. This identity-matched correspondence is then utilized to animate the source face using a novel feature transformation strategy. The generator network is further enhanced by the proposed geometry-aware skip connection. Once trained, our model can be applied to previously unseen faces without further training or fine-tuning. Through extensive experiments, we demonstrate the effectiveness of our method in face reenactment and show that our model outperforms state-of-the-art approaches both qualitatively and quantitatively. Additionally, the proposed PNCC reconstruction module can be easily inserted into other methods and improve their performance in cross-identity face reenactment. Jun Ling, Anni Tang, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | PTSEFormer: Progressive Temporal-Spatial Enhanced TransFormer Towards Video Object Detection
Shanyan Guan, Rong Xie 0004, Li Song 0001 |
ECCV (8) | 5 |
| 2022 | A Codec Information Assisted Framework for Efficient Compressed Video Super-Resolution
Hengsheng Zhang, Xueyi Zou, Jiaming Guo, Youliang Yan, Rong Xie 0004, Li Song 0001 |
ECCV (17) | 5 |
| 2022 | Generative Compression for Face Video: A Hybrid SchemeabstractAs the latest video coding standard, versatile video coding (VVC) has shown its ability in retaining pixel quality. To excavate more compression potential for video conference scenarios under ultra-low bitrate, this paper proposes a bitrate-adjustable hybrid compression scheme for face video. This hybrid scheme combines the pixel-level precise recovery capability of traditional coding with the generation capability of deep learning based on abridged information, where Pixel-wise Bi-Prediction, Low-Bitrate-FOM and Lossless Keypoint Encoder collaborate to achieve PSNR up to 36.23 dB at a low bitrate of 1.47 KB/s. Without introducing any additional bi-trate, our method has a clear advantage over VVC under a completely fair comparative experiment, which proves the effectiveness of our proposed scheme. Moreover, our scheme can adapt to any existing encoder/configuration to deal with different encoding requirements, and the bitrate can be dynamically adjusted according to the network condition. Anni Tang, Yan Huang 0033, Jun Ling, Zhiyu Zhang 0010, Rong Xie 0004, Li Song 0001 |
ICME | 6 |
| 2022 | Perceptual Video Coding Based on Semantic-Guided Texture Detection and SynthesisabstractVisually insensitive texture regions consume a large number of bitrate in hybrid video coding, leading to the waste of bandwidth resources. For this, we propose a semantic-guided texture synthesis framework (STSF). At encoder, high-level semantic information is adopted as texture features to detect texture regions and is sent to the decoder. Detected texture regions are coarsely encoded by hybrid codec. To generate realistic texture patterns, we design a multi-model semantic-guided texture synthesis generative adversarial network (STSGAN) at decoder, which works in a divide-and-conquer manner that semantically different texture regions are synthesized by different submodels in it. Experimental results show that STSF can achieve a −17.2% MOS BD-rate under the lowdelay_P configuration, compared with VVC. Guo Lu, Rong Xie 0004, Li Song 0001 |
PCS | 3 |
| 2022 | A Large-scale Sports Tracking Dataset and Progressive Re-detection Based Sports TrackingabstractRecent years have witnessed the great progress of Visual Object Tracking (VOT) which aims to predict the position of an object in each video frame given only its initial appearance. However, even the state-of-the-art methods are confronted with performance degradation, i.e., the tracker drift problem, in sports video scenes (e.g., soccer, basketball). There are two main causes that should be responsible for the tracker drift problem. First, the object of interest is often occluded by other objects that share a similar appearance. Such severe occlusion prevents the model from distinguishing the correct tracking object from other distractors in the future frames. Second, in sports videos, the objects often move fast from one place to another, which incurs severe blurry visual effects among consecutive frames. To address the issues of the tracker drift problem, we treat VOT as a tracking-by-re-detection task. Specifically, we detect candidate objects within a searching area (determined by object location in the previous frame) in the current frame and develop a progressive algorithm to filter out distractors in the area, which proves robust towards occlusion scenarios and tracker drift problems. Combining the advantages of our settings, the proposed framework method is robust to motion blur and object occlusion issues and achieves state-of-the-art tracking results on our challenging dataset. Qinyu Xu, Huaqiang Ren, Rong Xie 0004, Li Song 0001 |
VCIP | 5 |
| 2022 | L0 structure-prior assisted blur-intensity aware efficient video deblurring
Chen Li 0021, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
Neurocomputing | 3 |
| 2022 | IdentityDP: Differential private identification protection for face images
Yunqian Wen, Bo Liu 0001, Ming Ding 0001, Rong Xie 0004, Li Song 0001 |
Neurocomputing | 4 |
| 2022 | Multiview nonlinear discriminant structure learning for emotion recognition
Shuai Guo 0002, Li Song 0001, Rong Xie 0004, Lin Li 0062, Shenglan Liu 0001 |
Knowl. Based Syst. | 3 |
| 2022 | IdentityMask: Deep Motion Flow Guided Reversible Face Video De-IdentificationabstractUnprecedented video collection and sharing have exacerbated privacy concerns and led to increasing interest in privacy-preserving tools. A satisfactory video de-identification tool should be able to remove sensitive identity information from face videos while maintaining useful information for other identity-agnostic tasks. Meanwhile, it is necessary to allow the authority to inspect real identity when abnormal events are detected. Existing methods only focus on the study of de-identification, and lack the desired recovery ability when granting permissions. Furthermore, they all process the videos frame by frame, which hardly benefit from motion and inter-frame information. In this paper, we propose a modular architecture for reversible face video de-identification, called IdentityMask, which leverages deep motion flow to avoid per-frame evaluation. Our framework consists of two processes: the de-identification process provides a protective mask for identity information, while the recovery process can remove the protective mask if and only if the right key is provided. To this end, a Protection Module and a Recovery Module are built as two major functional modules, both based on an identity disentanglement network and guided by a crucial Motion Flow Module. An Affine Transformation Module provides simple but reliable assistance. Extensive experiments on a diverse natural video dataset (gender, ethnicity, age, etc.) demonstrate the effectiveness of the proposed framework for reversible face video de-identification. Yunqian Wen, Bo Liu 0001, Jingyi Cao, Rong Xie 0004, Li Song 0001, Zhu Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Edge-Based Video Compression Texture Synthesis Using Generative Adversarial NetworkabstractIt has been recognized that texture patterns with abundant high-frequency components, such as grass and water, produce visual masking effects, and the distortion in textures is hard to be perceived by human eyes than structure regions. However, modern video codecs in a rate-distortion optimized manner usually consume a lot of bits to encode textures, leading to the insufficiency in perceptual coding performance. Nowadays, with the rapid development of deep learning, learning based texture synthesis methods have been proposed to replace the coding process of prediction residuals to reduce the rate cost. In this paper, we present a deep texture synthesizer named edge-based texture synthesis framework (ETSF). At encoder side, the framework detects texture regions by semantic and fidelity classification criteria, and the detected regions are quantized coarsely by the hybrid coding framework. In texture characterization, ETSF extracts low-level edge features representing pixel intensity variation. Feature processing tools are developed to remove the spatiotemporal redundancy of edges. The processed edge information is compressed and transmitted. To effectively recover textures, we design an edge-based texture synthesis generative adversarial network (ETSGAN) at the decoder of ETSF, which can incorporate edge information into convolutional layers and generate realistic textures. Experimental results on a collected texture dataset show that the proposed ETSF can achieve an average of -12.8%, -14.2% and -9.6% MOS BD-rate under lowdelay_B, lowdelay_P and random_access configurations of VVC coding, respectively. Jun Xu 0040, Donghui Feng 0003, Rong Xie 0004, Li Song 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Region-Aware Adaptive Instance Normalization for Image HarmonizationabstractImage composition plays a common but important role in photo editing. To acquire photo-realistic composite images, one must adjust the appearance and visual style of the foreground to be compatible with the background. Existing deep learning methods for harmonizing composite images directly learn an image mapping network from the composite to real one, without explicit exploration on visual style consistency between the background and the foreground images. To ensure the visual style consistency between the foreground and the background, in this paper, we treat image harmonization as a style transfer problem. In particular, we propose a simple yet effective Region-aware Adaptive Instance Normalization (RAIN) module, which explicitly formulates the visual style from the background and adaptively applies them to the foreground. With our settings, our RAIN module can be used as a drop-in module for existing image harmonization networks and is able to bring significant improvements. Extensive experiments on the existing image harmonization benchmark datasets shows the superior capability of the proposed method. Code is available at https://github.com/junleen/RainNet. Jun Ling, Li Song 0001, Rong Xie 0004, Xiao Gu 0001 |
CVPR | 4 |
| 2021 | Dense 3D Coordinate Code Prior Guidance for High-Fidelity Face Swapping and Face ReenactmentabstractIn face synthesis tasks, commonly used 2D face representations (e.g. 2D landmarks, segmentation maps, etc.) are usually sparse and discontinuous. To combat these shortcomings, we utilize a dense and continuous representation, named Projected Normalized Coordinate Code (PNCC), as the guidance and develop a PNCC-Spatio-Normalization (PSN) method to achieve face synthesis regarding arbitrary head poses and expressions. Based on PSN, we provide an effective framework for face reenactment and face swapping task. To ensure a harmonious and seamless face swapping, a simple yet effective Appearance-Blending Module (ABM) is proposed to fit the synthesized face to the target face. Our method is subject-agnostic and can be applied to any pair of faces without extra fine-tuning. Both qualitative and quantitative experiments are conducted to demonstrate the superiority of the proposed method in comparisons to existing state-of-the-art systems. Anni Tang, Jun Ling, Rong Xie 0004, Li Song 0001 |
FG | 4 |
| 2021 | Personalized and Invertible Face De-identification by Disentangled Identity Information ManipulationabstractThe popularization of intelligent devices including smartphones and surveillance cameras results in more serious privacy issues. De-identification is regarded as an effective tool for visual privacy protection with the process of concealing or replacing identity information. Most of the existing de-identification methods suffer from some limitations since they mainly focus on the protection process and are usually non-reversible. In this paper, we propose a personalized and invertible de-identification method based on the deep generative model, where the main idea is introducing a user-specific password and an adjustable parameter to control the direction and degree of identity variation. Extensive experiments demonstrate the effectiveness and generalization of our proposed framework for both face de-identification and recovery. Jingyi Cao, Bo Liu 0001, Yunqian Wen, Rong Xie 0004, Li Song 0001 |
ICCV | 4 |
| 2021 | Blindly Predict Image and Video Quality in the WildabstractEmerging interests have been brought to blind quality assessment for images/videos captured in the wild, known as in-the-wild I/VQA. Prior deep learning based approaches have achieved considerable progress in I/VQA, but are intrinsically troubled with two issues. Firstly, most existing methods fine-tune the image-classification-oriented pre-trained models for the absence of large-scale I/VQA datasets. However, the task misalignment between I/VQA and image classification leads to degraded generalization performance. Secondly, existing VQA methods directly conduct temporal pooling on the predicted frame-wise scores, resulting in ambiguous inter-frame relation modeling. In this work, we propose a two-stage architecture to separately predict image and video quality in the wild. In the first stage, we resort to supervised contrastive learning to derive quality-aware representations that facilitate the prediction of image quality. Specifically, we propose a novel quality-aware contrastive loss to pull together samples of similar quality and push away quality-different ones in embedding space. In the second stage, we develop a Relation-Guided Temporal Attention (RTA) module for video quality prediction, which captures global inter-frame dependencies in embedding space to learn frame-wise attention weights for frame quality aggregation. Extensive experiments demonstrate that our approach performs favorably against state-of-the-art methods on both authentically distorted image benchmarks and video benchmarks. Jiapeng Tang, Yi Fang 0009, Rong Xie 0004, Xiao Gu 0001, Guangtao Zhai, Li Song 0001 |
MMAsia | 4 |
| 2021 | Deep Face Swapping via Cross-Identity Adversarial Training
Jun Ling, Li Song 0001, Rong Xie 0004 |
MMM (2) | 5 |
| 2021 | HEVC VMAF-oriented Perceptual Rate Distortion Optimization using CNNabstractVideo coding standards like HEVC and VVC have achieved significant coding performance. However, the RDO module in coding framework ignores the characteristics of human visual system (HVS), which leads to insufficiency for perceptual video coding. Recently, learning-based objective assessment metric VMAF is developed and has been demonstrated higher quality assessment accuracy than conventional metrics. To incorporate VMAF into RDO aiming at improving perceptual coding efficiency, in this paper, a perceptual RDO scheme is proposed. A CNN-based on-line training method is first explored to determine the VMAF-related distortion estimation coefficient. Based on the VMAF-related coefficient and R-D model, a VMAF-based Lagrangian multiplier is proposed to adjust the R-D performance of each coding block. Experiments demonstrate that the proposed method can achieve an average -2.80% VMAF-based BD-Rate compared with the original HEVC, which effectively improves the coding performance. Yan Huang 0033, Rong Xie 0004, Li Song 0001 |
PCS | 3 |
| 2021 | Deep Motion Flow Aided Face Video De-identificationabstractAdvances in cameras and web technology have made it easy to capture and share large amounts of face videos over to an unknown audience with uncontrollable purposes. These raise increasing concerns about unwanted identity-relevant computer vision devices invading the characters's privacy. Previous de-identification methods rely on designing novel neural networks and processing face videos frame by frame, which ignore the data feature in redundancy and continuity. Besides, these techniques are incapable of well-balancing privacy and utility, and per-frame evaluation is easy to cause flicker. In this paper, we present deep motion flow, which can create remarkable de-identified face videos with a good privacy-utility tradeoff. It calculates the relative dense motion flow between every two adjacent original frames and runs the high quality image anonymization only on the first frame. The de-identified video will be obtained based on the anonymous first frame via the relative dense motion flow. Extensive experiments demonstrate the effectiveness of our proposed de-identification method. Yunqian Wen, Bo Liu 0001, Rong Xie 0004, Jingyi Cao, Li Song 0001 |
VCIP | 3 |
| 2021 | Modeling Acceleration Properties for Flexible INTRA HEVC Complexity ControlabstractIt is a very well-known fact, that the high complexity of the High Efficiency Video Coding standard (HEVC) is the main hurdle for its wide deployment and use. To tackle this problem, a number of recent research outcomes exploit heuristic algorithms and machine learning, including deep learning, to reduce the coding complexity. However, in most cases, each encoder module, i.e., encoding process, is first accelerated individually, and then different acceleration algorithms are manually combined. Without a holistic strategy, the acceleration potential of multi-module combination is not exploited and the Rate-Distortion (RD) loss is generally not well controlled. To tackle these shortcomings, this paper exploits the acceleration properties of different modules, i.e., the numerical representation of potential time saving and possible RD loss, from which a heuristic model is explored. Then a Heuristic Model Oriented Framework (HMOF) is proposed which adapts the properties of modules to underlying acceleration algorithms. In the framework, two advanced acceleration algorithms, including Border Considered CNN (BC-CNN)-based Coding Unit (CU) partition and Naive Bayes-based Prediction Unit (PU) partition, are proposed for the CU and PU modules, respectively. Further, by leveraging the heuristic model as the guidance to combine the proposed acceleration algorithms, HMOF is globally optimized, where different time saving budgets are wisely allocated to different modules and a theoretically minimal RD loss is achieved. According to the experimental results, through fusing a suitable deep learning technique and a Bayes-Based prediction, the proposed acceleration framework HMOF enable multiple acceleration choices. Here the proposed joint optimization strategy help to make a choice leading to the best cost-performance. Furthermore, within the proposed framework, intra coding time can be precisely controlled with negligible Bjøntegaard delta bit-rate (BDBR) loss. In this context, as a complexity control method, HMOF outperforms the state-of-the-art complexity reduction algorithms under a similar complexity reduction ratio. These results partially demonstrate the superiority of the proposed technique. Yan Huang 0033, Li Song 0001, Rong Xie 0004, Ebroul Izquierdo, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | VMAF Oriented Perceptual Coding Based on Piecewise Metric CouplingabstractIt has been recognized that videos have to be encoded in a rate-distortion optimized manner for high coding performance. Therefore, operational coding methods have been developed for conventional distortion metrics such as Sum of Squared Error (SSE). Nowadays, with the rapid development of machine learning, the state-of-the-art learning based metric Video Multimethod Assessment Fusion (VMAF) has been proven to outperform conventional ones in terms of the correlation with human perception, and thus deserves integration into the coding framework. However, unlike conventional metrics, VMAF has no specific computational formulas and may be frequently updated by new training data, which invalidates the existing coding methods and makes it highly desired to develop a rate-distortion optimized method for VMAF. Moreover, VMAF is designed to operate at the frame level, which leads to further difficulties in its application to today's block based coding. In this paper, we propose a VMAF oriented perceptual coding method based on piecewise metric coupling. Firstly, we explore the correlation between VMAF and SSE in the neighborhood of a benchmark distortion. Then a rate-distortion optimization model is formulated based on the correlation, and an optimized block based coding method is presented for VMAF. Experimental results show that 3.61% and 2.67% bit saving on average can be achieved for VMAF under the low_delay_p and the random_access_main configurations of HEVC coding respectively. Zhengyi Luo 0001, Yan Huang 0033, Rong Xie 0004, Li Song 0001, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 4 |
| 2020 | FACT: Fused Attention for Clothing Transfer with Generative Adversarial NetworksabstractClothing transfer is a challenging task in computer vision where the goal is to transfer the human clothing style in an input image conditioned on a given language description. However, existing approaches have limited ability in delicate colorization and texture synthesis with a conventional fully convolutional generator. To tackle this problem, we propose a novel semantic-based Fused Attention model for Clothing Transfer (FACT), which allows fine-grained synthesis, high global consistency and plausible hallucination in images. Towards this end, we incorporate two attention modules based on spatial levels: (i) soft attention that searches for the most related positions in sentences, and (ii) self-attention modeling long-range dependencies on feature maps. Furthermore, we also develop a stylized channel-wise attention module to capture correlations on feature levels. We effectively fuse these attention modules in the generator and achieve better performances than the state-of-the-art method on the DeepFashion dataset. Qualitative and quantitative comparisons against the baselines demonstrate the effectiveness of our approach. Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
AAAI | 4 |
| 2020 | Toward Fine-Grained Facial Expression Manipulation
Jun Ling, Li Song 0001, Rong Xie 0004, Xiao Gu 0001 |
ECCV (28) | 5 |
| 2020 | Realistic Talking Face Synthesis With Geometry-Aware Feature TransformationabstractRecent studies have shown remarkable success in synthesizing realistic talking faces by exploiting generative adversarial networks. However, existing methods are mostly target specific that cannot generate images of previously unseen people, and they suffer from artifacts such as blurriness and mismatching of facial details. In this paper, we tackle these problems by proposing a target-agnostic framework. We introduce a geometry-aware feature transformation module to achieve shape transfer while preserving the appearance of the source face. To further improve image quality of synthesized results, we present a multi-scale spatially-consistent transfer unit to maintain spatial consistency between the encoder and decoder features. Experimental results show that our model is able to synthesize photo-realistic talking faces which are previously unseen, outperforming state-of-the-art methods both qualitatively and quantitatively. Jun Ling, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
ICIP | 4 |
| 2020 | Learning-Based Quality Enhancement For Scalable Coded Video Over Packet Lossy NetworksabstractThe layered feature of scalable video coding (SVC) offers a sufficient adaptation to unreliable transmission. When network condition drops sharply, enhancement layers will be abandoned, and only base layers are delivered. However, this will cause noticeable visual artifacts due to quality differences between different layers. To alleviate this problem, we novelly introduce a deep learning-based method into video reconstruction phase of scalable bitstreams. A super-resolution motivated recurrent network is proposed to extract and fuse features from both previous high-resolution frames and the current low-resolution frame. To the best of our knowledge, this is the first attempt to improve the performance of scalable bitstreams reconstruction by a specifically designed super-resolution network. By seamlessly integrating the accessible features, significant video quality improvements in terms of PSNR, SSIM, and VMAF are achieved. At the same time, the improvement of overall visual quality stability is apparent under packet lossy networks, indicating both efficiency and robustness of our approach. Shengwei Yu, Xun Tong, Yan Huang 0033, Rong Xie 0004, Li Song 0001 |
ICME | 4 |
| 2020 | A Deep Tracking and Segmentation Approach for Soccer Videos Visual Effects
Shenhui Peng, Li Song 0001, Jun Ling, Rong Xie 0004, Lin Li 0062 |
PRCV (2) | 4 |
| 2020 | Deep Blind Video Quality Assessment for User Generated VideosabstractAs short video industry grows up, quality assessment of user generated videos has become a hot issue. Existing no reference video quality assessment methods are not suitable for this type of application scenario since they are aimed at synthetic videos. In this paper, we propose a novel deep blind quality assessment model for user generated videos according to content variety and temporal memory effect. Content-aware features of frames are extracted through deep neural network, and a patch-based method is adopted to obtain frame quality score. Moreover, we propose a temporal memory-based pooling model considering temporal memory effect to predict video quality. Experimental results conducted on KoNViD-1k and LIVE-VQC databases demonstrate that the performance of our proposed method outperforms other state-of-the-art ones, and the comparative analysis proves the efficiency o f our temporal pooling model. Jiapeng Tang, Rong Xie 0004, Xiao Gu 0001, Li Song 0001, Lin Li 0062 |
VCIP | 3 |
| 2020 | A Hybrid Model for Natural Face De-Identiation with Adjustable PrivacyabstractAs more and more personal photos are shared and tagged in social media, security and privacy protection are becoming an unprecedentedly focus of attention. Avoiding privacy risks such as unintended verification, becomes increasingly challenging. To enable people to enjoy uploading photos without having to consider these privacy concerns, it is crucial to study techniques that allow individuals to limit the identity information leaked in visual data. In this paper, we propose a novel hybrid model consists of two stages to generate visually pleasing de-identified face images according to a single input. Meanwhile, we successfully preserve visual similarity with the original face to retain data usability. Our approach combines latest advances in GAN-based face generation with well-designed adjustable randomness. In our experiments we show visually pleasing de-identified output of our method while preserving a high similarity to the original image content. Moreover, our method adapts well to the verificator of unknown structure, which further improves the practical value in our real life. Yunqian Wen, Bo Liu 0001, Rong Xie 0004, Yunhui Zhu, Jingyi Cao, Li Song 0001 |
VCIP | 3 |
| 2020 | Quality of Experience Evaluation for Streaming Video Using CGNNabstractOne of the principal contradictions these days in the field of video i s lying between the booming demand for evaluating the streaming video quality and the low precision of the Quality of Experience prediction results. In this paper, we propose Convolutional Neural Network and Gate Recurrent Unit (CGNN)-QoE, a deep learning QoE model, that can predict overall and continuous scores of video streaming services accurately in real time. We further implement state-of-the-art models on the basis of their works and compare with our method on six public available datasets. In all considered scenarios, the CGNN-QoE outperforms existing methods. Zhiming Zhou 0001, Li Song 0001, Rong Xie 0004, Lin Li 0062 |
VCIP | 4 |
| 2019 | Gan Based Multi-Exposure Inverse Tone MappingabstractHigh dynamic range (HDR) imaging provide larger range of luminosity and wider color gamut than conventional low dynamic range (LDR) imaging. The method which transforms LDR contents to HDR contents is called inverse tone mapping. After deep neural networks are used in inverse tone mapping problem, researchers mostly focus on transforming normal exposure LDR images to HDR. However, when people use inverse tone mapping in practice, they get some ill-exposed images as well. The state-of-art algorithms can't transform these images to HDR well.In this work, we propose an end-to-end multi-exposure inverse tone mapping (MITM) framework based on existing generative adversarial network (GAN). This framework can transform a single LDR image not only at normal exposure, but also at unsuitable exposure to a normal exposure HDR image. We use histogram equalization to preprocess the luma of the input LDR images; when training the model, we use intrinsic image decomposition to divide the output HDR images into illuminance and reflectance components and use these two components to constrain the luminance information and the color information separately. This framework can adjust the unsuitable exposure and provide a better viewing experience than other state-of-art algorithms in the experimental results. Shiyu Ning, Rong Xie 0004, Li Song 0001 |
ICIP | 3 |
| 2019 | VMAF Oriented Perceptual Optimization for Video CodingabstractIn the light of low costs and automatic assessment, objective visual quality metrics enjoy many important applications such as perceptual coding. Recently multiple metrics obtain further improvement by means of machine learning. However, due to the absence of specific formulas, it's often hard to incorporate learning based metrics into video coding. In this paper, taking the state-of-the-art learning based metric VMAF for example, we propose a method of perceptual coding in an inferential manner for learning based metrics. The rate distortion optimization is adapted during coding as well. Experimental results show that compared with conventional methods, the proposed method can achieve obvious bitrate saving under HEVC coding. Zhengyi Luo 0001, Yan Huang 0033, Rong Xie 0004, Li Song 0001 |
ISCAS | 4 |
| 2019 | Reinforcement Learning Based Adaptive Bitrate Algorithm for Transmitting Panoramic VideosabstractPanoramic videos have become more and more popular now. 360-degree videos give users a better experience but put forward a higher request for network at the same time. Many kinds of solutions to meet the high need of bandwidth have been proposed, such as tiled-based transmission, layer-based transmission and so on. Then how to choose the most suitable bitrates to make full use of the network resource is the problem to be solved urgently. In this paper, we propose a method based on reinforcement learning(RL) algorithm to select the bitrates of the region-of-interest adaptively for panoramic videos. We also compare RL algorithm with three traditional algorithms when changing the Field of View(FOV) and find that RL algorithm proves to perform better in minimizing the rebuffer time and providing higher quality video contents under various network conditions. Xiaona Wu, Xun Tong, Rong Xie 0004, Li Song 0001 |
ISCAS | 4 |
| 2019 | JND-based Perceptual Rate Distortion Optimization for AV1 EncoderabstractAV1 is the next-generation open video coding format, and it can achieve significant coding efficiency with novel coding tools. It supports Lagrangian rate distortion optimization (RDO) method to optimize the coding performance. However, the distortion and the Lagrangian multiplier used in RDO ignore the characteristics of human visual system (HVS), which leads to insufficiency for perceptual video coding. To solve this problem, a perceptual RDO scheme based on the Just Noticeable Distortion (JND) threshold of HVS is proposed. The JND for each pixel is first measured according to three perceptual features: luminance adaptation, masking effects and structure sensitivity. Based on the observation that the regions with smaller distortion visibility thresholds are more sensitive to HVS, a JND-based Lagrangian multiplier is derived to adaptively adjust the rate-distortion (RD) performance for each coding block. Experiments demonstrate that the proposed method can achieve an average SSIM-based -3.93% BD-Rate saving compared with the original AV1 encoder, which effectively improve the coding performance. Li Song 0001, Rong Xie 0004, Jingning Han, Yaowu Xu |
PCS | 3 |
| 2019 | Identifying and Pruning Redundant Structures for Deep Neural NetworksabstractDeep convolutional neural networks have achieved considerable success in the field of computer vision. However, it is difficult to deploy state-of-the-art models on resource-constrained platforms due to their high storage, memory bandwidth, and computational costs. In this paper, we propose a structured pruning method which employs a three-step process to reduce the resource consumption of neural networks. First, we train an initial network on the training set and evaluate it on the validation set. Next, we introduce an iterative pruning and fine-tuning algorithm to identify and prune redundant structures, which results in a pruned network with a compact architecture. Finally, we train the pruned network from scratch on both the training set and validation set to obtain the final accuracy on the test set. In the experiments, our pruning method significantly reduces the model size (by 87.2% on CIFAR-10), saves inference time (53.3% on CIFAR-10), and achieves better performance as compared to recent state-of-the-art methods. Wenyao Gan, Li Song 0001, Li Chen 0021, Rong Xie 0004, Xiao Gu 0001 |
VCIP | 4 |
| 2019 | FPGA Based Video Transcoding System with 2K-4K Super-Resolution ConversionabstractWe present a FPGA-based system supporting video stream transcoding with 2k full high-definition (FHD) video to 4k ultra high-definition (UHD) video super- resolution(SR) conversion. Our system focuses on building a functional pipeline with convolutional neural network (CNN) accelerator and real-time video codec unit for converting H.264 video stream to H.265/HEVC video stream. The overall video processing system can be used as an important plug-in module in the video streaming network to improve the video stream service quality. Yuzhuo Wei, Li Chen 0021, Rong Xie 0004, Li Song 0001, Xiaoyun Zhang 0001 |
VCIP | 3 |
| 2019 | Deep Feature Guided Image RetargetingabstractImage retargeting is the technique to display images via devices with various aspect ratios and sizes. Traditional content-aware retargeting methods rely on low-level features to predict pixel-wise importance and can hardly preserve both the structure lines and salient regions of the source image. To address this problem, we propose a novel adaptive image warping approach which integrates with deep convolutional neural network. In the proposed method, a visual importance map and a foreground mask map are generated by a pre-trained network. The two maps and other constraints guide the warping process to yield retargeted results with less distortions. Extensive experiments in terms of visual quality and a user study are carried out on the widely used RetargetMe dataset. Experimental results show that our method outperforms current state-of-art image retargeting methods. Jinan Wu, Rong Xie 0004, Li Song 0001, Bo Liu 0001 |
VCIP | 2 |
| 2019 | An Improved QoE Evaluation Model for HTTP Adaptive StreamingabstractHTTP adaptive streaming (HAS) is an adaptive bitrate streaming technique that enables high quality streaming of media content over the Internet delivered from conventional HTTP web servers. ITU-T Rec. P.1203.3 is the first standardized Quality of Experience model for audiovisual HTTP Adaptive Streaming. It takes into account the subjective impact of HAS-typical effects(such as buffering, quality switches) on users. But the buffering inputs required for the model is too complex and redundant, and it does not consider the impact of the worst video quality on QoE. In the paper we optimize the ITU-T Rec. P.1203.3 model in the above two aspects, simplify model input and improve model accuracy and stability. Zaixin Yang, Tiantian He 0005, Li Song 0001, Rong Xie 0004, Xiao Gu 0001 |
VCIP | 4 |
| 2018 | Learning an Inverse Tone Mapping Network with a Generative Adversarial RegularizerabstractTransferring a low-dynamic-range (LDR) image to a high-dynamic-range (HDR) image, which is the so-called inverse tone mapping (iTM), is an important imaging technique to improve visual effects of imaging devices. In this paper, we propose a novel deep learning-based iTM method, which learns an inverse tone mapping network with a generative adversarial regularizer. In the framework of alternating optimization, we learn a U-Net-based HDR image generator to transfer input LDR images to HDR ones, and a simple CNN-based discriminator to classify the real HDR images and the generated ones. Specifically, when learning the generator we consider the content-related loss and the generative adversarial regularizer jointly to improve the stability and the robustness of the generated HDR images. Using the learned generator as the proposed inverse tone mapping network, we achieve superior iTM results to the state-of-the-art methods consistently. Shiyu Ning, Hongteng Xu, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
ICASSP | 4 |
| 2018 | GPU Based Motion-Compensated Frame Interpolation Acceleration for Future Video CodingabstractBeing developed by Joint Video Exploration Team (JVET), Future Video Coding (FVC) aims at higher resolutions and higher compression performance than the state-of-the-art HEVC standard, undoubtedly at the cost of further computing increases. As an efficient computing platform, Graphics Processing Unit (GPU) is often used to accelerate encoding. But with the adoption of instruction set acceleration in the reference software of FVC, previous methods often become less efficient or even lead to a lower speed. In this paper, based on the comparative analysis of the time consumption between HEVC and FVC, we propose a GPU based acceleration method for the most computation-intensive step - frame interpolation of FVC, where frame caching strategy and a multi-stream mechanism is designed to make the best of GPU resources. Experimental results show that compared with the instruction set accelerated reference software of FVC, our method could achieve average 67.12% speed-up gains on the interpolation module and average 6.35% speed-up gains on overall encoding with exactly the same performance as before. Jianlun Tang, Yan Huang 0033, Rong Xie 0004, Zhengyi Luo 0001, Li Song 0001 |
ICIP | 3 |
| 2018 | Frame Interpolation via Refined Deep Voxel FlowabstractTraditional frame interpolation methods first estimate motion between two consecutive frames and then synthesize intermediate frames. This problem is challenging because of complex motion and video scenes. In this paper, we present an end-to-end deep network for frame interpolation problem. Based on a video synthesis method deep voxel flow (DVF), refinement modules are designed to increase the accuracy of voxel flow, which we call Refined DVF (RDVF). A deeper architecture with more convolution and deconvolution layers is also utilized to help extract motion. Our results greatly improve the performance of original DVF and compare favorably to state-of-the-art methods both quantitatively and qualitatively. Zhifeng Zhang 0003, Li Chen 0021, Rong Xie 0004, Li Song 0001 |
ICIP | 3 |
| 2018 | An MCMC based Efficient Parameter Selection Model for x265 EncoderabstractAs an open-source and computationally efficient High Efficiency Video Coding (HEVC) encoder, x265 has been gaining increasing popularity in video applications. x265 provides numerous encoding parameters in view of flexibility. However, proper and efficient setting of parameters often becomes a great challenge in practice. In this paper, we deeply investigate the influence of x265 parameters based on the Slow preset and pick out important parameters in terms of efficiency and complexity. Then a Markov Chain Monte Carlo (MCMC) based algorithm is proposed for efficient parameter adaptation at the target encoding time. This paper shows that carefully selected low-complexity encoding configurations can achieve the coding efficiency comparable to that of high-complexity ones. Specifically, average 26.72% encoding time reduction can be achieved while maintaining similar Rate Distortion (RD) performance to x265 presets using the proposed algorithm. Yan Huang 0033, Li Song 0001, Rong Xie 0004, Zhengyi Luo 0001 |
ISCAS | 3 |
| 2018 | Masking Effects Based Rate Control Scheme for High Efficiency Video CodingabstractThis paper presents a masking effects based rate control scheme for high efficiency video coding (HEVC). Rate control is regarded as a very effective tool to improve the performance of video coding under the limited bandwidth. However, the state-of-the-art rate control algorithm based on R-X model ignores the characteristics of human visual system (HVS), which leads to poor performance in subjective quality. Moreover, some structural similarity (SSIM) or saliency based perceptual rate control algorithms only consider spatial characteristics. Since spatial and temporal visual masking effects can better reflect the characteristics of HVS, in this paper masking effects based perceptual factor for coding tree unit (CTU) is proposed, which takes both texture complexity and motion information into account. Then the proposed perceptual factor is utilized to guide bit allocation in CTU-level rate control. Experimental results show that the proposed scheme can effectively improve the coding performance compared with the R-λ algorithm. Hao Wang 0073, Li Song 0001, Rong Xie 0004, Zhengyi Luo 0001 |
ISCAS | 3 |
| 2018 | Rate-mixed HEVC Tile based 360 Video Streaming SystemabstractRecently, tile-based viewport adaptation is a popular method for 360 video streaming. Our demonstration adopts a rate-mixed transmission approach utilizing a VR adaptation agent at the server end for viewport-based streaming, which is client-compatible and can be scalable to different users. The FOV prediction is applied to improve the viewing experience. The feasibility of our system to head-mounted displays is verified, which can reduce bandwidth consumption by up to 36%. Xu Liu 0006, Rong Xie 0004, Li Song 0001 |
VCIP | 4 |
| 2018 | An improved Real-Time Video Communication SystemabstractIn this paper, we optimized the Linphone-based real-time video communication system. Firstly, we used HEVC to replace the H.264 in order to reduce the bandwidth pressure while reducing the buffer delay by configuring the appropriate encoding parameters. Secondly, we added an effective bitrate adaptive algorithm based on additive increase and multiplicative decrease (AIMD) in the system, which can effectively reduce the packet loss rate in the network with high bandwidth fluctuations. Zhaoliang Ma, Shengwei Yu, Yongcheng Huang, Rong Xie 0004, Li Song 0001 |
VCIP | 4 |
| 2017 | CNN based post-processing to improve HEVCabstractIn this paper, we propose a frame-based dynamic metadata post-processing scheme in HEVC. Video sequence is classified into different categories contains complexity of video content and quality indicator for each frame, an up-to-one byte flag embedded in the bitstream is transferred as side information. Meanwhile dynamic metadata contains classification information indicates the offline training of separate network models. Specifically, we adopt a 20-layers CNN (Con-volutional Neural networks) model to extract more meaningful information from the reconstructed error and improve the filtering performance. Experimental results shows that our proposed post-processing scheme leads on average 1.6% BD-rate reduction compared with HEVC baseline on the six sequences given in 2017 ICIP Grand Challenge. Chen Li 0021, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
ICIP | 3 |
| 2017 | Lagrangian method based Rate-Distortion Optimization revisited for dependent video codingabstractVideo encoding is based on the DPCM framework where temporal prediction coding introduces Rate-Distortion (RD) dependence. The RD operating point of the current unit depends on the particular choices of RD points of its reference units. Unfortunately, common Lagrangian optimization method based Rate-Distortion Optimization (RDO) for video coding is based on an independence assumption which omits the RD dependences, and thus compromises the RD performance. In this paper, we revisit the Lagrangian optimization method based RDO for dependent video coding. A theoretical RD dependence decoupling method based on independent distortion decomposition is firstly presented. After the discussion of reasonability of the theoretical decoupling method, the practical One Step Ahead Decoupling Strategy (OSADS) is proposed. After implemented on the HEVC encoder, the strategy achieves average 2.1% BD-rate saving compared with the HM encoder under the same low-delay P configuration. Li Song 0001, Zhengyi Luo 0001, Rong Xie 0004 |
ICIP | 4 |
| 2017 | Rate control model for high dynamic range videoabstractThis paper describes a luminance based rate control (RC) model for high dynamic range (HDR) video. A novel mathematical relationship between luminance and bit allocation of a coding tree unit (CTU) is presented. By adjusting the existing RC algorithm through the proposed model, a better balance between dark and bright areas can be achieved and -4.4% gains can be obtained in terms of average BD-Rate (tPSNR-XZY). Moreover, subjective assessment also shows that, compared with the existing RC model, the proposed method can convey a wider range of perceptible shadow and highlight more details. Lixun Bai, Li Song 0001, Rong Xie 0004, Liang Zhang 0026, Zhengyi Luo 0001 |
VCIP | 3 |
| 2017 | Weight-based bit allocation scheme for VR videos in HEVCabstractSince VR videos shown on the head-mounted display (HMD) is omnidirectional, the average distortion of VR videos in all directions shall be calculated in spherical domain. Several metrics have been proposed to calculate the coding loss of VR videos in spherical domain, including S-PSNR, WS-PSNR, CPP-PSNR. The above metrics are all improved based on PSNR by creating a weight map. This paper aims to optimize the rate control scheme on HEVC mostly for WS-PSNR. According to the weight map of WS-PSNR, regions with more weights in plain are more important, thus more bits shall be allocated to the important regions. We use weight-based rate control scheme to realize the above thoughts. Weight-based rate control scheme defines bit per weight (bpw) instead of bpp. Large values of bpw indicate that the important regions deserve high bitrates, thus probably achieving better quality. Consequently, the proposed rate control scheme improves the video quality of VR videos, which leads to average gain of 2.1%, 4.3% and 1.5% of S-PSNR, WS-PSNR and CPP-PSNR. BiJia Li, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
VCIP | 3 |
| 2017 | A generic method to improve no-reference image blur metric accuracy in video contentsabstractWe present in this work a generic and effective method to increase the prediction accuracy of no-reference image/video blur assessment facing the real-world content diversity. We demonstrate that benchmarking no reference image blur metrics, fitting a single logistic function to map the objective predictions to subjective scores in the well-known databases like LIVE or TID2008/2013, introduce biased fitting results towards better predictions only in the central part of the score scale. We find out that a multi-fitting approach, using the correlation parameters between subjective scores and objective predictions for content clustering and then conducting logistic fitting for each content type, can evidently improve the metric prediction accuracy in the full score scale. Besides, the overall prediction variance is also reduced with the proposed scheme, presenting more consistent results insensitive of content variation. We prove that the proposed method is of practical meaning to facilitate blur assessment techniques validated on limited databases to the vastly abundant real-life content types. Yankai Liu, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
VCIP | 3 |
| 2017 | Two-stream deep encoder-decoder architecture for fully automatic video object segmentationabstractWe propose a two-stream Deep Encoder-Decoder architecture to tackle the task of fully automatic video object segmentation. Both two streams, i.e., ImSeg-Stream (for static image segmentation) and MoSeg-Stream (for optical flow segmentation), hold the totally same Encoder-Decoder architecture. The Encoder part generates a low-resolution mask with accurate locations and smooth boundaries, while the Decoder part refines the details of initial mask and enlarges its resolution via integrating lower-level features progressively. At last two streams learn to integrate for better results. Moreover, to handle the problem of inadequate video object segmentation datasets, we propose a seeking strategy to generate a large-scale handcrafted dataset for training. Experiments on two standard datasets demonstrate that proposed method outperforms most state-of-the-art methods in both segmentation accuracy and run time. Li Song 0001, Rong Xie 0004 |
VCIP | 3 |
| 2016 | Improved intra angular prediction with novel interpolation filter and boundary filterabstractIn this paper, two improved intra angular prediction methods are proposed to enhance coding performance. The first method applies new four-tap interpolation filter algorithm. The reference samples at fractional position are interpolated by DCT-based or Gaussian interpolation filter. The second method proposes extended boundary prediction filter to reduce the prediction error. The experimental results show that for AI configuration, the overall coding gain is about 0.85% on average comparing to HEVC reference software while maintaining almost the same coding time. Rujun Wei, Rong Xie 0004, Li Song 0001, Liang Zhang 0026, Wenjun Zhang 0001 |
PCS | 2 |
| 2016 | Shot boundary detection using convolutional neural networksabstractVideo shot boundary detection (SBD) is necessary for further video analysis like video retrieval and annotation. Great efforts have been made to develop SBD algorithms for speed and accuracy. Most works implement frame histogram as features to measure similarity for detection. However, when changes between consecutive shot boundaries are small and backgrounds of them are highly similar, most state-of-the-art methods miss these boundaries thus cannot achieve high accuracy of detection. In this paper we propose a novel SBD framework with Convolutional Neural Networks (CNNs). Firstly we adopt a candidate segment selection method to locate the positions of shot boundaries coarsely using adaptive thresholds and eliminate most non-boundary frames. Then CNN is implemented to extract representative features of frames in candidate segments. Finally cut and gradual transitions can be obtained by using a novel pattern-matching method based on a new similarity strategy. Experiments on TRECVID 2001 test data demonstrate that the proposed scheme outperforms the state-of-the-art methods and achieves high accuracy of detection. Li Song 0001, Rong Xie 0004 |
VCIP | 3 |
| 2016 | Data-Driven Crowd Understanding: A Baseline for a Large-Scale Crowd DatasetabstractCrowd understanding has drawn increasing attention from the computer vision community, and its progress is driven by the availability of public crowd datasets. In this paper, we contribute a large-scale benchmark dataset collected from the Shanghai 2010 World Expo. It includes 2630 annotated video sequences captured by 245 surveillance cameras, far larger than any public dataset. It covers a large number of different scenes and is suitable for evaluating the performance of crowd segmentation and estimation of crowd density, collectiveness, and cohesiveness, all of which are universal properties of crowd systems. In total, 53 637 crowd segments are manually annotated with the three crowd properties. This dataset is released to the public to advance research on crowd understanding. The largescale annotated dataset enables using data-driven approaches for crowd understanding. In this paper, a data-driven approach is proposed as a baseline of crowd segmentation and estimation of crowd properties for the proposed dataset. Novel global and local crowd features are designed to retrieve similar training scenes and to match spatio-temporal crowd patches so that the labels of the training scenes can be accurately transferred to the query image. Extensive experiments demonstrate that the proposed method outperforms state-of-the-art approaches for crowd understanding. Cong Zhang 0005, Kai Kang 0006, Hongsheng Li 0001, Xiaogang Wang 0001, Rong Xie 0004, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 5 |
| 2010 | Backward Adaptive Pre/Post-Filtered DPCM with Near-Optimal Rate-Distortion PerformanceabstractIn this paper, we propose two backward adaptive coding algorithms based on a recently invented pre/post-filtered DPCM (Differential Pulse-Coded Modulation) codec. The pre/post-filters and the predictor are adapted jointly. Source statistics are assumed unknown a priori. One of the algorithms is based on power spectrum estimation, and the other gradient descent. Some properties of the algorithms are analyzed. Experiment results show that both algorithms achieve near-optimal rate-distortion performance, significantly outperforming the adaptive DPCM without adaptive pre/post-filtering. Yuhua Fan, Jia Wang 0004, Jun Sun 0005, Rong Xie 0004 |
ICC | 4 |
| 2009 | An improved block size selection method based on macroblock movement characteristic
Jun Sun 0005, Rong Xie 0004, Songyu Yu, Wenjun Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2008 | Variable block size selection for a transcoder based on MB movement informationabstractTo improve coding performance, the new video compression standard H.264 employs seven variable block sizes for one Macro Block (MB) to conduct motion estimation and compensation. MPEG-2 only has one size 16×16. This paper presents a novel fast variable block size selection method for an inter-MB in a video transcoder from MPEG-2 to H.264 with downscaling by a factor two in each dimension based on the MB motion information. Without conducting motion re-estimation, an optimal block size is decided. Experiment results show that this method saves the transcoder complexity dramatically with little compression performance degradation. Jun Sun 0005, Rong Xie 0004, Shibao Zheng, Songyu Yu |
ICME | 3 |
| 2007 | Bit Allocation for Fine-Granular SNR Scalability Coding with Hierarchical B PicturesabstractHierarchical B pictures are devised to achieve temporal scalability in the scalable extension of H.264/AVC (SVC) which is under standardization. The fine-granular SNR scalability (FGS) can be provided by progressive refinement (PR) slices in SVC. In this paper, we firstly investigate error propagation in the case of discarding PR slices and obtain a rate difference distortion optimization criterion to improve coding efficiency of base layer. Then we consider the full rate case and propose a rate distortion slope criterion to enhance FGS coding efficiency at high rate. Finally the criterion to boost coding efficiency in the whole range of FGS rate is derived by combing the criterions derived previously. The proposed method is compared to the approach in SVC test model and up to 0.3dB coding gains are achieved. Jun Xu 0040, Li Song 0001, Shibao Zheng, Xiaokang Yang 0001, Rong Xie 0004 |
ICME | 5 |