Li Song 0001

dblp:20/872-1 · DBLP profile ↗
← Back
199ranked-venue papers
2as first author
114since 2021 · last 2026
0000-0002-7124-5182ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 166 · 2 first-author · 95 since 2021Artificial intelligence and machine learning · 31 · 24 since 2021Computer networks · 13 · 8 since 2021Systems, architecture and hardware · 12 · 8 since 2021Databases, data management, data science and information retrieval · 10 · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021
YearPublicationVenuePosition
2026 D-FCGS: Feedforward Compression of Dynamic Gaussian Splatting for Free-Viewpoint Videos
abstract
Free-Viewpoint Video (FVV) enables immersive 3D experiences, but efficient compression of dynamic 3D representation remains a major challenge. Existing dynamic 3D Gaussian Splatting methods couple reconstruction with optimization-dependent compression and customized motion formats, limiting generalization and standardization. To address this, we propose D-FCGS, a novel Feedforward Compression framework for Dynamic Gaussian Splatting. Key innovations include: (1) a standardized Group-of-Frames (GoF) structure with I-P coding, leveraging sparse control points to extract inter-frame motion tensors; (2) a dual prior-aware entropy model that fuses hyperprior and spatial-temporal priors for accurate rate estimation; (3) a control-point-guided motion compensation mechanism and refinement network to enhance view-consistent fidelity. Trained on Gaussian frames derived from multi-view videos, D-FCGS generalizes across diverse scenes in a zero-shot fashion. Experiments show that it matches the rate-distortion performance of optimization-based methods, achieving over 40 times compression compared to the baseline while preserving visual quality across viewpoints. This work advances feedforward compression of dynamic 3DGS, facilitating scalable FVV transmission and storage for immersive applications.
Yan Zhao 0041, Qiang Wang 0061, Zhixin Xu, Li Song 0001, Zhengxue Cheng
AAAI5
2026 Volumetric Video on Demand System Based on Scalable Spacetime Gaussian Splatting
Jun Xu 0040, Bingcong Lu, Rong Xie 0004, Li Song 0001
ISCAS6
2026 LMM-VSC: Ultra-Low Bitrate Video Compression with Semantic Understanding
Chaolei Liu, Li Song 0001, Chuqin Zhou, Guo Lu
ISCAS2
2026 Diff-Band:Bandwidth Estimation in RTC via Diffusion-Based Offline Reinforcement Learning
Bingcong Lu, Zhengxue Cheng, Li Song 0001, Bingnan Duan, Jintao Fang
ISCAS4
2026 Unified Multimodal Retrieval Framework for Multimodal RAG
Tianyi Feng, Ruiyan Wang, Fei Huang 0002, Zhengxue Cheng, Rong Xie 0004, Li Song 0001
PAKDD (4)8
2026 Adaptive Bitrate Live Streaming over HTTP-FLV: A Practical System Perspective
Tong Meng, Bingcong Lu, Jinghao Yuan, Huanting Liu, Nailiang Wu, Zhou Sha, Changqing Yan, Jianrong Zhang, Jianxin Kuang, Li Song 0001
SIGCOMM13
2026 A fast video coding scheme based on perceptual rate distortion optimized preprocessing
Luheng Jia, Yifan Zang 0002, Jiyong Yu, Shuyuan Zhu, Li Song 0001, Kebin Jia
Multim. Syst.6
2026 OmniScaleSR: Unleashing Scale-Controlled Diffusion Prior for Faithful and Realistic Arbitrary-Scale Image Super-Resolution
abstract
Arbitrary-scale super-resolution (ASSR) overcomes the limitation of traditional super-resolution (SR) that works only at a fixed scale (e.g., ×4), enabling a single model to achieve arbitrary-scale SR. Most ASSR methods explicitly incorporate implicit neural representation (INR) to achieve ASSR, but INR’s inherently regression-driven feature extraction and aggregation nature restricts their capacity to synthesize meticulous details, leading to low realism. Recently, diffusion-based realistic image super-resolution (Real-ISR) methods leverage the pre-trained diffusion prior and have shown promising results at ×4 scale. We find that they could also achieve ASSR because the powerful pre-trained diffusion prior implicitly employs SR scale adaptation by encouraging the model to always generate high-realism images. However, due to the lack of explicit SR scale controls, the model fails to effectively manage the diffusion behavior according to different SR scales, causing either excessive hallucination or blurry results, especially for ultra-high magnification. To address these limitations, we proposeOmniScaleSR, a novel diffusion-based realistic arbitrary-scale super-resolution (Real-ASSR) method to achieve both high fidelity and high-realism ASSR. We introduce explicit diffusion-native SR scale controls, which could be elegantly coupled with the implicit scale adaptation, unleashing scale-controlled diffusion prior to dynamically managing the diffusion behavior in a content- and scale-aware manner. Furthermore, we incorporate multi-domain fidelity enhancement designs to achieve more faithful reconstruction. Extensive experiments on both bicubic degradation benchmarks and real-world datasets demonstrate that OmniScaleSR consistently outperforms state-of-the-art methods in terms of both fidelity and perceptual realism, with especially strong performance under high-magnification scenarios. Codes will be at https://github.com/chaixinning/OmniScaleSR.
Xinning Chai, Zhengxue Cheng, Hengsheng Zhang, Yingsheng Qin, Yucai Yang, Rong Xie 0004, Li Song 0001
IEEE Trans. Circuits Syst. Video Technol.8
2026 Distilling Complexity-Scalable Learned Image Compression Models via Neural Architecture Search
Shen Wang 0013, Zhengxue Cheng, Donghui Feng 0003, Cheems Wang, Qunshan Gu, Li Song 0001, Wenjun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2026 Diff-Restorer: Unleashing Visual Prompts for Diffusion-Based Universal Image Restoration
abstract
Image restoration aims to recover high-quality images from degraded observations, yet real-world degradations are complex, coupled, and difficult to model. Existing task-specific methods struggle to generalize beyond predefined degradation types, while recent all-in-one or prompt-based methods still face three key challenges: (1) they rely on task-specific training or fixed prompt pools, limiting adaptability to real-world and mixed degradations; (2) human-instruction or implicit-prompt mechanisms make them difficult to use in practice; and (3) they often fail to balance structural fidelity and perceptual realism. To address these issues, we propose Diff-Restorer, a diffusion-based universal image restoration framework that unifies diverse degradation handling within a single model. Diff-Restorer adaptively extracts decoupled visual prompts from a visual-language model (CLIP), including clear semantic and degradation embeddings. The clear semantic embeddings serve as content prompts to guide the diffusion model for generation, improving perceptual quality. The degradation embeddings as the task identifier modulate the Image-guided Control Module to generate structure control, ensuring faithfulness. Furthermore, we design a Task-aware Decoder to perform structural correction and convert the latent code to the pixel domain. Extensive experiments on various single, real-world, and mixed degradation tasks show that Diff-Restorer outperforms state-of-the-art methods in terms of generality, realism, and fidelity.
Hengsheng Zhang, Xinning Chai, Zhengxue Cheng, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2026 Exploring the Effect of Gaze and Distance Guidance in Room-Scale Virtual Reality
Haopeng Lu, Huiwen Ren, Li Song 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Hum. Mach. Syst.3
2026 Learned Image Compression via Local-to-Global Cross-Component Prior
abstract
Learned image compression (LIC) methods have shown promising results and achieved superior performance compared to traditional image compression methods. Due to the neglect of the utilization of cross-component correlations, there is still a potential for further performance improvement. In this paper, we first explore the inter-channel correlations of different color spaces and transform the image compression problem in RGB color space into that in YUV color space, which has cross-component prior information. We propose a novel image compression method that leverages local-to-global cross-component prior modeling, utilizing a cross-component attention mechanism to improve coding performance. First, we design the cross-component prior gate (CPG) to model the cross-component prior information based on attention mechanism. Inspired by common knowledge in data compression, luma component (Y) contains more details and textural/structural information compared to chroma components (UV). The proposed method can make full use of the cross-component guidance information from luma to chroma components to achieve effective image compression. Experimental results demonstrate that the proposed method can achieve superior performance compared to existing learned image compression methods. The proposed method can achieve 9.20% rate savings compared to the image compression standard Versatile Video Coding (VVC) Test Model (VTM-11.0) on Kodak dataset.
Wenhong Duan, Jiaye Fu, Li Song 0001, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Multim.5
2026 PoseTalk: Exploring Text- and Audio-Based Pose Control for One-Shot Talking Face Generation
abstract
Although audio-driven talking face generation has witnessed significant advances in recent years, two problems remain to be solved. First, the existing models cannot control the long-term head actions as humans expect, because the audio can only provide short-term cues, such as the rhythms and sentiments, for head movements. Second, generating long-term head poses and ensuring accurate lip motions remain challenging due to the difficulty in harnessing the optimization process for large-scale head movements and small-scale mouth motions. In this study, we propose a novel method to address these issues. First, to alleviate the limitations of audio conditions, we propose a Pose Latent Diffusion (PLD) model to generate head motions from two kinds of input modalities: the input audio and user-controlled text prompts. The audio provides short-term rhythm correspondence with the head movements, while the text prompts describe the long-term semantics of head motions. Second, we propose a refinement-based learning strategy to synthesize head movements and accurate lip motions using two cascaded networks, namely CoarseNet and RefineNet. The CoarseNet estimates coarse global motions to produce animated images with changed poses, and the RefineNet progressively estimates finer lip motions from low to high resolutions, yielding improved lip-synchronization performance. Experiments demonstrate that our method can achieve better pose diversity and realness compared to audio-based pose generation baselines, and our video generator model outperforms state-of-the-art methods in synthesizing natural head motions. Projects and demos are available at https://junleen.github.io/projects/posetalk .
Jun Ling, Rong Xie 0004, Li Song 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2026 A Hybrid Scheme for Face Video Compression
abstract
With the rapid development of social media, the amount of face video data has grown rapidly, making face video compression a hot research topic. Traditional video coding techniques do not discriminate video content and compress all videos in the same way, while talking head video compression should have more potential. Existing generative compression methods mostly adopt static reference frames, resulting in a decrease in fidelity caused by dynamic background or large pose change. In this article, we propose a hybrid compression scheme for face videos which combines traditional coding with generative compression. On the one hand, we sample and encode key frames with traditional codecs to provide dynamic reference frames which contain real-time background and motion information. On the other hand, we devise a deep video generation model to synthesize smooth video frames according to the extracted sparse keypoints. Combining the pixel-level recovery capability of traditional coding with the detail generation capability of deep generative models, our proposed hybrid scheme is able to implement high-fidelity face video compression at low bitrate in real time. Additionally, we also devise a Portrait Recovery module to recover the low-quality key frames, improving the reconstruction quality in low-bitrate scenarios. Extensive experiments show that our method has advantages over traditional codecs and existing generative compression methods in terms of both rate-distortion performance and coding complexity.
Anni Tang, Zhiyu Zhang 0010, Jun Ling, Rong Xie 0004, Li Song 0001
ACM Trans. Multim. Comput. Commun. Appl.6
2025 VRVVC: Variable-Rate NeRF-Based Volumetric Video Compression
abstract
Neural Radiance Field (NeRF)-based volumetric video has revolutionized visual media by delivering photorealistic Free-Viewpoint Video (FVV) experiences that provide audiences with unprecedented immersion and interactivity. However, the substantial data volumes pose significant challenges for storage and transmission. Existing solutions typically optimize NeRF representation and compression independently or focus on a single fixed rate-distortion (RD) tradeoff. In this paper, we propose VRVVC, a novel end-to-end joint optimization variable-rate framework for volumetric video compression that achieves variable bitrates using a single model while maintaining superior RD performance. Specifically, VRVVC introduces a compact tri-plane implicit residual representation for inter-frame modeling of long-duration dynamic scenes, effectively reducing temporal redundancy. We further propose a variable-rate residual representation compression scheme that leverages a learnable quantization and a tiny MLP-based entropy model. This approach enables variable bitrates through the utilization of predefined Lagrange multipliers to manage the quantization error of all latent representations. Finally, we present an end-to-end progressive training strategy combined with a multi-rate-distortion loss function to optimize the entire framework. Extensive experiments demonstrate that VRVVC achieves a wide range of variable bitrates within a single model and surpasses the RD performance of existing methods across various datasets.
Qiang Hu 0003, Houqiang Zhong, Zihan Zheng, Xiaoyun Zhang 0001, Zhengxue Cheng, Li Song 0001, Guangtao Zhai, Yanfeng Wang 0001
AAAI6
2025 L3TC: Leveraging RWKV for Learned Lossless Low-Complexity Text Compression
abstract
Learning-based probabilistic models can be combined with an entropy coder for data compression. However, due to the high complexity of learning-based models, their practical application as text compressors has been largely overlooked. To address this issue, our work focuses on a low-complexity design while maintaining compression performance. We introduce a novel Learned Lossless Low-complexity Text Compression method (L3TC). Specifically, we conduct extensive experiments demonstrating that RWKV models achieve the fastest decoding speed with a moderate compression ratio, making it the most suitable backbone for our method. Second, we propose an outlier-aware tokenizer that uses a limited vocabulary to cover frequent tokens while allowing outliers to bypass the prediction and encoding. Third, we propose a novel high-rank reparameterization strategy that enhances the learning capability during training without increasing complexity during inference. Experimental results validate that our method achieves 48% bit saving compared to gzip compressor. Besides, L3TC offers compression performance comparable to other learned compressors, with a 50x reduction in model parameters. More importantly, L3TC is the fastest among all learned compressors, providing real-time decoding speeds up to megabytes per second.
Junxuan Zhang, Zhengxue Cheng, Yan Zhao 0041, Dajiang Zhou, Guo Lu, Li Song 0001
AAAI7
2025 Controllable Distortion-Perception Tradeoff Through Latent Diffusion for Neural Image Compression
abstract
Neural image compression often faces a challenging trade-off among rate, distortion and perception. While most existing methods typically focus on either achieving high pixel-level fidelity or optimizing for perceptual metrics, we propose a novel approach that simultaneously addresses both aspects for a fixed neural image codec. Specifically, we introduce a plug-and-play module at the decoder side that leverages a latent diffusion process to transform the decoded features, enhancing either low distortion or high perceptual quality without altering the original image compression codec. Our approach facilitates fusion of original and transformed features without additional training, enabling users to flexibly adjust the balance between distortion and perception during inference. Extensive experimental results demonstrate that our method significantly enhances the pretrained codecs with a wide, adjustable distortion-perception range while maintaining their original compression capabilities. For instance, we can achieve more than 150% improvement in LPIPS-BDRate without sacrificing more than 1 dB in PSNR.
Chuqin Zhou, Guo Lu, Jiangchuan Li, Zhengxue Cheng, Li Song 0001, Wenjun Zhang 0001
AAAI6
2025 Linear Attention Modeling for Learned Image Compression
abstract
Recent years, learned image compression has made tremendous progress to achieve impressive coding efficiency. Its coding gain mainly comes from non-linear neural network-based transform and learnable entropy modeling. However, most studies focus on a strong backbone, and few studies consider a low complexity design. In this paper, we propose LALIC, a linear attention modeling for learned image compression. Specially, we propose to use Bi-RWKV blocks, by utilizing the Spatial Mix and Channel Mix modules to achieve more compact feature extraction, and apply the Conv based Omni-Shift module to adapt to two-dimensional latent representation. Furthermore, we propose a RWKV-based Spatial-Channel ConTeXt model (RWKV-SCCTX), that leverages the Bi-RWKV to modeling the correlation between neighboring features effectively. To our knowledge, our work is the first work to utilize efficient Bi-RWKV models with linear attention for learned image compression. Experimental results demonstrate that our method achieves competitive RD performances by outperforming VTM-9.1 by -15.26%, -15.41%, -17.63% in BD-rate on Kodak, CLIC and Tecnick datasets. The code is available at https://github.com/sjtu-medialab/RwkvCompress.
Donghui Feng 0003, Zhengxue Cheng, Shen Wang 0013, Ronghua Wu, Hongwei Hu, Guo Lu, Li Song 0001
CVPR7
2025 4DGC: Rate-Aware 4D Gaussian Compression for Efficient Streamable Free-Viewpoint Video
abstract
3D Gaussian Splatting (3DGS) has substantial potential for enabling photorealistic Free-Viewpoint Video (FVV) experiences. However, the vast number of Gaussians and their associated attributes poses significant challenges for storage and transmission. Existing methods typically handle dynamic 3DGS representation and compression separately, neglecting motion information and the rate-distortion (RD) trade-off during training, leading to performance degradation and increased model redundancy. To address this gap, we propose 4DGC, a novel rate-aware 4D Gaussian compression framework that significantly reduces storage size while maintaining superior RD performance for FVV. Specifically, 4DGC introduces a motion-aware dynamic Gaussian representation that utilizes a compact motion grid combined with sparse compensated Gaussians to exploit inter-frame similarities. This representation effectively handles large motions, preserving quality and reducing temporal redundancy. Furthermore, we present an end-to-end compression scheme that employs differentiable quantization and a tiny implicit entropy model to compress the motion grid and compensated Gaussians efficiently. The entire framework is jointly optimized using a rate-distortion trade-off. Extensive experiments demonstrate that 4DGC supports variable bitrates and consistently outperforms existing methods in RD performance across multiple datasets.
Qiang Hu 0003, Zihan Zheng, Houqiang Zhong, Sihua Fu, Li Song 0001, Xiaoyun Zhang 0001, Guangtao Zhai, Yanfeng Wang 0001
CVPR5
2025 Semantic and Temporal Integration in Latent Diffusion Space for High-Fidelity Video Super-Resolution
abstract
Recent advancements in video super-resolution (VSR) models have demonstrated impressive results in enhancing low-resolution videos. However, due to limitations in adequately controlling the generation process, achieving high fidelity alignment with the low-resolution input while maintaining temporal consistency across frames remains a significant challenge. In this work, we propose Semantic and Temporal Guided Video Super-Resolution (SeTe-VSR), a novel approach that incorporates both semantic and temporal-spatio guidance in the latent diffusion space to address these challenges. By incorporating high-level semantic information and integrating spatial and temporal information, our approach achieves a seamless balance between recovering intricate details and ensuring temporal coherence. Our method not only preserves high-reality visual content but also significantly enhances fidelity. Extensive experiments demonstrate that SeTe-VSR outperforms existing methods in terms of detail recovery and perceptual quality, highlighting its effectiveness for complex video super-resolution tasks.
Xinning Chai, Zhengxue Cheng, Rong Xie 0004, Li Song 0001
ICME7
2025 Serial Low-rank Adaptation of Vision Transformer
abstract
Fine-tuning large pre-trained vision foundation models in a parameter-efficient manner is critical for downstream vision tasks, considering the practical constraints of computational and storage costs. Low-rank adaptation (LoRA) is a well-established technique in this domain, achieving impressive efficiency by reducing the parameter space to a low-rank form. However, developing more advanced low-rank adaptation methods to reduce parameters and memory requirements remains a significant challenge in resource-constrained application scenarios. In this study, we consider on top of the commonly used vision transformer and propose Serial LoRA, a novel LoRA variant that introduces a shared low-rank matrix serially composite with the attention mechanism. Such a design extracts the underlying commonality of parameters in adaptation, significantly reducing redundancy. Notably, Serial LoRA uses only ${\color {Magenta}{\text{1/4}}}$ parameters of LoRA but achieves comparable performance in most cases. We conduct extensive experiments on a range of vision foundation models with the transformer structure, and the results confirm consistent superiority of our method.
Houqiang Zhong, Shaocheng Shen, Ke Cai, Zhenlong Wu, Jiangchao Yao, Xiaoyun Zhang 0001, Li Song 0001, Qiang Hu 0003
ICME9
2025 DiffFERV: Diffusion-based Facial Editing of Real Videos
abstract
Face video editing presents significant challenges, requiring precise preservation of facial identity, temporal consistency, and background details. Existing methods encounter three major challenges: difficulty in achieving accurate facial reconstruction, struggles with challenging real-world videos and reliance on a crop-edit-stitch paradigm that confines editing to localized facial regions. In response, we introduce DiffFERV, a novel diffusion-based framework for realistic face video editing that addresses these limitations through three core contributions. (1) A specialization stage that extends large Text-to-Image (T2I) models' general prior to faces while retaining their broad generative capabilities. This enables robust performance on non-aligned and challenging face images. (2) Temporal modeling, implemented through two distinct attention mechanisms, complements the specialization stage to ensure joint and temporally consistent processing of video frames. (3) Finally, we present a holistic editing pipeline and the concept of preservation features, which leverages our model’s enhanced priors and temporal mechanisms to achieve faithful edits of entire video frames without the need for cropping, excelling even in real-world scenarios. Extensive experiments demonstrate that DiffFERV achieves state-of-the-art performance in both reconstruction and editing tasks.
Xiangyi Chen, Li Song 0001
IJCAI3
2025 Rate-Aware Learned Speech Compression
abstract
The rapid rise of real-time communication and large language models has significantly increased the importance of speech compression. Deep learning-based neural speech codecs have outperformed traditional signal-level speech codecs in terms of rate-distortion (RD) performance. Typically, these neural codecs employ an encoder-quantizer-decoder architecture, where audio is first converted into latent code feature representations and then into discrete tokens. However, this architecture exhibits insufficient RD performance due to two main drawbacks: (1) the inadequate performance of the quantizer, challenging training processes, and issues such as codebook collapse; (2) the limited representational capacity of the encoder and decoder, making it difficult to meet feature representation requirements across various bitrates. In this paper, we propose a rate-aware learned speech compression scheme that replaces the quantizer with an advanced channel-wise entropy model to improve RD performance, simplify training, and avoid codebook collapse. We employ multi-scale convolution and linear attention mixture blocks to enhance the representational capacity and flexibility of the encoder and decoder. Experimental results demonstrate that the proposed method achieves state-of-the-art RD performance, obtaining 53.51% BD-Rate bitrate saving in average, and achieves 0.26 BD-VisQol and 0.44 BD-PESQ gains.
Zhengxue Cheng, Guangchuan Chi, Yuelin Hu, Li Song 0001
ISCAS6
2025 MoRLACS: A Monocular RGBD-based Locomotion Approach for CAVE Systems
abstract
Navigation within Cave Automatic Virtual Environment (CAVE) systems often faces challenges due to limited physical space and the necessity for seamless user interaction. Traditional solutions typically rely on multi-view tracking systems or constrained locomotion techniques, which can interrupt immersion and hinder usability. In this paper, we introduce MoRLACS, a novel locomotion approach for CAVE systems that leverages a single RGBD camera. This hybrid framework integrates small-scale physical walking with controller-based large-scale exploration through a tailored guidance method. By accurately tracking the user's head position in the real world and synchronizing it with the virtual camera, MoRLACS enables natural walking within confined CAVE spaces and supports extended interaction in larger virtual environments. Preliminary user experiments demonstrate the approach's effectiveness, revealing improvements in usability and a heightened sense of presence. These findings underscore the potential of MoRLACS to enrich user experiences in immersive CAVE settings and offer valuable design insights for integrating 3D sensor data into multimedia interaction frameworks.
Haopeng Lu, Qian Yin 0002, Li Song 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001
ICMR4
2025 MultiEgo: A Multi-View Egocentric Video Dataset for 4D Scene Reconstruction
abstract
Multi-view egocentric dynamic scene reconstruction holds significant research value for applications in holographic documentation of social interactions. However, existing reconstruction datasets focus on static multi-view or single-egocentric view setups, lacking multi-view egocentric datasets for dynamic scene reconstruction. Therefore, we present MultiEgo, the first multi-view egocentric dataset for 4D dynamic scene reconstruction. The dataset comprises five canonical social interaction scenes: meetings, performances, and a presentation. Each scene provides five authentic egocentric videos captured by participants wearing AR glasses. We design a hardware-based data acquisition system and processing pipeline, achieving sub-millisecond temporal synchronization across views, coupled with accurate pose annotations. Experiment validation demonstrates the practical utility and effectiveness of our dataset for free-viewpoint video (FVV) applications, establishing MultiEgo as a foundational resource for advancing multi-view egocentric dynamic scene reconstruction research.
Bate Li, Houqiang Zhong, Zhengxue Cheng, Qiang Hu 0003, Qiang Wang 0061, Li Song 0001, Wenjun Zhang 0001
ACM Multimedia6
2025 SemanticGarment: Semantic-Controlled Generation and Editing of 3D Gaussian Garments
abstract
3D digital garment generation and editing play a pivotal role in fashion design, virtual try-on, and gaming. Traditional methods struggle to meet the growing demand due to technical complexity and high resource costs. Learning-based approaches offer faster, more diverse garment synthesis based on specific requirements and reduce human efforts and time costs. However, they still face challenges such as inconsistent multi-view geometry or textures and heavy reliance on detailed garment topology and manual rigging. We propose SemanticGarment, a 3D Gaussian-based method that realizes high-fidelity 3D garment generation from text or image prompts and supports semantic-based interactive editing for flexible user customization. To ensure multi-view consistency and garment fitting, we propose to leverage structural human priors for the generative model by introducing a 3D semantic clothing model, which initializes the geometry structure and lays the groundwork for view-consistent garment generation and editing. Without the need to regenerate or rely on existing mesh templates, our approach allows for rapid and diverse modifications to existing Gaussians, either globally or within a local region. To address the artifacts caused by self-occlusion for garment reconstruction based on single image, we develop a self-occlusion optimization strategy to mitigate holes and artifacts that arise when directly animating self-occluded garments. Extensive experiments are conducted to demonstrate our superior performance in 3D garment generation and editing.
Ruiyan Wang, Zhengxue Cheng, Zonghao Lin, Jun Ling, Yanru An, Rong Xie 0004, Li Song 0001
ACM Multimedia8
2025 PA-HOI: A Physics-Aware Human and Object Interaction Dataset
Ruiyan Wang, Lin Zuo, Zonghao Lin, Qiang Wang 0061, Zhengxue Cheng, Rong Xie 0004, Jun Ling, Li Song 0001
ACM Multimedia8
2025 AnimeColor: Reference-based Animation Colorization with Diffusion Transformers
abstract
Animation colorization plays a vital role in animation production, yet existing methods struggle to achieve color accuracy and temporal consistency. To address these challenges, we propose AnimeColor, a novel reference-based animation colorization framework leveraging Diffusion Transformers (DiT). Our approach integrates sketch sequences into a DiT-based video diffusion model, enabling sketch-controlled animation generation. We introduce two key components: a High-level Color Extractor (HCE) to capture semantic color information and a Low-level Color Guider (LCG) to extract fine-grained color details from reference images. These components work synergistically to guide the video diffusion process. Additionally, we employ a multi-stage training strategy to maximize the utilization of reference image color information. Extensive experiments demonstrate that AnimeColor outperforms existing methods in color accuracy, sketch alignment, temporal consistency, and visual quality. Our framework not only advances the state of the art in animation colorization but also provides a practical solution for industrial applications. The code will be made publicly available at https://github.com/IamCreateAI/AnimeColor.
Liyao Wang, Danni Wu, Zuzeng Lin, Feng Wang 0015, Li Song 0001
ACM Multimedia7
2025 FreeInsert: Personalized Object Insertion with Geometric and Style Control
abstract
Text-to-image diffusion models have made significant progress in image generation, allowing for effortless customized generation. However, existing image editing methods still face certain limitations when dealing with personalized image composition tasks. First, there is the issue of lack of geometric control over the inserted objects. Current methods are confined to 2D space and typically rely on textual instructions, making it challenging to maintain precise geometric control over the objects. Second, there is the challenge of style consistency. Existing methods often overlook the style consistency between the inserted object and the background, resulting in a lack of realism. In addition, the challenge of inserting objects into images without extensive training remains significant. To address these issues, we propose FreeInsert, a novel training-free framework that customizes object insertion into arbitrary scenes by leveraging 3D geometric information. Benefiting from the advances in existing 3D generation models, we first convert the 2D object into 3D, perform interactive editing at the 3D level, and then re-render it into a 2D image from a specified view. This process introduces geometric controls such as shape or view. The rendered image, serving as geometric control, is combined with style and content control achieved through diffusion adapters, ultimately producing geometrically controlled, style-consistent edited images via the diffusion model.
Rong Xie 0004, Li Song 0001
ACM Multimedia5
2025 H3D-DGS: Exploring Heterogeneous 3D Motion Representation for Deformable 3D Gaussian Splatting
abstract
Dynamic scene reconstruction poses a persistent challenge in 3D vision. Deformable 3D Gaussian Splatting has emerged as an effective method for this task, offering real-time rendering and high visual fidelity. This approach decomposes a dynamic scene into a static representation in a canonical space and time-varying scene motion. Scene motion is defined as the collective movement of all Gaussian points, and for compactness, existing approaches commonly adopt implicit neural fields or sparse control points. However, these methods predominantly rely on gradient-based optimization for all motion information. Due to the high degree of freedom, they struggle to converge on real-world datasets exhibiting complex motion. To preserve the compactness of motion representation and address convergence challenges, this paper proposes heterogeneous 3D control points, termed \textbf{H3D control points}, whose attributes are obtained using a hybrid strategy combining optical flow back-projection and gradient-based methods. This design decouples directly observable motion components from those that are geometrically occluded. Specifically, components of 3D motion that project onto the image plane are directly acquired via optical flow back projection, while unobservable portions are refined through gradient-based optimization. Experiments on the Neu3DV and CMU-Panoptic datasets demonstrate that our method achieves superior performance over state-of-the-art deformable 3D Gaussian splatting techniques. Remarkably, our method converges within just 100 iterations and achieves a per-frame processing speed of 2 seconds on a single NVIDIA RTX 4070 GPU.
Yunuo Chen 0002, Guo Lu, Cheems Wang, Qunshan Gu, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001
NeurIPS7
2025 A Multi-Grid Implicit Neural Representation for Multi-View Videos
Qingyue Ling, Zhengxue Cheng, Donghui Feng 0003, Shen Wang 0013, Guo Lu, Heming Sun, Jiro Katto, Li Song 0001
PCS9
2025 AlignGS: Aligning Geometry and Semantics for Robust Indoor Reconstruction from Sparse Views
abstract
The demand for semantically rich 3D models of indoor scenes is rapidly growing, driven by applications in augmented reality, virtual reality, and robotics. However, creating them from sparse views remains a challenge due to geometric ambiguity. Existing methods often treat semantics as a passive feature painted on an already-formed, and potentially flawed, geometry. We posit that for robust sparse-view reconstruction, semantic understanding instead be an active, guiding force. This paper introduces AlignGS, a novel framework that actualizes this vision by pioneering a synergistic, end-to-end optimization of geometry and semantics. Our method distills rich priors from 2D foundation models and uses them to directly regularize the 3D representation through a set of novel semantic-to-geometry guidance mechanisms, including depth consistency and multifaceted normal regularization. Extensive evaluations on standard benchmarks demonstrate that our approach achieves state-of-the-art results in novel view synthesis and produces reconstructions with superior geometric accuracy. The results validate that leveraging semantic priors as a geometric regularizer leads to more coherent and complete 3D models from limited input views. Our code is avaliable at https://github.com/MediaX-SJTU/AlignGS.
Yijie Gao, Houqiang Zhong, Tianchi Zhu, Zhengxue Cheng, Qiang Hu 0003, Li Song 0001
VCIP6
2025 Lightweight High-Fidelity Low-Bitrate Talking Face Compression for 3D Video Conference
abstract
The demand for immersive and interactive communication has driven advancements in 3D video conferencing, yet achieving high-fidelity 3D talking face representation at low bitrates remains a challenge. Traditional 2D video compression techniques fail to preserve fine-grained geometric and appearance details, while implicit neural rendering methods like NeRF suffer from prohibitive computational costs. To address these challenges, we propose a lightweight, high-fidelity, low-bitrate 3D talking face compression framework that integrates FLAME-based parametric modeling with 3DGS neural rendering. Our approach transmits only essential facial metadata in real time, enabling efficient reconstruction with a Gaussian-based head model. Additionally, we introduce a compact representation and compression scheme, including Gaussian attribute compression and MLP optimization, to enhance transmission efficiency. Experimental results demonstrate that our method achieves superior rate-distortion performance, delivering high-quality facial rendering at extremely low bitrates, making it well-suited for real-time 3D video conferencing applications.
Jianglong Li, Bingcong Lu, Zhengxue Cheng, Hongwei Hu, Ronghua Wu, Li Song 0001
VCIP7
2025 PrismGS: Physically-Grounded Anti-Aliasing for High-Fidelity Large-Scale 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) has recently enabled real-time photorealistic rendering in compact scenes, but scaling to large urban environments introduces severe aliasing artifacts and optimization instability, especially under high-resolution (e.g., 4K) rendering. These artifacts, manifesting as flickering textures and jagged edges, arise from the mismatch between Gaussian primitives and the multi-scale nature of urban geometry. While existing "divide-and-conquer" pipelines address scalability, they fail to resolve this fidelity gap. In this paper, we propose PrismGS, a physically-grounded regularization framework that improves the intrinsic rendering behavior of 3D Gaussians. PrismGS integrates two synergistic regularizers. The first is pyramidal multi-scale supervision, which enforces consistency by supervising the rendering against a pre-filtered image pyramid. This compels the model to learn an inherently anti-aliased representation that remains coherent across different viewing scales, directly mitigating flickering textures. This is complemented by an explicit size regularization that imposes a physically-grounded lower bound on the dimensions of the 3D Gaussians. This prevents the formation of degenerate, view-dependent primitives, leading to more stable and plausible geometric surfaces and reducing jagged edges. Our method is plug-and-play and compatible with existing pipelines. Extensive experiments on MatrixCity, Mill-19, and UrbanScene3D demonstrate that PrismGS achieves state-of-the-art performance, yielding significant PSNR gains around 1.5 dB against CityGaussian, while maintaining its superior quality and robustness under demanding 4K rendering.
Houqiang Zhong, Zhenglong Wu, Sihua Fu, Zihan Zheng, Xin Jin 0014, Xiaoyun Zhang 0001, Li Song 0001, Qiang Hu 0003
VCIP7
2025 CPIG: Controlling the Portrait Image Generation by Distilling 3D GAN's Latent Directions
abstract
ABSTRACT Synthesis of 3D‐aware facial images from latent spaces has garnered significant attention in multimedia content generation due to its ability to model images with rich semantics and diverse appearances. However, existing methods often rely on labeled data or suffer from incomplete attribute control and ambiguous latent space semantics. This paper proposes an efficient semantic distillation method that learns attribute directions of pre‐trained 3D GAN models, without the supervised semantic labels. We consider the latent space of GAN models as the mixture of two featured subspaces, namely the geometry‐aware space and appearance‐aware space. Following this hypothesis, we define two sets of learnable latent bases and use linear composition to represent controllable geometry and appearance feature space, respectively. To learn semantic‐wise latent bases for attribute‐controllable image generation, we design a framework and propose a three‐staged training strategy, which optimizes the appearance‐aware and the geometry‐aware latent bases. With the two sets of latent bases, we obtain the combined latent vectors using different weights for those bases and synthesize images with specified attributes. Compared to existing methods, our approach eliminates the need for labeled data and enables more controllable attribute disentanglement while ensuring identity consistency, which can be directly applied to real‐world scenarios such as virtual avatars and augmented reality applications. Experiments demonstrate the effectiveness and insight of our approach in aiding a better understanding of the latent space of 3D GANs.
Ruiyan Wang, Jun Ling, Rong Xie 0004, Li Song 0001
IET Image Process.4
2025 Visual information fidelity based frame level rate control for H.265/HEVC
Luheng Jia, Haoqiang Ren, Zuhai Zhang, Li Song 0001, Kebin Jia
Signal Process. Image Commun.4
2025 SSP-IR: Semantic and Structure Priors for Diffusion-Based Realistic Image Restoration
abstract
Realistic image restoration is a crucial task in computer vision, and diffusion-based models for image restoration have garnered significant attention due to their ability to produce realistic results. Restoration can be seen as a controllable generation conditioning on priors. However, due to the severity of image degradation, existing diffusion-based restoration methods cannot fully exploit priors from low-quality images and still have many challenges in perceptual quality, semantic fidelity, and structure accuracy. Based on the challenges, we introduce a novel image restoration method, SSP-IR. Our approach aims to fully exploit semantic and structure priors from low-quality images to guide the diffusion model in generating semantically faithful and structurally accurate natural restoration results. Specifically, we integrate the visual comprehension capabilities of Multimodal Large Language Models (explicit) and the visual representations of the original image (implicit) to acquire accurate semantic prior. To extract degradation-independent structure prior, we introduce a Processor with RGB and FFT constraints to extract structure prior from the low-quality images, guiding the diffusion model and preventing the generation of unreasonable artifacts. Lastly, we employ a multi-level attention mechanism to integrate the acquired semantic and structure priors. The qualitative and quantitative results demonstrate that our method outperforms other state-of-the-art methods overall on both synthetic and real-world datasets. Our project page ishttps://zyhrainbow.github.io/projects/SSP-IR.
Hengsheng Zhang, Zhengxue Cheng, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Instance-Adaptive Spatial-Temporal Enhancement for Efficient Video Compression
abstract
Efficiently compressing HD/UHD content has long been challenging due to high bitrate costs. Instance-adaptive enhancement methods try to tackle this issue by compressing a video at reduced resolution and enhancing it using a neural model specifically overfitted for this video. However, existing methods focus solely on spatial super-resolution (SR) and under-utilize the videos' temporal redundancy. Their limited management of the model's updated parameters also causes excessive overfitting overheads. Therefore, this paper introduces IASTE, the first instance-adaptive enhancement method based on spatial-temporal enhancement (STE), and incorporates low-rank adaptation (LoRA) for efficient model overfitting. Specifically, we downscale videos spatially and temporally to reduce the data volume and achieve efficient video compression. Then, we overfit a specific STE model for each video and use it to enhance the decoded video's spatiotemporal resolution. Leveraging the video swin transformer's strong capability in capturing spatiotemporal correlations, we design a lightweight and efficient model to implement video STE. The model is overfitted for each video using LoRA. By freezing the pre-trained model and selectively updating a few low-rank matrices, the bitrate overhead for model storage can be mitigated. Experiments prove that compared to directly compressing high-frame-rate (HFR), high-resolution (HR) videos, our method achieves around 30% BD-Rate gains on the CTC and UVG datasets, about 15% gains on the YoutubeUGC dataset, and about 10% gains on the ultra-long videos in the Xiph dataset.
Yan Zhao 0041, Zhengxue Cheng, Jiangchuan Li, Donghui Feng 0003, Qunshan Gu, Cheems Wang, Guo Lu, Li Song 0001
IEEE Trans. Image Process.8
2025 Implicit-Explicit Integrated Representations for Multi-View Video Compression
abstract
With the increasing consumption of 3D displays and virtual reality, multi-view video has become a promising format. However, its high resolution and multi-camera shooting result in a substantial increase in data volume, making storage and transmission a challenging task. To tackle these difficulties, we propose an implicit-explicit integrated representation for multi-view video compression. Specifically, we first use the explicit representation-based 2D video codec to encode one of the source views. Subsequently, we propose employing the implicit neural representation (INR)-based codec to encode the remaining views. The implicit codec takes the time and view index of multi-view video as coordinate input and generates the corresponding implicit reconstruction frames. To enhance the compressibility, we introduce a multi-level feature grid embedding and a fully convolutional architecture into the implicit codec. These components facilitate coordinate-feature and feature-RGB mapping, respectively. To further enhance the reconstruction quality from the INR codec, we leverage the high-quality reconstructed frames from the explicit codec to achieve inter-view compensation. Finally, the compensated results are fused with the implicit reconstructions from the INR to obtain the final reconstructed frames. Our proposed framework combines the strengths of both implicit neural representation and explicit 2D codec. Extensive experiments conducted on public datasets demonstrate that the proposed framework can achieve comparable or even superior performance to the latest multi-view video compression standard MIV and other INR-based schemes in terms of view compression and scene modeling. The source code can be found at https://github.com/zc-lynen/MV-IERV.
Guo Lu, Rong Xie 0004, Li Song 0001
IEEE Trans. Image Process.5
2024 Depth-Guided Robust and Fast Point Cloud Fusion NeRF for Sparse Input Views
abstract
Novel-view synthesis with sparse input views is important for real-world applications like AR/VR and autonomous driving. Recent methods have integrated depth information into NeRFs for sparse input synthesis, leveraging depth prior for geometric and spatial understanding. However, most existing works tend to overlook inaccuracies within depth maps and have low time efficiency. To address these issues, we propose a depth-guided robust and fast point cloud fusion NeRF for sparse inputs. We perceive radiance fields as an explicit voxel grid of features. A point cloud is constructed for each input view, characterized within the voxel grid using matrices and vectors. We accumulate the point cloud of each input view to construct the fused point cloud of the entire scene. Each voxel determines its density and appearance by referring to the point cloud of the entire scene. Through point cloud fusion and voxel grid fine-tuning, inaccuracies in depth values are refined or substituted by those from other views. Moreover, our method can achieve faster reconstruction and greater compactness through effective vector-matrix decomposition. Experimental results underline the superior performance and time efficiency of our approach compared to state-of-the-art baselines.
Shuai Guo 0002, Qiuwen Wang, Yijie Gao, Rong Xie 0004, Li Song 0001
AAAI5
2024 Disentangled Clothed Avatar Generation from Text Descriptions
Jionghao Wang, Yuan Liu 0025, Zhiyang Dou, Zhengming Yu, Yongqing Liang 0001, Cheng Lin 0001, Rong Xie 0004, Li Song 0001, Xin Li 0003, Wenping Wang 0001
ECCV (52)8
2024 Hdrtvformer: Efficient Sdrtv-to-Hdrtv via Affine Transformation and Spatial-Aware Transformer
abstract
Recent works on reconstructing HDR videos in display format (HDRTV) suffer from high computational and memory requirements because they learn the SDRTV-to-HDRTV mapping directly in 4K resolution. This paper proposes an efficient SDRTV-to-HDRTV model (HDRTVFormer) that decomposes the HDRTV restoration into SDRTV-to-HDRTV Domain Mapping and HDRTV Refinement. SDRTV-to-HDRTV Domain Mapping is an affine transformation-based model that learns SDRTV-to-HDRTV affine coefficients in low-resolution space, achieving rapid processing times. To enhance the accuracy of the predicted affine coefficients, the model introduces global information-modulated feature extraction blocks and a detail guidance upsampling module. For HDRTV Refinement, we propose a spatial-aware Transformer to refine the luminance and color details. We modify the self-attention and feed-forward network of Transformer blocks to improve efficiency and feature representations. Experimental results have demonstrated that our method outperforms other state-of-the-art works in performance and efficiency.
Hengsheng Zhang, Xinning Chai, Rong Xie 0004, Li Song 0001
ICASSP5
2024 A New People-Object Interaction Dataset and NVS Benchmarks
abstract
Recently, NVS in human-object interaction scenes has received increasing attention. Existing human-object interaction datasets mainly consist of static data with limited views, offering only RGB images or videos, mostly containing interactions between a single person and objects. Moreover, these datasets exhibit complexities in lighting environments, poor synchronization, and low resolution, hindering high-quality human-object interaction studies. In this paper, we introduce a new people-object interaction dataset that comprises 38 series of 30-view multi-person or single-person RGB-D video sequences, accompanied by camera parameters, foreground masks, SMPL models, some point clouds, and mesh files. Video sequences are captured by 30 Kinect Azures, uniformly surrounding the scene, each in 4 K resolution 25 FPS, and lasting for 1~19 seconds. Meanwhile, we evaluate some SOTA NVS models on our dataset to establish the NVS benchmarks. We hope our work can inspire further research in human-object interaction.
Shuai Guo 0002, Houqiang Zhong, Qiuwen Wang, Yijie Gao, Jiajing Yuan, Rong Xie 0004, Li Song 0001
ICIP9
2024 JOINTRF: End-To-End Joint Optimization for Dynamic Neural Radiance Field Representation and Compression
abstract
Neural Radiance Field (NeRF) excels in photo-realistically static scenes, inspiring numerous efforts to facilitate volumetric videos. However, rendering dynamic and long-sequence radiance fields remains challenging due to the significant data required to represent volumetric videos. In this paper, we propose a novel end-to-end joint optimization scheme of dynamic NeRF representation and compression, called JointRF, thus achieving significantly improved quality and compression efficiency against the previous methods. Specifically, JointRF employs a compact residual feature grid and a coefficient feature grid to represent the dynamic NeRF. This representation handles large motions without compromising quality while concurrently diminishing temporal redundancy. We also introduce a sequential feature compression subnetwork to further reduce spatial-temporal redundancy. Finally, the representation and compression subnetworks are end-to-end trained combined within the JointRF. Extensive experiments demonstrate that JointRF can achieve superior compression performance across various datasets.
Zihan Zheng, Houqiang Zhong, Qiang Hu 0003, Xiaoyun Zhang 0001, Li Song 0001, Ya Zhang 0002, Yanfeng Wang 0001
ICIP5
2024 Neural Rate Control for Learned Video Compression
abstract
The learning-based video compression method has made significant progress in recent years, exhibiting promising compression performance compared with traditional video codecs. However, prior works have primarily focused on advanced compression architectures while neglecting the rate control technique. Rate control can precisely control the coding bitrate with optimal compression performance, which is a critical technique in practical deployment. To address this issue, we present a fully neural network-based rate control system for learned video compression methods. Our system accurately encodes videos at a given bitrate while enhancing the rate-distortion performance. Specifically, we first design a rate allocation model to assign optimal bitrates to each frame based on their varying spatial and temporal characteristics. Then, we propose a deep learning-based rate implementation network to perform the rate-parameter mapping, precisely predicting coding parameters for a given rate. Our proposed rate control system can be easily integrated into existing learning-based video compression methods. The extensive experimental results show that the proposed method achieves accurate rate control on several baseline methods while also improving overall rate-distortion performance.
Guo Lu, Yunuo Chen 0002, Shen Wang 0013, Yibo Shi, Jing Wang 0194, Li Song 0001
ICLR7
2024 SingAvatar: High-fidelity Audio-driven Singing Avatar Synthesis
abstract
Generating photo-realistic avatars from audio plays an important role in extended reality (XR) and metaverse. In this paper, we lift the input audio from speech to singing, which has been rarely studied. The significant distinction between singing and talking poses great challenges for adapting talking face generation methods to the singing regime. To address this, we propose a high-fidelity singing avatar synthesis method called SingAvatar. Besides the audio, we incorporate vocal conditions involving phonemes and variance to alleviate the ambiguity of learning the singing-to-face mapping. Concretely, we tailor a two-stage pipeline: singing voice synthesis and portrait generation from the synthesized audio and auxiliary vocal conditions. Further, we curate a fine-grained singing head dataset containing singing videos with synchronized audio and accurate vocal conditions. In experiments, SingAvatar outperforms competing methods regarding audio-mouth synchronization, the naturalness of head movements, and controllability over the results. The code and dataset will be made publicly available.
Anni Tang, Jun Ling, Huiheng Liao, Yunhui Zhu, Li Song 0001
ICME7
2024 Efficient Dynamic-NeRF Based Volumetric Video Coding with Rate Distortion Optimization
abstract
Volumetric videos, benefiting from immersive 3D realism and interactivity, hold vast potential for various applications, while the tremendous data volume poses significant challenges for compression. Recently, NeRF has demonstrated remarkable potential in volumetric video compression thanks to its simple representation and powerful 3D modeling capabilities, where a notable work is ReRF. However, ReRF separates the modeling from compression process, resulting in suboptimal compression efficiency. In contrast, in this paper, we propose a volumetric video compression method based on dynamic NeRF in a more compact manner. Specifically, we decompose the NeRF representation into the coefficient fields and the basis fields, incrementally updating the basis fields in the temporal domain to achieve dynamic modeling. Additionally, we perform end-to-end joint optimization on the modeling and compression process to further improve the compression efficiency. Extensive experiments demonstrate that our method achieves higher compression efficiency compared to ReRF on various datasets.
Zhiyu Zhang 0010, Guo Lu, Huanxiong Liang, Anni Tang, Qiang Hu 0003, Li Song 0001
ICME6
2024 Rate-aware Compression for NeRF-based Volumetric Video
abstract
The neural radiance fields (NeRF) have advanced the development of 3D volumetric video technology, but the large data volumes they involve pose significant challenges for storage and transmission. To address these problems, the existing solutions typically compress these NeRF representations after the training stage, leading to a separation between representation training and compression. In this paper, we try to directly learn a compact NeRF representation for volumetric video in the training stage based on the proposed rate-aware compression framework. Specifically, for volumetric video, we use a simple yet effective modeling strategy to reduce temporal redundancy for the NeRF representation. Then, during the training phase, an implicit entropy model is utilized to estimate the bitrate of the NeRF representation. This entropy model is then encoded into the bitstream to assist in the decoding of the NeRF representation. This approach enables precise bitrate estimation, thereby leading to a compact NeRF representation.Furthermore, we propose an adaptive quantization strategy and learn the optimal quantization step for the NeRF representations. Finally, the NeRF representation can be optimized by using the rate-distortion trade-off. Our proposed compression framework can be used for different representations and experimental results demonstrate that our approach significantly reduces the storage size with marginal distortion and achieves state-of-the-art rate-distortion performance for volumetric video on the HumanRF and ReRF datasets. Compared to the previous state-of-the-art method TeTriRF, we achieved an approximately -80% BD-rate on the HumanRF dataset and -60% BD-rate on the ReRF dataset.
Zhiyu Zhang 0010, Guo Lu, Huanxiong Liang, Zhengxue Cheng, Anni Tang, Li Song 0001
ACM Multimedia6
2024 HPC: Hierarchical Progressive Coding Framework for Volumetric Video
abstract
Volumetric video based on Neural Radiance Field (NeRF) holds vast potential for various 3D applications, but its substantial data volume poses significant challenges for compression and transmission. Current NeRF compression lacks the flexibility to adjust video quality and bitrate within a single model for various network and device capacities. To address these issues, we propose HPC, a novel hierarchical progressive volumetric video coding framework achieving variable bitrate using a single model. Specifically, HPC introduces a hierarchical representation with a multi-resolution residual radiance field to reduce temporal redundancy in long-duration sequences while simultaneously generating various levels of detail. Then, we propose an end-to-end progressive learning approach with a multi-rate-distortion loss function to jointly optimize both hierarchical representation and compression. Our HPC trained only once can realize multiple compression levels, while the current methods need to train multiple fixed-bitrate models for different rate-distortion (RD) tradeoffs. Extensive experiments demonstrate that HPC achieves flexible quality levels with variable bitrate by a single model and exhibits competitive RD performance, even outperforming fixed-bitrate models across various datasets.
Zihan Zheng, Houqiang Zhong, Qiang Hu 0003, Xiaoyun Zhang 0001, Li Song 0001, Ya Zhang 0002, Yanfeng Wang 0001
ACM Multimedia5
2024 Pioneer: Offline Reinforcement Learning based Bandwidth Estimation for Real-Time Communication
abstract
For Real-time Communication (RTC), Bandwidth Estimation (BWE) is crucial for enhancing user Quality of Experience (QoE) by ensuring efficient bandwidth utilization and low latency. Recent advancements have shifted towards machine learning based algorithms, particularly online reinforcement leanring (RL), to dynamically infer future bandwidth using statistical data. However, challenges such as dependency on training settings, the necessity for extensive trial and error, and instability in complex state spaces hinder their efficacy. To address these limitations, we propose Pioneer, a novel offline RL framework for BWE in RTC systems. Unlike its predecessors, Pioneer eliminates the need for real-time environment interaction during training and achieves good performance through lightweight training. Our framework consists of a Trajectory Sampler for state information preprocessing and a Bandwidth Estimator based on offline RL model. Our test results on offline datasets show that Pioneer can achieve better performance than expert algorithms. We also tested Pioneer on online simulation platforms, and Pioneer can improve QoE by 9% compared to other offline algorithm, demonstrating good robustness.
Bingcong Lu, Jun Xu 0040, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001
MMSys5
2024 AsymLLIC: Asymmetric Lightweight Learned Image Compression
abstract
Learned image compression (LIC) methods often employ symmetrical encoder and decoder architectures, evitably increasing decoding time. However, practical scenarios demand an asymmetric design, where the decoder requires low complexity to cater to diverse low-end devices, while the encoder can accommodate higher complexity to improve coding performance. In this paper, we propose an asymmetric lightweight learned image compression (AsymLLIC) architecture with a novel training scheme, enabling the gradual substitution of complex decoding modules with simpler ones. Building upon this approach, we conduct a comprehensive comparison of different decoder network structures to strike a better trade-off between complexity and compression performance. Experiment results validate the efficiency of our proposed method, which not only achieves comparable performance to VVC but also offers a lightweight decoder with only 51.47 GMACs computation and 19.65M parameters. Furthermore, this design methodology can be easily applied to any LIC models, enabling the practical deployment of LIC techniques.
Shen Wang 0013, Zhengxue Cheng, Donghui Feng 0003, Guo Lu, Li Song 0001, Wenjun Zhang 0001
VCIP5
2024 Coarse-to-fine Transformer For Lossless 3D Medical Image Compression
abstract
The rapid advancements in medical imaging have led to a growing demand for high-performance lossless compression of large 3D medical image datasets. Unlike natural images, medical images typically feature three-dimensional structures, and high bit-depth, necessitating specialized compression techniques. Based on a decoder-only transformer, we propose a learnable dual-decoder model for lossless compression of 3D medical images. Our approach packs voxels into patches, which are processed by a patch-level decoder to extract the patch feature. The voxels, along with the patch feature, are subsequently fed into a voxel-level decoder to model each voxel. This coarse-to-fine modeling strategy reduces the computational time for each voxel and enables long-range modeling dependencies. Experimental results demonstrate that our proposed model achieves state-of-the-art compression performance, with an approximately 15% improvement in compression performance over the traditional JP3D benchmark on various datasets.
Guo Lu, Donghui Feng 0003, Zhengxue Cheng, Guosheng Yu, Li Song 0001
VCIP6
2024 Content-Adaptive Rate-Quality Curve Prediction Model in Media Processing System
abstract
In streaming media services, video transcoding is a common practice to alleviate bandwidth demands. Unfortunately, traditional methods employing a uniform rate factor (RF) across all videos often result in significant inefficiencies. Content-adaptive encoding (CAE) techniques address this by dynamically adjusting encoding parameters based on video content characteristics. However, existing CAE methods are often tightly coupled with specific encoding strategies, leading to inflexibility. In this paper, we propose a model that predicts both RF-quality and RF-bitrate curves, which can be utilized to derive a comprehensive bitrate-quality curve. This approach facilitates flexible adjustments to the encoding strategy without necessitating model retraining. The model leverages codec features, content features, and anchor features to predict the bitrate-quality curve accurately. Additionally, we introduce an anchor suspension method to enhance prediction accuracy. Experiments confirm that the actual quality metric (VMAF) of the compressed video stays within ±1 of the target, achieving an accuracy of 99.14%. By incorporating our quality improvement strategy with the rate-quality curve prediction model, we conducted online A/B tests, obtaining both +0.107% improvements in video views and video completions and +0.064% app duration time. Our model has been deployed on the Xiaohongshu App.
Shibo Yin, Zhiyu Zhang 0010, Peirong Ning, Qiubo Chen, Guo Lu, Li Song 0001
VCIP8
2024 Efficient Bitrate Ladder Construction for Per-Shot Adaptive Encoding
abstract
HTTP adaptive streaming (HAS) constructs bitrate ladders to deliver videos with the best possible quality under varying network conditions. Though per-shot content adaptive encoding (CAE) largely improves the compression efficiency by constructing the optimal bitrate ladder for each video shot, it suffers from excessive encoding complexity as all the points in the operating space (typically resolution × bitrate) need to be encoded and compared. To address this issue, this paper proposes an efficient bitrate ladder construction method that encodes only a subset of operating points, then uses curve fitting and inter-curve prediction to estimate other points’ RD performance. The proposed method enables low-complexity ladder construction even for high-dimension operating spaces that incorporate dimensions like encoding presets. Experiments show that this method can achieve RD performance comparable to the original per-shot CAE with only 42% encoding points. Even when minimizing the encoding points to 3.6% of the original CAE, it achieves 15% BD-Rate improvements compared to using the fixed bitrate ladder.
Yan Zhao 0041, Zhengxue Cheng, Guo Lu, Rong Xie 0004, Li Song 0001
VCIP5
2024 Memories are One-to-Many Mapping Alleviators in Talking Face Generation
abstract
Talking face generation aims at generating photo-realistic video portraits of a target person driven by input audio. According to the nature of audio to lip motions mapping, the same speech content may have different appearances even for the same person at different occasions. Such one-to-many mapping problem brings ambiguity during training and thus causes inferior visual results. Although this one-to-many mapping could be alleviated in part by a two-stage framework (i.e., an audio-to-expression model followed by a neural-rendering model), it is still insufficient since the prediction is produced without enough information (e.g., emotions, wrinkles, etc.). In this paper, we propose MemFace to complement the missing information with an implicit memory and an explicit memory that follow the sense of the two stages respectively. More specifically, the implicit memory is employed in the audio-to-expression model to capture high-level semantics in the audio-expression shared space, while the explicit memory is employed in the neural-rendering model to help synthesize pixel-level details. Our experimental results show that our proposed MemFace surpasses all the state-of-the-art results across multiple scenarios consistently and significantly.
Anni Tang, Tianyu He, Xu Tan 0003, Jun Ling, Runnan Li, Sheng Zhao 0002, Jiang Bian 0002, Li Song 0001
IEEE Trans. Pattern Anal. Mach. Intell.8
2024 Depth-Guided Robust Point Cloud Fusion NeRF for Sparse Input Views
abstract
Novel-view synthesis with sparse input views is important for practical applications such as AR/VR and autonomous driving. Many works in this field have already integrated depth information into NeRF, utilizing depth priors for assistance in geometric and spatial understanding. However, most existing work tends to either overlook the inaccuracies in depth maps or only handle them roughly, limiting the effectiveness of the synthesis. To address this issue, we propose a depth-guided robust point cloud fusion NeRF for sparse input synthesis. We first construct a point cloud for each input view, with a novel point cloud representation based on learnable matrices and vectors. Then, through an additional lightweight scene fusion network, we fuse the point clouds from each input view to build a point cloud of the entire scene. By optimizing the point cloud representation and scene fusion network, inaccuracies in the depth map can be adjusted and refined, thereby achieving a more precise perception of the overall scene. Each voxel in the scene is determined by referencing the fused point cloud to establish its density and appearance. Experimental results demonstrate that our method outperforms state-of-the-art baselines.
Shuai Guo 0002, Qiuwen Wang, Yijie Gao, Rong Xie 0004, Lin Li 0062, Li Song 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 Fast Video Deduplication and Localization With Temporal Consistence Re-Ranking
abstract
The use of social media networks and mobile devices has experienced tremendous growth in recent years. This has led to a surge in the number of videos recorded and uploaded to social media platforms like TikTok and YouTube. However, this increase has also resulted in the rise of illegal duplicate videos, which are essentially the same as the original videos but with minor editing effects and variations in coding. In addition, the large number of duplicate videos is a major storage and communication efficiency issue. The task of finding duplicate videos from a large repository is referred to as video deduplication. Video deduplication is a crucial task for applications like saving storage space and detecting copyright infringement. This work proposes a fast and robust location-aware video deduplication system capable of retrieving duplicate videos from a large repository extremely quickly. In addition, the proposed system has the ability to find the precise location of the query video in the retrieved videos. To identify and localize short video clips against large video repositories, we utilize robust image-level features from keypoint aggregation and deep learning along with an efficient KNN search of query frames with a multiple k-d tree setup, giving us a set of candidate video clips. Then, a fast temporal consistence pruning algorithm re-ranks the clip-level candidates and identifies the matching clip along with its temporal location in a sequence in an efficient way. The system was tested on 1 million frame/145 hour and 4.5 million frame/636 hour repositories generated via the large-scale FIVR-200K and VCSL datasets, respectively. The proposed system achieves a recall of 98.8% and 94.1% for the FIVR-200K and VCSL datasets, respectively. A query frame is searched as fast as 83.96ms and 462.59ms from a 1 million frame/ 145 hour and a 4.5 million frame/636 hour repository, respectively. These experimental results demonstrate that our system is highly accurate and that the time consumption is extremely low for retrieving video along with its timestamp information from large-scale repositories.
Chris Henry, Li Song 0001, Zhu Li 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 A Character Position-Aware Compression Framework for Screen Text Image
abstract
Text patterns typically exhibit distinct boundaries and sparse color histograms. However, in current hybrid codec frameworks, the positions of coding units are often misaligned with the text patterns, resulting in prediction and color mapping tools consuming a large number of bits to indicate these patterns. Nowadays, some text detection and recognition methods have been proposed to accurately locate and analyze the text regions in screen images. Combined with these techniques, we propose a character position-aware compression framework for screen text image. On the encoder side, a low-complexity detection method is adopted to locate the text characters. Then it copies the detected characters to the position aligned with the coding unit (CU) grid to form a text layer. This text-layer representation can further increase the efficiency of existing screen content coding tools such as Intra Block Copy (IBC). Moreover, we design several compression tools based on this representation. We extend the two Motion Vector (MV) prediction modes: Adaptive Motion Vector Prediction (AMVP) and Merge. We modify the MV encoding syntax according to the layout characteristics of the text layer. We present a Gradient-guided In-loop Filter (GIF) to sharpen the text lines using a convolutional network. Experiments conducted on VVC reference software VTM all_intra configuration show that the proposed framework can achieve an average bitrate savings of 4.6% and 3.6% under the w/ GIF and w/o GIF versions, with a corresponding increase in CPU encoding complexity of 72% and 10%.
Guo Lu, Huanbang Chen, Donghui Feng 0003, Shen Wang 0013, Yan Zhao 0041, Rong Xie 0004, Li Song 0001
IEEE Trans. Circuits Syst. Video Technol.8
2024 Real-Time Free Viewpoint Video Synthesis System Based on DIBR and a Depth Estimation Network
abstract
Depth image-based rendering (DIBR) view synthesis is the most widely employed method in real-time FVV research. Despite recent progress, most DIBR-based FVV synthesis approaches are not sufficiently simple and effective in filling holes and artifacts. Additionally, they use RGB-D cameras, which are difficult to widely adopt or take considerable time to estimate high-quality depth images. This paper introduces a real-time FVV synthesis system based on DIBR and a depth estimation network. This system includes a 12-view synchronous camera system, a new multistage depth estimation network, a new GPU-accelerated DIBR algorithm, and a virtual view parameter generation method. This system provides the first real-time FVV solution for background-fixed fields based on DIBR and a depth estimation network. It can infer depth images for all camera views and synthesize any virtual view along the horizontal circular arc of the camera rig in real time. To our knowledge, we are the first to introduce background models and foreground masks and a refined multistage structure to address real-time high-quality depth estimation and DIBR FVV synthesis. We also build a high-quality multiview RGB-D synchronous dataset that has promising DIBR FVV synthesis performance to train and evaluate our system. The experimental results demonstrate the real-time and better performance of the proposed system.
Shuai Guo 0002, Jingchuan Hu, Kai Zhou 0016, Jionghao Wang, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001
IEEE Trans. Multim.5
2024 ViCoFace: Learning Disentangled Latent Motion Representations for Visual-Consistent Face Reenactment
abstract
Unsupervised face reenactment aims to animate a source image to imitate the motions of a target image while retaining the source portrait’s attributes like facial geometry, identity, hair texture, and background. While prior methods can extract the motion from the target image via compact representations (e.g., keypoints or latent motion bases [ 50 ]), they are not robust in predicting motions that are disentangled with portrait attributes, thus failing to preserve portrait attributes in the cross-subject reenactment. In this work, we propose an effective and cost-efficient face reenactment approach to address this issue. Our approach is highlighted by two major strengths. First, based on the theory of latent motion bases, we disentangle the full-head motion into two parts: the transferable motion and preservable motion and then compose the full motion representation using latent motions from the source image and the target image. Second, to optimize and learn disentangled motions, we introduce an efficient training framework, which features two training strategies: (1) a mixture training strategy that encompasses self-reenactment training and cross-subject training for better motion disentanglement and (2) a multi-path training strategy that improves the visual consistency of portrait attributes. Extensive experiments on widely used benchmarks demonstrate that our method exhibits a remarkable generalization ability compared to state-of-the-art baselines. Project and demos are available at https://junleen.github.io/projects/vicoface .
Jun Ling, Anni Tang, Rong Xie 0004, Li Song 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2023 Freestyle Layout-to-Image Synthesis
abstract
Typical layout-to-image synthesis (LIS) models generate images for a closed set of semantic classes, e.g., 182 common objects in COCO-Stuff. In this work, we explore the freestyle capability of the model, i.e., how far can it generate unseen semantics (e.g., classes, attributes, and styles) onto a given layout, and call the task Freestyle LIS (FLIS). Thanks to the development of large-scale pre-trained language-image models, a number of discriminative models (e.g., image classification and object detection) trained on limited base classes are empowered with the ability of unseen class prediction. Inspired by this, we opt to leverage large-scale pre-trained text-to-image diffusion models to achieve the generation of unseen semantics. The key challenge of FLIS is how to enable the diffusion model to synthesize images from a specific layout which very likely violates its pre-learned knowledge, e.g., the model never sees “a unicorn sitting on a bench” during its pre-training. To this end, we introduce a new module called Rectified Cross-Attention (RCA) that can be conveniently plugged in the diffusion model to integrate semantic masks. This “plug-in” is applied in each cross-attention layer of the model to rectify the attention maps between image and text tokens. The key idea of RCA is to enforce each text token to act on the pixels in a specified region, allowing us to freely put a wide variety of semantics from pre-trained knowledge (which is general) onto the given layout (which is specific). Extensive experiments show that the proposed diffusion network produces realistic and freestyle layout-to-image generation results with diverse text inputs, which has a high potential to spawn a bunch of interesting applications. Code is available at https://github.com/essunny310/FreestyleNet.
Zhiwu Huang, Qianru Sun, Li Song 0001, Wenjun Zhang 0001
CVPR4
2023 Boosting Video Object Segmentation via Space-Time Correspondence Learning
abstract
Current top-leading solutions for video object segmentation (VOS) typically follow a matching-based regime: for each query frame, the segmentation mask is inferred according to its correspondence to previously processed and the first annotated frames. They simply exploit the supervisory signals from the groundtruth masks for learning mask prediction only, without posing any constraint on the space-time correspondence matching, which, however, is the fundamental building block of such regime. To alleviate this crucial yet commonly ignored issue, we devise a correspondence-aware training framework, which boosts matching-based VOS solutions by explicitly encouraging robust correspondence matching during network learning. Through comprehensively exploring the intrinsic coherence in videos on pixel and object levels, our algorithm reinforces the standard, fully supervised training of mask segmentation with label-free, contrastive correspondence learning. Without neither requiring extra annotation cost during training, nor causing speed delay during deployment, nor incurring architectural modification, our algorithm provides solid performance gains on four widely used benchmarks, i.e., DAVIS2016&2017, and YouTube-VOS2018&2019, on the top of famous matching-based VOS solutions.
Liulei Li, Wenguan Wang, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001
CVPR5
2023 Dual-Head Fusion Network for Image Enhancement
abstract
Image enhancement algorithms have made great progress recently. However, most existing methods tend to construct a uniform enhancer for the color transformation of all pixels and ignore the local context information which is significant for photographs, causing unsatisfactory results. To solve these issues, we propose a novel dual-head fusion network for image enhancement, which synthetically considers both global scenario and local content information. Our network consists of four lightweight modules. We first develop a dual-head feature extraction module to extract the global condition vector and spatial context map. After that, we propose a context-aware retouching module and a global color rendering module to generate latent results. Finally, we employ the spatial attention based fusion module to adaptively aggregate the latent results. Experiments on public datasets show that our method consistently achieves the best results compared with SOTA methods both quantitatively and qualitatively.
Hengsheng Zhang, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001
ICASSP3
2023 Divide and Conquer: a Two-Step Method for High Quality Face De-identification with Model Explainability
abstract
Face de-identification involves concealing the true identity of a face while retaining other facial characteristics. Current target-generic methods typically disentangle identity features in the latent space, using adversarial training to balance privacy and utility. However, this pattern often leads to a trade-off between privacy and utility, and the latent space remains difficult to explain. To address these issues, we propose IDeudemon, which employs a "divide and conquer" strategy to protect identity and preserve utility step by step while maintaining good explainability. In Step I, we obfuscate the 3D disentangled ID code calculated by a parametric NeRF model to protect identity. In Step II, we incorporate visual similarity assistance and train a GAN with adjusted losses to preserve image utility. Thanks to the powerful 3D prior and delicate generative designs, our approach could protect the identity naturally, produce high quality details and is robust to different poses and expressions. Extensive experiments demonstrate that the proposed IDeudemon outperforms previous state-of-the-art methods.
Yunqian Wen, Bo Liu 0001, Jingyi Cao, Rong Xie 0004, Li Song 0001
ICCV5
2023 PACC: Perception Aware Congestion Control for Real-time Communication
abstract
Due to the network fluctuations, congestion control is indispensable to guarantee the quality of experience (QoE) for Real-Time Communication (RTC) users. This component adjusts the sending rate of media data, which determines the video encoding bitrate. However, existing control schemes either only focus on network numerical indicators or fail to adapt to various network environments. Logically, we propose PACC (Perception Aware Congestion Control) for RTC in this paper. Leveraging the convolutional neural network (CNN), we develop a quality sensor to infer the video quality increasing rate. Assisted with the variation trend analysis for user perception, PACC tunes the bitrate towards the direction of better QoE. Extensive tracedriven experiments demonstrate the effectiveness of PACC, which outperforms the existing landmark schemes by 8.2% to 32.4% and 6.8% to 18.0% in terms of transport and application layer QoE metrics, respectively.
Bingcong Lu, Li Song 0001, Rong Xie 0004, Yanmei Liu, Ying Chen 0011
ICME3
2023 Content Adaptive Checkerboard Context Model for Learned Image Compression
abstract
Learned image compression methods are becoming popular and have achieved excellent performance, of which joint context and hyperprior architectures are the mainstream. In order to avoid the time-consuming serial decoding pipeline introduced by the autoregressive context model, the checkerboard context model (CCM) is proposed to implement fast two-pass coding. However, CCM sets half of the latents as anchors to extract spatial context for the other non-anchors, which is rough and redundant. We propose a more precise and flexible content adaptive checkerboard context model to decrease the numbers and bit consumption of anchors. By introducing pseudo-anchors for simple regions in latents, our method can preserve the capability of fast two-pass coding and outperform CCM in Rate-Distortion performance on several baseline models with negligible computational overhead.
Guo Lu, Donghui Feng 0003, Li Song 0001
ISCAS5
2023 360-Degree Panorama Generation from Few Unregistered NFoV Images
abstract
360° panoramas are extensively utilized as environmental light sources in computer graphics. However, capturing a 360° × 180° panorama poses challenges due to the necessity of specialized and costly equipment, and additional human resources. Prior studies develop various learning-based generative methods to synthesize panoramas from a single Narrow Field-of-View (NFoV) image, but they are limited in alterable input patterns, generation quality, and controllability. To address these issues, we propose a novel pipeline called PanoDiff, which efficiently generates complete 360° panoramas using one or more unregistered NFoV images captured from arbitrary angles. Our approach has two primary components to overcome the limitations. Firstly, a two-stage angle prediction module to handle various numbers of NFoV inputs. Secondly, a novel latent diffusion-based panorama generation model uses incomplete panorama and text prompts as control signals and utilizes several geometric augmentation schemes to ensure geometric properties in generated panoramas. Experiments show that PanoDiff achieves state-of-the-art panoramic generation quality and high controllability, making it suitable for applications such as content editing.
Jionghao Wang, Jun Ling, Rong Xie 0004, Li Song 0001
ACM Multimedia5
2023 Achieving Privacy-Preserving Multi-View Consistency with Advanced 3D-Aware Face De-identification
abstract
The widespread application of face recognition technology has exacerbated privacy threats. Face de-identification is an effective means of protecting visual privacy by concealing identity information. While deep learning-based methods have greatly improved de-identification results, most existing algorithms rely on 2D generative models that struggle to produce identity-consistent results for multiple views. In this paper, we focus on identity disentanglement within the latest 3D-aware face generation model, and propose an advanced face de-identification framework that can be applied to various scenarios. Our proposed framework disentangles identity from other facial features, modifies only the former and generates the de-identified face using a 3D generator. This approach results in high-quality, identity-consistent de-identification that preserves other facial features. We demonstrate our approach on StyleNeRF, one of the most widely-used style-based neural radiation field models. Through extensive experiments, we demonstrate the effectiveness of our approach in achieving face de-identification both for a single image and group images with the same identity. Our work is a significant step forward in the field of face de-identification, opening up new possibilities for practical applications.
Jingyi Cao, Bo Liu 0001, Yunqian Wen, Rong Xie 0004, Li Song 0001
MMAsia5
2023 NeRF-SDP: Efficient Generalizable Neural Radiance Field with Scene Depth Perception
abstract
In recent years, neural radiance fields have exhibited impressive performance in novel view synthesis. However, exploiting complex network structures to achieve generalizable NeRF usually results in inefficient rendering. Existing methods for accelerating rendering directly employ simpler inference networks or fewer sampling points, leading to unsatisfactory synthesis quality. To address the challenge of balancing rendering speed and quality in generalizable NeRF, we propose a novel framework, NeRF-SDP, which achieves both efficiency and high fidelity by introducing scene depth perception. We incorporate more scene information into the radiance field by using our proposed geometry feature extraction and depth-encoded ray transformer to improve the model’s inference capabilities with sparse points. With the aid of scene depth perception, NeRF-SDP can better understand the scene’s structure, thus better reconstructing the objects’ edges with significantly fewer artifacts. Experimental results demonstrate that NeRF-SDP achieves comparable synthesis quality to state-of-the-art methods while significantly improving rendering efficiency. Furthermore, ablation studies confirm that the depth-encoded ray transformer enhances the model’s robustness to varying numbers of sampling points.
Qiuwen Wang, Shuai Guo 0002, Haoning Wu 0002, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001
MMAsia5
2023 High-Fidelity Free-View Talking Head Synthesis for Low-Bandwidth Video Conference
abstract
As video conferencing becomes an indispensable part of human’s daliy life, how to achieve a high-fidelity calling experience under low bandwidth has been a popular and challenging issue. Deep generative models have great potential in low-bandwidth facial video compression due to the excellent generation capability based on abridged information. Nevertheless, exsiting deep generation-based compression methods tend to handle motion information in pure 2D or pseudo 3D space, causing facial distortion when large head poses are encountered. In this paper, we propose a 3D-aware high-fidelity facial video conferencing system based on a parameterized NeRF-based face model. Through the compression of the parameterized face model and the transmisstion of extracted facial parameters, we implement high-fidelity talking head synthesis for video conferencing at an ultra-low bitrate. Additionally, the 3D perception capability of the system allows for viewpoint control over the head, achieving higher interactivity and practicability. Extensive experiments verify the effectiveness of the proposed 3D-aware high-fidelity free-view facial video conferencing system.
Zhiyu Zhang 0010, Anni Tang, Guo Lu, Rong Xie 0004, Li Song 0001
VCIP6
2023 Learned Image Compression Using Cross-Component Attention Mechanism
abstract
Learned image compression methods have achieved satisfactory results in recent years. However, existing methods are typically designed for RGB format, which are not suitable for YUV420 format due to the variance of different formats. In this paper, we propose an information-guided compression framework using cross-component attention mechanism, which can achieve efficient image compression in YUV420 format. Specifically, we design a dual-branch advanced information-preserving module (AIPM) based on the information-guided unit (IGU) and attention mechanism. On the one hand, the dual-branch architecture can prevent changes in original data distribution and avoid information disturbance between different components. The feature attention block (FAB) can preserve the important information. On the other hand, IGU can efficiently utilize the correlations between Y and UV components, which can further preserve the information of UV by the guidance of Y. Furthermore, we design an adaptive cross-channel enhancement module (ACEM) to reconstruct the details by utilizing the relations from different components, which makes use of the reconstructed Y as the textural and structural guidance for UV components. Extensive experiments show that the proposed framework can achieve the state-of-the-art performance in image compression for YUV420 format. More importantly, the proposed framework outperforms Versatile Video Coding (VVC) with 8.37% BD-rate reduction on common test conditions (CTC) sequences on average. In addition, we propose a quantization scheme for context model without model retraining, which can overcome the cross-platform decoding error caused by the floating-point operations in context model and provide a reference approach for the application of neural codec on different platforms.
Wenhong Duan, Zheng Chang 0002, Chuanmin Jia, Shanshe Wang, Siwei Ma 0001, Li Song 0001, Wen Gao 0001
IEEE Trans. Image Process.6
2023 Deep Online Video Stabilization Using IMU Sensors
abstract
In this paper, we propose a deep learning based sensor-driven method for online video stabilization. This method utilizes the Euler angles and acceleration values estimated from the gyroscope and accelerator to assist stable video reconstruction. We introduce two simple sub-networks for trajectory optimization. The first network exploits real unstable trajectories and camera acceleration values to detect shooting scenarios. This network also generates an attention mask to adaptively choose scenario-specific features. Then the second network predicts smooth camera paths based on real unstable trajectories using long short-term memory (LSTM) under the supervision of the above mask. The output of the trajectory optimization network is filtered with a two-step modification process to guarantee smoothness. The real and smoothed camera paths are then utilized as guidance to generate stable frames in a projective manner. We also capture videos with sensor data covering seven typical shooting scenarios and design a ground truth generation method to construct pseud-labels. Moreover, the trajectory smoothing network allows the use of 3- or 10-frame buffers as future information to construct a lookahead filter. Experimental results show that our online method could outperform other state-of-the-art offline methods in several shaky video clips with fewer buffer frames for both general and low-quality videos. Furthermore, our method could effectively reduce running times without performing image content analysis, and the stabilization efficiency reaches 25 fps on 1080p videos.
Chen Li 0021, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001
IEEE Trans. Multim.2
2023 Local Bidirection Recurrent Network for Efficient Video Deblurring with the Fused Temporal Merge Module
abstract
Video deblurring methods exploit the correlation between consecutive blurry inputs to generate sharp frames. However, designing an effective and efficient method is a challenging problem for video deblurring. To guarantee the effectiveness and further improve the deblurring performance, we adopt the recurrent-based method as the baseline and reconsider the recurrent mechanism as well as the temporal feature alignment in the state-of-the-art methods. For the recurrent mechanism, we add the local backward connection to the global forward recurrent backbone to effectively exploit accurate future information. For the temporal alignment, we adopt a fused temporal merge module that exploits the superiority of flow-based and kernel-based methods with progressive correlation volumes estimation. In addition, we evaluate our method with both synthetic datasets (GoPro, DVD) and a realistic dataset (BSD). The experimental results demonstrate that our method achieves significant performance improvement with a slight computational cost increase against the state-of-the-art video deblurring methods. The extended ablation studies verify the effectiveness of our model.
Chen Li 0021, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2023 High-Fidelity Face Reenactment Via Identity-Matched Correspondence Learning
abstract
Face reenactment aims to generate an animation of a source face using the poses and expressions from a target face. Although recent methods have made remarkable progress by exploiting generative adversarial networks, they are limited in generating high-fidelity and identity-preserving results due to the inappropriate driving information and insufficiently effective animating strategies. In this work, we propose a novel face reenactment framework that achieves both high-fidelity generation and identity preservation. Instead of sparse face representations (e.g., facial landmarks and keypoints), we utilize the Projected Normalized Coordinate Code (PNCC) to better preserve facial details. We propose to reconstruct the PNCC with the source identity parameters and the target pose and expression parameters estimated by 3D face reconstruction to factor out the target identity. By adopting the reconstructed representation as the driving information, we address the problem of identity mismatch. To effectively utilize the driving information, we establish the correspondence between the reconstructed representation and the source representation based on the features extracted by an encoder network. This identity-matched correspondence is then utilized to animate the source face using a novel feature transformation strategy. The generator network is further enhanced by the proposed geometry-aware skip connection. Once trained, our model can be applied to previously unseen faces without further training or fine-tuning. Through extensive experiments, we demonstrate the effectiveness of our method in face reenactment and show that our model outperforms state-of-the-art approaches both qualitatively and quantitatively. Additionally, the proposed PNCC reconstruction module can be easily inserted into other methods and improve their performance in cross-identity face reenactment.
Jun Ling, Anni Tang, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2022 PTSEFormer: Progressive Temporal-Spatial Enhanced TransFormer Towards Video Object Detection
Shanyan Guan, Rong Xie 0004, Li Song 0001
ECCV (8)6
2022 A Codec Information Assisted Framework for Efficient Compressed Video Super-Resolution
Hengsheng Zhang, Xueyi Zou, Jiaming Guo, Youliang Yan, Rong Xie 0004, Li Song 0001
ECCV (17)6
2022 Low-Complexity Multi-Model CNN in-Loop Filter for AVS3
abstract
Convolutional neural network (CNN) has demonstrated powerful capabilities in many image/video processing tasks. In this paper, a low-complexity multi-model CNN in-loop filtering scheme is proposed for AVS3. Firstly, we carefully choose simplified ResNet as the lightweight single model of our proposed network. Subsequently, based on the selected single model, the multi-model iterative training framework is proposed to train a multi-model filter, where the network depth and the number of multi-models are customized for different ranges of bit rate to achieve the trade-off between model performance and computational complexity. Experimental results show that our method achieves on average 6.06% BD-rate reduction on Y component under all intra configuration. Compared to other CNN filters with comparable performance, our proposed multi-model filter can significantly reduce the decoder complexity, and the experimental results indicate that the decoding time can be saved by 26.6% on average.
Shen Wang 0013, Yibing Fu, Li Song 0001, Wenjun Zhang 0001
ICASSP4
2022 MLS-GAN: Multi-Level Semantic Guided Image Colorization
abstract
Image colorization predicts plausible color versions of given grayscale images. Recently, several methods incorporate image semantics to assist image colorization and have shown impressive performance. To further exploit and take full advantage of more semantic information, in this paper, we propose a Multi-Level Semantic guided Generative Adversarial Network (MLS-GAN) for image colorization. Specifically, we utilize three different levels of semantics to guide the colorization process: image level, segmentation level and contextual level. Image-level classification semantics is used to learn category and high-level semantics, ensuring the reasonability of color results. At the segmentation level, multi-scale saliency map semantics is extracted to provide figure-background separation information, which can efficiently alleviate semantic confusion, especially for images with complex backgrounds. Furthermore, we novelly use non-local blocks to capture long-range semantic dependencies at the contextual level. Experiments show that our method enhances color consistency and can produce more vivid color in visually important regions, outperforming state-of-the-art methods qualitatively and quantitatively.
Xinning Chai, Xibei Liu, Hengsheng Zhang, Li Song 0001, Liean Cao
ICIP5
2022 CNN-Based Fast CU Partitioning Algorithm for VVC Intra Coding
abstract
Over a year has passed since the finalization of Versatile Video Coding (H.266/VVC), yet it is still far from practical deployment, a major reason being the excessive complexity. The flexible and sophisticated quad-tree with nested multi-type tree partitioning structure in VVC provides considerable performance gains while bringing about an exponential increase in encoding time. To reduce the coding complexity, this paper proposes a Convolutional Neural Network (CNN) based fast Coding Unit (CU) partitioning algorithm for intra coding, which accelerates CU partition through predicting the partition modes with texture information and terminating redundant modes in advance. Corresponding classifiers are designed for different CU sizes to improve prediction accuracy. Low rate-distortion performance degradation is guaranteed by introducing performance loss due to misclassification into the loss function. Experiments show that the proposed method can save encoding time ranging from 38.39% to 62.33% with 0.92% to 2.36% bit rate increase.
Jun Xu 0040, Yan Huang 0033, Li Song 0001
ICIP5
2022 A Multi-User Oriented Live Free-Viewpoint Video Streaming System Based on View Interpolation
abstract
As an important application form of immersive multimedia services, free-viewpoint video (FVV) enables users with great immersive experience by strong interaction. However, the computational complexity of virtual view synthesis algorithms poses a significant challenge to the real-time performance of an FVV system. Furthermore, the individuality of user interaction makes it difficult to serve multiple users simultaneously for a system with conventional architecture. In this paper, we novelly introduce a CNN-based view interpolation algorithm to synthesis dense virtual views in real time. Based on this, we also build an end-to-end live free-viewpoint system with a multi-user oriented streaming strategy. Our system can utilize a single edge server to serve multiple users at the same time without having to bring a large view synthesis load on the client side. We analyze the whole system and show that our approaches give the user a pleasant immersive experience, in terms of both visual quality and latency.
Jingchuan Hu, Shuai Guo 0002, Kai Zhou 0016, Jun Xu 0040, Li Song 0001
ICME6
2022 Generative Compression for Face Video: A Hybrid Scheme
abstract
As the latest video coding standard, versatile video coding (VVC) has shown its ability in retaining pixel quality. To excavate more compression potential for video conference scenarios under ultra-low bitrate, this paper proposes a bitrate-adjustable hybrid compression scheme for face video. This hybrid scheme combines the pixel-level precise recovery capability of traditional coding with the generation capability of deep learning based on abridged information, where Pixel-wise Bi-Prediction, Low-Bitrate-FOM and Lossless Keypoint Encoder collaborate to achieve PSNR up to 36.23 dB at a low bitrate of 1.47 KB/s. Without introducing any additional bi-trate, our method has a clear advantage over VVC under a completely fair comparative experiment, which proves the effectiveness of our proposed scheme. Moreover, our scheme can adapt to any existing encoder/configuration to deal with different encoding requirements, and the bitrate can be dynamically adjusted according to the network condition.
Anni Tang, Yan Huang 0033, Jun Ling, Zhiyu Zhang 0010, Rong Xie 0004, Li Song 0001
ICME7
2022 Complexity-Oriented Per-Shot Video Coding Optimization
abstract
Current per-shot encoding schemes aim to improve the compression efficiency by shot-level optimization. It splits a source video sequence into shots and imposes optimal sets of encoding parameters on each shot. Per-shot encoding achieved approximately 20% bitrate savings over baseline fixed QP encoding at the expense of pre-processing complexity. However, the adjustable parameter space of the current per-shot encoding schemes only has spatial resolution and QP/CRF, resulting in a lack of encoding flexibility. In this paper, we extend the per-shot encoding framework in the complexity dimension. We believe that per-shot encoding with flexible complexity will help in deploying user-generated content. We propose a rate-distortion-complexity optimization process for encoders and a methodology to determine the coding parameters under the constraints of complexities and bitrate ladders. Experimental results show that our proposed method achieves complexity constraints ranging from 100% to 3% in a dense form compared to the slowest per-shot anchor. With similar complexities of the per-shot scheme fixed in specific presets, our proposed method achieves BDrate gain up to −19.17%.
Hongcheng Zhong, Jun Xu 0040, Donghui Feng 0003, Li Song 0001
ICME5
2022 An Attention Based CNN with Temporal Hierarchical Deployment for AVS3 Inter In-loop Filtering
abstract
Convolutional Neural Network (CNN) based in-loop filter in video coding has demonstrated its superiority in benefiting coding efficiency and enhancing visual quality. In this paper, we develop a lightweight CNN-based in-loop filter for AVS3 encoder. The proposed network consists of several residual blocks with two attention branches, namely Dual Attention Network (DAN). The added channel attention branch and spatial attention branch can take advantage of the correlation between channels and pixels, improving the quality of reconstructed frames. In addition, by analyzing the inter prediction reference structure, we propose a temporal hierarchical deployment strategy to incorporate DAN into AVS3 video encoder. Therefore reconstructed frames with different distortions and referenced levels can be enhanced according to their temporal layer. Experiments prove the effectiveness of our strategy and results show our method achieves up to 6.57% and on average 3.64% BD-rate reduction on Y component under Random Access configuration.
Yibing Fu, Shen Wang 0013, Li Song 0001, Wenjun Zhang 0001
ISCAS4
2022 Intra Encoding Complexity Control with a Time-Cost Model for Versatile Video Coding
abstract
For the latest video coding standard Versatile Video Coding (VVC), the encoding complexity is much higher than previous video coding standards to achieve a better coding efficiency, especially for intra coding. The complexity becomes a major barrier of its deployment and use. Even with many fast encoding algorithms, it is still practically important to control the encoding complexity to a given level. Inspired by rate control algorithms, we propose a scheme to precisely control the intra encoding complexity of VVC. In the proposed scheme, a Time-PlanarCost (viz. Time-Cost, or T-C) model is utilized for CTU encoding time estimation. By combining a set of predefined parameters and the T-C model, CTU-level complexity can be roughly controlled. Then to achieve a precise picture-level complexity control, a framework is constructed including uneven complexity pre-allocation, preset selection and feedback. Experimental results show that, for the challenging intra coding scenario, the complexity error quickly converges to under 3.21%, while keeping a reasonable time saving and rate-distortion (RD) performance. This proves the efficiency of the proposed methods.
Yan Huang 0033, Jizheng Xu, Yan Zhao 0041, Li Song 0001
ISCAS5
2022 Multi-Scale Coarse-to-Fine Transformer for Frame Interpolation
abstract
The majority of prevailing video interpolation methods compute flows to estimate the intermediate motion. However, accurate estimation of the intermediate motion is difficult with low-order motion model hypothesis, which induces enormous difficulties for subsequent processing. To alleviate the limitation, we propose a two-stage flow-free video interpolation architecture. Rather than utilizing pre-defined motion models, our method represents complex motion through data-driven learning. In the first stage, we analyze spatial-temporal information and generate coarse anchor frame features. In the second stage, we employ transformers to transfer neighboring features to the intermediate time steps and enhance the spatial textures. To improve the quality of coarse anchor frame features and the robustness in dealing with the multi-scale textures with large-scale motion, we propose a multi-scale architecture and transformers with variable token sizes to progressively enhance the features. The experimental results demonstrate that our model outperforms state-of-the-art methods for both single frame and multi frames interpolation tasks, and the extended ablation studies verify the effectiveness of our model.
Chen Li 0021, Li Song 0001, Xueyi Zou, Jiaming Guo, Youliang Yan, Wenjun Zhang 0001
ACM Multimedia2
2022 A Cloud-based Free View Solution
abstract
We provide a real-time cloud-based free view solution composed of ingestion, processing, distribution, and interaction with optimization. Results show that our solution reduces the cost and complexity of free-viewpoint video (FVV) deployment and improves the performance by adopting the lightweight synthesis algorithm and the CMAF standard. Our demo video is available on https://github.com/Amygu1994/FreeView/raw/main/Demo_A%20Cloud-based%20Free%20View%20Solution.mp4
Yanying Sun, Joseph Ma, Li Song 0001
MMSP4
2022 A new free viewpoint video dataset and DIBR benchmark
abstract
Free viewpoint video (FVV) has drawn great attention in recent years, which provides viewers with strong interactive and immersive experience. Despite the developments made, further progress of FVV research is limited by existing datasets that mostly have too few number of camera views, or static scenes. To overcome the limitations, in this paper, we present a new dynamic RGB-D video dataset with up to 12 views. Our dataset consists of 13 groups of dynamic video sequences that are taken at the same scene, and a group of video sequences of the empty scene. Each group has 12 HD video sequences taken by synchronized cameras and 12 correspondingly estimated depth video sequences. Moreover, we also introduce a FVV synthesis benchmark on the basis of depth image based rendering (DIBR) to help researchers validate their data-driven methods. We hope our work will inspire more FVV synthesis methods with enhanced robustness, improved performance and deeper understanding.
Shuai Guo 0002, Kai Zhou 0016, Jingchuan Hu, Jionghao Wang, Jun Xu 0040, Li Song 0001
MMSys6
2022 Position-based Motion Vector Prediction for Textual Image Coding
abstract
Textual content is becoming increasingly important in video conferencing, while existing screen content encoding tools still produce a high bitrate in text regions. The main coding tool Intra Block Copy (IBC) inherits the MV prediction mechanism in inter-frame coding, but the adjacent text characters typically have irrelevant MVs, making it inefficient to predict MV using only neighbor MVs. To solve the problem, we propose the Position-based Motion Vector Prediction, to cache IBC AMVP PU positions as predictors. One character can find the previously encoded position to construct a good MV prediction. Experiment results show the effectiveness of the proposed prediction scheme.
Donghui Feng 0003, Guo Lu, Li Song 0001
PCS4
2022 Perceptual Video Coding Based on Semantic-Guided Texture Detection and Synthesis
abstract
Visually insensitive texture regions consume a large number of bitrate in hybrid video coding, leading to the waste of bandwidth resources. For this, we propose a semantic-guided texture synthesis framework (STSF). At encoder, high-level semantic information is adopted as texture features to detect texture regions and is sent to the decoder. Detected texture regions are coarsely encoded by hybrid codec. To generate realistic texture patterns, we design a multi-model semantic-guided texture synthesis generative adversarial network (STSGAN) at decoder, which works in a divide-and-conquer manner that semantically different texture regions are synthesized by different submodels in it. Experimental results show that STSF can achieve a −17.2% MOS BD-rate under the lowdelay_P configuration, compared with VVC.
Guo Lu, Rong Xie 0004, Li Song 0001
PCS4
2022 A Large-scale Sports Tracking Dataset and Progressive Re-detection Based Sports Tracking
abstract
Recent years have witnessed the great progress of Visual Object Tracking (VOT) which aims to predict the position of an object in each video frame given only its initial appearance. However, even the state-of-the-art methods are confronted with performance degradation, i.e., the tracker drift problem, in sports video scenes (e.g., soccer, basketball). There are two main causes that should be responsible for the tracker drift problem. First, the object of interest is often occluded by other objects that share a similar appearance. Such severe occlusion prevents the model from distinguishing the correct tracking object from other distractors in the future frames. Second, in sports videos, the objects often move fast from one place to another, which incurs severe blurry visual effects among consecutive frames. To address the issues of the tracker drift problem, we treat VOT as a tracking-by-re-detection task. Specifically, we detect candidate objects within a searching area (determined by object location in the previous frame) in the current frame and develop a progressive algorithm to filter out distractors in the area, which proves robust towards occlusion scenarios and tracker drift problems. Combining the advantages of our settings, the proposed framework method is robust to motion blur and object occlusion issues and achieves state-of-the-art tracking results on our challenging dataset.
Qinyu Xu, Huaqiang Ren, Rong Xie 0004, Li Song 0001
VCIP6
2022 RGBD-based Real-time Volumetric Reconstruction System: Architecture Design and Implementation
abstract
With the increasing popularity of commercial depth cameras, 3D reconstruction of dynamic scenes has aroused widespread interest. Although many novel 3D applications have been unlocked, real-time performance is still a big problem. In this paper, a low-cost, real-time system: LiveRecon3D, is presented, with multiple RGB-D cameras connected to one single computer. The goal of the system is to provide an interactive frame rate for 3D content capture and rendering at a reduced cost. In the proposed system, we adopt a scalable volume structure and employ ray casting technique to extract the surface of 3D content. Based on a pipeline design, all the modules in the system run in parallel and are designed to minimize the latency to achieve an interactive frame rate of 30 FPS. At last, experimental results corresponding to implementation with three Kinect v2 cameras are presented to verify the system's effectiveness in terms of visual quality and real-time performance.
Kai Zhou 0016, Shuai Guo 0002, Jingchuan Hu, Jionghao Wang, Qiuwen Wang, Li Song 0001
VCIP6
2022 L0 structure-prior assisted blur-intensity aware efficient video deblurring
Chen Li 0021, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001
Neurocomputing2
2022 IdentityDP: Differential private identification protection for face images
Yunqian Wen, Bo Liu 0001, Ming Ding 0001, Rong Xie 0004, Li Song 0001
Neurocomputing5
2022 Multiview nonlinear discriminant structure learning for emotion recognition
Shuai Guo 0002, Li Song 0001, Rong Xie 0004, Lin Li 0062, Shenglan Liu 0001
Knowl. Based Syst.2
2022 IdentityMask: Deep Motion Flow Guided Reversible Face Video De-Identification
abstract
Unprecedented video collection and sharing have exacerbated privacy concerns and led to increasing interest in privacy-preserving tools. A satisfactory video de-identification tool should be able to remove sensitive identity information from face videos while maintaining useful information for other identity-agnostic tasks. Meanwhile, it is necessary to allow the authority to inspect real identity when abnormal events are detected. Existing methods only focus on the study of de-identification, and lack the desired recovery ability when granting permissions. Furthermore, they all process the videos frame by frame, which hardly benefit from motion and inter-frame information. In this paper, we propose a modular architecture for reversible face video de-identification, called IdentityMask, which leverages deep motion flow to avoid per-frame evaluation. Our framework consists of two processes: the de-identification process provides a protective mask for identity information, while the recovery process can remove the protective mask if and only if the right key is provided. To this end, a Protection Module and a Recovery Module are built as two major functional modules, both based on an identity disentanglement network and guided by a crucial Motion Flow Module. An Affine Transformation Module provides simple but reliable assistance. Extensive experiments on a diverse natural video dataset (gender, ethnicity, age, etc.) demonstrate the effectiveness of the proposed framework for reversible face video de-identification.
Yunqian Wen, Bo Liu 0001, Jingyi Cao, Rong Xie 0004, Li Song 0001, Zhu Li 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 Edge-Based Video Compression Texture Synthesis Using Generative Adversarial Network
abstract
It has been recognized that texture patterns with abundant high-frequency components, such as grass and water, produce visual masking effects, and the distortion in textures is hard to be perceived by human eyes than structure regions. However, modern video codecs in a rate-distortion optimized manner usually consume a lot of bits to encode textures, leading to the insufficiency in perceptual coding performance. Nowadays, with the rapid development of deep learning, learning based texture synthesis methods have been proposed to replace the coding process of prediction residuals to reduce the rate cost. In this paper, we present a deep texture synthesizer named edge-based texture synthesis framework (ETSF). At encoder side, the framework detects texture regions by semantic and fidelity classification criteria, and the detected regions are quantized coarsely by the hybrid coding framework. In texture characterization, ETSF extracts low-level edge features representing pixel intensity variation. Feature processing tools are developed to remove the spatiotemporal redundancy of edges. The processed edge information is compressed and transmitted. To effectively recover textures, we design an edge-based texture synthesis generative adversarial network (ETSGAN) at the decoder of ETSF, which can incorporate edge information into convolutional layers and generate realistic textures. Experimental results on a collected texture dataset show that the proposed ETSF can achieve an average of -12.8%, -14.2% and -9.6% MOS BD-rate under lowdelay_B, lowdelay_P and random_access configurations of VVC coding, respectively.
Jun Xu 0040, Donghui Feng 0003, Rong Xie 0004, Li Song 0001
IEEE Trans. Circuits Syst. Video Technol.5
2022 Wireless Multiplayer Interactive Virtual Reality Game Systems With Edge Computing: Modeling and Optimization
abstract
Wireless multiplayer interactive virtual reality (VR) game has the high computing workload of VR and unpredictable interaction among players, which brings severe challenges to the design of wireless communication systems. In this paper, we propose a wireless multiplayer interactive VR game transmission framework based on mobile edge computing (MEC) that is able to model the interaction among players and compute the post-processing procedures at the MEC server or the mobile VR device. In the framework, the absolute delay of each player is used to avoid VR vertigo and the inter-player delay among players is used to model the fairness of the interactive game. Aiming to minimize the average inter-player delay, we optimize the computing resource allocation of the MEC server, the wireless bandwidth allocation and the post-processing decision policy subject to the constraints of the absolute delay requirements, the local energy limits of players, the total bandwidth limit and the computing resources limit. To tackle the non-convex problem efficiently, we design an iterative algorithm based on the NESTT-G algorithm which iteratively optimizes the truncated first-order Taylor approximation of the objective. Numerical results demonstrate the proposed algorithm can reduce the average inter-player delay significantly with lower complexity, and also reveal the impact of different parameters and the channel state conditions on the post-processing decision and edge resource allocation.
Zhiyong Chen 0002, Li Song 0001, Dazhi He, Bin Xia 0001
IEEE Trans. Wirel. Commun.3
2021 Dual Attention Guided Gaze Target Detection in the Wild
abstract
Gaze target detection aims to infer where each person in a scene is looking. Existing works focus on 2D gaze and 2D saliency, but fail to exploit 3D contexts. In this work, we propose a three-stage method to simulate the human gaze inference behavior in 3D space. In the first stage, we introduce a coarse-to-fine strategy to robustly estimate a 3D gaze orientation from the head. The predicted gaze is decomposed into a planar gaze on the image plane and a depth-channel gaze. In the second stage, we develop a Dual Attention Module (DAM), which takes the planar gaze to produce the filed of view and masks interfering objects regulated by depth information according to the depth-channel gaze. In the third stage, we use the generated dual attention as guidance to perform two sub-tasks: (1) identifying whether the gaze target is inside or out of the image; (2) locating the target if inside. Extensive experiments demonstrate that our approach performs favorably against state-of-the-art methods on GazeFollow and VideoAttentionTarget datasets.
Yi Fang 0009, Jiapeng Tang, Wang Shen, Wei Shen 0002, Xiao Gu 0001, Li Song 0001, Guangtao Zhai
CVPR6
2021 Region-Aware Adaptive Instance Normalization for Image Harmonization
abstract
Image composition plays a common but important role in photo editing. To acquire photo-realistic composite images, one must adjust the appearance and visual style of the foreground to be compatible with the background. Existing deep learning methods for harmonizing composite images directly learn an image mapping network from the composite to real one, without explicit exploration on visual style consistency between the background and the foreground images. To ensure the visual style consistency between the foreground and the background, in this paper, we treat image harmonization as a style transfer problem. In particular, we propose a simple yet effective Region-aware Adaptive Instance Normalization (RAIN) module, which explicitly formulates the visual style from the background and adaptively applies them to the foreground. With our settings, our RAIN module can be used as a drop-in module for existing image harmonization networks and is able to bring significant improvements. Extensive experiments on the existing image harmonization benchmark datasets shows the superior capability of the proposed method. Code is available at https://github.com/junleen/RainNet.
Jun Ling, Li Song 0001, Rong Xie 0004, Xiao Gu 0001
CVPR3
2021 Dense 3D Coordinate Code Prior Guidance for High-Fidelity Face Swapping and Face Reenactment
abstract
In face synthesis tasks, commonly used 2D face representations (e.g. 2D landmarks, segmentation maps, etc.) are usually sparse and discontinuous. To combat these shortcomings, we utilize a dense and continuous representation, named Projected Normalized Coordinate Code (PNCC), as the guidance and develop a PNCC-Spatio-Normalization (PSN) method to achieve face synthesis regarding arbitrary head poses and expressions. Based on PSN, we provide an effective framework for face reenactment and face swapping task. To ensure a harmonious and seamless face swapping, a simple yet effective Appearance-Blending Module (ABM) is proposed to fit the synthesized face to the target face. Our method is subject-agnostic and can be applied to any pair of faces without extra fine-tuning. Both qualitative and quantitative experiments are conducted to demonstrate the superiority of the proposed method in comparisons to existing state-of-the-art systems.
Anni Tang, Jun Ling, Rong Xie 0004, Li Song 0001
FG5
2021 Personalized and Invertible Face De-identification by Disentangled Identity Information Manipulation
abstract
The popularization of intelligent devices including smartphones and surveillance cameras results in more serious privacy issues. De-identification is regarded as an effective tool for visual privacy protection with the process of concealing or replacing identity information. Most of the existing de-identification methods suffer from some limitations since they mainly focus on the protection process and are usually non-reversible. In this paper, we propose a personalized and invertible de-identification method based on the deep generative model, where the main idea is introducing a user-specific password and an adjustable parameter to control the direction and degree of identity variation. Extensive experiments demonstrate the effectiveness and generalization of our proposed framework for both face de-identification and recovery.
Jingyi Cao, Bo Liu 0001, Yunqian Wen, Rong Xie 0004, Li Song 0001
ICCV5
2021 Video Multimethod Assessment Fusion Based Rate-Distortion Optimization for Versatile Video Coding
abstract
The emerging visual quality assessment metric VMAF that fuses several elementary metrics by SVM regression has shown a higher correlation with human perception. In this paper, we introduce VMAF into the traditional video coding task as the distortion metric, which needs to be optimized to explore the potential of perceptual quality improvement. Specifically, we propose a multi-granularity VMAF based rate-distortion optimization framework. A frame level visual quality adaption is first conducted by taking the quantization characteristics into account at the coarse-grained adjustment step. Within each frame, a CTU level Lagrangian multiplier and corresponding quantization parameter adaption are carried out based on the content of each CTU at the fine-grained adjustment step. The proposed method has been incorporated into the latest video coding standard – VVC. Experimental results show compared with the conventional rate-distortion optimization for SSE, the proposed method achieves an average 3.30% BD-rate reduction in VMAF.
Han Zhang 0030, Jizheng Xu, Li Song 0001
ICIP3
2021 SVM Based Fast CU Partitioning Algorithm for VVC Intra Coding
abstract
Recently, Joint Video Experts Team (JVET) has completed the new Versatile Video Coding (H.266/VVC) standard. VVC employs a new block partition structure named quad-tree with nested multi-type tree (QTMT) to improve coding efficiency. However, the new block partition structure increases huge encoding time compared with HEVC for brute-force ratedistortion (RD) optimization. To reduce encoding complexity, we propose a Support Vector Machine (SVM) based fast CU partitioning algorithm for VVC intra coding in this paper which terminates redundant partitions early by predicting the partition of CU using texture information. We trained classifiers for CUs of different sizes to improve accuracy and control the complexity of the classifiers themselves. Different thresholds are set for each classifier to achieve a trade-off between encoding complexity and RD performance. Experimental results show that the proposed method can save encoder time ranging from 30.78% to 63.16% with 1.10% to 2.71% BD-BR increase.
Yan Huang 0033, Li Song 0001, Wenjun Zhang 0001
ISCAS4
2021 Blindly Predict Image and Video Quality in the Wild
abstract
Emerging interests have been brought to blind quality assessment for images/videos captured in the wild, known as in-the-wild I/VQA. Prior deep learning based approaches have achieved considerable progress in I/VQA, but are intrinsically troubled with two issues. Firstly, most existing methods fine-tune the image-classification-oriented pre-trained models for the absence of large-scale I/VQA datasets. However, the task misalignment between I/VQA and image classification leads to degraded generalization performance. Secondly, existing VQA methods directly conduct temporal pooling on the predicted frame-wise scores, resulting in ambiguous inter-frame relation modeling. In this work, we propose a two-stage architecture to separately predict image and video quality in the wild. In the first stage, we resort to supervised contrastive learning to derive quality-aware representations that facilitate the prediction of image quality. Specifically, we propose a novel quality-aware contrastive loss to pull together samples of similar quality and push away quality-different ones in embedding space. In the second stage, we develop a Relation-Guided Temporal Attention (RTA) module for video quality prediction, which captures global inter-frame dependencies in embedding space to learn frame-wise attention weights for frame quality aggregation. Extensive experiments demonstrate that our approach performs favorably against state-of-the-art methods on both authentically distorted image benchmarks and video benchmarks.
Jiapeng Tang, Yi Fang 0009, Rong Xie 0004, Xiao Gu 0001, Guangtao Zhai, Li Song 0001
MMAsia7
2021 DVRCNN: Dark Video Post-processing Method for VVC
Donghui Feng 0003, Han Zhang 0030, Li Song 0001
MMM (1)5
2021 Deep Face Swapping via Cross-Identity Adversarial Training
Jun Ling, Li Song 0001, Rong Xie 0004
MMM (2)4
2021 HEVC VMAF-oriented Perceptual Rate Distortion Optimization using CNN
abstract
Video coding standards like HEVC and VVC have achieved significant coding performance. However, the RDO module in coding framework ignores the characteristics of human visual system (HVS), which leads to insufficiency for perceptual video coding. Recently, learning-based objective assessment metric VMAF is developed and has been demonstrated higher quality assessment accuracy than conventional metrics. To incorporate VMAF into RDO aiming at improving perceptual coding efficiency, in this paper, a perceptual RDO scheme is proposed. A CNN-based on-line training method is first explored to determine the VMAF-related distortion estimation coefficient. Based on the VMAF-related coefficient and R-D model, a VMAF-based Lagrangian multiplier is proposed to adjust the R-D performance of each coding block. Experiments demonstrate that the proposed method can achieve an average -2.80% VMAF-based BD-Rate compared with the original HEVC, which effectively improves the coding performance.
Yan Huang 0033, Rong Xie 0004, Li Song 0001
PCS4
2021 Video Compression based on Jointly Learned Down-Sampling and Super-Resolution Networks
abstract
With the blooming of deep learning technology in computer vision, the integration of deep learning and the traditional video coding has made significant improvements, especially applying the super-resolution neural network as the post-processing module in the down-sampling-based video compression framework. However, the pre-processing module lacks back-propagated gradients for jointly considering down-sampling and up-sampling due to the non-differentiability of the traditional video codec. In this paper, we propose an end- to-end down-sampling-based video compression framework applying convolutional neural networks both as down-sampling and upsampling. We use a virtual codec neural network to approximate the actual video codec so that the gradient can be effectively back-propagated for joint training. Experimental results show the superiority of our proposed framework compared with the predefined down-sampling-based video compression and various methods of joint training.
Yuzhuo Wei, Li Chen 0021, Li Song 0001
VCIP3
2021 Deep Motion Flow Aided Face Video De-identification
abstract
Advances in cameras and web technology have made it easy to capture and share large amounts of face videos over to an unknown audience with uncontrollable purposes. These raise increasing concerns about unwanted identity-relevant computer vision devices invading the characters's privacy. Previous de-identification methods rely on designing novel neural networks and processing face videos frame by frame, which ignore the data feature in redundancy and continuity. Besides, these techniques are incapable of well-balancing privacy and utility, and per-frame evaluation is easy to cause flicker. In this paper, we present deep motion flow, which can create remarkable de-identified face videos with a good privacy-utility tradeoff. It calculates the relative dense motion flow between every two adjacent original frames and runs the high quality image anonymization only on the first frame. The de-identified video will be obtained based on the anonymous first frame via the relative dense motion flow. Extensive experiments demonstrate the effectiveness of our proposed de-identification method.
Yunqian Wen, Bo Liu 0001, Rong Xie 0004, Jingyi Cao, Li Song 0001
VCIP5
2021 Fast and Context-Aware Framework for Space-Time Video Super-Resolution
abstract
Increasing the spatial resolution and frame rate of a video simultaneously has attracted attention in recent years. The current one-stage space-time video super-resolution (STVSR) methods are difficult to deal with large motion and complex scenes, and are time-consuming and memory intensive. We propose an efficient STVSR framework, which can correctly handle complicated scenes such as occlusion and large motion and generate results with clearer texture. In REDS dataset, our method outperforms all existing one-stage methods. Our method is lightweight and can generate 720p frames at 16fps on a NVIDIA GTX 1080 Ti GPU.
Xueheng Zhang, Li Chen 0021, Li Song 0001
VCIP3
2021 Mobile Edge Resource optimization for Multiplayer Interactive Virtual Reality Game
abstract
Edge computing has been regarded as an efficient approach to achieve the multiplayer interactive virtual reality (VR) game over wireless networks, where the game scenes can be rendered at the edge computing server and then the real-time video frames can be transmitted to the players. In order to ensure the fairness of the interacting players, we allocate the edge computing and wireless bandwidth resources for minimizing the average inter-player delay among different players. The proposed programming is under the constraints of the absolute delay requirements, the maximum frame per second (FPS) demands, the total bandwidth and rendering resources limits. The influence of prediction and pre-rendering field of views (FOVs) is also considered in the optimization problem to model the interaction among players. To tackle the non-convex problem efficiently, a sub-optimal algorithm which convert the original problem into several convex subproblems to optimize iteratively is designed. Finally, numerical results verify the proposed algorithm can reduce the average interplayer delay significantly with lower complexity, and it also reveals the impact of the content sizes and the channel state conditions on the edge resource allocation.
Yingjiao Li, Zhiyong Chen 0002, Li Song 0001
WCNC4
2021 Modeling Acceleration Properties for Flexible INTRA HEVC Complexity Control
abstract
It is a very well-known fact, that the high complexity of the High Efficiency Video Coding standard (HEVC) is the main hurdle for its wide deployment and use. To tackle this problem, a number of recent research outcomes exploit heuristic algorithms and machine learning, including deep learning, to reduce the coding complexity. However, in most cases, each encoder module, i.e., encoding process, is first accelerated individually, and then different acceleration algorithms are manually combined. Without a holistic strategy, the acceleration potential of multi-module combination is not exploited and the Rate-Distortion (RD) loss is generally not well controlled. To tackle these shortcomings, this paper exploits the acceleration properties of different modules, i.e., the numerical representation of potential time saving and possible RD loss, from which a heuristic model is explored. Then a Heuristic Model Oriented Framework (HMOF) is proposed which adapts the properties of modules to underlying acceleration algorithms. In the framework, two advanced acceleration algorithms, including Border Considered CNN (BC-CNN)-based Coding Unit (CU) partition and Naive Bayes-based Prediction Unit (PU) partition, are proposed for the CU and PU modules, respectively. Further, by leveraging the heuristic model as the guidance to combine the proposed acceleration algorithms, HMOF is globally optimized, where different time saving budgets are wisely allocated to different modules and a theoretically minimal RD loss is achieved. According to the experimental results, through fusing a suitable deep learning technique and a Bayes-Based prediction, the proposed acceleration framework HMOF enable multiple acceleration choices. Here the proposed joint optimization strategy help to make a choice leading to the best cost-performance. Furthermore, within the proposed framework, intra coding time can be precisely controlled with negligible Bjøntegaard delta bit-rate (BDBR) loss. In this context, as a complexity control method, HMOF outperforms the state-of-the-art complexity reduction algorithms under a similar complexity reduction ratio. These results partially demonstrate the superiority of the proposed technique.
Yan Huang 0033, Li Song 0001, Rong Xie 0004, Ebroul Izquierdo, Wenjun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 Compression Priors Assisted Convolutional Neural Network for Fractional Interpolation
abstract
Fractional interpolation has been extensively utilized in a series of video coding standards to generate fractional precision prediction to remove the temporal redundancy in consecutive frames. In the traditional interpolation filter based methods, the fractional samples are interpolated through a linear combination of the neighboring integer samples. This method is simple yet unable to accurately characterize the nonstationary video signals. Recently, convolutional neural network has been utilized in the fractional interpolation and shows superior performance compared with the traditional methods. However, only the reconstruction of the reference frame is used as the infer information source. All the other information contained in the bitstream or generated during the encoding/decoding procedure denoted as the compression prior is not utilized at all. In this paper, we give the first trial to involve some compression priors into the CNN to improve the performance of a in-loop coding tool. Specifically, we propose a Compression Priors assisted Convolutional Neural Network (CPCNN) to further improve the fractional interpolation efficiency. In addition to the reconstructed component, we additionally utilize two other compression priors - the corresponding residual component and col-located high quality component to boost the performance. Specifically, the residual component that indicates the prediction efficiency and contains effective texture information is utilized as a complementary input to the reconstructed one. While the col-located component provides more useful high quality information to help the reconstruction get rid of the quality fluctuation. Furthermore, a special network structure is designed to learn powerful representations of these triple input components. Comprehensive experiments have been conducted to demonstrate the effectiveness of our proposed CPCNN. The experimental results show that compared to HEVC, our proposed CPCNN achieves on average of 5.3%, 2.8% and 1.9% BD-Rate savings under LDP, LDB and RA configurations, respectively.
Han Zhang 0030, Li Song 0001, Li Li 0040, Zhu Li 0001, Xiaokang Yang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2021 VMAF Oriented Perceptual Coding Based on Piecewise Metric Coupling
abstract
It has been recognized that videos have to be encoded in a rate-distortion optimized manner for high coding performance. Therefore, operational coding methods have been developed for conventional distortion metrics such as Sum of Squared Error (SSE). Nowadays, with the rapid development of machine learning, the state-of-the-art learning based metric Video Multimethod Assessment Fusion (VMAF) has been proven to outperform conventional ones in terms of the correlation with human perception, and thus deserves integration into the coding framework. However, unlike conventional metrics, VMAF has no specific computational formulas and may be frequently updated by new training data, which invalidates the existing coding methods and makes it highly desired to develop a rate-distortion optimized method for VMAF. Moreover, VMAF is designed to operate at the frame level, which leads to further difficulties in its application to today's block based coding. In this paper, we propose a VMAF oriented perceptual coding method based on piecewise metric coupling. Firstly, we explore the correlation between VMAF and SSE in the neighborhood of a benchmark distortion. Then a rate-distortion optimization model is formulated based on the correlation, and an optimized block based coding method is presented for VMAF. Experimental results show that 3.61% and 2.67% bit saving on average can be achieved for VMAF under the low_delay_p and the random_access_main configurations of HEVC coding respectively.
Zhengyi Luo 0001, Yan Huang 0033, Rong Xie 0004, Li Song 0001, C.-C. Jay Kuo
IEEE Trans. Image Process.5
2020 FACT: Fused Attention for Clothing Transfer with Generative Adversarial Networks
abstract
Clothing transfer is a challenging task in computer vision where the goal is to transfer the human clothing style in an input image conditioned on a given language description. However, existing approaches have limited ability in delicate colorization and texture synthesis with a conventional fully convolutional generator. To tackle this problem, we propose a novel semantic-based Fused Attention model for Clothing Transfer (FACT), which allows fine-grained synthesis, high global consistency and plausible hallucination in images. Towards this end, we incorporate two attention modules based on spatial levels: (i) soft attention that searches for the most related positions in sentences, and (ii) self-attention modeling long-range dependencies on feature maps. Furthermore, we also develop a stylized channel-wise attention module to capture correlations on feature levels. We effectively fuse these attention modules in the generator and achieve better performances than the state-of-the-art method on the DeepFashion dataset. Qualitative and quantitative comparisons against the baselines demonstrate the effectiveness of our approach.
Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001
AAAI3
2020 Toward Fine-Grained Facial Expression Manipulation
Jun Ling, Li Song 0001, Rong Xie 0004, Xiao Gu 0001
ECCV (28)3
2020 Hiding Private Information in Images From AI
abstract
Privacy protection attracts increasing concerns these days. People tend to believe that large social platforms will comply with the agreement to protect their privacy. However, photos uploaded by people are usually not treated to achieve privacy protection. For example, Facebook, the world's largest social platform, was found leaking photos of millions of users to commercial organizations for big data analytics. A common analytical tool used by these commercial organizations is the Deep Neural Network (DNN). Today's DNN can accurately identify people's appearance, body shape, hobbies and even more sensitive personal information, such as addresses, phone numbers, emails, bank cards and so on. To enable people to enjoy sharing photos without worrying about their privacy, we propose an algorithm that allows users to selectively protect their privacy while preserving the contextual information contained in images. The results show that the proposed algorithm can select and perturb private objects to be protected among multiple optional objects so that the DNN can only identify non-private objects in images.
Hanyu Xue, Bo Liu 0001, Ming Ding 0001, Li Song 0001, Tianqing Zhu
ICC4
2020 Realistic Talking Face Synthesis With Geometry-Aware Feature Transformation
abstract
Recent studies have shown remarkable success in synthesizing realistic talking faces by exploiting generative adversarial networks. However, existing methods are mostly target specific that cannot generate images of previously unseen people, and they suffer from artifacts such as blurriness and mismatching of facial details. In this paper, we tackle these problems by proposing a target-agnostic framework. We introduce a geometry-aware feature transformation module to achieve shape transfer while preserving the appearance of the source face. To further improve image quality of synthesized results, we present a multi-scale spatially-consistent transfer unit to maintain spatial consistency between the encoder and decoder features. Experimental results show that our model is able to synthesize photo-realistic talking faces which are previously unseen, outperforming state-of-the-art methods both qualitatively and quantitatively.
Jun Ling, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001
ICIP3
2020 Learning-Based Quality Enhancement For Scalable Coded Video Over Packet Lossy Networks
abstract
The layered feature of scalable video coding (SVC) offers a sufficient adaptation to unreliable transmission. When network condition drops sharply, enhancement layers will be abandoned, and only base layers are delivered. However, this will cause noticeable visual artifacts due to quality differences between different layers. To alleviate this problem, we novelly introduce a deep learning-based method into video reconstruction phase of scalable bitstreams. A super-resolution motivated recurrent network is proposed to extract and fuse features from both previous high-resolution frames and the current low-resolution frame. To the best of our knowledge, this is the first attempt to improve the performance of scalable bitstreams reconstruction by a specifically designed super-resolution network. By seamlessly integrating the accessible features, significant video quality improvements in terms of PSNR, SSIM, and VMAF are achieved. At the same time, the improvement of overall visual quality stability is apparent under packet lossy networks, indicating both efficiency and robustness of our approach.
Shengwei Yu, Xun Tong, Yan Huang 0033, Rong Xie 0004, Li Song 0001
ICME5
2020 A Deep Tracking and Segmentation Approach for Soccer Videos Visual Effects
Shenhui Peng, Li Song 0001, Jun Ling, Rong Xie 0004, Lin Li 0062
PRCV (2)2
2020 Deep Blind Video Quality Assessment for User Generated Videos
abstract
As short video industry grows up, quality assessment of user generated videos has become a hot issue. Existing no reference video quality assessment methods are not suitable for this type of application scenario since they are aimed at synthetic videos. In this paper, we propose a novel deep blind quality assessment model for user generated videos according to content variety and temporal memory effect. Content-aware features of frames are extracted through deep neural network, and a patch-based method is adopted to obtain frame quality score. Moreover, we propose a temporal memory-based pooling model considering temporal memory effect to predict video quality. Experimental results conducted on KoNViD-1k and LIVE-VQC databases demonstrate that the performance of our proposed method outperforms other state-of-the-art ones, and the comparative analysis proves the efficiency o f our temporal pooling model.
Jiapeng Tang, Rong Xie 0004, Xiao Gu 0001, Li Song 0001, Lin Li 0062
VCIP5
2020 A Hybrid Model for Natural Face De-Identiation with Adjustable Privacy
abstract
As more and more personal photos are shared and tagged in social media, security and privacy protection are becoming an unprecedentedly focus of attention. Avoiding privacy risks such as unintended verification, becomes increasingly challenging. To enable people to enjoy uploading photos without having to consider these privacy concerns, it is crucial to study techniques that allow individuals to limit the identity information leaked in visual data. In this paper, we propose a novel hybrid model consists of two stages to generate visually pleasing de-identified face images according to a single input. Meanwhile, we successfully preserve visual similarity with the original face to retain data usability. Our approach combines latest advances in GAN-based face generation with well-designed adjustable randomness. In our experiments we show visually pleasing de-identified output of our method while preserving a high similarity to the original image content. Moreover, our method adapts well to the verificator of unknown structure, which further improves the practical value in our real life.
Yunqian Wen, Bo Liu 0001, Rong Xie 0004, Yunhui Zhu, Jingyi Cao, Li Song 0001
VCIP6
2020 Quality of Experience Evaluation for Streaming Video Using CGNN
abstract
One of the principal contradictions these days in the field of video i s lying between the booming demand for evaluating the streaming video quality and the low precision of the Quality of Experience prediction results. In this paper, we propose Convolutional Neural Network and Gate Recurrent Unit (CGNN)-QoE, a deep learning QoE model, that can predict overall and continuous scores of video streaming services accurately in real time. We further implement state-of-the-art models on the basis of their works and compare with our method on six public available datasets. In all considered scenarios, the CGNN-QoE outperforms existing methods.
Zhiming Zhou 0001, Li Song 0001, Rong Xie 0004, Lin Li 0062
VCIP3
2020 Rate Distortion Optimization: A Joint Framework and Algorithms for Random Access Hierarchical Video Coding
abstract
This paper revisits the problem of rate distortion optimization (RDO) with focus on inter-picture dependence. A joint RDO framework which incorporates the Lagrange multiplier as one of parameters to be optimized is proposed. Simplification strategies are demonstrated for practical applications. To make the problem tractable, we consider an approach where prediction residuals of pictures in a video sequence are assumed to be emitted from a finite set of sources. Consequently the RDO problem is formulated as finding optimal coding parameters for a finite number of sources, regardless of the length of the video sequence. Specifically, in cases where a hierarchical prediction structure is used, prediction residuals of pictures at the same prediction layer are assumed to be emitted from a common source. Following this approach, we propose an iterative algorithm to alternatively optimize the selections of quantization parameters (QPs) and the corresponding Lagrange multipliers. Based on the results of the iterative algorithm, we further propose two practical algorithms to compute QPs and the Lagrange multipliers for the RA(random access) hierarchical video coding: the first practical algorithm uses a fixed formula to compute QPs and the Lagrange multipliers, and the second practical algorithm adaptively adjusts both QPs and the Lagrange multipliers. Experimental results show that these three algorithms, integrated into the HM 16.20 reference software of HEVC, can achieve considerable RD improvements over the standard HM 16.20 encoder, in the common RA test configuration.
En-Hui Yang, Dake He, Li Song 0001, Xiang Yu 0001
IEEE Trans. Image Process.4
2019 Gan Based Multi-Exposure Inverse Tone Mapping
abstract
High dynamic range (HDR) imaging provide larger range of luminosity and wider color gamut than conventional low dynamic range (LDR) imaging. The method which transforms LDR contents to HDR contents is called inverse tone mapping. After deep neural networks are used in inverse tone mapping problem, researchers mostly focus on transforming normal exposure LDR images to HDR. However, when people use inverse tone mapping in practice, they get some ill-exposed images as well. The state-of-art algorithms can't transform these images to HDR well.In this work, we propose an end-to-end multi-exposure inverse tone mapping (MITM) framework based on existing generative adversarial network (GAN). This framework can transform a single LDR image not only at normal exposure, but also at unsuitable exposure to a normal exposure HDR image. We use histogram equalization to preprocess the luma of the input LDR images; when training the model, we use intrinsic image decomposition to divide the output HDR images into illuminance and reflectance components and use these two components to constrain the luminance information and the color information separately. This framework can adjust the unsuitable exposure and provide a better viewing experience than other state-of-art algorithms in the experimental results.
Shiyu Ning, Rong Xie 0004, Li Song 0001
ICIP4
2019 Advanced CNN Based Motion Compensation Fractional Interpolation
abstract
Fractional-sample precision motion compensation has been widely adopted in a series of video coding standards to further improve compression efficiency. Usually, signal decomposition based interpolation filters are used to generate fractional samples from integer pixels. However, the coefficients of these finite impluse response filters may not be suitable for varied video contents and coding conditions because of the assumption when designing these filters. In this paper, we regard the fractional interpolation process as an image generation task, which utilizes the real interger position samples at the reference block to predict and generate fractional samples that are much closer to current coding block. We use the con-volutional neural netwok (CNN) as the generator. Moreover, to make the best of CNN's powerful nonlinear learning ability, instead of inputting the reference block directly, we separately input the corresponding prediction and residual parts of reference block. The proposed dual-input CNN-based interpolation scheme has been incorporated into the HEVC framework and experimental results demonstrate our approach achieves average 0.9% bitrate reduction.
Han Zhang 0030, Li Li 0040, Li Song 0001, Xiaokang Yang 0001, Zhu Li 0001
ICIP3
2019 VMAF Oriented Perceptual Optimization for Video Coding
abstract
In the light of low costs and automatic assessment, objective visual quality metrics enjoy many important applications such as perceptual coding. Recently multiple metrics obtain further improvement by means of machine learning. However, due to the absence of specific formulas, it's often hard to incorporate learning based metrics into video coding. In this paper, taking the state-of-the-art learning based metric VMAF for example, we propose a method of perceptual coding in an inferential manner for learning based metrics. The rate distortion optimization is adapted during coding as well. Experimental results show that compared with conventional methods, the proposed method can achieve obvious bitrate saving under HEVC coding.
Zhengyi Luo 0001, Yan Huang 0033, Rong Xie 0004, Li Song 0001
ISCAS5
2019 Reinforcement Learning Based Adaptive Bitrate Algorithm for Transmitting Panoramic Videos
abstract
Panoramic videos have become more and more popular now. 360-degree videos give users a better experience but put forward a higher request for network at the same time. Many kinds of solutions to meet the high need of bandwidth have been proposed, such as tiled-based transmission, layer-based transmission and so on. Then how to choose the most suitable bitrates to make full use of the network resource is the problem to be solved urgently. In this paper, we propose a method based on reinforcement learning(RL) algorithm to select the bitrates of the region-of-interest adaptively for panoramic videos. We also compare RL algorithm with three traditional algorithms when changing the Field of View(FOV) and find that RL algorithm proves to perform better in minimizing the rebuffer time and providing higher quality video contents under various network conditions.
Xiaona Wu, Xun Tong, Rong Xie 0004, Li Song 0001
ISCAS5
2019 CNN Accelerated Intra Video Coding, Where Is the Upper Bound?
abstract
The very high complexity of the High Efficiency Video Coding standard (HEVC) is the main hurdle for its wide deployment and use. To tackle this problem, a number of recent research outcomes exploit Convolutional Neural Network (CNN) in each HEVC module for reducing the coding complexity. In this paper an effective method to analyse the potential of CNN techniques to reduce the computational cost of HEVC is proposed. A theoretical upper bound for the effectiveness of this approach in common HEVC modules is investigated. The theoretical maximum of learning-based complexity reduction in HEVC and possible reasons for Rate-Distortion (RD) loss are investigated. On the basis of this analysis, an Intra Video Coding Acceleration (IVCA) scheme is proposed, where Border Considered CNN (BC-CNN) based Coding Unit (CU) partition and heuristic Prediction Unit (PU) partition are seamlessly integrated. According to the experimental results, 66.7% of intra coding time can be saved with negligible 1.71% Bjøntegaard delta bit-rate (BDBR) loss. These results partially demonstrate the superiority of the proposed technique against other state-of-the-art approaches aiming at reducing HEVC complexity in intra mode.
Yan Huang 0033, Li Song 0001, Ebroul Izquierdo
PCS2
2019 JND-based Perceptual Rate Distortion Optimization for AV1 Encoder
abstract
AV1 is the next-generation open video coding format, and it can achieve significant coding efficiency with novel coding tools. It supports Lagrangian rate distortion optimization (RDO) method to optimize the coding performance. However, the distortion and the Lagrangian multiplier used in RDO ignore the characteristics of human visual system (HVS), which leads to insufficiency for perceptual video coding. To solve this problem, a perceptual RDO scheme based on the Just Noticeable Distortion (JND) threshold of HVS is proposed. The JND for each pixel is first measured according to three perceptual features: luminance adaptation, masking effects and structure sensitivity. Based on the observation that the regions with smaller distortion visibility thresholds are more sensitive to HVS, a JND-based Lagrangian multiplier is derived to adaptively adjust the rate-distortion (RD) performance for each coding block. Experiments demonstrate that the proposed method can achieve an average SSIM-based -3.93% BD-Rate saving compared with the original AV1 encoder, which effectively improve the coding performance.
Li Song 0001, Rong Xie 0004, Jingning Han, Yaowu Xu
PCS2
2019 Identifying and Pruning Redundant Structures for Deep Neural Networks
abstract
Deep convolutional neural networks have achieved considerable success in the field of computer vision. However, it is difficult to deploy state-of-the-art models on resource-constrained platforms due to their high storage, memory bandwidth, and computational costs. In this paper, we propose a structured pruning method which employs a three-step process to reduce the resource consumption of neural networks. First, we train an initial network on the training set and evaluate it on the validation set. Next, we introduce an iterative pruning and fine-tuning algorithm to identify and prune redundant structures, which results in a pruned network with a compact architecture. Finally, we train the pruned network from scratch on both the training set and validation set to obtain the final accuracy on the test set. In the experiments, our pruning method significantly reduces the model size (by 87.2% on CIFAR-10), saves inference time (53.3% on CIFAR-10), and achieves better performance as compared to recent state-of-the-art methods.
Wenyao Gan, Li Song 0001, Li Chen 0021, Rong Xie 0004, Xiao Gu 0001
VCIP2
2019 FPGA Based Video Transcoding System with 2K-4K Super-Resolution Conversion
abstract
We present a FPGA-based system supporting video stream transcoding with 2k full high-definition (FHD) video to 4k ultra high-definition (UHD) video super- resolution(SR) conversion. Our system focuses on building a functional pipeline with convolutional neural network (CNN) accelerator and real-time video codec unit for converting H.264 video stream to H.265/HEVC video stream. The overall video processing system can be used as an important plug-in module in the video streaming network to improve the video stream service quality.
Yuzhuo Wei, Li Chen 0021, Rong Xie 0004, Li Song 0001, Xiaoyun Zhang 0001
VCIP4
2019 Deep Feature Guided Image Retargeting
abstract
Image retargeting is the technique to display images via devices with various aspect ratios and sizes. Traditional content-aware retargeting methods rely on low-level features to predict pixel-wise importance and can hardly preserve both the structure lines and salient regions of the source image. To address this problem, we propose a novel adaptive image warping approach which integrates with deep convolutional neural network. In the proposed method, a visual importance map and a foreground mask map are generated by a pre-trained network. The two maps and other constraints guide the warping process to yield retargeted results with less distortions. Extensive experiments in terms of visual quality and a user study are carried out on the widely used RetargetMe dataset. Experimental results show that our method outperforms current state-of-art image retargeting methods.
Jinan Wu, Rong Xie 0004, Li Song 0001, Bo Liu 0001
VCIP3
2019 An Improved QoE Evaluation Model for HTTP Adaptive Streaming
abstract
HTTP adaptive streaming (HAS) is an adaptive bitrate streaming technique that enables high quality streaming of media content over the Internet delivered from conventional HTTP web servers. ITU-T Rec. P.1203.3 is the first standardized Quality of Experience model for audiovisual HTTP Adaptive Streaming. It takes into account the subjective impact of HAS-typical effects(such as buffering, quality switches) on users. But the buffering inputs required for the model is too complex and redundant, and it does not consider the impact of the worst video quality on QoE. In the paper we optimize the ITU-T Rec. P.1203.3 model in the above two aspects, simplify model input and improve model accuracy and stability.
Zaixin Yang, Tiantian He 0005, Li Song 0001, Rong Xie 0004, Xiao Gu 0001
VCIP3
2019 Gated-GAN: Adversarial Gated Networks for Multi-Collection Style Transfer
abstract
Style transfer describes the rendering of an image's semantic content as different artistic styles. Recently, generative adversarial networks (GANs) have emerged as an effective approach in style transfer by adversarially training the generator to synthesize convincing counterfeits. However, traditional GAN suffers from the mode collapse issue, resulting in unstable training and making style transfer quality difficult to guarantee. In addition, the GAN generator is only compatible with one style, so a series of GANs must be trained to provide users with choices to transfer more than one kind of style. In this paper, we focus on tackling these challenges and limitations to improve style transfer. We propose adversarial gated networks (Gated-GAN) to transfer multiple styles in a single model. The generative networks have three modules: an encoder, a gated transformer, and a decoder. Different styles can be achieved by passing input images through different branches of the gated transformer. To stabilize training, the encoder and decoder are combined as an auto-encoder to reconstruct the input images. The discriminative networks are used to distinguish whether the input image is a stylized or genuine image. An auxiliary classifier is used to recognize the style categories of transferred images, thereby helping the generative networks generate images in multiple styles. In addition, Gated-GAN makes it possible to explore a new style by investigating styles learned from artists or genres. Our extensive experiments demonstrate the stability and effectiveness of the proposed model for multi-style transfer.
Chang Xu 0002, Xiaokang Yang 0001, Li Song 0001, Dacheng Tao
IEEE Trans. Image Process.4
2018 Learning an Inverse Tone Mapping Network with a Generative Adversarial Regularizer
abstract
Transferring a low-dynamic-range (LDR) image to a high-dynamic-range (HDR) image, which is the so-called inverse tone mapping (iTM), is an important imaging technique to improve visual effects of imaging devices. In this paper, we propose a novel deep learning-based iTM method, which learns an inverse tone mapping network with a generative adversarial regularizer. In the framework of alternating optimization, we learn a U-Net-based HDR image generator to transfer input LDR images to HDR ones, and a simple CNN-based discriminator to classify the real HDR images and the generated ones. Specifically, when learning the generator we consider the content-related loss and the generative adversarial regularizer jointly to improve the stability and the robustness of the generated HDR images. Using the learned generator as the proposed inverse tone mapping network, we achieve superior iTM results to the state-of-the-art methods consistently.
Shiyu Ning, Hongteng Xu, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001
ICASSP3
2018 GPU Based Motion-Compensated Frame Interpolation Acceleration for Future Video Coding
abstract
Being developed by Joint Video Exploration Team (JVET), Future Video Coding (FVC) aims at higher resolutions and higher compression performance than the state-of-the-art HEVC standard, undoubtedly at the cost of further computing increases. As an efficient computing platform, Graphics Processing Unit (GPU) is often used to accelerate encoding. But with the adoption of instruction set acceleration in the reference software of FVC, previous methods often become less efficient or even lead to a lower speed. In this paper, based on the comparative analysis of the time consumption between HEVC and FVC, we propose a GPU based acceleration method for the most computation-intensive step - frame interpolation of FVC, where frame caching strategy and a multi-stream mechanism is designed to make the best of GPU resources. Experimental results show that compared with the instruction set accelerated reference software of FVC, our method could achieve average 67.12% speed-up gains on the interpolation module and average 6.35% speed-up gains on overall encoding with exactly the same performance as before.
Jianlun Tang, Yan Huang 0033, Rong Xie 0004, Zhengyi Luo 0001, Li Song 0001
ICIP5
2018 Frame Interpolation via Refined Deep Voxel Flow
abstract
Traditional frame interpolation methods first estimate motion between two consecutive frames and then synthesize intermediate frames. This problem is challenging because of complex motion and video scenes. In this paper, we present an end-to-end deep network for frame interpolation problem. Based on a video synthesis method deep voxel flow (DVF), refinement modules are designed to increase the accuracy of voxel flow, which we call Refined DVF (RDVF). A deeper architecture with more convolution and deconvolution layers is also utilized to help extract motion. Our results greatly improve the performance of original DVF and compare favorably to state-of-the-art methods both quantitatively and qualitatively.
Zhifeng Zhang 0003, Li Chen 0021, Rong Xie 0004, Li Song 0001
ICIP4
2018 An MCMC based Efficient Parameter Selection Model for x265 Encoder
abstract
As an open-source and computationally efficient High Efficiency Video Coding (HEVC) encoder, x265 has been gaining increasing popularity in video applications. x265 provides numerous encoding parameters in view of flexibility. However, proper and efficient setting of parameters often becomes a great challenge in practice. In this paper, we deeply investigate the influence of x265 parameters based on the Slow preset and pick out important parameters in terms of efficiency and complexity. Then a Markov Chain Monte Carlo (MCMC) based algorithm is proposed for efficient parameter adaptation at the target encoding time. This paper shows that carefully selected low-complexity encoding configurations can achieve the coding efficiency comparable to that of high-complexity ones. Specifically, average 26.72% encoding time reduction can be achieved while maintaining similar Rate Distortion (RD) performance to x265 presets using the proposed algorithm.
Yan Huang 0033, Li Song 0001, Rong Xie 0004, Zhengyi Luo 0001
ISCAS2
2018 Masking Effects Based Rate Control Scheme for High Efficiency Video Coding
abstract
This paper presents a masking effects based rate control scheme for high efficiency video coding (HEVC). Rate control is regarded as a very effective tool to improve the performance of video coding under the limited bandwidth. However, the state-of-the-art rate control algorithm based on R-X model ignores the characteristics of human visual system (HVS), which leads to poor performance in subjective quality. Moreover, some structural similarity (SSIM) or saliency based perceptual rate control algorithms only consider spatial characteristics. Since spatial and temporal visual masking effects can better reflect the characteristics of HVS, in this paper masking effects based perceptual factor for coding tree unit (CTU) is proposed, which takes both texture complexity and motion information into account. Then the proposed perceptual factor is utilized to guide bit allocation in CTU-level rate control. Experimental results show that the proposed scheme can effectively improve the coding performance compared with the R-λ algorithm.
Hao Wang 0073, Li Song 0001, Rong Xie 0004, Zhengyi Luo 0001
ISCAS2
2018 Rate-mixed HEVC Tile based 360 Video Streaming System
abstract
Recently, tile-based viewport adaptation is a popular method for 360 video streaming. Our demonstration adopts a rate-mixed transmission approach utilizing a VR adaptation agent at the server end for viewport-based streaming, which is client-compatible and can be scalable to different users. The FOV prediction is applied to improve the viewing experience. The feasibility of our system to head-mounted displays is verified, which can reduce bandwidth consumption by up to 36%.
Xu Liu 0006, Rong Xie 0004, Li Song 0001
VCIP5
2018 An improved Real-Time Video Communication System
abstract
In this paper, we optimized the Linphone-based real-time video communication system. Firstly, we used HEVC to replace the H.264 in order to reduce the bandwidth pressure while reducing the buffer delay by configuring the appropriate encoding parameters. Secondly, we added an effective bitrate adaptive algorithm based on additive increase and multiplicative decrease (AIMD) in the system, which can effectively reduce the packet loss rate in the network with high bandwidth fluctuations.
Zhaoliang Ma, Shengwei Yu, Yongcheng Huang, Rong Xie 0004, Li Song 0001
VCIP5
2017 CNN based post-processing to improve HEVC
abstract
In this paper, we propose a frame-based dynamic metadata post-processing scheme in HEVC. Video sequence is classified into different categories contains complexity of video content and quality indicator for each frame, an up-to-one byte flag embedded in the bitstream is transferred as side information. Meanwhile dynamic metadata contains classification information indicates the offline training of separate network models. Specifically, we adopt a 20-layers CNN (Con-volutional Neural networks) model to extract more meaningful information from the reconstructed error and improve the filtering performance. Experimental results shows that our proposed post-processing scheme leads on average 1.6% BD-rate reduction compared with HEVC baseline on the six sequences given in 2017 ICIP Grand Challenge.
Chen Li 0021, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001
ICIP2
2017 Lagrangian method based Rate-Distortion Optimization revisited for dependent video coding
abstract
Video encoding is based on the DPCM framework where temporal prediction coding introduces Rate-Distortion (RD) dependence. The RD operating point of the current unit depends on the particular choices of RD points of its reference units. Unfortunately, common Lagrangian optimization method based Rate-Distortion Optimization (RDO) for video coding is based on an independence assumption which omits the RD dependences, and thus compromises the RD performance. In this paper, we revisit the Lagrangian optimization method based RDO for dependent video coding. A theoretical RD dependence decoupling method based on independent distortion decomposition is firstly presented. After the discussion of reasonability of the theoretical decoupling method, the practical One Step Ahead Decoupling Strategy (OSADS) is proposed. After implemented on the HEVC encoder, the strategy achieves average 2.1% BD-rate saving compared with the HM encoder under the same low-delay P configuration.
Li Song 0001, Zhengyi Luo 0001, Rong Xie 0004
ICIP2
2017 Rate control model for high dynamic range video
abstract
This paper describes a luminance based rate control (RC) model for high dynamic range (HDR) video. A novel mathematical relationship between luminance and bit allocation of a coding tree unit (CTU) is presented. By adjusting the existing RC algorithm through the proposed model, a better balance between dark and bright areas can be achieved and -4.4% gains can be obtained in terms of average BD-Rate (tPSNR-XZY). Moreover, subjective assessment also shows that, compared with the existing RC model, the proposed method can convey a wider range of perceptible shadow and highlight more details.
Lixun Bai, Li Song 0001, Rong Xie 0004, Liang Zhang 0026, Zhengyi Luo 0001
VCIP2
2017 Weight-based bit allocation scheme for VR videos in HEVC
abstract
Since VR videos shown on the head-mounted display (HMD) is omnidirectional, the average distortion of VR videos in all directions shall be calculated in spherical domain. Several metrics have been proposed to calculate the coding loss of VR videos in spherical domain, including S-PSNR, WS-PSNR, CPP-PSNR. The above metrics are all improved based on PSNR by creating a weight map. This paper aims to optimize the rate control scheme on HEVC mostly for WS-PSNR. According to the weight map of WS-PSNR, regions with more weights in plain are more important, thus more bits shall be allocated to the important regions. We use weight-based rate control scheme to realize the above thoughts. Weight-based rate control scheme defines bit per weight (bpw) instead of bpp. Large values of bpw indicate that the important regions deserve high bitrates, thus probably achieving better quality. Consequently, the proposed rate control scheme improves the video quality of VR videos, which leads to average gain of 2.1%, 4.3% and 1.5% of S-PSNR, WS-PSNR and CPP-PSNR.
BiJia Li, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001
VCIP2
2017 A generic method to improve no-reference image blur metric accuracy in video contents
abstract
We present in this work a generic and effective method to increase the prediction accuracy of no-reference image/video blur assessment facing the real-world content diversity. We demonstrate that benchmarking no reference image blur metrics, fitting a single logistic function to map the objective predictions to subjective scores in the well-known databases like LIVE or TID2008/2013, introduce biased fitting results towards better predictions only in the central part of the score scale. We find out that a multi-fitting approach, using the correlation parameters between subjective scores and objective predictions for content clustering and then conducting logistic fitting for each content type, can evidently improve the metric prediction accuracy in the full score scale. Besides, the overall prediction variance is also reduced with the proposed scheme, presenting more consistent results insensitive of content variation. We prove that the proposed method is of practical meaning to facilitate blur assessment techniques validated on limited databases to the vastly abundant real-life content types.
Yankai Liu, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001
VCIP2
2017 Two-stream deep encoder-decoder architecture for fully automatic video object segmentation
abstract
We propose a two-stream Deep Encoder-Decoder architecture to tackle the task of fully automatic video object segmentation. Both two streams, i.e., ImSeg-Stream (for static image segmentation) and MoSeg-Stream (for optical flow segmentation), hold the totally same Encoder-Decoder architecture. The Encoder part generates a low-resolution mask with accurate locations and smooth boundaries, while the Decoder part refines the details of initial mask and enlarges its resolution via integrating lower-level features progressively. At last two streams learn to integrate for better results. Moreover, to handle the problem of inadequate video object segmentation datasets, we propose a seeking strategy to generate a large-scale handcrafted dataset for training. Experiments on two standard datasets demonstrate that proposed method outperforms most state-of-the-art methods in both segmentation accuracy and run time.
Li Song 0001, Rong Xie 0004
VCIP2
2017 Learning a convolutional neural network for fractional interpolation in HEVC inter coding
abstract
Motion compensated prediction (MCP) is an effective technology for video coding to improve compression efficiency. Fractional sample precision prediction is utilized in HEVC to further remove temporal redundancy, and finite impulse response (FIR) filters designed using decomposition of the discrete cosine transform are applied to generate samples that do not fall on the integer positions. However, the coefficients of these DCT-based interpolation filters are fixed, which may not be able to adapt to varied video content. Inspired by the remarkable success of convolutional neural network (CNN) in the single image super-resolution task, we propose to learn a convolutional neural network for fractional interpolation in HEVC inter prediction. Compared with super-resolution, there is one big difference in fractional interpolation - fractional interpolation needs to maintain samples at integer positions while super-resolution generates a whole high-resolution image. Another difference is no real ground truth is available in fractional interpolation process. To overcome these two challenges, we introduce a constraint strategy to the training phase of the original super-resolution network as well as a specially designed preprocessing step which reuses the DCTIF interpolation process. Unlike other previous work, our proposed approach simultaneously generating the fractional positions from one network and experimental results show our proposed approach achieves 0.45% BD-Rate reduction under the low-delay-P configuration on average.
Han Zhang 0030, Li Song 0001, Zhengyi Luo 0001, Xiaokang Yang 0001
VCIP2
2017 DRIMUX: Dynamic Rumor Influence Minimization with User Experience in Social Networks
abstract
With the soaring development of large scale online social networks, online information sharing is becoming ubiquitous everyday. Various information is propagating through online social networks including both the positive and negative. In this paper, we focus on the negative information problems such as the online rumors. Rumor blocking is a serious problem in large-scale social networks. Malicious rumors could cause chaos in society and hence need to be blocked as soon as possible after being detected. In this paper, we propose a model of dynamic rumor influence minimization with user experience (DRIMUX). Our goal is to minimize the influence of the rumor (i.e., the number of users that have accepted and sent the rumor) by blocking a certain subset of nodes. A dynamic Ising propagation model considering both the global popularity and individual attraction of the rumor is presented based on a realistic scenario. In addition, different from existing problems of influence minimization, we take into account the constraint of user experience utility. Specifically, each node is assigned a tolerance time threshold. If the blocking time of each user exceeds that threshold, the utility of the network will decrease. Under this constraint, we then formulate the problem as a network inference problem with survival theory, and propose solutions based on maximum likelihood principle. Experiments are implemented based on large-scale real world networks and validate the effectiveness of our method.
Luoyi Fu, Li Song 0001, Xinbing Wang
IEEE Trans. Knowl. Data Eng.4
2016 DRIMUX: Dynamic Rumor Influence Minimization with User Experience in Social Networks
abstract
Rumor blocking is a serious problem in large-scale social networks. Malicious rumors could cause chaos in society and hence need to be blocked as soon as possible after being detected. In this paper, we propose a model of dynamic rumor influence minimization with user experience (DRIMUX). Our goal is to minimize the influence of the rumor (i.e., the number of users that have accepted and sent the rumor) by blocking a certain subset of nodes. A dynamic Ising propagation model considering both the global popularity and individual attraction of the rumor is presented based on realistic scenario. In addition, different from existing problems of influence minimization, we take into account the constraint of user experience utility. Specifically, each node is assigned a tolerance time threshold. If the blocking time of each user exceeds that threshold, the utility of the network will decrease. Under this constraint, we then formulate the problem as a network inference problem with survival theory, and propose solutions based on maximum likelihood principle. Experiments are implemented based on large-scale real world networks and validate the effectiveness of our method.
Luoyi Fu, Li Song 0001, Xinbing Wang, Xue (Steve) Liu
AAAI4
2016 A parallel-fusion RNN-LSTM architecture for image caption generation
abstract
The models based on deep convolutional networks and recurrent neural networks have dominated in recent image caption generation tasks. Performance and complexity are still eternal topic. Inspired by recent work, by combining the advantages of simple RNN and LSTM, we present a novel parallel-fusion RNN-LSTM architecture, which obtains better results than a dominated one and improves the efficiency as well. The proposed approach divides the hidden units of RNN into several same-size parts, and lets them work in parallel. Then, we merge their outputs with corresponding ratios to generate final results. Moreover, these units can be different types of RNNs, for instance, a simple RNN and a LSTM. By training normally using NeuralTalk1platform on Flickr8k dataset, without additional training data, we get better results than that of dominated structure and particularly, the proposed model surpass GoogleNIC in image caption generation.
Minsi Wang, Li Song 0001, Xiaokang Yang 0001, Chuanfei Luo
ICIP2
2016 Improved intra angular prediction with novel interpolation filter and boundary filter
abstract
In this paper, two improved intra angular prediction methods are proposed to enhance coding performance. The first method applies new four-tap interpolation filter algorithm. The reference samples at fractional position are interpolated by DCT-based or Gaussian interpolation filter. The second method proposes extended boundary prediction filter to reduce the prediction error. The experimental results show that for AI configuration, the overall coding gain is about 0.85% on average comparing to HEVC reference software while maintaining almost the same coding time.
Rujun Wei, Rong Xie 0004, Li Song 0001, Liang Zhang 0026, Wenjun Zhang 0001
PCS3
2016 Shot boundary detection using convolutional neural networks
abstract
Video shot boundary detection (SBD) is necessary for further video analysis like video retrieval and annotation. Great efforts have been made to develop SBD algorithms for speed and accuracy. Most works implement frame histogram as features to measure similarity for detection. However, when changes between consecutive shot boundaries are small and backgrounds of them are highly similar, most state-of-the-art methods miss these boundaries thus cannot achieve high accuracy of detection. In this paper we propose a novel SBD framework with Convolutional Neural Networks (CNNs). Firstly we adopt a candidate segment selection method to locate the positions of shot boundaries coarsely using adaptive thresholds and eliminate most non-boundary frames. Then CNN is implemented to extract representative features of frames in candidate segments. Finally cut and gradual transitions can be obtained by using a novel pattern-matching method based on a new similarity strategy. Experiments on TRECVID 2001 test data demonstrate that the proposed scheme outperforms the state-of-the-art methods and achieves high accuracy of detection.
Li Song 0001, Rong Xie 0004
VCIP2
2016 Evaluation of beyond-HEVC entropy coding methods for DCT transform coefficients
abstract
Entropy coding, which acts as one of the most important compression tools in video coding standard, had been improved step by step for HEVC. There are also several advanced methods which provide better performance than current solutions of HEVC proposed during the standardization of HEVC. However, these methods are all tested in different conditions. Comprehensive evaluation of these advanced methods under a common scenario is desired to indicate where the potential improvement of entropy coding may come from for next generation video codec. In this paper, we first introduce several advanced entropy coding methods for DCT transform coefficients, which aim to improve CABAC performance from two aspects - context modeling and probability updating. Then some modifications based on these original ones are presented. Comprehensive comparison of these methods is conducted under common test conditions. Besides, some combined methods of these two aspects are also tested. Experimental results show that all individual approaches can achieve coding gain and two new combined methods can reduce the BD-Rate up to 1.7%, 1.2% and 1.0% on common test sequences and 1.4%, 1.0% and 1.1% on 4K sequences under all intra, random access and low delay configurations, respectively.
Han Zhang 0030, Li Song 0001, Xiaokang Yang 0001, Zhengyi Luo 0001
VCIP2
2016 Identifying effective initiators in OSNs: from the spectral radius perspective
abstract
Abstract In this paper, we focus on maximizing the influence of online social networks (OSNs). Particularly, we try to answer how to select proper information initiators such that information can propagate as widely as possible. We stress our attention on the susceptible‐infected model, a type of epidemic models, to describe the process of information diffusion. In general, OSNs can be classified into two categories, Facebook‐like OSNs and Twitter‐like OSNs. The former ones require bidirectional connections, while the latter do not, so we use the undirected unweighted graph and directed unweighted graph to describe them, respectively. We also pay additional attention to the nonidentity of the link probability on information transmission and build the weight graph, which can also cover both the two types of OSNs. In order to determine values of weight graph's weights, we introduce a learning method to obtain useful factors from raw data for assessing the true link probability on information transmission. Based on spectral analysis within the three graphs, our investigations on the information diffusion show that the spectral radius of the graph adjacency matrix can reflect the capability of information propagation, according to which we could determine effective initiators. We conduct our simulations on real OSNs. Experimental results show that our approach could effectively discover the initiators that spread information widely. Copyright © 2016 John Wiley & Sons, Ltd.
Songjun Ma, Weijie Wu, Li Song 0001, Xiaohua Tian, Xinbing Wang
Wirel. Commun. Mob. Comput.4
2015 An Optimized Pixel-Wise Weighting Approach for Patch-Based Image Denoising
abstract
Most existing patch-based image denoising algorithms filter overlapping image patches and aggregate multiple estimates for the same pixel via weighting. Current weighting approaches always assume the restored estimates as independent random variables, which is inconsistent with the reality. In this letter, we analyze the correlation among the estimates and propose a bias-variance model to estimate the Mean Squared Error (MSE) under various weights. The new model exploits the overlapping information of the patches; it then utilizes the optimization to try to minimize the estimated MSE. Under this model, we propose a new weighting approach based on Quadratic Programming (QP), which can be embedded into various denoising algorithms. Experimental results show that the Peak Signal to Noise Ratio (PSNR) of algorithms like K-SVD and EPLL can be improved by around 0.1 dB under a range of noise levels. This improvement is promising, since it is gained independent to which image model is used, especially when the gain from designing new image models becomes less and less.
Jianzhou Feng 0001, Li Song 0001, Xiaoming Huo, Xiaokang Yang 0001, Wenjun Zhang 0001
IEEE Signal Process. Lett.2
2014 Are we still friends: Kernel multivariate survival analysis
abstract
Online Social Network becomes the most prevalent platform for exchanging information between users, maintaining friendships online. As is well-known to us, however, some friendships even those intimate ones might vanish. Therefore, precisely modeling and predicting state of each online relationship is worthwhile in many respects. For social communication services such modeling permits new and novel online services. In addition, constructing this model might enlighten us in exploiting information spreading pattern in online social network. In this paper, we propose a model in determining a probability distribution which describes the ‘surviving time’ of each friendships by applying one commonly used method in sociology, survival analysis. We discuss a series of social explanatory variables that highly affect this probability distribution. Moreover, methods in the moving average process are devoted to determining the appropriate parameter in survival model. Furthermore, to avoid the high computational complexity in kernel learning we impose sparsity in our model. Finally, with the experiments on real data, the proposed survival model is proven to be of high accuracy, and thus of great potential for further applications.
Shiyu Liang, Ruotian Luo, Songjun Ma, Weijie Wu, Li Song 0001, Xiaohua Tian, Xinbing Wang
GLOBECOM6
2014 Blind image quality assessment based on a new feature of nature scene statistics
abstract
A recently proposed model, known as blind/referenceless image spatial quality evaluator (BRISQUE), achieves the state-of-the-art performance in context of blind image quality assessment (IQA). This model used the predefined generalized Gaussian distribution (GGD) to describe the regularity of natural scene statistics, introducing fitting errors due to variations of image contents. In this paper, a more generalized model is proposed to better characterize the regularity of extensive image contents, which is learned from the concatenated histograms of mean subtracted contrast normalized (MSCN) coefficients and pairwise products of MSCN coefficients of neighbouring pixels. The new feature based on MSCN shows its capability of preserving intrinsic distribution of image statistics. Consequently support vector machine regression (SVR) can map it to more accurate image quality scores. Experimental results show that the proposed approach achieves a slight gain from BRISQUE, which indicates the crafted GGD modelling step in BRISQUE is not essential for final performance.
Li Song 0001, Yi Xu 0001, Gengjian Xue, Yi Zhou 0003
VCIP1
2014 Evaluation of Different Algorithms of Nonnegative Matrix Factorization in Temporal Psychovisual Modulation
abstract
Temporal psychovisual modulation (TPVM) is a newly proposed information display paradigm, which can be implemented by nonnegative matrix factorization (NMF) with additional upper bound constraints on the variables. In this paper, we study all the state-of-the-art algorithms in NMF, extend them to incorporate the upper bounds and discuss their potential use in TPVM. By comparing all the NMF algorithms with their extended versions, we find that: 1) the factorization error of the truncated alternating least squares algorithm always fluctuates throughout the iterations, 2) the alternating nonnegative least squares based algorithms may slow down dramatically under the upper bound constraints, and 3) the hierarchical alternating least squares (HALS) algorithm converges the fastest and its final factorization error is often the smallest among all the algorithms. Based on the experimental results of the HALS, we propose a guideline of determining the parameter setting of TPVM, that is, the number of viewers to support and the scaling factor for adjusting the light intensity of the images formed by TPVM. This paper will facilitate the applications of TPVM.
Jianzhou Feng 0001, Xiaoming Huo, Li Song 0001, Xiaokang Yang 0001, Wenjun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2013 Image restoration via efficient Gaussian mixture model learning
abstract
Expected Patch Log Likelihood (EPLL) framework using Gaussian Mixture Model (GMM) prior for image restoration was recently proposed with its performance comparable to the state-of-the-art algorithms. However, EPLL uses generic prior trained from offline image patches, which may not correctly represent statistics of the current image patches. In this paper, we extend the EPLL framework to an adaptive one, named A-EPLL, which not only concerns the likelihood of restored patches, but also trains the GMM to fit for the degraded image. To efficiently estimate GMM parameters in A-EPLL framework, we improve a recent Expectation-Maximization (EM) algorithm by exploiting specific structures of GMM from image patches, like Gaussian Scale Models. Experiment results show that A-EPLL outperforms the original EPLL significantly on several image restoration problems, like inpainting, denoising and deblurring.
Jianzhou Feng 0001, Li Song 0001, Xiaoming Huo, Xiaokang Yang 0001, Wenjun Zhang 0001
ICIP2
2013 Paralleling variable block size motion estimation of HEVC on multi-core CPU plus GPU platform
abstract
Motion estimation with variable block sizes (VBSME) is one of the most complex models in the HEVC encoder. The HEVC standard supports up to 12 variable block sizes ranging from 4×8/8×4 to 64×64 for motion estimation (ME) and motion compensation (MC). This feature contributes substantial coding gain compared with 7 variable block sizes in H.264/AVC at the cost of huge computational complexity. The VBSME becomes the bottleneck for real time encoding. In this paper, we propose novel strategies for parallel acceleration the VBSME in HEVC encoder based on multi-core CPU plus many-core GPU platform. Firstly, a two-stage ME strategy is proposed for dividing ME task onto the CPU and the GPU. Then, a span-wavefront VBSME sequence is designed for efficient synchronization between the threads on the CPU and the threads on the GPU. Experimental results show that the speed of the HEVC encoder with the proposed strategies reaches about 28 fps for 1080P videos with a little compression performance degradation.
Li Song 0001, Min Chen 0011
ICIP2
2013 Foreground detection: Combining background subspace learning with object smoothing model
abstract
Foreground detection is a challenging problem in complex scenes. In this paper, a novel foreground detection method is proposed which combines background subspace learning with object smoothing model. Considering background scenes in consecutive frames are almost the same, they are approximated using an efficient subspace learning technique which is based on 2D images. Due to the pixels of objects are usually clustered, an object smoothing model is adopted where a spatial smoothing constraint is imposed on its values during the estimation, and then it can be solved as a regularized matrix restoration problem with a spatial smoothing constraint. As a result, isolated noises can be suppressed while clustered foreground pixels can be preserved. We test our method on some challenging sequences and compare it with some other techniques. Experimental results show its effectiveness and robustness.
Gengjian Xue, Li Song 0001, Jun Sun 0005, Jun Zhou 0007
ICME2
2013 Shaking video synthesis for video stabilization performance assessment
abstract
The goal of video stabilization is to remove the unwanted camera motion and obtain stable versions. Theoretically, a good stabilization algorithm should remove the unwanted motion without the loss of image qualities. However, due to the lack of ground-truth video frames, the accurate performance evaluation of different algorithms is hard. Most existing evaluation techniques usually synthesize stable videos from shaking ones, but they are not effective enough. Different from previous methods, in this paper we propose a novel method which synthesize shaking videos from stable frames. Based on the synthetic shaking videos, we perform preliminary video stabilization performance assessment on three stabilization algorithms. Our shaking video synthesis method can not only give a benchmark for full-reference video stabilization performance assessment, but also provide a basis for exploring the theoretical bound of video stabilization which may help to improve existing stabilization algorithms.
Li Song 0001, Gengjian Xue
VCIP2
2013 H.264/Advanced Video Control Perceptual Optimization Coding Based on JND-Directed Coefficient Suppression
abstract
The field of video coding has been exploring the compact representation of video data, where perceptual redundancies in addition to signal redundancies are removed for higher compression. Many research efforts have been dedicated to modeling the human visual system's characteristics. The resulting models have been integrated into video coding frameworks in different ways. Among them, coding enhancements with the just noticeable distortion (JND) model have drawn much attention in recent years due to its significant gains. A common application of the JND model is the adjustment of quantization by a multiplying factor corresponding to the JND threshold. In this paper, we propose an alternative perceptual video coding method to improve upon the current H.264/advanced video control (AVC) framework based on an independent JND-directed suppression tool. This new tool is capable of finely tuning the quantization using a JND-normalized error model. To make full use of this new rate distortion adjustment component the Lagrange multiplier for rate distortion optimization is derived in terms of the equivalent distortion. Because the H.264/AVC integer discrete cosine transform (DCT) is different from classic DCT, on which state-of-the-art JND models are computed, we analytically derive a JND mapping formula between the integer DCT domain and the classic DCT domain which permits us to reuse the JND models in a more natural way. In addition, the JND threshold can be refined by adopting a saliency algorithm in the coding framework and we reduce the complexity of the JND computation by reusing the motion estimation of the encoder. Another benefit of the proposed scheme is that it remains fully compliant with the existing H.264/AVC standard. Subjective experimental results show that significant bit saving can be obtained using our method while maintaining a similar visual quality to the traditional H.264/AVC coded video.
Zhengyi Luo 0001, Li Song 0001, Shibao Zheng, Nam Ling
IEEE Trans. Circuits Syst. Video Technol.2
2013 Foreground Estimation Based on Linear Regression Model With Fused Sparsity on Outliers
abstract
Foreground detection is an important task in computer vision applications. In this paper, we present an efficient foreground detection method based on a robust linear regression model. First, a novel framework is proposed where foreground detection has been cast as an outlier signal estimation problem in a linear regression model. We regularize this problem by imposing a so-called fused sparsity constraint, which encourages both sparsity and smoothness of vector coefficients, on the outlier signal. Second, we convert this outlier signal estimation problem into an equivalent Fused Lasso problem, and then use existing solutions to obtain an optimized solution. Third, a new foreground detection method is presented to apply this new model to the 2-D image domain by merging the results from different vectorizations. Experiments on a set of challenging sequences show that the proposed method is not only superior to many state-of-the-art techniques, but also robust to noise.
Gengjian Xue, Li Song 0001, Jun Sun 0005
IEEE Trans. Circuits Syst. Video Technol.2
2013 Reorder user's tweets
abstract
Twitter displays the tweets a user received in a reversed chronological order, which is not always the best choice. As Twitter is full of messages of very different qualities, many informative or relevant tweets might be flooded or displayed at the bottom while some nonsense buzzes might be ranked higher. In this work, we present a supervised learning method for personalized tweets reordering based on user interests. User activities on Twitter, in terms of tweeting, retweeting, and replying, are leveraged to obtain the training data for reordering models. Through exploring a rich set of social and personalized features, we model the relevance of tweets by minimizing the pairwise loss of relevant and irrelevant tweets. The tweets are then reordered according to the predicted relevance scores. Experimental results with real twitter user activities demonstrated the effectiveness of our method. The new method achieved above 30% accuracy gain compared with the default ordering in twitter based on time.
Keyi Shen, Jianmin Wu, Ya Zhang 0002, Yiping Han, Xiaokang Yang 0001, Li Song 0001, Xiao Gu 0001
ACM Trans. Intell. Syst. Technol.6
2013 Raptor Codes Based Unequal Protection for Compressed Video According to Packet Priority
abstract
Raptor codes are state-of-the-art forward error correction (FEC) solutions for multimedia transmission, which have been applied to unequal error protection (UEP) of multi-layered media such as scalable video coding. In this paper, we address the problem of UEP for single-layered video over packet erasure channels. By exploiting the different priorities of video packets inside a group of pictures (GOP) and making full use of the good characteristics of standardized Raptor codes at large block length, we propose an optimized UEP framework for single-layered video and develop an efficient algorithm to solve it. Simulation results show that significant gains can be obtained by our method in case of packet losses.
Zhengyi Luo 0001, Li Song 0001, Shibao Zheng, Nam Ling
IEEE Trans. Multim.2
2012 Feature Analysis of Spammers in Social Networks with Active Honeypots: A Case Study of Chinese Microblogging Networks
abstract
In this poster we report our study on the microblog spammers with samples attracted by 50 honeyspots from two popular Chinese microblogging networks: Sina Weibo (weibo.com), and Ten cent Weibo (t.QQ.com) in seven months. We studied their features such as social information, activity, account age and spamming strategy. Several distinguishing characteristics of spammers on these two social network communities are observed, which can be helpful to the further study on automatic detection of microblog spammers. To our best knowledge our work is the first of its kind on the analysis of features of Chinese micloblog spammers.
Yi Zhou 0003, Kai Chen 0006, Li Song 0001, Xiaokang Yang 0001, Jianhua He 0001
ASONAM3
2012 New bounds on image denoising: Viewpoint of sparse representation and non-local averaging
abstract
Image denoising plays a fundamental role in many image processing applications. Utilizing sparse representation and nonlocal averaging together is such a successful framework that leads to considerable progress in denoising. Almost all the newly proposed denoising algorithms are built base on it, different in detailed implementation, and the denoising performance seems converging. What is the denoising bound of this framework turns into a key question. In this paper, we assume all the possible algorithms under the framework can be approximated by a fixed two steps denoising process with different parameters. Step one cluster geometric similar image patches into groups so that patches within each group could be sparse represented under the basis of the group. Step two use the atoms of the group basis and radiometric similar patches of each patch for non-local averaging. The parameters of the process are the cluster number, the atoms and the number of radiometric similar patches for estimating each patch. Finally, the bound is derived as the minimum denoising error of all the possible parameters. Comparing with previous bounds, the new one is image specific and more practical. Experiment results show that there still exists room to improve the denoising performance for natural images.
Jianzhou Feng 0001, Li Song 0001, Xiaoming Huo, Xiaokang Yang 0001, Wenjun Zhang 0001
VCIP2
2012 Optimized nested protection for video Region of Interest with Raptor codes
abstract
Due to the best effort feature of many existing transmission channels, video streams often suffer from inevitable transmission errors. In this paper, we propose a scheme of robust video transmission based on the state-of-the-art Raptor codes, whose applications are in full swing now. And considering Region of Interest (ROI) often draws much attention in images, the scheme adopts a nested protection framework to show partialities to ROI areas for better protection. Different from many existing Raptor codes based UEP methods, our scheme is developed based on the easy-to-use standardized Raptor codes. Experimental results show that significant robustness can be obtained for the video streams, especially for the ROI areas.
Zhengyi Luo 0001, Li Song 0001, Shibao Zheng, Nam Ling
VCIP2
2012 Background subtraction based on phase feature and distance transform
Gengjian Xue, Jun Sun 0005, Li Song 0001
Pattern Recognit. Lett.3
2011 Building Artificial Identities in Social Network Using Semantic Information
abstract
As the popularity of social networking sites increase, so does their attractiveness for criminals. In this work, we show how an adversary can build artificial identities using semantic information in social network. Our method make the identities look more like real people, therefore can be used to support many kinds of attacks, such as ASE, profile cloning. A prototype of this method is implemented, includes following stages: Firstly, categories of virtual identity are predefined, and each category has multiple properties, such as geographical region, hobby, education, age, interested topic/keywords, etc. Secondly, based on category information, each identity will foster its own "life" semantically, such as edit profile and update status, find hot related news/topic from Google then post to wall, find related groups/networks then request to add in, and find/like/create/comment pages/posts, etc. Thirdly, artificial identity will evolve to multiple stages according to its status (for example, number of friends of real people), single identity with different evolutionary stages is linked together to a group that will help to ensure the number of attack edges.
Kai Chen 0006, Yi Zhou 0003, Li Song 0001, Xiaokang Yang 0001
ASONAM3
2011 Learning sparse dictionaries with a popularity-based model
abstract
Sparse signal representation based on overcomplete dictionaries has recently been extensively investigated, rendering the state-of-the-art results in signal, image and video processing. We propose a novel dictionary learning algorithm-the PK-SVD algorithm-which assumes prior probabilities on the dictionary atoms and learns a sparse dictionary under a popularity-based model. The prior distribution brings the flexibility that is desirable in applications. We examine our algorithm in both synthetic tests and image denoising experiments.
Jianzhou Feng 0001, Li Song 0001, Xiaoming Huo, Xiaokang Yang 0001, Wenjun Zhang 0001
ICASSP2
2011 Learning dictionary via subspace segmentation for sparse representation
abstract
Sparse signal representation based on redundant dictionaries contributed to much progress in image processing in the past decades. But the common overcomplete dictionary model is not well structured and there is still no guideline for selecting the proper dictionary size. In this paper, we propose a new algorithm for dictionary learning based on subspace segmentation. Our algorithm divides the training data into sub-spaces and constructs the dictionary by extracting the shared basis from multiple subspaces. The learned dictionary is well structured and its size is adaptive to the training data. We analyze this algorithm and demonstrate its ability on some initial supportive experiments using real image data.
Jianzhou Feng 0001, Li Song 0001, Xiaokang Yang 0001, Wenjun Zhang 0001
ICIP2
2011 Foreground estimation based on robust linear regression model
abstract
Background subtraction is a basic task for many computer vision applications, yet in dynamic scenes it is still a challenging problem. In this paper, we propose a new method to deal with this difficulty. Our approach is based on robust linear regression model and casts background subtraction as a outlier signal estimation problem. In our linear regression model, we explicitly model the error term as a combination of two components: foreground outlier and background noise. The foreground outlier is sparse and can be arbitrarily large in most cases, while the background noise is relatively small and dispersed. In order to reliably estimate the coefficients under the constraint of sparse foreground outlier, we propose a new objective function. Then we transform the function to fit our problem by only estimating the foreground outlier and give the solution method. Experimental results demonstrate the effectiveness of our method.
Gengjian Xue, Li Song 0001, Jun Sun 0005
ICIP2
2011 Hybrid center-symmetric local pattern for dynamic background subtraction
abstract
Effective foreground detection in dynamic scenes is a challenging task in computer vision applications. In this paper, we propose a novel background modeling method to tackle this problem. First, we propose a second-order center-symmetric local derivative pattern (CS-LDP) which extracts more detail information compared with the first-order center-symmetric local binary pattern (CS-LBP). Then by concatenating the CS-LBP and CS-LDP histograms, a new hybrid histogram feature is presented. The length of this histogram is much shorter than the local binary pattern (LBP) histogram. Based on this hybrid feature, a novel background modeling method is proposed where the pixel process is modeled with a group of adaptive hybrid histograms. The major advantage of our method is its low complexity. Experiments on three challenging sequences demonstrate that the proposed method is effective and fast, producing comparable results to state-of-art algorithm while reducing the computation time greatly.
Gengjian Xue, Li Song 0001, Jun Sun 0005
ICME2
2011 Prioritized Flow Optimization With Multi-Path and Network Coding Based Routing for Scalable Multirate Multicasting
abstract
In this paper, we study performance optimization for scalable video coding and multicast over networks. Multi-path video streaming, network coding based routing, and network flow control are jointly optimized to maximize a network utility function defined over heterogeneous receivers. Content priority of video coding layers is considered during the flow routing to determine the optimal multicast paths and associated data rates for each layer. Our optimization scheme attempts to find content distribution meshes with minimum path costs for each video coding layer while satisfying the inter-layer dependency during scalable video coding. Based on primal decomposition and primal-dual analysis, we develop a decentralized algorithm with two optimization levels to solve the performance optimization problem. We also prove the stability and convergence of the proposed iterative algorithm using Lyapunov theories. Extensive experimental results demonstrate that the proposed algorithm not only achieves the max-flow throughput using network coding, but also provides better video quality with balanced layered access for heterogeneous receivers.
Junni Zou, Hongkai Xiong, Li Song 0001, Zhihai He, Tsuhan Chen
IEEE Trans. Circuits Syst. Video Technol.4
2010 Multi-illumination Face Recognition from a Single Training Image per Person with Sparse Representation
Li Song 0001, Cheng Zhi
ACCV (2)2
2010 Background subtraction based on phase and distance transform under sudden illumination change
abstract
Effective foreground detection under sudden illumination change is an active research topic. However, most existing background subtraction approaches, which are intensity based, fail to handle this situation. In this paper, we propose a novel background modeling method that overcomes this limitation by relying on statistical models which use pixel phase instead of intensities. We first extract the phase feature of the pixel using Gabor filters. Then, a phase based background subtraction approach is proposed. In this approach, each phase feature is modeled independently by a mixture of Gaussian models and updated with a novel scheme. Since foreground pixels are scattered in the preliminary detection result, distance transform is implemented on the binary image which transforms the image into a distance map. We segment the distance image with a threshold and get the final result. Experiments on two challenging sequences demonstrate the effectiveness and robustness of our method.
Gengjian Xue, Jun Sun 0005, Li Song 0001
ICIP3
2010 Dynamic background subtraction based on spatial extended center-symmetric local binary pattern
abstract
Moving objects detection in dynamic scenes is a challenging task in many computer vision applications. Traditional background modeling methods do not work well in these situations since they assume a nearly static background. In this paper, a novel operator named spatial extended center-symmetric local binary pattern (SCS-LBP) for background modeling is proposed. It extracts spatial and temporal information simultaneously while has low complexity compared to the local binary pattern (LBP) operator. Then combining this operator with an improved temporal distribution estimation scheme, we propose a new background subtraction method. In our method, each pixel is modeled by a group of adaptive SCS-LBP histograms, which provides us with many advantages compared to traditional ones. Experimental results demonstrate the effectiveness and robustness of our method.
Gengjian Xue, Jun Sun 0005, Li Song 0001
ICME3
2010 Image denoising using local tangent space alignment
abstract
We propose a novel image denoising approach, which is based on exploring an underlying (nonlinear) lowdimensional manifold. Using local tangent space alignment (LTSA), we 'learn' such a manifold, which approximates the image content effectively. The denoising is performed by minimizing a newly defined objective function, which is a sum of two terms: (a) the difference between the noisy image and the denoised image, (b) the distance from the image patch to the manifold. We extend the LTSA method from manifold learning to denoising. We introduce the local dimension concept that leads to adaptivity to different kind of image patches, e.g. flat patches having lower dimension. We also plug in a basic denoising stage to estimate the local coordinate more accurately. It is found that the proposed method is competitive: its performance surpasses the K-SVD denoising method.
Jianzhou Feng 0001, Li Song 0001, Xiaoming Huo, Xiaokang Yang 0001, Wenjun Zhang 0001
VCIP2
2009 Generic video coding with abstraction and detail completion
abstract
This paper presents a generic video coding framework with the texture abstraction and completion, inspired by a strong grouping bias of local elements in Gestalt psychology. Abstracting imagery by grouping perceptual salience from anisotropic diffusion, it decomposes video images into two layers composing of semantic components and residual detail. The similarity between textures of abstraction layer is motivated to infer the restoration of missing detail, under the spatio-temporal variation regularity. Through a motion and spatial context of moton, hence, a group of pictures (GOP) is divided into key frames and abstracted frames to form the final compressed data. An abstraction refinement is tuned to improve matching of detail restoration based on bilateral filtering. The proposed approach is more generic without incurring any specific side information, and achieves up to 20% bit saving versus standard H.264 at similar visual quality levels.
Hongkai Xiong, Li Song 0001, Yuan F. Zheng
ICASSP3
2009 Graph Matching Based Side Information Generation for Distributed Multi-View Video Coding
abstract
In this paper, we adopt constrained relaxation for distributed multi-view video coding (DMVC). The novel framework integrates the graph-based segmentation and matching to generate inter-view correlated side information without knowing the camera parameters. Moreover, graph-based representations of multi-view images are incorporated to form more distinctive feature constraints. The sparse data as a good hypothesis space aim for a best matching optimization of inter-view side information with compact syndromes, from inferred relaxed coset. The plausible filling-in from a priori feature constraints between neighboring views could reinforce a promising compensation to inter-view side information generation for joint multi-view decoding. In order to find distinctive feature matching with a more stable approximation, PCA-SIFT and TPS (thin plate spline) are adopted to reduce the dimension of SIFT descriptors and construct a more accurate inter-view motion model. The experimental results validate the high estimation precision and the rate-distortion improvements.
Hui Lu 0001, Hongkai Xiong, Li Song 0001, Zhihai He, Tsuhan Chen
ICC3
2009 Prioritized Flow Optimization with Generalized Routing for Scalable Multirate Multicasting
abstract
This paper addresses the performance optimization for scalable video coding and multicast over networks. Multi-path video streaming, network coding based routing, and network flow control are jointly optimized to maximize a network utility function defined over heterogeneous receivers. Importantly, contextual priors of scalable video layers are imposed on the flow routing optimization problem, seeking to guarantee the transmission cost for each layer in an incremental order and find jointly optimal multicast paths and associated rates. Through a primal decomposition and the primal-dual approach, a decentralized algorithm with two-level optimization update is developed to solve the target convex optimization problem. Numerical and simulation results validate the convergence and network performance of the proposed algorithm.
Junni Zou, Hongkai Xiong, Li Song 0001, Zhihai He, Tsuhan Chen
ICC3
2009 Sub clustering K-SVD: Size variable dictionary learning for sparse representations
abstract
Sparse signal representation from overcomplete dictionaries have been extensively investigated in recent research, leading to state-of-the-art results in signal, image and video restoration. One of the most important issues is involved in selecting the proper size of dictionary. However, the related guidelines are still not established. In this paper, we tackle this problem by proposing a so-called sub clustering K-SVD algorithm. This approach incorporates the subtractive clustering method into K-SVD to retain the most important atom candidates. At the same time, the redundant atoms are removed to produce a well-trained dictionary. As for a given dataset and approximation error bound, the proposed approach can deduce the optimized size of dictionary, which is greatly compressed as compared with the one needed in the K-SVD algorithm.
Jianzhou Feng 0001, Li Song 0001, Xiaokang Yang 0001, Wenjun Zhang 0001
ICIP2
2009 Spatial non-stationary correlation noise modeling for Wyner-Ziv error resilience video coding
abstract
Most of the Wyner-Ziv (WZ) video coding schemes in literature model the correlation noise (CN) between original frame and side information (SI) by a given distribution whose parameters are estimated in an offline process. In this paper, an online CN modeling algorithm is proposed towards a more practical WZ-based error resilient video coding (WZ-ERVC). In ERVC scenario, the side-information is typically generated from the error concealed picture instead of bi-directional motion prediction. The proposed online CN modeling algorithm achieves the so-called classification gain by exploiting the spatially non-stationary characteristics of the motion field and texture. The CN between the source and error concealed SI is modeled by a Laplacian mixture model, where each mixture component represents the statistical distribution of prediction residuals and the mixing coefficients portray the motion vectors estimation error. Experimental results demonstrate significant performance gains both in rate and distortion versus the conventional Laplacian model.
Hongkai Xiong, Li Song 0001, Songyu Yu
ICIP3
2009 Robust Video Region-of-Interest Coding Based on Leaky Prediction
abstract
A video region-of-interest (ROI) scalable coding scheme can ensure the priority of ROI. Error protection schemes can be used to guarantee the correct receipt of the ROI stream when transporting ROI scalable video over an error-prone network. However, we find that the correct receipt of ROI bitstreams cannot ensure the correct decoding of ROI due to the unique issue of the cross error propagation between ROI and background in ROI scalable coding. In this letter, we propose an ROI scalable coding framework based on leaky prediction (LP) for robustly transporting video over an error-prone network. Although several LP approaches have been proposed to improve layered coding, they cannot be applied to ROI scalable coding straightforwardly due to the cross error propagation issue. We deploy a leaky factor to weigh the two predictions: one from the constrained motion estimation (ME) within the ROI layer of the reference frame, and the other from the unrestricted ME in the overall reference frame. Simulation results show that the proposed scheme enhances the robustness of ROI scalability while maintaining coding efficiency.
Qian Chen 0024, Xiaokang Yang 0001, Li Song 0001, Wenjun Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2008 On Non-sequential Context Modeling with Application to Executable Data Compression
abstract
The sequential context modeling framework is generalized to a non-sequential one by context relaxation from consecutive suffix of the subsequences of symbols to the permutation of the preceding symbols as result of considering complex context structures in such sources as video and program binaries. Context weighting tree is also extended to a series of context trees which are built according to the "model tree", in which the descendent relationship in the formation of non-sequential context sets is described. Model redundancy and maximum a posteriori model in the framework are discussed and compared. A decision method based on the greedy algorithm is proposed to customize sets of models fitting the concrete sources. Brief description of application to executable data files incorporating with the semantics and syntax constraints are given and experiment are made accordingly as a validation.
Wenrui Dai, Hongkai Xiong, Li Song 0001
DCC3
2008 Local Quaternionic Gabor Binary Patterns for color face recognition
abstract
In this paper, a novel color face recognition method is proposed based on local binary patterns (LBP) of quaternionic Gabor features (QGF). By introducing quaternion Gabor analysis into image representation, we make full use of the interrelationship among different color channels to enhance the performance of the face recognition system. Moreover, the QGF are used to encode the positions and attributes of the face elements. Non-parametric transformation is then imposed on these QGF using LBP method to obtain the robustness against variations of pose, illumination and facial expressions. Compared with the monochromatic face recognition systems, which nowadays dominate the marketplace and research field, this approach materializes the strong potential use of color face recognition system by establishing invariant quaternion wavelet features of color images. The experimental results on the open face database testify the validity of the proposed method under severe noise corruption and distinct variations of scale, illumination and facial expressions.
Wei Lu 0021, Yi Xu 0001, Xiaokang Yang 0001, Li Song 0001
ICASSP4
2008 Color image watermarking using local quaternion Fourier spectral analysis
abstract
We propose a watermarking scheme for color images based on local quaternion Fourier spectral analysis (LQFSA).The merits of the proposed scheme include: 1) Quaternion Fourier transform is defined in a 4D vector space and thus provides a larger embedding scope for watermark than conventional monochannel transformation techniques. 2) We improve the imperceptibility of watermark with regard to human color vision properties through LQFSA. 3) We introduce invariant feature transform (IFT) and geometric correction scheme so as to enhance the robustness to extensive attacks, which is another essential factor to evaluate a watermarking scheme. 4) We adopt the nearest-neighborhood search to ensure the correctness of watermark extraction. Extensive experiments on the Stirmark platform validate the aforementioned merits.
Yi Xu 0001, Li Song 0001, Xiaokang Yang 0001, Hans Burkhardt
ICME3
2008 2-D dual multiresolution decomposition through NUDFB and its application
abstract
This paper aims to attain sparser representation of a 2-D signal by introducing orientation resolution as a second multiresolution besides multiscale, which is formulated to achieve a dual multiresolution decomposition framework by nonuniform directional frequency decompositions (NUDFB) under arbitrary scales. In this scheme, NUDFB is fulfilled by changing the topology structure of a non-symmetric binary tree (NSBT). Through this nonuniform division, we can get arbitrary orientation resolution r at a direction of c2-runder a target scale. Every two-channel filter bank on each node of this NSBT is designed to be a paraunitary perfect reconstruction filter bank, so NUDFB is an orthogonal filter bank. This dual multiresolution decomposition will definitely have bright prospect in its application, such as texture analysis, image processing or video coding. A potential application is presented by applying NUDFB in wavelet domain.
Nannan Ma, Hongkai Xiong, Li Song 0001
MMSP3
2008 Contourlet-based image adaptive watermarking
Songyu Yu, Xiaokang Yang 0001, Li Song 0001, Chen Wang 0053
Signal Process. Image Commun.4
2007 2D Quaternion Fourier Transform: The Spectrum Properties and its Application in Color Image Registration
abstract
We first investigate 2D quatemion Fourier transform (QFT) spectrum relationship between an image and its geometrically transformed counterpart from the aspects of gray images and color images respectively, and then propose a 2D QFT-based color image registration approach, which is able to handle large translation, rotation and scaling. Fourier transform (FT) can be utilized in gray image registration but can not process color images naturally. As the extended version of FT in multidimensional signal processing, QFT is capable of processing three color components of color images together as pure quatemions and is able to deal with color image registration.
Zheng Lu 0003, Yi Xu 0001, Xiaokang Yang 0001, Li Song 0001, Leonardo Traversoni
ICME4
2007 Bit Allocation for Fine-Granular SNR Scalability Coding with Hierarchical B Pictures
abstract
Hierarchical B pictures are devised to achieve temporal scalability in the scalable extension of H.264/AVC (SVC) which is under standardization. The fine-granular SNR scalability (FGS) can be provided by progressive refinement (PR) slices in SVC. In this paper, we firstly investigate error propagation in the case of discarding PR slices and obtain a rate difference distortion optimization criterion to improve coding efficiency of base layer. Then we consider the full rate case and propose a rate distortion slope criterion to enhance FGS coding efficiency at high rate. Finally the criterion to boost coding efficiency in the whole range of FGS rate is derived by combing the criterions derived previously. The proposed method is compared to the approach in SVC test model and up to 0.3dB coding gains are achieved.
Jun Xu 0040, Li Song 0001, Shibao Zheng, Xiaokang Yang 0001, Rong Xie 0004
ICME2
2007 Cooperative Stereo Matching using Quaternion Wavlets and Top-Down Segmentation
abstract
We explore the principles of quaternion wavelet construction for achieving multiscale analysis of geometric image features. Then the quaternion wavelets are applied to propose a cooperative stereo matching algorithm using top-down segmentation-based disparity propagation. Without bidirectional matching to remove ambiguous outliers, uniqueness constraint is enforced on cost function by inhibiting the matches along similar sightlines. To produce smooth disparity maps with the discontinuities well-preserved, cost aggregation is performed in segmentation-based local support and high confidence matches serve as heavyweight seeds for disparity propagation in the supports. Compared with the current matching methods based on quatemion wavelets, the main merit of the proposed algorithm is that the matching results are encouraging in extensive comparison data, ranging from calibrated images to uncalibrated images, indoor images to aerial images.
Yi Xu 0001, Xiaokang Yang 0001, Peifeng Zhang, Li Song 0001, Leonardo Traversoni
ICME4
2006 A Context-Based Error Detection Strategy into H.264/AVC CABAC
abstract
Various error control schemes have been addressed in wireless video stream transmission. By combining an adaptive binary arithmetic coding technique with context modeling, CABAC as a normative part of H.264/AVC has achieved a high degree of adaptation and redundancy reduction. However, error propagation still remains a problem because of the property of arithmetic coding. The presented scheme compares the various error detection methods, and proposes an efficient error detection technique based on CABAC semantics, which is achieved by inserting detective markers denoting by syntax elements. The misdetection probability versus stream size expansion can be easily handled. In addition, placements of markers can vary with regard to specific video content, thus efficiency within this scheme is enhanced. Comparison with other detection scheme is also presented
Hongkai Xiong, Li Song 0001, Songyu Yu
ICME3
2006 A New Deblocking Algorithm Based on Adjusted Contourlet Transform
abstract
A new postprocessing method based on adjusted contourlet transform is introduced in this paper for suppressing blocking artifacts (BA) in block-based discrete cosine transform (BDCT) compressed images. To our best knowledge, this is the first time contourlet is applied to this field. By exploiting scale space edge detector (ss-edge detector), our algorithm can extract and protect blocking map (BM) and edge map (EM) in the compressed image respectively in the same time. By transforming the compressed image into adjusted contourlet domain, the adaptive thresholds are obtained according to BM. According to the adaptive thresholds, the contourlet coefficients in different subbands are filtered. Experimental results show that our deblocking algorithm achieves better performance than the other iterative and noniterative methods reported in the literature
Songyu Yu, Chen Wang 0053, Li Song 0001, Hongkai Xiong
ICME4
2005 Adaptive predict based on fading compensation for lifting-based motion compensated temporal filtering
abstract
A lifting implementation of the discrete wavelet transform applied along motion trajectories has recently gained a lot of attention in the video community as strong candidates in incoming scalable video coders. We generalize the coding scheme for classical lifting-based motion compensation temporal filtering and permit the codec to choose adaptively between the original reference frames and new fading-compensated reference frames to predict residuals while maintaining the invertibility of the inter-frame transform. Experimental results show that the proposed algorithm not only significantly improves subjective visual quality of the temporal low-pass frames, but also has 0.15-0.3 dB gain in PSNR performance compared with the normal (5, 3) lifting schemes.
Li Song 0001, Hongkai Xiong, Jizheng Xu, Feng Wu 0001, Hui Su
ICASSP (2)1