EDBT 2026 Demo / reviewers in the wild / expert
Bin Li 0012
dblp:89/6764-12
· DBLP profile ↗
70ranked-venue papers
12as first author
25since 2021 · last 2026
0000-0002-5635-2916ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 58 · 8 first-author · 19 since 2021Artificial intelligence and machine learning · 15 · 14 since 2021Computer networks · 6 · 5 since 2021Systems, architecture and hardware · 5 · 4 first-authorDatabases, data management, data science and information retrieval · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent TrainingabstractGUI agents that interact with graphical interfaces on behalf of users are a promising direction for practical AI assistants, yet training them is hindered by scarce suitable environments. We present InfiniteWeb, a system that automatically generates functional web environments at scale for GUI agent training. While LLMs perform well on generating a single webpage, building a realistic and functional website with many interconnected pages faces challenges. We address these challenges through unified specification, task-centric test-driven development, and combining website seed variation with reference design images. Our system also generates verifiable task evaluators enabling dense reward signals for reinforcement learning. Experiments show that our system surpasses commercial coding agents at realistic website construction, and GUI agents trained on our generated environments achieve significant performance improvements on OSWorld and Online-Mind2Web, demonstrating the effectiveness of the proposed system. Zezhou Wang, Zongyu Guo, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
ACL (1) | 6 |
| 2026 | A Channel Adaptive Encoding and Decoding Method for Unmanned Aerial Vehicle Image TransmissionabstractUnmanned Aerial Vehicles (UAVs) are an indispensable core component of low altitude economic networks. It is very critical for achieving efficient UAV image transmission of air-to-ground communication. Due to the changes in flight area and unstable channel conditions, the signal-to-noise and transmission rate change rapidly. To adapt to these changes, we propose a channel adaptive encoding and decoding method for UAV image transmission. The proposed method includes a lightweight feature extraction module, a channel-wise feature enhance module, a transmission rate adaptive module, and the corresponding decoding module. The lightweight feature extraction module can quickly extract local detailed features and long-range spatial dependencies via residual block and mobile mamba. The channel-wise feature enhance module can enhance channel wise useful features via the involution operation and the SNR adjustment block according to channel state SNRs. The transmission rate adaptive module can further adaptively adjust the size of transmission features according to the transmission rate via the rate adjustment block and the rate mask block. The extensive experimental results on the SIRI-WHU, WHU-RS19, AID, and UCMerced Land Use datasets demonstrate that our method obtains higher PSNR, MS-SSMI, LPIPS and ACC than state-of-the-art methods. Shuhang Zhang, Qi Qiu, Bin Li 0012, Chiya Zhang, Guangming Shi |
IEEE Internet Things J. | 4 |
| 2025 | Towards Practical Real-Time Neural Video CompressionabstractWe introduce a practical real-time neural video codec (NVC) designed to deliver high compression ratio, low latency and broad versatility. In practice, the coding speed of NVCs depends on 1) computational costs, and 2) non-computational operational costs, such as memory I/O and the number of function calls. While most efficient NVCs prioritize reducing computational cost, we identify operational cost as the primary bottleneck to achieving higher coding speed. Leveraging this insight, we introduce a set of efficiency-driven design improvements focused on minimizing operational costs. Specifically, we employ implicit temporal modeling to eliminate complex explicit motion modules, and use single low-resolution latent representations rather than progressive downsampling. These innovations significantly accelerate NVC without sacrificing compression quality. Additionally, we implement model integerization for consistent cross-device coding and a module-bank-based rate control scheme to improve practical adaptability. Experiments show our proposed DCVC-RT achieves an impressive average encoding/decoding speed at 125.2/112.8 fps (frames per second) for 1080p video, while saving an average of 21% in bitrate compared to H.266/VTM. The code is available at https://github.com/microsoft/DCVC. Zhaoyang Jia, Bin Li 0012, Jiahao Li 0001, Wenxuan Xie, Houqiang Li, Yan Lu 0001 |
CVPR | 2 |
| 2025 | PICD: Versatile Perceptual Image Compression with Diffusion RenderingabstractRecently, perceptual image compression has achieved significant advancements, delivering high visual quality at low bitrates for natural images. However, for screen content, existing methods often produce noticeable artifacts when compressing text. To tackle this challenge, we propose versatile perceptual screen image compression with diffusion rendering (PICD), a codec that works well for both screen and natural images. More specifically, we propose a compression framework that encodes the text and image separately, and renders them into one image using diffusion model. For this diffusion rendering, we integrate conditional information into diffusion models at three distinct levels: 1). Domain level: We fine-tune the base diffusion model using text content prompts with screen content. 2). Adaptor level: We develop an efficient adaptor to control the diffusion model using compressed image and text as input. 3). Instance level: We apply instance-wise guidance to further enhance the decoding process. Empirically, our PICD surpasses existing perceptual codecs in terms of both text accuracy and perceptual quality. Additionally, without text conditions, our approach serves effectively as a perceptual codec for natural images. Tongda Xu, Jiahao Li 0001, Bin Li 0012, Yan Wang 0105, Ya-Qin Zhang, Yan Lu 0001 |
CVPR | 3 |
| 2025 | DLF: Extreme Image Compression with Dual-Generative Latent Fusion
Naifu Xue, Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
ICCV | 4 |
| 2025 | One-Step Diffusion-Based Image Compression with Semantic DistillationabstractWhile recent diffusion-based generative image codecs have shown impressive performance, their iterative sampling process introduces unpleasant latency. In this work, we revisit the design of a diffusion-based codec and argue that multi-step sampling is not necessary for generative compression. Based on this insight, we propose OneDC, a One-step Diffusion-based generative image Codec—that integrates a latent compression module with a one-step diffusion generator. Recognizing the critical role of semantic guidance in one-step diffusion, we propose using the hyperprior as a semantic signal, overcoming the limitations of text prompts in representing complex visual content. To further enhance the semantic capability of the hyperprior, we introduce a semantic distillation mechanism that transfers knowledge from a pretrained generative tokenizer to the hyperprior codec. Additionally, we adopt a hybrid pixel- and latent-domain optimization to jointly enhance both reconstruction fidelity and perceptual realism. Extensive experiments demonstrate that OneDC achieves SOTA perceptual quality even with one-step generation, offering over 39% bitrate reduction and 20× faster decoding compared to prior multi-step diffusion-based codecs. Project: https://onedc-codec.github.io/ Naifu Xue, Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
NeurIPS | 4 |
| 2025 | Deep Video Discovery: Agentic Search with Tool Use for Long-form Video UnderstandingabstractLong-form video understanding presents significant challenges due to extensive temporal-spatial complexity and the difficulty of question answering under such extended contexts.
While Large Language Models (LLMs) have demonstrated considerable advancements in video analysis capabilities and long context handling, they continue to exhibit limitations when processing information-dense hour-long videos.
To overcome such limitations, we propose the $\textbf{D}eep \ \textbf{V}ideo \ \textbf{D}iscovery \ (\textbf{DVD})$ agent to leverage an $\textit{agentic search}$ strategy over segmented video clips. Different from previous video agents manually designing a rigid workflow, our approach emphasizes the autonomous nature of agents.
By providing a set of search-centric tools on multi-granular video database,
our DVD agent leverages the advanced reasoning capability of LLM to plan on its current observation state, strategically selects tools to orchestrate adaptive workflow for different queries in light of the gathered information.
We perform comprehensive evaluation on multiple long video understanding benchmarks that demonstrates our advantage.
Our DVD agent achieves state-of-the-art performance on the challenging LVBench dataset, reaching an accuracy of $\textbf{74.2\%}$, which substantially surpasses all prior works, and further improves to $\textbf{76.0\%}$ with transcripts. Zhaoyang Jia, Zongyu Guo, Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001 |
NeurIPS | 5 |
| 2025 | Generative Latent Coding for Ultra-Low Bitrate Image and Video CompressionabstractMost existing approaches for image and video compression perform transform coding in the pixel space to reduce redundancy. However, due to the misalignment between the pixel-space distortion and human perception, such schemes often face the difficulties in achieving both high-realism and high-fidelity at ultra-low bitrate. To solve this problem, we propose Generative Latent Coding (GLC) models for image and video compression, termed GLC-image and GLC-Video. The transform coding of GLC is conducted in the latent space of a generative vector-quantized variational auto-encoder (VQ-VAE). Compared to the pixel-space, such a latent space offers greater sparsity, richer semantics and better alignment with human perception, and show its advantages in achieving high-realism and high-fidelity compression. To further enhance performance, we improve the hyper prior by introducing a spatial categorical hyper module in GLC-image and a spatio-temporal categorical hyper module in GLC-video. Additionally, the code-prediction-based loss function is proposed to enhance the semantic consistency. Experiments demonstrate that our scheme shows high visual quality at ultra-low bitrate for both image and video compression. For image compression, GLC-image achieves an impressive bitrate of less than 0.04 bpp, achieving the same FID as previous SOTA model MS-ILLM while using 45% fewer bitrate on the CLIC 2020 test set. For video compression, GLC-video achieves 65.3% bitrate saving over PLVC in terms of DISTS. Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Neural Image Compression with Regional DecodingabstractAs advancements are made in technology such as AR/VR and high-resolution photography, there is a growing need for a function in image compression named regional decoding . This function lets an image be encoded as a whole, but allows for an arbitrary region to be decoded using only a small part of the bitstream. However, existing neural image compression methods lack support for this crucial functionality. In this article, we propose a novel approach called the slicing en/decoder , which addresses the need for regional decoding while maintaining performance on par with state-of-the-art methods. Our approach is based on the insight that, during the compression process, local information within pixels holds greater importance than global information. By leveraging this understanding, we divide the image into different bitstreams according to cross-boundary patterns. Consequently, for a selected region, our method can intelligently choose specific portions of the bitstreams to decode only that particular region of interest. Furthermore, we extend the application of our method to 360° image compression, allowing for efficient encoding and decoding of immersive visual content. Moreover, our proposed technique offers the capability to decode regions identically, which paves the way for future advancements in regional video decoding. Our experimental results demonstrate that our method maintains performance on par with state-of-the-art methods while providing the functionality of regional decoding . In conclusion, this article presents a significant step forward in image compression technology, offering enhanced flexibility and efficiency for emerging applications in digital media. Yili Jin 0001, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Generative Latent Coding for Ultra-Low Bitrate Image CompressionabstractMost existing image compression approaches perform transform coding in the pixel space to reduce its spatial re-dundancy. However, they encounter difficulties in achieving both high-realism and high-fidelity at low bitrate, as the pixel-space distortion may not align with human perception. To address this issue, we introduce a Generative Latent Coding (GLC) architecture, which performs transform coding in the latent space of a generative vector-quantized variational auto-encoder (VQ- VAE), instead of in the pixel space. The generative latent space is characterized by greater sparsity, richer semantic and better alignment with human perception, rendering it advantageous for achieving high-realism and high-fidelity compression. Additionally, we introduce a categorical hyper module to reduce the bit cost of hyper-information, and a code-prediction-based su-pervision to enhance the semantic consistency. Experiments demonstrate that our GLC maintains high visual quality with less than 0.04 bpp on natural images and less than 0.01 bpp on facial images. On the CLIC2020 test set, we achieve the same FID as MS-ILLM with 45% fewer bits. Furthermore, the powerful generative latent space enables various applications built on our GLC pipeline, such as image restoration and style transfer. Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001 |
CVPR | 3 |
| 2024 | Neural Video Compression with Feature ModulationabstractThe emerging conditional coding-based neural video codec (NVC) shows superiority over commonly-used resid-ual coding-based codec and the latest NVC already claims to outperform the best traditional codec. However, there still exist critical problems blocking the practicality of NVC. In this paper, we propose a powerful conditional coding- based NVC that solves two critical problems via feature modulation. The first is how to support a wide quality range in a single model. Previous NVC with this capability only supports about 3.8 dB PSNR range on average. To tackle this limitation, we modulate the latent feature of the cur-rent frame via the learnable quantization scaler. During the training, we specially design the uniform quantization pa-rameter sampling mechanism to improve the harmonization of encoding and quantization. This results in a better learning of the quantization scaler and helps our NVC support about 11.4 dB PSNR range. The second is how to make NVC still work under a long prediction chain. We expose that the previous SOTA NVC has an obvious quality degra-dation problem when using a large intra-period setting. To this end, we propose modulating the temporal feature with a periodically refreshing mechanism to boost the quality. Notably, under single intra-frame setting, our codec can achieve 29.7% bitrate saving over previous SOTA NVC with 16% MACs reduction. Our codec serves as a notable land-mark in the journey of NVC evolution. The codes are at https://github.com/microsoft/DCVC. Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
CVPR | 2 |
| 2024 | Long-Term Temporal Context Gathering for Neural Video Compression
Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001 |
ECCV (66) | 4 |
| 2024 | Uncertainty-Aware Deep Video Compression With EnsemblesabstractDeep learning-based video compression is a challenging task, and many previous state-of-the-art learning-based video codecs use optical flows to exploit the temporal correlation between successive frames and then compress the residual error. Although these two-stage models are end-to-end optimized, the epistemic uncertainty in the motion estimation and the aleatoric uncertainty from the quantization operation lead to errors in the intermediate representations and introduce artifacts in the reconstructed frames. This inherent flaw limits the potential for higher bit rate savings. To address this issue, we propose an uncertainty-aware video compression model that can effectively capture the predictive uncertainty with deep ensembles. Additionally, we introduce an ensemble-aware loss to encourage the diversity among ensemble members and investigate the benefits of incorporating adversarial training in the video compression task. Experimental results on 1080p sequences show that our model can effectively save bits by more than 20% compared to DVC Pro. Wufei Ma, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
IEEE Trans. Multim. | 3 |
| 2024 | A Universal Optimization Framework for Learning-based Image CodecabstractRecently, machine learning-based image compression has attracted increasing interests and is approaching the state-of-the-art compression ratio. But unlike traditional codec, it lacks a universal optimization method to seek efficient representation for different images. In this paper, we develop a plug-and-play optimization framework for seeking higher compression ratio, which can be flexibly applied to existing and potential future compression networks. To make the latent representation more efficient, we propose a novel latent optimization algorithm to adaptively remove the redundancy for each image. Additionally, inspired by the potential of side information for traditional codecs, we introduce side information into our framework, and integrate side information optimization with latent optimization to further enhance the compression ratio. In particular, with the joint side information and latent optimization, we can achieve fine rate control using only single model instead of training different models for different rate-distortion trade-offs, which significantly reduces the training and storage cost to support multiple bit rates. Experimental results demonstrate that our proposed framework can remarkably boost the machine learning-based compression ratio, achieving more than 10% additional bit rate saving on three different representative network structures. With the proposed optimization framework, we can achieve 7.6% bit rate saving against the latest traditional coding standard VVC on Kodak dataset, yielding the state-of-the-art compression ratio. Jing Zhao 0011, Bin Li 0012, Jiahao Li 0001, Ruiqin Xiong, Yan Lu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | Neural Video Compression with Diverse ContextsabstractFor any video codecs, the coding efficiency highly relies on whether the current signal to be encoded can find the relevant contexts from the previous reconstructed signals. Traditional codec has verified more contexts bring substantial coding gain, but in a time-consuming manner. However, for the emerging neural video codec (NVC), its contexts are still limited, leading to low compression ratio. To boost NVC, this paper proposes increasing the context diversity in both temporal and spatial dimensions. First, we guide the model to learn hierarchical quality patterns across frames, which enriches long-term and yet highquality temporal contexts. Furthermore, to tap the potential of optical flow-based coding framework, we introduce a group-based offset diversity where the cross-group interaction is proposed for better context mining. In addition, this paper also adopts a quadtree-based partition to increase spatial context diversity when encoding the latent representation in parallel. Experiments show that our codec obtains 23.5% bitrate saving over previous SOTA NVC. Better yet, our codec has surpassed the under-developing next generation traditional codec/ECM in both RGB and YUV420 colorspaces, in terms of PSNR. The codes are at https://github.com/microsoft/DCVC. Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
CVPR | 2 |
| 2023 | Motion Information Propagation for Neural Video CompressionabstractIn most existing neural video codecs, the information flow therein is uni-directional, where only motion coding provides motion vectors for frame coding. In this paper, we argue that, through information interactions, the synergy between motion coding and frame coding can be achieved. We effectively introduce bi-directional information interactions between motion coding and frame coding via our Motion Information Propagation. When generating the temporal contexts for frame coding, the high-dimension motion feature from the motion decoder serves as motion guidance to mitigate the alignment errors. Meanwhile, besides assisting frame coding at the current time step, the feature from context generation will be propagated as motion condition when coding the subsequent motion latent. Through the cycle of such interactions, feature propagation on motion coding is built, strengthening the capacity of exploiting long-range temporal correlation. In addition, we propose hybrid context generation to exploit the multiscale context features and provide better motion condition. Experiments show that our method can achieve 12.9% bit rate saving over the previous SOTA neural video codec. Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001 |
CVPR | 3 |
| 2023 | EVC: Towards Real-Time Neural Image Compression with Mask Decay
Guo-Hua Wang, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
ICLR | 3 |
| 2023 | Temporal Context Mining for Learned Video CompressionabstractApplying deep learning to video compression has attracted increasing attention in recent few years. In this work, we address end-to-end learned video compression with a special focus on better learning and utilizing temporal contexts. We propose to propagate not only the last reconstructed frame but also the feature before obtaining the reconstructed frame for temporal context mining. From the propagated feature, we learn multi-scale temporal contexts and re-fill the learned temporal contexts into the modules of our compression scheme, including the contextual encoder-decoder, the frame generator, and the temporal context encoder. We discard the parallelization-unfriendly auto-regressive entropy model to pursue a more practical encoding and decoding time. Experimental results show that our proposed scheme achieves a higher compression ratio than the existing learned video codecs. Our scheme also outperforms x264 and x265 (representing industrial software for H.264 and H.265, respectively) as well as the official reference software for H.264, H.265, and H.266 (JM, HM, and VTM, respectively). Specifically, when intra period is 32 and oriented to PSNR, our scheme outperforms H.265–HM by 14.4% bit rate saving; when oriented to MS-SSIM, our scheme outperforms H.266–VTM by 21.1% bit rate saving. Xihua Sheng, Jiahao Li 0001, Bin Li 0012, Li Li 0040, Dong Liu 0002, Yan Lu 0001 |
IEEE Trans. Multim. | 3 |
| 2022 | Neural Compression-Based Feature Learning for Video RestorationabstractHow to efficiently utilize the temporal features is crucial, yet challenging, for video restoration. The temporal features usually contain various noisy and uncorrelated information, and they may interfere with the restoration of the current frame. This paper proposes learning noiserobust feature representations to help video restoration. We are inspired by that the neural codec is a natural denoiser: In neural codec, the noisy and uncorrelated contents which are hard to predict but cost lots of bits are more inclined to be discarded for bitrate saving. Therefore, we design a neural compression module to filter the noise and keep the most useful information in features for video restoration. To achieve robustness to noise, our compression module adopts a spatial-channel-wise quantization mechanism to adaptively determine the quantization step size for each position in the latent. Experiments show that our method can significantly boost the performance on video denoising, where we obtain 0.13 dB improvement over BasicVSR++ with only 0.23x FLOPs. Meanwhile, our method also obtains SOTA results on video deraining and dehazing. Jiahao Li 0001, Bin Li 0012, Dong Liu 0002, Yan Lu 0001 |
CVPR | 3 |
| 2022 | Hybrid Spatial-Temporal Entropy Modelling for Neural Video CompressionabstractFor neural video codec, it is critical, yet challenging, to design an efficient entropy model which can accurately predict the probability distribution of the quantized latent representation. However, most existing video codecs directly use the ready-made entropy model from image codec to encode the residual or motion, and do not fully leverage the spatial-temporal characteristics in video. To this end, this paper proposes a powerful entropy model which efficiently captures both spatial and temporal dependencies. In particular, we introduce the latent prior which exploits the correlation among the latent representation to squeeze the temporal redundancy. Meanwhile, the dual spatial prior is proposed to reduce the spatial redundancy in a parallel-friendly manner. In addition, our entropy model is also versatile. Besides estimating the probability distribution, our entropy model also generates the quantization step at spatial-channel-wise. This content-adaptive quantization mechanism not only helps our codec achieve the smooth rate adjustment in single model but also improves the final rate-distortion performance by dynamic bit allocation. Experimental results show that, powered by the proposed entropy model, our neural codec can achieve 18.2% bitrate saving on UVG dataset when compared with H.266 (VTM) using the highest compression ratio configuration. It makes a new milestone in the development of neural video codec. The codes are at https://github.com/microsoft/DCVC. Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
ACM Multimedia | 2 |
| 2021 | Deep Contextual Video CompressionabstractMost of the existing neural video compression methods adopt the predictive coding framework, which first generates the predicted frame and then encodes its residue with the current frame. However, as for compression ratio, predictive coding is only a sub-optimal solution as it uses simple subtraction operation to remove the redundancy across frames. In this paper, we propose a deep contextual video compression framework to enable a paradigm shift from predictive coding to conditional coding. In particular, we try to answer the following questions: how to define, use, and learn condition under a deep video compression framework. To tap the potential of conditional coding, we propose using feature domain context as condition. This enables us to leverage the high dimension context to carry rich information to both the encoder and the decoder, which helps reconstruct the high-frequency contents for higher video quality. Our framework is also extensible, in which the condition can be flexibly designed. Experiments show that our method can significantly outperform the previous state-of-the-art (SOTA) deep video compression methods. When compared with x265 using veryslow preset, we can achieve 26.0% bitrate saving for 1080P standard test videos. Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
NeurIPS | 2 |
| 2021 | Deep Network-Based Frame Extrapolation With Reference Frame AlignmentabstractFrame extrapolation is to predict future frames from the past (reference) frames, which has been studied intensively in the computer vision research and has great potential in video coding. Recently, a number of studies have been devoted to the use of deep networks for frame extrapolation, which achieves certain success. However, due to the complex and diverse motion patterns in natural video, it is still difficult to extrapolate frames with high fidelity directly from reference frames. To address this problem, we introduce reference frame alignment as a key technique for deep network-based frame extrapolation. We propose to align the reference frames, e.g. using block-based motion estimation and motion compensation, and then to extrapolate from the aligned frames by a trained deep network. Since the alignment, a preprocessing step, effectively reduces the diversity of network input, we observe that the network is easier to train and the extrapolated frames are of higher quality. We verify the proposed technique in video coding, using the extrapolated frame for inter prediction in High Efficiency Video Coding (HEVC) and Versatile Video Coding (VVC). We investigate different schemes, including whether to align between the target frame and the reference frames, and whether to perform motion estimation on the extrapolated frame. We conduct a comprehensive set of experiments to study the efficiency of the proposed method and to compare different schemes. Experimental results show that our proposal achieves on average 5.3% and 2.8% BD-rate reduction in Y component compared to HEVC, under low-delay P and low-delay B configurations, respectively. Our proposal performs much better than the frame extrapolation without reference frame alignment. Dong Liu 0002, Bin Li 0012, Siwei Ma 0001, Feng Wu 0001, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | A Deep Reinforcement Learning Approach to Multiple Streams' Joint Bitrate AllocationabstractFor widely used real-time applications, encoding and transmitting multiple videos jointly over a limited bandwidth has become a popular topic. Allocating different bitrates for different sources is a better way to meet different demands from applications. In this paper, we focus on providing equal quality to users by minimizing the variance of distortion among sequences, which is denoted as the minVAR problem. The state-of-the-art Look-ahead and Feed-back Allocation Model (LFAM) allocates bitrate by taking both look-ahead complexity measures and feed-back information into consideration. However, LFAM brings additional delay to real-time applications. By taking the bitrate allocation problem as a time-series decision making problem, we propose a Deep-Reinforcement-Learning-based approach to allocate bitrate with only feed-back information to solve the two-source minVAR problem. Afterward, we introduce a binary-tree-based hierarchical approach to apply our model to arbitrary number of sources. Tested with the widely used open-source x264 encoder, our approach decreases the variance compared with LFAM in all experiments under two-, three- and four-source scenarios. Furthermore, the proposed approach also outperforms LFAM in the mean quality. The proposed approach is insensitive to the order of sequences and encoders with different complexities, showing its robustness and generalization capability. Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Residual Refinement Network with Attribute Guidance for Precise Saliency DetectionabstractAs an important topic in the multimedia and computer vision fields, salient object detection has been researched for years. Recently, state-of-the-art performance has been witnessed with the aid of the fully convolutional networks (FCNs) and the various pyramid-like encoder-decoder frameworks. Starting from a common encoder-decoder architecture, we enhance a residual refinement network with feature purification for better saliency estimation. To this end, we improve the global knowledge streams with intermediate supervisions for global saliency estimation and design a specific feature subtraction module for residual learning, respectively. On the basis of the strengthened network, we also introduce an attribute encoding sub-network (AENet) with a grid aggregation block (GAB) to guide the final saliency predictor to obtain more accurate saliency maps. Furthermore, the network is trained with a novel constraint loss besides the traditional cross-entropy loss to yield the finer results. Extensive experiments on five public benchmarks show our method achieves better or comparable performance compared with previous state-of-the-art methods. Feng Lin 0009, Wengang Zhou 0001, Jiajun Deng, Bin Li 0012, Yan Lu 0001, Houqiang Li |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | Affinity Derivation for Accurate Instance SegmentationabstractAffinity, which represents whether two pixels belong to a same instance, is an equivalent representation to the instance segmentation labels. Conventional works do not make an explicit exploration on the affinity. In this article, we present two instance segmentation schemes based on pixel affinity information and show the effectiveness of affinity in both aspects. For proposal-free method, we predict pixel affinity for each image and then propose a simple yet effective graph merge algorithm to cluster pixels into instances. It shows that the affinity is powerful as an instance-relevant information to guide the clustering procedure in proposal-free instance segmentation. For proposal-based methods, we extend conventional framework with affinity head and introduce affinity as attached supervision in training phase. Without any additional inference cost, we can improve the performance of existing proposal-based instance segmentation methods, which shows that the affinity can also be applied as an auxiliary loss and training with such extra loss is beneficial to the training progress. Experimental results show that our schemes achieve comparable performance to other state-of-the-art instance segmentation methods. With Cityscapes training data, the proposed proposal-free method achieves 28.8 AP and the proposal-based method gets 27.2 AP both on test sets. Siyu Yang 0006, Bin Li 0012, Wengang Zhou 0001, Jizheng Xu, Houqiang Li, Yan Lu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | Quadtree-Based Coding Framework for High-Density Camera Array-Based Light Field ImageabstractThe size of a high-density-camera-array (HDCA)-based light field image (LFI) is usually very large, containing hundreds of high-resolution views. Therefore, there is an urgent need to efficiently compress it. Currently, no compression algorithms, specially, for the HDCA-based LFI have been designed. In this paper, we propose an algorithm based on a quadtree-based 2D hierarchical coding framework for the HDCA-based LFI data compression. The proposed framework has the following contributions. First, we organize the views of the HDCA-based LFI into a quadtree-based coding structure. Under this structure, all of the views are divided into four quadrants at the first level. Each quadrant is further sub-divided into four quadrants at each subsequent level. The process continues until the desired depth is reached. This quadtree-based coding structure can make full use of the strong inter-view correlations to improve the coding efficiency. In addition, the proposed quadtree-based structure can be easily extended to a general 2D hierarchical structure with variable group of pictures (GOP) sizes to adapt to the reference frame buffer constraint. Second, we try to improve the performance of the 2D hierarchical coding framework using the distance-based criteria for both the reference frame selection and motion vector scaling. Third, a one-pass optimal bit allocation scheme is proposed to further optimize the performance by taking the quality dependencies among various views into consideration. The proposed framework is implemented in the newest video coding standard, high efficiency video coding (HEVC). The experimental results show that the proposed quadtree-based 2D hierarchical coding framework can achieve an average of over 25% bitrate saving compared with the 1D hierarchical coding structure. Li Li 0040, Zhu Li 0001, Bin Li 0012, Dong Liu 0002, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Single-stage Instance SegmentationabstractAlbeit the highest accuracy of object detection is generally acquired by multi-stage detectors, like R-CNN and its extension approaches, the single-stage object detectors also achieve remarkable performance with faster execution and higher scalability. Inspired by this, we propose a single-stage framework to tackle the instance segmentation task. Building on a single-stage object detection network in hand, our model outputs the detected bounding box of each instance, the semantic segmentation result, and the pixel affinity simultaneously. After that, we generate the final instance masks via a fast post-processing method with the help of the three outputs above. As far as we know, it is the first attempt to segment instances in a single-stage pipeline on challenging datasets. Extensive experiments demonstrate the efficiency of our post-processing method, and the proposed framework obtains competitive results as a single-stage instance segmentation method. We achieve 32.5 box AP and 26.0 mask AP on the COCO validation set with 500 pixels input scale and 22.9 mask AP on the Cityscapes test set. Feng Lin 0009, Bin Li 0012, Wengang Zhou 0001, Houqiang Li, Yan Lu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | Convolutional Neural Network-Based Fractional-Pixel Motion CompensationabstractFractional-pixel motion compensation (MC) improves the efficiency of inter prediction and has been utilized extensively in video coding standards. The traditional methods of fractional-pixel MC usually follow the approach of interpolation, i.e., they adopt different kinds of filters, either fixed or adaptive, to interpolate fractional-pixel values from integer-pixel values in a reference picture. Different from the interpolation approach, in this paper, we formulate the fractional-pixel MC as an inter-picture regression problem, which is to predict the pixel values of the current to-be-coded picture from the integer-pixel values of a reference picture, given a fractional-pixel motion vector that relates the two pictures. We then propose to adopt convolutional neural network (CNN) models to approach the regression problem, inspired by the recent advances of CNN. Accordingly, we propose fractional-pixel reference generation CNN (FRCNN) for both uni-directional and bi-directional MC in video coding. We further investigate how to train FRCNN by using encoded video sequences, and empirically study the effect of different training data and different CNN structures. Moreover, we propose to integrate FRCNN into the high efficiency video coding (HEVC) scheme, and perform a comprehensive set of experiments to evaluate the effectiveness of FRCNN. The experimental results show that our proposed FRCNN achieves on average 3.9%, 2.7%, and 1.3% bits saving compared with HEVC, under low-delay P, low-delay B, and random-access configurations, respectively. Ning Yan 0001, Dong Liu 0002, Houqiang Li, Bin Li 0012, Li Li 0040, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | A Hardware-Accelerated System for High Resolution Real-Time Screen SharingabstractEstablishing an interactive screen sharing system that supports ultra high resolution (such as 4k) is challenging, with latency and frame rate playing important roles in user experience. The screen frame needs to be compressed efficiently without consuming extensive computational resources. We present a hardware-accelerated system for real-time screen sharing, which decreases encoding workload by exploiting content redundancies between successive screen frames. We propose a multiple codec approach that utilizes several encoders with H.264 Advanced Video Coding (H.264/AVC) of different input sizes, creating savings in encoding time by selecting the appropriate one for updated screen content. An optimized metadata processing method is proposed as well. Small but distant updates within a frame can be split into independent frames for more efficient compression, which is also beneficial for interactive latency. In the evaluation, the proposed system takes less encoding time than general single codec implementation in common screen sharing scenarios. Measurement for latency shows that the end-to-end latency for 4K resolution screen sharing is only about 17-25 ms, which makes the proposed system suitable for various applications in local wired and wireless connections. Siyu Yang 0006, Bin Li 0012, You Song, Jizheng Xu, Yan Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Invertibility-Driven Interpolation Filter for Video CodingabstractMotion compensation with fractional motion vector has been widely utilized in the video coding standards. The fractional samples are usually generated by fractional interpolation filters. Traditional interpolation filters are usually designed based on the signal processing theory with the assumption of band-limited signal, which cannot effectively capture the non-stationary property of video content and cannot adapt to the variety of video quality. In this paper, we reveal an intuitive property of the fractional interpolation problem, named invertibility. That is, the fractional interpolation filters should not only generate fractional samples from integer samples but also recover the integer samples from the fractional samples in an invertible manner. We prove in theory that the invertibility in the spatial domain is equivalent to the constant magnitude in the Fourier transform domain. Driven by the invertibility, we then develop a learning-based method to solve the fractional interpolation problem. Inspired by the advances of convolutional neural network (CNN), we propose to establish an end-to-end scheme using CNN to train invertibility-driven interpolation filter (InvIF). Different from the previous learning-based methods, the proposed training scheme does not need hand-crafted "ground truth" of fractional samples. The proposed InvIF is integrated into high efficiency video coding (HEVC), and extensive experiments are conducted to verify its effectiveness. The experimental results show that the proposed method can achieve on average 4.7% and 3.6% BD-rate reduction compared with the HEVC anchor, under low-delay-B and random-access configurations, respectively. Ning Yan 0001, Dong Liu 0002, Houqiang Li, Bin Li 0012, Li Li 0040, Feng Wu 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Intra Block Copy for Screen Content in the Emerging AV1 Video CodecabstractScreen content coding plays an important role in many applications. To meet the growing demands of screen content coding, the emerging AV1 video codec incorporates several coding tools, which are specially designed for screen content utilizing its distinctive characteristics. Among these tools, the intra block copy utilizes the characteristic that repeating patterns frequently occur in screen content. This paper presents the technology of intra block copy in AV1. In particular, to efficiently search the predictor in the reconstructed regions of the current picture, AV1 uses the hash matching method at the encoder side. For the generation of hash table, a bottom-to-up manner is adopted to reduce the redundant computation and then decrease the encoding time. In addition, several constraints are involved to facilitate hardware design. Experimental results demonstrate that the intra block copy in AV1 can bring 27.1% bitrate saving for screen content. When compared with the non hash-based intra block copy, the hash-based method achieves 12.2% bitrate saving. Jiahao Li 0001, Hui Su, Alex Converse, Bin Li 0012, Roger Zhou, Bruce Lin, Jizheng Xu, Yan Lu 0001, Ruiqin Xiong |
DCC | 4 |
| 2018 | Affinity Derivation and Graph Merge for Instance Segmentation
Siyu Yang 0006, Bin Li 0012, Wengang Zhou 0001, Jizheng Xu, Houqiang Li, Yan Lu 0001 |
ECCV (3) | 3 |
| 2018 | Convolutional Neural Network-Based Invertible Half-Pixel Interpolation Filter for Video CodingabstractFractional-pixel interpolation has been widely used in the modern video coding standards to improve the accuracy of motion compensated prediction. Traditional interpolation filters are designed based on the signal processing theory. However, video signal is non-stationary, making the traditional methods less effective. In this paper, we reveal that the interpolation filter can not only generate the fractional pixels from the integer pixels, but also reconstruct the integer pixels from the fractional ones. This property is called invertibility. Inspired by the invertibility of fractional-pixel interpolation, we propose an end-to-end scheme based on convolutional neural network (CNN) to derive the invertible interpolation filter, termed CNNInvIF. CNNlnvIF does not need the “ground-truth” of fractional pixels for training. Experimental results show that the proposed CNNInvIF can achieve up to 4.6% and on average 2.2% BD-rate reduction than HEVC under the low-delay P configuration. Ning Yan 0001, Dong Liu 0002, Houqiang Li, Tong Xu 0001, Feng Wu 0001, Bin Li 0012 |
ICIP | 6 |
| 2018 | Efficient Multiple-Line-Based Intra Prediction for HEVCabstractTraditional intra prediction usually utilizes the nearest reference line to generate the predicted block when considering strong spatial correlation. However, this kind of single-line-based method does not always work well due to at least two issues. One is the incoherence caused by the signal noise or the texture of other objects, where this texture deviates from the inherent texture of the current block. The other reason is that the nearest reference line usually has worse reconstruction quality in block-based video coding. Due to these two issues, this paper proposes an efficient multiple-line-based intra-prediction scheme to improve coding efficiency. Besides the nearest reference line, further reference lines are also utilized. The further reference lines with a relatively higher quality can provide potentially better prediction. At the same time, the residue compensation is introduced to calibrate the prediction of boundary regions in a block when we utilize further reference lines. To speed up the encoding process, this paper designs several fast algorithms. The experimental results show that compared with HM-16.9, the proposed fast search method achieves a 2.0% bit saving on average and up to 3.7% by increasing the encoding time by 112%. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | λ-Domain Optimal Bit Allocation Algorithm for High Efficiency Video CodingabstractRate control typically involves two steps: bit allocation and bitrate control. The bit allocation step can be implemented in various fashions depending on how many levels of allocation are desired and whether or not an optimal rate- distortion (R-D) performance is pursued. The bitrate control step has a simple aim in achieving the target bitrate as precisely as possible. In our recent research, we have developed a λ-domain rate control algorithm that is capable of controlling the bitrate precisely for High Efficiency Video Coding (HEVC). The initial research showed that the bitrate control in the λ-domain can be more precise than the conventional schemes. However, the simple bit allocation scheme adopted in this initial research is unable to achieve an optimal R-D performance reflecting the inherent R-D characteristics governed by the video content. In order to achieve an optimal R-D performance, the bit allocation algorithms need to be developed taking into account the video content of a given sequence. The key issue in deriving the video-content-guided optimal bit allocation algorithm is to build a suitable R-D model to characterize the R-D behavior of the video content. In this paper, to complement the R-λ model developed in our initial work, a D-λ model is properly constructed to complete a comprehensive framework of λ-domain R-D analysis. Based on this comprehensive λ-domain R-D analysis framework, a suite of optimal bit allocation algorithms are developed. In particular, we design both picture-level and basic-unit-level bit allocation algorithms based on the fundamental R-D optimization theory to take full advantage of the content-guided principles. The proposed algorithms are implemented in HEVC reference software, and the experimental results demonstrate that they can achieve an obvious R-D performance improvement with a smaller bitrate control error. The proposed bit allocation algorithms have already been adopted by the Joint Collaborative Team on Video Coding and integrated into the HEVC reference software. Li Li 0040, Bin Li 0012, Houqiang Li, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Diversity-Based Reference Picture Management for Low Delay Screen Content CodingabstractScreen content coding plays an important role in many applications. Conventional reference picture management (RPM) strategies developed for natural content may not work well for screen content. This is because many regions in screen content remain static for a long time, causing a lot of repetitive contents to stay in the decoded picture buffer. The repetitive contents are not conducive to inter prediction, but still occupy valuable memory. This paper proposes a diversity-based RPM scheme for screen content coding. The concept of diversity is introduced for the reference picture set (RPS) to help formulate the RPM problem. By maximizing the diversity of RPS, more potentially better predictions are provided. Better compression performance can then be achieved. Meanwhile, the proposed scheme is nonnormative and compatible with existing video coding standards, such as High Efficiency Video Coding. The experimental results show that, for low delay screen content coding, the bit saving of the proposed scheme is 4.9% on average and up to 13.7%, without increasing encoding time. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Weighted Rate-Distortion Optimization for Screen Content CodingabstractUnlike camera-captured video, screen content (SC) often contains a lot of repeating patterns, which makes some blocks used as references much more important than others. However, conventional rate-distortion optimization (RDO) schemes in video coding do not consider the dependence among image blocks, which often leads to a locally optimal parameter selection, especially for SC. In this paper, we present a weighted RDO scheme for SC coding (SCC), in which the repeating characteristics are taken into account when deciding RD tradeoff for each block. For one block, the number being referenced by the current picture and following pictures is estimated and based on the number, we set a proper weight in the RDO process to reflect its importance from a global point of view. To estimate the number being referenced, we propose a hash-based method to approximate the results to avoid the complexity of direct search. Experimental results show that compared with the High Efficiency Video Coding SCC reference software, 10.1%, 14.5%, and 2.2% on average and up to 25.7%, 39.8%, and 4.6% bit saving can be achieved by considering weights provided by our scheme for hierarchical-B, IBBB, and all intra coding structures, respectively. Thanks to our hash-based design, the complexity increase brought by the proposed scheme is marginal. Bin Li 0012, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Fast Hash-Based Inter-Block Matching for Screen Content CodingabstractIn the latest High Efficiency Video Coding (HEVC) development, i.e., HEVC screen content coding extensions (HEVC-SCC), a hash-based inter-motion search/block matching scheme is adopted in the reference test model, which brings significant coding gains to code screen content. However, the hash table generation itself may take up to half the encoding time and is thus too complex for practical usage. In this paper, we propose a hierarchical hash design and the corresponding block matching scheme to significantly reduce the complexity of hash-based block matching. The hierarchical structure in the proposed scheme allows large block calculation to use the results of small blocks. Thus, we avoid redundant computation among blocks with different sizes, which greatly reduces complexity without compromising coding efficiency. The experimental results show that compared with the hash-based block matching scheme in the HEVC-SCC test model (SCM)-6.0, the proposed scheme reduces about 77% of hash processing time, which leads to 12% and 16% encoding time savings in random access (RA) and low-delay B coding structures. The proposed scheme has been adopted into the latest SCM. A parallel implementation of the proposed hash table generation on graphics processing unit (GPU) is also presented to show the high parallelism of the proposed scheme, which achieves more than 30 frames/s for 1080p sequences and 60 frames/s for 720p sequences. With the fast hash-based block matching integrated into x265 and the hash table generated on GPU, the encoder can achieve 11.8% and 14.0% coding gains on average for RA and low-delay P coding structures, respectively, for real-time encoding. Guangming Shi, Bin Li 0012, Jizheng Xu, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Fully Connected Network-Based Intra Prediction for Image CodingabstractThis paper proposes a deep learning method for intra prediction. Different from traditional methods utilizing some fixed rules, we propose using a fully connected network to learn an end-to-end mapping from neighboring reconstructed pixels to the current block. In the proposed method, the network is fed by multiple reference lines. Compared with traditional single line-based methods, more contextual information of the current block is utilized. For this reason, the proposed network has the potential to generate better prediction. In addition, the proposed network has good generalization ability on different bitrate settings. The model trained from a specified bitrate setting also works well on other bitrate settings. Experimental results demonstrate the effectiveness of the proposed method. When compared with high efficiency video coding reference software HM-16.9, our network can achieve an average of 3.4% bitrate saving. In particular, the average result of 4K sequences is 4.5% bitrate saving, where the maximum one is 7.4%. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong, Wen Gao 0001 |
IEEE Trans. Image Process. | 2 |
| 2017 | Pseudo Sequence Based 2-D Hierarchical Coding Structure for Light-Field Image CompressionabstractIn this paper, we present a novel pseudo sequence based 2-D hierarchical reference structure for light-field image compression. In the proposed scheme, we first decompose the light-field image into multiple views and organize them into a 2-D coding structure according to the spatial coordinates of the corresponding microlens. Then we mainly develop three technologies to optimize the 2-D coding structure. First, we divide all the views into four quadrants, and all the views are encoded one quadrant after another to reduce the reference buffer size as much as possible. Inside each quadrant, all the views are encoded hierarchically to fully exploit the correlations between different views. Second, we propose to use the distance between the current view and its reference views as the criteria for selecting better reference frames for each inter view. Third, we propose to use the spatial relative positions between different views to achieve more accurate motion vector scaling. The whole scheme is implemented in the reference software of High Efficiency Video Coding. The experimental results demonstrate that the proposed novel pseudo-sequence based 2-D hierarchical structure can achieve maximum 14.2% bit-rate savings compared with the state-of-the-art light-field image compression method. Li Li 0040, Zhu Li 0001, Bin Li 0012, Dong Liu 0002, Houqiang Li |
DCC | 3 |
| 2017 | Intra Prediction Using Multiple Reference Lines for Video CodingabstractTraditional intra prediction schemes usually only use the nearest adjacent reference line to generate the prediction. Although the nearest reference line generally has the strongest statistical correlation with current block, the farther non-adjacent reference lines can still provide potential better prediction in some cases. Thus, in this paper, not only the nearest reference line but also the farther reference lines are utilized to help intra prediction. When using the farther reference lines, an additional residue compensation procedure is introduced to further refine the prediction. In particular, this paper designs three solutions to meet different complexity requirements. They are multiple line-based intra prediction (MLIP), fast search for multiple line-based intra prediction (FS-MLIP), and dual line-based intra prediction (DLIP). Experimental results verify the effectiveness of the proposed methods. When compared with HM-16.9, the proposed MLIP achieves 2.4% bit saving on average with the encoding time increasing about 362%. The FS-MLIP achieves 2.0% bit saving on average with the encoding time increasing about 114%. The DLIP achieves 0.9% bit saving on average with the encoding time increasing about only 15%. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
DCC | 2 |
| 2017 | Intra prediction using fully connected network for video codingabstractTraditional intra prediction methods exploit some fixed rules to generate prediction, which might not be adaptive enough to handle complicated contents. In this paper, we investigate applying deep neural network to improve the state-of-the-art intra prediction. Considering the characteristics of block-based video coding framework, we propose a fully connected network for intra prediction where all layers except non-linear ones are fully connected. In the proposed network, the inputs are multiple reference lines of the current block and the output is the prediction for the block. When compared with the traditional intra prediction method, the richer context of current block is exploited. For this reason, the proposed network is capable of providing more accurate prediction. Experimental results demonstrate the effectiveness of proposed network. When integrated into the HEVC reference software, the proposed method can achieve up to 3.3% bitrate saving and an average of 1.6% bitrate saving for 4K sequences. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
ICIP | 2 |
| 2017 | Selective motion estimation strategy based on content classification for HEVC screen content codingabstractMotion estimation is one of the most time-consuming parts in video coding. This paper proposes a selective motion estimation strategy to reduce the complexity of the encoder for HEVC screen content coding. The proposed scheme is based on content classification. For each screen content inter frame, it is classified into screen content areas and camera-captured areas. Instead of using one single motion estimation search method for the whole frame, different methods are adaptively switched for different areas. Experimental results show that, compared with SCM-7.0 anchor, the proposed strategy achieves 15% and 13% encoding time saving on average merely with average 0.20% and 0.12% Y/G BD-rate loss for low delay and random access configurations, respectively. Bin Li 0012 |
ICIP | 3 |
| 2017 | A convolutional neural network-based approach to rate control in HEVC intra codingabstractRate control is an essential element for the practical use of video coding standards. A rate control scheme typically builds a model that characterizes the relationship between rate (R) and a coding parameter, e.g. quantization parameter or Lagrange multiplier (A). In such a scheme, the rate control performance depends highly on the modeling accuracy. For inter frames, the model parameters can be precisely updated to fit the video content, based on the information of previously coded frames. However, for intra frames, especially the first frame of a video sequence, there is no prior information to rely on. Therefore, intra frame rate control has remained a challenge. In this paper, we adopt the R-A model to characterize each coding tree unit (CTU) in an intra frame, and we propose a convolutional neural network (CNN) based approach to effectively predict the model parameters for every CTU. Then we develop a new CTU level bit allocation and bitrate control algorithm based on the R-A model for HEVC intra coding. The experimental results show that our proposed CNN-based approach outperforms the currently used rate control algorithm in HEVC reference software, leading to on average 0.46 percent decrease of rate control error and 0.7 percent BD-rate reduction. Bin Li 0012, Dong Liu 0002, Zhibo Chen 0001 |
VCIP | 2 |
| 2017 | Rate control with delay constraint for screen content codingabstractDifferent from conventional video, screen content video often contains large movements and abrupt changes between adjacent frames. These distinct characteristics bring great challenges to the implementation of rate control in screen content coding (SCC). This paper proposes a novel rate control scheme for SCC considering the delay constraint. First, a pre-analyzer is designed to collect the information of the proceeding frames, which are about to be encoded. Then, after the information collection, more rational bit allocation strategy is adopted to keep the encoder from both buffer overflow and underflow. Furthermore, a delay constraint condition is derived for buffer and pre-analyzer to avoid additional delay. Experimental results demonstrate that the proposed scheme achieves more accurate rate control accuracy and 3.10 dB gain on average when compared with the existing rate control scheme in HM-16.8+SCM-7.0. Junshi Xiao, Bin Li 0012, Songlin Sun, Jizheng Xu |
VCIP | 2 |
| 2016 | Diagonal motion partitions for inter prediction in HEVCabstractThis paper presents diagonal motion partitions (DMP) for inter prediction in HEVC. In addition to the square and rectangular partitions, we propose to add diagonal shaped partitions to match different motion parts with oblique boundaries. Considering the overlap of pixels along the partition boundaries, the calculation of sum of absolute differences (SAD) and the motion compensation for the pixels on the boundaries are weighted. Besides, the residues of a diagonal prediction unit (PU) are augmented to form a rectangular one, so as to perform Hadamard transform on the residues. We also revise the advanced motion vector prediction (AMVP) and merge candidates based on the diagonal motion partitions. Experimental results show that on average 0.8%-1.0% BD-rate reduction can be achieved by DMP. Ning Yan 0001, Bin Li 0012, Jizheng Xu, Houqiang Li, Feng Wu 0001 |
VCIP | 2 |
| 2016 | An Efficient Fast Mode Decision Method for Inter Prediction in HEVCabstractThe emerging High Efficiency Video Coding (HEVC) standard adopts many advanced techniques with flexible combinations, which enables HEVC to achieve about 50% bit-rate reduction for similar perceptual video quality relative to the prior video coding standard H.264/Advanced Video Coding. However, the enormously increased encoding complexity of HEVC inevitably becomes one of the greatest challenges for real-time applications. Among all the factors resulting in the increase in encoding complexity of HEVC, the quad-tree structure for coding units (CUs) with different sizes and accordingly a large number of prediction modes is one critical reason. Thus, it is greatly desired to develop a fast mode decision method for HEVC to reduce the computational complexity. In this paper, considering that HEVC employs the quad-tree structure, and the distortion of each sub-CU can indicate whether the current mode is suitable for current CU, we explore the relationship between the impossible modes and the distribution of the distortions to help the encoder skip checking the unnecessary modes. Besides, since the residual values can reflect the prediction result directly, we propose a method to skip some motion estimation operations according to the distribution of the residuals. Experimental results show that the proposed method can save about 77% of encoding time with only about a 4.1% bit-rate increase compared with HM16.4 anchor, while compared with the fast mode decision method adopted in HM16.4, the proposed algorithm can save about 48% of encoding time with only about a 2.9% bit-rate increase. Jinlei Zhang, Bin Li 0012, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | λ-Domain Rate Control Algorithm for HEVC Scalable ExtensionabstractIn this paper, we propose a λ-domain rate control algorithm for high efficiency video coding (HEVC) scalable extension. All the commonly used scalabilities including temporal, spatial, and quality scalability are taken into consideration. The proposed algorithm mainly has three key contributions. First, we propose an optimal initial target bits and initial encoding parameters determination algorithm for the first frame of each layer to achieve the best rate-distortion (R-D) performance. Second, an optimal bit allocation algorithm taking both the intra and inter layer dependence into consideration is proposed for the inter frames under spatial and quality scalability cases. Third, since the coding scheme of HEVC scalable extension with multiple layers is even more flexible than HEVC, an adaptive updating algorithm for R-λ model is proposed to control the bits per frame even more precisely. The experimental results demonstrate that the proposed λ-domain rate control algorithm can bring both more precise bitrate accuracy and better R-D performance compared with the previous rate control algorithms for HEVC scalable extension. Li Li 0040, Bin Li 0012, Dong Liu 0002, Houqiang Li |
IEEE Trans. Multim. | 2 |
| 2015 | A Fast Algorithm for Adaptive Motion Compensation Precision in Screen Content CodingabstractFractional-pel motion compensation is very good at improving video coding efficiency, especially for camera-captured content. But for screen content, which is obtained from a computer desktop, motion vectors with integer-precision may be enough to represent the motion in different pictures. Using fractional-pel motion compensation for such content is a waste of bits. Thus, adaptive motion compensation precision is helpful for improving coding efficiency, especially for screen content coding. Usually, to select suitable motion compensation precision, multi-pass encoding is introduced, which significantly increases the encoding time. This paper presents a fast encoding algorithm for adaptive motion compensation precision used in screen content coding by hash-based block matching. With the proposed method, multi-pass encoding is avoided and most of the benefits brought by adaptive motion compensation precision are preserved. The experimental results show that with the proposed method, up to 7.7% bit saving is obtained without a significant impact on encoding time. Bin Li 0012, Jizheng Xu |
DCC | 1 |
| 2015 | Rate control for screen content coding in HEVCabstractScreen content usually has much different motion characteristics compared with conventional videos, which makes exiting rate control schemes unsuitable. This paper analyses the motion characteristics of screen content and proposes an efficient rate control method for screen content coding, by improving the bit allocation and model parameter adaptation strategies. The proposed rate control algorithm is able to both control bitrate accurately and improve the quality of the entire video sequence. The experimental results demonstrate that the proposed algorithm achieves smaller bitrate errors and better coding performance. Compared with the existing rate control scheme in High Efficiency Video Coding (HEVC) reference software, the proposed algorithm improves the coding efficiency by 5.6% on average while obtaining smaller bitrate errors. Yaoyao Guo, Bin Li 0012, Songlin Sun, Jizheng Xu |
ISCAS | 2 |
| 2015 | Rate control for screen content coding based on picture classificationabstractEmerging screen content coding brings great challenges to rate control due to significantly different characteristics of screen content from conventional video, e.g. large motion, frequent scene changes. This paper proposes an efficient rate control scheme for screen content coding by considering the characteristics of screen content. We first classify pictures of a video sequence into different groups by comparing the current picture with its neighbours. Then for each group, we apply different strategies to bit allocation and parameter updating process based on the characteristics of each group. The proposed algorithm can control the bitrate accurately without introducing any additional encoding delay. Experiment results show that the proposed scheme can significantly improve PSNR with a more accurate bitrate compared with the existing rate control scheme. Yaoyao Guo, Bin Li 0012, Songlin Sun, Jizheng Xu |
VCIP | 2 |
| 2015 | An adaptive hierarchical QP setting for screen content codingabstractScreen content refers to computer generated content like text, graphics, and animations. In such video, many regions may remain static for a long period after a sudden change. Traditional hierarchical Quantization Parameter (QP) setting may not be able to handle these regions efficiently because the encoder probably needs to refine the quality of these static regions multiple times. It will cost more bits while the quality of the static regions may reach the expected degree which the flat QP setting is able to achieve. This paper proposes using different QP settings for different regions in a picture. Region classification algorithms are developed to determine whether a flat or hierarchical QP setting is used. Experimental results demonstrate that the proposed scheme can achieve an average bitrate reduction of 3.1%, and up to 8.1% bitrate reduction for IBBB coding. The proposed method improves coding efficiency without increasing encoding complexity. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
VCIP | 2 |
| 2015 | Weighted rate-distortion optimization for screen content intra codingabstractScreen content videos often have mixed content consisting of various types such as natural content, text and graphics in the same picture. To achieve high coding efficiency for this content, new coding tools are developed in the High Efficiency Video Coding (HEVC) standard Screen Content Coding (SCC) extension. Among them, intra block copy (IntraBC) allows a nonlocal intra prediction from the coded region of the same picture. However, the rate-distortion optimization (RDO) scheme of the screen content coding still follows that of the HEVC reference software and each block is optimized locally. The screen content characteristics are not fully utilized in the current RDO scheme. This paper presents a weighted RDO scheme for the intra coding of screen content videos, which measures the importance of each block within the picture first and then larger distortion weights are applied to the blocks with larger importance in the RDO. In this way, a better rate-distortion trade-off can be achieved for the picture instead of the blocks themselves. The experimental results show that up to 5.5% coding efficiency gain can be achieved compared with the reference software. Bin Li 0012, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
VCIP | 2 |
| 2015 | HEVC Encoding Optimization Using Multicore CPUs and GPUsabstractAlthough the High Efficiency Video Coding (HEVC) standard significantly improves the coding efficiency of video compression, it is unacceptable even in offline applications to spend several hours compressing 10 s of high-definition video. In this paper, we propose using a multicore central processing unit (CPU) and an off-the-shelf graphics processing unit (GPU) with 3072 streaming processors (SPs) for HEVC fast encoding, so that the speed optimization does not result in loss of coding efficiency. There are two key technical contributions in this paper. First, we propose an algorithm that is both parallel and fast for the GPU, which can utilize 3072 SPs in parallel to estimate the motion vector (MV) of every prediction unit (PU) in every combination of the coding unit (CU) and PU partitions. Furthermore, the proposed GPU algorithm can avoid coding efficiency loss caused by the lack of a MV predictor (MVP). Second, we propose a fast algorithm for the CPU, which can fully utilize the results from the GPU to significantly reduce the number of possible CU and PU partitions without any coding efficiency loss. Our experimental results show that compared with the reference software, we can encode high-resolution video that consumes 1.9% of the CPU time and 1.0% of the GPU time, with only a 1.4% rate increase. Bin Li 0012, Jizheng Xu, Guangming Shi, Feng Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2014 | Adaptive weighted distortion optimization for video coding in RGB color spaceabstractIn this paper, we propose an adaptive weighted distortion optimization algorithm for improving the coding performance of video represented in RGB color space. The YCbCr is the color space, which is usually adopted in video coding. And with 4:2:0 planar format, it generally uses the ratio 6:1:1 metrics objective peak signal-to-noise ratio (PSNR) for quality assessment. In this paper, we still use 6:1:1 combined PSNR as a universal criterion to optimize video coded in RGB domain and measure the final quality of decoded video as well. The new distortion metric in RGB coding is derived from the optimal target, and the modification is only applied on the encoder side. The proposed algorithm achieves about 26% bit-saving on average and maximum 44% bit-saving in HEVC range extensions reference software HM12.0 RExt4.1. Additionally, the visual quality of video coded by the proposed algorithm is also improved. Honggang Qi, Bin Li 0012, Jizheng Xu |
ICIP | 3 |
| 2014 | 1-D dictionary mode for screen content codingabstractThis paper introduces 1-D dictionary mode designed for screen content coding. Two 1-D dictionary modes are designed to improve the coding efficiency for screen content. The first one is called normal dictionary mode, in which a virtual dictionary should be maintained and all the prediction comes from the virtual dictionary. The other one is called reconstruction based dictionary mode, where no virtual dictionary is to be maintained and all the previously reconstructed pixels in the same picture can be used for prediction. Hash based search is designed to find matching for both dictionary modes efficiently. 1-D dictionary mode with variable block sizes are also supported in the proposed scheme. The experimental results show the proposed algorithm achieves about 10% ~ 18.4% bit saving for different coding structures. The bit saving is up to 60% for the proposed method. Bin Li 0012, Jizheng Xu, Feng Wu 0001 |
VCIP | 1 |
| 2014 | A unified framework of hash-based matching for screen content codingabstractThis paper introduces a unified framework of hash-based matching method for screen content coding. Screen content has some different characteristics from camera-captured content, such as large motion and repeating patterns. Hash-based matching is proposed to better explore the correlation in screen content, thus, improving the coding efficiency. The proposed method can handle both intra picture and inter picture block matching with variable block sizes in a unified framework. The proposed framework is also easy to be extended to handle other motion models to further improve the coding efficiency of screen content. We also develop fast encoding algorithms to make full use of the hash results. The experimental results show the proposed algorithm achieves about 12% bit saving while saving more than 25% encoding time. The bit saving is up to 57% and the encoding time saving is up to 60% for the proposed method. Bin Li 0012, Jizheng Xu, Feng Wu 0001 |
VCIP | 1 |
| 2014 | λ Domain Rate Control Algorithm for High Efficiency Video CodingabstractRate control is a useful tool for video coding, especially in real-time communication applications. Most of existing rate control algorithms are based on the R-Q model, which characterizes the relationship between bitrate R and quantization Q , under the assumption that Q is the critical factor on rate control. However, with the video coding schemes becoming more and more flexible, it is very difficult to accurately model the R-Q relationship. In fact, we find that there exists a more robust correspondence between R and the Lagrange multiplier λ . Therefore, in this paper, we propose a novel λ -domain rate control algorithm based on the R-λ model, and implement it in the newest video coding standard high efficiency video coding (HEVC). Experimental results show that the proposed λ -domain rate control can achieve the target bitrates more accurately than the original rate control algorithm in the HEVC reference software as well as obtain significant R-D performance gain. Thanks to the high accurate rate control algorithm, hierarchical bit allocation can be enabled in the implemented video coding scheme, which can bring additional R-D performance gain. Experimental results demonstrate that the proposed λ -domain rate control algorithm is effective for HEVC, which outperforms the R-Q model based rate control in HM-8.0 (HEVC reference software) by 0.55 dB on average and up to 1.81 dB for low delay coding structure, and 1.08 dB on average and up to 3.77 dB for random access coding structure. The proposed λ -domain rate control algorithm has already been adopted by Joint Collaborative Team on Video Coding and integrated into the HEVC reference software. Bin Li 0012, Houqiang Li, Li Li 0040, Jinlei Zhang |
IEEE Trans. Image Process. | 1 |
| 2013 | Refining QP to improve coding efficiency in AVSabstractThis paper presents a Lagrange multiplier determination method and Quantization Parameter (QP) refinement algorithm used in the Rate-Distortion Optimization (RDO) process to improve the coding efficiency for Audio Video Coding Standard (AVS) (IEEE 1857). This paper investigates the Lagrange multiplier setting problem for different kinds of pictures in the encoding process of AVS. When Lagrange multiplier is determined, to minimize the RD (Rate-Distortion) cost, the optimal QP can be selected by multiple-QP optimization. However, this kind of optimization increases the encoding time significantly. This paper analyzes the relationship between the QP value and Lagrange multiplier in this paper. And then this relationship is applied in the encoding process of AVS to refine the predetermined QP value. The experimental results show that proposed algorithm can lead to significant bit saving compared with the default anchor of AVS. The proposed method improves the coding efficiency without increasing the processing time. Bin Li 0012, Jizheng Xu, Houqiang Li |
ICIP | 1 |
| 2013 | Rate-distortion optimization with adaptive weighted distortion in high Efficiency Video CodingabstractThis paper presents an adaptive weighted distortion optimization algorithm used in the Rate-Distortion Optimization (RDO) process of the High Efficiency Video Coding (HEVC). RDO is an important tool to improve the coding efficiency. Usually the distortion weights of different color components are equal or predetermined. In this paper, an adaptive weighted distortion optimization algorithm is introduced to improve the coding efficiency. The distortion weight is estimated according to the previous coded pictures belonging to the same temporal level, such that encoding complexity is almost unchanged. With the proposed adaptive weighted distortion optimization method, on average about 3.3% and up to 10.6% bit-saving are obtained based on the latest HEVC reference software, HM-8.0 and the corresponding common test conditions. The proposed algorithm can also be applied to other coding schemes such as H.264/MPEG-4 AVC. Bin Li 0012, Jizheng Xu, Houqiang Li |
ISCAS | 1 |
| 2013 | QP refinement according to Lagrange multiplier for High Efficiency Video CodingabstractThis paper presents a Quantization Parameter (QP) refinement algorithm used in Rate-Distortion Optimization (RDO) process to improve the coding efficiency for High Efficiency Video Coding (HEVC). To minimize the RD cost, QP is one of the parameters that can be optimized. Usually, multiple-QP optimization can be applied to choose the best QP value. However, this kind of optimization increases the encoding complexity significantly. We analyze the relationship between the QP value and Lagrange multiplier in this paper firstly. And then we apply the relationship between QP and Lagrange multiplier to the encoding process of HEVC. The proposed QP refinement method has already been adopted into HM (the HEVC reference software). The experimental results show that proposed algorithm can achieve about 1.4%~1.9% bit saving for luma component on average compared with the default anchor of HM-8.0. The bit saving on chroma components are larger than that of luma. The proposed method improves the coding efficiency without increasing the processing time. Bin Li 0012, Jizheng Xu, Houqiang Li |
ISCAS | 1 |
| 2013 | A Video Communication System Based on Spatial Rewriting and ROI Rewriting
Fangdong Chen, Bin Li 0012, Houqiang Li |
MMM (2) | 3 |
| 2012 | Fast Transcoding from H.264 AVC to High Efficiency Video CodingabstractIn this paper, we present several strategies for transcoding from H.264/AVC bit streams to High Efficiency Video Coding (HEVC) bit streams. Because HEVC and AVC share the similar coding architecture, we try to exploit the information in AVC bit streams as much as possible. For inter picture, we utilize the power spectrum based rate-distortion optimization (PS-RDO) model as well as the input residual, modes and motion vectors to estimate the best coding unit (CU) split quad tree, the best prediction unit (PU) mode and the best motion vector of each PU partition For intra picture, we propose to reduce the CU and PU partition candidates. The proposed strategies can significantly reduce the transcoding complexity in terms of reduced processing for RDO evaluations, motion estimation, motion compensation as well as fractional pixel interpolation operations. Experiment results show that the proposed transcoding methods can achieve a good tradeoff between coding efficiency and transcoding complexity. Bin Li 0012, Jizheng Xu, Houqiang Li |
ICME | 2 |
| 2012 | Compression performance of high efficiency video coding (HEVC) working draft 4abstractThis paper presents the results of compression comparison tests between the current state of the emerging High Efficiency Video Coding (HEVC) draft standard and the current dominant standard H.264/MPEG-4 AVC (High Profile) as an anchor reference. The conditions used for the comparison tests were designed to reflect relevant application scenarios and to enable a fair comparison to the maximum extent feasible, i.e. using comparable quantization settings, reference frame buffering, etc. The testing was generally configured in favour of using a relatively strong H.264/MPEG-4 AVC anchor reference. Several of the encoder optimizations currently found in the HEVC software are tested and shown to be helpful to improve the H.264/MPEG-4 AVC anchor performance. When compared to the improved anchor encoder configurations, the HEVC draft design currently provides a bit rate savings for equal PSNR of about 39% for random access applications, 44% for low-delay use, and 25% for all-intra use. Bin Li 0012, Gary J. Sullivan, Jizheng Xu |
ISCAS | 1 |
| 2012 | Counter based adaptation for CAVLC in HEVCabstractCompared to the CAVLC in H.264/AVC, CAVLC in HEVC has been improved much. One mechanism that contributes to the improvement is adaptive code word swapping based on previous coded symbol in the early design of HEVC. However, such a instantly swapping strategy often makes a wrong estimation of the real statistics and may lead to inefficient coding that may even worse than without adaptation. In this paper, we analyze potential loss caused by instantly swapping and propose a counter-based adaptation scheme to further improve the entropy coding efficiency. We discuss the bit-depth of those counters that achieves a trade-off between the adaptation capability and a good estimation of the statistics. The proposed scheme introduces a sum counter for each code table and a counter for each code. The sum counter determines when to shift all the counters to prevent overflow and make sure symbols far away influence less. The experimental results show that the counter based algorithm brings about 0.7% bits saving compared with the instantly swapping adaptation algorithm. Compared with the CAVLC without adaptation, the performance gain of the counter adaptation algorithm is about 4.0% bits saving. A variant of the proposed counter-based adaptation has been adopted by HEVC design. Bin Li 0012, Jizheng Xu, Houqiang Li |
ISCAS | 1 |
| 2012 | Rate-Distortion Optimized Reference Picture Management for High Efficiency Video CodingabstractMotion compensation with multiple reference pictures has been widely used during the development of the emerging High Efficiency Video Coding (HEVC) standard, which greatly helps to improve the coding efficiency. Usually, a heuristic strategy is exploited to use the nearest reconstructed pictures as references. However, such a strategy may not be efficient on all occasions, especially when different content characteristics and coding settings are considered. In this paper, we investigate how to manage reference pictures so as to achieve better rate-distortion performance under the memory constraint of the decoded picture buffer at the decoder. We formulate the reference picture management as an optimization problem and approximate its optimal solution. Moreover, we explore how to adjust quality for each picture according to the reference structure to further improve coding efficiency. For some coding cases, where a complicated encoder optimization is unaffordable, we also develop fast algorithms to get the most benefit from reference picture selection. Among them, one strategy has been adopted by the HEVC software and common test conditions to generate the anchor. Experimental results show that the proposed full search algorithm and fast search algorithms achieve significant bitrate reduction. Houqiang Li, Bin Li 0012, Jizheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Optimized reference frame selection for video coding by cloudabstractWe investigate how to improve video coding efficiency via optimized reference frame selection using large-scale computation resources, e.g., a cloud. We first formulate the optimization problem for reference frame selection in video coding, which can be simplified to a manageable level. Given the maximum number of reference frames for encoding one frame, we give the upper bound of the coding efficiency on the High Efficiency Video Coding (HEVC) platform, which, although ideal, may require a huge amount of reference frame buffering at the decoder. Then we give a solution and the corresponding performance when the reference frame buffer size at the decoder is constrained. Experimental results show that when the number of reference frames is four, the proposed encoding scheme can achieve up to 16.9% bit-saving compared to HEVC, the state-of-the-art video coding system. The proposed encoding scheme is standard-compliant and can also be applied to H.264/AVC to improve coding efficiency. Bin Li 0012, Jizheng Xu, Houqiang Li, Feng Wu 0001 |
MMSP | 1 |
| 2011 | Intermedia: system and application for video adaptationabstractVideo adaptation has been considered as a promising technique to bridge the gap between network status, device capabilities and user preferences in pervasive media applications. However, conventional adaptation framework based on transcoding or multiple pre-transcoding is not competent for accommodating diversified usages and a large number of users. In this paper, the system design of a novel video adaptation system based on Intermediate video description called "Intermedia" is introduced. Intermedia is consist of multiple video signal components, such as texture, motion, structural characteristic, rate control information, ROI, as well as some semantic features. It is off-line generated and stored in the video server or media gateway. Our adaptation system is able to quickly and easily generate required bitstream from Intermedia with very low complexity to fulfill a specific adaptation requirement, e.g., bit-rate conversion, temporal/spatial resolution reduction, video summarization, region-of-interest browsing, and some multi-dimensional adaptation involving both signal level adaptation and semantic level adaptation. The satisfactory performance of our system demonstrates the effective and efficiency of our proposed video adaptation framework. Bin Li 0012, Houqiang Li |
MUM | 2 |
| 2011 | Parsing robustness in High Efficiency Video Coding - analysis and improvementabstractOne problem in the current High Efficiency Video Coding design is that the entropy decoding, or the parsing process of one frame may depend on syntax elements on other frames. Thus, a small error may lead to immediate failure of decoding of the current frame and all the following inter frames. This paper analyzes the cause of this vulnerability of HEVC bitstreams. Based on the analysis, we further present two solutions to make a tradeoff between the coding efficiency and the error robustness by cutting down the error propagation caused by temporal motion vector prediction. One solution is to redesign some code tables so that the dependency on other frames' syntax elements can be broken. The other solution is to insert recovery points for the coded bitstream, so that the decoding process can recover from errors. The pros and cons of these two solutions are discussed. Experimental results show that the proposed methods can solve the parsing problem in HEVC with a marginal influence on the coding performance and provide significant improvements for the decoded video when there are errors. Bin Li 0012, Jizheng Xu, Houqiang Li |
VCIP | 1 |
| 2010 | Hybrid bit-stream rewriting from scalable video coding to H.264/AVCabstractScalable Video Coding (SVC) is an extension of H.264/AVC standard. The base layer of SVC is compatible with H.264/AVC standard, while the enhancement layers provide desired temporal, quality and/or spatial scalabilities. Bit-stream rewriting in SVC standard allows an SVC bit-stream to be converted to an H.264/AVC bit-stream without quality loss and preferably with low computational complexity. However, current rewriting is only supported in quality scalability rather than spatial scalability, which limits the application in many practical scenarios. In this paper, a hybrid bit-stream rewriting approach to support both quality and spatial scalability is proposed based on the principle of residue upsampling in transform domain. The computational complexity of the proposed approach is much lower than the conventional scheme of cascading transcoding. Extensive experimental results demonstrate that the loss of the rate-distortion (RD) performance of the proposed rewritable SVC bit-stream is acceptable compared with the conventional SVC bit-stream, however, the RD performance is better than that of simulcast. Furthermore, the RD performance of the H.264/AVC bit-stream rewritten from the rewritable SVC bit-stream is even better than that of the input SVC bit-stream. Compared with the cascading transcoding scheme, the proposed hybrid rewriting can achieve 0.8 dB Y-PSNR gains while saving 80% processing time on average. Bin Li 0012, Houqiang Li, Chang Wen Chen |
VCIP | 1 |