EDBT 2026 Demo / reviewers in the wild / expert
Jiahao Li 0001
dblp:150/5524-1
· DBLP profile ↗
35ranked-venue papers
11as first author
28since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 29 · 10 first-author · 22 since 2021Artificial intelligence and machine learning · 20 · 3 first-author · 20 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent TrainingabstractGUI agents that interact with graphical interfaces on behalf of users are a promising direction for practical AI assistants, yet training them is hindered by scarce suitable environments. We present InfiniteWeb, a system that automatically generates functional web environments at scale for GUI agent training. While LLMs perform well on generating a single webpage, building a realistic and functional website with many interconnected pages faces challenges. We address these challenges through unified specification, task-centric test-driven development, and combining website seed variation with reference design images. Our system also generates verifiable task evaluators enabling dense reward signals for reinforcement learning. Experiments show that our system surpasses commercial coding agents at realistic website construction, and GUI agents trained on our generated environments achieve significant performance improvements on OSWorld and Online-Mind2Web, demonstrating the effectiveness of the proposed system. Zezhou Wang, Zongyu Guo, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
ACL (1) | 5 |
| 2025 | Towards Practical Real-Time Neural Video CompressionabstractWe introduce a practical real-time neural video codec (NVC) designed to deliver high compression ratio, low latency and broad versatility. In practice, the coding speed of NVCs depends on 1) computational costs, and 2) non-computational operational costs, such as memory I/O and the number of function calls. While most efficient NVCs prioritize reducing computational cost, we identify operational cost as the primary bottleneck to achieving higher coding speed. Leveraging this insight, we introduce a set of efficiency-driven design improvements focused on minimizing operational costs. Specifically, we employ implicit temporal modeling to eliminate complex explicit motion modules, and use single low-resolution latent representations rather than progressive downsampling. These innovations significantly accelerate NVC without sacrificing compression quality. Additionally, we implement model integerization for consistent cross-device coding and a module-bank-based rate control scheme to improve practical adaptability. Experiments show our proposed DCVC-RT achieves an impressive average encoding/decoding speed at 125.2/112.8 fps (frames per second) for 1080p video, while saving an average of 21% in bitrate compared to H.266/VTM. The code is available at https://github.com/microsoft/DCVC. Zhaoyang Jia, Bin Li 0012, Jiahao Li 0001, Wenxuan Xie, Houqiang Li, Yan Lu 0001 |
CVPR | 3 |
| 2025 | PICD: Versatile Perceptual Image Compression with Diffusion RenderingabstractRecently, perceptual image compression has achieved significant advancements, delivering high visual quality at low bitrates for natural images. However, for screen content, existing methods often produce noticeable artifacts when compressing text. To tackle this challenge, we propose versatile perceptual screen image compression with diffusion rendering (PICD), a codec that works well for both screen and natural images. More specifically, we propose a compression framework that encodes the text and image separately, and renders them into one image using diffusion model. For this diffusion rendering, we integrate conditional information into diffusion models at three distinct levels: 1). Domain level: We fine-tune the base diffusion model using text content prompts with screen content. 2). Adaptor level: We develop an efficient adaptor to control the diffusion model using compressed image and text as input. 3). Instance level: We apply instance-wise guidance to further enhance the decoding process. Empirically, our PICD surpasses existing perceptual codecs in terms of both text accuracy and perceptual quality. Additionally, without text conditions, our approach serves effectively as a perceptual codec for natural images. Tongda Xu, Jiahao Li 0001, Bin Li 0012, Yan Wang 0105, Ya-Qin Zhang, Yan Lu 0001 |
CVPR | 2 |
| 2025 | FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language ModelabstractCurrently, instruction-based image editing methods have made significant progress by leveraging the powerful cross-modal understanding capabilities of vision language models (VLMs). However, they still face challenges in three key areas: 1) complex scenarios; 2) semantic consistency; and 3) fine-grained editing. To address these issues, we propose FireEdit, an innovative Fine-grained Instruction-based image editing framework that exploits a REgion-aware VLM. FireEdit is designed to accurately comprehend user instructions and ensure effective control over the editing process. Specifically, we enhance the fine-grained visual perception capabilities of the VLM by introducing additional region tokens. Relying solely on the output of the LLM to guide the diffusion model may lead to suboptimal editing results. Therefore, we propose a Time-Aware Target Injection module and a Hybrid Visual Cross Attention module. The former dynamically adjusts the guidance strength at various denoising stages by integrating timestep embeddings with the text embeddings. The latter enhances visual details for image editing, thereby preserving semantic consistency between the edited result and the source image. By combining the VLM enhanced with fine-grained region tokens and the time-dependent diffusion model, FireEdit demonstrates significant advantages in comprehending editing instructions and maintaining high semantic consistency. Extensive experiments indicate that our approach surpasses the state-of-the-art instruction-based image editing methods. Jiahao Li 0001, Zunnan Xu, Yiji Cheng, Fa-Ting Hong, Qin Lin 0003, Qinglin Lu, Xiaodan Liang |
CVPR | 2 |
| 2025 | DLF: Extreme Image Compression with Dual-Generative Latent Fusion
Naifu Xue, Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
ICCV | 3 |
| 2025 | One-Step Diffusion-Based Image Compression with Semantic DistillationabstractWhile recent diffusion-based generative image codecs have shown impressive performance, their iterative sampling process introduces unpleasant latency. In this work, we revisit the design of a diffusion-based codec and argue that multi-step sampling is not necessary for generative compression. Based on this insight, we propose OneDC, a One-step Diffusion-based generative image Codec—that integrates a latent compression module with a one-step diffusion generator. Recognizing the critical role of semantic guidance in one-step diffusion, we propose using the hyperprior as a semantic signal, overcoming the limitations of text prompts in representing complex visual content. To further enhance the semantic capability of the hyperprior, we introduce a semantic distillation mechanism that transfers knowledge from a pretrained generative tokenizer to the hyperprior codec. Additionally, we adopt a hybrid pixel- and latent-domain optimization to jointly enhance both reconstruction fidelity and perceptual realism. Extensive experiments demonstrate that OneDC achieves SOTA perceptual quality even with one-step generation, offering over 39% bitrate reduction and 20× faster decoding compared to prior multi-step diffusion-based codecs. Project: https://onedc-codec.github.io/ Naifu Xue, Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
NeurIPS | 3 |
| 2025 | Omnidirectional 3D Scene Reconstruction from Single ImageabstractReconstruction of 3D scenes from a single image is a crucial step towards enabling next-generation AI-powered immersive experiences. However, existing diffusion-based methods often struggle with reconstructing omnidirectional scenes due to geometric distortions and inconsistencies across the generated novel views, hindering accurate 3D recovery. To overcome this challenge, we propose Omni3D, an approach designed to enhance the geometric fidelity of diffusion-generated views for robust omnidirectional reconstruction. Our method leverages priors from pose estimation techniques, such as MASt3R, to iteratively refine both the generated novel views and their estimated camera poses. Specifically, we minimize the 3D reprojection errors between paired views to optimize the generated images, and simultaneously, correct the pose estimation based on the refined views. This synergistic optimization process yields geometrically consistent views and accurate poses, which are then used to build an explicit 3D Gaussian Splatting representation capable of omnidirectional rendering. Experimental results validate the effectiveness of Omni3D, demonstrating significantly advanced 3D reconstruction quality in the omnidirectional space, compared to previous state-of-the-art methods. Project page: https://omni3d-neurips.github.io. Jiahao Li 0001, Yan Lu 0001 |
NeurIPS | 2 |
| 2025 | Deep Video Discovery: Agentic Search with Tool Use for Long-form Video UnderstandingabstractLong-form video understanding presents significant challenges due to extensive temporal-spatial complexity and the difficulty of question answering under such extended contexts.
While Large Language Models (LLMs) have demonstrated considerable advancements in video analysis capabilities and long context handling, they continue to exhibit limitations when processing information-dense hour-long videos.
To overcome such limitations, we propose the $\textbf{D}eep \ \textbf{V}ideo \ \textbf{D}iscovery \ (\textbf{DVD})$ agent to leverage an $\textit{agentic search}$ strategy over segmented video clips. Different from previous video agents manually designing a rigid workflow, our approach emphasizes the autonomous nature of agents.
By providing a set of search-centric tools on multi-granular video database,
our DVD agent leverages the advanced reasoning capability of LLM to plan on its current observation state, strategically selects tools to orchestrate adaptive workflow for different queries in light of the gathered information.
We perform comprehensive evaluation on multiple long video understanding benchmarks that demonstrates our advantage.
Our DVD agent achieves state-of-the-art performance on the challenging LVBench dataset, reaching an accuracy of $\textbf{74.2\%}$, which substantially surpasses all prior works, and further improves to $\textbf{76.0\%}$ with transcripts. Zhaoyang Jia, Zongyu Guo, Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001 |
NeurIPS | 4 |
| 2025 | Generative Latent Coding for Ultra-Low Bitrate Image and Video CompressionabstractMost existing approaches for image and video compression perform transform coding in the pixel space to reduce redundancy. However, due to the misalignment between the pixel-space distortion and human perception, such schemes often face the difficulties in achieving both high-realism and high-fidelity at ultra-low bitrate. To solve this problem, we propose Generative Latent Coding (GLC) models for image and video compression, termed GLC-image and GLC-Video. The transform coding of GLC is conducted in the latent space of a generative vector-quantized variational auto-encoder (VQ-VAE). Compared to the pixel-space, such a latent space offers greater sparsity, richer semantics and better alignment with human perception, and show its advantages in achieving high-realism and high-fidelity compression. To further enhance performance, we improve the hyper prior by introducing a spatial categorical hyper module in GLC-image and a spatio-temporal categorical hyper module in GLC-video. Additionally, the code-prediction-based loss function is proposed to enhance the semantic consistency. Experiments demonstrate that our scheme shows high visual quality at ultra-low bitrate for both image and video compression. For image compression, GLC-image achieves an impressive bitrate of less than 0.04 bpp, achieving the same FID as previous SOTA model MS-ILLM while using 45% fewer bitrate on the CLIC 2020 test set. For video compression, GLC-video achieves 65.3% bitrate saving over PLVC in terms of DISTS. Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Neural Image Compression with Regional DecodingabstractAs advancements are made in technology such as AR/VR and high-resolution photography, there is a growing need for a function in image compression named regional decoding . This function lets an image be encoded as a whole, but allows for an arbitrary region to be decoded using only a small part of the bitstream. However, existing neural image compression methods lack support for this crucial functionality. In this article, we propose a novel approach called the slicing en/decoder , which addresses the need for regional decoding while maintaining performance on par with state-of-the-art methods. Our approach is based on the insight that, during the compression process, local information within pixels holds greater importance than global information. By leveraging this understanding, we divide the image into different bitstreams according to cross-boundary patterns. Consequently, for a selected region, our method can intelligently choose specific portions of the bitstreams to decode only that particular region of interest. Furthermore, we extend the application of our method to 360° image compression, allowing for efficient encoding and decoding of immersive visual content. Moreover, our proposed technique offers the capability to decode regions identically, which paves the way for future advancements in regional video decoding. Our experimental results demonstrate that our method maintains performance on par with state-of-the-art methods while providing the functionality of regional decoding . In conclusion, this article presents a significant step forward in image compression technology, offering enhanced flexibility and efficiency for emerging applications in digital media. Yili Jin 0001, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Arbitrary-Scale Video Super-resolution Guided by Dynamic ContextabstractWe propose a Dynamic Context-Guided Upsampling (DCGU) module for video super-resolution (VSR) that leverages temporal context guidance to achieve efficient and effective arbitrary-scale VSR. While most VSR research focuses on backbone design, the importance of the upsampling part is often overlooked. Existing methods rely on pixelshuffle-based upsampling, which has limited capabilities in handling arbitrary upsampling scales. Recent attempts to replace pixelshuffle-based modules with implicit neural function-based and filter-based approaches suffer from slow inference speeds and limited representation capacity, respectively. To overcome these limitations, our DCGU module predicts non-local sampling locations and content-dependent filter weights, enabling efficient and effective arbitrary-scale VSR. Our proposed multi-granularity location search module efficiently identifies non-local sampling locations across the entire low-resolution grid, and the temporal bilateral filter modulation module integrates content information with the filter weight to enhance textual details. Extensive experiments demonstrate the superiority of our method in terms of performance and speed on arbitrary-scale VSR. Jiahao Li 0001, Dong Liu 0002, Yan Lu 0001 |
AAAI | 2 |
| 2024 | Implicit Motion FunctionabstractRecent advancements in video modeling extensively rely on optical flow to represent the relationships across frames, but this approach often lacks efficiency and fails to model the probability of the intrinsic motion of objects. In addition, conventional encoder-decoder frameworks in video processing focus on modeling the correlation in the encoder, leading to limited generative capabilities and redundant intermediate representations. To address these challenges, this paper proposes a novel Implicit Motion Function (IMF) method. Our approach utilizes a low-dimensional latent token as the implicit representation, along with the use of cross-attention, to implicitly model the correlation between frames. This enables the implicit modeling of temporal correlations and understanding of object motions. Our method not only improves sparsity and efficiency in representation but also explores the generative capabilities of the decoder by integrating correlation modeling within it. The IMF framework facilitates video editing and other generative tasks by allowing the direct manipulation of latent tokens. We validate the effectiveness of IMF through extensive experiments on multiple video tasks, demonstrating superior performance in terms of reconstructed video quality, compression efficiency and generation ability. Jiahao Li 0001, Yan Lu 0001 |
CVPR | 2 |
| 2024 | Generative Latent Coding for Ultra-Low Bitrate Image CompressionabstractMost existing image compression approaches perform transform coding in the pixel space to reduce its spatial re-dundancy. However, they encounter difficulties in achieving both high-realism and high-fidelity at low bitrate, as the pixel-space distortion may not align with human perception. To address this issue, we introduce a Generative Latent Coding (GLC) architecture, which performs transform coding in the latent space of a generative vector-quantized variational auto-encoder (VQ- VAE), instead of in the pixel space. The generative latent space is characterized by greater sparsity, richer semantic and better alignment with human perception, rendering it advantageous for achieving high-realism and high-fidelity compression. Additionally, we introduce a categorical hyper module to reduce the bit cost of hyper-information, and a code-prediction-based su-pervision to enhance the semantic consistency. Experiments demonstrate that our GLC maintains high visual quality with less than 0.04 bpp on natural images and less than 0.01 bpp on facial images. On the CLIC2020 test set, we achieve the same FID as MS-ILLM with 45% fewer bits. Furthermore, the powerful generative latent space enables various applications built on our GLC pipeline, such as image restoration and style transfer. Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001 |
CVPR | 2 |
| 2024 | Hierarchical Intra-Modal Correlation Learning for Label-Free 3D Semantic SegmentationabstractRecent methods for label-free 3D semantic segmentation aim to assist 3D model training by leveraging the open-world recognition ability of pre-trained vision language models. However, these methods usually suffer from in-consistent and noisy pseudo-labels provided by the vision language models. To address this issue, we present a hierarchical intra-modal correlation learning framework that captures visual and geometric correlations in 3D scenes at three levels: intra-set, intra-scene, and inter-scene, to help learn more compact 3D representations. We refine pseudo-labels using intra-set correlations within each geometric consistency set and align features of visually and geometrically similar points using intra-scene and inter-scene correlation learning. We also introduce a feedback mechanism to distill the correlation learning capability into the 3D model. Experiments on both indoor and outdoor datasets show the superiority of our method. We achieve a state-of-the-art 36.6% mIoU on the ScanNet dataset, and a 23.0% mIoU on the nuScenes dataset, with improvements of 7.8% mIoU and 2.2% mIoU compared with previous SOTA. We also provide theoretical analysis and qualitative visualization results to discuss the mechanism and conduct thorough ablation studies to support the effectiveness of our framework. Jiahao Li 0001, Xuejin Chen, Yan Lu 0001 |
CVPR | 3 |
| 2024 | Neural Video Compression with Feature ModulationabstractThe emerging conditional coding-based neural video codec (NVC) shows superiority over commonly-used resid-ual coding-based codec and the latest NVC already claims to outperform the best traditional codec. However, there still exist critical problems blocking the practicality of NVC. In this paper, we propose a powerful conditional coding- based NVC that solves two critical problems via feature modulation. The first is how to support a wide quality range in a single model. Previous NVC with this capability only supports about 3.8 dB PSNR range on average. To tackle this limitation, we modulate the latent feature of the cur-rent frame via the learnable quantization scaler. During the training, we specially design the uniform quantization pa-rameter sampling mechanism to improve the harmonization of encoding and quantization. This results in a better learning of the quantization scaler and helps our NVC support about 11.4 dB PSNR range. The second is how to make NVC still work under a long prediction chain. We expose that the previous SOTA NVC has an obvious quality degra-dation problem when using a large intra-period setting. To this end, we propose modulating the temporal feature with a periodically refreshing mechanism to boost the quality. Notably, under single intra-frame setting, our codec can achieve 29.7% bitrate saving over previous SOTA NVC with 16% MACs reduction. Our codec serves as a notable land-mark in the journey of NVC evolution. The codes are at https://github.com/microsoft/DCVC. Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
CVPR | 1 |
| 2024 | Long-Term Temporal Context Gathering for Neural Video Compression
Zhaoyang Jia, Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001 |
ECCV (66) | 3 |
| 2024 | Uncertainty-Aware Deep Video Compression With EnsemblesabstractDeep learning-based video compression is a challenging task, and many previous state-of-the-art learning-based video codecs use optical flows to exploit the temporal correlation between successive frames and then compress the residual error. Although these two-stage models are end-to-end optimized, the epistemic uncertainty in the motion estimation and the aleatoric uncertainty from the quantization operation lead to errors in the intermediate representations and introduce artifacts in the reconstructed frames. This inherent flaw limits the potential for higher bit rate savings. To address this issue, we propose an uncertainty-aware video compression model that can effectively capture the predictive uncertainty with deep ensembles. Additionally, we introduce an ensemble-aware loss to encourage the diversity among ensemble members and investigate the benefits of incorporating adversarial training in the video compression task. Experimental results on 1080p sequences show that our model can effectively save bits by more than 20% compared to DVC Pro. Wufei Ma, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | A Universal Optimization Framework for Learning-based Image CodecabstractRecently, machine learning-based image compression has attracted increasing interests and is approaching the state-of-the-art compression ratio. But unlike traditional codec, it lacks a universal optimization method to seek efficient representation for different images. In this paper, we develop a plug-and-play optimization framework for seeking higher compression ratio, which can be flexibly applied to existing and potential future compression networks. To make the latent representation more efficient, we propose a novel latent optimization algorithm to adaptively remove the redundancy for each image. Additionally, inspired by the potential of side information for traditional codecs, we introduce side information into our framework, and integrate side information optimization with latent optimization to further enhance the compression ratio. In particular, with the joint side information and latent optimization, we can achieve fine rate control using only single model instead of training different models for different rate-distortion trade-offs, which significantly reduces the training and storage cost to support multiple bit rates. Experimental results demonstrate that our proposed framework can remarkably boost the machine learning-based compression ratio, achieving more than 10% additional bit rate saving on three different representative network structures. With the proposed optimization framework, we can achieve 7.6% bit rate saving against the latest traditional coding standard VVC on Kodak dataset, yielding the state-of-the-art compression ratio. Jing Zhao 0011, Bin Li 0012, Jiahao Li 0001, Ruiqin Xiong, Yan Lu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Neural Video Compression with Diverse ContextsabstractFor any video codecs, the coding efficiency highly relies on whether the current signal to be encoded can find the relevant contexts from the previous reconstructed signals. Traditional codec has verified more contexts bring substantial coding gain, but in a time-consuming manner. However, for the emerging neural video codec (NVC), its contexts are still limited, leading to low compression ratio. To boost NVC, this paper proposes increasing the context diversity in both temporal and spatial dimensions. First, we guide the model to learn hierarchical quality patterns across frames, which enriches long-term and yet highquality temporal contexts. Furthermore, to tap the potential of optical flow-based coding framework, we introduce a group-based offset diversity where the cross-group interaction is proposed for better context mining. In addition, this paper also adopts a quadtree-based partition to increase spatial context diversity when encoding the latent representation in parallel. Experiments show that our codec obtains 23.5% bitrate saving over previous SOTA NVC. Better yet, our codec has surpassed the under-developing next generation traditional codec/ECM in both RGB and YUV420 colorspaces, in terms of PSNR. The codes are at https://github.com/microsoft/DCVC. Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
CVPR | 1 |
| 2023 | Motion Information Propagation for Neural Video CompressionabstractIn most existing neural video codecs, the information flow therein is uni-directional, where only motion coding provides motion vectors for frame coding. In this paper, we argue that, through information interactions, the synergy between motion coding and frame coding can be achieved. We effectively introduce bi-directional information interactions between motion coding and frame coding via our Motion Information Propagation. When generating the temporal contexts for frame coding, the high-dimension motion feature from the motion decoder serves as motion guidance to mitigate the alignment errors. Meanwhile, besides assisting frame coding at the current time step, the feature from context generation will be propagated as motion condition when coding the subsequent motion latent. Through the cycle of such interactions, feature propagation on motion coding is built, strengthening the capacity of exploiting long-range temporal correlation. In addition, we propose hybrid context generation to exploit the multiscale context features and provide better motion condition. Experiments show that our method can achieve 12.9% bit rate saving over the previous SOTA neural video codec. Jiahao Li 0001, Bin Li 0012, Houqiang Li, Yan Lu 0001 |
CVPR | 2 |
| 2023 | VideoTrack: Learning to Track Objects via Video TransformerabstractExisting Siamese tracking methods, which are built on pair-wise matching between two single frames, heavily rely on additional sophisticated mechanism to exploit temporal information among successive video frames, hindering them from efficiency and industrial deployments. In this work, we resort to sequence-level target matching that can encode temporal contexts into the spatial features through a neat feedforward video model. Specifically, we adapt the standard video transformer architecture to visual tracking by enabling spatiotemporal feature learning directly from frame-level patch sequences. To better adapt to the tracking task, we carefully blend the spatiotemporal information in the video clips through sequential multi-branch triplet blocks, which formulates a video transformer backbone. Our experimental study compares different model variants, such as tokenization strategies, hierarchical structures, and video attention schemes. Then, we propose a disentangled dual-template mechanism that decouples static and dynamic appearance clues over time, and reduces temporal redundancy in video frames. Extensive experiments show that our method, named as Video Track, achieves state-of-the-art results while running in real-time. Jiahao Li 0001, Yan Lu 0001, Chao Ma 0004 |
CVPR | 3 |
| 2023 | EVC: Towards Real-Time Neural Image Compression with Mask Decay
Guo-Hua Wang, Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
ICLR | 2 |
| 2023 | Disentangle Propagation and Restoration for Efficient Video RecoveryabstractWe propose the first framework for accelerating video recovery, which aims to efficiently recover high-quality videos from degraded inputs affected by various deteriorative factors. Although current video recovery methods have achieved excellent performance, their significant computational overhead limits their widespread application. To address this, we present a pioneering study on explicitly disentangling temporal and spatial redundant computation by decomposing the input frame into propagation and restoration regions, thereby achieving significant computational reduction. Specifically, we leverage contrastive learning to learn degradation-invariant features, which overcomes the disturbance of deteriorative factors and enables accurate disentanglement. For the propagation region, we introduce a split-fusion block to address inter-frame variations, efficiently generating high-quality output at a low cost and significantly reducing temporal redundant computation. For the restoration region, we propose an efficient adaptive halting mechanism that requires few extra parameters and can adaptively halt the patch processing, considerably reducing spatial redundant computation. Furthermore, we design patch-adaptive prior regularization to boost efficiency and performance. Our proposed method achieves outstanding results on various video recovery tasks, such as video denoising, video deraining, video dehazing, and video super-resolution, with a 50% ~ 60% reduction in GMAC over the state-of-the-art video recovery methods while maintaining comparable performance. Jiahao Li 0001, Dong Liu 0002, Yan Lu 0001 |
ACM Multimedia | 2 |
| 2023 | Temporal Context Mining for Learned Video CompressionabstractApplying deep learning to video compression has attracted increasing attention in recent few years. In this work, we address end-to-end learned video compression with a special focus on better learning and utilizing temporal contexts. We propose to propagate not only the last reconstructed frame but also the feature before obtaining the reconstructed frame for temporal context mining. From the propagated feature, we learn multi-scale temporal contexts and re-fill the learned temporal contexts into the modules of our compression scheme, including the contextual encoder-decoder, the frame generator, and the temporal context encoder. We discard the parallelization-unfriendly auto-regressive entropy model to pursue a more practical encoding and decoding time. Experimental results show that our proposed scheme achieves a higher compression ratio than the existing learned video codecs. Our scheme also outperforms x264 and x265 (representing industrial software for H.264 and H.265, respectively) as well as the official reference software for H.264, H.265, and H.266 (JM, HM, and VTM, respectively). Specifically, when intra period is 32 and oriented to PSNR, our scheme outperforms H.265–HM by 14.4% bit rate saving; when oriented to MS-SSIM, our scheme outperforms H.266–VTM by 21.1% bit rate saving. Xihua Sheng, Jiahao Li 0001, Bin Li 0012, Li Li 0040, Dong Liu 0002, Yan Lu 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Neural Compression-Based Feature Learning for Video RestorationabstractHow to efficiently utilize the temporal features is crucial, yet challenging, for video restoration. The temporal features usually contain various noisy and uncorrelated information, and they may interfere with the restoration of the current frame. This paper proposes learning noiserobust feature representations to help video restoration. We are inspired by that the neural codec is a natural denoiser: In neural codec, the noisy and uncorrelated contents which are hard to predict but cost lots of bits are more inclined to be discarded for bitrate saving. Therefore, we design a neural compression module to filter the noise and keep the most useful information in features for video restoration. To achieve robustness to noise, our compression module adopts a spatial-channel-wise quantization mechanism to adaptively determine the quantization step size for each position in the latent. Experiments show that our method can significantly boost the performance on video denoising, where we obtain 0.13 dB improvement over BasicVSR++ with only 0.23x FLOPs. Meanwhile, our method also obtains SOTA results on video deraining and dehazing. Jiahao Li 0001, Bin Li 0012, Dong Liu 0002, Yan Lu 0001 |
CVPR | 2 |
| 2022 | Hybrid Spatial-Temporal Entropy Modelling for Neural Video CompressionabstractFor neural video codec, it is critical, yet challenging, to design an efficient entropy model which can accurately predict the probability distribution of the quantized latent representation. However, most existing video codecs directly use the ready-made entropy model from image codec to encode the residual or motion, and do not fully leverage the spatial-temporal characteristics in video. To this end, this paper proposes a powerful entropy model which efficiently captures both spatial and temporal dependencies. In particular, we introduce the latent prior which exploits the correlation among the latent representation to squeeze the temporal redundancy. Meanwhile, the dual spatial prior is proposed to reduce the spatial redundancy in a parallel-friendly manner. In addition, our entropy model is also versatile. Besides estimating the probability distribution, our entropy model also generates the quantization step at spatial-channel-wise. This content-adaptive quantization mechanism not only helps our codec achieve the smooth rate adjustment in single model but also improves the final rate-distortion performance by dynamic bit allocation. Experimental results show that, powered by the proposed entropy model, our neural codec can achieve 18.2% bitrate saving on UVG dataset when compared with H.266 (VTM) using the highest compression ratio configuration. It makes a new milestone in the development of neural video codec. The codes are at https://github.com/microsoft/DCVC. Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
ACM Multimedia | 1 |
| 2021 | Deep Contextual Video CompressionabstractMost of the existing neural video compression methods adopt the predictive coding framework, which first generates the predicted frame and then encodes its residue with the current frame. However, as for compression ratio, predictive coding is only a sub-optimal solution as it uses simple subtraction operation to remove the redundancy across frames. In this paper, we propose a deep contextual video compression framework to enable a paradigm shift from predictive coding to conditional coding. In particular, we try to answer the following questions: how to define, use, and learn condition under a deep video compression framework. To tap the potential of conditional coding, we propose using feature domain context as condition. This enables us to leverage the high dimension context to carry rich information to both the encoder and the decoder, which helps reconstruct the high-frequency contents for higher video quality. Our framework is also extensible, in which the condition can be flexibly designed. Experiments show that our method can significantly outperform the previous state-of-the-art (SOTA) deep video compression methods. When compared with x265 using veryslow preset, we can achieve 26.0% bitrate saving for 1080P standard test videos. Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
NeurIPS | 1 |
| 2021 | A Deep Reinforcement Learning Approach to Multiple Streams' Joint Bitrate AllocationabstractFor widely used real-time applications, encoding and transmitting multiple videos jointly over a limited bandwidth has become a popular topic. Allocating different bitrates for different sources is a better way to meet different demands from applications. In this paper, we focus on providing equal quality to users by minimizing the variance of distortion among sequences, which is denoted as the minVAR problem. The state-of-the-art Look-ahead and Feed-back Allocation Model (LFAM) allocates bitrate by taking both look-ahead complexity measures and feed-back information into consideration. However, LFAM brings additional delay to real-time applications. By taking the bitrate allocation problem as a time-series decision making problem, we propose a Deep-Reinforcement-Learning-based approach to allocate bitrate with only feed-back information to solve the two-source minVAR problem. Afterward, we introduce a binary-tree-based hierarchical approach to apply our model to arbitrary number of sources. Tested with the widely used open-source x264 encoder, our approach decreases the variance compared with LFAM in all experiments under two-, three- and four-source scenarios. Furthermore, the proposed approach also outperforms LFAM in the mean quality. The proposed approach is insensitive to the order of sequences and encoders with different complexities, showing its robustness and generalization capability. Jiahao Li 0001, Bin Li 0012, Yan Lu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Intra Block Copy for Screen Content in the Emerging AV1 Video CodecabstractScreen content coding plays an important role in many applications. To meet the growing demands of screen content coding, the emerging AV1 video codec incorporates several coding tools, which are specially designed for screen content utilizing its distinctive characteristics. Among these tools, the intra block copy utilizes the characteristic that repeating patterns frequently occur in screen content. This paper presents the technology of intra block copy in AV1. In particular, to efficiently search the predictor in the reconstructed regions of the current picture, AV1 uses the hash matching method at the encoder side. For the generation of hash table, a bottom-to-up manner is adopted to reduce the redundant computation and then decrease the encoding time. In addition, several constraints are involved to facilitate hardware design. Experimental results demonstrate that the intra block copy in AV1 can bring 27.1% bitrate saving for screen content. When compared with the non hash-based intra block copy, the hash-based method achieves 12.2% bitrate saving. Jiahao Li 0001, Hui Su, Alex Converse, Bin Li 0012, Roger Zhou, Bruce Lin, Jizheng Xu, Yan Lu 0001, Ruiqin Xiong |
DCC | 1 |
| 2018 | Efficient Multiple-Line-Based Intra Prediction for HEVCabstractTraditional intra prediction usually utilizes the nearest reference line to generate the predicted block when considering strong spatial correlation. However, this kind of single-line-based method does not always work well due to at least two issues. One is the incoherence caused by the signal noise or the texture of other objects, where this texture deviates from the inherent texture of the current block. The other reason is that the nearest reference line usually has worse reconstruction quality in block-based video coding. Due to these two issues, this paper proposes an efficient multiple-line-based intra-prediction scheme to improve coding efficiency. Besides the nearest reference line, further reference lines are also utilized. The further reference lines with a relatively higher quality can provide potentially better prediction. At the same time, the residue compensation is introduced to calibrate the prediction of boundary regions in a block when we utilize further reference lines. To speed up the encoding process, this paper designs several fast algorithms. The experimental results show that compared with HM-16.9, the proposed fast search method achieves a 2.0% bit saving on average and up to 3.7% by increasing the encoding time by 112%. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Diversity-Based Reference Picture Management for Low Delay Screen Content CodingabstractScreen content coding plays an important role in many applications. Conventional reference picture management (RPM) strategies developed for natural content may not work well for screen content. This is because many regions in screen content remain static for a long time, causing a lot of repetitive contents to stay in the decoded picture buffer. The repetitive contents are not conducive to inter prediction, but still occupy valuable memory. This paper proposes a diversity-based RPM scheme for screen content coding. The concept of diversity is introduced for the reference picture set (RPS) to help formulate the RPM problem. By maximizing the diversity of RPS, more potentially better predictions are provided. Better compression performance can then be achieved. Meanwhile, the proposed scheme is nonnormative and compatible with existing video coding standards, such as High Efficiency Video Coding. The experimental results show that, for low delay screen content coding, the bit saving of the proposed scheme is 4.9% on average and up to 13.7%, without increasing encoding time. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Fully Connected Network-Based Intra Prediction for Image CodingabstractThis paper proposes a deep learning method for intra prediction. Different from traditional methods utilizing some fixed rules, we propose using a fully connected network to learn an end-to-end mapping from neighboring reconstructed pixels to the current block. In the proposed method, the network is fed by multiple reference lines. Compared with traditional single line-based methods, more contextual information of the current block is utilized. For this reason, the proposed network has the potential to generate better prediction. In addition, the proposed network has good generalization ability on different bitrate settings. The model trained from a specified bitrate setting also works well on other bitrate settings. Experimental results demonstrate the effectiveness of the proposed method. When compared with high efficiency video coding reference software HM-16.9, our network can achieve an average of 3.4% bitrate saving. In particular, the average result of 4K sequences is 4.5% bitrate saving, where the maximum one is 7.4%. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong, Wen Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Intra Prediction Using Multiple Reference Lines for Video CodingabstractTraditional intra prediction schemes usually only use the nearest adjacent reference line to generate the prediction. Although the nearest reference line generally has the strongest statistical correlation with current block, the farther non-adjacent reference lines can still provide potential better prediction in some cases. Thus, in this paper, not only the nearest reference line but also the farther reference lines are utilized to help intra prediction. When using the farther reference lines, an additional residue compensation procedure is introduced to further refine the prediction. In particular, this paper designs three solutions to meet different complexity requirements. They are multiple line-based intra prediction (MLIP), fast search for multiple line-based intra prediction (FS-MLIP), and dual line-based intra prediction (DLIP). Experimental results verify the effectiveness of the proposed methods. When compared with HM-16.9, the proposed MLIP achieves 2.4% bit saving on average with the encoding time increasing about 362%. The FS-MLIP achieves 2.0% bit saving on average with the encoding time increasing about 114%. The DLIP achieves 0.9% bit saving on average with the encoding time increasing about only 15%. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
DCC | 1 |
| 2017 | Intra prediction using fully connected network for video codingabstractTraditional intra prediction methods exploit some fixed rules to generate prediction, which might not be adaptive enough to handle complicated contents. In this paper, we investigate applying deep neural network to improve the state-of-the-art intra prediction. Considering the characteristics of block-based video coding framework, we propose a fully connected network for intra prediction where all layers except non-linear ones are fully connected. In the proposed network, the inputs are multiple reference lines of the current block and the output is the prediction for the block. When compared with the traditional intra prediction method, the richer context of current block is exploited. For this reason, the proposed network is capable of providing more accurate prediction. Experimental results demonstrate the effectiveness of proposed network. When integrated into the HEVC reference software, the proposed method can achieve up to 3.3% bitrate saving and an average of 1.6% bitrate saving for 4K sequences. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
ICIP | 1 |
| 2015 | An adaptive hierarchical QP setting for screen content codingabstractScreen content refers to computer generated content like text, graphics, and animations. In such video, many regions may remain static for a long period after a sudden change. Traditional hierarchical Quantization Parameter (QP) setting may not be able to handle these regions efficiently because the encoder probably needs to refine the quality of these static regions multiple times. It will cost more bits while the quality of the static regions may reach the expected degree which the flat QP setting is able to achieve. This paper proposes using different QP settings for different regions in a picture. Region classification algorithms are developed to determine whether a flat or hierarchical QP setting is used. Experimental results demonstrate that the proposed scheme can achieve an average bitrate reduction of 3.1%, and up to 8.1% bitrate reduction for IBBB coding. The proposed method improves coding efficiency without increasing encoding complexity. Jiahao Li 0001, Bin Li 0012, Jizheng Xu, Ruiqin Xiong |
VCIP | 1 |