VLDB 2026 Research / reviewers in the wild / expert
Binzhe Li
dblp:263/4643
· DBLP profile ↗
9ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-1535-0121ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | High Efficiency Image Compression for Large Visual-Language ModelsabstractIn recent years, large visual language models (LVLMs) have shown impressive performance and promising generalization capability in multi-modal tasks, thus replacing humans as receivers of visual information in various application scenarios. In this paper, we pioneer to propose a variable bitrate image compression scheme consisting of a pre-editing module and an end-to-end codec to achieve promising rate-accuracy performance for different LVLMs. In particular, instead of optimizing an adaptive pre-editing network towards a particular task or several representative tasks, we propose a new optimization strategy tailored for LVLMs, which is designed based on the representation and discrimination capability with token-level distortion and rank. The pre-editing module and the variable bitrate end-to-end image codec are jointly trained by the losses based on semantic tokens of the large model, which introduce enhanced generalization capability for various data and tasks. Experimental results demonstrate that the proposed framework could efficiently achieve much better rate-accuracy performance compared to the state-of-the-art coding standard, Versatile Video Coding. Meanwhile, experiments with multi-modal tasks have revealed the robustness and generalization capability of the proposed framework. Binzhe Li, Shurun Wang, Shiqi Wang 0001, Yan Ye 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Interactive Face Video Coding: A Generative Compression FrameworkabstractIn this paper, we propose a novel framework for Interactive Face Video Coding (IFVC), which allows humans to interact with the intrinsic visual representations instead of the signals. The proposed solution enjoys several distinct advantages, including ultra-compact representation, low delay interaction, and vivid expression/headpose animation. In particular, we propose the Internal Dimension Increase (IDI) based representation, greatly enhancing the fidelity and flexibility in rendering the appearance while maintaining reasonable representation cost. By leveraging strong statistical regularities, the visual signals can be effectively projected into controllable semantics in the three dimensional space (e.g., mouth motion, eye blinking, head rotation, head translation and head location), which are compressed and transmitted. The editable bitstream, which naturally supports the interactivity at the semantic level, can synthesize the face frames via the strong inference ability of the deep generative model. Experimental results have demonstrated the performance superiority and application prospects of our proposed IFVC scheme. In particular, the proposed scheme not only outperforms the state-of-the-art video coding standard Versatile Video Coding (VVC) and the latest generative compression schemes in terms of rate-distortion performance for face videos, but also enables the interactive coding without introducing additional manipulation processes. Furthermore, the proposed framework is expected to shed lights on the future design of the digital human communication in the metaverse. The project page can be found at https://github.com/Berlin0610/Interactive_Face_Video_Coding. Zhao Wang 0004, Binzhe Li, Shurun Wang, Shiqi Wang 0001, Yan Ye 0003 |
IEEE Trans. Image Process. | 3 |
| 2024 | Learned Image Compression for Both Humans and Machines via Dynamic AdaptationabstractRecent advancements in neural image compression have shown great potential in outperforming conventional standard codecs in terms of both rate-distortion and rate-analysis performance. However, there is an issue of divergent preferences in information preservation or reconstruction in the process of compression for humans and machines, respectively. Compression for humans tends to retain the signal fidelity or perceptual quality of visual appearance while compression for machines requires preserving critical semantic information, resulting in the limitation of the bitstream supporting only a single requirement during the compression. To bridge this gap, we propose a dynamic adaptation approach that generates a single bitstream serving both humans and machines. This approach aims to mitigate the domain gap among tasks, which facilitates maintaining the performance of out-of-scope tasks. Specifically, the proposed method concentrates on learning a dynamic adaptation process, i.e., optimizing the latent representation in the compressed domain in an end-to-end manner while adhering to the rate-performance constraint. Extensive results reveal that our paradigm significantly reduces the domain gap, surpassing existing codecs. Lingyu Zhu 0006, Binzhe Li, Riyu Lu, Peilin Chen 0001, Qi Mao 0002, Zhao Wang 0004, Wenhan Yang, Shiqi Wang 0001 |
ICIP | 2 |
| 2024 | Quality Harmonization for Virtual Composition in Online Video CommunicationsabstractRecent years have witnessed strong demands for video composition in online video communications, enabling a series of new functionalities for video conferencing including virtual conference rooms, virtual reunions, and virtual backgrounds. In video composition, typically the foreground videos including the human bodies and faces are subject to compression due to the constrained bandwidth, whereas the virtual background is uncompressed and in pristine quality. The disharmony caused by the incoherent quality of foreground and background, which may worsen the quality of experience, has not been extensively studied. In this paper, we focus on this particular problem and present an image quality harmonization framework. Our principle is to align the quality of the background with that of the foreground such that they share similar levels of distortion. This is achieved by inferring the quantization parameter for background compression based on the foreground information. In particular, we aim to learn the quality and compression parameters in a self-supervised manner without laborious human annotation. Furthermore, a large dataset is constructed to provide sufficient training samples and testing scenarios for validation. The composite videos show superior harmonized quality in both quantitative and qualitative comparisons, demonstrating the effectiveness of the proposed framework. Binzhe Li, Zhao Wang 0004, Baoliang Chen, Shiqi Wang 0001, Yan Ye 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Semantic Face Compression for Metaverse: A Compact 3D Descriptor Based ApproachabstractThe metaverse, a 3D virtual world, requires efficient interactive avatar communication. To achieve this goal, we envision a new metaverse paradigm for virtual avatar faces and develop semantic face compression with compact 3D facial descriptors. The paradigm comprises a compression framework that transmits 3D face descriptors for semantic compression and applications based on the semantic descriptors. The fundamental principle is that the communication of virtual avatar faces primarily emphasizes the conveyance of semantic information. In light of this, the proposed scheme offers the advantages of being highly flexible, efficient, and semantically meaningful. The promise of the proposed paradigm is also demonstrated by performance comparisons with the state-of-the-art video coding standard, Versatile Video Coding. A significant improvement in terms of rate-accuracy performance has been achieved. The proposed scheme is expected to enable numerous applications especially for real-time communication in the metaverse, such as digital human communication based on machine analysis, and to form the cornerstone of interactions. Binzhe Li, Zhao Wang 0004, Shiqi Wang 0001, Yan Ye 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Compact Temporal Trajectory Representation for Talking Face Video CompressionabstractIn this paper, we propose to compactly represent the nonlinear dynamics along the temporal trajectories for talking face video compression. By projecting the frames into a high dimensional space, the temporal trajectories of talking face frames, which are complex, non-linear and difficult to extrapolate, are implicitly modelled in an end-to-end inference framework based upon very compact feature representation. As such, the proposed framework is suitable for ultra-low bandwidth video communication and can guarantee the quality of the reconstructed video in such applications. The proposed compression scheme is also robust against large head-pose motions, due to the delicately designed dynamic reference refresh and temporal stabilization mechanisms. Experimental results demonstrate that compared to the state-of-the-art video coding standard Versatile Video Coding (VVC) as well as the latest generative compression schemes, our proposed scheme is superior in terms of both objective and subjective quality at the same bitrate. The project page can be found athttps://github.com/Berlin0610/CTTR. Zhao Wang 0004, Binzhe Li, Shiqi Wang 0001, Yan Ye 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Beyond Keypoint Coding: Temporal Evolution Inference with Compact Feature Representation for Talking Face Video CompressionabstractWe propose a talking face video compression framework by implicitly transforming the temporal evolution into compact feature representation. More specifically, the temporal evolution of faces, which is complex, non-linear and difficult to extrapolate, is modelled in an end-to-end inference framework based upon very compact features. This enables the high-quality rendering of the face videos, which benefits from the learning of dense motion map with compact feature representation. Therefore, the proposed framework can accommodate ultra-low bandwidth video communication and maintain the quality of the reconstructed videos. Experimental results demonstrate that compared with the state-of-the-art video coding standard Versatile Video Coding (VVC) as well as the latest generative compression scheme Face Video-to-Video Synthesis (Face_vid2vid), the proposed scheme is superior in terms of both objective and subjective quality assessment methods. Zhao Wang 0004, Binzhe Li, Rongqun Lin, Shiqi Wang 0001, Yan Ye 0003 |
DCC | 3 |
| 2022 | Towards Ultra Low Bit-Rate Digital Human Character Communication via Compact 3D Face DescriptorsabstractRecently, there has been a tremendous demand for high-efficiency face video communications, coinciding with the popularization of the digital human character in numerous applications. This paper demonstrates a new communication paradigm of 3D human digital characters in ultra low-bit-rate application scenarios. The paradigm is grounded on the mild assumption of the consistency and persistence of human ap-pearance, such that only the compact features that determine the pose and expression of the 3D character need to be transmitted. The proposed is also expected to benefit virtual-physical world interaction in Metaverse. Binzhe Li, Zhao Wang 0004, Shiqi Wang 0001, Yan Ye 0003 |
DCC | 1 |
| 2021 | Defending Against Noise by Characterizing the Rate-Distortion Functions in End-to-End Noisy Image CompressionabstractThere has been an increasing consensus that precise understanding of the rate-distortion (RD) characteristics plays a critical role in image and video coding. In this paper, we explore the RD behaviors of end-to-end image compression in the real-world application scenario that the images could be corrupted by noise at different levels. With the RD behaviors that all images share, we develop a deep learning driven pre-analytical model which fully exploits the properties of RD functions and allows us to improve the quality with economized coding bits. The proposed approach does not require any prior knowledge of the noise level, and could effectively defend against the noise through the end-to-end compression. Extensive experimental results show that the proposed scheme offers the best promise in predicting RD behaviors, and naturally avoids the unnecessary bits consumption. Binzhe Li, Shurun Wang, Shiqi Wang 0001 |
ICIP | 1 |