Qi Zhang 0042

dblp:52/323-42 · DBLP profile ↗
← Back
2ranked-venue papers in the field
0as first author
2since 2021 · last 2025
0000-0002-1189-8755ORCID · conflict

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2025 Point Cloud-Assisted Neural Image Compression
abstract
High-efficient image compression is a critical requirement. In several scenarios where multiple modalities of data are captured by different sensors, the auxiliary information from other modalities are not fully leveraged by existing image-only codecs, leading to suboptimal compression efficiency. In this paper, we increase image compression performance with the assistance of point cloud, which is widely adopted in the area of autonomous driving. As depicted in Figure 1 (a), we have unified the digital representation of image and point cloud, and propose the point cloud-assisted neural image codec (PCA-NIC) to enhance the preservation of image texture and structure by utilizing the high-dimensional point cloud information. As depicted in Figure 1 (b), we further introduce a multi-modal feature fusion transform module (MMFFT) to capture more representative image features, remove redundant information between channels and modalities that are not relevant to the image content.
Ziqun Li, Qi Zhang 0042, Xiaofeng Huang, Zhao Wang 0004, Siwei Ma 0001
DCC2
2025 LL-ICM: Image Compression for Low-Level Machine Vision via Large Vision-Language Model
abstract
Image Compression for Machines (ICM) aims to compress images for machine vision tasks, while current methods mostly focus on the demands for high-level tasks. However, the quality of original images is usually not guaranteed in the real world, leading to even worse downstream task performance after compression. Thus, lowlevel (LL) restoration tasks should also be considered in ICM. In this paper, we propose the first ICM framework for LL machine vision tasks, namely LL-ICM, which optimizes the compression and LL processing performance simultaneously. Moreover, LL-ICM leverages large vision-language model (VLM) to solve different LL task within a single model, which is particularly useful when the distortion type of the original image is uncertain. As illustrated in Fig. 1(a), LL-ICM consists of a neural image codec and a VLM-based LL processing module. Given an original image with distortions, LL-ICM firstly compress it as$\hat{\mathbf{X}}$. Then, we extract a generalized feature F from$\hat{\mathbf{X}}$, which is then encoded as two representations, distortion type$\varphi$and caption$\sigma$. After that, the LL processing module receives$\hat{\mathbf{X}}$and its representations to generate the restored version of$\hat{\mathbf{X}}$, i.e.,$\hat{\mathbf{X}}_{\mathbf{H}}$.
Qi Zhang 0042, Chuanmin Jia, Shiqi Wang 0001
DCC2