Qi Mao 0002

dblp:78/9363-2 · DBLP profile ↗
← Back
25ranked-venue papers
11as first author
18since 2021 · last 2026
0000-0001-9362-6237ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
abstract
Large language models, trained on extensive corpora, successfully unify diverse linguistic tasks within a single generative framework. Inspired by this, recent works like Large Vision Model (LVM) extend this paradigm to vision by organizing tasks into sequential visual sentences, where visual prompts serve as the context to guide outputs. However, such modeling requires task-specific pre-training across modalities and sources, which is costly and limits scalability to unseen tasks. Given that pre-trained video generation models inherently capture temporal sequence dependencies, we explore a more unified and scalable alternative: can a pre-trained video generation model adapt to diverse image and video tasks? To answer this, we propose UniVid, a framework that fine-tunes a video diffusion transformer to handle various vision tasks without task-specific modifications. Tasks are represented as visual sentences, where the context sequence defines both the task and the expected output modality. We evaluate the generalization of UniVid from two perspectives: (1) cross-modal inference with contexts composed of both images and videos, extending beyond LVM’s uni-modal setting; (2) cross-source tasks from natural to annotated data, without multi-source pre-training. Despite being trained solely on natural video data, UniVid generalizes well in both settings. Notably, understanding and generation tasks can easily switch by simply reversing the visual sentence order in this paradigm. These findings highlight the potential of pre-trained video generation models to serve as a scalable and unified foundation for vision modeling. Our code is released at https://github.com/CUC-MIPG/UniVid.
Yuchao Gu, Qi Mao 0002
WACV3
2026 EmoAgent: A Multi-Agent Framework for Diverse Affective Image Manipulation
abstract
Affective Image Manipulation (AIM) aims to alter visual elements within an image to evoke specific emotional responses from viewers. However, existing AIM approaches rely on rigidone-to-onemappings between emotions and visual cues, making them ill-suited for the inherently subjective and diverse ways in which humans perceive and express emotion. To address this, we introduce a novel task setting termedDiverse AIM (D-AIM), aiming to generate multiple visually distinct yet emotionally consistent image edits from a single source image and target emotion. We proposeEmoAgent, the first multi-agent framework tailored specifically for D-AIM. EmoAgent explicitly decomposes the manipulation process into three specialized phases executed by collaborative agents: a Planning Agent that generates diverse emotional editing strategies, an Editing Agent that precisely executes these strategies, and a Critic Agent that iteratively refines the results to ensure emotional accuracy. This collaborative design empowers EmoAgent to modelone-to-manyemotion-to-visual mappings, enabling semantically diverse and emotionally faithful edits. Extensive quantitative and qualitative evaluations demonstrate that EmoAgent substantially outperforms state-of-the-art approaches in both emotional fidelity and semantic diversity, effectively generating multiple distinct visual edits that convey the same target emotion.
Qi Mao 0002, Haobo Hu, Yujie She, Difei Gao, Libiao Jin
IEEE Trans. Affect. Comput.1
2026 HD-Custom: Efficient Hierarchical Disentanglement for Coarse-to-Fine Concept Customization in Subject Video Generation
Yuanhang Li, Qi Mao 0002, Xinyan Xiao, Libiao Jin, Siwei Ma 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 Stable Diffusion is a Natural Cross-Modal Decoder for Layered AI-Generated Image Compression
abstract
Recent advances in Artificial Intelligence Generated Content (AIGC) triggered an increasing need to transmit and compress the vast number of AI-generated images (AIGIs). However, there is a noticeable deficiency in research focused on compression methods for AIGIs. To address this critical gap, we advocate that Stable Diffusion serves as a natural cross-modal decoder by leveraging rich and scalable priors, and introduce a scalable cross-modal compression framework that incorporates multiple human-comprehensible modalities. As illustrated in Fig. 1(a), the proposed framework encodes images into a layered bitstream: a semantic prior that delivers high-level semantic information through text prompts; a structural prior that captures spatial details using edge or skeleton maps; and a texture prior that preserves local textures via a colormap. Utilizing Stable Diffusion as the backend, the decoder leverages multi-modal scalable priors to generate images with different levels of fidelity. Experiments show our method preserves realistic details and semantic fidelity at an extremely low bitrate (< 0.02 bpp), comparable with recent perceptual coding approaches and outperforming VVC. The R-D performance also demonstrate the scalability of our proposed multi-layered bitstream since image fidelity incrementally improves with structure and texture priors provided during decoding. Additionally, as illustrated in Fig. 1(b), our framework facilitates downstream editing applications such as Structure Manipulation, Texture Synthesis, and Object Erasing, without requiring full decoding, thereby paving a new direction for future research in AIGI compression.
Ruijie Chen, Qi Mao 0002, Zhengxue Cheng
DCC2
2025 Exploring Multimodal Knowledge for Image Compression via Large Foundation Models
abstract
Knowledge is an abstraction of factual principles of the physical world. Large foundation models encapsulate extensive multimodal knowledge into the parameters and thus invoke machine intelligence on various tasks. How to invoke the knowledge in these models to facilitate image compression lacks in-depth exploration. In this work, we aim to harness multimodal knowledge into ultra-low bitrate compression and propose Multimodal Knowledge-aware Image Compression (MKIC). Our key insight is that under the context of ultra-low bitrate compression, where the encoded representation is too sparse to represent enough information of the input signal, knowledge from the physical world is required to be incorporated into the compression. Thus, more shared patterns can be stored in the model together with sparse unique features also embedded into the bitstream. In light of two kinds of knowledge, namely natural visual knowledge and human language knowledge, we propose a novel Alternating Rate-Distortion Optimization to enhance the accuracy and compactness of global semantic text representation extraction, extract the local feature map that captures visual details, and integrate these multimodal representations into a large generative foundation model to achieve high-quality reconstruction. The proposed method relights the path of learned image coding, leveraging decoupled knowledge from large foundation models. Extensive experiments show that our proposed method achieves superior comprehensive performance compared to various methods and shows great potential for ultra-low bitrate image compression.
Junlong Gao, Zhimeng Huang, Qi Mao 0002, Siwei Ma 0001, Chuanmin Jia
IEEE Trans. Image Process.3
2024 Extreme Image Compression Using Fine-tuned VQGANs
abstract
Recent advances in generative compression methods have demonstrated remarkable progress in enhancing the perceptual quality of compressed data, especially in scenarios with low bitrates. However, their efficacy and applicability to achieve extreme compression ratios (< 0.05 bpp) remain constrained. In this work, we propose a simple yet effective coding framework by introducing vector quantization (VQ)–based generative models into the image compression domain. The main insight is that the codebook learned by the VQGAN model yields a strong expressive capacity, facilitating efficient compression of continuous information in the latent space while maintaining reconstruction quality. Specifically, an image can be represented as VQ-indices by finding the nearest codeword, which can be encoded using lossless compression methods into bitstreams. We propose clustering a pre-trained large-scale codebook into smaller codebooks through the K-means algorithm, yielding variable bitrates and different levels of reconstruction quality within the coding framework. Furthermore, we introduce a transformer to predict lost indices and restore images in unstable environments. Extensive qualitative and quantitative experiments on various benchmark datasets demonstrate that the proposed framework outperforms state-of-the-art codecs in terms of perceptual quality-oriented metrics and human perception at extremely low bitrates (≤ 0.04 bpp). Remarkably, even with the loss of up to 20% of indices, the images can be effectively restored with minimal perceptual loss.
Qi Mao 0002, Tinghan Yang, Meng Wang 0017, Shiqi Wang 0001, Libiao Jin, Siwei Ma 0001
DCC1
2024 Unrolled Decomposed Unpaired Learning for Controllable Low-Light Video Enhancement
Lingyu Zhu 0006, Wenhan Yang, Baoliang Chen, Hanwei Zhu, Zhangkai Ni, Qi Mao 0002, Shiqi Wang 0001
ECCV (23)6
2024 Learned Image Compression for Both Humans and Machines via Dynamic Adaptation
abstract
Recent advancements in neural image compression have shown great potential in outperforming conventional standard codecs in terms of both rate-distortion and rate-analysis performance. However, there is an issue of divergent preferences in information preservation or reconstruction in the process of compression for humans and machines, respectively. Compression for humans tends to retain the signal fidelity or perceptual quality of visual appearance while compression for machines requires preserving critical semantic information, resulting in the limitation of the bitstream supporting only a single requirement during the compression. To bridge this gap, we propose a dynamic adaptation approach that generates a single bitstream serving both humans and machines. This approach aims to mitigate the domain gap among tasks, which facilitates maintaining the performance of out-of-scope tasks. Specifically, the proposed method concentrates on learning a dynamic adaptation process, i.e., optimizing the latent representation in the compressed domain in an end-to-end manner while adhering to the rate-performance constraint. Extensive results reveal that our paradigm significantly reduces the domain gap, surpassing existing codecs.
Lingyu Zhu 0006, Binzhe Li, Riyu Lu, Peilin Chen 0001, Qi Mao 0002, Zhao Wang 0004, Wenhan Yang, Shiqi Wang 0001
ICIP5
2024 Unifying Generation and Compression: Ultra-low bitrate Image Coding Via Multi-stage Transformer
abstract
Recent progress in generative compression technology has significantly improved the perceptual quality of compressed data. However, these advancements primarily focus on producing high-frequency details, often overlooking the ability of generative models to capture the prior distribution of image content, thus impeding further bitrate reduction in extreme compression scenarios (< 0.05 bpp). Motivated by the capabilities of predictive language models for lossless compression, this paper introduces a novel Unified Image Generation-Compression (UIGC) paradigm, merging the processes of generation and compression. A key feature of the UIGC framework is the adoption of vector-quantized (VQ) image models for tokenization, alongside a multi-stage transformer designed to exploit spatial contextual information for modeling the prior distribution. As such, the dual-purpose framework effectively utilizes the learned prior for entropy estimation and assists in the regeneration of lost tokens. Extensive experiments demonstrate the superiority of the proposed UIGC framework over existing codecs in perceptual quality and human perception, particularly in ultra-low bitrate scenarios (≤0.03 bpp), pioneering a new direction in generative compression.
Naifu Xue, Qi Mao 0002, Yuan Zhang 0013, Siwei Ma 0001
ICME2
2024 MAG-Edit: Localized Image Editing in Complex Scenarios via Mask-Based Attention-Adjusted Guidance
Qi Mao 0002, Yuchao Gu, Zheng Shou 0001
ACM Multimedia1
2024 Scalable Face Image Coding via StyleGAN Prior: Toward Compression for Human-Machine Collaborative Vision
abstract
The accelerated proliferation of visual content and the rapid development of machine vision technologies bring significant challenges in delivering visual data on a gigantic scale, which shall be effectively represented to satisfy both human and machine requirements. In this work, we investigate how hierarchical representations derived from the advanced generative prior facilitate constructing an efficient scalable coding paradigm for human-machine collaborative vision. Our key insight is that by exploiting the StyleGAN prior, we can learn three-layered representations encoding hierarchical semantics, which are elaborately designed into the basic, middle, and enhanced layers, supporting machine intelligence and human visual perception in a progressive fashion. With the aim of achieving efficient compression, we propose the layer-wise scalable entropy transformer to reduce the redundancy between layers. Based on the multi-task scalable rate-distortion objective, the proposed scheme is jointly optimized to achieve optimal machine analysis performance, human perception experience, and compression ratio. We validate the proposed paradigm's feasibility in face image compression. Extensive qualitative and quantitative experimental results demonstrate the superiority of the proposed paradigm over the latest compression standard Versatile Video Coding (VVC) in terms of both machine analysis as well as human perception at extremely low bitrates (< 0.01 bpp), offering new insights for human-machine collaborative compression.
Qi Mao 0002, Chongyu Wang, Meng Wang 0017, Shiqi Wang 0001, Ruijie Chen, Libiao Jin, Siwei Ma 0001
IEEE Trans. Image Process.1
2023 Extreme Generative Human-Oriented Video Coding via Motion Representation Compression
abstract
The increasing popularity of video conferencing and live streaming raises the growing demand for encoding human-oriented videos at ultra-low bit rates. Recently, several ultra-low bitrate video codecs have proposed using inter-frame keypoints or landmarks to derive motion representations, which are then used to warp decoded frames in a generative manner. Despite its success, compression of the motion representation has been less investigated in the literature. In this work, we propose a novel principal component analysis (PCA)-based decomposing method to fully exploit the compression potential of motion representations. In particular, we decompose the derived motion affine matrices into three parts and apply quantization and entropy estimation to each part in a different way depending on its significance. Using such compressed-friendly motion representations allows for preserving most of the motion information and achieving lower coding costs. Extensive qualitatively and quantitatively experimental results on the human video datasets demonstrate the superiority of the proposed paradigm over existing video codecs under extreme compression ratios.
Qi Mao 0002, Chuanmin Jia, Ronggang Wang, Siwei Ma 0001
ISCAS2
2023 ZGaming: Zero-Latency 3D Cloud Gaming by Image Prediction
abstract
In cloud gaming, interactive latency is one of the most important factors in users' experience. Although the interactive latency can be reduced through typical network infrastructures like edge caching and congestion control, the interactive latency of current cloud-gaming platforms is still far from users' satisfaction.
Jiangkai Wu, Yu Guan 0005, Qi Mao 0002, Yong Cui 0001, Zongming Guo, Xinggong Zhang
SIGCOMM3
2023 Semantic-Aware Visual Decomposition for Image Coding
Jianhui Chang, Jian Zhang 0018, Jiguo Li 0002, Shiqi Wang 0001, Qi Mao 0002, Chuanmin Jia, Siwei Ma 0001, Wen Gao 0001
Int. J. Comput. Vis.5
2023 Enhancing Style-Guided Image-to-Image Translation via Self-Supervised Metric Learning
abstract
There has been significant success in recent image-to-image translation (I2I) approaches in translating the source image into the style of the target image. Existing techniques rely on the disentanglement of content and style representations, requiring a two-stage style mapping process: Reference images are used to extract style vectors, which are subsequently remapped into the translated images. However, when the target domain contains a variety of styles, such a two-stage style mapping cannot guarantee the translated image be style consistent with its guided reference image. In this work, we propose to explicitly employ metric learning to enhance the two-stage style mapping in style-guided image translation. The distance between deep features Gram matrices is utilized to construct the visual style metric as self-supervised similarity labels, guiding the embedding of style vectors using triplet loss with adaptive margins in the first stage. Furthermore, in the second stage, we consider generated images and their corresponding reference images as positive samples and anchors for each other, while the nearest negative sample is used to construct the triplet loss in the proposed metric space. The proposed learning algorithms can be applied to any I2I framework that uses disentangled representations without modifying the original network architectures. We evaluate the proposed method on three representative I2I translation baselines. Both qualitative and quantitative results demonstrate that the proposed approach enhances style alignment in style-guided translation compared to the baselines.
Qi Mao 0002, Siwei Ma 0001
IEEE Trans. Multim.1
2022 Disentangled Visual Representations for Extreme Human Body Video Compression
abstract
Recent years have witnessed the great promise of deep neural video compression codecs. However, there are still unprecedented challenges ahead when the videos are expected to be encoded with extremely low bitrate. Motivated by recent attempts of layered conceptual image compression, we make the first attempt to leverage the disentangled visual representations for extreme human body video compression. More specifically, to capture the main structure, we adopt the inferred human pose keypoints as the structure code of each frame, thereby deriving the motion information from structure codes of adjacent frames for further compression. To better exploit the texture redundancy, all frames share the same texture codes by incorporating the proposed texture contrastive learning to ensure texture consistency within a video. Two branches are consequently transmitted in a separable manner, and the generator synthesizes the reconstructed video with the combination of all decoded representations at the decoder side. Both qualitative and quantitative experimental results demonstrate that the proposed scheme can produce perceptually pleasing reconstruction results in ultra-low bitrates far below that can be reached by other video codecs.
Qi Mao 0002, Shiqi Wang 0001, Chuanmin Jia, Ronggang Wang, Siwei Ma 0001
ICME2
2022 Continuous and Diverse Image-to-Image Translation via Signed Attribute Vectors
Qi Mao 0002, Hung-Yu Tseng, Hsin-Ying Lee 0001, Jia-Bin Huang 0001, Siwei Ma 0001, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.1
2022 Conceptual Compression via Deep Structure and Texture Synthesis
abstract
Existing compression methods typically focus on the removal of signal-level redundancies, while the potential and versatility of decomposing visual data into compact conceptual components still lack further study. To this end, we propose a novel conceptual compression framework that encodes visual data into compact structure and texture representations, then decodes in a deep synthesis fashion, aiming to achieve better visual reconstruction quality, flexible content manipulation, and potential support for various vision tasks. In particular, we propose to compress images by a dual-layered model consisting of two complementary visual features: 1) structure layer represented by structural maps and 2) texture layer characterized by low-dimensional deep representations. At the encoder side, the structural maps and texture representations are individually extracted and compressed, generating the compact, interpretable, inter-operable bitstreams. During the decoding stage, a hierarchical fusion GAN (HF-GAN) is proposed to learn the synthesis paradigm where the textures are rendered into the decoded structural maps, leading to high-quality reconstruction with remarkable visual realism. Extensive experiments on diverse images have demonstrated the superiority of our framework with lower bitrates, higher reconstruction quality, and increased versatility towards visual analysis and content manipulation tasks.
Jianhui Chang, Zhenghui Zhao, Chuanmin Jia, Shiqi Wang 0001, Lingbo Yang, Qi Mao 0002, Jian Zhang 0018, Siwei Ma 0001
IEEE Trans. Image Process.6
2020 DRIT++: Diverse Image-to-Image Translation via Disentangled Representations
Hsin-Ying Lee 0001, Hung-Yu Tseng, Qi Mao 0002, Jia-Bin Huang 0001, Yu-Ding Lu, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001
Int. J. Comput. Vis.3
2019 Mode Seeking Generative Adversarial Networks for Diverse Image Synthesis
abstract
Most conditional generation tasks expect diverse outputs given a single conditional context. However, conditional generative adversarial networks (cGANs) often focus on the prior conditional information and ignore the input noise vectors, which contribute to the output variations. Recent attempts to resolve the mode collapse issue for cGANs are usually task-specific and computationally expensive. In this work, we propose a simple yet effective regularization term to address the mode collapse issue for cGANs. The proposed method explicitly maximizes the ratio of the distance between generated images with respect to the corresponding latent codes, thus encouraging the generators to explore more minor modes during training. This mode seeking regularization term is readily applicable to various conditional generation tasks without imposing training overhead or modifying the original network structures. We validate the proposed algorithm on three conditional image synthesis tasks including categorical generation, image-to-image translation, and text-to-image synthesis with different baseline models. Both qualitative and quantitative results demonstrate the effectiveness of the proposed regularization method for improving diversity without loss of quality.
Qi Mao 0002, Hsin-Ying Lee 0001, Hung-Yu Tseng, Siwei Ma 0001, Ming-Hsuan Yang 0001
CVPR1
2019 Layered Conceptual Image Compression Via Deep Semantic Synthesis
abstract
Motivated by the insight of Marr on generative image representations, we propose a layered conceptual image compression scheme by integrating the advantages of both variational auto-encoders (VAEs) and generative adversarial networks (GANs). In particular, the image is represented by two layers: the low-dimensional codes of the stochastic textures encoded by the VAE and the geometric structures characterized by edge maps. Subsequently, the edge maps and latent codes are compressed individually such that the final bit streams are formed in a combined manner. At the decoder side, the GAN synthesizes the decoded images on the basis of the latent codes and the reconstructed edge maps. Experimental results demonstrate that our proposed scheme achieves better visual reconstruction quality than the traditional image compression algorithms such as JPEG, JPEG2000 and HEVC (intra coding) in the low bit rate coding scenarios.
Jianhui Chang, Qi Mao 0002, Zhenghui Zhao, Shanshe Wang, Shiqi Wang 0001, Siwei Ma 0001
ICIP2
2019 Fidelity or Quality? A Region-Aware Framework for Enhanced Image Decoding via Hybrid Neural Networks
abstract
The generative deep learning models such as the generative adversarial networks (GAN) have been shown to efficiently generate visually appealing images by learning the natural scene statistics. However, the signal fidelity, instead of the visual quality, has been largely ignored in the generation process, especially for the highly structural regions. In this paper, we introduce a region-aware visual signal restoration scheme to achieve a good balance between visual quality and fidelity. As a specific example of this framework, we develop an enhanced decoding scheme with hybrid neural networks, such that the base fidelity layer and texture quality enhancement layer are combined adaptively to restore the compressed images. The efficiency of the proposed framework is demonstrated with extensive experimental results, which show favorable performance against the state-of-the-art methods.
Qi Mao 0002, Shiqi Wang 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001
ICIP1
2018 Enhanced Image Decoding via Edge-Preserving Generative Adversarial Networks
abstract
Lossy image compression usually introduces undesired compression artifacts, such as blocking, ringing and blurry effect{###} S, especially in low bit rate coding scenarios. Although many algorithms have been proposed to reduce these compression artifacts, most of them are based on image local smoothness prior, which usually leads to over-smoothing around the areas with distinct structures, e.g., edges and textures. In this paper, we propose a novel framework to enhance the perceptual quality of decoded images by well preserving the edge structures and predicting visually pleasing textures. Firstly, we propose an edge-preserving generative adversarial network (EP-GAN) to achieve edge restoration and texture generation simultaneously. Then, we elaborately design an edge fidelity regularization term to guide our network, which jointly utilizes the signal fidelity, feature fidelity and adversarial constraint to reconstruct high quality decoded images. Experimental results demonstrate that the proposed EP-GAN is able to efficiently enhance decoded images at low bit rate and reconstruct more perceptually pleasing images with abundant textures and sharp edges.
Qi Mao 0002, Shiqi Wang 0001, Shanshe Wang, Xinfeng Zhang 0001, Siwei Ma 0001
ICME1
2017 Local Disparity Vector Derivation Scheme in 3D-AVS2
Qi Mao 0002, Shanshe Wang, Siwei Ma 0001
ICIG (2)1
2016 A local-adapted disparity vector derivation scheme for 3D-AVS
abstract
In the 3D extension of Audio Video Coding Standard (AVS), i.e. 3D-AVS, the Global Disparity Vector (GDV) derivation technique has been proposed to provide an estimation for Disparity Vector (DV) in inter-view prediction, where the GDV is generated by averaging all Disparity Vectors (DVs) in the latest previously coded frame. The prediction accuracy of GDV may be however limited by the lack of local adaptivity. In this paper, we introduce a novel Local Disparity Vector (LDV) derivation scheme. Specifically, the DV of the current block is calculated from the DVs within a neighboring region, whose size can be adaptively expanded to increase the robustness and accuracy. Experimental results show that the proposed LDV derivation method can provide around 2.12% and 1.37% bitrate reductions for compressed views and synthesized views compared with the GDV scheme, respectively.
Qi Mao 0002, Shanshe Wang, Xiang Zhang 0004, Xinfeng Zhang 0001, Siwei Ma 0001
VCIP1