VLDB 2026 Research / reviewers in the wild / expert
Qi Mao 0002
dblp:78/9363-2
· DBLP profile ↗
25ranked-venue papers
11as first author
18since 2021 · last 2026
0000-0001-9362-6237ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 9 first-author · 13 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniVid: Unifying Vision Tasks with Pre-trained Video Generation ModelsabstractLarge language models, trained on extensive corpora, successfully unify diverse linguistic tasks within a single generative framework. Inspired by this, recent works like Large Vision Model (LVM) extend this paradigm to vision by organizing tasks into sequential visual sentences, where visual prompts serve as the context to guide outputs. However, such modeling requires task-specific pre-training across modalities and sources, which is costly and limits scalability to unseen tasks. Given that pre-trained video generation models inherently capture temporal sequence dependencies, we explore a more unified and scalable alternative: can a pre-trained video generation model adapt to diverse image and video tasks? To answer this, we propose UniVid, a framework that fine-tunes a video diffusion transformer to handle various vision tasks without task-specific modifications. Tasks are represented as visual sentences, where the context sequence defines both the task and the expected output modality. We evaluate the generalization of UniVid from two perspectives: (1) cross-modal inference with contexts composed of both images and videos, extending beyond LVM’s uni-modal setting; (2) cross-source tasks from natural to annotated data, without multi-source pre-training. Despite being trained solely on natural video data, UniVid generalizes well in both settings. Notably, understanding and generation tasks can easily switch by simply reversing the visual sentence order in this paradigm. These findings highlight the potential of pre-trained video generation models to serve as a scalable and unified foundation for vision modeling. Our code is released at https://github.com/CUC-MIPG/UniVid. Yuchao Gu, Qi Mao 0002 |
WACV | 3 |
| 2026 | EmoAgent: A Multi-Agent Framework for Diverse Affective Image ManipulationabstractAffective Image Manipulation (AIM) aims to alter visual elements within an image to evoke specific emotional responses from viewers. However, existing AIM approaches rely on rigidone-to-onemappings between emotions and visual cues, making them ill-suited for the inherently subjective and diverse ways in which humans perceive and express emotion. To address this, we introduce a novel task setting termedDiverse AIM (D-AIM), aiming to generate multiple visually distinct yet emotionally consistent image edits from a single source image and target emotion. We proposeEmoAgent, the first multi-agent framework tailored specifically for D-AIM. EmoAgent explicitly decomposes the manipulation process into three specialized phases executed by collaborative agents: a Planning Agent that generates diverse emotional editing strategies, an Editing Agent that precisely executes these strategies, and a Critic Agent that iteratively refines the results to ensure emotional accuracy. This collaborative design empowers EmoAgent to modelone-to-manyemotion-to-visual mappings, enabling semantically diverse and emotionally faithful edits. Extensive quantitative and qualitative evaluations demonstrate that EmoAgent substantially outperforms state-of-the-art approaches in both emotional fidelity and semantic diversity, effectively generating multiple distinct visual edits that convey the same target emotion. Qi Mao 0002, Haobo Hu, Yujie She, Difei Gao, Libiao Jin |
IEEE Trans. Affect. Comput. | 1 |
| 2026 | HD-Custom: Efficient Hierarchical Disentanglement for Coarse-to-Fine Concept Customization in Subject Video Generation
Yuanhang Li, Qi Mao 0002, Xinyan Xiao, Libiao Jin, Siwei Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Stable Diffusion is a Natural Cross-Modal Decoder for Layered AI-Generated Image CompressionabstractRecent advances in Artificial Intelligence Generated Content (AIGC) triggered an increasing need to transmit and compress the vast number of AI-generated images (AIGIs). However, there is a noticeable deficiency in research focused on compression methods for AIGIs. To address this critical gap, we advocate that Stable Diffusion serves as a natural cross-modal decoder by leveraging rich and scalable priors, and introduce a scalable cross-modal compression framework that incorporates multiple human-comprehensible modalities. As illustrated in Fig. 1(a), the proposed framework encodes images into a layered bitstream: a semantic prior that delivers high-level semantic information through text prompts; a structural prior that captures spatial details using edge or skeleton maps; and a texture prior that preserves local textures via a colormap. Utilizing Stable Diffusion as the backend, the decoder leverages multi-modal scalable priors to generate images with different levels of fidelity. Experiments show our method preserves realistic details and semantic fidelity at an extremely low bitrate (< 0.02 bpp), comparable with recent perceptual coding approaches and outperforming VVC. The R-D performance also demonstrate the scalability of our proposed multi-layered bitstream since image fidelity incrementally improves with structure and texture priors provided during decoding. Additionally, as illustrated in Fig. 1(b), our framework facilitates downstream editing applications such as Structure Manipulation, Texture Synthesis, and Object Erasing, without requiring full decoding, thereby paving a new direction for future research in AIGI compression. Ruijie Chen, Qi Mao 0002, Zhengxue Cheng |
DCC | 2 |
| 2025 | Exploring Multimodal Knowledge for Image Compression via Large Foundation ModelsabstractKnowledge is an abstraction of factual principles of the physical world. Large foundation models encapsulate extensive multimodal knowledge into the parameters and thus invoke machine intelligence on various tasks. How to invoke the knowledge in these models to facilitate image compression lacks in-depth exploration. In this work, we aim to harness multimodal knowledge into ultra-low bitrate compression and propose Multimodal Knowledge-aware Image Compression (MKIC). Our key insight is that under the context of ultra-low bitrate compression, where the encoded representation is too sparse to represent enough information of the input signal, knowledge from the physical world is required to be incorporated into the compression. Thus, more shared patterns can be stored in the model together with sparse unique features also embedded into the bitstream. In light of two kinds of knowledge, namely natural visual knowledge and human language knowledge, we propose a novel Alternating Rate-Distortion Optimization to enhance the accuracy and compactness of global semantic text representation extraction, extract the local feature map that captures visual details, and integrate these multimodal representations into a large generative foundation model to achieve high-quality reconstruction. The proposed method relights the path of learned image coding, leveraging decoupled knowledge from large foundation models. Extensive experiments show that our proposed method achieves superior comprehensive performance compared to various methods and shows great potential for ultra-low bitrate image compression. Junlong Gao, Zhimeng Huang, Qi Mao 0002, Siwei Ma 0001, Chuanmin Jia |
IEEE Trans. Image Process. | 3 |
| 2024 | Extreme Image Compression Using Fine-tuned VQGANsabstractRecent advances in generative compression methods have demonstrated remarkable progress in enhancing the perceptual quality of compressed data, especially in scenarios with low bitrates. However, their efficacy and applicability to achieve extreme compression ratios (< 0.05 bpp) remain constrained. In this work, we propose a simple yet effective coding framework by introducing vector quantization (VQ)–based generative models into the image compression domain. The main insight is that the codebook learned by the VQGAN model yields a strong expressive capacity, facilitating efficient compression of continuous information in the latent space while maintaining reconstruction quality. Specifically, an image can be represented as VQ-indices by finding the nearest codeword, which can be encoded using lossless compression methods into bitstreams. We propose clustering a pre-trained large-scale codebook into smaller codebooks through the K-means algorithm, yielding variable bitrates and different levels of reconstruction quality within the coding framework. Furthermore, we introduce a transformer to predict lost indices and restore images in unstable environments. Extensive qualitative and quantitative experiments on various benchmark datasets demonstrate that the proposed framework outperforms state-of-the-art codecs in terms of perceptual quality-oriented metrics and human perception at extremely low bitrates (≤ 0.04 bpp). Remarkably, even with the loss of up to 20% of indices, the images can be effectively restored with minimal perceptual loss. Qi Mao 0002, Tinghan Yang, Meng Wang 0017, Shiqi Wang 0001, Libiao Jin, Siwei Ma 0001 |
DCC | 1 |
| 2024 | Unrolled Decomposed Unpaired Learning for Controllable Low-Light Video Enhancement
Lingyu Zhu 0006, Wenhan Yang, Baoliang Chen, Hanwei Zhu, Zhangkai Ni, Qi Mao 0002, Shiqi Wang 0001 |
ECCV (23) | 6 |
| 2024 | Learned Image Compression for Both Humans and Machines via Dynamic AdaptationabstractRecent advancements in neural image compression have shown great potential in outperforming conventional standard codecs in terms of both rate-distortion and rate-analysis performance. However, there is an issue of divergent preferences in information preservation or reconstruction in the process of compression for humans and machines, respectively. Compression for humans tends to retain the signal fidelity or perceptual quality of visual appearance while compression for machines requires preserving critical semantic information, resulting in the limitation of the bitstream supporting only a single requirement during the compression. To bridge this gap, we propose a dynamic adaptation approach that generates a single bitstream serving both humans and machines. This approach aims to mitigate the domain gap among tasks, which facilitates maintaining the performance of out-of-scope tasks. Specifically, the proposed method concentrates on learning a dynamic adaptation process, i.e., optimizing the latent representation in the compressed domain in an end-to-end manner while adhering to the rate-performance constraint. Extensive results reveal that our paradigm significantly reduces the domain gap, surpassing existing codecs. Lingyu Zhu 0006, Binzhe Li, Riyu Lu, Peilin Chen 0001, Qi Mao 0002, Zhao Wang 0004, Wenhan Yang, Shiqi Wang 0001 |
ICIP | 5 |
| 2024 | Unifying Generation and Compression: Ultra-low bitrate Image Coding Via Multi-stage TransformerabstractRecent progress in generative compression technology has significantly improved the perceptual quality of compressed data. However, these advancements primarily focus on producing high-frequency details, often overlooking the ability of generative models to capture the prior distribution of image content, thus impeding further bitrate reduction in extreme compression scenarios (< 0.05 bpp). Motivated by the capabilities of predictive language models for lossless compression, this paper introduces a novel Unified Image Generation-Compression (UIGC) paradigm, merging the processes of generation and compression. A key feature of the UIGC framework is the adoption of vector-quantized (VQ) image models for tokenization, alongside a multi-stage transformer designed to exploit spatial contextual information for modeling the prior distribution. As such, the dual-purpose framework effectively utilizes the learned prior for entropy estimation and assists in the regeneration of lost tokens. Extensive experiments demonstrate the superiority of the proposed UIGC framework over existing codecs in perceptual quality and human perception, particularly in ultra-low bitrate scenarios (≤0.03 bpp), pioneering a new direction in generative compression. Naifu Xue, Qi Mao 0002, Yuan Zhang 0013, Siwei Ma 0001 |
ICME | 2 |
| 2024 | MAG-Edit: Localized Image Editing in Complex Scenarios via Mask-Based Attention-Adjusted Guidance
Qi Mao 0002, Yuchao Gu, Zheng Shou 0001 |
ACM Multimedia | 1 |
| 2024 | Scalable Face Image Coding via StyleGAN Prior: Toward Compression for Human-Machine Collaborative VisionabstractThe accelerated proliferation of visual content and the rapid development of machine vision technologies bring significant challenges in delivering visual data on a gigantic scale, which shall be effectively represented to satisfy both human and machine requirements. In this work, we investigate how hierarchical representations derived from the advanced generative prior facilitate constructing an efficient scalable coding paradigm for human-machine collaborative vision. Our key insight is that by exploiting the StyleGAN prior, we can learn three-layered representations encoding hierarchical semantics, which are elaborately designed into the basic, middle, and enhanced layers, supporting machine intelligence and human visual perception in a progressive fashion. With the aim of achieving efficient compression, we propose the layer-wise scalable entropy transformer to reduce the redundancy between layers. Based on the multi-task scalable rate-distortion objective, the proposed scheme is jointly optimized to achieve optimal machine analysis performance, human perception experience, and compression ratio. We validate the proposed paradigm's feasibility in face image compression. Extensive qualitative and quantitative experimental results demonstrate the superiority of the proposed paradigm over the latest compression standard Versatile Video Coding (VVC) in terms of both machine analysis as well as human perception at extremely low bitrates (< 0.01 bpp), offering new insights for human-machine collaborative compression. Qi Mao 0002, Chongyu Wang, Meng Wang 0017, Shiqi Wang 0001, Ruijie Chen, Libiao Jin, Siwei Ma 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Extreme Generative Human-Oriented Video Coding via Motion Representation CompressionabstractThe increasing popularity of video conferencing and live streaming raises the growing demand for encoding human-oriented videos at ultra-low bit rates. Recently, several ultra-low bitrate video codecs have proposed using inter-frame keypoints or landmarks to derive motion representations, which are then used to warp decoded frames in a generative manner. Despite its success, compression of the motion representation has been less investigated in the literature. In this work, we propose a novel principal component analysis (PCA)-based decomposing method to fully exploit the compression potential of motion representations. In particular, we decompose the derived motion affine matrices into three parts and apply quantization and entropy estimation to each part in a different way depending on its significance. Using such compressed-friendly motion representations allows for preserving most of the motion information and achieving lower coding costs. Extensive qualitatively and quantitatively experimental results on the human video datasets demonstrate the superiority of the proposed paradigm over existing video codecs under extreme compression ratios. Qi Mao 0002, Chuanmin Jia, Ronggang Wang, Siwei Ma 0001 |
ISCAS | 2 |
| 2023 | ZGaming: Zero-Latency 3D Cloud Gaming by Image PredictionabstractIn cloud gaming, interactive latency is one of the most important factors in users' experience. Although the interactive latency can be reduced through typical network infrastructures like edge caching and congestion control, the interactive latency of current cloud-gaming platforms is still far from users' satisfaction. Jiangkai Wu, Yu Guan 0005, Qi Mao 0002, Yong Cui 0001, Zongming Guo, Xinggong Zhang |
SIGCOMM | 3 |
| 2023 | Semantic-Aware Visual Decomposition for Image Coding
Jianhui Chang, Jian Zhang 0018, Jiguo Li 0002, Shiqi Wang 0001, Qi Mao 0002, Chuanmin Jia, Siwei Ma 0001, Wen Gao 0001 |
Int. J. Comput. Vis. | 5 |
| 2023 | Enhancing Style-Guided Image-to-Image Translation via Self-Supervised Metric LearningabstractThere has been significant success in recent image-to-image translation (I2I) approaches in translating the source image into the style of the target image. Existing techniques rely on the disentanglement of content and style representations, requiring a two-stage style mapping process: Reference images are used to extract style vectors, which are subsequently remapped into the translated images. However, when the target domain contains a variety of styles, such a two-stage style mapping cannot guarantee the translated image be style consistent with its guided reference image. In this work, we propose to explicitly employ metric learning to enhance the two-stage style mapping in style-guided image translation. The distance between deep features Gram matrices is utilized to construct the visual style metric as self-supervised similarity labels, guiding the embedding of style vectors using triplet loss with adaptive margins in the first stage. Furthermore, in the second stage, we consider generated images and their corresponding reference images as positive samples and anchors for each other, while the nearest negative sample is used to construct the triplet loss in the proposed metric space. The proposed learning algorithms can be applied to any I2I framework that uses disentangled representations without modifying the original network architectures. We evaluate the proposed method on three representative I2I translation baselines. Both qualitative and quantitative results demonstrate that the proposed approach enhances style alignment in style-guided translation compared to the baselines. Qi Mao 0002, Siwei Ma 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | Disentangled Visual Representations for Extreme Human Body Video CompressionabstractRecent years have witnessed the great promise of deep neural video compression codecs. However, there are still unprecedented challenges ahead when the videos are expected to be encoded with extremely low bitrate. Motivated by recent attempts of layered conceptual image compression, we make the first attempt to leverage the disentangled visual representations for extreme human body video compression. More specifically, to capture the main structure, we adopt the inferred human pose keypoints as the structure code of each frame, thereby deriving the motion information from structure codes of adjacent frames for further compression. To better exploit the texture redundancy, all frames share the same texture codes by incorporating the proposed texture contrastive learning to ensure texture consistency within a video. Two branches are consequently transmitted in a separable manner, and the generator synthesizes the reconstructed video with the combination of all decoded representations at the decoder side. Both qualitative and quantitative experimental results demonstrate that the proposed scheme can produce perceptually pleasing reconstruction results in ultra-low bitrates far below that can be reached by other video codecs. Qi Mao 0002, Shiqi Wang 0001, Chuanmin Jia, Ronggang Wang, Siwei Ma 0001 |
ICME | 2 |
| 2022 | Continuous and Diverse Image-to-Image Translation via Signed Attribute Vectors
Qi Mao 0002, Hung-Yu Tseng, Hsin-Ying Lee 0001, Jia-Bin Huang 0001, Siwei Ma 0001, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 1 |
| 2022 | Conceptual Compression via Deep Structure and Texture SynthesisabstractExisting compression methods typically focus on the removal of signal-level redundancies, while the potential and versatility of decomposing visual data into compact conceptual components still lack further study. To this end, we propose a novel conceptual compression framework that encodes visual data into compact structure and texture representations, then decodes in a deep synthesis fashion, aiming to achieve better visual reconstruction quality, flexible content manipulation, and potential support for various vision tasks. In particular, we propose to compress images by a dual-layered model consisting of two complementary visual features: 1) structure layer represented by structural maps and 2) texture layer characterized by low-dimensional deep representations. At the encoder side, the structural maps and texture representations are individually extracted and compressed, generating the compact, interpretable, inter-operable bitstreams. During the decoding stage, a hierarchical fusion GAN (HF-GAN) is proposed to learn the synthesis paradigm where the textures are rendered into the decoded structural maps, leading to high-quality reconstruction with remarkable visual realism. Extensive experiments on diverse images have demonstrated the superiority of our framework with lower bitrates, higher reconstruction quality, and increased versatility towards visual analysis and content manipulation tasks. Jianhui Chang, Zhenghui Zhao, Chuanmin Jia, Shiqi Wang 0001, Lingbo Yang, Qi Mao 0002, Jian Zhang 0018, Siwei Ma 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | DRIT++: Diverse Image-to-Image Translation via Disentangled Representations
Hsin-Ying Lee 0001, Hung-Yu Tseng, Qi Mao 0002, Jia-Bin Huang 0001, Yu-Ding Lu, Maneesh Kumar Singh 0001, Ming-Hsuan Yang 0001 |
Int. J. Comput. Vis. | 3 |
| 2019 | Mode Seeking Generative Adversarial Networks for Diverse Image SynthesisabstractMost conditional generation tasks expect diverse outputs given a single conditional context. However, conditional generative adversarial networks (cGANs) often focus on the prior conditional information and ignore the input noise vectors, which contribute to the output variations. Recent attempts to resolve the mode collapse issue for cGANs are usually task-specific and computationally expensive. In this work, we propose a simple yet effective regularization term to address the mode collapse issue for cGANs. The proposed method explicitly maximizes the ratio of the distance between generated images with respect to the corresponding latent codes, thus encouraging the generators to explore more minor modes during training. This mode seeking regularization term is readily applicable to various conditional generation tasks without imposing training overhead or modifying the original network structures. We validate the proposed algorithm on three conditional image synthesis tasks including categorical generation, image-to-image translation, and text-to-image synthesis with different baseline models. Both qualitative and quantitative results demonstrate the effectiveness of the proposed regularization method for improving diversity without loss of quality. Qi Mao 0002, Hsin-Ying Lee 0001, Hung-Yu Tseng, Siwei Ma 0001, Ming-Hsuan Yang 0001 |
CVPR | 1 |
| 2019 | Layered Conceptual Image Compression Via Deep Semantic SynthesisabstractMotivated by the insight of Marr on generative image representations, we propose a layered conceptual image compression scheme by integrating the advantages of both variational auto-encoders (VAEs) and generative adversarial networks (GANs). In particular, the image is represented by two layers: the low-dimensional codes of the stochastic textures encoded by the VAE and the geometric structures characterized by edge maps. Subsequently, the edge maps and latent codes are compressed individually such that the final bit streams are formed in a combined manner. At the decoder side, the GAN synthesizes the decoded images on the basis of the latent codes and the reconstructed edge maps. Experimental results demonstrate that our proposed scheme achieves better visual reconstruction quality than the traditional image compression algorithms such as JPEG, JPEG2000 and HEVC (intra coding) in the low bit rate coding scenarios. Jianhui Chang, Qi Mao 0002, Zhenghui Zhao, Shanshe Wang, Shiqi Wang 0001, Siwei Ma 0001 |
ICIP | 2 |
| 2019 | Fidelity or Quality? A Region-Aware Framework for Enhanced Image Decoding via Hybrid Neural NetworksabstractThe generative deep learning models such as the generative adversarial networks (GAN) have been shown to efficiently generate visually appealing images by learning the natural scene statistics. However, the signal fidelity, instead of the visual quality, has been largely ignored in the generation process, especially for the highly structural regions. In this paper, we introduce a region-aware visual signal restoration scheme to achieve a good balance between visual quality and fidelity. As a specific example of this framework, we develop an enhanced decoding scheme with hybrid neural networks, such that the base fidelity layer and texture quality enhancement layer are combined adaptively to restore the compressed images. The efficiency of the proposed framework is demonstrated with extensive experimental results, which show favorable performance against the state-of-the-art methods. Qi Mao 0002, Shiqi Wang 0001, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001 |
ICIP | 1 |
| 2018 | Enhanced Image Decoding via Edge-Preserving Generative Adversarial NetworksabstractLossy image compression usually introduces undesired compression artifacts, such as blocking, ringing and blurry effect{###} S, especially in low bit rate coding scenarios. Although many algorithms have been proposed to reduce these compression artifacts, most of them are based on image local smoothness prior, which usually leads to over-smoothing around the areas with distinct structures, e.g., edges and textures. In this paper, we propose a novel framework to enhance the perceptual quality of decoded images by well preserving the edge structures and predicting visually pleasing textures. Firstly, we propose an edge-preserving generative adversarial network (EP-GAN) to achieve edge restoration and texture generation simultaneously. Then, we elaborately design an edge fidelity regularization term to guide our network, which jointly utilizes the signal fidelity, feature fidelity and adversarial constraint to reconstruct high quality decoded images. Experimental results demonstrate that the proposed EP-GAN is able to efficiently enhance decoded images at low bit rate and reconstruct more perceptually pleasing images with abundant textures and sharp edges. Qi Mao 0002, Shiqi Wang 0001, Shanshe Wang, Xinfeng Zhang 0001, Siwei Ma 0001 |
ICME | 1 |
| 2017 | Local Disparity Vector Derivation Scheme in 3D-AVS2
Qi Mao 0002, Shanshe Wang, Siwei Ma 0001 |
ICIG (2) | 1 |
| 2016 | A local-adapted disparity vector derivation scheme for 3D-AVSabstractIn the 3D extension of Audio Video Coding Standard (AVS), i.e. 3D-AVS, the Global Disparity Vector (GDV) derivation technique has been proposed to provide an estimation for Disparity Vector (DV) in inter-view prediction, where the GDV is generated by averaging all Disparity Vectors (DVs) in the latest previously coded frame. The prediction accuracy of GDV may be however limited by the lack of local adaptivity. In this paper, we introduce a novel Local Disparity Vector (LDV) derivation scheme. Specifically, the DV of the current block is calculated from the DVs within a neighboring region, whose size can be adaptively expanded to increase the robustness and accuracy. Experimental results show that the proposed LDV derivation method can provide around 2.12% and 1.37% bitrate reductions for compressed views and synthesized views compared with the GDV scheme, respectively. Qi Mao 0002, Shanshe Wang, Xiang Zhang 0004, Xinfeng Zhang 0001, Siwei Ma 0001 |
VCIP | 1 |