VLDB 2026 Research / reviewers in the wild / expert
Zheng Gu 0001
dblp:13/6292-1
· DBLP profile ↗
12ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0001-9914-3922ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Retrieval-augmented diffusion with acoustic priors for high-fidelity sonar image generation
Shaocong Yang, Zheng Gu 0001, Hongye Cao, Xiaolong Qi, Jing Huo, Yang Gao 0001 |
Pattern Recognit. | 2 |
| 2026 | Semantic-Structural Alignment for Generative Pictorial ChartsabstractTraditional statistical graphics are precise but often lack the visual appeal, memorability, and engagement of pictorial charts. We present a generative framework for the automated synthesis of pictorial charts that bridges the gap between semantic expression and structural faithfulness. Rather than treating charts merely as images to be stylized, we frame the problem as a dual-conditioned generation task guided by two parallel external control signals: a text prompt capturing the semantic context of the editing intent, and a context image providing the abstract statistical chart's global structure. To reinforce these controls within a Multi-Modal Diffusion Transformer, we introduce two complementary feature-level mechanisms: structural alignment to anchor spatial layouts to the input chart, and semantic alignment to transfer expressive textures from reference images. Generalizing across major visual channels (i.e., length, area, angle, and position) and diverse semantic domains, our method produces pictorial charts that are both artistically compelling and structurally consistent. Extensive quantitative evaluations and perceptual user studies demonstrate that our framework outperforms traditional controllable generation and image editing baselines, providing a foundation for high-fidelity, data-driven generative modeling in expressive visual storytelling. Project page: https://ssalign.github.io/. Zhida Sun, Zheng Gu 0001, Min Lu 0002, Bongshin Lee, Daniel Cohen-Or, Hui Huang 0004 |
ACM Trans. Graph. | 3 |
| 2026 | MultiPaint: A Unified Framework for Multi-Task, Multi-Object, and Multi-Condition Video InpaintingabstractVideo inpainting modifies local regions in video while ensuring spatial and temporal coherence. However, existing methods-both traditional and recent diffusion-based ones-face key limitations: they lack unified support for both insertion and completion, and are restricted to single-object inpainting, making it difficult to handle multi-object scenarios involving grounding and interaction. In this article, we propose MultiPaint, a unified framework for multi-task, multi-object, and multi-condition video inpainting. First, we introduce dual-branch adapters to unify the insertion and completion tasks within a single model. Moreover, we propose a test-time scheduled feature composition strategy that enables multi-object inpainting with user-specified locations while better preserving interactions among objects, a setting that has been insufficiently addressed in prior work. Additionally, we introduce a multi-condition inpainting scheme that integrates text-guided, image-guided, and keyframe-guided modes via dynamic frame masking, providing more controllability in appearance customization. Extensive experiments show that MultiPaint achieves state-of-the-art performance on object insertion and scene completion among the recent works. We further demonstrate its versatility in downstream tasks including grounded video generation, object editing, object removal, image-guided inpainting, and long video inpainting. Zheng Gu 0001, Xin Tao 0001, Pengfei Wan 0001, Xiaodong Chen 0009, Jing Liao 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | REST: A resolution preserving network for photorealistic style transfer via semantic distillation
Jing Huo, Zheng Gu 0001, Jiulin Zhang, Xiangde Liu, Shiyin Jin, Pinzhuo Tian, Wenbin Li 0006, Jing Wu 0004, Yukun Lai, Yang Gao 0001 |
Comput. Vis. Image Underst. | 2 |
| 2025 | Few-shot exemplar-driven inpainting with parameter-efficient diffusion fine-tuningabstractText-to-image diffusion models have demonstrated impressive capabilities in image generation and have been effectively applied to image inpainting. While text prompt provides an intuitive guidance for conditional inpainting, users often seek the ability to inpaint a specific object with customized appearance by providing an exemplar image. Unfortunately, existing methods struggle to achieve high fidelity in exemplar-driven inpainting. To address this, we use a plug-and-play low-rank adaptation (LoRA) module based on a pretrained text-driven inpainting model. The LoRA module is dedicated to learn the exemplar-specific concepts through few-shot fine-tuning, bringing improved fitting capability to customized exemplar images, without intensive training on large-scale datasets. Additionally, we introduce GPT-4V prompting and prior noise initialization techniques to further facilitate the fidelity in inpainting results. In brief, the denoising diffusion process first starts with the noise derived from a composite exemplar–background image, and is subsequently guided by an expressive prompt generated from the exemplar using the GPT-4V model. Extensive experiments demonstrate that our method achieves state-of-the-art performance, qualitatively and quantitatively, offering users an exemplar-driven inpainting tool with enhanced customization capability. Zheng Gu 0001, Wenyue Hao, Yi Wang 0066, Huaiyu Cai, Xiaodong Chen 0009 |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2024 | Task-Aware Few-Shot Image Generation via Dynamic Local Distribution Estimation and Sampling
Zheng Gu 0001, Wenbin Li 0006, Tianyu Ding, Jing Huo, Kuihua Huang, Yang Gao 0001 |
PRCV (2) | 1 |
| 2024 | Analogist: Out-of-the-box Visual In-Context Learning with Image Diffusion ModelabstractVisual In-Context Learning (ICL) has emerged as a promising research area due to its capability to accomplish various tasks with limited example pairs through analogical reasoning. However, training-based visual ICL has limitations in its ability to generalize to unseen tasks and requires the collection of a diverse task dataset. On the other hand, existing methods in the inference-based visual ICL category solely rely on textual prompts, which fail to capture fine-grained contextual information from given examples and can be time-consuming when converting from images to text prompts. To address these challenges, we propose Analogist, a novel inference-based visual ICL approach that exploits both visual and textual prompting techniques using a text-to-image diffusion model pretrained for image inpainting. For visual prompting, we propose a self-attention cloning (SAC) method to guide the fine-grained structural-level analogy between image examples. For textual prompting, we leverage GPT-4V's visual reasoning capability to efficiently generate text prompts and introduce a cross-attention masking (CAM) operation to enhance the accuracy of semantic-level analogy guided by text prompts. Our method is out-of-the-box and does not require fine-tuning or optimization. It is also generic and flexible, enabling a wide range of visual tasks to be performed in an in-context manner. Extensive experiments demonstrate the superiority of our method over existing approaches, both qualitatively and quantitatively. Our project webpage is available at https://analogist2d.github.io. Zheng Gu 0001, Jing Liao 0001, Jing Huo, Yang Gao 0001 |
ACM Trans. Graph. | 1 |
| 2022 | CariMe: Unpaired Caricature Generation With Multiple ExaggerationsabstractCaricature generation aims to translate real photos into caricatures with artistic styles and shape exaggerations while maintaining the identity of the subject. Different from generic image-to-image translation, drawing caricatures automatically is a more challenging task due to the existence of various spatial deformations. Previous caricature generation methods are obsessed with predicting definite image warping from a given photo while ignoring the intrinsic representation and distribution of geometric exaggerations in caricatures. This limits their ability on diverse exaggeration generation. In this paper, we generalize the caricature generation problem from instance-level warping prediction to distribution-level deformation modeling. Based on this assumption, we present the first exploration forunpaired CARIcature generation with Multiple Exaggerations (CariMe). Technically, we propose a Multi-exaggeration Warper network to learn the distribution-level mapping from photos to facial exaggerations. This makes it possible to generate diverse and reasonable exaggerations from randomly sampled warp codes given one input photo. To better represent the facial exaggeration and produce fine-grained warping, a deformation-field-based warping method is also proposed, which captures more detailed exaggerations than previous point-based warping methods. Experiments and two perceptual studies prove the superiority of our method comparing with other state-of-the-art methods, showing the improvement of our work on caricature generation. The source code is available athttps://github.com/edward3862/CariMe-pytorch. Zheng Gu 0001, Chuanqi Dong, Jing Huo, Wenbin Li 0006, Yang Gao 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | LoFGAN: Fusing Local Representations for Few-shot Image GenerationabstractGiven only a few available images for a novel unseen category, few-shot image generation aims to generate more data for this category. Previous works attempt to globally fuse these images by using adjustable weighted coefficients. However, there is a serious semantic misalignment between different images from a global perspective, making these works suffer from poor generation quality and diversity. To tackle this problem, we propose a novel Local-Fusion Generative Adversarial Network (LoFGAN) for fewshot image generation. Instead of using these available images as a whole, we first randomly divide them into a base image and several reference images. Next, LoFGAN matches local representations between the base and reference images based on semantic similarities, and replaces the local features with the closest related local features. In this way, LoFGAN can produce more realistic and diverse images at a more fine-grained level, and simultaneously enjoy the characteristic of semantic alignment. Furthermore, a local reconstruction loss is also proposed, which can provide better training stability and generation quality. We conduct extensive experiments on three datasets, which successfully demonstrates the effectiveness of our proposed method for few-shot image generation and downstream visual applications with limited data. Code is available at https://github.com/edward3862/LoFGAN-pytorch. Zheng Gu 0001, Wenbin Li 0006, Jing Huo, Lei Wang 0001, Yang Gao 0001 |
ICCV | 1 |
| 2020 | Unsupervised Domain Attention Adaptation Network for Caricature Attribute Recognition
Kelei He, Jing Huo, Zheng Gu 0001, Yang Gao 0001 |
ECCV (8) | 4 |
| 2020 | Learning Task-aware Local Representations for Few-shot LearningabstractFew-shot learning for visual recognition aims to adapt to novel unseen classes with only a few images. Recent work, especially the work based on low-level information, has achieved great progress. In these work, local representations (LRs) are typically employed, because LRs are more consistent among the seen and unseen classes. However, most of them are limited to an individual image-to-image or image-to-class measure manner, which cannot fully exploit the capabilities of LRs, especially in the context of a certain task. This paper proposes an Adaptive Task-aware Local Representations Network (ATL-Net) to address this limitation by introducing episodic attention, which can adaptively select the important local patches among the entire task, as the process of human recognition. We achieve much superior results on multiple benchmarks. On the miniImagenet, ATL-Net gains 0.93% and 0.88% improvements over the compared methods under the 5-way 1-shot and 5-shot settings. Moreover, ATL-Net can naturally tackle the problem that how to adaptively identify and weight the importance of different key local parts, which is the major concern of fine-grained recognition. Specifically, on the fine-grained dataset Stanford Dogs, ATL-Net outperforms the second best method with 5.39% and 9.69% gains under the 5-way 1-shot and 5-shot settings. Chuanqi Dong, Wenbin Li 0006, Jing Huo, Zheng Gu 0001, Yang Gao 0001 |
IJCAI | 4 |
| 2019 | DeepMEF: A Deep Model Ensemble Framework for Video Based Multi-modal Person IdentificationabstractThe goal of video based multi-modal person identification is to identify a person of interest using multi-modal video features, such as person's face, body, audio or head features. This task is challenging due to many factors, for example, variant body or face poses, poor face image quality, low frame resolution, etc. To address these problems, we propose a deep model ensemble framework, namely DeepMEF. Specifically, the proposed framework includes three novel modules, i.e., the video feature fusion module, the multi-modal feature fusion module and the model ensemble module. The first and second module form the basic deep model for ensemble, with the video feature fusion module fuses facial features from different frames as one. Then the multi-modal feature fusion module further fuses the face feature and features of other modalities for identification. In this work, we adopt the scene feature extracted by ourselves as the additional input of the multi-modal module. At last, the model ensemble module promotes the overall performance by combining the predictions of multiple multi-modal learners. The proposed method achieves a competitive result of 89.86% in mAP on the iQIYI-VID-2019 dataset, which helps us win the third place in the 2019 iQIYI Celebrity Video Identification Challenge. Chuanqi Dong, Zheng Gu 0001, Zhonghao Huang, Jing Huo, Yang Gao 0001 |
ACM Multimedia | 2 |