VLDB 2026 Research / reviewers in the wild / expert
Zhaorui Gu
dblp:208/9141
· DBLP profile ↗
20ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-6673-7932ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | G2LFormer: Global-to-local token mixing transformer for blind image inpainting and beyond
Haoru Zhao, Zonghui Guo, Shishi Qiao, Zhaorui Gu, Junyu Dong, Haiyong Zheng |
Pattern Recognit. | 4 |
| 2025 | Context-aware mutual learning for blind image inpainting and beyond
Haoru Zhao, Zhaorui Gu, Haiyong Zheng |
Expert Syst. Appl. | 3 |
| 2025 | Image sterilization via overwriting
Mingyao Tan, Lei Huang 0010, Zhaorui Gu, Xiaodong Wang 0006, Haiyong Zheng |
Inf. Sci. | 4 |
| 2025 | Reference-then-supervision framework for infrared and visible image fusion
Guihui Li, Zhensheng Shi, Zhaorui Gu, Haiyong Zheng |
Pattern Recognit. | 3 |
| 2023 | Transformer for Image Harmonization and BeyondabstractImage harmonization, aiming to make composite images look more realistic, is an important and challenging task. The composite, synthesized by combining foreground from one image with background from another image, inevitably suffers from the issue of inharmonious appearance caused by distinct imaging conditions, i.e., lights. Current solutions mainly adopt an encoder-decoder architecture with convolutional neural network (CNN) to capture the context of composite images, trying to understand what it should look like in the foreground referring to surrounding background. In this work, we seek to solve image harmonization with Transformer, by leveraging its powerful ability of modeling long-range context dependencies, for adjusting foreground light to make it compatible with background light while keeping structure and semantics unchanged. We present the design of our two vision Transformer frameworks and corresponding methods, as well as comprehensive experiments and empirical study, demonstrating the power of Transformer and investigating the Transformer for vision. Our methods achieve state-of-the-art performance on the image harmonization as well as four additional vision and graphics tasks, i.e., image enhancement, image inpainting, white-balance editing, and portrait relighting, indicating the superiority of our work. Code, models, more results and details can be found at the project website http://ouc.ai/project/HarmonyTransformer. Zonghui Guo, Zhaorui Gu, Junyu Dong, Haiyong Zheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | TransCNN-HAE: Transformer-CNN Hybrid AutoEncoder for Blind Image InpaintingabstractBlind image inpainting is extremely challenging due to the unknown and multi-property complexity of contamination in different contaminated images. Current mainstream work decomposes blind image inpainting into two stages: mask estimating from the contaminated image and image inpainting based on the estimated mask, and this two-stage solution involves two CNN-based encoder-decoder architectures for estimating and inpainting separately. In this work, we propose a novel one-stage Transformer-CNN Hybrid AutoEncoder (TransCNN-HAE) for blind image inpainting, which intuitively follows the inpainting-then-reconstructing pipeline by leveraging global long-range contextual modeling of Transformer to repair contaminated regions and local short-range contextual modeling of CNN to reconstruct the repaired image. Moreover, a Cross-layer Dissimilarity Prompt (CDP) is devised to accelerate the identifying and inpainting of contaminated regions. Ablation studies validate the efficacy of both TransCNN-HAE and CDP, and extensive experiments on various datasets with multi-property contaminations show that our method achieves state-of-the-art performance with much lower computational cost on blind image inpainting. Our code is available at https://github.com/zhenglab/TransCNN-HAE. Haoru Zhao, Zhaorui Gu, Haiyong Zheng |
ACM Multimedia | 2 |
| 2021 | Intrinsic Image HarmonizationabstractCompositing an image usually inevitably suffers from inharmony problem that is mainly caused by incompatibility of foreground and background from two different images with distinct surfaces and lights, corresponding to material-dependent and light-dependent characteristics, namely, reflectance and illumination intrinsic images, respectively. Therefore, we seek to solve image harmonization via separable harmonization of reflectance and illumination, i.e., intrinsic image harmonization. Our method is based on an autoencoder that disentangles composite image into reflectance and illumination for further separate harmonization. Specifically, we harmonize reflectance through material-consistency penalty, while harmonize illumination by learning and transferring light from background to foreground, moreover, we model patch relations between foreground and background of composite images in an inharmony-free learning way, to adaptively guide our intrinsic image harmonization. Both extensive experiments and ablation studies demonstrate the power of our method as well as the efficacy of each component. We also contribute a new challenging dataset for benchmarking illumination harmonization. Code and dataset are at https://github.com/zhenglab/IntrinsicHarmony. Zonghui Guo, Haiyong Zheng, Zhaorui Gu |
CVPR | 4 |
| 2021 | Image Harmonization with TransformerabstractImage harmonization, aiming to make composite images look more realistic, is an important and challenging task. The composite, synthesized by combining foreground from one image with background from another image, inevitably suffers from the issue of inharmonious appearance caused by distinct imaging conditions, i.e., lights. Current solutions mainly adopt an encoder-decoder architecture with convolutional neural network (CNN) to capture the context of composite images, trying to understand what it looks like in the surrounding background near the foreground. In this work, we seek to solve image harmonization with Transformer, by leveraging its powerful ability of modeling long-range context dependencies, for adjusting foreground light to make it compatible with background light while keeping structure and semantics unchanged. We present the design of our harmonization Transformer frameworks without and with disentanglement, as well as comprehensive experiments and ablation study, demonstrating the power of Transformer and investigating the Transformer for vision. Our method achieves state-of-the-art performance on both image harmonization and image inpainting/enhancement, indicating its superiority. Our code and models are available at https://github.com/zhenglab/HarmonyTransformer. Zonghui Guo, Haiyong Zheng, Zhaorui Gu, Junyu Dong |
ICCV | 4 |
| 2021 | Painting from PartabstractThis paper studies the problem of painting the whole image from part of it, namely painting from part or part-painting for short, involving both inpainting and outpainting. To address the challenge of taking full advantage of both information from local domain (part) and knowledge from global domain (dataset), we propose a novel part-painting method according to the observations of relationship between part and whole, which consists of three stages: part-noise restarting, part-feature repainting, and part-patch refining, to paint the whole image by leveraging both feature-level and patch-level part as well as powerful representation ability of generative adversarial network. Extensive ablation studies show efficacy of each stage, and our method achieves state-of-the-art performance on both inpainting and outpainting benchmarks with free-form parts, including our new mask dataset for irregular outpainting. Our code and dataset are available at https://github.com/zhenglab/partpainting. Haoru Zhao, Yunhao Cheng, Haiyong Zheng, Zhaorui Gu |
ICCV | 5 |
| 2021 | Multi-Modal Multi-Action Video RecognitionabstractMulti-action video recognition is much more challenging due to the requirement to recognize multiple actions co-occurring simultaneously or sequentially. Modeling multi-action relations is beneficial and crucial to understand videos with multiple actions, and actions in a video are usually presented in multiple modalities. In this paper, we propose a novel multi-action relation model for videos, by leveraging both relational graph convolutional networks (GCNs) and video multi-modality. We first build multi-modal GCNs to explore modality-aware multi-action relations, fed by modality-specific action representation as node features, i.e., spatiotemporal features learned by 3D convolutional neural network (CNN), audio and textual embeddings queried from respective feature lexicons. We then joint both multi-modal CNN-GCN models and multi-modal feature representations for learning better relational action predictions. Ablation study, multi-action relation visualization, and boosts analysis, all show efficacy of our multi-modal multi-action relation modeling. Also our method achieves state-of-the-art performance on large-scale multi-action M-MiT benchmark. Our code is made publicly available at https://github.com/zhenglab/multi-action-video. Zhensheng Shi, Ju Liang, Haiyong Zheng, Zhaorui Gu, Junyu Dong |
ICCV | 5 |
| 2020 | Spiral Generative Network for Image Extrapolation
Hongzhi Liu 0002, Haoru Zhao, Yunhao Cheng, Qingwei Song, Zhaorui Gu, Haiyong Zheng |
ECCV (19) | 6 |
| 2020 | CoTeRe-Net: Discovering Collaborative Ternary Relations in Videos
Zhensheng Shi, Cheng Guan, Liangjie Cao, Ju Liang, Zhaorui Gu, Haiyong Zheng |
ECCV (6) | 6 |
| 2020 | Multi-Group Multi-Attention: Towards Discriminative Spatiotemporal RepresentationabstractLearning spatiotemporal features is very effective but challenging for video understanding especially action recognition. In this paper, we propose Multi-Group Multi-Attention, dubbed MGMA, paying more attention to "where and when" the action happens, for learning discriminative spatiotemporal representation in videos. The contribution of MGMA is three-fold: First, by devising a new spatiotemporal separable attention mechanism, it can learn temporal attention and spatial attention separately for fine-grained spatiotemporal representation. Second, through designing a novel multi-group structure, it can capture multi-attention rendered spatiotemporal features better. Finally, our MGMA module is lightweight and flexible yet effective, so that can be easily embedded into any 3D Convolutional Neural Network (3D-CNN) architecture. We embed multiple MGMA modules into 3D-CNN to train an end-to-end, RGB-only model and evaluate on four popular benchmarks: UCF101 and HMDB51, Something-Something V1 and V2. Ablation study and experimental comparison demonstrate the strength of our MGMA, which achieves superior performance compared to state-of-the-arts. Our code is available at https://github.com/zhenglab/mgma. Zhensheng Shi, Liangjie Cao, Cheng Guan, Ju Liang, Zhaorui Gu, Haiyong Zheng |
ACM Multimedia | 6 |
| 2020 | Discriminative Region Proposal Adversarial Network for High-Quality Image-to-Image Translation
Chao Wang 0022, Wenjie Niu, Haiyong Zheng, Zhibin Yu 0002, Zhaorui Gu |
Int. J. Comput. Vis. | 6 |
| 2020 | KA-Ensemble: towards imbalanced image classification ensembling under-sampling and over-sampling
Zhaorui Gu, Zhibin Yu 0002, Haiyong Zheng |
Multim. Tools Appl. | 3 |
| 2020 | Fine-grained facial image-to-image translation with an attention based pipeline generative adversarial framework
Ziqiang Zheng, Chao Wang 0022, Zhaorui Gu, Zhibin Yu 0002, Haiyong Zheng, Nan Wang 0013 |
Multim. Tools Appl. | 4 |
| 2018 | Discriminative Region Proposal Adversarial Networks for High-Quality Image-to-Image Translation
Chao Wang 0022, Haiyong Zheng, Zhibin Yu 0002, Ziqiang Zheng, Zhaorui Gu |
ECCV (1) | 5 |
| 2018 | Unsupervised pixel-wise classification for Chaetoceros image segmentation
Fei Zhou 0007, Zhaorui Gu, Haiyong Zheng, Zhibin Yu 0002 |
Neurocomputing | 3 |
| 2017 | Automatic plankton image classification combining multiple view features via multiple kernel learningabstractBACKGROUND: Plankton, including phytoplankton and zooplankton, are the main source of food for organisms in the ocean and form the base of marine food chain. As the fundamental components of marine ecosystems, plankton is very sensitive to environment changes, and the study of plankton abundance and distribution is crucial, in order to understand environment changes and protect marine ecosystems. This study was carried out to develop an extensive applicable plankton classification system with high accuracy for the increasing number of various imaging devices. Literature shows that most plankton image classification systems were limited to only one specific imaging device and a relatively narrow taxonomic scope. The real practical system for automatic plankton classification is even non-existent and this study is partly to fill this gap. RESULTS: Inspired by the analysis of literature and development of technology, we focused on the requirements of practical application and proposed an automatic system for plankton image classification combining multiple view features via multiple kernel learning (MKL). For one thing, in order to describe the biomorphic characteristics of plankton more completely and comprehensively, we combined general features with robust features, especially by adding features like Inner-Distance Shape Context for morphological representation. For another, we divided all the features into different types from multiple views and feed them to multiple classifiers instead of only one by combining different kernel matrices computed from different types of features optimally via multiple kernel learning. Moreover, we also applied feature selection method to choose the optimal feature subsets from redundant features for satisfying different datasets from different imaging devices. We implemented our proposed classification system on three different datasets across more than 20 categories from phytoplankton to zooplankton. The experimental results validated that our system outperforms state-of-the-art plankton image classification systems in terms of accuracy and robustness. CONCLUSIONS: This study demonstrated automatic plankton image classification system combining multiple view features using multiple kernel learning. The results indicated that multiple view features combined by NLMKL using three kernel functions (linear, polynomial and Gaussian kernel functions) can describe and use information of features better so that achieve a higher classification accuracy. Haiyong Zheng, Ruchen Wang, Zhibin Yu 0002, Nan Wang 0013, Zhaorui Gu |
BMC Bioinform. | 5 |
| 2017 | Robust and automatic cell detection and segmentation from microscopic images of non-setae phytoplankton speciesabstractSaliency‐based marker‐controlled watershed method was proposed to detect and segment phytoplankton cells from microscopic images of non‐setae species. This method first improved IG saliency detection method by combining saturation feature with colour and luminance feature to detect cells from microscopic images uniformly and then produced effective internal and external markers by removing various specific noises in microscopic images for efficient performance of watershed segmentation automatically. The authors built the first benchmark dataset for cell detection and segmentation, including 240 microscopic images across multiple phytoplankton species with pixel‐wise cell regions labelled by a taxonomist, to evaluate their method. They compared their cell detection method with seven popular saliency detection methods and their cell segmentation method with six commonly used segmentation methods. The quantitative comparison validates that their method performs better on cell detection in terms of robustness and uniformity and cell segmentation in terms of accuracy and completeness. The qualitative results show that their improved saliency detection method can detect and highlight all cells, and the following marker selection scheme can remove the corner noise caused by illumination, the small noise caused by specks, and debris, as well as deal with blurred edges. Haiyong Zheng, Nan Wang 0013, Zhibin Yu 0002, Zhaorui Gu |
IET Image Process. | 4 |