Xiaoguang Di

dblp:237/9979 · DBLP profile ↗
← Back
20ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0002-5709-6862ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
YearPublicationVenuePosition
2026 MAANet: A lightweight multi-axis adaptation network for efficient image super-resolution
Muyan He, Hang Cheng, Xiaoguang Di, Yuyu Ma
Neurocomputing3
2026 Advancing 3D scene generation through GCWA with dynamic upsampling using two-stage diffusion model
Shaoxun Ye, Xiaoguang Di
Multim. Syst.2
2026 L2G-Net: Local-to-global feature enhancement via cluster tokens for 3D place recognition
Ming Liao, Xiaoguang Di, Shaoxun Ye, Mao Zhen Liu 0002
Neural Networks2
2025 Feature-aligned distillation for dense object detection via refined semantic guidance and distribution consistency
Xiaoguang Di, Mao Zhen Liu 0002, Shaoxun Ye
Comput. Vis. Image Underst.2
2025 DRIR-Net: Dual-branch rotation invariant and robust network for 3D place recognition
Ming Liao, Xiaoguang Di, Mao Zhen Liu 0002, Teng Lv, Runwen Zhu
Neurocomputing2
2025 Dynamic-Aware and Static Context Network for large-scale 3D place recognition
Ming Liao, Xiaoguang Di, Mao Zhen Liu 0002, Teng Lv, Runwen Zhu
Knowl. Based Syst.2
2025 DBLDNet: dual branch low light object detector based on feature localization and multi-scale feature enhancement
Xiaoguang Di, Mao Zhen Liu 0002
Multim. Syst.2
2025 Enhancing realism in LiDAR scene generation with CSPA-DFN and linear cross-attention via Diffusion Transformer model
Shaoxun Ye, Xiaoguang Di, Ming Liao
Neural Networks2
2025 Image Inpainting Detection via Dual Guidance of Uncertainty and Precise Boundary Information
abstract
Deep-learning-based image inpainting technology has achieved remarkable visual consistency but is vulnerable to malicious use. Existing detection methods overlook semantic inconsistencies between targets and backgrounds, leading to ambiguous results due to low discriminability. To tackle these challenges, we draw inspiration from human strategies in visual tasks, which involve initially assigning uncertainty across the entire input and subsequently concentrating on highly uncertain regions using prior knowledge like boundary information. Building on this, we propose a Dual Information Guided Network (DIGNet). It combines object-background semantic modulation with uncertainty to precisely locate inpainting regions. This is the first work to address inpainting prediction inaccuracies by considering both edge uncertainty and semantic inconsistency. DIGNet consists of three key parts: the Edge Uncertainty Awareness Module (EUAM), the Edge Correction Module (ECM) based on semantic differences, and the Dual Information Guided Interaction Module (DIGIM). We use semantic inconsistency to get edge constraints and quantify uncertainty as feature variance to guide mainstream feature maps. The DIGIM effectively fuses guide information for accurate predictions. Comprehensive experiments show that our method outperforms existing CNN-based approaches. Specifically, it improves the F1 Score by at least 0.31% and the IOU by at least 0.19% on multiple datasets.
Mao Zhen Liu 0002, Xiaoguang Di, Ming Liao
IEEE Trans. Circuits Syst. Video Technol.2
2024 Towards to Human Intention: A few-shot open-set object detection for X-ray hazard inspection
Mao Zhen Liu 0002, Xiaoguang Di, Teng Lv, Ming Liao
Neurocomputing2
2024 Multi-level Symmetric Semantic Alignment Network for image-text matching
Wenzhuang Wang, Xiaoguang Di, Mao Zhen Liu 0002
Neurocomputing2
2024 HDNet: Human-like discrimination with visual key for few-shot cross-domain object detection
Mao Zhen Liu 0002, Xiaoguang Di, Wenzhuang Wang
Knowl. Based Syst.2
2023 Extraordinary MHNet: Military high-level camouflage object detection network and dataset
Mao Zhen Liu 0002, Xiaoguang Di
Neurocomputing2
2022 Global Context Parallel Attention for Anchor-Free Instance Segmentation in Remote Sensing Images
abstract
Segmenting objects in optical remote sensing images has always been a hot topic for remote sensing image researchers. However, many previous works used segmentation algorithms designed for common objects without modification, leading to slow and poor results. In this work, we exploit self-attention mechanism into anchor-free segmentation architectures to improve the segmentation accuracy for objects in high-resolution remote sensing images. The proposed module integrates the self-attention mechanism, namely the global context parallel attention module (GC-PAM). It is composed of a parallel global context channel self-attention block and a spatial self-attention block. By implementing our GC-PAM in an anchor-free network, the channel-wise and spatial-wise weights are both reassigned, which can improve the segmentation accuracy significantly.
Xinyu Liu 0001, Xiaoguang Di
IEEE Geosci. Remote. Sens. Lett.2
2022 Better Than Reference in Low-Light Image Enhancement: Conditional Re-Enhancement Network
abstract
Low-light images suffer from severe noise, low brightness, low contrast, etc. In previous researches, many image enhancement methods have been proposed, but few methods can deal with these problems simultaneously. In this paper, to solve these problems simultaneously, we propose a low-light image enhancement method that can be combined with supervised learning and previous HSV (Hue, Saturation, Value) or Retinex model-based image enhancement methods. First, we analyse the relationship between the HSV color space and the Retinex theory, and show that the V channel (V channel in HSV color space, equals the maximum channel in RGB color space) of the enhanced image can well represent the contrast and brightness enhancement process. Then, a data-driven conditional re-enhancement network (denoted as CRENet) is proposed. The network takes low-light images as input and the enhanced V channel (V channel of the enhanced image) as a condition during testing, and then it can re-enhance the contrast and brightness of the low-light image and at the same time reduce noise and color distortion. In addition, it takes 23 ms to process a color image with the resolution 400*600 on a 1080Ti GPU. Finally, some comparative experiments are implemented to prove the effectiveness of the method. The results show that the method proposed in this paper can significantly improve the quality of the enhanced image, and by combining it with other image contrast enhancement methods, the final enhancement result can even be better than the reference image in contrast and brightness when the contrast and brightness of the reference are not good.
Yu Zhang 0091, Xiaoguang Di, Ruihang Ji
IEEE Trans. Image Process.2
2021 TanhExp: A smooth activation function with high convergence speed for lightweight neural networks
abstract
Abstract Lightweight or mobile neural networks used for real‐time computer vision tasks contain fewer parameters than normal networks, which lead to a constrained performance. Herein, a novel activation function named as Tanh Exponential Activation Function (TanhExp) is proposed which can improve the performance for these networks on image classification task significantly. The definition of TanhExp is f ( x ) = x tanh( e x ). The simplicity, efficiency, and robustness of TanhExp on various datasets and network models is demonstrated and TanhExp outperforms its counterparts in both convergence speed and accuracy. Its behaviour also remains stable even with noise added and dataset altered. It is shown that without increasing the size of the network, the capacity of lightweight neural networks can be enhanced by TanhExp with only a few training epochs and no extra parameters added.
Xinyu Liu 0001, Xiaoguang Di
IET Comput. Vis.2
2020 Leveraging Undiagnosed Data for Glaucoma Classification with Teacher-Student Learning
Wenting Chen, Kai Ma 0002, Hanruo Liu, Xiaoguang Di, Yefeng Zheng 0001
MICCAI (1)7
2020 Learning an adaptive model for extreme low-light raw image processing
abstract
Low‐light images suffer from severe noise and low illumination. In this work, the authors propose an adaptive low‐light raw image enhancement network to avoid parameter‐handcrafting in current deep learning models and to improve image quality. The proposed method can be divided into two sub‐models: brightness prediction and exposure shifting (ES). The former is designed to control the brightness of the resulting image by estimating a guideline exposure time . The latter learns to approximate an exposure‐shifting operator ES, converting a low‐light image with real exposure time to a noise‐free image with guideline exposure time . Additionally, structural similarity loss and image enhancement vector are introduced to promote image quality, and a new campus image data set (CID) for training the proposed model is proposed to overcome the limitations of the existing data sets. In quantitative tests, it is shown that the proposed method has the lowest noise level estimation score compared with the state‐of‐the‐art low‐light algorithms, suggesting a superior denoising performance. Furthermore, those tests illustrate that the proposed method is able to adaptively control the global image brightness according to the content of the image scene. Lastly, the potential application in video processing is briefly discussed.
Qingxu Fu, Xiaoguang Di, Yu Zhang 0091
IET Image Process.2
2020 A novel convolutional neural network method for crowd counting
abstract
Crowd density estimation, in general, is a challenging task due to the large variation of head sizes in the crowds. Existing methods always use a multi-column convolutional neural network (MCNN) to adapt to this variation, which results in an average effect in areas with different densities and brings a lot of noise to the density map. To address this problem, we propose a new method called the segmentation-aware prior network (SAPNet), which generates a high-quality density map without noise based on a coarse head-segmentation map. SAPNet is composed of two networks, i.e., a foreground-segmentation convolutional neural network (FS-CNN) as the front end and a crowd-regression convolutional neural network (CR-CNN) as the back end. With only the single dot annotation, we generate the ground truth of segmentation masks in heads. Then, based on the ground truth, FS-CNN outputs a coarse head-segmentation map, which helps eliminate the noise in regions without people in the density map. By inputting the head-segmentation map generated by the front end, CR-CNN performs accurate crowd counting estimation and generates a high-quality density map. We demonstrate SAPNet on four datasets (i.e., ShanghaiTech, UCF-CC-50, WorldExpo’10, and UCSD), and show the state-of-the-art performances on ShanghaiTech part B and UCF-CC-50 datasets.
Jiehao Huang, Xiaoguang Di, Ai-yue Chen
Frontiers Inf. Technol. Electron. Eng.2
2020 Integrating Neural Networks Into the Blind Deblurring Framework to Compete With the End-to-End Learning-Based Methods
abstract
Recently, the end-to-end learning-based methods have been proven effective for the blind image deblurring. Without human-made assumptions or numerical algorithms, they are able to restore images with fewer artifacts and better perceptual quality. However, in practice, these methods suffer from limited performance under complex motion scenario and produces unnatural results sometimes. In this paper, in order to overcome their limitations, we propose to integrate deep convolution neural networks into a conventional deblurring framework. Specifically, we propose Stacked Estimation Residual Net (SEN) to estimate the motion flow map and Recurrent Prior Generative and Adversarial Net (RP-GAN) to learn the implicit image prior for the optimization. Comparing with the state-of-the-art end-to-end learning-based methods, the proposed method restores image content more naturally and shows better generalization ability.
Xiaoguang Di
IEEE Trans. Image Process.2