Qingji Guan

dblp:213/8038 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0002-0155-940XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 CLIP-driven Coarse-to-fine Semantic Guidance for Fine-grained Open-set Semi-supervised Learning
abstract
Fine-grained open-set semi-supervised learning (OSSL) investigates a practical scenario where unlabeled data may contain fine-grained out-of-distribution (OOD) samples. Due to the subtle visual differences among in-distribution (ID) samples, as well as between ID and OOD samples, it is extremely challenging to separate the ID and OOD samples. Recent Vision-Language Models, such as CLIP, have shown excellent generalization capabilities. However, it tends to focus on general attributes, and thus is insufficient to distinguish the fine-grained details. To tackle the issues, in this paper, we propose a novel CLIP-driven coarse-to-fine semantic-guided framework, named CFSG-CLIP, to progressively focus on the distinctive fine-grained clues. Specifically, CFSG-CLIP comprises a coarse-guidance branch and a fine-guidance branch derived from the pre-trained CLIP model. In the coarse-guidance branch, we design a semantic filtering module to initially filter and highlight local visual features guided by cross-modality features. Then, in the fine-guidance branch, we further design a visual-semantic injection strategy, which embeds category-related visual cues into the visual encoder to further refine the local visual features. By the designed dual-guidance framework, local subtle cues are progressively discovered to distinct the subtle difference between ID and OOD samples. Extensive experiments demonstrate that CFSG-CLIP achieves competitive performance on multiple fine-grained datasets. The source code is available at https://github.com/LxxxxK/CFSG-CLIP.
Xiaokun Li, Qingji Guan
CVPR3
2025 RINDNet++: Edge Detection for Discontinuity in Reflectance, Illumination, Normal, and Depth
Mengyang Pu, Qingji Guan, Haibin Ling
Int. J. Comput. Vis.3
2025 Improving Image Inpainting via Adversarial Collaborative Training
abstract
Image inpainting aims to restore visually realistic contents from a corrupted image, while inpainting forensic methods focus on locating the inpainted regions to fight against inpainting manipulations. Motivated by these two mutually interdependent tasks, in this paper, we propose a novel image inpainting network called Adversarial Collaborative Network (AdvColabNet), which leverages the contradictory and collaborative information from the two tasks of image inpainting and inpainting forensics to enhance the progress of the inpainting model through adversarial collaborative training. Specifically, the proposed AdvColabNet is a coarse-to-fine two-stage framework. In the coarse training stage, a simple generative adversarial model-based U-Net-style network generates initial coarse inpainting results. In the fine stage, the authenticity of inpainting results is assessed using the estimated forensic mask. A forensics-driven adaptive weighting refinement strategy is developed to emphasize learning from pixels with higher probabilities of being inpainted, which helps the network to focus on the challenging regions, resulting in more plausible inpainting results. Comprehensive evaluations on the CelebA-HQ and Places2 datasets demonstrate that our method achieves state-of-the-art robustness performance in terms of PSNR, SSIM, MAE, FID, and LPIPS metrics. We also show that our method effectively deceives the proposed inpainting forensic method compared to state-of-the-art inpainting methods, further demonstrating the superiority of the proposed method.
Qingji Guan
IEEE Trans. Multim.3
2025 ReFrame: A Resource-Friendly Cloud-Assisted On-Device Deep Learning Framework for Vision Services
abstract
Cloud-assisted Internet of Things (IoT) device deployment of deep neural networks (DNNs) promotes On-device deep learning to provide users with ubiquitous high-quality services by solving the contradiction between insufficient IoT device resources and intensive demand for high-performance DNN resources. However, most existing methods optimize DNNs by considering one or two terms of transmission, computation, and storage resources, but do not consider all three terms at the same time in cloud-assisted IoT device deployment and updating DNNs. To this end, we propose a non-learnable module-based ResNet and a cloud-assisted on-device deep learning framework, ReFrame, based on the consideration of three indicators: model transmission parameters, computation resources, and storage resources. In the proposed method, we first specify that some parameters in DNNs are non-learnable and randomly initialized, so that, these parameters can be saved and reproduced with a few random seeds. By doing so, the cloud only transmits random seeds and learnable parameters to reduce the number of parameter transmissions. Second, we reduce the computation resource consumption of the model by introducing computation-friendly operators, such as pooling, to replace vanilla convolutions. Finally, since random seeds are used to save non-learnable model parameters, on IoT devices we only need to store random seeds and learnable parameters to reproduce the well-trained model. Compared with saving the complete model, our method greatly reduces IoT device storage resource consumption. Experimental results on image classification, object detection, and semantic segmentation tasks demonstrate the effectiveness of the proposed method. Specifically, on the CIFAR-10, our proposed method reduces approximately 89% of FLOPs and 90% of transmitted data in the prototype system compared to ResNet-18.
Jianhang Xie, Chuntao Ding, Qingji Guan, Ao Zhou 0001, Yidong Li
IEEE Trans. Serv. Comput.3
2024 MuGE: Multiple Granularity Edge Detection
abstract
Edge segmentation is well-known to be subjective due to personalized annotation styles and preferred granular-ity. However, most existing deterministic edge detection methods produce only a single edge map for one input image. We argue that generating multiple edge maps is more reasonable than generating a single one considering the subjectivity and ambiguity of the edges. Thus motivated, in this paper we propose multiple granularity edge detection, called MuGE, which can produce a wide range of edge maps, from approximate object contours to fine texture edges. Specifically, we first propose to design an edge granularity network to estimate the edge granularity from an individual edge annotation. Subsequently, to guide the generation of diversified edge maps, we integrate such edge granularity into the multi-scale feature maps in the spatial domain. Meanwhile, we decompose the feature maps into low-frequency and high-frequency parts, where the encoded edge granularity is further fused into the high-frequency part to achieve more precise control over the details of the produced edge maps. Compared to previous methods, MuGE is able to not only generate multiple edge maps at different controllable granularities but also achieve a com-petitive performance on the BSDS500 and Multicue benchmark datasets.
Caixia Zhou, Mengyang Pu, Qingji Guan, Ruoxi Deng, Haibin Ling
CVPR4
2023 The Treasure Beneath Multiple Annotations: An Uncertainty-Aware Edge Detector
abstract
Deep learning-based edge detectors heavily rely on pixel-wise labels which are often provided by multiple annotators. Existing methods fuse multiple annotations using a simple voting process, ignoring the inherent ambiguity of edges and labeling bias of annotators. In this paper, we propose a novel uncertainty-aware edge detector (UAED), which employs uncertainty to investigate the subjectivity and ambiguity of diverse annotations. Specifically, we first convert the deterministic label space into a learnable Gaussian distribution, whose variance measures the degree of ambiguity among different annotations. Then we regard the learned variance as the estimated uncertainty of the predicted edge maps, and pixels with higher uncertainty are likely to be hard samples for edge detection. Therefore we design an adaptive weighting loss to emphasize the learning from those pixels with high uncertainty, which helps the network to gradually concentrate on the important pixels. UAED can be combined with various encoder-decoder backbones, and the extensive experiments demonstrate that UAED achieves superior performance consistently across multiple edge detection benchmarks. The source code is available at https://github.com/ZhouCX117/UAED.
Caixia Zhou, Mengyang Pu, Qingji Guan, Haibin Ling
CVPR4
2023 Joint representation and classifier learning for long-tailed image classification
Qingji Guan, Zhuangzhuang Li
Image Vis. Comput.1
2023 A Hard Knowledge Regularization Method with Probability Difference in Thorax Disease Images
Qingji Guan, Qinrun Chen, Zhun Zhong
Knowl. Based Syst.1
2022 EDTER: Edge Detection with Transformer
abstract
Convolutional neural networks have made significant progresses in edge detection by progressively exploring the context and semantic features. However, local details are gradually suppressed with the enlarging of receptive fields. Recently, vision transformer has shown excellent capability in capturing long-range dependencies. Inspired by this, we propose a novel transformer-based edge detector, Edge Detection TransformER (EDTER), to extract clear and crisp object boundaries and meaningful edges by exploiting the full image context information and detailed local cues simultaneously. EDTER works in two stages. In Stage I, a global transformer encoder is used to capture long-range global context on coarse-grained image patches. Then in Stage II, a local transformer encoder works on fine-grained patches to excavate the short-range local cues. Each transformer encoder is followed by an elaborately designed Bi-directional Multi-Level Aggregation decoder to achieve high-resolution features. Finally, the global context and local cues are combined by a Feature Fusion Module and fed into a decision head for edge prediction. Extensive experiments on BSDS500, NYUDv2, and Multicue demonstrate the superiority of EDTER in comparison with state-of-the-arts. The source code is available at https://github.com/MengyangPu/EDTER.
Mengyang Pu, Qingji Guan, Haibin Ling
CVPR4
2022 Learning From Pixel-Level Label Noise: A New Perspective for Semi-Supervised Semantic Segmentation
abstract
This paper addresses semi-supervised semantic segmentation by exploiting a small set of images with pixel-level annotations (strong supervisions) and a large set of images with only image-level annotations (weak supervisions). Most existing approaches aim to generate accurate pixel-level labels from weak supervisions. However, we observe that those generated labels still inevitably contain noisy labels. Motivated by this observation, we present a novel perspective and formulate this task as a problem of learning with pixel-level label noise. Existing noisy label methods, nevertheless, mainly aim at image-level tasks, which can not capture the relationship between neighboring labels in one image. Therefore, we propose a graph-based label noise detection and correction framework to deal with pixel-level noisy labels. In particular, for the generated pixel-level noisy labels from weak supervisions by Class Activation Map (CAM), we train a clean segmentation model with strong supervisions to detect the clean labels from these noisy labels according to the cross-entropy loss. Then, we adopt a superpixel-based graph to represent the relations of spatial adjacency and semantic similarity between pixels in one image. Finally we correct the noisy labels using a Graph Attention Network (GAT) supervised by detected clean labels. We comprehensively conduct experiments on PASCAL VOC 2012, PASCAL-Context, MS-COCO and Cityscapes datasets. The experimental results show that our proposed semi-supervised method achieves the state-of-the-art performances and even outperforms the fully-supervised models on PASCAL VOC 2012 and MS-COCO datasets in some cases.
Rumeng Yi, Qingji Guan, Mengyang Pu, Runsheng Zhang
IEEE Trans. Image Process.3
2021 RINDNet: Edge Detection for Discontinuity in Reflectance, Illumination, Normal and Depth
abstract
As a fundamental building block in computer vision, edges can be categorised into four types according to the discontinuity in surface-Reflectance, Illumination, surface-Normal or Depth. While great progress has been made in detecting generic or individual types of edges, it remains under-explored to comprehensively study all four edge types together. In this paper, we propose a novel neural network solution, RINDNet, to jointly detect all four types of edges. Taking into consideration the distinct attributes of each type of edges and the relationship between them, RINDNet learns effective representations for each of them and works in three stages. In stage I, RINDNet uses a common backbone to extract features shared by all edges. Then in stage II it branches to prepare discriminative features for each edge type by the corresponding decoder. In stage III, an independent decision head for each type aggregates the features from previous stages to predict the initial results. Additionally, an attention module learns attention maps for all types to capture the underlying relations between them, and these maps are combined with initial results to generate the final edge detection results. For training and evaluation, we construct the first public benchmark, BSDS-RIND, with all four types of edges carefully annotated. In our experiments, RINDNet yields promising results in comparison with state-of-the-art methods. Additional analysis is presented in supplementary material.
Mengyang Pu, Qingji Guan, Haibin Ling
ICCV3
2021 Discriminative Feature Learning for Thorax Disease Classification in Chest X-ray Images
abstract
This paper focuses on the thorax disease classification problem in chest X-ray (CXR) images. Different from the generic image classification task, a robust and stable CXR image analysis system should consider the unique characteristics of CXR images. Particularly, it should be able to: 1) automatically focus on the disease-critical regions, which usually are of small sizes; 2) adaptively capture the intrinsic relationships among different disease features and utilize them to boost the multi-label disease recognition rates jointly. In this paper, we propose to learn discriminative features with a two-branch architecture, named ConsultNet, to achieve those two purposes simultaneously. ConsultNet consists of two components. First, an information bottleneck constrained feature selector extracts critical disease-specific features according to the feature importance. Second, a spatial-and-channel encoding based feature integrator enhances the latent semantic dependencies in the feature space. ConsultNet fuses these discriminative features to improve the performance of thorax disease classification in CXRs. Experiments conducted on the ChestX-ray14 and CheXpert dataset demonstrate the effectiveness of the proposed method.
Qingji Guan, Yawei Luo, Ping Liu 0004, Mingliang Xu 0001, Yi Yang 0001
IEEE Trans. Image Process.1
2020 Looking ahead: Joint small group detection and tracking in crowd scenes
Qiulin Ma, Qi Zou 0001, Qingji Guan, Yanting Pei
J. Vis. Commun. Image Represent.4
2020 Multi-label chest X-ray image classification via category-wise residual attention learning
Qingji Guan
Pattern Recognit. Lett.1
2020 Thorax disease classification with attention guided convolutional neural network
Qingji Guan, Zhun Zhong, Zhedong Zheng, Liang Zheng 0001, Yi Yang 0001
Pattern Recognit. Lett.1
2020 Object Discovery From a Single Unlabeled Image by Mining Frequent Itemsets With Multi-Scale Features
abstract
The goal of our work is to discover dominant objects in a very general setting where only a single unlabeled image is given. This is far more challenge than typical colocalization or weakly-supervised localization tasks. To tackle this problem, we propose a simple but effective pattern mining-based method, called Object Location Mining (OLM), which exploits the advantages of data mining and feature representation of pretrained convolutional neural networks (CNNs). Specifically, we first convert the feature maps from a pre-trained CNN model into a set of transactions, and then discovers frequent patterns from transaction database through pattern mining techniques. We observe that those discovered patterns, i.e., co-occurrence highlighted regions, typically hold appearance and spatial consistency. Motivated by this observation, we can easily discover and localize possible objects by merging relevant meaningful patterns. Extensive experiments on a variety of benchmarks demonstrate that OLM achieves competitive localization performance compared with the state-of-the-art methods. We also evaluate our approach compared with unsupervised saliency detection methods and achieves competitive results on seven benchmark datasets. Moreover, we conduct experiments on finegrained classification to show that our proposed method can locate the entire object and parts accurately, which can benefit to improving the classification results significantly.
Runsheng Zhang, Mengyang Pu, Qingji Guan, Qi Zou 0001, Haibin Ling
IEEE Trans. Image Process.5
2018 GraphNet: Learning Image Pseudo Annotations for Weakly-Supervised Semantic Segmentation
abstract
Weakly-supervised semantic image segmentation suffers from lacking accurate pixel-level annotations. In this paper, we propose a novel graph convolutional network-based method, called GraphNet, to learn pixel-wise labels from weak annotations. Firstly, we construct a graph on the superpixels of a training image by combining the low-level spatial relation and high-level semantic content. Meanwhile, scribble or bounding box annotations are embedded into the graph, respectively. Then, GraphNet takes the graph as input and learns to predict high-confidence pseudo image masks by a convolutional network operating directly on graphs. At last, a segmentation network is trained supervised by these pseudo image masks. We comprehensively conduct experiments on the PASCAL VOC 2012 and PASCAL-CONTEXT segmentation benchmarks. Experimental results demonstrate that GraphNet is effective to predict the pixel labels with scribble or bounding box annotations. The proposed framework yields state-of-the-art results in the community.
Mengyang Pu, Qingji Guan, Qi Zou 0001
ACM Multimedia3