Yuqiu Kong

dblp:185/7895 · DBLP profile ↗
← Back
23ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0003-2168-204XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 11 since 2021Security and privacy · 2 · 2 since 2021
YearPublicationVenuePosition
2026 EipFormer: Enhancing 3D instance segmentation by emphasizing instance positions
Mengnan Zhao 0001, Lihe Zhang, Yuqiu Kong
Expert Syst. Appl.3
2025 PMCFNet: Prompt-Guided Multi-scale Cross-Modal Fusion Network for Referring Remote Sensing Image Segmentation
Yuqiu Kong, Shenglan Liu 0001
PRCV (5)1
2025 Elegantly Written V2: Next-Scale Prediction for Enhancing Online Chinese Handwriting
Yu Liu 0072, Yuqiu Kong, Lei Wang 0233, Cunrui Wang
PRCV (7)2
2025 CollabLearn: Propelling Weakly-Supervised Referring Image Segmentation Through Collaboration Between Semantics and Details
abstract
This work presents a weakly supervised referring image segmentation method, namedCollabLearn, that segments objects described by free-form referring expression utilizing solely image-text pairs. Existing methods suffer from incorrect localization of referring expressions due to the lack of high-level semantics in cross-modal alignment or rough segmentation of referenced objects stemming from the absence of low-level details. To address these issues, we propose an innovative framework for generating cross-modal features encompassing both high-level semantics and low-level details via two fusion modules: a semantic awareness module and a detail cognition module. Each of these modules generates an activation map, and they mutually correct each other through a collaborative learning strategy. Specifically, the semantic awareness module performs in-depth cross-modal interaction and achieves accurate localization in a top-down manner. The detail cognition module facilitates the segmentation of entire objects in a bottom-up manner. A collaborative learning strategy is designed to enable interaction between these two modules, enforcing sufficient vision-language alignment. Experiments on three benchmarks demonstrate that CollabLearn consistently outperforms state-of-the-art weakly supervised methods.
Yuqiu Kong, Mengnan Zhao 0001, Lihe Zhang
IEEE Trans. Multim.2
2024 Catastrophic Overfitting: A Potential Blessing in Disguise
Mengnan Zhao 0001, Lihe Zhang, Yuqiu Kong
ECCV (42)3
2024 Dual Context Perception Transformer for Referring Image Segmentation
Yuqiu Kong, Cuili Yao
PRCV (5)1
2024 Class correlation correction for unbiased scene graph generation
Mengnan Zhao 0001, Yuqiu Kong, Lihe Zhang
Pattern Recognit.2
2024 Adversarial Attacks on Scene Graph Generation
abstract
Scene graph generation (SGG) effectively improves semantic understanding of the visual world. However, the recent interest of researchers focuses on enhancing SGG in non-adversarial settings, which raises our curiosity about the adversarial robustness of SGG models. To bridge this gap, we perform adversarial attacks on two typical SGG tasks, Scene Graph Detection (SGDet) and Scene Graph Classification (SGCls). Specifically, we initially propose a bounding box relabeling method to reconstruct reasonable attack targets for SGCls. It solves the inconsistency between the specified bounding boxes and the scene graphs selected as attack targets. Subsequently, we introduce a two-step weighted attack by removing the predicted objects and relational triples that affect attack performance, which significantly increases the success rate of adversarial attacks on two SGG tasks. Extensive experiments demonstrate the effectiveness of our methods on five popular SGG models and four adversarial attacks. The Pytorch® implementation can be downloaded from an open-source Github project https://github.com/Dlut-lab-zmn/SGG_Attack.
Mengnan Zhao 0001, Lihe Zhang, Wei Wang 0025, Yuqiu Kong
IEEE Trans. Inf. Forensics Secur.4
2023 Referring Image Segmentation Using Text Supervision
abstract
Existing Referring Image Segmentation (RIS) methods typically require expensive pixel-level or box-level annotations for supervision. In this paper, we observe that the referring texts used in RIS already provide sufficient information to localize the target object. Hence, we propose a novel weakly-supervised RIS framework to formulate the target localization problem as a classification process to differentiate between positive and negative text expressions. While the referring text expressions for an image are used as positive expressions, the referring text expressions from other images can be used as negative expressions for this image. Our framework has three main novelties. First, we propose a bilateral prompt method to facilitate the classification process, by harmonizing the domain discrepancy between visual and linguistic features. Second, we propose a calibration method to reduce noisy background information and improve the correctness of the response maps for target object localization. Third, we propose a positive response map selection strategy to generate high-quality pseudo-labels from the enhanced response maps, for training a segmentation network for RIS inference. For evaluation, we propose a new metric to measure localization accuracy. Experiments on four benchmarks show that our framework achieves promising performances to existing fully-supervised RIS methods while outperforming state-of-the-art weakly-supervised methods adapted from related areas. Code is available at https://github.com/fawnliu/TRIS.
Fang Liu 0033, Yuhao Liu 0001, Yuqiu Kong, Ke Xu 0010, Lihe Zhang, Gerhard P. Hancke 0002, Rynson W. H. Lau
ICCV3
2023 Fast Adversarial Training with Smooth Convergence
abstract
Fast adversarial training (FAT) is beneficial for improving the adversarial robustness of neural networks. However, previous FAT work has encountered a significant issue known as catastrophic overfitting when dealing with large perturbation budgets, i.e. the adversarial robustness of models declines to near zero during training. To address this, we analyze the training process of prior FAT work and observe that catastrophic overfitting is accompanied by the appearance of loss convergence outliers. Therefore, we argue a moderately smooth loss convergence process will be a stable FAT process that solves catastrophic overfitting. To obtain a smooth loss convergence process, we propose a novel oscillatory constraint (dubbed ConvergeSmooth) to limit the loss difference between adjacent epochs. The convergence stride of ConvergeSmooth is introduced to balance convergence and smoothing. Likewise, we design weight centralization without introducing additional hyperparameters other than the loss balance coefficient. Our proposed methods are attack-agnostic and thus can improve the training stability of various FAT techniques. Extensive experiments on popular datasets show that the proposed methods efficiently avoid catastrophic overfitting and outperform all previous FAT methods. Code is available at https://github.com/FAT-CS/ConvergeSmooth.
Mengnan Zhao 0001, Lihe Zhang, Yuqiu Kong
ICCV3
2023 Temporal knowledge graph reasoning triggered by memories
Mengnan Zhao 0001, Lihe Zhang, Yuqiu Kong
Appl. Intell.3
2023 Local-global coordination with transformers for referring image segmentation
Fang Liu 0033, Yuqiu Kong, Lihe Zhang
Neurocomputing2
2023 Transformers and CNNs fusion network for salient object detection
Cuili Yao, Yuqiu Kong
Neurocomputing3
2022 Scale Adaptive Fusion Network for RGB-D Salient Object Detection
Yuqiu Kong, Yushuo Zheng, Cuili Yao, Yang Liu 0066
ACCV (3)1
2022 Vision Shared and Representation Isolated Network for Person Search
abstract
Person search is a widely-concerned computer vision task that aims to jointly solve the problems of pedestrian detection and person re-identification in panoramic scenes. However, the pedestrian detection focuses on the consistency of pedestrians, while the person re-identification attempts to extract the discriminative features of pedestrians. The inevitable conflict greatly restricts the researches on the one-stage person search methods. To address this issue, we propose a Vision Shared and Representation Isolated (VSRI) network to decouple the two conflicted subtasks simultaneously, through which two independent representations are constructed for the two subtasks. To enhance the discrimination of the re-ID representation, a Multi-Level Feature Fusion (MLFF) module is proposed. The MLFF adopts the Spatial Pyramid Feature Fusion (SPFF) module to obtain diverse features from the stem network. Moreover, the multi-head self-attention mechanism is employed to construct a Multi-head Attention Driven Extraction (MADE) module and the cascaded convolution unit is adopted to devise a Feature Decomposition and Cascaded Integration (FDCI) module, which facilitates the MLFF to obtain more discriminative representations of the pedestrians. The proposed method outperforms the state-of-the-art methods on the mainstream datasets.
Yang Liu 0066, Yingping Li, Chengyu Kong, Yuqiu Kong, Shenglan Liu 0001
IJCAI4
2022 Background Suppressed and Motion Enhanced Network for Weakly Supervised Video Anomaly Detection
Yang Liu 0066, Wanxiao Yang, Hangyou Yu, Lin Feng 0001, Yuqiu Kong, Shenglan Liu 0001
PRCV (3)5
2022 Double cross-modality progressively guided network for RGB-D salient object detection
Cuili Yao, Lin Feng 0001, Yuqiu Kong, Shengming Li
Image Vis. Comput.3
2022 Guided Erasable Adversarial Attack (GEAA) Toward Shared Data Protection
abstract
In recent years, there has been increasing interest in studying the adversarial attack, which poses potential risks to deep learning applications and has stimulated numerous researches, e.g. improving the robustness of deep neural networks. In this work, we propose a novel double-stream architecture – Guided Erasable Adversarial Attack (GEAA) – for protecting high-quality labeled data with high commercial values under data-sharing scenarios. GEAA contains three phases, the double-stream adversarial attack, denoising reconstruction, and watermark extraction. Specifically, the double-stream adversarial attack injects erasable perturbations into the training data to avoid database abuse. The denoising reconstruction rebuilds the traceable denoising data from adversarial examples. The watermark extraction recovers identity information from the denoised data for copyright protection. Additionally, we introduce the annealing optimization strategy to balance these phases and a boundary constraint to degrade the availability of adversarial examples. Through extensive experiments, we demonstrate the effectiveness of the proposed framework in data protection. The Pytorch® implementations of GEAA can be downloaded from an open-source Github project https://github.com/Dlut-lab-zmn/ GEAA-for-data-protection.
Mengnan Zhao 0001, Bo Wang 0024, Wei Wang 0025, Yuqiu Kong, Tianhang Zheng, Kui Ren 0001
IEEE Trans. Inf. Forensics Secur.4
2021 Label2im: Knowledge Graph Guided Image Generation from Labels
Hewen Xiao, Yuqiu Kong, Hongchen Tan, Xiuping Liu
BMVC2
2021 Bi-DAINet: Bi-Directional Discard-Accept-Integrate Network for salient object detection
Cuili Yao, Lin Feng 0001, Yuqiu Kong, Bo Jin 0001, Leheng Li
Neurocomputing3
2021 Spatial context-aware network for salient object detection
Yuqiu Kong, Mengyang Feng, Xin Li 0003, Huchuan Lu, Xiuping Liu
Pattern Recognit.1
2018 Exemplar-Aided Salient Object Detection via Joint Latent Space Embedding
abstract
Traditional unsupervised salient object detection methods majorly rely on pre-defined assumptions about saliency. However, these assumptions may not be sufficient for handling test images of varied content and context. Meanwhile, supervised models learn saliency knowledge from thousands of annotated images, which are usually expensive to obtain. In this paper, we propose an exemplar-aided salient object detection method, which can complement heuristic saliency assumptions by leveraging only a few exemplar images. This is a challenging task since the appearances between the query images and the exemplars can be quite different. We handle it by learning the matching relationship of the intra-class instances in a latent embedding space in an online fashion. Given a test image and an annotated reference image (retrieved from several exemplar images), our method transfers the foreground and background information of the reference image to the test image via a joint latent embedding of image superpixels. Extensive experiments show that our method can easily improve the performance of existing unsupervised methods even when a very small reference image dataset (e.g. one image) is used. In addition, our method is able to attain competitive performance against fully supervised methods.
Yuqiu Kong, Jianming Zhang 0001, Huchuan Lu, Xiuping Liu
IEEE Trans. Image Process.1
2016 Pattern Mining Saliency
Yuqiu Kong, Lijun Wang 0001, Xiuping Liu, Huchuan Lu, Xiang Ruan
ECCV (6)1