Qirui Wu

dblp:271/3136 · DBLP profile ↗
← Back
19ranked-venue papers
11as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Better Matching, Less Forgetting: A Quality-Guided Matcher for Transformer-based Incremental Object Detection
abstract
Incremental Object Detection (IOD) aims to continuously learn new object classes without forgetting previously learned ones. A persistent challenge is catastrophic forgetting, primarily attributed to background shift in conventional detectors. While pseudo-labeling mitigates this in dense detectors, we identify a novel, distinct source of forgetting specific to DETR-like architectures: background foregrounding. This arises from the exhaustiveness constraint of the Hungarian matcher, which forcibly assigns every ground truth target to one prediction, even when predictions primarily cover background regions (i.e., low IoU). This erroneous supervision compels the model to misclassify background features as specific foreground classes, disrupting learned representations and accelerating forgetting. To address this, we propose a Quality-guided Min-Cost Max-Flow (Q-MCMF) matcher. To avoid forced assignments, Q-MCMF builds a flow graph and prunes implausible matches based on geometric quality. It then optimizes for the final matching that minimizes cost and maximizes valid assignments. This strategy eliminates harmful supervision from background foregrounding while maximizing foreground learning signals. Extensive experiments on the COCO dataset under various incremental settings demonstrate that our method consistently outperforms existing state-of-the-art approaches.
Qirui Wu, Shizhou Zhang, De Cheng, Yinghui Xing, Lingyan Ran, Dahu Shi, Peng Wang 0015
AAAI1
2026 YOLO-IOD: Towards Real Time Incremental Object Detection
abstract
Current methodologies for incremental object detection (IOD) primarily rely on Faster R-CNN or DETR series detectors; however, these approaches do not accommodate the real-time YOLO detection frameworks. In this paper, we first identify three primary types of knowledge conflicts that contribute to catastrophic forgetting in YOLO-based incremental detectors: foreground-background confusion, parameter interference, and misaligned knowledge distillation. Subsequently, we introduce YOLO-IOD, a real-time Incremental Object Detection (IOD) framework that is constructed upon the pretrained YOLO-World model, facilitating incremental learning via a stage-wise parameter-efficient finetuning process. Specifically, YOLO-IOD encompasses three principal components: 1) Conflict-Aware Pseudo-Label Refinement (CPR), which mitigates the foreground-background confusion by leveraging the confidence levels of pseudo labels and identifying potential objects relevant to future tasks. 2) Importance-based Kernel Selection (IKS), which identifies and updates the pivotal convolution kernels pertinent to the current task during the current learning stage. 3)Cross-Stage Asymmetric Knowledge Distillation (CAKD), which addresses the misaligned knowledge distillation conflict by transmitting the features of the student target detector through the detection heads of both the previous and current teacher detectors, thereby facilitating asymmetric distillation between existing and newly introduced categories. We further introduce LoCo COCO, a more realistic benchmark that eliminates data leakage across stages. Experiments on both conventional and LoCo COCO benchmarks show that YOLO-IOD achieves superior performance with minimal forgetting.
Shizhou Zhang, Xueqiang Lv, Yinghui Xing, Qirui Wu, Di Xu 0010, Chen Zhao 0009, Yanning Zhang 0001
AAAI4
2026 Multi-type and multi-scale geological three-dimensional modeling using entity-relationship networks
Qirui Wu, Zhong Xie, Qinjun Qiu
Adv. Eng. Informatics1
2026 Distributed Maneuvering Target Tracking via Transformer-Based IMM and Adaptive Information Consensus Under Sensor Degradation
Qirui Wu, Heng Deng, Liguo Zhang 0001
IEEE Trans Autom. Sci. Eng.1
2025 Revisiting Generative Replay for Class Incremental Object Detection
abstract
Generative replay has gained significant attention in class-incremental learning; however, its application to Class Incremental Object Detection (CIOD) remains limited due to the challenges in generating complex images with precise spatial arrangements. In this study, motivated by the observation that the forgetting of prior knowledge is predominantly present in the classification sub-task as opposed to the localization sub-task, we revisit the generative replay method for class incremental object detection. Our method utilize a standard Stable Diffusion model to generate image-level replay data for all old and new tasks. Accordingly, the old detector and a stage-wise detector are conducted on the synthetic images respectively to determine the bounding box positions through pseudo-labeling. Furthermore, we propose to use a Similarity-based Cross Sampling mechanism to select valuable confusing data between old and new tasks to more effectively mitigate catastrophic forgetting and reduce the false alarm rate for the new task. Finally, all synthetic and real data are integrated for current-stage detector training, where the images generated for previous tasks are highly beneficial in minimizing the forgetting of existing knowledge, while those synthesized for the new task can help bridge the domain gap between real and synthetic images. We conducted extensive experiments on PASCAL VOC 2007 and MS COCO benchmark datasets in multiple settings to showcase the efficacy of our proposed approach, which achieves state-of-the-art results. The code is available at https://github.com/qiangzailv/RGR-IOD.
Shizhou Zhang, Xueqiang Lv, Yinghui Xing, Qirui Wu, Di Xu 0010, Yanning Zhang 0001
CVPR4
2025 Diorama: Unleashing Zero-Shot Single-View 3D Indoor Scene Modeling
Qirui Wu, Denys Iliash, Daniel Ritchie 0001, Manolis Savva, Angel X. Chang
ICCV1
2025 Demystifying Catastrophic Forgetting in Two-Stage Incremental Object Detector
abstract
Catastrophic forgetting is a critical chanllenge for incremental object detection (IOD). Most existing methods treat the detector monolithically, relying on instance replay or knowledge distillation without analyzing component-specific forgetting. Through dissection of Faster R-CNN, we reveal a key insight: Catastrophic forgetting is predominantly localized to the RoI Head classifier, while regressors retain robustness across incremental stages. This finding challenges conventional assumptions, motivating us to develop a framework termed NSGP-RePRE. Regional Prototype Replay (RePRE) mitigates classifier forgetting via replay of two types of prototypes: coarse prototypes represent class-wise semantic centers of RoI features, while fine-grained prototypes model intra-class variations. Null Space Gradient Projection (NSGP) is further introduced to eliminate prototype-feature misalignment by updating the feature extractor in directions orthogonal to subspace of old inputs via gradient projection, aligning RePRE with incremental learning dynamics. Our simple yet effective design allows NSGP-RePRE to achieve state-of-the-art performance on the Pascal VOC and MS COCO datasets under various settings. Our work not only advances IOD methodology but also provide pivotal insights for catastrophic forgetting mitigation in IOD. Code will be available soon.
Qirui Wu, Shizhou Zhang, De Cheng, Yinghui Xing, Di Xu 0010, Peng Wang 0015, Yanning Zhang 0001
ICML1
2025 Eliminating Retrieval Knowledge Conflicts: Cross-Validation Re-ranking with Large Language Models
abstract
In retrieval-augmented generation (RAG) systems, Large Language Models (LLMs) have been shown to be effective for re-ranking. However, existing research often prioritizes passage relevance over reliability, which can result in the incorporation of conflicting information and the generation of ambiguous responses. This issue becomes particularly pronounced when addressing inter-context knowledge conflicts, where candidate documents present contradictory information that may mislead the model. To mitigate this problem, we propose a novel cross-validation re-ranking technique designed specifically to resolve inter-context knowledge conflicts during the retrieval process. We also develop a new dataset, ContraPRT, to evaluate the ability of models to rank passages containing conflicting knowledge. Experimental results using GPT-4 and LlaMA3-70B demonstrate that our approach not only effectively filters out conflicting information but also ensures accurate passage rankings, thereby providing reliable supplementary knowledge for the generation module.
Qirui Wu, Lin Hai, Hai-Tao Zheng 0002, Ruobing Xie, Saiyong Yang, Xingwu Sun, Zhanhui Kang, Hong-Gee Kim
IJCNN1
2025 A broken-track association method for robust multi-target tracking adopting multi-view Doppler measurement information
Cao Zeng, Haihong Tao, Yuhong Zhang 0001, Shihua Zhao, Qirui Wu
Signal Process.6
2024 Generalizing Single-View 3D Shape Retrieval to Occlusions and Unseen Objects
abstract
Single-view 3D shape retrieval is a challenging task that is increasingly important with the growth of available 3D data. Prior work that has studied this task has not focused on evaluating how realistic occlusions impact performance, and how shape retrieval methods generalize to scenarios where either the target 3D shape database contains unseen shapes, or the input image contains unseen objects. In this paper, we systematically evaluate single-view 3D shape retrieval along three different axes: the presence of object occlusions and truncations, generalization to unseen 3D shape data, and generalization to unseen objects in the input images. We standardize two existing datasets of real images and propose a dataset generation pipeline to produce a synthetic dataset of scenes with multiple objects exhibiting realistic occlusions. Our experiments show that training on occlusion-free data as was commonly done in prior work leads to significant performance degradation for inputs with occlusion. We find that that by first pretraining on our synthetic dataset with occlusions and then finetuning on real data, we can significantly outperform models from prior work and demonstrate robustness to both unseen 3D shapes and unseen objects.
Qirui Wu, Daniel Ritchie 0001, Manolis Savva, Angel X. Chang
3DV1
2024 R3DS: Reality-Linked 3D Scenes for Panoramic Scene Understanding
Qirui Wu, Sonia Raychaudhuri, Daniel Ritchie 0001, Manolis Savva, Angel X. Chang
ECCV (63)1
2024 Dual-Branch Task Residual Enhancement with Parameter-Free Attention for Zero-Shot Multi-label Image Recognition
Shizhou Zhang, Kairui Dang, De Cheng, Yinghui Xing, Qirui Wu, Dexuan Kong, Yanning Zhang 0001
ICPR (22)5
2024 Dual Modality Prompt Tuning for Vision-Language Pre-Trained Model
abstract
With the emergence of large pretrained vison-language models such as CLIP, transferable representations can be adapted to a wide range of downstream tasks via prompt tuning. Prompt tuning probes for beneficial information for downstream tasks from the general knowledge stored in the pretrained model. A recently proposed method named Context Optimization (CoOp) introduces a set of learnable vectors as text prompts from the language side. However, tuning the text prompt alone can only adjust the synthesized “classifier”, while the computed visual features of the image encoder cannot be affected, thus leading to suboptimal solutions. In this article, we propose a novel dual-modality prompt tuning (DPT) paradigm through learning text and visual prompts simultaneously. To make the final image feature concentrate more on the target visual concept, a class-aware visual prompt tuning (CAVPT) scheme is further proposed in our DPT. In this scheme, the class-aware visual prompt is generated dynamically by performing the cross attention between text prompt features and image patch token embeddings to encode both the downstream task-related information and visual instance information. Extensive experimental results on 11 datasets demonstrate the effectiveness and generalization ability of the proposed method.
Yinghui Xing, Qirui Wu, De Cheng, Shizhou Zhang, Guoqiang Liang 0001, Peng Wang 0015, Yanning Zhang 0001
IEEE Trans. Multim.2
2024 Deep Channel Prediction-Based Energy-Efficient Intelligent Reflecting Surface-Aided Terahertz Communications
abstract
We propose a novel deep learning-based algorithm for channel prediction and energy efficiency (EE) optimisation in an intelligent reflecting surface (IRS) aided Terahertz communication system. Specifically, a multi-antenna base station with an IRS with massive reflecting elements is designed to serve multiple moving users. A deep learning-based prediction-optimisation scheme is presented where we first propose a transformer encoder with channel index embedding (TE-CIE) deep learning model for time-varying channel prediction. With the help of channel prediction, an EE optimisation algorithm is then designed to maximise the EE in advance based on the predicted channel state information (CSI). Finally, we combine both methods to construct a deep learning-based prediction-optimisation scheme. Specifically, the future CSI is predicted by TE-CIE and the IRS phase-shift and precoding matrices are optimised in advance. Simulation results demonstrate that our proposed channel prediction method achieves close-to-optimal performance in terms of low mean absolute error and a much faster speed than the two baseline models. We demonstrate that the proposed EE optimisation algorithm outperforms the baseline algorithms in terms of much better EE under diverse parameter settings. Finally, the proposed prediction-optimisation scheme achieves at least twice the EE improvement compared to the baseline methods in the literature.
Qirui Wu, Yirun Zhang, Zhaohui Yang 0001, Mohammad Shikh-Bahaei
IEEE Trans. Wirel. Commun.1
2024 Deep Learning for Secure UAV Swarm Communication Under Malicious Attacks
abstract
Unmanned aerial vehicle (UAV) swarms have become a promising solution to enhance modern wireless communication in complicated environments. However, due to the existence of real-world malicious attacks, the performance of prediction and optimisation methods used for UAV swarms are easily degraded. In this paper, we propose a novel deep learning-based user mobility prediction, user assignment and drone position optimisation scheme for a UAV swarm-enabled wireless communication system in the presence of malicious Global Navigation Satellite System (GNSS) spoofing attackers. Specifically, a robust deep learning-based user mobility prediction model, namely denoising autoencoder recurrent transformer (DART), is designed. Additionally, two efficient user assignment and drone position optimisation methods are proposed. The proposed deep learning model forecasts user locations, on which we construct and solve assignment and position optimisation problems. Simulation results show that the proposed deep learning-based prediction-optimisation scheme can provide up to 30% higher overall sum rate compared with the adversarially trained long short-term memory (LSTM) baseline and almost double the overall sum rate compared with the vanilla LSTM baseline.
Qirui Wu, Yirun Zhang, Zhaohui Yang 0001, Mohammad Shikh-Bahaei
IEEE Trans. Wirel. Commun.1
2022 D3Net: A Unified Speaker-Listener Architecture for 3D Dense Captioning and Visual Grounding
Dave Zhenyu Chen, Qirui Wu, Matthias Nießner, Angel X. Chang
ECCV (32)2
2021 Plan2Scene: Converting Floorplans to 3D Scenes
abstract
We address the task of converting a floorplan and a set of associated photos of a residence into a textured 3D mesh model, a task which we call Plan2Scene. Our system 1) lifts a floorplan image to a 3D mesh model; 2) synthesizes surface textures based on the input photos; and 3) infers textures for unobserved surfaces using a graph neural network architecture. To train and evaluate our system we create indoor surface texture datasets, and augment a dataset of floorplans and photos from prior work with rectified surface crops and additional annotations. Our approach handles the challenge of producing tileable textures for dominant surfaces such as floors, walls, and ceilings from a sparse set of unaligned photos that only partially cover the residence. Qualitative and quantitative evaluations show that our system produces realistic 3D interior models, outperforming baseline approaches on a suite of texture quality metrics and as measured by a holistic user study.
Madhawa Vidanapathirana, Qirui Wu, Yasutaka Furukawa, Angel X. Chang, Manolis Savva
CVPR2
2020 Can Deep Learning Predict Problematic Gaming?
abstract
How does one build a healthy gaming ecosystem? Recent evidence clearly demonstrates the existence of problematic gaming [1]. Predicting problematic gaming is still in its infancy. Here we focus on excessive gaming and model in-game behaviour as a means to continuously predict future play time. This can be used to help players maintain a healthy balance between the virtual and real worlds. To do this, we convert game log data into time-series and label such data with criteria of problematic gaming. Deep learning is then used to solve the resulting multi-class classification problem.
Qirui Wu, Jacques Carette
CoG1
2020 On Ensemble Learning-Based Secure Fusion Strategy for Robust Cooperative Sensing in Full-Duplex Cognitive Radio Networks
abstract
We propose an ensemble machine learning (EML) based robust cooperative spectrum sensing framework in full-duplex cognitive radio networks (FD-CRNs), which is robust with accuracy against malicious attacks and interference. FD communication improves the spectrum awareness capability of secondary users (SUs) by allowing them to sense and transmit simultaneously over the same frequency band. However, it also complicates the sensing environment by introducing self-interference and co-channel interference. Meanwhile, the presence of malicious attacks such as Primary User Emulation and Spectrum Sensing Data Falsification (SSDF) attacks also degrades the sensing performance. To alleviate the influence of interference and attacks, we design an EML framework that provides robust and accurate fusion performance. In such a context, we analyse the spectrum waste probability, collision probability and secondary throughput in both FD Listen-Before-Talk and Listen-And-Talk protocols. Simulation results show that our proposed EML framework can provide lower and more robust false-alarm probability than single-model based fusion methods with the same detection probability constraint for any size of training sets. It also outperforms the conventional majority vote based fusion strategy in terms of spectrum waste probability, collision probability and secondary throughput for any number of SSDF SUs, only at the cost of slightly higher inference time.
Yirun Zhang, Qirui Wu, Mohammad Shikh-Bahaei
IEEE Trans. Commun.2