Yuanrong Xu

dblp:161/8357 · DBLP profile ↗
← Back
28ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0002-9457-7956ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 5 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AIFIND: Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection
abstract
As forgery types continue to emerge consistently, Incremental Face Forgery Detection (IFFD) has become a crucial paradigm. However, existing methods typically rely on data replay or coarse binary supervision, which fails to explicitly constrain the feature space, leading to severe feature drift and catastrophic forgetting. To address this, we propose AIFIND, Artifact-Aware Interpreting Fine-Grained Alignment for Incremental Face Forgery Detection, which leverages semantic anchors to stabilize incremental learning. We design the Artifact-Driven Semantic Prior Generator to instantiate invariant semantic anchors, establishing a fixed coordinate system from low-level artifact cues. These anchors are injected into the image encoder via Artifact-Probe Attention, which explicitly constrains volatile visual features to align with stable semantic anchors. Adaptive Decision Harmonizer harmonizes the classifiers by preserving angular relationships of semantic anchors, maintaining geometric consistency across tasks. Extensive experiments on multiple incremental protocols validate the superiority of AIFIND.
Hao Wang 0035, Beichen Zhang 0006, Yanpei Gong, Shaoyi Fang, Zhaobo Qi, Yuanrong Xu, Xinyan Liu 0008, Weigang Zhang
ICMR6
2026 Multimodal-guided mixture-of-experts bias removal strategy for natural language video localization
Xiaowen Ruan, Zhaobo Qi, Ruisi Chen, Yuanrong Xu, Beichen Zhang 0006, Weigang Zhang
Multim. Syst.4
2025 Procedure Knowledge Decoupled Distillation Strategy for Procedure Planning in Instructional Videos
abstract
Procedure planning in instructional videos, producing a structured and plannable action sequence facilitating the transition from the start to the goal states, has achieved significant progress. The dominant single-branch non-autoregressive planning paradigm guides action sequence generation through action labels, overlooking the limitation of the absence of intermediate visual information. Hence, we introduce the procedure knowledge decoupled distillation strategy to address the above issue. This innovative strategy deliberately lets the teacher model see the real visual information among the start and goal states to enhance its action semantic understanding and relationship modeling ability, producing the potential probability distribution containing the real action class and other action classes that may occur. Accordingly, we introduce a decoupled intermediate information knowledge distillation loss, which comprises single action knowledge distillation and sequence distribution knowledge distillation for the student model. The former improves the student model's precise inference ability for individual actions by transferring knowledge of a single action target category using binary classification loss. Conversely, the latter uses MSE loss to constrain the student model to learn the action sequence probability distribution from the teacher model, thereby enhancing the student model's global planning capability. Extensive experiments on three datasets demonstrate that our strategy can improve the performance of multiple weakly supervised models, achieving promising procedure knowledge modeling ability and plug-and-play flexibility.
Xiaotian Pan, Zhaobo Qi, Yuanrong Xu, Weigang Zhang
AAAI4
2025 Multi-View Learning with Context-Guided Receptance for Image Denoising
abstract
Image denoising is essential in low-level vision applications such as photography and automated driving. Existing methods struggle with distinguishing complex noise patterns in real-world scenes and consume significant computational resources due to reliance on Transformer-based models. In this work, the Context-guided Receptance Weighted Key-Value (CRWKV) model is proposed, combining enhanced multi-view feature integration with efficient sequence modeling. The Context-guided Token Shift (CTS) mechanism is introduced to effectively capture local spatial dependencies and enhance the model's ability to model real-world noise distributions. Also, the Frequency Mix (FMix) module extracting frequency-domain features is designed to isolate noise in high-frequency spectra, and is integrated with spatial representations through a multi-view learning process. To improve computational efficiency, the Bidirectional WKV (BiWKV) mechanism is adopted, enabling full pixel-sequence interaction with linear complexity while overcoming the causal selection constraints. The model is validated on multiple real-world image denoising datasets, outperforming the state-of-the-art methods quantitatively and reducing inference time up to 40%. Qualitative results further demonstrate the ability of our model to restore fine details in various scenes. The code is publicly available at https://github.com/Seeker98/CRWKV.
Binghong Chen, Tingting Chai, Yuanrong Xu, Guanglu Zhou
IJCAI4
2025 IdTrPalm: Identity-Traceable Stylized Palmprint Image Generation
Longfa Liu, Lunke Fei, Shuyi Li 0003, Jian Zhu 0001, Yuanrong Xu, Shaohua Teng
PRCV (15)5
2025 Efficient U-shape invertible neural network for large-capacity image steganography
Le Zhang 0016, Yao Lu 0008, Yuanrong Xu, Guangming Lu 0002
J. Inf. Secur. Appl.4
2025 Dual-guided multi-modal bias removal strategy for temporal sentence grounding in video
Xiaowen Ruan, Zhaobo Qi, Yuanrong Xu, Weigang Zhang
Multim. Syst.3
2025 KN-VLM: KNowledge-guided Vision-and-Language Model for visual abductive reasoning
Kuo Tan, Zhaobo Qi, Jianping Zhong, Yuanrong Xu, Weigang Zhang
Multim. Syst.4
2025 VPA: Multi-Modal Virtual Point Augmentation for 3D Object Detection
abstract
Integrating LiDAR and camera data is crucial for precise 3D object detection. Existing methods resort to augmenting virtual points from 2D image space in a random manner to complete the appearance of 3D objects with sparse points. However, these augmented virtual points have unreasonable 3D positions and representations, which brings serious negative effects on accurate detection. To this end, we introduce a general 3D object detection framework called Virtual Point Augmenting (VPA) to enrich the 3D point cloud by controllably generating virtual points with accurate depth and position information as well as domain-gap-eliminated multi-modal representations from image and point cloud spaces. VPA contains two core designs, namely Hybrid Sampling Method (HSM) and Fine-Grained Cross-modal Fusion (FGCF). HSM uses the constructed seed point distribution map based on the edge score and mask score map to sample high-quality seed points, and employs a feature similarity function to sample withkneighbors’ depth to obtain more accurate depth for the seed points, thereby enhancing the quality of the virtual points’ 3D positions. FGCF fuses the multi-modal features,i.e., the semantic feature, the geometric feature from the image space, and the 3D position feature in an adaptive manner using self-attention mechanism, thereby further improving the representation of the virtual points. We apply VPA to the LiDAR-based method CenterPoint and fusion-based method Cross-modal transformer. Experimental results on the nuScenes, KITTI, and Waymo benchmarks validate the efficiency of our VPA, which achieves promising performance with 72.9% mAP and 74.8% NDS without using test-time augmentation and model ensemble techniques on the nuScenes test set. Code is available at https://github.com/jianpingZhonggit/vpa.git.
Jianping Zhong, Zhaobo Qi, Kaiwen Duan, Yuanrong Xu, Weigang Zhang, Qingming Huang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Multi-Modal 3D Object Detector with Object-Guided Fusion and Hierarchical Sample Selection
abstract
Accurately detecting objects in 3D scenes is crucial for autonomous driving. Although existing voxel-based methods have achieved remarkable progress, their performance on tail objects remains unsatisfactory. We identify two core issues contributing to this phenomenon: the detectors frequently misidentify some background elements as foreground objects, and there is a misalignment between the classification score and detection quality. To tackle these challenges, we introduce an object-level – guided multi-modal 3D object detector with an object-guided feature fusion (OFF) module and a hierarchical sample selection (HSS) strategy, named OGMMDet. Specifically, OFF introduces rich image features to enhance the representation of objects while using an object distribution heatmap to suppress the background. This approach provides geometry clues for tail objects while providing category priors to filter out the background. HSS uses a local-to-global ranking approach to calculate the relative classification loss weights of all proposals. It assigns higher weights to proposals with higher IoU when optimizing classification branches. This ensures that the model focuses its optimization on these higher-quality proposals. Consequently, there is a positive correlation between the classification score and IoU. This method alleviates the misalignment between the classification score and detection quality. Extensive experiments on the KITTI and nuScenes benchmarks demonstrate the effectiveness of our OGMMDet, which achieves 45.61% and 68.96% mean average precision (mAP) on pedestrians and cyclists on the KITTI benchmark, respectively. Code is available at https://github.com/ZhongJianPing1/ogmmdet.git .
Jianping Zhong, Zhaobo Qi, Kaiwen Duan, Yuanrong Xu, Weigang Zhang, Qingming Huang
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Hierarchical Retrieval of High-Resolution Fingerprints Based on Pore Feature
abstract
Faced with an escalating number of fingerprint images, most existing retrieval approachs suffer from a common problem: diminishing computational efficiency. This paper presents a hierarchical retrieval system tailored for high-resolution fingerprint images that utilizes abundant pore features and robust recognizability to improve retrieval performance. The framework comprises two core components. Firstly, a CNN-based feature extraction network is established, incorporating an attention mechanism to capture pore features in fingerprint images comprehensively. Subsequently, a hierarchical fingerprint retrieval approach is introduced, involving connection graph construction and a hierarchy of jump table structures for efficient retrieval of query pores. Empirical experiments conducted on high-resolution fingerprint image datasets underscore the system’s effectiveness. Compared with other advanced pore-based fingerprint retrieval methods, the proposed method exhibits a notable rise in the hit rate with reduced penetration rates, significantly reducing the retrieval time.
Yuanrong Xu, Suyu Dong, Wei Wang 0169
BIBM2
2024 Improving Sequential DeepFake Detection with Local information enhancement
Longyun Dong, Yuanrong Xu, Jianping Zhong, Zhaobo Qi, Weigang Zhang
MMAsia2
2022 Multi-modal Finger Feature Fusion Algorithms on Large-Scale Dataset
Chuhao Zhou, Yuanrong Xu, Fanglin Chen 0001, Guangming Lu 0002
PRCV (2)2
2022 Multiscale Conditional Regularization for Convolutional Neural Networks
abstract
With the increased model size of convolutional neural networks (CNNs), overfitting has become the main bottleneck to further improve the performance of networks. Currently, the weighting regularization methods have been proposed to address the overfitting problem and they perform satisfactorily. Since these regularization methods cannot be used in all the networks and they are usually not flexible enough in different phases of the training and test processes, this article proposes a multiscale conditional (MSC) regularization method. MSC divides the intermediate features into different scales and then generates new data for each scale features, respectively. In addition, the new data are generated by employing the information from two conditions: 1) each sample feature and 2) each layer pattern. Finally, a self-identity structure is proposed to supplement the features with the generated data. Therefore, MSC can adaptively and efficiently generate much finer and individualized data to make the entire regularization more flexible. Furthermore, MSC is more general and can be applied to all kinds of networks through the proposed self-identity structure. The experimental results on all the benchmark datasets showed that the proposed MSC regularization method achieves the best performances in all the networks.
Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, Yuanrong Xu, Zheng Zhang 0006, David Zhang 0001
IEEE Trans. Cybern.4
2022 High Resolution Fingerprint Retrieval Based on Pore Indexing and Graph Comparison
abstract
Fingerprint retrieval aims to identify a query fingerprint image in a large database using indexing algorithms. Because of the abundant level 3 pore features within high-resolution fingerprint images, pore-based fingerprint retrieval algorithms have been rapidly developed. These retrieval algorithms, however, suffer from severe calculation-consuming problems with the pores increasing. This paper proposes a pore-based fingerprint retrieval method for high-resolution fingerprint images. The proposed method consists of two main steps. 1) In the pore indexing step, an indexing space is constructed using the binary codes of pores in enrolled images. Then, a designed graph-based searching algorithm searches the nearest neighbors of pores from the query image to construct one-to-many correspondences. 2) In the refinement step, the one-to-many correspondences are refined by a random walker-based graph comparison algorithm to remove the false correspondences. The remained nearest neighbors are used to calculate the similarities between the query image and the enrolled images. The proposed method is evaluated on two databases, showing that our method achieves better retrieval accuracies with a higher speed than the existing pore-based retrieval algorithms.
Yuanrong Xu, Yao Lu 0008, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001
IEEE Trans. Inf. Forensics Secur.1
2021 Highly shared Convolutional Neural Networks
Yao Lu 0008, Guangming Lu 0002, Yicong Zhou, Jinxing Li 0003, Yuanrong Xu, David Zhang 0001
Expert Syst. Appl.5
2021 Fully shared convolutional neural networks
Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, Zheng Zhang 0006, Yuanrong Xu
Neural Comput. Appl.5
2021 Fast Pore Comparison for High Resolution Fingerprint Images Based on Multiple Co-Occurrence Descriptors and Local Topology Similarities
abstract
Pore-based fingerprint recognition has been researched for decades. Many algorithms have been proposed to improve the recognition accuracy of the system. However, the accuracies are always improved at the cost of speed. This article proposes a novel method to compare the pores in high-resolution fingerprint images using the popular coarse-to-fine strategy. A multiple spatial pairwise local co-occurrence descriptor is proposed to improve the calculation of the similarities between pores. It calculates multiple local co-occurrence statistics for each pore using its neighbors. The proposed method can establish correspondences between pores more accurately. The refinement of the correspondences is then achieved by using a local topology-preserving matching algorithm. The algorithm uses rotational invariant local structures and pore pair local topology similarities to calculate the cost of each correspondence. It can remove the mismatches more accurately and efficiently. The experimental results on two high-resolution fingerprint image databases show that the proposed algorithm perform well in both accuracy and speed comparing to the existing algorithms.
Yuanrong Xu, Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, David Zhang 0001
IEEE Trans. Syst. Man Cybern. Syst.1
2020 High-parameter-efficiency convolutional neural networks
Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, Yuanrong Xu, David Zhang 0001
Neural Comput. Appl.4
2019 Super Sparse Convolutional Neural Networks
abstract
To construct small mobile networks without performance loss and address the over-fitting issues caused by the less abundant training datasets, this paper proposes a novel super sparse convolutional (SSC) kernel, and its corresponding network is called SSC-Net. In a SSC kernel, every spatial kernel has only one non-zero parameter and these non-zero spatial positions are all different. The SSC kernel can effectively select the pixels from the feature maps according to its non-zero positions and perform on them. Therefore, SSC can preserve the general characteristics of the geometric and the channels’ differences, resulting in preserving the quality of the retrieved features and meeting the general accuracy requirements. Furthermore, SSC can be entirely implemented by the “shift” and “group point-wise” convolutional operations without any spatial kernels (e.g., “3×3”). Therefore, SSC is the first method to remove the parameters’ redundancy from the both spatial extent and the channel extent, leading to largely decreasing the parameters and Flops as well as further reducing the img2col and col2img operations implemented by the low leveled libraries. Meanwhile, SSC-Net can improve the sparsity and overcome the over-fitting more effectively than the other mobile networks. Comparative experiments were performed on the less abundant CIFAR and low resolution ImageNet datasets. The results showed that the SSC-Nets can significantly decrease the parameters and the computational Flops without any performance losses. Additionally, it can also improve the ability of addressing the over-fitting problem on the more challenging less abundant datasets.
Yao Lu 0008, Guangming Lu 0002, Bob Zhang 0001, Yuanrong Xu, Jinxing Li 0003
AAAI4
2019 Stable Pore Detection for High-Resolution Fingerprint based on a CNN Detector
abstract
High-resolution fingerprint images contain three levels of features. Pores, as one of the level 3 features, have wide attention due to its significant contribution to the recognition accuracy. An accurate and stable pore detection algorithm plays a key role on the pore-based fingerprint recognition system. This paper proposes a pore detection method for high-resolution fingerprint images. The method uses fully convolutional network combined with the focal loss and shortcut structure to detect pores. The proposed algorithm is tested on the high-resolution fingerprint database. Experimental results show that our method outperforms the existing algorithms in accuracy, stability and matching performance.
Zuolin Shen, Yuanrong Xu, Jinxing Li 0003, Guangming Lu 0002
ICIP2
2019 Weighted Channel-Wise Decomposed Convolutional Neural Networks
Yao Lu 0008, Guangming Lu 0002, Yuanrong Xu
Neural Process. Lett.3
2019 High resolution fingerprint recognition using pore and edge descriptors
Yuanrong Xu, Guangming Lu 0002, Yao Lu 0008, David Zhang 0001
Pattern Recognit. Lett.1
2019 Fingerprint Pore Comparison Using Local Features and Spatial Relations
abstract
High-resolution fingerprint recognition has been a hot topic for many years. Compared with a traditional fingerprint image, a high-resolution fingerprint image can provide more features, such as pores and ridge contours. Introducing these features into fingerprint comparison and recognition can improve the recognition accuracy and reduce the risk of identification errors. This paper proposes a novel method for comparing pores on high-resolution fingerprint images. The method can be divided into two steps. In the first step, fingerprints are aligned using the pixel-category-distance-based data-driven descending algorithm. Traditionally, fingerprints are aligned based on feature points, such as minutiae and singular points. Such alignment methods are not suitable when dealing with partial fingerprints because small overlapping areas often do not contain enough features to guarantee a correct alignment. In this research, the ridges and valleys on fingerprints are used in combination with the orientation field for alignment. The proposed algorithm performs well when aligning both partial and full fingerprints. The common areas between the two images can be estimated based on the alignment result. In the second step, pores lying in the common areas are selected for comparison. To improve the comparison accuracy, pores are compared using local features and spatial relations. A graph comparison algorithm is designed in this step. The experimental results show that the proposed method is more accurate than other state-of-the-art pore comparison algorithms.
Yuanrong Xu, Guangming Lu 0002, Yao Lu 0008, Feng Liu 0013, David Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2018 AAR-CNNs: Auto Adaptive Regularized Convolutional Neural Networks
abstract
In order to address the overfitting problem caused by the small or simple training datasets and the large model’s size in Convolutional Neural Networks (CNNs), a novel Auto Adaptive Regularization (AAR) method is proposed in this paper. The relevant networks can be called AAR-CNNs. AAR is the first method using the “abstraction extent” (predicted by AE net) and a tiny learnable module (SE net) to auto adaptively predict more accurate and individualized regularization information. The AAR module can be directly inserted into every stage of any popular networks and trained end to end to improve the networks’ flexibility. This method can not only regularize the network at both the forward and the backward processes in the training phase, but also regularize the network on a more refined level (channel or pixel level) depending on the abstraction extent’s form. Comparative experiments are performed on low resolution ImageNet, CIFAR and SVHN datasets. Experimental results show that the AAR-CNNs can achieve state-of-the-art performances on these datasets.
Yao Lu 0008, Guangming Lu 0002, Yuanrong Xu, Bob Zhang 0001
IJCAI3
2018 Inter-class sparsity based discriminative least square regression
Jie Wen 0001, Yong Xu 0001, Zhongli Ma, Yuanrong Xu
Neural Networks5
2017 Fast pore matching method based on deterministic annealing algorithm
abstract
High‐resolution fingerprint identification system (HRFIS) has become a hot topic in the field of academic research. Compared to traditional automatic fingerprint identification system, HRFIS reduces the risk of being faked by using level 3 features, such as pores, which cannot be detected in lower resolution images. However, there is a serious problem in HRFIS: there are hundreds of sweat pores in one fingerprint image, which will spend a considerable amount of time for direct fingerprint matching. The authors propose a method to match pores in two fingerprint images based on deterministic annealing algorithm. In this method, fingerprints are aligned using singular points. Then minutiae are matched based on the alignment result. To reduce the impact of deformation, they build a convex hull for each of these fingerprints. Pores in these convex hulls are used for matching. In the experiments, their method is compared with random sample consensus method, minutia and ICP‐based method, and direct pore matching method. The results show that the proposed method is more efficient.
Guangming Lu 0002, Yuanrong Xu
IET Image Process.2
2016 Fingerprint Pores Matching based on Improved Deterministic Annealing Algorithm
Guangming Lu 0002, Yuanrong Xu, Liying Ye
ICPRAM2