Wanying Xu

dblp:90/9454 · DBLP profile ↗
← Back
18ranked-venue papers
4as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Image recognition and object detection · 63% Efficient and distributed learning · 16% Vision and language · 16%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Emerging computing paradigms · 82% Cloud and datacenter computing · 18%
Network and information security
1 paper
Cryptographic primitives and cryptanalysis · 67% Blockchain and cryptocurrency security · 33%
Computer graphics and multimedia
1 paper
Image and video processing · 100%

Topics — the 18 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object detection
2.022026
Completing Missing Entities: Exploring Consistency Reasoning for Remote Sensing Object Detection · IEEE Trans. Image Process. 2026
VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection · AAAI 2026
Computer vision › Image recognition and object detection › object detection › robust object detection
occluded object detection
1.012026
Completing Missing Entities: Exploring Consistency Reasoning for Remote Sensing Object Detection · IEEE Trans. Image Process. 2026
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection
1.012026
VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection · AAAI 2026
Computer vision › Image recognition and object detection › object detection
remote sensing object detection
1.012026
Completing Missing Entities: Exploring Consistency Reasoning for Remote Sensing Object Detection · IEEE Trans. Image Process. 2026
Machine learning › Efficient and distributed learning › model compression
sparse neural network
1.012026
S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing · AAAI 2026
Computer vision › Vision and language
vision-language model
1.012026
VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection · AAAI 2026
Emerging computing paradigms › neuromorphic computing › neuromorphic vision
event-based vision
1.012026
S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing · AAAI 2026
Emerging computing paradigms
neuromorphic computing
1.012026
S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing · AAAI 2026
Image and video processing › image warping
image rectification
0.912025
Lifting the Structural Morphing for Wide-Angle Images Rectification: Unified Content and Boundary Modeling · ICCV 2025
Image and video processing › image warping › image rectification
wide-angle image rectification
0.912025
Lifting the Structural Morphing for Wide-Angle Images Rectification: Unified Content and Boundary Modeling · ICCV 2025
Cryptographic primitives and cryptanalysis › searchable encryption › searchable symmetric encryption
dynamic searchable symmetric encryption
0.812024
Themis: Robust and Light-Client Dynamic Searchable Symmetric Encryption · IEEE Trans. Inf. Forensics Secur. 2024
Blockchain and cryptocurrency security › blockchain network
light client
0.812024
Themis: Robust and Light-Client Dynamic Searchable Symmetric Encryption · IEEE Trans. Inf. Forensics Secur. 2024
Cryptographic primitives and cryptanalysis
searchable encryption
0.812024
Themis: Robust and Light-Client Dynamic Searchable Symmetric Encryption · IEEE Trans. Inf. Forensics Secur. 2024
Machine learning › Efficient and distributed learning
model compression
0.312026
S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing · AAAI 2026
Computer vision › Vision and language
multimodal reasoning
0.312026
Completing Missing Entities: Exploring Consistency Reasoning for Remote Sensing Object Detection · IEEE Trans. Image Process. 2026
Natural language and speech › Question answering and dialogue systems
reasoning consistency
0.312026
Completing Missing Entities: Exploring Consistency Reasoning for Remote Sensing Object Detection · IEEE Trans. Image Process. 2026
Cloud and datacenter computing › cloud storage
encrypted data search
0.212024
Themis: Robust and Light-Client Dynamic Searchable Symmetric Encryption · IEEE Trans. Inf. Forensics Secur. 2024
Cloud and datacenter computing › data outsourcing
outsourced storage
0.212024
Themis: Robust and Light-Client Dynamic Searchable Symmetric Encryption · IEEE Trans. Inf. Forensics Secur. 2024

Methods — techniques the papers use, named apart from their topics

sparse encoding · 2.0dual-branch architecture · 2.0asynchronous processing · 2.0secret sharing · 1.5oblivious map · 1.5relation proposals · 1.0pseudo-labeling · 1.0prototype learning · 1.0mask-reconstruction relation learning · 1.0knowledge distillation · 1.0consistency-reasoning transformer · 1.0structural morphing · 0.9
YearPublicationVenuePosition
2026 S3Net: Spatiotemporally Separated Sparse Network for Neuromorphic Vision Processing
abstract
Dynamic Vision Sensor (DVS) asynchronously records sparse events triggered by changes in pixel intensity, offering high temporal resolution and low latency. Existing frame-based methods process event data densely, violating its inherent sparsity and introducing computational redundancy. While asynchronous models preserve the event stream's native format, they often neglect spatial information, compromising their adaptability and efficiency. To address these limitations, we propose a Spatiotemporally Separated Sparse Network (S3Net) for efficient event stream encoding and learning. Specifically, we employ a learnable sparse encoding scheme to construct a voxel-structured representation that effectively extracts spatiotemporal relationships among event data. After that, we propose a dual-branch architecture to capture localized spatial dependencies and dynamic temporal patterns of event data. By explicitly decoupling spatial and temporal modeling, S3Net enables end-to-end asynchronous processing of variable-length event sequences, achieving both strong representational capacity and high computational efficiency. Experimental results on six event-based datasets demonstrate that S3Net achieves state-of-the-art performance. Compared to frame-based methods, it significantly reduces computational overhead and model complexity, while also outperforming existing asynchronous approaches in inference speed without compromising accuracy. Extensive experiments across six event-based datasets show that S3Net establishes new state-of-the-art performance. Our method reduces computational costs by 35% and model parameters by 27% compared to frame-based approaches, while delivering 1.58× faster inference than existing point-based methods at comparable accuracy levels.
Rong Xiao 0001, Wanying Xu, Chenwei Tang, Shudong Huang, Huajin Tang
AAAI3
2026 VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection
abstract
To identify objects beyond predefined categories, open-vocabulary aerial object detection (OVAD) leverages the zero-shot capabilities of visual-language models (VLMs) to generalize from base to novel categories. Existing approaches typically utilize self-learning mechanisms with weak text supervision to generate region-level pseudo-labels to align detectors with VLMs semantic spaces. However, text dependence induces semantic bias, restricting open-vocabulary expansion to text-specified concepts. We propose VK-Det, a visual knowledge-guided open-vocabulary object detection framework without extra supervision. First, we discover and leverage vision encoder's inherent informative region perception to attain fine-grained localization and adaptive distillation. Second, we introduce a novel prototype-aware pseudo-labeling strategy. It models inter-class decision boundaries through feature clustering and maps detection regions to latent categories via prototype matching. This enhances attention to novel objects while compensating for missing supervision. Extensive experiments show state-of-the-art performance, achieving 30.1 mAPᴺ on DIOR and 23.3 mAPᴺ on DOTA, outperforming even extra supervised methods.
Jianhang Yao, Yongbin Zheng, Siqi Lu, Wanying Xu
AAAI4
2026 Lane detection with vanishing box based dynamic anchor generation mechanism
Zhixiong Nan, Wanying Xu, Tao Xiang 0001
Neurocomputing2
2026 MSER: Multi-scale event representation model for enhanced spatio-temporal feature extraction
Wanying Xu, Rong Xiao 0001, Chenwei Tang, Jiancheng Lv 0001, Huajin Tang
Neurocomputing1
2026 MACTrack: Spatiotemporal context propagation with motion compensation for anti-UAV tracking
Wanying Xu, Haijiang Sun
Neural Networks1
2026 A lane detection model with knowledge guided anchor feature enhancement mechanism
Zhixiong Nan, Wanying Xu, Fulin Luo, Tao Xiang 0001
Pattern Recognit.2
2026 Completing Missing Entities: Exploring Consistency Reasoning for Remote Sensing Object Detection
abstract
Recent studies in remote sensing object detection have made excellent progress and shown promising performance. However, most current detectors only explore rotation-invariant feature extraction but disregard the valuable spatial and semantic prior knowledge in remote sensing images (RSIs), which limits the detection performance when encountering blurred or heavy occluded objects. To address this issue, we propose a mask-reconstruction relation learning (MRRL) framework to learn such prior knowledge among objects and a consistency-reasoning transformer over relation proposals (CTRP) to recognize objects with limited visual features via consistency reasoning. Specifically, MRRL framework applies random mask to some objects in the training dataset and performs masked objects reconstruction to guide the network to learn the distribution consistency of objects. CTRP is the core component of the MRRL framework, which models the interaction between spatial and semantic priors, and uses easy detected objects to reason hard detected objects. The trained CTRP can be integrated into the existing detector to improve the ability of object detection with limited visual features in RSIs. Extensive experiments on widely-used datasets for two distinct tasks, namely remote sensing object detection task and occluded object detection task, demonstrate the effectiveness of the proposed method. Source code is available at https://github.com/sunpeng96/CTRP_mmrotate.
Yongbin Zheng, Wanying Xu, Jian Li 0003, Jiansong Yang
IEEE Trans. Image Process.3
2025 Lifting the Structural Morphing for Wide-Angle Images Rectification: Unified Content and Boundary Modeling
Wenting Luan, Siqi Lu, Yongbin Zheng, Wanying Xu, Lang Nie, Zongtan Zhou, Kang Liao
ICCV4
2024 A novel and efficient model pruning method for deep convolutional neural networks by evaluating the direct and indirect effects of filters
Yongbin Zheng, Wanying Xu
Neurocomputing4
2024 Themis: Robust and Light-Client Dynamic Searchable Symmetric Encryption
abstract
Dynamic searchable symmetric encryption (DSSE), as one of the promising cryptographic tools in cloud-based services, faces two crying needs at the age of multi-device. One is a lightweight client, and the other is robustness. A lightweight client facilitates seamless synchronization among multiple devices allowing users to feel as if they are operating on a single device, even on resource-constrained devices. Robustness ensures a reliable system that can tolerate misoperations. DSSE requires both of them to achieve a leap in practicability. However, to our best knowledge, lightweight client and robustness have not been effectively combined thus far. Most existing DSSE schemes maintain a substantial amount of state information on the client for sub-linear search efficiency, but they fail to guarantee security even correctness, after executing the client’s misoperations (e.g., duplicate addition or deletion operation and deleting non-existent targets). The seminal work on robustness, ROSE (TIFS’22), leverages a heavy primitive to preserve security and correctness during post-processing and requires a heavy client storage burden. To guarantee robustness and constant client storage simultaneously, we devise a novel method to preserve robustness timely in the process of misoperations. Specifically, we introduce an alarm mechanism to promptly eliminate the effects of misoperations. Based on the misoperation alarm mechanism and thevORAM+HIRBoblivious map (S&P’16), we propose a new DSSE schemeThemis. In addition to satisfying robustness and constant client storage, it has competitive search and update performance compared to prior representative DSSE schemes. Moreover, it is superior to existing robust schemes in search.
Yubo Zheng, Peng Xu 0003, Wanying Xu, Wei Wang 0088, Hai Jin 0001
IEEE Trans. Inf. Forensics Secur.4
2023 How does business-IT alignment influence supply chain resilience?
Shaobo Wei, Wanying Xu, Xiayu Chen
Inf. Manag.2
2022 A Robust and Accurate End-to-End Template Matching Method Based on the Siamese Network
abstract
Template matching is an important and challenging task in remote sensing and computer vision. Existing template matching methods often fail in the presence of complex nonrigid deformation, occlusion, and background clutter. In this letter, inspired by Siamese trackers, we propose an end-to-end template matching method that is based on the Siamese network. Different from the traditional template matching methods, our method treats the template matching task as a classification-regression task. It is more robust to background clutter, occlusion, and nonrigid deformation. Moreover, we introduce a channel-attention mechanism in the cross correlation operation and replace the commonly used intersection-over-union (IoU) with distance-IoU (DIoU) to build a new regression loss, which further improves the performance of our method. Extensive experiments on the commonly used public benchmark demonstrate that our method achieves the state-of-the-art performance.
Yongbin Zheng, Wanying Xu
IEEE Geosci. Remote. Sens. Lett.4
2021 Unsupervised Entity Resolution Method Based on Random Forest
Wanying Xu, Chenchen Sun, Zhijiang Hou
WISA1
2020 R4 Det: Refined single-stage detector with feature recursion and refinement for rotating object detection in aerial images
Yongbin Zheng, Zongtan Zhou, Wanying Xu
Image Vis. Comput.4
2017 Visual Tracking via Probabilistic Hypergraph Ranking
abstract
Online object tracking is a challenging issue because the appearance of an object tends to change due to intrinsic or extrinsic factors. In this paper, we propose a tracking algorithm based on probabilistic hypergraph ranking. First, three types of hypergraphs are constructed to encode local affinity information. Then, a probabilistic hypergraph is built by combining three distinct hypergraphs linearly. Second, an adaptive template constraint is proposed to effectively use the discriminative information of different templates. Third, object tracking is formulated as a transductive learning issue, and the optimal target location is determined by maximum a posteriori estimation on the ranking scores. Finally, a dynamic updating scheme of positive and negative template sets provides the proposed tracker with robustness against appearance variations. A series of experiments and evaluations on various challenging image sequences is performed, and the results show that the proposed algorithm performs favorably against other state-of-the-art methods.
Ruitao Lu, Wanying Xu, Yongbin Zheng, Xinsheng Huang
IEEE Trans. Circuits Syst. Video Technol.2
2011 Robust Visual Tracking by Integrating Lucas-Kanade into Mean-Shift
abstract
The mean-shift algorithm has achieved considerable success in object tracking due to its simplicity and robustness. However, the lack of template update often leads to out of adaptation to affine transformation of the object. The Lucas-Kanade algorithm has some advantages in obtaining the affine parameters. In this paper, we introduce the inverse compositional algorithm, which is equivalent to but more efficient than Lucas-Kanade algorithm, to complement the traditional mean-shift algorithm. In this method, the average of squared error (ASE) between the initial template and the object image which is warped through the obtained affine parameters is computed to decide whether to update the current template. Experimental results show that the mean-shift tracking with Lucas-Kanade algorithm (MSLK) has high tracking accuracy and good robustness to the change of appearance of the object.
Lurong Shen, Xinsheng Huang, Wanying Xu, Yongbin Zheng
ICIG3
2011 Local Dominant Orientation Based Mutual Information for Multisensor Template Matching
abstract
Mutual information (MI) has been very successful in multisensor or multimodal image matching. However, it may lead to mismatching due to lack of spacial information. In this paper, based on a local dominant orientation (LDO), which is a stable nature among images of different sensors and is widely used in the relative rotation estimation, an improved MI for multisensor images matching is proposed. Firstly, the frequently used intensity images are converted to a LDO represented form, where the LDO for each pixel is calculated by cumulating the surrounding gradient vectors within a disk like region. Next, we introduce a simple clustering to cluster each transformed image, thus the joint histogram of MI in the matching stage can be reduced significantly, and hence the computations, memory consumption. Our approach is evaluated by 10 groups of multisensor images, and the results have demonstrated its outstanding performances.
Yuzhuang Yan, Yongbin Zheng, Wanying Xu, Xinsheng Huang
ICIG3
2010 An affine invariant interest point and region detector based on Gabor filters
abstract
This paper presents a novel approach for interest point and region detection which is invariant to affine transformations. Such transformations introduce significant changes in the point location as well as in the scale and the shape of the neighborhood of an interest point. Our approach allows to solve for these problems simultaneously. The approach is based on three key ideas: 1) Interest points can be extracted based on local maxima of the normalized local energy maps. 2) Local extrema over scale of the normalized energy function indicate the presence of characteristic local structures. 3) The maximum response along all the orientations indicates the principle orientation of the local structure. We first extract interest points at multi-scales from the local energy map constructed by Gabor filter responses, and then select points at which a local measure is maximal over scales. This allows a selection of distinctive points for which the characteristic scale is known. We then estimate the principle orientation through the orientational responses of Gabor filters and extend the detector to affine invariance by estimating the affine shape of a point neighborhood. The characteristic scale and the affine shape of neighborhood determine an affine invariant region for each point. Experimental results with synthetic images and natural images show the affine invariance performance of our approach. Comparative evaluation using the repeatability criteria demonstrates the comparable performance in the presence of large viewpoint changes.
Wanying Xu, Xinsheng Huang, Xingwei Li, Ying Zhang 0032, Wei Zhang 0027
ICARCV1