VLDB 2026 Research / reviewers in the wild / expert
Xiaolong Sun
dblp:169/3312
· DBLP profile ↗
15ranked-venue papers
4as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DFA-FedDG:A federated domain generalization network anomaly detection system based on dynamic feature alignment
Hao Zhang 0078, Xiaoping Wen, Junwei Ye, Xiaolong Sun, Wei Huang 0037 |
Comput. Networks | 4 |
| 2026 | Select and assign: anchor-guided proposals for temporal sentence grounding
Ziyue Zou, Xiaolong Sun |
Multim. Syst. | 5 |
| 2025 | Diversifying Query: Region-Guided Transformer for Temporal Sentence GroundingabstractTemporal sentence grounding is a challenging task that aims to localize the moment spans relevant to a language description. Although recent DETR-based models have achieved notable progress by leveraging multiple learnable moment queries, they suffer from overlapped and redundant proposals, leading to inaccurate predictions. We attribute this limitation to the lack of task-related guidance for the learnable queries to serve a specific mode. Furthermore, the complex solution space generated by variable and open-vocabulary language descriptions complicates optimization, making it harder for learnable queries to adaptively distinguish each other, leading to more severe overlapped proposals. To address this limitation, we present the Region-Guided TRansformer (RGTR) for temporal sentence grounding, which introduces regional guidance to increase query diversity and eliminate overlapped proposals. Instead of using learnable queries, RGTR adopts a set of anchor pairs as moment queries to introduce explicit regional guidance. Each moment query takes charge of moment prediction for a specific temporal region, which reduces the optimization difficulty and ensures the diversity of the proposals. In addition, we design an IoU-aware scoring head to improve proposal quality. Extensive experiments demonstrate the effectiveness of RGTR, outperforming state-of-the-art methods on three public benchmarks and exhibiting good generalization and robustness on out-of-distribution splits. Xiaolong Sun, Liushuai Shi, Le Wang 0003, Sanping Zhou, Gang Hua 0001 |
AAAI | 1 |
| 2025 | Boosting Point-Supervised Temporal Action Localization through Integrating Query Reformation and Optimal TransportabstractPoint-supervised Temporal Action Localization poses significant challenges due to the difficulty of identifying complete actions with a single-point annotation per action. Existing methods typically employ Multiple Instance Learning, which struggles to capture global temporal context and requires heuristic post-processing. In research on fully-supervised tasks, DETR-based structures have effectively addressed these limitations. However, it is nontrivial to merely adapt DETR to this task, encountering two major bottlenecks. (1) How to integrate point label information into the model and (2) How to select optimal decoder proposals for training in the absence of complete action segment annotations. To address this issue, we introduce an end-to-end framework by integrating Query Reformation and Optimal Transport (QROT). Specifically, we encode point labels through a set of semantic consensus queries, enabling effective focus on action-relevant snippets. Furthermore, we integrate an optimal transport mechanism to generate high-quality pseudo labels. These pseudo-labels facilitate precise proposals selection based on the Hungarian algorithm, significantly enhancing localization accuracy in point-supervised settings. Extensive experiments on the THUMOS14 and ActivityNet-v1.3 datasets demonstrate that our method outperforms existing MIL-based approaches, offering more stable and accurate temporal action localization in point-level supervision. Mengnan Liu 0001, Le Wang 0003, Sanping Zhou, Xiaolong Sun, Gang Hua 0001 |
CVPR | 5 |
| 2025 | Moment Quantization for Video Temporal GroundingabstractVideo temporal grounding is a critical video understanding task, which aims to localize moments relevant to a language description. The challenge of this task lies in distinguishing relevant and irrelevant moments. Previous methods focused on learning continuous features exhibit weak differentiation between foreground and background features. In this paper, we propose a novel Moment-Quantization based Video Temporal Grounding method (MQVTG), which quantizes the input video into various discrete vectors to enhance the discrimination between relevant and irrelevant moments. Specifically, MQVTG maintains a learnable moment codebook, where each video moment matches a codeword. Considering the visual diversity, i.e., various visual expressions for the same moment, MQVTG treats moment-codeword matching as a clustering process without using discrete vectors, avoiding the loss of useful information from direct hard quantization. Additionally, we employ effective prior-initialization and joint-projection strategies to enhance the maintained moment codebook. With its simple implementation, the proposed method can be integrated into existing temporal grounding models as a plug-and-play component. Extensive experiments on six popular benchmarks demonstrate the effectiveness and generalizability of MQVTG, significantly outperforming state-of-the-art methods. Further qualitative analysis shows that our method effectively groups relevant features and separates irrelevant ones, aligning with our goal of enhancing discrimination. Xiaolong Sun, Le Wang 0003, Sanping Zhou, Liushuai Shi, Mengnan Liu 0001, Gang Hua 0001 |
ICCV | 1 |
| 2025 | An Indoor Moving Target Detection Method Based on Doppler Chirp Rate Profile and Range Gating FilterabstractThe detection of indoor moving targets using autonomous-aerial-vehicle (AAV)-mounted through-the-wall radar (TWR) has been widely applied in both military and civilian fields. However, the imaging results of moving targets often suffer from defocusing and are prone to be submerged in strong stationary clutter, which reduces the detection rate of moving targets. To address this issue, this article proposes a detection method for indoor moving targets based on Doppler chirp rate profile and range gating filter using AAV-mounted TWR. In the proposed method, the characteristics of Doppler chirp rate for both clutter and moving targets are analyzed. The range migration of the target echo is corrected in the range-Doppler (RD) domain through sinc interpolation. Then, the fractional Fourier transform (FrFT) is applied to estimate the chirp rate at each range bin. Subsequently, the estimated chirp rate values are compared with the theoretical chirp rate values of stationary objects. Based on the differences of Doppler chirp rates, a gating filter in fast time domain is designed to suppress clutter originated from walls and stationary objects. Finally, an azimuth compression filter is constructed to achieve focused imaging of moving targets. Both simulation and field experiments demonstrate the feasibility and effectiveness of the proposed method. The proposed method shows strengths over other methods in terms of improvement factor, image entropy, and peak-sidelobe ratio. Hao Zhang 0138, Xiaolong Sun, Xiaopeng Yang 0002 |
IEEE Internet Things J. | 4 |
| 2025 | Target Tracking Method Based on Scale-Adaptive Rotation Kernelized Correlation Filter for Through-the-Wall RadarabstractThrough-the-wall radar (TWR) can track moving human targets in obscured spaces, providing real-time information to operators. Changes in the position and angle of targets during movement will lead to diversity in the scale and orientation of the radar images, causing difficulty in target tracking. In order to solve this problem, this letter proposes a TWR target tracking method based on scale-adaptive rotation kernelized correlation filter (SA-RKCF). In the proposed method, the size and angle of the target region are both estimated, which are used for training samples generation and correlation filter updating. The candidate target region of current frame is delineated based on the geometric information of the previous frame, and then the correlation filter is used to localize the target position. The effectiveness of the proposed method is verified using both numerical simulation and experiment. Hao Zhang 0138, Xiaolong Sun, Xiaopeng Yang 0002 |
IEEE Signal Process. Lett. | 4 |
| 2025 | Moving Target Coherent Integration Method Based on TRCM-KT for UAV-Mounted Through-the-Wall RadarabstractMoving target detection in through-the-wall scenario can be achieved by utilizing unmanned aerial vehicle (UAV) radar. However, considering indoor human target detection, the distance between the radar and the moving target is relatively short, and the ratio of target's velocity to the UAV's velocity varies greatly, which makes the range migration and Doppler phase much more complex and poses huge challenges for coherent integration. To address this issue, this letter proposes a moving target coherent integration method based on time reverse conjugate multiply and keystone transform (TRCM-KT) for UAV-mounted through-the-wall radar. In the proposed method, the reference signal is constructed by azimuthal time reverse and conjugation. Then, by multiplying the echo signal and the reference signal, the second-order range migration and doppler phase terms are eliminated. Next, keystone transform is employed to correct the remained first-order range migration. Finally, the fully coherent integration results can be obtained by performing Fourier transform along the azimuth direction. The effectiveness of the proposed method is verified by both simulation and experiment results. Xiaolong Sun, Hao Zhang 0138, Xiaopeng Yang 0002 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Network intrusion detection based on feature fusion of attack dimension
Xiaolong Sun, Zhengyao Gu, Hao Zhang 0078, Jason Gu, Chen Dong 0002, Junwei Ye |
J. Supercomput. | 1 |
| 2024 | A Novel DDoS Detection Model for SDN Using Single-Class Cluster Oversampling and Weighted Ensemble MethodabstractThe centralized control plane characteristic of Software Defined Networking (SDN) makes it a prime target for Distributed Denial of Service (DDoS) attacks. A significant issue in network traffic data is the severe imbalance between background traffic and various types of DDoS attack traffic. To address this challenge, we propose a novel DDoS detection model based on single-class cluster oversampling and weighted ensemble method. Initially, we design a multi-strategy feature selection using variance inflation factor and recursive feature elimination. Next, we construct a single-class cluster model based on the spatial distribution of samples, and create minority class samples independently within each cluster. Finally, detection results are determined using the weight vector of each base classifier in ensemble learning. The proposed DDoS detection model's performance is validated through simulations of various DDoS attack scenarios in an SDN environment. Experimental results indicate that the proposed strategy exhibits superior performance, showing improvements in precision, recall, and F1 score compared to state-of-the-art techniques. Hao Zhang 0078, Shuqi Wu, Xiaolong Sun |
ICNP | 5 |
| 2024 | Adaptive Micro-Doppler Corner Feature Extraction Method Based on Difference of Gaussian Filter and Deformable ConvolutionabstractThrough-the-wall radar (TWR) utilizes range and Doppler information to achieve indoor human activity recognition. However, traditional recognition methods are developed based on range-time maps (RTM) and Doppler-time maps (DTM), resulting in low accuracy and poor robustness. In order to solve these problems, this letter proposes to use micro-Doppler corner feature to achieve activity recognition and gives an adaptive corner feature extraction method based on difference of Gaussian (DoG) filter and deformable convolution. Micro-Doppler corner feature is defined as the points on the radar squared-range and squared-Doppler images where the gray scale changes sharply in different directions, reflecting the inflection, stationing, intersection, and boundaries of the motion trajectory curves of the human limb nodes. The proposed corner feature extraction method utilizes the DoG filter to extract the micro-Doppler corner supervisory labels on simulated data. The labels are then used to train the$\boldsymbol{\mu }$D-CornerDet, which is constructed based on deformable convolution network (DCN), task-adaptive deformable convolution network (TDCN), feature pyramid network (FPN) and learnable regression global attention module (LRGA). For predictions, only$\mu$D-CornerDet is used on measured data to obatin the corner feature maps. Both numerical simulations and experiments are conducted to verify the effectiveness and robustness of the proposed method. Weicheng Gao, Haoyu Meng, Xiaolong Sun, Xiaopeng Yang 0002 |
IEEE Signal Process. Lett. | 4 |
| 2022 | Ship Segmentation via Encoder-Decoder Network With Global Attention in High-Resolution SAR ImagesabstractShip detection in the synthetic aperture radar (SAR) image is of great significance in the fields of military and coastal defense. Most ship detection methods are designed based on the object detection framework, which can only provide the vertices’ coordinates of the bounding box covering the ship targets but cannot provide more detailed contour information. Target segmentation can further explore the shape and edge information of the objects, which can be used as a blazing novel means for automatic object detection. In this letter, a 3-D atrous encoder–decoder neural network with global attention modules (GAM-EDNet) is proposed to achieve ship segmentation in SAR images. The encoder–decoder structure with atrous convolution is developed as the network body to fully exploit the structural information of the ship targets with various sizes. To increase the structural information of the single-polarization SAR images, a 3-D image cube is designed as the input of the GAM-EDNet. A global attention module is proposed to further improve the segmentation performance by integrating the high-level semantic features with the low-level location features. Besides, an SAR ship segmentation dataset (SAR-HR4) is built to evaluate the segmentation performance, and the experimental results show that the proposed GAM-EDNet achieves better performance than other state-of-the-art methods. Jichao Li 0003, Shuiping Gou, Jiawei Chen 0001, Xiaolong Sun |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2018 | Local tangent space alignment via nuclear norm regularization for incomplete data
Jing Wang 0049, Xiaolong Sun, Jixiang Du |
Neurocomputing | 2 |
| 2017 | Novel hybrid CNN-SVM model for recognition of functional magnetic resonance imagesabstractThis paper proposes a novel hybrid model that integrates the synergy of two superior classifiers for functional magnetic resonance imaging (fMRI) recognition, namely, convolutional neural networks (CNNs) and support vector machines (SVMs), both of which have proven results in the field of image recognition. In the proposed model, the CNN functions as a trainable feature extractor and the SVM functions as a recognizer. This hybrid model extracts features from raw images and generates predictions for fMRI recognition. We conducted experiments on Haxby's 2001 fMRI dataset. Comparisons with Haxby's study using the same database indicated that the proposed fusion achieved superior recognition accuracy of 99.5% compared to the Haxby's approach. Further, when the CNN was used as a feature extractor, the SVM classifier was demonstrated to be the best combining counterpart, providing the best synergy effect in terms of accuracy. This is compared with other classifiers based on learning algorithms such as decision tree, neural network, K-nearest neighbor, random forest, and AdaBoost. Xiaolong Sun, Juyoung Park, Kyungtae Kang, Junbeom Hur |
SMC | 1 |
| 2015 | Multi-scale saliency detection via inter-regional shortest colour pathabstractSaliency detection has attracted considerable attention, and numerous approaches aimed at locating meaningful regions in images have been presented. Nevertheless, accurate saliency detection algorithms remain in urgent demand. Many algorithms work well when dealing with simple images, but work poorly with complex images that contain small‐scale and high‐contrast structures. Moreover, most existing local and global regional saliency detection methods measure image saliency through region contrast. Such measurement is achieved by directly computing the difference between non‐adjacent regions. In this study, the authors introduce a new perspective for evaluating region contrast. We propose a novel multi‐scale saliency region detection method by optimising the shortest path of two non‐adjacent regions in the colour space and by measuring the region contrast from different scales. The final saliency maps indicate that the proposed method can work well with images containing small patches, but with high contrast. The proposed approach can also make the foreground significantly more uniform. Experimental results on three public benchmark datasets show that the proposed method achieves better precision–recall curve than some state‐of‐the‐art methods. Wenzhong Guo, Xiaolong Sun, Yuzhen Niu |
IET Comput. Vis. | 2 |