EDBT 2026 Demo / reviewers in the wild / expert
Xi Chen 0078
dblp:16/3283-78
· DBLP profile ↗
18ranked-venue papers
0as first author
16since 2021 · last 2026
0000-0002-0735-9919ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 9 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | URF-Loc: Uncertainty-Aware Frequency-Spatial Network for GNSS-Denied UAV Geo-Localization
Xi Chen 0078, Chu He |
ICIC (18) | 3 |
| 2025 | Wavelet-based Feature Representation Framework for Event Stream RecognitionabstractEvent streams generated by event cameras exhibit low data redundancy and preserve precise temporal information through Address Event Representation (AER), which differs significantly from the outputs of traditional frame-based cameras. However, traditional Artificial Neural Networks (ANNs) and Spiking Neural Networks (SNNs) often convert these streams into static frames, which can lead to some loss of temporal features. In this work, we propose a novel wavelet-based framework, Tempo-Spatial Wavelet Transform (TSWT) framework, designed to optimize both ANNs and SNNs. The Temporal Wavelet Transform (TWT) Module within this framework produces a wavelet-based event representation that captures more precise temporal correlations from the event streams. Additionally, the Spatial Wavelet Transform (SWT) Module provides a wavelet-based method for selecting frequency features. By integrating these two modules with a tailored loss function, TSWT enhances the network’s ability to learn tempo-spatial information effectively. Experimental results demonstrate that both ANNs and SNNs equipped with the proposed TSWT framework achieve superior performance across various datasets. Xingyu Pan, Xi Chen 0078, Shenghua Fan |
ICME | 3 |
| 2024 | Textual Tokens Classification for Multi-Modal Alignment in Vision-Language TrackingabstractMost vision-language (VL) trackers rely on coarse-grained information from sentences to achieve multi-modal alignment. However, this information is insufficient for accurately describing the target in each frame due to the inherent ambiguity, summarization, and invariance of sentences, thereby making multi-modal alignment challenging. This paper introduces TTCTrack, a novel VL tracker that employs textual token classification to address this challenge. Specifically, we exploit the multi-modal cross-relations to classify textual tokens into various types and employ diverse operations on them, enabling multi-modal alignment in the tracking process. This enables the textual tokens to accurately describe the target and dynamically reflect the scene changes. Moreover, we introduce a dual-encoder structure that effectively handles the multi-modal input and fusion. Extensive experiments on four datasets demonstrate the effectiveness of our proposed tracking method. Zhongjie Mao, Xi Chen 0078, Jia Yan 0006 |
ICASSP | 3 |
| 2023 | Point-to-RBox Network for Oriented Object Detection via Single Point Supervision
Chu He, Xi Chen 0078 |
BMVC | 3 |
| 2023 | Wavelet-Based Frequency-Dividing Interactive CNN for Image ClassificationabstractThe vanilla tensor in convolutional neural networks (CNNs) can be seen as a mixture of feature information at different frequencies, which is currently only used as a carrier of information. However, few people notice that vanilla tensor is spatial information redundant and information interaction of different frequency bands is beneficial for CNNs. In this paper, we design a novel Wavelet-based frequency-dividing interactive block (WFDI) to factorize a vanilla tensor into a pair of tensors with complementary information to reduce redundancy. Based on this, we embed it into the CNN (WFDI-CNN) for image classification. Specifically, the WFDI-CNN factorizes the vanilla tensor into a low-frequency tensor with lower spatial resolution and a high-frequency tensor with complementary information. Then, the information interaction and forward propagation between the high-frequency and low-frequency tensors not only save computational resources but also improve the network performance. Experimental results on CIFAR10 and CIFAR100 datasets all demonstrate the effectiveness of the proposed WFDI block. Jidong Cao, Chu He, Xi Chen 0078 |
ICIP | 5 |
| 2023 | Weakly Semi-Supervised Oriented Object Detection with PointsabstractOriented object detection based on deep learning has received extensive attention while marking the Oriented Bounding Box (OBB) is time-consuming and laborious. For the oriented object detection task, this paper proposes a point-based weakly semi-supervised training strategy, which only requires the training set with an extremely small number (10%) of fully labeled images with OBB and the other with points. Specifically, following the common self-training pipeline, we propose Point to OBB Network (P2ONet) as the teacher model to generate the high-quality pseudo OBB for each point-annotated object. Inspired by channel attention, we introduce Group Attention to P2ONet to better tackle the opposite assignment for similar proposals in the training task leading by the point label assignment strategy. Furthermore, by exploring the constraint in normal self-training pipeline, we propose Confidence-Aware Loss to alleviate the impact of inaccurate pseudo-boxes. Experiments on the DOTA dataset show the close performance between our method and normal oriented object detection training methods with remarkably lower labeling costs. Chu He, Xi Chen 0078 |
ICIP | 5 |
| 2023 | Deep Homography Estimation With Feature Correlation TransformerabstractHomography estimation is an important image alignment method that has been widely used in computer vision applications. Traditional methods heavily rely on the distribution of features and usually fail in low-texture and large-baseline scenes. Most learning-based methods use convolutional neural networks(CNNs) to extract features. However, the dense features extracted in this way have a limited receptive field, leading to poor accuracy of results. In this paper, we propose a novel method for homography estimation. We first estimate the projective transformation between the reference image and the target image at a coarse level and then refine the estimated homography at the fine level. Unlike approaches that use simple CNNs or global correlations to search correspondences, we add self- and cross-attention layers in the transformer to enhance the feature correlations. The experiments show that our method significantly outperforms the existing solutions in challenging large-baseline scenes. Haoyu Zhou, Chu He, Xi Chen 0078 |
ICME | 5 |
| 2023 | Cross-scale content-based full Transformer network with Bayesian inference for object tracking
Shenghua Fan, Xi Chen 0078, Chu He |
Multim. Tools Appl. | 2 |
| 2023 | Superpixel-based foreground-preserving image stitching
Xinpeng Miao, Tao Qu, Xi Chen 0078, Chu He |
Mach. Vis. Appl. | 3 |
| 2023 | Multiple frequency-spatial network for RGBT tracking in the presence of motion blur
Shenghua Fan, Xi Chen 0078, Chu He, Lei Yu 0006, Zhongjie Mao, Yujin Zheng |
Neural Comput. Appl. | 2 |
| 2023 | Bayesian Dumbbell Diffusion Model for RGBT Object Tracking With Enriched PriorsabstractRGBT tracking can be accomplished by constructing Bayesian estimators that incorporate fusion prior distributions for the visible (RGB) and thermal (T) modalities. Such estimators enable the computation of a posterior distribution for the variables of interest to locate the target. Incorporating rich prior information can improve the performance of predictors. However, current RGBT trackers face limited fusion prior data. To mitigate this issue, we propose a novel tracker, BD$^{2}$Track, which employs a diffusion model. Firstly, this paper introduces a dumbbell diffusion model, and employ convolution networks and the dumbbell model to derive the fusion feature prior information from various index frames in the same tracking video sequence. Secondly, we propose a plug-and-play channel augmented joint learning strategy to derive the images prior distribution. This strategy not only homogeneously generates modality-relevant prior information but also increases the distance between positive and negative samples within the modality, while reducing the distance between modalities during fusion. Results demonstrate promising performance in the GTOT, RGBT234, LasHeR, and VTUAV-ST datasets, surpassing other state-of-the-art trackers. Shenghua Fan, Chu He, Chenxia Wei, Yujin Zheng, Xi Chen 0078 |
IEEE Signal Process. Lett. | 5 |
| 2023 | Tiny Object Tracking With Proposal Position EnhancementabstractTiny object tracking is challenging due to the target's weak appearance and features. The current state-of-theart approach for this task uses probabilistic regression based on discriminative correlation filters (DCF) to predict regression scores. However, the probabilistic regression equally relies on the position and size of the proposal box. In this letter, we point out that the position is of greater importance relative to the size when regressing the tiny object. To this end, we introduce the receptive field distance to define the quality of the proposal box, which places more emphasis on the position. By incorporating quality as an additional important factor for Monte Carlo sampling, the number of high quality proposals can be effectively increased, leading to regression optimization. Moreover, we propose a network to generate quality scores for proposals. The combination of probability and quality scores serves as a selection criterion for the optimal proposal, and can boost the tiny object tracking performance. Extensive experiments demonstrate the effectiveness of the proposed method. The code and model will be available athttps://github.com/jankin987/track-mcsq Zhongjie Mao, Chenxia Wei, Xi Chen 0078, Jia Yan 0006 |
IEEE Signal Process. Lett. | 4 |
| 2022 | Fuzzy Information Measures Feature Selection Using Descriptive Statistics Data
Omar A. M. Salem, Yi-Ping Phoebe Chen, Xi Chen 0078 |
KSEM (3) | 5 |
| 2022 | Fuzzy joint mutual information feature selection based on ideal vector
Omar A. M. Salem, Yi-Ping Phoebe Chen, Ahmed Hamed Attia, Xi Chen 0078 |
Expert Syst. Appl. | 5 |
| 2022 | Effective fuzzy joint mutual information feature selection based on uncertainty region for classification problem
Omar A. M. Salem, Yi-Ping Phoebe Chen, Ahmed Hamed Attia, Xi Chen 0078 |
Knowl. Based Syst. | 5 |
| 2021 | Feature selection and threshold method based on fuzzy joint mutual information
Omar A. M. Salem, Yi-Ping Phoebe Chen, Xi Chen 0078 |
Int. J. Approx. Reason. | 4 |
| 2015 | Visual object tracking via online sparse instance learning
Jia Yan 0006, Xi Chen 0078, Dexiang Deng, Qiuping Zhu |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Structured partial least squares based appearance model for visual tracking
Jia Yan 0006, Xi Chen 0078, Dexiang Deng, Qiuping Zhu |
Neurocomputing | 2 |