Xu Si

dblp:179/3204 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Source-free domain adaptation for unsupervised radar-based human activity recognition
Xu Si
Pattern Recognit.1
2025 Target recognition via discriminant information and geometrical structure co-learning using radar sensor network
Xu Si, Peikun Zhu, Jing Liang 0002
Pattern Recognit.2
2024 TACO: Benchmarking Generalizable Bimanual Tool-ACtion-Object Understanding
abstract
Humans commonly work with multiple objects in daily life and can intuitively transfer manipulation skills to novel objects by understanding object functional regularities. However, existing technical approaches for analyzing and synthesizing hand-object manipulation are mostly limited to handling a single hand and object due to the lack of data support. To address this, we construct TACO, an extensive bimanual hand-object-interaction dataset spanning a large variety of tool-action-object compositions for daily human activities. TACO contains 2.5K motion sequences paired with third-person and egocentric views, precise hand-object 3D meshes, and action labels. To rapidly expand the data scale, we present a fully automatic data acquisition pipeline combining multi-view sensing with an optical motion capture system. With the vast research fields provided by TACO, we benchmark three generalizable hand-object-interaction tasks: compositional action recognition, generalizable hand-object motion forecasting, and cooperative grasp synthesis. Extensive experiments re-veal new insights, challenges, and opportunities for advancing the studies of generalizable hand-object motion anal-ysis and synthesis. Our data and code are available at https://taco2024.github.io.
Yun Liu 0018, Xu Si, Yuxiang Zhang 0006, Yebin Liu, Li Yi 0001
CVPR3
2024 Flight Trajectory Change Prediction Via Patched Spatial-Temporal Transformer
abstract
The change prediction of flight traces is a long-range multi-variant forecasting task aiming at detecting abnormal behavior in the airway. Directly implementing two-dimensional attention in Transformer-based models faces unaffordable computation costs when they handle long-range trajectories. This study proposes a patched spatial-temporal attention framework based on the Transformer architecture to reduce memory consumption. The propensity for error propagation in auto-regressive models limits the accuracy of remote predictions. To address this problem, a single-step decoder is designed to mitigate the accumulation of errors. Experiments prove that this approach enhances the precision of predictions compared to an AR model while ensuring a modest uptick in memory usage across two distinct routes from an ADS-B trace dataset.
Haiyang Hao, Xu Si
IGARSS2
2024 Learning from Noisy Label for HRRP Signal Recognition
abstract
Supervised machine learning technology has greatly improved the accuracy of radar target recognition based on HRRP signals. It relies on complete dataset labels, but the dataset is prone to noise labels due to the instability of data collection and the abstract nature of the HRRP signal itself. This problem will affect the model training robustness and testing accuracy. In this paper, we propose a noisy label learning method, "Contrastive Learning with or without Freeze, CLwF" to solve this problem. CLwF proposes a self-supervised learning algorithm to train models for efficient representation of HRRP signals. We also propose a clean rate estimation module to select the correct fine-tuning strategy for the model training with noisy labels. The experiment verified that CLwF can achieve excellent results under different noise rates.
Xu Si, Peikun Zhu, Jing Liang 0002
IGARSS1
2024 Contextual Distillation Model for Diversified Recommendation
abstract
The diversity of recommendation is equally crucial as accuracy in improving user experience. Existing studies, e.g., Determinantal Point Process (DPP) and Maximal Marginal Relevance (MMR), employ a greedy paradigm to iteratively select items that optimize both accuracy and diversity. However, prior methods typically exhibit quadratic complexity, limiting their applications to the re-ranking stage and are not applicable to other recommendation stages with a larger pool of candidate items, such as the pre-ranking and ranking stages. In this paper, we propose Contextual Distillation Model (CDM), an efficient recommendation model that addresses diversification, suitable for the deployment in all stages of industrial recommendation pipelines. Specifically, CDM utilizes the candidate items in the same user request as context to enhance the diversification of the results. We propose a contrastive context encoder that employs attention mechanisms to model both positive and negative contexts. For the training of CDM, we compare each target item with its context embedding and utilize the knowledge distillation framework to learn the win probability of each target item under the MMR algorithm, where the teacher is derived from MMR outputs. During inference, ranking is performed through a linear combination of the recommendation and student model scores, ensuring both diversity and efficiency. We perform offline evaluations on two industrial datasets and conduct online A/B test of CDM on the short-video platform KuaiShou. The considerable enhancements observed in both recommendation quality and diversity, as shown by metrics, provide strong superiority for the effectiveness of CDM.
Fan Li 0017, Xu Si, Shisong Tang, Dingmin Wang, Kunyan Han, Guorui Zhou, Yang Song 0008, Hechang Chen
KDD2
2024 A micro-Doppler spectrogram denoising algorithm for radar human activity recognition
Xu Si, Peikun Zhu, Jing Liang 0002
Signal Process.1
2024 SeisCLIP: A Seismology Foundation Model Pre-Trained by Multimodal Data for Multipurpose Seismic Feature Extraction
abstract
In seismology, while training a specific deep learning model for each task is common, it often faces challenges such as the scarcity of labeled data and limited regional generalization. Addressing these issues, we introduce SeisCLIP: a foundation model for seismology, leveraging contrastive learning during pre-training on multi-modal data of seismic waveform spectra and the corresponding local and global event information. SeisCLIP consists of a transformer-based spectrum encoder and an MLP-based information encoder that are jointly pre-trained on massive data. During pre-training, contrastive learning aims to enhance representations by training two encoders to bring corresponding waveform spectra and event information closer in the feature space, while distancing uncorrelated pairs. Remarkably, the pre-trained spectrum encoder offers versatile features, enabling its application across diverse tasks and regions. Thus, it requires only modest datasets for fine-tuning to specific downstream tasks. Our evaluations demonstrate SeisCLIP’s superior performance over baseline methods in tasks like event classification, localization, and focal mechanism analysis, even when using distinct datasets from various regions. In essence, SeisCLIP emerges as a promising foundational model for seismology, potentially revolutionizing foundation-model-based research in the domain.
Xu Si, Xinming Wu, Hanlin Sheng, Zefeng Li
IEEE Trans. Geosci. Remote. Sens.1
2024 Fast Global Self-Attention for Seismic Image Fault Identification
abstract
Fault identification is one of the challenging techniques for reservoir characterization in seismic exploration and development. Because of the widespread implementation of deep learning, automatic fault identification has developed rapidly. Recently, researchers have begun exploring the application of Transformer-based neural networks from language tasks to image recognition tasks, which provide promising results among various research fields. However, large 3-D seismic data applications bring some challenges to conventional self-attention, such as large memory and expensive computational cost for pixel-level dense identification tasks. We present a new scheme called fast global self-attention (FGSA), whose computation is achieved by cyclic multiplication. The cyclic multiplication brings greater efficiency and less memory through the fast Fourier transform (FFT). Instead of generating large attention matrices, FFT-based computations can directly weigh the features with attention. Compared with windowed attention (e.g., Swin Transformer), this FGSA architecture offers flexible global attention, saves nearly 50% of the memory, and has less computational complexity for the fault detection tasks in this article. These advantages of FGSA allow it to provide more distinct and interpretable fault identification results than conventional fault identification neural networks. Several synthetic and field seismic data examples show that the neural network based on our FGSA architecture has better applicability than some baseline methods in fault identification tasks.
Shenghou Wang, Xu Si, Zhongxian Cai, Leiming Sun, Zirun Jiang
IEEE Trans. Geosci. Remote. Sens.2
2024 Completing Any Borehole Images
abstract
Borehole images contain the physical information and chemical properties of geological formations, which are crucial for high-resolution interpretation of subsurface stratigraphic and structural features and geological modeling of the subsurface. However, due to the special design of borehole tools and variations in borehole diameter, all kinds of borehole images (FMI, Earth-imager, OMRI, and OBMI) obtained from scanning the borehole walls exhibit varying degrees of data missing, with OBMI data missing up to 70%. We propose a deep-learning approach with a hybrid CNN and Transformer architecture to fill in the gaps in borehole images, addressing the challenges of missing training labels and filling large-scale gaps. To solve the challenge of missing labels of complete borehole images, our deep-learning model is pretrained on a vast collection of complete natural and seismic images and then fine-tuned with a partial loss function on incomplete borehole images. A multistage completion strategy is further introduced into the inference stage to enhance the continuity and textural features of the completed areas. In addition, by incorporating the circular consistency constraint between the left and right sides of the borehole image, our method can reasonably complete the gaps with highly consistent features on both sides of the image. During the tests on borehole images from multiple wells in different work areas with various geological features, our model is capable of completing any type of borehole image with masks of any size, ultimately yielding complete images free of any artifacts, while also possessing richer and more reasonable textures and semantic information. We have open-sourced the code and the fine-tuned models, which are available athttps://github.com/zgyustc/LogMAT/tree/master.
Xinming Wu, Xu Pang, Hanlin Sheng, Xu Si
IEEE Trans. Geosci. Remote. Sens.5
2023 Exemplar-free Incremental Learning For Micro-Doppler Signature Classification
abstract
The utilization of machine learning techniques has greatly improved the accuracy of micro-Doppler(m-D) signatures-based radar signal recognition. However, the "catastrophic forgetting" problem commonly exists in data-driven algorithms severely limits the adaptability of recognition algorithms in real-world applications, as models cannot incrementally train and learn new categories. In this paper, we propose an incremental learning method, "Boundary Transfer and Uncertainty Augmentation (BTUA)" for continuous learning of m-D signatures. BTUA utilizes boundary transfer(BT) to generate pseudo-decision boundaries for old categories and avoid the forgetting problem. It also employs an uncertainty augmentation(UA) algorithm to enhance the model’s generalization and improve the correctness of feature extraction. Finally, the validation demonstrates the advantages of our algorithm in terms of both accuracy and practicality.
Xu Si, Peikun Zhu, Jing Liang 0002
IGARSS1
2023 A Nonlinear Waveform Selection Method for Cognitive Radar Target Tracking Based on Reinforcement Learning
abstract
Cognitive radar automatically adjusts its waveform via ceaseless interaction with the environment and learning from the experience. The waveform development of cognitive radar has been attracting much attention in improving tracking performance. In this paper, we propose an intelligent radar target tracking strategy based on variable nonlinear frequency modulated waveforms (NLFM). The strategy considers the combination of constant velocity (CV), constant acceleration (CA), and constant turning (CT) motion for high maneuvering targets. A library of NLFM is constructed and the entropy reward Q-Learning (ERQL) method is designed to perform joint waveform parameters selection. It merges the radar and target into a closed loop to provide the optimum target tracking performance, updating the waveform in real-time as the target state changes. Numerical results show that the tracking performance of our proposed method is much better than that of the linear frequency modulated waveform (LFM) pure parameter selection method.
Peikun Zhu, Xu Si, Jing Liang 0002
IGARSS2
2017 Exploiting Content Delivery Networks for covert channel communications
Yongzhi Wang 0001, Yulong Shen 0001, Xiaopeng Jiao, Tao Zhang 0029, Xu Si, Ahmed Salem 0003, Jia Liu 0009
Comput. Commun.5