VLDB 2026 Research / reviewers in the wild / expert
Xiaodong Mu
dblp:92/5972
· DBLP profile ↗
23ranked-venue papers
0as first author
17since 2021 · last 2025
0009-0009-9375-0727ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | IAMTrack: interframe appearance and modality tokens propagation with temporal modeling for RGBT tracking
Huiwei Shi, Xiaodong Mu, Hao He 0005, Chengliang Zhong |
Appl. Intell. | 2 |
| 2024 | A quantum-enhanced solution method for multi classification problems
Xiaodong Mu, Dao Zhao |
Neurocomputing | 2 |
| 2023 | 3D Implicit Transporter for Temporally Consistent Keypoint DiscoveryabstractKeypoint-based representation has proven advantageous in various visual and robotic tasks. However, the existing 2D and 3D methods for detecting keypoints mainly rely on geometric consistency to achieve spatial alignment, neglecting temporal consistency. To address this issue, the Transporter method was introduced for 2D data, which reconstructs the target frame from the source frame to incorporate both spatial and temporal information. However, the direct application of the Transporter to 3D point clouds is infeasible due to their structural differences from 2D images. Thus, we propose the first 3D version of the Transporter, which leverages hybrid 3D representation, cross attention, and implicit reconstruction. We apply this new learning system on 3D articulated objects and non-rigid animals (humans and rodents) and show that learned keypoints are spatio-temporally consistent. Additionally, we propose a closed-loop control strategy that utilizes the learned keypoints for 3D object manipulation and demonstrate its superior performance. Codes are available at https://github.com/zhongcl-thu/3D-Implicit-Transporter. Chengliang Zhong, Yuhang Zheng 0004, Yupeng Zheng, Hao Zhao 0002, Li Yi 0001, Xiaodong Mu, Ling Wang 0001, Pengfei Li 0007, Guyue Zhou, Chao Yang 0026, Jian Zhao 0006 |
ICCV | 6 |
| 2023 | Urban Surface Emission Longwave Radiation Estimation from High Spatial Resolution Image Using a Hybrid MethodabstractAccurate estimation of the surface emission longwave radiation (SELR) has important scientific significance for understanding its spatiotemporal dynamics and surface thermal environment. High spatial resolution thermal infrared images provide better data support for studying SELR of complex surfaces such as urban surface. This paper focus on proposing a new urban-oriented hybrid method to estimate urban surface emission longwave radiation from top-of-atmosphere thermal radiance images, by taking the GF-5/VIMI thermal image as an example, and conduct the parameter sensitive analysis of the model as well as application over Beijing city. The experimental results of the simulation dataset showed that the developed method has relatively high precision, with SELR errors of less than 12.0 W/m2under low water vapor conditions and less than 17.0 W/m2under high water vapor conditions. The application of method in GF-5 image also demonstrated the rationality and effectiveness of the method. Songyi Lin, Rongyuan Liu, Qiming Qin, Wenjie Fan 0001, Xiaodong Mu, Baozhen Wang, Yunzhu Tao |
IGARSS | 6 |
| 2023 | Hierarchical Neural Network with Serial Attention Mechanism for Review Sentiment Classifification
Xiaodong Mu |
Neural Process. Lett. | 4 |
| 2022 | Sim2Real Object-Centric Keypoint Detection and DescriptionabstractKeypoint detection and description play a central role in computer vision. Most existing methods are in the form of scene-level prediction, without returning the object classes of different keypoints. In this paper, we propose the object-centric formulation, which, beyond the conventional setting, requires further identifying which object each interest point belongs to. With such fine-grained information, our framework enables more downstream potentials, such as object-level matching and pose estimation in a clustered environment. To get around the difficulty of label collection in the real world, we develop a sim2real contrastive learning mechanism that can generalize the model trained in simulation to real-world applications. The novelties of our training method are three-fold: (i) we integrate the uncertainty into the learning framework to improve feature description of hard cases, e.g., less-textured or symmetric patches; (ii) we decouple the object descriptor into two independent branches, intra-object salience and inter-object distinctness, resulting in a better pixel-wise description; (iii) we enforce cross-view semantic consistency for enhanced robustness in representation learning. Comprehensive experiments on image matching and 6D pose estimation verify the encouraging generalization ability of our method. Particularly for 6D pose estimation, our method significantly outperforms typical unsupervised/sim2real methods, achieving a closer gap with the fully supervised counterpart. Chengliang Zhong, Chao Yang 0026, Fuchun Sun 0001, Jinshan Qi, Xiaodong Mu, Huaping Liu 0001, Wenbing Huang 0001 |
AAAI | 5 |
| 2022 | SNAKE: Shape-aware Neural 3D Keypoint FieldabstractDetecting 3D keypoints from point clouds is important for shape reconstruction, while this work investigates the dual question: can shape reconstruction benefit 3D keypoint detection? Existing methods either seek salient features according to statistics of different orders or learn to predict keypoints that are invariant to transformation. Nevertheless, the idea of incorporating shape reconstruction into 3D keypoint detection is under-explored. We argue that this is restricted by former problem formulations. To this end, a novel unsupervised paradigm named SNAKE is proposed, which is short for shape-aware neural 3D keypoint field. Similar to recent coordinate-based radiance or distance field, our network takes 3D coordinates as inputs and predicts implicit shape indicators and keypoint saliency simultaneously, thus naturally entangling 3D keypoint detection and shape reconstruction. We achieve superior performance on various public benchmarks, including standalone object datasets ModelNet40, KeypointNet, SMPL meshes and scene-level datasets 3DMatch and Redwood. Intrinsic shape awareness brings several advantages as follows. (1) SNAKE generates 3D keypoints consistent with human semantic annotation, even without such supervision. (2) SNAKE outperforms counterparts in terms of repeatability, especially when the input point clouds are down-sampled. (3) the generated keypoints allow accurate geometric registration, notably in a zero-shot setting. Codes and models are available at https://github.com/zhongcl-thu/SNAKE. Chengliang Zhong, Peixing You, Xiaoxue Chen, Hao Zhao 0002, Fuchun Sun 0001, Guyue Zhou, Xiaodong Mu, Chuang Gan 0001, Wenbing Huang 0001 |
NeurIPS | 7 |
| 2022 | Hierarchical attention and feature projection for click-through rate prediction
Chengliang Zhong, Shouxiang Fan, Xiaodong Mu, Zhen Ni |
Appl. Intell. | 4 |
| 2022 | Combining feature importance and neighbor node interactions for cold start recommendation
Chenhui Ma, Chengliang Zhong, Xiaodong Mu |
Eng. Appl. Artif. Intell. | 5 |
| 2022 | Multi-scale and multi-channel neural network for click-through rate prediction
Chenhui Ma, Chengliang Zhong, Xiaodong Mu |
Neurocomputing | 5 |
| 2022 | Enhance Tensor RPCA-Based Mahalanobis Distance Method for Hyperspectral Anomaly DetectionabstractThis letter proposes a spectral–spatial anomaly detection method based on tensor decomposition. First, tensor data are used to represent hyperspectral data to retain the original spectral and spatial information. Second, hyperspectral image (HSI) data are decomposed into low-rank and sparse tensors. The proposed method uses weighted tensor Schatten$p$-norm minimization (WTSNM) instead of rank minimization. WTSNM assigns different weights to singular values to retain the important information and filter out noise. It efficiently solves the tensor decomposition problem using Fourier transform, generalized soft thresholding, and a tensor singular value decomposition (T-SVD) method. Finally, the obtained low-rank tensor is used to estimate the background statistics, and a Mahalanobis distance-based anomaly detector is developed using the background statistics. The experimental results on three real datasets show that the proposed method outperforms several state-of-the-art algorithms. A. Ruhan, Xiaodong Mu, Jingyuan He |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | Hausdorff IoU and Context Maximum Selection NMS: Improving Object Detection in Remote Sensing Images With a Novel Metric and Postprocessing ModuleabstractThe object detectors based on deep convolution neural network have achieved significant success in the field of remote sensing images. Intersection over Union (IoU) and No-maximum suppression (NMS) are the essential components of state-of-the-art anchor-based object detectors. However, as a localization evaluation metric, IoU does not precisely match the boundary box regression, leading to inaccurate regression of the object detector. Therefore, we introduce Hausdorff distance and combine it with IoU as a new evaluation metric (HIoU). NMS is an integral part of the object detection pipeline. However, it may lose relatively small object information in the case of high overlap. Because of the denseness of objects, this defect is more prominent in remote sensing image object detection. Therefore, we consider the context information of location confidence and propose the context maximum selection NMS (Cms-NMS) algorithm. Finally, we integrate HIoU and Cms-NMS into state-of-the-art object detectors, respectively. The performance of these object detectors is improved on the benchmark datasets NWPUVHR-10 and RSOD without any additional hyperparameters. The experiments show that HIoU and Cms-NMS are compatible, and using them together can further improve the detectors’ accuracy. Lizhi Wang 0009, Xiaodong Mu, Chenhui Ma |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | MBPI: Mixed behaviors and preference interaction for session-based recommendation
Chenhui Ma, Chengliang Zhong, Xiaodong Mu, Lizhi Wang 0009 |
Appl. Intell. | 4 |
| 2021 | Improving current interest with item and review sequential patterns for sequential recommendation
Xiaodong Mu, Chenhui Ma |
Eng. Appl. Artif. Intell. | 2 |
| 2021 | Recurrent convolutional neural network for session-based recommendation
Chenhui Ma, Xiaodong Mu, Chengliang Zhong, A. Ruhan |
Neurocomputing | 3 |
| 2021 | Multilayer Feature Fusion With Weight Adjustment Based on a Convolutional Neural Network for Remote Sensing Scene ClassificationabstractRemote sensing scene classification is still a challenging task. Extracting features effectively from restricted existing labeled data is key to scene classification. Convolutional neural networks (CNNs) are an effective method of constructing discriminating feature representation. However, CNNs usually utilize the feature map from the last layer and ignore additional layers with valuable feature information. In addition, the direct integration of multiple layers brings only a small improvement due to feature redundancy and destruction. To explore the potential information from additional layers and improve the effect of feature fusion, we propose multilayer feature fusion accesses with weight adjustment based on a CNN. We construct access to deliver additional features to one layer to achieve feature fusion and set weight factors to adjust the fusion degree to reduce feature redundancy and destruction. We perform experiments on two common data sets, which indicate improved accuracies and advantages of the extraction capability of our method. Chenhui Ma, Xiaodong Mu, Renpu Lin |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2021 | Object Detection Based on Efficient Multiscale Auto-Inference in Remote Sensing ImagesabstractObject detection in remote sensing images has important applications in various aspects. Object detection algorithms with deep convolutional neural networks (DCNNs) have made remarkable progress. However, when processing objects on vastly multiple scales in high-resolution optical remote sensing images, there is a high computational cost. Therefore, to simplify neural network multiscale training and inference, an automatic multiscale inference framework is proposed to balance the speed and accuracy of object detection. We use an attention mechanism that uses a key-point network to predict regions with small objects on a coarse scale and only process regions obtained from the first stage on finer scales instead of processing an entire larger scale image. The fully convolutional neural network (CNN) that is used in training and detecting is not affected by the image input resolution. The experiments are carried out using the NWPUVHR-10 data set, and the experimental results show that these methods can improve the training efficiency and detection accuracy in remote sensing images. Shaojing Zhang, Xiaodong Mu, Guangjie Kou, Jingyu Zhao 0007 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2019 | SAR Target Image Classification Based on Transfer Learning and Model CompressionabstractWhen convolutional neural networks (CNNs) are applied to the synthetic aperture radar (SAR) image classification, they are prone to overfitting due to scarce SAR image data, and CNNs require a large amount of storage and long computing time, so it is difficult to deploy them on resource constrained devices. This letter proposes a simple and feasible approach that can effectively solve these problems. First, the convolutional layers of the pretrained model on the ImageNet data set are transferred, and a new convolutional layer and global pooling layer are added afterward. Then, fine-tuning is performed on the new network from the SAR image data set. Finally, a filterbased pruning method is used on the convolutional layers to obtain a compact network. Compared with the all-convolutional network (A-ConvNets) which is the state-of-the-art method on the moving and stationary target acquisition and recognition data set, our method achieves about 3.6× speedup during forward propagation and 3.7× compression of the parameters, with only a 1.42% decrease in the accuracy. Chengliang Zhong, Xiaodong Mu, Xiangchen He |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2018 | Classification for SAR Scene Matching Areas Based on Convolutional Neural NetworksabstractThe selection of scene matching areas is a difficult problem in the field of matching guidance. Compared with the traditional methods of matching feature extraction and pattern classification, this letter applies convolutional neural networks (CNN) to the extraction of synthetic aperture radar (SAR) scene matching regions for the first time. First of all, we match the SAR images of the same land taken by satellites from different angles and in different phases, and then automatically label the matching suitability of the images as the output of the network according to the matching results. Next, the digital elevation model data reflecting the elevation information and the SAR image grayscale information are fused as the input to the network. Finally, CNN is used to automatically extract the matching features and classify the suitability of the SAR images. The proposed method avoids the steps of extracting features manually and improves the classification performance of SAR scene matching area. Compared with the support vector machine method, the classification accuracy increases from 86.1% to 93.3%. Chengliang Zhong, Xiaodong Mu, Xiangchen He, Bichao Zhan |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2017 | Adapting Remote Sensing to New Domain With ELM Parameter TransferabstractIt is time consuming to annotate unlabeled remote sensing images. One strategy is taking the labeled remote sensing images from another domain as training samples, and the target remote sensing labels are predicted by supervised classification. However, this may lead to negative transfer due to the distribution difference between the two domains. To address this issue, we propose a novel domain adaptation method through transferring the parameters of extreme learning machine (ELM). The core of this method is learning a transformation to map the target ELM parameters to the source, making the classifier parameters of the target domain maximally aligned with the source. Our method has several advantages which was previously unavailable within a single method: multiclass adaptation through parameter transferring, learning the final classifier and transformation simultaneously, and avoiding negative transfer. We perform experiments on three data sets that indicate improved accuracy and computational advantages compared to baseline approaches. Suhui Xu, Xiaodong Mu, Dong Chai |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2008 | A matrix negative selection algorithm for anomaly detectionabstractThis paper presents a matrix negative selection algorithm for anomaly detection. The proposed algorithm is a twofold improvement over conventional negative selection algorithms. In matrix representation, characteristics of the self set are emerged by multiple vectors to distinctly express the boundary of self and non-self. On the other hand, based on the matrix matching coefficient, separate match rules for generating detectors and monitoring anomaly are designed to avoid the sharp distinction caused by threshold. Results have demonstrated that the matrix negative selection algorithm is effective and reliable for anomaly detection and suitable for small sample problems of complex systems. Zhaoxiang Yi, Xiaodong Mu |
IEEE Congress on Evolutionary Computation | 2 |
| 2008 | The Research on Image Classification of Remote Sensing Based on an Improved Neural NetworkabstractWith higher spatial resolution, the image classification of remote sensing is always a hot research field. Besides spectral information, texture information from remote sensing image of higher spatial resolution has become an important data source to improve the classification accuracy. The image classification approach adopts an improved neural network, which contains two steps connected by the refusal principle. Two steps of input neurons are spectral information, using 3times3 window size, and texture information from gray co-occurrence matrix, which is selected by the genetic algorithms. The final result which is to overlay of above results get higher accuracy that the traditional method that ANN combine simply all of information from different source as input neurons. Mu Bai, Huiping Liu, Xiaoluo Zhou, Xiaodong Mu |
IGARSS (2) | 5 |
| 2008 | Monitoring Urban Expansion in Beijing, China by Multi-Temporal TM and SPOT ImagesabstractWe select Beijing City as the study area, focusing on the urban expansion in the past decade based on remotely sensed data (TM and SPOT of year 1997, 2001, 2004 and 2007) and GIS technology. Land use change is detected by means of post classification comparison. We obtain the distribution map of urban land during three periods for ten years. Then, spatial model of urban expansion is analyzed by GIS technology. Markov chains is used to gain the percentage for each type of land use convert to urban land; and urban expansion in given size of window is discussed. The main results are: 1) The work flow of multi-temporal land use change detection proved to be efficient. The outcomes were maps of land use pattern, urban ratio and expansion intensity maps; 2) From central city to suburb, the land use variation takes on strong spatio-temporal change. Urban expansion cores and traffic are two urban expansion spatial influence factors. Huiping Liu, Qingzu Luan, Mu Bai, Xiaodong Mu |
IGARSS (4) | 5 |