Xutao Li 0003

dblp:64/1774-3 · DBLP profile ↗
← Back
116ranked-venue papers
12as first author
73since 2021 · last 2026
0000-0003-1894-984XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 65 · 5 first-author · 38 since 2021Applied, interdisciplinary, general and emerging computing · 30 · 2 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 2 first-author · 18 since 2021Databases, data management, data science and information retrieval · 15 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2026 LMcast: A pretrained language model guided long-term memory transformer for precipitation nowcasting
Feifan Gao, Chuyao Luo, Guangbo Deng, Xutao Li 0003, Baoquan Zhang, Demin Yu, Yunming Ye
Neural Networks4
2026 Advection-diffusion spatiotemporal recurrent network for regional wind speed prediction
Shidong Chen, Baoquan Zhang, Xutao Li 0003, Yunming Ye, Kenghong Lin, Rui Ye 0002
Pattern Recognit.4
2026 Influence Strength Estimation in Hyperbolic Space for Social Influence Maximization
Hongliang Qiao, Shanshan Feng 0001, Min Zhou 0006, Xutao Li 0003, Yunming Ye, Fan Li 0015, Shuo Shang, Yew-Soon Ong
IEEE Trans. Knowl. Data Eng.4
2025 Integrating Multi-Source Data for Long Sequence Precipitation Forecasting
abstract
Long-sequence precipitation forecasting is critical for both meteorological science and smart city applications. The primary objective of this task is to predict future radar echo sequences, which provide high resolution and timely references for atmospheric precipitation distribution based on current observations. However, the chaotic nature of precipitation systems poses significant challenges in extending reliable forecast horizons. Most existing methods struggle with accuracy and clarity when extended to long-sequence predictions, such as three-hour forecasts. This is primarily due to the insufficiency of spatio-temporal information within a single modality over time. In this paper, we propose a cascading forecasting framework that adaptively extracts and integrates multimodal spatio-temporal information to support accurate and realistic long-sequence radar forecasting. Our framework includes a temporal adaptive predictor and a flow-based precipitation distribution adaptor. The predictor utilizes a multi-branch encoder-decoder architecture. This design allows it to extract meteorological sequences from multiple sources at varying scales, resulting in an initial global precipitation estimate. The core component is a carefully designed cross-attention module with a temporal adaptive layer to enhance multi-modality alignment. The initial estimate is then refined by the flow-based adaptor, which adjusts the prediction to match the target precipitation distribution, enhancing local details and correcting extreme precipitation patterns. We validated our method using real multi-source dataset for long-sequence forecasting, and the experimental results demonstrate that our approach outperforms existing state-of-the-art methods.
Demin Yu, Wenzhi Feng, Kenghong Lin, Xutao Li 0003, Yunming Ye, Chuyao Luo, Wenchuan Du
AAAI4
2025 Sensitivity-Aware Efficient Fine-Tuning via Compact Dynamic-Rank Adaptation
abstract
Parameter-Efficient Fine-Tuning (PEFT) is a fundamental research problem in computer vision, which aims to tune a few parameters for efficient storage and adaptation of pre-trained vision models. Recently, sensitivity-aware parameter efficient fine-tuning method (SPT) addresses this problem by identifying sensitive parameters and then leveraging its sparse characteristic to combine unstructured and structured tuning for PEFT. However, existing methods only focus on the sparse characteristic of sensitive parameters but overlook its distribution characteristic, which results in additional storage burden and limited performance improvement. In this paper, we find that the distribution of sensitive parameters is not chaotic, but concentrates on a small number of rows or columns in each parameter matrix. Inspired by this fact, we propose a Compact Dynamic-Rank Adaptation-based tuning method for Sensitivity-aware Parameter efficient fine-Tuning, called CDRA-SPT. Specifically, we first identify the sensitive parameters that require tuning for each down-stream task. Then, we reorganize the sensitive parameters by following its row and column into a compact sub-parameter matrix. Finally, a dynamic-rank adaptation is designed and applied at sub-parameter matrix level for PEFT. Its advantage is that the dynamic-rank characteristic of sub-parameter matrix can be fully exploited for PEFT. Extensive experiments show that our method achieves superior performance over previous state-of-the-art methods.
Tianran Chen, Jiarui Chen, Baoquan Zhang, Zhehao Yu, Shidong Chen, Rui Ye 0002, Xutao Li 0003, Yunming Ye
CVPR7
2025 AlphaPre: Amplitude-Phase Disentanglement Model for Precipitation Nowcasting
abstract
Precipitation nowcasting involves using current radar observation sequences to predict future radar sequences and determine future precipitation distribution, which is crucial for disaster warning, traffic planning, and agricultural production. Despite numerous advancements, challenges persist in accurately predicting both the location and intensity of precipitation, as these factors are often interdependent, with complex atmospheric dynamics and moisture distribution causing position and intensity changes to be intricately coupled. Inspired by the fact that in the frequency domain, phase variations are shown to correspond to changes in the position of precipitation, while amplitude variations are linked to intensity changes, we propose an amplitude-phase disentanglement model called AlphaPre, which separately learn the position and intensity changes of precipitation. AlphaPre comprises three key components: a phase network, an amplitude network, and an AlphaMixer. The phase network captures positional changes by learning phase variations, and the amplitude network models intensity changes by alternating between the frequency and spatial domains. The AlphaMixer then integrates these components to produce a refined precipitation forecast. Extensive experiments on four datasets demonstrate the effectiveness and superiority of our method over state-of-the-art approaches. Our code is publicly available at https://github.com/linkenghong/AlphaPre.
Kenghong Lin, Baoquan Zhang, Demin Yu, Wenzhi Feng, Shidong Chen, Feifan Gao, Xutao Li 0003, Yunming Ye
CVPR7
2025 Perceptually Constrained Precipitation Nowcasting Model
abstract
Most current precipitation nowcasting methods aim to capture the underlying spatiotemporal dynamics of precipitation systems by minimizing the mean square error (MSE). However, these methods often neglect effective constraints on the data distribution, leading to unsatisfactory prediction accuracy and image quality, especially for long forecast sequences. To address this limitation, we propose a precipitation nowcasting model incorporating perceptual constraints. This model reformulates precipitation nowcasting as a posterior MSE problem under such constraints. Specifically, we first obtain the posteriori mean sequences of precipitation forecasts using a precipitation estimator. Subsequently, we construct the transmission between distributions using rectified flow. To enhance the focus on distant frames, we design a frame sampling strategy that gradually increases the corresponding weights. We theoretically demonstrate the reliability of our solution, and experimental results on two publicly available radar datasets demonstrate that our model is effective and outperforms current state-of-the-art models.
Wenzhi Feng, Xutao Li 0003, Zhe Wu 0006, Kenghong Lin, Demin Yu, Yunming Ye, Yaowei Wang 0001
ICML2
2025 Extreme Weather Nowcasting With Second-Order State Spaces
abstract
Nowcasting weather extremes poses significant challenges due to the complex evolution of their dynamical systems. State space models (SSMs) excel in sequence modeling, offering a promising avenue to address this issue. However, existing SSMs typically rely on first-order ordinary differential equations (ODEs), limiting their capacity to capture higher order dynamics. To eliminate these limitations, we propose a second-order state space (S3) architecture, where the weather motion is decomposed as position and momentum in latent spaces to find multi-order behaviors. For sequential weather observations, we provide a recurrent counterpart of S3, which can be further parallelized through fast Fourier transform (FFT) for computational efficiency. Building on these S3 blocks, we develop a unified framework, S3Cast, for spatiotemporal sequence extrapolation. Empirical results demonstrate that S3Cast matches or exceeds the performance of state-of-the-art methods in the prediction of weather extremes, such as lightning, hail, and heavy precipitation.
Jianlun Liu, Xutao Li 0003, Yunming Ye
IEEE Trans. Geosci. Remote. Sens.2
2025 RePA: Rebalance Parameter Adaptation for Precipitation Nowcasting
Youran Wang, Baoquan Zhang, Yunming Ye, Xutao Li 0003
IEEE Trans. Geosci. Remote. Sens.5
2025 HPCR: Holistic Proxy-Based Contrastive Replay for Online Continual Learning
abstract
Online continual learning (OCL), aimed at developing a neural network that continuously learns new data from a single pass over an online data stream, generally suffers from catastrophic forgetting (CF). Existing replay-based methods alleviate forgetting by replaying partial old data in a proxy-based or contrastive-based replay manner, each with its own shortcomings. Our previous work proposes a novel replay-based method called proxy-based contrastive replay (PCR), which handles the shortcomings by achieving complementary advantages of both replay manners. In this work, we further conduct gradient and limitation analysis of PCR. The analysis results show that PCR still can be further improved in feature extraction, generalization, and anti-forgetting capabilities of the model. Hence, we developed a more advanced method named holistic PCR (HPCR). HPCR consists of three components, each tackling one of the limitations of PCR. The contrastive component conditionally incorporates anchor-to-sample pairs to PCR, improving the feature extraction ability. The second is a temperature component that decouples the temperature coefficient into two parts based on their gradient impacts and sets different values for them to enhance the generalization ability. The third is a distillation component that constrains the learning process with additional loss terms to improve the anti-forgetting ability. Experiments on four datasets consistently demonstrate the superiority of HPCR over various state-of-the-art methods.
Huiwei Lin, Shanshan Feng 0001, Baoquan Zhang, Xutao Li 0003, Yunming Ye
IEEE Trans. Neural Networks Learn. Syst.4
2024 iTrendRNN: An Interpretable Trend-Aware RNN for Meteorological Spatiotemporal Prediction
abstract
Accurate prediction of meteorological elements, such as temperature and relative humidity, is important to human livelihood, early warning of extreme weather, and urban governance. Recently, neural network-based methods have shown impressive performance in this field. However, most of them are overcomplicated and impenetrable. In this paper, we propose a straightforward and interpretable differential framework, where the key lies in explicitly estimating the evolutionary trends. Specifically, three types of trends are exploited. (1) The proximity trend simply uses the most recent changes. It works well for approximately linear evolution. (2) The sequential trend explores the global information, aiming to capture the nonlinear dynamics. Here, we develop an attention-based trend unit to help memorize long-term features. (3) The flow trend is motivated by the nature of evolution, i.e., the heat or substance flows from one region to another. Here, we design a flow-aware attention unit. It can reflect the interactions via performing spatial attention over flow maps. Finally, we develop a trend fusion module to adaptively fuse the above three trends. Extensive experiments on two datasets demonstrate the effectiveness of our method.
Chuyao Luo, Bowen Zhang 0005, Huiwei Lin, Xutao Li 0003, Yunming Ye
AAAI5
2024 MetaDiff: Meta-Learning with Conditional Diffusion for Few-Shot Learning
abstract
Equipping a deep model the ability of few-shot learning (FSL) is a core challenge for artificial intelligence. Gradient-based meta-learning effectively addresses the challenge by learning how to learn novel tasks. Its key idea is learning a deep model in a bi-level optimization manner, where the outer-loop process learns a shared gradient descent algorithm (called meta-optimizer), while the inner-loop process leverages it to optimize a task-specific base learner with few examples. Although these methods have shown superior performance on FSL, the outer-loop process requires calculating second-order derivatives along the inner-loop path, which imposes considerable memory burdens and the risk of vanishing gradients. This degrades meta-learning performance. Inspired by recent diffusion models, we find that the inner-loop gradient descent process can be viewed as a reverse process (i.e., denoising) of diffusion where the target of denoising is the weight of base learner but origin data. Based on this fact, we propose to model the gradient descent algorithm as a diffusion model and then present a novel conditional diffusion-based meta-learning, called MetaDiff, that effectively models the optimization process of base learner weights from Gaussian initialization to target weights in a denoising manner. Thanks to the training efficiency of diffusion models, our MetaDiff does not need to differentiate through the inner-loop path such that the memory burdens and the risk of vanishing gradients can be effectively alleviated for improving FSL. Experimental results show that our MetaDiff outperforms state-of-the-art gradient-based meta-learning family on FSL tasks.
Baoquan Zhang, Chuyao Luo, Demin Yu, Xutao Li 0003, Huiwei Lin, Yunming Ye, Bowen Zhang 0005
AAAI4
2024 DiffCast: A Unified Framework via Residual Diffusion for Precipitation Nowcasting
abstract
Precipitation nowcasting is an important spatiotemporal prediction task to predict the radar echoes sequences based on current observations, which can serve both meteorological science and smart city applications. Due to the chaotic evolution nature of the precipitation systems, it is a very challenging problem. Previous studies address the problem either from the perspectives of deterministic modeling or probabilistic modeling. However, their predictions suffer from the blurry, high-value echoes fading away and position inaccurate issues. The root reason of these issues is that the chaotic evolutionary precipitation systems are not appropriately modeled. Inspired by the nature of the systems, we propose to decompose and model them from the perspective of global deterministic motion and local stochastic variations with residual mechanism. A unified and flexible framework that can equip any type of spatio-temporal models is proposed based on residual diffusion, which effectively tackles the shortcomings of previous methods. Extensive experimental results on four publicly available radar datasets demonstrate the effectiveness and superiority of the proposed framework, compared to state-of-the-art techniques. Our code is publicly available at https://github.com/DeminYu98/DiffCast.
Demin Yu, Xutao Li 0003, Yunming Ye, Baoquan Zhang, Chuyao Luo, Kuai Dai, Xunlai Chen
CVPR2
2024 Codebook Transfer with Part-of-Speech for Vector-Quantized Image Modeling
abstract
Vector-Quantized Image Modeling (VQIM) is a fundamental research problem in image synthesis, which aims to represent an image with a discrete token sequence. Existing studies effectively address this problem by learning a discrete codebook from scratch and in a code-independent manner to quantize continuous representations into discrete tokens. However, learning a codebook from scratch and in a code-independent manner is highly challenging, which may be a key reason causing codebook collapse, i.e., some code vectors can rarely be optimized without regard to the relationship between codes and good codebook priors such that die off finally. In this paper, inspired by pretrained language models, we find that these language models have actually pretrained a superior codebook via a large number of text corpus, but such information is rarely exploited in VQIM. To this end, we propose a novel codebook transfer framework with part-of-speech, called VQCT, which aims to transfer a well-trained codebook from pretrained language models to VQIM for robust codebook learning. Specifically, we first introduce a pretrained codebook from language models and part-of-speech knowledge as priors. Then, we construct a vision-related codebook with these priors for achieving codebook transfer. Finally, a novel codebook transfer network is designed to exploit abundant semantic relationships between codes contained in pretrained codebooks for robust VQIM codebook learning. Experimental results on four datasets show that our VQCT method achieves superior VQIM performance over previous state-of-the-art methods.
Baoquan Zhang, Huaibin Wang, Chuyao Luo, Xutao Li 0003, Guotao Liang, Yunming Ye, Xiaochen Qi
CVPR4
2024 LG-VQ: Language-Guided Codebook Learning
abstract
Vector quantization (VQ) is a key technique in high-resolution and high-fidelity image synthesis, which aims to learn a codebook to encode an image with a sequence of discrete codes and then generate an image in an auto-regression manner. Although existing methods have shown superior performance, most methods prefer to learn a single-modal codebook (\emph{e.g.}, image), resulting in suboptimal performance when the codebook is applied to multi-modal downstream tasks (\emph{e.g.}, text-to-image, image captioning) due to the existence of modal gaps. In this paper, we propose a novel language-guided codebook learning framework, called LG-VQ, which aims to learn a codebook that can be aligned with the text to improve the performance of multi-modal downstream tasks. Specifically, we first introduce pre-trained text semantics as prior knowledge, then design two novel alignment modules (\emph{i.e.}, Semantic Alignment Module, and Relationship Alignment Module) to transfer such prior knowledge into codes for achieving codebook text alignment. In particular, our LG-VQ method is model-agnostic, which can be easily integrated into existing VQ models. Experimental results show that our method achieves superior performance on reconstruction and various multi-modal downstream tasks.
Guotao Liang, Baoquan Zhang, Yaowei Wang 0001, Yunming Ye, Xutao Li 0003, Huaibin Wang, Chuyao Luo, Kola Ye, Linfeng Luo
NeurIPS5
2024 AMANet: An Adaptive Memory Attention Network for video cloud detection
Shanshan Feng 0001, Yingling Quan, Yunming Ye, Yong Xu 0001, Xutao Li 0003, Baoquan Zhang
Pattern Recognit.6
2024 Exploring and Exploiting High-Order Spatial-Temporal Dynamics for Long-Term Frame Prediction
abstract
Long-term spatial-temporal frame prediction focuses on predicting future image frames precisely, which has numerous applications in real-world scenarios. Existing deep learning prediction models mainly rely on advanced neural network architectures to model complicated spatial-temporal features, which make few efforts to explore high-order correlations to better capture long-term dynamics. Their prediction on long-term frames suffers from inaccurate visual and motion detail issue. In this article, we propose a high-order prediction model for long-term frame prediction, which improves the appearance and motion details by designing special high-order correlation modules in two aspects. First, to enhance the appearance details of predicted frames, we propose a high-order appearance encoder module, where high-order appearance features can be effectively captured with a carefully designed Non-local ConvLSTM. Second, to guarantee the motion accuracy of predicted sequences, we carefully design a high-order motion encoder module, which can accurately capture and preserve the high-order motion patterns with adaptive motion extractors and progressive memory banks, respectively. Comprehensive experiments are conducted on six challenging datasets from real-world scenarios, which demonstrate the effectiveness and superiority of our proposed method over state-of-the-art methods.
Kuai Dai, Xutao Li 0003, Yunming Ye, Yaowei Wang 0001, Shanshan Feng 0001, Di Xian
IEEE Trans. Circuits Syst. Video Technol.2
2024 Spherical Neural Operator Network for Global Weather Prediction
abstract
Global weather forecast is an important spatial-temporal prediction problem, which can provide numerous societal benefits such as extreme weather forewarning, traffic scheduling, and agricultural planning. Though many spatial-temporal prediction models have been proposed, they suffer from two drawbacks for global weather forecasts, namely (i) ignoring the physical mechanism and spherical characteristics and (ii) not effectively exploiting the global and local correlations. To address the above drawbacks, in this paper, we formalize global weather state dynamics as partial differential equations (PDEs) in spherical space and infer the state of the global weather system by solving these PDEs. Specifically, we use Green’s function method to solve the PDEs and find that the solution of the spherical PDEs can be obtained by the spherical convolution. We further proposed a novel Spherical Neural Operator, SNO, which consists of spherical convolution and vanilla convolution. The former is used to solve these PDEs and model the global correlations in spherical space, and the latter is used to capture the local correlations. Upon the operator, a global weather prediction model is developed. Extensive experimental results demonstrate the effectiveness and superiority of our method over state-of-the-art approaches.
Kenghong Lin, Xutao Li 0003, Yunming Ye, Shanshan Feng 0001, Baoquan Zhang, Guangning Xu
IEEE Trans. Circuits Syst. Video Technol.2
2024 TLS-MWP: A Tensor-Based Long- and Short-Range Convolution for Multiple Weather Prediction
abstract
Weather prediction plays a crucial role in human development. Recently, deep learning has demonstrated promising prospects in weather forecasting by integrating convolutional neural networks (CNNs) and recurrent neural networks (RNNs). However, two main challenges still exist in multiple weather condition prediction. The first challenge considers multiple weather condition correlations in predictions. The second challenge is how to model long- and short-range spatial dependencies under multiple weather conditions. A novel operator named as tensor-based long- and short-range convolution (TLS-Conv) is proposed to address these challenges. Within this operator, the node & relation attention is utilized to identify the contributions of spatial grid points and weather conditions for prediction. Additionally, the adaptive tensor graph convolution (ATGCN) is tailored to dynamically capture long-range spatial dependencies within multiple weather conditions. Finally, the traditional convolution is integrated with the ATGCN to model both long- and short-range spatial dependencies and weather condition correlations. Building upon the TLS-Conv, the tensor-based long- and short-range convolution for multiple weather prediction (TLS-MWP) model is proposed to predict multiple weather conditions. Extensive experiments are conducted under real-world weather conditions to evaluate its performance. These results unequivocally demonstrate that TLS-MWP surpasses previous methods. The code is available on GitHub at: https://github.com/xuguangning1218/TLS_MWP.
Guangning Xu, Michael Kwok-Po Ng, Yunming Ye, Xutao Li 0003, Bowen Zhang 0005, Zhichao Huang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 LGCNet: A Cloud Detection Method in Remote Sensing Images Using Local and Global Semantics
abstract
Detecting and eliminating clouds is a crucial step in remote sensing image (RSI) preprocessing. The removal of clouds can significantly enhance the performance of subsequent remote sensing applications. Existing deep learning (DL)-based cloud detection methods extract semantic information to improve feature representation and, subsequently, detection performance. However, these methods do not fully utilize the potential of context semantic information. Besides, to capture semantics from large receptive fields, they employ convolution operators with large kernel sizes, which results in high computational costs. Thus, these computationally heavy models are not suitable for resource-limited devices, particularly satellites. To address this issue, we propose a cloud detection model, LGCNet. This model efficiently extracts both local and global contextual information, fully utilizing semantics while reducing resource usage. LGCNet is built on an encoder-decoder structure. Specifically, the encoder extracts local scale-aware semantics through proposed local semantic blocks (LSBs), which are then skip-connected to the decoder. This approach provides adaptive and diverse local contextual information. On the top of the encoder, the high-level global semantics are captured via the proposed global feature TransBlock (GFTB). A variety of extracted semantics ensure improved detection performance. We evaluate the proposed method using two public datasets: LandSat8 and Moderate-Resolution Imaging Spectroradiometer (MODIS). We conducted experiments on both a server and an edge computing device. Our extensive experiments revealed that LGCNet outperforms other lightweight cloud detection and semantic segmentation methods in terms of performance and computational load.
Shanshan Feng 0001, Huaibin Wang, Baoquan Zhang, Pengjuan Yao, Chuyao Luo, Yunming Ye, Yong Xu 0001, Xutao Li 0003
IEEE Trans. Geosci. Remote. Sens.9
2024 A Practical Online Incremental Learning Framework for Precipitation Nowcasting
abstract
Precipitation nowcasting plays an important role in our life. Many deep learning-based methods are proposed for precipitation nowcasting by predicting radar echo sequence over the past years, and achieving better performance than traditional approaches. However, all of them are based on a static model, which is trained in offline learning and does not adapt to real-time changing precipitation data. Recently, online incremental learning (OIL) has been proposed to dynamically update the model by continually learning new data and preventing the forgetting of historical knowledge in an online fashion. While effective, existing OIL approaches that focus on a classification task are not suitable for the regression task of precipitation nowcasting. To fill this gap, we try to propose a novel OIL framework for precipitation nowcasting. By analyzing its characteristics, we find three challenges: 1) the distributions of radar echo maps in different rainfall events are different; 2) in each rainfall event, there is always an inevitable delay between the timestamps of the training and testing samples; and 3) the real-time requirement for model prediction is very high, which has strict limitations on the training speed of the model. Based on these observations, we propose a practical OIL framework based on gradient activation mapping (GAM). It can be mainly divided into three components: 1) recall training strategy (RTS) is used to eliminate the interference caused by distributions of different events; 2) iterative approximation training (IAT) is designed to align the timestamps of the training and testing; and 3) moreover, we propose gradient activation mapping weight (GAMW) to improve the training effectiveness. Extensive experiments show that the proposed framework can improve the performance of the model stably and effectively. Especially in those heavy rainfall regions where it usually causes more threat to human activity, the improvement is significant.
Chuyao Luo, Zheng Zhang 0046, Huiwei Lin, Baoquan Zhang, Xutao Li 0003, Yunming Ye
IEEE Trans. Geosci. Remote. Sens.5
2024 Cross-Modal Hashing With Feature Semi-Interaction and Semantic Ranking for Remote Sensing Ship Image Retrieval
abstract
Cross-modal hashing plays a pivotal role in large-scale remote sensing (RS) ship image retrieval. RS ship images often exhibit similar overall appearance with subtle differences. Existing hashing methods typically employ feature non-interaction strategies to generate common hash codes, which may not effectively capture the correlations between cross-modal ship images to reduce intermodality discrepancies. To address this issue, we propose a novel cross-modal hashing approach based on feature semi-interaction and semantic ranking (FSISR) for RS ship image retrieval. Our FSISR approach not only captures intricate correlations between different ship image modalities, but also enables the construction of hash tables for large-scale retrieval. FSISR comprises a feature semi-interaction module and a semantic ranking objective function. The semi-interaction module utilizes clustering centers from one modality to learn the correlations between two modalities and generate robust shared representations. The objective function optimizes these representations in a common Hamming space, consisting of a shared semantic alignment loss and a margin-free ranking loss. The alignment loss employs a shared semantic layer to preserve label-level similarity, while the ranking loss incorporates hard examples to establish a margin-free loss that captures similarity ranking relationships. We evaluate the performance of our method on benchmark datasets and demonstrate its effectiveness for cross-modal RS ship image retrieval.https://github.com/sunyuxi/FSISR.
Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Yifang Ban, Sebastian Hafner, Xutao Li 0003, Chuyao Luo, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.7
2024 Multiscale and Multilevel Feature Fusion Network for Quantitative Precipitation Estimation With Passive Microwave
abstract
Passive microwave (PMW) radiometers have been widely utilized for quantitative precipitation estimation (QPE) by leveraging the relationship between brightness temperature (Tb) and rain rate. Nevertheless, accurate precipitation estimation remains a challenge due to the intricate relationship between them, which is influenced by a diverse range of complex atmospheric and surface properties. In addition, the inherent skew distribution of rainfall values prevents models from correctly addressing extreme precipitation events, leading to a significant underestimation. This article presents a novel model called the multiscale and multilevel feature fusion network (MSMLNet), consisting of two essential components: a multiscale feature extractor and a multilevel regression predictor. The feature extractor is specifically designed to extract characteristics from multiple scales, enabling the model to incorporate various meteorological conditions, as well as atmospheric and surface information in the surrounding environment. The regression predictor first assesses the probabilities of multiple rainfall levels for each observed pixel and then extracts features of different levels separately. The multilevel features are fused according to the predicted probabilities. This approach allows each submodule only to focus on a specific range of precipitation, avoiding the undesirable effects of skew distributions. To evaluate the performance of MSMLNet, various deep learning methods are adapted for the precipitation retrieval task, and a PWM-based product from the global precipitation measurement (GPM) mission is also used for comparison. Extensive experiments show that MSMLNet surpasses GMI-based products and the most advanced deep learning approaches by 17.9% and 2.5% in root mean square error (RMSE), and 54.2% and 4.0% in CSI-10, respectively. Moreover, we demonstrate that MSMLNet significantly mitigates the propensity for underestimating heavy precipitation events and has a consistent and outstanding performance in estimating precipitation across various levels.
Xutao Li 0003, Kenghong Lin, Chuyao Luo, Yunming Ye, Xiuqing Hu
IEEE Trans. Geosci. Remote. Sens.2
2024 MCSDNet: Mesoscale Convective System Detection Network via Multiscale Spatiotemporal Information
abstract
The accurate detection of mesoscale convective systems (MCSs) is crucial for meteorological monitoring due to their potential to cause significant destruction through severe weather phenomena, such as hail, thunderstorms, and heavy rainfall. However, the existing methods for MCS detection mostly targets on single-frame detection, which just considers the static characteristics and ignores the temporal evolution in the life cycle of MCS. In this article, we propose a novel encoder-decoder neural network named mesoscale convective system detection network (MCSDNet) to detect MCS regions. MCSDNet has a simple architecture and is easy to expand. Different from the previous models, MCSDNet targets on multiframes detection and leverages multiscale spatiotemporal information in remote sensing imagery (RSI). As far as we know, it is the first work to utilize multiscale spatiotemporal information to detect MCS regions. First, we design a multiscale spatiotemporal information module to extract multilevel semantic from different encoder levels, which makes our models can extract more detail spatiotemporal features. Second, spatiotemporal mix unit (STMU), a dual spatiotemporal attention, is introduced to MCSDNet to capture both intraframe features and interframe. Finally, we present MCS remote sensing image (MCSRSI) the first publicly available dataset for multiframes MCS detection based on FY-4A satellite. We also conduct several experiments on MCSRSI and find that our proposed MCSDNet achieves the best performance on MCS detection task when comparing with other baseline methods. We hope that the combination of our open-access dataset and promising results will encourage the future research for MCS detection task and provide a robust framework for related tasks in atmospheric science. Our code is available at:https://github.com/250HandsomeLiang/MCSDNet.git
Baoquan Zhang, Jiajun Liang, Rui Ye 0002, Chuyao Luo, Xutao Li 0003, Yunming Ye, Xukai Fu
IEEE Trans. Geosci. Remote. Sens.5
2024 Toward a Variation-Aware and Interpretable Model for Radar Image Sequence Prediction
abstract
Radar image sequence prediction (RISP) aims to predict future radar images based on historical observations. In the past few years, neural network-based methods have shown impressive performance for RISP. However, two limitations stills exist. 1) They fail to exploit variation information when capturing spatial dependencies. 2) They neglect to analyze and interpret the model. In this article, we propose a variation-aware prediction model for the first limitation, and develop a relevance propagation technique for the second one. Specifically, 1) we recustomize the vanilla convolution by introducing a variation-aware item. The new convolution unit yields two advantages when capturing spatial dependencies, i.e., exploiting variation information and offering spatially-varying kernels. As a result, it can learn the diverse and complex radar echo patterns. By equipping the unit into a typical network (PredRNN), we propose a novel prediction model, dubbed as VA-PredRNN. 2) As for analyzing our model, we propagate the output backward layer by layer till the input. Hence, we can reveal the relevance between the output and the intermediate states. To the best of the authors' knowledge, this is the first work to study the interpretability of a multilayer RISP model. We conduct extensive experiments on two datasets, and the results demonstrate the effectiveness of our VA-PredRNN. We also carry out a series of analyses using the proposed relevance propagation technique. According to the results, we discover the importance of different states.
Yunming Ye, Bowen Zhang 0005, Huiwei Lin, Yuxi Sun 0002, Xutao Li 0003, Chuyao Luo
IEEE Trans. Ind. Informatics6
2023 RotDiff: A Hyperbolic Rotation Representation Model for Information Diffusion Prediction
abstract
The massive amounts of online user behavior data on social networks allow for the investigation of information diffusion prediction, which is essential to comprehend how information propagates among users. The main difficulty in diffusion prediction problem is to effectively model the complex social factors in social networks and diffusion cascades. However, existing methods are mainly based on Euclidean space, which cannot well preserve the underlying hierarchical structures that could better reflect the strength of user influence. Meanwhile, existing methods cannot accurately model the obvious asymmetric features of the diffusion process. To alleviate these limitations, we utilize rotation transformation in the hyperbolic to model complex diffusion patterns. The modulus of representations in the hyperbolic space could effectively describe the strength of the user's influence. Rotation transformations could represent a variety of complex asymmetric features. Further, rotation transformation could model various social factors without changing the strength of influence. In this paper, we propose a novel hyperbolic rotation representation model RotDiff for the diffusion prediction problem. Specifically, we first map each social user to a Lorentzian vector and use two groups of transformations to encode global social factors in the social graph and the diffusion graph. Then, we combine attention mechanism in the hyperbolic space with extra rotation transformations to capture local diffusion dependencies within a given cascade. Experimental results on five real-world datasets demonstrate that the proposed model RotDiff outperforms various state-of-the-art diffusion prediction models.
Hongliang Qiao, Shanshan Feng 0001, Xutao Li 0003, Huiwei Lin, Han Hu 0003, Wei Wei 0002, Yunming Ye
CIKM3
2023 UER: A Heuristic Bias Addressing Approach for Online Continual Learning
abstract
Online continual learning aims to continuously train neural networks from a continuous data stream with a single pass-through data. As the most effective approach, the rehearsal-based methods replay part of previous data. Commonly used predictors in existing methods tend to generate biased dot-product logits that prefer to the classes of current data, which is known as a bias issue and a phenomenon of forgetting. Many approaches have been proposed to overcome the forgetting problem by correcting the bias; however, they still need to be improved in online fashion. In this paper, we try to address the bias issue by a more straightforward and more efficient method. By decomposing the dot-product logits into an angle factor and a norm factor, we empirically find that the bias problem mainly occurs in the angle factor, which can be used to learn novel knowledge as cosine logits. On the contrary, the norm factor abandoned by existing methods helps remember historical knowledge. Based on this observation, we intuitively propose to leverage the norm factor to balance the new and old knowledge for addressing the bias. To this end, we develop a heuristic approach called unbias experience replay (UER). UER learns current samples only by the angle factor and further replays previous samples by both the norm and angle factors. Extensive experiments on three datasets show that UER achieves superior performance over various state-of-the-art methods. The code is in https://github.com/FelixHuiweiLin/UER.
Huiwei Lin, Shanshan Feng 0001, Baoquan Zhang, Hongliang Qiao, Xutao Li 0003, Yunming Ye
ACM Multimedia5
2023 Multi-view knowledge graph fusion via knowledge-aware attentional graph neural network
Zhichao Huang 0001, Xutao Li 0003, Yunming Ye, Baoquan Zhang, Guangning Xu, Wensheng Gan
Appl. Intell.2
2023 Spatiotemporal prediction in three-dimensional space by separating information interactions
Bowen Zhang 0005, Yunming Ye, Shanshan Feng 0001, Xutao Li 0003
Appl. Intell.5
2023 TFG-Net: Tropical Cyclone Intensity Estimation from a Fine-grained perspective with the Graph convolution neural network
Guangning Xu, Yan Li 0040, Xutao Li 0003, Yunming Ye, Qingquan Lin, Zhichao Huang 0001, Shidong Chen
Eng. Appl. Artif. Intell.4
2023 Exploiting Spatial-Temporal Dynamics for Satellite Image Sequence Prediction
abstract
Satellite image sequence prediction is a challenging and significant task. Existing deep learning methods for the task make predictions mainly based on low-level pixel-wise features, which fail to model the sophisticated spatial-temporal features of satellite image sequences and deliver unsatisfactory performance. In this paper, we present a Hierarchical Spatial-Temporal network (HSTnet) for satellite image sequence prediction. With a carefully designed hierarchical feature extraction mechanism, HSTnet can learn effective spatial-temporal features from both pixel level and patch level. In addition, to better capture patch-level spatial-temporal dynamics, a dual-branch Transformer is proposed to model patch-level spatial and temporal features, respectively. Comprehensive experiments on the FY-4A satellite dataset demonstrate the superiority and effectiveness of our proposed method HSTnet over state-of-the-art approaches.
Kuai Dai, Yongshen Long, Xutao Li 0003, Shanshan Feng 0001, Yunming Ye
IEEE Geosci. Remote. Sens. Lett.5
2023 Correction to: AM-ConvGRU: a spatio-temporal model for typhoon path prediction
Guangning Xu, Di Xian, Philippe Fournier-Viger, Xutao Li 0003, Yunming Ye, Xiuqing Hu
Neural Comput. Appl.4
2023 UNIMEMnet: Learning long-term motion and appearance dynamics for video prediction with a unified memory network
Kuai Dai, Xutao Li 0003, Chuyao Luo, Wuqiao Chen, Yunming Ye, Shanshan Feng 0001
Neural Networks2
2023 Interpretable local flow attention for multi-step traffic flow prediction
Bowen Zhang 0005, Shanshan Feng 0001, Yunming Ye, Xutao Li 0003
Neural Networks5
2023 Multi-relational graph convolutional networks: Generalization guarantees and experiments
Xutao Li 0003, Michael Kwok-Po Ng, Guangning Xu, Andy M. Yip
Neural Networks1
2023 WDMNet: Modeling diverse variations of regional wind speed for multi-step predictions
Rui Ye 0002, Shanshan Feng 0001, Xutao Li 0003, Yunming Ye, Baoquan Zhang, Yaowei Wang 0001
Neural Networks3
2023 PEPNet: A barotropic primitive equations-based network for wind speed prediction
Rui Ye 0002, Baoquan Zhang, Xutao Li 0003, Yunming Ye
Neural Networks3
2023 Prototype Completion for Few-Shot Learning
abstract
Few-shot learning (FSL) aims to recognize novel classes with few examples. Pre-training based methods effectively tackle the problem by pre-training a feature extractor and then fine-tuning it through the nearest centroid based meta-learning. However, results show that the fine-tuning step makes marginal improvements. In this paper, 1) we figure out the reason, i.e., in the pre-trained feature space, the base classes already form compact clusters while novel classes spread as groups with large variances, which implies that fine-tuning feature extractor is less meaningful; 2) instead of fine-tuning feature extractor, we focus on estimating more representative prototypes. Consequently, we propose a novel prototype completion based meta-learning framework. This framework first introduces primitive knowledge (i.e., class-level part or attribute annotations) and extracts representative features for seen attributes as priors. Second, a part/attribute transfer network is designed to learn to infer the representative features for unseen attributes as supplementary priors. Finally, a prototype completion network is devised to learn to complete prototypes with these priors. Moreover, to avoid the prototype completion error, we further develop a Gaussian based prototype fusion strategy that fuses the mean-based and completed prototypes by exploiting the unlabeled samples. At last, we also develop an economic prototype completion version for FSL, which does not need to collect primitive knowledge, for a fair comparison with existing FSL methods without external knowledge. Extensive experiments show that our method: i) obtains more accurate prototypes; ii) achieves superior performance on both inductive and transductive FSL settings.
Baoquan Zhang, Xutao Li 0003, Yunming Ye, Shanshan Feng 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 NPDN-3D: A 3D neural partial differential network for spatiotemporal prediction
Shanshan Feng 0001, Yunming Ye, Xutao Li 0003, Bowen Zhang 0005, Shidong Chen
Pattern Recognit.4
2023 Cross-Domain Aspect-Based Sentiment Classification by Exploiting Domain- Invariant Semantic-Primary Feature
abstract
Aspect-based sentiment analysis is an important task in fine-grained sentiment analysis, which aims to infer the sentiment towards a given aspect. Previous studies have shown notable success when sufficient labeled training data is available. However, annotating adequate data is labor-intensive, which sets substantial barriers for generalizing the sentiment predictor to the new domain. Two main challenges exist in cross-domain aspect-based sentiment analysis. One challenge is acquiring the domain-invariant knowledge; the other challenge is mining the syntactic-related words towards the aspect-term. In this article, we propose a transformer-based semantic-primary knowledge transferring network (TSPKT) for cross-domain aspect-term sentiment analysis, which utilizes semantic-primary knowledge as a bridge to enable knowledge transfer across domains. Specifically, we first build an S-Graph from external semantic lexicons, and extract the semantic-primary knowledge from the S-Graph. Second, AoaGraphormer is proposed to learn the syntactically relevant words towards the aspect-term. Third, we extend the standard biLSTM classifier to fully integrate the semantic-primary knowledge by adding a novel knowledge-aware memory unit (KAMU) to the biLSTM cell. Extensive experiments on six cross-domain setups demonstrate the superiority of TSPKT against the state-of-the-art baseline methods.
Bowen Zhang 0005, Xianghua Fu, Chuyao Luo, Yunming Ye, Xutao Li 0003, Liwen Jing 0001
IEEE Trans. Affect. Comput.5
2023 Adaptive Transfer of Graph Neural Networks for Few-Shot Molecular Property Prediction
abstract
Few-Shot Molecular Property Prediction (FSMPP) is an improtant task on drug discovery, which aims to learn transferable knowledge from base property prediction tasks with sufficient data for predicting novel properties with few labeled molecules. Its key challenge is how to alleviate the data scarcity issue of novel properties. Pretrained Graph Neural Network (GNN) based FSMPP methods effectively address the challenge by pre-training a GNN from large-scale self-supervised tasks and then finetuning it on base property prediction tasks to perform novel property prediction. However, in this paper, we find that the GNN finetuning step is not always effective, which even degrades the performance of pretrained GNN on some novel properties. This is because these molecule-property relationships among molecules change across different properties, which results in the finetuned GNN overfits to base properties and harms the transferability performance of pretrained GNN on novel properties. To address this issue, in this paper, we propose a novel Adaptive Transfer framework of GNN for FSMPP, called ATGNN, which transfers the knowledge of pretrained and finetuned GNNs in a task-adaptive manner to adapt novel properties. Specifically, we first regard the pretrained and finetuned GNNs as model priors of target-property GNN. Then, a task-adaptive weight prediction network is designed to leverage these priors to predict target GNN weights for novel properties. Finally, we combine our ATGNN framework with existing FSMPP methods for FSMPP. Extensive experiments on four real-world datasets, i.e., Tox21, SIDER, MUV, and ToxCast, show the effectiveness of our ATGNN framework.
Baoquan Zhang, Chuyao Luo, Hao Jiang 0051, Shanshan Feng 0001, Xutao Li 0003, Bowen Zhang 0005, Yunming Ye
IEEE ACM Trans. Comput. Biol. Bioinform.5
2023 On Understanding of Spatiotemporal Prediction Model
abstract
Recently, explainable artificial intelligence has received considerable attention. Most existing studies are focusing on the tasks of CNNs-based image classification and RNNs-based time series analysis. In this paper, we pay attention to the more complicated spatiotemporal predictive learning task (SPLT), where both the spatial and temporal information play important roles. To explain the internal mechanism of spatiotemporal prediction models, we propose a comprehensive analysis method. Specifically, with a typical encoder-decoder framework, we focus on two core issues of SPLT: image generation and spatiotemporal dynamics. For the first issue, we develop aquantitative channel perturbationmethod to explore the importance of features to prediction. Furthermore, we propose a technique called thesynthesis of multiple independent componentsto analyze how these features generate the prediction. According to the experimental results, thecoarse- and fine-grainedsynthesis (CFGS) mechanism is drawn for image generation in SPLT. For the second issue, we propose astate decompositiontechnique and astate expansiontechnique to disentangle coupled signals in the spatiotemporal dynamical system. This helps us to explore the mechanism of forming motion. Moreover, to diagnose the movement of a particular region during analysis, we propose a fluorescent stamp-based technique. By observing extensive experimental results, we summarize a collaboration mechanism to explain how the motion is formed in SPLT, namely, theextending the present and erasing the past (EPEP)mechanism. To the best of our knowledge, this is the first work to interpret the internal mechanism of SPLT models.
Xutao Li 0003, Yunming Ye, Shanshan Feng 0001, Chuyao Luo, Bowen Zhang 0005
IEEE Trans. Circuits Syst. Video Technol.2
2023 Anchor Assisted Experience Replay for Online Class-Incremental Learning
abstract
Online class-incremental learning (OCIL) studies the problem of mitigating the phenomenon of catastrophic forgetting while learning new classes from a continuously non-stationary data stream. Existing approaches mainly constrain the updating of parameters to prevent the drift of previous classes that reflects the movement of samples in the embedding space. Although this kind of drift can be relieved to some extent by existing approaches, it is usually inevitable. Therefore, only prevention of drift is not enough, and we also need to further compensate for it. To this end, for each previous class, we exploit the sample with the smallest loss value as its anchor, which can representatively characterize the corresponding class. Based on the assistance of anchors, we present a novel Anchor Assisted Experience Replay (AAER) method that not only prevents the drift but also compensates for the inevitable drift to overcome the catastrophic forgetting. Specifically, we design a Drift-Prevention with Anchor (DPA) operation, which plays a preventive role by reducing the drift implicitly as well as encouraging the samples with the same label cluster tightly. Moreover, we propose a Drift-Compensation with Anchor (DCA) operation that contains two remedy mechanisms: one is Forward-offset which keeps embedding of previous data but estimates new classification centers; the other is just the opposite named Backward-offset, which keeps the old classification centers unchanged but updates the embedding of previous data. We conduct extensive experiments on three real-world datasets, and empirical results consistently demonstrate the superior performance of AAER over various state-of-the-art methods.
Huiwei Lin, Shanshan Feng 0001, Xutao Li 0003, Yunming Ye
IEEE Trans. Circuits Syst. Video Technol.3
2023 MetaDT: Meta Decision Tree With Class Hierarchy for Interpretable Few-Shot Learning
abstract
Few-Shot Learning (FSL) is a challenging task, which aims to recognize novel classes with few examples. Recently, lots of methods have been proposed from the perspective of meta-learning and representation learning. However, few works focus on the interpretability of FSL decision process. In this paper, we take a step towards the interpretable FSL by proposing a novel meta-learning based decision tree framework, namely, MetaDT. In particular, the FSL interpretability is achieved from two aspects, i.e., a concept aspect and a visual aspect. On the concept aspect, we first introduce a tree-like concept hierarchy as FSL prior. Then, resorting to the prior, we split each few-shot task to a set of subtasks with different concept levels and then perform class prediction via a model of decision tree. The advantage of such design is that a sequence of high-level concept decisions that lead up to a final class prediction can be obtained, which clarifies the FSL decision process. On the visual aspect, a set of subtask-specific classifiers with visual attention mechanism is designed to perform decision at each node of the decision tree. As a result, a subtask-specific heatmap visualization can be obtained to achieve the decision interpretability of each tree node. At last, to alleviate the data scarcity issue of FSL, we regard the prior of concept hierarchy as an undirected graph, and then design a graph convolution-based decision tree inference network as our meta-learner to infer parameters of the decision tree. Extensive experiments on performance comparison and interpretability analysis show superiority of our MetaDT.
Baoquan Zhang, Hao Jiang 0051, Xutao Li 0003, Shanshan Feng 0001, Yunming Ye, Rui Ye 0002
IEEE Trans. Circuits Syst. Video Technol.3
2023 Learning Spatial-Temporal Consistency for Satellite Image Sequence Prediction
abstract
As an extremely challenging spatial-temporal sequence prediction task, satellite image sequence prediction has various and significant applications in real-world scenarios. Although lots of deep learning prediction models are developed for spatial-temporal sequence prediction, the methods still deliver unsatisfactory performance in terms of keeping spatial-temporal consistency, which leads to inaccurate and blurry satellite image sequence predictions. To maintain spatial-temporal consistency and achieve high-quality satellite image sequence prediction, we propose a novel and effective spatial-temporal consistency network (STCNet). In STCNet, a multi-level motion memory-based predictor is proposed to accurately predict motion patterns of satellite image sequences to ensure temporal consistency. Then, a time-variant frame discriminator is carefully designed and proposed, which can enhance the perception quality of predicted frames to guarantee spatial consistency and simultaneously maintain the motion coherency of predicted sequences. Moreover, a scheduled sampling strategy is proposed to reduce the optimizing difficulty and better train the proposed method. Comprehensive experiments conducted on satellite image sequences from FY-4A meteorological satellite verify the effectiveness, applicability, and adaptability of our proposed method compared to state-of-the-art approaches under challenging scenarios.
Kuai Dai, Xutao Li 0003, Shenyuan Lu, Yunming Ye, Di Xian, Danyu Qin
IEEE Trans. Geosci. Remote. Sens.2
2023 TRCDNet: A Transformer Network for Video Cloud Detection
abstract
In Remote Sensing Image (RSI) pre-processing steps, detecting and removing cloudy areas is a critical task. Recently, cloud detection methods based on deep neural networks achieve outstanding performance over traditional methods. Current approaches mostly focus on cloud detection on a single image captured by polar-orbiting satellites. However, there is another type of meteorological satellite - geostationary satellite, which can capture temporal consecutive frames of a particular location. Therefore, the cloud detection task targeting at geostationary satellite can be treated as a video cloud detection task. And in addition to extracting features on a single image, extracting and making full use of the relations between sequential frames is also important. To tackle this problem, we design a deep learning video cloud detection model: Transformer Network for Video Cloud Detection (TRCDNet). The proposed network is based on the encoder-decoder structure. In the encoder, the module ContextGhostLayer is proposed to encode more semantic information to tackle the challenging problems like thin cloud in RSIs. Besides, we design a transformer-based Video Sequence Transformer (VSTR) block. Based on attention mechanism, VSTR can fully extract the across-frame relations. In the proposed decoder, the cloud masks are recovered gradually to the same scale as the input image. To evaluate the methods, we create a Video Cloud Detection dataset based on the captured videos from Fengyun 4 (FY-4) satellite: Fengyun4aCloud. Extensive experiments of current cloud detection methods, semantic segmentation methods, and video semantic segmentation methods indicate that the designed TRCDNet achieves state-of-art performance in video cloud detection.
Shanshan Feng 0001, Yingling Quan, Yunming Ye, Xutao Li 0003, Yong Xu 0001, Baoquan Zhang, Zhihao Chen 0010
IEEE Trans. Geosci. Remote. Sens.5
2023 Cross-View Object Geo-Localization in a Local Region With Satellite Imagery
abstract
Cross-view geo-localization is a critical task in various applications, such as smart city management and disaster monitoring. Current methods typically divide a satellite image into patches and use these patches to identify the geographic location of a query image. However, these methods can only provide the location of an image rather than the location of a specific object of interest. This makes it difficult to link these methods to GeoDatabases to obtain detailed information about a target object, such as its name and construction time. To overcome this limitation, we propose a novel problem of cross-view object geo-localization in a local region with high-resolution satellite images. This problem includes two main challenges: accurately identifying the location of an object and distinguishing the target object from others in satellite images. To address these challenges, we present a new Detection-based Geo-localization method called DetGeo, which consists of an object detection-based framework with a two-branch encoder and a query-aware cross-view fusion module. DetGeo uses cross-view images as input to the detector to provide object-level geo-localization. The fusion module employs cross-view spatial attention to focus on relevant areas of target objects during cross-view feature fusion. To evaluate our method, we constructed a new Cross-View Object Geo-Localization dataset called CVOGL, which comprises ground-view or drone-view images as query images and satellite-view images as geo-tagged reference images. Comprehensive experiments are conducted to demonstrate the effectiveness of our method on CVOGL. https://github.com/sunyuxi/DetGeo.
Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Shanshan Feng 0001, Xutao Li 0003, Chuyao Luo, Puzhao Zhang, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.6
2023 Consistency Center-Based Deep Cross-Modal Hashing for Multisource Remote Sensing Image Retrieval
abstract
Cross-modal hashing aims to retrieve similar images from large-scale Earth Observation (EO) data archives, which typically contain multiple satellite sources of remote sensing (RS) images. However, existing cross-modal hashing methods primarily focus on dual-source RS images and often face two main limitations when retrieving multi-source RS images. Firstly, these methods exhibit significant redundancy as they require handling all possible dual-source combinations in multi-source RS images. Secondly, they often rely on pairwise or triplet image sources to construct objective functions, which are not significantly effective in reducing the discrepancies among multiple RS image sources. To address these limitations, we propose a novel Consistency Center-based deep cross-modal Hashing method called C2Hash for multi-source RS image retrieval. Our C2Hash employs a multi-branch hashing network to directly encode multi-source RS images into unified hash codes, thereby offering higher processing efficiency. Furthermore, C2Hash introduces consistency centers to construct a novel objective function. The consistency center represents the shared semantic features among similar multi-source RS images and is generated by a label hashing network. The objective function encourages similar multi-source RS images to approach the same consistency center to align all image sources in a unified Hamming space. Our method can effectively reduce the discrepancies across multiple image sources and generate unified hash codes. To evaluate its effectiveness, we construct a new Multi-Source RS Image dataset called MSRSI, comprising five different types of image sources. We conduct comprehensive experiments to demonstrate the superior performance of our method on the MSRSI dataset. https://github.com/sunyuxi/C2Hash.
Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Xutao Li 0003, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.5
2023 H-Diffu: Hyperbolic Representations for Information Diffusion Prediction
abstract
With the proliferation of online social networks, a great deal of online user action data has been generated. Such data has enabled the study of information diffusion prediction, which is a fundamental problem for understanding the propagation of information on social media platforms. In diffusion prediction models, there are two standard components, i.e., a social graph and information diffusion cascades. We observe that both components exhibit latent hierarchical structures. However, most existing models are designed based on euclidean spaces, and hence cannot effectively capture complex patterns, especially hierarchical structures. Therefore, we investigate a novel research problem to learn hyperbolic representations for information diffusion prediction. To reflect the different characteristics of social graphs and diffusion cascades, we encode them into two latent hyperbolic spaces with different trainable curvatures. In addition, to model influence dependencies, we propose a co-attention mechanism to capture the processes of diffusion cascades using positional embeddings. Given a set of activated seed users, we jointly exploit diffusion cascades and social links to predict which users will be influenced. We conduct extensive experiments on four real-world datasets. Empirical results demonstrate that the proposed H-Diffu model significantly outperforms several state-of-the-art diffusion prediction frameworks.
Shanshan Feng 0001, Kaiqi Zhao 0001, Lanting Fang, Kaiyu Feng, Wei Wei 0002, Xutao Li 0003, Ling Shao 0001
IEEE Trans. Knowl. Data Eng.6
2023 Multisource Heterogeneous Domain Adaptation With Conditional Weighting Adversarial Network
abstract
Heterogeneous domain adaptation (HDA) tackles the learning of cross-domain samples with both different probability distributions and feature representations. Most of the existing HDA studies focus on the single-source scenario. In reality, however, it is not uncommon to obtain samples from multiple heterogeneous domains. In this article, we study the multisource HDA problem and propose a conditional weighting adversarial network (CWAN) to address it. The proposed CWAN adversarially learns a feature transformer, a label classifier, and a domain discriminator. To quantify the importance of different source domains, CWAN introduces a sophisticated conditional weighting scheme to calculate the weights of the source domains according to the conditional distribution divergence between the source and target domains. Different from existing weighting schemes, the proposed conditional weighting scheme not only weights the source domains but also implicitly aligns the conditional distributions during the optimization process. Experimental results clearly demonstrate that the proposed CWAN performs much better than several state-of-the-art methods on four real-world datasets.
Yuan Yao 0016, Xutao Li 0003, Yu Zhang 0006, Yunming Ye
IEEE Trans. Neural Networks Learn. Syst.2
2022 MetaNODE: Prototype Optimization as a Neural ODE for Few-Shot Learning
abstract
Few-Shot Learning (FSL) is a challenging task, i.e., how to recognize novel classes with few examples? Pre-training based methods effectively tackle the problem by pre-training a feature extractor and then predicting novel classes via a cosine nearest neighbor classifier with mean-based prototypes. Nevertheless, due to the data scarcity, the mean-based prototypes are usually biased. In this paper, we attempt to diminish the prototype bias by regarding it as a prototype optimization problem. To this end, we propose a novel meta-learning based prototype optimization framework to rectify prototypes, i.e., introducing a meta-optimizer to optimize prototypes. Although the existing meta-optimizers can also be adapted to our framework, they all overlook a crucial gradient bias issue, i.e., the mean-based gradient estimation is also biased on sparse data. To address the issue, we regard the gradient and its flow as meta-knowledge and then propose a novel Neural Ordinary Differential Equation (ODE)-based meta-optimizer to polish prototypes, called MetaNODE. In this meta-optimizer, we first view the mean-based prototypes as initial prototypes, and then model the process of prototype optimization as continuous-time dynamics specified by a Neural ODE. A gradient flow inference network is carefully designed to learn to estimate the continuous gradient flow for prototype dynamics. Finally, the optimal prototypes can be obtained by solving the Neural ODE. Extensive experiments on miniImagenet, tieredImagenet, and CUB-200-2011 show the effectiveness of our method.
Baoquan Zhang, Xutao Li 0003, Shanshan Feng 0001, Yunming Ye, Rui Ye 0002
AAAI2
2022 Hyperbolic Knowledge Transfer with Class Hierarchy for Few-Shot Learning
abstract
Few-shot learning (FSL) aims to recognize a novel class with very few instances, which is a challenging task since it suffers from a data scarcity issue. One way to effectively alleviate this issue is introducing explicit knowledge summarized from human past experiences to achieve knowledge transfer for FSL. Based on this idea, in this paper, we introduce the explicit knowledge of class hierarchy (i.e., the hierarchy relations between classes) as FSL priors and propose a novel hyperbolic knowledge transfer framework for FSL, namely, HyperKT. Our insight is, in the hyperbolic space, the hierarchy relation between classes can be well preserved by resorting to the exponential growth characters of hyperbolic volume, so that better knowledge transfer can be achieved for FSL. Specifically, we first regard the class hierarchy as a tree-like structure. Then, 1) a hyperbolic representation learning module and a hyperbolic prototype inference module are employed to encode/infer each image and class prototype to the hyperbolic space, respectively; and 2) a novel hierarchical classification and relation reconstruction loss are carefully designed to learn the class hierarchy. Finally, the novel class prediction is performed in a nearest-prototype manner. Extensive experiments on three datasets show our method achieves superior performance over state-of-the-art methods, especially on 1-shot tasks.
Baoquan Zhang, Hao Jiang 0051, Shanshan Feng 0001, Xutao Li 0003, Yunming Ye, Rui Ye 0002
IJCAI4
2022 Visual Grounding in Remote Sensing Images
abstract
Ground object retrieval from a large-scale remote sensing image is very important for lots of applications. We present a novel problem of visual grounding in remote sensing images. Visual grounding aims to locate the particular objects (in the form of the bounding box or segmentation mask) in an image by a natural language expression. The task already exists in the computer vision community. However, existing benchmark datasets and methods mainly focus on natural images rather than remote sensing images. Compared with natural images, remote sensing images contain large-scale scenes and the geographical spatial information of ground objects (e.g., longitude, latitude). The existing method cannot deal with these challenges. In this paper, we collect a new visual grounding dataset, called RSVG, and design a new method, namely GeoVG. In particular, the proposed method consists of a language encoder, image encoder, and fusion module. The language encoder is used to learn numerical geospatial relations and represent a complex expression as a geospatial relation graph. The image encoder is applied to learn large-scale remote sensing scenes with adaptive region attention. The fusion module is used to fuse the text and image feature for visual grounding. We evaluate the proposed method by comparing it to the state-of-the-art methods on RSVG. Experiments show that our method outperforms the previous methods on the proposed datasets. https://sunyuxi.github.io/publication/GeoVG
Yuxi Sun 0002, Shanshan Feng 0001, Xutao Li 0003, Yunming Ye, Jian Kang 0005
ACM Multimedia3
2022 The reconstitution predictive network for precipitation nowcasting
Chuyao Luo, Guangning Xu, Xutao Li 0003, Yunming Ye
Neurocomputing3
2022 SPLNet: A sequence-to-one learning network with time-variant structure for regional wind speed prediction
Rui Ye 0002, Shanshan Feng 0001, Xutao Li 0003, Yunming Ye, Baoquan Zhang, Chuyao Luo
Inf. Sci.3
2022 Unsupervised deep hashing through learning soft pseudo label for remote sensing image retrieval
Yuxi Sun 0002, Yunming Ye, Xutao Li 0003, Shanshan Feng 0001, Bowen Zhang 0005, Jian Kang 0005, Kuai Dai
Knowl. Based Syst.3
2022 GCDB-UNet: A novel robust cloud detection approach for remote sensing images
Xian Li 0007, Xiaofei Yang 0002, Xutao Li 0003, Shijian Lu, Yunming Ye, Yifang Ban
Knowl. Based Syst.3
2022 PredRANN: The spatiotemporal attention Convolution Recurrent Neural Network for precipitation nowcasting
Chuyao Luo, Xinyue Zhao, Yuxi Sun 0002, Xutao Li 0003, Yunming Ye
Knowl. Based Syst.4
2022 Better Visual Interpretation for Remote Sensing Scene Classification
abstract
Deep learning-based methods have been widely applied in remote sensing scene classification tasks. Recently, researchers focus more on clarifying the basis of a decision. For example, class activation mapping (CAM) can provide us the evidence by highlighting the related area in an image. However, the interpretability of remote sensing scene classification is more challenging than natural images, since remote sensing images usually contain more complicated objects. As a result, the CAM visual interpretation with traditional convolutional neural networks cannot accurately locate all target objects, which leads to some important objects are ignored. In this letter, we propose a novel model, named encoder-classifier-reconstruction CAM (ECR-CAM) neural network, to provide a more precise visual explanation. Specifically, ECR-CAM consists of four modules: an encoder module, a classifier module, a reconstruction module, and a CAM module. Encoder module is utilized to extract image features, and classifier module accounts for generating predictions. The reconstruction module is the key to locate more target objects. It employs the extracted features to reconstruct the input images, which is a pixel-level process. The reconstruction process allows the features to retain important information about all objects, which cannot be achieved by the classification task alone. Finally, the CAM module can show more target objects with more informative features. Experimental results show that our model not only improves the classification performance but also can locate the target objects more accurately.
Yuxi Sun 0002, Shanshan Feng 0001, Yunming Ye, Xutao Li 0003
IEEE Geosci. Remote. Sens. Lett.5
2022 An Energy-Based Generative Adversarial Forecaster for Radar Echo Map Extrapolation
abstract
Precipitation nowcasting is an important task in weather forecast. The key challenge of the task lies at radar echo map extrapolation. Recent studies show that a convolutional recurrent neural network (ConvRNN) is a promising direction to solve the problem. However, the extrapolation results of the existing ConvRNN methods tend to be blurring and unrealistic. Recent studies show that generative adversarial network (GAN) is a promising tool to address the drawback, while it suffers from the instability for training. In this letter, we build a novel ConvRNN model based on the energy-based GAN for radar echo map extrapolation. The method can alleviate the blurring and unrealistic issues and is more stable. We have conducted experiments on a real-world data set, and the results show that the proposed method outperforms several existing models, including optical flow, convolution gated recurrent unit (ConvGRU), and generative adversarial ConvGRU (GA-ConvGRU).
Xutao Li 0003, Xiyang Ji, Xunlai Chen, Yuanzhao Chen, Yunming Ye
IEEE Geosci. Remote. Sens. Lett.2
2022 AM-ConvGRU: a spatio-temporal model for typhoon path prediction
Guangning Xu, Di Xian, Philippe Fournier-Viger, Xutao Li 0003, Yunming Ye, Xiuqing Hu
Neural Comput. Appl.4
2022 LS-NTP: Unifying long- and short-range spatial correlations for near-surface temperature prediction
Guangning Xu, Xutao Li 0003, Shanshan Feng 0001, Yunming Ye, Zhihua Tu, Kenghong Lin, Zhichao Huang 0001
Neural Networks2
2022 DynamicNet: A time-variant ODE network for multi-step wind speed prediction
Rui Ye 0002, Xutao Li 0003, Yunming Ye, Baoquan Zhang
Neural Networks2
2022 Multisensor Fusion and Explicit Semantic Preserving-Based Deep Hashing for Cross-Modal Remote Sensing Image Retrieval
abstract
Cross-modal hashing is an important tool for retrieving useful information from very-high-resolution (VHR) optical images and synthetic aperture radar (SAR) images. Dealing with the intermodal discrepancies, including both spatial–spectral and visual semantic aspects, between VHR and SAR images is extremely vital to generate high-quality common hash codes in the Hamming space. However, existing cross-modal hashing methods ignore the spatial–spectral discrepancy when representing VHR and SAR images. Moreover, existing methods employ derived supervised signals, such as pairwise training images, to implicitly guide hashing learning, which fails to effectively deal with the visual semantic discrepancy, i.e., cannot adequately preserve the intraclass similarity and interclass discrimination between VHR and SAR images. To address these drawbacks, this article proposes a multisensor fusion and explicit semantic preserving-based deep Hashing method, termed as MsEspH, which can effectively deal with the discrepancies. Specifically, we design a novel cross-modal hashing network to eliminate the spatial–spectral discrepancies by fusing extra multispectral images (MSIs), which are generated in real time by a generative adversarial network. Then, we propose an explicit semantic preserving-based objective function by analyzing the connection between classification and hash learning. The objective function can preserve the intraclass similarity and interclass discrimination with class labels directly. Moreover, we theoretically verify that hash learning and classification can be unified into a learning framework under certain conditions. To evaluate our method, we construct and release a large-scale VHR-SAR image dataset. Extensive experiments on the dataset demonstrate that our method outperforms various state-of-the-art cross-modal hashing methods.
Yuxi Sun 0002, Shanshan Feng 0001, Yunming Ye, Xutao Li 0003, Jian Kang 0005, Zhichao Huang 0001, Chuyao Luo
IEEE Trans. Geosci. Remote. Sens.4
2022 MSTCGAN: Multiscale Time Conditional Generative Adversarial Network for Long-Term Satellite Image Sequence Prediction
abstract
Satellite image sequence prediction is a crucial and challenging task. Previous studies leverage optical flow methods or existing deep learning methods on spatial-temporal sequence models for the task. However, they suffer from either oversimplified model assumptions or blurry predictions and sequential error accumulation issue, for a long-term forecast requirement. In this paper, we propose a novel Multi-Scale Time Conditional Generative Adversarial Network (MSTCGAN). To address the sequential error accumulation issue, MSTCGAN adopts a parallel prediction framework to produce the future image sequences by a one-hot time condition input. In addition, a powerful multi-scale generator is designed with the multi-head axial attention, which helps to carefully preserve the fine-grained details for appearance consistency. Moreover, we develop a temporal discriminator to address the blurry issue and maintain the motion consistency in prediction. Extensive experiments have been conducted on FengYun-4A satellite data set, and the results demonstrate the effectiveness and superiority of the proposed method over state-of-the-art approaches.
Kuai Dai, Xutao Li 0003, Yunming Ye, Shanshan Feng 0001, Danyu Qin, Rui Ye 0002
IEEE Trans. Geosci. Remote. Sens.2
2022 A Convolutional Neural Network-Based Relative Radiometric Calibration Method
abstract
Due to the degeneration problem of sensors, calibration becomes a prerequisite step to retrieve consistent satellite images, especially for the ones from long-term time series. Relative calibration is an economic manner to address the problem. Previous studies leverage the identified no-change pixels (NCPs) between two images for relative calibration. However, the identification of NCPs itself is a very hard task and the inferior detection quality affects the performances significantly. Inspired by the great success of deep learning techniques, in this article, we first develop a convolutional neural network (CNN)-based relative calibration method, which bypasses the NCP detection. In particular, the ratio of sensor sensitivity coefficients at two time points is directly estimated by feeding the corresponding image pair into our developed CNN regressor. A polynomial function is fitted upon the estimated ratios in time series. We train the CNN regressor based on the multisite calibration results and then conduct experiments on FengYun-3A (FY-3A), FengYun-3B (FY-3B), and FengYun-3C (FY-3C). The results validate the effectiveness of the proposed method, and it outperforms state-of-the-art NCP-based methods.
Xutao Li 0003, Zhizi Ye, Yunming Ye, Xiuqing Hu
IEEE Trans. Geosci. Remote. Sens.1
2022 LWCDnet: A Lightweight Network for Efficient Cloud Detection in Remote Sensing Images
abstract
Cloud detection is the task of detecting cloud areas in remote sensing images, and it has attracted extensive research interest. Recently, deep learning-based methods have been proposed and achieved great performance for cloud detection. However, due to the satellite’s limitation in storage and memory, existing deep learning approaches, which suffer from extensive computation and large model size, are almost impossible to be deployed on satellites. To fill this gap, we target at studying effective and efficient cloud detection solutions that are suitable for satellites. In this paper, we develop a lightweight autoencoder-based cloud detection method, namely LWCDnet. In the encoder part, the designed novel lightweight dual-branch block (LWDBB) in the backbone extracts spatial and contextual information concurrently. Moreover, a lightweight feature pyramid module (LWFPM) is proposed to capture high-level multi-scale contextual information. In the decoder part, the lightweight feature fusion module (LWFFM) compensates for the missing spatial and detail information from the encoder to the high-level feature maps. We evaluate the proposed method on two public datasets: LandSat8 and MODIS. Extensive experiments demonstrate that the proposed LWCDnet achieves comparable accuracy as the-state-of-art cloud detection methods and lightweight semantic segmentation algorithms. Meantime LWCDnet has much less computation burden with smaller model size.
Shanshan Feng 0001, Xiaofei Yang 0002, Yunming Ye, Xutao Li 0003, Baoquan Zhang, Zhihao Chen 0010, Yingling Quan
IEEE Trans. Geosci. Remote. Sens.5
2022 Experimental Study on Generative Adversarial Network for Precipitation Nowcasting
abstract
Precipitation nowcasting is an important task, which can be used in numerous applications. The key challenge of the task lies in radar echo map prediction. Previous studies leverage Convolutional Recurrent Neural Network (ConvRNN) to address the problem. However, the approaches are built upon mean square losses and the results tend to have inaccurate appearances, shapes and positions for predictions. To alleviate this problem, we explore the idea of adversarial regularization, and systematically compare four types of Generative Adversarial Networks (GANs), which are the combinations of GAN/Wasserstein GAN and its multi-scale version. Extensive experiments on a real-world radar data set and four typical meteorological examples are conducted. The results validate the effectiveness of adversarial regularization. The developed models show superior performances over the existing prediction approaches in the majority circumstances. Moreover, we find that the Wasserstein GAN regularization often delivers better results than the GAN regularization due to its robustness, and the Multi-scale Wasserstein GAN, in general, performs the best among all the methods. To reproduce the results, we release the source code at: https://github.com/luochuyao/MultiScaleGAN and the test system at: http://39.97.217.145:80/.
Chuyao Luo, Xutao Li 0003, Yunming Ye, Shanshan Feng 0001, Michael Kwok-Po Ng
IEEE Trans. Geosci. Remote. Sens.2
2022 Multisource Data Reconstruction-Based Deep Unsupervised Hashing for Unisource Remote Sensing Image Retrieval
abstract
Unsupervised hashing for remote sensing (RS) image retrieval first extracts image features and then use these features to construct supervised information (e.g., pseudo-labels) to train hashing networks. Existing methods usually regard RS images as natural images to extract unisource features. However, these features only contain partial information about ground objects and cannot produce reliable pseudo-labels. In addition, existing methods only generate a pseudo single-label to annotate each RS image, which cannot accurately represent multiple scenes in a RS image. To address these drawbacks, this paper proposes a new Multisource data reconstruction-based deep unsupervised Hashing method, called MrHash, which explores the characteristics of RS images to construct reliable pseudo-labels. In particular, we first use geographic coordinates to obtain different satellite images and develop a novel autoencoder network to extract multisource features from these images. Then pseudo multi-labels are designed to deal with the coexistence of multiple scenes in a single image. These labels are generated by a custom probability function with extracted multisource features. Finally, we propose a novel multi-semantic hash loss by using the Kull-back–Leibler (KL) divergence to preserve the semantic similarity of these pseudo multi-labels in Hamming space. Our newly developed MrHash only uses multisource images to construct supervised information, and hash code generation still relies on a unisource input image. Experiments on benchmark datasets clearly show the superiority of the proposed method over state-of-the-art baselines. https://github.com/sunyuxi/MrHash.
Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Yifang Ban, Xutao Li 0003, Bowen Zhang 0005, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.6
2022 SGMNet: Scene Graph Matching Network for Few-Shot Remote Sensing Scene Classification
abstract
Few-Shot Remote Sensing Scene Classification (FSRSSC) is an important task, which aims to recognize novel scene classes with few examples. Recently, several studies attempt to address the FSRSSC problem by following few-shot natural image classification methods. These existing methods have made promising progress and achieved superior performance. However, they all overlook two unique characteristics of remote sensing images: (i)object co-occurrencethat multiple objects tend to appear together in a scene image and (ii)object spatial correlationthat these co-occurrence objects are distributed in the scene image following some spatial structure patterns. Such unique characteristics are very beneficial for FSRSSC, which can effectively alleviate the scarcity issue of labeled remote sensing images since they can provide more refined descriptions for each scene class. To fully exploit these characteristics, we propose a novel scene graph matching-based meta-learning framework for FSRSSC, called SGMNet. In this framework, a scene graph construction module is carefully designed to represent each test remote sensing image or each scene class as a scene graph, where the nodes reflect these co-occurrence objects meanwhile the edges capture the spatial correlations between these co-occurrence objects. Then, a scene graph matching module is further developed to evaluate the similarity score between each test remote sensing image and each scene class. Finally, based on the similarity scores, we perform the scene class prediction via a nearest neighbor classifier. We conduct extensive experiments on UCMerced LandUse, WHU19, AID, and NWPU-RESISC45 datasets. The experimental results show that our method obtains superior performance over the previous state-of-the-art methods.
Baoquan Zhang, Shanshan Feng 0001, Xutao Li 0003, Yunming Ye, Rui Ye 0002, Hao Jiang 0051
IEEE Trans. Geosci. Remote. Sens.3
2021 Prototype Completion With Primitive Knowledge for Few-Shot Learning
abstract
Few-shot learning is a challenging task, which aims to learn a classifier for novel classes with few examples. Pre-training based meta-learning methods effectively tackle the problem by pre-training a feature extractor and then fine-tuning it through the nearest centroid based meta-learning. However, results show that the fine-tuning step makes very marginal improvements. In this paper, 1) we figure out the key reason, i.e., in the pre-trained feature space, the base classes already form compact clusters while novel classes spread as groups with large variances, which implies that fine-tuning the feature extractor is less meaningful; 2) instead of fine-tuning the feature extractor, we focus on estimating more representative prototypes during meta-learning. Consequently, we propose a novel prototype completion based meta-learning framework. This framework first introduces primitive knowledge (i.e., class-level part or attribute annotations) and extracts representative attribute features as priors. Then, we design a prototype completion network to learn to complete prototypes with these priors. To avoid the prototype completion error caused by primitive knowledge noises or class differences, we further develop a Gaussian based prototype fusion strategy that combines the mean-based and completed prototypes by exploiting the unlabeled samples. Extensive experiments show that our method: (i) can obtain more accurate prototypes; (ii) out-performs state-of-the-art techniques by 2%~9% in terms of classification accuracy. Our code is available online1.
Baoquan Zhang, Xutao Li 0003, Yunming Ye, Zhichao Huang 0001, Lisai Zhang
CVPR2
2021 MLCE: A Multi-Label Crotch Ensemble Method for Multi-Label Classification
abstract
Multi-label classification addresses the problem that each instance is associated with multiple labels simultaneously. In this paper, we propose a multi-label crotch ensemble (MLCE) model for multi-label classification, which takes label correlations into consideration. In MLCE, a multi-label cluster tree is first constructed. Then, we incorporate all multi-label crotch predictors of the tree into a classifier, where the multi-label crotch predictor is the crotch formed by an inner node of the tree and its children. Finally, a flexible weighted voting scheme is designed to produce the classification output. We perform experiments on 11 benchmark datasets. Experimental results clearly demonstrate the MLCE significantly outperforms six well-established multi-label classification approaches, in terms of the widely used evaluation metrics.
Yuan Yao 0016, Yan Li 0040, Yunming Ye, Xutao Li 0003
Int. J. Pattern Recognit. Artif. Intell.4
2021 An Investigation on Deep Learning Approaches to Combining Nighttime and Daytime Satellite Imagery for Poverty Prediction
abstract
Poverty prediction is an important task for developing countries that lack the key measures of economic development. The prediction can help governments to allocate scarce resources for sustainable development. Nighttime satellite imagery offers an opportunity to address the task. However, as the nighttime satellite data contain a large amount of noise, directly leveraging it is not very effective. Previous studies have shown that relying on deep learning techniques nighttime satellite data can be a good proxy between daytime satellite imagery and the poverty index. In this letter, based on the proxy, we leverage four deep learning approaches, namely, VGG-Net, Inception-Net, ResNet, and DenseNet, to extract deep features from daytime satellite imagery and then apply least absolute shrinkage and selection operator (LASSO) regression for poverty prediction. To further enhance the performance, we also integrate the squeeze and excitation (SE) module and focal loss into ResNet and DenseNet. Experimental results demonstrate the effectiveness of the investigated approaches, and the DenseNet with SE module and focal loss performs the best.
Ye Ni, Xutao Li 0003, Yunming Ye, Yan Li 0040, Chunshan Li
IEEE Geosci. Remote. Sens. Lett.2
2020 Enhancing Cross-target Stance Detection with Transferable Semantic-Emotion Knowledge
abstract
Stance detection is an important task, which aims to classify the attitude of an opinionated text towards a given target. Remarkable success has been achieved when sufficient labeled training data is available. However, annotating sufficient data is labor-intensive, which establishes significant barriers for generalizing the stance classifier to the data with new targets. In this paper, we proposed a Semantic-Emotion Knowledge Transferring (SEKT) model for cross-target stance detection, which uses the external knowledge (semantic and emotion lexicons) as a bridge to enable knowledge transfer across different targets. Specifically, a semantic-emotion heterogeneous graph is constructed from external semantic and emotion lexicons, which is then fed into a graph convolutional network to learn multi-hop semantic connections between words and emotion tags. Then, the learned semantic-emotion graph representation, which serves as prior knowledge bridging the gap between the source and target domains, is fully integrated into the bidirectional long short-term memory (BiLSTM) stance classifier by adding a novel knowledge-aware memory unit to the BiLSTM cell. Extensive experiments on a large real-world dataset demonstrate the superiority of SEKT against the state-of-the-art baseline methods.
Bowen Zhang 0005, Min Yang 0007, Xutao Li 0003, Yunming Ye, Xiaofei Xu 0001, Kuai Dai
ACL3
2020 Cascade SEIRD: Forecasting the Spread of COVID-19 with Dynamic Parameters Update
abstract
The SEIR model is widely used in simulating the spread of infectious diseases. COVID-19 virus is a very severe infectious disease. Some studies leverage the SEIR or SEIRD model to simulate the spread and estimate the number of infected and recovered people as time goes on. However, these models suffer from two key deficiencies: (i) conventional SEIRD does not update its model parameters w.r.t. time; (ii) it focuses on predicting the trend, instead of the actual number of infections in the future. In this paper, we propose a cascade SEIRD model. The model learns and updates its parameters every day. Moreover, it is able to predict the number of infection cases, recovered cases and deaths. Specifically, we leverage a machine learning like approach to dynamically estimate the parameters of infection rate, incubation rate, recovery rate and death rate, which can be updated by gradient descent algorithm. Once the nature of the parameters w.r.t. time is determined, ARIMA model is adopted to characterize the dynamics of the parameters and predict their future changes. To validate the effectiveness of the proposed cascade SEIRD model, we conduct experiments on five data sets of different scales of regions (China, Hubei, Wuhan, Shenzhen, US). Experimental results show that the proposed cascade SEIRD achieves the most accurate prediction and outperforms state-of-the-art techniques.
Yongliang Wen, Jiangnan Xu, Yunming Ye, Xutao Li 0003, Chuyao Luo, Tianlun Zhu
BIBM4
2020 MR-GCN: Multi-Relational Graph Convolutional Networks based on Generalized Tensor Product
abstract
Graph Convolutional Networks (GCNs) have been extensively studied in recent years. Most of existing GCN approaches are designed for the homogenous graphs with a single type of relation. However, heterogeneous graphs of multiple types of relations are also ubiquitous and there is a lack of methodologies to tackle such graphs. Some previous studies address the issue by performing conventional GCN on each single relation and then blending their results. However, as the convolutional kernels neglect the correlations across relations, the strategy is sub-optimal. In this paper, we propose the Multi-Relational Graph Convolutional Network (MR-GCN) framework by developing a novel convolution operator on multi-relational graphs. In particular, our multi-dimension convolution operator extends the graph spectral analysis into the eigen-decomposition of a Laplacian tensor. And the eigen-decomposition is formulated with a generalized tensor product, which can correspond to any unitary transform instead of limited merely to Fourier transform. We conduct comprehensive experiments on four real-world multi-relational graphs to solve the semi-supervised node classification task, and the results show the superiority of MR-GCN against the state-of-the-art competitors.
Zhichao Huang 0001, Xutao Li 0003, Yunming Ye, Michael Kwok-Po Ng
IJCAI2
2020 A Noise Adaptive Model for Distantly Supervised Relation Extraction
Bowen Zhang 0005, Yunming Ye, Xiaojun Chen 0006, Xutao Li 0003
NLPCC (1)5
2020 End-to-End Deep Reinforcement Learning based Recommendation with Supervised Embedding
abstract
The research of reinforcement learning (RL) based recommendation method has become a hot topic in recommendation community, due to the recent advance in interactive recommender systems. The existing RL recommendation approaches can be summarized into a unified framework with three components, namely embedding component (EC), state representation component (SRC) and policy component (PC). We find that EC cannot be nicely trained with the other two components simultaneously. Previous studies bypass the obstacle through a pre-training and fixing strategy, which makes their approaches unlike a real end-to-end fashion. More importantly, such pre-trained and fixed EC suffers from two inherent drawbacks: (1) Pre-trained and fixed embeddings are unable to model evolving preference of users and item correlations in the dynamic environment; (2) Pre-training is inconvenient in the industrial applications. To address the problem, in this paper, we propose an End-to-end Deep Reinforcement learning based Recommendation framework (EDRR). In this framework, a supervised learning signal is carefully designed for smoothing the update gradients to EC, and three incorporating ways are introduced and compared. To the best of our knowledge, we are the first to address the training compatibility between the three components in RL based recommendations. Extensive experiments are conducted on three real-world datasets, and the results demonstrate the proposed EDRR effectively achieves the end-to-end training purpose for both policy-based and value-based RL models, and delivers better performance than state-of-the-art methods.
Feng Liu 0034, Huifeng Guo, Xutao Li 0003, Ruiming Tang, Yunming Ye, Xiuqiang He 0001
WSDM3
2020 Top-aware reinforcement learning based recommendation
Feng Liu 0034, Ruiming Tang, Huifeng Guo, Xutao Li 0003, Yunming Ye, Xiuqiang He 0001
Neurocomputing4
2020 State representation modeling for deep reinforcement learning based recommendation
Feng Liu 0034, Ruiming Tang, Xutao Li 0003, Weinan Zhang 0001, Yunming Ye, Huifeng Guo, Xiuqiang He 0001
Knowl. Based Syst.3
2020 A memory network based end-to-end personalized task-oriented dialogue generation
Bowen Zhang 0005, Xiaofei Xu 0001, Xutao Li 0003, Yunming Ye, Xiaojun Chen 0006, Zhongjie Wang 0003
Knowl. Based Syst.3
2020 A Generative Adversarial Gated Recurrent Unit Model for Precipitation Nowcasting
abstract
Precipitation nowcasting is an important task in operational weather forecasts. The key challenge of the task is the radar echo map extrapolation. The problem is mainly solved by an optical-flow method in existing systems. However, the method cannot model rapid and nonlinear movements. Recently, a convolutional gated recurrent unit (ConvGRU) method is developed, which aims to model such movements based on deep learning techniques. Despite the promising performance, ConvGRU tends to yield blurring extrapolation images and fails to multi-modal and skewed intensity distribution. To overcome the limitations, we propose in this letter a generative adversarial ConvGRU (GA-ConvGRU) model. The model is composed of two adversarial learning systems, which are a ConvGRU-based generator and a convolution neural network-based discriminator. The two systems are trained by playing a minimax game. With the adversarial learning scheme, GA-ConvGRU can yield more realistic and more accurate extrapolation. Experiments on real data sets have been conducted and the results demonstrate that the proposed GA-ConvGRU significantly outperforms state-of-the-art extrapolation methods ConvGRU and optical flow.
Xutao Li 0003, Yunming Ye, Yan Li 0040
IEEE Geosci. Remote. Sens. Lett.2
2020 TLVANE: a two-level variation model for attributed network embedding
Zhichao Huang 0001, Xutao Li 0003, Yunming Ye, Feng Li 0022, Feng Liu 0034, Yuan Yao 0016
Neural Comput. Appl.2
2020 Discriminative distribution alignment: A unified framework for heterogeneous domain adaptation
Yuan Yao 0016, Yu Zhang 0006, Xutao Li 0003, Yunming Ye
Pattern Recognit.3
2020 Knowledge Guided Capsule Attention Network for Aspect-Based Sentiment Analysis
abstract
Aspect-based (aspect-level) sentiment analysis is an important task in fine-grained sentiment analysis, which aims to automatically infer the sentiment towards an aspect in its context. Previous studies have shown that utilizing the attention-based method can effectively improve the accuracy of the aspect-based sentiment analysis. Despite the outstanding progress, aspect-based sentiment analysis in the real-world remains several challenges. (1) The current attention-based method may cause a given aspect to incorrectly focus on syntactically unrelated words. (2) Conventional methods fail to identify the sentiment with the special sentence structure, such as double negatives. (3) Most of the studies leverage only one vector to represent context and target. However, utilizing one vector to represent the sentence is limited, as the natural languages are delicate and complex. In this paper, we propose a knowledge guided capsule network (KGCapsAN), which can address the above deficiencies. Our method is composed of two parts, a Bi-LSTM network and a capsule attention network. The capsule attention network implements the routing method by attention mechanism. Moreover, we utilize two prior knowledge to guide the capsule attention process, which are syntactical and n-gram structures. Extensive experiments are conducted on six datasets, and the results show that the proposed method yields the state-of-the-art.
Bowen Zhang 0005, Xutao Li 0003, Xiaofei Xu 0001, Ka-Cheong Leung, Zhiyao Chen, Yunming Ye
IEEE ACM Trans. Audio Speech Lang. Process.2
2019 Road Detection via Deep Residual Dense U-Net
abstract
Road extraction from aerial images is a hot research topic. With the advancement of convolutional neural network (CNN), several CNN-based road detection methods have been developed. However, most of them do not make full use of the hierarchical features from the original aerial images. In this paper, we propose a novel residual dense U-Net (RDUN), a semantic segmentation network which combines the strengths of residual learning, DenseNet, and U-Net, to overcome the drawback. Our proposed RDUN can fully exploit the hierarchical features from all the convolutional layers, which utilizes the residual dense blocks (RDB) to build up a U-Net architecture. The benefits of our model are two-fold. First, by using the RDB abundant local features can be extracted and fused effectively. Second, based the local features, hierarchical features are constructed by shortcut connections between layers in RDB. Extensive experiments are carried out on a real-world road detection dataset and the results demonstrate the proposed RDUN outperforms state-of-the-art competitors.
Xiaofei Yang 0002, Xutao Li 0003, Yunming Ye, Xiaofeng Zhang 0002, Haijun Zhang 0002, Xiaohui Huang 0003, Bowen Zhang 0005
IJCNN2
2019 Heterogeneous Domain Adaptation via Soft Transfer Network
abstract
Heterogeneous domain adaptation (HDA) aims to facilitate the learning task in a target domain by borrowing knowledge from a heterogeneous source domain. In this paper, we propose a Soft Transfer Network (STN), which jointly learns a domain-shared classifier and a domain-invariant subspace in an end-to-end manner, for addressing the HDA problem. The proposed STN not only aligns the discriminative directions of domains but also matches both the marginal and conditional distributions across domains. To circumvent negative transfer, STN aligns the conditional distributions by using the soft-label strategy of unlabeled target data, which prevents the hard assignment of each unlabeled target data to only one category that may be incorrect. Further, STN introduces an adaptive coefficient to gradually increase the importance of the soft-labels since they will become more and more accurate as the number of iterations increases. We perform experiments on the transfer tasks of image-to-image, text-to-image, and text-to-text. Experimental results testify that the STN significantly outperforms several state-of-the-art approaches.
Yuan Yao 0016, Yu Zhang 0006, Xutao Li 0003, Yunming Ye
ACM Multimedia3
2019 Learning Personalized End-to-End Task-Oriented Dialogue Generation
Bowen Zhang 0005, Xiaofei Xu 0001, Xutao Li 0003, Yunming Ye, Xiaojun Chen 0006, Lianjie Sun
NLPCC (1)3
2019 Learning Stance Classification with Recurrent Neural Capsule Network
Lianjie Sun, Xutao Li 0003, Bowen Zhang 0005, Yunming Ye, Baoxun Xu
NLPCC (1)2
2019 Sentiment analysis through critic learning for optimizing convolutional neural networks with rules
Bowen Zhang 0005, Xiaofei Xu 0001, Xutao Li 0003, Xiaojun Chen 0006, Yunming Ye, Zhongjie Wang 0003
Neurocomputing3
2019 Low-resolution image categorization via heterogeneous domain adaptation
Yuan Yao 0016, Xutao Li 0003, Yunming Ye, Feng Liu 0034, Michael Kwok-Po Ng, Zhichao Huang 0001, Yu Zhang 0006
Knowl. Based Syst.2
2019 Road Detection and Centerline Extraction Via Deep Recurrent Convolutional Neural Network U-Net
abstract
Road information extraction based on aerial images is a critical task for many applications, and it has attracted considerable attention from researchers in the field of remote sensing. The problem is mainly composed of two subtasks, namely, road detection and centerline extraction. Most of the previous studies rely on multistage-based learning methods to solve the problem. However, these approaches may suffer from the well-known problem of propagation errors. In this paper, we propose a novel deep learning model, recurrent convolution neural network U-Net (RCNN-UNet), to tackle the aforementioned problem. Our proposed RCNN-UNet has three distinct advantages. First, the end-to-end deep learning scheme eliminates the propagation errors. Second, a carefully designed RCNN unit is leveraged to build our deep learning architecture, which can better exploit the spatial context and the rich low-level visual features. Thereby, it alleviates the detection problems caused by noises, occlusions, and complex backgrounds of roads. Third, as the tasks of road detection and centerline extraction are strongly correlated, a multitask learning scheme is designed so that two predictors can be simultaneously trained to improve both effectiveness and efficiency. Extensive experiments were carried out based on two publicly available benchmark data sets, and nine state-of-the-art baselines were used in a comparative evaluation. Our experimental results demonstrate the superiority of the proposed RCNN-UNet model for both the road detection and the centerline extraction tasks.
Xiaofei Yang 0002, Xutao Li 0003, Yunming Ye, Raymond Y. K. Lau, Xiaofeng Zhang 0002, Xiaohui Huang 0003
IEEE Trans. Geosci. Remote. Sens.2
2018 Novel Approaches to Accelerating the Convergence Rate of Markov Decision Process for Search Result Diversification
Feng Liu 0034, Ruiming Tang, Xutao Li 0003, Yunming Ye, Huifeng Guo, Xiuqiang He 0001
DASFAA (2)3
2018 Block principal component analysis for tensor objects with frequency or time information
Xutao Li 0003, Michael Kwok-Po Ng, Xiaofei Xu 0001, Yunming Ye
Neurocomputing1
2018 Multi-attribute and relational learning via hypergraph regularized generative model
Shaokai Wang, Xutao Li 0003, Yunming Ye, Xiaohui Huang 0003, Yan Li 0040
Neurocomputing2
2018 Hyperspectral Image Classification With Deep Learning Models
abstract
Deep learning has achieved great successes in conventional computer vision tasks. In this paper, we exploit deep learning techniques to address the hyperspectral image classification problem. In contrast to conventional computer vision tasks that only examine the spatial context, our proposed method can exploit both spatial context and spectral correlation to enhance hyperspectral image classification. In particular, we advocate four new deep learning models, namely, 2-D convolutional neural network (2-D-CNN), 3-D-CNN, recurrent 2-D CNN (R-2-D-CNN), and recurrent 3-D-CNN (R-3-D-CNN) for hyperspectral image classification. We conducted rigorous experiments based on six publicly available data sets. Through a comparative evaluation with other state-of-the-art methods, our experimental results confirm the superiority of the proposed deep learning models, especially the R-3-D-CNN and the R-2-D-CNN deep learning models.
Xiaofei Yang 0002, Yunming Ye, Xutao Li 0003, Raymond Y. K. Lau, Xiaofeng Zhang 0002, Xiaohui Huang 0003
IEEE Trans. Geosci. Remote. Sens.3
2017 A generative model with hypergraph regularizers for protein function prediction
abstract
Heterogeneous data sources and multi-label are two important characteristics of protein function prediction. They describe protein data from two different aspects. However, it is of considerable challenge to integrate multiple data sources and multi-label simultaneously for predicting protein functions, especially when there are only a limited number of labeled proteins. In this paper, we propose a generative model with hypergraph regularizers algorithm, called GMHR, for predicting proteins with multiple functions. The GMHR algorithm integrates all data sources that are available, including protein attribute features, interaction networks, label correlations, and unlabeled data. Experimental results on the real-world datasets predicting the functions of proteins demonstrate the superiority of our proposed method compared with the state-of-the-art baselines.
Shaokai Wang, Xutao Li 0003, Yunming Ye, Yan Li 0040, Xiaohui Huang 0003, Xiaolin Du
IJCNN2
2017 Joint Weighted Nonnegative Matrix Factorization for Mining Attributed Graphs
Zhichao Huang 0001, Yunming Ye, Xutao Li 0003, Feng Liu 0034, Huajie Chen
PAKDD (1)3
2017 A General Model for Out-of-town Region Recommendation
abstract
With the rapid growth of location-based social networks (LBSNs), it is now available to analyze and understand user mobility behavior in real world. Studies show that users usually visit nearby points of interest (POIs), located in small regions, especially when they travel out of their hometowns. However, previous out-of-town recommendation systems mainly focus on recommending individual POIs that may reside far from each other, which makes the recommendation results less useful. In this paper, we introduce a novel problem called Region Recommendation, which aims to recommend an out-of-town region of POIs that are likely to be visited by a user. The proximity characteristic of user mobility behavior implies that the probability of visiting one POI depends on those of nearby POIs. Thus, to make accurate region recommendation, our proposed model exploits the influence between POIs, instead of treating them individually. Moreover, to overcome the efficiency problem of searching the best region, we propose a sweeping line-based method, and subsequently an constant-bounded algorithm for better efficiency. Experiments on two real-world datasets demonstrate the improved effectiveness of our models over baseline methods and efficiency of the approximate algorithm.
Tuan-Anh Nguyen Pham, Xutao Li 0003, Gao Cong
WWW2
2017 Block linear discriminant analysis for visual tensor objects with frequency or time information
Xutao Li 0003, Michael Kwok-Po Ng, Yunming Ye, Ke Wang 0068, Xiaofei Xu 0001
J. Vis. Commun. Image Represent.1
2017 Multi-view learning via multiple graph regularized generative model
Shaokai Wang, Ke Wang 0068, Xutao Li 0003, Yunming Ye, Raymond Y. K. Lau, Xiaolin Du
Knowl. Based Syst.3
2017 Semi-supervised Collective Classification in Multi-attribute Network Data
Shaokai Wang, Yunming Ye, Xutao Li 0003, Xiaohui Huang 0003, Raymond Y. K. Lau
Neural Process. Lett.3
2017 MR-NTD: Manifold Regularization Nonnegative Tucker Decomposition for Tensor Data Dimension Reduction and Representation
abstract
With the advancement of data acquisition techniques, tensor (multidimensional data) objects are increasingly accumulated and generated, for example, multichannel electroencephalographies, multiview images, and videos. In these applications, the tensor objects are usually nonnegative, since the physical signals are recorded. As the dimensionality of tensor objects is often very high, a dimension reduction technique becomes an important research topic of tensor data. From the perspective of geometry, high-dimensional objects often reside in a low-dimensional submanifold of the ambient space. In this paper, we propose a new approach to perform the dimension reduction for nonnegative tensor objects. Our idea is to use nonnegative Tucker decomposition (NTD) to obtain a set of core tensors of smaller sizes by finding a common set of projection matrices for tensor objects. To preserve geometric information in tensor data, we employ a manifold regularization term for the core tensors constructed in the Tucker decomposition. An algorithm called manifold regularization NTD (MR-NTD) is developed to solve the common projection matrices and core tensors in an alternating least squares manner. The convergence of the proposed algorithm is shown, and the computational complexity of the proposed method scales linearly with respect to the number of tensor objects and the size of the tensor objects, respectively. These theoretical results show that the proposed algorithm can be efficient. Extensive experimental results have been provided to further demonstrate the effectiveness and efficiency of the proposed MR-NTD algorithm.
Xutao Li 0003, Michael Kwok-Po Ng, Gao Cong, Yunming Ye, Qingyao Wu
IEEE Trans. Neural Networks Learn. Syst.1
2016 MultiVCRank With Applications to Image Retrieval
abstract
In this paper, we propose and develop a multi-visual-concept ranking (MultiVCRank) scheme for image retrieval. The key idea is that an image can be represented by several visual concepts, and a hypergraph is built based on visual concepts as hyperedges, where each edge contains images as vertices to share a specific visual concept. In the constructed hypergraph, the weight between two vertices in a hyperedge is incorporated, and it can be measured by their affinity in the corresponding visual concept. A ranking scheme is designed to compute the association scores of images and the relevance scores of visual concepts by employing input query vectors to handle image retrieval. In the scheme, the association and relevance scores are determined by an iterative method to solve limiting probabilities of a multi-dimensional Markov chain arising from the constructed hypergraph. The convergence analysis of the iteration method is studied and analyzed. Moreover, a learning algorithm is also proposed to set the parameters in the scheme, which makes it simple to use. Experimental results on the MSRC, Corel, and Caltech256 data sets have demonstrated the effectiveness of the proposed method. In the comparison, we find that the retrieval performance of MultiVCRank is substantially better than those of HypergraphRank, ManifoldRank, TOPHITS, and RankSVM.
Xutao Li 0003, Yunming Ye, Michael Kwok-Po Ng
IEEE Trans. Image Process.1
2016 A General Recommendation Model for Heterogeneous Networks
abstract
Heterogeneous networks refer to the networks comprising multiple types of entities as well as their interaction relationships. They arise in a great variety of domains, for example, event-based social networks Meetup and Plancast, and DBLP. Recommendation is a useful task in these heterogeneous network systems. Although many recommendation algorithms are proposed for heterogeneous data, none of them is able to explicitly model the influence strength between different types of entities, which is useful not only for achieving higher recommendation accuracy but also better understanding the role of each entity type in recommendation problems. Moreover, many of those algorithms are designed for a particular task, and hence it is challenging to apply them in other problems. In this paper, we propose a graph-based model, called HeteRS, which can solve general recommendation problems on heterogeneous networks. Our method models the rich information with a heterogeneous graph and considers the recommendation problem as a query-dependent node proximity problem. To address the challenging issue of weighting the influences between different types of entities, we propose a learning scheme to set the influence weights between different types of entities in recommendation. Experimental results on real-world datasets demonstrate that our proposed method significantly outperforms the baseline methods in our experiments for all the recommendation tasks, and the learned influence weights help understanding user behaviors.
Tuan-Anh Nguyen Pham, Xutao Li 0003, Gao Cong
IEEE Trans. Knowl. Data Eng.2
2015 Where you Instagram?: Associating Your Instagram Photos with Points of Interest
abstract
Instagram, an online photo-sharing platform, has gained increasing popularity. It allows users to take photos, apply digital filters and share them with friends instantaneously by using mobile devices.Instagram provides users with the functionality to associate their photos with points of interest, and it thus becomes feasible to study the association between points of interest and Instagram photos. However, no previous work studies the association. In this paper, we propose to study the problem of mapping Instagram photos to points of interest. To understand the problem, we analyze Instagram datasets, and report our findings, which also characterize the challenges of the problem. To address the challenges, we propose to model the mapping problem as a ranking problem, and develop a method to learn a ranking function by exploiting the textual, visual and user information of photos. To maximize the prediction effectiveness for textual and visual information, and incorporate the users' visiting preferences, we propose three subobjectives for learning the parameters of the proposed ranking function. Experimental results on two sets of Instagram data show that the proposed method substantially outperforms existing methods that are adapted to handle the problem.
Xutao Li 0003, Tuan-Anh Nguyen Pham, Gao Cong, Quan Yuan 0001, Xiaoli Li 0001, Shonali Krishnaswamy
CIKM1
2015 A general graph-based model for recommendation in event-based social networks
abstract
Event-based social networks (EBSNs), such as Meetup and Plancast, which offer platforms for users to plan, arrange, and publish events, have gained increasing popularity and rapid growth. EBSNs capture not only the online social relationship, but also the offline interactions from offline events. They contain rich heterogeneous information, including multiple types of entities, such as users, events, groups and tags, and their interaction relations. Three recommendation tasks, namely recommending groups to users, recommending tags to groups, and recommending events to users, have been explored in three separate studies. However, none of the proposed methods can handle all the three recommendation tasks. In this paper, we propose a general graph-based model, called HeteRS, to solve the three recommendation problems on EBSNs in one framework. Our method models the rich information with a heterogeneous graph and considers the recommendation problem as a query-dependent node proximity problem. To address the challenging issue of weighting the influences between different types of entities, we propose a learning scheme to set the influence weights between different types of entities. Experimental results on two real-world datasets demonstrate that our proposed method significantly outperforms the state-of-the-art methods for all the three recommendation tasks, and the learned influence weights help understanding user behaviors.
Tuan-Anh Nguyen Pham, Xutao Li 0003, Gao Cong
ICDE2
2015 Personalized Ranking Metric Embedding for Next New POI Recommendation
Shanshan Feng 0001, Xutao Li 0003, Yifeng Zeng, Gao Cong, Yeow Meng Chee, Quan Yuan 0001
IJCAI2
2015 Rank-GeoFM: A Ranking based Geographical Factorization Method for Point of Interest Recommendation
abstract
With the rapid growth of location-based social networks, Point of Interest (POI) recommendation has become an important research problem. However, the scarcity of the check-in data, a type of implicit feedback data, poses a severe challenge for existing POI recommendation methods. Moreover, different types of context information about POIs are available and how to leverage them becomes another challenge. In this paper, we propose a ranking based geographical factorization method, called Rank-GeoFM, for POI recommendation, which addresses the two challenges. In the proposed model, we consider that the check-in frequency characterizes users' visiting preference and learn the factorization by ranking the POIs correctly. In our model, POIs both with and without check-ins will contribute to learning the ranking and thus the data sparsity problem can be alleviated. In addition, our model can easily incorporate different types of context information, such as the geographical influence and temporal influence. We propose a stochastic gradient descent based algorithm to learn the factorization. Experiments on publicly available datasets under both user-POI setting and user-time-POI setting have been conducted to test the effectiveness of the proposed method. Experimental results under both settings show that the proposed method outperforms the state-of-the-art methods significantly in terms of recommendation accuracy.
Xutao Li 0003, Gao Cong, Xiaoli Li 0001, Tuan-Anh Nguyen Pham, Shonali Krishnaswamy
SIGIR1
2014 Multi-label collective classification via Markov chain based learning method
Qingyao Wu, Michael Kwok-Po Ng, Yunming Ye, Xutao Li 0003, Ruichao Shi, Yan Li 0040
Knowl. Based Syst.4
2014 MultiComm: Finding Community Structurein Multi-Dimensional Networks
abstract
The main aim of this paper is to develop a community discovery scheme in a multi-dimensional network for data mining applications. In online social media, networked data consists of multiple dimensions/entities such as users, tags, photos, comments, and stories. We are interested in finding a group of users who interact significantly on these media entities. In a co-citation network, we are interested in finding a group of authors who relate to other authors significantly on publication information in titles, abstracts, and keywords as multiple dimensions/entities in the network. The main contribution of this paper is to propose a framework (MultiComm)to identify a seed-based community in a multi-dimensional network by evaluating the affinity between two items in the same type of entity (same dimension)or different types of entities (different dimensions)from the network. Our idea is to calculate the probabilities of visiting each item in each dimension, and compare their values to generate communities from a set of seed items. In order to evaluate a high quality of generated communities by the proposed algorithm, we develop and study a local modularity measure of a community in a multi-dimensional network. Experiments based on synthetic and real-world data sets suggest that the proposed framework is able to find a community effectively. Experimental results have also shown that the performance of the proposed algorithm is better in accuracy than the other testing algorithms in finding communities in multi-dimensional networks.
Xutao Li 0003, Michael Kwok-Po Ng, Yunming Ye
IEEE Trans. Knowl. Data Eng.1
2013 Stratified sampling for feature subspace selection in random forests for high dimensional data
Yunming Ye, Qingyao Wu, Joshua Zhexue Huang, Michael Kwok-Po Ng, Xutao Li 0003
Pattern Recognit.5
2012 MultiFacTV: Finding modules from higher-order gene expression profiles with time dimension
abstract
Module detection is an important task in bioinformatics which aims at finding a set of cells/genes that interact together to be responsible for some biological functionalities. In this paper, we propose a novel tensor factorization approach to finding modules from higher-order gene expression profiles with the time dimension, e.g., gene × condition × time data. The main idea is to incorporate a total variation regularization term for the time dimension during the tensor factorization, and then use the factorization results to identify the modules. Experimental results on two real gene × condition × time datasets have shown the effectiveness of the proposed method.
Xutao Li 0003, Yunming Ye, Qingyao Wu, Michael Kwok-Po Ng
BIBM1
2012 HAR: Hub, Authority and Relevance Scores in Multi-Relational Data for Query Search
abstract
In this paper, we propose a framework HAR to study the hub and authority scores of objects, and the relevance scores of relations in multi-relational data for query search. The basic idea of our framework is to consider a random walk in multi-relational data, and study in such random walk, limiting probabilities of relations for relevance scores, and of objects for hub scores and authority scores. The main contribution of this paper is to (i) propose a framework (HAR) that can compute the hub, authority and relevance scores by solving limiting probabilities arising from multi-relational data, and can incorporate input query vectors to handle query-specific search; (ii) show existence and uniqueness of such limiting probabilities so that they can be used for query search effectively; and (iii) develop an iterative algorithm to solve a set of tensor (multivariate polynomial) equations to obtain such probabilities. Extensive experimental results on TREC and DBLP data sets suggest that the proposed method is very effective in obtaining relevant results to the querying inputs. In the comparison, we find that the performance of HAR is better than those of HITS, SALSA and TOPHITS.
Xutao Li 0003, Michael Kwok-Po Ng, Yunming Ye
SDM1
2011 MultiRank: co-ranking for objects and relations in multi-relational data
abstract
The main aim of this paper is to design a co-ranking scheme for objects and relations in multi-relational data. It has many important applications in data mining and information retrieval. However, in the literature, there is a lack of a general framework to deal with multi-relational data for co-ranking. The main contribution of this paper is to (i) propose a framework (MultiRank) to determine the importance of both objects and relations simultaneously based on a probability distribution computed from multi-relational data; (ii) show the existence and uniqueness of such probability distribution so that it can be used for co-ranking for objects and relations very effectively; and (iii) develop an efficient iterative algorithm to solve a set of tensor (multivariate polynomial) equations to obtain such probability distribution. Extensive experiments on real-world data suggest that the proposed framework is able to provide a co-ranking scheme for objects and relations successfully. Experimental results have also shown that our algorithm is computationally efficient, and effective for identification of interesting and explainable co-ranking results.
Michael Kwok-Po Ng, Xutao Li 0003, Yunming Ye
KDD2
2010 On cluster tree for nested and multi-density data clustering
Xutao Li 0003, Yunming Ye, Mark Junjie Li, Michael Kwok-Po Ng
Pattern Recognit.1