EDBT 2026 Demo / reviewers in the wild / expert
Chuyao Luo
dblp:210/4784
· DBLP profile ↗
30ranked-venue papers
4as first author
29since 2021 · last 2026
0000-0003-4848-609XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Satellite-Text-Prompted Large Language Model for Photovoltaic Power ForecastingabstractPhotovoltaic (PV) power forecasting is critical for the operation of solar power plants and the coordination of energy within power grids. This work aims to predict future PV power time series by leveraging multimodal data. While recent studies have incorporated numerical modalities such as satellite image sequences and numerical weather prediction (NWP) time series, they often overlook textual modalities—such as the spatio-temporal context of PV plants—and the potential of pretrained large language models (LLMs). In this paper, we build upon existing numerical inputs and further explore the use of spatio-temporal text prompts, generated based on plant coordinates and forecast start time, to enhance the forecasting process. We propose PV-LLM, a satellite-text-prompted framework that integrates a pretrained LLM to improve PV power forecasting. The framework consists of three key components: Text Prompt Construction, Modality-Specific Encoding, and Adaptive Prompt Tuning. First, the Text Prompt Construction module generates spatio-temporal prompts that offer high-level semantic guidance. Next, the Modality-Specific Encoding module encodes each modality according to its unique characteristics, capturing modality-specific patterns while managing varying context lengths. Finally, the Adaptive Prompt Tuning module fine-tunes the LLM to integrate multimodal embeddings, while an adaptive gating mechanism retains its pretrained knowledge. We validate the effectiveness of the proposed framework on a real-world dataset containing multiple PV plants. Experimental results demonstrate that our approach outperforms existing state-of-the-art methods. Jianghong Ma, Baoquan Zhang, Kenghong Lin, Chuyao Luo, Xutao Li 0001, Yunming Ye |
AAAI | 6 |
| 2026 | LMcast: A pretrained language model guided long-term memory transformer for precipitation nowcasting
Feifan Gao, Chuyao Luo, Guangbo Deng, Xutao Li 0003, Baoquan Zhang, Demin Yu, Yunming Ye |
Neural Networks | 2 |
| 2025 | Integrating Multi-Source Data for Long Sequence Precipitation ForecastingabstractLong-sequence precipitation forecasting is critical for both meteorological science and smart city applications. The primary objective of this task is to predict future radar echo sequences, which provide high resolution and timely references for atmospheric precipitation distribution based on current observations. However, the chaotic nature of precipitation systems poses significant challenges in extending reliable forecast horizons. Most existing methods struggle with accuracy and clarity when extended to long-sequence predictions, such as three-hour forecasts. This is primarily due to the insufficiency of spatio-temporal information within a single modality over time. In this paper, we propose a cascading forecasting framework that adaptively extracts and integrates multimodal spatio-temporal information to support accurate and realistic long-sequence radar forecasting. Our framework includes a temporal adaptive predictor and a flow-based precipitation distribution adaptor. The predictor utilizes a multi-branch encoder-decoder architecture. This design allows it to extract meteorological sequences from multiple sources at varying scales, resulting in an initial global precipitation estimate. The core component is a carefully designed cross-attention module with a temporal adaptive layer to enhance multi-modality alignment. The initial estimate is then refined by the flow-based adaptor, which adjusts the prediction to match the target precipitation distribution, enhancing local details and correcting extreme precipitation patterns. We validated our method using real multi-source dataset for long-sequence forecasting, and the experimental results demonstrate that our approach outperforms existing state-of-the-art methods. Demin Yu, Wenzhi Feng, Kenghong Lin, Xutao Li 0003, Yunming Ye, Chuyao Luo, Wenchuan Du |
AAAI | 6 |
| 2025 | PiMMNet: Introducing Multi-Modal Precipitation Nowcasting via a Physics-informed PerspectiveabstractPrecipitation nowcasting plays a pivotal role in urban planning and disaster mitigation, where extending forecast horizons offers critical advantages for proactive decision-making. Most data-driven methods focus on modeling radar echo sequences through end-to-end spatiotemporal predictive learning, yielding precise short-term predictions; however, they fundamentally neglect the inherent physical mechanism governing precipitation system. Moreover, approaches relying solely on single-modality radar observations suffer from persistent information bottlenecks, severely limiting their temporal generalizability for extended forecasting. To address these challenges, we propose PiMMNet, a Physics-informed Multi-Modal Network. It is constructed based on the advection-diffusion principle from fluid dynamics, explicitly modeling the precipitation evolution as a spatiotemporal transport processes characterized by the deterministic advection and the stochastic source. We carefully design a multi-model motion estimation network and a motion-guided diffusion model to describe the deterministic and stochastic terms, respectively. The core innovation of our method lies in jointly estimating a physics-constrained velocity field from multi-modal inputs (radar and satellite data). In this case, we naturally align the motion evolution among modalities into a unified representation, inherently mitigating cross-modal distribution biases. Experimental evaluations on two real-world multi-modal meteorological datasets demonstrate the efficacy of our approach, showcasing significant improvements in accuracy and robustness for longer-range precipitation nowcasting. Our code are available at https://github.com/DeminYu98/PiMMNet. Demin Yu, Wenchuan Du, Kenghong Lin, Xutao Li 0001, Yunming Ye, Chuyao Luo, Xunlai Chen |
ACM Multimedia | 6 |
| 2025 | ZooKT: Task-adaptive knowledge transfer of Model Zoo for few-shot learning
Baoquan Zhang, Bingqi Shan, Aoxue Li, Chuyao Luo, Yunming Ye, Zhenguo Li |
Pattern Recognit. | 4 |
| 2025 | RS-MoE: A Vision-Language Model With Mixture of Experts for Remote Sensing Image Captioning and Visual Question AnsweringabstractRemote sensing image captioning (RSIC) presents unique challenges and plays a critical role in applications such as environmental monitoring, urban planning, and disaster management. Traditional RSIC methods often struggle to produce rich and diverse descriptions. Recently, with significant advancements in vision-language models (VLMs), efforts have emerged to integrate these models into the remote sensing domain and to introduce richly descriptive datasets specifically designed to enhance VLM training. However, most current RSIC models generally apply only fine-tuning to these datasets without developing models tailored to the unique characteristics of remote sensing imagery. This article proposes RS-MoE, the first mixture of expert (MoE)-based VLM specifically customized for remote sensing domain. Unlike traditional MoE models, the core of RS-MoE is the MoE block, which incorporates a novel instruction router and multiple lightweight large language models (LLMs) as expert models. The instruction router is designed to generate specific prompts tailored for each corresponding LLM, guiding them to focus on distinct aspects of the RSIC task. This design not only allows each expert LLM to concentrate on a specific subset of the task, thereby enhancing the specificity and accuracy of the generated captions, but also improves the scalability of the model by facilitating parallel processing of subtasks. In addition, we present a two-stage training strategy for tuning our RS-MoE model to prevent performance degradation due to sparsity. We fine-tuned our model on the RSICap dataset using our proposed training strategy. Experimental results on the RSICap dataset, along with evaluations on other traditional datasets where no additional fine-tuning was applied, demonstrate that our model achieves state-of-the-art performance in generating precise and contextually relevant captions. Notably, our RS-MoE-1B variant achieves performance comparable to 13B VLMs, demonstrating the efficiency of our model design. Moreover, our model demonstrates promising generalization capabilities by consistently achieving state-of-the-art performance on the remote sensing visual question answering (RSVQA) task. Danfeng Hong, Shuhang Ge, Chuyao Luo, Congcong Wen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | iTrendRNN: An Interpretable Trend-Aware RNN for Meteorological Spatiotemporal PredictionabstractAccurate prediction of meteorological elements, such as temperature and relative humidity, is important to human livelihood, early warning of extreme weather, and urban governance. Recently, neural network-based methods have shown impressive performance in this field. However, most of them are overcomplicated and impenetrable. In this paper, we propose a straightforward and interpretable differential framework, where the key lies in explicitly estimating the evolutionary trends. Specifically, three types of trends are exploited. (1) The proximity trend simply uses the most recent changes. It works well for approximately linear evolution. (2) The sequential trend explores the global information, aiming to capture the nonlinear dynamics. Here, we develop an attention-based trend unit to help memorize long-term features. (3) The flow trend is motivated by the nature of evolution, i.e., the heat or substance flows from one region to another. Here, we design a flow-aware attention unit. It can reflect the interactions via performing spatial attention over flow maps. Finally, we develop a trend fusion module to adaptively fuse the above three trends. Extensive experiments on two datasets demonstrate the effectiveness of our method. Chuyao Luo, Bowen Zhang 0005, Huiwei Lin, Xutao Li 0003, Yunming Ye |
AAAI | 2 |
| 2024 | MetaDiff: Meta-Learning with Conditional Diffusion for Few-Shot LearningabstractEquipping a deep model the ability of few-shot learning (FSL) is a core challenge for artificial intelligence. Gradient-based meta-learning effectively addresses the challenge by learning how to learn novel tasks. Its key idea is learning a deep model in a bi-level optimization manner, where the outer-loop process learns a shared gradient descent algorithm (called meta-optimizer), while the inner-loop process leverages it to optimize a task-specific base learner with few examples. Although these methods have shown superior performance on FSL, the outer-loop process requires calculating second-order derivatives along the inner-loop path, which imposes considerable memory burdens and the risk of vanishing gradients. This degrades meta-learning performance. Inspired by recent diffusion models, we find that the inner-loop gradient descent process can be viewed as a reverse process (i.e., denoising) of diffusion where the target of denoising is the weight of base learner but origin data. Based on this fact, we propose to model the gradient descent algorithm as a diffusion model and then present a novel conditional diffusion-based meta-learning, called MetaDiff, that effectively models the optimization process of base learner weights from Gaussian initialization to target weights in a denoising manner. Thanks to the training efficiency of diffusion models, our MetaDiff does not need to differentiate through the inner-loop path such that the memory burdens and the risk of vanishing gradients can be effectively alleviated for improving FSL. Experimental results show that our MetaDiff outperforms state-of-the-art gradient-based meta-learning family on FSL tasks. Baoquan Zhang, Chuyao Luo, Demin Yu, Xutao Li 0003, Huiwei Lin, Yunming Ye, Bowen Zhang 0005 |
AAAI | 2 |
| 2024 | DiffCast: A Unified Framework via Residual Diffusion for Precipitation NowcastingabstractPrecipitation nowcasting is an important spatiotemporal prediction task to predict the radar echoes sequences based on current observations, which can serve both meteorological science and smart city applications. Due to the chaotic evolution nature of the precipitation systems, it is a very challenging problem. Previous studies address the problem either from the perspectives of deterministic modeling or probabilistic modeling. However, their predictions suffer from the blurry, high-value echoes fading away and position inaccurate issues. The root reason of these issues is that the chaotic evolutionary precipitation systems are not appropriately modeled. Inspired by the nature of the systems, we propose to decompose and model them from the perspective of global deterministic motion and local stochastic variations with residual mechanism. A unified and flexible framework that can equip any type of spatio-temporal models is proposed based on residual diffusion, which effectively tackles the shortcomings of previous methods. Extensive experimental results on four publicly available radar datasets demonstrate the effectiveness and superiority of the proposed framework, compared to state-of-the-art techniques. Our code is publicly available at https://github.com/DeminYu98/DiffCast. Demin Yu, Xutao Li 0003, Yunming Ye, Baoquan Zhang, Chuyao Luo, Kuai Dai, Xunlai Chen |
CVPR | 5 |
| 2024 | Codebook Transfer with Part-of-Speech for Vector-Quantized Image ModelingabstractVector-Quantized Image Modeling (VQIM) is a fundamental research problem in image synthesis, which aims to represent an image with a discrete token sequence. Existing studies effectively address this problem by learning a discrete codebook from scratch and in a code-independent manner to quantize continuous representations into discrete tokens. However, learning a codebook from scratch and in a code-independent manner is highly challenging, which may be a key reason causing codebook collapse, i.e., some code vectors can rarely be optimized without regard to the relationship between codes and good codebook priors such that die off finally. In this paper, inspired by pretrained language models, we find that these language models have actually pretrained a superior codebook via a large number of text corpus, but such information is rarely exploited in VQIM. To this end, we propose a novel codebook transfer framework with part-of-speech, called VQCT, which aims to transfer a well-trained codebook from pretrained language models to VQIM for robust codebook learning. Specifically, we first introduce a pretrained codebook from language models and part-of-speech knowledge as priors. Then, we construct a vision-related codebook with these priors for achieving codebook transfer. Finally, a novel codebook transfer network is designed to exploit abundant semantic relationships between codes contained in pretrained codebooks for robust VQIM codebook learning. Experimental results on four datasets show that our VQCT method achieves superior VQIM performance over previous state-of-the-art methods. Baoquan Zhang, Huaibin Wang, Chuyao Luo, Xutao Li 0003, Guotao Liang, Yunming Ye, Xiaochen Qi |
CVPR | 3 |
| 2024 | LG-VQ: Language-Guided Codebook LearningabstractVector quantization (VQ) is a key technique in high-resolution and high-fidelity image synthesis, which aims to learn a codebook to encode an image with a sequence of discrete codes and then generate an image in an auto-regression manner.
Although existing methods have shown superior performance, most methods prefer to learn a single-modal codebook (\emph{e.g.}, image), resulting in suboptimal performance when the codebook is applied to multi-modal downstream tasks (\emph{e.g.}, text-to-image, image captioning) due to the existence of modal gaps.
In this paper, we propose a novel language-guided codebook learning framework, called LG-VQ, which aims to learn a codebook that can be aligned with the text to improve the performance of multi-modal downstream tasks. Specifically, we first introduce pre-trained text semantics as prior knowledge, then design two novel alignment modules (\emph{i.e.}, Semantic Alignment Module, and Relationship Alignment Module) to transfer such prior knowledge into codes for achieving codebook text alignment.
In particular, our LG-VQ method is model-agnostic, which can be easily integrated into existing VQ models. Experimental results show that our method achieves superior performance on reconstruction and various multi-modal downstream tasks. Guotao Liang, Baoquan Zhang, Yaowei Wang 0001, Yunming Ye, Xutao Li 0003, Huaibin Wang, Chuyao Luo, Kola Ye, Linfeng Luo |
NeurIPS | 7 |
| 2024 | LGCNet: A Cloud Detection Method in Remote Sensing Images Using Local and Global SemanticsabstractDetecting and eliminating clouds is a crucial step in remote sensing image (RSI) preprocessing. The removal of clouds can significantly enhance the performance of subsequent remote sensing applications. Existing deep learning (DL)-based cloud detection methods extract semantic information to improve feature representation and, subsequently, detection performance. However, these methods do not fully utilize the potential of context semantic information. Besides, to capture semantics from large receptive fields, they employ convolution operators with large kernel sizes, which results in high computational costs. Thus, these computationally heavy models are not suitable for resource-limited devices, particularly satellites. To address this issue, we propose a cloud detection model, LGCNet. This model efficiently extracts both local and global contextual information, fully utilizing semantics while reducing resource usage. LGCNet is built on an encoder-decoder structure. Specifically, the encoder extracts local scale-aware semantics through proposed local semantic blocks (LSBs), which are then skip-connected to the decoder. This approach provides adaptive and diverse local contextual information. On the top of the encoder, the high-level global semantics are captured via the proposed global feature TransBlock (GFTB). A variety of extracted semantics ensure improved detection performance. We evaluate the proposed method using two public datasets: LandSat8 and Moderate-Resolution Imaging Spectroradiometer (MODIS). We conducted experiments on both a server and an edge computing device. Our extensive experiments revealed that LGCNet outperforms other lightweight cloud detection and semantic segmentation methods in terms of performance and computational load. Shanshan Feng 0001, Huaibin Wang, Baoquan Zhang, Pengjuan Yao, Chuyao Luo, Yunming Ye, Yong Xu 0001, Xutao Li 0003 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | A Practical Online Incremental Learning Framework for Precipitation NowcastingabstractPrecipitation nowcasting plays an important role in our life. Many deep learning-based methods are proposed for precipitation nowcasting by predicting radar echo sequence over the past years, and achieving better performance than traditional approaches. However, all of them are based on a static model, which is trained in offline learning and does not adapt to real-time changing precipitation data. Recently, online incremental learning (OIL) has been proposed to dynamically update the model by continually learning new data and preventing the forgetting of historical knowledge in an online fashion. While effective, existing OIL approaches that focus on a classification task are not suitable for the regression task of precipitation nowcasting. To fill this gap, we try to propose a novel OIL framework for precipitation nowcasting. By analyzing its characteristics, we find three challenges: 1) the distributions of radar echo maps in different rainfall events are different; 2) in each rainfall event, there is always an inevitable delay between the timestamps of the training and testing samples; and 3) the real-time requirement for model prediction is very high, which has strict limitations on the training speed of the model. Based on these observations, we propose a practical OIL framework based on gradient activation mapping (GAM). It can be mainly divided into three components: 1) recall training strategy (RTS) is used to eliminate the interference caused by distributions of different events; 2) iterative approximation training (IAT) is designed to align the timestamps of the training and testing; and 3) moreover, we propose gradient activation mapping weight (GAMW) to improve the training effectiveness. Extensive experiments show that the proposed framework can improve the performance of the model stably and effectively. Especially in those heavy rainfall regions where it usually causes more threat to human activity, the improvement is significant. Chuyao Luo, Zheng Zhang 0046, Huiwei Lin, Baoquan Zhang, Xutao Li 0003, Yunming Ye |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Cross-Modal Hashing With Feature Semi-Interaction and Semantic Ranking for Remote Sensing Ship Image RetrievalabstractCross-modal hashing plays a pivotal role in large-scale remote sensing (RS) ship image retrieval. RS ship images often exhibit similar overall appearance with subtle differences. Existing hashing methods typically employ feature non-interaction strategies to generate common hash codes, which may not effectively capture the correlations between cross-modal ship images to reduce intermodality discrepancies. To address this issue, we propose a novel cross-modal hashing approach based on feature semi-interaction and semantic ranking (FSISR) for RS ship image retrieval. Our FSISR approach not only captures intricate correlations between different ship image modalities, but also enables the construction of hash tables for large-scale retrieval. FSISR comprises a feature semi-interaction module and a semantic ranking objective function. The semi-interaction module utilizes clustering centers from one modality to learn the correlations between two modalities and generate robust shared representations. The objective function optimizes these representations in a common Hamming space, consisting of a shared semantic alignment loss and a margin-free ranking loss. The alignment loss employs a shared semantic layer to preserve label-level similarity, while the ranking loss incorporates hard examples to establish a margin-free loss that captures similarity ranking relationships. We evaluate the performance of our method on benchmark datasets and demonstrate its effectiveness for cross-modal RS ship image retrieval.https://github.com/sunyuxi/FSISR. Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Yifang Ban, Sebastian Hafner, Xutao Li 0003, Chuyao Luo, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2024 | Multiscale and Multilevel Feature Fusion Network for Quantitative Precipitation Estimation With Passive MicrowaveabstractPassive microwave (PMW) radiometers have been widely utilized for quantitative precipitation estimation (QPE) by leveraging the relationship between brightness temperature (Tb) and rain rate. Nevertheless, accurate precipitation estimation remains a challenge due to the intricate relationship between them, which is influenced by a diverse range of complex atmospheric and surface properties. In addition, the inherent skew distribution of rainfall values prevents models from correctly addressing extreme precipitation events, leading to a significant underestimation. This article presents a novel model called the multiscale and multilevel feature fusion network (MSMLNet), consisting of two essential components: a multiscale feature extractor and a multilevel regression predictor. The feature extractor is specifically designed to extract characteristics from multiple scales, enabling the model to incorporate various meteorological conditions, as well as atmospheric and surface information in the surrounding environment. The regression predictor first assesses the probabilities of multiple rainfall levels for each observed pixel and then extracts features of different levels separately. The multilevel features are fused according to the predicted probabilities. This approach allows each submodule only to focus on a specific range of precipitation, avoiding the undesirable effects of skew distributions. To evaluate the performance of MSMLNet, various deep learning methods are adapted for the precipitation retrieval task, and a PWM-based product from the global precipitation measurement (GPM) mission is also used for comparison. Extensive experiments show that MSMLNet surpasses GMI-based products and the most advanced deep learning approaches by 17.9% and 2.5% in root mean square error (RMSE), and 54.2% and 4.0% in CSI-10, respectively. Moreover, we demonstrate that MSMLNet significantly mitigates the propensity for underestimating heavy precipitation events and has a consistent and outstanding performance in estimating precipitation across various levels. Xutao Li 0003, Kenghong Lin, Chuyao Luo, Yunming Ye, Xiuqing Hu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | MCSDNet: Mesoscale Convective System Detection Network via Multiscale Spatiotemporal InformationabstractThe accurate detection of mesoscale convective systems (MCSs) is crucial for meteorological monitoring due to their potential to cause significant destruction through severe weather phenomena, such as hail, thunderstorms, and heavy rainfall. However, the existing methods for MCS detection mostly targets on single-frame detection, which just considers the static characteristics and ignores the temporal evolution in the life cycle of MCS. In this article, we propose a novel encoder-decoder neural network named mesoscale convective system detection network (MCSDNet) to detect MCS regions. MCSDNet has a simple architecture and is easy to expand. Different from the previous models, MCSDNet targets on multiframes detection and leverages multiscale spatiotemporal information in remote sensing imagery (RSI). As far as we know, it is the first work to utilize multiscale spatiotemporal information to detect MCS regions. First, we design a multiscale spatiotemporal information module to extract multilevel semantic from different encoder levels, which makes our models can extract more detail spatiotemporal features. Second, spatiotemporal mix unit (STMU), a dual spatiotemporal attention, is introduced to MCSDNet to capture both intraframe features and interframe. Finally, we present MCS remote sensing image (MCSRSI) the first publicly available dataset for multiframes MCS detection based on FY-4A satellite. We also conduct several experiments on MCSRSI and find that our proposed MCSDNet achieves the best performance on MCS detection task when comparing with other baseline methods. We hope that the combination of our open-access dataset and promising results will encourage the future research for MCS detection task and provide a robust framework for related tasks in atmospheric science. Our code is available at:https://github.com/250HandsomeLiang/MCSDNet.git Baoquan Zhang, Jiajun Liang, Rui Ye 0002, Chuyao Luo, Xutao Li 0003, Yunming Ye, Xukai Fu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Toward a Variation-Aware and Interpretable Model for Radar Image Sequence PredictionabstractRadar image sequence prediction (RISP) aims to predict future radar images based on historical observations. In the past few years, neural network-based methods have shown impressive performance for RISP. However, two limitations stills exist. 1) They fail to exploit variation information when capturing spatial dependencies. 2) They neglect to analyze and interpret the model. In this article, we propose a variation-aware prediction model for the first limitation, and develop a relevance propagation technique for the second one. Specifically, 1) we recustomize the vanilla convolution by introducing a variation-aware item. The new convolution unit yields two advantages when capturing spatial dependencies, i.e., exploiting variation information and offering spatially-varying kernels. As a result, it can learn the diverse and complex radar echo patterns. By equipping the unit into a typical network (PredRNN), we propose a novel prediction model, dubbed as VA-PredRNN. 2) As for analyzing our model, we propagate the output backward layer by layer till the input. Hence, we can reveal the relevance between the output and the intermediate states. To the best of the authors' knowledge, this is the first work to study the interpretability of a multilayer RISP model. We conduct extensive experiments on two datasets, and the results demonstrate the effectiveness of our VA-PredRNN. We also carry out a series of analyses using the proposed relevance propagation technique. According to the results, we discover the importance of different states. Yunming Ye, Bowen Zhang 0005, Huiwei Lin, Yuxi Sun 0002, Xutao Li 0003, Chuyao Luo |
IEEE Trans. Ind. Informatics | 7 |
| 2023 | STGV-Similarity between trend generating vectors: A new sample weighting scheme for stock trend prediction using financial features of companies
Yueyue Yao, Chuyao Luo, Ka-Cheong Leung, Yunming Ye |
Expert Syst. Appl. | 2 |
| 2023 | UNIMEMnet: Learning long-term motion and appearance dynamics for video prediction with a unified memory network
Kuai Dai, Xutao Li 0003, Chuyao Luo, Wuqiao Chen, Yunming Ye, Shanshan Feng 0001 |
Neural Networks | 3 |
| 2023 | Cross-Domain Aspect-Based Sentiment Classification by Exploiting Domain- Invariant Semantic-Primary FeatureabstractAspect-based sentiment analysis is an important task in fine-grained sentiment analysis, which aims to infer the sentiment towards a given aspect. Previous studies have shown notable success when sufficient labeled training data is available. However, annotating adequate data is labor-intensive, which sets substantial barriers for generalizing the sentiment predictor to the new domain. Two main challenges exist in cross-domain aspect-based sentiment analysis. One challenge is acquiring the domain-invariant knowledge; the other challenge is mining the syntactic-related words towards the aspect-term. In this article, we propose a transformer-based semantic-primary knowledge transferring network (TSPKT) for cross-domain aspect-term sentiment analysis, which utilizes semantic-primary knowledge as a bridge to enable knowledge transfer across domains. Specifically, we first build an S-Graph from external semantic lexicons, and extract the semantic-primary knowledge from the S-Graph. Second, AoaGraphormer is proposed to learn the syntactically relevant words towards the aspect-term. Third, we extend the standard biLSTM classifier to fully integrate the semantic-primary knowledge by adding a novel knowledge-aware memory unit (KAMU) to the biLSTM cell. Extensive experiments on six cross-domain setups demonstrate the superiority of TSPKT against the state-of-the-art baseline methods. Bowen Zhang 0005, Xianghua Fu, Chuyao Luo, Yunming Ye, Xutao Li 0003, Liwen Jing 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2023 | Knowledge-enhanced Prompt-tuning for Stance DetectionabstractInvestigating public attitudes on social media is important in opinion mining systems. Stance detection aims to analyze the attitude of an opinionated text (e.g., favor, neutral, or against) toward a given target. Existing methods mainly address this problem from the perspective of fine-tuning. Recently, prompt-tuning has achieved success in natural language processing tasks. However, conducting prompt-tuning methods for stance detection in real-world remains a challenge for several reasons: (1) The text form of stance detection is usually short and informal, which makes it difficult to design label words for the verbalizer. (2) The tweet text may not explicitly give the attitude. Instead, users may use various hashtags or background knowledge to express stance-aware perspectives. In this article, we first propose a prompt-tuning-based framework that performs stance detection in a cloze question manner. Specifically, a knowledge-enhanced prompt-tuning framework (KEprompt) method is designed, which consists of an automatic verbalizer (AutoV) and background knowledge injection (BKI). Specifically, in AutoV, we introduce a semantic graph to build a better mapping from the predicted word of the pretrained language model and detection labels. In BKI, we first propose a topic model for learning hashtag representation and introduce ConceptGraph as the supplement of the target. At last, we present a challenging dataset for stance detection, where all stance categories are expressed in an implicit manner. Extensive experiments on a large real-world dataset demonstrate the superiority of KEprompt over state-of-the-art methods. Hu Huang 0009, Bowen Zhang 0005, Xiang-Yang Li 0001, Baoquan Zhang, Yuxi Sun 0002, Chuyao Luo, Cheng Peng 0003 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2023 | Adaptive Transfer of Graph Neural Networks for Few-Shot Molecular Property PredictionabstractFew-Shot Molecular Property Prediction (FSMPP) is an improtant task on drug discovery, which aims to learn transferable knowledge from base property prediction tasks with sufficient data for predicting novel properties with few labeled molecules. Its key challenge is how to alleviate the data scarcity issue of novel properties. Pretrained Graph Neural Network (GNN) based FSMPP methods effectively address the challenge by pre-training a GNN from large-scale self-supervised tasks and then finetuning it on base property prediction tasks to perform novel property prediction. However, in this paper, we find that the GNN finetuning step is not always effective, which even degrades the performance of pretrained GNN on some novel properties. This is because these molecule-property relationships among molecules change across different properties, which results in the finetuned GNN overfits to base properties and harms the transferability performance of pretrained GNN on novel properties. To address this issue, in this paper, we propose a novel Adaptive Transfer framework of GNN for FSMPP, called ATGNN, which transfers the knowledge of pretrained and finetuned GNNs in a task-adaptive manner to adapt novel properties. Specifically, we first regard the pretrained and finetuned GNNs as model priors of target-property GNN. Then, a task-adaptive weight prediction network is designed to leverage these priors to predict target GNN weights for novel properties. Finally, we combine our ATGNN framework with existing FSMPP methods for FSMPP. Extensive experiments on four real-world datasets, i.e., Tox21, SIDER, MUV, and ToxCast, show the effectiveness of our ATGNN framework. Baoquan Zhang, Chuyao Luo, Hao Jiang 0051, Shanshan Feng 0001, Xutao Li 0003, Bowen Zhang 0005, Yunming Ye |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | On Understanding of Spatiotemporal Prediction ModelabstractRecently, explainable artificial intelligence has received considerable attention. Most existing studies are focusing on the tasks of CNNs-based image classification and RNNs-based time series analysis. In this paper, we pay attention to the more complicated spatiotemporal predictive learning task (SPLT), where both the spatial and temporal information play important roles. To explain the internal mechanism of spatiotemporal prediction models, we propose a comprehensive analysis method. Specifically, with a typical encoder-decoder framework, we focus on two core issues of SPLT: image generation and spatiotemporal dynamics. For the first issue, we develop aquantitative channel perturbationmethod to explore the importance of features to prediction. Furthermore, we propose a technique called thesynthesis of multiple independent componentsto analyze how these features generate the prediction. According to the experimental results, thecoarse- and fine-grainedsynthesis (CFGS) mechanism is drawn for image generation in SPLT. For the second issue, we propose astate decompositiontechnique and astate expansiontechnique to disentangle coupled signals in the spatiotemporal dynamical system. This helps us to explore the mechanism of forming motion. Moreover, to diagnose the movement of a particular region during analysis, we propose a fluorescent stamp-based technique. By observing extensive experimental results, we summarize a collaboration mechanism to explain how the motion is formed in SPLT, namely, theextending the present and erasing the past (EPEP)mechanism. To the best of our knowledge, this is the first work to interpret the internal mechanism of SPLT models. Xutao Li 0003, Yunming Ye, Shanshan Feng 0001, Chuyao Luo, Bowen Zhang 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Cross-View Object Geo-Localization in a Local Region With Satellite ImageryabstractCross-view geo-localization is a critical task in various applications, such as smart city management and disaster monitoring. Current methods typically divide a satellite image into patches and use these patches to identify the geographic location of a query image. However, these methods can only provide the location of an image rather than the location of a specific object of interest. This makes it difficult to link these methods to GeoDatabases to obtain detailed information about a target object, such as its name and construction time. To overcome this limitation, we propose a novel problem of cross-view object geo-localization in a local region with high-resolution satellite images. This problem includes two main challenges: accurately identifying the location of an object and distinguishing the target object from others in satellite images. To address these challenges, we present a new Detection-based Geo-localization method called DetGeo, which consists of an object detection-based framework with a two-branch encoder and a query-aware cross-view fusion module. DetGeo uses cross-view images as input to the detector to provide object-level geo-localization. The fusion module employs cross-view spatial attention to focus on relevant areas of target objects during cross-view feature fusion. To evaluate our method, we constructed a new Cross-View Object Geo-Localization dataset called CVOGL, which comprises ground-view or drone-view images as query images and satellite-view images as geo-tagged reference images. Comprehensive experiments are conducted to demonstrate the effectiveness of our method on CVOGL. https://github.com/sunyuxi/DetGeo. Yuxi Sun 0002, Yunming Ye, Jian Kang 0005, Rubén Fernández-Beltran, Shanshan Feng 0001, Xutao Li 0003, Chuyao Luo, Puzhao Zhang, Antonio Plaza |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | The reconstitution predictive network for precipitation nowcasting
Chuyao Luo, Guangning Xu, Xutao Li 0003, Yunming Ye |
Neurocomputing | 1 |
| 2022 | SPLNet: A sequence-to-one learning network with time-variant structure for regional wind speed prediction
Rui Ye 0002, Shanshan Feng 0001, Xutao Li 0003, Yunming Ye, Baoquan Zhang, Chuyao Luo |
Inf. Sci. | 6 |
| 2022 | PredRANN: The spatiotemporal attention Convolution Recurrent Neural Network for precipitation nowcasting
Chuyao Luo, Xinyue Zhao, Yuxi Sun 0002, Xutao Li 0003, Yunming Ye |
Knowl. Based Syst. | 1 |
| 2022 | Multisensor Fusion and Explicit Semantic Preserving-Based Deep Hashing for Cross-Modal Remote Sensing Image RetrievalabstractCross-modal hashing is an important tool for retrieving useful information from very-high-resolution (VHR) optical images and synthetic aperture radar (SAR) images. Dealing with the intermodal discrepancies, including both spatial–spectral and visual semantic aspects, between VHR and SAR images is extremely vital to generate high-quality common hash codes in the Hamming space. However, existing cross-modal hashing methods ignore the spatial–spectral discrepancy when representing VHR and SAR images. Moreover, existing methods employ derived supervised signals, such as pairwise training images, to implicitly guide hashing learning, which fails to effectively deal with the visual semantic discrepancy, i.e., cannot adequately preserve the intraclass similarity and interclass discrimination between VHR and SAR images. To address these drawbacks, this article proposes a multisensor fusion and explicit semantic preserving-based deep Hashing method, termed as MsEspH, which can effectively deal with the discrepancies. Specifically, we design a novel cross-modal hashing network to eliminate the spatial–spectral discrepancies by fusing extra multispectral images (MSIs), which are generated in real time by a generative adversarial network. Then, we propose an explicit semantic preserving-based objective function by analyzing the connection between classification and hash learning. The objective function can preserve the intraclass similarity and interclass discrimination with class labels directly. Moreover, we theoretically verify that hash learning and classification can be unified into a learning framework under certain conditions. To evaluate our method, we construct and release a large-scale VHR-SAR image dataset. Extensive experiments on the dataset demonstrate that our method outperforms various state-of-the-art cross-modal hashing methods. Yuxi Sun 0002, Shanshan Feng 0001, Yunming Ye, Xutao Li 0003, Jian Kang 0005, Zhichao Huang 0001, Chuyao Luo |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Experimental Study on Generative Adversarial Network for Precipitation NowcastingabstractPrecipitation nowcasting is an important task, which can be used in numerous applications. The key challenge of the task lies in radar echo map prediction. Previous studies leverage Convolutional Recurrent Neural Network (ConvRNN) to address the problem. However, the approaches are built upon mean square losses and the results tend to have inaccurate appearances, shapes and positions for predictions. To alleviate this problem, we explore the idea of adversarial regularization, and systematically compare four types of Generative Adversarial Networks (GANs), which are the combinations of GAN/Wasserstein GAN and its multi-scale version. Extensive experiments on a real-world radar data set and four typical meteorological examples are conducted. The results validate the effectiveness of adversarial regularization. The developed models show superior performances over the existing prediction approaches in the majority circumstances. Moreover, we find that the Wasserstein GAN regularization often delivers better results than the GAN regularization due to its robustness, and the Multi-scale Wasserstein GAN, in general, performs the best among all the methods. To reproduce the results, we release the source code at: https://github.com/luochuyao/MultiScaleGAN and the test system at: http://39.97.217.145:80/. Chuyao Luo, Xutao Li 0003, Yunming Ye, Shanshan Feng 0001, Michael Kwok-Po Ng |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2020 | Cascade SEIRD: Forecasting the Spread of COVID-19 with Dynamic Parameters UpdateabstractThe SEIR model is widely used in simulating the spread of infectious diseases. COVID-19 virus is a very severe infectious disease. Some studies leverage the SEIR or SEIRD model to simulate the spread and estimate the number of infected and recovered people as time goes on. However, these models suffer from two key deficiencies: (i) conventional SEIRD does not update its model parameters w.r.t. time; (ii) it focuses on predicting the trend, instead of the actual number of infections in the future. In this paper, we propose a cascade SEIRD model. The model learns and updates its parameters every day. Moreover, it is able to predict the number of infection cases, recovered cases and deaths. Specifically, we leverage a machine learning like approach to dynamically estimate the parameters of infection rate, incubation rate, recovery rate and death rate, which can be updated by gradient descent algorithm. Once the nature of the parameters w.r.t. time is determined, ARIMA model is adopted to characterize the dynamics of the parameters and predict their future changes. To validate the effectiveness of the proposed cascade SEIRD model, we conduct experiments on five data sets of different scales of regions (China, Hubei, Wuhan, Shenzhen, US). Experimental results show that the proposed cascade SEIRD achieves the most accurate prediction and outperforms state-of-the-art techniques. Yongliang Wen, Jiangnan Xu, Yunming Ye, Xutao Li 0003, Chuyao Luo, Tianlun Zhu |
BIBM | 5 |