Qian Li 0014

dblp:69/5902-14 · DBLP profile ↗
← Back
23ranked-venue papers
2as first author
16since 2021 · last 2025
0000-0002-9530-4925ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Text Prompted Spatiotemporal Sequence Prediction with Text-Vision Prompt Refiner and Masked Diffusion Transformers
abstract
Classical spatiotemporal sequence prediction tasks are designed to forecast future image sequences based on historical observations. However, the inherent unpredictability of future events often renders this process uncontrollable due to infinite possibilities in nature, limiting broader applicability of this technology. In this study, we explore the utilization of text prompts to constrain probabilistic space of future outcomes, resulting more controllable future prediction complying with user intent. We primarily address two critical challenges in this research setting: (i) text-vision misalignment, where embeddings extracted by text pre-trained models are not strictly aligned with visual embeddings, leading to predictions semantically irrelevant to text prompts. (ii) Spatiotemporal modeling distortion, where the fixed observation interval during training causes the model to produce unrealistic results when reasoning longer time dimensions. To tackle these issues, we propose a text-prompted spatiotemporal sequence prediction (TPS2P) model, leveraging historical observations and textual prompts to predict probabilistic future outcomes. In this model, a text-vision prompt refiner (TV-Refiner) is introduced to provide aligned textual and historical visual embeddings for integrating the denoising diffusion prediction process. Additionally, a spatiotemporal-masked diffusion transformer (StMDiT) is proposed by exploiting masked attention in constituting spatial and temporal self-attention modules within latent diffusion processes, enabling the model to observe more sequences of varying spatiotemporal patterns during training. We conduct extensive experiments on Something-Something V2 (Sthv2) and BridgeData datasets. Reported results demonstrate that our TPS2P predicts more accurate and high-quality future sequences, more user-intent compliant by textual controllability.
Yechao Xu, Zhengxing Sun, Qian Li 0014, Yunhan Sun
ACM Multimedia3
2025 A Deep Contrastive Model for Radar Echo Extrapolation
abstract
Weather radar echo extrapolation is one of the essential means for weather nowcasting. It has been considerably inspired over the last decade by deep learning. However, the internal similarity of the echo evolution process has little been exploited. To investigate this merit, a deep contrastive model with an encoder–projector structure is proposed in this letter, which projects the subsequences sampled from the same evolution process into the neighborhood of latent space by contrastive learning. Thus, the internal evolution similarity of the input echo sequence itself can be discovered and exploited for promoting prediction. To make the training smoother, we also adopt a cumulative sampling strategy that follows a simple-to-hard manner. Experimental results on two real-world radar datasets demonstrate the superiority of our model in comparison to state-of-the-art. The effectiveness of the sampling strategy and extrapolation ability on limited input is also analyzed and verified. Training code and pretrained models are available athttps://github.com/tolearnmuch/ESCL.
Qian Li 0014, Jinrui Jing, Leiming Ma, Shiqing Guo, Hanxing Chen, Tianying Wang, Yechao Xu
IEEE Geosci. Remote. Sens. Lett.1
2025 Heterogeneous Spatiotemporal Graph Learning for Localized Sparse Meteorological Forecasting
abstract
Localized meteorological forecasting holds significant practical value for human socio-economic activities. Current deep learning-based localized meteorological prediction methods trained with structured reanalysis data have made great progress. However, the assimilation process from observational data to reanalysis data may introduce uncertainty errors. Meanwhile, the inherent sparsity and unstructured nature of observational data fail to reflect localized sparse meteorological conditions accurately. Therefore, these methods struggle to deliver precise localized meteorological forecasts. To address these issues, we propose a Heterogeneous Spatio-Temporal Graph Neural Network (HSTGNN) that integrates observational and reanalysis data for in-situ prediction of meteorological variables at localized sparse stations. Specifically, we first design a Hybrid Graph Constructor (HGC) to map sparse observational data and reanalysis data into heterogeneous nodes, establishing node association topologies through a dual-channel mechanism combining geographical graphs and adaptive learning graphs. We then develop a Hierarchical Heterogeneous Message Passing Module (H2MPM) to facilitate differentiated information interaction between heterogeneous data, effectively capturing spatial dependencies across data sources. Finally, we introduce a temporal linear unit based on single-layer linear regression to extract temporal dependencies and perform sequence prediction of meteorological variables at each station. Experimental results on two datasets demonstrate that HSTGNN comprehensively models localized meteorological pattern and achieves outstanding performance in in-situ meteorological variables prediction tasks.
Qian Li 0014, Zhencai Du, Zeming Zhou
IEEE Trans. Geosci. Remote. Sens.2
2025 PMDNet: Polarimetric Multivariable Decoupling Network for Enhancing Nowcasting
Qian Li 0014, Tinger Hu
IEEE Trans. Geosci. Remote. Sens.2
2024 MPFNet: Multiproduct Fusion Network for Radar Echo Extrapolation
abstract
Radar echo extrapolation (REE) plays a crucial role in convective nowcasting. Existing deep learning (DL)-based methods for REE are predominantly based on the analysis of echo composite reflectivity (CR). However, CR product solely offers single-layered echo intensity information, thereby losing vertical details of convective systems such as echo top heights, resulting in lower accuracy in REE. To address these limitations, this article proposes a multiproduct fusion network (i.e., MPFNet) for REE. First, residual convolutional encoders (RCEs) are designed, which adopt the ResNet to reuse features and combine attention mechanisms to improve focus on convective features. In addition, to leverage the correlations and complementarities among multiproduct features, a multiproduct fusion module (MPFM) that adopts multihead attention for modeling the interrelations among multiproduct features and depthwise separable convolution (DSC) for feature fusion is proposed. Finally, a residual decoder (RD) is designed instead of a conventional deconvolution decoder to aggregate fused features for the restoration of predicted echo sequences. The proposed MPFNet is verified by convective nowcasting experiments, and the experimental results demonstrate that it can effectively utilize multiple radar products for guiding REE. It significantly outperforms the state-of-the-art (SOTA) methods, such as Earthformer and PreDiff. Compared to the previously best-performing Earthformer, MPFNet achieves an average improvement of 1.3% and 2.4% in critical success index (CSI) and Heidke skill score (HSS), respectively, in convective nowcasting experiments on the SWAN dataset, and an average improvement of 1.2% and 3.2% in CSI and HSS on the MeteoNet dataset.
Yanle Pei, Qian Li 0014, Nengli Sun, Jinrui Jing, Yuhong Ding, Tianying Wang
IEEE Trans. Geosci. Remote. Sens.2
2024 High-to-low-level feature matching and complementary information fusion for reference-based image super-resolution
Shuang Wang 0009, Zhengxing Sun, Qian Li 0014
Vis. Comput.3
2023 Image super-resolution based on self-similarity generative adversarial networks
abstract
Abstract Self‐attention has been successfully leveraged for long‐range feature‐wise similarities in deep learning super‐resolution (SR) methods. However, most of the SR methods only explore the features on the original scale, but do not take full advantage of self‐similarities features on different scales especially in generative adversarial networks (GAN). In this paper, self‐similarity generative adversarial networks (SSGAN) are proposed as the SR framework. The framework establishes the multi‐scale feature correlation by adding two modules to the generative network: downscale attention block (DAB) and upscale attention block (UAB). Specifically, DAB is designed to restore the repetitive details from the corresponding downsampled image, which achieves multi‐scale feature restoration through self‐similarity. And UAB improves the baseline up‐sampling operations and captures low‐resolution to high‐resolution feature mapping, which enhances the cross‐scale repetitive features to reconstruct the high‐resolution image. Experimental results demonstrate that the proposed SSGAN achieve better visual performance especially in the similar pattern details.
Shuang Wang 0009, Zhengxing Sun, Qian Li 0014
IET Image Process.3
2023 Fine-grained traffic video vehicle recognition based orientation estimation and temporal information
Anqi Hu, Zhengxing Sun, Qian Li 0014, Yechao Xu, Yihuan Zhu
Multim. Tools Appl.3
2022 Learning Semantic Segmentation on Unlabeled Real-World Indoor Point Clouds via Synthetic Data
abstract
The data-hungry nature of deep learning and the high cost of annotating point-level labels for point clouds make it difficult to apply semantic segmentation methods to unlabeled real-world indoor scenes. Therefore, label-efficient point cloud segmentation has become a promising research topic. We noticed that the online housing design platforms can provide a large number of synthetic indoor 3D scenes, which are created with semantic labels. In this paper, we propose to learn semantic segmentation on synthetic point clouds and adapt the model for unlabeled real-world data. The main challenge is that directly using models trained on synthetic data for real-world data produces poor results due to the large domain gap between synthetic and real-world data. We design a point cloud style transfer network and a feature discrimination network to reduce the domain gap in both the input space and the feature space. Experiments show that our approach significantly improves the performance on real-world data for models learned from synthetic data.
Youcheng Song, Zhengxing Sun, Yunjie Wu, Yunhan Sun, Shoutong Luo, Qian Li 0014
ICPR6
2022 Weakly Supervised Fine-grained Recognition based on Combined Learning for Small Data and Coarse Label
abstract
Learning with weak supervision already becomes one of the research trends in fine-grained image recognition. These methods aim to learn feature representation in the case of less manual cost or expert knowledge. Most existing weakly supervised methods are based on incomplete annotation or inexact annotation, which is difficult to perform well limited by supervision information. Therefore, using these two kind of annotations for training at the same time could mine more relevance while the annotating burden will not increase much. In this paper, we propose a combined learning framework by coarse-grained large data and fine-grained small data for weakly supervised fine-grained recognition. Combined learning contains two significant modules: 1) a discriminant module, which maintains the structure information consistent between coarse label and fine label by attention map and part sampling, 2) a cluster division strategy, which mines the detail differences between fine categories by feature subtraction. Experiment results show that our method outperforms weakly supervised methods and achieves the performance close to fully supervised methods in CUB-200-2011 and Stanford Cars datasets.
Anqi Hu, Zhengxing Sun, Qian Li 0014
ICMR3
2022 Active Patterns Perceived for Stochastic Video Prediction
abstract
Predicting future scenes based on historical frames is challenging, especially when it comes to the complex uncertainty in nature. We observe that there is a divergence between spatial-temporal variations of active patterns and non-active patterns in a video, where these patterns constitute visual content and the former ones implicate more violent movement. This divergence enables active patterns the higher potential to act with more severe future uncertainty. Meanwhile, the existence of non-active patterns provides an opportunity for machines to examine some underlying rules with a mutual constraint between non-active patterns and active patterns. In order to solve this divergence, we provide a method called active patterns-perceived stochastic video prediction (ASVP) which allows active patterns to be perceived by neural networks during training. Our method starts with separating active patterns along with non-active ones from a video. Then, both scene-based prediction and active pattern-perceived prediction are conducted to respectively capture the variations within the whole scene and active patterns. Specially for active pattern-perceived prediction, a conditional generative adversarial network (CGAN) is exploited to model active patterns as conditions, with a variational autoencoder (VAE) for predicting the complex dynamics of active patterns. Additionally, a mutual constraint is designed to improve the learning procedure for the network to better understand underlying interacting rules among these patterns. Extensive experiments are conducted on both KTH human action and BAIR action-free robot pushing datasets with comparison to state-of-the-art works. Experimental results demonstrate the competitive performance of the proposed method as we expected. The released code and models are at https://github.com/tolearnmuch/ASVP.
Yechao Xu, Zhengxing Sun, Qian Li 0014, Yunhan Sun, Shoutong Luo
ACM Multimedia3
2022 NLED: Nonlocal Echo Dynamics Network for Radar Echo Extrapolation
abstract
Radar echo extrapolation is a common approach to weather nowcasting, which has become a significant support to detect potential disastrous weather a few hours ahead. The dynamics pattern inside an echo intensity sequence is beneficial for echo prediction. However, existing extrapolation methods have limited ability to consider entire time series echo context in an entire echo sequence from a given historical timestamp to a given future timestamp, leading to low long-term extrapolation accuracy. To solve this issue, we introduce spatiotemporal self-attention and propose a deep learning model named the nonlocal echo dynamics (NLED) network to capture the dependencies of the entire time domain. The NLED network has an encoder-decoder architecture for extrapolation. The encoder decomposes historical echoes into features of multiple spatial scales, which makes it better at learning echo dynamics from the global scale to the local scale. The decoder employs nonlocal blocks with sparse self-attention related to echo dynamics to learn correlations in the entire echo event, which is beneficial for predicting long-term echo distributions. Our model is evaluated on radar reflectivity datasets from Shanghai and Hong Kong. The experimental results indicate that the NLED model achieves a more accurate long-term forecast, and alleviates the forgetting of stronger echo dynamics, validating the effectiveness of entire time series modeling by the NLED network.
Taisong Chen, Qian Li 0014, Jinrui Jing
IEEE Trans. Geosci. Remote. Sens.2
2022 REMNet: Recurrent Evolution Memory-Aware Network for Accurate Long-Term Weather Radar Echo Extrapolation
abstract
Weather radar echo extrapolation, which predicts future echoes based on historical observations, is one of the complicated spatial–temporal sequence prediction tasks and plays a prominent role in severe convection and precipitation nowcasting. However, existing extrapolation methods mainly focus on a defective echo-motion extrapolation paradigm based on finite observational dynamics, neglecting that the actual echo sequence has a more complicated evolution process that contains both nonlinear motions and the lifecycle from initiation to decay, resulting in poor prediction precision and limited application ability. To complement this paradigm, we propose to incorporate a novel long-term evolution regularity memory (LERM) module into the network, which can memorize long-term echo-evolution regularities during training and be recalled for guiding extrapolation. Moreover, to resolve the blurry prediction problem and improve forecast accuracy, we also adopt a coarse–fine hierarchical extrapolation strategy and compositive loss function. We separate the extrapolation task into coarse and fine two levels which can reduce the downsampling loss and retain echo fine details. Except for the average reconstruction loss, we additionally employ adversarial loss and perceptual similarity loss to further improve the visual quality. Experimental results from two real radar echo datasets demonstrate the effectiveness of our methodology and show that it can accurately extrapolate the echo evolution while ensuring the echo details are realistic enough, even for the long term. Our method can further be improved in the future by integrating multimodal radar variables or introducing certain domain prior knowledge of physical mechanisms. It can also be applied to other spatial–temporal sequence prediction tasks, such as the prediction of satellite cloud images and wind field figures.
Jinrui Jing, Qian Li 0014, Leiming Ma, Lei Ding 0008
IEEE Trans. Geosci. Remote. Sens.2
2022 CNGAT: A Graph Neural Network Model for Radar Quantitative Precipitation Estimation
abstract
Radar quantitative precipitation estimation (RQPE) is the most common measurement for area rainfall estimation with high spatial and temporal resolution. The radar reflectivity ($Z$) measured by the Doppler weather radar is strongly related to precipitation rate ($R$). However, conventional RQPE methods have limited capability of modeling the complex relationship between the radar echoes and the precipitation field. In this article, we propose a graph neural network (GNN)-based RQPE model named categorical node graph attention network (CNGAT) to model complex spatial–temporal features of precipitation field reflected by the radar echo field. CNGAT is derived from graph attention networks (GAT) utilizing attention mechanism to learn the importance of neighboring points to the central point, which is beneficial for learning varying local spatial patterns. Furthermore, CNGAT can handle multiple types of graph nodes by using different transform functions for different types of nodes, which makes it better at capturing diverse features of precipitation field indicated by strong and weak radar echo areas. The proposed model was trained and tested on radar network and rain gauge data distributed in East China during 2017 and 2018. The results of several experiments show that CNGAT greatly improves the estimation precision and detection rate than$Z$–$R$relation models and conventional data-driven RQPE methods, and alleviates under-estimation of higher precipitation rates, which validates that CNGAT can effectively represent complex spatial–temporal features of precipitation field.
Qian Li 0014, Jinrui Jing
IEEE Trans. Geosci. Remote. Sens.2
2022 Learning indoor point cloud semantic segmentation from image-level labels
Youcheng Song, Zhengxing Sun, Qian Li 0014, Yunjie Wu, Yunhan Sun, Shoutong Luo
Vis. Comput.3
2021 AdaAttN: Revisit Attention Mechanism in Arbitrary Neural Style Transfer
abstract
Fast arbitrary neural style transfer has attracted widespread attention from academic, industrial and art communities due to its flexibility in enabling various applications. Existing solutions either attentively fuse deep style feature into deep content feature without considering feature distributions, or adaptively normalize deep content feature according to the style such that their global statistics are matched. Although effective, leaving shallow feature unexplored and without locally considering feature statistics, they are prone to unnatural output with unpleasing local distortions. To alleviate this problem, in this paper, we propose a novel attention and normalization module, named Adaptive Attention Normalization (AdaAttN), to adaptively perform attentive normalization on per-point basis. Specifically, spatial attention score is learnt from both shallow and deep features of content and style images. Then perpoint weighted statistics are calculated by regarding a style feature point as a distribution of attention-weighted output of all style feature points. Finally, the content feature is normalized so that they demonstrate the same local feature statistics as the calculated per-point weighted style feature statistics. Besides, a novel local feature loss is derived based on AdaAttN to enhance local visual quality. We also extend AdaAttN to be ready for video style transfer with slight modifications. Experiments demonstrate that our method achieves state-of-the-art arbitrary image/video style transfer. Codes and models are available on https://github.com/wzmsltw/AdaAttN.
Songhua Liu, Dongliang He, Fu Li 0003, Xin Li 0106, Zhengxing Sun, Qian Li 0014, Errui Ding
ICCV8
2020 HPRNN: A Hierarchical Sequence Prediction Model for Long-Term Weather Radar Echo Extrapolation
abstract
Weather radar echo extrapolation has been one of the most important means for weather forecasting and precipitation nowcasting. However, the effective forecasting time of the most current extrapolation methods is usually short. In this paper, to meet the demand for long-term extrapolation in actual forecasting practice, we propose a hierarchical prediction recurrent neural network (HPRNN) for long-term radar echo extrapolation. HPRNN is composed of hierarchically stacked RNN modules and a refinement module, it employs both a hierarchical prediction strategy and a recurrent coarse-to-fine mechanism to alleviate the accumulation of prediction error with time and contribute to making long-term extrapolation. The extrapolation experiments conducted on the HKO-7 radar echo dataset demonstrate the effectiveness of our model.
Jinrui Jing, Qian Li 0014, Shaoen Tang
ICASSP2
2019 Group-Wise Deep Object Co-Segmentation With Co-Attention Recurrent Neural Network
abstract
Effective feature representations which should not only express the images individual properties, but also reflect the interaction among group images are essentially crucial for real-world co-segmentation. This paper proposes a novel end-to-end deep learning approach for group-wise object co-segmentation with a recurrent network architecture. Specifically, the semantic features extracted from a pre-trained CNN of each image are first processed by single image representation branch to learn the unique properties. Meanwhile, a specially designed Co-Attention Recurrent Unit (CARU) recurrently explores all images to generate the final group representation by using the co-attention between images, and simultaneously suppresses noisy information. The group feature which contains synergetic information is broadcasted to each individual image and fused with multi-scale fine-resolution features to facilitate the inferring of co-segmentation. Moreover, we propose a groupwise training objective to utilize the co-object similarity and figure-ground distinctness as the additional supervision. The whole modules are collaboratively optimized in an end-to-end manner, further improving the robustness of the approach. Comprehensive experiments on three benchmarks can demonstrate the superiority of our approach in comparison with the state-of-the-art methods.
Zhengxing Sun, Qian Li 0014, Yunjie Wu, Anqi Hu
ICCV3
2019 Co-saliency Detection Based on Hierarchical Consistency
abstract
As an interesting and emerging topic, co-saliency detection aims at discovering common and salient objects in a group of related images, which is useful to variety of visual media applications. Although a number of approaches have been proposed to address this problem, many of them are designed with the misleading assumption, suboptimal image representation, or heavy supervision cost and thus still suffer from certain limitations, which reduces their capability in the real-world scenarios. To alleviate these limitations, we propose a novel unsupervised co-saliency detection method, which successively explores the hierarchical consistency in the image group including background consistency, high-level and low-level objects consistency in a unified framework. We first design a novel superpixel-wise variational autoencoder (SVAE) network to precisely distinguish the salient objects from the background collection based on the reconstruction errors. Then, we propose a two-stage clustering strategy to explore the multi-level salient objects consistency by using high-level and low-level features separately. Finally, the co-saliency results are refined by applying a CRF based refinement method with the multi-level salient objects consistency. Extensive experiments on three widely datasets show that our method achieves superior or competitive performance compared to the state-of-the-art methods.
Zhengxing Sun, Qian Li 0014
ACM Multimedia4
2019 Direction-aware neural style transfer with texture enhancement
Zhengxing Sun, Yan Zhang 0007, Qian Li 0014
Neurocomputing4
2018 A Method of Weather Radar Echo Extrapolation Based on Convolutional Neural Networks
En Shi, Qian Li 0014, Daquan Gu, Zhangming Zhao
MMM (1)2
2016 Relevance feedback for human motion retrieval using a boosting approach
Song-Le Chen, Zhengxing Sun, Yan Zhang 0007, Qian Li 0014
Multim. Tools Appl.4
2016 Dynamic node selection in camera networks based on approximate reinforcement learning
Qian Li 0014, Zhengxing Sun, Song-Le Chen, Shi-ming Xia
Multim. Tools Appl.1