VLDB 2026 Research / reviewers in the wild / expert
Liang Xie 0001
dblp:81/2806-1
· DBLP profile ↗
37ranked-venue papers
11as first author
14since 2021 · last 2026
0000-0003-1718-7556ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 10 first-author · 1 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LOSTFormer: Linear Orthogonal Spatio-Temporal Transformer with Learnable Rotation
Yechen Yu, Liang Xie 0001, Jiankai Zheng, Peilin Tan |
DASFAA (4) | 2 |
| 2026 | GHOST: Sentiment-gated mamba and stock-wise tokenization attention for enhanced stock prediction
Liang Xie 0001, Hanyu Fu, Jiangling Zhang |
Expert Syst. Appl. | 2 |
| 2026 | LiFedST: A linearized federated split-attention transformer for spatio-temporal forecasting
Chengjie Ying, Zhiqiang Ru, Yuan Wan, Liang Xie 0001 |
Knowl. Based Syst. | 5 |
| 2026 | LLM-Augmented Stiefel Graph Neural Networks for Zero-Shot Spatio-Temporal ForecastingabstractSpatio-temporal time series (STTS) play a crucial role in domains such as traffic forecasting, energy scheduling, and financial analysis. However, accurate and efficient prediction remains challenging due to the complex dynamic dependencies across temporal and spatial dimensions. Existing Graph Neural Networks (GNNs) often struggle to balance effectiveness and efficiency when modeling dynamic spatio-temporal relations. Meanwhile, Large Language Models (LLMs) exhibit strong capabilities in modeling long-range dependencies and generalization under few-shot and zero-shot conditions, yet their ability to capture spatio-temporal structures remains limited. To address this, we propose LAD-SGNN, a unified framework that integrates the advantages of graph-based and language-based modeling. Specifically, LAD-SGNN employs Spectral Graph Convolution on the Stiefel manifold (SGSC) together with Linear Dynamic Graph Optimization (LDGOSM) to efficiently extract dynamic spatio-temporal features with nearly linear complexity. Structured prompts and a spatio-temporal alignment mechanism are then designed to fuse the extracted dynamic spectral information with task semantics, which are fed into a lightweight LLM to achieve unified modeling of spatio-temporal structures and semantic reasoning. In this framework, SGSC efficiently represents dynamic spatial dependencies in the spectral domain, while the prompt-based alignment strategy explicitly injects spatio-temporal information into the LLM, thereby enhancing generalization across datasets. Extensive experiments on multiple real-world spatio-temporal datasets demonstrate that LAD-SGNN significantly improves prediction accuracy in both zero-shot and few-shot scenarios. Jiankai Zheng, Liang Xie 0001, Lei Zhu 0002, Guoli Yang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | Hard-Token Spatiotemporal Pre-Training
Mingxu Jiang, Yiyan Xu, Haoyang Hu, Liang Xie 0001 |
IEEE Big Data | 5 |
| 2025 | A Dynamic Stiefel Graph Neural Network for Efficient Spatio-Temporal Time Series ForecastingabstractSpatio-temporal time series (STTS) have been widely used in many applications. However, accurately forecasting STTS is challenging due to complex dynamic correlations in both time and space dimensions. Existing graph neural networks struggle to balance effectiveness and efficiency in modeling dynamic spatio-temporal relations. To address this problem, we propose the Dynamic Spatio-Temporal Stiefel Graph Neural Network (DST-SGNN) to efficiently process STTS. For DST-SGNN, we first introduce the novel Stiefel Graph Spectral Convolution (SGSC) and Stiefel Graph Fourier Transform (SGFT). The SGFT matrix in SGSC is constrained to lie on the Stiefel manifold, and SGSC can be regarded as a filtered graph spectral convolution. We also propose the Linear Dynamic Graph Optimization on Stiefel Manifold (LDGOSM), which can efficiently learn the SGFT matrix from the dynamic graph and significantly reduce the computational complexity. Finally, we propose a multi-layer SGSC (MSGSC) that efficiently captures complex spatio-temporal correlations. Extensive experiments on seven spatio-temporal datasets show that DST-SGNN outperforms state-of-the-art methods while maintaining relatively low computational costs. Jiankai Zheng, Liang Xie 0001 |
IJCAI | 2 |
| 2025 | Deep reinforcement learning portfolio model based on mixture of experts
Ziqiang Wei, Deng Chen, Yanduo Zhang, Dawei Wen, Liang Xie 0001 |
Appl. Intell. | 6 |
| 2025 | Transformer based lifetime interval prediction for dynamic operating proton exchange membrane fuel cells
Liang Xie 0001, Dongqi Zhao, Ze Zhou 0001, Liyan Zhang 0003, Qihong Chen |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Multi-resolution Patch-based Fourier Graph Spectral Network for spatiotemporal time series forecasting
Jiankai Zheng, Liang Xie 0001, Haijiao Xu |
Neurocomputing | 2 |
| 2025 | DSTG: Distillation Swin Transformer for Cross-View GeolocalizationabstractIn cross-view image geolocation tasks, traditional CNN models are limited in performance due to their inability to effectively capture global correlations. Although Transformer-based methods can address this shortcoming, their computational complexity and GPU memory consumption are relatively high. To address these limitations, we propose Swin Transformer for Cross-view Geo-localization (STG), a novel model based on the Swin Transformer (ST) that leverages its advantage of linear computational complexity scaling with image resolution. We design an adaptive window shift mechanism and introduce a pre-training strategy for STG, which can enhance STGs global modeling capability and its ability to learn general features. To further address the redundancy issues in our STG model, we propose the Distillation Swin Transformer for Cross-view Geo-localization model (DSTG), which is obtained through knowledge distillation using the STG model as the teacher model. To tackle the issue that the student model strugglings to learn from the teacher model’s output during distillation directly, we propose a multi-scale logit standardization distillation method. The proposed STG and DSTG models can obtain superior global embedding descriptions without relying on polar coordinate transformations. Experimental results on the CVUSA, CVACT_val, and CVACT_test datasets indicate that the DSTG model has significantly lower computational cost and GPU memory usage compared to the STG model. At the same time, both models achieve state-of-the-art performance. The source code is available at https://github.com/liangjxiong/DSTG. Junxiong Liang, Mengwei Bao, Liang Xie 0001, Ryan Wen Liu, Nengcheng Chen |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | DTSMLA: A dynamic task scheduling multi-level attention model for stock ranking
Yuanchuang Du, Liang Xie 0001, Sihao Liao, Shengshuang Chen, Haijiao Xu |
Expert Syst. Appl. | 2 |
| 2023 | Stock trend prediction based on industry relationships driven hypergraph attention networks
Haodong Han, Liang Xie 0001, Shengshuang Chen, Haijiao Xu |
Appl. Intell. | 2 |
| 2021 | Multi-modal discrete tensor decomposition hashing for efficient multimedia retrieval
Xize Wu, Lei Zhu 0002, Liang Xie 0001, Zheng Zhang 0006, Huaxiang Zhang 0001 |
Neurocomputing | 3 |
| 2021 | Label propagation with structured graph learning for semi-supervised dimension reduction
Fei Wang 0055, Lei Zhu 0002, Liang Xie 0001, Zheng Zhang 0006, Mingyang Zhong |
Knowl. Based Syst. | 3 |
| 2019 | Discrete semi-supervised learning for multi-label image classification and large-scale image retrieval
Liang Xie 0001, Haohao Shu, Shengyuan Hu 0002 |
Multim. Tools Appl. | 2 |
| 2019 | Color and depth image registration algorithm based on multi-vector-fields constraints
Daoqing Li, Li Peng 0003, Huabing Zhou, Deng Chen, Yanduo Zhang, Liang Xie 0001 |
Multim. Tools Appl. | 7 |
| 2019 | Large scale image retrieval with DCNN and local geometrical constraint model
Huabing Zhou, Yiwei Tao, Jinshu Shi, Deng Chen, Yanduo Zhang, Liang Xie 0001 |
Multim. Tools Appl. | 7 |
| 2018 | Large-scale semantic web image retrieval using bimodal deep learning techniques
Changqin Huang, Haijiao Xu, Liang Xie 0001, Jia Zhu 0003, Chunyan Xu, Yong Tang 0001 |
Inf. Sci. | 3 |
| 2018 | Exploring Auxiliary Context: Discrete Semantic Transfer Hashing for Scalable Image RetrievalabstractUnsupervised hashing can desirably support scalable content-based image retrieval for its appealing advantages of semantic label independence, memory, and search efficiency. However, the learned hash codes are embedded with limited discriminative semantics due to the intrinsic limitation of image representation. To address the problem, in this paper, we propose a novel hashing approach, dubbed as discrete semantic transfer hashing (DSTH). The key idea is to directly augment the semantics of discrete image hash codes by exploring auxiliary contextual modalities. To this end, a unified hashing framework is formulated to simultaneously preserve visual similarities of images and perform semantic transfer from contextual modalities. Furthermore, to guarantee direct semantic transfer and avoid information loss, we explicitly impose the discrete constraint, bit-uncorrelation constraint, and bit-balance constraint on hash codes. A novel and effective discrete optimization method based on augmented Lagrangian multiplier is developed to iteratively solve the optimization problem. The whole learning process has linear computation complexity and desirable scalability. Experiments on three benchmark data sets demonstrate the superiority of DSTH compared with several state-of-the-art approaches. Lei Zhu 0002, Zi Huang, Zhihui Li 0001, Liang Xie 0001, Heng Tao Shen |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Dynamic Multi-View Hashing for Online Image RetrievalabstractAdvanced hashing technique is essential to facilitate effective large scale online image organization and retrieval, where image contents could be frequently changed. Traditional multi-view hashing methods are developed based on batch-based learning, which leads to very expensive updating cost. Meanwhile, existing online hashing methods mainly focus on single-view data and thus can not achieve promising performance when searching real online images, which are multiple view based data. Further, both types of hashing methods can only produce hash code with fixed length. Consequently they suffer from limited capability to comprehensive characterization of streaming image data in the real world. In this paper, we propose dynamic multi-view hashing (DMVH), which can adaptively augment hash codes according to dynamic changes of image. Meanwhile, DMVH leverages online learning to generate hash codes. It can increase the code length when current code is not able to represent new images effectively. Moreover, to gain further improvement on overall performance, each view is assigned with a weight, which can be efficiently updated during the online learning process. In order to avoid the frequent updating of code length and view weights, an intelligent buffering scheme is also specifically designed to preserve significant data to maintain good effectiveness of DMVH. Experimental results on two real-world image datasets demonstrate superior performance of DWVH over several state-of-the-art hashing methods. Liang Xie 0001, Jialie Shen 0001, Jungong Han, Lei Zhu 0002, Ling Shao 0001 |
IJCAI | 1 |
| 2017 | Multi-Task Multi-modal Semantic Hashing for Web Image Retrieval with Limited Supervision
Liang Xie 0001, Lei Zhu 0002, Zhiyong Cheng 0001 |
MMM (1) | 1 |
| 2017 | Unsupervised Topic Hypergraph Hashing for Efficient Mobile Image RetrievalabstractHashing compresses high-dimensional features into compact binary codes. It is one of the promising techniques to support efficient mobile image retrieval, due to its low data transmission cost and fast retrieval response. However, most of existing hashing strategies simply rely on low-level features. Thus, they may generate hashing codes with limited discriminative capability. Moreover, many of them fail to exploit complex and high-order semantic correlations that inherently exist among images. Motivated by these observations, we propose a novel unsupervised hashing scheme, called topic hypergraph hashing (THH), to address the limitations. THH effectively mitigates the semantic shortage of hashing codes by exploiting auxiliary texts around images. In our method, relations between images and semantic topics are first discovered via robust collective non-negative matrix factorization. Afterwards, a unified topic hypergraph, where images and topics are represented with independent vertices and hyperedges, respectively, is constructed to model inherent high-order semantic correlations of images. Finally, hashing codes and functions are learned by simultaneously enforcing semantic consistence and preserving the discovered semantic relations. Experiments on publicly available datasets demonstrate that THH can achieve superior performance compared with several state-of-the-art methods, and it is more suitable for mobile image retrieval. Lei Zhu 0002, Jialie Shen 0001, Liang Xie 0001, Zhiyong Cheng 0001 |
IEEE Trans. Cybern. | 3 |
| 2017 | Unsupervised Visual Hashing with Semantic Assistant for Content-Based Image RetrievalabstractAs an emerging technology to support scalable content-based image retrieval (CBIR), hashing has recently received great attention and became a very active research domain. In this study, we propose a novel unsupervised visual hashing approach called semantic-assisted visual hashing (SAVH). Distinguished from semi-supervised and supervised visual hashing, its core idea is to effectively extract the rich semantics latently embedded in auxiliary texts of images to boost the effectiveness of visual hashing without any explicit semantic labels. To achieve the target, a unified unsupervised framework is developed to learn hash codes by simultaneously preserving visual similarities of images, integrating the semantic assistance from auxiliary texts on modeling high-order relationships of inter-images, and characterizing the correlations between images and shared topics. Our performance study on three publicly available image collections: Wiki, MIR Flickr, and NUS-WIDE indicates that SAVH can achieve superior performance over several state-of-the-art techniques. Lei Zhu 0002, Jialie Shen 0001, Liang Xie 0001, Zhiyong Cheng 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2016 | Online Cross-Modal Hashing for Web Image RetrievalabstractCross-modal hashing (CMH) is an efficient technique for the fast retrieval of web image data, and it has gained a lot of attentions recently. However, traditional CMH methods usually apply batch learning for generating hash functions and codes. They are inefficient for the retrieval of web images which usually have streaming fashion. Online learning can be exploited for CMH. But existing online hashing methods still cannot solve two essential problems: efficient updating of hash codes and analysis of cross-modal correlation. In this paper, we propose Online Cross-modal Hashing (OCMH) which can effectively address the above two problems by learning the shared latent codes (SLC). In OCMH, hash codes can be represented by the permanent SLC and dynamic transfer matrix. Therefore, inefficient updating of hash codes is transformed to the efficient updating of SLC and transfer matrix, and the time complexity is irrelevant to the database size. Moreover, SLC is shared by all the modalities, and thus it can encode the latent cross-modal correlation, which further improves the overall cross-modal correlation between heterogeneous data. Experimental results on two real-world multi-modal web image datasets: MIR Flickr and NUS-WIDE, demonstrate the effectiveness and efficiency of OCMH for online cross-modal web image retrieval. Liang Xie 0001, Jialie Shen 0001, Lei Zhu 0002 |
AAAI | 1 |
| 2016 | Learning Compact Visual Representation with Canonical Views for Robust Mobile Landmark Search
Lei Zhu 0002, Jialie Shen 0001, Xiaobai Liu, Liang Xie 0001, Liqiang Nie |
IJCAI | 4 |
| 2016 | Unsupervised multi-graph cross-modal hashing for large-scale multimedia retrieval
Liang Xie 0001, Lei Zhu 0002, Guoqi Chen |
Multim. Tools Appl. | 1 |
| 2016 | Cross-Modal Self-Taught Hashing for large-scale image retrieval
Liang Xie 0001, Lei Zhu 0002, Peng Pan 0001, Yansheng Lu |
Signal Process. | 1 |
| 2015 | Topic Hypergraph Hashing for Mobile Image RetrievalabstractHashing is one of the promising solutions to support efficient Mobile Image Retrieval (MIR). However, most of existing hashing strategies simply rely on low-level features, which inevitably makes the generated hashing codes less semantic. Moreover, many of them fail to exploit complex and high-order semantic correlations of images. Motivated by these observations, we propose a novel unsupervised hashing scheme, \emph{Topic Hypergraph Hashing} (THH), to address the limitations. A unified topic hypergraph, where images and topics are represented with independent vertices and hyperedges respectively, is first constructed to model latent semantics of images and their correlations. With topic hypergraph model, hashing codes and functions are then learned by simultaneously preserving similarity consistence and semantic correlation. Experiments on standard datasets demonstrate that THH can achieve superior performance compared with several state-of-the-art techniques, and it is more suitable for MIR. Lei Zhu 0002, Jialie Shen 0001, Liang Xie 0001 |
ACM Multimedia | 3 |
| 2015 | Cross-Modal Self-Taught Learning for Image Retrieval
Liang Xie 0001, Peng Pan 0001, Yansheng Lu |
MMM (1) | 1 |
| 2015 | Analyzing semantic correlation for cross-modal retrieval
Liang Xie 0001, Peng Pan 0001, Yansheng Lu |
Multim. Syst. | 1 |
| 2015 | Improving cross-modal and multi-modal retrieval combining content and semantics similarities with probabilistic model
Shixun Wang, Peng Pan 0001, Yansheng Lu, Liang Xie 0001 |
Multim. Tools Appl. | 4 |
| 2015 | Markov random field based fusion for supervised and semi-supervised multi-modal image classification
Liang Xie 0001, Peng Pan 0001, Yansheng Lu |
Multim. Tools Appl. | 1 |
| 2015 | Content-Based Visual Landmark Search via Multimodal Hypergraph LearningabstractWhile content-based landmark image search has recently received a lot of attention and became a very active domain, it still remains a challenging problem. Among the various reasons, high diverse visual content is the most significant one. It is common that for the same landmark, images with a wide range of visual appearances can be found from different sources and different landmarks may share very similar sets of images. As a consequence, it is very hard to accurately estimate the similarities between the landmarks purely based on single type of visual feature. Moreover, the relationships between landmark images can be very complex and how to develop an effective modeling scheme to characterize the associations still remains an open question. Motivated by these concerns, we propose multimodal hypergraph (MMHG) to characterize the complex associations between landmark images. In MMHG, images are modeled as independent vertices and hyperedges contain several vertices corresponding to particular views. Multiple hypergraphs are firstly constructed independently based on different visual modalities to describe the hidden high-order relations from different aspects. Then, they are integrated together to involve discriminative information from heterogeneous sources. We also propose a novel content-based visual landmark search system based on MMHG to facilitate effective search. Distinguished from the existing approaches, we design a unified computational module to support query-specific combination weight learning. An extensive experiment study on a large-scale test collection demonstrates the effectiveness of our scheme over state-of-the-art approaches. Lei Zhu 0002, Jialie Shen 0001, Hai Jin 0001, Liang Xie 0001 |
IEEE Trans. Cybern. | 5 |
| 2015 | Landmark Classification With Hierarchical Multi-Modal Exemplar FeatureabstractLandmark image classification attracts increasing research attention due to its great importance in real applications, ranging from travel guide recommendation to 3-D modelling and visualization of geolocation. While large amount of efforts have been invested, it still remains unsolved by academia and industry. One of the key reasons is the large intra-class variance rooted from the diverse visual appearance of landmark images. Distinguished from most existing methods based on scalable image search, we approach the problem from a new perspective and model landmark classification as multi-modal categorization , which enjoys advantages of low storage overhead and high classification efficiency. Toward this goal, a novel and effective feature representation, called hierarchical multi-modal exemplar (HMME) feature, is proposed to characterize landmark images. In order to compute HMME, training images are first partitioned into the regions with hierarchical grids to generate candidate images and regions. Then, at the stage of exemplar selection, hierarchical discriminative exemplars in multiple modalities are discovered automatically via iterative boosting and latent region label mining. Finally, HMME is generated via a region-based locality-constrained linear coding (RLLC), which effectively encodes semantics of the discovered exemplars into HMME. Meanwhile, dimension reduction is applied to reduce redundant information by projecting the raw HMME into lower-dimensional space. The final HMME enjoys advantages of discriminative and linearly separable. Experimental study has been carried out on real world landmark datasets, and the results demonstrate the superior performance of the proposed approach over several state-of-the-art techniques. Lei Zhu 0002, Jialie Shen 0001, Hai Jin 0001, Liang Xie 0001 |
IEEE Trans. Multim. | 4 |
| 2014 | A Cross-modal Multi-task Learning Framework for Image AnnotationabstractWith the advance of internet, multi-modal data can be easily collected from many social websites such as Wikipedia, Flickr, YouTube, etc. Images shared on the web are usually associated with social tags or other textual information. Although existing multi-modal methods can make use of associated text to improve image annotation, the disadvantages of them are that associated text is also required for a new image to be predicted. In this paper, we propose the cross-modal multi-task learning (CMMTL) framework for image annotation. Labeled and unlabeled multi-modal data are both levaraged for training in CMMTL, and it finally obtains visual classifiers which can predict concepts for a single image without any associated information. CMMTL integrates graph learning, multi-task learning and cross-modal learning into a joint framework, where a shared subspace is learned to preserve both cross-modal correlation and concept correlation. The optimal solution of the proposed framework can be obtained by solving a generalized eigenvalue problem. We conduct comprehensive experiments on two real world image datasets: MIR Flickr and NUS-WIDE, to evaluate the performance of the proposed framework. Experimental results demonstrate that CMMTL obtains a significant improvement over several representative methods for cross-modal image annotation. Liang Xie 0001, Peng Pan 0001, Yansheng Lu, Shixun Wang |
CIKM | 1 |
| 2013 | A Two-Phase Generation Model for Automatic Image AnnotationabstractAutomatic image annotation is an important task for multimedia retrieval. By allocating relevant words to un-annotated images, these images can be retrieved in response to textual queries. There are many researches on the problem of image annotation and most of them construct models based on joint probability or posterior probabilities of words. In this paper we estimate the probabilities that words generate the images, and propose a two-phase generation model for the generation procedure. Each word first generates its related words, then these words generate an un-annotated image, and the relation between the words and the un-annotated image is obtained by the probability of the two-phase generation. The textual words usually contain more semantic information than visual content of images, thus the probabilities that words generate images is more reliable than the probability that images generate words. As a result, our model estimates the more reliable probability than other probabilistic methods for image annotation. The other advantage of our model is the relation of words is taken into consideration. The experimental results on Corel 5K and MIR Flickr demonstrate that our model performs better than other previous methods. And two-phase generation which considering word's relation for annotation is better than one-phase generation which only consider the relation between words and images. Moreover, the methods which estimate the generative probability obtain better performance than SVM which estimates the posterior probability. Liang Xie 0001, Peng Pan 0001, Yansheng Lu, Shixun Wang, Haijiao Xu, Deng Chen |
ISM | 1 |
| 2013 | A semantic model for cross-modal and multi-modal retrievalabstractIn this paper, a semantic model for cross-modal and multi-modal retrieval is studied. We assume that the semantic correlation of multimedia data from different modalities can be depicted in a probabilistic generation framework. Media data from different modalities can be generated by the same semantic concepts, and the generation process of each media data is conditional independent under the semantic concepts. The semantic generation model (SGM) for cross-modal and multi-modal analysis is proposed based on this assumption. We study two types of methods: direct method Gaussian distribution and indirect method random forest, to estimate the semantic conditional distribution of SGM. Then methods for cross-modal and multi-modal retrieval are derived from SGM. Experimental results show that SGM based methods for cross-modal retrieval improve the accuracy over the state-of-the-art cross-modal method, but don't increase the time consuming, and the SGM multimodal retrieval methods also outperform traditional methods in image retrieval. Moreover, indirect SGM based method outperforms direct SGM method in the two types of retrieval, which proves that indirect SGM can better describe the semantic distribution. Liang Xie 0001, Peng Pan 0001, Yansheng Lu |
ICMR | 1 |