VLDB 2026 Research / reviewers in the wild / expert
Xiang Ma 0006
dblp:52/7004-6
· DBLP profile ↗
19ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0002-4963-8705ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReCast: Reliability-aware Codebook-assisted Lightweight Time Series ForecastingabstractTime series forecasting is crucial for applications in various domains. Conventional methods often rely on global decomposition into trend, seasonal, and residual components, which become ineffective for real-world series dominated by local, complex, and highly dynamic patterns. Moreover, the high model complexity of such approaches limits their applicability in real-time or resource-constrained environments. In this work, we propose a novel reliability-aware codebook-assisted time series forecasting framework (ReCast) that enables lightweight and robust prediction by exploiting recurring local shapes. ReCast encodes local patterns into discrete embeddings through patch-wise quantization using a learnable codebook, thereby compactly capturing stable regular structures. To compensate for residual variations not preserved by quantization, ReCast employs a dual-path architecture comprising a quantization path for efficient modeling of regular structures and a residual path for reconstructing irregular fluctuations. A central contribution of ReCast is a reliability-aware codebook update strategy, which incrementally refines the codebook via weighted corrections. These correction weights are derived by fusing multiple reliability factors from complementary perspectives by a distributionally robust optimization (DRO) scheme, ensuring adaptability to non-stationarity and robustness to distribution shifts. Extensive experiments demonstrate that ReCast outperforms state-of-the-art (SOTA) models in accuracy, efficiency, and adaptability to distribution shifts. Xiang Ma 0006, Taihua Chen, Caiming Zhang 0001 |
AAAI | 1 |
| 2026 | Aligning the True Semantics: Constrained Decoupling and Distribution Sampling for Cross-Modal AlignmentabstractCross-modal alignment is a crucial task in multimodal learning aimed at achieving semantic consistency between vision and language. This requires that image-text pairs exhibit similar semantics. Traditional algorithms pursue embedding consistency to achieve semantic consistency, ignoring the non-semantic information present in the embedding. An intuitive approach is to decouple the embeddings into semantic and modality components, aligning only the semantic component. However, this introduces two main challenges: (1) There is no established standard for distinguishing semantic and modal information. (2) The modality gap can cause semantic alignment deviation or information loss. To align the true semantics, we propose a novel cross-modal alignment algorithm via Constrained Decoupling and Distribution Sampling (CDDS). Specifically, (1) A dual-path UNet is introduced to adaptively decouple the embeddings, applying multiple constraints to ensure effective separation. (2) A distribution sampling method is proposed to bridge the modality gap, ensuring the rationality of the alignment process. Extensive experiments on various benchmarks and model backbones demonstrate the superiority of CDDS, outperforming state-of-the-art methods by 6.6% to 14.2%. Xiang Ma 0006, Lexin Fang, Litian Xu, Caiming Zhang 0001 |
AAAI | 1 |
| 2026 | RTS-LLM: Restoring time structure for time series forecasting with LLMs
Taihua Chen, Xiang Ma 0006, Yanyu Xu 0001, Shuyuan Qian, Li-Zhen Cui 0001 |
Expert Syst. Appl. | 2 |
| 2026 | TSCG: Efficient grouped channel interaction for multivariate time series forecasting
Xiang Ma 0006, Li-Zhen Cui 0001 |
Expert Syst. Appl. | 2 |
| 2025 | Minding Fuzzy Regions: A Data-driven Alternating Learning Paradigm for Stable Lesion SegmentationabstractDeep learning has achieved significant advancements in medical image segmentation, but existing models still face challenges in accurately segmenting lesion regions. The main reason is that some lesion regions in medical images have unclear boundaries, irregular shapes, and small tissue density differences, leading to label ambiguity. However, the existing model treats all data equally without taking quality differences into account in the training process, resulting in noisy labels negatively impacting model training and unstable feature representations. In this paper, a data-driven alternating learning (DALE) paradigm is proposed to optimize the model’s training process, achieving stable and high-precision segmentation. The paradigm focuses on two key points: (1) reducing the impact of noisy labels, and (2) calibrating unstable representations. To mitigate the negative impact of noisy labels, a loss consistency-based collaborative optimization method is proposed, and its effectiveness is theoretically demonstrated. Specifically, the label confidence parameters are introduced to dynamically adjust the influence of labels of different confidence levels during model training, thus reducing the influence of noise labels. To calibrate the learning bias of unstable representations, a distribution alignment method is proposed. This method restores the underlying distribution of unstable representations, thereby enhancing the discriminative capability of fuzzy region representations. Extensive experiments on various benchmarks and model backbones demonstrate the superiority of the DALE paradigm, achieving an average performance improvement of up to 7.16%. Lexin Fang, Yunyang Xu, Xiang Ma 0006, Caiming Zhang 0001 |
CVPR | 3 |
| 2025 | Reliable Cross-modal Alignment via Prototype Iterative ConstructionabstractCross-modal alignment is an important multi-modal task, aiming to bridge the semantic gap between different modalities. The most reliable fundamention for achieving this objective lies in the semantic consistency between matched pairs. Conventional methods implicitly assume embeddings contain solely semantic information, ignoring the impact of non-semantic information during alignment, which inevitably leads to information bias or even loss. These non-semantic information primarily manifest as stylistic variations in the data, which we formally define as style information. An intuitive approach is to separate style from semantics, aligning only the semantic information. However, most existing methods distinguish them based on feature columns, which cannot represent the complex coupling relationship between semantic and style information. In this paper, we propose PICO, a novel framework for suppressing style interference during embedding interaction. Specifically, we quantify the probability of each feature column representing semantic information, and regard it as the weight during the embedding interaction. To ensure the reliability of the semantic probability, we propose a prototype iterative construction method. The key operation of this method is a performance feedback-based weighting function, and we have theoretically proven that the function can assign higher weight to prototypes that bring higher performance improvements. Extensive experiments on various benchmarks and model backbones demonstrate the superiority of PICO, outperforming state-of-the-art methods by 5.2%-14.1%. Xiang Ma 0006, Litian Xu, Lexin Fang, Caiming Zhang 0001, Li-Zhen Cui 0001 |
ACM Multimedia | 1 |
| 2025 | Reinforcement learning-based portfolio optimization with deterministic state transition
Guangle Song, Tianlong Zhao, Xiang Ma 0006, Peiguang Lin, Chaoran Cui |
Inf. Sci. | 3 |
| 2025 | MDWConv:CNN based on multi-scale atrous pyramid and depthwise separable convolution for long time series forecasting
Guangpo Tian, Yunyang Xu, Xiang Ma 0006, Xuemei Li 0001, Caiming Zhang 0001 |
Neural Networks | 3 |
| 2025 | TFformer: A time-frequency domain bidirectional sequence-level attention based transformer for interpretable long-term sequence forecasting
Tianlong Zhao, Lexin Fang, Xiang Ma 0006, Xuemei Li 0001, Caiming Zhang 0001 |
Pattern Recognit. | 3 |
| 2024 | U-Mixer: An Unet-Mixer Architecture with Stationarity Correction for Time Series ForecastingabstractTime series forecasting is a crucial task in various domains. Caused by factors such as trends, seasonality, or irregular fluctuations, time series often exhibits non-stationary. It obstructs stable feature propagation through deep layers, disrupts feature distributions, and complicates learning data distribution changes. As a result, many existing models struggle to capture the underlying patterns, leading to degraded forecasting performance. In this study, we tackle the challenge of non-stationarity in time series forecasting with our proposed framework called U-Mixer. By combining Unet and Mixer, U-Mixer effectively captures local temporal dependencies between different patches and channels separately to avoid the influence of distribution variations among channels, and merge low- and high-levels features to obtain comprehensive data representations. The key contribution is a novel stationarity correction method, explicitly restoring data distribution by constraining the difference in stationarity between the data before and after model processing to restore the non-stationarity information, while ensuring the temporal dependencies are preserved. Through extensive experiments on various real-world time series datasets, U-Mixer demonstrates its effectiveness and robustness, and achieves 14.5% and 7.7% improvements over state-of-the-art (SOTA) methods. Xiang Ma 0006, Xuemei Li 0001, Lexin Fang, Tianlong Zhao, Caiming Zhang 0001 |
AAAI | 1 |
| 2024 | Bridging the Modality Gap: Dimension Information Alignment and Sparse Spatial Constraint for Image-Text MatchingabstractMany contrastive learning based models have achieved advanced performance in image-text matching tasks. The key of these models lies in analyzing the correlation between image-text pairs, which involves cross-modal interaction of embeddings in corresponding dimensions. However, the embeddings of different modalities are from different models or modules, and there is a significant modality gap. Directly interacting such embeddings lacks rationality and may capture inaccurate correlation. Therefore, we propose a novel method called DIAS to bridge the modality gap from two aspects: (1) We align the information representation of embeddings from different modalities in corresponding dimension to ensure the correlation calculation is based on interactions of similar information. (2) The spatial constraints of inter- and intra-modalities unmatched pairs are introduced to ensure the effectiveness of semantic alignment of the model. Besides, a sparse correlation algorithm is proposed to select strong correlated spatial relationships, enabling the model to learn more significant features and avoid being misled by weak correlation. Extensive experiments demonstrate the superiority of DIAS, achieving 4.3%-10.2% rSum improvements on Flickr30k and MSCOCO benchmarks. Xiang Ma 0006, Xuemei Li 0001, Lexin Fang, Caiming Zhang 0001 |
ACM Multimedia | 1 |
| 2023 | COVID19-MLSF: A multi-task learning-based stock market forecasting framework during the COVID-19 pandemic
Chenxun Yuan, Xiang Ma 0006, Hua Wang 0012, Caiming Zhang 0001, Xuemei Li 0001 |
Expert Syst. Appl. | 2 |
| 2023 | Asset correlation based deep reinforcement learning for the portfolio selection
Tianlong Zhao, Xiang Ma 0006, Xuemei Li 0001, Caiming Zhang 0001 |
Expert Syst. Appl. | 2 |
| 2023 | Dynamic graph construction via motif detection for stock prediction
Xiang Ma 0006, Xuemei Li 0001, Wenzhi Feng, Lexin Fang, Caiming Zhang 0001 |
Inf. Process. Manag. | 1 |
| 2022 | A hierarchical attention network for stock prediction based on attentive multi-view news learning
Xingtong Chen, Xiang Ma 0006, Hua Wang 0012, Xuemei Li 0001, Caiming Zhang 0001 |
Neurocomputing | 2 |
| 2022 | Fuzzy hypergraph network for recommending top-K profitable stocks
Xiang Ma 0006, Tianlong Zhao, Qiang Guo 0003, Xuemei Li 0001, Caiming Zhang 0001 |
Inf. Sci. | 1 |
| 2022 | A stock price prediction method based on meta-learning and variational mode decomposition
Tengteng Liu, Xiang Ma 0006, Xuemei Li 0001, Caiming Zhang 0001 |
Knowl. Based Syst. | 2 |
| 2021 | Image smoothing based on global sparsity decomposition and a variable parameterabstractSmoothing images, especially with rich texture, is an important problem in computer vision. Obtaining an ideal result is difficult due to complexity, irregularity, and anisotropicity of the texture. Besides, some properties are shared by the texture and the structure in an image. It is a hard compromise to retain structure and simultaneously remove texture. To create an ideal algorithm for image smoothing, we face three problems. For images with rich textures, the smoothing effect should be enhanced. We should overcome inconsistency of smoothing results in different parts of the image. It is necessary to create a method to evaluate the smoothing effect. We apply texture pre-removal based on global sparse decomposition with a variable smoothing parameter to solve the first two problems. A parametric surface constructed by an improved Bessel method is used to determine the smoothing parameter. Three evaluation measures: edge integrity rate, texture removal rate, and gradient value distribution are proposed to cope with the third problem. We use the alternating direction method of multipliers to complete the whole algorithm and obtain the results. Experiments show that our algorithm is better than existing algorithms both visually and quantitatively. We also demonstrate our method's ability in other applications such as clip-art compression artifact removal and content-aware image manipulation. Xiang Ma 0006, Xuemei Li 0001, Yuanfeng Zhou, Caiming Zhang 0001 |
Comput. Vis. Media | 1 |
| 2020 | Two-stage image smoothing based on edge-patch histogram equalisation and patch decompositionabstractPart of important structural edges in the image is smoothed due to the small gradients, while the others are preserved with greater gradients. Therefore, the authors propose a two‐stage image smoothing method based on edge‐patch histogram equalisation and patch decomposition. The authors' purpose is to increase the gradient of important structural edges while reducing the gradient of the texture region. Therefore, they divide the image into edge‐patches where the structural edges are concentrated or non‐edge‐patches where the texture details are concentrated by image segmentation. The edge‐patch needs to be equalised by the histograms for increasing the gradient of the edge pixels. All patches are decomposed to extract the smooth component for reducing the gradient of pixels. The smooth component of each patch is smoothed via gradient minimisation. In order to ensure the continuity of the patch boundaries, the edge‐patch is inversely equalised. Finally, the whole image is smoothed via gradient minimisation for removing residual textures and seams. Experimental results demonstrate that the proposed method is more competitive in maintaining important structural edges and removing texture details than the state‐of‐the‐art approaches. The proposed method can be applied to many areas of image processing. Yepeng Liu 0003, Xiang Ma 0006, Xuemei Li 0001, Caiming Zhang 0001 |
IET Image Process. | 2 |