VLDB 2026 Research / reviewers in the wild / expert
Shaomin Xie
dblp:313/9605
· DBLP profile ↗
9ranked-venue papers
0as first author
9since 2021 · last 2026
0009-0007-0181-2587ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Geometry-Aware Noisy Correspondence Mitigation for Cross-Modal Text-Based Person RetrievalabstractText-Based Person Retrieval (TBPR) aims to accurately retrieve target individuals from large-scale image databases using only textual descriptions. Existing methods typically assume a ground-truth correspondence between text and images (i.e., strongly correlated). However, in real-world scenarios, this assumption may not be able to hold for the cross-modal matching due to weak or even corrupted correlations between textual descriptions and visual content, referred to as noisy correspondence (NC). Such NC largely disrupts the correspondence learning between visual and semantic modalities. Though prior works have improved single-modal robustness against noisy labels, systematic modeling of both cross-modal and intra-modal geometric structures in TBPR remains limited attention. In this paper, we propose Geometric Structure Consistency Alignment (GSCA) to TBPR, which leverages cross-modal cosine similarity and intra-modal nearest-neighbor affinity to learn visual-semantic consistency under noisy correspondence. To mitigate the structural corruption caused by noisy pairs, we introduce the Structure Refinement and Mining (SRAM) module. By partitioning training data into clean, ambiguous, and noisy subsets, SRAM enables the model to strategically refine the cross-modal correspondence by mining reliable pairs, thus enhancing the reliability of positive or negative samples discrimination and preserving structural consistency across modalities. Extensive experiments demonstrate that our method achieves state-of-the-art performance across three public datasets. On CUHK-PEDES, it boosts Rank-1 by 1.42% in noise-free conditions, sustaining a robust 74.25% Rank-1 under a 50% noise ratio. Xinpan Yuan, Shaomin Xie, Liujie Hua, Chengyuan Zhang 0001, Guihu Zhao, Lin Wu 0001 |
AAAI | 2 |
| 2026 | Syntax-Aware Dependency Parsing for Dual-Origin Noisy Correspondence in Text-Based Person Search
Xinpan Yuan, Wenguang Gan, Shaomin Xie, Chengyuan Zhang 0001, Liujie Hua |
DASFAA (1) | 4 |
| 2025 | VAMP: Visual Attribute-Guided Multi-dimensional Perception for Nasal Endoscopy Report Generation
Xinpan Yuan, Jianuo Ju, Liujie Hua, Mingzhu Huang, Shaomin Xie, Wenguang Gan |
ICIC (5) | 6 |
| 2025 | MAS-ZSAS: A Zero-Shot Anomaly Segmentation Framework with Multi-attribute Guided Text Prompts
Xinpan Yuan, Guorong Liang, Liujie Hua, Shaomin Xie, Wenguang Gan |
ICIC (6) | 4 |
| 2025 | Advanced Font-Aware Document Hierarchy Reconstruction for Enhanced Structured Parsing
Xinpan Yuan, Gan Li, Liujie Hua, Guihu Zhao, Shaomin Xie |
ICIC (16) | 5 |
| 2025 | CLIO: A Unified Framework for Consistency-Aware Learning and Intra-Modal Optimization in Text-Based Person Re-identification
Xinpan Yuan, Shaomin Xie, Guihu Zhao, Liujie Hua, Wenguang Gan |
ICIC (5) | 2 |
| 2025 | Configurable Platform for Biomedical Literature Mining via Multimodal-Driven Extraction
Xinpan Yuan, Bozhao Li, Guihu Zhao, Liujie Hua, Junhua Kuang, Shaomin Xie, Gan Li |
MICCAI (5) | 8 |
| 2024 | LMGAN: A Progressive End-to-End Chinese Landscape Painting Generation ModelabstractAI Generated Content (AIGC) has caught the attention of researchers around the world. Very good results have been achieved in the generation of Western art painting, but little research has been done on the generation of Chinese Landscape Painting (CLP). In addition, the quality of CLP generated by existing studies is still not high enough, they only imitate the color and often lack of details in their scenes. To overcome those weaknesses, we propose a progressive end-to-end Long Memory Adversarial Generation Network (LMGAN) to generate CLP. First, we create a new high-quality CLP dataset. Then we design a memory module and a mapping module to capture more low-dimensional features, as well as a layer extraction module to improve the layers of the generated results. Experiments show that LMGAN can generate richer and more detailed CLP compared to former methods. Yingqian Zhang 0002, Shaomin Xie, Xiangrong Liu |
IJCNN | 2 |
| 2022 | Noise Removal in Embedded Image With Bit ApproximationabstractStego-images are often contaminated by interchannel noise or active noise attack when communicating on the Web. And it is challenging to restore embedded image from corrupted stego-image. This paper studies akNN-bit approximation algorithm to remove noises in embedded image. The proposed algorithm distinguishes reliable bits from extracted bits, and estimates pixel values by keeping reliable bits unchanged and correcting unreliable bits. Specifically, the 8th (highest) unreliable bit of a pixel can be approximated with its nearest neighbor pixels. And then, if an unreliable bit locates at any one of the$5^{th}\sim 7^{th}$bits of a pixel, it is adjusted with two nearest neighbors of the pixel, where the pixel is in-between these two nearest neighbors. Finally, for other unreliable bits, each one is approximated by the maximum and minimum possible values of nearest neighbors of its pixel. We conduct experiments for illustrating the efficiency, and demonstrate that the proposed algorithm can recover the embedded images with good visual quality from corrupted stego-images. Xianquan Zhang, Xuelong Li 0001, Zhenjun Tang, Shichao Zhang 0001, Shaomin Xie |
IEEE Trans. Knowl. Data Eng. | 5 |