VLDB 2026 Research / reviewers in the wild / expert
Xueqin Xiang
dblp:27/9555
· DBLP profile ↗
12ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0004-4767-8078ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GaborNet: attention based Gabor convolutional networks for contactless palmprint recognition
Xueqin Xiang, Wanzeng Kong, Yong Peng 0001 |
Multim. Tools Appl. | 2 |
| 2026 | Hybrid feature selection for cross-domain few shot learning
Xueqin Xiang, Wanzeng Kong, Jinliang Yao, Haihong Wu |
Pattern Recognit. | 2 |
| 2026 | LAE-Net: Large Pretrained Models Assistant Text-Guided Image Editing Adversarial NetworkabstractAutomatic real image editing offers unprecedented freedom to modify the appearance of the image or to edit a few objects through natural language. Recent scalable model families such as diffusion models have showcased remarkable proficiency in editing highly realistic images due to the introduction of vast amounts of training data and large pretrained language models. However, these large diffusion models require iterative evaluation that would significantly hinder the pace of image editing. Moreover, the pioneering work in this field necessitates the learning of a unique textual token that corresponds to each input image, or a group of images containing the same object, leading to the generation of redundant and fragmented models. Given the aforementioned problems, we suggest a novel Large pretrained models Assistant text-guided image Editing adversarial Network (LAE-Net) in this paper. More concretely, we introduce a deep semantic editing network to globally transfer text information among different isolated editing blocks, which would extract features from the source image to differentiate text-required areas from text-irrelevant ones. Furthermore, based on idea that the multi-modal CLIP model, leveraging vision-language alignment, captures comprehensive global semantic cues, whereas the vision-centric DINO model specializes in delivering intricate, fine-grained pixel-level details, the powerful discriminator of LAE-Net is designed by harnessing the visual embeddings derived from both the CLIP and DINO models separately to boost the visual discriminant capability and facilitate training a strong generator for conditioning image generation. Comprehensive experimental evaluations show that our LAE-Net not only delivers outstanding performance but also surpasses several cutting-edge models. Xueqin Xiang, Yong Peng 0001, Wanzeng Kong, Jinliang Yao |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | SF-GAN: Semantic fusion generative adversarial networks for text-to-image synthesis
Xueqin Xiang, Wanzeng Kong, Jinliang Yao |
Expert Syst. Appl. | 2 |
| 2025 | Hybrid Feature Integrated Transformer for 3D Hand Reconstruction from a Single RGB ImageabstractReconstructing a 3D hand from a single RGB image is a very challenging task. Most of the existing Transformer-based 3D hand reconstructing methods do not fully consider the local spatial information from low-level image features, which would be crucial for capturing fine details and accurate shapes of the hand. Consequently, this oversight often leads to reconstructed hands that lack the precision and realism necessary for many applications, such as augmented reality, and hand gesture recognition. To address this limitation, in this paper, we propose a novel and efficient method named HybridMETRO to both utilize low-level and high-level image features for accurate reconstructing 3D hand pose and mesh vertices from a single RGB image. Specifically, we introduce the deformable attention into the encoder of Transformer, making it no longer limited by the length of the image feature sequence. Based on the above mechanism, we further propose an interleaved updating multi-scale feature encoder to fuse low-level and high-level features. Moreover, we incorporate the Graph Convolutional Residual (GCR) module to build a novel decoder to capture explicit semantic connections between grid vertices and thus improve spatial locality of extracted features. Experimental results demonstrate that, when compared with state-of-the-art methods, our proposed HybridMETRO could achieve better performance with significantly smaller model parameters that are about half of METRO’s and a quarter of HandOccNet’s. Xueqin Xiang, Wanzeng Kong, Jinliang Yao |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | DMF-GAN: Deep Multimodal Fusion Generative Adversarial Networks for Text-to-Image SynthesisabstractText-to-image synthesis aims to generate highquality realistic images conditioned on text description. The great challenge of this task depends on deeply and seamlessly integrating image and text information. Thus, in this paper, we propose a deep multimodal fusion generative adversarial networks (DMF-GAN) that allows effective semantic interactions for finegrained text-to-image generation. Specifically, through a novel recurrent semantic fusion network, DMF-GAN could consistently manipulate global assignment of text information among isolated fusion blocks. With the assistance of a multi-head attention module, DMF-GAN could model word information from different perspectives and further improve the semantic consistency. In addition, a word-level discriminator is proposed to provide the generator with fine-grained feedback related to each word. Compared with current state-of-the-art methods, our proposed DMFGAN could efficiently synthesize realistic and text-alignment images and achieve better performance on challenging benchmarks. The code link:https://github.com/xueqinxiang/DMF-GAN Xueqin Xiang, Wanzeng Kong, Yong Peng 0001 |
IEEE Trans. Multim. | 2 |
| 2022 | Adaptive multi-task learning using lagrange multiplier for automatic art analysis
Xueqin Xiang, Wanzeng Kong, Yong Peng 0001, Jinliang Yao |
Multim. Tools Appl. | 2 |
| 2020 | 3D palmprint recognition using complete block wise descriptor
Xueqin Xiang, Jinliang Yao, Duanqing Xu |
Multim. Tools Appl. | 2 |
| 2017 | 3D palmprint recognition using shape index representation and fragile bits
Xueqin Xiang, Duanqing Xu, Xin Yang 0011 |
Multim. Tools Appl. | 2 |
| 2012 | Real-time stereo matching based on fast belief propagation
Xueqin Xiang, Guangxia Li, Yuyong He |
Mach. Vis. Appl. | 1 |
| 2011 | 3D body scanning with hairstyle using one time-of-flight cameraabstractAbstract Capturing realistic 3D shapes of human bodies is very useful for many computer graphics applications. However, for existing 3D shape capturing devices, such as structured light, laser scanner or multi‐view methods, problems arise when dealing with 3D hairstyle scanning and body deformation. To solve these problems, a novel approach is proposed to scan 3D body with hairstyle using only one time‐of‐flight (TOF) camera. By capturing depth data at video rate, temporal average meshes can be obtained from different views. After some analysis, we found that the local geometric details of real surfaces, after the hair scanning process and low‐frequency body deformation, are still preserved in the average meshes. Utilising the restriction that the corresponding surfaces in different views should overlap, a global optimisation process is proposed to iteratively improve the average meshes, while still preserving the geometric details. The proposed system is compact, and can scan 3D human body easily. Copyright © 2011 John Wiley & Sons, Ltd. Jing Tong, Xueqin Xiang, Huaqing Shen |
Comput. Animat. Virtual Worlds | 3 |
| 2010 | Fast and Simple Super Resolution for Range DataabstractCurrent active 3D range sensors, such as time-of-flight cameras, enable acquiring of range maps at video frame rate. Unfortunately, the resolution of the range maps is quite limited and the captured data are typically contaminated by noise. We therefore present a simple pipeline to enhance the quality as well as improve the spatial and depth resolution of range data in real time by up sampling the depth information with the data from high resolution video camera and utilizing a new strategy to increase the sub-pixel accuracy. Our algorithm can greatly improve the reconstruction quality, boost the resolution of the range data to that of video sensor while achieving high computational efficiency for a real-time application. Xueqin Xiang, Guangxia Li, Jing Tong |
CW | 1 |