Xueqin Xiang

dblp:27/9555 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
7since 2021 · last 2026
0009-0004-4767-8078ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 GaborNet: attention based Gabor convolutional networks for contactless palmprint recognition
Xueqin Xiang, Wanzeng Kong, Yong Peng 0001
Multim. Tools Appl.2
2026 Hybrid feature selection for cross-domain few shot learning
Xueqin Xiang, Wanzeng Kong, Jinliang Yao, Haihong Wu
Pattern Recognit.2
2026 LAE-Net: Large Pretrained Models Assistant Text-Guided Image Editing Adversarial Network
abstract
Automatic real image editing offers unprecedented freedom to modify the appearance of the image or to edit a few objects through natural language. Recent scalable model families such as diffusion models have showcased remarkable proficiency in editing highly realistic images due to the introduction of vast amounts of training data and large pretrained language models. However, these large diffusion models require iterative evaluation that would significantly hinder the pace of image editing. Moreover, the pioneering work in this field necessitates the learning of a unique textual token that corresponds to each input image, or a group of images containing the same object, leading to the generation of redundant and fragmented models. Given the aforementioned problems, we suggest a novel Large pretrained models Assistant text-guided image Editing adversarial Network (LAE-Net) in this paper. More concretely, we introduce a deep semantic editing network to globally transfer text information among different isolated editing blocks, which would extract features from the source image to differentiate text-required areas from text-irrelevant ones. Furthermore, based on idea that the multi-modal CLIP model, leveraging vision-language alignment, captures comprehensive global semantic cues, whereas the vision-centric DINO model specializes in delivering intricate, fine-grained pixel-level details, the powerful discriminator of LAE-Net is designed by harnessing the visual embeddings derived from both the CLIP and DINO models separately to boost the visual discriminant capability and facilitate training a strong generator for conditioning image generation. Comprehensive experimental evaluations show that our LAE-Net not only delivers outstanding performance but also surpasses several cutting-edge models.
Xueqin Xiang, Yong Peng 0001, Wanzeng Kong, Jinliang Yao
IEEE Trans. Vis. Comput. Graph.2
2025 SF-GAN: Semantic fusion generative adversarial networks for text-to-image synthesis
Xueqin Xiang, Wanzeng Kong, Jinliang Yao
Expert Syst. Appl.2
2025 Hybrid Feature Integrated Transformer for 3D Hand Reconstruction from a Single RGB Image
abstract
Reconstructing a 3D hand from a single RGB image is a very challenging task. Most of the existing Transformer-based 3D hand reconstructing methods do not fully consider the local spatial information from low-level image features, which would be crucial for capturing fine details and accurate shapes of the hand. Consequently, this oversight often leads to reconstructed hands that lack the precision and realism necessary for many applications, such as augmented reality, and hand gesture recognition. To address this limitation, in this paper, we propose a novel and efficient method named HybridMETRO to both utilize low-level and high-level image features for accurate reconstructing 3D hand pose and mesh vertices from a single RGB image. Specifically, we introduce the deformable attention into the encoder of Transformer, making it no longer limited by the length of the image feature sequence. Based on the above mechanism, we further propose an interleaved updating multi-scale feature encoder to fuse low-level and high-level features. Moreover, we incorporate the Graph Convolutional Residual (GCR) module to build a novel decoder to capture explicit semantic connections between grid vertices and thus improve spatial locality of extracted features. Experimental results demonstrate that, when compared with state-of-the-art methods, our proposed HybridMETRO could achieve better performance with significantly smaller model parameters that are about half of METRO’s and a quarter of HandOccNet’s.
Xueqin Xiang, Wanzeng Kong, Jinliang Yao
ACM Trans. Multim. Comput. Commun. Appl.2
2024 DMF-GAN: Deep Multimodal Fusion Generative Adversarial Networks for Text-to-Image Synthesis
abstract
Text-to-image synthesis aims to generate highquality realistic images conditioned on text description. The great challenge of this task depends on deeply and seamlessly integrating image and text information. Thus, in this paper, we propose a deep multimodal fusion generative adversarial networks (DMF-GAN) that allows effective semantic interactions for finegrained text-to-image generation. Specifically, through a novel recurrent semantic fusion network, DMF-GAN could consistently manipulate global assignment of text information among isolated fusion blocks. With the assistance of a multi-head attention module, DMF-GAN could model word information from different perspectives and further improve the semantic consistency. In addition, a word-level discriminator is proposed to provide the generator with fine-grained feedback related to each word. Compared with current state-of-the-art methods, our proposed DMFGAN could efficiently synthesize realistic and text-alignment images and achieve better performance on challenging benchmarks. The code link:https://github.com/xueqinxiang/DMF-GAN
Xueqin Xiang, Wanzeng Kong, Yong Peng 0001
IEEE Trans. Multim.2
2022 Adaptive multi-task learning using lagrange multiplier for automatic art analysis
Xueqin Xiang, Wanzeng Kong, Yong Peng 0001, Jinliang Yao
Multim. Tools Appl.2
2020 3D palmprint recognition using complete block wise descriptor
Xueqin Xiang, Jinliang Yao, Duanqing Xu
Multim. Tools Appl.2
2017 3D palmprint recognition using shape index representation and fragile bits
Xueqin Xiang, Duanqing Xu, Xin Yang 0011
Multim. Tools Appl.2
2012 Real-time stereo matching based on fast belief propagation
Xueqin Xiang, Guangxia Li, Yuyong He
Mach. Vis. Appl.1
2011 3D body scanning with hairstyle using one time-of-flight camera
abstract
Abstract Capturing realistic 3D shapes of human bodies is very useful for many computer graphics applications. However, for existing 3D shape capturing devices, such as structured light, laser scanner or multi‐view methods, problems arise when dealing with 3D hairstyle scanning and body deformation. To solve these problems, a novel approach is proposed to scan 3D body with hairstyle using only one time‐of‐flight (TOF) camera. By capturing depth data at video rate, temporal average meshes can be obtained from different views. After some analysis, we found that the local geometric details of real surfaces, after the hair scanning process and low‐frequency body deformation, are still preserved in the average meshes. Utilising the restriction that the corresponding surfaces in different views should overlap, a global optimisation process is proposed to iteratively improve the average meshes, while still preserving the geometric details. The proposed system is compact, and can scan 3D human body easily. Copyright © 2011 John Wiley & Sons, Ltd.
Jing Tong, Xueqin Xiang, Huaqing Shen
Comput. Animat. Virtual Worlds3
2010 Fast and Simple Super Resolution for Range Data
abstract
Current active 3D range sensors, such as time-of-flight cameras, enable acquiring of range maps at video frame rate. Unfortunately, the resolution of the range maps is quite limited and the captured data are typically contaminated by noise. We therefore present a simple pipeline to enhance the quality as well as improve the spatial and depth resolution of range data in real time by up sampling the depth information with the data from high resolution video camera and utilizing a new strategy to increase the sub-pixel accuracy. Our algorithm can greatly improve the reconstruction quality, boost the resolution of the range data to that of video sensor while achieving high computational efficiency for a real-time application.
Xueqin Xiang, Guangxia Li, Jing Tong
CW1