Xinguang Xiang

dblp:80/2205 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
10since 2021 · last 2026
0000-0002-2344-6174ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 7 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 LAPTG: Length-Aligned Attentive Prefix-Target Graph for Sequential Recommendation
Boying Li, Danlu Chen, Xinguang Xiang, Xiaoyu Du 0002
PAKDD (1)6
2026 FineG-RAG: Fine-Grained Retrieval-Augmented Generation for Multimodal Large Language Models
abstract
Fine-grained visual recognition refers to the ability to distinguish subtle differences between visually similar objects— a fundamental yet challenging capability for Multimodal Large Language Models (MLLMs). In this paper, we observe that even strong open-source MLLMs, such as Qwen2-VL and InternVL2, still struggle with accurately identifying fine-grained categories. These models often fail to attend to subtle but critical details for precise discrimination. To unlock this potential, we propose FineG-RAG, a retrieval-augmented generation pipeline designed to enhance the fine-grained recognition capabilities of MLLMs. FineG-RAG integrates external fine-grained knowledge into the recognition process via a generalized retriever. To support this, we construct fine-grained visual-language knowledge database containing representative images with wide visual diversity and expert-crafted attribute descriptions from multiple perspectives. Relevant fine-grained knowledge is retrieved from this database and fed into a visual-language augmented prompt, which provides rich multimodal context to guide MLLMs in generating accurate labels. To better evaluate the fine-grained recognition capabilities of MLLMs, we design a multiple-choice evaluation strategy based on publicly four fine-grained datasets. Extensive experiments demonstrate that FineG-RAG consistently outperforms baseline methods, achieving superior recognition accuracy across a range of off-the-shelf, open-source MLLMs.
Lu Jin 0001, Xinguang Xiang, Yanpeng Sun, Zechao Li, Jinhui Tang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2025 Enriching Responses with Crowd-Sourced Knowledge for Task-Oriented Conversational Agents
abstract
Task-oriented conversational agents strive to aid users across various tasks by concentrating on generating suitable responses to guarantee successful task accomplishment. Nonetheless, several factors have a substantial influence on user contentment beyond task fulfillment, requiring further investigation. Within this work, we aim to analyze diverse behavioral patterns of conversational agents with the goal of enhancing user satisfaction. Our findings lead to the exploration of three different enriched response generation schemes: EnRG-ATT, EnRG-TIP, and EnRG-SIM. Specifically, EnRG-ATT is designed to integrate the model's capabilities with a dual attention mechanism across two distinct modalities of external resources. It employs a pair of gates to regulate the utilization of such sources efficiently. More elegantly, we introduce EnRG-TIP, which simplifies response enrichment as a sequence prediction problem and exploits the pre-trained language model to capture user tips related to the conversation. Moreover, building on the efficiency of grounding on similar responses, EnRG-SIM further enhances response generation by inserting similar responses into the training sequences, to direct the pre-trained model's attention towards this additional knowledge. Our comprehensive experiments demonstrate that our three proposed methods not only achieve good task completion but also generate responses that yield higher user satisfaction.
Zhaohui Wei, Lizi Liao, Xinguang Xiang, Xiaoyu Du 0002
ACM Trans. Intell. Syst. Technol.3
2025 Enhancing Item-level Bundle Representation for Bundle Recommendation
abstract
Bundle recommendation approaches offer users a set of related items on a particular topic. The current state-of-the-art (SOTA) method utilizes contrastive learning to learn representations at both the bundle and item levels. However, due to the inherent difference between the bundle-level and item-level preferences, the item-level representations may not receive sufficient information from the bundle affiliations to make accurate predictions. In this article, we propose a novel approach, Enhanced Bundle Recommendation (EBRec), which incorporates two enhanced modules to explore inherent item-level bundle representations. First, we propose to incorporate the bundle-user-item (B-U-I) high-order correlations to explore more collaborative information, thus to enhance the previous bundle representation that solely relies on the bundle-item affiliation information. Second, we further enhance the B-U-I correlations by augmenting the observed user-item interactions with interactions generated from pre-trained models, thus improving the item-level bundle representations. We conduct extensive experiments on three public datasets, and the results justify the effectiveness of our approach as well as the two core modules. Codes and datasets are available at https://github.com/answermycode/EBRec .
Xiaoyu Du 0002, Yunshan Ma 0002, Xinguang Xiang
Trans. Recomm. Syst.4
2024 Multi-scale Transformer with Prompt Learning for Remote Sensing Image Dehazing
abstract
Recently, Transformers have obtained decent performance in remote sensing (RS) image dehazing. However, most existing methods do not adequately consider the multi-scale properties of RS images and the non-homogeneous distribution of haze. To this end, we develop an effective RS image dehazing method based on multi-scale Transformer with prompt learning, named MPDformer. Specifically, our method contains two key designs: inter-scale prompt branch (ISPB) and crossscale dynamic feature fusion (CDFF). The ISPB dynamically provides guidance for the dehazing process to perceive spatially-varying haze distribution, thus facilitating the transfer of useful information across different scales. Simultaneously, the CDFF flexibly aggregates multi-scale features to facilitate representations learned at different scales to communicate with each other for better image restoration. Extensive experimental results on several RS dehazing benchmarks show that MPDformer achieves favorable performance against state-of-the-art approaches.
Xinguang Xiang
ICME2
2024 Alleviating Over-Fitting in Hashing-Based Fine-Grained Image Retrieval: From Causal Feature Learning to Binary-Injected Hash Learning
abstract
Hashing-based fine-grained image retrieval pursues learning diverse local features to generate inter-class discriminative hash codes. However, existing fine-grained hash methods with attention mechanisms usually tend to just focus on a few obvious areas, which misguides the network to over-fit some salient features. Such a problem raises two main limitations. 1) It overlooks some subtle local features, degrading the generalization capability of learned embedding. 2) It causes the over-activation of some hash bits correlated to salient features, which breaks the binary code balance and further weakens the discrimination abilities of hash codes. To address these limitations of the over-fitting problem, we propose a novel hash framework fromCausalFeature learning toBinary-injectedHash learning (CFBH), which captures various local information and suppresses over-activated hash bits simultaneously. For causal feature learning, we adopt causal inference theory to alleviate the bias towards the salient regions in fine-grained images. In detail, we obtain local features from the feature map and combine this local information with original image information followed by this theory. Theoretically, these fused embeddings help the network to re-weight the retrieval effort of each local feature and exploit more subtle variations without observational bias. For binary-injected hash learning, we propose a Binary Noise Injection (BNI) module inspired by Dropout. The BNI module not only mitigates over-activation to particular bits, but also makes hash codes uncorrelated and balanced in the Hamming space. Extensive experimental results on six popular fine-grained image datasets demonstrate the superiority of CFBH over several State-of-the-Art methods.
Xinguang Xiang, Xinhao Ding, Lu Jin 0001, Zechao Li, Jinhui Tang 0001, Ramesh Jain 0001
IEEE Trans. Multim.1
2022 Unabridged adjacent modulation for clothing parsing
Chengting Zuo, Qianhao Wu, Liyong Fu, Xinguang Xiang
Pattern Recognit.5
2022 Self-Guided Image Dehazing Using Progressive Feature Fusion
abstract
We propose an effective image dehazing algorithm which explores useful information from the input hazy image itself as the guidance for the haze removal. The proposed algorithm first uses a deep pre-dehazer to generate an intermediate result, and takes it as the reference image due to the clear structures it contains. To better explore the guidance information in the generated reference image, it then develops a progressive feature fusion module to fuse the features of the hazy image and the reference image. Finally, the image restoration module takes the fused features as input to use the guidance information for better clear image restoration. All the proposed modules are trained in an end-to-end fashion, and we show that the proposed deep pre-dehazer with progressive feature fusion module is able to help haze removal. Extensive experimental results show that the proposed algorithm performs favorably against state-of-the-art methods on the widely-used dehazing benchmark datasets as well as real-world hazy images.
Haoran Bai 0001, Jinshan Pan, Xinguang Xiang, Jinhui Tang 0001
IEEE Trans. Image Process.3
2022 Sub-Region Localized Hashing for Fine-Grained Image Retrieval
abstract
Fine-grained image hashing is challenging due to the difficulties of capturing discriminative local information to generate hash codes. On the one hand, existing methods usually extract local features with the dense attention mechanism by focusing on dense local regions, which cannot contain diverse local information for fine-grained hashing. On the other hand, hash codes of the same class suffer from large intra-class variation of fine-grained images. To address the above problems, this work proposes a novel sub-Region Localized Hashing (sRLH) to learn intra-class compact and inter-class separable hash codes that also contain diverse subtle local information for efficient fine-grained image retrieval. Specifically, to localize diverse local regions, a sub-region localization module is developed to learn discriminative local features by locating the peaks of non-overlap sub-regions in the feature map. Different from localizing dense local regions, these peaks can guide the sub-region localization module to capture multifarious local discriminative information by paying close attention to dispersive local regions. To mitigate intra-class variations, hash codes of the same class are enforced to approach one common binary center. Meanwhile, the gram-schmidt orthogonalization is performed on the binary centers to make the hash codes inter-class separable. Extensive experimental results on four widely used fine-grained image retrieval datasets demonstrate the superiority of sRLH to several state-of-the-art methods. The source code of sRLH will be released at https://github.com/ZhangYajie-NJUST/sRLH.git.
Xinguang Xiang, Lu Jin 0001, Zechao Li, Jinhui Tang 0001
IEEE Trans. Image Process.1
2021 Self-Adaptive Hashing for Fine-Grained Image Retrieval
abstract
The main challenge of fine-grained image hashing is how to learn highly discriminative hash codes to distinguish the within and between class variations. On the one hand, most of the existing methods treat sample pairs as equivalent in hash learning, ignoring the more discriminative information contained in hard sample pairs. On the other hand, in the testing phase, these methods ignore the influence of outliers on retrieval performance. In order to solve the above issues, this paper proposes a novel Self-Adaptive Hashing method, which learns discriminative hash codes by mining hard sample pairs, and improves retrieval performance by correcting outliers in the testing phase. In particular, to improve the discriminability of hash codes, a pair-weighted based loss function is proposed to enhance the learning of hash functions of hard sample pairs. Furthermore, in the testing phase, a self-adaptive module is proposed to discover and correct outliers by generating self-adaptive boundaries, thereby improving the retrieval performance. Experimental results on two widely-used fine-grained datasets demonstrate the effectiveness of the proposed method.
Yuxuan Dai, Wei Tang 0011, Lu Jin 0001, Xinguang Xiang
MMAsia5
2020 Decomposed Cyclegan for Single Image Deraining With Unpaired Data
abstract
Most previous learning-based methods required paired rain image data. In practice, however, paired rain data cannot be collected. Inspired by adopting unpaired data in task of translation, in this paper we present a new method for rain removal using unpaired data. We noticed that direct use of unpaired training data may have problems, such as color shifts and background blurs. Thus, we formulate DCycleGAN, a new deep framework that decomposes the input rain image into the foreground and background parts, then produces a rain mask to guide the rain generation via re-formulated cycle-consistency constraints. Particular, the framework can simultaneously learn the foreground and background portions of the rain image, which can better remove the rain streak. Experimental results demonstrate the effectiveness of our method when trained on unpaired data.
Kewen Han, Xinguang Xiang
ICASSP2
2020 Deep Video Deblurring Using Sharpness Features From Exemplars
abstract
Video deblurring is a challenging problem as the blur in videos is usually caused by camera shake, object motion, depth variation, etc. Existing methods usually impose handcrafted image priors or use end-to-end trainable networks to solve this problem. However, using image priors usually leads to highly non-convex problems while directly using end-to-end trainable networks in a regression generates over-smoothes details in the restored images. In this paper, we explore the sharpness features from exemplars to help the blur removal and details restoration. We first estimate optical flow to explore the temporal information which can help to make full use of neighboring information. Then, we develop an encoder and decoder network and explore the sharpness features from exemplars to guide the network for better image restoration. We train the proposed algorithm in an end-to-end manner and show that using sharpness features from exemplars can help blur removal and details restoration. Both quantitative and qualitative evaluations demonstrate that our method performs favorably against state-of-the-art approaches on the benchmark video deblurring datasets and real-world images.
Xinguang Xiang, Hao Wei 0005, Jinshan Pan
IEEE Trans. Image Process.1
2019 Cascaded and Dual: Discrimination Oriented Network for Brain Tumor Classification
abstract
Medical image classification is one of the fundamental research topics in the domain of computer-aided diagnosis. Although existing classification models of the natural image can produce promising results using deep convolutional neural networks in some cases, it is difficult to guarantee that these models can generate promising performance for medical images. To bridge such a gap, we propose a novel medical image classification method for brain tumors in this paper, termed as Discrimination Oriented Network (DONet). Inspired by the attention learning mechanism of the human brain, we first propose two categories of attention learning modules, i.e., the Cascaded Attention Learning (CAL) and the Dual Attention Learning (DAL), which can learn the discrimination information in both the spatial-wise and the channel-wise dimensions in a fine-grained manner. By the CAL and the DAL, the attention information of different dimensions is calculated in a series manner (for cascaded) and a parallel manner (for dual), respectively. To demonstrate the superiority of our proposed modules, we implement the CAL and the DAL on the Deep Residual Network (ResNet) for brain tumor classification. Compared with the ResNet, experimental results show that the DONet has a significant improvement in accuracy. Moreover, compared with state-of-the-art classification methods, the DONet can also achieve better performance.
Xinguang Xiang
ACML3
2018 Tracking the evolution of overlapping communities in dynamic social networks
Zechao Li, Guan Yuan, Yunlian Sun, Xiaobin Rui, Xinguang Xiang
Knowl. Based Syst.6
2017 Light-Field Depth Estimation via Epipolar Plane Image Analysis and Locally Linear Embedding
abstract
In this paper, we propose a novel method for 4D light-field (LF) depth estimation exploiting the special linear structure of an epipolar plane image (EPI) and locally linear embedding (LLE). Without high computational complexity, depth maps are locally estimated by locating the optimal slope of each line segmentation on the EPIs, which are projected by the corresponding scene points. For each pixel to be processed, we build and then minimize the matching cost that aggregates the intensity pixel value, gradient pixel value, spatial consistency, as well as reliability measure to select the optimal slope from a predefined set of directions. Next, a subangle estimation method is proposed to further refine the obtained optimal slope of each pixel. Furthermore, based on a local reliability measure, all the pixels are classified into reliable and unreliable pixels. For the unreliable pixels, LLE is employed to propagate the missing pixels by the reliable pixels based on the assumption of manifold preserving property maintained by natural images. We demonstrate the effectiveness of our approach on a number of synthetic LF examples and real-world LF data sets, and show that our experimental results can achieve higher performance than the typical and recent state-of-the-art LF stereo matching methods.
Yongbing Zhang 0002, Huijin Lv, Yebin Liu, Haoqian Wang, Xingzheng Wang, Qian Huang 0008, Xinguang Xiang, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.7
2015 Color face recognition by PCA-like approach
Xinguang Xiang, Qiuping Chen
Neurocomputing1
2014 Local structure based sparse representation for face recognition with single sample per person
abstract
In this paper, we propose local structure based sparse representation classification (LS SRC) to solve single sample per person (SSPP) problem. By adopting the “divide-conquer-aggregate” strategy, we successfully alleviate the dilemma of high data dimensionality and small samples, where we first divide the face into local blocks, and classify each local block, and then integrate all the classification results by voting. For each block, we further divide it into overlapped patches and assume that these patches lie in a linear subspace. This subspace assumption reflects local structure relationship of the overlapped patches and makes SRC feasible for SSPP problem. To lighten the computing burden, we further propose local structure based collaborative representation classification (LS CRC). Experimental results on three public face databases show that our methods not only generalize well to SSPP problem but also have strong robustness to expression, illumination, little pose variation, occlusion and time variation.
Fan Liu 0003, Jinhui Tang 0001, Yan Song 0005, Xinguang Xiang, Zhenmin Tang
ICIP4
2012 Packet Video Error Concealment With Auto Regressive Model
abstract
In this paper, auto regressive (AR) model is applied to error concealment for block-based packet video coding. In the proposed error concealment scheme, the motion vector for each corrupted block is first derived by any kind of recovery algorithms. Then each pixel within the corrupted block is replenished as the weighted summation of pixels within a square centered at the pixel indicated by the derived motion vector in a regression manner. Two block-dependent AR coefficient derivation algorithms under spatial and temporal continuity constraints are proposed respectively. The first one derives the AR coefficients via minimizing the summation of the weighted square errors within all the available neighboring blocks under the spatial continuity constraint. The confidence weight of each pixel sample within the available neighboring blocks is inversely proportional to the distance between the sample and the corrupted block. The second one derives the AR coefficients by minimizing the summation of the weighted square errors within an extended block in the previous frame along the motion trajectory under the temporal continuity constraint. The confidence weight of each extended sample is inversely proportional to the distance toward the corresponding motion aligned block whereas the confidence weight of each sample within the motion aligned block is set to be one. The regression results generated by the two algorithms are then merged to form the ultimate restorations. Various experimental results demonstrate that the proposed error concealment strategy is able to improve both the objective and subjective quality of the replenished blocks compared to other methods.
Yongbing Zhang 0002, Xinguang Xiang, Debin Zhao, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2011 Auto-regressive model based error concealment scheme for stereoscopic video coding
abstract
Stereoscopic video is an important manner for 3-D video applications, and robust stereoscopic video transmission has posed a technical challenge for stereoscopic video coding. In this paper, an auto-regressive (AR) model based error concealment scheme is proposed for stereoscopic video coding to address the challenging problem. The proposed error concealment scheme includes a temporal AR model for independent view, and a temporal-interview AR model for inter-view predicted view. First, appropriate motions and disparities for lost blocks are derived. Then, the proposed AR model coefficients are computed according to the spatial neighboring pixels and their temporal-correlated and interview-correlated pixels indicated by the selected prediction directions. Finally, applying the AR model, each pixel of the lost block is interpolated as a weighted summation of pixels in the reference frame along the selected prediction directions. Simulation results show that the performance of the proposed scheme is superior to conventional temporal error concealment methods for stereoscopic video coding.
Xinguang Xiang, Debin Zhao, Siwei Ma 0001, Wen Gao 0001
ICASSP1
2010 Auto Regressive Model and Weighted Least Squares Based Packet Video Error Concealment
abstract
In this paper, auto regressive (AR) model is applied to error concealment for block-based packet video encoding. Each pixel within the corrupted block is restored as the weighted summation of corresponding pixels within the previous frame in a linear regression manner. Two novel algorithms using weighted least squares method are proposed to derive the AR coefficients. First, we present a coefficient derivation algorithm under the spatial continuity constraint, in which the summation of the weighted square errors within the available neighboring blocks is minimized. The confident weight of each sample is inversely proportional to the distance between the sample and the corrupted block. Second, we provide a coefficient derivation algorithm under the temporal continuity constraint, where the summation of the weighted square errors around the target pixel within the previous frame is minimized. The confident weight of each sample is proportional to the similarity of geometric proximity as well as the intensity gray level. The regression results generated by the two algorithms are then merged to form the ultimate restorations. Various experimental results demonstrate that the proposed error concealment strategy is able to increase the peak signal-to-noise ratio (PSNR) compared to other methods.
Yongbing Zhang 0002, Xinguang Xiang, Siwei Ma 0001, Debin Zhao, Wen Gao 0001
DCC2
2010 A joint encoder-decoder error control framework for stereoscopic video coding
Xinguang Xiang, Debin Zhao, Qiang Wang 0011, Siwei Ma 0001, Wen Gao 0001
J. Vis. Commun. Image Represent.1
2009 A high efficient error concealment scheme based on auto-regressive model for video coding
abstract
In this paper, a high efficient temporal error concealment scheme based on auto-regressive (AR) model is proposed for video coding. The proposed AR based error concealment scheme includes a forward AR model for P slice, and a bi-direction AR model for B slice. First, we utilize the block matching algorithm (BMA) to select the best motions for lost blocks from the motions of available neighboring blocks. Then, the proposed AR model coefficients are computed according to the spatial neighboring pixels and their temporal-correlated pixels indicated by the selected best motions. Finally, applying the AR model, each pixel of the lost block is interpolated as a weighted summation of pixels in the reference frame along the selected best motions. Simulation results show that the performance of the proposed scheme is superior to conventional temporal error concealment methods.
Xinguang Xiang, Yongbing Zhang 0002, Debin Zhao, Siwei Ma 0001, Wen Gao 0001
PCS1
2007 A Novel Error Concealment Method for Stereoscopic Video Coding
abstract
A novel error concealment method is proposed for two-view based stereoscopic video coding to address the challenging problem of adaptively combining inter-view correlation and temporal correlation. First, the disparity vectors of the lost macroblocks' neighboring macroblocks are used to recover the lost or erroneously received motion or disparity vectors. Then we propose a novel error concealment method based on overlapped block motion and disparity compensation, whose weights are determined by the side match criterion and viewpoints. Simulation results show that the subjective and objective performances of the proposed technique are both superior to those of conventional temporal error concealment methods for stereoscopic video coding.
Xinguang Xiang, Debin Zhao, Qiang Wang 0011, Xiangyang Ji, Wen Gao 0001
ICIP (5)1