Guangqi Jiang

dblp:221/0785 · DBLP profile ↗
← Back
33ranked-venue papers
10as first author
33since 2021 · last 2026
0000-0001-9748-0407ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 18 since 2021Artificial intelligence and machine learning · 14 · 5 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Hybrid anchor graph learning and tensorized spectral embedding fusion for multi-view clustering
Guangqi Jiang, Wangjie Chen, Yi Liu 0038, Lin Shi 0007, Jinjia Peng, Huibing Wang
Neurocomputing1
2026 Dynamic routing towards few-shot point cloud semantic segmentation
Guangqi Jiang, Zhengyao Li, Gengshen Wu, Yi Liu 0038, Shoukun Xu
Neural Networks1
2026 Visual In-Context Learning for Underwater Image Restoration
abstract
Underwater images often exhibit common visual degradations, such as color distortion, loss of details, and reduced sharpness, which inevitably compromise the effectiveness of underwater vision tasks. However, most underwater image restoration methods solely focus on learning degradation features from raw images, neglecting the incorporation of additional contextual information to guide restoration, which limits the capability of deep models to restore image quality. In this paper, we propose Visual In-Context Learning (VICL) for underwater image restoration, which leverages degradation information from context to improve image quality. In VICL, Degraded Context Extraction Block (DCEB) employs a self-attention mechanism to extract degradation information from context. In addition, Context Spatial Feature Fusion Block (CSFFB) consists of a Degraded Context Guidance Block (DCGB) and a Multi-Feature Fusion Block (MFFB). DCGB employs a cross-attention mechanism to fuse degraded context with spatial features for guiding underwater image restoration. MFFB replaces traditional encoder-decoder skip connections to better coordinate feature fusion. Extensive experiments on multiple underwater image benchmarks demonstrate that VICL outperforms state-of-the-art methods both quantitatively and visually. The code is available at:https://github.com/zhangao668/VICL.
Guangqi Jiang, Yi Liu 0038, Huibing Wang, Shoukun Xu
IEEE Signal Process. Lett.1
2026 Reliable Feature Imputation With Cross-View Relation Transfer for Deep Incomplete Multi-View Classification
abstract
Incomplete Multi-view Classification has sparked widespread interest in recent years, since multi-view data suffering from missing values are ubiquitous in real-world scenarios. While many imputation-based methods recover missing data by exploiting inter-sample structural information within individual views, they are inherently susceptible to unreliable or noisy samples, which can lead to low-quality imputation and degrade classification accuracy. Therefore, it is a challenge to effectively mine the multi-stage complex correlations for incomplete multi-view data to achieve reliable imputation and obtain discriminative representation. To address these issues, we present a novel imputation-based approach called Reliable Feature Imputation with Cross-view Relation Transfer for Deep Incomplete Multiview Classification (RFI-IMvC). Our framework fully exploits inter-view and intra-view structural information in multi-stage manner. Specifically, we propose a novel cross-view relation transfer strategy to recover reliable neighbor relationships and achieve high-quality imputation for missing data. Besides, to fully exploit the structural information in reconstructed multi-view data, we develop a dual graph learning module to mine high-order semantic correlation and facilitate interactions of complementary information from instances linked by hyperedge. Finally, inspired by prototype learning, we incorporate a class--level representation loss to further promote intra-class compactness. Extensive experiments on 7 real-world datasets demonstrate that our method outperforms state-of-the-art methods.
Guangqi Jiang, Haodong Hou, Yi Liu 0038, Jinjia Peng, Huibing Wang
IEEE Trans. Circuits Syst. Video Technol.1
2025 Robots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot Datasets
abstract
The pre-training of visual representations has enhanced the efficiency of robot learning. Due to the lack of large-scale in-domain robotic datasets, prior works utilize in-the-wild human videos to pre-train robotic visual representation. Despite their promising results, representations from human videos are inevitably subject to distribution shifts and lack the dynamics information crucial for task completion. We first evaluate various pre-trained representations in terms of their correlation to the downstream robotic manipulation tasks (i.e., manipulation centricity). Interestingly, we find that the ''manipulation centricity'' is a strong indicator of success rates when applied to downstream tasks. Drawing from these findings, we propose **M**anipulation **C**entric **R**epresentation (**MCR**), a foundation representation learning framework capturing both visual features and the dynamics information such as actions and proprioceptions of manipulation tasks to improve manipulation centricity. Specifically, we pre-train a visual encoder on the DROID robotic dataset and leverage motion-relevant data such as robot proprioceptive states and actions. We introduce a novel contrastive loss that aligns visual observations with the robot's proprioceptive state-action dynamics, combined with an action prediction loss and a time contrastive loss during pre-training. Empirical results across four simulation domains with 20 robotic manipulation tasks demonstrate that **MCR** outperforms the strongest baseline by 14.8\%. Additionally, **MCR** significantly boosts the success rate in three real-world manipulation tasks by 76.9\%. Project website: robots-pretrain-robots.github.io
Guangqi Jiang, Huanyu Li 0010, Yongyuan Liang, Huazhe Xu
ICLR1
2025 Gradient Balanced Part-Whole Relational Weakly Supervised Semantic Segmentation
Zhuang Yao, Guangqi Jiang, Lin Shi 0007, Gengshen Wu, Shoukun Xu, Yi Liu 0038
KSEM (1)2
2025 Scalable Multi-view Clustering based on Tight Anchor Distribution
Yawei Chen, Huibing Wang, Mingze Yao, Jinjia Peng, Guangqi Jiang, Jiqing Zhang
ACM Multimedia5
2025 Prior-oriented Anchor Learning with Coalesced Semantics for Multi-View Clustering
abstract
Anchor-based multi-view clustering has received a lot of attention due to its efficiency in handling large-scale datasets. However, existing methods rely on penalty-based regularization terms in anchor graphs to handle noise and outliers but overlook the role of consistent semantics in label contributions, failing to effectively mitigate their impact and potentially deviating from actual data distributions. In addition, most strategies use adaptive anchor learning without considering the veracity of anchor selection and the lack of sufficient semantic support in modeling semantic consistency, which leads to anchors deviating from the clustering center. To solve the above problems, we propose a novel method called Prior-oriented Anchor Learning with Coalesced Semantics for Multi-View Clustering (PALCS). Specifically, PALCS strips out inconsistent semantics from anchor graphs to be processed separately through coalesced semantics and highlights consistent semantics to reveal the underlying shared structure of the data. Moreover, PALCS enhances the semantic consistency and discriminative properties of anchors by directing them to be evenly distributed across clusters via the prior matrix. Finally, the clustering labels are directly obtained by non-negative matrix decomposition, avoiding additional post-processing steps. Extensive experimental evidence demonstrates the superiority of our method compared to state-of-the-art methods.
Jinjia Peng, Tianhang Cheng, Guangqi Jiang, Huibing Wang
ACM Multimedia3
2025 Consensus guided incomplete multi-view clustering via geometric consistency learning
Huibing Wang, Mingze Yao, Yawei Chen, Jinjia Peng, Guangqi Jiang, Xianping Fu
Appl. Intell.6
2025 Tensorial multiview low-rank high-order graph learning for context-enhanced domain adaptation
Chenyang Zhu 0001, Lanlan Zhang, Weibin Luo, Guangqi Jiang, Qian Wang 0017
Neural Networks4
2025 MVG-FD: Multi-Modal Visual Guidance and Feature Decomposition for Underwater Image Restoration
abstract
Underwater images are frequently affected by light absorption and scattering, which lead to color distortion, reduced contrast, and blurred details, significantly degrading overall image quality. Most underwater image restoration methods are confined to the pixel space of the raw modality, overlooking the important role of other modalities and different frequency-domain features. As a result, the representational capacity of deep learning models is not fully realized, affecting the generation of high-quality images. To address the above issues, we propose Multi-modal Visual Guidance and Feature Decomposition (MVG-FD) method for underwater image restoration. Specifically, we introduce Modality Visual Guidance (MVG) module, which integrates the complementary information provided by depth modality features into the raw features to guide the model in restoring the color of underwater images. Meanwhile, we design Feature Decomposition (FD) module, which utilizes Learnable Wavelet Decomposition (LWD) to decompose and extract the high-frequency bands of the raw features to help restore the texture details of the image. MVG-FD significantly improves PSNR and SSIM on existing datasets. The code is available at:https://github.com/zhangao668/MVG-FD.
Guangqi Jiang, Yi Liu 0038, Huibing Wang, Shoukun Xu
IEEE Signal Process. Lett.1
2024 Diffusion Reward: Learning Rewards via Conditional Video Diffusion
Guangqi Jiang, Yanjie Ze, Huazhe Xu
ECCV (42)2
2024 Spatial-Frequency Integration Network with Dual Prompt Learning for Few-shot Image Classification
abstract
Few-shot image classification is a challenging task that aims to recognize image classes based on only a few training images. However, existing methods face the following two main challenges: (1) Ignoring the frequency domain information during image feature extraction. (2) It does not take the semantic gap between multiple modalities into consideration, which limits the classification performance. To overcome these limitations, we propose a novel method named Spatial-Frequency Integration Network with Dual Prompt Learning for few-shot image classification. Firstly, we introduce a spatial-frequency integration module that combines spatial domain and low-frequency information to extract discriminative image features from the image modality. Secondly, we design a dual prompting module, which integrates learnable prompts and hand-crafted prompts to improve the generalization of applications to new classes. Thirdly, we propose an image-text interaction module to enhance inter-modal complementary and consistency. Both theoretical and experimental validations confirm the effectiveness of the proposed method in few-shot image classification.
Yi Liu 0038, Shoukun Xu, Guangqi Jiang
ISPA4
2024 Pose Convolutional Routing Towards Lightweight Capsule Networks
abstract
An important branch of artificial intelligence systems and architectures is Capsule Networks (CapsNets) have been known extremely large amount of parameters and computation because of the complex capsule routing algorithm, making it difficult deep architectures in the era of deep learning. To address this challenge, in this paper, we propose a simple yet effective capsule routing algorithm. Specifically, we activate the pose of the entity using its activation probability. On top of that, a convolution on the activated pose matrix to learn the high-level capsules’ pose matrices. Activations of the high-level capsules can be digged from their pose matrices via convolution and activation. Such mechanism generates fewer network parameters and lightweight computation, which make it practitable a deep CapsNets architecture. Experiments on CIFAR-10/100, Small-NORB, MINIST and even large-scale benchmark PASCAL VOC 2007, demonstrate the effectiveness of the proposed method.
Chengxin Lv, Guangqi Jiang, Shoukun Xu, Yi Liu 0038
ISPA2
2024 Make-An-Agent: A Generalizable Policy Network Generator with Behavior-Prompted Diffusion
abstract
Can we generate a control policy for an agent using just one demonstration of desired behaviors as a prompt, as effortlessly as creating an image from a textual description? In this paper, we present **Make-An-Agent**, a novel policy parameter generator that leverages the power of conditional diffusion models for behavior-to-policy generation. Guided by behavior embeddings that encode trajectory information, our policy generator synthesizes latent parameter representations, which can then be decoded into policy networks. Trained on policy network checkpoints and their corresponding trajectories, our generation model demonstrates remarkable versatility and scalability on multiple tasks and has a strong generalization ability on unseen tasks to output well-performed policies with only few-shot demonstrations as inputs. We showcase its efficacy and efficiency on various domains and tasks, including varying objectives, behaviors, and even across different robot manipulators. Beyond simulation, we directly deploy policies generated by **Make-An-Agent** onto real-world robots on locomotion tasks. Project page: https://cheryyunl.github.io/make-an-agent/.
Yongyuan Liang, Tingqiang Xu, Kaizhe Hu, Guangqi Jiang, Furong Huang, Huazhe Xu
NeurIPS4
2024 Adaptive Unified Framework with Global Anchor Graph for Large-Scale Multi-view Clustering
Lin Shi 0007, Wangjie Chen, Yi Liu 0038, Lihua Zhuang, Guangqi Jiang
PRCV (1)5
2024 Graph-Collaborated Auto-Encoder Hashing for Multiview Binary Clustering
abstract
Unsupervised hashing methods have attracted widespread attention with the explosive growth of large-scale data, which can greatly reduce storage and computation by learning compact binary codes. Existing unsupervised hashing methods attempt to exploit the valuable information from samples, which fails to take the local geometric structure of unlabeled samples into consideration. Moreover, hashing based on auto-encoders aims to minimize the reconstruction loss between the input data and binary codes, which ignores the potential consistency and complementarity of multiple sources data. To address the above issues, we propose a hashing algorithm based on auto-encoders for multiview binary clustering, which dynamically learns affinity graphs with low-rank constraints and adopts collaboratively learning between auto-encoders and affinity graphs to learn a unified binary code, called graph-collaborated auto-encoder (GCAE) hashing for multiview binary clustering. Specifically, we propose a multiview affinity graphs' learning model with low-rank constraint, which can mine the underlying geometric information from multiview data. Then, we design an encoder-decoder paradigm to collaborate the multiple affinity graphs, which can learn a unified binary code effectively. Notably, we impose the decorrelation and code balance constraints on binary codes to reduce the quantization errors. Finally, we use an alternating iterative optimization scheme to obtain the multiview clustering results. Extensive experimental results on five public datasets are provided to reveal the effectiveness of the algorithm and its superior performance over other state-of-the-art alternatives.
Huibing Wang, Mingze Yao, Guangqi Jiang, Zetian Mi, Xianping Fu
IEEE Trans. Neural Networks Learn. Syst.3
2023 Learning interpretable shared space via rank constraint for multi-view clustering
Guangqi Jiang, Huibing Wang, Jinjia Peng, Dongyan Chen, Xianping Fu
Appl. Intell.1
2023 Joint learning with diverse knowledge for re-identification
Jinjia Peng, Jiazuo Yu 0002, Guangqi Jiang, Huibing Wang
Signal Process. Image Commun.3
2023 Adaptive Memorization With Group Labels for Unsupervised Person Re-Identification
abstract
Re-identification (re-ID) aims to identify a person’s images across different cameras. However, the domain differences between different datasets make it a challenge for re-ID models trained on one dataset to be adapted to another. A variety of unsupervised domain adaptation methods tend to transfer learned knowledge from one domain to another by optimizing with pseudo-labels. Though impressive performances have been achieved, there are still some limitations. To be specific, these methods always generate one pseudo label for each unlabeled sample, which is hard to describe a person accurately and introduces a large number of noisy labels by one-shot clustering, thus hindering the retraining process and limiting generalization. To build more comprehensive descriptions of samples and mitigate the effects of noisy pseudo labels, this paper proposes an Adaptive Memorization with Group labels (AdaMG) framework for unsupervised person re-ID, which resists noisy labels and exploits the diversity of samples by developing a multi-branch structure with the adaptive memorization. In particular, group labels are generated for one sample in the unseen domain to learn more complementary and diverse features through clustering. Meanwhile, to better optimize the neural networks with noisy data, multiple memory structures are designed in AdaMG, which are updated adaptively according to the confidence of samples. Comprehensive experimental results have demonstrated that our proposed method can achieve excellent performances on benchmark datasets.
Jinjia Peng, Guangqi Jiang, Huibing Wang
IEEE Trans. Circuits Syst. Video Technol.2
2023 Towards Adaptive Consensus Graph: Multi-View Clustering via Graph Collaboration
abstract
Multi-view clustering is a long-standing important task, however, it remains challenging to exploit valuable information from the complex multi-view data located in diverse high-dimensional spaces. The core issue is the effective collaboration of multiple views to holistically uncover the essential correlations between multi-view data through graph learning. Furthermore, it is indispensable for most existing methods to introduce an additional clustering step to produce the final clusters, which evidently reduces the uniform relationship between graph learning and clustering. Based on the above considerations, in this paper, we present a novel method named multi-view clustering via graph collaboration (MCGC). Based on the low-dimensional representation space developed by MCGC, it first perceives the correlations between samples in each individual view under the supervision of the Hilbert-Schmidt independence criterion (HSIC). Then, MCGC proposes learning a consensus graph by adaptively collaborating between all the views, which is able to uncover the essential structure of the multi-view data. Meanwhile, by imposing the rank constraint on the Laplacian matrix of the consensus graph to partition the multi-view data naturally into the required number of clusters, the optimal clustering results can be obtained directly without any postprocessing steps. Finally, the resulting optimization problem is solved by an alternating optimization scheme with guaranteed fast convergence. Extensive experiments on 5 benchmark multi-view datasets demonstrate that MCGC markedly outperforms the state-of-the-art baselines.
Huibing Wang, Guangqi Jiang, Jinjia Peng, Ruoxi Deng, Xianping Fu
IEEE Trans. Multim.2
2022 Parallelism Network with Partial-aware and Cross-correlated Transformer for Vehicle Re-identification
abstract
Vehicle re-identification (ReID) aims to identify a specific vehicle in the dataset captured by non-overlapping cameras, which plays a great significant role in the development of intelligent transportation systems. Even though CNN-based model achieves impressive performance for the ReID task, its Gaussian distribution of effective receptive fields has limitations in capturing the long-term dependence between features. Moreover, it is crucial to capture fine-grained features and the relationship between features as much as possible from vehicle images.
Guangqi Jiang, Huibing Wang, Jinjia Peng, Xianping Fu
ICMR1
2022 Learning latent features with local channel drop network for vehicle re-identification
Xianping Fu, Jinjia Peng, Guangqi Jiang, Huibing Wang
Eng. Appl. Artif. Intell.3
2022 Cooperative Refinement Learning for domain adaptive person Re-identification
Jinjia Peng, Guangqi Jiang, Huibing Wang
Knowl. Based Syst.2
2022 Eliminating cross-camera bias for vehicle re-identification
Jinjia Peng, Guangqi Jiang, Dongyan Chen, Huibing Wang, Xianping Fu
Multim. Tools Appl.2
2022 A natural-based fusion strategy for underwater image enhancement
Xiaohong Yan, Guangxin Wang, Guangqi Jiang, Yafei Wang 0004, Zetian Mi, Xianping Fu
Multim. Tools Appl.3
2022 Conditional generative adversarial network with dual-branch progressive generator for underwater image enhancement
Yafei Wang 0004, Guangyuan Wang, Xiaohong Yan, Guangqi Jiang, Xianping Fu
Signal Process. Image Commun.5
2022 Tensorial Multi-View Clustering via Low-Rank Constrained High-Order Graph Learning
abstract
Multi-view clustering aims to partition multi-view data into different categories by optimally exploring the consistency and complementary information from multiple sources. However, most existing multi-view clustering algorithms heavily rely on the similarity graphs from respective views and fail to comprehend multiple views holistically. Moreover, due to the noise and redundancy maintained in the original data, the original errors of multiple similarity graphs will continue to accumulate in the process of constructing consistent graphs. These situations always lead to the limitation to effective fuse the essential information from multiple views, which always influences the clustering performance and cries out for reliable solutions. Based on the above considerations, we propose a novel method termed Tensorial Multi-view Clustering (TMvC), which learns high-order graph by low-rank tensor constraint to uncover the essential information stored in multiple views. TMvC first learns the Laplacian graphs of all views and stacks them into a tensor which can be viewed as a high-order graph. With the high-order graph, consistency and complementary information from different views can be propagated smoothly across all views. Then, based on low-rank constraint, high-order graph is constrained in the horizontal and vertical directions to better uncover the inter-view and inter-class correlations between multi-view data, which is of vital importance for multi-view clustering. Extensive experiments on document and image datasets demonstrate that TMvC can achieve the state-of-the-art performance for multi-view clustering.
Guangqi Jiang, Jinjia Peng, Huibing Wang, Zetian Mi, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.1
2021 Learning Multiple Semantic Knowledge For Cross-Domain Unsupervised Vehicle Re-Identification
abstract
Unsupervised Vehicle re-identification (reID) aims at searching the similar vehicles’ images from large unlabelled datasets captured in a multiple camera network, which is still a challenging task. In this paper, a multiple semantic knowledge learning approach is proposed to exploit the potential similarity of unlabeled samples, which builds multiple clusters from different views automatically with different cues. Specially, different from some existing works focus on the knowledge of one view, for each vehicle in the target domain, different semantic knowledge could be learned with the proposed focal drop network and several different labels can be assigned according these knowledge, which would be employed to train the vehicle reID model jointly. In addition, due to the unreliability of pseudo labels assigned by the clustering, the hard triplet center loss is proposed to take the difference of intra-cluster and inter-cluster into consideration for better training the unsupervised framework to adapt the unknown domain. Comprehensive experimental results clearly demonstrate that our method achieves excellent performance on both VehicleID dataset and VeRi-776 dataset.
Huibing Wang, Jinjia Peng, Guangqi Jiang, Xianping Fu
ICME3
2021 MSAV: An Unified Framework for Multi-view Subspace Analysis with View Consistence
abstract
With the development of multimedia period, information is always caputred with multiple views, which causes a research upsurge on multi-view learning. It is obvious that multi-view data contains more information than those single view ones. Therefore, it is crucial to develop the multi-view algorithms to adapt the demand of many applications. Even though some excellent multi-view algorithms were proposed, most of them can only deal with the specific problems. To tacle this problem, this paper proposes an unified framework named Multi-view Subspace Analysis with View Consistence (MSAV), which provides an unified means to extend those single-view dimension reduciton algorithms into multi-view versions. MSAV first extends multi-view data into kernel space to avoid the problem caused by different dimensions of the data from multiple views. Then, we introduced a self-weighted learning strategy to automatically assign weights for all views according to their importance. Finally, in order to promote the consistence of all views, Hilbert-Schmidt Independence Criterion is adopted by MSAV. Furthermore, We conducted experiments on several benchmark datasets to verify the performance of MSAV.
Huibing Wang, Guangqi Jiang, Jinjia Peng, Xianping Fu
ICMR2
2021 Graph-based Multi-view Binary Learning for image clustering
Guangqi Jiang, Huibing Wang, Jinjia Peng, Dongyan Chen, Xianping Fu
Neurocomputing1
2021 Discriminative feature and dictionary learning with part-aware model for vehicle re-identification
Huibing Wang, Jinjia Peng, Guangqi Jiang, Fengqiang Xu, Xianping Fu
Neurocomputing3
2021 Generalized multiple sparse information fusion for vehicle re-identification
Jinjia Peng, Guangqi Jiang, Huibing Wang
J. Vis. Commun. Image Represent.2