VLDB 2026 Research / reviewers in the wild / expert
Jinjia Peng
dblp:191/5342
· DBLP profile ↗
78ranked-venue papers
18as first author
68since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 49 · 9 first-author · 41 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 12 first-author · 34 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Transferable adversarial queries on person re-identification via dynamic bilateral tuning
Zeze Tao, Jinjia Peng, Huibing Wang |
Expert Syst. Appl. | 3 |
| 2026 | Instance-Guided Scene Adaptation for Unsupervised Person SearchabstractUnsupervised Domain Adaptation (UDA) is a challenging task in person search. It adapts a well-trained model from a labeled source domain to an unlabeled target domain for privacy and efficiency. Currently, most of the state-of-the-art UDA person search methods adopt multi-scale feature alignment techniques to learn domain-invariant representations. However, person search is a multi-granularity task, and such an indiscriminate method of bridging the differences between domains misleads the identity learning process, which significantly limits the model's performance. In this paper, we propose an Instance-Guided Scene Adaptation (IGSA) framework by eradicating scene disparities and focusing the tasks on instances, effectively eliminating the contradiction between person search and domain adaptation. In IGSA, a Scene-Aware Bidirectional Filter (SABF) is designed to divide the image features into background and foreground to perform bidirectional modulations, thereby achieving simultaneous scene elimination and instance enhancement. To further improve the reliability of identity learning, we also propose an Instance Consistency Contrastive Learning (ICCL) method. By performing cross-epoch updates on the instance-level memory bank and re-initializing the cluster-level memory bank, the problem of inconsistent training across epochs caused by instance identity drift can be alleviated. Through the above designs, our method can achieve state-of-the-art performance on two benchmark datasets, with 82.1% mAP and 83.8% top-1 on the CUHK-SYSU dataset and 41.1% mAP and 82.3% top-1 on the PRW dataset, which is even better than some supervised methods. Huibing Wang, Jinjia Peng, Xianping Fu, Jiqing Zhang |
AAAI | 3 |
| 2026 | Localization-Anchored Instance Discrimination for Domain Adaptive Person SearchabstractDomain-adaptive person search (DAPS) aims to transfer pedestrian detection and re-identification capabilities from a labeled source domain to an unlabeled target domain, yet faces critical challenges from domain shift: semantic confusion among overlapping instances, over-reliance on shallow features for look-alike targets, and poor discriminability of small-scale instances. To address these issues, we propose the Localization-Anchored Instance Discrimination (LAID) framework, which leverages spatial relationships between bounding boxes as auxiliary signals to enhance instance identity learning. LAID integrates three complementary strategies: 1) Cost-Aware Instance Matching (CAIM) uses IoU-based global optimal assignment to align current detections with historical identities, reducing overlap-induced misassociations; 2) Dual-Scope Contrastive Learning (DSCL) combines spatial separation constraints (for geometrically distant pairs) with global contrastive learning, prompting the model to learn deep discriminative features beyond superficial similarities; 3) Task-Sensitivity Alignment (TSA) aligns confidence distributions of detection and ReID heads via KL divergence, ensuring consistent pseudo-label generation. Extensive experiments on CUHK-SYSU and PRW datasets demonstrate that LAID outperforms state-of-the-art DAPS methods, validating its effectiveness in mitigating domain shift and narrowing the performance gap between supervised and domain-adaptive person search. Linfeng Qi 0001, Huibing Wang, Jinjia Peng, Jiqing Zhang |
AAAI | 3 |
| 2026 | Prompting Adversarial Transferability via Path Flatness AttackabstractDeep neural networks are susceptible to adversarial examples, which induce incorrect predictions through imperceptible perturbations. Transfer-based attacks create adversarial examples for surrogate models and transfer these examples to target models under black-box scenarios. Recent studies have established a strong correlation between the geometric properties of loss landscapes and the transferability of adversarial examples, demonstrating that flatter loss surfaces consistently yield superior transferability. However, we identify that these methods fail to account for the loss landscape flatness along the path from the current point to local minima, resulting in poor transferability. To address this, this paper constructs a novel Path Flatness Attack (PFA) method to significantly enhance the transferability of adversarial examples. Specifically, this paper proposes a novel path flatness indicator that not only evaluates the flatness in local minima regions but also explicitly quantifies the loss surface geometry along the trajectory from the current point to the minimum. Furthermore, we incorporate the path flatness indicator into the attack process, integrating penalties over low-loss points along the path while maximizing the loss function, thereby explicitly flattening the loss landscape. Extensive experiments demonstrate that PFA consistently achieves state-of-the-art attack performance across all experimental settings. Zeze Tao, Jinjia Peng, Huibing Wang |
AAAI | 2 |
| 2026 | A novel road damage detection model with efficient attention and Dynamic Snake Convolution
Zhen Wang 0017, Zhengyao Ma, Shunqi Gao, Jinjia Peng |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | FedDSKI: Improving server-side model via dual-stage knowledge isolation in personalized federated learning
Xiaorui He, Jinjia Peng, Zhen Wang 0017, Huibing Wang |
Future Gener. Comput. Syst. | 2 |
| 2026 | Hybrid anchor graph learning and tensorized spectral embedding fusion for multi-view clustering
Guangqi Jiang, Wangjie Chen, Yi Liu 0038, Lin Shi 0007, Jinjia Peng, Huibing Wang |
Neurocomputing | 5 |
| 2026 | Mitigating modal discrepancies for visible-infrared person re-identification via high-order nonlinear constraint
Junyu Liu, Yanzhen Xiong, Jinjia Peng, Huibing Wang |
Knowl. Based Syst. | 3 |
| 2026 | Enhancing adversarial transferability via curvature-aware penalization
Zeze Tao, Junyu Liu, Jinjia Peng |
Neural Networks | 4 |
| 2026 | BCDnet: Balanced coupling and decoupling network for person search
Zhengjie Lu, Jinjia Peng, Huibing Wang, Xianping Fu |
Pattern Recognit. | 2 |
| 2026 | Bridging the gap : Learning adaptive knowledge transition for lifelong person re-identification
Jinjia Peng, Jican Tan, Huibing Wang, Xianping Fu |
Pattern Recognit. | 1 |
| 2026 | Unsupervised text-based person retrieval via Adaptive Uncertainty-Aware Cross-Modal Learning
Weijia Sun, Yida Qi, Jinjia Peng |
Pattern Recognit. | 5 |
| 2026 | Global aggregated gradient-guided adversarial attacks for person re-identification
Zeze Tao, Jinjia Peng, Huibing Wang |
Pattern Recognit. | 3 |
| 2026 | RDNet: Dynamic filtering guided transformer with cross-batch feature retention for camouflaged object detection
Songxiao Geng, Jinjia Peng, Weibin Liu |
Pattern Recognit. | 3 |
| 2026 | Reliable Feature Imputation With Cross-View Relation Transfer for Deep Incomplete Multi-View ClassificationabstractIncomplete Multi-view Classification has sparked widespread interest in recent years, since multi-view data suffering from missing values are ubiquitous in real-world scenarios. While many imputation-based methods recover missing data by exploiting inter-sample structural information within individual views, they are inherently susceptible to unreliable or noisy samples, which can lead to low-quality imputation and degrade classification accuracy. Therefore, it is a challenge to effectively mine the multi-stage complex correlations for incomplete multi-view data to achieve reliable imputation and obtain discriminative representation. To address these issues, we present a novel imputation-based approach called Reliable Feature Imputation with Cross-view Relation Transfer for Deep Incomplete Multiview Classification (RFI-IMvC). Our framework fully exploits inter-view and intra-view structural information in multi-stage manner. Specifically, we propose a novel cross-view relation transfer strategy to recover reliable neighbor relationships and achieve high-quality imputation for missing data. Besides, to fully exploit the structural information in reconstructed multi-view data, we develop a dual graph learning module to mine high-order semantic correlation and facilitate interactions of complementary information from instances linked by hyperedge. Finally, inspired by prototype learning, we incorporate a class--level representation loss to further promote intra-class compactness. Extensive experiments on 7 real-world datasets demonstrate that our method outperforms state-of-the-art methods. Guangqi Jiang, Haodong Hou, Yi Liu 0038, Jinjia Peng, Huibing Wang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Hierarchical Sequential Context Modeling for High-Fidelity Image InpaintingabstractImage inpainting aims to restore missing regions by leveraging surrounding spatial context, where nearby pixels provide crucial structural cues and distant regions offer complementary semantic guidance. To jointly model these complementary dependencies, this paper proposes Hierarchical Sequential Context Modeling (HSCM), a novel inpainting framework that employs state-space models for multi-scale autoregressive sequence modeling. Unlike existing single-scale SSM-based approaches, HSCM explicitly separates pixel-level and semanticlevel modeling into two complementary branches. The Local Perception Unit preserves fine-grained textures, and the Global Compensation Unit propagates high-level semantics across patches to enhance overall coherence. The asynchronous hierarchical design first reconstructs local textures and then performs semantic compensation, achieving notable performance gains with minimal computational overhead. Leveraging its four-directional architecture, HSCM maintains linear computational growth with spatial resolution and effectively establishes a comprehensive global receptive field. Furthermore, a Cross-Gated Feedforward Network is proposed to alleviate patch boundary artifacts and enhance inter-channel feature consistency. Built upon a multi-scale encoder–decoder architecture, HSCM delivers state-of-the-art inpainting quality and robust generalization across diverse benchmarks, including CelebA-HQ, FFHQ, Paris Street View, and Places2. Zexuan Sun, Jinjia Peng, Mengkai Li, Huibing Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Bi-Level Inter-Modality Modulation for Unsupervised Visible-Infrared Person Re-IdentificationabstractThe task of unsupervised visible–infrared person re-identification (USL-VI-ReID) aims to retrieve cross-modal pedestrian images without manual annotations. The key challenge lies in achieving semantic alignment to resolve modality bias in the absence of real labels. However, existing methods overly rely on single-modal information in the process of pseudo-label generation without considering cross-modal associations, making it difficult to bridge the modality gap between visible and infrared images. To address these issues, this paper proposes a Bi-level Inter-Modal Modulation Network (BIMM-Net), which employs multi-level cluster structure optimization as a core strategy to drive the establishment of cross-modal semantic associations, ultimately achieving cross-modal alignment at the feature representation level. Specifically, we construct a novel intermediary modality GrayMix from visible images to enhance model robustness against color variations and alleviate modality gaps. To filter out noise in cross-modal matching and establish a shared semantic space between visible and infrared modalities, we further develop a Ternary Pairs Calibration-Convergence module designed for filtering noise from visible-infrared cluster matching, on this basis constructing fused mixture clusters. Building on this mixture cluster space, an Heterogeneous-Isomorphic Alignment Loss is also designed to align the feature distributions of the three modalities, reinforcing cross-modal semantic consistency. In addition, we present a Cross-modal Neighborhood Consistency Clustering method, which facilitates the formation and propagation of cross-modal clusters by selecting high-confidence cross-modal neighbor pairs and refining feature distances. Ultimately, BIMM-Net through the joint modeling of bi-level clustering enables multiple levels to guide each other in refining cross-modal structures, thereby effectively establishing the semantic associations between visible and infrared modalities. Extensive experiments validate the superior performance of the proposed framework, achieving state-of-the-art results in USL-VI-ReID. The source code of this paper is available at: https://github.com/liujuny5920/DIMM-Net. Jinjia Peng, Junyu Liu, Xutao Zuo, Zeze Tao, Huibing Wang |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2026 | Unsupervised Lifelong Person Re-Identification via Affinity HarmonizationabstractLifelong Person Re-Identification (LReID) seeks to continuously train models across multiple target domains, enabling effective generalization in both known and unseen domains. Achieving a balance between “plasticity” (the ability to adapt to new knowledge) and “stability” (the capacity to prevent forgetting) is crucial in lifelong learning. However, most existing LReID methods primarily focus on enhancing model stability or plasticity, often neglecting the critical balance between them. Moreover, current LReID approaches largely rely on supervised learning, which necessitates large-scale pre-labeled datasets—a process that is both time-consuming and labor-intensive in practical applications. To address these challenges, this article proposes an Unsupervised LReID approach called the Affinity Harmonization Network (AHN). AHN includes an Old Domain Affinity Constraint (ODAC) module, which builds an expert model for the old domain to provide affinity relationships as references. This helps limit changes among old representations, enabling the model to integrate new knowledge while preserving compatibility with previous representations. To harmonize stability and plasticity while guiding the model in acquiring new knowledge, AHN incorporates a Current Domain Affinity Guidance (CDAG) module. This module builds an expert model for the new domain and uses the generated affinity relationships to assist in training the model. Furthermore, this article proposes the Old Domain Intra-class Variance Constraint (OIVC) module, which mitigates potential deviations in the intra-class variance of legacy samples by limiting the distance between replay samples and old domain camera prototypes. Extensive experiments demonstrate that our method achieves significant performance improvements over existing unsupervised lifelong ReID methods, with an average gain of 5.3% in mAP and 5.2% in Rank-1 accuracy. Jican Tan, Jinjia Peng, Songyu Zhang, Zhen Wang 0017, Huibing Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Anchor Learning with Potential Cluster Constraints for Multi-view ClusteringabstractAnchor-based multi-view clustering has received extensive attention due to its efficient performance. Existing methods only focus on how to dynamically learn anchors from the original data and simultaneously construct anchor graphs describing the relationships between samples and perform clustering, while ignoring the reality of anchors, i.e., high-quality anchors should be generated uniformly from different clusters of data rather than scattered outside the clusters. To deal with this problem, we propose a noval method termed Anchor Learning with Potential Cluster Constraints for Multi-view Clustering (ALPC) method. Specifically, ALPC first establishes a shared latent semantic module to constrain anchors to be generated from specific clusters, and subsequently, ALPC improves the representativeness and discriminability of anchors by adapting the anchor graph to capture the common clustering center of mass from samples and anchors, respectively. Finally, ALPC combines anchor learning and graph construction into a unified framework for collaborative learning and mutual optimization to improve the clustering performance. Extensive experiments demonstrate the effectiveness of our proposed method compared to some state-of-the-art MVC methods. Yawei Chen, Huibing Wang, Jinjia Peng, Yang Wang 0023 |
AAAI | 3 |
| 2025 | CDE-Learning: Camera Deviation Elimination Learning for Unsupervised Person Re-identificationabstractUnsupervised Person Re-identification (Re-ID) aims to identify the same person shot from non-overlapping cameras without any annotated data. In this task, attributes such as contrast, saturation, and resolution of the camera cause the deviation in target features. Since the camera label is readily available, they are employed to achieve the constraints across cameras and smooth the deviations during the model training phase. However, features from the same camera are prone to generating false positives due to the identical camera properties, which induce camera deviations on pseudo-label assignment. To address this problem, this paper proposes a novel camera-unbiased method named Camera Deviation Elimination Learning (CDE-Learning). In CDE-Learning, the Camera Deviation Compensation (CDC) module is designed to align data distributions from disparate cameras to decouple camera information from identity information during the pseudo-label allocation. Our Camera Deviation Balancing (CDB) module integrates different camera constraints in a united loss and adjusts camera constraints by constructing contrastive pairs between intra-camera and inter-camera. After explicit constraints, the Camera Attribution Auxiliary (CAA) task predicts whether a pair of images originates from the same camera to implicitly enhance the capacity to distinguish the camera deviation. We demonstrated the superior performance of the proposed CDE-Learning on benchmark datasets. Jinjia Peng, Songyu Zhang, Huibing Wang |
AAAI | 1 |
| 2025 | Unsupervised Domain Adaptive Person Search via Dual Self-CalibrationabstractUnsupervised Domain Adaptive (UDA) person search focuses on employing the model trained on a labeled source domain dataset to a target domain dataset without any additional annotations. Most effective UDA person search methods typically utilize the ground truth of the source domain and pseudo-labels derived from clustering during the training process for domain adaptation. However, the performance of these approaches will be significantly restricted by the disrupting pseudo-labels resulting from inter-domain disparities. In this paper, we propose a Dual Self-Calibration (DSCA) framework for UDA person search that effectively eliminates the interference of noisy pseudo-labels by considering both the image-level and instance-level features perspectives. Specifically, we first present a simple yet effective Perception-Driven Adaptive Filter (PDAF) to adaptively predict a dynamic filter threshold based on input features. This threshold assists in eliminating noisy pseudo-boxes and other background interference, allowing our approach to focus on foreground targets and avoid indiscriminate domain adaptation. Besides, we further propose a Cluster Proxy Representation (CPR) module to enhance the update strategy of cluster representation, which mitigates the pollution of clusters from misidentified instances and effectively streamlines the training process for unlabeled target domains. With the above design, our method can achieve state-of-the-art (SOTA) performance on two benchmark datasets, with 80.2% mAP and 81.7% top-1 on the CUHK-SYSU dataset, with 39.9% mAP and 81.6% top-1 on the PRW dataset, which is comparable to or even exceeds the performance of some fully supervised methods. Linfeng Qi 0001, Huibing Wang, Jiqing Zhang, Jinjia Peng |
AAAI | 4 |
| 2025 | Boosting Adversarial Transferability via Residual Perturbation AttackabstractDeep neural networks are susceptible to adversarial examples while suffering from incorrect predictions via imperceptible perturbations. Transfer-based attacks create adversarial examples for surrogate models and transfer these examples to target models under black-box scenarios. Recent studies reveal that adversarial examples in flat loss landscapes exhibit superior transferability to alleviate overfitting on surrogate models. However, the prior arts overlook the influence of perturbation directions, resulting in limited transferability. In this paper, we propose a novel attack method, named Residual Perturbation Attack (ResPA), relying on the residual gradient as the perturbation direction to guide the adversarial examples toward the flat regions of the loss function. Specifically, ResPA conducts an exponential moving average on the input gradients to obtain the first moment as the reference gradient, which encompasses the direction of historical gradients. Instead of heavily relying on the local flatness that stems from the current gradients as the perturbation direction, ResPA further considers the residual between the current gradient and the reference gradient to capture the changes in the global perturbation direction. The experimental results demonstrate the better transferability of ResPA than the existing typical transfer-based attack methods, while the transferability can be further improved by combining ResPA with the current input transformation methods. The code is available at https://github.com/ZezeTao/ResPA. Jinjia Peng, Zeze Tao, Huibing Wang, Meng Wang 0001, Yang Wang 0023 |
ICCV | 1 |
| 2025 | Consensus-Guided Incomplete Multi-view Clustering via Cross-view Affinities LearningabstractIncomplete multi-view clustering (IMC) has garnered substantial attention due to its capacity to handle unlabeled data. Existing methods predominantly explore pairwise consistency between every two views. However, such consistency is highly susceptible to missing samples and outliers within a certain view and thus deviates from the true clustering distribution. Moreover, dual-view interaction neglects the collaboration effects of multiple views, making it challenging to capture the holistic characteristics across views. In response to these issues, we propose a novel Consensus-Guided Incomplete Multi-view Clustering via Cross-view Affinities Learning (CAL). Specifically, CAL reconstructs views with available instances to mine sample-wise affinities and harness comprehensive content information within views. Subsequently, to extract clean structural information, CAL imposes a structured sparse constraint on the representation tensor to eliminate biased errors. Furthermore, by integrating the consensus representation into a representation tensor, CAL can employ high-order interaction of multiple views to depict the semantic correlation between views while acquiring a unified structural graph across multiple views. Extensive experiments on seven benchmark datasets demonstrate that CAL outperforms some state-of-the-art methods in clustering performance. The code is available at https://github.com/whbdmu/CAL. Huibing Wang, Jinjia Peng, Yawei Chen, Mingze Yao, Xianping Fu, Yang Wang 0023 |
IJCAI | 3 |
| 2025 | Stabilizing Holistic Semantics in Diffusion Bridge for Image InpaintingabstractImage inpainting aims to restore the original image from a damaged version. Recently, a special type of diffusion bridge model has achieved promising performance by directly mapping the degradation process and restoring corrupted images through the corresponding reverse process. However, due to the lack of explicit semantic priors during the denoising process, the inpainted results typically exhibit inferior context-stability and semantic consistency. To this end, this paper proposes a novel Global Structure-Guided Diffusion Bridge framework (GSGDiff), which incorporates an additional structure restorer to stabilize the generation of holistic semantics. Specifically, to acquire richer semantic structure priors, this paper proposes a posterior sampling approach that captures semantically global and consistent structures at each timestep, efficiently integrating them into the texture generation through the corresponding guidance module. Additionally, considering the characteristics of diffusion models with low denoising levels at larger timesteps, this paper proposes a semantic fusion schedule to avoid noise interference by reducing the weight of ineffective guided semantics in the early stages. By applying the proposed posterior sampling to the texture denoising process, GSGDiff can achieve more stable and superior inpainting results over competitive baselines. Experiments on Places2, Paris Street View and CelebA-HQ datasets validate the efficacy of the proposed method. Jinjia Peng, Mengkai Li, Huibing Wang |
IJCAI | 1 |
| 2025 | Scalable Multi-view Clustering based on Tight Anchor Distribution
Yawei Chen, Huibing Wang, Mingze Yao, Jinjia Peng, Guangqi Jiang, Jiqing Zhang |
ACM Multimedia | 4 |
| 2025 | Prior-oriented Anchor Learning with Coalesced Semantics for Multi-View ClusteringabstractAnchor-based multi-view clustering has received a lot of attention due to its efficiency in handling large-scale datasets. However, existing methods rely on penalty-based regularization terms in anchor graphs to handle noise and outliers but overlook the role of consistent semantics in label contributions, failing to effectively mitigate their impact and potentially deviating from actual data distributions. In addition, most strategies use adaptive anchor learning without considering the veracity of anchor selection and the lack of sufficient semantic support in modeling semantic consistency, which leads to anchors deviating from the clustering center. To solve the above problems, we propose a novel method called Prior-oriented Anchor Learning with Coalesced Semantics for Multi-View Clustering (PALCS). Specifically, PALCS strips out inconsistent semantics from anchor graphs to be processed separately through coalesced semantics and highlights consistent semantics to reveal the underlying shared structure of the data. Moreover, PALCS enhances the semantic consistency and discriminative properties of anchors by directing them to be evenly distributed across clusters via the prior matrix. Finally, the clustering labels are directly obtained by non-negative matrix decomposition, avoiding additional post-processing steps. Extensive experimental evidence demonstrates the superiority of our method compared to state-of-the-art methods. Jinjia Peng, Tianhang Cheng, Guangqi Jiang, Huibing Wang |
ACM Multimedia | 1 |
| 2025 | Spatiotemporal Consensus with Scene Prior for Unsupervised Domain Adaptive Person SearchabstractPerson Search aims to locate query persons in gallery scene images, but faces severe performance degradation under domain shifts. Unsupervised domain adaptation transfers knowledge from the labeled source domain to the unlabeled target domain and iteratively rectifies the pseudo-labels. However, the pseudo-labels are inevitably contaminated by the source-biased model, which misleads the training process. This, in turn, reduces the quality of the pseudo-labels themselves and ultimately affects the search performance. In this paper, we propose a Spatiotemporal Consensus with Scene Prior (STCSP) framework that effectively eliminates the interference of noise on pseudo-labels, establishes positive feedback, and thus gradually bridging the domain gap. Firstly, STCSP uses a Spatiotemporal Consensus pipeline to suppress the noise from being mixed into the pseudo-labels. Secondly, leveraging the scene prior, STCSP employs our designed Iterative Bilateral Extremum Matching method to prevent the occurrence of some incorrect pseudo-labels. Thirdly, we propose a Scene Prior Contrastive Learning module, which encourages the model to directly acquire the scene prior knowledge from the target domain, thereby mitigating the generation of noise. By suppressing noise contamination, avoiding noise occurrence and mitigating noise generation, our framework achieves state-of-the-art performance on two benchmark datasets, PRW with 50.2% mAP and CUHK-SYSU with 87.0% mAP. Huibing Wang, Jinjia Peng |
NeurIPS | 3 |
| 2025 | Bidirectional Knowledge Distillation for Unsupervised Lifelong Person Re-identification
Jican Tan, Kunze Li, Jinjia Peng |
PRCV (16) | 4 |
| 2025 | Consensus guided incomplete multi-view clustering via geometric consistency learning
Huibing Wang, Mingze Yao, Yawei Chen, Jinjia Peng, Guangqi Jiang, Xianping Fu |
Appl. Intell. | 5 |
| 2025 | Region-guided spatial feature aggregation network for vehicle re-identification
Yanzhen Xiong, Jinjia Peng, Zeze Tao, Huibing Wang |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Scene-intra deep mining via shifted instance refinement for weakly supervised person search
Shenghui Yin, Jinjia Peng, Zhen Wang 0017, Huibing Wang |
Knowl. Based Syst. | 2 |
| 2025 | Omni Contextual Aggregation Networks for High-Fidelity Image InpaintingabstractImage inpainting aims to restore a realistic image from a damaged or incomplete version. Although Transformer-based methods have achieved impressive results by modeling long-range dependencies, the inherent quadratic complexity of canonical self-attention has typically led to these approaches adopting uni-dimensional modeling, which limits the model’s ability to capture complex relationships from both spatial and channel dimensions. To this end, this paper exploits a novel attention paradigm termed Dynamic Omni-Attention Mechanism (DOAM) for simultaneously modeling pixel-interaction from both spatial and channel dimensions, and implements the information interaction across the omni-axis (i.e., spatial and channel) with linear computational complexity. In addition, to handle large-scale degradation, this paper proposes a Multi-band Feature Enhancement (MFE) module to enhance feature representation in downsampling, thus unlocking the potential of subsequent attentional interactions. Moreover, motivated by recent advances in image restoration, this paper incorporates a domain-related prior representation from CNN-based Network to modulate the features during proposed attention mechanism and feed-forward networks. Integrating the above designs into an encoder-decoder architecture, the proposed Omni Contextual Aggregation Networks (OCANet) achieve superior performance at lower parameters and time costs than the competitive baselines. Extensive experiments on CelebA-HQ, Paris Street View, FFHQ and Dunhuang datasets validate the efficacy of the proposed method. Jinjia Peng, Mengkai Li, Huibing Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Tensor Completion Framework by Graph Refinement for Incomplete Multi-View ClusteringabstractIncomplete Multi-view Clustering (IMVC) endeavors to harness information from multiple incomplete views to partition multi-view data into their respective clusters. How to recover missing information with lossless fidelity is the core of IMVC, which is of vital importance but challenging. Most of the existing methods include a feature recovery step to mitigate the negative impact of missing samples on the feature graph, however, these IMVC algorithms simply utilize the correlation between samples to recover the relationship between the unmissing instances and the missing instances while ignoring the consistency between views, which leads to often unsatisfactory recovery results. In addition, previous IMVC algorithms focus more on the recovery of incomplete data, ignoring the effect of the error term on incomplete graphs. This can mislead the recovery process of IMVC algorithm and the feature graph can be affected by anomalous information, which leads to degradation of clustering performance. To address this gap, this paper introduces the Tensor Completion Framework by Graph Refinement for Incomplete Multi-view Clustering (IMVC-TGR). IMVC-TGR separates the redundant information in each affine graph by graph refinement operation, aiming to mitigate the negative impact of error terms and redundant information on the feature graph during the recovery process. Meanwhile, IMVC-TGR stacks the feature graphs into tensors to explore intra-view correlation and inter-view consistency, so as to recover the relationship between missing samples and non-missing samples, and improve the quality of the feature graphs. Finally, IMVC-TGR introduces semantic consistency constraints and self-weighted fusion strategies into the high-quality feature graphs, aiming at preserving the complementary information between different views while balancing the contributions of the refined representation matrices of different views. The experimental results on multiple different datasets indicate that IMVC-TGR can achieve state-of-the-art performance. Huibing Wang, Yawei Chen, Mingze Yao, Jinjia Peng, Xianping Fu |
IEEE Trans. Multim. | 5 |
| 2024 | Multi-pattern Joint Denoising Diffusion Model for Sequential Recommendation
Hancheng Lu, Liang Wang 0010, Jinjia Peng |
ICONIP (5) | 5 |
| 2024 | Fast One-Stage Unsupervised Domain Adaptive Person Search
Tianxiang Cui, Huibing Wang, Jinjia Peng, Ruoxi Deng, Xianping Fu, Yang Wang 0023 |
IJCAI | 3 |
| 2024 | Scene-Adaptive Person Search via Bilateral Modulations
Huibing Wang, Jinjia Peng, Xianping Fu, Yang Wang 0023 |
IJCAI | 3 |
| 2024 | Multi-attribute Semantic Adversarial Attack Based on Cross-layer Interpolation for Face RecognitionabstractWith the extensive research and application on Face Recognition (FR) model in daily life, the security of FR has attracted much attention as it is easily attacked by adversarial examples. Specifically, adversarial attacks can cause the model to make completely erroneous judgments by making very subtle changes to the source image. Therefore, it is of great significance for studying adversarial attacks that can improve the robustness and security of FR models. However, most of the existing attacks have low transferability of attack and high vulnerability to denoising defense models. To solve the above problems, a multi-attribute semantic adversarial attack based on cross-layer interpolation(C&A Adv) is proposed, which can generate imperceptible adversarial images whith high success rate and robustness to denoising defense methods. Particularly, C&A Adv semantically edit images by cross-layer feature space interpolation, which not only generates high quality adversarial images, but also has the robustness to partial denoising defense methods. In addition, to improve the success rate of the attack, several attributes are selected to edit instead of just one. According to the marginal gain of each attribute calculated in different face images, several attributes with the greatest marginal gain are selected to edit. Comparison and verification on CelebA dataset show that the C&A Adv achieves good experimental result. Ruizhong Du, Yidan Li, Jinjia Peng, Caixia Ma |
IJCNN | 4 |
| 2024 | HSMnet: Hybrid Sampling and Matching Network for DETR-based Person Search
Zhengjie Lu, Jinjia Peng, Huibing Wang, Qingxuan Shi, Bin Wang 0044 |
MMAsia | 2 |
| 2024 | ABC-Learning: Attention-Boosted Contrastive Learning for unsupervised person re-identification
Songyu Zhang, Jinjia Peng |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Large-scale multi-view subspace clustering via embedding space and partition matrix
Tianhang Cheng, Jinjia Peng, Huibing Wang |
Neurocomputing | 2 |
| 2024 | AGS: Transferable adversarial attack for person re-identification by adaptive gradient similarity attack
Zeze Tao, Zhengjie Lu, Jinjia Peng, Huibing Wang |
Knowl. Based Syst. | 3 |
| 2024 | Generating high-quality texture via panoramic feature aggregation for large mask inpainting
Jinjia Peng, Huibing Wang |
Knowl. Based Syst. | 2 |
| 2024 | Uncertainty-guided Robust labels refinement for unsupervised person re-identification
Jinjia Peng, Zeze Tao, Huibing Wang |
Neural Comput. Appl. | 2 |
| 2024 | Adapt only once: Fast unsupervised person re-identification via relevance-aware guidance
Jinjia Peng, Jiazuo Yu 0002, Huibing Wang, Xianping Fu |
Pattern Recognit. | 1 |
| 2024 | ReFID: Reciprocal Frequency-aware Generalizable Person Re-identification via Decomposition and FilteringabstractDomain generalization of person re-identification aims to conduct testing across domains that have not been previously encountered, without utilizing target domain data during the training stage. As the number of source domains increases, the relationships between training samples become more complex. This can lead to domain-invariant features that include certain instance-level spurious correlations, which can impact the model’s ability to generalize further. To overcome this limitation, the Reciprocal Frequency-aware Generalizable Person Re-identification method is proposed in this article, which aims to utilize spectral feature correlation learning to transmit frequency component information and generate more discriminative hybrid features. A module called Bilateral Frequency Component-guided Attention is developed to help the network understand high-level semantic and texture information from various frequency features. Furthermore, to reduce the impact of noise from the frequency domain, this article proposes an innovative module called Fourier Noise Masquerade Filtering. This module enhances the portability of frequency domain components while simultaneously suppressing elements that do not contribute to generalization. Extensive experimental results on various datasets demonstrate that our method is effective and superior to the state-of-the-art methods. Jinjia Peng, Song Pengpeng, Huibing Wang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Dynamic context-driven progressive image inpainting with auxiliary generative units
Kai Li 0038, Jinjia Peng |
Vis. Comput. | 3 |
| 2023 | Learning Frequency-Based Disentanglement and Filtering for Generalizable Person Re-identification
Pengpeng Song, Jinjia Peng |
PRCV (12) | 2 |
| 2023 | Deep Multi-task Image Clustering with Attention-Guided Patch Filtering and Correlation Mining
Zhongyao Tian, Jinjia Peng |
PRCV (4) | 3 |
| 2023 | One Step Large-Scale Multi-view Subspace Clustering Based on Orthogonal Matrix Factorization with Consensus Graph Learning
Jinjia Peng |
PRCV (4) | 3 |
| 2023 | Distribution-based Learnable Filters with Side Information for Sequential RecommendationabstractSequential Recommendation aims to predict the next item by mining out the dynamic preference from user previous interactions. However, most methods represent each item as a single fixed vector, which is incapable of capturing the uncertainty of item-item transitions that result from time-dependent and multifarious interests of users. Besides, they struggle to effectively exploit side information that helps to better express user preferences. Finally, the noise in user’s access sequence, which is due to accidental clicks, can interfere with the next item prediction and lead to lower recommendation performance. To deal with these issues, we propose DLFS-Rec, a simple and novel model that combines Distribution-based Learnable Filters with Side information for sequential Recommendation. Specifically, items and their side information are represented by stochastic Gaussian distribution, which is described by mean and covariance embeddings, and then the corresponding embeddings are fused to generate a final representation for each item. To attenuate noise, stacked learnable filter layers are applied to smooth the fused embeddings. Extensive experiments on four public real-world datasets demonstrate the superiority of the proposed model over state-of-the-art baselines, especially on cold start users and items. Codes are available at https://github.com/zxiang30/DLFS-Rec. Zhixiang Deng, Liang Wang 0010, Jinjia Peng, Shi Feng 0001 |
RecSys | 4 |
| 2023 | Learning interpretable shared space via rank constraint for multi-view clustering
Guangqi Jiang, Huibing Wang, Jinjia Peng, Dongyan Chen, Xianping Fu |
Appl. Intell. | 3 |
| 2023 | Hybrid partial-constrained learning with orthogonality regularization for unsupervised person re-identification
Jiazuo Yu 0002, Jinjia Peng, Kai Li 0038, Huibing Wang |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Joint learning with diverse knowledge for re-identification
Jinjia Peng, Jiazuo Yu 0002, Guangqi Jiang, Huibing Wang |
Signal Process. Image Commun. | 1 |
| 2023 | Adaptive Memorization With Group Labels for Unsupervised Person Re-IdentificationabstractRe-identification (re-ID) aims to identify a person’s images across different cameras. However, the domain differences between different datasets make it a challenge for re-ID models trained on one dataset to be adapted to another. A variety of unsupervised domain adaptation methods tend to transfer learned knowledge from one domain to another by optimizing with pseudo-labels. Though impressive performances have been achieved, there are still some limitations. To be specific, these methods always generate one pseudo label for each unlabeled sample, which is hard to describe a person accurately and introduces a large number of noisy labels by one-shot clustering, thus hindering the retraining process and limiting generalization. To build more comprehensive descriptions of samples and mitigate the effects of noisy pseudo labels, this paper proposes an Adaptive Memorization with Group labels (AdaMG) framework for unsupervised person re-ID, which resists noisy labels and exploits the diversity of samples by developing a multi-branch structure with the adaptive memorization. In particular, group labels are generated for one sample in the unseen domain to learn more complementary and diverse features through clustering. Meanwhile, to better optimize the neural networks with noisy data, multiple memory structures are designed in AdaMG, which are updated adaptively according to the confidence of samples. Comprehensive experimental results have demonstrated that our proposed method can achieve excellent performances on benchmark datasets. Jinjia Peng, Guangqi Jiang, Huibing Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Towards Adaptive Consensus Graph: Multi-View Clustering via Graph CollaborationabstractMulti-view clustering is a long-standing important task, however, it remains challenging to exploit valuable information from the complex multi-view data located in diverse high-dimensional spaces. The core issue is the effective collaboration of multiple views to holistically uncover the essential correlations between multi-view data through graph learning. Furthermore, it is indispensable for most existing methods to introduce an additional clustering step to produce the final clusters, which evidently reduces the uniform relationship between graph learning and clustering. Based on the above considerations, in this paper, we present a novel method named multi-view clustering via graph collaboration (MCGC). Based on the low-dimensional representation space developed by MCGC, it first perceives the correlations between samples in each individual view under the supervision of the Hilbert-Schmidt independence criterion (HSIC). Then, MCGC proposes learning a consensus graph by adaptively collaborating between all the views, which is able to uncover the essential structure of the multi-view data. Meanwhile, by imposing the rank constraint on the Laplacian matrix of the consensus graph to partition the multi-view data naturally into the required number of clusters, the optimal clustering results can be obtained directly without any postprocessing steps. Finally, the resulting optimization problem is solved by an alternating optimization scheme with guaranteed fast convergence. Extensive experiments on 5 benchmark multi-view datasets demonstrate that MCGC markedly outperforms the state-of-the-art baselines. Huibing Wang, Guangqi Jiang, Jinjia Peng, Ruoxi Deng, Xianping Fu |
IEEE Trans. Multim. | 3 |
| 2022 | Parallelism Network with Partial-aware and Cross-correlated Transformer for Vehicle Re-identificationabstractVehicle re-identification (ReID) aims to identify a specific vehicle in the dataset captured by non-overlapping cameras, which plays a great significant role in the development of intelligent transportation systems. Even though CNN-based model achieves impressive performance for the ReID task, its Gaussian distribution of effective receptive fields has limitations in capturing the long-term dependence between features. Moreover, it is crucial to capture fine-grained features and the relationship between features as much as possible from vehicle images. Guangqi Jiang, Huibing Wang, Jinjia Peng, Xianping Fu |
ICMR | 3 |
| 2022 | Progressive learning with multi-scale attention network for cross-domain vehicle re-identification
Yang Wang 0023, Jinjia Peng, Huibing Wang, Meng Wang 0001 |
Sci. China Inf. Sci. | 2 |
| 2022 | Learning latent features with local channel drop network for vehicle re-identification
Xianping Fu, Jinjia Peng, Guangqi Jiang, Huibing Wang |
Eng. Appl. Artif. Intell. | 2 |
| 2022 | Cooperative Refinement Learning for domain adaptive person Re-identification
Jinjia Peng, Guangqi Jiang, Huibing Wang |
Knowl. Based Syst. | 1 |
| 2022 | Eliminating cross-camera bias for vehicle re-identification
Jinjia Peng, Guangqi Jiang, Dongyan Chen, Huibing Wang, Xianping Fu |
Multim. Tools Appl. | 1 |
| 2022 | Reference-guided face inpainting with reference attention network
Jiazuo Yu 0002, Kai Li 0038, Jinjia Peng |
Neural Comput. Appl. | 3 |
| 2022 | Tensorial Multi-View Clustering via Low-Rank Constrained High-Order Graph LearningabstractMulti-view clustering aims to partition multi-view data into different categories by optimally exploring the consistency and complementary information from multiple sources. However, most existing multi-view clustering algorithms heavily rely on the similarity graphs from respective views and fail to comprehend multiple views holistically. Moreover, due to the noise and redundancy maintained in the original data, the original errors of multiple similarity graphs will continue to accumulate in the process of constructing consistent graphs. These situations always lead to the limitation to effective fuse the essential information from multiple views, which always influences the clustering performance and cries out for reliable solutions. Based on the above considerations, we propose a novel method termed Tensorial Multi-view Clustering (TMvC), which learns high-order graph by low-rank tensor constraint to uncover the essential information stored in multiple views. TMvC first learns the Laplacian graphs of all views and stacks them into a tensor which can be viewed as a high-order graph. With the high-order graph, consistency and complementary information from different views can be propagated smoothly across all views. Then, based on low-rank constraint, high-order graph is constrained in the horizontal and vertical directions to better uncover the inter-view and inter-class correlations between multi-view data, which is of vital importance for multi-view clustering. Extensive experiments on document and image datasets demonstrate that TMvC can achieve the state-of-the-art performance for multi-view clustering. Guangqi Jiang, Jinjia Peng, Huibing Wang, Zetian Mi, Xianping Fu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Learning Multiple Semantic Knowledge For Cross-Domain Unsupervised Vehicle Re-IdentificationabstractUnsupervised Vehicle re-identification (reID) aims at searching the similar vehicles’ images from large unlabelled datasets captured in a multiple camera network, which is still a challenging task. In this paper, a multiple semantic knowledge learning approach is proposed to exploit the potential similarity of unlabeled samples, which builds multiple clusters from different views automatically with different cues. Specially, different from some existing works focus on the knowledge of one view, for each vehicle in the target domain, different semantic knowledge could be learned with the proposed focal drop network and several different labels can be assigned according these knowledge, which would be employed to train the vehicle reID model jointly. In addition, due to the unreliability of pseudo labels assigned by the clustering, the hard triplet center loss is proposed to take the difference of intra-cluster and inter-cluster into consideration for better training the unsupervised framework to adapt the unknown domain. Comprehensive experimental results clearly demonstrate that our method achieves excellent performance on both VehicleID dataset and VeRi-776 dataset. Huibing Wang, Jinjia Peng, Guangqi Jiang, Xianping Fu |
ICME | 2 |
| 2021 | MSAV: An Unified Framework for Multi-view Subspace Analysis with View ConsistenceabstractWith the development of multimedia period, information is always caputred with multiple views, which causes a research upsurge on multi-view learning. It is obvious that multi-view data contains more information than those single view ones. Therefore, it is crucial to develop the multi-view algorithms to adapt the demand of many applications. Even though some excellent multi-view algorithms were proposed, most of them can only deal with the specific problems. To tacle this problem, this paper proposes an unified framework named Multi-view Subspace Analysis with View Consistence (MSAV), which provides an unified means to extend those single-view dimension reduciton algorithms into multi-view versions. MSAV first extends multi-view data into kernel space to avoid the problem caused by different dimensions of the data from multiple views. Then, we introduced a self-weighted learning strategy to automatically assign weights for all views according to their importance. Finally, in order to promote the consistence of all views, Hilbert-Schmidt Independence Criterion is adopted by MSAV. Furthermore, We conducted experiments on several benchmark datasets to verify the performance of MSAV. Huibing Wang, Guangqi Jiang, Jinjia Peng, Xianping Fu |
ICMR | 3 |
| 2021 | Graph-based Multi-view Binary Learning for image clustering
Guangqi Jiang, Huibing Wang, Jinjia Peng, Dongyan Chen, Xianping Fu |
Neurocomputing | 3 |
| 2021 | Discriminative feature and dictionary learning with part-aware model for vehicle re-identification
Huibing Wang, Jinjia Peng, Guangqi Jiang, Fengqiang Xu, Xianping Fu |
Neurocomputing | 2 |
| 2021 | Generalized multiple sparse information fusion for vehicle re-identification
Jinjia Peng, Guangqi Jiang, Huibing Wang |
J. Vis. Commun. Image Represent. | 1 |
| 2021 | Scale-aware feature pyramid architecture for marine object detection
Fengqiang Xu, Huibing Wang, Jinjia Peng, Xianping Fu |
Neural Comput. Appl. | 3 |
| 2020 | Unsupervised Vehicle Re-identification with Progressive AdaptationabstractVehicle re-identification (reID) aims at identifying vehicles across different non-overlapping cameras views. The existing methods heavily relied on well-labeled datasets for ideal performance, which inevitably causes fateful drop due to the severe domain bias between the training domain and the real-world scenes; worse still, these approaches required full annotations, which is labor-consuming. To tackle these challenges, we propose a novel Progressive Adaptation Learning method for vehicle reID, named PAL, which infers from the abundant data without annotations. For PAL, a data adaptation module is employed for source domain, which generates the images with similar data distribution to unlabeled target domain as “pseudo target samples”. These pseudo samples are combined with the unlabeled samples that are selected by a dynamic sampling strategy to make training faster. We further proposed a weighted label smoothing (WLS) loss, which considers the similarity between samples with different clusters to balance the confidence of pseudo labels. Comprehensive experimental results validate the advantages of PAL on both VehicleID and VeRi-776 dataset. Jinjia Peng, Yang Wang 0023, Huibing Wang, Zhao Zhang 0001, Xianping Fu, Meng Wang 0001 |
IJCAI | 1 |
| 2020 | Purifying real images with an attention-guided style transfer network for gaze estimation
Xianping Fu, Yuxiao Yan, Jinjia Peng, Huibing Wang |
Eng. Appl. Artif. Intell. | 4 |
| 2020 | Cross domain knowledge learning with dual-branch adversarial network for vehicle re-identification
Jinjia Peng, Huibing Wang, Fengqiang Xu, Xianping Fu |
Neurocomputing | 1 |
| 2020 | Vehicle re-identification using multi-task deep learning network and spatio-temporal model
Jinjia Peng, Fengqiang Xu, Xianping Fu |
Multim. Tools Appl. | 1 |
| 2019 | Learning multi-region features for vehicle re-identification with context-based ranking method
Jinjia Peng, Huibing Wang, Xianping Fu |
Neurocomputing | 1 |
| 2019 | Co-regularized multi-view sparse reconstruction embedding for dimension reduction
Huibing Wang, Jinjia Peng, Xianping Fu |
Neurocomputing | 2 |
| 2019 | Multi-feature distance metric learning for non-rigid 3D shape retrieval
Huibing Wang, Haohao Li, Jinjia Peng, Xianping Fu |
Multim. Tools Appl. | 3 |
| 2019 | Guiding intelligent surveillance system by learning-by-synthesis gaze estimation
Yuxiao Yan, Jinjia Peng, Zetian Mi, Xianping Fu |
Pattern Recognit. Lett. | 3 |
| 2018 | Refining Synthetic Images with Semantic Layouts by Adversarial TrainingabstractRecently, progress in learning-by-synthesis has proposed training models on synthetic images, which can effectively reduce the cost of manpower and material resources. However, learning from synthetic images still fails to achieve the desired performance compared to naturalistic images due to the different distribution of synthetic images. In an attempt to address this issue, previous methods were to improve the realism of synthetic images by learning a model. However, the disadvantage of the method is that the distortion has not been improved and the authenticity level is unstable. To solve this problem, we put forward a new structure to improve synthetic images, via the reference to the idea of style transformation, through which we can efficiently reduce the distortion of pictures and minimize the need of real data annotation. We estimate that this enables generation of highly realistic images, which we demonstrate both qualitatively and with a user study. We quantitatively evaluate the generated images by training models for gaze estimation. We show a significant improvement over using synthetic images, and achieve state-of-the-art results on various datasets including MPIIGaze dataset. Yuxiao Yan, Jinjia Peng, HaoHui Wei, Xianping Fu |
ACML | 3 |
| 2018 | Learning a gaze estimator with neighbor selection from large-scale synthetic eye images
Yafei Wang 0004, Xueyan Ding, Jinjia Peng, Jiming Bian, Xianping Fu |
Knowl. Based Syst. | 4 |