Huibing Wang

dblp:150/0713 · DBLP profile ↗
← Back
120ranked-venue papers
19as first author
95since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 69 · 9 first-author · 51 since 2021Graphics, computer vision, multimedia, augmented reality and games · 55 · 9 first-author · 49 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Transferable adversarial queries on person re-identification via dynamic bilateral tuning
Zeze Tao, Jinjia Peng, Huibing Wang
Expert Syst. Appl.4
2026 Instance-Guided Scene Adaptation for Unsupervised Person Search
abstract
Unsupervised Domain Adaptation (UDA) is a challenging task in person search. It adapts a well-trained model from a labeled source domain to an unlabeled target domain for privacy and efficiency. Currently, most of the state-of-the-art UDA person search methods adopt multi-scale feature alignment techniques to learn domain-invariant representations. However, person search is a multi-granularity task, and such an indiscriminate method of bridging the differences between domains misleads the identity learning process, which significantly limits the model's performance. In this paper, we propose an Instance-Guided Scene Adaptation (IGSA) framework by eradicating scene disparities and focusing the tasks on instances, effectively eliminating the contradiction between person search and domain adaptation. In IGSA, a Scene-Aware Bidirectional Filter (SABF) is designed to divide the image features into background and foreground to perform bidirectional modulations, thereby achieving simultaneous scene elimination and instance enhancement. To further improve the reliability of identity learning, we also propose an Instance Consistency Contrastive Learning (ICCL) method. By performing cross-epoch updates on the instance-level memory bank and re-initializing the cluster-level memory bank, the problem of inconsistent training across epochs caused by instance identity drift can be alleviated. Through the above designs, our method can achieve state-of-the-art performance on two benchmark datasets, with 82.1% mAP and 83.8% top-1 on the CUHK-SYSU dataset and 41.1% mAP and 82.3% top-1 on the PRW dataset, which is even better than some supervised methods.
Huibing Wang, Jinjia Peng, Xianping Fu, Jiqing Zhang
AAAI2
2026 Localization-Anchored Instance Discrimination for Domain Adaptive Person Search
abstract
Domain-adaptive person search (DAPS) aims to transfer pedestrian detection and re-identification capabilities from a labeled source domain to an unlabeled target domain, yet faces critical challenges from domain shift: semantic confusion among overlapping instances, over-reliance on shallow features for look-alike targets, and poor discriminability of small-scale instances. To address these issues, we propose the Localization-Anchored Instance Discrimination (LAID) framework, which leverages spatial relationships between bounding boxes as auxiliary signals to enhance instance identity learning. LAID integrates three complementary strategies: 1) Cost-Aware Instance Matching (CAIM) uses IoU-based global optimal assignment to align current detections with historical identities, reducing overlap-induced misassociations; 2) Dual-Scope Contrastive Learning (DSCL) combines spatial separation constraints (for geometrically distant pairs) with global contrastive learning, prompting the model to learn deep discriminative features beyond superficial similarities; 3) Task-Sensitivity Alignment (TSA) aligns confidence distributions of detection and ReID heads via KL divergence, ensuring consistent pseudo-label generation. Extensive experiments on CUHK-SYSU and PRW datasets demonstrate that LAID outperforms state-of-the-art DAPS methods, validating its effectiveness in mitigating domain shift and narrowing the performance gap between supervised and domain-adaptive person search.
Linfeng Qi 0001, Huibing Wang, Jinjia Peng, Jiqing Zhang
AAAI2
2026 Prompting Adversarial Transferability via Path Flatness Attack
abstract
Deep neural networks are susceptible to adversarial examples, which induce incorrect predictions through imperceptible perturbations. Transfer-based attacks create adversarial examples for surrogate models and transfer these examples to target models under black-box scenarios. Recent studies have established a strong correlation between the geometric properties of loss landscapes and the transferability of adversarial examples, demonstrating that flatter loss surfaces consistently yield superior transferability. However, we identify that these methods fail to account for the loss landscape flatness along the path from the current point to local minima, resulting in poor transferability. To address this, this paper constructs a novel Path Flatness Attack (PFA) method to significantly enhance the transferability of adversarial examples. Specifically, this paper proposes a novel path flatness indicator that not only evaluates the flatness in local minima regions but also explicitly quantifies the loss surface geometry along the trajectory from the current point to the minimum. Furthermore, we incorporate the path flatness indicator into the attack process, integrating penalties over low-loss points along the path while maximizing the loss function, thereby explicitly flattening the loss landscape. Extensive experiments demonstrate that PFA consistently achieves state-of-the-art attack performance across all experimental settings.
Zeze Tao, Jinjia Peng, Huibing Wang
AAAI3
2026 Conditional Prompt Learning via Degradation Perception for Underwater Image Enhancement
abstract
Underwater Image Enhancement (UIE) focuses on improving visual quality from various underwater scenes. Existing methods simplistically treat various degradations as homogeneous, disregarding their intrinsic connections and causing models to blindly learn, resulting in conflicting optimization goals and visual distortions. To address above limitations, we propose a Conditional Prompt Learning via Degradation Perception (CPLDP) model, which employs conditional prompt as degradation perception priors and guides underwater image enhancement. Specifically, we show that the natural language prompts not only promote distinguishing different degraded images, but also aid in exploring more details with semantic information. Therefore, our method generates five key degradation prompts (green/blue/green-blue color casts, uneven illumination and haze) with conditional prompt learning. Subsequently, considering the intrinsic relationships among different degradations, we employ degradation perceptions as priors and fine-tune the learning strategy to enhance underwater images. During training, an adaptive loss function with multi-degradations is designed, allowing it to effectively handle the task conflicts among multiple underwater degradations. Additionally, we conduct a human visual-based underwater dataset with various degradation types by subjective statistics. Extensive experiments on both full-reference and non-reference datasets demonstrate that our CPLDP can achieve better visual results and outperforms state-of-the-art UIE methods across various degradation scenarios.
Mingze Yao, Zhiying Jiang, Xianping Fu, Huibing Wang
AAAI4
2026 T-Control: An Efficient Dynamic Tensor Rematerialization System for DNN Training
Junmin Xiao, Xiaochuan Deng, Huibing Wang, Yunfei Pang, Guangming Tan
ASPLOS (2)4
2026 Hierarchical structure-guided incomplete multi-view tensor clustering
Chunyan Yang, Wengeng Chen, Huibing Wang
Appl. Intell.7
2026 Partially view-aligned clustering via data recoupling and elastic bi-consistency learning
Jiongcheng Zhu, Jingbo Tan, Huibing Wang, Yong Zhang 0030
Expert Syst. Appl.4
2026 FedDSKI: Improving server-side model via dual-stage knowledge isolation in personalized federated learning
Xiaorui He, Jinjia Peng, Zhen Wang 0017, Huibing Wang
Future Gener. Comput. Syst.5
2026 Efficient Vision Transformer with Token Sparsification for Event-Based Object Tracking
Jiqing Zhang, Xin Yang 0011, Haoming Tang, Yuanchen Wang, Huibing Wang, Xianping Fu
Int. J. Comput. Vis.6
2026 Hybrid anchor graph learning and tensorized spectral embedding fusion for multi-view clustering
Guangqi Jiang, Wangjie Chen, Yi Liu 0038, Lin Shi 0007, Jinjia Peng, Huibing Wang
Neurocomputing6
2026 Mitigating modal discrepancies for visible-infrared person re-identification via high-order nonlinear constraint
Junyu Liu, Yanzhen Xiong, Jinjia Peng, Huibing Wang
Knowl. Based Syst.4
2026 One-step multi-view graph clustering via bottom-up structural learning
Huibing Wang, Yong Zhang 0030
Pattern Recognit.3
2026 BCDnet: Balanced coupling and decoupling network for person search
Zhengjie Lu, Jinjia Peng, Huibing Wang, Xianping Fu
Pattern Recognit.3
2026 Bridging the gap : Learning adaptive knowledge transition for lifelong person re-identification
Jinjia Peng, Jican Tan, Huibing Wang, Xianping Fu
Pattern Recognit.4
2026 Global aggregated gradient-guided adversarial attacks for person re-identification
Zeze Tao, Jinjia Peng, Huibing Wang
Pattern Recognit.4
2026 Visual In-Context Learning for Underwater Image Restoration
abstract
Underwater images often exhibit common visual degradations, such as color distortion, loss of details, and reduced sharpness, which inevitably compromise the effectiveness of underwater vision tasks. However, most underwater image restoration methods solely focus on learning degradation features from raw images, neglecting the incorporation of additional contextual information to guide restoration, which limits the capability of deep models to restore image quality. In this paper, we propose Visual In-Context Learning (VICL) for underwater image restoration, which leverages degradation information from context to improve image quality. In VICL, Degraded Context Extraction Block (DCEB) employs a self-attention mechanism to extract degradation information from context. In addition, Context Spatial Feature Fusion Block (CSFFB) consists of a Degraded Context Guidance Block (DCGB) and a Multi-Feature Fusion Block (MFFB). DCGB employs a cross-attention mechanism to fuse degraded context with spatial features for guiding underwater image restoration. MFFB replaces traditional encoder-decoder skip connections to better coordinate feature fusion. Extensive experiments on multiple underwater image benchmarks demonstrate that VICL outperforms state-of-the-art methods both quantitatively and visually. The code is available at:https://github.com/zhangao668/VICL.
Guangqi Jiang, Yi Liu 0038, Huibing Wang, Shoukun Xu
IEEE Signal Process. Lett.4
2026 Dynamic Fusion Network Driven Private-Consensus Learning for Multiview Clustering
Jiongcheng Zhu, Yong Zhang 0030, Huibing Wang
IEEE Trans. Comput. Soc. Syst.5
2026 Reliable Feature Imputation With Cross-View Relation Transfer for Deep Incomplete Multi-View Classification
abstract
Incomplete Multi-view Classification has sparked widespread interest in recent years, since multi-view data suffering from missing values are ubiquitous in real-world scenarios. While many imputation-based methods recover missing data by exploiting inter-sample structural information within individual views, they are inherently susceptible to unreliable or noisy samples, which can lead to low-quality imputation and degrade classification accuracy. Therefore, it is a challenge to effectively mine the multi-stage complex correlations for incomplete multi-view data to achieve reliable imputation and obtain discriminative representation. To address these issues, we present a novel imputation-based approach called Reliable Feature Imputation with Cross-view Relation Transfer for Deep Incomplete Multiview Classification (RFI-IMvC). Our framework fully exploits inter-view and intra-view structural information in multi-stage manner. Specifically, we propose a novel cross-view relation transfer strategy to recover reliable neighbor relationships and achieve high-quality imputation for missing data. Besides, to fully exploit the structural information in reconstructed multi-view data, we develop a dual graph learning module to mine high-order semantic correlation and facilitate interactions of complementary information from instances linked by hyperedge. Finally, inspired by prototype learning, we incorporate a class--level representation loss to further promote intra-class compactness. Extensive experiments on 7 real-world datasets demonstrate that our method outperforms state-of-the-art methods.
Guangqi Jiang, Haodong Hou, Yi Liu 0038, Jinjia Peng, Huibing Wang
IEEE Trans. Circuits Syst. Video Technol.5
2026 A Lightweight Polarization-Guided Plug-In for Underwater Image Enhancement
abstract
Underwater images play a vital role in marine exploration, but are often severely degraded due to complex imaging conditions, including color distortion, haze effects, and non-uniform illumination. Existing deep learning-based enhancement methods predominantly rely on conventional RGB sensors, which struggle to distinguish between scattered and reflected light, thereby limiting enhancement performance. Polarization imaging, with its capability to capture directional light information, offers promising potential for underwater image enhancement. In this paper, we propose a lightweight yet effective polarization feature extractor that captures global spatial cues from polarization images. Additionally, we design a polarization-guided feature integration module that adaptively enhances the representational capacity of RGB features. Notably, the proposed module is plug-in and can be seamlessly integrated into existing RGB-based enhancement networks. Extensive experiments across multiple datasets demonstrate that incorporating polarization information significantly improves enhancement performance, highlighting its effectiveness as a valuable cue for underwater image enhancement. The code and pretrained models are at https://github.com/jgy0/UPGD.
Guangyao Ju, Jiqing Zhang, Jingqi Zang, Zetian Mi, Xin Yang 0011, Huibing Wang, Jiarui Fan, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.8
2026 Interaction-Driven Edge Crisping for Underwater Salient Object Detection
abstract
Underwater salient object detection (USOD) faces greater challenges than general scenes due to the edge blurring which is caused by light absorption and scattering in water. Existing methods employ unrefined edge feature to perform unidirectional guidance on saliency feature, resulting in the coarse edge of saliency map. To address this issue, we propose a novel interaction-driven edge crisping network (IDENet) for underwater salient object detection. IDENet facilitates the bidi-rectional modulation of inter-features and the self-refinement of intra-feature, generates crisp saliency map and edge map. In IDENet, the interaction-driven edge guidance module (IDEGM) is designed to utilize cross-feature interaction by leveraging their correlations, facilitating saliency feature’s awareness of edge information, mitigating the interference of non-salient objects in edge feature. To learn more accurate edge region of the salient object, the edge intersection-and-union loss function (EIUL) is introduced to restrict the intersection and union of predicted saliency maps and edge maps to prevent over-expansion or under-contraction. Experimental results on two latest underwater datasets demonstrate the superiority of the proposed method over the state-of-the-art models. The source code of our method will be made available at https://github.com/ UnderwaterVisionMZTdlmu/IDENet.
Zetian Mi, Shuaiyong Jiang, Guanxi Li, Jiqing Zhang, Huibing Wang, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.6
2026 Hierarchical Sequential Context Modeling for High-Fidelity Image Inpainting
abstract
Image inpainting aims to restore missing regions by leveraging surrounding spatial context, where nearby pixels provide crucial structural cues and distant regions offer complementary semantic guidance. To jointly model these complementary dependencies, this paper proposes Hierarchical Sequential Context Modeling (HSCM), a novel inpainting framework that employs state-space models for multi-scale autoregressive sequence modeling. Unlike existing single-scale SSM-based approaches, HSCM explicitly separates pixel-level and semanticlevel modeling into two complementary branches. The Local Perception Unit preserves fine-grained textures, and the Global Compensation Unit propagates high-level semantics across patches to enhance overall coherence. The asynchronous hierarchical design first reconstructs local textures and then performs semantic compensation, achieving notable performance gains with minimal computational overhead. Leveraging its four-directional architecture, HSCM maintains linear computational growth with spatial resolution and effectively establishes a comprehensive global receptive field. Furthermore, a Cross-Gated Feedforward Network is proposed to alleviate patch boundary artifacts and enhance inter-channel feature consistency. Built upon a multi-scale encoder–decoder architecture, HSCM delivers state-of-the-art inpainting quality and robust generalization across diverse benchmarks, including CelebA-HQ, FFHQ, Paris Street View, and Places2.
Zexuan Sun, Jinjia Peng, Mengkai Li, Huibing Wang
IEEE Trans. Circuits Syst. Video Technol.4
2026 Bi-Level Inter-Modality Modulation for Unsupervised Visible-Infrared Person Re-Identification
abstract
The task of unsupervised visible–infrared person re-identification (USL-VI-ReID) aims to retrieve cross-modal pedestrian images without manual annotations. The key challenge lies in achieving semantic alignment to resolve modality bias in the absence of real labels. However, existing methods overly rely on single-modal information in the process of pseudo-label generation without considering cross-modal associations, making it difficult to bridge the modality gap between visible and infrared images. To address these issues, this paper proposes a Bi-level Inter-Modal Modulation Network (BIMM-Net), which employs multi-level cluster structure optimization as a core strategy to drive the establishment of cross-modal semantic associations, ultimately achieving cross-modal alignment at the feature representation level. Specifically, we construct a novel intermediary modality GrayMix from visible images to enhance model robustness against color variations and alleviate modality gaps. To filter out noise in cross-modal matching and establish a shared semantic space between visible and infrared modalities, we further develop a Ternary Pairs Calibration-Convergence module designed for filtering noise from visible-infrared cluster matching, on this basis constructing fused mixture clusters. Building on this mixture cluster space, an Heterogeneous-Isomorphic Alignment Loss is also designed to align the feature distributions of the three modalities, reinforcing cross-modal semantic consistency. In addition, we present a Cross-modal Neighborhood Consistency Clustering method, which facilitates the formation and propagation of cross-modal clusters by selecting high-confidence cross-modal neighbor pairs and refining feature distances. Ultimately, BIMM-Net through the joint modeling of bi-level clustering enables multiple levels to guide each other in refining cross-modal structures, thereby effectively establishing the semantic associations between visible and infrared modalities. Extensive experiments validate the superior performance of the proposed framework, achieving state-of-the-art results in USL-VI-ReID. The source code of this paper is available at: https://github.com/liujuny5920/DIMM-Net.
Jinjia Peng, Junyu Liu, Xutao Zuo, Zeze Tao, Huibing Wang
IEEE Trans. Inf. Forensics Secur.5
2026 Unsupervised Lifelong Person Re-Identification via Affinity Harmonization
abstract
Lifelong Person Re-Identification (LReID) seeks to continuously train models across multiple target domains, enabling effective generalization in both known and unseen domains. Achieving a balance between “plasticity” (the ability to adapt to new knowledge) and “stability” (the capacity to prevent forgetting) is crucial in lifelong learning. However, most existing LReID methods primarily focus on enhancing model stability or plasticity, often neglecting the critical balance between them. Moreover, current LReID approaches largely rely on supervised learning, which necessitates large-scale pre-labeled datasets—a process that is both time-consuming and labor-intensive in practical applications. To address these challenges, this article proposes an Unsupervised LReID approach called the Affinity Harmonization Network (AHN). AHN includes an Old Domain Affinity Constraint (ODAC) module, which builds an expert model for the old domain to provide affinity relationships as references. This helps limit changes among old representations, enabling the model to integrate new knowledge while preserving compatibility with previous representations. To harmonize stability and plasticity while guiding the model in acquiring new knowledge, AHN incorporates a Current Domain Affinity Guidance (CDAG) module. This module builds an expert model for the new domain and uses the generated affinity relationships to assist in training the model. Furthermore, this article proposes the Old Domain Intra-class Variance Constraint (OIVC) module, which mitigates potential deviations in the intra-class variance of legacy samples by limiting the distance between replay samples and old domain camera prototypes. Extensive experiments demonstrate that our method achieves significant performance improvements over existing unsupervised lifelong ReID methods, with an average gain of 5.3% in mAP and 5.2% in Rank-1 accuracy.
Jican Tan, Jinjia Peng, Songyu Zhang, Zhen Wang 0017, Huibing Wang
ACM Trans. Multim. Comput. Commun. Appl.5
2026 Tensorized Fine-Grained Incomplete Multiview Clustering via Intrinsic Structure Recovery
abstract
Incomplete multiview clustering (IMVC) focuses on exploiting the complementary and consistent information from multiple incomplete views for dividing unlabeled multiview data into corresponding clusters. Most existing methods seek to recover the missing samples of the views while inevitably having an influence on the intrinsic structure of the original space. Moreover, previous IMVC algorithms treat the samples of each view equally, which learns the consensus representation in a view-level manner and thus neglects that different views contribute to each individual sample diversely. These situations are a limitation to effective recovery of absent samples and fuse the heterogeneous information, which leads to suboptimal clustering performance. To tackle the above issues, this article proposes a novel approach called tensorized fine-grained IMVC via intrinsic structure recovery (TFIR), which flexibly recovers the intrinsic structure of the original data and attains the unified representation based on the fine-grained fusion strategy. Specifically, TFIR infers the incomplete data via available instances’ relations to improve the accuracy of learned intersample-specific representations. Afterward, TFIR stacks the specific representations of multiple views into a tensor to preserve the intrasample consistency. Finally, TFIR explores the complementarity among various views from the fine-grained sample perspective to obtain a rich underlying structure for clustering. Experimental results on the eight benchmark datasets clearly verify the remarkable superiority of TFIR compared with the most state-of-the-art methods.
Huibing Wang, Luyan Cui, Mingze Yao, Xianping Fu
IEEE Trans. Syst. Man Cybern. Syst.1
2025 Anchor Learning with Potential Cluster Constraints for Multi-view Clustering
abstract
Anchor-based multi-view clustering has received extensive attention due to its efficient performance. Existing methods only focus on how to dynamically learn anchors from the original data and simultaneously construct anchor graphs describing the relationships between samples and perform clustering, while ignoring the reality of anchors, i.e., high-quality anchors should be generated uniformly from different clusters of data rather than scattered outside the clusters. To deal with this problem, we propose a noval method termed Anchor Learning with Potential Cluster Constraints for Multi-view Clustering (ALPC) method. Specifically, ALPC first establishes a shared latent semantic module to constrain anchors to be generated from specific clusters, and subsequently, ALPC improves the representativeness and discriminability of anchors by adapting the anchor graph to capture the common clustering center of mass from samples and anchors, respectively. Finally, ALPC combines anchor learning and graph construction into a unified framework for collaborative learning and mutual optimization to improve the clustering performance. Extensive experiments demonstrate the effectiveness of our proposed method compared to some state-of-the-art MVC methods.
Yawei Chen, Huibing Wang, Jinjia Peng, Yang Wang 0023
AAAI2
2025 CDE-Learning: Camera Deviation Elimination Learning for Unsupervised Person Re-identification
abstract
Unsupervised Person Re-identification (Re-ID) aims to identify the same person shot from non-overlapping cameras without any annotated data. In this task, attributes such as contrast, saturation, and resolution of the camera cause the deviation in target features. Since the camera label is readily available, they are employed to achieve the constraints across cameras and smooth the deviations during the model training phase. However, features from the same camera are prone to generating false positives due to the identical camera properties, which induce camera deviations on pseudo-label assignment. To address this problem, this paper proposes a novel camera-unbiased method named Camera Deviation Elimination Learning (CDE-Learning). In CDE-Learning, the Camera Deviation Compensation (CDC) module is designed to align data distributions from disparate cameras to decouple camera information from identity information during the pseudo-label allocation. Our Camera Deviation Balancing (CDB) module integrates different camera constraints in a united loss and adjusts camera constraints by constructing contrastive pairs between intra-camera and inter-camera. After explicit constraints, the Camera Attribution Auxiliary (CAA) task predicts whether a pair of images originates from the same camera to implicitly enhance the capacity to distinguish the camera deviation. We demonstrated the superior performance of the proposed CDE-Learning on benchmark datasets.
Jinjia Peng, Songyu Zhang, Huibing Wang
AAAI3
2025 Unsupervised Domain Adaptive Person Search via Dual Self-Calibration
abstract
Unsupervised Domain Adaptive (UDA) person search focuses on employing the model trained on a labeled source domain dataset to a target domain dataset without any additional annotations. Most effective UDA person search methods typically utilize the ground truth of the source domain and pseudo-labels derived from clustering during the training process for domain adaptation. However, the performance of these approaches will be significantly restricted by the disrupting pseudo-labels resulting from inter-domain disparities. In this paper, we propose a Dual Self-Calibration (DSCA) framework for UDA person search that effectively eliminates the interference of noisy pseudo-labels by considering both the image-level and instance-level features perspectives. Specifically, we first present a simple yet effective Perception-Driven Adaptive Filter (PDAF) to adaptively predict a dynamic filter threshold based on input features. This threshold assists in eliminating noisy pseudo-boxes and other background interference, allowing our approach to focus on foreground targets and avoid indiscriminate domain adaptation. Besides, we further propose a Cluster Proxy Representation (CPR) module to enhance the update strategy of cluster representation, which mitigates the pollution of clusters from misidentified instances and effectively streamlines the training process for unlabeled target domains. With the above design, our method can achieve state-of-the-art (SOTA) performance on two benchmark datasets, with 80.2% mAP and 81.7% top-1 on the CUHK-SYSU dataset, with 39.9% mAP and 81.6% top-1 on the PRW dataset, which is comparable to or even exceeds the performance of some fully supervised methods.
Linfeng Qi 0001, Huibing Wang, Jiqing Zhang, Jinjia Peng
AAAI2
2025 Boosting Adversarial Transferability via Residual Perturbation Attack
abstract
Deep neural networks are susceptible to adversarial examples while suffering from incorrect predictions via imperceptible perturbations. Transfer-based attacks create adversarial examples for surrogate models and transfer these examples to target models under black-box scenarios. Recent studies reveal that adversarial examples in flat loss landscapes exhibit superior transferability to alleviate overfitting on surrogate models. However, the prior arts overlook the influence of perturbation directions, resulting in limited transferability. In this paper, we propose a novel attack method, named Residual Perturbation Attack (ResPA), relying on the residual gradient as the perturbation direction to guide the adversarial examples toward the flat regions of the loss function. Specifically, ResPA conducts an exponential moving average on the input gradients to obtain the first moment as the reference gradient, which encompasses the direction of historical gradients. Instead of heavily relying on the local flatness that stems from the current gradients as the perturbation direction, ResPA further considers the residual between the current gradient and the reference gradient to capture the changes in the global perturbation direction. The experimental results demonstrate the better transferability of ResPA than the existing typical transfer-based attack methods, while the transferability can be further improved by combining ResPA with the current input transformation methods. The code is available at https://github.com/ZezeTao/ResPA.
Jinjia Peng, Zeze Tao, Huibing Wang, Meng Wang 0001, Yang Wang 0023
ICCV3
2025 Consensus-Guided Incomplete Multi-view Clustering via Cross-view Affinities Learning
abstract
Incomplete multi-view clustering (IMC) has garnered substantial attention due to its capacity to handle unlabeled data. Existing methods predominantly explore pairwise consistency between every two views. However, such consistency is highly susceptible to missing samples and outliers within a certain view and thus deviates from the true clustering distribution. Moreover, dual-view interaction neglects the collaboration effects of multiple views, making it challenging to capture the holistic characteristics across views. In response to these issues, we propose a novel Consensus-Guided Incomplete Multi-view Clustering via Cross-view Affinities Learning (CAL). Specifically, CAL reconstructs views with available instances to mine sample-wise affinities and harness comprehensive content information within views. Subsequently, to extract clean structural information, CAL imposes a structured sparse constraint on the representation tensor to eliminate biased errors. Furthermore, by integrating the consensus representation into a representation tensor, CAL can employ high-order interaction of multiple views to depict the semantic correlation between views while acquiring a unified structural graph across multiple views. Extensive experiments on seven benchmark datasets demonstrate that CAL outperforms some state-of-the-art methods in clustering performance. The code is available at https://github.com/whbdmu/CAL.
Huibing Wang, Jinjia Peng, Yawei Chen, Mingze Yao, Xianping Fu, Yang Wang 0023
IJCAI2
2025 Stabilizing Holistic Semantics in Diffusion Bridge for Image Inpainting
abstract
Image inpainting aims to restore the original image from a damaged version. Recently, a special type of diffusion bridge model has achieved promising performance by directly mapping the degradation process and restoring corrupted images through the corresponding reverse process. However, due to the lack of explicit semantic priors during the denoising process, the inpainted results typically exhibit inferior context-stability and semantic consistency. To this end, this paper proposes a novel Global Structure-Guided Diffusion Bridge framework (GSGDiff), which incorporates an additional structure restorer to stabilize the generation of holistic semantics. Specifically, to acquire richer semantic structure priors, this paper proposes a posterior sampling approach that captures semantically global and consistent structures at each timestep, efficiently integrating them into the texture generation through the corresponding guidance module. Additionally, considering the characteristics of diffusion models with low denoising levels at larger timesteps, this paper proposes a semantic fusion schedule to avoid noise interference by reducing the weight of ineffective guided semantics in the early stages. By applying the proposed posterior sampling to the texture denoising process, GSGDiff can achieve more stable and superior inpainting results over competitive baselines. Experiments on Places2, Paris Street View and CelebA-HQ datasets validate the efficacy of the proposed method.
Jinjia Peng, Mengkai Li, Huibing Wang
IJCAI3
2025 Scalable Multi-view Clustering based on Tight Anchor Distribution
Yawei Chen, Huibing Wang, Mingze Yao, Jinjia Peng, Guangqi Jiang, Jiqing Zhang
ACM Multimedia2
2025 Dual-Constraint Multi-view Fuzzy Clustering with Scalable Anchor Graph Learning
Luyan Cui, Huibing Wang, Yawei Chen, Mingze Yao, Xianping Fu, Jiqing Zhang
ACM Multimedia2
2025 Prior-oriented Anchor Learning with Coalesced Semantics for Multi-View Clustering
abstract
Anchor-based multi-view clustering has received a lot of attention due to its efficiency in handling large-scale datasets. However, existing methods rely on penalty-based regularization terms in anchor graphs to handle noise and outliers but overlook the role of consistent semantics in label contributions, failing to effectively mitigate their impact and potentially deviating from actual data distributions. In addition, most strategies use adaptive anchor learning without considering the veracity of anchor selection and the lack of sufficient semantic support in modeling semantic consistency, which leads to anchors deviating from the clustering center. To solve the above problems, we propose a novel method called Prior-oriented Anchor Learning with Coalesced Semantics for Multi-View Clustering (PALCS). Specifically, PALCS strips out inconsistent semantics from anchor graphs to be processed separately through coalesced semantics and highlights consistent semantics to reveal the underlying shared structure of the data. Moreover, PALCS enhances the semantic consistency and discriminative properties of anchors by directing them to be evenly distributed across clusters via the prior matrix. Finally, the clustering labels are directly obtained by non-negative matrix decomposition, avoiding additional post-processing steps. Extensive experimental evidence demonstrates the superiority of our method compared to state-of-the-art methods.
Jinjia Peng, Tianhang Cheng, Guangqi Jiang, Huibing Wang
ACM Multimedia4
2025 Eye-based Emotion Recognition via Event-Driven Sparse Transformers
abstract
Event-driven eye-based emotion recognition has attracted increasing attention due to the high temporal resolution and dynamic range inherent to event cameras. The intrinsic spatial sparsity of event data, combined with the eye-based emotion recognition task's reliance on localized features such as eyebrows and eyelids, makes it intuitive and efficient to discard less informative regions. However, integrating such sparsification into CNNs remains challenging due to their reliance on dense grid-based operations. In this paper, we propose an efficient vision transformer framework for eye-based emotion recognition with event cameras. Specifically, we present window selection and token selection schemes tailored for event data and eye-based emotion recognition, which can diminish computing demands while enhancing performance. Firstly, we estimate the importance of all local windows and discard those with limited information, reducing computational cost while emphasizing attention on the periocular region. Secondly, we further introduce an adaptive token pruning mechanism that jointly evaluates the input event data and tokens to predict a binary decision mask, identifying and discarding uninformative tokens. Extensive experiments validate that the proposed approach outperforms existing state-of-the-art methods in accuracy by a significant margin.
Zixuan Wan, Jiqing Zhang, Yafei Wang 0004, Zetian Mi, Xin Yang 0011, Xianping Fu, Huibing Wang
ACM Multimedia9
2025 Spatiotemporal Consensus with Scene Prior for Unsupervised Domain Adaptive Person Search
abstract
Person Search aims to locate query persons in gallery scene images, but faces severe performance degradation under domain shifts. Unsupervised domain adaptation transfers knowledge from the labeled source domain to the unlabeled target domain and iteratively rectifies the pseudo-labels. However, the pseudo-labels are inevitably contaminated by the source-biased model, which misleads the training process. This, in turn, reduces the quality of the pseudo-labels themselves and ultimately affects the search performance. In this paper, we propose a Spatiotemporal Consensus with Scene Prior (STCSP) framework that effectively eliminates the interference of noise on pseudo-labels, establishes positive feedback, and thus gradually bridging the domain gap. Firstly, STCSP uses a Spatiotemporal Consensus pipeline to suppress the noise from being mixed into the pseudo-labels. Secondly, leveraging the scene prior, STCSP employs our designed Iterative Bilateral Extremum Matching method to prevent the occurrence of some incorrect pseudo-labels. Thirdly, we propose a Scene Prior Contrastive Learning module, which encourages the model to directly acquire the scene prior knowledge from the target domain, thereby mitigating the generation of noise. By suppressing noise contamination, avoiding noise occurrence and mitigating noise generation, our framework achieves state-of-the-art performance on two benchmark datasets, PRW with 50.2% mAP and CUHK-SYSU with 87.0% mAP.
Huibing Wang, Jinjia Peng
NeurIPS2
2025 Consensus guided incomplete multi-view clustering via geometric consistency learning
Huibing Wang, Mingze Yao, Yawei Chen, Jinjia Peng, Guangqi Jiang, Xianping Fu
Appl. Intell.1
2025 Region-guided spatial feature aggregation network for vehicle re-identification
Yanzhen Xiong, Jinjia Peng, Zeze Tao, Huibing Wang
Eng. Appl. Artif. Intell.4
2025 Detail-focused and polarization-guided multi-modality fusion for underwater image clarity enhancing
Mingze Yao, Huibing Wang, Xianping Fu
Eng. Appl. Artif. Intell.2
2025 Dynamic multiobjective optimization via an improved r-dominance relation and a novel prediction approach
Yaru Hu, Junwei Ou, Huibing Wang, Jinhua Zheng, Shengxiang Yang
Expert Syst. Appl.3
2025 Parameter-free multi-view clustering via refined tensor learning
Yuemeng Huang, Chunyan Yang, Wengeng Chen, Huibing Wang
Neurocomputing9
2025 Scene-intra deep mining via shifted instance refinement for weakly supervised person search
Shenghui Yin, Jinjia Peng, Zhen Wang 0017, Huibing Wang
Knowl. Based Syst.4
2025 MVG-FD: Multi-Modal Visual Guidance and Feature Decomposition for Underwater Image Restoration
abstract
Underwater images are frequently affected by light absorption and scattering, which lead to color distortion, reduced contrast, and blurred details, significantly degrading overall image quality. Most underwater image restoration methods are confined to the pixel space of the raw modality, overlooking the important role of other modalities and different frequency-domain features. As a result, the representational capacity of deep learning models is not fully realized, affecting the generation of high-quality images. To address the above issues, we propose Multi-modal Visual Guidance and Feature Decomposition (MVG-FD) method for underwater image restoration. Specifically, we introduce Modality Visual Guidance (MVG) module, which integrates the complementary information provided by depth modality features into the raw features to guide the model in restoring the color of underwater images. Meanwhile, we design Feature Decomposition (FD) module, which utilizes Learnable Wavelet Decomposition (LWD) to decompose and extract the high-frequency bands of the raw features to help restore the texture details of the image. MVG-FD significantly improves PSNR and SSIM on existing datasets. The code is available at:https://github.com/zhangao668/MVG-FD.
Guangqi Jiang, Yi Liu 0038, Huibing Wang, Shoukun Xu
IEEE Signal Process. Lett.4
2025 Focus More on What? Guiding Multi-Task Training for End-to-End Person Search
abstract
End-to-end person search is a research domain that executes the task of locating and identifying a target individual from a large number of scene images via a multi-task framework. However, a major challenge for the learning of end-to-end methods is the inherent conflicts between the two sub-tasks: person detection focuses on identifying generic features of persons, while person re-identification (Re-ID) strives to find unique, distinguishing features for matching the target person. Unlike previous research focusing on model architectures, this paper delves into the end-to-end person search training process. We find that the unbalanced and conflicting training issues significantly impair the learning efficiency of the Re-ID sub-task, which directly influences person search accuracy. To address this, we propose a novel Guiding Multi-Task Training (GMT) framework that facilitates end-to-end balanced learning for person search. We introduce a Guiding Multi-Task Harmonious Learning (GMHL) module, which decouple the features and then performs intra- and cross-task feature interaction to enhance the learning of each sub-task. Moreover, GMT employs a Balancing Multi-Task Oriented Fusing (BMOF) method to explicitly enhance Re-ID sub-task learning through additional Re-ID training and target-guided multi-model parameters fusion. Extensive experiments on 2 benchmark datasets, CUHK-SYSU and PRW, show that GMT achieves leading performance with 96.0% mAP and 61.3% mAP, respectively.
Boyu Cai, Huibing Wang, Mingze Yao, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.2
2025 Fusion-Based Channel-Wise Isotropic Convergent Real-Time Underwater Image Enhancement
abstract
Existing underwater image enhancement (UIE) methods typically prioritize improving image quality at the expense of algorithmic efficiency. In this paper, we propose a fusion-based, channel-wise isotropic convergent UIE method designed for real-time performance. The proposed approach comprises three key modules: (i) a non-linear transformation module that corrects color casts and aligns the pixel distribution with the gray-world assumption (GWA); (ii) a channel-wise isotropic convergence scheme that reduces intensity distribution disparities across channels, promoting balanced convergence; and (iii) a patch-based enhancement strategy that divides the image into smaller patches to better capture local features and improve adaptability to non-uniform degradation. Moreover, certain critical steps in our method are optimized to achieve O(1) time complexity, allowing it to meet real-time requirements. Extensive experiments validate the effectiveness of each module in the proposed method, showcasing its superiority when compared to the existing state-of-the-art (SOTA) approaches. Code has been released at https://github.com/JohnChenS/FCICE_UnderwaterImageEnhancement.
Yuehan Chen, Jiqing Zhang, Haoming Tang, Huibing Wang, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.6
2025 Omni Contextual Aggregation Networks for High-Fidelity Image Inpainting
abstract
Image inpainting aims to restore a realistic image from a damaged or incomplete version. Although Transformer-based methods have achieved impressive results by modeling long-range dependencies, the inherent quadratic complexity of canonical self-attention has typically led to these approaches adopting uni-dimensional modeling, which limits the model’s ability to capture complex relationships from both spatial and channel dimensions. To this end, this paper exploits a novel attention paradigm termed Dynamic Omni-Attention Mechanism (DOAM) for simultaneously modeling pixel-interaction from both spatial and channel dimensions, and implements the information interaction across the omni-axis (i.e., spatial and channel) with linear computational complexity. In addition, to handle large-scale degradation, this paper proposes a Multi-band Feature Enhancement (MFE) module to enhance feature representation in downsampling, thus unlocking the potential of subsequent attentional interactions. Moreover, motivated by recent advances in image restoration, this paper incorporates a domain-related prior representation from CNN-based Network to modulate the features during proposed attention mechanism and feed-forward networks. Integrating the above designs into an encoder-decoder architecture, the proposed Omni Contextual Aggregation Networks (OCANet) achieve superior performance at lower parameters and time costs than the competitive baselines. Extensive experiments on CelebA-HQ, Paris Street View, FFHQ and Dunhuang datasets validate the efficacy of the proposed method.
Jinjia Peng, Mengkai Li, Huibing Wang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Tensor Completion Framework by Graph Refinement for Incomplete Multi-View Clustering
abstract
Incomplete Multi-view Clustering (IMVC) endeavors to harness information from multiple incomplete views to partition multi-view data into their respective clusters. How to recover missing information with lossless fidelity is the core of IMVC, which is of vital importance but challenging. Most of the existing methods include a feature recovery step to mitigate the negative impact of missing samples on the feature graph, however, these IMVC algorithms simply utilize the correlation between samples to recover the relationship between the unmissing instances and the missing instances while ignoring the consistency between views, which leads to often unsatisfactory recovery results. In addition, previous IMVC algorithms focus more on the recovery of incomplete data, ignoring the effect of the error term on incomplete graphs. This can mislead the recovery process of IMVC algorithm and the feature graph can be affected by anomalous information, which leads to degradation of clustering performance. To address this gap, this paper introduces the Tensor Completion Framework by Graph Refinement for Incomplete Multi-view Clustering (IMVC-TGR). IMVC-TGR separates the redundant information in each affine graph by graph refinement operation, aiming to mitigate the negative impact of error terms and redundant information on the feature graph during the recovery process. Meanwhile, IMVC-TGR stacks the feature graphs into tensors to explore intra-view correlation and inter-view consistency, so as to recover the relationship between missing samples and non-missing samples, and improve the quality of the feature graphs. Finally, IMVC-TGR introduces semantic consistency constraints and self-weighted fusion strategies into the high-quality feature graphs, aiming at preserving the complementary information between different views while balancing the contributions of the refined representation matrices of different views. The experimental results on multiple different datasets indicate that IMVC-TGR can achieve state-of-the-art performance.
Huibing Wang, Yawei Chen, Mingze Yao, Jinjia Peng, Xianping Fu
IEEE Trans. Multim.1
2025 Between/Within View Information Completing for Tensorial Incomplete Multi-View Clustering
abstract
Incomplete Multi-view Clustering (IMvC) receives increasing attention due to its effectiveness in solving data-missing problems. With the information loss in incomplete situations, the core of IMvC needs to consider effectively overcoming the challenge of missing views, that is, exploring the underlying correlations from available data and recovering the missing information. However, most existing IMvC methods overemphasize the recovery-first principle with integrating the existing data from different views while neglecting the influence of view consistency in IMvC task together with valuable within view information. In this paper, a novel Between/Within View Information Completing for Tensorial Incomplete Multi-view Clustering (BWIC-TIMC) has been proposed, in which between/within view information is jointly exploited for effectively completing the missing views. Specifically, the proposed method designs a dual tensor constraint module, which focuses on simultaneously exploring the view-specific correlations of incomplete views and enforcing the between view consistency across different views. With the dual tensor constraint, between/within view information can be effectively integrated for completing missing views for IMvC task. Furthermore, in order to balance different contributions of multiple views and alleviate the problem of feature degeneration, BWIC-TIMC implements an adaptive fusion graph learning strategy for consensus representation learning. Extensive comparative experiments with the-state-of-art baselines can demonstrate the effectiveness of BWIC-TIMC.
Mingze Yao, Huibing Wang, Yawei Chen, Xianping Fu
IEEE Trans. Multim.2
2025 Feature decomposition and structural learning for multi-diverse and multi-view data clustering
Huibing Wang
Vis. Comput.4
2024 Fast One-Stage Unsupervised Domain Adaptive Person Search
Tianxiang Cui, Huibing Wang, Jinjia Peng, Ruoxi Deng, Xianping Fu, Yang Wang 0023
IJCAI2
2024 Scene-Adaptive Person Search via Bilateral Modulations
Huibing Wang, Jinjia Peng, Xianping Fu, Yang Wang 0023
IJCAI2
2024 HSMnet: Hybrid Sampling and Matching Network for DETR-based Person Search
Zhengjie Lu, Jinjia Peng, Huibing Wang, Qingxuan Shi, Bin Wang 0044
MMAsia3
2024 Cascade transformers with dynamic attention for video question answering
abstract
Visual question answering (VQA) has become a hot study topic with challenging motivation of correctly answering the videos or images questions in recent years. However, the existing VQA model mostly aimed at answering questions about images and performed poorly in the video question answering (VideoQA) domain. VideoQA needs to simultaneously consider the correlations between video frames and the dynamic information of multiple objects in video. Therefore, we propose a novel Cascade Transformers with Dynamic Attention for Video Question Answering (CTDA-QA), which aims to simultaneously solve the above considerations. Specifically, the proposed CTDA-QA model utilizes multiple transformers structure to encode videos for reasoning complex spatial and temporal information, which is different from the previous recurrent neural network methods. Besides, in order to effectively capture the dynamic information from various scenarios in videos, a flexible attention module has been proposed to explore the essential relations between objects in a dynamic timeline. Finally, to avoid spurious answers and fully explore the cross-modal relationships, a mixed-supervised learning strategy is designed for optimizing the reasoning tasks. The experiments on several benchmark video question–answer datasets clearly verify the performance and effectiveness of CTDA-QA, which contains the results in contrast to the state-of-the-art methods. Besides, the provided ablation study and visualization results further reveal the potential of CTDA-QA.
Tingfei Yan, Mingze Yao, Huibing Wang
Comput. Vis. Image Underst.4
2024 A depth map stitching framework based on salient region matching
Zetian Mi, Haixia Qi, Huibing Wang, Xianping Fu
Eng. Appl. Artif. Intell.6
2024 Large-scale multi-view subspace clustering via embedding space and partition matrix
Tianhang Cheng, Jinjia Peng, Huibing Wang
Neurocomputing4
2024 AGS: Transferable adversarial attack for person re-identification by adaptive gradient similarity attack
Zeze Tao, Zhengjie Lu, Jinjia Peng, Huibing Wang
Knowl. Based Syst.4
2024 Generating high-quality texture via panoramic feature aggregation for large mask inpainting
Jinjia Peng, Huibing Wang
Knowl. Based Syst.4
2024 High-strength synergic-calibration attention system in YOLO for underwater object detection application
GuoLiang Yuan 0001, Huibing Wang, Xianping Fu
Multim. Syst.3
2024 Criss-cross global interaction-based selective attention in YOLO for underwater object detection
Huibing Wang, Tianzhu Gao, Xianping Fu
Multim. Tools Appl.2
2024 Uncertainty-guided Robust labels refinement for unsupervised person re-identification
Jinjia Peng, Zeze Tao, Huibing Wang
Neural Comput. Appl.4
2024 Adapt only once: Fast unsupervised person re-identification via relevance-aware guidance
Jinjia Peng, Jiazuo Yu 0002, Huibing Wang, Xianping Fu
Pattern Recognit.4
2024 Manifold-Based Incomplete Multi-View Clustering via Bi-Consistency Guidance
abstract
Incomplete multi-view clustering primarily focuses on dividing unlabeled data into corresponding categories with missing instances, and has received intensive attention due to its superiority in real applications. Considering the influence of incomplete data, the existing methods mostly attempt to recover data by adding extra terms. However, for the unsupervised methods, a simple recovery strategy will cause errors and outlying value accumulations, which will affect the performance of the methods. Broadly, the previous methods have not taken the effectiveness of recovered instances into consideration, or cannot flexibly balance the discrepancies between recovered data and original data. To address these problems, we propose a novel method termed Manifold-based Incomplete Multi-view clustering via Bi-consistency guidance (MIMB), which flexibly recovers incomplete data among various views, and attempts to achieve biconsistency guidance via reverse regularization. In particular, MIMB adds reconstruction terms to representation learning by recovering missing instances, which dynamically examines the latent consensus representation. Moreover, to preserve the consistency information among multiple views, MIMB implements a biconsistency guidance strategy with reverse regularization of the consensus representation and proposes a manifold embedding measure for exploring the hidden structure of the recovered data. Notably, MIMB aims to balance the importance of different views, and introduces an adaptive weight term for each view. Finally, an optimization algorithm with an alternating iteration optimization strategy is designed for final clustering. Extensive experimental results on 6 benchmark datasets are provided to confirm that MIMB can significantly obtain superior results as compared with several state-of-the-art baselines.
Huibing Wang, Mingze Yao, Yawei Chen, Yunqiu Xu, Haipeng Liu 0004, Wei Jia 0001, Xianping Fu, Yang Wang 0023
IEEE Trans. Multim.1
2024 A Unified Framework Based on Graph Consensus Term for Multiview Learning
abstract
In recent years, multiview learning technologies have attracted a surge of interest in the machine learning domain. However, when facing complex and diverse applications, most multiview learning methods mainly focus on specific fields rather than provide a scalable and robust proposal for different tasks. Moreover, most conventional methods used in these tasks are based on single view, which cannot be readily extended into the multiview scenario. Therefore, how to provide an efficient and scalable multiview framework is very necessary yet full of challenges. Inspired by the fact that most of the existing single view algorithms are graph-based ones to learn the complex structures within given data, this article aims at leveraging most existing graph embedding works into one formula via introducing the graph consensus term and proposes a unified and scalable multiview learning framework, termed graph consensus multiview framework (GCMF). GCMF attempts to make full advantage of graph-based works and rich information in the multiview data at the same time. On one hand, the proposed method explores the graph structure in each view independently to preserve the diversity property of graph embedding methods; on the other hand, learned graphs can be flexibly chosen to construct the graph consensus term, which can more stably explore the correlations among multiple views. To this end, GCMF can simultaneously take the diversity and complementary information among different views into consideration. To further facilitate related research, we provide an implementation of the multiview extension for locality linear embedding (LLE), named GCMF-LLE, which can be efficiently solved by applying the alternating optimization strategy. Empirical validations conducted on six benchmark datasets can show the effectiveness of our proposed method.
Xiangzhu Meng, Lin Feng 0001, Chonghui Guo, Huibing Wang
IEEE Trans. Neural Networks Learn. Syst.4
2024 Graph-Collaborated Auto-Encoder Hashing for Multiview Binary Clustering
abstract
Unsupervised hashing methods have attracted widespread attention with the explosive growth of large-scale data, which can greatly reduce storage and computation by learning compact binary codes. Existing unsupervised hashing methods attempt to exploit the valuable information from samples, which fails to take the local geometric structure of unlabeled samples into consideration. Moreover, hashing based on auto-encoders aims to minimize the reconstruction loss between the input data and binary codes, which ignores the potential consistency and complementarity of multiple sources data. To address the above issues, we propose a hashing algorithm based on auto-encoders for multiview binary clustering, which dynamically learns affinity graphs with low-rank constraints and adopts collaboratively learning between auto-encoders and affinity graphs to learn a unified binary code, called graph-collaborated auto-encoder (GCAE) hashing for multiview binary clustering. Specifically, we propose a multiview affinity graphs' learning model with low-rank constraint, which can mine the underlying geometric information from multiview data. Then, we design an encoder-decoder paradigm to collaborate the multiple affinity graphs, which can learn a unified binary code effectively. Notably, we impose the decorrelation and code balance constraints on binary codes to reduce the quantization errors. Finally, we use an alternating iterative optimization scheme to obtain the multiview clustering results. Extensive experimental results on five public datasets are provided to reveal the effectiveness of the algorithm and its superior performance over other state-of-the-art alternatives.
Huibing Wang, Mingze Yao, Guangqi Jiang, Zetian Mi, Xianping Fu
IEEE Trans. Neural Networks Learn. Syst.1
2024 ReFID: Reciprocal Frequency-aware Generalizable Person Re-identification via Decomposition and Filtering
abstract
Domain generalization of person re-identification aims to conduct testing across domains that have not been previously encountered, without utilizing target domain data during the training stage. As the number of source domains increases, the relationships between training samples become more complex. This can lead to domain-invariant features that include certain instance-level spurious correlations, which can impact the model’s ability to generalize further. To overcome this limitation, the Reciprocal Frequency-aware Generalizable Person Re-identification method is proposed in this article, which aims to utilize spectral feature correlation learning to transmit frequency component information and generate more discriminative hybrid features. A module called Bilateral Frequency Component-guided Attention is developed to help the network understand high-level semantic and texture information from various frequency features. Furthermore, to reduce the impact of noise from the frequency domain, this article proposes an innovative module called Fourier Noise Masquerade Filtering. This module enhances the portability of frequency domain components while simultaneously suppressing elements that do not contribute to generalization. Extensive experimental results on various datasets demonstrate that our method is effective and superior to the state-of-the-art methods.
Jinjia Peng, Song Pengpeng, Huibing Wang
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Multiple information perception-based attention in YOLO for underwater object detection
Huibing Wang, Tianxiang Cui, Zhicheng Guo, Xianping Fu
Vis. Comput.2
2024 Publisher Correction: Multiple information perception-based attention in YOLO for underwater object detection
Huibing Wang, Tianxiang Cui, Zhicheng Guo, Xianping Fu
Vis. Comput.2
2023 Target-Oriented Multi-criteria Band Selection for Hyperspectral Image
Huijuan Pang, Xudong Sun 0009, Xianping Fu, Huibing Wang
PRCV (7)4
2023 Learning interpretable shared space via rank constraint for multi-view clustering
Guangqi Jiang, Huibing Wang, Jinjia Peng, Dongyan Chen, Xianping Fu
Appl. Intell.2
2023 Hybrid partial-constrained learning with orthogonality regularization for unsupervised person re-identification
Jiazuo Yu 0002, Jinjia Peng, Kai Li 0038, Huibing Wang
Eng. Appl. Artif. Intell.4
2023 OPM2L: An optimal instance partition-based multi-metric learning method for heterogeneous dataset classification
abstract
Multi-metric learning -a method to learn multiple local metrics to reveal the feature's correlations of samples from different local regions-has become an essential tool to measure the similarities between instances from heterogeneous datasets. However, most existing cluster-based MML methods first partition the training data with a predefined metric and then learn multiple metrics via the local instances, leading to these two independent procedures fail to cooperate with each other. In this paper, we propose an Optimal instance Partition-based Multi-Metric Learning (OPM 2 L) method for heterogeneous dataset classification by unifying the instance partition and multiple local metrics learning into a single objective. In particular, multiple anchor centers together with a global metric are employed to assist the instance partition process. During the training, the shared information contained in local metrics is aggregated into the global metric by a dedicated regularizer, which improves the instance partition process and offers the subsequent multiple local metrics learning with more informative instances. Moreover, an efficient alternating direction technology is employed to seek a feasible solution to the proposed method. We further confirmed that the sub-problems can be settled with closed-form solutions, while the superiority of the proposed method is also proved by experimental results on extensive datasets.
Huiyuan Deng, Xiangzhu Meng, Huibing Wang, Lin Feng 0001
Inf. Sci.3
2023 Domain adaptive person search via GAN-based scene synthesis for cross-scene videos
Huibing Wang, Tianxiang Cui, Mingze Yao, Huijuan Pang, Yushan Du
Image Vis. Comput.1
2023 Multi-dimensional, multi-functional and multi-level attention in YOLO for underwater object detection
Xudong Sun 0009, Huibing Wang, Xianping Fu
Neural Comput. Appl.3
2023 Joint learning with diverse knowledge for re-identification
Jinjia Peng, Jiazuo Yu 0002, Guangqi Jiang, Huibing Wang
Signal Process. Image Commun.4
2023 Adaptive Memorization With Group Labels for Unsupervised Person Re-Identification
abstract
Re-identification (re-ID) aims to identify a person’s images across different cameras. However, the domain differences between different datasets make it a challenge for re-ID models trained on one dataset to be adapted to another. A variety of unsupervised domain adaptation methods tend to transfer learned knowledge from one domain to another by optimizing with pseudo-labels. Though impressive performances have been achieved, there are still some limitations. To be specific, these methods always generate one pseudo label for each unlabeled sample, which is hard to describe a person accurately and introduces a large number of noisy labels by one-shot clustering, thus hindering the retraining process and limiting generalization. To build more comprehensive descriptions of samples and mitigate the effects of noisy pseudo labels, this paper proposes an Adaptive Memorization with Group labels (AdaMG) framework for unsupervised person re-ID, which resists noisy labels and exploits the diversity of samples by developing a multi-branch structure with the adaptive memorization. In particular, group labels are generated for one sample in the unseen domain to learn more complementary and diverse features through clustering. Meanwhile, to better optimize the neural networks with noisy data, multiple memory structures are designed in AdaMG, which are updated adaptively according to the confidence of samples. Comprehensive experimental results have demonstrated that our proposed method can achieve excellent performances on benchmark datasets.
Jinjia Peng, Guangqi Jiang, Huibing Wang
IEEE Trans. Circuits Syst. Video Technol.3
2023 Towards Adaptive Consensus Graph: Multi-View Clustering via Graph Collaboration
abstract
Multi-view clustering is a long-standing important task, however, it remains challenging to exploit valuable information from the complex multi-view data located in diverse high-dimensional spaces. The core issue is the effective collaboration of multiple views to holistically uncover the essential correlations between multi-view data through graph learning. Furthermore, it is indispensable for most existing methods to introduce an additional clustering step to produce the final clusters, which evidently reduces the uniform relationship between graph learning and clustering. Based on the above considerations, in this paper, we present a novel method named multi-view clustering via graph collaboration (MCGC). Based on the low-dimensional representation space developed by MCGC, it first perceives the correlations between samples in each individual view under the supervision of the Hilbert-Schmidt independence criterion (HSIC). Then, MCGC proposes learning a consensus graph by adaptively collaborating between all the views, which is able to uncover the essential structure of the multi-view data. Meanwhile, by imposing the rank constraint on the Laplacian matrix of the consensus graph to partition the multi-view data naturally into the required number of clusters, the optimal clustering results can be obtained directly without any postprocessing steps. Finally, the resulting optimization problem is solved by an alternating optimization scheme with guaranteed fast convergence. Extensive experiments on 5 benchmark multi-view datasets demonstrate that MCGC markedly outperforms the state-of-the-art baselines.
Huibing Wang, Guangqi Jiang, Jinjia Peng, Ruoxi Deng, Xianping Fu
IEEE Trans. Multim.1
2022 Parallelism Network with Partial-aware and Cross-correlated Transformer for Vehicle Re-identification
abstract
Vehicle re-identification (ReID) aims to identify a specific vehicle in the dataset captured by non-overlapping cameras, which plays a great significant role in the development of intelligent transportation systems. Even though CNN-based model achieves impressive performance for the ReID task, its Gaussian distribution of effective receptive fields has limitations in capturing the long-term dependence between features. Moreover, it is crucial to capture fine-grained features and the relationship between features as much as possible from vehicle images.
Guangqi Jiang, Huibing Wang, Jinjia Peng, Xianping Fu
ICMR2
2022 Progressive learning with multi-scale attention network for cross-domain vehicle re-identification
Yang Wang 0023, Jinjia Peng, Huibing Wang, Meng Wang 0001
Sci. China Inf. Sci.3
2022 Learning latent features with local channel drop network for vehicle re-identification
Xianping Fu, Jinjia Peng, Guangqi Jiang, Huibing Wang
Eng. Appl. Artif. Intell.4
2022 Adaptive multi-view multiple-means clustering via subspace reconstruction
Luyao Liu 0001, Yong Zhang 0030, Huibing Wang, Lin Feng 0001
Eng. Appl. Artif. Intell.4
2022 Hierarchical multi-view metric learning with HSIC regularization
Huiyuan Deng, Xiangzhu Meng, Huibing Wang, Lin Feng 0001
Neurocomputing3
2022 Cooperative Refinement Learning for domain adaptive person Re-identification
Jinjia Peng, Guangqi Jiang, Huibing Wang
Knowl. Based Syst.3
2022 Eliminating cross-camera bias for vehicle re-identification
Jinjia Peng, Guangqi Jiang, Dongyan Chen, Huibing Wang, Xianping Fu
Multim. Tools Appl.5
2022 Refined marine object detector with attention-based spatial pyramid pooling networks and bidirectional feature fusion strategy
Fengqiang Xu, Huibing Wang, Xudong Sun 0009, Xianping Fu
Neural Comput. Appl.2
2022 Tensorial Multi-View Clustering via Low-Rank Constrained High-Order Graph Learning
abstract
Multi-view clustering aims to partition multi-view data into different categories by optimally exploring the consistency and complementary information from multiple sources. However, most existing multi-view clustering algorithms heavily rely on the similarity graphs from respective views and fail to comprehend multiple views holistically. Moreover, due to the noise and redundancy maintained in the original data, the original errors of multiple similarity graphs will continue to accumulate in the process of constructing consistent graphs. These situations always lead to the limitation to effective fuse the essential information from multiple views, which always influences the clustering performance and cries out for reliable solutions. Based on the above considerations, we propose a novel method termed Tensorial Multi-view Clustering (TMvC), which learns high-order graph by low-rank tensor constraint to uncover the essential information stored in multiple views. TMvC first learns the Laplacian graphs of all views and stacks them into a tensor which can be viewed as a high-order graph. With the high-order graph, consistency and complementary information from different views can be propagated smoothly across all views. Then, based on low-rank constraint, high-order graph is constrained in the horizontal and vertical directions to better uncover the inter-view and inter-class correlations between multi-view data, which is of vital importance for multi-view clustering. Extensive experiments on document and image datasets demonstrate that TMvC can achieve the state-of-the-art performance for multi-view clustering.
Guangqi Jiang, Jinjia Peng, Huibing Wang, Zetian Mi, Xianping Fu
IEEE Trans. Circuits Syst. Video Technol.3
2021 Learning Multiple Semantic Knowledge For Cross-Domain Unsupervised Vehicle Re-Identification
abstract
Unsupervised Vehicle re-identification (reID) aims at searching the similar vehicles’ images from large unlabelled datasets captured in a multiple camera network, which is still a challenging task. In this paper, a multiple semantic knowledge learning approach is proposed to exploit the potential similarity of unlabeled samples, which builds multiple clusters from different views automatically with different cues. Specially, different from some existing works focus on the knowledge of one view, for each vehicle in the target domain, different semantic knowledge could be learned with the proposed focal drop network and several different labels can be assigned according these knowledge, which would be employed to train the vehicle reID model jointly. In addition, due to the unreliability of pseudo labels assigned by the clustering, the hard triplet center loss is proposed to take the difference of intra-cluster and inter-cluster into consideration for better training the unsupervised framework to adapt the unknown domain. Comprehensive experimental results clearly demonstrate that our method achieves excellent performance on both VehicleID dataset and VeRi-776 dataset.
Huibing Wang, Jinjia Peng, Guangqi Jiang, Xianping Fu
ICME1
2021 MSAV: An Unified Framework for Multi-view Subspace Analysis with View Consistence
abstract
With the development of multimedia period, information is always caputred with multiple views, which causes a research upsurge on multi-view learning. It is obvious that multi-view data contains more information than those single view ones. Therefore, it is crucial to develop the multi-view algorithms to adapt the demand of many applications. Even though some excellent multi-view algorithms were proposed, most of them can only deal with the specific problems. To tacle this problem, this paper proposes an unified framework named Multi-view Subspace Analysis with View Consistence (MSAV), which provides an unified means to extend those single-view dimension reduciton algorithms into multi-view versions. MSAV first extends multi-view data into kernel space to avoid the problem caused by different dimensions of the data from multiple views. Then, we introduced a self-weighted learning strategy to automatically assign weights for all views according to their importance. Finally, in order to promote the consistence of all views, Hilbert-Schmidt Independence Criterion is adopted by MSAV. Furthermore, We conducted experiments on several benchmark datasets to verify the performance of MSAV.
Huibing Wang, Guangqi Jiang, Jinjia Peng, Xianping Fu
ICMR1
2021 Learning to Decode Contextual Information for Efficient Contour Detection
abstract
Contour detection plays an important role in both academic research and real-world applications. As the basic building block of many applications, its accuracy and efficiency highly influence the subsequent stages. In this work, we propose a novel lightweight system for contour detection that achieves state-of-the-art performance while keeps ultra-slim model size. The proposed method is built on an efficient encoder in a bottom-up/top-down fashion. Specially, we propose a novel decoder that compresses side features from an encoder and effectively decodes compact contextual information for high-accurate boundary localization. Besides, we propose a novel loss function that is able to assist a model to produce crisp object boundaries.
Ruoxi Deng, Shengjun Liu 0002, Huibing Wang, Hanli Zhao, Xiaoqin Zhang 0002
ACM Multimedia4
2021 Multi-view Low-rank Preserving Embedding: A novel method for multi-view representation
Xiangzhu Meng, Lin Feng 0001, Huibing Wang
Eng. Appl. Artif. Intell.3
2021 Graph-based Multi-view Binary Learning for image clustering
Guangqi Jiang, Huibing Wang, Jinjia Peng, Dongyan Chen, Xianping Fu
Neurocomputing2
2021 Discriminative feature and dictionary learning with part-aware model for vehicle re-identification
Huibing Wang, Jinjia Peng, Guangqi Jiang, Fengqiang Xu, Xianping Fu
Neurocomputing1
2021 Generalized multiple sparse information fusion for vehicle re-identification
Jinjia Peng, Guangqi Jiang, Huibing Wang
J. Vis. Commun. Image Represent.3
2021 Bottom-up broadcast neural network for music genre classification
Caifeng Liu, Lin Feng 0001, Guochao Liu, Huibing Wang, Shenglan Liu 0001
Multim. Tools Appl.4
2021 Scale-aware feature pyramid architecture for marine object detection
Fengqiang Xu, Huibing Wang, Jinjia Peng, Xianping Fu
Neural Comput. Appl.2
2021 Kernelized Multiview Subspace Analysis By Self-Weighted Learning
abstract
With the popularity of multimedia technology, information is always represented from multiple views. Even though multiview data can reflect the same sample from different perspectives, multiple views are consistent to some extent because they are representations of the same sample. Most of the existing algorithms are graph-based ones to learn the complex structures within multiview data but overlook the information within data representations. Furthermore, many existing works treat multiple views discriminatively by introducing some hyperparameters, which is undesirable in practice. To this end, abundant multiview-based methods have been proposed for dimension reduction. However, there is still no research that leverages the existing work into a unified framework. In this paper, we propose a general framework for multiview data dimension reduction, named kernelized multiview subspace analysis (KMSA) to handle multiview feature representation in the kernel space, providing a feasible channel for multiview data with different dimensions. Compared with the graph-based methods, KMSA can fully exploit information from multiview data with nothing to lose. Since different views have different influences on KMSA, we propose a self-weighted strategy to treat different views discriminatively. A co-regularized term is proposed to promote the mutual learning from multiviews. KMSA combines self-weighted learning with the co-regularized term to learn the appropriate weights for all views. We evaluate our proposed framework on 6 multiview datasets for classification and image retrieval. The experimental results validate the advantages of our proposed method.
Huibing Wang, Yang Wang 0023, Zhao Zhang 0001, Xianping Fu, Li Zhuo 0001, Mingliang Xu 0001, Meng Wang 0001
IEEE Trans. Multim.1
2020 Unsupervised Vehicle Re-identification with Progressive Adaptation
abstract
Vehicle re-identification (reID) aims at identifying vehicles across different non-overlapping cameras views. The existing methods heavily relied on well-labeled datasets for ideal performance, which inevitably causes fateful drop due to the severe domain bias between the training domain and the real-world scenes; worse still, these approaches required full annotations, which is labor-consuming. To tackle these challenges, we propose a novel Progressive Adaptation Learning method for vehicle reID, named PAL, which infers from the abundant data without annotations. For PAL, a data adaptation module is employed for source domain, which generates the images with similar data distribution to unlabeled target domain as “pseudo target samples”. These pseudo samples are combined with the unlabeled samples that are selected by a dynamic sampling strategy to make training faster. We further proposed a weighted label smoothing (WLS) loss, which considers the similarity between samples with different clusters to balance the confidence of pseudo labels. Comprehensive experimental results validate the advantages of PAL on both VehicleID and VeRi-776 dataset.
Jinjia Peng, Yang Wang 0023, Huibing Wang, Zhao Zhang 0001, Xianping Fu, Meng Wang 0001
IJCAI3
2020 Purifying real images with an attention-guided style transfer network for gaze estimation
Xianping Fu, Yuxiao Yan, Jinjia Peng, Huibing Wang
Eng. Appl. Artif. Intell.5
2020 An incrementally cascaded broad learning framework to facial landmark tracking
Caifeng Liu, Lin Feng 0001, Shuai Guo 0002, Huibing Wang, Shenglan Liu 0001, Hong Qiao
Neurocomputing4
2020 Cross domain knowledge learning with dual-branch adversarial network for vehicle re-identification
Jinjia Peng, Huibing Wang, Fengqiang Xu, Xianping Fu
Neurocomputing2
2020 Multi-view Locality Low-rank Embedding for Dimension Reduction
Lin Feng 0001, Xiangzhu Meng, Huibing Wang
Knowl. Based Syst.3
2020 The similarity-consensus regularized multi-view learning for dimension reduction
Xiangzhu Meng, Huibing Wang, Lin Feng 0001
Knowl. Based Syst.2
2020 Multi-view reconstructive preserving embedding for dimension reduction
Huibing Wang, Lin Feng 0001, Adong Kong, Bo Jin 0001
Soft Comput.1
2020 Structural Analysis of Attributes for Vehicle Re-Identification and Retrieval
abstract
Vehicle re-identification plays an important role in video surveillance applications. Despite the efforts made on this problem in the past few years, it remains a challenging task due to various factors such as pose variation, illumination changes, and subtle inter-class difference. We believe that the key information for identification has not been well explored in the literature. In this paper, we first collect a vehicle dataset `VAC21' which contains 7129 images of five types of vehicles. Then, we carefully label the 21 classes of structural attributes hierarchically with bounding boxes. To our knowledge, this is the first dataset with several detailed attributes labeled. Based on this dataset, we use the state-of-the-art one-stage detection method, Single-shot Detection, as a baseline model for detecting attributes. Subsequently, we make a few important modifications tailored for this application to improve accuracy: 1) adding more proposals from low-level layers to improve the accuracy of detecting small objects and 2) employing the focal loss to improve the mean average precision. Furthermore, the results of the attribute detection can be applied to a series of vision tasks that focus on analyzing the images of vehicles. Finally, we propose a novel region of interests (ROIs)-based vehicle re-identification and retrieval method in which the ROIs' deep features are used as discriminative identifiers, encoding the structure information of a vehicle. These deep features are input to a boosting model to improve the accuracy. A set of experiments are conducted on the dataset VehicleID and the experimental results show that our method outperforms the state-of-the-art methods.
Yanzhu Zhao, Chunhua Shen, Huibing Wang, Shengyong Chen
IEEE Trans. Intell. Transp. Syst.3
2019 Purifying naturalistic images through a real-time style transfer semantics network
Yuxiao Yan, Ibrahim Shehi Shehu, Xianping Fu, Huibing Wang
Eng. Appl. Artif. Intell.5
2019 Learning multi-region features for vehicle re-identification with context-based ranking method
Jinjia Peng, Huibing Wang, Xianping Fu
Neurocomputing2
2019 Co-regularized multi-view sparse reconstruction embedding for dimension reduction
Huibing Wang, Jinjia Peng, Xianping Fu
Neurocomputing1
2019 Auto-weighted Mutli-view Sparse Reconstructive Embedding
Huibing Wang, Haohao Li, Xianping Fu
Multim. Tools Appl.1
2019 Multi-feature distance metric learning for non-rigid 3D shape retrieval
Huibing Wang, Haohao Li, Jinjia Peng, Xianping Fu
Multim. Tools Appl.1
2019 Robust tracking via weighted online extreme learning machine
Huibing Wang, Yonggong Ren
Multim. Tools Appl.2
2019 Learning a Distance Metric by Balancing KL-Divergence for Imbalanced Datasets
abstract
In many real-world domains, datasets with imbalanced class distributions occur frequently, which may confuse various machine learning tasks. Among all these tasks, learning classifiers from imbalanced datasets is an important topic. To perform this task well, it is crucial to train a distance metric which can accurately measure similarities between samples from imbalanced datasets. Unfortunately, existing distance metric methods, such as large margin nearest neighbor, information-theoretic metric learning, etc., care more about distances between samples and fail to take imbalanced class distributions into consideration. Traditional distance metrics have natural tendencies to favor the majority classes, which can more easily satisfy their objective function. Those important minority classes are always neglected during the construction process of distance metrics, which severely affects the decision system of most classifiers. Therefore, how to learn an appropriate distance metric which can deal with imbalanced datasets is of vital importance, but challenging. In order to solve this problem, this paper proposes a novel distance metric learning method named distance metric by balancing KL-divergence (DMBK). DMBK defines normalized divergences using KL-divergence to describe distinctions between different classes. Then it combines geometric mean with normalized divergences and separates samples from different classes simultaneously. This procedure separates all classes in a balanced way and avoids inaccurate similarities incurred by imbalanced class distributions. Various experiments on imbalanced datasets have verified the excellent performance of our novel method.
Lin Feng 0001, Huibing Wang, Bo Jin 0001, Haohao Li, Mingliang Xue
IEEE Trans. Syst. Man Cybern. Syst.2
2018 Learning to Predict Crisp Boundaries
Ruoxi Deng, Chunhua Shen, Shengjun Liu 0002, Huibing Wang
ECCV (6)4
2018 Clustering algorithms based on correlation coefficients for probabilistic linguistic term sets
abstract
As a novel and powerful tool, the notion of probabilistic linguistic term sets (PLTSs) can efficiently model this kind of qualitative assessment information utilizing several possible linguistic terms associated with probabilities or weights over alternatives. Considering that there are no investigation and research on the correlation coefficient and clustering analysis for the concept of PLTSs. Therefore, some correlation coefficient formulas are put forward to measure the relationship between two PLTSs and then they are utilized to develop two novel clustering algorithms to group PLTSs in this paper. We first define some correlation coefficient formulas and their weighted forms to measure the relationship between PLTSs. Then, we extend a fuzzy clustering algorithm for PLTSs and also propose a novel orthogonal clustering algorithm for PLTSs. Finally, we provide a practical example, which performs cluster analysis on the levels of general higher education in different regions of China, to test and verify the usability of our proposed clustering algorithm.
Mingwei Lin, Huibing Wang, Zeshui Xu, Jinli Huang
Int. J. Intell. Syst.2
2017 Multi-view metric learning based on KL-divergence for similarity measurement
Huibing Wang, Lin Feng 0001, Xiangzhu Meng, Zhaofeng Chen, Laihang Yu
Neurocomputing1
2017 Deep CNNs With Spatially Weighted Pooling for Fine-Grained Car Recognition
abstract
Fine-grained car recognition aims to recognize the category information of a car, such as car make, car model, or even the year of manufacture. A number of recent studies have shown that a deep convolutional neural network (DCNN) trained on a large-scale data set can achieve impressive results at a range of generic object classification tasks. In this paper, we propose a spatially weighted pooling (SWP) strategy, which considerably improves the robustness and effectiveness of the feature representation of most dominant DCNNs. More specifically, the SWP is a novel pooling layer, which contains a predefined number of spatially weighted masks or pooling channels. The SWP pools the extracted features of DCNNs with the guidance of its learnt masks, which measures the importance of the spatial units in terms of discriminative power. As the existing methods that apply uniform grid pooling on the convolutional feature maps of DCNNs, the proposed method can extract the convolutional features and generate the pooling channels from a single DCNN. Thus minimal modification is needed in terms of implementation. Moreover, the parameters of the SWP layer can be learned in the end-to-end training process of the DCNN. By applying our method to several fine-grained car recognition data sets, we demonstrate that the proposed method can achieve better performance than recent approaches in the literature. We advance the state-of-the-art results by improving the accuracy from 92.6% to 93.1% on the Stanford Cars-196 data set and 91.2% to 97.6% on the recent CompCars data set. We have also tested the proposed method on two additional large-scale data sets with impressive results observed.
Qichang Hu, Huibing Wang, Chunhua Shen
IEEE Trans. Intell. Transp. Syst.2
2016 Multi-view Sparsity Preserving Projection for dimension reduction
Huibing Wang, Lin Feng 0001, Laihang Yu, Jing Zhang 0028
Neurocomputing1
2016 Extend semi-supervised ELM and a frame work
Shenglan Liu 0001, Lin Feng 0001, Huibing Wang, Xiao Yao 0001
Neural Comput. Appl.3
2016 Metric learning with geometric mean for similarities measurement
Huibing Wang, Lin Feng 0001, Yang Liu 0066
Soft Comput.1
2016 Semantic Discriminative Metric Learning for Image Similarity Measurement
abstract
Along with the arrival of multimedia time, multimedia data has replaced textual data to transfer information in various fields. As an important form of multimedia data, images have been widely utilized by many applications, such as face recognition and image classification. Therefore, how to accurately annotate each image from a large set of images is of vital importance but challenging. To perform these tasks well, it is crucial to extract suitable features to character the visual contents of images and learn an appropriate distance metric to measure similarities between all images. Unfortunately, existing feature operators, such as histogram of gradient, local binary pattern, and color histogram, care more about the visual character of images and lack the ability to distinguish semantic information. Similarities between those features cannot reflect the real category correlations due to the well-known semantic gap. In order to solve this problem, this paper proposes a regularized distance metric framework called semantic discriminative metric learning (SDML). SDML combines geometric mean with normalized divergences and separates images from different classes simultaneously. The learned distance metric can treat all images from different classes equally. And distinctions between similar classes with entirely different semantic contents are emphasized by SDML. This procedure ensures the consistency between dissimilarities and semantic distinctions and avoids inaccuracy similarities incurred by unbalanced locations of samples. Various experiments on benchmark image datasets show the excellent performance of the novel method.
Huibing Wang, Lin Feng 0001, Jing Zhang 0028, Yang Liu 0066
IEEE Trans. Multim.1
2015 Locality Structured Sparsity Preserving Embedding
abstract
In recent years, the theory of sparse representation (SR) has been widely exploited in sparse subspace learning (SSL). Among all these methods, SR is a parameter-free global algorithm in nature which is mostly utilized to construct the correlations between samples to avoid some negative effects incurred by k-nearest neighbor (KNN) or some other methods. However, these SSL algorithms always lack obvious discrimination because of the ignorance of samples distribution. Meanwhile, some incorrect correlations are taken into consideration owing to the global feature of SR. To solve these two problems, a new SSL algorithm called locality structured sparsity preserving embedding (LSPE) is proposed in this paper. We add the local structured information to SR and construct correlations between samples. However, LSPE is an unsupervised method which wastes all label information. Therefore, LSPE is extended to semi-supervised LSPE (SLSPE) in this paper. SLSPE not only makes good use of the label information but also enhances the discriminative power of LSPE. Extensive experiments have been performed on three image datasets (CMU, COIL20, ORL) and two UCI datasets (Glass, Segment) to prove the efficiency of the LSPE and SLSPE.
Lin Feng 0001, Huibing Wang, Shenglan Liu 0001
Int. J. Pattern Recognit. Artif. Intell.2
2014 Robust activation function and its application: Semi-supervised kernel extreme learning method
Shenglan Liu 0001, Lin Feng 0001, Xiao Yao 0001, Huibing Wang
Neurocomputing4