EDBT 2026 Demo / reviewers in the wild / expert
Guangming Lu 0002
dblp:78/4785-2
· DBLP profile ↗
205ranked-venue papers
3as first author
157since 2021 · last 2026
0000-0003-1578-2634ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 108 · 2 first-author · 93 since 2021Artificial intelligence and machine learning · 92 · 1 first-author · 69 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 10 since 2021Human-computer interaction and ubiquitous computing · 11 · 5 since 2021Databases, data management, data science and information retrieval · 9 · 7 since 2021Computer networks · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | S2PD: Physically grounded augmentations and stable parameter updates for cross-domain few-shot learning
Shuai Kang, Yunyu Zou, Ziteng Hong, Bingzhi Chen, Guangming Lu 0002 |
Neurocomputing | 6 |
| 2026 | Domain Information Removal With Decision Region Enlargement for Unseen Conditions Fault DiagnosisabstractMost existing domain generalization methods for fault diagnosis focus on extracting domain-invariant features from multiple source domains. However, the lack of target data severely restricts these domain-invariant features to only the source domain distribution, leading to inadequate generalization performance under unseen working conditions. To address this critical limitation, we propose a novel method named domain information removal with decision region enlargement. Specifically, for domain information elimination, we design dual encoders and dual classifiers to separately extract and classify fault-related features and domain-related features. A distribution discriminator is then introduced to minimize the mutual information between fault-features and domain-features, thereby yielding purer fault-discriminative representations. To further enhance the discriminability of fault features, learnable comparative anchors are employed to strengthen intra-class compactness and inter-class separability. This process clarifies the decision region for each fault category, enabling more robust adaptation to accurate identification of diverse faults in the target domain. Extensive experiments demonstrate the superiority of our proposed method. Yu Gao 0024, Zhanpei Zhang, Shilong Sun 0001, Jinxing Li 0003, Guangming Lu 0002 |
IEEE Internet Things J. | 5 |
| 2026 | Multi -directional decision fusion for black-box source-free anomaly detection
Yu Gao 0024, Shilong Sun 0001, Zhanpei Zhang, Jinxing Li 0003, Guangming Lu 0002 |
Pattern Recognit. | 5 |
| 2026 | Hierarchical Multi-Criteria Representation Fusion for Robust Incomplete Multimodal Sentiment Analysis
Yijing Dai, Yingjian Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
IEEE Trans. Affect. Comput. | 4 |
| 2026 | Rethinking the Knowledge Gap Between Cloud and Device Models for Effective Co-AdaptationabstractBy collaboratively updating the cloud (large-scale) and device (small-scale) models, co-adaptation aims to enhance the generalization performance of device models in response to the distribution shifts in the incoming data. Existing methods often rely on low-entropy samples that are selected by thedevice modelfor co-adaptation, which ignores the differences between the predictions of the cloud and device models that are caused by the knowledge gap. As a result, some of the selected samples are redundant and contribute limited value to cloud model updating and knowledge distillation. To this end, we propose a test-time co-adaptation method by Rethinking the Knowledge Gap (RKG) between cloud and device models, which effectively updates the models by informative sample selection and targeted knowledge distillation for image-based classification tasks. Specifically, we design a sample selection module that integrates semantic prediction entropy with object structure cues to identify valuable samples, which effectively alleviates the redundancy problem. Based on these selected samples, we further construct a reweighting module that measures the prediction consistency between the two models and assigns greater emphasis to samples with larger prediction discrepancies, i.e., larger knowledge gaps, to improve knowledge distillation. Furthermore, by jointly leveraging these two modules, RKG enables efficient and effective co-adaptation, thereby achieving robust model generalization to continuously changing data in classification scenarios. Extensive experiments demonstrate that RKG outperforms state-of-the-art methods while requiring fewer uploaded samples. Yingjian Li 0001, Yushi Zeng, Dongmei Jiang, Yaowei Wang 0001, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | DS2VP: Dynamically-Selected Spatially Visual PromptingabstractThe significant effectiveness of prompt tuning for computer vision tasks has been extensively demonstrated in numerous studies. As a widely feasible solution, the spatial modeling paradigm aims to overcome the limitations of sequence modeling paradigm in capturing spatial relationships within images by learning a prompt token map and aligning it spatially with the image token map. However, such spatial modeling paradigms of visual prompt tuning still face two potential challenges: 1) Most existing methods fail to design individual prompts for different images, and the learned prompts have the same static effect on all images. 2) The strategy of existing methods overlooks the selection of key spatial information and indiscriminately prompts all information within the image. In this work, we propose a novel Dynamically-Selected and Spatial Visual Prompting, termed as DS2VP, which aims to effectively utilize the key spatial information of the input image and enable dynamic visual prompt selection. Specifically, our DS2VP approach is meticulously designed to leverage the key index generator to filter key regions of the image for determining the spatial target of prompts, thus enabling dynamic selection of prompts for different images. By adding prompt tokens at selected key locations, an image prompt fusion module is deployed by adapting the learnable prompt tokens into the input image tokens, further achieving a fine-grained spatial alignment. Moreover, we propose a multi-level prompt interaction module that facilitates interactions between visual prompts at different levels to enhance feature representations across various semantic levels. Extensive experiments on two challenging benchmarks for image classification have demonstrated the superiority of DS2VP over other state-of-the-art methods for visual prompt tuning. Yishu Liu 0001, Bingzhi Chen, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Soft Supervision-Guided Spatial-Temporal Refinement Network for Video-Based Visible-Infrared Person Re-IdentificationabstractThanks to automatic switch between visible and infrared modes, person re-identification (Re-ID) in 24-hour has been possible through cross-modal retrieval. Instead of exploiting still images, video-based cross-modal person Re-ID is studied in this paper. Specifically, a large-scale dataset 'HITSZ-PVCM' is first collected, consisting of as many as 1,681 identities and 839,632 frames. Generally, videos contain much richer pedestrian appearances. However, most existing works only generate temporal representations by whole frames, inevitably losing fine-grained details. Furthermore, training a network by metric losses (e.g., center loss) is a common strategy, while such point-to-point constraints are too strong and limit model generalization due to existing diversity among intra-class samples. Here, we propose a Soft Supervision guided Spatial-Temporal Refinement (S3TR) network to tackle these problems. Specifically, S3TR refines each frame guided by a coarse temporal feature, so that more discriminative features are extracted and transformed to a sequential representation. Followed by a global-local mutual learning module, the modality gap is then erased without losing fine-grained details. Furthermore, we propose a novel soft-clustering center loss to measure intra-/inter-class similarity/dissimilarity in a group-to-group way, efficiently improving model generalization. To the best of our knowledge, HITSZ-PVCM is the largest dataset and S3TR achieves superior performances compared with state-of-the-arts. Jinxing Li 0003, Chuhao Zhou, Huafeng Li 0001, Guangming Lu 0002, Yong Xu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2025 | Towards Robust Visual Question Answering via Prompt-Driven Geometric HarmonizationabstractVisual Question Answering (VQA) has garnered significant attention as a crucial link between vision and language, aimed at generating accurate responses to visual queries. However, current VQA models still struggle with the challenges of minority class collapse and spurious semantic correlations posed by language bias and imbalanced distributions. To address these challenges, this paper proposes a novel Prompt-Driven Geometric Harmonization (PDGH) paradigm, which integrates both geometric structure and information entropy principles to enhance the ability of VQA models to generalize effectively across diverse scenarios. Specifically, our PDGH approach is meticulously designed to generate image-generated prompts that are guided by specific question cues, facilitating a more accurate and context-aware understanding of the visual content. Moreover, we project the prompt-visual-question and visual-question joint representations into a unified hypersphere space, applying feature weight self-orthogonality and prompt-information entropy correction constraints to optimize the margin, further alleviating minority class collapse and correcting language bias. To maintain the geometric integrity of the representation space, we introduce multi-space geometric contrast constraints to minimize the impact of spurious priors introduced during training. Finally, a semantic matrix is constructed for the coordinated joint representation to ensure that the learned instances are semantically consistent and improve reasoning ability. Extensive experiments on various general and medical VQA datasets demonstrate the consistent superiority of our PDGH approach over existing state-of-the-art baselines. Yishu Liu 0001, Congcong Wen, Guangming Lu 0002, Bingzhi Chen |
AAAI | 4 |
| 2025 | ALRMR-GEC: Adjusting Learning Rate Based on Memory Rate to Optimize the Edit Scorer for Grammatical Error CorrectionabstractEdit-based approaches for Grammatical Error Correction (GEC) have attracted volume attention due to their outstanding explanations of the correction process and rapid inference. Through exploring the characteristics of the generalized and specific knowledge learning for GEC, we discover that efficiently training GEC systems with satisfactory generalization capacity prefers more generalized knowledge rather than specific knowledge. Current gradient-based methods for training GEC systems, however, usually prioritize minimizing training loss over generalization loss. This paper proposes the strategy of Adjusting Learning Rate Based on Mermory Rate to optimize the edit-based GEC scorer (ALRMR-GEC). Specifically, we introduce the memory rate, a novel metric, to provide an explicit indicator for the model’s state of learning generalized and specific knowledge, which can effectively guide the GEC system to adjust the learning rate timely. Extensive experiments, conducted by optimizing the published edit scorer on the BEA2019 dataset, have shown our ALRMR-GEC significantly enhances the model generalization ability with stable and satisfactory performance nearly irrespective of the initial learning rate selection. Also, our method can accelerate the training over tenfold faster in certain cases. Finally, the experiments indicate the memory rate introduced in our ALRMR-GEC guides the GEC editscorer to learn more generalized knowledge. Zhixiao Wu, Yao Lu 0008, Jie Wen 0001, Guangming Lu 0002 |
AAAI | 4 |
| 2025 | Among General Spine Segmentation with Multi-scale and DiscriminateFeature Fusion
Tingwei Wen, Yao Lu 0008, Xiaosheng Chen, Xinhai Lu, Guangming Lu 0002 |
CVM (1) | 5 |
| 2025 | Recognition-Synergistic Scene Text EditingabstractScene text editing aims to modify text content within scene images while maintaining style consistency. Traditional methods achieve this by explicitly disentangling style and content from the source image and then fusing the style with the target content, while ensuring content consistency using a pre-trained recognition model. Despite notable progress, these methods suffer from complex pipelines, leading to suboptimal performance in complex scenarios. In this work, we introduce Recognition-Synergistic Scene Text Editing (RS-STE), a novel approach that fully exploits the intrinsic synergy of text recognition for editing. Our model seamlessly integrates text recognition with text editing within a unified framework, and leverages the recognition model’s ability to implicitly disentangle style and content while ensuring content consistency. Specifically, our approach employs a multi-modal parallel decoder based on transformer architecture, which predicts both text content and stylized images in parallel. Additionally, our cyclic self-supervised fine-tuning strategy enables effective training on unpaired real-world data without ground truth, enhancing style and content consistency through a twice-cyclic generation process. Built on a relatively simple architecture, RS-STE achieves state-of-the-art performance on both synthetic and real-world benchmarks, and further demonstrates the effectiveness of leveraging the generated hard cases to boost the performance of downstream recognition tasks. Code is available at https://github.com/ZhengyaoFang/RS-STE. Zhengyao Fang, Pengyuan Lv, Chengquan Zhang, Jun Yu 0002, Guangming Lu 0002, Wenjie Pei |
CVPR | 6 |
| 2025 | Learning Compatible Multi-Prize Subnetworks for Asymmetric RetrievalabstractAsymmetric retrieval is a typical scenario in real-world retrieval systems, where compatible models of varying capacities are deployed on platforms with different resource configurations. Existing methods generally train pre-defined networks or subnetworks with capacities specifically designed for pre-determined platforms, using compatible learning. Nevertheless, these methods suffer from limited flexibility for multi-platform deployment. For example, when introducing a new platform into the retrieval systems, developers have to train an additional model at an appropriate capacity that is compatible with existing models via backward-compatible learning. In this paper, we propose a Prunable Network with self-compatibility, which allows developers to generate compatible subnetworks at any desired capacity through post-training pruning. Thus it allows the creation of a sparse subnetwork matching the resources of the new platform without additional training. Specifically, we optimize both the architecture and weight of subnetworks at different capacities within a dense network in compatible learning. We also design a conflict-aware gradient integration scheme to handle the gradient conflicts between the dense network and subnetworks during compatible learning. Extensive experiments on diverse benchmarks and visual backbones demonstrate the effectiveness of our method. The code will be made publicly available. Yushuai Sun, Zikun Zhou, Dongmei Jiang, Yaowei Wang 0001, Jun Yu 0002, Guangming Lu 0002, Wenjie Pei |
CVPR | 6 |
| 2025 | Enhancing Spatial Reasoning in Multimodal Large Language Models Through Reasoning-Based SegmentationabstractRecent advances in point cloud perception have demonstrated remarkable progress in scene understanding through vision-language alignment leveraging large language models (LLMs). However, existing methods may still encounter challenges in handling complex instructions that require accurate spatial reasoning, even if the 3D point cloud data provides detailed spatial cues such as size and position for identifying the targets. To tackle this issue, we propose Relevant Reasoning Segmentation (R$^2$S), a reasoning-based segmentation framework. The framework emulates human cognitive processes by decomposing spatial reasoning into two sequential stages: first identifying relevant elements, then processing instructions guided by their associated visual priors. Furthermore, acknowledging the inadequacy of existing datasets in complex reasoning tasks, we introduce 3D ReasonSeg, a reasoning-based segmentation dataset comprising 25,185 training samples and 3,966 validation samples with precise annotations. Both quantitative and qualitative experiments demonstrate that the R$^2$S and 3D ReasonSeg effectively endow 3D point cloud perception with stronger spatial reasoning capabilities, and we hope that they can serve as a new baseline and benchmark for future work. Zhenhua Ning, Zhuotao Tian, Shaoshuai Shi, Guangming Lu 0002, Daojing He, Wenjie Pei, Li Jiang 0009 |
ICCV | 4 |
| 2025 | D2 ST-Adapter: Disentangled-and-Deformable Spatio-Temporal Adapter for Few-Shot Action Recognition
Wenjie Pei, Qizhong Tan, Guangming Lu 0002, Jiandong Tian, Jun Yu 0002 |
ICCV | 3 |
| 2025 | SSHR: More Secure Generative Steganography with High-Quality Revealed Secret ImagesabstractImage steganography ensures secure information transmission and storage by concealing secret messages within images. Recently, the diffusion model has been incorporated into the generative image steganography task, with text prompts being employed to guide the entire process. However, existing methods are plagued by three problems: (1) the restricted control exerted by text prompts causes generated stego images resemble the secret images and seem unnatural, raising the severe detection risk; (2) inconsistent intermediate states between Denoising Diffusion Implicit Models and its inversion, coupled with limited control of text prompts degrade the revealed secret images; (3) the descriptive text of images(i.e. text prompts) are also deployed as the keys, but this incurs significant security risks for both the keys and the secret images.To tackle these drawbacks, we systematically propose the SSHR, which joints the Reference Images with the adaptive keys to govern the entire process, enhancing the naturalness and imperceptibility of stego images. Additionally, we methodically construct an Exact Reveal Process to improve the quality of the revealed secret images. Furthermore, adaptive Reference-Secret Image Related Symmetric Keys are generated to enhance the security of both the keys and the concealed secret images. Various experiments indicate that our model outperforms existing methods in terms of recovery quality and secret image security. Jiannian Wang, Yao Lu 0008, Guangming Lu 0002 |
ICML | 3 |
| 2025 | Efficient and Separate Authentication Image Steganography NetworkabstractImage steganography hides multiple images for multiple recipients into a single cover image. All secret images are usually revealed without authentication, which reduces security among multiple recipients. It is elegant to design an authentication mechanism for isolated reception. We explore such mechanism through sufficient experiments, and uncover that additional authentication information will affect the distribution of hidden information and occupy more hiding space of the cover image. This severely decreases effectiveness and efficiency in large-capacity hiding. To overcome such a challenge, we first prove the authentication feasibility within image steganography. Then, this paper proposes an image steganography network collaborating with separate authentication and efficient scheme. Specifically, multiple pairs of lock-key are generated during hiding and revealing. Unlike traditional methods, our method has two stages to make appropriate distribution adaptation between locks and secret images, simultaneously extracting more reasonable primary information from secret images, which can release hiding space of the cover image to some extent. Furthermore, due to separate authentication, fused information can be hidden in parallel with a single network rather than traditional serial hiding with multiple networks, which can largely decrease the model size. Extensive experiments demonstrate that the proposed method achieves more secure, effective, and efficient image steganography. Code is available at https://github.com/Revive624/Authentication-Image-Steganography. Junchao Zhou, Yao Lu 0008, Jie Wen 0001, Guangming Lu 0002 |
ICML | 4 |
| 2025 | Cause-Effect Driven Optimization for Robust Medical Visual Question Answering with Language BiasesabstractExisting Medical Visual Question Answering (Med-VQA) models often suffer from language biases, where spurious correlations between question types and answer categories are inadvertently established. To address these issues, we propose a novel Cause-Effect Driven Optimization framework called CEDO, that incorporates three well-established mechanisms, i.e., Modality-driven Heterogeneous Optimization (MHO), Gradient-guided Modality Synergy (GMS), and Distribution-adapted Loss Rescaling (DLR), for comprehensively mitigating language biases from both causal and effectual perspectives. Specifically, MHO employs adaptive learning rates for specific modalities to achieve heterogeneous optimization, thus enhancing robust reasoning capabilities. Additionally, GMS leverages the Pareto optimization method to foster synergistic interactions between modalities and enforce gradient orthogonality to eliminate bias updates, thereby mitigating language biases from the effect side, i.e., shortcut bias. Furthermore, DLR is designed to assign adaptive weights to individual losses to ensure balanced learning across all answer categories, effectively alleviating language biases from the cause side, i.e., imbalance biases within datasets. Extensive experiments on multiple traditional and bias-sensitive benchmarks consistently demonstrate the robustness of CEDO over state-of-the-art competitors. Huanjia Zhu, Yishu Liu 0001, Xiaozhao Fang, Guangming Lu 0002, Bingzhi Chen |
IJCAI | 4 |
| 2025 | Med-BiasX: Robust Medical Visual Question Answering with Language Biases
Huanjia Zhu, Yishu Liu 0001, Chengju Zhou, Guangming Lu 0002, Bingzhi Chen |
MICCAI (14) | 4 |
| 2025 | CauRDG: Enhancing Domain Generalization with Causal-Driven Semantic Consistency ReasoningabstractDomain generalization (DG) plays a pivotal role in enabling models to maintain robust performance across heterogeneous environments. However, existing DG methods are fundamentally constrained by two intertwined limitations: (1) causal misalignment, which stems from undifferentiated feature encoding that entangles causal mechanisms with environmental biases; (2)semantic conflict arises when conventional adaptation methods find it challenging to balance the preservation of class discriminability with the mitigation of domain-specific distribution discrepancies. To address these challenges of DG, we propose a novel Causal-Driven Semantic Consistency Reasoning (CauRDG) method, which synergistically integrates Prototype-Guided Causal Disentanglement (PGCD) and Dual-Space Semantic Disambiguation (DSSD). Specifically, PGCD constructs a causal framework that identifies stable relationships and decouples invariant mechanisms from domain-specific variations, preserving causal consistency while adapting to contextual differences. DSSD harnesses a dual-space paradigm, enhancing local categorical clarity and maintaining global conceptual unity, thus balancing domain-specific precision with cross-domain coherence. The robustness provided by CauRDG ensures robust extraction and interpretation of essential features by preserving invariant causal structures, thereby harmonizing discriminative semantics with domain-varying contexts. Extensive experiments on multiple benchmark datasets consistently demonstrate the effectiveness and superiority of our CauRDG over state-of-the-art baselines. Zongxin Liu 0003, Yishu Liu 0001, Guangming Lu 0002, Xiaoling Luo 0001, Bingzhi Chen |
ACM Multimedia | 3 |
| 2025 | PET-GPRA: Rethinking PET with Gradient-Aware Prompting and Router-Free Adapters for Few-shot Class-Incremental LearningabstractFew-Shot Class-Incremental Learning (FSCIL) aims to continuously learn novel concepts from limited training samples without forgetting previously encountered classes. Recent advancements have leveraged Parameter-Efficient Tuning (PET) strategies on pre-trained models to enhance FSCIL performance. However, current PET-based FSCIL approaches still suffer from the challenges posed by catastrophic collapse of general prompt and limited adaptability of specific prompt . To this end, we redefine the function of the PET paradigm with both gradient-aware prompting (GAP) and router-free adapters (RFA) to boost the performance of FSCIL, termed as "PET-GPRA". To dynamically balance the retention of previously learned general knowledge and the acquisition of novel class information across sessions, the GAP paradigm adaptively adjusts the updated gradient of the general prompt by leveraging the angular relationship between the general knowledge gradient and the novel knowledge gradient. Meanwhile, the RFA mechanism utilizes the semantic similarity between class attributes to replace the routing network, guiding the integration of adapter information, in which adapters serve as specific prompts to enhance the adaptability. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority and effectiveness of our proposed PET-GPRA framework over state-of-the-art baselines. Yishu Liu 0001, Desen Wang, Xiaoling Luo 0001, Bingzhi Chen, Guangming Lu 0002 |
ACM Multimedia | 6 |
| 2025 | EditInfinity: Image Editing with Binary-Quantized Generative ModelsabstractAdapting pretrained diffusion-based generative models for text-driven image editing with negligible tuning overhead has demonstrated remarkable potential. A classical adaptation paradigm, as followed by these methods, first infers the generative trajectory inversely for a given source image by image inversion, then performs image editing along the inferred trajectory guided by the target text prompts. However, the performance of image editing is heavily limited by the approximation errors introduced during image inversion by diffusion models, which arise from the absence of exact supervision in the intermediate generative steps. To circumvent this issue, we investigate the parameter-efficient adaptation of binary-quantized generative models for image editing, and leverage their inherent characteristic that the exact intermediate quantized representations of a source image are attainable, enabling more effective supervision for precise image inversion. Specifically, we propose EditInfinity, which adapts Infinity, a binary-quantized generative model, for image editing. We propose an efficient yet effective image inversion mechanism that integrates text prompting rectification and image style preservation, enabling precise image inversion. Furthermore, we devise a holistic smoothing strategy which allows our EditInfinity to perform image editing with high fidelity to source images and precise semantic alignment to the text prompts. Extensive experiments on the PIE-Bench benchmark across add, change, and delete editing operations, demonstrate the superior performance of our model compared to state-of-the-art diffusion-based baselines. Code available at: https://github.com/yx-chen-ust/EditInfinity. Jun Yu 0002, Guangming Lu 0002, Wenjie Pei |
NeurIPS | 4 |
| 2025 | A Set of Generalized Components to Achieve Effective Poison-only Clean-label Backdoor Attacks with Collaborative Sample Selection and TriggersabstractPoison-only Clean-label Backdoor Attacks (PCBAs) aim to covertly inject attacker-desired behavior into DNNs by merely poisoning the dataset without changing the labels. To effectively implant a backdoor, multiple triggers are proposed for various attack requirements of Attack Success Rate (ASR) and stealthiness. Additionally, sample selection enhances clean-label backdoor attacks' ASR by meticulously selecting "hard'' samples instead of random samples to poison. Current methods, however, 1) usually handle the sample selection and triggers in isolation, leading to severely limited improvements on both ASR and stealthiness. Consequently, attacks exhibit unsatisfactory performance on evaluation metrics when converted to PCBAs via a mere stacking of methods. Therefore, we seek to explore the bi-directional collaborative relations between the sample selection and triggers to address the above dilemma. 2) Since the strong specificity within triggers, the simple combination of sample selection and triggers fails to substantially enhance both evaluation metrics, with generalization preserved among various attacks. Therefore, we seek to propose a set of components to significantly improve both stealthiness and ASR based on the commonalities of attacks. Specifically, Component A ascertains two critical selection factors, and then makes them an appropriate combination based on the trigger scale to select more reasonable "hard'' samples for improving ASR. Component B is proposed to select samples with similarities to relevant trigger implanted samples to promote stealthiness. Component C reassigns trigger poisoning intensity on RGB colors through distinct sensitivity of the human visual system to RGB for higher ASR, with stealthiness ensured by sample selection including Component B. Furthermore, all components can be strategically integrated into diverse PCBAs, enabling tailored solutions that balance ASR and stealthiness enhancement for specific attack requirements. Extensive experiments demonstrate the superiority of our components in stealthiness, ASR, and generalization. Our code will be released as soon as possible. Zhixiao Wu, Yao Lu 0008, Jie Wen 0001, Guangming Lu 0002 |
NeurIPS | 6 |
| 2025 | BiTA: Bi-directional tuning for lossless acceleration in large language models
Feng Lin 0009, Hanling Yi, Xiaotian Yu, Guangming Lu 0002, Rong Xiao 0003 |
Expert Syst. Appl. | 6 |
| 2025 | Efficient U-shape invertible neural network for large-capacity image steganography
Le Zhang 0016, Yao Lu 0008, Yuanrong Xu, Guangming Lu 0002 |
J. Inf. Secur. Appl. | 5 |
| 2025 | Individualized image steganography method with Dynamic Separable Key and Adaptive Redundancy Anchor
Junchao Zhou, Yao Lu 0008, Guangming Lu 0002 |
Knowl. Based Syst. | 3 |
| 2025 | Caption Assisted Multimodal Large Language Model for Video Moment RetrievalabstractMultimodal Large Language Models (MLLMs) have demonstrated significant potential across various multimodal tasks, including retrieval, summarization, and reasoning. However, it remains a substantial challenge for MLLMs to understand and precisely retrieve specific moments from a video, which require fine-grained spatial and temporal understanding of a video. To overcome this, we propose the Caption Assisted MLLM from Coarse to finE (CALCE), a novel two-stage framework designed for enhanced moment retrieval. Our pipeline begins with a first stage where captions extracted from the audio are utilized to assist the MLLM to provide a robust foundation for precise moment retrieval. To efficiently manage memory consumption from this additional data, a clustering algorithm is applied to the sparsely sampled video frames, categorizing them into key frames and non-key frames. The second stage focuses on recalling missed moments and achieving more fine-grained moment boundaries by adopting a higher sampling rate. In this process, predictions from the first stage cast votes for their correlated densely sampled frames, thereby filtering out less relevant frames. By repeating the process of the first stage with these selected frames, CALCE progressively retrieves video moments from coarse to precise. Experiments on QVHighlights and Charades-STA demonstrate the effectiveness of CALCE, which outperforms existing state-of-the-art methods. The code is available at https://github.com/tjhd1475/CALCE. Peiyu Xie, Jinxing Li 0003, Guangming Lu 0002, Yong Xu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | LBF-VQA: Towards Language Bias-Free Visual Question Answering With Multi-Space Collaborative Debiasing
Yishu Liu 0001, Huanjia Zhu, Bingzhi Chen, Xiaozhao Fang, Guangming Lu 0002, Shengli Xie 0001 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | Toward Robust Semi-Supervised Distribution Alignment Against Label Distribution Shift With Noisy AnnotationsabstractDeep learning-based AI models typically require a large amount of high-quality annotated data to achieve optimal performance. However, thelabel distribution shiftcaused by noisy annotations can lead to perturbations in the classification boundary, reducing the robustness and generalization capabilities of deep learning models. To mitigate this issue, we transform the problem of learning from noisy labels into a semi-supervised learning problem, and propose a novel Semi-Supervised Distribution Alignment (SSDA) framework that strategically integrates noise-robust distribution alignment within a unified semi-supervised learning paradigm for combating noisy labels. By leveraging the similarity distribution between historical predictions, the proposed SSDA approach benefits from a flexible multi-historical regression modeling strategy, which aims to identify high-confidence samples/pairs and recalibrate the label shift through pseudo-labels. Furthermore, our approach employs a comprehensive multi-granularity distribution adaptation strategy, incorporating both instance-wise and class-aware distribution alignment to quantitatively minimize semantic discrepancies across different mixed feature domains. In this way, our SSDA approach ultimately achieves more resilient and generalizable performance against label noise, even in the presence of substantial noise. Extensive experiments conducted on multiple simulated and real-world noisy benchmark datasets consistently demonstrate the superiority and effectiveness of our SSDA method compared to existing state-of-the-art baselines. Bingzhi Chen, Zhanhao Ye, Yishu Liu 0001, Xiaozhao Fang, Guangming Lu 0002, Shengli Xie 0001, Xuelong Li 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Focus Affinity Perception and Super-Resolution Embedding for Multifocus Image FusionabstractDespite the fact that there is a remarkable achievement on multifocus image fusion, most of the existing methods only generate a low-resolution image if the given source images suffer from low resolution. Obviously, a naive strategy is to independently conduct image fusion and image super-resolution. However, this two-step approach would inevitably introduce and enlarge artifacts in the final result if the result from the first step meets artifacts. To address this problem, in this article, we propose a novel method to simultaneously achieve image fusion and super-resolution in one framework, avoiding step-by-step processing of fusion and super-resolution. Since a small receptive field can discriminate the focusing characteristics of pixels in detailed regions, while a large receptive field is more robust to pixels in smooth regions, a subnetwork is first proposed to compute the affinity of features under different types of receptive fields, efficiently increasing the discriminability of focused pixels. Simultaneously, in order to prevent from distortion, a gradient embedding-based super-resolution subnetwork is also proposed, in which the features from the shallow layer, the deep layer, and the gradient map are jointly taken into account, allowing us to get an upsampled image with high resolution. Compared with the existing methods, which implemented fusion and super-resolution independently, our proposed method directly achieves these two tasks in a parallel way, avoiding artifacts caused by the inferior output of image fusion or super-resolution. Experiments conducted on the real-world dataset substantiate the superiority of our proposed method compared with state of the arts. Huafeng Li 0001, Jinxing Li 0003, Yu Liu 0023, Guangming Lu 0002, Yong Xu 0001, Zhengtao Yu 0001, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | CariesXrays: Enhancing Caries Detection in Hospital-Scale Panoramic Dental X-rays via Feature Pyramid Contrastive LearningabstractDental caries has been widely recognized as one of the most prevalent chronic diseases in the field of public health. Despite advancements in automated diagnosis across various medical domains, it remains a substantial challenge for dental caries detection due to its inherent variability and intricacies. To bridge this gap, we release a hospital-scale panoramic dental X-ray benchmark, namely “CariesXrays”, to facilitate the advancements in high-precision computer-aided diagnosis for dental caries. It comprises 6,000 panoramic dental X-ray images, with a total of 13,783 instances of dental caries, all meticulously annotated by dental professionals. In this paper, we propose a novel Feature Pyramid Contrastive Learning (FPCL) framework, that jointly incorporates feature pyramid learning and contrastive learning within a unified diagnostic paradigm for automated dental caries detection. Specifically, a robust dual-directional feature pyramid network (D2D-FPN) is designed to adaptively capture rich and informative contextual information from multi-level feature maps, thus enhancing the generalization ability of caries detection across different scales. Furthermore, our model is augmented with an effective proposals-prototype contrastive regularization learning (P2P-CRL) mechanism, which can flexibly bridge the semantic gaps among diverse dental caries with varying appearances, resulting in high-quality dental caries proposals. Extensive experiments on our newly-established CariesXrays benchmark demonstrate the potential of FPCL to make a significant social impact on caries diagnosis. Bingzhi Chen, Sisi Fu, Yishu Liu 0001, Jiahui Pan 0003, Guangming Lu 0002, Zheng Zhang 0006 |
AAAI | 5 |
| 2024 | SA²VP: Spatially Aligned-and-Adapted Visual PromptabstractAs a prominent parameter-efficient fine-tuning technique in NLP, prompt tuning is being explored its potential in computer vision. Typical methods for visual prompt tuning follow the sequential modeling paradigm stemming from NLP, which represents an input image as a flattened sequence of token embeddings and then learns a set of unordered parameterized tokens prefixed to the sequence representation as the visual prompts for task adaptation of large vision models. While such sequential modeling paradigm of visual prompt has shown great promise, there are two potential limitations. First, the learned visual prompts cannot model the underlying spatial relations in the input image, which is crucial for image encoding. Second, since all prompt tokens play the same role of prompting for all image tokens without distinction, it lacks the fine-grained prompting capability, i.e., individual prompting for different image tokens. In this work, we propose the Spatially Aligned-and-Adapted Visual Prompt model (SA^2VP), which learns a two-dimensional prompt token map with equal (or scaled) size to the image token map, thereby being able to spatially align with the image map. Each prompt token is designated to prompt knowledge only for the spatially corresponding image tokens. As a result, our model can conduct individual prompting for different image tokens in a fine-grained manner. Moreover, benefiting from the capability of preserving the spatial structure by the learned prompt token map, our SA^2VP is able to model the spatial relations in the input image, leading to more effective prompting. Extensive experiments on three challenging benchmarks for image classification demonstrate the superiority of our model over other state-of-the-art methods for visual prompt tuning. Code is available at https://github.com/tommy-xq/SA2VP. Wenjie Pei, Tongqi Xia, Fanglin Chen 0001, Jiandong Tian, Guangming Lu 0002 |
AAAI | 6 |
| 2024 | Robust 3D Tracking with Quality-Aware Shape Completionabstract3D single object tracking remains a challenging problem due to the sparsity and incompleteness of the point clouds. Existing algorithms attempt to address the challenges in two strategies. The first strategy is to learn dense geometric features based on the captured sparse point cloud. Nevertheless, it is quite a formidable task since the learned dense geometric features are with high uncertainty for depicting the shape of the target object. The other strategy is to aggregate the sparse geometric features of multiple templates to enrich the shape information, which is a routine solution in 2D tracking. However, aggregating the coarse shape representations can hardly yield a precise shape representation. Different from 2D pixels, 3D points of different frames can be directly fused by coordinate transform, i.e., shape completion. Considering that, we propose to construct a synthetic target representation composed of dense and complete point clouds depicting the target shape precisely by shape completion for robust 3D tracking. Specifically, we design a voxelized 3D tracking framework with shape completion, in which we propose a quality-aware shape completion mechanism to alleviate the adverse effect of noisy historical predictions. It enables us to effectively construct and leverage the synthetic target representation. Besides, we also develop a voxelized relation modeling module and box refinement module to improve tracking performance. Favorable performance against state-of-the-art algorithms on three benchmarks demonstrates the effectiveness and generalization ability of our method. Zikun Zhou, Guangming Lu 0002, Jiandong Tian, Wenjie Pei |
AAAI | 3 |
| 2024 | Progressive Stepwise Diffusion Model with Dual Decoders for Semi-Supervised Medical Image SegmentationabstractSemi-supervised medical image segmentation tasks aim to harness the potential of vast amounts of unlabeled data using a limited amount of annotated data. Denoising Diffusion Probabilistic Models, which have achieved significant success in image generation, are gradually being explored for their potential in semantic image segmentation. However, their application in semi-supervised medical image segmentation is still in its early stages. Initially, due to the high randomness of diffusion models, the pseudo-labels generated during the early training phase may mislead the processing of unlabeled data. Additionally, the use of fixed-time steps for random sampling during training limits the ability of the model to learn effective denoising functions at an early stage. To address these issues, we propose an innovative framework named Progressive Stepwise Diffusion Network with Dual Decoders (PSDD) for semi-supervised medical image segmentation. This framework incorporates an additional normal decoder into the denoising diffusion encoder-decoder structure to provide more accurate labels and employs a Progressive Incremental Step strategy to gradually train the model for longer generation processes. Evaluated on two 2D colon polyp segmentation datasets and a 3D Left Atrium dataset, the experimental results demonstrate significant performance improvements over current advanced methods, thereby validating the effectiveness and potential of this framework in handling complex semi-supervised learning scenarios. Xiaolin Huang, Jingchun Lin, Bingzhi Chen, Guangming Lu 0002 |
BIBM | 6 |
| 2024 | Domain-Rectifying Adapter for Cross-Domain Few-Shot SegmentationabstractFew-shot semantic segmentation (FSS) has achieved great success on segmenting objects of novel classes, supported by only a few annotated samples. However, existing FSS methods often underperform in the presence of domain shifts, especially when encountering new domain styles that are unseen during training. It is suboptimal to directly adapt or generalize the entire model to new domains in the few-shot scenario. Instead, our key idea is to adapt a small adapter for rectifying diverse target domain styles to the source domain. Consequently, the rectified target domain features can fittingly benefit from the well-optimized source domain segmentation model, which is intently trained on sufficient source domain data. Training domain-rectifying adapter requires sufficiently diverse target domains. We thus propose a novel local-global style perturbation method to simulate diverse potential target domains by perturbating the feature channel statistics of the individual images and collective statistics of the entire source domain, respectively. Additionally, we propose a cyclic domain alignment module to facilitate the adapter effectively rectifying domains using a reverse domain rectification supervision. The adapter is trained to rectify the image features from diverse synthesized target domains to align with the source domain. During testing on target domains, we start by rectifying the image features and then conduct few-shot segmentation on the domain-rectified features. Extensive experiments demonstrate the effectiveness of our method, achieving promising results on cross-domain few-shot semantic segmentation tasks. Our code is available at https://github.com/Matt-Su/DR-Adapter. Jiapeng Su, Wenjie Pei, Guangming Lu 0002, Fanglin Chen 0001 |
CVPR | 4 |
| 2024 | WeCromCL: Weakly Supervised Cross-Modality Contrastive Learning for Transcription-Only Supervised Text Spotting
Zhengyao Fang, Pengyuan Lv, Chengquan Zhang, Fanglin Chen 0001, Guangming Lu 0002, Wenjie Pei |
ECCV (31) | 6 |
| 2024 | UniVoxel: Fast Inverse Rendering by Unified Voxelization of Scene Representation
Songlin Tang, Guangming Lu 0002, Jianzhuang Liu, Wenjie Pei |
ECCV (71) | 3 |
| 2024 | Decoupled Self-Adaptive Distribution Regularization for Few-Shot Image ClassificationabstractThe feature dispersion, arising from the inherent constraints of data scarcity, has emerged as a prominent challenge in the domain of few-shot learning. In this paper, we propose a novel Self-adaptive Distribution Regularization (SADR) approach, which can adaptively bridge the semantic gaps across distribution patterns and boundaries for learning from limited labeled data. Specifically, the technical core of our SADR approach is to decouple the feature embeddings into two discrete spaces: the intra-class and inter-class distributions, leading to robust and discriminative feature representations in a self-adaptive manner. To achieve meticulous similarity measurements while mitigating redundant feature information, an innovative regularized Brownian Distance Covariance (R-BDC) metric is strategically designed to simultaneously explore both the joint and marginal distributions present among diverse input samples. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our SADR approach over state-of-the-art baselines. Bingzhi Chen, Haoming Zhou, Yishu Liu 0001, Guangming Lu 0002, Zheng Zhang 0006 |
ICASSP | 5 |
| 2024 | Rethinking Adversarial Robustness Distillation VIA Strength-Dependent Adaptive RegularizationabstractDespite the progress achieved by existing adversarial distillation (AD) approaches, most mainstream models suffer from inadequate adversarial robustness, due to the challenges of fixed attack strength and unreliable teacher guidance. In this paper, we propose a novel Strength-Dependent Adaptive Regularization (SDAR) paradigm to reinforce the function of adversarial distillation with strength-adaptive adversarial attack (SAA) and multi-dimensional knowledge distillation (MKD). Different from the traditional adversarial training (AT) methods, the proposed SAA scheme dynamically assigns an adaptive and efficient attack strength for each instance, which aims to facilitate smoother classification boundaries. By incorporating dynamic strength coefficients, a comprehensive MKD strategy is designed to fully explore the valuable context information and narrow distribution discrepancies across teacher-student domains. Particularly, our SDAR paradigm can seamlessly integrate with the current AD frameworks, further enhancing the adversarial robustness of deep learning models. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of SDAR over state-of-the-art baselines. Bingzhi Chen, Shuobin Lin, Yishu Liu 0001, Zheng Zhang 0006, Guangming Lu 0002, Lewei He |
ICME | 5 |
| 2024 | Enhancing Few-Shot Classification without Forgetting Through Multi-level Contrastive ConstraintsabstractMost recent few-shot learning approaches are based on meta-learning with episodic training. However, prior studies encounter two crucial problems: (1) the presence of inductive bias, and (2) the occurrence of catastrophic forgetting. In this paper, we propose a novel Multi-Level Contrastive Constraints (MLCC) framework, that jointly integrates within-episode learning and across-episode learning into a unified interactive learning paradigm to solve these issues. Specifically, we employ a space-aware interaction modeling scheme to explore the correct inductive paradigms for each class between within-episode similarity/dis-similarity distributions. Additionally, with the aim of better utilizing former prior knowledge, a cross-stage distribution adaption strategy is designed to align the across-episode distributions from different time stages, thus reducing the semantic gap between existing and past prediction distribution. Extensive experiments on multiple few-shot datasets demonstrate the consistent superiority of MLCC approach over the existing state-of-the-art baselines. Bingzhi Chen, Haoming Zhou, Yishu Liu 0001, Jiahui Pan 0003, Guangming Lu 0002 |
ICME | 6 |
| 2024 | Efficient U-Shape Invertible Neural Network for Image SteganographyabstractCurrently, it is challenging to recover high-quality secret images from highly secure stego images while maintaining a limited computational cost for image steganography. This paper proposes an Efficient U-shape Invertible Neural Network (EUIN-Net) for image steganography. Due to the gradual fusion and separation properties of the U-shape invertible mechanism, our EUIN-Net comprehensively couples and decouples the secret-cover information on different scales and depths. Besides, using the skip connections between each pair of U-shape invertible blocks, the long-range dependency can be retrieved. Such above two factors can drive our EUIN-Net to promote the quality of both stego and revealed secret images. Furthermore, the shared and multi-scale characteristics of the U-shaped invertible blocks during the hiding and revealing stages contribute to significant reductions of our EUIN-Net in the model size and Flops. Extensive experiments demonstrate that the proposed EUIN-Net is efficient and can achieve state-of-the-art performances for image steganography. Le Zhang 0016, Yao Lu 0008, Mi-Xiao Hou, Guangming Lu 0002 |
ICME | 5 |
| 2024 | Robust Visual Question Answering With Contrastive-Adversarial Consistency ConstraintsabstractVisual cues and question semantics contribute to final answer predictions from distinct perspectives. However, inherent language bias confounds the relationship between visual and question cues, leading to a misguided preference for question semantics. Different from the existing studies that focus on inter-class discrimination, this paper proposes a robust visual question answering framework with contrastive-adversarial consistency constraints (CACC) at both inter- and intra-instance levels. From a fine-grained instance-level perspective, our approach initially introduces an effective inter-instance contrastive constraint to perform adaptive bias rectification. To enhance intra-instance invariance and reduce information redundancy, we refine the concept of semantic structure relationships by constructing intra-instance adversarial constraints using the Hilbert-Schmidt Independence Criterion (HSIC) independence criterion. Benefitting from both inter- and intra-instance perspectives, our method can effectively alleviate these language biases, enhancing the overall robustness of the representation. Extensive experiments on multiple benchmark datasets consistently demonstrate the superiority of our CACC over state-of-the-art baselines. Meirong Ding, Yishu Liu 0001, Guangming Lu 0002, Bingzhi Chen |
ICME | 5 |
| 2024 | Enhancing Cross-Modal Retrieval via Visual-Textual Prompt Hashing
Bingzhi Chen, Zhongqi Wu, Yishu Liu 0001, Guangming Lu 0002, Zheng Zhang 0006 |
IJCAI | 5 |
| 2024 | QFormer: An Efficient Quaternion Transformer for Image Denoising
Bo Jiang 0017, Yao Lu 0008, Guangming Lu 0002, Bob Zhang 0001 |
IJCAI | 3 |
| 2024 | Implicit Prompt Learning for Image Denoising
Yao Lu 0008, Bo Jiang 0017, Guangming Lu 0002, Bob Zhang 0001 |
IJCAI | 3 |
| 2024 | Medical Cross-Modal Prompt Hashing with Robust Noisy Correspondence Learning
Yishu Liu 0001, Zhongqi Wu, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002 |
MICCAI (3) | 5 |
| 2024 | Stay Focused is All You Need for Adversarial Robustness
Bingzhi Chen, Ruihan Liu, Yishu Liu 0001, Xiaozhao Fang, Jiahui Pan 0003, Guangming Lu 0002, Zheng Zhang 0006 |
ACM Multimedia | 6 |
| 2024 | Prototype-Guided Dual-Transformer Reasoning for Video Individual CountingabstractVideo Individual Counting (VIC), which focuses on accurately tallying the total number of individuals in a video without duplication, is crucial for urban public space management and densely-populated areas planning. Existing methods suffer from limitations in terms of expensive manual annotation, and the efficiency of location or detection algorithms. In this work, we contribute a novel Prototype-guided Dual-Transformer Reasoning framework, termed PDTR, which takes both similarity and difference of adjacent frames into account to achieve accurate counting in an end-to-end regression manner. Specifically, we first design a multi-receptive field feature fusion module to acquire initial comprehensive representations. Subsequently, the dynamic prototype generation module memorizes consistent representations of similar information to generate prototypes. Additionally, to further dig out the shared and private features from different frames, a prototype cross-guided decoder and a privacy-decoupling module are designed. Extensive experiments conducted on two existing VIC datasets, consistently demonstrate the superiority of PDTR over state-of-the-art baselines. Yishu Liu 0001, Huafeng Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
ACM Multimedia | 5 |
| 2024 | Frequency Adapter and Spatial Prompt Network for All-in-One Blind Image Restoration
Shuoming Chen, Wenjie Pei, Yao Lu 0008, Guangming Lu 0002 |
PRCV (8) | 4 |
| 2024 | Saliency-aware regularized graph neural network
Wenjie Pei, Weina Xu, Weichao Li 0001, Jinfan Wang, Guangming Lu 0002, Xiangrong Wang 0002 |
Artif. Intell. | 6 |
| 2024 | A coarse-to-fine registration network based on affine transformation and multi-scale pyramid
Dongming Li 0002, Yingjian Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
Expert Syst. Appl. | 4 |
| 2024 | Universal Object Detection with Large Vision Model
Feng Lin 0009, Wenze Hu, Yaowei Wang 0001, Yonghong Tian 0001, Guangming Lu 0002, Fanglin Chen 0001, Yong Xu 0007, Xiaoyu Wang 0002 |
Int. J. Comput. Vis. | 5 |
| 2024 | Context-aware graph embedding with gate and attention for session-based recommendation
Junlong Chi, Peilin Hong, Guangming Lu 0002, David Zhang 0001, Bingzhi Chen |
Neurocomputing | 4 |
| 2024 | Multi-modal graph context extraction and consensus-aware learning for emotion recognition in conversationabstractMulti-modal emotion recognition in conversation is challenging because of the difficulty to jointly leverage the information from heterogeneous text, acoustic, and visual modalities . Recent context-aware methods usually design a graph structure to model dependencies of utterances and speakers, or integrate Multi-modal information. However, they typically lack a sufficient extraction of unimodal context, and rarely explore the emotion consensus prototypes among different samples with the same label. For solving these problems, in this paper, we propose a Graph Context extraction and Consensus-aware Learning (GCCL) framework to excavate context-sensitive fusion features and simulate the emotion evocation process during the emotion consensus learning. Specifically, GCCL contains a well-designed graph-based module to capture speaker, temporal and modality dependencies and integrate information from different modalities. Then, we design an emotion consensus learning unit to mine the most typical feature of each category in each modality. A speaker-guided contrastive learning loss is further proposed to guarantee the diversity between different individuals and the semantic consistency between distinct modalities. Moreover, we construct a consensus-aware unit with an attention-based memory mechanism to preserve semantic correlations among different samples on the category-level. Extensive experimental results on two conversational datasets demonstrate that the proposed GCCL outperforms the state-of-art methods. Code is available at https://github.com/gityider/GCCL . Yijing Dai, Jinxing Li 0003, Yingjian Li 0001, Guangming Lu 0002 |
Knowl. Based Syst. | 4 |
| 2024 | Exploring the complementarity between convolution and transformer matching for visual tracking
Zheng'ao Wang, Ming Li 0028, Wenjie Pei, Guangming Lu 0002, Fanglin Chen 0001 |
Knowl. Based Syst. | 4 |
| 2024 | Contrastive feature decomposition for single image layer separation
Xin Feng 0005, Haobo Ji, Wenjie Pei, Guangming Lu 0002, David Zhang 0001 |
Neural Comput. Appl. | 5 |
| 2024 | Unifying emotion-oriented and cause-oriented predictions for emotion-cause pair extraction
Guimin Hu, Yi Zhao 0007, Guangming Lu 0002 |
Neural Networks | 3 |
| 2024 | Improving Representation With Hierarchical Contrastive Learning for Emotion-Cause Pair ExtractionabstractEmotion-cause pair extraction (ECPE) aims to extract emotions and their corresponding cause from a document. The previous works have made great progress. However, there exist two major issues in existing works. First, most existing works mainly focus on the semantic relation between the emotion clause and cause clause, ignoring their inner statistical relation in representation space. Second, the existing works are sensitive to the relative position between the emotion clause and cause clause, which damages the model's robustness. To address the two issues, we propose a hierarchical contrastive learning framework (HCL-ECPE), which hierarchically performs contrastive learning on representation from two levels. The first level is inter-clause contrastive learning (ICCL), which performs between emotion clause and cause clause through mutual information maximization. The second level is intra-pair contrastive learning (IPCL), which performs between clause representation and pair representation through contrastive predictive coding (CPC). HCL-ECPE integrates ICCL and IPCL modules to explore the statistical relations between the emotion clause, cause clause, and their constructed emotion-cause pair from the perspective of mutual information, thereby improving the model performance and robustness. Experimental results on two public datasets, ECPED and RECCON, demonstrate that HCL-ECPE outperforms the most competitive baselines. Furthermore, ICCL and IPCL are orthogonal to the existing model, and introducing them into the current models updates state-of-the-art performance. Guimin Hu, Yi Zhao 0007, Guangming Lu 0002 |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | Multimodal Decoupled Distillation Graph Neural Network for Emotion Recognition in ConversationabstractGraph Neural Networks (GNNs) have attracted increasing attentions for multimodal Emotion Recognition in Conversation (ERC) due to their good performance in contextual understanding. However, most existing GNN-based methods suffer from two challenges: 1) How to explore and propagate appropriate information in a conversational graph. Typical GNNs in ERC neglect to mine the emotion commonality and discrepancy in the local neighborhood, leading to learn similar embbedings for connected nodes. However, the embeddings of these connected nodes are supposed to be distinguishable as they belong to different speakers with different emotions. 2) Most existing works apply simple concatenation or co-occurrence prior for modality combination, failing to fully capture the emotional information of multiple modalities in relationship modeling. In this paper, we propose a multimodal Decoupled Distillation Graph Neural Network (D2GNN) to address the above challenges. Specifically, D2GNN decouples the input features into emotion-aware and emotion-agnostic ones on the emotion category-level, aiming to capture emotion commonality and implicit emotion information, respectively. Moreover, we design a new message passing mechanism to separately propagate emotion-aware and -agnostic knowledge between nodes according to speaker dependency in two GNN-based modules, exploring the correlations of utterances and alleviating the similarities of embeddings. Furthermore, a multimodal distillation unit is performed to obtain the distinguishable embeddings by aggregating unimodal decoupled features. Experimental results on two ERC benchmarks demonstrate the superiority of the proposed model. Code is available at https://github.com/gityider/D2GNN. Yijing Dai, Yingjian Li 0001, Dongpeng Chen, Jinxing Li 0003, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | U²-Former: Nested U-Shaped Transformer for Image Restoration via Multi-View Contrastive LearningabstractWhile Transformer has achieved remarkable performance in various high-level vision tasks, it is still challenging to exploit the full potential of Transformer in image restoration. The crux lies in the limited depth of applying Transformer in the typical encoder-decoder framework for image restoration, resulting from heavy self-attention computation load and inefficient communications across different depth (scales) of layers. In this paper, we present a deep and effective Transformer-based network for image restoration, termed as U2-Former, which is able to employ self-attention of Transformer as the core operation for feature learning to perform image restoration in a deep encoding and decoding space. Specifically, it leverages the nested U-shaped structure to facilitate the interactions across different layers with different scales of feature maps. Furthermore, we optimize the computational efficiency for the basic Transformer block by introducing a simple yet effective feature-filtering mechanism to compress the token representation. Apart from the typical supervision ways for image restoration, our U2-Former also performs multi-view contrastive learning, which constructs positive pairs in various aspects, to learn noise-sensitive but content-irrelevant features and further decouple the noise component from the background image. Extensive experiments on various image restoration tasks, including reflection removal, rain streak removal and dehazing respectively, demonstrate the effectiveness of the proposed U2-Former. Xin Feng 0005, Haobo Ji, Wenjie Pei, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Robust Tracking via Fully Exploring Background Prior KnowledgeabstractTypical Siamese-based trackers focus on the target region and pay less attention to the background area. However, the background area can provide the tracker with prior knowledge about the target surroundings. Nonetheless, since the tracker can naturally utilize the target template for localization, importing additional background knowledge requires proper design so that the background area prior knowledge can be fully explored. Furthermore, the introduction of the entire background regions is redundant. Instead, the part background distractors in the regions are more meaningful for the discrimination of the tracker. In this work, we propose a background prior knowledge fully explored tracker for robust tracking. Firstly, we present a Transformer-based explicitly and fully background-utilizing scheme by boosting the tracker to independently exploit the background for localization. Specifically, a target-distractor independent decoder explicitly utilizes the background knowledge by making the target and the distractors independently perform fusion with the search feature. Secondly, we design a simple yet efficient discriminative distractors mining module to refine the background prior knowledge by replacing the whole background region with the mined background distractors. Extensive experiments demonstrate that the proposed method performs favorably against state-of-the-art trackers on nine benchmarks. Zheng'ao Wang, Zikun Zhou, Fanglin Chen 0001, Jun Xu 0008, Wenjie Pei, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Text Position-Aware Pixel Aggregation Network With Adaptive Gaussian Threshold: Detecting Text in the WildabstractOver recent years, deep learning has significantly boosted scene text detection performance, and current segmentation-based scene text detectors can achieve compact bounding boxes for irregular texts. However, it is also challenging to tackle crowded or overlapping texts for these existing methods due to conglutination between adjacent text instances in segmentation results. To address these issues, we propose a more accurate scene text detector, Text Position-Aware Pixel Aggregation Network, termed TPPAN. Specifically, a Gaussian threshold representation is adaptively learned instead of a constant setting in Adaptively Text Kernel Thresholding (ATKT) module to obtain more accurate text kernels. Then Text Position-Aware Region Pixel Aggregation (TPAR-PA) module predicts the text regions in relative positions and generates more accurate text contours. Adequate experiments have demonstrated that the resulting detector has achieved state-of-the-art performance on multi-oriented and curved scene text benchmarks. Jiayu Xu 0002, Ailiang Lin, Jinxing Li 0003, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Deep Fuzzy Multiteacher Distillation Network for Medical Visual Question AnsweringabstractMedical visual question answering (medical VQA) is a critical cross-modal interaction task that garnered considerable attention in the medical domain. Several existing methods commonly leverage the vision-and-language pretraining paradigms to mitigate the limitation of small-scale data. Nevertheless, most of them still suffer from two challenges that remain for further research: 1) limited research focuses on distilling representation from a complete modality to guide the representation learning of masked data in other modalities. 2) Multimodal fusion based on self-attention mechanisms cannot effectively handle the inherent uncertainty and vagueness of information interaction across modalities. To mitigate these issues, in this article, we propose a novel deep fuzzy multiteacher distillation (DFMD) network for medical VQA, which can take advantage of fuzzy logic to model the uncertainties from vison-language representations across modalities in a multiteacher framework. Specifically, a multiteacher knowledge distillation module is conceived to assist in reconstructing the missing semantics under the supervision signal generated by teachers from the other complete modality, achieving more robust semantic interaction across modalities. Incorporating insights from the fuzzy logic theory, we propose a noise-robust encoder called FuzBERT that enables our DFMD model to reduce the imprecision and ambiguity in feature representation during the multimodal interaction process. To the best of our knowledge, our work isthe first attemptto combine the fuzzy logic theory with the transformer-based encoder to effectively learn multimodal representation for medical VQA. Experimental results on the VQA-RAD and SLAKE datasets consistently demonstrate the superiority of our proposed DFMD method over state-of-the-art baselines. Yishu Liu 0001, Bingzhi Chen, Shuihua Wang, Guangming Lu 0002, Zheng Zhang 0006 |
IEEE Trans. Fuzzy Syst. | 4 |
| 2024 | Disentanglement Learning With Adaptive Centroid Alignment for Multiple Target Domains Fault DiagnosisabstractMost current domain adaptation methods for fault diagnosis focus on single target domain. However, test data often comes from multiple target domains, as machines work under different operating conditions, subsequently generating a more complex and extensive distribution of target data. Unfortunately, single target domain adaptation methods are not adaptive for multiple target domains adaptation (MTDA), which results in transfer performance degradation. To this end, a novel disentanglement learning with adaptive centroid alignment is proposed for MTDA. Specifically for disentanglement learning, two encoders and two classifiers are constructed independently for fault-related and domain-related feature extractions and classifications. Followed by the dual-adversarial strategy, only fault-related but domain-irrelevant features are extracted. Furthermore, to achieve the category alignment, we also propose an adaptive centroid alignment strategy, so that the feature centroids of the same fault category in different domains are enforced to be close to each other. Extensive experiments demonstrate the superiority of our proposed method compared with other popular approaches. Yu Gao 0024, Xutao Zheng, Jinxing Li 0003, Lijun Zong, Hongpeng Yin, Huafeng Li 0001, Guangming Lu 0002 |
IEEE Trans. Ind. Informatics | 7 |
| 2024 | AGP-Net: Adaptive Graph Prior Network for Image DenoisingabstractImage denoising is a critical problem in industrial information applications since noisy images can have adverse effects on the performance of many industrial tasks. Currently, Transformer structures and graph convolutional networks (GCNs) have been widely employed in image denoising to capture long-range dependencies for the performance promotion. These methods, however, severely suffer from three major problems. Initially, the long-range dependencies captured by Transformers and GCNs are only focused on the pixel level and patch level, respectively. This leads to the coarse retrieved feature, hindering further performance promotion. In addition, due to the limited training data, especially for the noisy images with highly diverse and complex noise, the denoising process may lack sufficient feature for reconstructing denoised images. Eventually, the limited training data may also results in over-fitting, leading to poor generalization in the denoising process. This article first proposes adaptive graph prior network (AGP-Net) using a novel graph construction method to capture the long-range dependencies on both the pixel and patch levels. Then, we propose graph supplementary prior and graph noise prior in AGP-Net to adaptively generate supplementary feature and regularization noise for improving the performance and generalization of image denoising. Extensive ablation and benchmark tests show our AGP-Net achieve the most advanced image denoising performance. Bo Jiang 0017, Yao Lu 0008, Bob Zhang 0001, Guangming Lu 0002 |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Joint Adaptive Robust Steganography NetworkabstractTransmission distortions within steganography systems easily cause dramatic degradations of revealing and invisibility performances. Previous works lacked sufficient adaptation for different distortions, which hinders the performance improvement of robust image steganography. This article proposes joint adaptive robust steganography network (JARS-Net). Specifically, the hierarchical attentive invertible (HAI) mechanism is first proposed to achieve adaptive feature tuning by gradually adjusting and fusing the cover-secret information from different depths and scales. Moreover, adaptive key learning (AKL) is proposed as an adaptive steganography strategy to generate adaptive keys for secret recovery under different distortions. Furthermore, benefiting from the joint of reversible HAI and the soft AKL, revealed secret images can be progressively decoupled from the received stego images along the backward HAI flow. Extensive experiments demonstrate that the proposed JARS-Net can significantly promote the invisibility and revealing performances of covert communication under different distortions. Le Zhang 0016, Yao Lu 0008, Guangming Lu 0002 |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Learning Content-Weighted Pseudocylindrical Representation for 360° Image CompressionabstractLearned 360° image compression methods using equirectangular projection (ERP) often confront a non-uniform sampling issue, inherent to sphere-to-rectangle projection. While uniformly or nearly uniformly sampling representations, along with their corresponding convolution operations, have been proposed to mitigate this issue, these methods often concentrate solely on uniform sampling rates, thus neglecting the content of the image. In this paper, we urge that different contents within 360° images have varying significance and advocate for the adoption of a content-adaptive parametric representation in 360° image compression, which takes into account both the content and sampling rate. We first introduce the parametric pseudocylindrical representation and corresponding convolution operation, upon which we build a learned 360° image codec. Then, we model the hyperparameter of the representation as the output of a network, derived from the image's content and its spherical coordinates. We treat the optimization of hyperparameters for different 360° images as distinct compression tasks and propose a meta-learning algorithm to jointly optimize the codec and the metaknowledge, i.e., the hyperparameter estimation network. A significant challenge is the lack of a direct derivative from the compression loss to the hyperparameter network. To address this, we present a novel method to relax the rate-distortion loss as a function of the hyperparameters, enabling gradient-based optimization of the metaknowledge. Experimental results on omnidirectional images demonstrate that our method achieves state-of-the-art performance and superior visual quality. Mu Li 0005, Youneng Bao, Xiaohang Sui, Jinxing Li 0003, Guangming Lu 0002, Yong Xu 0001 |
IEEE Trans. Image Process. | 5 |
| 2024 | Discrepancy and Structure-Based Contrast for Test-Time Adaptive RetrievalabstractDomain adaptive hashing has received increasing attention since it is capable of enhancing the performance of retrieval if the target domain for testing meets domain shift. However, owing to data security and transmission constraints nowadays, abundant source data is often not available. Towards this end, this paper investigates a novel yet practical problem named test-time adaptive hashing, which aims to enhance the performance of hashing models without access to the source domain data when tested on the target domain with domain shift. This problem is challenging due to both fugacious domain shift and label scarcity on the target domain. In this paper, we propose a novel hashing approach namedDiscrepancy andStructure-basedContrast (DISC) for effective test-time adaptive retrieval. In particular, DISC first trains the hashing model using the source domain data and stores the distribution of each class in the hidden space. During test-time adaptation, we generate simulated source features based on stored distributions and compare class-specific distributions across domains using maximum mean discrepancy (MMD) to overcome potential domain shift. Furthermore, to tackle the label scarcity, we estimate the graph structure using deep features on the target domain, which guides effective hashing contrastive learning for generating discriminative and domain-invariant hash codes. Extensive experiments on various benchmark datasets validate the superiority of our proposed DISC compared with a range of competing baselines. Zeyu Ma 0001, Yizhi Luo, Xiao Luo 0001, Jinxing Li 0003, Chong Chen 0002, Xian-Sheng Hua 0001, Guangming Lu 0002 |
IEEE Trans. Multim. | 8 |
| 2024 | HARR: Learning Discriminative and High-Quality Hash Codes for Image RetrievalabstractThis article studies deep unsupervised hashing, which has attracted increasing attention in large-scale image retrieval. The majority of recent approaches usually reconstruct semantic similarity information, which then guides the hash code learning. However, they still fail to achieve satisfactory performance in reality for two reasons. On the one hand, without accurate supervised information, these methods usually fail to produce independent and robust hash codes with semantics information well preserved, which may hinder effective image retrieval. On the other hand, due to discrete constraints, how to effectively optimize the hashing network in an end-to-end manner with small quantization errors remains a problem. To address these difficulties, we propose a novel unsupervised hashing method called HARR to learn discriminative and high-quality hash codes. To comprehensively explore semantic similarity structure, HARR adopts the Winner-Take-All hash to model the similarity structure. Then similarity-preserving hash codes are learned under the reliable guidance of the reconstructed similarity structure. Additionally, we improve the quality of hash codes by a bit correlation reduction module, which forces the cross-correlation matrix between a batch of hash codes under different augmentations to approach the identity matrix. In this way, the generated hash bits are expected to be invariant to disturbances with minimal redundancy, which can be further interpreted as an instantiation of the information bottleneck principle. Finally, for effective hashing network training, we minimize the cosine distances between real-value network outputs and their binary codes for small quantization errors. Extensive experiments demonstrate the effectiveness of our proposed HARR. Zeyu Ma 0001, Siwei Wang 0010, Xiao Luo 0001, Zhonghui Gu, Chong Chen 0002, Jinxing Li 0003, Xian-Sheng Hua 0001, Guangming Lu 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 8 |
| 2023 | Hierarchical Contrastive Learning for Pattern-Generalizable Image Corruption DetectionabstractEffective image restoration with large-size corruptions, such as blind image inpainting, entails precise detection of corruption region masks which remains extremely challenging due to diverse shapes and patterns of corruptions. In this work, we present a novel method for automatic corruption detection, which allows for blind corruption restoration without known corruption masks. Specifically, we develop a hierarchical contrastive learning framework to detect corrupted regions by capturing the intrinsic semantic distinctions between corrupted and uncorrupted regions. In particular, our model detects the corrupted mask in a coarse-to-fine manner by first predicting a coarse mask by contrastive learning in low-resolution feature space and then refines the uncertain area of the mask by high-resolution contrastive learning. A specialized hierarchical interaction mechanism is designed to facilitate the knowledge propagation of contrastive learning in different scales, boosting the modeling performance substantially. The detected multi-scale corruption masks are then leveraged to guide the corruption restoration. Detecting corrupted regions by learning the contrastive distinctions rather than the semantic patterns of corruptions, our model has well generalization ability across different corruption patterns. Extensive experiments demonstrate following merits of our model: 1) the superior performance over other methods on both corruption detection and various image restoration tasks including blind inpainting and watermark removal, and 2) strong generalization across different corruption patterns such as graffiti, random noise or other image content. Codes and trained weights are available at https://github.com/xyfJASON/HCL. Xin Feng 0005, Guangming Lu 0002, Wenjie Pei |
ICCV | 3 |
| 2023 | Differential Enhanced Siamese Segmentation Network for Printed Label Defect DetectionabstractMany vision-based methods have been widely used to detect defects in industrial printed labels. However, most of them still face challenges of detecting unseen defects, low-contrast defects, and false detections caused by artifacts. To address these problems, we propose a differential enhanced Siamese segmentation network (DESS-Net) for defect detection. This method is based on Siamese similarity comparison which has a better generalization ability for unseen defects. Moreover, we introduce the differential feature enhancement (DFE) modules into the Siamese network to focus on multiple differential feature information which contributes to identifying defects and reducing false detections caused by artifacts. Additionally, a multi-scale feature fusion (MFF) module is further designed to fuse multiple low-level differential features, which is conducive to recovering fine boundaries of low-contrast defects. Experimental results show that our DESS-Net outperforms other compared methods. Dongming Li 0002, Yingjian Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
ICIP | 4 |
| 2023 | Combating Medical Label Noise via Robust Semi-supervised Contrastive Learning
Bingzhi Chen, Zhanhao Ye, Yishu Liu 0001, Zheng Zhang 0006, Jiahui Pan 0003, Guangming Lu 0002 |
MICCAI (1) | 7 |
| 2023 | Multi-Granularity Interactive Transformer Hashing for Cross-modal RetrievalabstractWith the powerful representation ability and privileged efficiency, deep cross-modal hashing (DCMH) has become an emerging fast similarity search technique. Prior studies primarily focus on exploring pairwise similarities across modalities, but fail to comprehensively capture the multi-grained semantic correlations during intra- and inter-modal negotiation. To tackle this issue, this paper proposes a novel Multi-granularity Interactive Transformer Hashing (MITH) network, which hierarchically considers both coarse- and fine-grained similarity measurements across different modalities in one unified transformer-based framework. To the best of our knowledge, this is the first attempt for multi-granularity transformer-based cross-modal hashing. Specifically, a well-designed distilled intra-modal interaction module is deployed to excavate modality-specific concept knowledge with global-local knowledge distillation under the guidance of implicit conceptual category-level representations. Moreover, we construct a contrastive inter-modal alignment module to mine modality-independent semantic concept correspondences with instance- and token-wise contrastive learning, respectively. Such a collaborative learning paradigm can jointly alleviate the heterogeneity and semantic gaps among different modalities from a multi-granularity perspective, yielding discriminative modality-invariant hash codes. Extensive experiments on multiple representative cross-modal datasets demonstrate the consistent superiority of MITH over the existing state-of-the-art baselines. The codes are available at https://github.com/DarrenZZhang/MITH. Yishu Liu 0001, Qingpeng Wu, Zheng Zhang 0006, Guangming Lu 0002 |
ACM Multimedia | 5 |
| 2023 | Boosting Few-shot 3D Point Cloud Segmentation via Query-Guided EnhancementabstractAlthough extensive research has been conducted on 3D point cloud segmentation, effectively adapting generic models to novel categories remains a formidable challenge. This paper proposes a novel approach to improve point cloud few-shot segmentation (PC-FSS) models. Unlike existing PC-FSS methods that directly utilize categorical information from support prototypes to recognize novel classes in query samples, our method identifies two critical aspects that substantially enhance model performance by reducing contextual gaps between support prototypes and query features. Specifically, we (1) adapt support background prototypes to match query context while removing extraneous cues that may obscure foreground and background in query samples, and (2) holistically rectify support prototypes under the guidance of query features to emulate the latter having no semantic gap to the query targets. Our proposed designs are agnostic to the feature extractor, rendering them readily applicable to any prototype-based methods. The experimental results on S3DIS and ScanNet demonstrate notable practical benefits, as our approach achieves significant improvements while still maintaining high efficiency. The code for our approach is available at https://github.com/AaronNZH/Boosting-Few-shot-3D-Point-Cloud-Segmentation-via-Query-Guided-Enhancement Zhenhua Ning, Zhuotao Tian, Guangming Lu 0002, Wenjie Pei |
ACM Multimedia | 3 |
| 2023 | Scene-Generalizable Interactive Segmentation of Radiance FieldsabstractExisting methods for interactive segmentation in radiance fields entail scene-specific optimization and thus cannot generalize across different scenes, which greatly limits their applicability. In this work we make the first attempt at Scene-Generalizable Interactive Segmentation in Radiance Fields (SGISRF) and propose a novel SGISRF method, which can perform 3D object segmentation for novel (unseen) scenes represented by radiance fields, guided by only a few interactive user clicks in a given set of multi-view 2D images. In particular, the proposed SGISRF focuses on addressing three crucial challenges with three specially designed techniques. First, we devise the Cross-Dimension Guidance Propagation to encode the scarce 2D user clicks into informative 3D guidance representations. Second, the Uncertainty-Eliminated 3D Segmentation module is designed to achieve efficient yet effective 3D segmentation. Third, Concealment-Revealed Supervised Learning scheme is proposed to reveal and correct the concealed 3D segmentation errors resulted from the supervision in 2D space with only 2D mask annotations. Extensive experiments on two real-world challenging benchmarks covering diverse scenes demonstrate 1) effectiveness and scene-generalizability of the proposed method, 2) favorable performance compared to classical method requiring scene-specific optimization. Songlin Tang, Wenjie Pei, Xin Tao 0001, Tanghui Jia, Guangming Lu 0002, Yu-Wing Tai |
ACM Multimedia | 5 |
| 2023 | Multimodal Emotion Interaction and Visualization PlatformabstractIn this paper, we present a multimodal emotion analysis platform, which can flexibly capture, detect and analyze the emotions of video object with multiple modalities under different situations, including offline and online application scenarios. This system can visualize the dynamic effects of different types of emotions from both multimodal and unimodal circumstances. The presented emotion analysis results show instant and time series states in both specific modality and multiple modalities. Our system fills the current research and application gaps in multimodal emotion analysis with an interactive interface. Notably, the constructed system can adaptively process pre-recorded video clips as well as collected real-world data with excellent practicality and interactivity. Zheng Zhang 0006, Songling Chen, Mi-Xiao Hou, Guangming Lu 0002 |
ACM Multimedia | 4 |
| 2023 | Video-based Visible-Infrared Person Re-Identification via Style Disturbance Defense and Dual InteractionabstractVideo-based visible-infrared person re-identification (VVI-ReID) aims to retrieve video sequences of the same pedestrian from different modalities. The key of VVI-ReID is to learn discriminative sequence-level representations that are invariant to both intra- and inter-modal discrepancies. However, most works only focus on the elimination of modality-gap while ignore the distractors within the modality. Moreover, existing sequence-level representation learning approaches are limited to a single video, failing to mine the correlations among multiple videos of the same pedestrian. In this paper, we propose a Style Augmentation, Attack and Defense network with Graph-based dual interaction (SAADG) to guarantee the semantic consistency against both intra-modal discrepancies and inter-modal gap. Specifically, we first generate diverse styles for video frames by random style variation in image spaces. Followed by the style attack and defense, the intra- and inter-modal discrepancies are modeled as different types of style disturbance (attack), and our model achieves to keep the id-related content invariant under such attack. Besides, a graph-based dual interaction module is further introduced to fully explore the cross-view and cross-modal correlations among various videos of the same identity, which are then transferred to the sequence-level representations. Extensive experiments on the public SYSU-MM01 and HITSZ-VCM datasets show that our approach achieves the remarkable performance compared with state-of-the-arts. The code is available at https://github.com/ChuhaoZhou99/SAADG_VVIReID. Chuhao Zhou, Jinxing Li 0003, Huafeng Li 0001, Guangming Lu 0002, Yong Xu 0001, Min Zhang 0005 |
ACM Multimedia | 4 |
| 2023 | Lightweight image denoising network with four-channel interaction transform
Yao Lu 0008, Guangming Lu 0002 |
Image Vis. Comput. | 3 |
| 2023 | Deep adaptive hiding network for image hiding using attentive frequency extraction and gradual depth extraction
Le Zhang 0016, Yao Lu 0008, Jinxing Li 0003, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001 |
Neural Comput. Appl. | 5 |
| 2023 | Adapt-Infomap: Face clustering with adaptive graph refinement in infomapabstractFace clustering is a critical task in computer vision due to the increasing number of applications such as augmented reality or photo album management. The primary challenge in this task arises from the imperfections in image feature representations. Given image features extracted from an existing pre-trained representation model, it remains an unresolved problem that how to leverage the inherent characteristics of similarities among unlabelled images to improve the clustering performance. In order to solve face clustering in an unsupervised manner , we develop an effective and robust framework named as Adapt-Infomap. First, we reformulate face clustering as a process of non-overlapping community detection. Specially, Adapt-Infomap achieves face clustering by minimizing the entropy of information flows (also known as the map equation) on an affinity graph of images. Since the affinity graph of images might contain noisy edges, we develop an outlier detection strategy in Adapt-Infomap to adaptively refine the affinity graph. Experiments with ablation studies demonstrate that Adapt-Infomap significantly outperforms existing methods and achieves new state-of-the-arts on three popular large-scale datasets for face clustering, e.g. , an absolute improvement of more than 10 % and 3 % comparing with prior unsupervised and supervised methods respectively in terms of average of Pairwise F-score. Xiaotian Yu, Aibo Wang, Haokui Zhang, Hanling Yi, Guangming Lu 0002, Xiaoyu Wang 0002 |
Pattern Recognit. | 7 |
| 2023 | Deep collaborative graph hashing for discriminative image retrieval
Zheng Zhang 0006, Jianning Wang, Lei Zhu 0002, Yadan Luo, Guangming Lu 0002 |
Pattern Recognit. | 5 |
| 2023 | Joint adjustment image steganography networks
Le Zhang 0016, Yao Lu 0008, Guangming Lu 0002 |
Signal Process. Image Commun. | 4 |
| 2023 | Facial Expression Recognition in the Wild Using Multi-Level Features and Attention MechanismsabstractLearning discriminative features is of vital importance for automatic facial expression recognition (FER) in the wild. In this article, we propose a novel Slide-Patch and Whole-Face Attention model with SE blocks (SPWFA-SE), which jointly perceives the discriminative locality characteristics and informative global features of the face for effective FER. Specifically, the well-designed slide patches are proposed to extract local features. Different from the existing methods, our slide patches not only can maintain the information at the edge area of patches, but also do not need to detect facial landmarks. Moreover, to make the model adaptively focus on the distinguishable regions, an attention module is proposed in the patch level to learn the weight of each patch. Furthermore, squeeze-and-excitation blocks are explored in the channel level to learn the weight of each channel. As such, the proposed multi-level feature extraction and attention mechanisms can enhance the representative ability of the learned features. Extensive experiments on five challenging datasets demonstrate that our method can achieve state-of-the-art performance. Cross database experiments on another three databases show the superior generalization performance of our model. Furthermore, complexity analysis results show that our model contains fewer parameters with fast training advantages than other competing models. Yingjian Li 0001, Guangming Lu 0002, Jinxing Li 0003, Zheng Zhang 0006, David Zhang 0001 |
IEEE Trans. Affect. Comput. | 2 |
| 2023 | Emotion Prediction Oriented Method With Multiple Supervisions for Emotion-Cause Pair ExtractionabstractEmotion-cause pair extraction (ECPE) task aims to extract all the pairs of emotions and their causes from an unannotated emotion text. The previous works usually extract the emotion-cause pairs from two perspectives of emotion and cause. However, emotion extraction is more crucial to the ECPE task than cause extraction. Motivated by this analysis, we propose an end-to-end emotion-cause extraction approach oriented toward emotion prediction (EPO-ECPE), aiming to fully exploit the potential of emotion prediction to enhance emotion-cause pair extraction. Considering the strong dependence between emotion prediction and emotion-cause pair extraction, we propose a synchronization mechanism to share their improvement in the training process. That is, the improvement of emotion prediction can facilitate the emotion-cause pair extraction, and then the results of emotion-cause pair extraction can also be used to improve the accuracy of emotion prediction simultaneously. For the emotion-cause pair extraction, we divide it into genuine pair supervision and fake pair supervision, where the genuine pair supervision learns from the pairs with more possibility to be emotion-cause pairs. In contrast, fake pair supervision learns from other pairs. In this way, the emotion-cause pairs can be extracted directly from the genuine pair, thereby reducing the difficulty of extraction. Experimental results show that our approach outperforms the 13 compared systems and achieves new state-of-the-art performance. Guimin Hu, Yi Zhao 0007, Guangming Lu 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Semantic Alignment Network for Multi-Modal Emotion RecognitionabstractModality alignment can maintain the consistency of semantics in multi-modal emotion recognition tasks, ensuring that features from different modalities accurately represent the emotion-related information in an encoding space. However, current alignment models either focus only on the local fusion of different modal representations or lack a mining process for unimodal specificity information. We design a Semantic Alignment network based on Multi-Spatial learning (SAMS) for multi-modal emotion recognition, which achieves local and global alignment between modalities using high-level emotion representations of different modalities as supervisory signals. SAMS builds a multi-spatial learning framework for each modality, and constructs a self-modal interaction module under this framework based on cross-modal semantic learning. SAMS provides two learning spaces for each modality, one to detect the affective information for a specific modality, and the other to learn semantic knowledge from other modalities. Subsequently, the features of these two spaces are aligned in temporal and utterance levels by homologous encoding and different target constraints. Based on the alignment characteristics of these two spaces, a self-modal interaction is built to investigate the fusion representation by exploring the global correlation between the alignment features in unimodal multi-spatial learning. In experiments, our proposed model yields consistent improvements on two standard multi-modal benchmarks, and outperforms state-of-the-art approaches. The code of our SAMS is available at:https://github.com/xiaomi1024/code_SAMS. Mi-Xiao Hou, Zheng Zhang 0006, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Few-Shot Learning for Image DenoisingabstractDeep Neural Networks (DNNs) have achieved impressive results on the task of image denoising, but there are two serious problems. First, the denoising ability of DNNs-based image denoising models using traditional training strategies heavily relies on extensive training on clean-noise image pairs. Second, image denoising models based on DNNs usually have large parameters and high computational complexity. To address these issues, this paper proposes a two-stage Few-Shot Learning for Image Denoising (FSLID). Our FSLID is a two-stage denoising strategy integrating Basic Feature Learner (BFL), Denoising Feature Inducer (DFI), and Shared Image Reconstructor (SIR). BFL and SIR are first jointly unsupervised to train on the base image dataset$\mathcal {D}_{base}$consisting of easily collected high-quality clean images. Following this, the trained BFL extracts the guided features and constraint features for the noisy and corresponding clean images in the novel image dataset$\mathcal {D}_{novel}$, respectively. Furthermore, DFI encodes the noisy features of the noisy images in$\mathcal {D}_{novel}$. Then, inducing both the guided features and noisy features, DFI can generate the denoising prior features for the SIR with frozen weights to adaptively denoise the noisy images. Furthermore, we propose refined, low-channel-count, recursive multi-branch Multi-Scale Feature Recursive (MSFR) to modularly formulate an efficient DFI to capture more diverse contextual features information under a limited number of feature channels. Thus, compared with the baseline models, the FSLID composed of the proposed MSFR can significantly reduce the number of model parameters and computational complexity. Extensive experimental results demonstrate our FSLID significantly outperforms well-established baselines on multiple datasets and settings. We hope that our work will encourage further research to explore the field of few-shot image denoising. Bo Jiang 0017, Yao Lu 0008, Bob Zhang 0001, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Image-Text Retrieval With Cross-Modal Semantic Importance ConsistencyabstractCross-modal image-text retrieval is an important area of Vision-and-Language task that models the similarity of image-text pairs by embedding features into a shared space for alignment. To bridge the heterogeneous gap between the two modalities, current approaches achieve inter-modal alignment and intra-modal semantic relationship modeling through complex weighted combinations between items. In the intra-modal association and inter-modal interaction processes, the higher-weight items have a higher contribution to the global semantics. However, the same item always produces different contributions in the two processes, since most traditional approaches only focus on the alignment. This usually results in semantic changes and misalignment. To address this issue, this paper proposes Cross-modal Semantic Importance Consistency (CSIC) which achieves invariance in the semantic of items during aligning. The proposed technique measures the semantic importance of items obtained from intra-modal and inter-modal self-attention and learns a more reasonable representation vector by inter-calibrating the importance distribution to improve performance. We conducted extensive experiments on the Flickr30K and MS COCO datasets. The results show that our approach can significantly improve retrieval performance, proving the proposed approach’s superiority and rationality. Zejun Liu, Fanglin Chen 0001, Jun Xu 0008, Wenjie Pei, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Generalized Nonconvex Low-Rank Tensor Representation for Hyperspectral Anomaly DetectionabstractLow-rank tensor representation (LRTR) methods have attracted great interest for their powerful ability to separate backgrounds and anomalies. However, most of the current LRTR models use the popular and convex surrogate tensor nuclear norm to solve optimization problems, which results in a loose approximation and suboptimal solver for the original problem. Besides, most existing methods solve the nonconvex optimization problems case-by-case, consequently losing one unified solver. To solve the above issues, we propose the Generalized Nonconvex Low-rank Tensor Representation (GNLTR) for hyperspectral anomaly detection (HAD), a unified solver not case-by-case one of existing nonconvex optimization problems. Compared to the tensor nuclear norm, GNLTR contains many popular nonconvex penalty functions as tighter regularizers of the tensor tubal rank to constrain the low rank of the background. Moreover, theL2,1norm has been integrated into the GNLTR model for the sparse anomalies. For the optimization problem, it is handled quickly and efficiently through a well-organized alternating direction method of multipliers (ADMM). The experiments on several real-world hyperspectral data sets demonstrate the superior performance of the GNLTR model in comparison with some state-of-the-art anomaly detection models. Qiangqiang Shen, Haijin Zeng, Yongyong Chen, Guangming Lu 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Cross-Domain Facial Expression Recognition via Contrastive Warm up and Complexity-Aware Self-TrainingabstractUnsupervised cross-domain Facial Expression Recognition (FER) aims to transfer the knowledge from a labeled source domain to an unlabeled target domain. Existing methods strive to reduce the discrepancy between source and target domain, but cannot effectively explore the abundant semantic information of the target domain due to the absence of target labels. To this end, we propose a novel framework via Contrastive Warm up and Complexity-aware Self-Training (namely CWCST), which facilitates source knowledge transfer and target semantic learning jointly. Specifically, we formulate a contrastive warm up strategy via features, momentum features, and learnable category centers to concurrently learn discriminative representations and narrow the domain gap, which benefits domain adaptation by generating more accurate target pseudo labels. Moreover, to deal with the inevitable noise in pseudo labels, we develop complexity-aware self-training with a label selection module based on prediction entropy, which iteratively generates pseudo labels and adaptively chooses the reliable ones for training, ultimately yielding effective target semantics exploration. Furthermore, by jointly using the two mentioned components, our framework enables to effectively utilize the source knowledge and target semantic information by source-target co- training. In addition, our framework can be easily incorporated into other baselines with consistent performance improvements. Extensive experimental results on seven databases show the superior performance of the proposed method against various baselines. Yingjian Li 0001, Jiaxing Huang 0001, Shijian Lu, Zheng Zhang 0006, Guangming Lu 0002 |
IEEE Trans. Image Process. | 5 |
| 2023 | From Global to Local: Multi-Patch and Multi-Scale Contrastive Similarity Learning for Unsupervised Defocus Blur DetectionabstractDefocus blur detection (DBD), which aims to detect out-of-focus or in-focus pixels from a single image, has been widely applied to many vision tasks. To remove the limitation on the abundant pixel-level manual annotations, unsupervised DBD has attracted much attention in recent years. In this paper, a novel deep network named Multi-patch and Multi-scale Contrastive Similarity (M2CS) learning is proposed for unsupervised DBD. Specifically, the predicted DBD mask from a generator is first exploited to re-generate two composite images by transporting the estimated clear and unclear areas from the source image to realistic full-clear and full-blurred images, respectively. To encourage these two composite images to be completely in-focus or out-of-focus, a global similarity discriminator is exploited to measure the similarity of each pair in a contrastive way, through which each two positive samples (two clear images or two blurred images) are enforced to be close while each two negative samples (a clear image and a blurred image) are inversely far. Since the global similarity discriminator only focuses on the blur-level of a whole image and there do exist some fail-detected pixels which only cover a small part of areas, a set of local similarity discriminators are further designed to measure the similarity of image patches in multiple scales. Thanks to this joint global and local strategy, as well as the contrastive similarity learning, the two composite images are more efficiently moved to be all-clear or all-blurred. Experimental results on real-world datasets substantiate the superiority of our proposed method both in quantification and visualization. The source code is released at: https://github.com/jerysaw/M2CS. Jinxing Li 0003, Beicheng Liang, Xiangwei Lu, Mu Li 0005, Guangming Lu 0002, Yong Xu 0001 |
IEEE Trans. Image Process. | 5 |
| 2023 | Pedestrian Detection by Exemplar-Guided Contrastive LearningabstractTypical methods for pedestrian detection focus on either tackling mutual occlusions between crowded pedestrians, or dealing with the various scales of pedestrians. Detecting pedestrians with substantial appearance diversities such as different pedestrian silhouettes, different viewpoints or different dressing, remains a crucial challenge. Instead of learning each of these diverse pedestrian appearance features individually as most existing methods do, we propose to perform contrastive learning to guide the feature learning in such a way that the semantic distance between pedestrians with different appearances in the learned feature space is minimized to eliminate the appearance diversities, whilst the distance between pedestrians and background is maximized. To facilitate the efficiency and effectiveness of contrastive learning, we construct an exemplar dictionary with representative pedestrian appearances as prior knowledge to construct effective contrastive training pairs and thus guide contrastive learning. Besides, the constructed exemplar dictionary is further leveraged to evaluate the quality of pedestrian proposals during inference by measuring the semantic distance between the proposal and the exemplar dictionary. Extensive experiments on both daytime and nighttime pedestrian detection validate the effectiveness of the proposed method. Zebin Lin, Wenjie Pei, Fanglin Chen 0001, David Zhang 0001, Guangming Lu 0002 |
IEEE Trans. Image Process. | 5 |
| 2023 | Modality-Invariant Asymmetric Networks for Cross-Modal HashingabstractCross-modal hashing has garnered considerable attention and gained great success in many cross-media similarity search applications due to its prominent computational efficiency and low storage overhead. However, it still remains challenging how to effectively take multilevel advantages of semantics on the entire database to jointly bridge the semantic and heterogeneity gaps across different modalities. In this paper, we propose a novel Modality-Invariant Asymmetric Networks (MIAN) architecture, which explores the asymmetric intra- and inter-modal similarity preservation under a probabilistic modality alignment framework. Specifically, an intra-modal asymmetric network is conceived to capture the query-vs-all internal pairwise similarities for each modality in a probabilistic asymmetric learning manner. Moreover, an inter-modal asymmetric network is deployed to fully harness the cross-modal semantic similarities supported by the maximum inner product search formula between two distinct hash embeddings. Particularly, the pairwise, piecewise and transformed semantics are jointly considered into one unified semantic-preserving hash codes learning scheme. Furthermore, we construct a modality alignment network to distill the redundancy-free visual features and maximize the conditional bottleneck information between different modalities. Such a network could close the heterogeneity and domain shift across different modalities. Extensive experiments evidence that our MIAN approach can outperform the state-of-the-art cross-modal hashing methods. Zheng Zhang 0006, Haoyang Luo, Lei Zhu 0002, Guangming Lu 0002, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Deep Margin-Sensitive Representation Learning for Cross-Domain Facial Expression RecognitionabstractCross-domain Facial Expression Recognition (FER) aims to safely transfer the learned knowledge from labeled source data to unlabeled target data, which is challenging due to the subtle difference between various expressions and the large discrepancy between domains. Existing methods mainly focus on reducing the domain shift for transferable features but fail to learn discriminative representations for recognizing facial expression, which may result in negative transfer under cross-domain settings. To this end, we propose a novel Deep Margin-Sensitive Representation Learning (DMSRL) framework, which can extract multi-level discriminative features during sematic-aware domain adaptation. Specifically, we design a semantic metric learning module based on the category prior of source data and generated pseudo labels of target data, which can facilitate discriminative intra-domain representation learning and transferable inter-domain knowledge discovery by enlarging the category margin. Moreover, we develop a mutual information minimization module by simultaneously distilling the domain-invariant components and eliminating the domain-sensitive ones, which benefits discriminative transferable feature learning by generating accurate pseudo target labels. Furthermore, instead of only utilizing the global features, we formulate a multi-level feature extracting module to concurrently get the local ones, which contain detailed information to distinguish the small changes among different expressions. These modules are jointly utilized in our DMSRL in an end-to-end manner to ensure the positive transfer of source knowledge. Extensive experimental results on seven databases demonstrate that our DMSRL can achieve superior performance against state-of-the-art baselines. Yingjian Li 0001, Zheng Zhang 0006, Bingzhi Chen, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Graph Attention in Attention Network for Image DenoisingabstractImage denoising aims to remove the noise from noisy images. With the increasing complexity of the noise within the noisy images, current denoising methods cannot satisfactorily address this issue. This article proposes a graph attention in attention network (GAiA-Net) for image denoising. First, we introduce a novel approach to graph construction for the GAiA-Net. In the process of such graph construction, the noisy images are divided into patches to formulate the nodes in a graph. The edges are initialized using$k $-nearest neighbors. Hence, through iterative transformation and learning, both the pixel-level and structure-level features can be captured by different information exchanges and aggregation within (pixel-level) and outside (structure-level) of the nodes, respectively. Second, we propose the graph attention in attention (GAiA) in the GAiA-Net. The proposed GAiA produces the pixel-level attention within nodes to be further induced to the nodes with various distances to generate the final attention. Therefore, our GAiA-Net can capture the long dependencies on both the pixel-level and structure-level features, which can effectively reduce the complex noise in the denoising process. Comprehensive experiments demonstrate that the proposed GAiA-Net produces state-of-the-art performances on both synthetic noise image and real noise image datasets. Especially, when experimenting on complex noisy Nam datasets, our GAiA-Net achieves a PSNR of 40.40 dB and SSIM of 0.989. These results prove the satisfactory potential and effectiveness of our GAiA-Net. Bo Jiang 0017, Yao Lu 0008, Xiaosheng Chen, Xinhai Lu, Guangming Lu 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2022 | PPR-Net: Patch-Based Multi-scale Pyramid Registration Network for Defect Detection of Printed Label
Dongming Li 0002, Yingjian Li 0001, Jinxing Li 0003, Guangming Lu 0002 |
ACCV (2) | 4 |
| 2022 | Learning Modal-Invariant and Temporal-Memory for Video-based Visible-Infrared Person Re-IdentificationabstractThanks for the cross-modal retrieval techniques, visible-infrared (RGB-IR) person re-identification (Re-ID) is achieved by projecting them into a common space, allowing person Re-ID in 24-hour surveillance systems. However, with respect to the probe-to- gallery, almost all existing RGB-IR based cross-modal person Re-ID methods focus on image-to-image matching, while the video-to-video matching which contains much richer spatial- and temporal-information remains under-explored. In this paper, we primarily study the video-based cross-modal per-son Re-ID method. To achieve this task, a video-based RGB-IR dataset is constructed, in which 927 valid identities with 463,259 frames and 21,863 tracklets captured by 12 RGB/IR cameras are collected. Based on our constructed dataset, we prove that with the increase of frames in a tracklet, the performance does meet more enhancement, demonstrating the significance of video-to-video matching in RGB-IR person Re-ID. Additionally, a novel method is further proposed, which not only projects two modalities to a modal-invariant subspace, but also extracts the temporal-memory for motion-invariant. Thanks to these two strategies, much better results are achieved on our video-based cross-modal person Re-ID. The code and dataset are released at: https://github.com/VCM-project233/MITML. Jinxing Li 0003, Zeyu Ma 0001, Huafeng Li 0001, Kaixiong Xu, Guangming Lu 0002, David Zhang 0001 |
CVPR | 7 |
| 2022 | Few-Shot Object Detection by Knowledge Distillation Using Bag-of-Visual-Words Representations
Wenjie Pei, Dianwen Mei, Fanglin Chen 0001, Jiandong Tian, Guangming Lu 0002 |
ECCV (10) | 6 |
| 2022 | Multi-faceted Distillation of Base-Novel Commonality for Few-Shot Object Detection
Wenjie Pei, Dianwen Mei, Fanglin Chen 0001, Jiandong Tian, Guangming Lu 0002 |
ECCV (9) | 6 |
| 2022 | UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion RecognitionabstractMultimodal sentiment analysis (MSA) and emotion recognition in conversation (ERC) are key research topics for computers to understand human behaviors.From a psychological perspective, emotions are the expression of affect or feelings during a short period, while sentiments are formed and held for a longer period.However, most existing works study sentiment and emotion separately and do not fully exploit the complementary knowledge behind the two.In this paper, we propose a multimodal sentiment knowledge-sharing framework (UniMSE) that unifies MSA and ERC tasks from features, labels, and models.We perform modality fusion at the syntactic and semantic levels and introduce contrastive learning between modalities and samples to better capture the difference and consistency between sentiments and emotions.Experiments on four public benchmark datasets, MOSI, MOSEI, MELD, and IEMO-CAP, demonstrate the effectiveness of the proposed method and achieve consistent improvements compared with state-of-the-art methods. Guimin Hu, Ting-En Lin, Yi Zhao 0007, Guangming Lu 0002, Yuchuan Wu |
EMNLP | 4 |
| 2022 | Correlation-Based Transformer Tracking
Minghan Zhong, Fanglin Chen 0001, Jun Xu 0008, Guangming Lu 0002 |
ICANN (1) | 4 |
| 2022 | Multi-Modal Emotion Recognition with Self-Guided Modality CalibrationabstractMulti-modal emotion recognition aims to extract sentiment-related information from multiple sources and integrate different modal representations for sentiment analysis. Alignment is an effective strategy to achieve semantically consistent representations for multi-modal emotion recognition, while the current alignment models are jointly unable to maintain the dependence of word-to-sentence and independence of unimodal learning. In this paper, we propose a Self-guided Modality Calibration Network (SMCN) to realize multi-modal alignment which can capture the global connections without interfering with unimodal learning. While preserving unimodal learning without interference, our model leverages semantic sentiment-related features to guide modality-specific representation learning. On one hand, SMCN simulates human thinking by deriving a branch for acquiring knowledge of other modalities in unimodal learning. This branch aims to lean high-level semantic information of other modalities for realizing semantic alignment between modalities. On the other hand, we also provide an indirect interaction manner to integrate unimodal feature and calibrate features in different levels for avoiding unimodal features mixed with other clues. Experiments demonstrate that our approach outperforms the state-of-the-art methods on both IEMOCAP and MELD datasets. Mi-Xiao Hou, Zheng Zhang 0006, Guangming Lu 0002 |
ICASSP | 3 |
| 2022 | DHWP: Learning High-Quality Short Hash Codes Via Weight PruningabstractHashing is widely used in large-scale image retrieval because of its efficiency in storage and computation. Although longer hash codes can lead to higher search accuracy, the retrieval cost increases linearly with the increase of the number of hash bits. Most deep hashing methods suffer from the problem of trivial solutions and usually result in highly correlated redundant hash bits in practice, which limits the performance. To obtain short hash codes with high quality for fast and accurate image retrieval, we propose a novel framework named Deep Hashing via Weight Pruning (DHWP). DHWP first trains the model with relatively long hash codes. Then it obtains shorter codes gradually by weight pruning based on four different criteria. The framework of DHWP can be applied to most deep supervised hashing models, which helps remove those redundant hash bits while retaining the representation ability of long hash codes. Extensive experimental results on two widely used benchmark datasets show that DHWP outperforms the existing state-of-the-art methods, especially for short hash codes. Zeyu Ma 0001, Yuhang Guo 0002, Xiao Luo 0001, Chong Chen 0002, Minghua Deng, Wei Cheng 0002, Guangming Lu 0002 |
ICASSP | 7 |
| 2022 | Pruning Based Training-Free Neural Architecture SearchabstractNeural Architecture Search (NAS) plays an important role in searching for high-performance neural networks. How-ever, NAS algorithms are slow and require a terrific amount of computing resources, because they need to be trained on supernet or dense candidate networks to obtain information for evaluation. If the high-performance network architecture could be selected without training, it would eliminate a signif-icant part of the computational cost. Therefore, we propose a zero-cost metric called EX-score, which can represent the ex-pressivity of the network and rank the untrained architectures. To further reduce cost, we design a pruning based zero-cost neural architecture search framework (PZ-NAS) using EX-score. PZ-NAS can prune the initialised supernet rapidly and obtains hundreds of times faster speed performance, whilst archieving comparable accuracy property on CIFAR-IO and ImageNet. Jiawang Zhou, Fanglin Chen 0001, Guangming Lu 0002 |
ICME | 3 |
| 2022 | Improved Deep Unsupervised Hashing with Fine-grained Semantic Similarity Mining for Multi-Label Image RetrievalabstractIn this paper, we study deep unsupervised hashing, a critical problem for approximate nearest neighbor research. Most recent methods solve this problem by semantic similarity reconstruction for guiding hashing network learning or contrastive learning of hash codes. However, in multi-label scenarios, these methods usually either generate an inaccurate similarity matrix without reflection of similarity ranking or suffer from the violation of the underlying assumption in contrastive learning, resulting in limited retrieval performance. To tackle this issue, we propose a novel method termed HAMAN, which explores semantics from a fine-grained view to enhance the ability of multi-label image retrieval. In particular, we reconstruct the pairwise similarity structure by matching fine-grained patch features generated by the pre-trained neural network, serving as reliable guidance for similarity preserving of hash codes. Moreover, a novel conditional contrastive learning on hash codes is proposed to adopt self-supervised learning in multi-label scenarios. According to extensive experiments on three multi-label datasets, the proposed method outperforms a broad range of state-of-the-art methods. Zeyu Ma 0001, Xiao Luo 0001, Yingjie Chen 0002, Mi-Xiao Hou, Jinxing Li 0003, Minghua Deng, Guangming Lu 0002 |
IJCAI | 7 |
| 2022 | Domain Adaptive Nuclei Instance Segmentation and Classification via Category-Aware Feature Alignment and Pseudo-Labelling
Canran Li, Dongnan Liu, Haoran Li 0024, Zheng Zhang 0006, Guangming Lu 0002, Xiaojun Chang, Tom Weidong Cai |
MICCAI (8) | 5 |
| 2022 | ConTrans: Improving Transformer with Convolutional Attention for Medical Image Segmentation
Ailiang Lin, Jiayu Xu 0002, Jinxing Li 0003, Guangming Lu 0002 |
MICCAI (5) | 4 |
| 2022 | Learning Generalizable Latent Representations for Novel Degradations in Super-ResolutionabstractTypical methods for blind image super-resolution (SR) focus on dealing with unknown degradations by directly estimating them or learning the degradation representations in a latent space. A potential limitation of these methods is that they assume the unknown degradations can be simulated by the integration of various handcrafted degradations (e.g., bicubic downsampling), which is not necessarily true. The real-world degradations can be beyond the simulation scope by the handcrafted degradations, which are referred to as novel degradations. In this work, we propose to learn a latent representation space for degradations, which can be generalized from handcrafted (base) degradations to novel degradations. Furthermore, we perform variational inference to match the posterior of degradations in latent representation space with a prior distribution (e.g., Gaussian distribution). Consequently, we are able to sample more high-quality representations for a novel degradation to augment the training data for SR model. We conduct extensive experiments on both synthetic and real-world datasets to validate the effectiveness and advantages of our method for blind super-resolution with novel degradations. Fengjun Li, Xin Feng 0005, Fanglin Chen 0001, Guangming Lu 0002, Wenjie Pei |
ACM Multimedia | 4 |
| 2022 | Improved Deep Unsupervised Hashing via Prototypical LearningabstractHashing has become increasingly popular in approximate nearest neighbor search in recent years due to its storage and computational efficiency. While deep unsupervised hashing has shown encouraging performance recently, its efficacy in the more realistic unsupervised situation is far from satisfactory due to two limitations. On one hand, they usually neglect the underlying global semantic structure in the deep feature space. On the other hand, they also ignore reconstructing the global structure in the hash code space. In this research, we develop a simple yet effective approach named deeP U nsupeR vised hashing via P rototypical LEarning.. Specifically, introduces both feature prototypes and hashing prototypes to model the underlying semantic structures of the images in both deep feature space and hash code space. Then we impose a smoothness constraint to regularize the consistency of the global structures in two spaces through our semantic prototypical consistency learning. Moreover, our method encourages the prototypical consistency for different augmentations of each image via contrastive prototypical consistency learning. Comprehensive experiments on three benchmark datasets demonstrate that our proposed performs better than a variety of state-of-the-art retrieval methods. Zeyu Ma 0001, Wei Ju 0001, Xiao Luo 0001, Chong Chen 0002, Xian-Sheng Hua 0001, Guangming Lu 0002 |
ACM Multimedia | 6 |
| 2022 | Decoupling Recognition from Detection: Single Shot Self-Reliant Scene Text SpotterabstractTypical text spotters follow the two-stage spotting strategy: detect the precise boundary for a text instance first and then perform text recognition within the located text region. While such strategy has achieved substantial progress, there are two underlying limitations. 1) The performance of text recognition depends heavily on the precision of text detection, resulting in the potential error propagation from detection to recognition. 2) The RoI cropping which bridges the detection and recognition brings noise from background and leads to information loss when pooling or interpolating from feature maps. In this work we propose the single shot Self-Reliant Scene Text Spotter (SRSTS), which circumvents these limitations by decoupling recognition from detection. Specifically, we conduct text detection and recognition in parallel and bridge them by the shared positive anchor point. Consequently, our method is able to recognize the text instances correctly even though the precise text boundaries are challenging to detect. Additionally, our method reduces the annotation cost for text detection substantially. Extensive experiments on regular-shaped benchmark and arbitrary-shaped benchmark demonstrate that our SRSTS compares favorably to previous state-of-the-art spotters in terms of both accuracy and efficiency. Pengyuan Lv, Guangming Lu 0002, Chengquan Zhang, Wenjie Pei |
ACM Multimedia | 3 |
| 2022 | AIA: Attention in Attention Within Collaborate Domains
Le Zhang 0016, Yao Lu 0008, Guangming Lu 0002 |
PRCV (1) | 5 |
| 2022 | Multi-modal Finger Feature Fusion Algorithms on Large-Scale Dataset
Chuhao Zhou, Yuanrong Xu, Fanglin Chen 0001, Guangming Lu 0002 |
PRCV (2) | 4 |
| 2022 | Printed label defect detection using twice gradient matching based on improved cosine similarity measure
Dongming Li 0002, Jinxing Li 0003, Yuanyi Fan, Guangming Lu 0002, Jie Ge |
Expert Syst. Appl. | 4 |
| 2022 | Learning Sequence Representations by Non-local Recurrent Neural Memory
Wenjie Pei, Xin Feng 0005, Canmiao Fu, Qiong Cao, Guangming Lu 0002, Yu-Wing Tai |
Int. J. Comput. Vis. | 5 |
| 2022 | A survey of crowd counting and density estimation based on convolutional neural network
Zizhu Fan, Zheng Zhang 0006, Guangming Lu 0002, Yudong Zhang 0001, Yaowei Wang 0001 |
Neurocomputing | 4 |
| 2022 | Cognitive multi-modal consistent hashing with flexible semantic transformation
Junfeng An, Haoyang Luo, Zheng Zhang 0006, Lei Zhu 0002, Guangming Lu 0002 |
Inf. Process. Manag. | 5 |
| 2022 | Multiscale feature fusion for surveillance video diagnosis
Fanglin Chen 0001, Weihang Wang 0005, Huiyuan Yang, Wenjie Pei, Guangming Lu 0002 |
Knowl. Based Syst. | 5 |
| 2022 | Neighborhood-Exact Nearest Neighbor Search for face retrieval
Fanglin Chen 0001, Wenjie Pei, Guangming Lu 0002 |
Knowl. Based Syst. | 3 |
| 2022 | An exploration of mutual information based on emotion-cause pair extraction
Guimin Hu, Yi Zhao 0007, Guangming Lu 0002, Fanghao Yin, Jiashan Chen |
Knowl. Based Syst. | 3 |
| 2022 | Real noise image adjustment networks for saliency-aware stylistic color retouch
Bo Jiang 0017, Yao Lu 0008, Guangming Lu 0002, David Zhang 0001 |
Knowl. Based Syst. | 3 |
| 2022 | Recursive Feature Diversity Network for audio super-resolution
Bo Jiang 0017, Mi-Xiao Hou, Yao Lu 0008, David Zhang 0001, Guangming Lu 0002 |
Speech Commun. | 6 |
| 2022 | Multi-View Speech Emotion Recognition Via Collective Relation ConstructionabstractAutomatic emotion recognition from speech plays a fundamental role towards advanced emotional intelligence in human-machine interaction systems. The discriminative knowledge from speech for effective emotion recognition may come from multiple physical properties such as energy spectrum, frequency, prosody, which could be collected as multi-view representations. However, the current works fail to fully explore the underlying interactive relations among multiple speech representations for emotion recognition. In this paper, we propose a novel Collective Multi-view Relation Network (CMRN) to exploit the intrinsic characteristics of multi-view speech representations for discriminative speech emotion recognition. Generally, the proposed CMRN consists of three sub-networks,i.e.,view-specific attention network, multi-view shared attention network and collective relation network. Specifically, the view-specific attention network is designed to excavate the distinguishable view-specific features deduced from the original speech. By contrast, the multi-view shared attention network is conceived to capture the collaborative knowledge from multiple views. Moreover, a well-designed collective relation network is explicitly constructed to characterize the shared-specific correlations, which could reflect the underlying physical interaction capabilities. As such, the decision phase can comprehensively leverage the shared and view-specific information of multiple representations, such that the final privileged deciding principle can aggregate the heterogeneous information of multi-view features to make accurate emotion recognition. Extensive experiments on two benchmark datasets demonstrate the superb performance of the proposed method in comparison with some state-of-the-art methods. Mi-Xiao Hou, Zheng Zhang 0006, David Zhang 0001, Guangming Lu 0002 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | Stepwise-Refining Speech Separation Network via Fine-Grained Encoding in High-Order Latent DomainabstractThe crux of single-channel speech separation is how to encode the mixture of signals into such a latent embedding space that the signals from different speakers can be precisely separated. Existing methods for speech separation either transform the speech signals into frequency domain to perform separation or seek to learn a separable embedding space by constructing a latent domain based on convolutional filters. While the latter type of methods learning an embedding space achieves substantial improvement for speech separation, we argue that the embedding space defined by only one latent domain does not suffice to provide a thoroughly separable encoding space for speech separation. In this paper, we propose the Stepwise-Refining Speech Separation Network (SRSSN), which follows a coarse-to-fine separation framework. It first learns a 1-order latent domain to define an encoding space and thereby performs a rough separation in the coarse phase. Then the proposedSRSSNlearns a new latent domain along each basis function of the existing latent domain to obtain a high-order latent domain in the refining phase, which enables our model to perform a refining separation to achieve a more precise speech separation. We demonstrate the effectiveness of ourSRSSNby conducting extensive experiments, including speech separation in a clean (noise-free) setting on WSJ0-2/3mix datasets as well as in noisy/reverberant settings on WHAM!/WHAMR! datasets. Furthermore, we also perform experiments of speech recognition on separated speech signals by our model to evaluate the performance of speech separation indirectly. Zengwei Yao, Wenjie Pei, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Low-Rank Tensor Graph Learning for Multi-View Subspace ClusteringabstractGraph and subspace clustering methods have become the mainstream of multi-view clustering due to their promising performance. However, (1) since graph clustering methods learn graphs directly from the raw data, when the raw data is distorted by noise and outliers, their performance may seriously decrease; (2) subspace clustering methods use a “two-step” strategy to learn the representation and affinity matrix independently, and thus may fail to explore their high correlation. To address these issues, we propose a novel multi-view clustering method via learning aLow-RankTensorGraph (LRTG). Different from subspace clustering methods, LRTG simultaneously learns the representation and affinity matrix in a single step to preserve their correlation. We apply Tucker decomposition and$l_{2,1}$-norm to the LRTG model to alleviate noise and outliers for learning a “clean” representation. LRTG then learns the affinity matrix from this “clean” representation. Additionally, an adaptive neighbor scheme is proposed to find the$K$largest entries of the affinity matrix to form a flexible graph for clustering. An effective optimization algorithm is designed to solve the LRTG model based on the alternating direction method of multipliers. Extensive experiments on different clustering tasks demonstrate the effectiveness and superiority of LRTG over seventeen state-of-the-art clustering methods. Yongyong Chen, Xiaolin Xiao, Chong Peng 0001, Guangming Lu 0002, Yicong Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Multi-Label Chest X-Ray Image Classification via Semantic Similarity Graph EmbeddingabstractAutomated multi-label chest X-ray (CXR) image classification has recently made significant progress in clinical diagnosis based on the advanced deep learning techniques. However, most existing methods mainly focus on analyzing locality visual cues from a single image but fail to leverage the underlying explicit correlations among different images for precise disease diagnosis. By contrast, an experienced radiologist expertizes in transferring knowledge from previous tasks to diagnose the present radiograph. To enable the machine like a radiologist, this paper proposes a novel Semantic Similarity Graph Embedding (SSGE) framework, which explicitly explores the semantic similarities among images to optimize the visual feature embedding for improving the performance of multi-label CXR images classification. Specifically, the proposed SSGE framework contains three main components: the image feature embedding (IFE) module, similarity graph construction (SGC) module, and semantic similarity learning (SSL) module. To realize interactive teaching and learning between visual and semantic information, the proposed SSGE framework is built on the “Teacher-Student” (semantic-visual) learning mechanism. With the guidance and supervision of the cross-image similarity graph generated by the SGC module, the SSL module leverages Graph Convolutional Network (GCN) to adaptively recalibrate the multi-image feature representations extracted from the IFE module, which guarantees their semantic consistency. Furthermore, we propose a novel re-weighting strategy to learn a more optimal semantic-similarity graph for the information propagation of the GCN layers. Extensive experiments on two benchmark datasets demonstrate the effectiveness of the proposed method in comparison with some state-of-the-art baselines. Bingzhi Chen, Zheng Zhang 0006, Yingjian Li 0001, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Generative Memory-Guided Semantic Reasoning Model for Image InpaintingabstractThe critical challenge of single image inpainting stems from accurate semantic inference via limited information while maintaining image quality. Typical methods for semantic image inpainting train an encoder-decoder network by learning a one-to-one mapping from the corrupted image to the inpainted version. While such methods perform well on images with small corrupted regions, it is challenging for these methods to deal with images with large corrupted area due to two potential limitations. 1) Such one-to-one mapping paradigm tends to overfit each single training pair of images; 2) The inter-image prior knowledge about the general distribution patterns of visual semantics, which can be transferred across images sharing similar semantics, is not explicitly exploited. In this paper, we propose the Generative Memory-guided Semantic Reasoning Model (GM-SRM), which infers the content of corrupted regions based on not only the known regions of the corrupted image, but also the learned inter-image reasoning priors characterizing the generalizable semantic distribution patterns between similar images. In particular, the proposed GM-SRM first pre-learns a generative memory from the whole training data to explicitly learn the distribution of different semantic patterns. Then the learned memory are leveraged to retrieve the matching semantics for the current corrupted image to perform semantic reasoning during image inpainting. While the encoder-decoder network is used for guaranteeing the pixel-level content consistency, our generative priors are favorable for performing high-level semantic reasoning, which is particularly effective for inferring semantic content for large corrupted area. Extensive experiments on Paris Street View, CelebA-HQ, and Places2 benchmarks demonstrate that our GM-SRM outperforms the state-of-the-art methods for image inpainting in terms of both visual quality and quantitative metrics. Xin Feng 0005, Wenjie Pei, Fengjun Li, Fanglin Chen 0001, David Zhang 0001, Guangming Lu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Deep Image Denoising With Adaptive PriorsabstractImage denoising methods using deep neural networks have achieved a great progress in the image restoration. However, the recovered images restored by these deep denoising methods usually suffer from severe over-smoothness, artifacts, and detail loss. To improve the quality of restored images, we first propose Supplemental Priors (SP) method to adaptively predict depth-directed and sample-directed prior information for the reconstruction (decoder) networks. Furthermore, the over-parameterized deep neural networks and too precise supplemental prior information may cause an over-fitting, restricting the performance promotion. To improve the generalization of denoising networks, we further propose Regularization Priors (RP) method to flexibly learn depth-directed and dataset-directed regularization noise for the retrieving (encoder) networks. By respectively integrating the encoder and decoder with these plug-and-play RP block and SP block, we propose the final Adaptive Prior Denoising Networks, called APD-Nets. APD-Nets is the first attempt to simultaneously regularize and supplement denoising networks from the adaptive priors’ view with drawing learning-based mechanism into producing adaptive regularization noise and supplemental information. Extensive experiment results demonstrate our method significantly improves the generalization of denoising networks and the quality of restored images with greatly outperforming the traditional deep denoising methods both quantitatively and visually.The code will be released athttps://github.com/JiangBoCS/APD-Nets. Bo Jiang 0017, Yao Lu 0008, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Self-Supervised Exclusive-Inclusive Interactive Learning for Multi-Label Facial Expression Recognition in the WildabstractFacial Expression Recognition (FER) is a long-standing but challenging research problem in computer vision. Existing approaches mainly focus on single-label emotional prediction, which cannot handle the complex multi-label FER task because of the coupling behavior of multiple emotions on a single facial image. To this end, in this paper, we propose a novel Self-supervised Exclusive-Inclusive Interactive Learning (SEIIL) method to facilitate discriminative multi-label FER in the wild, which can effectively handle the coupled multiple sentiments with limited unconstrained training data. Specifically, we construct an emotion disentangling module to capture the inclusive and exclusive characteristics of facial expressions, which can decouple the compound numerous emotions on an image. Moreover, an adaptively-weighted ensemble technique is conceived to aggregate category-level latent exclusive embeddings, and then a conditional adversarial interactive learning module is designed to fully leverage the complementary between the inclusive and formulated latent representations. Furthermore, to tackle the insufficient data for training, we introduce a self-supervised learning strategy to augment the amount and diversity of facial images, which can endow the model with advanced generalization ability. Under this strategy, the proposed two modules can be concurrently utilized in our SEIIL to jointly handle the coupled emotions and alleviate the overfitting problem. Extensive experimental results on six databases illustrate the superb performance of our method against state-of-the-art baselines. Yingjian Li 0001, Yingnan Gao, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Learning Informative and Discriminative Features for Facial Expression Recognition in the WildabstractThe informativeness and discriminativeness of features collaboratively ensure high-accuracy Facial Expression Recognition (FER) in the wild. Most of existing methods use the single-path deep convolutional neural network with softmax loss for basic FER, while they cannot deal with the challenging situations of the compound FER in the wild, because they fail to learn informative and discriminative features in a targeted manner. To this end, we present an Informative and Discriminative Feature Learning (IDFL) framework that consists of two key components: the Multi-Path Attention Convolutional Neural Network (MPACNN) and Balanced Separate loss (BS loss), for both basic and compound high-accuracy FER in the wild. Specifically, MPACNN leverages different paths to learn diverse features. These features are then adaptively fused into informative ones via an attention module, such that the model can adequately capture detailed information for both basic and compound FER. The BS loss maximizes the inter-class distance of features and minimizes the intra-class one. In this way, the features are discriminative enough for high-accuracy FER in the wild. Particularly, the BS loss is invoked as the objective function of MPACNN, so the model can learn informative and discriminative features at the same time, yielding better performance. Seven databases are utilized to evaluate the proposed method, and the results demonstrate that our method achieves state-of-the-art performance on both basic and compound expressions with good generalization ability. Moreover, our model contains fewer parameters and can be trained faster than other related models. Yingjian Li 0001, Yao Lu 0008, Bingzhi Chen, Zheng Zhang 0006, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Multiscale Conditional Regularization for Convolutional Neural NetworksabstractWith the increased model size of convolutional neural networks (CNNs), overfitting has become the main bottleneck to further improve the performance of networks. Currently, the weighting regularization methods have been proposed to address the overfitting problem and they perform satisfactorily. Since these regularization methods cannot be used in all the networks and they are usually not flexible enough in different phases of the training and test processes, this article proposes a multiscale conditional (MSC) regularization method. MSC divides the intermediate features into different scales and then generates new data for each scale features, respectively. In addition, the new data are generated by employing the information from two conditions: 1) each sample feature and 2) each layer pattern. Finally, a self-identity structure is proposed to supplement the features with the generated data. Therefore, MSC can adaptively and efficiently generate much finer and individualized data to make the entire regularization more flexible. Furthermore, MSC is more general and can be applied to all kinds of networks through the proposed self-identity structure. The experimental results on all the benchmark datasets showed that the proposed MSC regularization method achieves the best performances in all the networks. Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, Yuanrong Xu, Zheng Zhang 0006, David Zhang 0001 |
IEEE Trans. Cybern. | 2 |
| 2022 | Addi-Reg: A Better Generalization-Optimization Tradeoff Regularization Method for Convolutional Neural NetworksabstractIn convolutional neural networks (CNNs), generating noise for the intermediate feature is a hot research topic in improving generalization. The existing methods usually regularize the CNNs by producing multiplicative noise (regularization weights), called multiplicative regularization (Multi-Reg). However, Multi-Reg methods usually focus on improving generalization but fail to jointly consider optimization, leading to unstable learning with slow convergence. Moreover, Multi-Reg methods are not flexible enough since the regularization weights are generated from a definite manual-design distribution. Besides, most popular methods are not universal enough, because these methods are only designed for the residual networks. In this article, we, for the first time, experimentally and theoretically explore the nature of generating noise in the intermediate features for popular CNNs. We demonstrate that injecting noise in the feature space can be transformed to generating noise in the input space, and these methods regularize the networks in a Mini-batch in Mini-batch (MiM) sampling manner. Based on these observations, this article further discovers that generating multiplicative noise can easily degenerate the optimization due to its high dependence on the intermediate feature. Based on these studies, we propose a novel additional regularization (Addi-Reg) method, which can adaptively produce additional noise with low dependence on intermediate feature in CNNs by employing a series of mechanisms. Particularly, these well-designed mechanisms can stabilize the learning process in training, and our Addi-Reg method can pertinently learn the noise distributions for every layer in CNNs. Extensive experiments demonstrate that the proposed Addi-Reg method is more flexible and universal, and meanwhile achieves better generalization performance with faster convergence against the state-of-the-art Multi-Reg methods. Yao Lu 0008, Zheng Zhang 0006, Guangming Lu 0002, Yicong Zhou, Jinxing Li 0003, David Zhang 0001 |
IEEE Trans. Cybern. | 3 |
| 2022 | High Resolution Fingerprint Retrieval Based on Pore Indexing and Graph ComparisonabstractFingerprint retrieval aims to identify a query fingerprint image in a large database using indexing algorithms. Because of the abundant level 3 pore features within high-resolution fingerprint images, pore-based fingerprint retrieval algorithms have been rapidly developed. These retrieval algorithms, however, suffer from severe calculation-consuming problems with the pores increasing. This paper proposes a pore-based fingerprint retrieval method for high-resolution fingerprint images. The proposed method consists of two main steps. 1) In the pore indexing step, an indexing space is constructed using the binary codes of pores in enrolled images. Then, a designed graph-based searching algorithm searches the nearest neighbors of pores from the query image to construct one-to-many correspondences. 2) In the refinement step, the one-to-many correspondences are refined by a random walker-based graph comparison algorithm to remove the false correspondences. The remained nearest neighbors are used to calculate the similarities between the query image and the enrolled images. The proposed method is evaluated on two databases, showing that our method achieves better retrieval accuracies with a higher speed than the existing pore-based retrieval algorithms. Yuanrong Xu, Yao Lu 0008, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2022 | Targeted Attack of Deep Hashing Via Prototype-Supervised Adversarial NetworksabstractDue to its powerful capability of representation learning and efficient computation, deep hashing has made significant progress in large-scale image retrieval. It has been recognized that deep neural networks are vulnerable to adversarial examples, which is a practical secure problem but seldom studied in deep hashing-based retrieval field. In this paper, we propose a novel prototype-supervised adversarial network (ProS-GAN), which formulates a flexible generative architecture for efficient and effective targeted hashing attack.To the best of our knowledge, this is one of the first generation-based methods to attack deep hashing networks. Generally, our proposed framework consists of three parts,i.e., a PrototypeNet, a Generator and a Discriminator. Specifically, the designed PrototypeNet embeds the target label into the semantic representation and learns the prototype code as the category-level representative of the target label. Moreover, the semantic representation and the original image are jointly fed into the generator for flexible targeted attack. Particularly, the prototype code is adopted to supervise the generator to construct the targeted adversarial example by minimizing the Hamming distance between the hash code of the adversarial example and the prototype code. Furthermore, the generator fools the discriminator to simultaneously encourage the adversarial examples visually realistic and the semantic representation informative. Extensive experiments demonstrate that the proposed framework can efficiently produce adversarial examples with better targeted attack performance and transferability over state-of-the-art targeted attack methods of deep hashing. The source code is available athttps://github.com/xunguangwang/ProS-GAN_Trans. Zheng Zhang 0006, Xunguang Wang, Guangming Lu 0002, Fumin Shen, Lei Zhu 0002 |
IEEE Trans. Multim. | 3 |
| 2022 | Discriminative Visual Similarity Search with Semantically Cycle-consistent Hashing NetworksabstractDeep hashing has great potential in large-scale visual similarity search due to its preferable efficiency in storage and computation. Technically, deep hashing for visual similarity search inherits the powerful representation capability of deep neural networks, and it encodes visual features into compact binary codes by preserving representative semantic visual features. Works in this field mainly focus on building the relationship between the visual and objective hash spaces, while they seldom study the triadic cross-domain semantic knowledge transfer among visual, semantic, and hashing spaces, leading to a serious semantic ignorance problem during space transformation. In this article, we propose a novel deep tripartite semantically interactive hashing framework, dubbed Semantically Cycle-consistent Hashing Networks (SCHNs), for discriminative hash code learning. Particularly, we construct a flexible semantic space and a transitive latent space, in conjunction with the visual space, to jointly deduce the privileged discriminative hash space. Specifically, a new semantic space is conceived to strengthen the flexibility and completeness of categories in the semantic feature inference phase. At the same time, a transitive latent space is formulated to explore and uncover the shared semantic interactivity embedded in visual and semantic features. Moreover, to further ensure semantic consistency across multiple spaces, we propose to build a cyclic adversarial learning module to preserve and keep their semantic concurrence during space transformation. Notably, our SCHN, for the first time, establishes the cyclic principle of deep semantic-preserving hashing by adaptive semantic parsing across different spaces in a single-modal visual similarity search. In addition, the entire learning framework is jointly optimized in an end-to-end manner. Extensive experiments performed on diverse large-scale datasets evidence the superiority of our method against other state-of-the-art deep hashing algorithms. The source codes of this article are available at https://github.com/JalinWang/SCHN. Zheng Zhang 0006, Jianning Wang, Lei Zhu 0002, Guangming Lu 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Semantic-Interactive Graph Convolutional Network for Multilabel Image RecognitionabstractMultilabel image recognition, a critically practical task in computer vision, aims to predict multiple objects present in each image. The existing studies mainly focus on conceptual visual cues but fail to reconcile the visual information with their semantic guidance. Intuitively, humans can not only associate extra topological concepts but also imagine other approximate scenes based on a semantic description. Inspired by such semantic-interactive capability, two different types of semantic priors, i.e., the concept correlations of the same scene and semantic similarities among different scenes, should be further explored for the recognition decisions. To efficiently interact with these semantic relationships, in this article, we propose a novel semantic-interactive graph convolutional network (SI-GCN), which can leverage the topological information learned from knowledge graphs to boost the performance of multilabel recognition. Specifically, the proposed SI-GCN framework consists of two different GCN-based branches in parallel, i.e., concept correlations learning (CCL) branch and semantic similarity learning (SSL) branch. Inputting the semantic-embedding vectors of all the concepts, the CCL branch maps the label co-occurrence graph into a set of interdependent concept classifiers. Recalibrating the image feature embedding with the standardized supervision of the semantic similarity graph, the SSL branch learns the semantically consistent in-batch visual representations. Finally, a well-established interactive learning scheme is formulated to concurrently optimize the obtained concept classifiers and the visual representation learning in an end-to-end manner. Extensive experiments on the MS-COCO and Pascal VOC 2007 & 2012 benchmarks demonstrate the superiorities of the proposed SI-GCN method compared to the state-of-the-art baselines. Bingzhi Chen, Zheng Zhang 0006, Yao Lu 0008, Fanglin Chen 0001, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2022 | Innovative Contactless Palmprint Recognition System Based on Dual-Camera AlignmentabstractRecently, contactless bimodal palmprint recognition technology has attracted increased attention due to the COVID-19 pandemic. Many dual-camera-based sensors have been proposed to capture palm vein and palmprint images synchronously. However, translations between captured palmprint and palm vein images differ depending on the distance between the hand and the sensors. To address this issue, we designed a low-cost method to align the bimodal palm regions for current dual-camera systems. In this study, we first implemented a contactless palm image acquisition device with a dual-camera module and a single-point time of flight (TOF) ranging sensor. Using this device, we collected a dataset named DCPD under different distances and light source intensities from 271 different palms. Then, a bimodal palm image alignment method is proposed based on the imaging and ranging models. After the system model is calibrated, the translation between the visible light and infrared light palm regions can be estimated quickly based on the palm distance. Finally, we designed a convolutional neural network (CNN) to effectively extract the fine- and coarse-grained palm features. Compared to widely used existing methods, the proposed networks achieved the lowest equal error rate (EER) on the Tongji, IITD, and DCPD datasets, and the average time cost of the system to perform one-time identification is approximately 0.15 s. The experimental results indicate that the proposed methods achieved high efficiency and comparable accuracy. In addition, the system’s EER and rank-1 on the DCPD dataset were 0.304% and 98.66%, respectively. Zhaoqun Li, Bob Zhang 0001, Guangming Lu 0002, David Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2021 | Prototype-Supervised Adversarial Network for Targeted Attack of Deep HashingabstractDue to its powerful capability of representation learning and high-efficiency computation, deep hashing has made significant progress in large-scale image retrieval. However, deep hashing networks are vulnerable to adversarial examples, which is a practical secure problem but seldom studied in hashing-based retrieval field. In this paper, we propose a novel prototype-supervised adversarial network (ProS-GAN), which formulates a flexible generative architecture for efficient and effective targeted hashing attack. To the best of our knowledge, this is the first generation-based method to attack deep hashing networks. Generally, our proposed framework consists of three parts, i.e., a PrototypeNet, a generator and a discriminator. Specifically, the designed PrototypeNet embeds the target label into the semantic representation and learns the prototype code as the category-level representative of the target label. Moreover, the semantic representation and the original image are jointly fed into the generator for flexible targeted attack. Particularly, the prototype code is adopted to supervise the generator to construct the targeted adversarial example by minimizing the Hamming distance between the hash code of the adversarial example and the prototype code. Furthermore, the generator is against the discriminator to simultaneously encourage the adversarial examples visually realistic and the semantic representation informative. Extensive experiments verify that the proposed framework can efficiently produce adversarial examples with better targeted attack performance and transferability over state-of-the-art targeted attack methods of deep hashing. Xunguang Wang, Zheng Zhang 0006, Baoyuan Wu, Fumin Shen, Guangming Lu 0002 |
CVPR | 5 |
| 2021 | Hierarchical Network Based on the Fusion of Static and Dynamic Features for Speech Emotion RecognitionabstractMany studies on automatic speech emotion recognition (SER) have been devoted to extracting meaningful emotional features for generating emotion-relevant representations. However, they generally ignore the complementary learning of static and dynamic features, leading to limited performances. In this paper, we propose a novel hierarchical network called HNSD that can efficiently integrate the static and dynamic features for SER. Specifically, the proposed HNSD framework consists of three different modules. To capture the discriminative features, an effective encoding module is firstly designed to simultaneously encode both static and dynamic features. By taking the obtained features as inputs, the Gated Multi-features Unit (GMU) is conducted to explicitly determine the emotional intermediate representations for frame-level features fusion, instead of directly fusing these acoustic features. In this way, the learned static and dynamic features can jointly and comprehensively generate the unified feature representations. Benefiting from a well-designed attention mechanism, the last classification module is applied to predict the emotional states at the utterance level. Extensive experiments on the IEMOCAP benchmark dataset demonstrate the superiority of our method in comparison with state-of-the-art baselines. Mi-Xiao Hou, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002 |
ICASSP | 5 |
| 2021 | Contrastive Feature Decomposition for Image Reflection RemovalabstractThe crux of image reflection removal stems from the difficulty of recognizing the diverse reflection patterns. Typical methods optimize the modeling of background restoration by performing low-level supervision on the restored image to minimize its per-pixel difference from the groundtruth, which re-lies on substantial training samples to learn diverse reflection patterns robustly and avoid overfitting spurious reflection patterns. In this work, we perform supervision on the contrastive distribution between the predicted background and the reflection image. Specifically, our proposed method restores the background and the reflection images in parallel, and seeks to maximize the distribution consistency between the predicted background-reflection contrast and the groundtruth contrast in the latent space. Such supervision pushes the model to focus on contrastive modeling between the background and reflection image. Extensive experiments on four real-world bench-marks demonstrate that our method consistently outperforms state-of-the-art methods. Xin Feng 0005, Haobo Ji, Bo Jiang 0017, Wenjie Pei, Fanglin Chen 0001, Guangming Lu 0002 |
ICME | 6 |
| 2021 | Partial Tubal Nuclear Norm Regularized Multi-view LearningabstractMulti-view clustering and multi-view dimension reduction explore ubiquitous and complementary information between multiple features to enhance the clustering, recognition performance. However, multi-view clustering and multi-view dimension reduction are treated independently, ignoring the underlying correlations between them. In addition, previous methods mainly focus on using the tensor nuclear norm for low-rank representation to explore the high correlation of multi-view features, which often causes the estimation bias of the tensor rank. To overcome these limitations, we propose the partial tubal nuclear norm regularized multi-view learning (PTN2ML) method, in which the partial tubal nuclear norm as a non-convex surrogate of the tensor tubal multi-rank, only minimizes the partial sum of the smaller tubal singular values to preserve the low-rank property of the self-representation tensor. PTN2ML pursues the latent representation from the projection space rather than from the input space to reveal the structural consensus and suppress the disturbance of noisy data. The proposed method can be efficiently optimized by the alternating direction method of multipliers. Extensive experiments, including multi-view clustering and multi-view dimension reduction substantiate the superiority of the proposed methods beyond state-of-the-arts. Yongyong Chen, Shuqin Wang 0001, Chong Peng 0001, Guangming Lu 0002, Yicong Zhou |
ACM Multimedia | 4 |
| 2021 | JDMAN: Joint Discriminative and Mutual Adaptation Networks for Cross-Domain Facial Expression RecognitionabstractCross-domain Facial Expression Recognition (FER) is challenging due to the difficulty of concurrently handling the domain shift and semantic gap during domain adaptation. Existing methods mainly focus on reducing the domain discrepancy for transferable features but fail to decrease the semantic one, which may result in negative transfer. To this end, we propose Joint Discriminative and Mutual Adaptation Networks (JDMAN), which collaboratively bridge the domain shift and semantic gap by domain- and category-level co-adaptation based on mutual information and discriminative metric learning techniques. Specifically, we design a mutual information minimization module for domain-level adaptation, which narrows the domain shift by simultaneously distilling the domain-invariant components and eliminating the untransferable ones lying in different domains. Moreover, we propose a semantic metric learning module for category-level adaptation, which can close the semantic discrepancy during discriminative intra-domain representation learning and transferable inter-domain knowledge discovery. These two modules are jointly leveraged in our JDMAN to safely transfer the source knowledge to target data in an end-to-end manner. Extensive experimental results on six databases show that our method achieves state-of-the-art performance. The code of our JDMAN is available at https://github.com/YingjianLi/JDMAN. Yingjian Li 0001, Yingnan Gao, Bingzhi Chen, Zheng Zhang 0006, Lei Zhu 0002, Guangming Lu 0002 |
ACM Multimedia | 6 |
| 2021 | Towards Discriminative Visual Search via Semantically Cycle-consistent Hashing NetworksabstractDeep hashing has shown great potentials in large-scale visual similarity search due to preferable storage and computation efficiency. Typically, deep hashing encodes visual features into compact binary codes by preserving representative semantic visual features. Works in this area mainly focus on building the relationship between the visual and objective hash space, while they seldom study the triadic cross-domain semantic knowledge transfer among visual, semantic and hashing spaces, leading to serious semantic ignorance problem during space transformation. In this paper, we propose a novel deep tripartite semantically interactive hashing framework, dubbed Semantically Cycle-consistent Hashing Networks (SCHN), for discriminative hash code learning. Particularly, we construct a flexible semantic space and a transitive latent space, in conjunction with the visual space, to jointly deduce the privileged discriminative hash space. Specifically, a semantic space is conceived to strengthen the flexibility and completeness of categories in feature inference. Moreover, a transitive latent space is formulated to explore the shared semantic interactivity embedded in visual and semantic features. Our SCHN, for the first time, establishes the cyclic principle of deep semantic-preserving hashing by adaptive semantic parsing across different spaces in visual similarity search. In addition, the entire learning framework is jointly optimized in an end-to-end manner. Extensive experiments performed on diverse large-scale datasets evidence the superiority of our method against other state-of-the-art deep hashing algorithms. Zheng Zhang 0006, Jianning Wang, Guangming Lu 0002 |
MMAsia | 3 |
| 2021 | An Embarrassingly Simple Approach to Discrete Supervised HashingabstractPrior hashing works typically learn a projection function from high-dimensional visual feature space to low-dimensional latent space. However, such a projection function remains several crucial bottlenecks: 1) information loss and coding redundancy are inevitable; 2) the available information of semantic labels is not well-explored; 3) the learned latent embedding lacks explicit semantic meaning. To overcome these limitations, we propose a novel supervised Discrete Auto-Encoder Hashing (DAEH) framework, in which a linear auto-encoder can effectively project the semantic labels of images into a latent representation space. Instead of using the visual feature projection, the proposed DAEH framework skillfully explores the semantic information of supervised labels to refine the latent feature embedding and further optimizes hashing function. Meanwhile, we reformulate the objective and relax the discrete constraints for the binary optimization problem. Extensive experiments on Caltech-256, CIFAR-10, and MNIST datasets demonstrate that our method can outperform the state-of-the-art hashing baselines. Shuguang Zhao, Bingzhi Chen, Zheng Zhang 0006, Guangming Lu 0002 |
MMAsia | 4 |
| 2021 | Targeted Attack and Defense for Deep HashingabstractDeep hashing methods have been intensively studied and successfully applied in massive fast image retrieval. However, inherited from the deficiency of deep neural networks, deep hashing models can be easily fooled by adversarial examples, which brings a serious security risk to hashing based retrieval. In this paper, we propose a novel targeted attack method and the first defense scheme for deep hashing based retrieval. Specifically, a simple yet effective PrototypeNet is designed to generate category-level semantic embedding (dubbed prototype code) regarded as the semantic representative of the target label, which preserves the semantic similarity with relevant labels and dissimilarity with irrelevant labels. Subsequently, we conduct the targeted attack by minimizing the Hamming distance between the hash code of the adversarial sample and the prototype code. Moreover, we provide an adversarial training algorithm to improve the adversarial robustness of deep hashing networks. Extensive experiments demonstrate our method can produce high-quality adversarial samples with the benefit of superior targeted attack performance over state-of-the-arts. Importantly, our adversarial defense framework can significantly boost the robustness of hashing networks against adversarial attacks on deep hashing based retrieval. The code is available at https://github.com/xunguangwang/Targeted-Attack-and-Defense-for-Deep-Hashing. Xunguang Wang, Zheng Zhang 0006, Guangming Lu 0002, Yong Xu 0001 |
SIGIR | 3 |
| 2021 | Highly shared Convolutional Neural Networks
Yao Lu 0008, Guangming Lu 0002, Yicong Zhou, Jinxing Li 0003, Yuanrong Xu, David Zhang 0001 |
Expert Syst. Appl. | 2 |
| 2021 | FSS-GCN: A graph convolutional networks with fusion of semantic and structure for emotion cause analysis
Guimin Hu, Guangming Lu 0002, Yi Zhao 0007 |
Knowl. Based Syst. | 2 |
| 2021 | Fully shared convolutional neural networks
Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, Zheng Zhang 0006, Yuanrong Xu |
Neural Comput. Appl. | 2 |
| 2021 | CompNet: Competitive Neural Network for Palmprint Recognition Using Learnable Gabor KernelsabstractContactless palmprint recognition has recently made significant progress in palm-scanning payment and social security. However, most existing methods are based on handcrafted kernels and are sensitive to illumination and scale variations. To address this problem, a competitive convolutional neural network (CompNet) with constrained learnable Gabor filters is proposed for contactless palmprint recognition. The proposed CompNet is built on multisize competitive blocks, which are applied to effectively exploit the rich direction ordering information of the palmprint patterns by means of the ad-hoc softmax and channel-wise convolution operations. Compared to the current deep neural networks, the backbone of the proposed network contains only very few parameters, making it quite easy to train, especially on small-scale datasets. Experimental results obtained on four popular contactless palmprint datasets demonstrate that the proposed CompNet achieves the lowest equal error rate compared to the most commonly used methods. Jinyang Yang, Guangming Lu 0002, David Zhang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Multimodal Emotion Recognition With Temporal and Semantic ConsistencyabstractAutomated multimodal emotion recognition has become an emerging but challenging research topic in the fields of affective learning and sentiment analysis. The existing works mainly focus on developing multimodal fusion strategies to incorporate different emotion-related features. However, they fail to explore the inherent contextual consistency to reconcile the emotional information across modalities. In this paper, we propose a novel Time and Semantic Interaction Network (TSIN), which concurrently incorporates the advantages of temporal and semantic consistency into the multimodal emotion recognition task. Specifically, a well-designed Speech and Text Embedding (STE) module is devoted to formulating the initial embedding spaces by respectively building the modality-specific representations of speech and text. Instead of separately learning or directly fusing the acoustic and textual features, we propose a well-defined Time and Semantic Interaction (TSI) module to conduct the emotional parsing and sentiment refining by performing the fine-grained temporal alignment and cross-modal semantic interaction. Benefitting from temporal and semantic consistency constraints, both speech-text embeddings can be interactively optimized and fine-tuned in the learning process. In this way, the learnt acoustics and textual features can jointly and efficiently predict the final emotional state. Extensive experiments on the IEMOCAP dataset demonstrate the superiorities of our TSIN framework in comparison with state-of-the-art baselines. Bingzhi Chen, Mi-Xiao Hou, Zheng Zhang 0006, Guangming Lu 0002, David Zhang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Shared Linear Encoder-Based Multikernel Gaussian Process Latent Variable Model for Visual ClassificationabstractMultiview learning has been widely studied in various fields and achieved outstanding performances in comparison to many single-view-based approaches. In this paper, a novel multiview learning method based on the Gaussian process latent variable model (GPLVM) is proposed. In contrast to existing GPLVM methods which only assume that there are transformations from the latent variable to the multiple observed inputs, our proposed method simultaneously takes a back constraint into account, encoding multiple observations to the latent variable by enjoying the Gaussian process (GP) prior. Particularly, to overcome the difficulty of the covariance matrix calculation in the encoder, a linear projection is designed to map different observations to a consistent subspace first. The obtained variable in this subspace is then projected to the latent variable in the manifold space with the GP prior. Furthermore, different from most GPLVM methods which strongly assume that the covariance matrices follow a certain kernel function, for example, radial basis function (RBF), we introduce a multikernel strategy to design the covariance matrix, being more reasonable and adaptive for the data representation. In order to apply the presented approach to the classification, a discriminative prior is also embedded to the learned latent variables to encourage samples belonging to the same category to be close and those belonging to different categories to be far. Experimental results on three real-world databases substantiate the effectiveness and superiority of the proposed method compared with state-of-the-art approaches. Jinxing Li 0003, Guangming Lu 0002, Bob Zhang 0001, Jane You, David Zhang 0001 |
IEEE Trans. Cybern. | 2 |
| 2021 | Deep-Masking Generative Network: A Unified Framework for Background Restoration From Superimposed ImagesabstractRestoring the clean background from the superimposed images containing a noisy layer is the common crux of a classical category of tasks on image restoration such as image reflection removal, image deraining and image dehazing. These tasks are typically formulated and tackled individually due to diverse and complicated appearance patterns of noise layers within the image. In this work we present the Deep-Masking Generative Network (DMGN), which is a unified framework for background restoration from the superimposed images and is able to cope with different types of noise. Our proposed DMGN follows a coarse-to-fine generative process: a coarse background image and a noise image are first generated in parallel, then the noise image is further leveraged to refine the background image to achieve a higher-quality background image. In particular, we design the novel Residual Deep-Masking Cell as the core operating unit for our DMGN to enhance the effective information and suppress the negative information during image generation via learning a gating mask to control the information flow. By iteratively employing this Residual Deep-Masking Cell, our proposed DMGN is able to generate both high-quality background image and noisy image progressively. Furthermore, we propose a two-pronged strategy to effectively leverage the generated noise image as contrasting cues to facilitate the refinement of the background image. Extensive experiments across three typical tasks for image background restoration, including image reflection removal, image rain steak removal and image dehazing, show that our DMGN consistently outperforms state-of-the-art methods specifically designed for each single task. Xin Feng 0005, Wenjie Pei, Zihui Jia, Fanglin Chen 0001, David Zhang 0001, Guangming Lu 0002 |
IEEE Trans. Image Process. | 6 |
| 2021 | Layer-Output Guided Complementary Attention Learning for Image Defocus Blur DetectionabstractDefocus blur detection (DBD), which has been widely applied to various fields, aims to detect the out-of-focus or in-focus pixels from a single image. Despite the fact that the deep learning based methods applied to DBD have outperformed the hand-crafted feature based methods, the performance cannot still meet our requirement. In this paper, a novel network is established for DBD. Unlike existing methods which only learn the projection from the in-focus part to the ground-truth, both in-focus and out-of-focus pixels, which are completely and symmetrically complementary, are taken into account. Specifically, two symmetric branches are designed to jointly estimate the probability of focus and defocus pixels, respectively. Due to their complementary constraint, each layer in a branch is affected by an attention obtained from another branch, effectively learning the detailed information which may be ignored in one branch. The feature maps from these two branches are then passed through a unique fusion block to simultaneously get the two-channel output measured by a complementary loss. Additionally, instead of estimating only one binary map from a specific layer, each layer is encouraged to estimate the ground truth to guide the binary map estimation in its linked shallower layer followed by a top-to-bottom combination strategy, gradually exploiting the global and local information. Experimental results on released datasets demonstrate that our proposed method remarkably outperforms state-of-the-art algorithms. Jinxing Li 0003, Lingxiao Yang, Shuhang Gu, Guangming Lu 0002, Yong Xu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2021 | Push for Center Learning via Orthogonalization and Subspace Masking for Person Re-IdentificationabstractPerson re-identification aims to identify whether pairs of images belong to the same person or not. This problem is challenging due to large differences in camera views, lighting and background. One of the mainstream in learning CNN features is to design loss functions which reinforce both the class separation and intra-class compactness. In this paper, we propose a novel Orthogonal Center Learning method with Subspace Masking for person re-identification. We make the following contributions: 1) we develop a center learning module to learn the class centers by simultaneously reducing the intra-class differences and inter-class correlations by orthogonalization; 2) we introduce a subspace masking mechanism to enhance the generalization of the learned class centers; and 3) we propose to integrate the average pooling and max pooling in a regularizing manner that fully exploits their powers. Extensive experiments show that our proposed method consistently outperforms the state-of-the-art methods on large-scale ReID datasets including Market-1501, DukeMTMC-ReID, CUHK03 and MSMT17. Weinong Wang, Wenjie Pei, Qiong Cao, Shu Liu 0005, Guangming Lu 0002, Yu-Wing Tai |
IEEE Trans. Image Process. | 5 |
| 2021 | Probability Ordinal-Preserving Semantic Hashing for Large-Scale Image RetrievalabstractSemantic hashing enables computation and memory-efficient image retrieval through learning similarity-preserving binary representations. Most existing hashing methods mainly focus on preserving the piecewise class information or pairwise correlations of samples into the learned binary codes while failing to capture the mutual triplet-level ordinal structure in similarity preservation. In this article, we propose a novel Probability Ordinal-preserving Semantic Hashing (POSH) framework, which for the first time defines the ordinal-preserving hashing concept under a non-parametric Bayesian theory. Specifically, we derive the whole learning framework of the ordinal similarity-preserving hashing based on the maximum posteriori estimation, where the probabilistic ordinal similarity preservation, probabilistic quantization function, and probabilistic semantic-preserving function are jointly considered into one unified learning framework. In particular, the proposed triplet-ordering correlation preservation scheme can effectively improve the interpretation of the learned hash codes under an economical anchor-induced asymmetric graph learning model. Moreover, the sparsity-guided selective quantization function is designed to minimize the loss of space transformation, and the regressive semantic function is explored to promote the flexibility of the formulated semantics in hash code learning. The final joint learning objective is formulated to concurrently preserve the ordinal locality of original data and explore potentials of semantics for producing discriminative hash codes. Importantly, an efficient alternating optimization algorithm with the strictly proof convergence guarantee is developed to solve the resulting objective problem. Extensive experiments on several large-scale datasets validate the superiority of the proposed method against state-of-the-art hashing-based retrieval methods. Zheng Zhang 0006, Xiaofeng Zhu 0001, Guangming Lu 0002, Yudong Zhang 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2021 | Harmonization Shared Autoencoder Gaussian Process Latent Variable Model With Relaxed Hamming DistanceabstractMultiview learning has shown its superiority in visual classification compared with the single-view-based methods. Especially, due to the powerful representation capacity, the Gaussian process latent variable model (GPLVM)-based multiview approaches have achieved outstanding performances. However, most of them only follow the assumption that the shared latent variables can be generated from or projected to the multiple observations but fail to exploit the harmonization in the back constraint and adaptively learn a classifier according to these learned variables, which would result in performance degradation. To tackle these two issues, in this article, we propose a novel harmonization shared autoencoder GPLVM with a relaxed Hamming distance (HSAGP-RHD). Particularly, an autoencoder structure with the Gaussian process (GP) prior is first constructed to learn the shared latent variable for multiple views. To enforce the agreement among various views in the encoder, a harmonization constraint is embedded into the model by making consistency for the view-specific similarity. Furthermore, we also propose a novel discriminative prior, which is directly imposed on the latent variable to simultaneously learn the fused features and adaptive classifier in a unit model. In detail, the centroid matrix corresponding to the centroids of different categories is first obtained. A relaxed Hamming distance (RHD)-based measurement is subsequently presented to measure the similarity and dissimilarity between the latent variable and centroids, not only allowing us to get the closed-form solutions but also encouraging the points belonging to the same class to be close, while those belonging to different classes to be far. Due to this novel prior, the category of the out-of-sample is also allowed to be simply assigned in the testing phase. Experimental results conducted on three real-world data sets demonstrate the effectiveness of the proposed method compared with state-of-the-art approaches. Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, Yong Xu 0001, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Inductive Structure Consistent Hashing via Flexible Semantic CalibrationabstractSemantic-preserving hashing establishes efficient multimedia retrieval by transferring knowledge from original data to hash codes so that the latter can preserve the underlying visual and semantic similarities. However, it becomes a crucial bottleneck: how to effectively bridge the trilateral domain gaps (i.e., the visual, semantic, and hashing spaces) to further improve the retrieval accuracy. In this article, we propose an inductive structure consistent hashing (ISCH) method, which can interactively coordinate the semantic correlations between the visual feature space, the binary class space, and the discrete hashing space. Specifically, an inductive semantic space is formulated by a simple multilayer stacking class-encoder, which transforms the naive class information into flexible semantic embeddings. Meanwhile, we design a semantic dictionary learning model to facilitate the bilateral visual-semantic bridging and guide the class-encoder toward reliable semantics, which could well alleviate the visual-semantic bias problem. In particular, the visual descriptors and respective semantic class representations are regularized with a coinciding alignment module. In order to generate privileged hash codes, we further explore semantic and prototype binary code learning to jointly quantify the semantic and latent visual representations into unified discrete hash codes. Moreover, an efficient optimization algorithm is developed to address the resulting discrete programming problem. Comprehensive experiments conducted on four large-scale data sets, i.e., CIFAR-10, NUSWIDE, ImageNet, and MSCOCO, demonstrate the superiority of our method over the state-of-the-art alternatives against different evaluation protocols. Zheng Zhang 0006, Luyao Liu 0002, Yadan Luo, Zi Huang, Fumin Shen, Heng Tao Shen, Guangming Lu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2021 | Deep Active Context Estimation for Automated COVID-19 DiagnosisabstractMany studies on automated COVID-19 diagnosis have advanced rapidly with the increasing availability of large-scale CT annotated datasets. Inevitably, there are still a large number of unlabeled CT slices in the existing data sources since it requires considerable consuming labor efforts. Notably, cinical experience indicates that the neighboring CT slices may present similar symptoms and signs. Inspired by such wisdom, we propose DACE, a novel CNN-based deep active context estimation framework, which leverages the unlabeled neighbors to progressively learn more robust feature representations and generate a well-performed classifier for COVID-19 diagnosis. Specifically, the backbone of the proposed DACE framework is constructed by a well-designed Long-Short Hierarchical Attention Network (LSHAN), which effectively incorporates two complementary attention mechanisms, i.e., short-range channel interactions (SCI) module and long-range spatial dependencies (LSD) module, to learn the most discriminative features from CT slices. To make full use of such available data, we design an efficient context estimation criterion to carefully assign the additional labels to these neighbors. Benefiting from two complementary types of informative annotations from -nearest neighbors, i.e., the majority of high-confidence samples with pseudo labels and the minority of low-confidence samples with hand-annotated labels, the proposed LSHAN can be fine-tuned and optimized in an incremental learning manner. Extensive experiments on the Clean-CC-CCII dataset demonstrate the superior performance of our method compared with the state-of-the-art baselines. Bingzhi Chen, Yishu Liu 0001, Zheng Zhang 0006, Yingjian Li 0001, Zhao Zhang 0001, Guangming Lu 0002, Hongbing Yu |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2021 | A Novel Multicamera System for High-Speed Touchless Palm RecognitionabstractPalm-related biometrics have been widely studied for a long time, as the palm contains many distinctive patterns. However, most of the existing systems are designed to work within an ideal environment, such as in front of a unicolor background or in a large enclosure. Those preconditions can avoid influences of ambient light and hand distance change, but at the same time, they also limit the applications of palm recognition. In the work reported in this paper, we designed a novel red-green-blue and depth-based four-camera system that can capture the palm-related images separately in real time. The techniques of region-of-interest (ROI) location, ROI alignment, and light-source intensity optimization were studied. The ROI location method is modified to increase the robustness of hand gesture variation. Based on the depth information, we proposed the coordinate mapping and inclination rectification methods to obtain aligned ROI pairs. Using this device, we collected a video-based multimodal palm image database. After the parameter optimization and information fusion, the equal-error-rate of our approach on this database is lower than 0.47%. The recognition rate obtained from the support-vector-machine-based fusion is higher than 99.8%. The experimental results prove that the proposed system achieves advantages of anti-spoofing, high speed, high accuracy, and small size. David Zhang 0001, Guangming Lu 0002, Zhenhua Guo 0001, Nan Luo |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2021 | Fast Pore Comparison for High Resolution Fingerprint Images Based on Multiple Co-Occurrence Descriptors and Local Topology SimilaritiesabstractPore-based fingerprint recognition has been researched for decades. Many algorithms have been proposed to improve the recognition accuracy of the system. However, the accuracies are always improved at the cost of speed. This article proposes a novel method to compare the pores in high-resolution fingerprint images using the popular coarse-to-fine strategy. A multiple spatial pairwise local co-occurrence descriptor is proposed to improve the calculation of the similarities between pores. It calculates multiple local co-occurrence statistics for each pore using its neighbors. The proposed method can establish correspondences between pores more accurately. The refinement of the correspondences is then achieved by using a local topology-preserving matching algorithm. The algorithm uses rotational invariant local structures and pore pair local topology similarities to calculate the cost of each correspondence. It can remove the mismatches more accurately and efficiently. The experimental results on two high-resolution fingerprint image databases show that the proposed algorithm perform well in both accuracy and speed comparing to the existing algorithms. Yuanrong Xu, Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, David Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2020 | Filter Grafting for Deep Neural NetworksabstractThis paper proposes a new learning paradigm called filter grafting, which aims to improve the representation capability of Deep Neural Networks (DNNs). The motivation is that DNNs have unimportant (invalid) filters (e.g., l1norm close to 0). These filters limit the potential of DNNs since they are identified as having little effect on the network. While filter pruning removes these invalid filters for efficiency consideration, filter grafting re-activates them from an accuracy boosting perspective. The activation is processed by grafting external information (weights) into invalid filters. To better perform the grafting process, we develop an entropy-based criterion to measure the information of filters and an adaptive weighting strategy for balancing the grafted information among networks. After the grafting operation, the network has very few invalid filters compared with its untouched state, empowering the model with more representation capacity. We also perform extensive experiments on the classification and recognition tasks to show the superiority of our method. For example, the grafted MobileNetV2 outperforms the non-grafted MobileNetV2 by about 7 percent on CIFAR-100 dataset. Fanxu Meng 0003, Hao Cheng 0012, Ke Li 0015, Zhixin Xu, Rongrong Ji, Xing Sun 0001, Guangming Lu 0002 |
CVPR | 7 |
| 2020 | Pruning Filter in FilterabstractPruning has become a very powerful and effective technique to compress and accelerate modern neural networks. Existing pruning methods can be grouped into two categories: filter pruning (FP) and weight pruning (WP). FP wins at hardware compatibility but loses at the compression ratio compared with WP. To converge the strength of both methods, we propose to prune the filter in the filter. Specifically, we treat a filter F, whose size is CKK, as KK stripes, i.e., 11 filters, then by pruning the stripes instead of the whole filter, we can achieves finer granularity than traditional FP while being hardware friendly. We term our method as SWP (Stripe-Wise Pruning). SWP is implemented by introducing a novel learnable matrix called Filter Skeleton, whose values reflect the optimal shape of each filter. As some recent work has shown that the pruned architecture is more crucial than the inherited important weights, we argue that the architecture of a single filter, i.e., the Filter Skeleton, also matters. Through extensive experiments, we demonstrate that SWP is more effective compared to the previous FP-based methods and achieves the state-of-art pruning ratio on CIFAR-10 and ImageNet datasets without obvious accuracy drop. Fanxu Meng 0003, Hao Cheng 0012, Ke Li 0015, Huixiang Luo, Guangming Lu 0002, Xing Sun 0001 |
NeurIPS | 6 |
| 2020 | Emotion-Cause Joint Detection: A Unified Network with Dual Interaction for Emotion Cause Analysis
Guimin Hu, Guangming Lu 0002, Yi Zhao 0007 |
NLPCC (1) | 2 |
| 2020 | Similarity and diversity induced paired projection for cross-modal retrieval
Jinxing Li 0003, Mu Li 0005, Guangming Lu 0002, Bob Zhang 0001, Hongpeng Yin, David Zhang 0001 |
Inf. Sci. | 3 |
| 2020 | High-parameter-efficiency convolutional neural networks
Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, Yuanrong Xu, David Zhang 0001 |
Neural Comput. Appl. | 2 |
| 2020 | Two-stream collaborative network for multi-label chest X-ray Image classification with lung segmentation
Bingzhi Chen, Zheng Zhang 0006, Jianyong Lin, Yi Chen 0023, Guangming Lu 0002 |
Pattern Recognit. Lett. | 5 |
| 2020 | A supervised non-negative matrix factorization model for speech emotion recognition
Mi-Xiao Hou, Jinxing Li 0003, Guangming Lu 0002 |
Speech Commun. | 3 |
| 2020 | DRPL: Deep Regression Pair Learning for Multi-Focus Image FusionabstractIn this paper, a novel deep network is proposed for multi-focus image fusion, named Deep Regression Pair Learning (DRPL). In contrast to existing deep fusion methods which divide the input image into small patches and apply a classifier to judge whether the patch is in focus or not, DRPL directly converts the whole image into a binary mask without any patch operation, subsequently tackling the difficulty of the blur level estimation around the focused/defocused boundary. Simultaneously, a pair learning strategy, which takes a pair of complementary source images as inputs and generates two corresponding binary masks, is introduced into the model, greatly imposing the complementary constraint on each pair and making a large contribution to the performance improvement. Furthermore, as the edge or gradient does exist in the focus part while there is no similar property for the defocus part, we also embed a gradient loss to ensure the generated image to be all-in-focus. Then the structural similarity index (SSIM) is utilized to make a trade-off between the reference and fused images. Experimental results conducted on the synthetic and real-world datasets substantiate the effectiveness and superiority of DRPL compared with other state-of-the-art approaches. The testing code can be found in https://github.com/sasky1/DPRL. Jinxing Li 0003, Xiaobao Guo, Guangming Lu 0002, Bob Zhang 0001, Yong Xu 0001, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | Lesion Location Attention Guided Network for Multi-Label Thoracic Disease Classification in Chest X-RaysabstractTraditional clinical experiences have shown the benefit of lesion location attention for improving clinical diagnosis tasks. Inspired by this point of interest, in this paper we propose a novel lesion location attention guided network named LLAGnet to focus on the discriminative features from lesion locations for multi-label thoracic disease classification in chest X-rays (CXRs). By revealing the equivalence of the region-level attention (RLA) and channel-level attention (CLA), we find that the RLA is available as priors for object localization while the CLA implicitly provides high weights to the attractive channels, which both enable lesion location attention excitation. To integrate the advantages from both mechanisms, the proposed LLAGnet is structured with two corresponding attention modules, i.e., the RLA and CLA modules. Specifically, the RLA module consists of the global and local branches. And the weakly supervised attention mechanism embedded in the global branch can obtain visual regions of lesion locations by back-propagating gradients. Then the optimal attention region is amplified and applied to the local branch to provide more fine-grained features for the image classification. Finally, the CLA module adaptively enhances the weights of channel-wise features from the lesion locations by modeling interdependencies among channels. Extensive experiments on the ChestX-ray14 dataset clearly substantiate the effectiveness of LLAGnet as compared with the state-of-the-art baselines. Bingzhi Chen, Jinxing Li 0003, Guangming Lu 0002, David Zhang 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Label Co-Occurrence Learning With Graph Convolutional Networks for Multi-Label Chest X-Ray Image ClassificationabstractExisting multi-label medical image learning tasks generally contain rich relationship information among pathologies such as label co-occurrence and interdependency, which is of great importance for assisting in clinical diagnosis and can be represented as the graph-structured data. However, most state-of-the-art works only focus on regression from the input to the binary labels, failing to make full use of such valuable graph-structured information due to the complexity of graph data. In this paper, we propose a novel label co-occurrence learning framework based on Graph Convolution Networks (GCNs) to explicitly explore the dependencies between pathologies for the multi-label chest X-ray (CXR) image classification task, which we term the "CheXGCN". Specifically, the proposed CheXGCN consists of two modules, i.e., the image feature embedding (IFE) module and label co-occurrence learning (LCL) module. Thanks to the LCL model, the relationship between pathologies is generalized into a set of classifier scores by introducing the word embedding of pathologies and multi-layer graph information propagation. During end-to-end training, it can be flexibly integrated into the IFE module and then adaptively recalibrate multi-label outputs with these scores. Extensive experiments on the ChestX-Ray14 and CheXpert datasets have demonstrated the effectiveness of CheXGCN as compared with the state-of-the-art baselines. Bingzhi Chen, Jinxing Li 0003, Guangming Lu 0002, Hongbing Yu, David Zhang 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Relaxed Asymmetric Deep Hashing Learning: Point-to-Angle MatchingabstractDue to the powerful capability of the data representation, deep learning has achieved a remarkable performance in supervised hash function learning. However, most of the existing hashing methods focus on point-to-point matching that is too strict and unnecessary. In this article, we propose a novel deep supervised hashing method by relaxing the matching between each pair of instances to a point-to-angle way. Specifically, an inner product is introduced to asymmetrically measure the similarity and dissimilarity between the real-valued output and the binary code. Different from existing methods that strictly enforce each element in the real-valued output to be either +1 or -1, we only encourage the output to be close to its corresponding semantic-related binary code under the cross-angle. This asymmetric product not only projects both the real-valued output and the binary code into the same Hamming space but also relaxes the output with wider choices. To further exploit the semantic affinity, we propose a novel Hamming-distance-based triplet loss, efficiently making a ranking for the positive and negative pairs. An algorithm is then designed to alternatively achieve optimal deep features and binary codes. Experiments on four real-world data sets demonstrate the effectiveness and superiority of our approach to the state of the art. Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, Jane You, Yong Xu 0001, Feng Wu 0001, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | SRGC-Nets: Sparse Repeated Group Convolutional Neural NetworksabstractGroup convolution is widely used in many mobile networks to remove the filter's redundancy from the channel extent. In order to further reduce the redundancy of group convolution, this article proposes a novel repeated group convolutional (RGC) kernel, which has M primary groups, and each primary group includes N tiny groups. In every primary group, the same convolutional kernel is repeated in all the tiny groups. The RGC filter is the first kernel to remove the redundancy from group extent. Based on RGC, a sparse RGC (SRGC) kernel is also introduced in this article, and its corresponding network is called SRGC neural networks (SRGC-Net). The SRGC kernel is the summation of RGC kernel and pointwise group convolutional (PGC) kernel. The number of PGC's groups is M . Accordingly, in each primary group, besides the center locations in all channels, the values of parameters located in other N-1 tiny groups are all zero. Therefore, SRGC can significantly reduce the parameters. Moreover, it can also effectively retrieve spatial and channel-difference features by utilizing RGC and PGC to preserve the richness of produced features. Comparative experiments were performed on the benchmark classification data sets. Compared with the traditional popular networks, SRGC-Nets can perform better with timely reducing the model size and computational complexity. Furthermore, it can also achieve better performances than other latest state-of-the-art mobile networks on most of the databases and effectively decrease the test and training runtime. Yao Lu 0008, Guangming Lu 0002, Jinxing Li 0003, David Zhang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Super Sparse Convolutional Neural NetworksabstractTo construct small mobile networks without performance loss and address the over-fitting issues caused by the less abundant training datasets, this paper proposes a novel super sparse convolutional (SSC) kernel, and its corresponding network is called SSC-Net. In a SSC kernel, every spatial kernel has only one non-zero parameter and these non-zero spatial positions are all different. The SSC kernel can effectively select the pixels from the feature maps according to its non-zero positions and perform on them. Therefore, SSC can preserve the general characteristics of the geometric and the channels’ differences, resulting in preserving the quality of the retrieved features and meeting the general accuracy requirements. Furthermore, SSC can be entirely implemented by the “shift” and “group point-wise” convolutional operations without any spatial kernels (e.g., “3×3”). Therefore, SSC is the first method to remove the parameters’ redundancy from the both spatial extent and the channel extent, leading to largely decreasing the parameters and Flops as well as further reducing the img2col and col2img operations implemented by the low leveled libraries. Meanwhile, SSC-Net can improve the sparsity and overcome the over-fitting more effectively than the other mobile networks. Comparative experiments were performed on the less abundant CIFAR and low resolution ImageNet datasets. The results showed that the SSC-Nets can significantly decrease the parameters and the computational Flops without any performance losses. Additionally, it can also improve the ability of addressing the over-fitting problem on the more challenging less abundant datasets. Yao Lu 0008, Guangming Lu 0002, Bob Zhang 0001, Yuanrong Xu, Jinxing Li 0003 |
AAAI | 2 |
| 2019 | Separate Loss for Basic and Compound Facial Expression Recognition in the WildabstractIn the past few years, facial expression recognition has made great progress because of the development of convolutional neural networks. However, the features learned only using the softmax loss are not discriminative enough for highly accurate facial expression recognition in the wild, especially for the compound facial expression recognition. To enhance the discriminative power of the learned features, we propose the separate loss for both basic and compound facial expression recognition in the wild in this paper. Such loss maximizes intra-class similarity while minimizing the similarity between different classes. The qualitative and quantitative analysis shows that the features learned using such loss function are characterized by intra-class compactness and inter-class separation. Experiments are performed on two databases in the wild and the proposed method achieves state-of-the-art results on both basic and compound expressions. Furthermore, another two databases are used to perform cross database experiments to show the generalization ability of our method. Yingjian Li 0001, Yao Lu 0008, Jinxing Li 0003, Guangming Lu 0002 |
ACML | 4 |
| 2019 | Stable Pore Detection for High-Resolution Fingerprint based on a CNN DetectorabstractHigh-resolution fingerprint images contain three levels of features. Pores, as one of the level 3 features, have wide attention due to its significant contribution to the recognition accuracy. An accurate and stable pore detection algorithm plays a key role on the pore-based fingerprint recognition system. This paper proposes a pore detection method for high-resolution fingerprint images. The method uses fully convolutional network combined with the focal loss and shortcut structure to detect pores. The proposed algorithm is tested on the high-resolution fingerprint database. Experimental results show that our method outperforms the existing algorithms in accuracy, stability and matching performance. Zuolin Shen, Yuanrong Xu, Jinxing Li 0003, Guangming Lu 0002 |
ICIP | 4 |
| 2019 | Mask-Most Net: Mask Approximation Based Multi-oriented Scene Text Detection NetworkabstractIn this paper, a novel multi-task cascade framework, which jointly takes the detection and the segmentation into account, is presented for the scene text detection. To address the issue of multi-oriented scene text detection, we propose an instance-level mask approximation method through the auxiliary regression task on center and corner points. Specifically, the text instance in the image is first coarsely detected, followed by a contextual module which can capture more accurate instances. To cope with the scale variation existing in these detected instances, a combination of high-level semantic and low-level features is further exploited, achieving more robust and better performance. A series of experiments conducted on different benchmark datasets demonstrate the effectiveness of the proposed method. Xiaobao Guo, Jinxing Li 0003, Bingzhi Chen, Guangming Lu 0002 |
ICME | 4 |
| 2019 | Multi-label Chest X-Ray Image Classification via Label Co-occurrence Learning
Bingzhi Chen, Yao Lu 0008, Guangming Lu 0002 |
PRCV (2) | 3 |
| 2019 | Single Image Reflection Removal Based on Deep Residual Learning
Zhixin Xu, Xiaobao Guo, Guangming Lu 0002 |
PRCV (2) | 3 |
| 2019 | Body surface feature-based multi-modal Learning for Diabetes Mellitus detection
Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, Jane You, David Zhang 0001 |
Inf. Sci. | 3 |
| 2019 | Weighted Channel-Wise Decomposed Convolutional Neural Networks
Yao Lu 0008, Guangming Lu 0002, Yuanrong Xu |
Neural Process. Lett. | 2 |
| 2019 | Joint learning for voice based disease detection
Kebin Wu, David Zhang 0001, Guangming Lu 0002, Zhenhua Guo 0001 |
Pattern Recognit. | 3 |
| 2019 | High resolution fingerprint recognition using pore and edge descriptors
Yuanrong Xu, Guangming Lu 0002, Yao Lu 0008, David Zhang 0001 |
Pattern Recognit. Lett. | 2 |
| 2019 | Fingerprint Pore Comparison Using Local Features and Spatial RelationsabstractHigh-resolution fingerprint recognition has been a hot topic for many years. Compared with a traditional fingerprint image, a high-resolution fingerprint image can provide more features, such as pores and ridge contours. Introducing these features into fingerprint comparison and recognition can improve the recognition accuracy and reduce the risk of identification errors. This paper proposes a novel method for comparing pores on high-resolution fingerprint images. The method can be divided into two steps. In the first step, fingerprints are aligned using the pixel-category-distance-based data-driven descending algorithm. Traditionally, fingerprints are aligned based on feature points, such as minutiae and singular points. Such alignment methods are not suitable when dealing with partial fingerprints because small overlapping areas often do not contain enough features to guarantee a correct alignment. In this research, the ridges and valleys on fingerprints are used in combination with the orientation field for alignment. The proposed algorithm performs well when aligning both partial and full fingerprints. The common areas between the two images can be estimated based on the alignment result. In the second step, pores lying in the common areas are selected for comparison. To improve the comparison accuracy, pores are compared using local features and spatial relations. A graph comparison algorithm is designed in this step. The experimental results show that the proposed method is more accurate than other state-of-the-art pore comparison algorithms. Yuanrong Xu, Guangming Lu 0002, Yao Lu 0008, Feng Liu 0013, David Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Visual Classification With Multikernel Shared Gaussian Process Latent Variable ModelabstractMultiview learning methods often achieve improvement compared with single-view-based approaches in many applications. Due to the powerful nonlinear ability and probabilistic perspective of Gaussian process (GP), some GP-based multiview efforts were presented. However, most of these methods make a strong assumption on the kernel function (e.g., radial basis function), which limits the capacity of the real data modeling. In order to address this issue, in this paper, we propose a novel multiview approach by combining a multikernel and GP latent variable model. Instead of designing a deterministic kernel function, multiple kernel functions are established to automatically adapt various types of data. Considering a simple way of obtaining latent variables at the testing stage, a projection from the observed space to the latent space as a back constraint has also been simultaneously introduced into the proposed method. Additionally, different from some existing methods which apply the classifiers off-line, a hinge loss is embedded into the model to jointly learn the classification hyperplane, encouraging the latent variables belonging to the different classes to be separated. An efficient algorithm based on the gradient decent technique is constructed to optimize our method. Finally, we apply the proposed approach to three real-world datasets and the associated results demonstrate the effectiveness and superiority of our model compared with other state-of-the-art methods. Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, Hu Ren, David Zhang 0001 |
IEEE Trans. Cybern. | 3 |
| 2019 | Feature Extraction Methods for Palmprint Recognition: A Survey and EvaluationabstractPalmprint processes a number of unique features for reliable personal recognition. However, different types of palmprint images contain different dominant features. Instead, only some features of the palmprint are visible in a palmprint image, whereas the other features may not be notable. For example, the low-resolution palmprint image has visible principal lines and wrinkles. By contrast, the high-resolution palmprint image contains clear ridge patterns and minutiae points. In addition, the three dimensional (3-D) palmprint image possesses curvatures of the palmprint surface. So far, there is no work to summarize the feature extraction of different types of palmprint images. In this paper, we have an aim to completely study the feature extraction and recognition of palmprint. We propose to use a unified framework to classify palmprint images into four categories: (1) the contact-based; (2) contactless; (3) high-resolution; and (4) 3-D palmprint images. Then, we analyze the motivations and theories of the representative extraction and matching methods for different types of palmprint images. Finally, we compare and test the state-of-the-art methods via the widely used palmprint databases, and point out some potential directions for future research. Lunke Fei, Guangming Lu 0002, Wei Jia 0001, Shaohua Teng, David Zhang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2018 | AAR-CNNs: Auto Adaptive Regularized Convolutional Neural NetworksabstractIn order to address the overfitting problem caused by the small or simple training datasets and the large model’s size in Convolutional Neural Networks (CNNs), a novel Auto Adaptive Regularization (AAR) method is proposed in this paper. The relevant networks can be called AAR-CNNs. AAR is the first method using the “abstraction extent” (predicted by AE net) and a tiny learnable module (SE net) to auto adaptively predict more accurate and individualized regularization information. The AAR module can be directly inserted into every stage of any popular networks and trained end to end to improve the networks’ flexibility. This method can not only regularize the network at both the forward and the backward processes in the training phase, but also regularize the network on a more refined level (channel or pixel level) depending on the abstraction extent’s form. Comparative experiments are performed on low resolution ImageNet, CIFAR and SVHN datasets. Experimental results show that the AAR-CNNs can achieve state-of-the-art performances on these datasets. Yao Lu 0008, Guangming Lu 0002, Yuanrong Xu, Bob Zhang 0001 |
IJCAI | 2 |
| 2018 | Shared Linear Encoder-based Gaussian Process Latent Variable Model for Visual ClassificationabstractMulti-view learning has shown its powerful potential in many applications and achieved outstanding performances compared with the single-view based methods. In this paper, we propose a novel multi-view learning model based on the Gaussian Process Latent Variable Model (GPLVM) to learn a shared latent variable in the manifold space with a linear and gaussian process prior based back projection. Different from existing GPLVM methods which only consider a mapping from the latent space to the observed space, the proposed method simultaneously takes a back projection from the observation to the latent variable into account. Concretely, due to the various dimensions of different views, a projection for each view is first learned to linearly map its observation to a subspace. The gaussian process prior is then imposed on another transformation to non-linearly and efficiently map the learned subspace to a shared manifold space. In order to apply the proposed approach to the classification, a discriminative regularization is also embedded to exploit the label information. Experimental results on three real-world databases substantiate the effectiveness and superiority of the proposed approach as compared with several state-of-the-art approaches. Jinxing Li 0003, Bob Zhang 0001, Guangming Lu 0002, David Zhang 0001 |
ACM Multimedia | 3 |
| 2018 | Pyramidal Combination of Separable Branches for Deep Short Connected Neural Networks
Yao Lu 0008, Guangming Lu 0002 |
PRCV (2) | 2 |
| 2018 | Learning acoustic features to detect Parkinson's disease
Kebin Wu, David Zhang 0001, Guangming Lu 0002, Zhenhua Guo 0001 |
Neurocomputing | 3 |
| 2018 | Facial beauty analysis based on features prediction and beautification models
Bob Zhang 0001, Xihua Xiao, Guangming Lu 0002 |
Pattern Anal. Appl. | 3 |
| 2017 | Fast pore matching method based on deterministic annealing algorithmabstractHigh‐resolution fingerprint identification system (HRFIS) has become a hot topic in the field of academic research. Compared to traditional automatic fingerprint identification system, HRFIS reduces the risk of being faked by using level 3 features, such as pores, which cannot be detected in lower resolution images. However, there is a serious problem in HRFIS: there are hundreds of sweat pores in one fingerprint image, which will spend a considerable amount of time for direct fingerprint matching. The authors propose a method to match pores in two fingerprint images based on deterministic annealing algorithm. In this method, fingerprints are aligned using singular points. Then minutiae are matched based on the alignment result. To reduce the impact of deformation, they build a convex hull for each of these fingerprints. Pores in these convex hulls are used for matching. In the experiments, their method is compared with random sample consensus method, minutia and ICP‐based method, and direct pore matching method. The results show that the proposed method is more efficient. Guangming Lu 0002, Yuanrong Xu |
IET Image Process. | 1 |
| 2017 | Generalized Feature Extraction for Wrist Pulse Analysis: From 1-D Time Series to 2-D MatrixabstractTraditional Chinese pulse diagnosis, known as an empirical science, depends on the subjective experience. Inconsistent diagnostic results may be obtained among different practitioners. A scientific way of studying the pulse should be to analyze the objectified wrist pulse waveforms. In recent years, many pulse acquisition platforms have been developed with the advances in sensor and computer technology. And the pulse diagnosis using pattern recognition theories is also increasingly attracting attentions. Though many literatures on pulse feature extraction have been published, they just handle the pulse signals as simple 1-D time series and ignore the information within the class. This paper presents a generalized method of pulse feature extraction, extending the feature dimension from 1-D time series to 2-D matrix. The conventional wrist pulse features correspond to a particular case of the generalized models. The proposed method is validated through pattern classification on actual pulse records. Both quantitative and qualitative results relative to the 1-D pulse features are given through diabetes diagnosis. The experimental results show that the generalized 2-D matrix feature is effective in extracting both the periodic and nonperiodic information. And it is practical for wrist pulse analysis. Dimin Wang, David Zhang 0001, Guangming Lu 0002 |
IEEE J. Biomed. Health Informatics | 3 |
| 2017 | An Adaptive Background Modeling Method for Foreground SegmentationabstractBackground modeling has played an important role in detecting the foreground for video analysis. In this paper, we presented a novel background modeling method for foreground segmentation. The innovations of the proposed method lie in the joint usage of the pixel-based adaptive segmentation method and the background updating strategy, which is performed in both pixel and object levels. Current pixel-based adaptive segmentation method only updates the background at the pixel level and does not take into account the physical changes of the object, which may result in a series of problems in foreground detection, e.g., a static or low-speed object is updated too fast or merely a partial foreground region is properly detected. To avoid these deficiencies, we used a counter to place the foreground pixels into two categories (illumination and object). The proposed method extracted a correct foreground object by controlling the updating time of the pixels belonging to an object or an illumination region respectively. Extensive experiments showed that our method is more competitive than the state-of-the-art foreground detection methods, particularly in the intermittent object motion scenario. Moreover, we also analyzed the efficiency of our method in different situations to show that the proposed method is available for real-time applications. Zuofeng Zhong, Bob Zhang 0001, Guangming Lu 0002, Yong Zhao 0010, Yong Xu 0001 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2017 | Door Knob Hand Recognition SystemabstractBiometric applications have been used globally in everyday life. However, conventional biometrics is created and optimized for high-security scenarios. Being used in daily life by ordinary untrained people is a new challenge. Facing this challenge, designing a biometric system with prior constraints of ergonomics, we propose ergonomic biometrics design model, which attains the physiological factors, the psychological factors, and the conventional security characteristics. With this model, a novel hand-based biometric system, door knob hand recognition system (DKHRS), is proposed. DKHRS has the identical appearance of a conventional door knob, which is an optimum solution in both physiological factors and psychological factors. In this system, a hand image is captured by door knob imaging scheme, which is a tailored omnivision imaging structure and is optimized for this predetermined door knob appearance. Then features are extracted by local Gabor binary pattern histogram sequence method and classified by projective dictionary pair learning. In the experiment on a large data set including 12 000 images from 200 people, the proposed system achieves competitive recognition performance comparing with conventional biometrics like face and fingerprint recognition systems, with an equal error rate of 0.091%. This paper shows that a biometric system could be built with a reliable recognition performance under the ergonomic constraints. Xiaofeng Qu, David Zhang 0001, Guangming Lu 0002, Zhenhua Guo 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2016 | Fingerprint Pores Matching based on Improved Deterministic Annealing Algorithm
Guangming Lu 0002, Yuanrong Xu, Liying Ye |
ICPRAM | 1 |
| 2016 | iPEEH: Improving pitch estimation by enhancing harmonics
Kebin Wu, David Zhang 0001, Guangming Lu 0002 |
Expert Syst. Appl. | 3 |
| 2016 | Approximately symmetrical face images for image preprocessing in face recognition and sparse representation based classification
Yong Xu 0001, Zheng Zhang 0006, Guangming Lu 0002, Jian Yang 0003 |
Pattern Recognit. | 3 |
| 2016 | An Optimal Pulse System Design by Multichannel Sensors FusionabstractPulse diagnosis, recognized as an important branch of traditional Chinese medicine (TCM), has a long history for health diagnosis. Certain features in the pulse are known to be related with the physiological status, which have been identified as biomarkers. In recent years, an electronic equipment is designed to obtain the valuable information inside pulse. Single-point pulse acquisition platform has the benefit of low cost and flexibility, but is time consuming in operation and not standardized in pulse location. The pulse system with a single-type sensor is easy to implement, but is limited in extracting sufficient pulse information. This paper proposes a novel system with optimal design that is special for pulse diagnosis. We combine a pressure sensor with a photoelectric sensor array to make a multichannel sensor fusion structure. Then, the optimal pulse signal processing methods and sensor fusion strategy are introduced for the feature extraction. Finally, the developed optimal pulse system and methods are tested on pulse database acquired from the healthy subjects and the patients known to be afflicted with diabetes. The experimental results indicate that the classification accuracy is increased significantly under the optimal design and also demonstrate that the developed pulse system with multichannel sensors fusion is more effective than the previous pulse acquisition platforms. Dimin Wang, David Zhang 0001, Guangming Lu 0002 |
IEEE J. Biomed. Health Informatics | 3 |
| 2016 | A Novel Line-Scan Palmprint Acquisition SystemabstractBiometric recognition systems have been widely used globally. However, one effective and highly accurate biometric authentication method, palmprint recognition, has not been popularly applied as it should have been, which could be due to the lack of small, flexible and user-friendly acquisition systems. To expand the use of palmprint biometrics, we propose a novel palmprint acquisition system based on the line-scan image sensor. The proposed system consists of a customized and highly integrated line-scan sensor, a self-adaptive synchronizing unit, and a field-programmable gate array controller with a cross-platform interface. The volume of the proposed system is over 94% smaller than the volume of existing palmprint systems, without compromising its verification performance. The verification performance of the proposed system was tested on a database of 8000 samples collected from 250 people, and the equal error rate is 0.048%, which is comparable to the best area camera-based systems. Xiaofeng Qu, David Zhang 0001, Guangming Lu 0002 |
IEEE Trans. Syst. Man Cybern. Syst. | 3 |
| 2012 | A Novel 3-D Palmprint Acquisition SystemabstractPalmprints have been widely studied for personal authentication because they are highly accurate and incur low costs. Most of the previous work has focused on two-dimensional (2-D) palmprint identification. However, the inner surfaces of palms contain not only texture information but also shape information. Unfortunately, 2-D palmprint systems lose the shape information when capturing palmprint images. Hence, three-dimensional (3-D) information is important for palmprint systems. In this paper, we have designed and developed a novel 3-D palmprint acquisition system based on structured-light imaging technology. The acquisition system can obtain 3-D palmprint information and, at the same time, the corresponding 2-D texture, which are used for personal authentication. A 3-D palmprint database that contains 8000 samples has been established by using the developed acquisition system, and the test results illustrate the effectiveness of our system. Wei Li 0016, David Zhang 0001, Guangming Lu 0002, Nan Luo |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 2011 | Online joint palmprint and palmvein verification
David Zhang 0001, Zhenhua Guo 0001, Guangming Lu 0002, Lei Zhang 0006, Wangmeng Zuo |
Expert Syst. Appl. | 3 |
| 2011 | Empirical study of light source selection for palmprint recognition
Zhenhua Guo 0001, David Zhang 0001, Lei Zhang 0006, Wangmeng Zuo, Guangming Lu 0002 |
Pattern Recognit. Lett. | 5 |
| 2011 | 3-D Palmprint Recognition With Joint Line and Orientation Featuresabstract2-D palmprint has been recognized as an effective biometric identifier in the past decade. Recently, 3-D palmprint recognition was proposed to further improve the performance of palmprint systems. This paper presents a simple yet efficient scheme for 3-D palmprint recognition. After calculating and enhancing the mean-curvature image of the 3-D palmprint data, we extract both line and orientation features from it. The two types of features are then fused at either score level or feature level for the final 3-D palmprint recognition. The experiments on The Hong Kong Polytechnic University 3-D palmprint database, which contains 8000 samples from 400 palms show that the proposed feature extraction and fusion methods lead to promising performance. Wei Li 0016, David Zhang 0001, Lei Zhang 0006, Guangming Lu 0002, Jingqi Yan |
IEEE Trans. Syst. Man Cybern. Part C | 4 |
| 2010 | Efficient joint 2D and 3D palmprint matching with alignment refinementabstractPalmprint verification is a relatively new but promising personal authentication technique for its high accuracy and fast matching speed. Two dimensional (2D) palmprint recognition has been well studied in the past decade, and recently three dimensional (3D) palmprint recognition techniques were also proposed. The 2D and 3D palmprint data can be captured simultaneously and they provide different and complementary information. 3D palmprint contains the depth information of the palm surface, while 2D palmprint contains plenty of textures. How to efficiently extract and fuse the 2D and 3D palmprint features to improve the recognition performance is a critical issue for practical palmprint systems. In this paper, an efficient joint 2D and 3D palmprint matching scheme is proposed. The principal line features and palm shape features are extracted and used to accurately align the palmprint, and a couple of matching rules are defined to efficiently use the 2D and 3D features for recognition. The experiments on a 2D+3D palmprint database which contains 8000 samples show that the proposed scheme can greatly improve the performance of palmprint verification. Wei Li 0016, Lei Zhang 0006, David Zhang 0001, Guangming Lu 0002, Jingqi Yan |
CVPR | 4 |
| 2009 | Palmprint Recognition Using 3-D InformationabstractPalmprint has proved to be one of the most unique and stable biometric characteristics. Almost all the current palmprint recognition techniques capture the 2-D image of the palm surface and use it for feature extraction and matching. Although 2-D palmprint recognition can achieve high accuracy, the 2-D palmprint images can be counterfeited easily and much 3-D depth information is lost in the imaging process. This paper explores a 3-D palmprint recognition approach by exploiting the 3-D structural information of the palm surface. The structured light imaging is used to acquire the 3-D palmprint data, from which several types of unique features, including mean curvature image, Gaussian curvature image, and surface type, are extracted. A fast feature matching and score-level fusion strategy are proposed for palmprint matching and classification. With the established 3-D palmprint database, a series of verification and identification experiments is conducted to evaluate the proposed method. The results demonstrate that 3-D palmprint technique has high recognition performance. Although its recognition rate is a little lower than 2-D palmprint recognition, 3-D palmprint recognition has higher anticounterfeiting capability and is more robust to illumination variations and serious scrabbling in the palm surface. Meanwhile, by fusing the 2-D and 3-D palmprint information, much higher recognition rate can be achieved. David Zhang 0001, Guangming Lu 0002, Wei Li 0016, Lei Zhang 0006, Nan Luo |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2006 | A study of identical twins' palmprints for personal verification
Adams Wai-Kin Kong, David Zhang 0001, Guangming Lu 0002 |
Pattern Recognit. | 3 |
| 2005 | Online Palmprint Identification System for Civil Applications
David Zhang 0001, Guangming Lu 0002, Adams Wai-Kin Kong |
J. Comput. Sci. Technol. | 2 |
| 2003 | Palmprint recognition using eigenpalms features
Guangming Lu 0002, David Zhang 0001, Kuanquan Wang |
Pattern Recognit. Lett. | 1 |