Lianghua He

dblp:24/2365 · DBLP profile ↗
← Back
65ranked-venue papers
5as first author
38since 2021 · last 2026
0000-0002-5250-170XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 3 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 21 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 7 since 2021Computer networks · 9 · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 3 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Dynamic MAsk-Pruning Strategy for Source-Free Model Intellectual Property Protection
Boyang Peng, Sanqing Qu, Yong Wu 0007, Tianpei Zou, Lianghua He, Alois C. Knoll, Guang Chen 0001, Changjun Jiang 0002
Int. J. Comput. Vis.5
2026 SAM foundation model and expert model cross prompting framework for semi-supervised medical image segmentation
Zerong Zhang, Lianghua He
J. Vis. Commun. Image Represent.2
2026 HSENet: Hierarchical semantic-enriched network for multi-modal image fusion
Rui Ming, Songlin Du, Lianghua He, Guobao Xiao
Pattern Recognit.4
2026 Spatially Aware Adaptive Diffusion: Unifying Low-Resolution Image Fusion and Super-Resolution
abstract
Low-resolution visible-infrared image fusion and super-resolution (LRVIF) are critical for enhancing image quality in low-resolution scenarios, yet limited information in the input images often constrains performance. To address these challenges, we propose SaDiff, a spatially-aware adaptive diffusion model that introduces diffusion processes into LRVIF for the first time, representing a major breakthrough in the field. Leveraging the generative capabilities of diffusion models, our approach unifies and enhances image fusion and super-resolution within a cohesive framework. A key component of SaDiff is the Spatial Residual Adaptation Block, which extends the diffusion process by dynamically adapting feature representations to spatial variations in the local regions of the input images. This module maximally preserves crucial information from the input images, such as texture details and contrast, while effectively suppressing noise, ensuring robust and context-aware feature refinement. Then we further propose Direct Diffusion Synthesis, a novel mechanism that utilizes noise predictions during diffusion to generate fused images, enabling joint training of the fusion and super-resolution networks. Additionally, a Cross-Feature Fusion Module integrates texture and contrast details, producing super-resolution fused images with improved clarity and structural integrity. Extensive experiments show that SaDiff achieves state-of-the-art performance, offering a robust and unified solution to infrared-visible image fusion and super-resolution. The code for the proposed method will be made available at https://github.com/guobaoxiao/SaDiff.
Jiajia Fu, Zhenni Yu, Haosheng Chen 0001, Songlin Du, Changcai Yang, Lianghua He, Guobao Xiao
IEEE Trans. Circuits Syst. Video Technol.6
2026 A General Framework for Efficient Medical Image Analysis via Shared Attention Vision Transformer
abstract
Vision Transformers (ViTs) demonstrate significant promise in medical image analysis but face two critical challenges: 1) their limited ability to capture local features in data-scarce scenarios, leading to data inefficiency, and 2) their high computational and storage demands of the full fine-tuning process in transfer learning, resulting in parameter inefficiency. To achieve efficient and accurate medical image analysis, we propose Shared Attention Vision Transformer (SAViT) that comprises three innovative modules: i) Shared Prior Attention (SPA) that enhances data efficiency by innovatively employing a visual prompt to sequentially share consistent attention weights across local image regions, thereby enabling the learning of translational invariance to capture locality; ii) MixPool that preserves global modeling ability by aggregating local features after SPA through a multi-pooling mechanism, thus effectively facilitating long-range dependency across local image regions; and iii) Low-rank Multi-head Self-Attention (Lr-MSA) that improves parameter efficiency by using low-rank weights of multi-head self-attention, hence reducing computational complexity while maintaining accuracy in medical image analysis. SAViT demonstrates strong generalization across multiple medical imaging modalities, including retinopathy, dermoscopy, and radiography. Extensive experiments are conducted. The results indicate its high data efficiency and outstanding performance in comparison with more than 20 medical-specific and ViT-based models when all of them are trained from scratch. It excels in parameter-efficient tuning by surpassing 17 models across 6 datasets in transfer learning, with only ${0}.{17}$ M/ ${0}.{23}$ M trainable parameters on ViT-B/SwinViT-B backbones requiring ${86}.{60}$ M/ ${88}.{00}$ M parameters. Source code can be found at: https://github.com/LYH-hh/SAViT.
Ying Wen 0003, Longzhen Yang, Lianghua He, MengChu Zhou
IEEE Trans. Medical Imaging4
2025 Enhancing Generalized Few-Shot Semantic Segmentation via Effective Knowledge Transfer
abstract
Generalized few-shot semantic segmentation (GFSS) aims to segment objects of both base and novel classes, using sufficient samples of base classes and few samples of novel classes. Representative GFSS approaches typically employ a two-phase training scheme, involving base class pre-training followed by novel class fine-tuning, to learn the classifiers for base and novel classes respectively. Nevertheless, distribution gap exists between base and novel classes in this process. To narrow this gap, we exploit effective knowledge transfer from base to novel classes. First, a novel prototype modulation module is designed to modulate novel class prototypes by exploiting the correlations between base and novel classes. Second, a novel classifier calibration module is proposed to calibrate the weight distribution of the novel classifier according to that of the base classifier. Furthermore, existing GFSS approaches suffer from a lack of contextual information for novel classes due to their limited samples, we thereby introduce a context consistency learning scheme to transfer the contextual knowledge from base to novel classes. Extensive experiments on PASCAL-5i and COCO-20i demonstrate that our approach significantly enhances the state of the art in the GFSS setting.
Xinyue Chen 0009, Miaojing Shi, Zijian Zhou 0002, Lianghua He, Sophia Tsoka
AAAI4
2025 M²RL-Net: Multi-View and Multi-Level Relation Learning Network for Weakly-Supervised Image Forgery Detection
abstract
As digital media manipulation becomes increasingly sophisticated, accurately detecting and localizing image forgeries with minimal supervision has become a critical challenge. Existing weakly supervised image forgery detection (W-IFD) methods often rely on convolutional neural networks (CNNs) and limited exploration of internal relationships, leading to poor detection and localization performance with only image-level labels. To address these limitations, we introduce a novel Multi-View and Multi-Level Relation Learning Network (M²RL-Net) for W-IFD. M²RL-Net effectively identifies forged images using only image-level annotations by exploring relationships between different views and hierarchical levels within images. Specifically, M²RL-Net achieves patch-level self-consistency learning (PSL) and feature-level contrastive learning (FCL) across different views, facilitating more generalized self-supervised learning of forgery features. In detail, PSL employs self-supervised learning to distinguish consistent and inconsistent regions within images, enhancing its ability to accurately locate tampered areas. FCL utilizes feature-level self-view and multi-view contrastive learning to differentiate between genuine and tampered image features, thereby improving the recognition of authentic and manipulated content across different views. Extensive experiments on various datasets demonstrate that M²RL-Net outperforms existing weakly-supervised methods in both detection and localization accuracy. This research sets a new benchmark for weakly-supervised image forgery detection and lays a robust foundation for future studies in this field.
Jiafeng Li 0005, Ying Wen 0003, Lianghua He
AAAI3
2025 AFiRe: Anatomy-Driven Self-Supervised Learning for Fine-Grained Representation in Radiographic Images
abstract
Current self-supervised methods, such as contrastive learning, predominantly focus on global discrimination, neglecting the critical fine-grained anatomical details required for accurate radiographic analysis. To address this challenge, we propose the Anatomy-driven self-supervised framework for enhancing Fine-grained Representation in radiographic image analysis (AFiRe). The core idea of AFiRe is to align the anatomical consistency with the unique token-processing characteristics of Vision Transformer. Specifically, AFiRe synergistically performs two self-supervised schemes: (i) Token-wise anatomy-guided contrastive learning, which aligns image tokens based on structural and categorical consistency to enhance fine-grained spatial-anatomical discrimination; (ii) Pixel-level anomaly-removal restoration, which particularly focuses on local anomalies, thereby refining the learned discrimination with detailed geometrical information. Additionally, we propose the Synthetic Lesion Mask to enhance anatomical diversity while preserving intra-consistency, which is typically corrupted by traditional data augmentations, such as Cropping and Affine transformations. Experimental results show that AFiRe: (i) provides robust anatomical discrimination, achieving more cohesive feature clusters compared to state-of-the-art contrastive learning methods; (ii) demonstrates superior generalization, surpassing 7 radiography-specific self-supervised methods in multi-label classification tasks with limited labeling; and (iii) integrates fine-grained information, enabling precise anomaly detection using only image-level annotations.
Lianghua He, Ying Wen 0003, Longzhen Yang, Hongzhou Chen
AAAI2
2025 CoSMIC: Continual Self-Supervised Learning for Multi-Domain Medical Imaging Via Conditional Mutual Information Maximization
Ying Wen 0003, Longzhen Yang, Lianghua He, Heng Tao Shen
ICCV4
2025 CFII-Net: Explicit Class Embeddings and Feature Maps Through Iterative Interaction for Boosting Medical Image Segmentation
abstract
Prior knowledge of category structure is essential in medical image segmentation, especially with significant organ structure differences. However, current hybrid architectures primarily focus on enhancing pixel-level representation learning, often neglecting or weakening the key prior knowledge of categorical structures, which poses challenges in capturing category relationships and accurate segmenting. To address this concern, we propose a novel network using Explicit Class Embeddings and Feature Maps through Iterative Interaction (CFII-Net) for boosting medical image segmentation. CFII-Net effectively segments images by exploring the relationship between explicit class embeddings and pixels in images. Specifically, we propose an Explicit Class Embedding Generator (ECEG) to obtain high-quality class semantic embeddings, incorporating category structure priors, which are used to guide high-accuracy segmentation. We then introduce an iterative Interactor, which utilizes transformers to facilitate the interaction between feature maps and class embeddings, thereby exploring pixel-to-class relationships. Furthermore, we propose updating strategies to refine the class embeddings and feature maps during the iteration process for achieving refined image segmentation. Extensive empirical evidence shows that any codec can be easily integrated into CFII-Net and yields improvements over the state-of-the-art methods in four public benchmarks.
Lianghua He
IJCAI3
2025 RadLAS: A Foundation Model for Interpretable Radiography Image Analysis with Lesion-Aware Self-Supervised Pre-training
abstract
Medical Foundation Models (MFMs) are revolutionizing radiography image analysis with scalable and generalized diagnostic capabilities. However, their effectiveness in real-world clinical practice is limited due to insufficient interpretability. To address this limitation, we propose RadLAS, a novel MFM for interpretable Radiographic image analysis by introducing Lesion-Aware Self-supervised pre-training. Unlike conventional MFMs that rely on post-hoc explanations, RadLAS innovates by directly emulating human diagnostic reasoning to first grounding lesion evidence and then making decisions accordingly. Specifically, RadLAS introduces two self-supervised tasks: (I) Lesion-grounded Reconstruction, which learns structured anatomical representations by restoring lesion-aware image patches into their healthy counterparts, thereby facilitating pixel-level grounding of lesion evidence via input-normal contrast. (II) Lesion-discrimination Contrastive Learning, which enhances lesion-aware pattern in representations by explicitly decoupling grounded lesion evidence as clinical cues and aligning them with global semantics, thereby enabling direct lesion-oriented diagnosis while preserving global context. RadLAS demonstrates excellent performance across diverse downstream radiographic datasets, offering verifiable explanations by deriving specific diagnoses (Task II) based on grounded lesion evidence (Task I), while preserving generalized representations essential for high diagnostic accuracy. Extensive experiments demonstrate that RadLAS (i) achieves superior interpretability with highly correlated lesion prediction and localization, surpassing 11 interpretable medical models; (ii) delivers scalable representation learning, outperforming 14 SOTA supervised and self-supervised MFMs.
Ying Wen 0003, Longzhen Yang, Lianghua He, Heng Tao Shen
ACM Multimedia4
2025 Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation
abstract
Vision-grounded medical report generation aims to produce clinically accurate descriptions of medical images, anchored in explicit visual evidence to improve interpretability and facilitate integration into clinical workflows. However, existing methods often rely on separately trained detection modules that require extensive expert annotations, introducing high labeling costs and limiting generalizability due to pathology distribution bias across datasets. To address these challenges, we propose Self-Supervised Anatomical Consistency Learning (SS-ACL)-a novel and annotation-free framework that aligns generated reports with corresponding anatomical regions using simple textual prompts. SS-ACL constructs a hierarchical anatomical graph inspired by the invariant top-down inclusion structure of human anatomy, organizing entities by spatial location. It recursively reconstructs fine-grained anatomical regions to enforce intra-sample spatial alignment, inherently guiding attention maps toward visually relevant areas prompted by text. To further enhance inter-sample semantic alignment for abnormality recognition, SS-ACL introduces a region-level contrastive learning based on anatomical consistency. These aligned embeddings serve as priors for report generation, enabling attention maps to provide interpretable visual evidence. Extensive experiments demonstrate that SS-ACL, without relying on expert annotations, (i) generates accurate and visually grounded reports-outperforming state-of-the-art methods by 10% in lexical accuracy and 25% in clinical efficacy, and (ii) achieves competitive performance on various downstream visual tasks, surpassing current leading visual foundation models by 8% in zero-shot visual grounding. Our code is available at https://github.com/kaelsunkiller/ssacl.
Longzhen Yang, Zhangkai Ni, Ying Wen 0003, Lianghua He, Heng Tao Shen
ACM Multimedia5
2025 Joint Objective and Subjective Fuzziness Denoising for Multimodal Sentiment Analysis
abstract
Multimodal Sentiment Analysis (MSA) aims at teaching computers or robotics to understand human sentiment with diverse multimodal signals, including audio, vision, and text. Current MSA approaches primarily concentrate on devising fusion strategies for multimodal signals and trying to learn better multimodal joint representations. However, employing multimodal signals directly is not appropriate since the human psychological states are fuzzy and can not be categorized easily, which undermines the effectiveness of existing methods. In this paper, we regard the natural fuzziness of human sentiments can be observed as two types: objective fuzziness introduced by human expression and subjective fuzziness caused by the complexity of human affection. Based on the assumption, we proposed a novel method termedJoint Objective and Subjective Fuzziness Denoising (JOSFD), which introduced fuzzy logic into the multimodal fusion process and sentiment decision process to overcome the objective and subjective fuzziness. Specifically, our JOSFD method contains two key modules: (1) Modality-Specific Fuzzification Module leveraging uncertainty estimation and fuzzy logic to overcome the influence of objective fuzziness in different modalities in multimodal fusion. (2) Attitude-Intensity Representation Disentangling that learns joint representations for human attitude and sentiment strength separately and further employs fuzzy logic to decide the sentiment analysis results. We evaluate our proposed JOSFD method on three widely used MSA benchmark datasets, CMU-MOSI, CMU-MOSEI, and CH-SIMS. Extensive experiments demonstrate our proposed JOSFD method outperforms recent state-of-the-art methods.
Xun Jiang 0001, Xing Xu 0001, Huimin Lu 0001, Lianghua He, Heng Tao Shen
IEEE Trans. Fuzzy Syst.4
2025 Rethinking Artifact Mitigation in HDR Reconstruction: From Detection to Optimization
abstract
Artifact remains a long-standing challenge in High Dynamic Range (HDR) reconstruction. Existing methods focus on model designs for artifact mitigation but ignore explicit detection and suppression strategies. Because artifact lacks clear boundaries, distinct shapes, and semantic consistency, and there is no existing dedicated dataset for HDR artifact, progress in direct artifact detection and recovery is impeded. To bridge the gap, we propose a unified HDR reconstruction framework that integrates artifact detection and model optimization. Firstly, we build the first HDR artifact dataset (HADataset), comprising 1,213 diverse multi-exposure Low Dynamic Range (LDR) image sets and 1,765 HDR image pairs with per-pixel artifact annotations. Secondly, we develop an effective HDR artifact detector (HADetector), a robust artifact detection model capable of accurately localizing HDR reconstruction artifact. HADetector plays two pivotal roles: (1) enhancing existing HDR reconstruction models through fine-tuning, and (2) serving as a non-reference image quality assessment (NR-IQA) metric, the Artifact Score (AS), which aligns closely with human visual perception for reliable quality evaluation. Extensive experiments validate the effectiveness and generalizability of our framework, including the HADataset, HADetector, fine-tuning paradigm, and AS metric. The code and datasets are available at: https://github.com/xinyueliii/hdr-artifact-detect-optimize.
Zhangkai Ni, Wenhan Yang, Hanli Wang, Lianghua He, Sam Kwong
IEEE Trans. Image Process.6
2025 MambaMatch: Establishing Reliable Correspondences via Multi-Scale State Space Model
abstract
Correspondence pruning aims to identify inliers from correspondences severely disturbed by outliers. Although Transformers and graph neural networks have shown impressive results in this field, they are either limited by a narrow receptive field or encounter quadratic computational complexity. To tackle this challenge, this work pioneers the integration of state space model into correspondence pruning task, proposing a Mamba-based framework named MambaMatch. Specifically, to address the limitations of the Mamba architecture in local consensus modeling, we proposes a multi-scale scanning strategy. It first employs an adaptive clustering algorithm to map origin correspondences into spatially coherent feature clusters, constructing a dual-representation space encompassing both full-scale and clustered-scale features. Bidirectional scan operations are then performed at both scales: 1) full-scale scan preserves global structural context, and 2) clustered-scale scan enhances local consistency. Subsequently, a Multi-Scale Interaction layer is designed to dynamically fuse dual-scale features via a cross-attention mechanism, further integrated with a Gated Feed-Forward Network to significantly improve the network's feature discrimination capability. Extensive experiments validate that MambaMatch surpasses state-of-the-art approaches across multiple benchmarks for two-view geometry estimation. Furthermore, MambaMatch exhibits robust generalization across diverse scenarios, tasks, and feature extractors. The source code is available at: https://github.com/mxyttkx/MambaMatch.
Xiangyang Miao, Shunxing Chen, Shiping Wang, Songlin Du, Lianghua He, Guobao Xiao
IEEE Trans. Image Process.6
2025 SFM-Net: Semantic Feature-Based Multi-Stage Network for Unsupervised Image Registration
abstract
It is difficult for general registration methods to establish the fine correspondence between images with complex anatomical structures. To overcome the above problem, this work presents SFM-Net, an unsupervised multi-stage semantic feature-based network. In addition to using the pixel-based similarity metrics, we propose a feature operator and emphasize a feature registration to improve the alignment of semantic related areas. Specifically, we design a two-stage training strategy, the intensity image registration stage and the semantic feature registration stage. The former is for valid semantic features learning and intensity-based coarse registration, while the latter is for semantic areas alignment, achieving fine transformation of anatomical structure. The same structure of both stages is composed of a dual-stream feature extraction module (DFEM) and a refined deformation field generation module (RDGM). Unlike the deep learning-based approaches that utilizing down-sampled encoder to extract features, DFEM constructed by dual-stream U-Net structure can capture semantic information in decoder feature for structural alignment. Different with approaches applying cascaded networks to learn deformation field, our proposed RDGM generates multi-scale deformation fields by performing a coarse-to-fine registration within a single network. Experiments on 3D brain MRI and liver CT datasets confirm that the proposed SFM-Net achieves accurate and diffeomorphic registration results, outperforming other state-of-the-art methods.
Tai Ma, Xinru Dai, Suwei Zhang, Haidong Zou, Lianghua He, Ying Wen 0003
IEEE J. Biomed. Health Informatics5
2025 FeaInfNet: Diagnosis of Medical Images With Feature-Driven Inference and Visual Explanations
abstract
Interpretable deep-learning models have received widespread attention in the field of image recognition. However, owing to the coexistence of medical-image categories and the challenge of identifying subtle decision-making regions, many proposed interpretable deep-learning models suffer from insufficient accuracy and interpretability in diagnosing images of medical diseases. Therefore, this study proposed a feature-driven inference network (FeaInfNet) that incorporates a feature-based network reasoning structure. Specifically, local feature masks (LFM) were developed to extract feature vectors, thereby providing global information for these vectors and enhancing the expressive ability of FeaInfNet. Second, FeaInfNet compares the similarity of the feature vector corresponding to each subregion image patch with the disease and normal prototype templates that may appear in the region. It then combines the comparison of each subregion when making the final diagnosis. This strategy simulates the diagnosis process of doctors, making the model interpretable during the reasoning process, while avoiding misleading results caused by the participation of normal areas during reasoning. Finally, we proposed adaptive dynamic masks (Adaptive-DM) to interpret feature vectors and prototypes into human-understandable image patches to provide an accurate visual interpretation. Extensive experiments on multiple publicly available medical datasets, including RSNA, iChallenge-PM, COVID-19, ChinaCXRSet, MontgomerySet, and CBIS-DDSM, demonstrated that our method achieves state-of-the-art classification accuracy and interpretability compared with baseline methods in the diagnosis of medical images. Additional ablation studies were performed to verify the effectiveness of each component.
Yitao Peng, Lianghua He, Die Hu 0002, Longzhen Yang, Shaohua Shang
IEEE J. Biomed. Health Informatics2
2025 Variational Transformer: A Framework Beyond the Tradeoff Between Accuracy and Diversity for Image Captioning
abstract
Accuracy and diversity represent two critical quantifiable performance metrics in the generation of natural and semantically accurate captions. While efforts are made to enhance one of them, the other suffers due to the inherent conflicting and complex relationship between them. In this study, we demonstrate that the suboptimal accuracy levels derived from human annotations are unsuitable for machine-generated captions. To boost diversity while maintaining high accuracy, we propose an innovative variational transformer (VaT) framework. By integrating "invisible information prior (IIP)" and "auto-selectable Gaussian mixture model (AGMM)," we enable its encoder to learn precise linguistic information and object relationships in various scenes, thus ensuring high accuracy. By incorporating the "range-median reward (RMR)" baseline into it, we preserve a wider range of candidates with higher rewards during the reinforcement-learning-based training process, thereby guaranteeing outstanding diversity. Experimental results indicate that our method achieves simultaneous improvements in accuracy and diversity by up to 1.1% and 4.8%, respectively, over the state-of-the-art. Furthermore, our approach demonstrates its performance that is the closest to human annotations in semantic retrieval, with its score of 50.3 versus the human score of 50.6. Thus, the method can be readily put into industrial use.
Longzhen Yang, Lianghua He, Die Hu 0002, Yitao Peng, Hongzhou Chen, MengChu Zhou
IEEE Trans. Neural Networks Learn. Syst.2
2025 R-HMF: A Relation-enhanced Hierarchical Multimodal Framework for Few-shot Knowledge Graph Completion
abstract
Knowledge graph completion (KGC) , which aims at inferring the missing fact triples, has shown an essential role in constructing a complete knowledge graph to enhance downstream applications. However, most KGC techniques require a large number of labeled training instances, and the performance drops dramatically when only a few triples are available. The primary challenge lies in the insufficient information that the few-shot annotated triples provided. Recently, several works have utilized multimodal entity contexts to enrich the entity representation, but their performance remains constrained by (1) overlooking the challenges of modality heterogeneity, (2) introducing the redundant multimodal noise of entities that is irrelevant to the corresponding relation, and (3) the difficulty in learning relation representation with only a few labeled cases. To address the above issues, we propose a novel Relation-enhanced Hierarchical Multimodal Framework (R-HMF) for the few-shot KGC. Specifically, to take the modality heterogeneity into account, we first conduct the modality-specific few-shot relation learning to capture the correlation between entities and relations within each modality. Subsequently, a multimodal fact assessment module is designed to validate the correctness of given triples by considering the entity contexts in a joint multimodal representation space. Notably, to avoid the involvement of redundant multimodal noise of entities, adaptive entity features are extracted to fit for different relations. In addition, to strengthen the relation representation in the few-shot setting, we also take full advantage of the large language models (LLMs) to generate the corresponding relation names and descriptions, which will be utilized to align different modalities as well. Extensive experimental results on two multimodal knowledge graph datasets, MM-FB15K237 and MM-DBpedia, show that our framework achieves better performance than previous state-of-the-art methods by improving 3.23% Hits@10 score under the 1-shot setting and 6.45% Hits@10 score under the 5-shot setting on average.
Chengmei Yang, Qian Li 0033, Chen Ma 0001, Lianghua He
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Hierarchical Salient Patch Identification for Interpretable Fundus Disease Localization
abstract
With the widespread application of deep learning technology in medical image analysis, the effective explanation of model predictions and improvement of diagnostic accuracy have become urgent problems that need to be solved. Attribution methods have become key tools to help doctors better understand the diagnostic basis of models, and are used to explain and localize diseases in medical images. However, previous methods suffer from inaccurate and incomplete localization problems for fundus diseases with complex and diverse structures. To solve these problems, we propose a weakly supervised interpretable fundus disease localization method called hierarchical salient patch identification (HSPI) that can achieve interpretable disease localization using only image-level labels and a neural network classifier (NNC). First, we propose salient patch identification (SPI), which divides the image into several patches and optimizes consistency loss to identify which patch in the input image is most important for the network’s prediction, in order to locate the disease. Second, we propose a hierarchical identification strategy to force SPI to analyze the importance of different areas to neural network classifier’s prediction to comprehensively locate disease areas. Conditional peak focusing is then introduced to ensure that the mask vector can accurately locate the disease area. Finally, we propose patch selection based on multi-sized intersections to filter out incorrectly or additionally identified non-disease regions. We conduct disease localization experiments on fundus image datasets and achieve the best performance on multiple evaluation metrics compared to previous interpretable attribution methods. Additional ablation studies are conducted to verify the effectiveness of each method.
Yitao Peng, Lianghua He, Die Hu 0002
BIBM2
2024 MAP: MAsk-Pruning for Source-Free Model Intellectual Property Protection
abstract
Deep learning has achieved remarkable progress in various applications, heightening the importance of safeguarding the intellectual property (IP) of well-trained models. It entails not only authorizing usage but also ensuring the deployment of models in authorized data domains, i.e., making models exclusive to certain target domains. Previous methods necessitate concurrent access to source training data and target unauthorized data when performing IP protection, making them risky and inefficient for decentralized private data. In this paper, we target a practical setting where only a well-trained source model is available and investigate how we can realize IP protection. To achieve this, we propose a novel MAsk Pruning (MAP) framework. MAP stems from an intuitive hypothesis, i.e., there are target-related parameters in a well-trained model, locating and pruning them is the key to IP protection. Technically, MAP freezes the source model and learns a target-specific binary mask to prevent unauthorized data usage while minimizing performance degradation on authorized data. Moreover, we introduce a new metric aimed at achieving a better balance between source and target performance degradation. To verify the effectiveness and versatility, we have evaluated MAP in a variety of scenarios, including vanilla source-available, practical source-free, and challenging data-free. Extensive experiments indicate that MAP yields new state-of-the-art performance. Code will be available at https://github.com/ispc-lab/MAP.
Boyang Peng, Sanqing Qu, Yong Wu 0007, Tianpei Zou, Lianghua He, Alois C. Knoll, Guang Chen 0001, Changjun Jiang 0002
CVPR5
2024 LEAD: Learning Decomposition for Source-free Universal Domain Adaptation
abstract
Universal Domain Adaptation (UniDA) targets knowledge transfer in the presence of both covariate and label shifts. Recently, Source-free Universal Domain Adaptation (SF-UniDA) has emerged to achieve UniDA without access to source data, which tends to be more practical due to data protection policies. The main challenge lies in determining whether covariate-shifted samples belong to target-private unknown categories. Existing methods tackle this either through hand-crafted thresholding or by developing time-consuming iterative clustering strategies. In this paper, we propose a new idea of LEArning Decomposition (LEAD), which decouples features into source-known and-unknown components to identify target-private data. Technically, LEAD initially leverages the or-thogonal decomposition analysis for feature decomposition. Then, LEAD builds instance-level decision boundaries to adaptively identify target-private data. Extensive experiments across various UniDA scenarios have demonstrated the effectiveness and superiority of LEAD. Notably, in the OPDA scenario on VisDA dataset, LEAD outperforms GLC by 3.5% overall H-score and reduces 75% time to derive pseudo-labeling decision boundaries. Besides, LEAD is also appealing in that it is complementary to most existing methods. The code is available at https://github.com/ispc-lab/LEAD.
Sanqing Qu, Tianpei Zou, Lianghua He, Florian Röhrbein, Alois C. Knoll, Guang Chen 0001, Changjun Jiang 0002
CVPR3
2024 HGL: Hierarchical Geometry Learning for Test-Time Adaptation in 3D Point Cloud Segmentation
Tianpei Zou, Sanqing Qu, Zhijun Li 0001, Alois C. Knoll, Lianghua He, Guang Chen 0001, Changjun Jiang 0002
ECCV (55)5
2024 PSAM: Prompt-based Segment Anything Model Adaption for Medical Image Segmentation
abstract
In the current landscape where large models are increasingly becoming the norm for task solving, maximizing the utilization of these models has emerged as a focal point of research. The Segment Anything Model (SAM), an eminent large-scale image segmentation model using a new task, model, and dataset, has garnered recognition for its efficacy across different scenarios. However, the effectiveness of SAM is hindered in the medical domain due to the scarcity of available medical images, leading to suboptimal training and inadequate adaptation of its feature extractor to medical imagery. In this work, we propose PSAM, which built upon SAM to explore a new research paradigm of customizing large-scale models to meet the demands of medical image segmentation tasks. This is achieved through a two-fold strategy: Firstly, we incorporate parallel feature extraction branches into SAM, guided by task-specific prompts derived from CLIP, enhancing its ability to extract relevant features. Secondly, we introduce an enhanced, visually task-friendly adapter mechanism, which effectively injects medical knowledge into SAM's ViT image encoder for facilitating adaptive task execution in medical image scenarios. Our experimental findings demonstrate the effectiveness of PSAM in accurately segmenting medical images, underscoring its potential as a valuable tool in the medical imaging domain.
Chengyi Wen, Lianghua He
SMC2
2024 Dual Contrastive Learning with Mutual Correction for Semi-Supervised Medical Image Segmentation
abstract
Semi-supervised learning has garnered considerable attention from researchers due to its capacity to utilize extensive amounts of unlabeled data, thus reducing the reliance of deep learning models on annotated datasets. However, in medical image segmentation, this method still encounters challenges such as suboptimal pseudo-labeling and insufficient feature extraction because of the confirmation bias problem induced by erroneously fitting unlabeled data. To tackle these challenges, we propose a novel approach that integrates correction modules and contrastive learning. First, our method exploits the difference in the output predictions from two different decoders and employs two rectification losses in the inconsistent regions for labeled and unlabeled data respectively, which mitigates the confirmation bias problem. Additionally, we incorporate two uncertainty-guided pixel-prototype contrastive learning modules, which are designed to perceive complete sample distribution information and optimize the features of pixels with low-uncertainty pseudo labels. Both modules complement each other and enable the encoder to generate class-discriminative features, thereby enhancing the final segmentation performance. Finally, extensive experiments are conducted on the two widely used medical image datasets to demonstrate the effectiveness of our method.
Jiazhe Zhu, Lianghua He
SMC2
2024 Hierarchical Dynamic Masks for Visual Explanation of Neural Networks
abstract
Despite the remarkable accomplishments of deep neural networks in computer vision tasks, the inherent opacity of their operations remains a pressing concern. Attribution methods generating visual explanatory maps representing the importance of image pixels for model classification are popular for explaining neural network decisions. However, the small and diverse decision regions in fine-grained or medical images limit the precision and comprehensiveness of the existing attribution methods when explaining decisions made for such a data type. This paper introduces a novel attribution method called hierarchical dynamic masks (HDM) to overcome these concerns to generate saliency maps with high recognition reliability and localization capability. Specifically, we suggest dynamic masks (DM), which enable multiple small-sized benchmark mask vectors to learn the image's critical information roughly through an optimization method. The benchmark mask vectors guide the learning of the large-sized combination mask vectors so that their overlay mask accurately learns detailed pixel importance information. Additionally, we construct the HDM by hierarchically concatenating DM modules. These DM modules search and combine the regions of interest in the remaining neural network classification decisions within the masked image in a learning-based way. Since HDM forces DM to perform importance analysis in different areas, it makes the fused saliency map more comprehensive. The experiments reveal that the proposed method outperforms existing approaches significantly regarding recognition credibility and positioning ability when qualitatively and quantitatively tested on CUB-200-2011 and iChallenge-PM datasets.
Yitao Peng, Lianghua He, Die Hu 0002, Longzhen Yang, Shaohua Shang
IEEE Trans. Multim.2
2024 Explainability of Speech Recognition Transformers via Gradient-Based Attention Visualization
abstract
In vision Transformers, attention visualization methods are used to generate heatmaps highlighting the class-corresponding areas in input images, which offers explanations on how the models make predictions. However, it is not so applicable for explaining automatic speech recognition (ASR) Transformers. An ASR Transformer makes a particular prediction for every input token to form a sentence, but a vision Transformer only makes an overall classification for the input data. Therefore, traditional attention visualization methods may fail in ASR Transformers. In this work, we propose a novel attention visualization method in ASR Transformers and try to explain which frames of the audio result in the output text. Inspired by the model explainability, we also explore ways of improving the effectiveness of the ASR model. Comparing with other Transformer attention visualization methods, our method is more efficient and intuitively understandable, which unravels the attention calculation from information flow of Transformer attention modules. In addition, we demonstrate the utilization of visualization result in three ways: (1) We visualize attention with respect to connectionist temporal classification (CTC) loss to train an ASR model with adversarial attention erasing regularization, which effectively decreases the word error rate (WER) of the model and improves its generalization capability. (2) We visualize the attention on some specific words, interpreting the model by effectively demonstrating the semantic and grammar relationships between these words. (3) Similarly, we analyze how the model manage to distinguish homophones, using contrastive explanation with respect to homophones.
Tianli Sun, Haonan Chen 0003, Guosheng Hu, Lianghua He, Cairong Zhao
IEEE Trans. Multim.4
2024 Decoupling Deep Learning for Enhanced Image Recognition Interpretability
abstract
The quest for enhancing the interpretability of neural networks has become a prominent focus in recent research endeavors. Prototype-based neural networks have emerged as a promising avenue for imbuing models with interpretability by gauging the similarity between image components and category prototypes to inform decision-making. However, these networks face challenges as they share similarity activations during both the inference and explanation processes, creating a tradeoff between accuracy and interpretability. To address this issue and ensure that a network achieves high accuracy and robust interpretability in the classification process, this article introduces a groundbreaking prototype-based neural network termed the “Decoupling Prototypical Network” (DProtoNet). This novel architecture comprises encoder, inference, and interpretation modules. In the encoder module, we introduce decoupling feature masks to facilitate the generation of feature vectors and prototypes, enhancing the generalization capabilities of the model. The inference module leverages these feature vectors and prototypes to make predictions based on similarity comparisons, thereby preserving an interpretable inference structure. Meanwhile, the interpretation module advances the field by presenting a novel approach: a “multiple dynamic masks decoder” that replaces conventional upsampling similarity activations. This decoder operates by perturbing images with mask vectors of varying sizes and learning saliency maps through consistent activation. This methodology offers a precise and innovative means of interpreting prototype-based networks. DProtoNet effectively separates the inference and explanation components within prototype-based networks. By eliminating the constraints imposed by shared similarity activations during the inference and explanation phases, our approach concurrently elevates accuracy and interpretability. Experimental evaluations on diverse public natural datasets, including CUB-200-2011, Stanford Cars, and medical datasets like RSNA and iChallenge-PM, corroborate the substantial enhancements achieved by our method compared to previous state-of-the-art approaches. Furthermore, ablation studies are conducted to provide additional evidence of the effectiveness of our proposed components.
Yitao Peng, Lianghua He, Die Hu 0002, Longzhen Yang, Shaohua Shang
ACM Trans. Multim. Comput. Commun. Appl.2
2023 SCConv: Spatial and Channel Reconstruction Convolution for Feature Redundancy
abstract
Convolutional Neural Networks (CNNs) have achieved remarkable performance in various computer vision tasks but this comes at the cost of tremendous computational resources, partly due to convolutional layers extracting redundant features. Recent works either compress well-trained large-scale models or explore well-designed lightweight models. In this paper, we make an attempt to exploit spatial and channel redundancy among features for CNN compression and propose an efficient convolution module, called SCConv (Spatial and Channel reconstruction Convolution), to decrease redundant computing and facilitate representative feature learning. The proposed SCConv consists of two units: spatial reconstruction unit (SRU) and channel reconstruction unit (CRU). SRU utilizes a separate-and-reconstruct method to suppress the spatial redundancy while CRU uses a split-transform-and-fuse strategy to diminish the channel redundancy. In addition, SCConv is a plug-and-play architectural unit that can be used to replace standard convolution in various convolutional neural networks directly. Experimental results show that SCConv-embedded models are able to achieve better performance by reducing redundant features with significantly lower complexity and computational costs.
Jiafeng Li 0005, Ying Wen 0003, Lianghua He
CVPR3
2023 Few-Shot Relational Triple Extraction Based on Evaluation of Token-Level Semantic Similarity
Jiazhe Zhu, Lianghua He
ICANN (8)3
2023 Mutually Guided Few-Shot Learning For Relational Triple Extraction
abstract
Knowledge graphs (KGs), containing many entity-relation-entity triples, provide rich information for downstream applications. Although extracting triples from unstructured texts has been widely explored, most of them require a large number of labeled instances. The performance will drop dramatically when only few labeled data are available. To tackle this problem, we propose the Mutually Guided Few-shot learning framework for Relational Triple Extraction (MG-FTE). Specifically, our method consists of an entity-guided relation proto-decoder to classify the relations firstly and a relation-guided entity proto-decoder to extract entities based on the classified relations. To draw the connection between entity and relation, we design a proto-level fusion module to boost the performance of both entity extraction and relation classification. Moreover, a new cross-domain few-shot triple extraction task is introduced. Extensive experiments show that our method outperforms many state-of-the-art methods by 12.6 F1 score on FewRel 1.0 (single-domain) and 20.5 F1 score on FewRel 2.0 (cross-domain).
Chengmei Yang, Bowei He, Chen Ma 0001, Lianghua He
ICASSP5
2023 EEG-Based Emotion Analysis Using Person-Event Network
abstract
Brain-computer interface (BCI) technology has attracted a lot of attention in recent years. Emotion recognition which based on electroencephalography is a typical application of BCI. Traditional methods on emotion recognition are mainly focusing on time domain feature and frequency domain feature while spatial information is often been ignored. In this paper, to make use of spatial feature, we propose a new convolutional neural network using not only temporal feature but also person related feature and event related feature. Depthwise convolution and separable convolution are also used for feature extraction. To verify the effectiveness of our method, we conduct extensive experiments on the public dataset DEAP and DREAMER. Compared with other methods, our method has achieved the state-of-the-art effect.
Liwei Tang, Lianghua He
SMC2
2023 MMEL: A Joint Learning Framework for Multi-Mention Entity Linking
abstract
Entity linking, bridging mentions in the contexts with their corresponding entities in the knowledge bases, has attracted wide attention due to many potential applications. Recently, plenty of multimodal entity linking approaches have been proposed to take full advantage of the visual information rather than solely the textual modality. Although feasible, these methods mainly focus on the single-mention scenarios and neglect the scenarios where multiple mentions exist simultaneously in the same context, which limits the performance. In fact, such multi-mention scenarios are pretty common in public datasets and real-world applications. To solve this challenge, we first propose a joint feature extraction module to learn the representations of context and entity candidates, from both the visual and textual perspectives. Then, we design a pairwise training scheme (for training) and a multi-mention collaborative ranking method (for testing) to model the potential connections between different mentions. We evaluate our method on a public dataset and a self-constructed dataset, NYTimes-MEL, under both text-only and multimodal scenarios. The experimental results demonstrate that our method can largely outperform the state-of-the-art methods, especially in multi-mention scenarios. Our dataset and source code are publicly available at https://github.com/ycm094/MMEL-main.
Chengmei Yang, Bowei He, Yimeng Wu, Lianghua He, Chen Ma 0001
UAI5
2023 A multi-grained unsupervised domain adaptation approach for semantic segmentation
Tai Ma, Yue Lu 0001, Qingli Li, Lianghua He, Ying Wen 0003
Pattern Recognit.5
2023 Learning Smooth Representation for Unsupervised Domain Adaptation
abstract
Typical adversarial-training-based unsupervised domain adaptation (UDA) methods are vulnerable when the source and target datasets are highly complex or exhibit a large discrepancy between their data distributions. Recently, several Lipschitz-constraint-based methods have been explored. The satisfaction of Lipschitz continuity guarantees a remarkable performance on a target domain. However, they lack a mathematical analysis of why a Lipschitz constraint is beneficial to UDA and usually perform poorly on large-scale datasets. In this article, we take the principle of utilizing a Lipschitz constraint further by discussing how it affects the error bound of UDA. A connection between them is built, and an illustration of how Lipschitzness reduces the error bound is presented. A local smooth discrepancy is defined to measure the Lipschitzness of a target distribution in a pointwise way. When constructing a deep end-to-end model, to ensure the effectiveness and stability of UDA, three critical factors are considered in our proposed optimization strategy, i.e., the sample amount of a target domain, dimension, and batchsize of samples. Experimental results demonstrate that our model performs well on several standard benchmarks. Our ablation study shows that the sample amount of a target domain, the dimension, and batchsize of samples, indeed, greatly impact Lipschitz-constraint-based methods' ability to handle large-scale datasets. Code is available at https://github.com/CuthbertCai/SRDA.
Guanyu Cai, Lianghua He, MengChu Zhou, Hesham Alhumade, Die Hu 0002
IEEE Trans. Neural Networks Learn. Syst.2
2021 Ask&Confirm: Active Detail Enriching for Cross-Modal Retrieval with Partial Query
abstract
Text-based image retrieval has seen considerable progress in recent years. However, the performance of existing methods suffers in real life since the user is likely to provide an incomplete description of an image, which often leads to results filled with false positives that fit the incomplete description. In this work, we introduce the partial-query problem and extensively analyze its influence on text-based image retrieval. Previous interactive methods tackle the problem by passively receiving users’ feedback to supplement the incomplete query iteratively, which is time-consuming and requires heavy user effort. Instead, we propose a novel retrieval framework that conducts the interactive process in an Ask-and-Confirm fashion, where AI actively searches for discriminative details missing in the current query, and users only need to confirm AI’s proposal. Specifically, we propose an object-based interaction to make the interactive retrieval more user-friendly and present a reinforcement-learning-based policy to search for discriminative objects. Furthermore, since fully-supervised training is often infeasible due to the difficulty of obtaining human-machine dialog data, we present a weakly-supervised training strategy that needs no human-annotated dialogs other than a text-image dataset. Experiments show that our framework significantly improves the performance of text-based image retrieval. Code is available at https://github.com/CuthbertCai/Ask-Confirm.
Guanyu Cai, Jun Zhang 0018, Xinyang Jiang, Yifei Gong, Lianghua He, Fufu Yu, Feiyue Huang, Xing Sun 0001
ICCV5
2021 Sparse Common Feature Representation for Undersampled Face Recognition
abstract
This work investigates the problem of undersampled face recognition (i.e., insufficient training data) encountered in practical Internet-of-Things (IoT) applications. Insufficient and uncertain samples captured by IoT devices may include background and facial disguise that makes face recognition more challenging than that with sufficient and reliable images. Many models work well in face recognition on a big data set, but when training data are insufficient, they achieve unsatisfactory performance. This work proposes a novel method named sparse common feature-based representation (SCFR) that provides a unique and stable result and completely avoids very time-consuming training required by a deep learning model. Specially, it constructs a common feature dictionary using both training and test images. Thereinto, a common feature is based on a discriminative common vector and learned by a Gaussian mixture model for both training and test images in a semisupervised learninig manner, which would reduce the difference among samples in each class. In the optimization, the latent indicator of test data is initialized by the estimated label. This can avoid learning invalid information and lead to good prototype images. A new variation dictionary characterizes variables that can be shared by different classes. Finally, this work adopts minimum reconstruction residuals to recognize test images, thus bringing about a substantial improvement in SCFR's performance. Extensive results on benchmark face databases demonstrate that the proposed method is better than the state-of-the-art methods handling undersampled face recognition.
Shicheng Yang, Ying Wen 0003, Lianghua He, MengChu Zhou
IEEE Internet Things J.3
2021 Sparse Individual Low-Rank Component Representation for Face Recognition in the IoT-Based System
abstract
The performance of face recognition has been greatly improved by deep neural network algorithms when a dataset is large. However, when face data are insufficient as in practical Internet of Things (IoT) applications and captured by IoT devices under the same intrasubject variation, both data quantity and quality bring big challenges to construct a model or representation, and most of the time it becomes infeasible to build a deep neural network model. This work proposes a sparse individual low-rank component-based representation (SILR) such that the representation of testing images can be based on individual subjects’ low-rank component. Theoretically, we put the$l_{2}$-norm constraint on intrasubject coefficients to represent testing images, thus making intrasubject coefficients dense. Hence, we alleviate the impact of an undersampled training dataset and its same intersubject variation on classification performance. We solve a convex minimization problem in polynomial time via an augmented lagrange multiplier scheme to get the solution of SILR. The scheme can reduce the influences from the same intersubject variation and contribute to an accurate recognition of the undersampled training dataset. We adopt sparse individual low-rank component representation and minimum reconstruction residual to recognize testing images. Extensive results on various databases show that SILR outperforms the other state-of-the-art methods for face recognition.
Shicheng Yang, Ying Wen 0003, Lianghua He, MengChu Zhou, Abdullah Abusorrah
IEEE Internet Things J.3
2020 Segmenting Medical MRI via Recurrent Decoding Cell
abstract
The encoder-decoder networks are commonly used in medical image segmentation due to their remarkable performance in hierarchical feature fusion. However, the expanding path for feature decoding and spatial recovery does not consider the long-term dependency when fusing feature maps from different layers, and the universal encoder-decoder network does not make full use of the multi-modality information to improve the network robustness especially for segmenting medical MRI. In this paper, we propose a novel feature fusion unit called Recurrent Decoding Cell (RDC) which leverages convolutional RNNs to memorize the long-term context information from the previous layers in the decoding phase. An encoder-decoder network, named Convolutional Recurrent Decoding Network (CRDN), is also proposed based on RDC for segmenting multi-modality medical MRI. CRDN adopts CNN backbone to encode image features and decode them hierarchically through a chain of RDCs to obtain the final high-resolution score map. The evaluation experiments on BrainWeb, MRBrainS and HVSMR datasets demonstrate that the introduction of RDC effectively improves the segmentation accuracy as well as reduces the model size, and the proposed CRDN owns its robustness to image noise and intensity non-uniformity in medical MRI.
Ying Wen 0003, Lianghua He
AAAI3
2020 Gabor Feature-Based LogDemons With Inertial Constraint for Nonrigid Image Registration
abstract
Nonrigid image registration plays an important role in the field of computer vision and medical application. The methods based on Demons algorithm for image registration usually use intensity difference as similarity criteria. However, intensity based methods can not preserve image texture details well and are limited by local minima. In order to solve these problems, we propose a Gabor feature based LogDemons registration method in this paper, called GFDemons. We extract Gabor features of the registered images to construct feature similarity metric since Gabor filters are suitable to extract image texture information. Furthermore, because of the weak gradients in some image regions, the update fields are too small to transform the moving image to the fixed image correctly. In order to compensate this deficiency, we propose an inertial constraint strategy based on GFDemons, named IGFDemons, using the previous update fields to provide guided information for the current update field. The inertial constraint strategy can further improve the performance of the proposed method in terms of accuracy and convergence. We conduct experiments on three different types of images and the results demonstrate that the proposed methods achieve better performance than some popular methods.
Ying Wen 0003, Yue Lu 0001, Qingli Li, Haibin Cai, Lianghua He
IEEE Trans. Image Process.6
2020 Unsupervised Domain Adaptation With Adversarial Residual Transform Networks
abstract
Domain adaptation (DA) is widely used in learning problems lacking labels. Recent studies show that deep adversarial DA models can make markable improvements in performance, which include symmetric and asymmetric architectures. However, the former has poor generalization ability, whereas the latter is very hard to train. In this article, we propose a novel adversarial DA method named adversarial residual transform networks (ARTNs) to improve the generalization ability, which directly transforms the source features into the space of target features. In this model, residual connections are used to share features and adversarial loss is reconstructed, thus making the model more generalized and easier to train. Moreover, a special regularization term is added to the loss function to alleviate a vanishing gradient problem, which enables its training process stable. A series of experiments based on Amazon review data set, digits data sets, and Office-31 image data sets are conducted to show that the proposed ARTN can be comparable with the methods of the state of the art.
Guanyu Cai, Lianghua He, MengChu Zhou
IEEE Trans. Neural Networks Learn. Syst.3
2019 Sparse Low-Rank Component-Based Representation for Face Recognition With Low-Quality Images
abstract
Sparse-representation-based classification (SRC) has been showing a good performance for face recognition in recent years. But SRC is not good at face recognition with low quality images (e.g., disguised, corrupted, occluded, and so on) which often appear in practical applications. To solve the problem, in this paper, we propose a novel SRC-based method for face recognition with low quality images named sparse low-rank component-based representation (SLCR). In SLCR, we utilize low-rank matrix recovery on the training data set to obtain low-rank components and non-low-rank components, which are used to construct the dictionary. The new dictionary is capable of describing facial features better, especially for low quality face samples. Furthermore, the minimum class-wise reconstruction residual is used as the recognition rule, leading to a substantial improvement on the proposed SLCR's performance. Extensive experiments on benchmark face databases demonstrate that the proposed method is consistently superior to other sparse-representation-based approaches for face recognition with low quality images.
Shicheng Yang, Lianghua He, Ying Wen 0003
IEEE Trans. Inf. Forensics Secur.3
2019 Incorporation of Structural Tensor and Driving Force Into Log-Demons for Large-Deformation Image Registration
abstract
Large-deformation image registration is important in theory and application in computer vision, but is a difficult task for non-rigid registration methods. In this paper, we propose a structural Tensor and Driving force-based Log-Demons algorithm for it, named TDLog-Demons for short. The structural tensor of an image is proposed to obtain a highly accurate deformation field. The driving force is proposed to solve the registration issue of large-deformation that often causes Log-Demons to trap into local minima. It is defined as a point correspondence obtained via multisupport-region-order-based gradient histogram descriptor matching on image's boundary points. It is integrated into an exponentially decreasing form with the velocity field of Log-Demons to move the points accurately and to speed up a registration process. Consequently, the driving force-based Log-Demons can well deal with large-deformation image registration. Extensive experiments demonstrate that the TDLog-Demons not only captures large deformations at a high accuracy but also yields a smooth deformation.
Ying Wen 0003, Lianghua He, MengChu Zhou
IEEE Trans. Image Process.3
2018 Sparse Low-Rank Component Coding for Face Recognition with Illumination And Corruption
abstract
Sparse representation-based classification shows a good performance for face recognition in recent years, but it can not be suitable for face recognition with illumination and corruption, which are often presented in the practical applications. To solve the problem, in this paper, we propose a novel SRC based method for face recognition named sparse low-rank component coding (SLC). In SLC, we utilize the low-rank component from training dataset to construct dictionary. The dictionary composed of low-rank component is able to describe the face feature better, especially for training samples with illumination and corruption. Our recognition rule is based on the minimum class-wise reconstruction residual which leads to a substantial improvement on the performance of SLC. Extensive experiments on benchmark face databases demonstrate that the proposed method consistently outperforms the other sparse representation based approaches for face recognition with illumination and corruption.
Shicheng Yang, Ying Wen 0003, Lianghua He
ICASSP3
2018 Effective Sample Synthesizing in Kernel Space for Imbalanced Classification
abstract
Imbalance data are common and result in major challenges when classifying big data. Researchers have found that synthetic samples in the kernel space are better than those in the original data space for imbalanced classification. Although previous kernel-based methods can be defective when generating samples far from the classification boundary, they have limited contributions to revising classification boundary. Therefore, a robust synthetic method is proposed in this paper to generate an effective sample based on three real samples with given constraints. Furthermore, we research how to determine the stopping criteria. The hyperplane is reconstructed on the condition of real samples and synthetic samples, and a series of experiments is designed to test the validity, robustness and effectiveness of the synthetic samples. All the experimental results show that the proposed method is feasible and effective, especially when the dataset is heavily imbalanced.
Wenwen Mo, Lianghua He
SMC2
2018 A Novel Forward-Link Multiplexed Scheme in Satellite-Based Internet of Things
abstract
Satellite communication has the potential to play a key role in many applications of Internet of Things (IoT). In this paper, we consider a satellite-based IoT and investigate the technology that can improve the spectral efficiency. In general, one beam in satellite systems serves one user. To serve multiple users, time division multiplexing or frequency division multiplexing is usually used. In this paper, we propose a novel forward-link multiplexed scheme, by which the signals of different users can be transmitted simultaneously using the same frequency band. Specifically, at the transmitter, we first map each combination of the users' constellation points to a higher-order constellation point, which is referred to as constellation coding, and then transmit such higher-order modulation signals. At the user side, after receiving and detecting the transmitted signal, each user obtain its own signal by the corresponding demapping, which is referred to as constellation decoding. The total system capacity over an additive white Gaussian noise channel is analyzed in this paper. Simulation results demonstrate that the proposed scheme can greatly improve the spectral efficiency.
Die Hu 0002, Lianghua He, Jun Wu 0006
IEEE Internet Things J.2
2017 Channel Estimation for FDD Massive MIMO OFDM Systems
abstract
For frequency-division duplexing (FDD) massive multiple-input multiple-output (MIMO) systems, it is challenging for the base station (BS) to acquire downlink channel state information (CSI) since the number of channel parameters is proportional to the number of BS antennas. To reduce the downlink training and the uplink feedback overhead for orthogonal frequency-division multiplexing (OFDM)-based FDD massive MIMO systems, we propose an effective and practical downlink channel estimation approach that does not require statistical information of the channels. In the proposed method, we parameterize the wideband MIMO channel by a limited number of distinct paths, each characterized by path delay, path angle and path gain. Based on this parametric channel model, we first estimate the reciprocal channel properties, i.e., the number of paths, path delays and path angles at the BS in uplink, and then estimate the non-reciprocal channel properties, i.e., the path gains at the user-side in downlink. A compressive sensing (CS)-based algorithm is proposed to obtain the estimates of the path number and the path delays, and based on these estimates, a low- complexity algorithm is presented to estimate the path angles. Simulation results demonstrate that the proposed method can provide reliable channel estimation for FDD massive MIMO OFDM systems.
Die Hu 0002, Lianghua He
VTC Fall2
2016 Motor imagery EEG signals analysis based on Bayesian network with Gaussian distribution
Lianghua He, Bin Liu 0018, Die Hu 0002, Ying Wen 0003, Meng Wan
Neurocomputing1
2016 Common Bayesian Network for Classification of EEG-Based Multiclass Motor Imagery BCI
abstract
Modeling and learning of brain activity patterns represent a huge challenge to the brain-computer interface (BCI) based on electroencephalography (EEG). Many existing methods estimate the uncorrelated instantaneous demixing of EEG signals to classify multiclass motor imagery (MI). However, the condition of uncorrelation does not hold true in practice, because the brain regions work with partial or complete collaboration. This work proposes a novel method, termed as a common Bayesian network (CBN), to discriminate multiclass MI EEG signals. First, with the constraints of a Gaussian mixture model on every channel, only related channels are selected to construct a normal Bayesian network. Second, the nodes that have both common and varying edges are selected to construct a CBN. Third, the probabilities on common edges are used to learn about the support vector machine for classification. To validate the proposed method, we conduct experiments on two well-known BCI datasets and perform a numerical analysis of the propose algorithm for EEG classification in a multiclass MI BCI. Experimental results show that the proposed CBN method not only has excellent classification performance, but also is highly efficient. Hence, it is suitable for the cases where a system is required to respond within a second.
Lianghua He, Die Hu 0002, Meng Wan, Ying Wen 0003, Karen M. von Deneen, MengChu Zhou
IEEE Trans. Syst. Man Cybern. Syst.1
2016 Semi-Blind Pilot Decontamination for Massive MIMO Systems
abstract
In multicell multiuser massive multi-input multi-output (MIMO) systems, pilot contamination degrades the uplink (UL) channel estimation performance. To mitigate the effect of pilot contamination, we propose a semiblind channel estimation method that does not require cell cooperation or statistical information of the channels. In the proposed method, we first sequentially estimate the UL data from different users in the target cell. To do that, for each user, we solve a constrained minimization problem to obtain an extracting vector and then use it to extract the desired data source from the observed mixture signal. An efficient algorithm is presented to solve the optimization problem. After the ambiguities in the extracted source are corrected with the aid of the pilot sequence, the estimates of the user UL data can be obtained. Based on the demodulated UL data of all users in the target cell, we finally obtain the least squares (LS) estimate of the channel. The pilot contamination effect is shown to be reduced as the UL data length grows. Simulation results demonstrate that the proposed method significantly outperforms some existing channel estimation methods that do not require cell cooperation or channel statistics.
Die Hu 0002, Lianghua He, Xiaodong Wang 0001
IEEE Trans. Wirel. Commun.2
2015 Null space based discriminant sparse representation large margin for face recognition
abstract
In this paper, we propose a novel subspace learning algorithm, termed as null space based discriminant sparse representation large margin (NDSLM). There are two contributions in the paper. First, we propose a new expectation to obtain the neighborhood information for large margin subspace learning, i.e., the within-neighborhood scatter and betweenneighborhood scatter are modeled by the sparse reconstruction weights of the samples from the same class and different classes, respectively. Since the neighborhood information formed by sparse representation can capture non-linearities in the data, the proposed method possesses more discriminative information than the traditional large margin learning methods with the expectation using Euclidean distance, etc. Second, the large margin information integrated into the model of Fisher criterion makes the discriminating power of NDSLM further boosted. NDSLM addresses the small sample size problem by solving an eigenvalue problem in null space. Experiments on ORL, Yale, AR, Extended Yale B and CMU PIE five face databases are performed to evaluate the proposed algorithm and the results demonstrate the effectiveness of NDSLM.
Ying Wen 0003, Lili Hou, Lianghua He
IJCNN3
2014 A Deep Learning Method for Classification of EEG Data Based on Motor Imagery
Xiu An, Deping Kuang, Xiaojiao Guo, Yilu Zhao, Lianghua He
ICIC (3)5
2014 ADHD-200 Classification Based on Social Network Method
Xiaojiao Guo, Xiu An, Deping Kuang, Yilu Zhao, Lianghua He
ICIC (3)5
2014 Motor Imagery EEG Signals Analysis Based on Bayesian Network with Gaussian Distribution
Lianghua He, Bin Liu 0018
ICIC (3)1
2014 Discrimination of ADHD Based on fMRI Data with Deep Belief Network
Deping Kuang, Xiaojiao Guo, Xiu An, Yilu Zhao, Lianghua He
ICIC (3)5
2014 Multi-channel features based automated segmentation of diffusion tensor imaging using an improved FCM with spatial constraints
Lianghua He, Ying Wen 0003, Meng Wan
Neurocomputing1
2012 Discriminative common vectors based on the Gram-Schmidt reorthogonalization for the small sample size problem
abstract
The discriminative common vectors (DCV) algorithm shows better face recognition effects than some commonly used linear discriminant algorithms, which uses the subspace methods and the Gram-Schmidt orthogonalization (GSO) procedure to obtain the DCV. However, the Gram-Schmidt technique may produce a set of vectors which is far from orthogonal so that sometimes the orthogonality may be lost completely. Hence, the effectiveness of the DCV is also decreased. In this paper, we proposed an improved DCV method based on the GSO. For obtaining an accurate projection onto the corresponding space, the orthogonal basis problem is usually solved with the Gram-Schmidt process with reorthogonalization. Thus, the effectiveness of the DCV can be improved and the experimental results show that the proposed method is better for the small sample size problem as compared to the DCV.
Ying Wen 0003, Lianghua He, Yue Lu 0001
ICASSP2
2012 A Co-adaptive Training Paradigm for Motor Imagery Based Brain-Computer Interface
Bin Xia 0004, Qingmei Zhang, Hong Xie 0001, Jie Li 0016, Lianghua He
ISNN (1)6
2012 A classifier for Bangla handwritten numeral recognition
Ying Wen 0003, Lianghua He
Expert Syst. Appl.2
2011 An Improved Locally Linear Embedding for Sparse Data Sets
abstract
Locally linear embedding is often invalid for sparse data sets because locally linear embedding simply takes the reconstruction weights obtained from the data space as the weights of the embedding space. This paper proposes an improved method for sparse data sets, a united locally linear embedding, to make the reconstruction more robust to sparse data sets. In the proposed method, the neighborhood correlation matrix presenting the position information of the points constructed from the embedding space is added to the correlation matrix in the original space, thus the reconstruction weights can be adjusted. As the reconstruction weights adjusted gradually, the position information of sparse points can also be changed continually and the local geometry of the data manifolds in the embedding space can be well preserved. Experimental results on both synthetic and real-world data show that the proposed approach is very robust against sparse data sets.
Ying Wen 0003, Lianghua He
Int. J. Pattern Recognit. Artif. Intell.2
2011 An Efficient Pilot Design Method for OFDM-Based Cognitive Radio Systems
abstract
In orthogonal frequency-division multiplexing (OFDM)-based cognitive radio (CR) systems, the subcarriers already occupied by the primary users cannot be used by the secondary users. This leads to possibly non-contiguous positions of the available subcarriers for the secondary users. The conventional pilot design methods are no longer effective for such systems. In this paper, we propose a new practical pilot design method for OFDM-based CR systems. We first formulate the pilot design as a new optimization problem. Instead of minimizing the mean-square error (MSE) of the least-squares (LS) channel estimator, we minimize an upper bound which is related to this MSE. We then propose an efficient scheme to solve the optimization problem. Specifically, the pilot indices are obtained sequentially by solving a series of one-dimensional optimization problems of significantly lower complexity. The computational complexity of the proposed scheme is low since it only involves real additions. Simulation results show that the pilot index sequences obtained by the proposed method exhibit significantly better performance than those obtained by existing pilot design methods.
Die Hu 0002, Lianghua He, Xiaodong Wang 0001
IEEE Trans. Wirel. Commun.2
2010 Pilot Design for Channel Estimation in OFDM-Based Cognitive Radio Systems
abstract
In orthogonal frequency division multiplexing (OFDM)-based cognitive radio (CR) systems, the subcarriers used by the licensed users (LUs) have to be deactivated to avoid interference. Thus only part of subcarriers can be used for transmission and the positions of the activated subcarriers may be non-contiguous. The conventional pilot design methods are no longer effective for such systems. In this paper, we propose a new practical pilot design method for OFDM-based CR systems. We first formulate the pilot design as a new optimization problem, where a simple objective function related to the mean-square error (MSE) of the least-squares (LS) channel estimation method is minimized. We then propose an efficient scheme to solve the optimization problem. Specifically, the pilot tones are obtained sequentially by solving some one-dimensional optimization problems. Only real additions are needed in the proposed scheme. The simulation results show that the pilot sequence obtained by the proposed method exhibits better performance than those obtained by existing pilot design methods.
Die Hu 0002, Lianghua He
GLOBECOM2
2006 Estimation of Rapidly Time-Varying Channels for OFDM Systems
abstract
Channel estimation for OFDM systems in rapidly time-varying environments is challenging. In this paper, relying on a basis expansion channel model, we propose a scheme for estimating channel parameters varying within a transmission block. Along with the estimation scheme, we also derive the optimal pilot sequence and optimal placement of pilot tones with respect to the mean square error (MSE) of the channel estimate. It is shown that the optimal pilot sequence consists of some adjacent equipowered and equispaced subsequences that are constrained by certain phase conditions. Simulation results demonstrate the performance of the proposed scheme in rapidly time-varying scenarios.
Die Hu 0002, Lianghua He, Luxi Yang
ICASSP (4)2
2005 Optimal pilot sequence design for multiple-input multiple-output OFDM systems
abstract
In orthogonal frequency division multiplexing (OFDM) systems, some subcarriers at the borders of the allocated bandwidth are usually used as guard band. Since these subcarriers, which are often referred to as virtual subcarriers, are not used for transmission, approach of conventional uniformly placed pilot tones is not applicable any more in some situations. Therefore, it is necessary to derive the optimal pilot sequences based on the nonuniform placement of pilot tones. In this paper, we first compute the mean square error (MSE) of the least squares (LS) channel estimate for multiple-input multiple-output (MIMO) OFDM systems. Then, based on nonuniform pilot tone placement, we derive the optimal pilot sequences with respect to this MSE. Simulation results demonstrate the effectiveness of the proposed approach
Die Hu 0002, Luxi Yang, Lianghua He, Yuhui Shi 0001
GLOBECOM3
2005 Boosted Independent Features for Face Expression Recognition
Lianghua He, Die Hu 0002, Cairong Zou, Li Zhao 0003
ISNN (2)1