Ying Wen 0003

dblp:41/4203-3 · DBLP profile ↗
← Back
54ranked-venue papers
12as first author
29since 2021 · last 2026
0000-0002-6974-5110ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 8 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 4 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3Computer networks · 2 · 2 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 A General Framework for Efficient Medical Image Analysis via Shared Attention Vision Transformer
abstract
Vision Transformers (ViTs) demonstrate significant promise in medical image analysis but face two critical challenges: 1) their limited ability to capture local features in data-scarce scenarios, leading to data inefficiency, and 2) their high computational and storage demands of the full fine-tuning process in transfer learning, resulting in parameter inefficiency. To achieve efficient and accurate medical image analysis, we propose Shared Attention Vision Transformer (SAViT) that comprises three innovative modules: i) Shared Prior Attention (SPA) that enhances data efficiency by innovatively employing a visual prompt to sequentially share consistent attention weights across local image regions, thereby enabling the learning of translational invariance to capture locality; ii) MixPool that preserves global modeling ability by aggregating local features after SPA through a multi-pooling mechanism, thus effectively facilitating long-range dependency across local image regions; and iii) Low-rank Multi-head Self-Attention (Lr-MSA) that improves parameter efficiency by using low-rank weights of multi-head self-attention, hence reducing computational complexity while maintaining accuracy in medical image analysis. SAViT demonstrates strong generalization across multiple medical imaging modalities, including retinopathy, dermoscopy, and radiography. Extensive experiments are conducted. The results indicate its high data efficiency and outstanding performance in comparison with more than 20 medical-specific and ViT-based models when all of them are trained from scratch. It excels in parameter-efficient tuning by surpassing 17 models across 6 datasets in transfer learning, with only ${0}.{17}$ M/ ${0}.{23}$ M trainable parameters on ViT-B/SwinViT-B backbones requiring ${86}.{60}$ M/ ${88}.{00}$ M parameters. Source code can be found at: https://github.com/LYH-hh/SAViT.
Ying Wen 0003, Longzhen Yang, Lianghua He, MengChu Zhou
IEEE Trans. Medical Imaging2
2025 M²RL-Net: Multi-View and Multi-Level Relation Learning Network for Weakly-Supervised Image Forgery Detection
abstract
As digital media manipulation becomes increasingly sophisticated, accurately detecting and localizing image forgeries with minimal supervision has become a critical challenge. Existing weakly supervised image forgery detection (W-IFD) methods often rely on convolutional neural networks (CNNs) and limited exploration of internal relationships, leading to poor detection and localization performance with only image-level labels. To address these limitations, we introduce a novel Multi-View and Multi-Level Relation Learning Network (M²RL-Net) for W-IFD. M²RL-Net effectively identifies forged images using only image-level annotations by exploring relationships between different views and hierarchical levels within images. Specifically, M²RL-Net achieves patch-level self-consistency learning (PSL) and feature-level contrastive learning (FCL) across different views, facilitating more generalized self-supervised learning of forgery features. In detail, PSL employs self-supervised learning to distinguish consistent and inconsistent regions within images, enhancing its ability to accurately locate tampered areas. FCL utilizes feature-level self-view and multi-view contrastive learning to differentiate between genuine and tampered image features, thereby improving the recognition of authentic and manipulated content across different views. Extensive experiments on various datasets demonstrate that M²RL-Net outperforms existing weakly-supervised methods in both detection and localization accuracy. This research sets a new benchmark for weakly-supervised image forgery detection and lays a robust foundation for future studies in this field.
Jiafeng Li 0005, Ying Wen 0003, Lianghua He
AAAI2
2025 AFiRe: Anatomy-Driven Self-Supervised Learning for Fine-Grained Representation in Radiographic Images
abstract
Current self-supervised methods, such as contrastive learning, predominantly focus on global discrimination, neglecting the critical fine-grained anatomical details required for accurate radiographic analysis. To address this challenge, we propose the Anatomy-driven self-supervised framework for enhancing Fine-grained Representation in radiographic image analysis (AFiRe). The core idea of AFiRe is to align the anatomical consistency with the unique token-processing characteristics of Vision Transformer. Specifically, AFiRe synergistically performs two self-supervised schemes: (i) Token-wise anatomy-guided contrastive learning, which aligns image tokens based on structural and categorical consistency to enhance fine-grained spatial-anatomical discrimination; (ii) Pixel-level anomaly-removal restoration, which particularly focuses on local anomalies, thereby refining the learned discrimination with detailed geometrical information. Additionally, we propose the Synthetic Lesion Mask to enhance anatomical diversity while preserving intra-consistency, which is typically corrupted by traditional data augmentations, such as Cropping and Affine transformations. Experimental results show that AFiRe: (i) provides robust anatomical discrimination, achieving more cohesive feature clusters compared to state-of-the-art contrastive learning methods; (ii) demonstrates superior generalization, surpassing 7 radiography-specific self-supervised methods in multi-label classification tasks with limited labeling; and (iii) integrates fine-grained information, enabling precise anomaly detection using only image-level annotations.
Lianghua He, Ying Wen 0003, Longzhen Yang, Hongzhou Chen
AAAI3
2025 Self-Optimization Training for Weakly Supervised Image Manipulation Localization
abstract
With the continuous evolution of image manipulation techniques, there is an urgent need for an effective method to detect and localize manipulated images. However, existing fully supervised methods require large amounts of costly pixel-level annotations, whereas weakly supervised methods often fall short in localization performance due to their inability to accurately localize tampered regions with precise boundaries. To tackle this issue, we propose a Self-Optimization Weakly Supervised Localization (SO-WSL) framework, which consists of two main components: a Pseudo-Label Generator (PLG) and a Self Iterative Optimization (SIO) module. The PLG employs Class Activation Maps (CAM) to guide the Segment Anything Model (SAM) in generating pseudo-labels with distinct edges, while the SIO module enhances the model’s focus on forgery-specific features by applying masks to suspected tampered regions and iteratively refining the pseudo-labels, thereby improving localization accuracy. Extensive experiments have shown that our SO-WSL framework significantly outperforms existing weakly supervised methods and can even compete with some fully supervised approaches.
Zhangchen Zhu, Jiafeng Li 0005, Ying Wen 0003
ICASSP3
2025 CoSMIC: Continual Self-Supervised Learning for Multi-Domain Medical Imaging Via Conditional Mutual Information Maximization
Ying Wen 0003, Longzhen Yang, Lianghua He, Heng Tao Shen
ICCV2
2025 RadLAS: A Foundation Model for Interpretable Radiography Image Analysis with Lesion-Aware Self-Supervised Pre-training
abstract
Medical Foundation Models (MFMs) are revolutionizing radiography image analysis with scalable and generalized diagnostic capabilities. However, their effectiveness in real-world clinical practice is limited due to insufficient interpretability. To address this limitation, we propose RadLAS, a novel MFM for interpretable Radiographic image analysis by introducing Lesion-Aware Self-supervised pre-training. Unlike conventional MFMs that rely on post-hoc explanations, RadLAS innovates by directly emulating human diagnostic reasoning to first grounding lesion evidence and then making decisions accordingly. Specifically, RadLAS introduces two self-supervised tasks: (I) Lesion-grounded Reconstruction, which learns structured anatomical representations by restoring lesion-aware image patches into their healthy counterparts, thereby facilitating pixel-level grounding of lesion evidence via input-normal contrast. (II) Lesion-discrimination Contrastive Learning, which enhances lesion-aware pattern in representations by explicitly decoupling grounded lesion evidence as clinical cues and aligning them with global semantics, thereby enabling direct lesion-oriented diagnosis while preserving global context. RadLAS demonstrates excellent performance across diverse downstream radiographic datasets, offering verifiable explanations by deriving specific diagnoses (Task II) based on grounded lesion evidence (Task I), while preserving generalized representations essential for high diagnostic accuracy. Extensive experiments demonstrate that RadLAS (i) achieves superior interpretability with highly correlated lesion prediction and localization, surpassing 11 interpretable medical models; (ii) delivers scalable representation learning, outperforming 14 SOTA supervised and self-supervised MFMs.
Ying Wen 0003, Longzhen Yang, Lianghua He, Heng Tao Shen
ACM Multimedia2
2025 Self-Supervised Anatomical Consistency Learning for Vision-Grounded Medical Report Generation
abstract
Vision-grounded medical report generation aims to produce clinically accurate descriptions of medical images, anchored in explicit visual evidence to improve interpretability and facilitate integration into clinical workflows. However, existing methods often rely on separately trained detection modules that require extensive expert annotations, introducing high labeling costs and limiting generalizability due to pathology distribution bias across datasets. To address these challenges, we propose Self-Supervised Anatomical Consistency Learning (SS-ACL)-a novel and annotation-free framework that aligns generated reports with corresponding anatomical regions using simple textual prompts. SS-ACL constructs a hierarchical anatomical graph inspired by the invariant top-down inclusion structure of human anatomy, organizing entities by spatial location. It recursively reconstructs fine-grained anatomical regions to enforce intra-sample spatial alignment, inherently guiding attention maps toward visually relevant areas prompted by text. To further enhance inter-sample semantic alignment for abnormality recognition, SS-ACL introduces a region-level contrastive learning based on anatomical consistency. These aligned embeddings serve as priors for report generation, enabling attention maps to provide interpretable visual evidence. Extensive experiments demonstrate that SS-ACL, without relying on expert annotations, (i) generates accurate and visually grounded reports-outperforming state-of-the-art methods by 10% in lexical accuracy and 25% in clinical efficacy, and (ii) achieves competitive performance on various downstream visual tasks, surpassing current leading visual foundation models by 8% in zero-shot visual grounding. Our code is available at https://github.com/kaelsunkiller/ssacl.
Longzhen Yang, Zhangkai Ni, Ying Wen 0003, Lianghua He, Heng Tao Shen
ACM Multimedia3
2025 SFM-Net: Semantic Feature-Based Multi-Stage Network for Unsupervised Image Registration
abstract
It is difficult for general registration methods to establish the fine correspondence between images with complex anatomical structures. To overcome the above problem, this work presents SFM-Net, an unsupervised multi-stage semantic feature-based network. In addition to using the pixel-based similarity metrics, we propose a feature operator and emphasize a feature registration to improve the alignment of semantic related areas. Specifically, we design a two-stage training strategy, the intensity image registration stage and the semantic feature registration stage. The former is for valid semantic features learning and intensity-based coarse registration, while the latter is for semantic areas alignment, achieving fine transformation of anatomical structure. The same structure of both stages is composed of a dual-stream feature extraction module (DFEM) and a refined deformation field generation module (RDGM). Unlike the deep learning-based approaches that utilizing down-sampled encoder to extract features, DFEM constructed by dual-stream U-Net structure can capture semantic information in decoder feature for structural alignment. Different with approaches applying cascaded networks to learn deformation field, our proposed RDGM generates multi-scale deformation fields by performing a coarse-to-fine registration within a single network. Experiments on 3D brain MRI and liver CT datasets confirm that the proposed SFM-Net achieves accurate and diffeomorphic registration results, outperforming other state-of-the-art methods.
Tai Ma, Xinru Dai, Suwei Zhang, Haidong Zou, Lianghua He, Ying Wen 0003
IEEE J. Biomed. Health Informatics6
2024 EAN: An Efficient Attention Module Guided by Normalization for Deep Neural Networks
abstract
Deep neural networks (DNNs) have achieved remarkable success in various fields, and two powerful techniques, feature normalization and attention mechanisms, have been widely used to enhance model performance. However, they are usually considered as two separate approaches or combined in a simplistic manner. In this paper, we investigate the intrinsic relationship between feature normalization and attention mechanisms and propose an Efficient Attention module guided by Normalization, dubbed EAN. Instead of using costly fully-connected layers for attention learning, EAN leverages the strengths of feature normalization and incorporates an Attention Generation (AG) unit to re-calibrate features. The proposed AG unit exploits the normalization component as a measure of the importance of distinct features and generates an attention mask using GroupNorm, L2 Norm, and Adaptation operations. By employing a grouping, AG unit and aggregation strategy, EAN is established, offering a unified module that harnesses the advantages of both normalization and attention, while maintaining minimal computational overhead. Furthermore, EAN serves as a plug-and-play module that can be seamlessly integrated with classic backbone architectures. Extensive quantitative evaluations on various visual tasks demonstrate that EAN achieves highly competitive performance compared to the current state-of-the-art attention methods while sustaining lower model complexity.
Jiafeng Li 0005, Zelin Li 0005, Ying Wen 0003
AAAI3
2024 IIRP-Net: Iterative Inference Residual Pyramid Network for Enhanced Image Registration
abstract
Deep learning-based image registration (DLIR) meth-ods have achieved remarkable success in deformable im-age registration. We observe that iterative inference can exploit the well-trained registration network to the fullest extent. In this work, we propose a novel Iterative Inference Residual Pyramid Network (IIRP-Net) to enhance registration performance without any additional training costs. In IIRP-Net, we construct a streamlined pyramid registration network consisting of a feature extractor and residual flow estimators (RP-Net) to achieve generalized capabilities in feature extraction and registration. Then, in the inference phase, IIRP-Net employs an iterative inference strategy to enhance RP-Net by iteratively reutilizing residual flow es-timators from coarse to fine. The number of iterations is adaptively determined by the proposed IterStop mecha-nism. We conduct extensive experiments on the FLARE and Mindboggle datasets and the results verify the effectiveness of the proposed method, outperforming state-of-the-art de-formable image registration methods. Our code is available at https://github.com/Torbjorn1997/IIRP-Net.
Tai Ma, Suwei Zhang, Jiafeng Li 0005, Ying Wen 0003
CVPR4
2024 RC-Block: Refinement Coefficient for Rectifying Deformation Field
abstract
In deformable image registration, learning-based methods have demonstrated impressive performance. However, previous methods mainly focus on enhancing the capability to predict the deformation field, without fully exploring the optimal ways to update the deformation field. As a result, errors can easily accumulate during the process of updating the deformation field, thus leading to inaccurate registration results. To solve this problem, we propose RC-Block (Refinement Coefficient Block) to find the optimal way to update the deformation field, thus achieving high accuracy registration results. The proposed module introduces a novel refinement coefficient, which can leverage information from the current scale to continuously rectify the previous deformation field. In addition, RC-Block is a plug-and-play module that can be seamlessly and easily integrated into most registration methods. Extensive experimental results show that our module can effectively improve existing advanced methods’ performance with minimal additional burden.
Suwei Zhang, Tai Ma, Ying Wen 0003
ICME3
2024 KD loss: Enhancing discriminability of features with kernel trick for object detection in VHR remote sensing images
Xi Chen 0004, Liyue Li, Qingli Li, Honggang Qi, Ying Wen 0003, Guitao Cao, Philip L. H. Yu
Eng. Appl. Artif. Intell.8
2023 ScaleNet: Rethinking Feature Interaction from a Scale-Wise Perspective for Medical Image Segmentation
Tai Ma, Zhengke Xu, Suwei Zhang, Ying Wen 0003
CGI (4)6
2023 MSINET: Multi-scale Interconnection Network for Medical Image Segmentation
Zhengke Xu, Xinxin Shan, Ying Wen 0003
CGI (4)3
2023 SCConv: Spatial and Channel Reconstruction Convolution for Feature Redundancy
abstract
Convolutional Neural Networks (CNNs) have achieved remarkable performance in various computer vision tasks but this comes at the cost of tremendous computational resources, partly due to convolutional layers extracting redundant features. Recent works either compress well-trained large-scale models or explore well-designed lightweight models. In this paper, we make an attempt to exploit spatial and channel redundancy among features for CNN compression and propose an efficient convolution module, called SCConv (Spatial and Channel reconstruction Convolution), to decrease redundant computing and facilitate representative feature learning. The proposed SCConv consists of two units: spatial reconstruction unit (SRU) and channel reconstruction unit (CRU). SRU utilizes a separate-and-reconstruct method to suppress the spatial redundancy while CRU uses a split-transform-and-fuse strategy to diminish the channel redundancy. In addition, SCConv is a plug-and-play architectural unit that can be used to replace standard convolution in various convolutional neural networks directly. Experimental results show that SCConv-embedded models are able to achieve better performance by reducing redundant features with significantly lower complexity and computational costs.
Jiafeng Li 0005, Ying Wen 0003, Lianghua He
CVPR2
2023 MSAANet: Multi-scale Axial Attention Network for medical image segmentation
abstract
U-Net and its variants have achieved impressive results in medical image segmentation. However, the downsampling operation of such U-shaped networks causes the feature maps to lose a certain degree of spatial information, and most existing methods use convolution and transformer sequentially, it is hard to extract more comprehensive feature representation of the image. In this paper, we propose a novel U-shaped segmentation network named Multi-scale Axial Attention Network (MSAANet) to solve the above problems. Specifically, we propose a cross-scale interactive attention: multi-scale axial attention (MSAA), which achieves direction-perception attention of different scales interaction. So that the downsampling deep features and the shallow features can maintain context spatial consistency. Besides, we propose a Convolution-Transformer (CT) block, which makes transformer and convolution complement each other to enhance comprehensive feature representation. We evaluate the proposed method on the public datasets Synapse and ACDC. Experimental results demonstrate that MSAANet effectively improves segmentation accuracy.
Xinxin Shan, Ying Wen 0003
ICME4
2023 PIViT: Large Deformation Image Registration with Pyramid-Iterative Vision Transformer
Tai Ma, Xinru Dai, Suwei Zhang, Ying Wen 0003
MICCAI (10)4
2023 Prediction of common labels for universal domain adaptation
Xinxin Shan, Tai Ma, Ying Wen 0003
Neural Networks3
2023 A multi-grained unsupervised domain adaptation approach for semantic segmentation
Tai Ma, Yue Lu 0001, Qingli Li, Lianghua He, Ying Wen 0003
Pattern Recognit.6
2022 Unsupervised Hierarchical Translation-Based Model for Multi-Modal Medical Image Registration
abstract
Deformable registration of multi-modal medical images is a challenging task in medical image processing due to the differences in both appearance and structure. We propose an unsupervised hierarchical translation-based model to perform a coarse to fine registration of multi-modal medical images. The proposed model consists of three parts: a coarse registration network, a modal translation network and a fine registration network. First, the coarse registration network learns to obtain the coarse deformation field, which is applied as structure-preserving information to generate a translated image by the modal translation network. Then, the translated image as enhancing information combined with the original images are used to derive a fine deformation field in the fine registration network. Furthermore, the final deformation field is composed from the coarse and the fine deformation fields. In this way, the proposed model can learn high accurate deformation field to implement multi-modal medical image registration. Experiments on two multi-modal brain image datasets demonstrate the effectiveness of this model.
Xinru Dai, Tai Ma, Haibin Cai, Ying Wen 0003
ICASSP4
2022 TCRNet: Make Transformer, CNN and RNN Complement Each Other
abstract
Recently, several Transformer-based methods have been presented to improve image segmentation. However, since Transformer needs regular square images and has difficulty in obtaining local feature information, the performance of image segmentation is seriously affected. In this paper, we propose a novel encoder-decoder network named TCRNet, which makes Transformer, Convolutional neural network (CNN) and Recurrent neural network (RNN) complement each other. In the encoder, we extract and concatenate the feature maps from Transformer and CNN to effectively capture global and local feature information of images. Then in the decoder, we utilize convolutional RNN in the proposed recurrent decoding unit to refine the feature maps from the decoder for finer prediction. Experimental results on three medical datasets demonstrate that TCRNet effectively improves the segmentation precision.
Xinxin Shan, Tai Ma, Anqi Gu, Haibin Cai, Ying Wen 0003
ICASSP5
2022 KAConv: Kernel attention convolutions
Xinxin Shan, Tai Ma, YuTao Shen, Jiafeng Li 0005, Ying Wen 0003
Neurocomputing5
2021 A New Framework Based on Transfer Learning for Cross-Database Pneumonia Detection
abstract
Cross-database classification means that the model is able to apply to the serious disequilibrium of data distributions, and it is trained by one database while tested by another database. Thus, cross-database pneumonia detection is a challenging task. In this paper, we proposed a new framework based on transfer learning for cross-database pneumonia detection. First, based on transfer learning, we fine-tune a backbone that pre-trained on non-medical data by using a small amount of pneumonia images, which improves the detection performance on homogeneous dataset. Then in order to make the fine-tuned model applicable to cross-database classification, the adaptation layer combined with a self-learning strategy is proposed to retrain the model. The adaptation layer is to make the heterogeneous data distributions approximate and the self-learning strategy helps to tweak the model by generating pseudo-labels. Experiments on three pneumonia databases show that our proposed model completes the cross-database detection of pneumonia and shows good performance.
Xinxin Shan, Ying Wen 0003
ICASSP2
2021 A Coherent Cooperative Learning Framework Based on Transfer Learning for Unsupervised Cross-Domain Classification
Xinxin Shan, Ying Wen 0003, Qingli Li, Yue Lu 0001, Haibin Cai
MICCAI (5)2
2021 Gaussian Mixture Model Based Semi-supervised Sparse Representation for Face Recognition
Xinxin Shan, Ying Wen 0003
MMM (1)2
2021 Sparse Common Feature Representation for Undersampled Face Recognition
abstract
This work investigates the problem of undersampled face recognition (i.e., insufficient training data) encountered in practical Internet-of-Things (IoT) applications. Insufficient and uncertain samples captured by IoT devices may include background and facial disguise that makes face recognition more challenging than that with sufficient and reliable images. Many models work well in face recognition on a big data set, but when training data are insufficient, they achieve unsatisfactory performance. This work proposes a novel method named sparse common feature-based representation (SCFR) that provides a unique and stable result and completely avoids very time-consuming training required by a deep learning model. Specially, it constructs a common feature dictionary using both training and test images. Thereinto, a common feature is based on a discriminative common vector and learned by a Gaussian mixture model for both training and test images in a semisupervised learninig manner, which would reduce the difference among samples in each class. In the optimization, the latent indicator of test data is initialized by the estimated label. This can avoid learning invalid information and lead to good prototype images. A new variation dictionary characterizes variables that can be shared by different classes. Finally, this work adopts minimum reconstruction residuals to recognize test images, thus bringing about a substantial improvement in SCFR's performance. Extensive results on benchmark face databases demonstrate that the proposed method is better than the state-of-the-art methods handling undersampled face recognition.
Shicheng Yang, Ying Wen 0003, Lianghua He, MengChu Zhou
IEEE Internet Things J.2
2021 Sparse Individual Low-Rank Component Representation for Face Recognition in the IoT-Based System
abstract
The performance of face recognition has been greatly improved by deep neural network algorithms when a dataset is large. However, when face data are insufficient as in practical Internet of Things (IoT) applications and captured by IoT devices under the same intrasubject variation, both data quantity and quality bring big challenges to construct a model or representation, and most of the time it becomes infeasible to build a deep neural network model. This work proposes a sparse individual low-rank component-based representation (SILR) such that the representation of testing images can be based on individual subjects’ low-rank component. Theoretically, we put the$l_{2}$-norm constraint on intrasubject coefficients to represent testing images, thus making intrasubject coefficients dense. Hence, we alleviate the impact of an undersampled training dataset and its same intersubject variation on classification performance. We solve a convex minimization problem in polynomial time via an augmented lagrange multiplier scheme to get the solution of SILR. The scheme can reduce the influences from the same intersubject variation and contribute to an accurate recognition of the undersampled training dataset. We adopt sparse individual low-rank component representation and minimum reconstruction residual to recognize testing images. Extensive results on various databases show that SILR outperforms the other state-of-the-art methods for face recognition.
Shicheng Yang, Ying Wen 0003, Lianghua He, MengChu Zhou, Abdullah Abusorrah
IEEE Internet Things J.2
2021 Model-Based Transfer Learning and Sparse Coding for Partial Face Recognition
abstract
With the growing needs of practical applications such as security monitoring, partial face recognition is a challenging but important issue, because the captured faces in real-world surveillance videos may be occluded or with variations. Though current face recognition methods perform well in relatively constrained scenes, they may suffer from degradation for partial faces. In this paper, we propose a framework of model-based transfer learning and sparse coding (MTLSC) for partial face recognition. First, due to less information in partial face image, we exploit the mirrored image of an original probe sample as sample augment to provide further information. Considering the inadequacy of training face samples, we obtain face features based on model-based transfer learning VGGNet that is pre-trained on VGGFace dataset. Then we reconstruct face features by sliding window in view of different sizes of partial face hard to extract the same feature dimension. Finally we carry out sparse coding with rectification and calculate the minimum score of the probe and mirrored samples among all classes to get the results. Thus, by model-based transfer learning, sliding window for feature reconstruction and sparse coding with rectification, the proposed framework improves partial face recognition performance. Experimental results on three face databases (LFW, AR and NIR), and two person re-identification databases (iLIDS-VID and PKU-Reid) demonstrate our method is effective for partial face recognition.
Xinxin Shan, Yue Lu 0001, Qingli Li, Ying Wen 0003
IEEE Trans. Circuits Syst. Video Technol.4
2021 Identification of Melanoma From Hyperspectral Pathology Image Using 3D Convolutional Networks
abstract
Skin biopsy histopathological analysis is one of the primary methods used for pathologists to assess the presence and deterioration of melanoma in clinical. A comprehensive and reliable pathological analysis is the result of correctly segmented melanoma and its interaction with benign tissues, and therefore providing accurate therapy. In this study, we applied the deep convolution network on the hyperspectral pathology images to perform the segmentation of melanoma. To make the best use of spectral properties of three dimensional hyperspectral data, we proposed a 3D fully convolutional network named Hyper-net to segment melanoma from hyperspectral pathology images. In order to enhance the sensitivity of the model, we made a specific modification to the loss function with caution of false negative in diagnosis. The performance of Hyper-net surpassed the 2D model with the accuracy over 92%. The false negative rate decreased by nearly 66% using Hyper-net with the modified loss function. These findings demonstrated the ability of the Hyper-net for assisting pathologists in diagnosis of melanoma based on hyperspectral pathology images.
Qian Wang 0046, Li Sun 0012, Yan Wang 0033, Mei Zhou, Menghan Hu, Ying Wen 0003, Qingli Li
IEEE Trans. Medical Imaging7
2020 Segmenting Medical MRI via Recurrent Decoding Cell
abstract
The encoder-decoder networks are commonly used in medical image segmentation due to their remarkable performance in hierarchical feature fusion. However, the expanding path for feature decoding and spatial recovery does not consider the long-term dependency when fusing feature maps from different layers, and the universal encoder-decoder network does not make full use of the multi-modality information to improve the network robustness especially for segmenting medical MRI. In this paper, we propose a novel feature fusion unit called Recurrent Decoding Cell (RDC) which leverages convolutional RNNs to memorize the long-term context information from the previous layers in the decoding phase. An encoder-decoder network, named Convolutional Recurrent Decoding Network (CRDN), is also proposed based on RDC for segmenting multi-modality medical MRI. CRDN adopts CNN backbone to encode image features and decode them hierarchically through a chain of RDCs to obtain the final high-resolution score map. The evaluation experiments on BrainWeb, MRBrainS and HVSMR datasets demonstrate that the introduction of RDC effectively improves the segmentation accuracy as well as reduces the model size, and the proposed CRDN owns its robustness to image noise and intensity non-uniformity in medical MRI.
Ying Wen 0003, Lianghua He
AAAI1
2020 Gabor Feature-Based LogDemons With Inertial Constraint for Nonrigid Image Registration
abstract
Nonrigid image registration plays an important role in the field of computer vision and medical application. The methods based on Demons algorithm for image registration usually use intensity difference as similarity criteria. However, intensity based methods can not preserve image texture details well and are limited by local minima. In order to solve these problems, we propose a Gabor feature based LogDemons registration method in this paper, called GFDemons. We extract Gabor features of the registered images to construct feature similarity metric since Gabor filters are suitable to extract image texture information. Furthermore, because of the weak gradients in some image regions, the update fields are too small to transform the moving image to the fixed image correctly. In order to compensate this deficiency, we propose an inertial constraint strategy based on GFDemons, named IGFDemons, using the previous update fields to provide guided information for the current update field. The inertial constraint strategy can further improve the performance of the proposed method in terms of accuracy and convergence. We conduct experiments on three different types of images and the results demonstrate that the proposed methods achieve better performance than some popular methods.
Ying Wen 0003, Yue Lu 0001, Qingli Li, Haibin Cai, Lianghua He
IEEE Trans. Image Process.1
2019 Sparse Low-Rank Component-Based Representation for Face Recognition With Low-Quality Images
abstract
Sparse-representation-based classification (SRC) has been showing a good performance for face recognition in recent years. But SRC is not good at face recognition with low quality images (e.g., disguised, corrupted, occluded, and so on) which often appear in practical applications. To solve the problem, in this paper, we propose a novel SRC-based method for face recognition with low quality images named sparse low-rank component-based representation (SLCR). In SLCR, we utilize low-rank matrix recovery on the training data set to obtain low-rank components and non-low-rank components, which are used to construct the dictionary. The new dictionary is capable of describing facial features better, especially for low quality face samples. Furthermore, the minimum class-wise reconstruction residual is used as the recognition rule, leading to a substantial improvement on the proposed SLCR's performance. Extensive experiments on benchmark face databases demonstrate that the proposed method is consistently superior to other sparse-representation-based approaches for face recognition with low quality images.
Shicheng Yang, Lianghua He, Ying Wen 0003
IEEE Trans. Inf. Forensics Secur.4
2019 Incorporation of Structural Tensor and Driving Force Into Log-Demons for Large-Deformation Image Registration
abstract
Large-deformation image registration is important in theory and application in computer vision, but is a difficult task for non-rigid registration methods. In this paper, we propose a structural Tensor and Driving force-based Log-Demons algorithm for it, named TDLog-Demons for short. The structural tensor of an image is proposed to obtain a highly accurate deformation field. The driving force is proposed to solve the registration issue of large-deformation that often causes Log-Demons to trap into local minima. It is defined as a point correspondence obtained via multisupport-region-order-based gradient histogram descriptor matching on image's boundary points. It is integrated into an exponentially decreasing form with the velocity field of Log-Demons to move the points accurately and to speed up a registration process. Consequently, the driving force-based Log-Demons can well deal with large-deformation image registration. Extensive experiments demonstrate that the TDLog-Demons not only captures large deformations at a high accuracy but also yields a smooth deformation.
Ying Wen 0003, Lianghua He, MengChu Zhou
IEEE Trans. Image Process.1
2018 Sparse Low-Rank Component Coding for Face Recognition with Illumination And Corruption
abstract
Sparse representation-based classification shows a good performance for face recognition in recent years, but it can not be suitable for face recognition with illumination and corruption, which are often presented in the practical applications. To solve the problem, in this paper, we propose a novel SRC based method for face recognition named sparse low-rank component coding (SLC). In SLC, we utilize the low-rank component from training dataset to construct dictionary. The dictionary composed of low-rank component is able to describe the face feature better, especially for training samples with illumination and corruption. Our recognition rule is based on the minimum class-wise reconstruction residual which leads to a substantial improvement on the performance of SLC. Extensive experiments on benchmark face databases demonstrate that the proposed method consistently outperforms the other sparse representation based approaches for face recognition with illumination and corruption.
Shicheng Yang, Ying Wen 0003, Lianghua He
ICASSP2
2018 Adaptive Fuzzy Clustering Algorithm with Local Information and Markov Random Field for Image Segmentation
Jialiang Hu, Ying Wen 0003
ICONIP (4)2
2018 Inertial Constrained Hierarchical Belief Propagation for Optical Flow
Zixing Zhang 0003, Ying Wen 0003
PRICAI (1)2
2016 Face recognition using locality sparsity preserving projections
abstract
In this paper, we present a new and effective dimensionality reduction method called locality sparsity preserving projections (LSPP). Locality preserving projections (LPP) and sparsity preserving projections (SPP) only focus on an aspect of local structure and sparse reconstructive information of the dataset, respectively. The proposed method integrates the sparse reconstructive information and local structure of data. The projection of LSPP is sought such that the sparse reconstructive weights and local preserving weights can be best preserved and integrated. Extensive experiments on ORL, Yale, Yale B, AR and CMU PIE face databases show the effectiveness of the proposed LSPP.
Ying Wen 0003, Shicheng Yang, Lili Hou, Hongda Zhang
IJCNN1
2016 Motor imagery EEG signals analysis based on Bayesian network with Gaussian distribution
Lianghua He, Bin Liu 0018, Die Hu 0002, Ying Wen 0003, Meng Wan
Neurocomputing4
2016 Discriminant Sparsity Preserving Analysis for Face Recognition
abstract
Sparse subspace learning has drawn more and more attentions recently, however, most of them are unsupervised and unsuitable for classification tasks. In this paper, a new discriminant sparsity preserving analysis (DSPA) method by integrating sparse reconstructive weighting into Fisher criterion is proposed for face recognition. We first get sparsity preserving space spanned by the eigenvectors of sparsity preserving projections (SPP). Then, the optimal projection can be obtained by solving an eigenvalue and eigenvector problem of the between-class scatter matrix in sparsity preserving space. The method not only preserves the sparse reconstructive relationship of the data, but also encodes the discriminant information. Extensive experiments on four face image datasets (Yale, ORL, AR and CMU PIE) demonstrate the effectiveness of the proposed DSPA method.
Ying Wen 0003, Lili Hou
Int. J. Pattern Recognit. Artif. Intell.1
2016 Recognition of handwritten Chinese address with writing variations
Xiaohua Wei, Shujing Lu, Ying Wen 0003, Yue Lu 0001
Pattern Recognit. Lett.3
2016 Common Bayesian Network for Classification of EEG-Based Multiclass Motor Imagery BCI
abstract
Modeling and learning of brain activity patterns represent a huge challenge to the brain-computer interface (BCI) based on electroencephalography (EEG). Many existing methods estimate the uncorrelated instantaneous demixing of EEG signals to classify multiclass motor imagery (MI). However, the condition of uncorrelation does not hold true in practice, because the brain regions work with partial or complete collaboration. This work proposes a novel method, termed as a common Bayesian network (CBN), to discriminate multiclass MI EEG signals. First, with the constraints of a Gaussian mixture model on every channel, only related channels are selected to construct a normal Bayesian network. Second, the nodes that have both common and varying edges are selected to construct a CBN. Third, the probabilities on common edges are used to learn about the support vector machine for classification. To validate the proposed method, we conduct experiments on two well-known BCI datasets and perform a numerical analysis of the propose algorithm for EEG classification in a multiclass MI BCI. Experimental results show that the proposed CBN method not only has excellent classification performance, but also is highly efficient. Hence, it is suitable for the cases where a system is required to respond within a second.
Lianghua He, Die Hu 0002, Meng Wan, Ying Wen 0003, Karen M. von Deneen, MengChu Zhou
IEEE Trans. Syst. Man Cybern. Syst.4
2015 Using multiple sequence alignment and statistical language model to integrate multiple Chinese address recognition outputs
abstract
Different recognizers may result in different mistakes when they are used to recognize a Chinese address. In this paper, we present a method of combining multiple Chinese address recognition outputs to improve Chinese address recognition accuracy. The method first employs multiple sequence alignment to generate a lattice of candidate hypotheses from multiple different recognizer outputs and then applies statistical language model to choose the maximum likelihood candidate sequence. Taking the maximum as the final decision, the performance of our method is superior, compared to the single recognizers and Miyao's method. The experiments on the address images of real envelopes demonstrate that the proposed method increases the character recognition accuracy rate from 95.80% to 98.38%, with 61.30% error reduction. Furthermore, the corrected sorting rate of an automatic mail sorting system increases from 84.11% to 93.72%.
Shengchang Chen, Shujing Lu, Ying Wen 0003, Yue Lu 0001
ICDAR3
2015 Text-independent writer identification using SIFT descriptor and contour-directional feature
abstract
This paper presents a method for text-independent writer identification using SIFT descriptor and contour-directional feature (CDF). The proposed method contains two stages. In the first stage, a codebook of local texture patterns is constructed by clustering a set of SIFT descriptors extracted from images. Using this codebook, the occurrence histograms are calculated to determine the similarities between different images. For each image, we obtain a candidate list of reference images. The next stage is to refine the candidate list using the contour-directional feature and SIFT descriptor. The proposed method is evaluated with two datasets: the ICFHR2012-Latin dataset and the ICDAR2013 dataset. Experimental results show that the proposed method outperforms the state-of-the-art algorithms and archives the best performance.
Yujie Xiong, Ying Wen 0003, Patrick Shen-Pei Wang, Yue Lu 0001
ICDAR2
2015 HoG based two-directional Dynamic Time Warping for handwritten word spotting
abstract
We present a Histogram of Oriented Gradient (HoG) based two-directional Dynamic Time Warping (DTW) matching method for handwritten word spotting. Firstly, we extract HoG descriptors from each cell in the normalized images. Then we connect the HoG descriptors in the same column and get a sequence of feature vectors. We do the same operation for the HoG descriptors in the same row. We then apply the two-directional DTW method to calculate the distance between the feature vectors sequences extracted from the query word and the candidate one. The experimental results show that the two-directional DTW is more robust to word deformation than the traditional DTW. And the local features such as HoG, LBP and SIFT combined with the two-directional DTW method outperform the method using the local feature descriptors directly. The HoG based two-directional DTW get the highest mean average precision on both the George Washington dataset and the CASIA-HWDB 2.1 dataset.
Shunyi Yao, Ying Wen 0003, Yue Lu 0001
ICDAR2
2015 Scene text detection using sequential nontext filtering
abstract
We present a scene text detection method based on sequential nontext filtering. Firstly, we start our work with multi-channel maximally stable extremal region (MSER) detection. Then nontext components are eliminated by a four-stage sequential nontext filtering strategy which consists of inner-channel MSER pruning, between-channel MSER pruning, unary feature-based nontext filtering, and binary feature-based nontext filtering. Finally, text components are grouped into words and false positives are eliminated. The proposed method achieves the state-of-the-art on the ICDAR2013 database when compared with some existing methods.
Yue Lu 0001, Ying Wen 0003
ICIP3
2015 An extended fuzzy local information C-means clustering algorithm
abstract
Fuzzy c-means clustering algorithm (FCM) is often used for image segmentation but it is sensitive to noise. This paper presents an extended fuzzy local information c-means clustering algorithm for robust image segmentation. In this method, a novel fuzzy factor created by the neighborhood spatial and gray information is integrated into the objective function of FCM. The fuzzy factor can enhance the algorithm's clustering performance by adjusting the influence of neighboring pixels to the center pixel. The proposed method can not only preserve the image details but also enhance the robustness to noise. Experiments implemented on synthetic images and real images demonstrate that the proposed method achieves better performance for image segmentation, especially for images corrupted by strong noise, compared to the traditional FCM and its extended methods.
Lili Hou, Qiuying Yang, Ying Wen 0003
IJCNN4
2015 Null space based discriminant sparse representation large margin for face recognition
abstract
In this paper, we propose a novel subspace learning algorithm, termed as null space based discriminant sparse representation large margin (NDSLM). There are two contributions in the paper. First, we propose a new expectation to obtain the neighborhood information for large margin subspace learning, i.e., the within-neighborhood scatter and betweenneighborhood scatter are modeled by the sparse reconstruction weights of the samples from the same class and different classes, respectively. Since the neighborhood information formed by sparse representation can capture non-linearities in the data, the proposed method possesses more discriminative information than the traditional large margin learning methods with the expectation using Euclidean distance, etc. Second, the large margin information integrated into the model of Fisher criterion makes the discriminating power of NDSLM further boosted. NDSLM addresses the small sample size problem by solving an eigenvalue problem in null space. Experiments on ORL, Yale, AR, Extended Yale B and CMU PIE five face databases are performed to evaluate the proposed algorithm and the results demonstrate the effectiveness of NDSLM.
Ying Wen 0003, Lili Hou, Lianghua He
IJCNN1
2014 Multi-channel features based automated segmentation of diffusion tensor imaging using an improved FCM with spatial constraints
Lianghua He, Ying Wen 0003, Meng Wan
Neurocomputing2
2013 A highly accurate, optical flow-based algorithm for nonlinear spatial normalization of diffusion tensor images
abstract
Spatial normalization plays a key role in voxel-based analyses of diffusion tensor images (DTI). We propose a highly accurate algorithm for high-dimensional spatial normalization of DTI data based on the technique of 3D optical flow. The theory of conventional optic flow assumes consistency of intensity and consistency of the gradient of intensity under a constraint of discontinuity-preserving spatio-temporal smoothness. By employing a hierarchical strategy ranging from coarse to fine scales of resolution and a method of Euler-Lagrange numerical analysis, our algorithm is capable of registering DTI data. Experiments using both simulated and real datasets demonstrated that the accuracy of our algorithm is better not only than that of those traditional optical flow algorithms or using affine alignment, but also better than the results using popular tools such as the statistical parametric mapping (SPM) software package. Moreover, our registration algorithm is fully automated, requiring a very limited number of parameters and no manual intervention.
Ying Wen 0003, Bradley S. Peterson, Dongrong Xu
IJCNN1
2012 Discriminative common vectors based on the Gram-Schmidt reorthogonalization for the small sample size problem
abstract
The discriminative common vectors (DCV) algorithm shows better face recognition effects than some commonly used linear discriminant algorithms, which uses the subspace methods and the Gram-Schmidt orthogonalization (GSO) procedure to obtain the DCV. However, the Gram-Schmidt technique may produce a set of vectors which is far from orthogonal so that sometimes the orthogonality may be lost completely. Hence, the effectiveness of the DCV is also decreased. In this paper, we proposed an improved DCV method based on the GSO. For obtaining an accurate projection onto the corresponding space, the orthogonal basis problem is usually solved with the Gram-Schmidt process with reorthogonalization. Thus, the effectiveness of the DCV can be improved and the experimental results show that the proposed method is better for the small sample size problem as compared to the DCV.
Ying Wen 0003, Lianghua He, Yue Lu 0001
ICASSP1
2012 A classifier for Bangla handwritten numeral recognition
Ying Wen 0003, Lianghua He
Expert Syst. Appl.1
2011 An Improved Locally Linear Embedding for Sparse Data Sets
abstract
Locally linear embedding is often invalid for sparse data sets because locally linear embedding simply takes the reconstruction weights obtained from the data space as the weights of the embedding space. This paper proposes an improved method for sparse data sets, a united locally linear embedding, to make the reconstruction more robust to sparse data sets. In the proposed method, the neighborhood correlation matrix presenting the position information of the points constructed from the embedding space is added to the correlation matrix in the original space, thus the reconstruction weights can be adjusted. As the reconstruction weights adjusted gradually, the position information of sparse points can also be changed continually and the local geometry of the data manifolds in the embedding space can be well preserved. Experimental results on both synthetic and real-world data show that the proposed approach is very robust against sparse data sets.
Ying Wen 0003, Lianghua He
Int. J. Pattern Recognit. Artif. Intell.1
2011 An Algorithm for License Plate Recognition Applied to Intelligent Transportation System
abstract
An algorithm for license plate recognition (LPR) applied to the intelligent transportation system is proposed on the basis of a novel shadow removal technique and character recognition algorithms. This paper has two major contributions. One contribution is a new binary method, i.e., the shadow removal method, which is based on the improved Bernsen algorithm combined with the Gaussian filter. Our second contribution is a character recognition algorithm known as support vector machine (SVM) integration. In SVM integration, character features are extracted from the elastic mesh, and the entire address character string is taken as the object of study, as opposed to a single character. This paper also presents improved techniques for image tilt correction and image gray enhancement. Our algorithm is robust to the variance of illumination, view angle, position, size, and color of the license plates when working in a complex environment. The algorithm was tested with 9026 images, such as natural-scene vehicle images using different backgrounds and ambient illumination particularly for low-resolution images. The license plates were properly located and segmented as 97.16% and 98.34%, respectively. The optical character recognition system is the SVM integration with different character features, whose performance for numerals, Kana, and address recognition reached 99.5%, 98.6%, and 97.8%, respectively. Combining the preceding tests, the overall performance of success for the license plate achieves 93.54% when the system is used for LPR in various complex conditions.
Ying Wen 0003, Yue Lu 0001, Jingqi Yan, Karen M. von Deneen
IEEE Trans. Intell. Transp. Syst.1
2007 Handwritten Bangla numeral recognition system and its application to postal automation
Ying Wen 0003, Yue Lu 0001
Pattern Recognit.1