VLDB 2026 Research / reviewers in the wild / expert
Liyan Ma
dblp:56/5572
· DBLP profile ↗
46ranked-venue papers
3as first author
38since 2021 · last 2027
0000-0001-8377-0163ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | URG-MRI: Unpaired reference guided MRI reconstruction
Xuanmin Chen, Jingtian Gu, Shihui Ying, Liyan Ma |
Expert Syst. Appl. | 5 |
| 2026 | Robust Group Activity Recognition via Hierarchical Topology Reasoning and Error-Aware Contrastive DenoisingabstractGroup Activity Recognition (GAR) aims to infer collective behaviors from the motions and interactions of multiple people in a scene. Compared with RGB-based pipelines, skeleton-based methods are more efficient and less sensitive to background appearance, but they usually rely on fixed graph topologies and are vulnerable to pose estimation noise. These two issues are especially harmful in crowded sports scenes, where interaction patterns change rapidly and upstream pose estimators may produce missing, swapped, or jittered joints. To address this problem, we propose a condensed presentation of our Contrastive Memory-Augmented Adaptive Graph Convolutional Network (CMAA-GCN). The framework is built on two key ideas. First, a Dual-Level Adaptive Spatial Graph Convolution (DASGC) module separates joint-level motion modeling from group-level interaction modeling and learns adaptive adjacency matrices for both levels. Second, an error-aware contrastive supervision mechanism simulates realistic pose failures and uses dual memory banks to align noisy and clean representations. This design improves both adaptability and robustness. Experiments on the reannotated Volleyball dataset, the original Volleyball dataset, the NBA benchmark, and Kinetics-400 show that the proposed method achieves strong accuracy while preserving the efficiency advantage of skeleton-based recognition. The model reaches 97.8% on reannotated Volleyball, 96.2% on original Volleyball, 76.1% on NBA, and 50.6% on Kinetics-400. Jinlong Lv, Liyan Ma, Xiangfeng Luo |
ICIC | 2 |
| 2026 | ConceptCap: Parameter-Efficient Image Captioning via Type-Aware Retrieval
Liyan Ma, Xiangfeng Luo, Ruoxin Zheng |
ICIC (12) | 2 |
| 2026 | ProDe: Prototype-based pattern decoupling for scene-dependent multi-modality video anomaly detection
Feiran Liu, Liyan Ma, Xiangfeng Luo |
Expert Syst. Appl. | 3 |
| 2026 | FedAPEX: Find flatter minima via federated accelerated perturbation exploration
Xuesong Chen 0003, Liyan Ma, Fang Li 0004 |
Neurocomputing | 3 |
| 2026 | GIGAS: Adversarial Attacks on Visual Question Answering With Multi-Modal Generative ModelsabstractVQA models, which answer questions about images by combining both visual and textual information, have been proven susceptible to adversarial attacks. These attacks introduce subtle perturbations to the input data to manipulate the model’s predictions. This paper focuses on adversarial attacks targeting VQA models that follow the “pre-training & fine-tuning” paradigm, an area that remains under-explored. We have identified two key issues in the current field. On one hand, existing multi-modal attacks have low ASR due to inter-modal semantic inconsistency from insufficient cross-modal interaction. On the other hand, the dilemma between attack effectiveness and stealthiness limits the practical applicability of adversarial texts. To address these issues, we propose GIGAS, an innovative attack that uses multi-modal generative models to explore multi-modal interaction through three key modules tailored to solve above-mentioned problems. MIGA aligns adversarial visual features with semantics of misleading images generated by multi-modal generative models to mitigate cross-modal inconsistencies. GSA employs MLLMs to generate natural adversarial texts with greater variation and evaluate similarity to filter based on clean images, balancing effectiveness and stealthiness. Iteration Allocation dynamically adjusts attack iterations based on image-text similarity, maximizing the utility of the limited iterations. Experiments conducted on various VL models and VQA datasets demonstrate superior attack performance, with an average ASR of 89.09% on VQAv2.0. Furthermore, our GIGAS exhibits outstanding transferability, around 60% ASR, across diverse models and specific domains. Our code will be available at: https://github.com/Yvonna-cloud/GIGAS. Yunxuan Li, Jing Yu 0007, Tieyong Zeng, Liyan Ma |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | MSAFed: Generalized Multi-Stage and Adaptive Federated Learning for Test-Time Medical SegmentationabstractFederated learning (FL) enables collaborative model training across multiple medical centers without sharing data, offering significant promise for privacy-preserving AI in healthcare. However, FL models often lack generalization across all participating clients (inside FL) and perform poorly when deployed to unseen clients (outside FL), particularly in heterogeneous domains. Current test-time adaptation methods for outside FL fail to address biases in personalized models toward source distributions, limiting their clinical applications. To tackle these challenges, we propose MSAFed, a generalized multi-stage adaptive FL framework that enhances both inside generalization and outside test-time adaptation. During pretraining, intra-client and inter-client contrastive learning with prototype-aware aggregation produces a generalized global model. An adaptive learning rate strategy further improves inside FL generalization. For unseen clients, source knowledge, including adaptive learning rates and prototypes, is leveraged to dynamically adapt the network architecture during test time. Experiments on three real-world multi-center medical datasets demonstrate the effectiveness of MSAFed, achieving superior performance on both inside and outside FL tasks. Jiajie Jin, Xuanmin Chen, Liyan Ma, Shihui Ying, Guang Yang 0006, Tieyong Zeng |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | Zero-Shot Scene Graph Generation with Bias Correction and Unseen Space Optimization
Yinsai Guo, Liyan Ma, Shaorong Xie |
ICIC (6) | 3 |
| 2025 | Dual-level Concept-aware Network for Video Captioning with Augmented SemanticsabstractVideo captioning aims to generate natural language descriptions of events in a video. While several methods have made efforts by utilizing detection modules to obtain principal concepts like objects and then capitalize on their semantics. Most of them lack prior semantic information about scenes and ignore the structural semantics of concept combinations, causing semantic deviation when generating captions in different scenes. To address these problems, we propose a method called Dual-level Concept-aware Network (DCN) for video captioning, which is capable of mining comprehensive semantics via two concept-aware modules. Specifically, the Concept Detection Module (CDM) is designed to detect concepts in diverse scenes precisely by exploiting scene semantic information from external prior knowledge according to video content. While the Concept Alignment Module (CAM) is constructed to represent concept combinations at multiple granularities by fusing visual features hierarchically and mapping them into the semantic space. With pseudo concept-related supervision signals extracted from ground-truth texts, both concept-aware modules get fully optimized explicitly. Hence they do well in obtaining useful semantic information about scenes and concept combinations when identifying concepts to augment caption semantics. Given competitive results on MSVD and MSR-VTT datasets, our method is demonstrated to generate more accurate and plausible captions. Sijia Lu, Yinsai Guo, Liyan Ma |
IJCNN | 4 |
| 2025 | Unsupervised adaptive learning method for salient object detection under weak observation conditions
Ying Tong, Xiangfeng Luo, Liyan Ma, Shaorong Xie |
Appl. Intell. | 3 |
| 2025 | Saliency and correlation learning for co-salient object detection
Ying Tong, Xiangfeng Luo, Liyan Ma, Shaorong Xie |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Co-saliency guided multi-modal learning for referring video object segmentation
Ying Tong, Xiangfeng Luo, Liyan Ma, Shaorong Xie |
Knowl. Based Syst. | 3 |
| 2025 | MPGB: Learning discriminative embeddings with multi-prototype and gradient balancing strategy for multi-modal 3D open world object detection
Liyan Ma, Zhi Li 0080, Tieyong Zeng |
Knowl. Based Syst. | 2 |
| 2025 | Open set label noise learning with robust sample selection and margin-guided module
Yuandi Zhao, Qianxi Xia, Zhijie Wen, Liyan Ma, Shihui Ying |
Knowl. Based Syst. | 5 |
| 2025 | Shadow defense against gradient inversion attack in federated learningabstractFederated learning (FL) has emerged as a transformative framework for privacy-preserving distributed training, allowing clients to collaboratively train a global model without sharing their local data. This is especially crucial in sensitive fields like healthcare, where protecting patient data is paramount. However, privacy leakage remains a critical challenge, as the communication of model updates can be exploited by potential adversaries. Gradient inversion attacks (GIAs), for instance, allow adversaries to approximate the gradients used for training and reconstruct training images, thus stealing patient privacy. Existing defense mechanisms obscure gradients, yet lack a nuanced understanding of which gradients or types of image information are most vulnerable to such attacks. These indiscriminate calibrated perturbations result in either excessive privacy protection degrading model accuracy, or insufficient one failing to safeguard sensitive information. Therefore, we introduce a framework that addresses these challenges by leveraging a shadow model with interpretability for identifying sensitive areas. This enables a more targeted and sample-specific noise injection. Specially, our defensive strategy achieves discrepancies of 3.73 in PSNR and 0.2 in SSIM compared to the circumstance without defense on the ChestXRay dataset, and 2.78 in PSNR and 0.166 in the EyePACS dataset. Moreover, it minimizes adverse effects on model performance, with less than 1% F1 reduction compared to SOTA methods. Our extensive experiments, conducted across diverse types of medical images, validate the generalization of the proposed framework. The stable defense improvements for FedAvg are consistently over 1.5% times in LPIPS and SSIM. It also offers a universal defense against various GIA types, especially for these sensitive areas in images. Liyan Ma, Guang Yang 0006 |
Medical Image Anal. | 2 |
| 2025 | LMS-Net: A learned Mumford-Shah network for binary few-shot medical image segmentation
Shengdong Zhang, Hao Zhang 0026, Jun Shi 0004, Liyan Ma, Shihui Ying |
Medical Image Anal. | 6 |
| 2024 | SRENet: Structure recovery ensemble network for single image deraining
Yingbing Xu, Liyan Ma, Yaoran Chen |
Appl. Intell. | 3 |
| 2024 | DP-DDCL: A discriminative prototype with dual decoupled contrast learning method for few-shot object detection
Yinsai Guo, Liyan Ma, Xiangfeng Luo, Shaorong Xie |
Knowl. Based Syst. | 2 |
| 2024 | Saliency information and mosaic based data augmentation method for densely occluded object recognition
Ying Tong, Xiangfeng Luo, Liyan Ma, Shaorong Xie, Yinsai Guo |
Pattern Anal. Appl. | 3 |
| 2024 | BSDP: Brain-inspired Streaming Dual-level Perturbations for Online Open World Object Detection
Liyan Ma, Liping Jing, Jian Yu 0001 |
Pattern Recognit. | 2 |
| 2024 | DSCA: A Dual Semantic Correlation Alignment Method for domain adaptation object detection
Yinsai Guo, Hang Yu 0006, Shaorong Xie, Liyan Ma, Xinzhi Cao, Xiangfeng Luo |
Pattern Recognit. | 4 |
| 2024 | DQDG: Data-Free Quantization With Dual Generators for Keyword SpottingabstractData-free quantization effectively compresses deep learning models with privacy guarantees. However, previous data-free quantization methods applied to keyword spotting models have the following issues: (1) The synthesized samples are excessively similar, leading to severe homogenization problems; (2) The low-quality samples during the initial training hinder model fine-tuning. To address these issues, this paper proposes a novel framework called Data-Free Quantization with Dual Generator (DQDG). Our framework introduces Dual Generators with Center Distance Constraint (DGCDC) to enhance the intra-class heterogeneity of synthesized samples, and utilizes a selector to select high-quality samples to assist in model fine-tuning. Additionally, allowing the quantized model to infer complete data from masked data, we adopt Time Masking Quantization Distillation (TMQD) to improve the understanding of the data distribution. Experimental results demonstrate that DQDG outperforms existing data-free quantization methods by a large margin. Xinbiao Xu, Liyan Ma, Fan Jia 0007, Tieyong Zeng |
IEEE Signal Process. Lett. | 2 |
| 2024 | DIE-CDK: A Discriminative Information Enhancement Method With Cross-Modal Domain Knowledge for Fine-Grained Ship DetectionabstractDue to the overarching similarities of ships, subtle information is imperative for fine-grained ship detection. However, this information is easily lost in adverse weather (e.g., fog, rain, snow, and cloud) or occlusion scenarios. Experts can quickly and accurately recognize fine-grained objects because they have the domain knowledge to help them find the most discriminative information (e.g., edge, structure, texture, and class semantics); thus, they do not need a lot of information to make an identification. Motivated by it, we propose a discriminative information enhancement method with cross-modal domain knowledge (DIE-CDK) for fine-grained ship detection. The core idea behind DIE-CDK is to enhance the discriminative information about fine-grained ships by fusing cross-modal domain knowledge. The introduced cross-modal domain knowledge comprises local and global knowledge: 1) local knowledge is the knowledge of visual shape (e.g., edge contour) which is extracted from the image domain; and 2) global knowledge is the knowledge of the class semantics which is obtained from the common sense domain. In addition, to further study fine-grained ship detection, we introduce a Fine-grained ship dataset (called FgShips). Experiments show that our proposed DIE-CDK method achieves impressive gains in detection performance and outperforms state-of-the-art methods on fine-grained ship and public datasets. Yinsai Guo, Hang Yu 0006, Liyan Ma, Xiangfeng Luo, Shaorong Xie |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | FEFA: Frequency Enhanced Multi-Modal MRI Reconstruction With Deep Feature AlignmentabstractIntegrating complementary information from multiple magnetic resonance imaging (MRI) modalities is often necessary to make accurate and reliable diagnostic decisions. However, the different acquisition speeds of these modalities mean that obtaining information can be time consuming and require significant effort. Reference-based MRI reconstruction aims to accelerate slower, under-sampled imaging modalities, such as T2-modality, by utilizing redundant information from faster, fully sampled modalities, such as T1-modality. Unfortunately, spatial misalignment between different modalities often negatively impacts the final results. To address this issue, we propose FEFA, which consists of cascading FEFA blocks. The FEFA block first aligns and fuses the two modalities at the feature level. The combined features are then filtered in the frequency domain to enhance the important features while simultaneously suppressing the less essential ones, thereby ensuring accurate reconstruction. Furthermore, we emphasize the advantages of combining the reconstruction results from multiple cascaded blocks, which also contributes to stabilizing the training process. Compared to existing registration-then-reconstruction and cross-attention-based approaches, our method is end-to-end trainable without requiring additional supervision, extensive parameters, or heavy computation. Experiments on the public fastMRI, IXI and in-house datasets demonstrate that our approach is effective across various under-sampling patterns and ratios. Xuanmin Chen, Liyan Ma, Shihui Ying, Dinggang Shen, Tieyong Zeng |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Memory-Based Cross-Modal Semantic Alignment Network for Radiology Report GenerationabstractGenerating radiology reports automatically reduces the workload of radiologists and helps the diagnoses of specific diseases. Many existing methods take this task as modality transfer process. However, since the key information related to disease accounts for a small proportion in both image and report, it is hard for the model to learn the latent relation between the radiology image and its report, thus failing to generate fluent and accurate radiology reports. To tackle this problem, we propose a memory-based cross-modal semantic alignment model (MCSAM) following an encoder-decoder paradigm. MCSAM includes a well initialized long-term clinical memory bank to learn disease-related representations as well as prior knowledge for different modalities to retrieve and use the retrieved memory to perform feature consolidation. To ensure the semantic consistency of the retrieved cross modal prior knowledge, a cross-modal semantic alignment module (SAM) is proposed. SAM is also able to generate semantic visual feature embeddings which can be added to the decoder and benefits report generation. More importantly, to memorize the state and additional information while generating reports with the decoder, we use learnable memory tokens which can be seen as prompts. Extensive experiments demonstrate the promising performance of our proposed method which generates state-of-the-art performance on the MIMIC-CXR dataset. Yitian Tao, Liyan Ma, Jing Yu 0007, Han Zhang 0002 |
IEEE J. Biomed. Health Informatics | 2 |
| 2023 | TriGait: Aligning and Fusing Skeleton and Silhouette Gait Data via a Tri-Branch NetworkabstractGait recognition is a promising biometric technology for identification due to its non-invasiveness and long-distance. However, external variations such as clothing changes and viewpoint differences pose significant challenges to gait recognition. Silhouette-based methods preserve body shape but neglect internal structure information, while skeleton-based methods preserve structure information but omit appearance. To fully exploit the complementary nature of the two modalities, a novel triple branch gait recognition framework, TriGait, is proposed in this paper. It effectively integrates features from the skeleton and silhouette data in a hybrid fusion manner, including a two-stream network to extract static and motion features from appearance, a simple yet effective module named JSA-TC to capture dependencies between all joints, and a third branch for cross-modal learning by aligning and fusing low-level features of two modalities. Experimental results demonstrate the superiority and effectiveness of TriGait for gait recognition. The proposed method achieves a mean rank-1 accuracy of 96.0% over all conditions on CASIA-B dataset and 94.3% accuracy for CL, significantly outperforming all the state-of-the-art methods. The source code will be available at https://github.com/feng-xueling/TriGait/. Xueling Feng, Liyan Ma, Long Hu, Mark S. Nixon |
IJCB | 3 |
| 2023 | Multi-scale and Multi-stage Deraining Network with Fourier Space Loss
Zhaoyong Yan, Liyan Ma, Xiangfeng Luo |
MMM (2) | 2 |
| 2023 | New Insights on the Generation of Rain Streaks: Generating-Removing United Unpaired Image Deraining Network
Zhaoyong Yan, Liyan Ma |
PRCV (11) | 3 |
| 2023 | Global Consistency Enhancement Network for Weakly-Supervised Semantic Segmentation
Liyan Ma |
PRCV (9) | 3 |
| 2023 | Object-Level Contrast Learning for 3D Sparse Object Detection in Ocean SceneabstractLiDAR-based 3D object detection provides the necessary high-precision environmental sensing information for the safe navigation of smart ships.However, relying on viewpoint projections, voxelized point clouds, or using inefficient point sampling methods, current LiDAR 3D object detection methods treat all objects uniformly and quantitatively while ignoring the specificity of sparse objects in the scene, which leaves less useful information about sparse objects.In this paper, we propose an end-to-end two-stage architecture, Object-Level Contrast Learning 3D Object Detection network (OCL), for better construction of sparse object features and improving the ability of model to detect sparse objects.In the first stage, the Contrast Learning based Sparse Object Feature Enhancement training strategy is proposed to decrease the feature discrepancy between sparse and regular objects in object-level.In the second stage, we use the Point-level Feature Multiple Aggregation Strategy to aggregate finer point-level features of sparse objects.Extensive experiments show that OCL achieves excellent performance on both Ship dataset and KITTI dataset.Furthermore, our work proposes a promising new idea for applying contrast learning to 3D object detection. Yuheng He, Xiangfeng Luo, Liyan Ma, Shaorong Xie |
SEKE | 4 |
| 2023 | Few-Shot Object Detection via Instance-wise and Prototypical Contrastive LearningabstractFew-shot object detection (FSOD), which involves training the detector with few annotated data to detect novel objects, has aroused a wide range of research interests.However, the performance of FSOD is still limited by insufficient data.Existing works usually adopt fine-tuning paradigm, which first uses rich base classes for pre-training and then uses them to carve the novel class feature space.In the fine-tuning phase, the balance space learned by the pre-trained model will be broken leading to an intersection between the feature space of novel and base classes, which makes it difficult to distinguish the difference between them.Contrastive learning has been shown to learn a balanced feature space and enhance the discriminability of the learned features.Here, we present Few-Shot object detection via Instance-wise and Prototype Contrastive Learning (FS-IPCL), which introduces contrastive learning to learn a balanced feature space.FS-IPCL uses instance-wise and prototype contrastive loss during feature learning to enhance the intra-class compactness and inter-class separability of samples.In this way, the base and novel classes can be evenly distributed in the feature space, improving the class boundary to alleviate the confusion problem of the novel classes.Extensive experimental results on the PASCAL VOC and MS-COCO datasets demonstrate the effectiveness of the proposed method and achieve state-of-theart performance. Qiaoning Lei, Yinsai Guo, Liyan Ma, Xiangfeng Luo |
SEKE | 3 |
| 2023 | THFE: A Triple-hierarchy Feature Enhancement method for tiny boat detection
Yinsai Guo, Hang Yu 0006, Liyan Ma, Xiangfeng Luo |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | Prior Semantic Harmonization Network for Few-Shot Semantic SegmentationabstractFew-shot semantic segmentation(FSS) is intended to segment a foreground object from a query image with a novel object using only a few annotated support images. Although attracting the attention of many researchers, this challenging problem remains to be not well solved due to two critical issues: (1)The information mismatching between support and query features leads to model distraction. (2)The key feature of query images is not activated well. In this paper, we introduce the Prior Semantic Harmonization Network(PSHNet) to tackle these limitations. PSHNet is composed of three effective modules. The Semantic Harmonization Module(SHM) corrects the information matching between support and query images, while the Feature Activation Module(FAM) activates the key feature of query images. Furthermore, we introduce a Hierarchical Aggregation Module(HAM) to refine each output of the multi-scale module. Experiments show that our model achieves an excellent performance on both PASCAL-5iand COCO-20idatasets. Liyan Ma, Yan Peng 0001, Shaorong Xie |
ICIP | 2 |
| 2022 | Open-World Object Detection via Discriminative Class Prototype LearningabstractOpen-world object detection (OWOD) is a challenging problem that combines object detection with incremental learning and open-set learning. Compared to standard object detection, the OWOD setting is task to: 1) detect objects seen during training while identifying unseen classes, and 2) incrementally learn the knowledge of the identified unknown objects when the corresponding annotations is available. We propose a novel and efficient OWOD solution from a prototype perspective, which we call OCPL: Open-world object detection via discriminative Class Prototype Learning, which consists of a Proposal Embedding Aggregator (PEA), an Embedding Space Compressor (ESC) and a Cosine Similarity-based Classifier (CSC). All our proposed modules aim to learn the discriminative embeddings of known classes in the feature space to minimize the overlapping distributions of known and unknown classes, which is beneficial to differentiate known and unknown classes. Extensive experiments performed on PASCAL VOC and MS-COCO benchmark demonstrate the effectiveness of our proposed method. Jinan Yu, Liyan Ma, Yan Peng 0001, Shaorong Xie |
ICIP | 2 |
| 2022 | A Transformer-based Cascade Network with Boundary Enhancement Loss for Retinal Vessel SegmentationabstractAnalyzing the retinal vessel structure is of great importance that it can assist doctors to diagnose eye diseases including hypertension and glaucoma. Existing deep learning methods for retinal vessel segmentation assign equal importance to all vessel pixels with the guidance of cross entropy loss. How-ever, the receptive field of the near-boundary vessel pixel samples and extremely thin vessel pixel samples contain pixels that belong to background and these pixels tend to be wrongly classified if do not penalize them more. Thus, we propose CasUTNet with a new boundary enhancement loss to segment retinal vessels accurately. Aim to pay more attention to thin capillaries and microvascular structure, we propose a new pixel-level boundary enhancement loss so that the neural network is able to learn more distinguishable features for these tinny vessels and boundary. In addition, We introduce a cascade structures based on UTNet to refine the segmentation result. The experiments on two retinal image databases: DRIVE and CHASEDB1 demonstrate the effectiveness of our proposed approach which obtains state-of-the-art performance on both retinal vessel datasets. Binke Cai, Liyan Ma |
ICPR | 2 |
| 2021 | Few-Shot Classification With Intra-Class Unrelated Multi-Prototype Representation and Episode Adaptation StrategyabstractThe episode training strategy, which trains models by many episodes to recognize unseen object categories using one or a few samples, is used by many existing approaches to solve the few-shot classification. An episode can be regarded as a classification task for the few-shot classification. However, this training strategy will be tricky when the task feature differences between the training episodes are big. To improve on this shortcomimg, we propose an Episode Adaptation Loss (EAL) to reduce these gaps for optimizing the models efficiently. Furthermore, to enhance the ability of each class feature representation and preserve diversities of the intra-class features, we propose Multi-Prototypical Representations (MPR) and a Multi-Prototype Unrelated Loss (MPUL). In addition, we introduce the multi-prototypes inductive inference as a classification strategy for our method. We evaluate our method on miniImageNet and Fewshot-CIFAR100 benchmarks. Experimental results demonstrate that our method outperforms the baseline approaches. The ablation study validates that the components of the proposed method all provide positive effects on few-shot learning. Zhijie Wen, Liyan Ma |
ICTAI | 4 |
| 2021 | Learning Discriminative Representations for Fine-Grained Diabetic Retinopathy GradingabstractDiabetic retinopathy is one of the leading causes of blindness. However, no specific symptoms of early DR lead to a delayed diagnosis, which results in disease progression in patients. To determine the disease severity levels, ophthalmologists need to focus on the discriminative parts of the retinal images. In recent years, deep learning has achieved great success in medical image analysis. However, most works directly employ algorithms based on convolutional neural networks (CNNs), which ignore the fact that the difference among classes is subtle and gradual. Hence, we consider automatic image grading of DR as a fine-grained classification task, and construct a bilinear model to identify the pathologically discriminative areas. In order to leverage the ordinal information among classes, we put the soft labels with ordinal information among classes into the loss function rather than the most commonly used one-hot labels for the diabetic retinopathy classification. In addition, other than only using a categorical loss to train our network, we also introduce the metric loss to learn a more discriminative feature space which is beneficial to locate the finer discriminative lesion parts. Experimental results demonstrate the superior performance of the proposed method on publicly available IDRiD, DeepDRiD and FGADR datasets. Liyan Ma, Zhijie Wen, Shaorong Xie, Yupeng Xu |
IJCNN | 2 |
| 2021 | Pixel-Attention CNN With Color Correlation Loss for Color Image DenoisingabstractConvolutional neural networks (CNNs) have been applied to many image processing tasks and achieve great successes. In order to extract common features, every pixel in an image shares the same filters. However, pixels in different regions of an image varies dramatically and shared filters may lose some important local information. Rather than shared filters, smart filters which can be adapted to image context should be designed to better remove noise which occurs randomly in noisy image. Meanwhile, current CNN architectures compute the loss of each color channel independently, regardless of the potential color information. In this letter, we proposed a pixel-attention convolutional neural network (PACNN) with color correlation loss for the color image denoising task. The pixel-attention mechanism could generate pixel-wise attention maps which help remove random noise. The color correlation loss exploits color correlation to further improve denoising performance on color noisy images. The experimental results on several standard datasets demonstrate the state-of-the-art (SOTA) performance and the superiority of the proposed method. Fan Jia 0007, Liyan Ma, Yijin Yang, Tieyong Zeng |
IEEE Signal Process. Lett. | 2 |
| 2020 | Context-Aware Hierarchical Feature Attention Network For Multi-Scale Object DetectionabstractMulti-scale object detection involves classification and regression assignments of objects with variable scales from an image. How to extract discriminative features is a key point for multi-scale object detection. Recent detectors simply fuse pyramidal features extracted from ConvNets, which does not take full advantage of useful features and drop out redundant features. To address this problem, we propose Context-Aware Hierarchical Feature Attention Network (CHFANet) to focus on effective multi-scale feature extraction for object detection. Based on single shot multibox detector (SSD) framework, the CHFANet consists of two components: the context-aware feature extraction (CFE) module to capture rich multi-scale context features and the hierarchical feature fusion (HFF) module followed with the channel-wise attention model to generate deeply fused attentive features. On the Pascal VOC benchmark, our CHFANet can achieve 82.6% mAP. Extensive experiments demonstrate that the CHFANet outperforms a lot of state-of-the-art object detectors in accuracy without any bells and whistles. Xuelong Xu, Xiangfeng Luo, Liyan Ma |
ICIP | 3 |
| 2020 | Realistic Style-Transfer Generative Adversarial Network With a Weight-Sharing StrategyabstractStyle transfer aims to generate images by combining the style of one image and the content of another. Though valuable efforts have been made in generating high-quality style transferred images, the resulting images are far from the distribution of real images. This greatly limits the application of style transfer such as improving the diversity of training set in computer vision task. We find that the reason of style transfer failing to generate realistic images is lack of reference targets and neglect of preserving data distribution. To solve the problem, we propose a Style-transfer Generative Adversarial Network with a weight-sharing strategy to make the stylized images be resemblance to the real images. The experimental results demonstrate that the proposed method can generate images with satisfying style transfers and high visual quality. Moreover, we apply our stylized images to augment the training set of object detection task, and improve the average precision faithfully. We believe that our method can enhance the performance of style transfer on computer vision tasks. Shixiong Zhu, Xiangfeng Luo, Liyan Ma, Shaorong Xie |
ICTAI | 3 |
| 2020 | Multi-metric Joint Discrimination Network for Few-Shot Classification
Zhijie Wen, Liyan Ma, Shihui Ying |
PRCV (3) | 3 |
| 2018 | Multi-image matching for object recognitionabstractOne of the central problems in object recognition is to develop appropriate representations for the objects in images. The authors present a novel approach for image representation that is based on graphs. In the proposed image graph, each node represents a patch and edges are added between neighbouring nodes. First, class‐specific match‐set graphs are generated by matching the image graphs that are in the same categories, and the multi‐image matching problem is solved by applying a seed‐expansion strategy. Then, the matches between the match‐set graphs and an image graph are considered to be the object patches in the image. Finally, the features extracted from these patches are used for the image representation. Extensive experiments are conducted to demonstrate that their approach can obtain state‐of‐the‐art results on several challenging datasets. Shufang Wu, Xizhao Wang, Liyan Ma |
IET Comput. Vis. | 5 |
| 2016 | Stereoscopic view synthesis based on region-wise rendering and sparse representation
Wei Liu 0023, Liyan Ma, Mingyue Cui |
Signal Process. Image Commun. | 2 |
| 2013 | Sparse Representation Prior and Total Variation-Based Image Deblurring under Impulse NoiseabstractIn this paper, we study the image recovery problem where the observed image is simultaneously corrupted by blur and impulse noise. Our proposed patch-based model contains three terms: the sparse representation prior, the total variation regularization, and the data-fidelity term. We are interested in the two-phase approach. The first phase is to identify the possible impulse noise positions; the second phase is to recover the image via the patch-based model using noise position information. An alternating minimization method is then applied to solve the model. This approach works extremely well for image deblurring under salt-and-pepper noise. However, as the detection for random-valued noise is usually unreliable, extra work is then needed. Indeed, to get better recovery results for the latter case, we combine the two separate phases to simultaneously detect the random-valued noise positions and to recover the image. The numerical experiments clearly demonstrate the super performance of the proposed methods. Liyan Ma, Jian Yu 0001, Tieyong Zeng |
SIAM J. Imaging Sci. | 1 |
| 2013 | A Dictionary Learning Approach for Poisson Image DeblurringabstractThe restoration of images corrupted by blur and Poisson noise is a key issue in medical and biological image processing. While most existing methods are based on variational models, generally derived from a maximum a posteriori (MAP) formulation, recently sparse representations of images have shown to be efficient approaches for image recovery. Following this idea, we propose in this paper a model containing three terms: a patch-based sparse representation prior over a learned dictionary, the pixel-based total variation regularization term and a data-fidelity term capturing the statistics of Poisson noise. The resulting optimization problem can be solved by an alternating minimization technique combined with variable splitting. Extensive experimental results suggest that in terms of visual quality, peak signal-to-noise ratio value and the method noise, the proposed algorithm outperforms state-of-the-art methods. Liyan Ma, Lionel Moisan, Jian Yu 0001, Tieyong Zeng |
IEEE Trans. Medical Imaging | 1 |
| 2011 | Texture segmentation based on local feature histogramsabstractThis paper presents a convex vector-valued active contour model for texture segmentation. This model uses histograms of the semi-local region descriptor and image intensity for measuring the similarity of image regions. We use the Quadratic-Chi histogram distance to compare the dissimilarity of histograms. Quadratic-Chi histogram distance is a cross-bin distance that matches perceptual similarity better than the bin-to-bin distance (such as Kullback-Leibler divergence and Bhattacharyya distance). Then we use a primal-dual method to solve the minimization problem. Experimental results for real images show the effective of the proposed method. Liyan Ma, Jian Yu 0001 |
ICIP | 1 |