EDBT 2026 Demo / reviewers in the wild / expert
Yixiong Liang
dblp:79/1649
· DBLP profile ↗
49ranked-venue papers
11as first author
34since 2021 · last 2026
0000-0003-0407-5838ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 4 first-author · 16 since 2021Artificial intelligence and machine learning · 17 · 6 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Ultrasound-based Reliable Disease Diagnosis Using Causal InferenceabstractAligning the decision-making process of deep learning models with that of experienced sonographers is essential for ultrasound-based reliable disease diagnosis. Although existing methods have made significant progress in this aspect, their alignments are primarily associational rather than causal, leading to pseudo-correlations between features and diagnostic results. Such a biased diagnosis blindly models the sonographer's diagnostic skills and attention to specific patterns, which we argue hardly produces an AI diagnoser that is comparable to human experts. To address this issue, we propose a causality-based diagnostic framework to align the model's diagnostic behaviors with those of experts. Specifically, by delving into both conspicuous and inconspicuous confounders within the ultrasound images, the back-door and front-door adjustment causal learning modules are proposed to promote unbiased learning by mitigating potential pseudo-correlations. In addition, we integrate causal inference into a well-designed dual-branch model with feature interaction bridges for compatibility with multimodal ultrasound inputs. To fully evaluate our method, we conduct comparative studies on different diseases and ultrasound modalities. In particular, we publish a carefully constructed multimodal ultrasound dataset for breast lesion diagnosis and segmentation. Sufficient comparative and ablation studies on this dataset emphasize that our method outperforms state-of-the-art methods. Bolei Chen, Jiaxu Kang, Haonan Yang 0001, Ping Zhong 0002, Yixiong Liang, Rui Fan 0001, Jianxin Wang 0001 |
AAAI | 5 |
| 2026 | MAPL: Enhancing Visual Prompt Encoding for Robust Open-Set Blood Cell Detection
Wenzhuo Xu, Jianfeng Liu 0001, Shichao Kan, Yixiong Liang |
ICIC (29) | 5 |
| 2026 | Integrating spatial features and dynamically learned temporal features via contrastive learning for video temporal grounding in LLM
Peifu Wang, Yixiong Liang, Yi-Gang Cen, Jin Liu 0012, Shichao Kan |
Image Vis. Comput. | 2 |
| 2025 | EPCPE: A Real-time End-to-End Pipeline for RGB-based Category-level 6D Pose EstimationabstractRGB-based category-level 6D pose estimation methods have faced significant challenges in achieving real-time performance, primarily due to the design of two-stage pipeline. To address this issue, we propose a novel end-to-end pipeline named EPCPE. In detail, we first extract implicit rotation features via Large Visual Model (LVM), and then adaptively obtain pose-specific features with a fined-tuned Lite Feature Extractor. Finally, we introduce a novel Pose Decoder with two parallel branches, enabling simultaneous 6D pose estimation and 2D object detection. We also propose a novel rotation loss function to further enhance the performance. Extensive experiments on the CAMERA25 and REAL275 datasets demonstrate that our pipeline is concise, achieves state-of-the-art (SOTA) and real-time performance. Xiaofeng Fan, Shichao Kan, Yixiong Liang |
ICASSP | 4 |
| 2025 | SAM 2-Driven Self-Training for Mammogram Segmentation: Zero-Shot Mask Generation Via Pseudo-VideoabstractAccurate mammogram segmentation is crucial for breast cancer diagnosis. However, existing Deep Learning methods often require large, annotated datasets, which are time-consuming and expensive to obtain. We introduce SAM 2-driven self-training, a novel approach that leverages SAM 2 for efficient and accurate mammogram segmentation. By constructing a pseudo-video sequence from static mammograms, we use SAM 2 video inference to generate initial masks, subsequently applying them for parameter-efficient adaptation of SAM, focusing on the mask decoder, and an automatic point prompt generator for enhanced usability. Our method significantly reduces the need for manual annotation while maintaining high accuracy. Evaluations on the mini-MIAS and CBIS-DDSM datasets demonstrate significant improvements in accuracy and efficiency compared to established techniques and the original SAM. This robust solution facilitates rapid mammogram segmentation and aids creating annotated datasets with minimal user intervention. The code is available at: https://github.com/MauricioFernandezM/Self-TrainingSAM. Mauricio Fernandez M, Yixiong Liang, Christopher A. Cochran |
ICIP | 2 |
| 2025 | Dynamic Residual Encoding with Slide-Level Contrastive Learning for End-to-End Whole Slide Image RepresentationabstractWhole Slide Image (WSI) representation is critical for cancer subtyping, cancer recognition and mutation prediction.Training an end-to-end WSI representation model poses significant challenges, as a standard gigapixel slide can contain tens of thousands of image tiles, making it difficult to compute gradients of all tiles in a single mini-batch due to current GPU limitations. To address this challenge, we propose a method of dynamic residual encoding with slide-level contrastive learning (DRE-SLCL) for end-to-end WSI representation. Our approach utilizes a memory bank to store the features of tiles across all WSIs in the dataset. During training, a mini-batch usually contains multiple WSIs. For each WSI in the batch, a subset of tiles is randomly sampled and their features are computed using a tile encoder. Then, additional tile features from the same WSI are selected from the memory bank. The representation of each individual WSI is generated using a residual encoding technique that incorporates both the sampled features and those retrieved from the memory bank. Finally, the slide-level contrastive loss is computed based on the representations and histopathology reports ofthe WSIs within the mini-batch. Experiments conducted over cancer subtyping, cancer recognition, and mutation prediction tasks proved the effectiveness of the proposed DRE-SLCL method. Te Gao, Zhihong Shi, Yixiong Liang, Ruiqing Zheng, Hulin Kuang, Min Zeng 0004, Shichao Kan |
ACM Multimedia | 5 |
| 2025 | Contrastive Regularization over LoRA for Multimodal Biomedical Image Incremental LearningabstractMultimodal Biomedical Image Incremental Learning (MBIIL) is essential for handling diverse tasks and modalities in the biomedical domain, as training separate models for each modality or task significantly increases inference costs. Existing incremental learning methods focus on task expansion within a single modality, whereas MBIIL seeks to train a unified model incrementally across modalities. The MBIIL faces two challenges: I) How to preserve previously learned knowledge during incremental updates? II) How to effectively leverage knowledge acquired from existing modalities to support new modalities? To address these challenges, we propose MSLoRA-CR, a method that fine-tunes Modality-Specific LoRA modules while incorporating Contrastive Regularization to enhance intra-modality knowledge sharing and promote inter-modality knowledge differentiation. Our approach builds upon a large vision-language model (LVLM), keeping the pretrained model frozen while incrementally adapting new LoRA modules for each modality or task. Experiments on the incremental learning of biomedical images demonstrate that MSLoRA-CR outperforms both the state-of-the-art (SOTA) approach of training separate models for each modality and the general incremental learning method (incrementally fine-tuning LoRA). Specifically, MSLoRA-CR achieves a 1.88% improvement in overall performance compared to unconstrained incremental learning methods while maintaining computational efficiency. Our code is publicly available at https://github.com/VentusAislant/MSLoRA_CR. Yixiong Liang, Hulin Kuang, Yi-Gang Cen, Min Zeng 0004, Shichao Kan |
ACM Multimedia | 2 |
| 2025 | A unified Personalized Federated Learning framework ensuring Domain Generalization
Yuan Liu 0038, Chengchao Shen, Yixiong Liang, Jianxin Wang 0001 |
Expert Syst. Appl. | 5 |
| 2024 | Aligning Multimodal Biomedical Images and Language via One Large Vision-Language ModelabstractLarge Vision-Language Models (LVLMs) have garnered substantial attention in the biomedical image analysis domain due to their robust vision understanding capabilities. However, current methods rely heavily on dataset- and modality-specific fine-tuning. This involves tuning separate models for each dataset and biomedical modality. In this paper, we introduce a method for aligning multimodal biomedical images and language using a single LVLM, dubbed UniMed-LVLM. Specifically, we devise a General Projection Module (GPM) by integrating multiple image projection branches and implementing dynamic routing between the vision encoder and language decoder within the LLaVA-Med framework. Subsequently, we progressively align multiple biomedical modalities using a Parameter-Efficient Fine-Tuning (PEFT) technique known as Low-Rank Adaptation (LoRA). The model is initially trained on the LLaVA-Med dataset and then fine-tuned on four biomedical image analysis datasets: PathVqa, Slake, VqaRad, and Fitzpatrick17k, enabling the simultaneous analysis of radiology, pathology, and dermatology images. A single model is fine-tuned on three modalities across these datasets and evaluated on all test sets. Experimental results show that UniMed-LVLM improves the average evaluation score by 1.88% across the four datasets, validating its effectiveness in handling multimodal biomedical images. Min Zeng 0004, Jinfeng Ding, Yixiong Liang, Ruiqing Zheng, Min Li 0007, Shichao Kan |
BIBM | 4 |
| 2024 | FedMMR: Multi-Modal Federated Learning via Missing Modality ReconstructionabstractFederated Learning (FL) presents a decentralized learning approach for privacy preservation. While many research focuses on uni-modal FL, a more intricate version, multimodal FL, uncovers fundamental attributes. Clients might only gather specific modalities, incurring missing modalities in multi-modal FL. The missing modality problem incurs modality heterogeneity among clients and loses inter-modal connection within one client. To address these issues, we introduce a novel multi-modal FL method called Federated Missing Modality Reconstruction (FedMMR), which tackles these dual challenges through two distinct facets. First, we devise a cross-modal reconstruction policy for synthesizing absent modalities from other modalities. This aligns the data feature spaces across clients, thereby alleviating bias in resultant local models. Subsequently, we steer local models to maintain an inter-modal awareness of both existing and reconstructed modalities, recognizing their potential as complementary components. Comprehensive results show that FedMMR outperforms existing FL baselines. Yuan Liu 0038, Shichao Kan, Yixiong Liang, Jianxin Wang 0001 |
ICME | 5 |
| 2024 | Task-Aware Transformer For Partially Supervised Retinal Fundus Image SegmentationabstractThe segmentation of retinal fundus images plays a critical role in ophthalmic disease screening and diagnosis. Due to the considerable human labor and expertise required for data annotation, fundus image datasets are often only partially labeled. Existing approaches typically employ separate networks for each specific task, resulting in resource-intensive processes and limited generalization capability. To address this, we propose a novel task-aware transformer (TAFormer) that learns to segment lesions and anatomic structures on multiple partially labeled datastes. Our method is built upon the query-based learning approaches. Specifically, we propose a task-specific query grouping strategy, which enables each group of queries to learn structural information related to the corresponding task. Queries from different groups can then capture relationships between the current task and other tasks through self-attention mechanisms, thereby enhancing their own representations. Moreover, recognizing that the classification task and mask prediction task involve distinct features, we choose to decouple these tasks. This strategic separation enables each query to focus on the most relevant features for its respective task. To address challenges stemming from partial label absence and domain shift, we incorporate pseudo-labels as an additional supervision signal. Experimental results obtained from partially supervised segmentation tasks on retinal fundus images validate the effectiveness of TAFormer, surpassing the performance of state-of-the-art methods. Hailong Zeng, Jianfeng Liu 0001, Yixiong Liang |
IJCNN | 3 |
| 2024 | HRDecoder: High-Resolution Decoder Network for Fundus Image Lesion Segmentation
Ziyuan Ding, Yixiong Liang, Shichao Kan, Qing Liu 0003 |
MICCAI (9) | 2 |
| 2024 | DES-SAM: Distillation-Enhanced Semantic SAM for Cervical Nuclear Segmentation with Box Annotation
Lina Huang, Yixiong Liang, Jianfeng Liu 0001 |
MICCAI (9) | 2 |
| 2024 | Embodied Contrastive Learning with Geometric Consistency and Behavioral Awareness for Object NavigationabstractObject Navigation (ObjcetNav), which enables an agent to seek any instance of an object category specified by a semantic label, has shown great advances. However, current agents are built upon occlusion-prone visual observations or compressed 2D semantic maps, which hinder their embodied perception of 3D scene geometry and easily lead to ambiguous object localization and blind exploration. To address these limitations, we present an Embodied Contrastive Learning (ECL) method with Geometric Consistency (GC) and Behavioral Awareness (BA), which motivates agents to actively encode 3D scene layouts and semantic cues. Driven by our embodied exploration strategy, BA is modeled by predicting navigational actions based on multi-frame visual images, as behaviors that cause differences between adjacent visual sensations are crucial for learning correlations among continuous visions. The GC is modeled as the alignment of behavior-aware visual stimulus with 3D semantic shapes by employing unsupervised contrastive learning. The aligned behavior-aware visual features and geometric invariance priors are injected into a modular ObjectNav framework to enhance object recognition and exploration capabilities. As expected, our ECL method performs well on object detection and instance segmentation tasks. Our ObjectNav strategy outperforms state-of-the-art methods on MP3D and Gibson datasets, showing the potential of our ECL in embodied navigation. Bolei Chen, Jiaxu Kang, Ping Zhong 0002, Yixiong Liang, Yu Sheng, Jianxin Wang 0001 |
ACM Multimedia | 4 |
| 2024 | SemNav-HRO: A target-driven semantic navigation strategy with human-robot-object ternary fusion
Bolei Chen, Siyi Lu, Ping Zhong 0002, Yongzheng Cui, Yixiong Liang, Jianxin Wang 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Many birds, one stone: Medical image segmentation with multiple partially labeled datasets
Qing Liu 0003, Hailong Zeng, Zhaodong Sun, Guoying Zhao 0001, Yixiong Liang |
Pattern Recognit. | 6 |
| 2024 | Think Holistically, Act Down-to-Earth: A Semantic Navigation Strategy With Continuous Environmental Representation and Multi-Step Forward PlanningabstractThe Object goal Navigation (ObjectNav) task requires an agent to navigate through a previously unknown domestic scenario using spatial and semantic contextual information, where the goal is specified by a semantic label (e.g., find a TV). Such a task is especially challenging as it requires formulating and understanding the complex co-occurrence relations among objects in diverse settings, which is critical for long-sequence navigational decision-making. Existing methods learn to either explicitly represent co-occurrence relationships as discrete semantic priors, or implicitly encode them from raw observations, thus can not benefit from the rich environmental semantics. In this work, we propose a novel Deep Reinforcement Learning (DRL) based ObjectNav strategy by actively imagining spatial and semantic clues outside the agent’s Field of View (FoV) and further mining Continuous Environmental Representations (CER) using self-supervised learning. Additionally, the illusion of spatial and semantic patterns allows the agent to perform Multi-Step Forward-Looking Planning (MSFLP) by considering the temporal evolution of egocentric local observations. Our approach is thoroughly evaluated and ablated in the visually realistic environments of the Matterport3D (MP3D) dataset. The experimental results reflect that our method combining CER and imagination-based MSFLP facilitates learning complicated semantic priors and navigation skills, thus achieving state-of-the-art performance on the ObjectNav task. In addition, adequate quantitative and qualitative analyses validate the excellent generalization ability and superiority of our method. Bolei Chen, Jiaxu Kang, Ping Zhong 0002, Yongzheng Cui, Siyi Lu, Yixiong Liang, Jianxin Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | CCBox: Improving Box-supervised Nuclei Segmentation with Consistency ConstraintabstractNuclei segmentation is a critical step in the automated analysis of digitized microscopic images, which facilitates analysis of pathological images. The current state-of-the-art (SOTA) methods for nuclei segmentation require significant time and resources to provide pixel-level annotations for training. To reduce the labor-intensive annotation cost, we propose CCBox, a high-quality nuclei segmentation network that only requires bounding box annotations. We first generate hard pseudo labels for each nuclei within the bounding boxes using the traditional methods and then train a nuclei segmentation network with these hard pseudo labels. The major challenge lies in the presence of significant noise in the boundary regions of the nuclei due to the susceptibility of traditional methods to complex texture information in pathological images, impacting the model’s segmentation performance at the nucleus boundaries. To address this challenge, we propose the consistency constraint for similarity maps (CSM) strategy, which aggregates pixel features of the same semantics and enhances the discriminability of foreground and background features at the nucleus boundaries. Furthermore, to mitigate the overfitting of the model to noisy samples in the hard pseudo labels, we employ the hard-soft joint supervision strategy to supervise the student network. Extensive experiment results demonstrate that our CCBox significantly narrows the gap between box-supervised and fully-supervised nuclei segmentation methods. Chaojun Zhang, Yixiong Liang, Qing Liu 0003 |
BIBM | 2 |
| 2023 | A Novel Transformer-Based Pipeline for Lung Cytopathological Whole Slide Image ClassificationabstractWe propose a novel three-stage Transformer-based methodology for entire cytopathological whole slide image (WSI) classification. The key idea is to leverage Transformer to extract the fine-grained lesion-level features and then progressively aggregate them into intermediate-grained patch-level features and coarse-grained WSI-level features for classification. Specifically, we first extract multi-scale lesion features from each patch image via Transformer-based lesion detection, and then adaptively aggregate the extracted lesion features into the corresponding patch feature with an MLP-Mixer. Finally, we select the most representative patch features and feed them into the Vision Transformer (ViT) for the final WSI-level classification. We collect a dataset consisting of 961 lung cytopathological WSIs of pleural effusions cytology specimens and conduct extensive experiments on it. The experimental results demonstrate that the proposed method outperforms existing state-of-the-art (SOTA) methods for cytopathological WSI classification. Gaojie Li, Qing Liu 0003, Yixiong Liang |
ICASSP | 4 |
| 2023 | Exploring Effective Knowledge Distillation for Tiny Object DetectionabstractDetecting tiny objects is a long-standing and critical problem in object detection, with broad real-world applications such as autonomous driving, surveillance, and medical diagnosis. Recent studies for tiny object detection often cause extra computational costs during inference due to introducing feature maps with increased resolution or additional network modules. This scarifies the inference speed for better detection accuracy and may heavily limit their availability to real-world applications. Therefore, this paper turns to knowledge distillation to improve the representation learning of a small model regarding both superior detection accuracy and fast inference speed. The masked scale-aware feature distillation and local attention distillation are proposed to address the critical issues in the distillation of tiny objects. Experimental results on two tiny benchmarks indicate that our method can bring noticeable performance gains to different detectors while keeping their original inference speeds. Our method also shows competitive performance compared to state-of-the-art methods for tiny object detection. Our code is available at https://github.com/haotianll/TinyKD. Qing Liu 0003, Yang Liu 0182, Yixiong Liang, Guoying Zhao 0001 |
ICIP | 4 |
| 2023 | Confidence-Aware Contrastive Learning for Semantic SegmentationabstractRecently supervised contrastive learning (SCL) has achieved remarkable progress in semantic segmentation. Nevertheless, prior works have often necessitated a substantial number of samples to attain satisfactory performance, leading to a significant increase in training overhead. In this work, we leverage the idea of reweighting each pair to reduce the demand for large numbers of training samples in contrastive learning and propose a novel loss, dubbed confidence-aware contrastive (CAC) loss, which adaptively reweights each pair according to the predicted confidence for semantic segmentation. To alleviate the misalignment between supervised learning and contrastive learning, we further introduce an extra weight branch with a stop-gradient operator to generate the pair weights. Moreover, we present a confidence-aware marginal anchor sampling method for the calculation of supervised contrastive loss which focuses on marginal rather than the hardest pairs. Coupling with our method consistently improves the performance of various models (e.g. HRNet, OCRNet, SegFormer) on Cityscapes, ADE20K, PASCAL-Context, and COCO-Stuff datasets. Compared to existing SCL-based methods, the proposed method achieves competitive or even better results without relying on a memory bank or a large number of samples. Our code is at https://github.com/CVIU-CSU/Confidence-Aware-Contrastive-Loss. Lele Lv, Qing Liu 0003, Shichao Kan, Yixiong Liang |
ACM Multimedia | 4 |
| 2023 | Adaptive Cluster Assignment for Unsupervised Semantic Segmentation
Shengqi Li, Qing Liu 0003, Chaojun Zhang, Yixiong Liang |
PRCV (4) | 4 |
| 2023 | Automated lesion segmentation in fundus images with many-to-many reassembly of features
Qing Liu 0003, Wei Ke 0003, Yixiong Liang |
Pattern Recognit. | 4 |
| 2023 | Modeling global distribution for federated learning with label distribution skew
Chengchao Shen, Yuan Liu 0038, Yeyu Ou, Yixiong Liang, Jianxin Wang 0001 |
Pattern Recognit. | 6 |
| 2023 | Weakly Supervised Deep Nuclei Segmentation With Sparsely Annotated Bounding Boxes for DNA Image CytometryabstractNuclei segmentation is an essential step in DNA ploidy analysis by image-based cytometry (DNA-ICM) which is widely used in cytopathology and allows an objective measurement of DNA content (ploidy). The routine fully supervised learning-based method requires often tedious and expensive pixel-wise labels. In this paper, we propose a novel weakly supervised nuclei segmentation framework which exploits only sparsely annotated bounding boxes, without any segmentation labels. The key is to integrate the traditional image segmentation and self-training into fully supervised instance segmentation. We first leverage the traditional segmentation to generate coarse masks for each box-annotated nucleus to supervise the training of a teacher model, which is then responsible for both the refinement of these coarse masks and pseudo labels generation of unlabeled nuclei. These pseudo labels and refined masks along with the original manually annotated bounding boxes jointly supervise the training of student model. Both teacher and student share the same architecture and especially the student is initialized by the teacher. We have extensively evaluated our method with both our DNA-ICM dataset and public cytopathological dataset. Without bells and whistles, our method outperforms all existing weakly supervised entries on both datasets. Code and our DNA-ICM dataset are publicly available at https://github.com/CVIU-CSU/Weakly-Supervised-Nuclei-Segmentation. Yixiong Liang, Zhihua Yin, Hailong Zeng, Jianxin Wang 0001, Jianfeng Liu 0001, Nanying Che |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2023 | STExplorer: A Hierarchical Autonomous Exploration Strategy with Spatio-temporal Awareness for Aerial RobotsabstractThe autonomous exploration task we consider requires Unmanned Aerial Vehicles (UAVs) to actively navigate through unknown environments with the goal of fully perceiving and mapping the environments. Some existing exploration strategies suffer from rough cost budgets, ambiguous Information Gain (IG), and unnecessary backtracking exploration caused by Fragmented Regions (FRs). In our work, a hierarchical spatio-temporal-aware exploration framework is proposed to alleviate these problems. At the local exploration level, the Asymmetrical Traveling Salesman Problem (ATSP) is solved by comprehensively considering exploration time, IG, and heading consistency to avoid blindly exploring. Specifically, the exploration time is reasonably budgeted by fast marching in an artificial potential field. Meanwhile, a transformer-based map occupancy predictor is designed to assist in IG calculation by imagining spatial clues out of the Field of View (FoV), facilitating the prescient exploration. We verify that our local exploration is effective in alleviating the unnecessary back-and-forth movements caused by FRs and the interference of potential obstacle occlusion on the IG calculation. At the global exploration level, the classical Next Best View Points (NBVP) are generalized to Next Best Sub-Regions (NBSR) to choose informative sub-regions for further forward-looking exploration based on a well-designed utility function. Safe flight paths and dynamically feasible trajectories are reasonably generated throughout the exploration process by fast marching and B-spline curve optimization. Comparative simulations and benchmark tests demonstrate that our proposed exploration strategy is quite competitive in terms of exploration path length, total exploration time, and exploration ratio. Bolei Chen, Yongzheng Cui, Ping Zhong 0002, Wang Yang 0002, Yixiong Liang, Jianxin Wang 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2023 | BEA-Net: Body and Edge Aware Network With Multi-Scale Short-Term Concatenation for Medical Image SegmentationabstractMedical image segmentation is indispensable for diagnosis and prognosis of many diseases. To improve the segmentation performance, this study proposes a new 2D body and edge aware network with multi-scale short-term concatenation for medical image segmentation. Multi-scale short-term concatenation modules which concatenate successive convolution layers with different receptive fields, are proposed for capturing multi-scale representations with fewer parameters. Body generation modules with feature adjustment based on weight map computing via enlarging the receptive fields, and edge generation modules with multi-scale convolutions using Sobel kernels for edge detection, are proposed to separately learn body and edge features from convolutional features in decoders, making the proposed network be body and edge aware. Based on the body and edge modules, we design parallel body and edge decoders whose outputs are fused to achieve the final segmentation. Besides, deep supervision from the body and edge decoders is applied to ensure the effectiveness of the generated body and edge features and further improve the final segmentation. The proposed method is trained and evaluated on six public medical image segmentation datasets to show its effectiveness and generality. Experimental results show that the proposed method achieves better average Dice similarity coefficient and 95% Hausdorff distance than several benchmarks on all used datasets. Ablation studies validate the effectiveness of the proposed multi-scale representation learning modules, body and edge generation modules and deep supervision. Hulin Kuang, Yixiong Liang, Jin Liu 0012, Jianxin Wang 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Exploring Contextual Relationships for Cervical Abnormal Cell DetectionabstractCervical abnormal cell detection is a challenging task as the morphological discrepancies between abnormal and normal cells are usually subtle. To determine whether a cervical cell is normal or abnormal, cytopathologists always take surrounding cells as references to identify its abnormality. To mimic these behaviors, we propose to explore contextual relationships to boost the performance of cervical abnormal cell detection. Specifically, both contextual relationships between cells and cell-to-global images are exploited to enhance features of each region of interest (RoI) proposal. Accordingly, two modules, dubbed as RoI-relationship attention module (RRAM) and global RoI attention module (GRAM), are developed and their combination strategies are also investigated. We establish a strong baseline by using Double-Head Faster R-CNN with a feature pyramid network (FPN) and integrate our RRAM and GRAM into it to validate the effectiveness of the proposed modules. Experiments conducted on a large cervical cell detection dataset reveal that the introduction of RRAM and GRAM both achieves better average precision (AP) than the baseline methods. Moreover, when cascading RRAM and GRAM, our method outperforms the state-of-the-art (SOTA) methods. Furthermore, we show that the proposed feature-enhancing scheme can facilitate image- and smear-level classification. Yixiong Liang, Qing Liu 0003, Hulin Kuang, Jianfeng Liu 0001, Liyan Liao, Jianxin Wang 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | Learning Deep Pathological Features for WSI-Level Cervical Cancer GradingabstractFully automated cervical cancer grading on the level of Whole Slide Images (WSI) is a challenge task. As WSIs are in gigapixel resolution, it is impossible to train a deep classification neural network with the entire WSIs as inputs. To bypass this problem, we propose a two-stage learning framework. In detail, we propose to first learn patch-level deep pathological features for smear patches via a patch-level feature learning module, which is trained via leveraging the cell instance detection task. Then, we propose to learn WSI-level pathological features from patch-level features for cervical cancer grading. We conduct extensive experiments on our private dataset and make comparisons with rule-based cervical cancer grading methods. Experimental results demonstrate that our proposed deep feature-based WSI-level cervical cancer grading method achieves state-of-the-art performance. Ruixiang Geng, Qing Liu 0003, Yixiong Liang |
ICASSP | 4 |
| 2022 | MEJIGCLU: More Effective Jigsaw Clustering For Unsupervised Visual Representation LearningabstractUnsupervised visual representation learning aims to learn general features from unlabelled data. Early methods design intra-image pretext tasks as learning targets and can be achieved with low computational overhead but unsatisfactory performance. Recent methods introduce contrastive learning and achieve surprising performance, but multiple views of training data are required in one batch, resulting in high computational overhead. To achieve competitive results to contrastive learning with low computational overhead, we propose a new unsupervised representation learning method with jigsaw clustering and classification as pretext tasks motivate the network to learn discriminative feature. To increase the data diversity, we propose to partition each training image into patches with random overlap, then randomly permute and stitch them into new training batch. Comparing with SOTAs, our method achieves state-of-the-art performance on both image classification/semi-classification on ImageNet and object detection on COCO. Qing Liu 0003, Yixiong Liang |
ICASSP | 4 |
| 2022 | Coded Residual Transform for Generalizable Deep Metric LearningabstractA fundamental challenge in deep metric learning is the generalization capability of the feature embedding network model since the embedding network learned on training classes need to be evaluated on new test classes. To address this challenge, in this paper, we introduce a new method called coded residual transform (CRT) for deep metric learning to significantly improve its generalization capability. Specifically, we learn a set of diversified prototype features, project the feature map onto each prototype, and then encode its features using their projection residuals weighted by their correlation coefficients with each prototype. The proposed CRT method has the following two unique characteristics. First, it represents and encodes the feature map from a set of complimentary perspectives based on projections onto diversified prototypes. Second, unlike existing transformer-based feature representation approaches which encode the original values of features based on global correlation analysis, the proposed coded residual transform encodes the relative differences between the original features and their projected prototypes. Embedding space density and spectral decay analysis show that this multi perspective projection onto diversified prototypes and coded residual representation are able to achieve significantly improved generalization capability in metric learning. Finally, to further enhance the generalization performance, we propose to enforce the consistency on their feature similarity matrices between coded residual transforms with different sizes of projection prototypes and embedding dimensions. Our extensive experimental results and ablation studies demonstrate that the proposed CRT method outperform the state-of-the-art deep metric learning methods by large margins and improving upon the current best method by up to 4.28% on the CUB dataset. Shichao Kan, Yixiong Liang, Min Li 0007, Yi-Gang Cen, Jianxin Wang 0001, Zhihai He |
NeurIPS | 2 |
| 2022 | Dual-Branch Network With Dual-Sampling Modulated Dice Loss for Hard Exudate Segmentation in Color Fundus ImagesabstractAutomated segmentation of hard exudates in colour fundus images is a challenge task due to issues of extreme class imbalance and enormous size variation. This paper aims to tackle these issues and proposes a dual-branch network with dual-sampling modulated Dice loss. It consists of two branches: large hard exudate biased segmentation branch and small hard exudate biased segmentation branch. Both of them are responsible for their own duties separately. Furthermore, we propose a dual-sampling modulated Dice loss for the training such that our proposed dual-branch network is able to segment hard exudates in different sizes. In detail, for the first branch, we use a uniform sampler to sample pixels from predicted segmentation mask for Dice loss calculation, which leads to this branch naturally be biased in favour of large hard exudates as Dice loss generates larger cost on misidentification of large hard exudates than small hard exudates. For the second branch, we use a re-balanced sampler to oversample hard exudate pixels and undersample background pixels for loss calculation. In this way, cost on misidentification of small hard exudates is enlarged, which enforces the parameters in the second branch fit small hard exudates well. Considering that large hard exudates are much easier to be correctly identified than small hard exudates, we propose an easy-to-difficult learning strategy by adaptively modulating the losses of two branches. We evaluate our proposed method on two public datasets and the results demonstrate that ours achieves state-of-the-art performance. Qing Liu 0003, Yixiong Liang |
IEEE J. Biomed. Health Informatics | 4 |
| 2021 | BEA-SegNet: Body and Edge Aware Network for Medical Image SegmentationabstractMedical image segmentation is a fundamental step for diagnosis and prognosis. This study proposes a new body and edge aware network for automated 2D medical image segmentation (called BEA-SegNet). The proposed BEA-SegNet consists of a shared encoder, a body and edge decouple (BEdecouple) module, two parallel decoders for body and edge segmentation. In the encoder and decoders, short-term multi-scale concatenation (STMSC) modules are utilized to implement multi-scale representation. We design a BEdecouple module to decouple the convolutional features into the body and edge features, making the proposed method be body and edge aware. The body and edge decoders utilize Bedecouple modules in each level to learn more effective features for the body and edge segmentation respectively, and their outputs are fused to generate the final segmentation. Besides, the body and edge supervision are applied to improve the final segmentation. The proposed BEA-SegNet is trained and evaluated on the International Skin Imaging Collaboration challenge 2018 dataset (ISIC2018). Experimental results show that the proposed BEA-SegNet achieves an average Dice similarity coefficient of 90.3% and an average Hausdorff distance of 15.9 for the skin lesion segmentation task and outperforms five benchmarks for skin lesion segmentation. Hulin Kuang, Yixiong Liang, Jin Liu 0012, Jianxin Wang 0001 |
BIBM | 2 |
| 2021 | Comparison detector for cervical cell/clumps detection in the limited data scenario
Yixiong Liang, Zhihong Tang, Meng Yan 0009, Qing Liu 0003, Yao Xiang |
Neurocomputing | 1 |
| 2020 | A Deep Gradient Boosting Network for Optic Disc and Cup SegmentationabstractSegmentation of optic disc (OD) and optic cup (OC) is critical in automated fundus image analysis system. Existing state-of-the-arts focus on designing deep neural networks with one or multiple dense prediction branches. Such kind of designs ignore connections among prediction branches and their learning capacity is limited. To build connections among prediction branches, this paper introduces gradient boosting framework to deep classification model and proposes a gradient boosting network called BoostNet. Specifically, deformable side-output unit and aggregation unit with deep supervisions are proposed to learn base functions and expansion coefficients in gradient boosting framework. By stacking aggregation units in a deep-to-shallow manner, models’ performances are gradually boosted along deep to shallow stages. BoostNet achieves superior results to existing deep OD and OC segmentation networks on the public dataset ORIGA. Qing Liu 0003, Beiji Zou 0001, Yixiong Liang |
ICASSP | 4 |
| 2020 | A Bidirectional Context Propagation Network for Urine Sediment Particle Detection in Microscopic ImagesabstractThe microscopic urine sediment examination is a crucial part in the evaluation of renal and urinary tract diseases. Recently, there are emerging CNNs-based detectors to detect the urine sediment particles in an end-to-end manner. However, it is not very compatible to transfer CNNs-based detector directly from natural images application to microscopic images, especially in which small objects are in majority. This paper proposes a bidirectional context propagation network called BCPNet for urine sediment particle detection. In BCPNet, spatial details encoded by shallow convolutional layers are propagated upward to improve the localisation ability of deep features. On the contrary, high semantic information encoded by deep convolutional layers is propagated downward to enhance the distinctiveness of shallow features. With the refinement by convolutional block attention modules, the enriched features are more powerful to both localisation and classification. Experimental results on urine sediment particle dataset USE demonstrate effectiveness of the proposed BCPNet. Meng Yan 0009, Qing Liu 0003, Zhihua Yin, Du Wang, Yixiong Liang |
ICASSP | 5 |
| 2020 | Disentangled Representation Learning Based Multidomain Stain Normalization For Histological ImagesabstractColor variations of histological images due to multi-factor hinder the performance of computer-aided diagnosis (CAD) systems. Previous stain normalization methods have achieved excellent results. While in practice, a multidomain stain normalization method is still be needed when more than two color variations exist in dataset. In this paper, we propose a multidomain stain normalization model inspired by MUNIT [1], with the idea of disentangling the representations of content and style. We assume that the latent space of histological images can be decomposed into domain-shared content space and domain-specific style space. The stain normalization aims to transfer the styles cross domains and maintain the contents. In addition, we propose to use the earth mover’s distance(EMD) to evaluate the effectiveness of stain normalization. We evaluate our approach against the state-of-the-art methods quantitatively and qualitatively. Yao Xiang, Qing Liu 0003, Yixiong Liang |
ICIP | 4 |
| 2019 | Scale-invariant structure saliency selection for fast image fusion
Yixiong Liang, Yuan Mao, Jiazhi Xia, Yao Xiang, Jianfeng Liu 0001 |
Neurocomputing | 1 |
| 2019 | Combining static and dynamic features for real-time moving pedestrian detection
Ying-Jun Jiang, Jianxin Wang 0001, Yixiong Liang, Jiazhi Xia |
Multim. Tools Appl. | 3 |
| 2019 | Efficient misalignment-robust multi-focus microscopical images fusion
Yixiong Liang, Yuan Mao, Zhihong Tang, Meng Yan 0009, Jianfeng Liu 0001 |
Signal Process. | 1 |
| 2017 | Automatic Retinal Image Registration Using Blood Vessel Segmentation and SIFT FeatureabstractAutomatic retinal image registration is still a great challenge in computer aided diagnosis and screening system. In this paper, a new retinal image registration method is proposed based on the combination of blood vessel segmentation and scale invariant feature transform (SIFT) feature. The algorithm includes two stages: retinal image segmentation and registration. In the segmentation stage, the blood vessel is segmented by using the guided filter to enhance the vessel structure and the bottom-hat transformation to extract blood vessel. In the registration stage, the SIFT algorithm is adopted to detect the feature of vessel segmentation image, complemented by using a random sample consensus (RANSAC) algorithm to eliminate incorrect matches. We evaluate our method from both segmentation and registration aspects. For segmentation evaluation, we test our method on DRIVE database, which provides manually labeled images from two specialists. The experimental results show that our method achieves 0.9562 in accuracy (Acc), which presents competitive performance compare to other existing segmentation methods. For registration evaluation, we test our method on STARE database, and the experimental results demonstrate the superior performance of the proposed method, which makes the algorithm a suitable tool for automated retinal image analysis. Fan Guo 0001, Beiji Zou 0001, Yixiong Liang |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2011 | Exploring regularized feature selection for person specific face verificationabstractIn this paper, we explore the regularized feature selection method for person specific face verification in unconstrained environments. We reformulate the generalization of the single-task sparsity-enforced feature selection method to multi-task cases as a simultaneous sparse approximation problem. We also investigate two feature selection strategies in the multi-task generalization based on the positive and negative feature correlation assumptions across different persons. Simultaneous orthogonal matching pursuit (SOMP) is adopted and modified to solve the corresponding optimization problems. We further proposed a named simultaneous subspace pursuit (SSP) methods which generalize the subspace pursuit method to solve the corresponding optimization problems. The performance of different feature selection strategies and different solvers for face verification are compared on the challenging LFW face database. Our experimental results show that 1) the selected subsets based on positive correlation assumption are more effective than those based on the negative correlation assumption; 2) the OMP-based solvers outperform SP-based solvers in terms of feature selection and 3) the regularized methods with OMP-based solvers can outperform state-of-the-art feature selection methods. Yixiong Liang, Lei Wang 0017, Beiji Zou 0001 |
ICCV | 1 |
| 2011 | Multi-task GLOH feature selection for human age estimationabstractIn this paper, we propose a novel age estimation method based on gradient location and orientation histogram (GLOH) descriptor and multi-task learning (MTL). The GLOH, one of the state-of-the-art local descriptor, is used to capture the age- related local and spatial information of face image. As the extracted GLOH features are often redundant, MTL is designed to select the most informative GLOH bins for age estimation problem, while the corresponding weights are determined by ridge regression. This approach largely reduces the dimensions of feature, which can not only improve performance but also decrease the computational burden. Experiments on the public available FG-NET database show that the proposed method can achieve comparable performance over previous approaches while using much fewer features. Yixiong Liang, Lingbo Liu, Yao Xiang, Beiji Zou 0001 |
ICIP | 1 |
| 2011 | Feature selection via simultaneous sparse approximation for person specific face verificationabstractThere is an increasing use of some imperceivable and redundant local features for face recognition. While only a relatively small fraction of them is relevant to the final recognition task, the feature selection is a crucial and necessary step to select the most discriminant ones to obtain a compact face representation. In this paper, we investigate the sparsity-enforced regularization-based feature selection methods and propose a multi-task feature selection method for building person specific models for face verification. We assume that the person specific models share a common subset of features and novelly reformulated the common subset selection problem as a simultaneous sparse approximation problem. The effectiveness of the proposed methods is verified with the challenging LFW face databases. Yixiong Liang, Lei Wang 0017, Beiji Zou 0001 |
ICIP | 1 |
| 2011 | Practical craniofacial surgery simulator based on GPU accelerated lattice shape matchingabstractAbstract This paper presents an intuitive and practical craniofacial surgery simulation system, which is suitable for daily clinical practice. The key component of the system is a GPU accelerated lattice shape matching method, to estimate individual patient post‐operative appearance interactively. The lattice model can be set up from individual CT data in an easy and robust way, incorporating CT mapping physical inhomogeneous material as well as orthotropic behavior of soft tissue. Instead of simulating dynamic behavior, an iterative optimization is used for direct computation of soft tissue deformation. In addition, a GPU acceleration framework for the lattice shape matching method is further exploited based on the lattice regularity, thus the simulation operations can be performed at interactive or real time. Finally, a craniofacial surgery simulation system is developed for daily clinical practice, which is capable of simulating a variety of surgeries including osteotomies, bone fragment repositioning, and insertion of implants. The treatment of more than 50 patients was found to provide a good correlation between simulation and post‐operative outcome. Copyright © 2011 John Wiley & Sons, Ltd. Yixiong Liang, Lingzhi Li 0006, Beiji Zou 0001, Xing-Hao Zhu |
Comput. Animat. Virtual Worlds | 2 |
| 2008 | Null space discriminant locality preserving projections for face recognition
Liping Yang 0001, Weiguo Gong, Xiaohua Gu, Weihong Li 0001, Yixiong Liang |
Neurocomputing | 5 |
| 2007 | Uncorrelated linear discriminant analysis based on weighted pairwise Fisher criterion
Yixiong Liang, Chengrong Li, Weiguo Gong, Yingjun Pan |
Pattern Recognit. | 1 |
| 2005 | Gabor Features-Based Classification Using SVM for Face Recognition
Yixiong Liang, Weiguo Gong, Yingjun Pan, Weihong Li 0001, Zhenjiang Hu 0002 |
ISNN (2) | 1 |
| 2005 | Generalizing relevance weighted LDA
Yixiong Liang, Weiguo Gong, Yingjun Pan, Weihong Li 0001 |
Pattern Recognit. | 1 |