Yepeng Liu 0002

dblp:184/0225-2 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0009-0002-1706-2681ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Leveraging Textual Anatomical Knowledge for Class-Imbalanced Semi-Supervised Multi-Organ Segmentation
abstract
Imbalanced class distributions among different organs pose significant challenges in real-world semi-supervised multi-organ segmentation. Integrating anatomical priors offers a promising research direction to mitigate these imbalances. In this paper, we explore the capabilities of Multimodal Large Language Models (MLLM) to extract robust, generic textual anatomical insights serving as prior knowledge for segmentation model. Specifically, we employ GPT-4o to generate detailed textual descriptions of anatomical priors-including both inter-organ relative positional relationships and organ shape characteristics. These priors generated only once for the whole training and testing are then seamlessly integrated into the segmentation model as parameters within the segmentation head. Furthermore, we align the textual priors with visual features using contrastive learning. The inter-organ positional priors guide the model in localizing smaller organs relative to larger ones, while the organ shape priors help ensure that the learned morphological structures are more anatomically plausible. Extensive experiments demonstrate that our method significantly outperforms some state-of-the-art approaches. The source code is available at: https://github.com/Lunn88/TAK-Semi.
Yuliang Gu, Weilun Tsao, Yepeng Liu 0002, Lianming Wu, Thierry Géraud, Bo Du 0001, Yongchao Xu
IEEE Trans. Medical Imaging3
2025 LiftFeat: 3D Geometry-Aware Local Feature Matching
abstract
Robust and efficient local feature matching plays a crucial role in applications such as SLAM and visual localization for robotics. Despite great progress, it is still very challenging to extract robust and discriminative visual features in scenarios with drastic lighting changes, low texture areas, or repetitive patterns. In this paper, we propose a new lightweight network called LiftFeat, which lifts the robustness of raw descriptor by aggregating 3D geometric feature. Specifically, we first adopt a pre-trained monocular depth estimation model to generate pseudo surface normal label, supervising the extraction of 3D geometric feature in terms of predicted surface normal. We then design a 3D geometry-aware feature lifting module to fuse surface normal feature with raw 2D descriptor feature. Integrating such 3D geometric feature enhances the discriminative ability of 2D feature description in extreme conditions. Extensive experimental results on relative pose estimation, homography estimation, and visual localization tasks, demonstrate that our LiftFeat outperforms some lightweight state-of-the-art methods. Code will be released at: https://github.com/lyp-deeplearning/LiftFeat.
Yepeng Liu 0002, Wenpeng Lai, Yuxuan Xiong, Jinchi Zhu, Jun Cheng 0003, Yongchao Xu
ICRA1
2025 Dual structure-aware image filterings for semi-supervised medical image segmentation
Yuliang Gu, Zhichao Sun 0004, Xin Xiao 0010, Yepeng Liu 0002, Yongchao Xu, Laurent Najman
Medical Image Anal.5
2025 Learning Modality-Invariant Feature for Multimodal Image Matching via Knowledge Distillation
abstract
Multimodal remote sensing image matching is essential for multi-source information fusion. Recently, learning-based feature matching networks have significantly enhanced the performance of unimodal image matching tasks through data-driven approaches. However, progress in applying these learning-based methods to multimodal image matching has been slower. A major obstacle is the substantial nonlinear radiometric differences between modalities, which require networks to learn modality-invariant features from large amounts of paired data. To address this, we propose EMINet, an efficient method for learning modality-invariant features from limited data to improve matching performance. Our approach constructs a high-performance teacher network by combining the DINOv2 foundational model, the keypoint and descriptor extraction network SuperPoint, and the feature matching network SuperGlue. Leveraging the strong semantic representation capability of DINOv2, the teacher network achieves excellent cross-modality matching ability. To meet low-latency requirements in practical applications, we introduce two novel knowledge distillation strategies: Semantic Window Relation Distillation (SWRD) and Cross-Triplet Descriptor Distillation (CTDD). SWRD improves the discriminative power of the student network’s descriptors by learning patch-level distributions from DINOv2, while CTDD enforces cross-modality triplet constraints to enhance modality invariance of the student network. Experimental results demonstrate that EMINet outperforms several state-of-the-art methods on various datasets, including Optical-SAR, Optical-NIR, and Optical-IR datasets.
Yepeng Liu 0002, Wenpeng Lai, Yuliang Gu, Gui-Song Xia, Bo Du 0001, Yongchao Xu
IEEE Trans. Geosci. Remote. Sens.1
2025 MIFNet: Learning Modality-Invariant Features for Generalizable Multimodal Image Matching
abstract
Many keypoint detection and description methods have been proposed for image matching or registration. While these methods demonstrate promising performance for single-modality image matching, they often struggle with multimodal data because the descriptors trained on single-modality data tend to lack robustness against the non-linear variations present in multimodal data. Extending such methods to multimodal image matching often requires well-aligned multimodal data to learn modality-invariant descriptors. However, acquiring such data is often costly and impractical in many real-world scenarios. To address this challenge, we propose a modality-invariant feature learning network (MIFNet) to compute modality-invariant features for keypoint descriptions in multimodal image matching using only single-modality training data. Specifically, we propose a novel latent feature aggregation module and a cumulative hybrid aggregation module to enhance the base keypoint descriptors trained on single-modality data by leveraging pre-trained features from Stable Diffusion models. We validate our method with recent keypoint detection and description methods in three multimodal retinal image datasets (CF-FA, CF-OCT, EMA-OCTA) and two remote sensing datasets (Optical-SAR and Optical-NIR). Extensive experiments demonstrate that the proposed MIFNet is able to learn modality-invariant feature for multimodal image matching without accessing the targeted modality and has good zero-shot generalization ability. The code will be released at https://github.com/lyp-deeplearning/MIFNet.
Yepeng Liu 0002, Zhichao Sun 0004, Baosheng Yu, Yitian Zhao, Bo Du 0001, Yongchao Xu, Jun Cheng 0003
IEEE Trans. Image Process.1
2024 Shape Transformation Driven by Active Contour for Class-Imbalanced Semi-Supervised Medical Image Segmentation
abstract
Annotating 3D medical images demands expert knowledge and is time-consuming. As a result, semi-supervised learning (SSL) approaches have gained significant interest in 3D medical image segmentation. The significant size differences among various organs in the human body lead to imbalanced class distribution, which is a major challenge in the real-world application of these SSL approaches. To address this issue, we develop a novel Shape Transformation driven by Active Contour (STAC), that enlarges smaller organs to alleviate imbalanced class distribution across different organs. Inspired by curve evolution theory in active contour methods, STAC employs a signed distance function (SDF) as the level set function, to implicitly represent the shape of organs, and deforms voxels in the direction of the steepest descent of SDF (i.e., the normal vector). To ensure that the voxels far from expansion organs remain unchanged, we design an SDF-based weight function to control the degree of deformation for each voxel. We then use STAC as a data-augmentation process during the training stage. Experimental results on two benchmark datasets demonstrate that the proposed method significantly outperforms some state-of-the-art methods. Source code is publicly available at https://github.com/GuGuLL123/STAC.
Yuliang Gu, Yepeng Liu 0002, Zhichao Sun 0004, Jinchi Zhu, Yongchao Xu, Laurent Najman
BIBM2
2024 Progressive Retinal Image Registration via Global and Local Deformable Transformations
abstract
Retinal image registration plays an important role in the ophthalmological diagnosis process. Since there exist variances in viewing angles and anatomical structures across different retinal images, keypoint-based approaches become the mainstream methods for retinal image registration thanks to their robustness and low latency. These methods typically assume the retinal surfaces are planar, and adopt feature matching to obtain the homography matrix that represents the global transformation between images. Yet, such a planar hypothesis inevitably introduces registration errors since retinal surface is approximately curved. This limitation is more prominent when registering image pairs with significant differences in viewing angles. To address this problem, we propose a hybrid registration framework called HybridRetina, which progressively registers retinal images with global and local deformable transformations. For that, we use a keypoint detector and a deformation network called GAMorph to estimate the global transformation and local deformable transformation, respectively. Specifically, we integrate multi-level pixel relation knowledge to guide the training of GAMorph. Additionally, we utilize an edge attention module that includes the geometric priors of the images, ensuring the deformation field focuses more on the vascular regions of clinical interest. Experiments on two widely-used datasets, FIRE and FLoRI21, show that our proposed HybridRetina significantly outperforms some state-of-the-art methods. The code is available at https://github.com/lyp-deeplearning/awesome-retinal-registration.
Yepeng Liu 0002, Baosheng Yu, Yuliang Gu, Bo Du 0001, Yongchao Xu, Jun Cheng 0003
BIBM1
2024 Position-Guided Prompt Learning for Anomaly Detection in Chest X-Rays
Zhichao Sun 0004, Yuliang Gu, Yepeng Liu 0002, Yongchao Xu
MICCAI (1)3
2021 MOS: A Low Latency and Lightweight Framework for Face Detection, Landmark Localization, and Head Pose Estimation
Yepeng Liu 0002, Zaiwang Gu, Shenghua Gao, Yusheng Zeng, Jun Cheng 0003
BMVC1