Yimin Luo

dblp:206/8690 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0002-8032-371XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 tHPM-LDM: Integrating Individual Historical Record with Population Memory in Latent Diffusion-Based Glaucoma Forecasting
Jianyang Xie, Yimin Luo, Yanda Meng, Savita Madhusudhan, Gregory Yoke Hong Lip, Li Cheng 0001, Yalin Zheng, He Zhao 0002
MICCAI (1)3
2025 WASABI: A Metric for Evaluating Morphometric Plausibility of Synthetic Brain MRIs
Bahram Jafrasteh, Wei Peng 0009, Yimin Luo, Ehsan Adeli-Mosabbeb, Qingyu Zhao
MICCAI (2)4
2025 Hierarchical Multi-Scale Enhanced Transformer for Medical Image Segmentation
abstract
Segmentation is an important prerequisite for developing model healthcare systems, particularly for disease diagnosis and treatment planning. In the field of medical image segmentation, the U-shaped architecture, commonly referred to as U-Net, has emerged as the de facto standard and achieved remarkable success. However, due to the intrinsic locality of convolution operations, U-Net generally demonstrates limitations in explicitly modeling long-range dependency. Recent transformer-based models, designed for sequence-to-sequence prediction, have emerged as an alternative to traditional architectures, featuring innate global self-attention mechanisms. Unfortunately, they may sometimes suffer from limited localization abilities due to a lack of sufficient low-level details. To merit both Transformers and U-Net, in this paper, we propose a novel two-channel self-attention mechanism U-network, which performs feature extraction from two channels, CNN and Transformer, respectively. Compared to previous models, we propose two hierarchical feature fusion strategies from both spatial and channel dimensions. Moreover, to further promote the model performance, a loss function that can dynamically adjust the weights according to the output of each layer is constructed. Experimental results on five different datasets show that our method performs consistently outperforms state-of-the-art methods, and it also has an outstanding generalization ability to various medical image modalities.
Yantao Song, Yunli Lu, Lu Chen 0003, Yimin Luo
IEEE J. Biomed. Health Informatics4
2024 OphNet: A Large-Scale Video Benchmark for Ophthalmic Surgical Workflow Understanding
Peng Xia 0005, Lin Wang 0027, Siyuan Yan, Zhongxing Xu, Yimin Luo, Kaimin Song, Jürgen Leitner, Xuelian Cheng, Chi Liu 0002, Kaijing Zhou, ZongYuan Ge
ECCV (4)7
2024 An Online Two-Phase Workload Management Scheme for Collaborative Edge Computing System
abstract
The rapid expansion of Internet of Things (IoT) de-vices poses substantial challenges to the Quality of Service (QoS) provided by edge servers, which are often constrained by limited resources, particularly when handling numerous time-sensitive tasks. To address these challenges cost-effectively, collaborative edge computing has been introduced. This approach enhances system performance by enabling edge servers to offload tasks to neighboring edge servers or remote cloud servers, thereby alleviating the computational burden. However, existing research often falls short in providing robust long-term solutions for real-time online scenarios where tasks arrive unpredictably, lacking prior information. To this end, this paper explores the problem of online workload management in collaborative edge computing systems, aiming to improve the long-term utility of edge servers. We formulate the problem as a non-linear optimization problem. To solve this problem efficiently within polynomial time, we intro-duce an online two-phase workload management scheme called OTWMS. This scheme breaks down the workload management problem into several distributed and parallelizable local resource allocation problems and a centralized task migration matching problem. Through evaluation experiments, we demonstrate that the proposed scheme surpasses several benchmarks in terms of long-term utility and service performance optimization.
Zhennan Zhang, Yimin Luo, Shi Zhu, Fangliao Yang, Lailong Luo
MSN3
2024 Measurement Guidance in Diffusion Models: Insight from Medical Image Synthesis
abstract
In the field of healthcare, the acquisition of sample is usually restricted by multiple considerations, including cost, labor- intensive annotation, privacy concerns, and radiation hazards, therefore, synthesizing images-of-interest is an important tool to data augmentation. Diffusion models have recently attained state-of-the-art results in various synthesis tasks, and embedding energy functions has been proved that can effectively guide the pre-trained model to synthesize target samples. However, we notice that current method development and validation are still limited to improving indicators, such as Fréchet Inception Distance score (FID) and Inception Score (IS), and have not provided deeper investigations on downstream tasks, like disease grading and diagnosis. Moreover, existing classifier guidance which can be regarded as a special case of energy function can only has a singular effect on altering the distribution of the synthetic dataset. This may contribute to in-distribution synthetic sample that has limited help to downstream model optimization. All these limitations remind that we still have a long way to go to achieve controllable generation. In this work, we first conducted an analysis on previous guidance as well as its contributions on further applications from the perspective of data distribution. To synthesize samples which can help downstream applications, we then introduce uncertainty guidance in each sampling step and design an uncertainty-guided diffusion models. Extensive experiments on four medical datasets, with ten classic networks trained on the augmented sample sets provided a comprehensive evaluation on the practical contributions of our methodology. Furthermore, we provide a theoretical guarantee for general gradient guidance in diffusion models, which would benefit future research on investigating other forms of measurement guidance for specific generative tasks.
Yimin Luo, Qinyu Yang, Haikun Qi, Menghan Xia
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Target-Guided Diffusion Models for Unpaired Cross-Modality Medical Image Translation
abstract
In a clinical setting, the acquisition of certain medical image modality is often unavailable due to various considerations such as cost, radiation, etc. Therefore, unpaired cross-modality translation techniques, which involve training on the unpaired data and synthesizing the target modality with the guidance of the acquired source modality, are of great interest. Previous methods for synthesizing target medical images are to establish one-shot mapping through generative adversarial networks (GANs). As promising alternatives to GANs, diffusion models have recently received wide interests in generative tasks. In this paper, we propose a target-guided diffusion model (TGDM) for unpaired cross-modality medical image translation. For training, to encourage our diffusion model to learn more visual concepts, we adopted a perception prioritized weight scheme (P2W) to the training objectives. For sampling, a pre-trained classifier is adopted in the reverse process to relieve modality-specific remnants from source data. Experiments on both brain MRI-CT and prostate MRI-US datasets demonstrate that the proposed method achieves a visually realistic result that mimics a vivid anatomical section of the target organ. In addition, we have also conducted a subjective assessment based on the synthesized samples to further validate the clinical value of TGDM.
Yimin Luo, Qinyu Yang, Ziyi Liu 0010, Zenglin Shi, Weimin Huang 0002, Guoyan Zheng, Jun Cheng 0003
IEEE J. Biomed. Health Informatics1
2023 Cervical OCT Image Synthesis via Noise Adaptation Diffusion Models
abstract
The utilization of optical coherence tomography (OCT) holds significant promise in the realm of cervical cancer screening. This study considers the characteristics of OCT, such as sample scarcity and noise, and introduces a noise adaptive diffusion model (NADM) for data augmentation. The NADM comprises two components: the Adaptive Blind De-noising Module (ABD) and the Class-Guided Diffusion Model (CDM). The CDM is responsible for generating high-quality cervical samples, while the ABD is designed to adaptively suppress noise. Extensive experiments are done to establish the performance and generalizability of NADM, based on data collected from three distinct hospitals. To the best of our knowledge, we are the first to employ a data augmentation technique utilizing generative models for the purpose of diagnosing cervical OCT images.
Yuxuan Xiong, Yimin Luo, Yan Zhang 0123, Bo Du 0001
BIBM2
2020 Multi-Scale Progressive Fusion Network for Single Image Deraining
abstract
Rain streaks in the air appear in various blurring degrees and resolutions due to different distances from their positions to the camera. Similar rain patterns are visible in a rain image as well as its multi-scale (or multi-resolution) versions, which makes it possible to exploit such complementary information for rain streak representation. In this work, we explore the multi-scale collaborative representation for rain streaks from the perspective of input image scales and hierarchical deep features in a unified framework, termed multi-scale progressive fusion network (MSPFN) for single image rain streak removal. For the similar rain streaks at different positions, we employ recurrent calculation to capture the global texture, thus allowing to explore the complementary and redundant information at the spatial dimension to characterize target rain streaks. Besides, we construct multi-scale pyramid structure, and further introduce the attention mechanism to guide the fine fusion of these correlated information from different scales. This multi-scale progressive fusion strategy not only promotes the cooperative representation, but also boosts the end-to-end training. Our proposed method is extensively evaluated on several benchmark datasets and achieves the state-of-the-art results. Moreover, we conduct experiments on joint deraining, detection, and segmentation tasks, and inspire a new research direction of vision task driven image deraining. The source code is available at https://github.com/kuihua/MSPFN.
Kui Jiang, Zhongyuan Wang 0001, Peng Yi 0002, Chen Chen 0001, Baojin Huang, Yimin Luo, Jiayi Ma 0001, Junjun Jiang
CVPR6
2020 DEM Extraction from Airborne Lidar Point Cloud in Thick-Forested Areas via Convolutional Neural Network
abstract
Digital Elevation Model (DEM), representing the height of the earth terrain, is one of the crucial geographic information products. One of the main data source of DEM is the airborne LiDAR point cloud with its non-ground-reflections filtered out. Point cloud filtering in thick-forested areas is difficult without enough ground control points when using conventional methods. In this paper, a supervised method is proposed to handle the problem of automatic DEM extraction with little ground control points. The design of the method is inspired by the successful application of the convolutional neural networks (CNN) in the image super resolution (SR) process. First, with the given LiDAR point cloud, the digital surface model (DSM) is resampled with regular grid. Then, by learning the spatial autocorrelation between the DSM and its corresponding DEM, a robust CNN model is established. Finally, the DEM in thick-forested areas can be generated from the DSM with the trained model. Experimental results at two different mountain sites in China validate the effectiveness of the proposed method of high-precision DEM generation.
Yongjun Zhang 0002, Sizhe Xiang, Yi Wan 0001, Yimin Luo
IGARSS5
2019 Separability and Compactness Network for Image Recognition and Superresolution
abstract
Convolutional neural networks (CNNs) have wide applications in pattern recognition and image processing. Despite recent advances, much remains to be done for CNNs to learn a better representation of image samples. Therefore, constant optimizations should be provided on CNNs. To achieve a good performance on classification, intuitively, samples' interclass separability, or intraclass compactness should be simultaneously maximized. Accordingly, in this paper, we propose a new network, named separability and compactness network (SCNet) to rectify this problem. SCNet minimizes the softmax loss and the distance between features of samples from the same class under a jointly supervised framework, resulting in simultaneous maximization of interclass separability and intraclass compactness of samples. Furthermore, considering the convenience and the efficiency of the cosine similarity in face recognition tasks, we incorporate it into SCNet's distance metric to enable sample features from the same class to line up in the same direction and those from different classes to have a large angle of separation. We apply SCNet to three different tasks: visual classification, face recognition, and image superresolution. Experiments on both public data sets and real-world satellite images validate the effectiveness of our SCNet.
Liguo Zhou, Zhongyuan Wang 0001, Yimin Luo, Zixiang Xiong
IEEE Trans. Neural Networks Learn. Syst.3
2018 Improving Convolutional Neural Networks Via Compacting Features
abstract
Convolutional neural networks (CNNs) have shown great advantages in computer vision fields, and loss functions are of great significance to their gradient descent algorithms. Softmax loss, a combination of cross-entropy loss and Softmax function, is the most commonly used one for CNNs. Hence, it can continuously increase the discernibility of sample features in classification tasks. Intuitively, to promote the discrimination of CNNs, the learned features are desirable when the inter-class separability and intra-class compactness are maximized simultaneously. Since Softmax loss hardly motivates this inter-class separability and intra-class compactness simultaneously and explicitly, we propose a new method to achieve this simultaneous maximization. This method minimizes the distance between features of homogeneous samples along with Softmax loss and thus improves CNNs' performance on vision-related tasks. Experiments on both visual classification and face verification datasets validate the effectiveness and advantages of our method.
Liguo Zhou, Yimin Luo, Zhongyuan Wang 0001
ICASSP3
2018 Coarse-to-Fine Image Super-Resolution Using Convolutional Neural Networks
Liguo Zhou, Zhongyuan Wang 0001, Yimin Luo
MMM (2)4
2017 DEM Retrieval From Airborne LiDAR Point Clouds in Mountain Areas via Deep Neural Networks
abstract
Airborne light detection and ranging (LiDAR) remote sensing enables accurate estimation and monitoring of terrain and vegetation, and digital surface model (DSM) and digital elevation model (DEM) are vital analytical tools to achieve this estimation and monitoring. Among them, DSM can be directly acquired from airborne LiDAR point clouds; nevertheless, for the production of DEM, point clouds representing a surface of ground objects should be accurately filtered out at first. In some mountain forest areas, due to the limited penetration of airborne LiDAR, ground points sustain a serious lack, which results in the difficulty in producing accurate DEMs. To reduce the intricacy and subjectivity caused by the manual supplement to ground points, this letter proposes a new DEM retrieval method from airborne LiDAR point clouds in mountain areas based on deep neural networks (DNNs). With a DNN model trained by accurate DEMs and DSMs, DEM retrieval becomes much easier by inputting their DSM into this model for prediction. Experiments on Fujian and Hainan mountain data sets demonstrate the effectiveness of this supervised method.
Yimin Luo, Hongchao Ma, Liguo Zhou
IEEE Geosci. Remote. Sens. Lett.1
2017 Video Satellite Imagery Super Resolution via Convolutional Neural Networks
abstract
Video satellite imagery is a new technique for earth dynamic observation and has a wide range of uses in environmental fields. Despite its capability of dynamic targets' detection, it sustains a serious restriction of the image quality due to the degradation and compression in its imaging process. Hence, the super-resolution (SR) reconstruction on these compressed low-spatial-resolution images is of significance to afterward ground objects recognition and detection tasks. Based on the recent proposed state-of-the-art convolutional neural networks (CNNs) SR methods, we proposed an SR method which could get more precise reconstructed high-spatial-resolution images. Trained with Gaofen-2 satellite images, a robust CNN model specified in satellite image SR is obtained. Experimentally, the reconstruction results on Jilin-1 mission satellite images validate the effectiveness of our method.
Yimin Luo, Liguo Zhou, Zhongyuan Wang 0001
IEEE Geosci. Remote. Sens. Lett.1