Yidan Feng

dblp:260/2533 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
13since 2021 · last 2025
0000-0001-7208-0458ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Asynchronous Multi-modal Learning for Dynamic Risk Monitoring of Acute Respiratory Distress Syndrome in Intensive Care Units
Yidan Feng, Zhanli Hu, Harry Qin
MICCAI (15)1
2025 Bridging MRI Cross-Modality Synthesis and Multi-Contrast Super-Resolution by Fine-Grained Difference Learning
abstract
In multi-modal magnetic resonance imaging (MRI), the tasks of imputing or reconstructing the target modality share a common obstacle: the accurate modeling of fine-grained inter-modal differences, which has been sparingly addressed in current literature. These differences stem from two sources: 1) spatial misalignment remaining after coarse registration and 2) structural distinction arising from modality-specific signal manifestations. This paper integrates the previously separate research trajectories of cross-modality synthesis (CMS) and multi-contrast super-resolution (MCSR) to address this pervasive challenge within a unified framework. Connected through generalized down-sampling ratios, this unification not only emphasizes their common goal in reducing structural differences, but also identifies the key task distinguishing MCSR from CMS: modeling the structural distinctions using the limited information from the misaligned target input. Specifically, we propose a composite network architecture with several key components: a label correction module to align the coordinates of multi-modal training pairs, a CMS module serving as the base model, an SR branch to handle target inputs, and a difference projection discriminator for structural distinction-centered adversarial training. When training the SR branch as the generator, the adversarial learning is enhanced with distinction-aware incremental modulation to ensure better-controlled generation. Moreover, the SR branch integrates deformable convolutions to address cross-modal spatial misalignment at the feature level. Experiments conducted on three public datasets demonstrate that our approach effectively balances structural accuracy and realism, exhibiting overall superiority in comprehensive evaluations for both tasks over current state-of-the-art approaches. The code is available at https://github.com/papshare/FGDL.
Yidan Feng, Jing Cai 0001, Mingqiang Wei, Harry Qin
IEEE Trans. Medical Imaging1
2024 Semi-supervised TEE Segmentation via Interacting with SAM Equipped with Noise-Resilient Prompting
abstract
Semi-supervised learning (SSL) is a powerful tool to address the challenge of insufficient annotated data in medical segmentation problems. However, existing semi-supervised methods mainly rely on internal knowledge for pseudo labeling, which is biased due to the distribution mismatch between the highly imbalanced labeled and unlabeled data. Segmenting left atrial appendage (LAA) from transesophageal echocardiogram (TEE) images is a typical medical image segmentation task featured by scarcity of professional annotations and diverse data distributions, for which existing SSL models cannot achieve satisfactory performance. In this paper, we propose a novel strategy to mitigate the inherent challenge of distribution mismatch in SSL by, for the first time, incorporating a large foundation model (i.e. SAM in our implementation) into an SSL model to improve the quality of pseudo labels. We further propose a new self-reconstruction mechanism to generate both noise-resilient prompts to demonically improve SAM’s generalization capability over TEE images and self-perturbations to stabilize the training process and reduce the impact of noisy labels. We conduct extensive experiments on an in-house TEE dataset; experimental results demonstrate that our method achieves better performance than state-of-the-art SSL models.
Yidan Feng, Haoneng Lin, Yiting Fan, Alex Pui-Wai Lee, Xiaowei Hu 0001, Harry Qin
AAAI2
2024 Unified Multi-modal Learning for Any Modality Combinations in Alzheimer's Disease Diagnosis
Yidan Feng, Bingchen Gao, Anqi Qiu, Harry Qin
MICCAI (3)1
2022 MBA-RainGAN: A Multi-Branch Attention Generative Adversarial Network for Mixture of Rain Removal
abstract
Rain severely degrades the visibility of scene objects, especially when images are captured through the glass under rainy weather. We observe three intriguing phenomena: 1) rain is a mixture of raindrops, rain streaks and rainy haze; 2) the depth from the camera determines the degree of object visibility, where objects nearby and far away are visually blocked by rain streaks and rainy haze, respectively; and 3) raindrops on the glass randomly affect the object visibility of the whole image space. However, existing solutions and benchmark datasets lack full consideration of the mixture of rain (MOR). In this paper, we originally consider that the overall object visibility is determined by MOR, and enrich the RainCityscapes by considering real-world raindrops to construct the MOR dataset, named RainCityscapes++. To solve the practical rain removal problem arisen from MOR, we formulate a new rain imaging model and propose a multi-branch attention generative adversarial network (MBA-RainGAN). Extensive experiments show clear improvements of our approach over SOTAs on RainCityscapes++.
Yiyang Shen, Yidan Feng, Weiming Wang 0002, Dong Liang 0008, Harry Qin, Haoran Xie 0001, Mingqiang Wei
ICASSP2
2022 Towards Robust Part-aware Instance Segmentation for Industrial Bin Picking
abstract
Industrial bin picking is a challenging task that requires accurate and robust segmentation of individual object instances. Particularly, industrial objects can have irregular shapes, that is, thin and concave, whereas in bin-picking scenarios, objects are often closely packed with strong occlusion. To address these challenges, we formulate a novel part-aware instance segmentation pipeline. The key idea is to decompose industrial objects into correlated approximate convex parts and enhance the object-level segmentation with part-level segmentation. We design a part-aware network to predict part masks and part-to-part offsets, followed by a part aggregation module to assemble the recognized parts into instances. To guide the network learning, we also propose an automatic label decoupling scheme to generate ground-truth part-level labels from instance-level labels. Finally, we contribute the first instance segmentation dataset, which contains a variety of industrial objects that are thin and have non-trivial shapes. Extensive experimental results on various industrial objects demonstrate that our method can achieve the best segmentation results compared with the state-of-the-art approaches.
Yidan Feng, Biqi Yang, Xianzhi Li 0001, Chi-Wing Fu, Kai Chen 0028, Qi Dou 0001, Mingqiang Wei, Yun-Hui Liu 0001, Pheng-Ann Heng
ICRA1
2022 SESR: Self-Ensembling Sim-to-Real Instance Segmentation for Auto-Store Bin Picking
abstract
Instance segmentation is an important task for supporting robotic grasping in auto-store scenarios. Accurate segmentation usually relies on the quantity and quality of available annotated training data. However, it requires tremendous cost to obtain these labels. In this work, without requiring any human annotations on real data, our proposed self-ensembling sim-to-real network, namely SESR, is able to generate precise instance masks for a wide variety of supermarket goods. We design our SESR with a teacher model and a student model trained with a self-ensembling strategy. We adopt different levels of consistency to bridge the sim-to-real gap and boost the model generalization ability. Also, we compile an auto-store bin-picking dataset covering various goods. Extensive experiments on both unseen scenarios and unseen objects validate the effectiveness and superiority of our method over others, and the robot arm demonstrations further show that our segmentation results can support real-time auto-store bin picking.
Biqi Yang, Kai Chen 0028, Yidan Feng, Xianzhi Li 0001, Qi Dou 0001, Chi-Wing Fu, Yun-Hui Liu 0001, Pheng-Ann Heng
IROS5
2022 Contrastive Semantic-Guided Image Smoothing Network
abstract
Abstract Image smoothing is a fundamental low‐level vision task that aims to preserve salient structures of an image while removing insignificant details. Deep learning has been explored in image smoothing to deal with the complex entanglement of semantic structures and trivial details. However, current methods neglect two important facts in smoothing: 1) naive pixel‐level regression supervised by the limited number of high‐quality smoothing ground‐truth could lead to domain shift and cause generalization problems towards real‐world images; 2) texture appearance is closely related to object semantics, so that image smoothing requires awareness of semantic difference to apply adaptive smoothing strengths. To address these issues, we propose a novel Contrastive Semantic‐Guided Image Smoothing Network (CSGIS‐Net) that combines both contrastive prior and semantic prior to facilitate robust image smoothing. The supervision signal is augmented by leveraging undesired smoothing effects as negative teachers, and by incorporating segmentation tasks to encourage semantic distinctiveness. To realize the proposed network, we also enrich the original VOC dataset with texture enhancement and smoothing labels, namely VOC‐smooth, which first bridges image smoothing and semantic segmentation. Extensive experiments demonstrate that the proposed CSGIS‐Net outperforms state‐of‐the‐art algorithms by a large margin. Code and dataset are available at https://github.com/wangjie6866/CSGIS-Net .
Jie Wang 0069, Yongzhen Wang 0001, Yidan Feng, Lina Gong, Xuefeng Yan 0001, Haoran Xie 0001, Fu Lee Wang, Mingqiang Wei
Comput. Graph. Forum3
2022 Easy2Hard: Learning to Solve the Intractables From a Synthetic Dataset for Structure-Preserving Image Smoothing
abstract
Image smoothing is a prerequisite for many computer vision and graphics applications. In this article, we raise an intriguing question whether a dataset that semantically describes meaningful structures and unimportant details can facilitate a deep learning model to smooth complex natural images. To answer it, we generate ground-truth labels from easy samples by candidate generation and a screening test and synthesize hard samples in structure-preserving smoothing by blending intricate and multifarious details with the labels. To take full advantage of this dataset, we present a joint edge detection and structure-preserving image smoothing neural network (JESS-Net). Moreover, we propose the distinctive total variation loss as prior knowledge to narrow the gap between synthetic and real data. Experiments on different datasets and real images show clear improvements of our method over the state of the arts in terms of both the image cleanness and structure-preserving ability. Code and dataset are available at https://github.com/YidFeng/Easy2Hard.
Yidan Feng, Xuefeng Yan 0001, Xin Yang 0011, Mingqiang Wei, Ligang Liu 0001
IEEE Trans. Neural Networks Learn. Syst.1
2021 Adaptive Graph Convolution for Point Cloud Analysis
abstract
Convolution on 3D point clouds that generalized from 2D grid-like domains is widely researched yet far from perfect. The standard convolution characterises feature correspondences indistinguishably among 3D points, presenting an intrinsic limitation of poor distinctive feature learning. In this paper, we propose Adaptive Graph Convolution (AdaptConv) which generates adaptive kernels for points according to their dynamically learned features. Compared with using a fixed/isotropic kernel, AdaptConv improves the flexibility of point cloud convolutions, effectively and precisely capturing the diverse relations between points from different semantic parts. Unlike popular attentional weight schemes, the proposed AdaptConv implements the adaptiveness inside the convolution operation instead of simply assigning different weights to the neighboring points. Extensive qualitative and quantitative evaluations show that our method outperforms state-of-the-art point cloud classification and segmentation approaches on several benchmark datasets. Our code is available at https://github.com/hrzhou2/AdaptConv-master.
Yidan Feng, Mingsheng Fang, Mingqiang Wei, Harry Qin, Tong Lu 0002
ICCV2
2021 Direction-aware Feature-level Frequency Decomposition for Single Image Deraining
abstract
We present a novel direction-aware feature-level frequency decomposition network for single image deraining. Compared with existing solutions, the proposed network has three compelling characteristics. First, unlike previous algorithms, we propose to perform frequency decomposition at feature-level instead of image-level, allowing both low-frequency maps containing structures and high-frequency maps containing details to be continuously refined during the training procedure. Second, we further establish communication channels between low-frequency maps and high-frequency maps to interactively capture structures from high-frequency maps and add them back to low-frequency maps and, simultaneously, extract details from low-frequency maps and send them back to high-frequency maps, thereby removing rain streaks while preserving more delicate features in the input image. Third, different from existing algorithms using convolutional filters consistent in all directions, we propose a direction-aware filter to capture the direction of rain streaks in order to more effectively and thoroughly purge the input images of rain streaks. We extensively evaluate the proposed approach in three representative datasets and experimental results corroborate our approach consistently outperforms state-of-the-art deraining algorithms.
Yidan Feng, Mingqiang Wei, Haoran Xie 0001, Yiping Chen 0002, Jonathan Li 0001, Xiao-Ping Zhang 0002, Harry Qin
IJCAI2
2021 Multi-scale selective image texture smoothing via intuitive single clicks
Chong Liu 0005, Yidan Feng, Cui Yang, Mingqiang Wei, Jun Wang 0039
Signal Process. Image Commun.2
2021 Selective Guidance Normal Filter for Geometric Texture Removal
abstract
There is typically a trade-off between removing the detailed appearance (i.e., geometric textures) and preserving the intrinsic properties (i.e., geometric structures) of 3D surfaces. The conventional use of mesh vertex/facet-centered patches in many filters leads to side-effects including remnant textures, improperly filtered structures, and distorted shapes. We propose a selective guidance normal filter (SGNF) which adapts the Relative Total Variation (RTV) to a maximal/minimal scheme (mmRTV). The mmRTV measures the geometric flatness of surface patches, which helps in finding adaptive patches whose boundaries are aligned with the facet being processed. The adaptive patches provide selective guidance normals, which are subsequently used for normal filtering. The filtering smooths out the geometric textures by using guidance normals estimated from patches with maximal RTV (the least flatness), and preserves the geometric structures by using normals estimated from patches with minimal RTV (the most flatness). This simple yet effective modification of the RTV makes our SGNF specialized rather than trade off between texture removal and structure preservation, which is distinct from existing mesh filters. Experiments show that our approach is visually and numerically comparable to the state-of-the-art mesh filters, in most cases. In addition, the mmRTV is generally applicable to bas-relief modeling and image texture removal.
Mingqiang Wei, Yidan Feng, Honghua Chen
IEEE Trans. Vis. Comput. Graph.2
2020 Detail-recovery Image Deraining via Context Aggregation Networks
abstract
This paper looks at this intriguing question: are single images with their details lost during deraining, reversible to their artifact-free status? We propose an end-to-end detail-recovery image deraining network (termed a DRDNet) to solve the problem. Unlike existing image deraining approaches that attempt to meet the conflicting goal of simultaneously deraining and preserving details in a unified framework, we propose to view rain removal and detail recovery as two seperate tasks, so that each part could specialize rather than trade-off between two conflicting goals. Specifically, we introduce two parallel sub-networks with a comprehensive loss function which synergize to derain and recover the lost details caused by deraining. For complete rain removal, we present a rain residual network with the squeeze-and-excitation (SE) operation to remove rain streaks from the rainy images. For detail recovery, we construct a specialized detail repair network consisting of welldesigned blocks, named structure detail context aggregation block (SDCAB), to encourage the lost details to return for eliminating image degradations. Moreover, the detail recovery branch of our proposed detail repair framework is detachable and can be incorporated into existing deraining methods to boost their performances. DRD-Net has been validated on several well-known benchmark datasets in terms of deraining robustness and detail accuracy. Comparisons show clear visual and numerical improvements of our method over the state-of-the-arts.
Mingqiang Wei, Jun Wang 0039, Yidan Feng, Luming Liang, Haoran Xie 0001, Fu Lee Wang, Meng Wang 0001
CVPR4
2020 Geometry and Learning Co-Supported Normal Estimation for Unstructured Point Cloud
abstract
In this paper, we propose a normal estimation method for unstructured point cloud. We observe that geometric estimators commonly focus more on feature preservation but are hard to tune parameters and sensitive to noise, while learning-based approaches pursue an overall normal estimation accuracy but cannot well handle challenging regions such as surface edges. This paper presents a novel normal estimation method, under the co-support of geometric estimator and deep learning. To lowering the learning difficulty, we first propose to compute a suboptimal initial normal at each point by searching for a best fitting patch. Based on the computed normal field, we design a normal-based height map network (NH-Net) to fine-tune the suboptimal normals. Qualitative and quantitative evaluations demonstrate the clear improvements of our results over both traditional methods and learning-based methods, in terms of estimation accuracy and feature recovery.
Honghua Chen, Yidan Feng, Qiong Wang 0001, Harry Qin, Haoran Xie 0001, Fu Lee Wang, Mingqiang Wei, Jun Wang 0039
CVPR3
2020 NormalF-Net: Normal Filtering Neural Network for Feature-preserving Mesh Denoising
Zhiqi Li 0002, Yingkui Zhang, Yidan Feng, Xingyu Xie, Qiong Wang 0001, Mingqiang Wei, Pheng-Ann Heng
Comput. Aided Des.3