EDBT 2026 Demo / reviewers in the wild / expert
Fei Luo 0004
dblp:83/1192-4
· DBLP profile ↗
49ranked-venue papers
6as first author
37since 2021 · last 2026
0000-0001-7320-5144ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 2 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HumanPro: Single-view 3D Clothed Human Reconstruction with Progressive Normal GuidanceabstractReconstructing fine-grained geometry of clothed human from single-view image is a challenging task, particularly in accurately recovering complex shapes and generating clothes details. To address these limitations, we propose a novel approach named HumanPro, which estimates high-quality human normals via a generative model, and progressively deforms a parametric body into the final clothed human mesh guided by normals. First, we propose a geometry-aware latent diffusion model with a normal enhancer to estimate high-quality human normals from four views. Then, we propose a progressive mesh optimization consisting of shape-aware deformation alignment and global-to-patch detail refinement for human mesh reconstruction. The shape-aware deformation alignment applies image morphing to learn the shape-level gap of normals, addressing large-scale deformation of complex clothes. It can recover the overall silhouette of a clothed human, and serves as an initialization for the global-to-patch detail refinement. Our detail refinement combines global and patch-wise optimization strategies to iteratively produce the clothed human mesh by minimizing the pixel-level difference of normals. This way effectively recovers fine-grained details while avoiding local minima. Extensive experiments demonstrate that HumanPro can deal with various challenging scenarios and outperforms state-of-the-art methods. Jianchi Sun, Fei Luo 0004, Wenzhuo Fan, Yu Jiang 0007, Chunxia Xiao |
AAAI | 2 |
| 2026 | AdaEndoGS: An Adaptive Enlightening Model for Endoscopy Based on 3D Gaussian SplattingabstractEndoscopic reconstruction and rendering technology are crucial for minimally invasive diagnosis and treatment. Due to physical constraints such as limited light source placement and narrow operational space, endoscopic scenes often contain some dark regions that compromise clinical observation and diagnostic accuracy. To address this issue, we propose AdaEndoGS for endoscopic 3D reconstruction with physically consistent dark region illumination enhancement. Specifically, built upon 3D Gaussian Splatting, AdaEndoGS enriches each Gaussian with surface attributes including normal, roughness, and reflectance, thereby constructing an illuminatable 3D scene representation. AdaEndoGS automatically identifies the dark region via ray tracing, leveraging Gaussian opacity, surface normal, and viewpoint information to adaptively plan the placement of a virtual light source. Supervised by a carefully designed loss function incorporating multiple illumination-related terms, AdaEndoGS optimizes light source intensity and attenuation parameters to generate harmonious illumination enhancement in the dark regions, while avoiding overexposure in originally bright areas. We comprehensively evaluate our method on the public dataset C3VD and a novel endoscopy simulation dataset created by us in Unity3D, which includes paired low-lit and well-lit images to facilitate quantitative evaluation. Experimental results show that AdaEndoGS more accurately simulates light-matter interactions compared to existing methods. It significantly improves visual quality and enhances detail visibility, offering an effective technical solution for advancing endoscopic image-based reconstruction and rendering. Project repository: https://webstermorton.github.io/adaendogs-website. Yiding Wen, Yuanfan Liu, Huanmei Guan, Fei Luo 0004 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | MSFP-Net: Multi-Scale Fusion of SAM-Derived Priors for Medical Image SegmentationabstractThe Segment Anything Model (SAM) has demonstrated strong zero-shot segmentation performance on natural images and has gained increasing attention in medical image applications. However, most existing approaches fine-tune SAM or incorporate task-specific adapters to adapt it to medical data, resulting in high computational cost and strong dependence on large-scale annotated datasets. In this paper, we propose MSFP-Net, a medical image segmentation framework that uses the Segment Anything Model (SAM) as a frozen prior generator. Unlike existing methods that fine-tune SAM or insert task-specific adapters, we apply SAM before training to generate coarse segmentation masks. These masks serve as priors and are injected into a U-shaped segmentation network through multiple prior learning blocks within skip connections. Each prior learning block includes a self-update module that applies self-attention to refine the priors, and a multi-scale learning block that performs cross-attention at multiple resolutions to enhance encoder features under the guidance of SAM priors. A dynamic learning block further fuses the original encoder features and the multi-scale updated features using learnable weights, enabling the network to adaptively balance task-specific representations with SAM-guided priors. By injecting SAM priors through prior learning blocks, MSFP- Net avoids fine-tuning and achieves superior segmentation accuracy with reduced computational cost, outperforming state-of-the-art SAM-adapted methods (e.g., SAMUS) by up to 4.9% in Dice score and 9.0% in IoU across four public datasets. Qian Zhou 0001, Hua Zou 0002, Fei Luo 0004 |
BIBM | 4 |
| 2025 | VPI-Depth: Endoscopic Dark Image Depth Estimation via Virtual Physical IlluminationabstractEndoscopic imaging often encounters challenges, including narrow fields of view and uneven illumination, which significantly impair depth perception and diagnostic accuracy. Current representative methods, such as PPSNet, Monodepth2, and EndoSLAM, often struggle with dark regions, leading to unsatisfactory performance. To address these limitations, we introduce VPI-Depth, an innovative depth estimation method that leverages multi-light source rendering to simulate additional illumination in dark areas. Our approach integrates a physically spatial awareness and illumination model to enhance visual cues, enabling joint optimization of depth and Albedo estimation. By combining realistic shading with depth constraints, VPI-Depth achieves better depth estimation accuracy and robustness in challenging endoscopic environments. Extensive experiments on the C3VD dataset and our synthetic dataset demonstrate that our method outperforms state-of-the-art methods across multiple metrics. Yuanfan Liu, Chi Kong Tam, Jun Kui Zhang, Fei Luo 0004 |
BIBM | 6 |
| 2025 | Anatomy-Aware Adaptation of Pre-Trained Models for Medical Difference Visual Question AnsweringabstractMedical Difference Visual Question Answering (Med-Diff-VQA) is a challenging and clinically significant task that requires identifying and interpreting subtle anatomical differences between pairs of medical images, such as chest X-rays, in response to domain-specific questions. Unlike traditional Visual Question Answering (VQA) tasks, Med-DiffVQA is characterized by high visual similarity, limited data availability, and a strong requirement for anatomically precise reasoning. To tackle these challenges, we propose$\mathbf{A}^{\mathbf{2}} \mathbf{M}$-Diff, an Anatomy-Aware adaptation framework that leverages pretrained vision and language models to meet the specific demands of Med-Diff-VQA. Specifically,$\mathbf{A}^{\mathbf{2}}$M-Diff utilizes Medical Masked Autoencoders (MedMAE) and the Medical Segment Anything Model (MedSAM) to extract both global contextual and anatomyfocused visual features. These features are token-compressed, projected, and injected into a pre-trained LLaMA2 language model via prompt-guided multimodal alignment. To efficiently adapt the language model to the medical domain with minimal additional parameters, we adopt Low-Rank Adaptation (LoRA), which updates only a small subset of model parameters. Experimental results on the MIMIC-Diff-VQA dataset demonstrate that$\mathbf{A}^{\mathbf{2}}$M-Diff outperforms existing methods, achieving a BLEU4 score of 0.542, METEOR of 0.412, ROUGE-L of 0.734, and CIDEr of 2.162. These results validate the effectiveness of anatomy-aware representation and lightweight adaptation in finegrained medical reasoning. The code is publicly available at: https://github.com/liyiersan/Med-Diff-VQA. Qian Zhou 0001, Hua Zou 0002, Fei Luo 0004, Xiwen Bai |
BIBM | 4 |
| 2025 | ReDACT: Reconstructing Detailed Avatar with Controllable Texture
Zezheng Chen, Huizhi Zhu, Fei Luo 0004, Chunxia Xiao |
CASA | 3 |
| 2025 | iG-6DoF: Model-free 6DoF Pose Estimation for Unseen Object via Iterative 3D Gaussian SplattingabstractTraditional methods in pose estimation often rely on precise 3D models or additional data such as depth and normals, limiting their generalization, especially when objects undergo large translations or rotations. We propose iG6DoF, a novel model-free 6D pose estimation method using iterative 3D Gaussian Splatting to estimate the pose of unseen objects. We first estimates an initial pose by leveraging multi-scale data augmentation and the rotation-equivariant features to create a better pose hypothesis from a set of candidates. Then, we propose an iterative 3DGS approach through iteratively rendering and comparing the rendered image with the input image to further progressively improve pose estimation accuracy. The proposed method consists of an object detector, a multi-scale rotation-equivariant feature based initial pose estimator, and a coarse-to-fine pose refiner. Such combination allows our method to focus on the target object in a complex scene dealing with large movement and weak textures. Our method achieves state-of-the-art results on the LINEMOD, OnePose-LowTexture, GenMOP datasets and our self-captured data, demonstrating its strong generalization to unseen objects and robustness across various scenes. Tuo Cao, Fei Luo 0004, Jiongming Qin, Yu Jiang 0007, Yusen Wang 0002, Chunxia Xiao |
CVPR | 2 |
| 2025 | JumpingGS: Level-jump 3D Gaussian Representation for Delicate Textures in Aerial Large-scale Scene RenderingabstractExisting 3D Gaussian (3DGS) based methods tend to produce blurriness and artifacts on delicate textures (small objects and high-frequency textures) in aerial large-scale scenes. The reason is that the delicate textures usually occupy a relatively small number of pixels, and the accumulated gradients from loss function are difficult to promote the splitting of 3DGS. To minimize the rendering error, the model will use a small number of large Gaussians to cover these details, resulting in blurriness and artifacts. To solve the above problem, we propose a novel hierarchical Gaussian: JumpingGS. JumpingGS assigns different levels to Gaussians to establish a hierarchical representation. Low-level Gaussians are responsible for the coarse appearance, while high-level Gaussians are responsible for the details. First, we design a splitting strategy that allows low-level Gaussians to skip intermediate levels and directly split the appropriate high-level Gaussians for delicate textures. This level-jump splitting ensures that the weak gradients of delicate textures can always activate a higher level instead of being ignored by the intermediate levels. Second, JumpingGS reduces the gradient and opacity thresholds for density control according to the representation levels, which improves the sensitivity of high-level Gaussians to delicate textures. Third, we design a novel training strategy to detect training views in hard-to-observe regions, and train the model multiple times on these views to alleviate underfitting. Experiments on aerial large-scale scenes demonstrate that JumpingGS outperforms existing 3DGS-based methods, accurately and efficiently recovering delicate textures in large scenes. Jiongming Qin, Kaixuan Zhou, Yu Jiang 0007, Huizhi Zhu, Fei Luo 0004, Chunxia Xiao |
ACM Trans. Graph. | 5 |
| 2025 | HumanIR-MGI: human inverse rendering via jointly optimizing geometry, material, and illumination
Ruhao Wang, Yu Jiang 0007, Huizhi Zhu, Fei Luo 0004, Chunxia Xiao |
Vis. Comput. | 4 |
| 2025 | EE-Head: emotion estimation for precise facial expression in NeRF head avatars
Enxu Zhao, Jianchi Sun, Fei Luo 0004, Chunxia Xiao |
Vis. Comput. | 3 |
| 2024 | DLCA-Recon: Dynamic Loose Clothing Avatar Reconstruction from Monocular VideosabstractReconstructing a dynamic human with loose clothing is an important but difficult task. To address this challenge, we propose a method named DLCA-Recon to create human avatars from monocular videos. The distance from loose clothing to the underlying body rapidly changes in every frame when the human freely moves and acts. Previous methods lack effective geometric initialization and constraints for guiding the optimization of deformation to explain this dramatic change, resulting in the discontinuous and incomplete reconstruction surface.To model the deformation more accurately, we propose to initialize an estimated 3D clothed human in the canonical space, as it is easier for deformation fields to learn from the clothed human than from SMPL.With both representations of explicit mesh and implicit SDF, we utilize the physical connection information between consecutive frames and propose a dynamic deformation field (DDF) to optimize deformation fields. DDF accounts for contributive forces on loose clothing to enhance the interpretability of deformations and effectively capture the free movement of loose clothing. Moreover, we propagate SMPL skinning weights to each individual and refine pose and skinning weights during the optimization to improve skinning transformation. Based on more reasonable initialization and DDF, we can simulate real-world physics more accurately. Extensive experiments on public and our own datasets validate that our method can produce superior results for humans with loose clothing compared to the SOTA methods. Chunjie Luo, Fei Luo 0004, Yusen Wang 0002, Enxu Zhao, Chunxia Xiao |
AAAI | 2 |
| 2024 | MFSegDiff: A Multi-Frequency Diffusion Model for Medical Image SegmentationabstractMedical image segmentation accurately identifies and delineates diagnostic regions, which is a crucial step in the early detection and accurate diagnosis of diseases. Diffusion models have demonstrated remarkable performance in preserving image details and structures by gradually adding noise followed by a reverse denoising process, making them widely explored in image segmentation tasks. In this study, we propose a novel approach for medical image segmentation utilizing diffusion models, termed MFSegDiff, which frames the segmentation task as an iterative denoising process. To address the inconsistency between image semantic features and noise embeddings, we introduce a Cross-Attention Alignment Module(CAAM). This module enhances the original image features and integrates noise and semantic information into the network through a linear attention mechanism. Additionally, we employ a Global-Local Multi-Frequency Module (GLMFM) to extract global contextual information from the spatial to the frequency domain by integrating multi-frequency features with local features. We evaluate the proposed on three datasets, including ISIC-2017, ISIC-2018, and ROSE. Results demonstrate strong generalization capabilities and achieve state-of-the-art segmentation performance, highlighting significant potential for clinical applications. Zidi Shi, Hua Zou 0002, Fei Luo 0004, Zhiyu Huo |
BIBM | 3 |
| 2024 | EnlightenDepth: a Novel Self-supervised Monocular Depth Estimation with Low-light Enhancement in EndoscopyabstractIn minimally invasive surgery, depth estimation is important for enhancing the perceptual abilities of surgeons. Current monocular self-supervised depth estimation methods suffer from the problems of poor illumination and narrow-view field in endoscopy. To address them, we propose a novel self-supervised monocular depth estimation with low-light enhancement, named EnlightenDepth. First, we introduce a GAN-based module to improve the local brightness and visibility of endoscopic images. Then, we leverage the appearance flow to handle the illumination variance between the adjacent image frames, assisting the supervision. Furthermore, to better extract global and local features, we apply MonoVit to combine the advantages of CNN and ViT. Extensive experiments are conducted on the Hamlyn dataset, proving that low-light enhancement can improve depth estimation performance in endoscopic illumination conditions. Compared with other monocular self-supervised methods designed for endoscopy, our method achieves around 10% improvement. Chi Kong Tam, Arafat Shantu, Jun Kui Zhang, Fei Luo 0004 |
BIBM | 6 |
| 2024 | Multi-Scale Implicit Surface Reconstruction for Outdoor Scenes
Ruhao Wang, Fei Luo 0004, Chunxia Xiao |
CVM (1) | 3 |
| 2024 | Diffusion-FOF: Single-View Clothed Human Reconstruction via Diffusion-Based Fourier Occupancy FieldabstractReconstructing a clothed human from a single-view image has several challenging issues, including flexibly representing various body shapes and poses, estimating complete 3D geometry and consistent texture, and achieving more fine-grained details. To address them, we propose a new diffusion-based Fourier occupancy field method to improve the human representing ability and the geometry generating ability. First, we estimate the back-view image from the given reference image by incorporating a style consistency constraint. Then, we extract multi-scale features of the two images as conditional and design a diffusion model to generate the Fourier occupancy field in the wavelet domain. We refine the initial estimated Fourier occupancy field with image features as conditions to improve the geometric accuracy. Finally, the reference and estimated back-view images are mapped onto the human model, creating a textured clothed human model. Substantial experiments are conducted, and the experimental results show that our method outperforms the state-of-the-art methods in geometry and texture reconstruction performance. Yuanzhen Li, Fei Luo 0004, Chunxia Xiao |
CVPR | 2 |
| 2024 | HS-Surf: A Novel High-Frequency Surface Shell Radiance Field to Improve Large-Scale Scene RenderingabstractPrevious neural radiance fields often struggle to preserve high-frequency textures in urban and aerial large-scale scenes due to insufficient model capacity on the scene surface. This is attributed to their sampling locations or grid vertices falling in empty areas. Additionally, most models do not consider the drastic changes in distances. To address these issues, we propose a novel high-frequency surface shell radiance field, which uses depth-guided information to create a shell enveloping the scene surface under the current view, and then samples conic frustums on this shell to render high-frequency textures. Specifically, our method comprises three parts. Initially, we propose a strategy to fuse voxel grids and information of distance scales to generate a coarse scene at different distance scales. Subsequently, we construct a shell based on the depth information to carry out compensation to incorporate texture details not captured by voxels. Finally, the smooth and denoise post-processing further improves the rendering quality. Substantial scene experiments and ablation experiments demonstrate that our method achieves the obvious improvement of high-frequency textures at different distance scales and outperforms the state-of-the-art methods. Jiongming Qin, Fei Luo 0004, Tuo Cao, Wenju Xu, Chunxia Xiao |
ACM Multimedia | 2 |
| 2024 | DGECN++: A Depth-Guided Edge Convolutional Network for End-to-End 6D Pose Estimation via Attention MechanismabstractMonocular object 6D pose estimation is a fundamental yet challenging task in computer vision. Recently, deep learning has been proven to be capable of predicting remarkable results in this task. Existing works often adopt a two-stage pipeline with establishing 2D-3D correspondences and utilizing a PnP/RANSAC or differentiable PnP algorithm to recover 6 degrees-of-freedom (6DoF) pose parameters. However, most of them hardly consider the geometric features in 3D space, and ignore the topological cues when performing differentiable PnP algorithms. To this end, we present an improved end-to-end monocular 6D pose estimation method (DGECN++) that incorporates depth estimation and a geometric-aware learnable PnP network. Our method is based on keypoints. First we detect the 2D keypoints that correspond to the 3D model. We then integrate differentiable PnP/RANSAC algorithm to create an end-to-end pipeline for 6D pose estimation. We focuses on the following three key aspects: 1) We utilize the estimated depth information to guide the process of extracting 2D-3D correspondences and refine the results using a cascaded differentiable PnP/RANSAC algorithm that incorporates geometric information. 2) We leverage the uncertainty of the estimated depth map to enhance the accuracy and robustness of the predicted 6D pose. 3) We propose a differentiable Perspective-n-Point (PnP) algorithm based on edge convolution and self-attention to explore the topological relationships between 2D-3D correspondences. Experimental results demonstrate that our proposed network surpasses existing methods in terms of both effectiveness and efficiency. Tuo Cao, Yanping Fu, Shengjie Zheng, Fei Luo 0004, Chunxia Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Towards High-Quality Photorealistic Image Style TransferabstractPreserving important textures of the content image and achieving prominent style transfer results remains a challenge in the field of image style transfer. This challenge arises from the entanglement between color and texture during the style transfer process. To address this challenge, we propose an end-to-end network that incorporates adaptive weighted least squares (AWLS) filter, iterative least squares (ILS) filter, and channel separation. Given a content image ($\mathcal {C}$) and a reference style image ($\mathcal {S}$), we begin by separating the RGB channels and utilizing ILS filter to decompose them into structure and texture layers. We then perform style transfer on the structural layers using WCT$^{2}$(incorporating wavelet pooling and unpooling techniques for whitening and coloring transforms) in the R, G, and B channels, respectively. We address the texture distortion caused by WCT$^{2}$with a texture enhancing (TE) module in the structural layer. Furthermore, we propose an estimating and compensating for the structure loss (ECSL) module. In the ECSL module, with the AWLS filter and the ILS filter, we estimate the texture loss caused by TE, convert the loss of the structural layer to the loss of the texture layer, and compensate for the loss in the texture layer. The final structural layer and the texture layer are merged into the channel style transfer results in the separated R, G, and B channels into the final style transfer result. Thereby, this enables a more complete texture preservation and a significant style transfer process. To evaluate our method, we utilize quantitative experiments using various metrics, including NIQE, AG, SSIM, PSNR, and a user study. The experimental results demonstrate the superiority of our approach over the previous state-of-the-art methods. Haimin Zhang 0001, Gang Fu 0003, Caoqing Jiang, Fei Luo 0004, Chunxia Xiao, Min Xu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | CPAConv-POCO:a Continuous Position Adaptive Convolution based POCO for lung nodule 3D ReconstructionabstractThe development of CT technology has played a crucial role in assisting doctors in the diagnosis of lung nodules. However, due to their three-dimensional morphology and complex structure, two-dimensional medical images are insufficient for intuitive visualization and analysis of lung nodules, making three-dimensional reconstruction necessary to address this issue. In this study, we propose a novel three-dimensional reconstruction framework called CPAConv-POCO, which is based on point convolution. This framework employs continuous convolution instead of discrete convolution commonly used in traditional image processing tasks to handle unstructured data such as point cloud. Additionally, we introduce a new convolution kernel construction method called Position Adaptive Convolution (PAConv). PAConv dynamically assembles convolution kernels by combining basic weight matrices stored in Weight Bank. The coefficients of these weight matrices are adaptively learned from the positions of points using ScoreNet. This data-driven construction of kernels provides the flexibility of PAConv, allowing it to better handle irregular and unordered point cloud data. The convolution module obtained by combining these two components is named Continuous Position Adaptive Convolution (CPAConv). In our experiments, we extensively evaluated our method on the publicly available LUNA16 and LNDb datasets. In terms of reconstruction accuracy, the Intersection over Union(IoU) of CPAConv-POCO’s lung nodule reconstruction reaches 85.70%, which is 1.51% higher than POCO. On the LNDb dataset, the IoU of our method reaches 71.33%, which is 3.33% higher than POCO. We also evaluated the performance of our model on lung nodules of different diameters. On the LUNA16 dataset, the IoU of our method reaches 90.19% on lung nodules with a diameter smaller than 5mm, demonstrating superior reconstruction performance for small lung nodules. Ao Jiang, Ruoshan Kong, Fei Luo 0004, Wen Cai Huang, Jia Ni Zou |
BIBM | 5 |
| 2023 | ConvUNET: a Novel Depthwise Separable ConvNet for Lung Nodule SegmentationabstractLung nodule segmentation is usually considered a 3D semantic segmentation task. Due to the small size, diverse morphology, and low recognition of lung nodules, it is hard to segment any nodule precisely. To solve this problem, we propose a lightweight depthwise separable convolutional network named ConvUNET, which consists of a hierarchical encoder and a U-shaped decoder. Compared with some Transformer-based models (e.g., SwinUNETR) and ConvNeXt-based models (e.g., 3D UX-Net), our model has the advantages of fewer parameters, faster inference speed, and higher accuracy. We test the segmentation performance on the LUNA-16 and LNDb-19 datasets using standard 5-fold cross-validations, and the proposed method achieves competitive dice scores of 88.90% and 84.16%, respectively. Besides, it also shows considerable precision in segmenting lung nodules with diverse characteristics. Our source code is available at https://github.com/Xinkai-Tang/ConvUNET. Xinkai Tang, Ruoshan Kong, Fei Luo 0004, Wen Cai Huang, Jia Ni Zou |
BIBM | 4 |
| 2023 | DFNodule: a Novel Deformable Faster R-CNN for Lung Nodule DetectionabstractIn computer-aided diagnosis systems, lung nodule detection plays a crucial role in the overall framework. In this work, we propose a new three-dimensional deformable convolutional neural network (dcnn) method for lung nodule detection based on the Faster R-CNN framework. We incorporate deformable convolutions to design a hybrid convolutional module, which enhances feature extraction in the lung nodule detection model. By leveraging the deformable convolutions’ characteristics, the network is capable of capturing the diverse morphological variations of lung nodules, addressing challenges such as large morphological variations and the inability to capture unified image features. This improves the accuracy of the lung nodule detection algorithm. Additionally, we employ a second-stage network to further discriminate suspected nodules, which enhances the recognition of non-nodule tissues and reduces false positive nodules. To comprehensively evaluate the performance of various lung nodule detection models, we conducted experiments using the publicly available LUNA16 dataset. Our method surpasses other detection algorithms in terms of CPM, achieving a 1.5% improvement. Particularly, the nodule recognition rate is significantly improved at lower false positive rates. In addition, in other metrics such as F1-score, AP, we also achieved 0.6%, 2% improvement. GuoWei Tao, Fu Zhou, Hao Gui, Fei Luo 0004, Wen Cai Huang, Jia Ni Zou, Yi-Ping Phoebe Chen |
BIBM | 5 |
| 2023 | RHViT: A Robust Hierarchical Transformer for 3D Multimodal Brain Tumor Segmentation Using Biased Masked Image Modeling Pre-trainingabstractAccurate brain tumor segmentation in medical image analysis is crucial for diagnosis and treatment planning. While computer-aided methods have shown promise, several challenges persist. Most existing methods struggle with smaller tumors, treating all regions uniformly. Additionally, they lack robustness when dealing with data corruption and handling missing modalities, common in clinical settings. In this paper, we present a robust hierarchical vision transformer (RHViT) for 3D multimodal brain tumor segmentation, employing an encoder-decoder structure. Our approach combines 3D convolutions and self-attention, offering efficient and effective training. 3D convolutions help capture local information and generate hierarchical features, improving tumor segmentation accuracy. To enhance robustness, we pre-train the encoder using masked image modeling (MIM). This pre-training equips the model to handle data corruption, resulting in improved segmentation even in challenging scenarios. Furthermore, we introduce a novel biased masking strategy during MIM to focus the model's attention on tumor regions. This facilitates better tumor representations and effective fusion of multimodal features. Importantly, our biased masking technique strengthens the model's resilience when dealing with incomplete multimodal data during testing, making it a practical choice. Extensive experiments confirm the superiority of our model over existing approaches. Qian Zhou 0001, Hua Zou 0002, Fei Luo 0004, Yishi Qiu |
BIBM | 3 |
| 2023 | NeTO: Neural Reconstruction of Transparent Objects with Self-Occlusion Aware Refraction-TracingabstractWe present a novel method called NeTO, for capturing the 3D geometry of solid transparent objects from 2D images via volume rendering. Reconstructing transparent objects is a very challenging task, which is ill-suited for general-purpose reconstruction techniques due to the specular light transport phenomena. Although existing refraction-tracing-based methods, designed especially for this task, achieve impressive results, they still suffer from unstable optimization and loss of fine details since the explicit surface representation they adopted is difficult to be optimized, and the self-occlusion problem is ignored for refraction-tracing. In this paper, we propose to leverage implicit Signed Distance Function (SDF) as surface representation and optimize the SDF field via volume rendering with a self-occlusion aware refractive ray tracing. The implicit representation enables our method to be capable of reconstructing high-quality reconstruction even with a limited set of views, and the self-occlusion aware strategy makes it possible for our method to accurately reconstruct the self-occluded regions. Experiments show that our method achieves faithful reconstruction results and outperforms prior works by a large margin. Visit our project page at https://www.xxlong.site/NeTO/. Zongcheng Li, Xiaoxiao Long, Yusen Wang 0002, Tuo Cao, Wenping Wang 0001, Fei Luo 0004, Chunxia Xiao |
ICCV | 6 |
| 2023 | Learning Long-range Information with Dual-Scale Transformers for Indoor Scene CompletionabstractDue to the limited resolution of 3D sensors and the inevitable mutual occlusion between objects, 3D scans of real scenes are commonly incomplete. Previous scene completion methods struggle to capture long-range spatial context, resulting in unsatisfactory completion results. To alleviate the problem, we propose a novel Dual-Scale Transformer Network (DST-Net) that efficiently utilizes both long-range and short-range spatial context information to improve the quality of 3D scene completion. To reduce the heavy computation cost of extracting long-range features via transformers, DST-Net adopts a self-supervised two-stage completion strategy. In the first stage, we split the input scene into blocks and perform completion on individual blocks. In the second stage, the blocks are merged together as a whole and then further refined to improve completeness. More importantly, we propose a contrastive attention training strategy to encourage the transformers to learn distinguishable features for better scene completion. Experiments on datasets of Matterport3D, ScanNet, and ICL-NUIM demonstrate that our method can generate better completion results, and our method outperforms the state-of-the-art methods quantitatively and qualitatively. Fei Luo 0004, Xiaoxiao Long, Chunxia Xiao |
ICCV | 2 |
| 2023 | Self-Supervised Monocular Depth Estimation by Digging into Uncertainty Quantification
Yuanzhen Li, Shengjie Zheng, Zi-Xin Tan, Tuo Cao, Fei Luo 0004, Chunxia Xiao |
J. Comput. Sci. Technol. | 5 |
| 2023 | Monocular human depth estimation with 3D motion flow and surface normals
Yuanzhen Li, Fei Luo 0004, Chunxia Xiao |
Vis. Comput. | 2 |
| 2023 | Sparse RGB-D images create a real thing: A flexible voxel based 3D reconstruction pipeline for single objectabstractReconstructing 3D models for single objects with complex backgrounds has wide applications like 3D printing, AR/VR, and so on. It is necessary to consider the tradeoff between capturing data at low cost and getting high-quality reconstruction results. In this work, we propose a voxel-based modeling pipeline with sparse RGB-D images to effectively and efficiently reconstruct a single real object without the geometrical post-processing operation on background removal. First, referring to the idea of VisualHull, useless and inconsistent voxels of a targeted object are clipped. It helps focus on the target object and rectify the voxel projection information. Second, a modified TSDF calculation and voxel filling operations are proposed to alleviate the problem of depth missing in the depth images. They can improve TSDF value completeness for voxels on the surface of the object. After the mesh is generated by the MarchingCube, texture mapping is optimized with view selection, color optimization, and camera parameters fine-tuning. Experiments on Kinect capturing dataset, TUM public dataset, and virtual environment dataset validate the effectiveness and flexibility of our proposed pipeline. Fei Luo 0004, Yongqiong Zhu, Yanping Fu, Huajian Zhou, Zezheng Chen, Chunxia Xiao |
Vis. Informatics | 1 |
| 2022 | ViFUNet: a Vision Flash based UNet for lung nodules segmentation taskabstractAs an efficient computer-aided diagnosis technology, pulmonary nodule image segmentation plays an important role in greatly improving the identification efficiency. Due to their different shapes and sizes, there is always tremendous difficulty in the segmentation of lung nodules. The attention and Vision Transformer mechanisms are used in such scenarios, such as TransUNet, to solve the problem faced by traditional CNNs to obtain local information as well as have a good segmentation result. However, the addition to a network leads to it being bloated, increases the number of network parameters, slows down the training speed as well as makes it more prone to overfitting. In this paper, we propose an image segmentation network ViFUNet which is based on the Flash attention mechanism. Compared with the traditional Transformer, Flash adopts the fusion attention mechanism, which greatly reduces the number of network parameters and improves the network training speed. Experiments show that the convergence speed of ViFUNet is faster than that of the ViT-based network. In terms of segmentation accuracy, the Dice coefficient of ViFUNet’s lung nodule segmentation reaches 86.03%,which is 1.49% higher than U-Net and 2.66% higher than TransUNet, and it also performs well for different types of lung nodules. The average DSC score of pulmonary nodules less than 4mm in diameter reaches 85.20%, which is 1.37% higher than Atten-UNet, 5.03% higher than TransUNet, and 4.31% higher than U-Net. In addition, its parameter size is only 12. 7M, which is about half of the UNet and one-tenth of the TransUNet. Taoyu Chen, Fei Luo 0004, Fu Zhou, Yunfei Zha |
BIBM | 2 |
| 2022 | DGECN: A Depth-Guided Edge Convolutional Network for End-to-End 6D Pose EstimationabstractMonocular 6D pose estimation is a fundamental task in computer vision. Existing works often adopt a two-stage pipeline by establishing correspondences and utilizing a RANSAC algorithm to calculate 6 degrees-of-freedom (6DoF) pose. Recent works try to integrate differentiable RANSAC algorithms to achieve an end-to-end 6D pose estimation. However, most of them hardly consider the geometric features in 3D space, and ignore the topology cues when performing differentiable RANSAC algorithms. To this end, we proposed a Depth-Guided Edge Convolutional Network (DGECN) for 6D pose estimation task. We have made efforts from the following three aspects: 1) We take advantages of estimated depth information to guide both the correspondences-extraction process and the cascaded differentiable RANSAC algorithm with geometric information. 2) We leverage the uncertainty of the estimated depth map to improve accuracy and robustness of the output 6D pose. 3) We propose a differentiable Perspective-n-Point(PnP) algorithm via edge convolution to explore the topology relations between 2D-3D correspondences. Experiments demonstrate that our proposed network outperforms current works on both effectiveness and efficiency. Tuo Cao, Fei Luo 0004, Yanping Fu, Shengjie Zheng, Chunxia Xiao |
CVPR | 2 |
| 2022 | PhraseGAN: Phrase-Boost Generative Adversarial Network for Text-to-Image GenerationabstractA phrase contains an object-orienting noun and some attribution-associating words. Therefore, focusing on phrases could better generate images with the objects and their tightly relevant characteristics. We propose a Phrase-boost Gener-ative Adversarial Network (PhraseGAN) with threefold im-provement for scene level text-to-image generation. First, we propose a Transformer-based encoder to encode the in-put words and sentences and encode related words and their targeting nouns into phrases by text correlation analysis. Sec-ond, we utilize Graph Convolution Networks to measure fine-grained text-image similarity, which could gain constraints on relative positions between different objects. Finally, we de-sign a phrase-region discriminator to discriminate the qual-ity of the generated objects and the consistency between the phrases and their corresponding objects. Experimental results on the Microsoft COCO dataset demonstrate that PhraseGAN can generate better images from texts than state-of-the-art methods. Fei Luo 0004, Chengjiang Long, Shenghong Hu, Chunxia Xiao |
ICME | 3 |
| 2022 | Discriminator Modification in GAN for Text-to-Image GenerationabstractThe existing Generative Adversarial Network-based text-to-image generation methods suffer from mode collapse and training instability. This paper relieves these problems by improving the discriminator ability from three aspects. First, we propose a diversity-sensitive conditional discriminator (D-SCD), which increases the diversity of the generated images by judging the combination of the generated image and mismatched text as false. Second, for the unconditional discriminator, we propose a contrastive searching gradient penalty (CSGP) strategy to measure the realism of the generated images and to penalize the gradients for stabilizing the training process. Finally, we introduce a multi-level images similarity (MLIS) loss for the discriminator feature extractor to further promote the high-level feature similarity between the real and generated images and objects. Extensive experimental results and ablation studies demonstrate that our modifications on the discriminators can effectively improve the quality of the generated images. Fei Luo 0004, Chunxia Xiao |
ICME | 3 |
| 2022 | Photorealistic Style Transfer via Adaptive Filtering and Channel SeperationabstractThe problem of color and texture distortion remains unsolved in the photorealistic style transfer task. It is mainly caused by the interference between color and texture during transferring. To address this problem, we propose a end-to-end network via adaptive filtering and channel separation. Given a pair of content image and reference image, we firstly decompose them into two structure layers through adaptive weighted least squares filter (AWLSF), which could better perceive the color structure and illumination. Then, we carry out RGB transfer in a channel separation way on the two generated structure layers. To deal with texture in a relatively independent manner, we use a module and a subtraction operation to get more complete and clear content features. Finally, we merge the color structure and texture detail into the ultimate result. We conduct solid quantitative experiments on four metrics NIQE, AG, SSIM, and PSNR, and make a user study. The experimental results demonstrate that our method is able to produce better results than previous state-of-the-art methods, and validate the effectiveness and superiority of our method. Fei Luo 0004, Caoqing Jiang, Gang Fu 0003, Zipei Chen, Shenghong Hu, Chunxia Xiao |
ACM Multimedia | 2 |
| 2022 | Self-supervised coarse-to-fine monocular depth estimation using a lightweight attention moduleabstractSelf-supervised monocular depth estimation has been widely investigated and applied in previous works. However, existing methods suffer from texture-copy, depth drift, and incomplete structure. It is difficult for normal CNN networks to completely understand the relationship between the object and its surrounding environment. Moreover, it is hard to design the depth smoothness loss to balance depth smoothness and sharpness. To address these issues, we propose a coarse-to-fine method with a normalized convolutional block attention module (NCBAM). In the coarse estimation stage, we incorporate the NCBAM into depth and pose networks to overcome the texture-copy and depth drift problems. Then, we use a new network to refine the coarse depth guided by the color image and produce a structure-preserving depth result in the refinement stage. Our method can produce results competitive with state-of-the-art methods. Comprehensive experiments prove the effectiveness of our two-stage method using the NCBAM. Yuanzhen Li, Fei Luo 0004, Chunxia Xiao |
Comput. Vis. Media | 2 |
| 2021 | HAUNet-3D: a Novel Hierarchical Attention 3D UNet for Lung Nodule SegmentationabstractUNet and its extended versions are the most used networks in the lung nodule segmentation from CT images. However, current UNet-like methods still suffer from some problems: 1) The heterogeneity of lung nodules affect the segmentation performance; 2) the mixture of lung nodules and their surrounding tissues in the CT image increases the segmentation difficulty. To address these issues, we propose a novel hierarchical attention 3D UNet named HAUNet-3D. It introduces the attention mechanism at multiple scales and organizes them in a bottom-up hierarchical connection way. Such a proposition could better capture features with various sizes and guide the fusion of features from adjacent attention outputs without losing the advantages of 3D UNet. In experiment, our method has been extensively evaluated on the public LUNA16 dataset. It achieves competitive segmentation performance on dice similarity coefficient of 83.34% and average surface distance of 0.28 mm. More importantly, our method is proven to be more robust to the heterogeneous types of lung nodules and shows better segmentation performance on small lung nodules. Fu Zhou, Fei Luo 0004, Kafui Efio-Akolly, Ronald Bbosa, Wen Cai Huang, Jia Ni Zou, Yi-Ping Phoebe Chen |
BIBM | 2 |
| 2021 | Stable Depth Estimation Within Consecutive Video Frames
Fei Luo 0004, Chunxia Xiao |
CGI | 1 |
| 2021 | Adaptive depth estimation for pyramid multi-view stereo
Yanping Fu, Qingan Yan, Fei Luo 0004, Chunxia Xiao |
Comput. Graph. | 4 |
| 2021 | Self-supervised monocular depth estimation based on image texture detail enhancement
Yuanzhen Li, Fei Luo 0004, Shenjie Zheng, Huanhuan Wu, Chunxia Xiao |
Vis. Comput. | 2 |
| 2020 | A Comprehensive Pipeline for Complex Text-to-Image Synthesis
Fei Luo 0004, Hongpan Zhang, Hua-Jian Zhou, Alix L. H. Chow, Chunxia Xiao |
J. Comput. Sci. Technol. | 2 |
| 2019 | A systematic evaluation of copy number alterations detection methods on real SNP array and deep sequencing dataabstractBACKGROUND: The Copy Number Alterations (CNAs) are discovered to be tightly associated with cancers, so accurately detecting them is one of the most important tasks in the cancer genomics. A series of CNAs detection methods have been proposed and new ones are still being developed. Due to the complexity of CNAs in cancers, no CNAs detection method has been accepted as the gold standard caller. Several evaluation works have made attempts to reveal typical CNAs detection methods' performance. Limited by the scale of evaluation data, these different comparison works don't reach a consensus and the researchers are still confused on how to choose one proper CNAs caller for their analysis. Therefore, it needs a more comprehensive evaluation of typical CNAs detection methods' performance. RESULTS: In this work, we use a large-scale real dataset from CAGEKID consortium to evaluate total 12 typical CNAs detection methods. These methods are most widely used in cancer researches and always used as benchmark for the newly proposed CNAs detection methods. This large-scale dataset comprises of SNP array data on 94 samples and the whole genome sequencing data on 10 samples. Evaluations are comprehensively implemented in current scenarios of CNAs detection, which include that detect CNAs on SNP array data, on sequencing data with tumor and normal matched samples and on sequencing data with single tumor sample. Three SNP based methods are firstly ranked. Subsequently, the best SNP based method's results are used as benchmark to compare six matched samples based methods and three single tumor sample based methods in terms of the preprocessing, recall rate, Jaccard index and segmentation characteristics. CONCLUSIONS: Our survey thoroughly reveals 12 typical methods' superiority and inferiority. We explain why methods show specific characteristics from a methodological standpoint. Finally, we present the guiding principle for choosing one proper CNAs detection method under specific conditions. Some unsolved problems and expectations are also addressed for upcoming CNAs detection methods. Fei Luo 0004 |
BMC Bioinform. | 1 |
| 2019 | Illumination animating and editing in a single picture using scene structure estimation
Bin Liao 0006, Chao Liang 0001, Fei Luo 0004, Chunxia Xiao |
Comput. Graph. | 4 |
| 2019 | Joint bilateral propagation upsampling for unstructured multi-view stereo
Mengqiang Wei, Qingan Yan, Fei Luo 0004, Chengfang Song, Chunxia Xiao |
Vis. Comput. | 3 |
| 2018 | Compare Copy Number Alterations Detection Methods on Real Cancer Data
Fei Luo 0004, Yongqiong Zhu |
ICIC (1) | 1 |
| 2017 | Predicting potential drug-drug interactions by integrating chemical, biological, phenotypic and network dataabstractBACKGROUND: Drug-drug interactions (DDIs) are one of the major concerns in drug discovery. Accurate prediction of potential DDIs can help to reduce unexpected interactions in the entire lifecycle of drugs, and are important for the drug safety surveillance. RESULTS: Since many DDIs are not detected or observed in clinical trials, this work is aimed to predict unobserved or undetected DDIs. In this paper, we collect a variety of drug data that may influence drug-drug interactions, i.e., drug substructure data, drug target data, drug enzyme data, drug transporter data, drug pathway data, drug indication data, drug side effect data, drug off side effect data and known drug-drug interactions. We adopt three representative methods: the neighbor recommender method, the random walk method and the matrix perturbation method to build prediction models based on different data. Thus, we evaluate the usefulness of different information sources for the DDI prediction. Further, we present flexible frames of integrating different models with suitable ensemble rules, including weighted average ensemble rule and classifier ensemble rule, and develop ensemble models to achieve better performances. CONCLUSIONS: The experiments demonstrate that different data sources provide diverse information, and the DDI network based on known DDIs is one of most important information for DDI prediction. The ensemble methods can produce better performances than individual methods, and outperform existing state-of-the-art methods. The datasets and source codes are available at https://github.com/zw9977129/drug-drug-interaction/ . Wen Zhang 0008, Yanlin Chen 0002, Fei Luo 0004, Gang Tian, Xiaohong Li 0003 |
BMC Bioinform. | 4 |
| 2016 | A genetic algorithm-based weighted ensemble method for predicting transposon-derived piRNAsabstractBACKGROUND: Predicting piwi-interacting RNA (piRNA) is an important topic in the small non-coding RNAs, which provides clues for understanding the generation mechanism of gamete. To the best of our knowledge, several machine learning approaches have been proposed for the piRNA prediction, but there is still room for improvements. RESULTS: In this paper, we develop a genetic algorithm-based weighted ensemble method for predicting transposon-derived piRNAs. We construct datasets for three species: Human, Mouse and Drosophila. For each species, we compile the balanced dataset and imbalanced dataset, and thus obtain six datasets to build and evaluate prediction models. In the computational experiments, the genetic algorithm-based weighted ensemble method achieves 10-fold cross validation AUC of 0.932, 0.937 and 0.995 on the balanced Human dataset, Mouse dataset and Drosophila dataset, respectively, and achieves AUC of 0.935, 0.939 and 0.996 on the imbalanced datasets of three species. Further, we use the prediction models trained on the Mouse dataset to identify piRNAs of other species, and the models demonstrate the good performances in the cross-species prediction. CONCLUSIONS: Compared with other state-of-the-art methods, our method can lead to better performances. In conclusion, the proposed method is promising for the transposon-derived piRNA prediction. The source codes and datasets are available in https://github.com/zw9977129/piRNAPredictor . Dingfang Li, Longqiang Luo, Wen Zhang 0008, Fei Luo 0004 |
BMC Bioinform. | 5 |
| 2016 | Multi-fields model for predicting target-ligand interaction
Caihua Wang, Juan Liu 0007, Fei Luo 0004, Qian-Nan Hu |
Neurocomputing | 3 |
| 2015 | A novel two-stage method for identifying microRNA-gene regulatory modules in breast cancerabstractIn this paper, we propose a two-stage method for identifying miRNA-gene regulatory modules by integrating miRNA/mRNA expression profiles and miRNA genomic cluster data. We first adopt a Multiple-output Sparse Group Lasso (MSGL) regression model to predict the miRNA-gene regulatory network. Further, we propose a L0-penalized Singular Value Decomposition (L0-SVD) model to identify modules from the predicted network. We apply this method to miRNA and mRNA expression profiles of the breast cancer data from TCGA databases and identify ten miRNA-gene regulatory modules. We find that (1) the modules are significantly associated in a predicted miRNA-gene regulatory network; (2) the modules are significantly enriched in GO biological processes and KEGG pathways, respectively; (3) many miRNAs and genes in the modules are related with breast cancer. On average, 51% of the miRNAs and 30% of the genes are related with breast cancer. The results demonstrate that miRNA-gene regulatory modules provide insights into the mechanisms of the combinatorial regulation between miRNAs and genes. Wenwen Min, Juan Liu 0007, Fei Luo 0004 |
BIBM | 3 |
| 2014 | Pairwise input neural network for target-ligand interaction predictionabstractPrediction the interactions between proteins (targets) and small molecules (ligands) is a critical task for the drug discovery in silico. In this work, we consider the target binding site instead of the whole target and propose a pairwise input neural network (PINN) for constructing the site-ligand interaction prediction model. Different with the ordinary artificial neural network (ANN) with one vector as input, the proposed PINN can accept a pair of vectors as the input, corresponding to a binding site and a ligand respectively. The 5-CV evaluation results show that PINN outperforms other representative target-ligand interaction prediction methods. Caihua Wang, Juan Liu 0007, Fei Luo 0004, Yafang Tan, Zixin Deng, Qian-Nan Hu |
BIBM | 3 |
| 2013 | Integrating peptides' sequence and energy of contact residues information improves prediction of peptide and HLA-I binding with unknown allelesabstractBACKGROUND: The HLA (human leukocyte antigen) class I is a kind of molecule encoded by a large family of genes and is characteristic of high polymorphism. Now the number of the registered HLA-I molecules has exceeded 3000. Slight differences in the amino acid sequences of HLAs would make them bind to different sets of peptides. In the past decades, although many methods have been proposed to predict the binding between peptides and HLA-I molecules and achieved good performance, most experimental data used by them is limited to the HLAs with a small number of alleles. Thus they are inclined to obtain high prediction accuracy only for data with similar alleles. Because the peptides and HLAs together determine the binding, it's necessary to consider their contribution meanwhile. RESULTS: By taking into account the features of the peptides sequence and the energy of contact residues, in this paper a method based on the artificial neural network is proposed to predict the binding of peptides and HLA-I even when the HLAs' potential alleles are unknown. Two experiments in the allele-specific and super-type cases are performed respectively to validate our method. In the first case, we collect 14 HLA-A and 14 HLA-B molecules on Bjoern Peters dataset, and compare our method with the ARB, SMM, NetMHC and other 16 online methods. Our method gets the best average AUC (Area under the ROC) value as 0.909. In the second one, we use leave one out cross validation on MHC-peptide binding data that has different alleles but shares the common super-type. Compared to gold standard methods like NetMHC and NetMHCpan, our method again achieves the best average AUC value as 0.847. CONCLUSIONS: Our method achieves satisfactory results. Whenever it's tested on the HLA-I with single definite gene or with super-type gene locus, it gets better classification accuracy. Especially, when the training set is small, our method still works better than the other methods in the comparison. Therefore, we could make a conclusion that by combining the peptides' information, HLAs amino acid residues' interaction information and contact energy, our method really could improve prediction of the peptide HLA-I binding even when there aren't the prior experimental dataset for HLAs with various alleles. Fei Luo 0004, Yangyang Gao, Yongqiong Zhu, Juan Liu 0007 |
BMC Bioinform. | 1 |
| 2012 | Predicting Binding-Peptide of HLA-I on Unknown Alleles by Integrating Sequence Information and Energies of Contact Residues
Fei Luo 0004, Yangyang Gao, Yongqiong Zhu, Juan Liu 0007 |
ICIC (3) | 1 |