EDBT 2026 Demo / reviewers in the wild / expert
Wenkang Fan
dblp:287/8310
· DBLP profile ↗
23ranked-venue papers
6as first author
23since 2021 · last 2026
0000-0002-8364-5159ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 5 first-author · 16 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Virtually Supervised Depth-Aware Registration for Monocular Endoscopic NavigationabstractRobotic-assisted endoscopy commonly uses various endoscopes for early detection and treatment of tumors or cancers. Tracking the endoscope 3D location from monocular endoscopic video sequences is the key to develop surgical navigation. This work proposes a new virtually supervised depth-aware registration method to track the monocular endoscope 3D location in the preoperative image space. Specifically, a virtually supervised monocular endoscopic structural depth estimation model is proposed and trained on virtual or synthesis endoscopic data, without using any manually annotated real endoscopic video images. This model can accomplish zero-shot generalization to precisely estimate dense depth maps of real endoscopic images with artifacts and illumination variations. Moreover, a new structure-invariant and depth-aware similarity function and spatial constraint are introduced for 2D-3D registration. We validated our method on clinical data from different medical centers, with the experimental results showing that the average tracked position and direction errors were reduced to (3.10±2.99mm, 7.22±7.13°), which significantly outperforms current vision-based surgical navigation methods. Guangcheng Luo, Ming Wu 0009, Wenkang Fan, Xiangxing Chen, Xióngbiao Luó |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Deep Bilateral Intensity-Position Registration for Autonomous Ureteroscopic NavigationabstractFlexible ureteroscopy is a routinely performed surgical procedure to treat renal disorders such as tumors and stones, but it gets trapped in precisely orientating the ureteroscope in the complex kidneys. To facilitate ureteroscopic procedures, we propose employing deep learning techniques for preliminary data processing and propose a new ureteroscopic navigation framework that uses deeply learned bilateral 2D-3D registration. Specifically, a new structural intensity-position similarity function is formulated to characterize the difference between 2D ureteroscopic video sequences and preoperative computed tomography urography images. While we propose a small deep learning model of deformable large-kernel convolutional networks without any transformer blocks to segment the urinary collecting system from preoperative images, we employ dense prediction transformers and a color model of hue-saturation-value to extract structural regions from ureteroscopic video sequences. The new cost function is designed by the dice similarity coefficient and structural similarity index to calculate the pixel intensity and position (coordinate) differences. We validated our method on clinical data collected from different patients in the operating room, with the experimental results showing that our method outperforms state-of-the-art registration approaches, reducing the navigation errors from (7.8 mm, 10.7°) to (7.1 mm, 9.7°). Xiangtao Du, Wenkang Fan, Guangcheng Luo, Xióngbiao Luó |
ECAI | 2 |
| 2025 | Deep Support Vein Machine for Lung ParcellationabstractPulmonary segments parcellation is essential to thoracoscopic segmentectomy. Surgeons manually outline pulmonary segments from preoperative images before surgery, which is a time-consuming, labor-intensive and mental-stress procedure. This work proposes a novel small learning model of deep support vein machine without using annotated pulmonary segments data for automatic lung parcellation. Specifically, this machine can learn anatomical structures of pulmonary lobe, bronchus, artery, and vein by two cascade multilayer perceptrons to automatically divides the lung into eighteen segments. The perceptron module typically smooths the boundary of the pulmonary segment to attain robust and precise parcellation. Additionally, three new metrics are defined to quantitatively evaluate the quality of lung parcellation. We validate our methods on 108 clinical pulmonary computed tomography scans, with the experimental results showing that our proposed machine certainly outperforms current methods and provides a promising way to fully automated lung parcellation. Particularly, the dice similarity coefficient of lung parcellation was significantly improved from 0.886 to 0.918. Haichao Peng, Wenkang Fan, Sunkui Ke, Xióngbiao Luó |
ICASSP | 3 |
| 2025 | Dual-Triple Transformer Networks for Accurate CT Pleural Effusion SegmentationabstractPleural effusion segmentation in computed tomography images is essential to its precise diagnosis and treatment but remains challenging due to blurred boundaries, heterogeneous morphology, and low contrast with adjacent anatomical structures. This work shows a first study on pleural effusion segmentation by introducing a new deep learning architecture of dual-triple transformer networks. Specifically, this architecture builds a dual encoder of swin transformer and 3D deformable convolution, leveraging the multiscale representation capability to capture global contextual information and model complex deformations. Moreover, a triple decoder with a fusion module, a boundary-awareness mechanism, and a transposed-residual convolution block is introduced to effectively propagate these global and local features and refine the segmentation by mitigating ambiguity at the interfaces with surrounding tissues. We validate our method on 143 chest computed tomography scans. The experimental results demonstrate that our proposed model significantly outperforms state-of-the-art segmentation approaches, greatly improving the dice similarity coefficient and mean intersection over union while critically reducing both Hausdorff distance 95 and average symmetric surface distance. Wenkang Fan, Xióngbiao Luó |
ICASSP | 2 |
| 2025 | Hybrid Attention-Residual Networks for Hepatic and Portal Veins Semantic Segmentation in MR Images
Wenkang Fan, Xióngbiao Luó |
ICIC (2) | 3 |
| 2025 | Dense Depth-Supervised Simultaneous Localization and Mapping for Robust Bronchoscopic Navigation
Xiuling Huang, Wenkang Fan, Xióngbiao Luó |
ICIC (27) | 2 |
| 2025 | Spatially Constrained and Deeply Learned Bilateral Structural Intensity-Depth Registration Autonomously Navigates a Flexible EndoscopeabstractEndoscope tracking is commonly utilized to provide surgeons with in-body camera poses and visual fields during invasive procedures. The fundamental aspect of endoscopic navigation lies in precisely and continuously tracing the position and orientation of the endoscope within monocular endoscopic video sequences in a preoperative data space. This work proposes a new spatially constrained and deeply learned bilateral structural intensity-depth 2D-3D registration framework for autonomously navigating a flexible endoscope. Concretely, a novel bilateral structural intensity-depth similarity function is defined to tackle the deficiency of using image intensity, while a cross-domain monocular depth estimation model trained on virtual image data is used to accurately predict real image dense depth. Additionally, a spatial constraint is introduced to precisely reinitialize an optimizer to reduce accumulative tracking errors. We validate our method on clinical data, with the experimental results showing that our method significantly outperforms current vision-based navigation methods. Particularly, the average of position and orientation errors were reduced from (4.59mm, 9.22°) to (1.65mm, 4.67°). Ming Wu 0009, Wenkang Fan, Guangcheng Luo, Xióngbiao Luó |
ICRA | 3 |
| 2025 | Anatomy-Aware Frequency-Attention Transformer Networks for Liver Couinaud CT/MR Segmentation
Wenkang Fan, Yanduan Lin, Chao An, Xióngbiao Luó |
MICCAI (1) | 1 |
| 2025 | Unsupervised Structure-Geometric Consistency for Monocular Endoscopic Depth Overestimation
Wenkang Fan, Enqi Qiu, Hongzhi Xu, Xióngbiao Luó |
MICCAI (9) | 1 |
| 2025 | Structure-Aware Cross-Modal Prompt Tuning for Autonomous Bronchoscopic Navigation
Zhuo Zeng, Wenkang Fan, Xióngbiao Luó |
MICCAI (11) | 4 |
| 2025 | U-bilateral attention gate nested U-transformers for medical image segmentation
Wenkang Fan, Haichao Peng, Xióngbiao Luó |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | DCTAN: Densely Convolved Transformer Aggregation Networks for Monocular Dense Depth Prediction in Robotic EndoscopyabstractAccurate dense depth prediction for 3-D reconstruction of monocular endoscopic images plays an essential role in expanding the surgical field in robotic surgery. However, it is generally a challenge to precisely estimate dense depth due to complex surgical fields with limited field of viewing, illumination variations, and variable texture structure. This work explores the performance of convolutional networks and transformer-based networks for endoscopic depth prediction, and proposes a new architecture called densely convolved transformer aggregation networks (DCTAN) that can aggregate local texture features and global spatial-temporal features for endoscopic dense depth recovery. Specifically, DCTAN creates a new hybrid encoder that combines dense convolution and scalable transformers to parallel extract local texture features and global spatial-temporal features from monocular endoscopic video sequences. Then, a local and global aggregation decoder is established to assemble the tokens of each frame to generate the global feature maps, that are integrated with the corresponding local feature maps to predict depth from coarse to fine. We trained and evaluated DCTAN through self-supervised learning on monocular synthesis (ground-truth) data and colonoscopic video images, with the experimental results demonstrating that our new architecture can extract more accurate local and global features for depth prediction and achieve more accurate depth range, more complete depth structure, and more sufficient texture information than other networks. In particular, all qualitative and quantitative assessment results of our method are better than current monocular dense depth estimation models. Wenkang Fan, Wenjing Jiang, Xióngbiao Luó |
ECAI | 1 |
| 2024 | Chat: Cascade Hole-Aware Transformers with Geometric Spatial Consistency for Accurate Monocular Endoscopic Depth EstimationabstractMonocular endoscopic depth estimation is essential for surgical navigation. Current deeply learned estimation methods still suffer from lack of real data labels and porous, artifacts (e.g., bubbles), illumination variations (e.g., specular highlight), and weak texture in endoscopic video images. This paper proposes a new deep learning framework of cascade hole-aware transformers with geometric spatial consistency for accurate endoscopic depth estimation without using any image annotation. Specifically, this framework employs cascade hole-aware encoders to powerfully extract structural features of deep and shallow holes, while it further introduces multiscale filtering decoders to suppress non-hole region features, addressing the problems of specular highlights, weak textures or bubbles. Additionally, a geometric spatial consistency loss can strongly perceive geometric information and suppress the color difference between virtual and real images. We generated virtual endoscopic image data to train our network architecture and test it on both virtual and real endoscopic video images, with the experimental results showing that our method is robust to zero-shot evaluation of real data. Particularly, our method can attain lower root mean square error 1.551±1.147 mm and mean absolute error 1.004±0.632 mm than state-of-the-art deep learning approaches. Ming Wu 0009, Wenkang Fan, Sunkui Ke, Hui-Qing Zeng, Yinran Chen, Xióngbiao Luó |
ICASSP | 3 |
| 2024 | Deformable Dual-Path Networks for Chronic Obstructive Pulmonary Disease Staging in CT ImagesabstractChronic obstructive pulmonary disease (COPD) is a respiratory disease that progresses over time and can significantly affect a person’s quality of life. Our proposed method for assessing the severity of COPD involves using lung computed tomography (CT) scans and a deep learning model called Deformable Dual-Path Networks. This model incorporates a side path that is densely connected to learn as many pathological features as possible. Our study provides a methodological idea for making full use of the results of vascular and airway tree segmentation in the diagnosis and monitoring of COPD patients, which can provide a reference for future researchers. We have created a network that uses chest CT scans to classify patients with COPD. Our method was tested on a dataset of 70 patients, and the results showed that it performs better than existing methods. This approach has the potential to enhance the diagnosis and monitoring of COPD patients. In summary, our method shows promise for improving the accuracy and efficiency of COPD severity assessment through medical imaging. Xióngbiao Luó, Wenkang Fan, Zhuo Zeng, Xiangxing Chen, Haichao Peng |
IJCNN | 3 |
| 2024 | Simultaneous Monocular Endoscopic Dense Depth and Odometry Estimation Using Local-Global Integration Networks
Wenkang Fan, Wenjing Jiang, Xióngbiao Luó |
MICCAI (6) | 1 |
| 2024 | Localization and Local Motion Magnification of Pulsatile Regions in Endoscopic Surgery Videos
Honglei Zheng, Wenkang Fan, Yinran Chen, Xióngbiao Luó |
MMM (3) | 2 |
| 2023 | Deep Triple-Supervision Learning Unannotated Surgical Endoscopic Video Data for Monocular Dense Depth EstimationabstractSurface reconstruction is an essential way to expand surgical field of view during endoscopic surgery, but it certainly requires dense depth estimation of endoscopic video sequences. Unfortunately, such a dense depth recovery suffers from illumination variation, weak texture, and occlusion. To address these problems, this work proposes a new triple-supervision self-learning strategy that uses unannotated endoscopic video data to predict monocular endoscopic dense depth information. This strategy first employs an effective conventional method to estimate camera poses and sparse depth maps to establishing a sparse data self-supervision. Furthermore, our strategy still combines two consistency measures to supervise dense depth and photometric information. We evaluated our method on collected colonoscopic videos, with the experimental results showing that our triple-supervision learning framework works more effective and accurate than some current self-supervised and unsupervised learning methods. Wenkang Fan, Kaiyun Zhang, Yinran Chen, Xióngbiao Luó |
ICASSP | 1 |
| 2023 | DGN: Descriptor Generation Network for Feature Matching in Monocular Endoscopy 3D ReconstructionabstractEndoscopy 3D reconstruction can provide more intuitive perception of the lesions in minimally invasive surgery. The success of 3D reconstruction highly relies on high-quality feature matches between the monocular image pairs, which remains challenging in the textureless endoscopic scenario. In this paper, we propose an effective feature matching framework for monocular endoscopy 3D reconstruction. The framework contains a descriptor generation network (DGN) to generate high-quality feature descriptors in a local-to-global manner, and a local region expansion to fine tune the initial matches obtained from the DGN module. We evaluated our method on the public Hamlyn Centre Laparoscopic/Endoscopic Video Datasets. The experimental results demonstrated that our method can generate sufficient accurate feature matches. Particularly, our method performed better in sparse depth estimation of the endoscopic scenario when compared with the current conventional and deep-learning methods. Kaiyun Zhang, Wenkang Fan, Yinran Chen, Xióngbiao Luó |
ICASSP | 2 |
| 2023 | DUP-Net: Double U-PoolFormer Networks for Renal Artery Segmentation in CT UrographyabstractRenal artery segmentation plays a fundamental role in nephrectomy, which can help surgeons get a better under-standing of vascular structures. However, the similar intensity between the renal arteries and cortex, the complex variations and tiny structures of arteries, bring challenges to accurate segmentation. To address these issues, we construct double U-PoolFormer networks (DUP-Net) to establish a coarse-to-fine framework for renal artery segmentation. Specifically, we use 2- D U - N et for the kidney extraction and then create 3-D DUP-Net for artery segmentation. DUP-Net is a serial network architecture that uses two U-PoolFormer modules to extract long-range spatial dependencies to create tree-like constraints while removing mis-segmentation of renal cortex through the serial structure. While DUP-Net improving the segmentation accuracy, it reduces memory cost during segmengtation. We evaluated our method on 70 cases of computed tomography urography data, with the experimental results showing that our proposed method certainly outperforms current 2-D and 3-D network models. Particularly, the average dice similarity coefficient of our method was improved from 81.51 % to 88.35%. Wenkang Fan, Mingxian Yang, Yinran Chen, Xióngbiao Luó |
IJCNN | 2 |
| 2023 | Cascade Transformer Encoded Boundary-Aware Multibranch Fusion Networks for Real-Time and Accurate Colonoscopic Lesion Segmentation
Ming Wu 0009, Wenkang Fan, Sunkui Ke, Yinran Chen, Xióngbiao Luó |
MICCAI (9) | 4 |
| 2023 | Self-supervised Cascade Training for Monocular Endoscopic Dense Depth Recovery
Wenjing Jiang, Wenkang Fan, Xióngbiao Luó |
PRCV (5) | 2 |
| 2022 | Contrastive Translation Learning For Medical Image SegmentationabstractUnsupervised domain adaptation commonly uses cycle generative networks to produce synthesis data from source to target domains. Unfortunately, translated samples cannot effectively preserve semantic information from input sources, resulting in bad or low adaptability of the network to segment target data. This work proposes an advantageous domain translation mechanism to improve the perceptual ability of the network for accurate unlabeled target data segmentation. Our domain translation employs patchwise contrastive learning to improve the semantic content consistency between input and translated images. Our approach was applied to unsupervised domain adaptation based abdominal organ segmentation. The experimental results demonstrate the effectiveness of our framework that outperforms other methods. Wankang Zeng, Wenkang Fan, Dongfang Shen, Yinran Chen, Xióngbiao Luó |
ICASSP | 2 |
| 2022 | Residual U-Structure Nested Conditional Adversarial Nets Colorized CT Improves Deep Learning Based Abdominal Multi-Organ SegmentationabstractSegmentation of abdominal organs such as the liver, pancreas, spleen, and kidneys plays an essential role in diagnosing and treating abdominal diseases. Although numerous deeply learned segmentation methods work well, they still suffer from partial volume effects, image noise, and data imbalance. This study aims to colorize CT images to boost or augment these segmentation approaches. We propose new residual U-structure nested generative adversarial nets that use residual U-blocks and spectral normalization for CT image colorization. Generated color CT images were introduced to train and validate V-Net and DenseV-Net for multiple abdominal organ segmentation. The experimental results demonstrate that colorized CT images can improve the dice similarity coefficient and reduce the Hausdorff distance from (0.32, 302.7) to (0.67, 78.2), significantly boosting the performance of V-Net and Dense V-Net for multiple abdominal organ segmentation. Vincent Chandra, Wenkang Fan, Yinran Chen, Xióngbiao Luó |
ICIP | 2 |