EDBT 2026 Demo / reviewers in the wild / expert
Qi Dou 0001
dblp:165/7846
· DBLP profile ↗
183ranked-venue papers
8as first author
138since 2021 · last 2026
0000-0002-3416-9950ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 101 · 5 first-author · 68 since 2021Graphics, computer vision, multimedia, augmented reality and games · 79 · 3 first-author · 59 since 2021Artificial intelligence and machine learning · 74 · 2 first-author · 64 since 2021Systems, architecture and hardware · 28 · 26 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Concepts from Representations: Post-hoc Concept Bottleneck Models via Sparse Decomposition of Visual RepresentationsabstractDeep learning has achieved remarkable success in image recognition, yet their inherent opacity poses challenges for deployment in critical domains. Concept-based interpretations aim to address this by explaining model reasoning through human-understandable concepts. However, existing post-hoc methods and ante-hoc concept bottleneck models (CBMs), suffer from limitations such as unreliable concept relevance, non-visual or labor-intensive concept definitions, and model/data-agnostic assumptions. This paper introduces Post-hoc Concept Bottleneck Model via Representation Decomposition (PCBM-ReD), a novel pipeline that retrofits interpretability onto pretrained opaque models. PCBM-ReD automatically extracts visual concepts from a pre-trained encoder, employs multimodal large language models (MLLMs) to label and filter concepts based on visual identifiability and task relevance, and selects an independent subset via reconstruction-guided optimization. Leveraging CLIP’s visual-text alignment, it decomposes image representations into linear combination of concept embeddings to fit into the CBMs abstraction. Extensive experiments across 11 image classification tasks show PCBM-ReD achieves state-of-the-art accuracy, narrows the performance gap with end-to-end models, and exhibits better interpretability. Shizhan Gong, Xiaofan Zhang 0002, Qi Dou 0001 |
AAAI | 3 |
| 2026 | Extreme cardiac MRI analysis under respiratory motion: Results of the CMRxMotion challenge
Kang Wang 0017, Chen Qin, Zhang Shi, Haoran Wang 0009, Chen Chen 0042, Cheng Ouyang, Chengliang Dai, Yuanhan Mo, Chenchen Dai, Xutong Kuang, Ruizhe Li 0005, Xin Chen 0003, Xiuzheng Yue, Song Tian, Alejandro Mora-Rubio, Kumaradevan Punithakumar, Shizhan Gong, Qi Dou 0001, Sina Amirrajab, Yasmina Alkhalil, Cian M. Scannell, Lexiaozi Fan, Huili Yang, Xiaowu Sun, Rob J. van der Geest, Tewodros Weldebirhan Arega, Fabrice Mériaudeau, Caner Ozer, Amin Ranem, John Kalkhof, Ilkay Öksüz, Anirban Mukhopadhyay 0003, Abdul Qayyum 0002, Moona Mazher, Steven A. Niederer, Carles García-Cabrera, Eric Arazo Sanchez, Michal K. Grzeszczyk, Szymon Plotka, Wanqin Ma, Xiaomeng Li 0001, Rongjun Ge, Yongqing Kou, Xinrong Chen, He Wang 0016, Chengyan Wang, Wenjia Bai, Shuo Wang 0011 |
Medical Image Anal. | 19 |
| 2026 | Toward Modality- and Sampling-Universal Learning Strategies for Accelerating Cardiovascular Imaging: Summary of the CMRxRecon2024 ChallengeabstractCardiovascular health is vital to human well-being, and cardiac magnetic resonance (CMR) imaging is considered the clinical reference standard for diagnosing cardiovascular disease. However, its adoption is hindered by long scan times, complex contrasts, and inconsistent quality. While deep learning methods perform well on specific CMR imaging sequences, they often fail to generalize across modalities and sampling schemes. The lack of benchmarks for high-quality, fast CMR image reconstruction further limits technology comparison and adoption. The CMRxRecon2024 challenge, attracting over 200 teams from 18 countries, addressed these issues with two tasks: generalization to unseen modalities and robustness to diverse undersampling patterns. We introduced the largest public multi-modality CMR raw dataset, an open benchmarking platform, and shared code. Analysis of the best-performing solutions revealed that prompt-based adaptation and enhanced physics-driven consistency enabled strong cross-scenario performance. These findings establish principles for generalizable reconstruction models and advance clinically translatable AI in cardiovascular imaging. Fanwen Wang, Zi Wang 0005, Yan Li 0064, Chen Qin, Shuo Wang 0011, Kunyuan Guo, Mengting Sun, Mingkai Huang, Michael Tänzer, Qirong Li, Yinzhe Wu 0001, Haosen Zhang, Kian Anvari Hamedani, Yuntong Lyu, Longyu Sun, Tianxing He, Lizhen Lan, Qiong Yao, Bingyu Xin, Dimitris N. Metaxas, Narges Razizadeh, Shahabedin Nabavi, George Yiasemis, Jonas Teuwen, Daniel B. Ennis, Zhihao Xue, Ruru Xu, Ilkay Öksüz, Donghang Lyu, Yanxin Huang, Xinrui Guo, Ruqian Hao, Jaykumar H. Patel, Guanke Cai, Binghua Chen, Sha Hua, Zhensen Chen, Qi Dou 0001, Xiahai Zhuang, Wenjia Bai, Harry Qin, He Wang 0016, Claudia Prieto, Michael Markl 0001, Alistair A. Young, Hao Li 0082, Xihong Hu, Lianming Wu, Xiaobo Qu 0001, Guang Yang 0006, Chengyan Wang |
IEEE Trans. Medical Imaging | 49 |
| 2026 | CustomVideo: Customizing Text-to-Video Generation With Multiple SubjectsabstractCustomized text-to-video generation aims to generate high-quality videos guided by text prompts and subject references. Current approaches for personalizing text-to-video generation suffer from tackling multiple subjects, which is a more challenging and practical scenario. In this work, our aim is to promote multi-subject guided text-to-video customization. We propose CustomVideo, a novel framework that can generate identity-preserving videos with the guidance of multiple subjects. To be specific, firstly, we encourage the co-occurrence of multiple subjects via composing them in a single image. Further, upon a basic text-to-video diffusion model, we design a simple yet effective attention control strategy to disentangle different subjects in the latent space of diffusion model. Moreover, to help the model focus on the specific area of the object, we segment the object from given reference images and provide a corresponding object mask for attention learning. Also, we collect a multi-subject text-to-video generation dataset as a comprehensive benchmark. Extensive qualitative, quantitative, and user study results demonstrate the superiority of our method compared to previous state-of-the-art approaches. Zhao Wang 0006, Aoxue Li, Lingting Zhu, Qi Dou 0001, Zhenguo Li |
IEEE Trans. Multim. | 5 |
| 2026 | Real-Time Monocular 2-D and 3-D Perception of Endoluminal Scenes for Controlling Flexible Robotic Endoscopic InstrumentsabstractEndoluminal surgery offers a minimally invasive option for early-stage gastrointestinal and urinary tract, but is limited by basic surgical tools and a steep learning curve. Robotic systems, particularly continuum robots, provide flexible instruments that enable precise, intuitive tissue resection in confined spaces, potentially improving outcomes. This paper presents an integrated visual perception platform for a continuum robotic system in endoluminal surgery. Our objective is to leverage monocular endoscopic image-based perception algorithms to accurately identify the position and orientation of flexible instruments and measure their distances from surrounding tissues. This thorough understanding of continuum robots and surgical scenes enhances the robustness of robotic procedures. We introduce 2D and 3D learning-based perception algorithms and develop a physically-realistic simulator that models the dynamics of flexible instruments. This simulator features a pipeline for generating realistic endoluminal scenes, enabling control of flexible robots in a realistic environment and substantial data collection. Using a continuum robot prototype, we conducted extensive evaluations, including module assessments and system-level evaluation of the perception platform. Results demonstrate that our perception algorithms significantly improve control of flexible instruments, reducing manipulation time by over 70% for trajectory-following tasks and enhancing the understanding of complex surgical scenarios, leading to robust endoluminal surgeries. Ruofeng Wei, Kai Chen 0024, Yui-Lun Ng, Yiyao Ma, Justin D. L. Ho, Hon-Sing Tong, Ka-Wai Kwok, Qi Dou 0001 |
IEEE Trans. Robotics | 10 |
| 2025 | DDxTutor: Clinical Reasoning Tutoring System with Differential Diagnosis-Based Structured ReasoningabstractClinical diagnosis education requires students to master both systematic reasoning processes and comprehensive medical knowledge.While recent advances in Large Language Models (LLMs) have enabled various medical educational applications, these systems often provide direct answers that could reduce students' cognitive engagement and lead to fragmented learning.Motivated by these challenges, we propose DDxTutor, a framework that follows differential diagnosis principles to decompose clinical reasoning into teachable components.It consists of a structured reasoning module that analyzes clinical clues and synthesizes diagnostic conclusions, and an interactive dialogue framework that guides students through this process.To enable such tutoring, we construct DDxReasoning, a dataset of 933 clinical cases with fine-grained diagnostic steps verified by doctors.Our experiments demonstrate that fine-tuned LLMs achieve strong performance in generating structured teaching references and conducting interactive diagnostic tutoring dialogues.Human evaluation by medical educators and students validates the framework's potential and effectiveness for clinical diagnosis education.Our project is available at https://github.com/med-air/DDxTutor. Zheyao Gao, Longfei Gou, Qi Dou 0001 |
ACL (1) | 4 |
| 2025 | Hybrid Reciprocal Transformer with Triplet Feature Alignment for Scene Graph GenerationabstractScene graph generation is a pivotal task in computer vision, focusing on comprehensive identification of visual relation tuples embedded within images. The advancement of methods involving triplets has sought to enhance task performance by integrating triplets as contextual features for more precise predicate identification from component level. However, challenges remain due to interference from multi-role objects in overlapping tuples within complex environments, which impairs the model’s ability to distinguish and align specific triplet features for reasoning diverse semantics of multi-role objects. To address these issues, we introduce a novel framework that incorporates a triplet alignment model into a hybrid reciprocal transformer architecture, starting from using triplet mask features to guide the learning of component-level relation graphs. To effectively distinguish multi-role objects characterized by overlapping visual relation tuples, we introduce a triplet alignment loss, which provides multi-role objects with aligned features from triplet and helps customize them. Additionally, we explore the inherent connectivity between hybrid aligned triplet and component features through a bidirectional refinement module, which enhances feature interaction and reciprocal reinforcement. Experimental results demonstrate that our model achieves state-of-the-art performance on the Visual Genome and Action Genome datasets, underscoring its effectiveness and adaptability. Project page: hq-sg.github.io. Jiawei Fu 0001, Kai Chen 0028, Qi Dou 0001 |
CVPR | 4 |
| 2025 | LM-Searcher: Cross-domain Neural Architecture Search with LLMs via Unified Numerical EncodingabstractYuxuan Hu, Jihao Liu, Ke Wang, Jinliang Zheng, Weikang Shi, Manyuan Zhang, Qi Dou, Rui Liu, Aojun Zhou, Hongsheng Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jihao Liu, Ke Wang 0036, Jinliang Zheng, Weikang Shi, Manyuan Zhang, Qi Dou 0001, Rui Liu 0019, Aojun Zhou, Hongsheng Li 0001 |
EMNLP | 7 |
| 2025 | HealthCards: Exploring Text-to-Image Generation as Visual Aids for Healthcare Knowledge Democratizing and EducationabstractThe evolution of text-to-image (T2I) generation techniques has introduced new capabilities for information visualization, with the potential to advance knowledge democratization and education. In this paper, we investigate how T2I models can be adapted to generate educational health knowledge contents, exploring their potential to make healthcare information more visually accessible and engaging. We explore methods to harness recent T2I models for generating health knowledge flashcards—visual educational aids that present healthcare information through appealing and concise imagery. To support this goal, we curated a diverse, high-quality healthcare knowledge flashcard dataset containing 2,034 samples sourced from credible medical resources. We further validate the effectiveness of fine-tuning open-source models with our dataset, demonstrating their promise as specialized health flashcard generators. Our code and dataset are available at: https://github.com/med-air/HealthCards. Zheyao Gao, Longfei Gou, Ann Sin Nga Lau, Qi Dou 0001 |
EMNLP | 6 |
| 2025 | Test-Time Retrieval-Augmented Adaptation for Vision-Language Models
Xinqi Fan, Luoxiao Yang, Chuin Hong Yap, Rizwan Qureshi, Qi Dou 0001, Moi Hoon Yap, Mubarak Shah |
ICCV | 6 |
| 2025 | Boosting the visual interpretability of CLIP via adversarial fine-tuningabstractCLIP has achieved great success in visual representation learning and is becoming an important plug-in component for many large multi-modal models like LLaVA and DALL-E. However, the lack of interpretability caused by the intricate image encoder architecture and training process restricts its wider use in high-stake decision making applications. In this work, we propose an unsupervised adversarial fine-tuning (AFT) with norm-regularization to enhance the visual interpretability of CLIP. We provide theoretical analysis showing that AFT has implicit regularization that enforces the image encoder to encode the input features sparsely, directing the network's focus towards meaningful features. Evaluations by both feature attribution techniques and network dissection offer convincing evidence that the visual interpretability of CLIP has significant improvements. With AFT, the image encoder prioritizes pertinent input features, and the neuron within the encoder exhibits better alignment with human-understandable concepts. Moreover, these effects are generalizable to out-of-distribution datasets and can be transferred to downstream tasks. Additionally, AFT enhances the visual interpretability of derived large vision-language models that incorporate the pre-trained CLIP an integral component. The code of this paper is available at [the CLIP_AFT GitHub repository](https://github.com/peterant330/CLIP_AFT). Shizhan Gong, Haoyu Lei, Qi Dou 0001, Farzan Farnia |
ICLR | 3 |
| 2025 | Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language ModelsabstractVision-language models, such as CLIP, have achieved significant success in aligning visual and textual representations, becoming essential components of many multi-modal large language models (MLLMs) like LLaVA and OpenFlamingo. However, numerous studies have identified CLIP’s limited fine-grained perception as a critical drawback, leading to substantial failures in downstream MLLMs. In contrast, vision-centric foundation models like DINOv2 demonstrate remarkable capabilities in capturing fine details from images. In this work, we propose a novel kernel-based method to align CLIP’s visual representation with that of DINOv2, ensuring that the resulting embeddings maintain compatibility with text embeddings while enhancing perceptual capabilities. Our alignment objective is designed for efficient stochastic optimization. Following this image-only alignment fine-tuning, the visual encoder retains compatibility with the frozen text encoder and exhibits significant improvements in zero-shot object recognition, fine-grained spatial reasoning, and localization. By integrating the aligned visual encoder, downstream MLLMs also demonstrate enhanced performance. The code and models are available at https://github.com/peterant330/KUEA. Shizhan Gong, Yankai Jiang 0003, Qi Dou 0001, Farzan Farnia |
ICML | 3 |
| 2025 | Gaussian Splatting with Reflectance Regularization for Endoscopic Scene ReconstructionabstractEndoscopic reconstruction plays a crucial role in surgical robotics. The dynamic lighting conditions and integrated camera-light source in endoscopic scenes create a distinct reconstruction challenge: shape ambiguity. To mitigate this, we propose a Gaussian Splatting (GS) based framework for endoscopic scene reconstruction, enhanced with reflectance regularization. We embed every 3D Gaussian point with physical reflective attributes and combine this representation with a physically based inverse rendering framework. By jointly training 3DGS for view synthesis with this reflectance regularization, we are able to attain high-quality geometry without changing the volume rendering pipeline. Our experiments demonstrate the superiority in both geometry representation and rendering performance compared to existing GS approaches, making it a practical solution for endoscopic applications. Project is available at: https://med-air.github.io/GSR2. Chengkun Li, Kai Chen 0028, Shi Qiu 0001, Jason Ying-Kuen Chan, Qi Dou 0001 |
IROS | 5 |
| 2025 | ColaDex: Contact-guided Optimization and VLM-assisted Selection for Task-oriented Dexterous Grasp GenerationabstractTask-oriented dexterous grasp generation aims to generate stable and functional grasps that enable a robotic hand to effectively interact with objects to accomplish specific tasks. However, generating high-dimensional hand configurations that seamlessly adapt to diverse task requirements and object geometries remains a significant challenge. In this paper, we propose a novel pipeline called ColaDex to address this challenging problem. The core idea of ColaDex is to leverage a vision-language models (VLMs) to select the dexterous grasp from a set of candidates that aligns well with the task description. To this end, we first introduce a contact-guided optimization method to generate a set of high-quality grasp candidates around the object through analytical optimization. Subsequently, to effectively prompt VLMs with the sampled numerous grasp candidates, we propose an object-centric approach that adaptively represents a group of candidates as prototypical contact maps, learned based on the geometric relationships between the grasping hand and object shape. We then feed the task requirement and the generated prototypical contact maps into the VLM, enabling it to reason about grasp-object interactions and assess their alignment with the given task, ultimately selecting the grasp that best aligns with the task requirement. Extensive experiments demonstrate that our prototypical contact map is a more informative prompting mechanism than conventional RGB images, enabling ColaDex to consistently generate high-quality task-oriented grasps and achieve a high success rate across diverse objects and tasks. Yiyao Ma, Kai Chen 0028, Xuecheng Xu, Zhongxiang Zhou, Rong Xiong, Qi Dou 0001 |
IROS | 7 |
| 2025 | Endo3R: Unified Online Reconstruction from Dynamic Monocular Endoscopic Video
Wenzhen Dong, Hao Ding 0021, Ziyi Wang 0006, Haomin Kuang, Qi Dou 0001, Yun-Hui Liu 0001 |
MICCAI (9) | 7 |
| 2025 | ClipGS: Clippable Gaussian Splatting for Interactive Cinematic Visualization of Volumetric Medical Data
Chengkun Li, Yuqi Tong, Kai Chen 0028, Zhenya Yang, Shi Qiu 0001, Jason Ying-Kuen Chan, Pheng-Ann Heng, Qi Dou 0001 |
MICCAI (10) | 9 |
| 2025 | Surgical Action Planning with Large Language Models
Mengya Xu, Zhongzhen Huang, Xiaofan Zhang 0002, Qi Dou 0001 |
MICCAI (9) | 5 |
| 2025 | Medical Large Vision Language Models with Multi-image Visual Ability
Xikai Yang, Juzheng Miao, Yuchen Yuan, Qi Dou 0001, Jinpeng Li 0004, Pheng-Ann Heng |
MICCAI (5) | 5 |
| 2025 | CSAP-Assist: Instrument-Agent Dialogue Empowered Vision-Language Models for Collaborative Surgical Action Planning
Mengya Xu, Qi Dou 0001 |
MICCAI (9) | 4 |
| 2025 | Contact Map Transfer with Conditional Diffusion Model for Generalizable Dexterous Grasp GenerationabstractDexterous grasp generation is a fundamental challenge in robotics, requiring both grasp stability and adaptability across diverse objects and tasks. Analytical methods ensure stable grasps but are inefficient and lack task adaptability, while generative approaches improve efficiency and task integration but generalize poorly to unseen objects and tasks due to data limitations. In this paper, we propose a transfer-based framework for dexterous grasp generation, leveraging a conditional diffusion model to transfer high-quality grasps from shape templates to novel objects within the same category. Specifically, we reformulate the grasp transfer problem as the generation of an object contact map, incorporating object shape similarity and task specifications into the diffusion process. To handle complex shape variations, we introduce a dual mapping mechanism, capturing intricate geometric relationship between shape templates and novel objects. Beyond the contact map, we derive two additional object-centric maps, the part map and direction map, to encode finer contact details for more stable grasps. We then develop a cascaded conditional diffusion model framework to jointly transfer these three maps, ensuring their intra-consistency. Finally, we introduce a robust grasp recovery mechanism, identifying reliable contact points and optimizing grasp configurations efficiently. Extensive experiments demonstrate the superiority of our proposed method. Our approach effectively balances grasp quality, generation efficiency, and generalization performance across various tasks. Project homepage: https://cmtdiffusion.github.io/ Yiyao Ma, Kai Chen 0028, Kexin Zheng, Qi Dou 0001 |
NeurIPS | 4 |
| 2025 | Learning dissection trajectories from expert surgical videos via imitation learning with equivariant diffusion
Yonghao Long 0001, Yueyao Chen, Hon-Chi Yip, Markus Scheppach, Philip W. Y. Chiu, Yeung Yam, Helen M. Meng, Qi Dou 0001 |
Medical Image Anal. | 9 |
| 2025 | Illuminating the unseen: Advancing MRI domain generalization through causality
Tianjiao Zeng, Furui Liu, Qi Dou 0001, Hing-Chiu Chang, Edward S. Hui |
Medical Image Anal. | 4 |
| 2025 | Absolute Monocular Depth Estimation on Robotic Visual and Kinematics Data via Self-Supervised LearningabstractAccurate estimation of absolute depth from a monocular endoscope is a fundamental task for automatic navigation systems in robotic surgery. Previous works solely rely on uni-modal data (i.e., monocular images), which can only estimate depth values arbitrarily scaled with the real world. In this paper, we present a novel framework, SADER, which explores vision and robot kinematics to estimate the high-quality absolute depth for monocular surgical scenes. To jointly learn the multi-modal data, we introduce a self-distillation based two-stage training policy in the framework. In the first stage, a boosting depth module based on vision transformer is proposed to improve the relative depth estimation network that is trained in a self-supervised method. Then, we develop an algorithm to automatically compute the scale from robot kinematics. By coupling the scale and relative depth data, pseudo absolute depth labels for all images are yielded. In the second stage, we re-train the network with 3D loss supervised by pseudo labels. To make our method generalize to different endoscopes, the learning of endoscopic intrinsics is integrated into the network. In addition, we did cadaver experiments to collect new surgical depth estimation data about robotic laparoscopy for evaluation. Experimental results on public SCARED and cadaver data demonstrate that the SADER outperforms previous state-of-art even stereo-based methods with an accuracy error under 1.90 mm, proving the feasibility of our approach to recover the absolute depth with monocular inputs. Note to Practitioners—This paper aims to solve the problem of absolute monocular depth estimation in automatic surgical navigation by leveraging the multi-modal data from the robot-based endoscopic system. Accurate depth perception with real scales of the monocular scene is essential for the control of surgical robots in automatic navigation. However, current methods can only predict the relative depth of the surgical scene using monocular images. In this article, we propose a self-supervised learning-based method to achieve high-quality absolute depth estimation of monocular endoscopic images. It neither needs manual data annotation, nor other imaging modalities. The experiments extensively validate the feasibility and high performance of our framework for absolute depth estimation on monocular endoscopes. This absolute depth perception framework can be potentially encapsulated into the automatic navigation system in the near future. Ruofeng Wei, Bin Li 0082, Fangxun Zhong, Hangjie Mo, Qi Dou 0001, Yun-Hui Liu 0001, Dong Sun 0001 |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2025 | Canonical Shape Reconstruction With SE(3) Equivariance Learning for Weakly-Supervised Object Pose Estimationabstract6D object pose estimation from a single RGB-D image is a fundamental problem in computer vision and robot manipulation. Despite recent advancements, existing methods still suffer several limitations. First of all, the object shape representation extracted from the depth map is often less expressive because the object point cloud parsed from the depth map is highly incomplete due to the object self-occlusion and noisy due to the sensor artifacts. This shape representation issue further intensifies when lacking sufficient labeled data for model training, which unfortunately is another typical problem for object pose estimation considering the heavy annotation cost for real-world pose labeling. In this study, we propose to tackle the above issues in a unified way. First, we enhance the object shape representation from the partial point cloud with a novel canonical shape reconstruction module, in which an implicit canonical frame is established by incorporating the SE(3) equivariance, achieving implicit feature alignment of the partial point cloud inputs, leading to robust shape recovery. Second, based on the enhanced object representation, we further utilize the de-canonicalized and pose-dependent completed object shape as the training signal, and develop a novel weakly-supervised learning framework to leverage both labeled synthetic data and unlabeled real data to train the pose estimation model in a label-efficient way. Extensive experiments on three widely used benchmarks demonstrate the effectiveness, and superiority of our framework over state-of-the-art methods. Jun Zhou 0029, Kai Chen 0024, Mingqiang Wei, Xiao-Ping Zhang 0002, Qi Dou 0001, Harry Qin |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Multi-Organ Segmentation From Partially Labeled and Unaligned Multi-Modal MRI in Thyroid-Associated OrbitopathyabstractThyroid-associated orbitopathy (TAO) is a prevalent inflammatory autoimmune disorder, leading to orbital disfigurement and visual disability. Automatic comprehensive segmentation tailored for quantitative multi-modal MRI assessment of TAO holds enormous promise but is still lacking. In this paper, we propose a novel method, named cross-modal attentive self-training (CMAST), for the multi-organ segmentation in TAO using partially labeled and unaligned multi-modal MRI data. Our method first introduces a dedicatedly designed cross-modal pseudo label self-training scheme, which leverages self-training to refine the initial pseudo labels generated by cross-modal registration, so as to complete the label sets for comprehensive segmentation. With the obtained pseudo labels, we further devise a learnable attentive fusion module to aggregate multi-modal knowledge based on learned cross-modal feature attention, which relaxes the requirement of pixel-wise alignment across modalities. A prototypical contrastive learning loss is further incorporated to facilitate cross-modal feature alignment. We evaluate our method on a large clinical TAO cohort with 100 cases of multi-modal orbital MRI. The experimental results demonstrate the promising performance of our method in achieving comprehensive segmentation of TAO-affected organs on both T1 and T1c modalities, outperforming previous methods by a large margin. Our code is available at: https://github.com/cchen-cc/CMAST. Cheng Chen 0013, Yuan Zhong 0003, Jinyue Cai, Karen Kar Wun Chan, Qi Dou 0001, Kelvin Kam Lung Chong, Pheng-Ann Heng, Winnie Chiu-Wing Chu |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Improving Foundation Model for Endoscopy Video Analysis via Representation Learning on Long SequencesabstractRecent advancements in endoscopy video analysis have relied on the utilization of relatively short video clips extracted from longer videos or millions of individual frames. However, these approaches tend to neglect the domain-specific characteristics of endoscopy data, which is typically presented as a long stream containing valuable semantic spatial and temporal information. To address this limitation, we propose EndoFM-LV, a foundation model developed under a minute-level pre-training framework upon long endoscopy video sequences. To be specific, we propose a novel masked token modeling scheme within a teacher-student framework for self-supervised video pre-training, which is tailored for learning representations from long video sequences. For pre-training, we construct a large-scale long endoscopy video dataset comprising 6,469 long endoscopic video samples, each longer than 1 minute and totaling over 13 million frames. Our EndoFM-LV is evaluated on four types of endoscopy tasks, namely classification, segmentation, detection, and workflow recognition, serving as the backbone or temporal module. Extensive experimental results demonstrate that our framework outperforms previous state-of-the-art video-based and frame-based approaches by a significant margin, surpassing Endo-FM (5.6% F1, 9.3% Dice, 8.4% F1, and 3.3% accuracy for classification, segmentation, detection, and workflow recognition) and EndoSSL (5.0% F1, 8.1% Dice, 9.3% F1 and 3.1% accuracy for classification, segmentation, detection, and workflow recognition). Zhao Wang 0006, Lingting Zhu, Shaoting Zhang 0001, Qi Dou 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | IPNet: An Interpretable Network With Progressive Loss for Whole-Stage Colorectal Disease DiagnosisabstractColorectal cancer plays a dominant role in cancer-related deaths, primarily due to the absence of obvious early-stage symptoms. Whole-stage colorectal disease diagnosis is crucial for assessing lesion evolution and determining treatment plans. However, locality difference and disease progression lead to intra-class disparities and inter-class similarities for colorectal lesion representation. In addition, interpretable algorithms explaining the lesion progression are still lacking, making the prediction process a "black box". In this paper, we propose IPNet, a dual-branch interpretable network with progressive loss for whole-stage colorectal disease diagnosis. The dual-branch architecture captures unbiased features representing diverse localities to suppress intra-class variation. The progressive loss function considers inter-class relationship, using prior knowledge of disease evolution to guide classification. Furthermore, a novel Grain-CAM is designed to interpret IPNet by visualizing pixel-wise attention maps from shallow to deep layers, providing regions semantically related to IPNet's progressive classification. We conducted whole-stage diagnosis on two image modalities, i.e., colorectal lesion classification on 129,893 endoscopic optical images and rectal tumor T-staging on 11,072 endoscopic ultrasound images. IPNet is shown to surpass other state-of-the-art algorithms, accordingly achieving an accuracy of 93.15% and 89.62%. Especially, it establishes effective decision boundaries for challenges like polyp vs. adenoma and T2 vs. T3. The results demonstrate an explainable attempt for colorectal lesion classification at a whole-stage level, and rectal tumor T-staging by endoscopic ultrasound is also unprecedentedly explored. IPNet is expected to be further applied, assisting physicians in whole-stage disease diagnosis and enhancing diagnostic interpretability. Junhu Fu, Qi Dou 0001, Yiping He, Pinghong Zhou, Shengli Lin, Yuanyuan Wang 0001, Yi Guo 0002 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | UC-NeRF: Uncertainty-Aware Conditional Neural Radiance Fields From Endoscopic Sparse ViewsabstractVisualizing surgical scenes is crucial for revealing internal anatomical structures during minimally invasive procedures. Novel View Synthesis is a vital technique that offers geometry and appearance reconstruction, enhancing understanding, planning, and decision-making in surgical scenes. Despite the impressive achievements of Neural Radiance Field (NeRF), its direct application to surgical scenes produces unsatisfying results due to two challenges: endoscopic sparse views and significant photometric inconsistencies. In this paper, we propose uncertainty-aware conditional NeRF for novel view synthesis to tackle the severe shape-radiance ambiguity from sparse surgical views. The core of UC-NeRF is to incorporate the multi-view uncertainty estimation to condition the neural radiance field for modeling the severe photometric inconsistencies adaptively. Specifically, our UC-NeRF first builds a consistency learner in the form of multi-view stereo network, to establish the geometric correspondence from sparse views and generate uncertainty estimation and feature priors. In neural rendering, we design a base-adaptive NeRF network to exploit the uncertainty estimation for explicitly handling the photometric inconsistencies. Furthermore, an uncertainty-guided geometry distillation is employed to enhance geometry learning. Experiments on the SCARED and Hamlyn datasets demonstrate our superior performance in rendering appearance and geometry, consistently outperforming the current state-of-the-art approaches. Our code will be released at https://github.com/wrld/UC-NeRF. Jiangliu Wang, Ruofeng Wei, Qi Dou 0001, Yun-Hui Liu 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Toward Reliable AR-Guided Surgical Navigation: Interactive Deformation Modeling With Data-Driven Biomechanics and PromptsabstractIn augmented reality (AR)-guided surgical navigation, preoperative organ models are superimposed onto the patient's intraoperative anatomy to visualize critical structures such as vessels and tumors. Accurate deformation modeling is essential to maintain the reliability of AR overlays by ensuring alignment between preoperative models and the dynamically changing anatomy. Although the finite element method (FEM) offers physically plausible modeling, its high computational cost limits intraoperative applicability. Moreover, existing algorithms often fail to handle large anatomical changes, such as those induced by pneumoperitoneum or ligament dissection, leading to inaccurate anatomical correspondences and compromised AR guidance. To address these challenges, we propose a data-driven biomechanics algorithm that preserves FEM-level accuracy while improving computational efficiency. In addition, we introduce a novel human-in-the-loop mechanism into the deformation modeling process. This enables surgeons to interactively provide prompts to correct anatomical misalignments, thereby incorporating clinical expertise and allowing the model to adapt dynamically to complex surgical scenarios. Experiments on a publicly available dataset demonstrate that our algorithm achieves a mean target registration error of 3.42 mm. Incorporating surgeon prompts through the interactive framework further reduces the error to 2.78 mm, surpassing state-of-the-art methods in volumetric accuracy. These results highlight the ability of our framework to deliver efficient and accurate deformation modeling while enhancing surgeon-algorithm collaboration, paving the way for safer and more reliable computer-assisted surgeries. Jun Zhou 0029, Jialun Pei, Harry Qin, Yingfang Fan, Qi Dou 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | On-the-Fly Improving Segment Anything for Medical Image Segmentation Using Auxiliary Online LearningabstractThe current variants of the Segment Anything Model (SAM), which include the original SAM and Medical SAM, still lack the capability to produce sufficiently accurate segmentation for medical images. In medical imaging contexts, it is not uncommon for human experts to rectify segmentations of specific test samples after SAM generates its segmentation predictions. These rectifications typically entail manual or semi-manual corrections employing state-of-the-art annotation tools. Motivated by this process, we introduce a novel approach that leverages the advantages of online machine learning to enhance Segment Anything (SA) during test time. We employ rectified annotations to perform online learning, with the aim of improving the segmentation quality of SA on medical images. To ensure the effectiveness and efficiency of online learning when integrated with large-scale vision models like SAM, we propose a new method called Auxiliary Online Learning (AuxOL), which entails adaptive online-batch and adaptive segmentation fusion. Experiments conducted on eight datasets covering four medical imaging modalities validate the effectiveness of the proposed method. Our work proposes and validates a new, practical, and effective approach for enhancing SA on downstream segmentation tasks (e.g., medical image segmentation). The code is publicly available at https://sam-auxol.github.io/AuxOL/. Tao Zhou 0002, Weidi Xie, Shuo Wang 0011, Qi Dou 0001, Yizhe Zhang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | A Dexterous and Compliant (DexCo) Hand Based on Soft Hydraulic Actuation for Human-Inspired Fine In-Hand ManipulationabstractHuman beings possess a remarkable skill for fine in-hand manipulation, utilizing both intrafinger interactions (in-finger) and finger–environment interactions across a wide range of daily tasks. These tasks range from skilled activities like screwing light bulbs, picking and sorting pills, and in-hand rotation, to more complex tasks such as opening plastic bags, cluttered bin picking, and counting cards. Despite its prevalence in human activities, replicating these fine motor skills in robotics remains a substantial challenge. This study tackles the challenge of fine in-hand manipulation by introducing the dexterous and compliant (DexCo) hand system. The DexCo hand mimics human dexterity, replicating the intricate interaction between the thumb, index, and middle fingers, with a contractable palm. The key to maneuverable fine in-hand manipulation lies in its innovative soft hydraulic actuation, which strikes a balance between control complexity, dexterity, compliance, and motion accuracy within a compact structure, enhancing the overall performance of the system. The model of soft hydraulic actuation, based on hydrostatic force analysis, reveals the compliance of hand joints, which is also further extended to a dedicated robot operating system (ROS) package for DexCo hand simulation, considering both motion and stiffness aspects. Dedicated velocity and position teleoperation controllers are designed for implementing real physical manipulation tasks. The benchmark results show that the fingertip achieves a maximum repeatable finger strength of 34.4 N, a grasp cycle time of less than 2.04 s, and a maximum repeatability accuracy of 0.03 mm. Experimental results demonstrate the DexCo hand successfully performs complex fine in-hand manipulation tasks, providing a promising solution for advancing robotic manipulation capabilities toward the human level. Jianshu Zhou, Junda Huang, Qi Dou 0001, Pieter Abbeel, Yun-Hui Liu 0001 |
IEEE Trans. Robotics | 3 |
| 2024 | ANEDL: Adaptive Negative Evidential Deep Learning for Open-Set Semi-supervised LearningabstractSemi-supervised learning (SSL) methods assume that labeled data, unlabeled data and test data are from the same distribution. Open-set semi-supervised learning (Open-set SSL) con- siders a more practical scenario, where unlabeled data and test data contain new categories (outliers) not observed in labeled data (inliers). Most previous works focused on out- lier detection via binary classifiers, which suffer from insufficient scalability and inability to distinguish different types of uncertainty. In this paper, we propose a novel framework, Adaptive Negative Evidential Deep Learning (ANEDL) to tackle these limitations. Concretely, we first introduce evidential deep learning (EDL) as an outlier detector to quantify different types of uncertainty, and design different uncertainty metrics for self-training and inference. Furthermore, we propose a novel adaptive negative optimization strategy, making EDL more tailored to the unlabeled dataset containing both inliers and outliers. As demonstrated empirically, our proposed method outperforms existing state-of-the-art methods across four datasets. Yang Yu 0070, Danruo Deng, Furui Liu, Qi Dou 0001, Yueming Jin, Guangyong Chen, Pheng-Ann Heng |
AAAI | 4 |
| 2024 | A Super-pixel-based Approach to the Stable Interpretation of Neural Networks
Shizhan Gong, Qi Dou 0001, Farzan Farnia |
BMVC | 3 |
| 2024 | Structured Gradient-Based Interpretations via Norm-Regularized Adversarial TrainingabstractGradient-based saliency maps have been widely used to explain the decisions of deep neural network classifiers. However, standard gradient-based interpretation maps, including the simple gradient and integrated gradient algorithms, often lack desired structures such as sparsity and connectedness in their application to real-world computer vision models. A frequently used approach to inducing sparsity structures into gradient-based saliency maps is to alter the simple gradient scheme using sparsification or norm-based regularization. A drawback with such post-processing methods is their frequently-observed significant loss in fidelity to the original simple gradient map. In this work, we propose to apply adversarial training as an inprocessing scheme to train neural networks with structured simple gradient maps. We show a duality relation between the regularized norms of the adversarial perturbations and gradient-based maps, based on which we design adversarial training loss functions promoting sparsity and group-sparsity properties in simple gradient maps. We present several numerical results to show the influence of our proposed norm-based adversarial training methods on the standard gradient-based maps of standard neural network architectures on benchmark image datasets11The paper's code is available at: https://github.com/peterant330/AdvGrad. Shizhan Gong, Qi Dou 0001, Farzan Farnia |
CVPR | 2 |
| 2024 | Ensemble Diversity Facilitates Adversarial TransferabilityabstractWith the advent of ensemble-based attacks, the transfer-ability of generated adversarial examples is elevated by a noticeable margin despite many methods only employing superficial integration yet ignoring the diversity between ensemble models. However, most of them compromise the latent value of the diversity between generated perturbation from distinct models which we argue is also able to increase the adversarial transferability, especially heterogeneous at-tacks. To address the issues, we propose a novel method of Stochastic Mini-batch black-box attack with Ensemble Reweighing using reinforcement learning (SMER) to produce highly transferable adversarial examples. We emphasize the diversity between surrogate models achieving indi-vidual perturbation iteratively. In order to customize the individual effect between surrogates, ensemble reweighing is introduced to refine ensemble weights by maximizing attack loss based on reinforcement learning which functions on the ultimate transferability elevation. Extensive exper-iments demonstrate our superiority to recent ensemble at-tacks with a significant margin across different black-box attack scenarios, especially on heterogeneous conditions. https://github.com/tangbwb/SMER Zheng Wang 0044, Yi Bin, Qi Dou 0001, Yang Yang 0002, Heng Tao Shen |
CVPR | 4 |
| 2024 | Shape-Guided Configuration-Aware Learning for Endoscopic-Image-Based Pose Estimation of Flexible Robotic Instruments
Yiyao Ma, Kai Chen 0028, Hon-Sing Tong, Ruofeng Wei, Yui-Lun Ng, Ka-Wai Kwok, Qi Dou 0001 |
ECCV (22) | 7 |
| 2024 | Towards Real-World Adverse Weather Image Restoration: Enhancing Clearness and Semantics with Vision-Language Models
Mengyang Wu, Xiaohu You 0001, Chi-Wing Fu, Qi Dou 0001, Pheng-Ann Heng |
ECCV (18) | 5 |
| 2024 | Heterogeneous Personalized Federated Learning by Local-Global Updates Mixing via Convergence RateabstractPersonalized federated learning (PFL) has emerged as a promising technique for addressing the challenge of data heterogeneity. While recent studies have made notable progress in mitigating heterogeneity associated with label distributions, the issue of effectively handling feature heterogeneity remains an open question. In this paper, we propose a personalization approach by Local-global updates Mixing (LG-Mix) via Neural Tangent Kernel (NTK)-based convergence. The core idea is to leverage the convergence rate induced by NTK to quantify the importance of local and global updates, and subsequently mix these updates based on their importance. Specifically, we find the trace of the NTK matrix can manifest the convergence rate, and propose an efficient and effective approximation to calculate the trace of a feature matrix instead of the NTK matrix. Such approximation significantly reduces the cost of computing NTK, and the feature matrix explicitly considers the heterogeneous features among samples. We have theoretically analyzed the convergence of our method in the over-parameterize regime, and experimentally evaluated our method on five datasets. These datasets present heterogeneous data features in natural and medical images. With comprehensive comparison to existing state-of-the-art approaches, our LG-Mix has consistently outperformed them across all datasets (largest accuracy improvement of 5.01\%), demonstrating the outstanding efficacy of our method for model personalization. Code is available at \url{https://github.com/med-air/HeteroPFL}. Meirui Jiang, Anjie Le, Qi Dou 0001 |
ICLR | 4 |
| 2024 | RGBManip: Monocular Image-based Robotic Manipulation through Active Object Pose EstimationabstractRobotic manipulation requires accurate perception of the environment, which poses a significant challenge due to its inherent complexity and constantly changing nature. In this context, RGB image and point-cloud observations are two commonly used modalities in visual-based robotic manipulation, but each of these modalities have their own limitations. Commercial point-cloud observations often suffer from issues like sparse sampling and noisy output due to the limits of the emission-reception imaging principle. On the other hand, RGB images, while rich in texture information, lack essential depth and 3D information crucial for robotic manipulation. To mitigate these challenges, we propose an image-only robotic manipulation framework that leverages an eye-on-hand monocular camera installed on the robot’s parallel gripper. By moving with the robot gripper, this camera gains the ability to actively perceive the object from multiple perspectives during the manipulation process. This enables the estimation of 6D object poses, which can be utilized for manipulation. While, obtaining images from more and diverse viewpoints typically improves pose estimation, it also increases the manipulation time. To address this trade-off, we employ a reinforcement learning policy to synchronize the manipulation strategy with active perception, achieving a balance between 6D pose accuracy and manipulation efficiency. Our experimental results in both simulated and real-world environments showcase the state-of-the-art effectiveness of our approach. We believe that our method will inspire further research on real-world-oriented robotic manipulation. See https://rgbmanip.github.io/ for more details. Boshi An, Yiran Geng, Kai Chen 0028, Xiaoqi Li 0020, Qi Dou 0001, Hao Dong 0003 |
ICRA | 5 |
| 2024 | Multi-objective Cross-task Learning via Goal-conditioned GPT-based Decision Transformers for Surgical Robot Task AutomationabstractSurgical robot task automation has been a promising research topic for improving surgical efficiency and quality. Learning-based methods have been recognized as an interesting paradigm and been increasingly investigated. However, existing approaches encounter difficulties in long-horizon goal-conditioned tasks due to the intricate compositional structure, which requires decision-making for a sequence of sub-steps and understanding of inherent dynamics of goal-reaching tasks. In this paper, we propose a new learning-based framework by leveraging the strong reasoning capability of the GPT-based architecture to automate surgical robotic tasks. The key to our approach is developing a goal-conditioned decision transformer to achieve sequential representations with goal-aware future indicators in order to enhance temporal reasoning. Moreover, considering to exploit a general understanding of dynamics inherent in manipulations, thus making the model’s reasoning ability to be task-agnostic, we also design a cross-task pretraining paradigm that uses multiple training objectives associated with data from diverse tasks. We have conducted extensive experiments on 10 tasks using the surgical robot learning simulator SurRoL [1]. The results show that our new approach achieves promising performance and task versatility compared to existing methods. The learned trajectories can be deployed on the da Vinci Research Kit (dVRK) for validating its practicality in real surgical robot settings. Our project website is at: https://med-air.github.io/SurRoL. Jiawei Fu 0001, Yonghao Long 0001, Kai Chen 0028, Qi Dou 0001 |
ICRA | 5 |
| 2024 | Ada-Tracker: Soft Tissue Tracking via Inter-Frame and Adaptive-template MatchingabstractSoft tissue tracking is crucial for computer-assisted interventions. Existing approaches mainly rely on extracting discriminative features from the template and videos to recover corresponding matches. However, it is difficult to adopt these techniques in surgical scenes, where tissues are changing in shape and appearance throughout the surgery. To address this problem, we exploit optical flow to naturally capture the pixel-wise tissue deformations and adaptively correct the tracked template. Specifically, we first implement an inter-frame matching mechanism to extract a coarse region of interest based on optical flow from consecutive frames. To accommodate appearance change and alleviate drift, we then propose an adaptive-template matching method, which updates the tracked template based on the reliability of the estimates. Our approach, Ada-Tracker, enjoys both short-term dynamics modeling by capturing local deformations and long-term dynamics modeling by introducing global temporal compensation. We evaluate our approach on the public SurgT benchmark, which is generated from Hamlyn, SCARED, and Kidney boundary datasets. The experimental results show that Ada-Tracker achieves superior accuracy and performs more robustly against prior works. Code is available at https://github.com/wrld/Ada-Tracker. Jiangliu Wang, Zhaoshuo Li, Tongyu Jia, Qi Dou 0001, Yun-Hui Liu 0001 |
ICRA | 5 |
| 2024 | Simultaneous Estimation of Shape and Force along Highly Deformable Surgical Manipulators Using Sparse FBG MeasurementabstractRecently, fiber optic sensors such as fiber Bragg gratings (FBGs) have been widely investigated for shape reconstruction and force estimation of flexible surgical robots. However, most existing approaches need precise model parameters of FBGs inside the fiber and their alignments with the flexible robots for accurate sensing results. Another challenge lies in online acquiring external forces at arbitrary locations along the flexible robots, which is highly required when with large deflections in robotic surgery. In this paper, we propose a novel data-driven paradigm for simultaneous estimation of shape and force along highly deformable flexible robots by using sparse strain measurement from a single-core FBG fiber. A thin-walled soft sensing tube helically embedded with FBG sensors is designed for a robotic-assisted flexible ureteroscope with large deflection up to 270° and a bend radius under 10 mm. We introduce and study three learning models by incorporating spatial strain encoders, and compare their performances in both free space without interactions as well as constrained environments with contact forces at different locations. The experimental results in terms of dynamic shape-force sensing accuracy demonstrate the effectiveness and superiority of the proposed methods. Yiang Lu, Bin Li 0082, Wei Chen 0068, Junyan Yan, Shing Shin Cheng, Jiangliu Wang, Jianshu Zhou, Qi Dou 0001, Yun-Hui Liu 0001 |
ICRA | 8 |
| 2024 | Interactive Navigation in Environments with Traversable Obstacles Using Large Language and Vision-Language ModelsabstractThis paper proposes an interactive navigation framework by using large language and vision-language models, allowing robots to navigate in environments with traversable obstacles. We utilize the large language model (GPT-3.5) and the open-set Vision-language Model (Grounding DINO) to create an action-aware costmap to perform effective path planning without fine-tuning. With the large models, we can achieve an end-to-end system from textual instructions like "Can you pass through the curtains to deliver medicines to me?", to bounding boxes (e.g., curtains) with action-aware attributes. They can be used to segment LiDAR point clouds into two parts: traversable and untraversable parts, and then an action-aware costmap is constructed for generating a feasible path. The pre-trained large models have great generalization ability and do not require additional annotated data for training, allowing fast deployment in the interactive navigation tasks. We choose to use multiple traversable objects such as curtains and grasses for verification by instructing the robot to traverse them. Besides, traversing curtains in a medical scenario was tested. All experimental results demonstrated the proposed framework’s effectiveness and adaptability to diverse environments. Zhen Zhang 0066, Anran Lin, Chun Wai Wong, Xiangyu Chu, Qi Dou 0001, K. W. Samuel Au |
ICRA | 5 |
| 2024 | EndoGSLAM: Real-Time Dense Reconstruction and Tracking in Endoscopic Surgeries Using Gaussian Splatting
Kailing Wang, Chen Yang 0023, Yuehao Wang, Sikuang Li, Yan Wang 0033, Qi Dou 0001, Xiaokang Yang 0001, Wei Shen 0002 |
MICCAI (6) | 6 |
| 2024 | Enhanced Scale-Aware Depth Estimation for Monocular Endoscopic Scenes with Geometric Modeling
Ruofeng Wei, Bin Li 0082, Kai Chen 0028, Yiyao Ma, Yun-Hui Liu 0001, Qi Dou 0001 |
MICCAI (6) | 6 |
| 2024 | Enhancing Federated Learning Performance Fairness via Collaboration Graph-Based Reinforcement Learning
Yuexuan Xia, Benteng Ma, Qi Dou 0001, Yong Xia 0001 |
MICCAI (10) | 3 |
| 2024 | Deform3DGS: Flexible Deformation for Fast Surgical Scene Reconstruction with Gaussian Splatting
Shuojue Yang, Qian Li 0036, Daiyun Shen, Bingchen Gong, Qi Dou 0001, Yueming Jin |
MICCAI (6) | 5 |
| 2024 | Language-Enhanced Local-Global Aggregation Network for Multi-organ Trauma Detection
Jianxun Yu, Qixin Hu, Meirui Jiang, Chin Ting Wong, Huimao Zhang, Qi Dou 0001 |
MICCAI (5) | 8 |
| 2024 | Incorporating Clinical Guidelines Through Adapting Multi-modal Large Language Model for Prostate Cancer PI-RADS Scoring
Manxi Lin, Hongda Guo, Xiaofan Zhang 0002, Ka Fung Peter Chiu, Aasa Feragen, Qi Dou 0001 |
MICCAI (5) | 7 |
| 2024 | Robust Semi-supervised Multimodal Medical Image Segmentation via Cross Modality Collaboration
Xiaogen Zhon, Yiyou Sun, Winnie Chiu-Wing Chu, Qi Dou 0001 |
MICCAI (1) | 5 |
| 2024 | Weakly-Supervised Medical Image Segmentation with Gaze Annotations
Yuan Zhong 0003, Chenhui Tang, Ruoxi Qi, Yuqi Gong, Pheng-Ann Heng, Janet Hui-wen Hsiao, Qi Dou 0001 |
MICCAI (3) | 9 |
| 2024 | Vision Foundation Model Enables Generalizable Object Pose EstimationabstractObject pose estimation plays a crucial role in robotic manipulation, however, its practical applicability still suffers from limited generalizability. This paper addresses the challenge of generalizable object pose estimation, particularly focusing on category-level object pose estimation for unseen object categories. Current methods either require impractical instance-level training or are confined to predefined categories, limiting their applicability. We propose VFM-6D, a novel framework that explores harnessing existing vision and language models, to elaborate object pose estimation into two stages: category-level object viewpoint estimation and object coordinate map estimation. Based on the two-stage framework, we introduce a 2D-to-3D feature lifting module and a shape-matching module, both of which leverage pre-trained vision foundation models to improve object representation and matching accuracy. VFM-6D is trained on cost-effective synthetic data and exhibits superior generalization capabilities. It can be applied to both instance-level unseen object pose estimation and category-level object pose estimation for novel categories. Evaluations on benchmark datasets demonstrate the effectiveness and versatility of VFM-6D in various real-world scenarios. Kai Chen 0028, Yiyao Ma, Stephen James, Jianshu Zhou, Yun-Hui Liu 0001, Pieter Abbeel, Qi Dou 0001 |
NeurIPS | 8 |
| 2024 | Local Superior Soups: A Catalyst for Model Merging in Cross-Silo Federated LearningabstractFederated learning (FL) is a learning paradigm that enables collaborative training of models using decentralized data.
Recently, the utilization of pre-trained weight initialization in FL has been demonstrated to effectively improve model performance.
However, the evolving complexity of current pre-trained models, characterized by a substantial increase in parameters, markedly intensifies the challenges associated with communication rounds required for their adaptation to FL.
To address these communication cost issues and increase the performance of pre-trained model adaptation in FL, we propose an innovative model interpolation-based local training technique called ``Local Superior Soups.''
Our method enhances local training across different clients, encouraging the exploration of a connected low-loss basin within a few communication rounds through regularized model interpolation.
This approach acts as a catalyst for the seamless adaptation of pre-trained models in in FL.
We demonstrated its effectiveness and efficiency across diverse widely-used FL datasets. Meirui Jiang, Xin Zhang 0054, Qi Dou 0001 |
NeurIPS | 4 |
| 2024 | FairMedFM: Fairness Benchmarking for Medical Imaging Foundation ModelsabstractThe advent of foundation models (FMs) in healthcare offers unprecedented opportunities to enhance medical diagnostics through automated classification and segmentation tasks. However, these models also raise significant concerns about their fairness, especially when applied to diverse and underrepresented populations in healthcare applications. Currently, there is a lack of comprehensive benchmarks, standardized pipelines, and easily adaptable libraries to evaluate and understand the fairness performance of FMs in medical imaging, leading to considerable challenges in formulating and implementing solutions that ensure equitable outcomes across diverse patient populations. To fill this gap, we introduce FairMedFM, a fairness benchmark for FM research in medical imaging. FairMedFM integrates with 17 popular medical imaging datasets, encompassing different modalities, dimensionalities, and sensitive attributes. It explores 20 widely used FMs, with various usages such as zero-shot learning, linear probing, parameter-efficient fine-tuning, and prompting in various downstream tasks -- classification and segmentation. Our exhaustive analysis evaluates the fairness performance over different evaluation metrics from multiple perspectives, revealing the existence of bias, varied utility-fairness trade-offs on different FMs, consistent disparities on the same datasets regardless FMs, and limited effectiveness of existing unfairness mitigation methods. Furthermore, FairMedFM provides an open-sourced codebase at https://github.com/FairMedFM/FairMedFM, supporting extendible functionalities and applications and inclusive for studies on FMs in medical imaging over the long term. Ruinan Jin, Yuan Zhong 0003, Qingsong Yao, Qi Dou 0001, Shaohua Kevin Zhou, Xiaoxiao Li 0001 |
NeurIPS | 5 |
| 2024 | Efficient Transferability Assessment for Selection of Pre-trained DetectorsabstractLarge-scale pre-training followed by downstream finetuning is an effective solution for transferring deeplearning-based models. Since finetuning all possible pretrained models is computational costly, we aim to predict the transferability performance of these pre-trained models in a computational efficient manner. Different from previous work that seek out suitable models for downstream classification and segmentation tasks, this paper studies the efficient transferability assessment of pre-trained object detectors. To this end, we build up a detector transferability benchmark which contains a large and diverse zoo of pre-trained detectors with various architectures, source datasets and training schemes. Given this zoo, we adopt 7 target datasets from 5 diverse domains as the downstream target tasks for evaluation. Further, we propose to assess classification and regression sub-tasks simultaneously in a unified framework. Additionally, we design a complementary metric for evaluating tasks with varying objects. Experimental results demonstrate that our method outperforms other state-of-the-art approaches in assessing transferability under different target domains while efficiently reducing wall-clock time 32× and requires a mere 5.2% memory footprint compared to brute-force fine-tuning of all pretrained detectors. Our assessment code and benchmark will be publicly available. Zhao Wang 0006, Aoxue Li, Zhenguo Li, Qi Dou 0001 |
WACV | 4 |
| 2024 | Editorial for the Special Issue on the 2022 Medical Imaging with Deep Learning Conference
Shadi Albarqouni, Christian F. Baumgartner, Qi Dou 0001, Ender Konukoglu, Bjoern Menze, Archana Venkataraman |
Medical Image Anal. | 3 |
| 2024 | 3DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation
Shizhan Gong, Yuan Zhong 0003, Wenao Ma, Jinpeng Li 0004, Zhao Wang 0006, Jingyang Zhang, Pheng-Ann Heng, Qi Dou 0001 |
Medical Image Anal. | 8 |
| 2024 | Where is VALDO? VAscular Lesions Detection and segmentatiOn challenge at MICCAI 2021
Carole H. Sudre, Kimberlin M. H. van Wijnen, Florian Dubost, Hieab Adams, David Atkinson, Frederik Barkhof, Mahlet A. Birhanu, Esther Bron, Robin Camarasa, Nish Chaturvedi, Qi Dou 0001, Tavia E. Evans, Ivan Ezhov, Haojun Gao, Marta Gironés-Sangüesa, Juan Domingo Gispert, Beatriz Gomez Anson, Alun D. Hughes, Mohammad Arfan Ikram, Silvia Ingala, Hans Rolf Jäger, Florian Kofler, Hugo J. Kuijf, Denis Kutnar, Bo Li 0088, Luigi Lorenzini, Bjoern Menze, José Luis Molinuevo, Yiwei Pan, Élodie Puybareau, Rafael Rehwald, Ruisheng Su, Lorna Smith, Therese Tillin, Guillaume Tochon, Hélène Urien, Bas H. M. van der Velden, Isabelle F. van der Velpen, Benedikt Wiestler, Frank J. Wolters, Pinar Yilmaz, Marius de Groot, Meike W. Vernooij, Marleen de Bruijne |
Medical Image Anal. | 14 |
| 2024 | FedDBL: Communication and Data Efficient Federated Deep-Broad Learning for Histopathological Tissue ClassificationabstractHistopathological tissue classification is a fundamental task in computational pathology. Deep learning (DL)-based models have achieved superior performance but centralized training suffers from the privacy leakage problem. Federated learning (FL) can safeguard privacy by keeping training samples locally, while existing FL-based frameworks require a large number of well-annotated training samples and numerous rounds of communication which hinder their viability in real-world clinical scenarios. In this article, we propose a lightweight and universal FL framework, named federated deep-broad learning (FedDBL), to achieve superior classification performance with limited training samples and only one-round communication. By simply integrating a pretrained DL feature extractor, a fast and lightweight broad learning inference system with a classical federated aggregation approach, FedDBL can dramatically reduce data dependency and improve communication efficiency. Five-fold cross-validation demonstrates that FedDBL greatly outperforms the competitors with only one-round communication and limited training samples, while it even achieves comparable performance with the ones under multiple-round communications. Furthermore, due to the lightweight design and one-round communication, FedDBL reduces the communication burden from 4.6 GB to only 138.4 KB per client using the ResNet-50 backbone at 50-round training. Extensive experiments also show the scalability of FedDBL on model generalization to the unseen dataset, various client numbers, model personalization and other image modalities. Since no data or deep model sharing across different clients, the privacy issue is well-solved and the model security is guaranteed with no model inversion attack risk. Code is available at https://github.com/tianpeng-deng/FedDBL. Tianpeng Deng, Guoqiang Han 0002, Zhenwei Shi 0002, Jiatai Lin, Qi Dou 0001, Zaiyi Liu, Xiao-jing Guo, C. L. Philip Chen, Chu Han |
IEEE Trans. Cybern. | 6 |
| 2024 | Self-Supervised Cyclic Diffeomorphic Mapping for Soft Tissue Deformation Recovery in Robotic Surgery ScenesabstractThe ability to recover tissue deformation from surgical video is fundamental for many downstream applications in robotic surgery. Despite noticeable advancements, this task remains under-explored due to the complex dynamics of soft tissues manipulated by surgical instruments. Achieving dense and accurate tissue tracking is further complicated by ambiguous pixel correspondence in regions with homogeneous texture. In this paper, we introduce a novel self-supervised framework to recover tissue deformations from stereo surgical videos. Our approach integrates semantics, cross-frame motion flow, and long-range temporal dependencies to accurately represent tissue dynamics for deformation recovery. Moreover, we incorporate diffeomorphic mapping to regularize the warping field to be physically more realistic. To comprehensively evaluate our method, we collected stereo surgical video clips containing three types of tissue manipulation (i.e., pushing, dissection and retraction) from two surgical procedures (i.e., hemicolectomy and mesorectal excision). Our method demonstrates promising results in capturing tissue 3D deformation, and generalizes well across different actions and procedures. It also outperforms current state-of-the-art approaches based on non-rigid registration and optical flow estimation. To the best of our knowledge, this is the first work on self-supervised learning for dense tissue deformation modeling from stereo surgical videos. The paper's code is available at: https://github.com/ med-air/RecoverTissueDeform. Shizhan Gong, Yonghao Long 0001, Kai Chen 0024, Yuliang Xiao, Alexis Cheng, Zerui Wang, Qi Dou 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2024 | Causal Effect Estimation on Imaging and Clinical Data for Treatment Decision Support of Aneurysmal Subarachnoid HemorrhageabstractAneurysmal subarachnoid hemorrhage is a medical emergency of brain that has high mortality and poor prognosis. Causal effect estimation of treatment strategies on patient outcomes is crucial for aneurysmal subarachnoid hemorrhage treatment decision-making. However, most existing studies on treatment decision-making support of this disease are unable to simultaneously compare the potential outcomes of different treatments for a patient. Furthermore, these studies fail to harmoniously integrate the imaging data with non-imaging clinical data, both of which are useful in clinical scenarios. In this paper, we estimate the causal effect of various treatments on patients with aneurysmal subarachnoid hemorrhage by integrating plain CT with non-imaging clinical data, which is represented using structured tabular data. Specifically, we first propose a novel scheme that uses multi-modality confounders distillation architecture to predict the treatment outcome and treatment assignment simultaneously. With these distilled confounder features, we design an imaging and non-imaging interaction representation learning strategy to use the complementary information extracted from different modalities to balance the feature distribution of different treatment groups. We have conducted extensive experiments using a clinical dataset of 656 subarachnoid hemorrhage cases, which was collected from the Hospital Authority Data Collaboration Laboratory in Hong Kong. Our method shows consistent improvements on the evaluation metrics of treatment effect estimation, achieving state-of-the-art results over strong competitors. Code is released at https://github.com/med-air/TOP-aSAH. Wenao Ma, Cheng Chen 0013, Yuqi Gong, Nga Yan Chan, Meirui Jiang, Calvin Hoi-Kwan Mak, Jill M. Abrigo, Qi Dou 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2024 | Breast Cancer Classification From Digital Pathology Images via Connectivity-Aware Graph TransformerabstractAutomated classification of breast cancer subtypes from digital pathology images has been an extremely challenging task due to the complicated spatial patterns of cells in the tissue micro-environment. While newly proposed graph transformers are able to capture more long-range dependencies to enhance accuracy, they largely ignore the topological connectivity between graph nodes, which is nevertheless critical to extract more representative features to address this difficult task. In this paper, we propose a novel connectivity-aware graph transformer (CGT) for phenotyping the topology connectivity of the tissue graph constructed from digital pathology images for breast cancer classification. Our CGT seamlessly integrates connectivity embedding to node feature at every graph transformer layer by using local connectivity aggregation, in order to yield more comprehensive graph representations to distinguish different breast cancer subtypes. In light of the realistic intercellular communication mode, we then encode the spatial distance between two arbitrary nodes as connectivity bias in self-attention calculation, thereby allowing the CGT to distinctively harness the connectivity embedding based on the distance of two nodes. We extensively evaluate the proposed CGT on a large cohort of breast carcinoma digital pathology images stained by Haematoxylin & Eosin. Experimental results demonstrate the effectiveness of our CGT, which outperforms state-of-the-art methods by a large margin. Codes are released on https://github.com/wang-kang-6/CGT. Kang Wang 0004, Feiyang Zheng, Hongning Dai, Qi Dou 0001, Harry Qin |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Efficient Deformable Tissue Reconstruction via Orthogonal Neural PlaneabstractIntraoperative imaging techniques for reconstructing deformable tissues in vivo are pivotal for advanced surgical systems. Existing methods either compromise on rendering quality or are excessively computationally intensive, often demanding dozens of hours to perform, which significantly hinders their practical application. In this paper, we introduce Fast Orthogonal Plane (Forplane), a novel, efficient framework based on neural radiance fields (NeRF) for the reconstruction of deformable tissues. We conceptualize surgical procedures as 4D volumes, and break them down into static and dynamic fields comprised of orthogonal neural planes. This factorization discretizes the four-dimensional space, leading to a decreased memory usage and faster optimization. A spatiotemporal importance sampling scheme is introduced to improve performance in regions with tool occlusion as well as large motions and accelerate training. An efficient ray marching method is applied to skip sampling among empty regions, significantly improving inference speed. Forplane accommodates both binocular and monocular endoscopy videos, demonstrating its extensive applicability and flexibility. Our experiments, carried out on two in vivo datasets, the EndoNeRF and Hamlyn datasets, demonstrate the effectiveness of our framework. In all cases, Forplane substantially accelerates both the optimization process (by over 100 times) and the inference process (by over 15 times) while maintaining or even improving the quality across a variety of non-rigid deformations. This significant performance improvement promises to be a valuable asset for future intraoperative surgical applications. The code of our project is now available at https://github.com/Loping151/ForPlane. Chen Yang 0023, Kailing Wang, Yuehao Wang, Qi Dou 0001, Xiaokang Yang 0001, Wei Shen 0002 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Multicontrast MRI Super-Resolution via Transformer-Empowered Multiscale Contextual Matching and AggregationabstractMagnetic resonance imaging (MRI) possesses the unique versatility to acquire images under a diverse array of distinct tissue contrasts, which makes multicontrast super-resolution (SR) techniques possible and needful. Compared with single-contrast MRI SR, multicontrast SR is expected to produce higher quality images by exploiting a variety of complementary information embedded in different imaging contrasts. However, existing approaches still have two shortcomings: 1) most of them are convolution-based methods and, hence, weak in capturing long-range dependencies, which are essential for MR images with complicated anatomical patterns and 2) they ignore to make full use of the multicontrast features at different scales and lack effective modules to match and aggregate these features for faithful SR. To address these issues, we develop a novel multicontrast MRI SR network via transformer-empowered multiscale feature matching and aggregation, dubbed McMRSR$^{++}$. First, we tame transformers to model long-range dependencies in both reference and target images at different scales. Then, a novel multiscale feature matching and aggregation method is proposed to transfer corresponding contexts from reference features at different scales to the target features and interactively aggregate them Furthermore, a texture-preserving branch and a contrastive constraint are incorporated into our framework for enhancing the textural details in the SR images. Experimental results on both public and clinical in vivo datasets show that McMRSR$^{++}$outperforms state-of-the-art methods under peak signal to noise ratio (PSNR), structure similarity index measure (SSIM), and root mean square error (RMSE) metrics significantly. Visual results demonstrate the superiority of our method in restoring structures, demonstrating its great potential to improve scan efficiency in clinical practice. Chengyan Wang, Qi Dou 0001, David Zhang 0001, Harry Qin |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Open-Vocabulary Object Detection with Meta Prompt Representation and Instance Contrastive Optimization
Zhao Wang 0006, Aoxue Li, Fengwei Zhou, Zhenguo Li, Qi Dou 0001 |
BMVC | 5 |
| 2023 | Why is the Winner the Best?abstractInternational benchmarking competitions have become fundamental for the comparative performance assessment of image analysis methods. However, little attention has been given to investigating what can be learnt from these competitions. Do they really generate scientific progress? What are common and successful participation strategies? What makes a solution superior to a competing method? To address this gap in the literature, we performed a multicenter study with all 80 competitions that were conducted in the scope of IEEE ISBI 2021 and MICCAI 2021. Statistical analyses performed based on comprehensive descriptions of the submitted algorithms linked to their rank as well as the underlying participation strategies revealed common characteristics of winning solutions. These typically include the use of multi-task learning (63%) and/or multi-stage pipelines (61%), and a focus on augmentation (100%), image preprocessing (97%), data curation (79%), and post-processing (66%). The “typical” lead of a winning team is a computer scientist with a doctoral degree, five years of experience in biomedical image analysis, and four years of experience in deep learning. Two core general development strategies stood out for highly-ranked teams: the reflection of the metrics in the method design and the focus on analyzing and handling failure cases. According to the organizers, 43% of the winning algorithms exceeded the state of the art but only 11% completely solved the respective domain problem. The insights of our study could help researchers (1) improve algorithm development strategies when approaching new problems, and (2) focus on open research questions revealed by this work. Matthias Eisenmann, Annika Reinke, Vivienn Weru, Minu Tizabi, Fabian Isensee, Tim Adler, Sharib Ali, Vincent Andrearczyk, Marc Aubreville, Ujjwal Baid, Spyridon Bakas, Niranjan Balu, Sophia Bano, Jorge Bernal, Sebastian Bodenstedt, Alessandro Casella, Veronika Cheplygina, Marie Daum, Marleen de Bruijne, Adrien Depeursinge, Reuben Dorent, Jan Egger, David Gage Ellis, Sandy Engelhardt, Melanie Ganz-Benjaminsen, Noha M. Ghatwary, Gabriel Girard, Patrick Godau, Anubha Gupta, Lasse Hansen, Kanako Harada, Mattias P. Heinrich, Nicholas Heller, Alessa Hering, Arnaud Huaulmé, Pierre Jannin, A. Emre Kavur, Oldrich Kodym, Michal Kozubek 0001, Jianning Li 0002, Hongwei Li 0004, Jun Ma 0016, Carlos Martín-Isla, Bjoern Menze, J. Alison Noble, Valentin Oreiller, Nicolas Padoy, Sarthak Pati, Kelly Payette, Tim Rädsch, Jonathan Rafael-Patino, Vivek Singh Bawa, Stefanie Speidel, Carole H. Sudre, Kimberlin M. H. van Wijnen, Martin Wagner 0001, D. Wei, Amine Yamlahi, Moi Hoon Yap, C. Yuan, Maximilian Zenk, A. Zia, David Zimmerer, Dogu Baran Aydogan, Binod Bhattarai, Louise Bloch, Raphael Brüngel, J. Cho, C. Choi, Qi Dou 0001, Ivan Ezhov, Christoph M. Friedrich, C. Fuller, Rebati Raman Gaire, Adrian Galdran, Álvaro García-Faura, Maria Grammatikopoulou, S. Hong, Mostafa Jahanifar, I. Jang, Abdolrahim Kadkhodamohammadi, I. Kang, Florian Kofler, S. Kondo, Hugo J. Kuijf, M. Luu, Tomaz Martincic, Pedro Morais, Mohamed A. Naser, Bruno Oliveira 0002, David Owen 0001, S. Pang, Szymon Plotka, Élodie Puybareau, Nasir M. Rajpoot, K. Ryu, Numan Saeed, Adam J. Shephard, Dejan Stepec, Ronast Subedi, Guillaume Tochon, Helena R. Torres, Hélène Urien, João L. Vilaça, Kareem A. Wahid, Benedikt Wiestler, Marek Wodzinski, F. Xia, J. Xie, Z. Xiong, Sen Yang 0006, Klaus H. Maier-Hein, Paul F. Jaeger, Annette Kopp-Schneider, Lena Maier-Hein |
CVPR | 70 |
| 2023 | Fair Federated Medical Image Segmentation via Client Contribution EstimationabstractHow to ensure fairness is an important topic in federated learning (FL). Recent studies have investigated how to reward clients based on their contribution (collaboration fairness), and how to achieve uniformity of performance across clients (performance fairness). Despite achieving progress on either one, we argue that it is critical to consider them together, in order to engage and motivate more diverse clients joining FL to derive a high-quality global model. In this work, we propose a novel method to optimize both types of fairness simultaneously. Specifically, we propose to estimate client contribution in gradient and data space. In gradient space, we monitor the gradient direction differences of each client with respect to others. And in data space, we measure the prediction error on client data using an auxiliary model. Based on this contribution estimation, we propose a FL method, federated training via contribution estimation (FedCE), i.e., using estimation as global model aggregation weights. We have theoretically analyzed our method and empirically evaluated it on two real-world medical datasets. The effectiveness of our approach has been validated with significant performance improvements, better collaboration fairness, better performance fairness, and comprehensive analytical studies. Code is available at https://nvidia.github.io/NVFlare/research/fed-ce Meirui Jiang, Holger Roth, Wenqi Li 0001, Dong Yang 0005, Can Zhao 0001, Vishwesh Nath, Daguang Xu, Qi Dou 0001, Ziyue Xu 0001 |
CVPR | 8 |
| 2023 | Video Dehazing via a Multi-Range Temporal Alignment Network with Physical PriorabstractVideo dehazing aims to recover haze-free frames with high visibility and contrast. This paper presents a novel framework to effectively explore the physical haze priors and aggregate temporal information. Specifically, we design a memory-based physical prior guidance module to encode the prior-related features into long-range memory. Besides, we formulate a multi-range scene radiance recovery module to capture space-time dependencies in multiple space-time ranges, which helps to effectively aggregate temporal information from adjacent frames. Moreover, we construct the first large-scale outdoor video dehazing benchmark dataset, which contains videos in various real-world scenarios. Experimental results on both synthetic and real conditions show the superiority of our proposed method. Xiaowei Hu 0001, Lei Zhu 0003, Qi Dou 0001, Jifeng Dai, Yu Qiao 0001, Pheng-Ann Heng |
CVPR | 4 |
| 2023 | Deep Fusion Transformer Network with Weighted Vector-Wise Keypoints Voting for Robust 6D Object Pose EstimationabstractOne critical challenge in 6D object pose estimation from a single RGBD image is efficient integration of two different modalities, i.e., color and depth. In this work, we tackle this problem by a novel Deep Fusion Transformer (DFTr) block that can aggregate cross-modality features for improving pose estimation. Unlike existing fusion methods, the proposed DFTr can better model cross-modality semantic correlation by leveraging their semantic similarity, such that globally enhanced features from different modalities can be better integrated for improved information extraction. Moreover, to further improve robustness and efficiency, we introduce a novel weighted vector-wise voting algorithm that employs a non-iterative global optimization strategy for precise 3D keypoint localization while achieving near real-time inference. Extensive experiments show the effectiveness and strong generalization capability of our proposed 3D keypoint voting algorithm. Results on four widely used benchmarks also demonstrate that our method outperforms the state-of-the-art methods by large margins. Code is available at https://github.com/junzastar/DFTr_Voting. Jun Zhou 0007, Kai Chen 0028, Linlin Xu, Qi Dou 0001, Harry Qin |
ICCV | 4 |
| 2023 | Two-Stage Grasping: A New Bin Picking Framework for Small ObjectsabstractThis paper proposes a novel bin picking framework, two-stage grasping, aiming at precise grasping of cluttered small objects. Object density estimation and rough grasping are conducted in the first stage. Fine segmentation, detection, grasping, and pushing are performed in the second stage. A small object bin picking system has been realized to exhibit the concept of two-stage grasping. Experiments have shown the effectiveness of the proposed framework. Unlike traditional bin picking methods focusing on vision-based grasping planning using classic frameworks, the challenges of picking cluttered small objects can be solved by the proposed new framework with simple vision detection and planning. Jianshu Zhou, Junda Huang, Yichuan Li 0002, Ng Cheng Meng, Qi Dou 0001, Yun-Hui Liu 0001 |
ICRA | 7 |
| 2023 | StereoPose: Category-Level 6D Transparent Object Pose Estimation from Stereo Images via Back-View NOCSabstractMost existing methods for category-level pose estimation rely on object point clouds. However, when considering transparent objects, depth cameras are usually not able to capture high-quality data, resulting in point clouds with severe artifacts. Without a complete point cloud, existing methods are not applicable to challenging transparent objects. To tackle this problem, we present StereoPose, a novel stereo image based framework for category-level object pose estimation, ideally suited for transparent objects. For a robust estimation from pure stereo images, we develop a pipeline that decouples category-level pose estimation into object size estimation, initial pose estimation, and pose refinement. StereoPose then estimates object pose based on representation in the normalized object coordinate space (NOCS). To address the issue of image content aliasing, we further define a back-view NOCS map for the transparent object. The back-view NOCS aims to reduce the network learning ambiguity caused by content aliasing, and leverage informative cues on the back of the transparent object for more accurate pose estimation. To further improve the performance of the stereo framework, StereoPose is equipped with a parallax attention module for stereo feature fusion and an epipolar loss for improving the stereo-view consistency of network predictions. Extensive experiments on the public TOD dataset demonstrate the superiority of the proposed StereoPose framework for category-level 6D transparent object pose estimation. Code and demos will be available on the project homepage: www.cse.cuhk.edu.hk/~kaichen/stereopose.html. Kai Chen 0028, Stephen James, Congying Sui, Yun-Hui Liu 0001, Pieter Abbeel, Qi Dou 0001 |
ICRA | 6 |
| 2023 | Demonstration-Guided Reinforcement Learning with Efficient Exploration for Task Automation of Surgical RobotabstractTask automation of surgical robot has the potentials to improve surgical efficiency. Recent reinforcement learning (RL) based approaches provide scalable solutions to surgical automation, but typically require extensive data collection to solve a task if no prior knowledge is given. This issue is known as the exploration challenge, which can be alleviated by providing expert demonstrations to an RL agent. Yet, how to make effective use of demonstration data to improve exploration efficiency still remains an open challenge. In this work, we introduce Demonstration-guided EXploration (DEX), an efficient reinforcement learning algorithm that aims to overcome the exploration problem with expert demonstrations for surgical automation. To effectively exploit demonstrations, our method estimates expert-like behaviors with higher values to facilitate productive interactions, and adopts non-parametric regression to enable such guidance at states unobserved in demonstration data. Extensive experiments on 10 surgical manipulation tasks from SurRoL, a comprehensive surgical simulation platform, demonstrate significant improvements in the exploration efficiency and task success rates of our method. Moreover, we also deploy the learned policies to the da Vinci Research Kit (dVRK) platform to show the effectiveness on the real robot. Code is available at https://github.com/med-air/DEX. Kai Chen 0028, Bin Li 0082, Yun-Hui Liu 0001, Qi Dou 0001 |
ICRA | 5 |
| 2023 | Autonomous Intelligent Navigation for Flexible Endoscopy Using Monocular Depth Guidance and 3-D Shape PlanningabstractRecent advancements toward perception and decision-making of flexible endoscopes have shown great potential in computer-aided surgical interventions. However, owing to modeling uncertainty and inter-patient anatomical variation in flexible endoscopy, the challenge remains for efficient and safe navigation in patient-specific scenarios. This paper presents a novel data-driven framework with self-contained visual-shape fusion for autonomous intelligent navigation of flexible endoscopes requiring no priori knowledge of system models and global environments. A learning-based adaptive visual servoing controller is proposed to online update the eye-in-hand vision-motor configuration and steer the endoscope, which is guided by monocular depth estimation via a vision transformer (ViT). To prevent unnecessary and excessive interactions with surrounding anatomy, an energy-motivated shape planning algorithm is introduced through entire endoscope 3-D proprioception from embedded fiber Bragg grating (FBG) sensors. Furthermore, a model predictive control (MPC) strategy is developed to minimize the elastic potential energy flow and simultaneously optimize the steering policy. Dedicated navigation experiments on a robotic-assisted flexible endoscope with an FBG fiber in several phantom environments demonstrate the effectiveness and adaptability of the proposed framework. Yiang Lu, Ruofeng Wei, Bin Li 0082, Wei Chen 0068, Jianshu Zhou, Qi Dou 0001, Dong Sun 0001, Yun-Hui Liu 0001 |
ICRA | 6 |
| 2023 | Value-Informed Skill Chaining for Policy Learning of Long-Horizon Tasks with Surgical RobotabstractReinforcement learning is still struggling with solving long-horizon surgical robot tasks which involve multiple steps over an extended duration of time due to the policy exploration challenge. Recent methods try to tackle this problem by skill chaining, in which the long-horizon task is decomposed into multiple subtasks for easing the exploration burden and subtask policies are temporally connected to complete the whole long-horizon task. However, smoothly connecting all subtask policies is difficult for surgical robot scenarios. Not all states are equally suitable for connecting two adjacent subtasks. An undesired terminate state of the previous subtask would make the current subtask policy unstable and result in a failed execution. In this work, we introduce value-informed skill chaining (ViSkill), a novel reinforcement learning framework for long-horizon surgical robot tasks. The core idea is to distinguish which terminal state is suitable for starting all the following subtask policies. To achieve this target, we introduce a state value function that estimates the expected success probability of the entire task given a state. Based on this value function, a chaining policy is learned to instruct subtask policies to terminate at the state with the highest value so that all subsequent policies are more likely to be connected for accomplishing the task. We demonstrate the effectiveness of our method on three complex surgical robot tasks from SurRoL, a comprehensive surgical simulation platform, achieving high task success rates and execution efficiency. Code is available at https: / /github. com/med-air/ViSkill. Kai Chen 0028, Jianan Li 0006, Yonghao Long 0001, Qi Dou 0001 |
IROS | 6 |
| 2023 | End-to-End Learning of Deep Visuomotor Policy for Needle PickingabstractNeedle picking is a challenging manipulation task in robot-assisted surgery due to the characteristics of small slender shapes of needles, needles' variations in shapes and sizes, and demands for millimeter-level control. Prior works, heavily relying on the prior of needles (e.g., geometric models), are hard to scale to unseen needles' variations. In this paper, we present the first end- to-end learning method to train deep visuomotor policy for needle picking. Concretely, we propose DreamerfD to maximally leverage demonstrations to improve the learning efficiency of a state-of-the-art model-based reinforcement learning method, DreamerV2; Since Variational Auto-Encoder (VAE) in DreamerV2 is difficult to scale to high-resolution images, we propose Dynamic Spotlight Adaptation to represent control-related visual signals in a low-resolution image space; Virtual Clutch is also proposed to reduce per-formance degradation due to significant error between prior and posterior encoded states at the beginning of a rollout. We conducted extensive experiments in simulation to evaluate the performance, robustness, in-domain variation adaptation, and effectiveness of individual components of our method. Our method, trained by 8k demonstration timesteps and 140k online policy timesteps, can achieve a remarkable success rate of 80%. Furthermore, our method effectively demonstrated its superiority in generalization to unseen in-domain variations including needle variations and image disturbance, highlighting its robustness and versatility. Codes and videos are available at https://sites.google.com/view/DreamerfD. Bin Li 0082, Xiangyu Chu, Qi Dou 0001, Yun-Hui Liu 0001, K. W. Samuel Au |
IROS | 4 |
| 2023 | Visual-Kinematics Graph Learning for Procedure-Agnostic Instrument Tip Segmentation in Robotic SurgeriesabstractAccurate segmentation of surgical instrument tip is an important task for enabling downstream applications in robotic surgery, such as surgical skill assessment, tool-tissue interaction and deformation modeling, as well as surgical autonomy. However, this task is very challenging due to the small sizes of surgical instrument tips, and significant variance of surgical scenes across different procedures. Although much effort has been made on visual-based methods, existing segmentation models still suffer from low robustness thus not usable in practice. Fortunately, kinematics data from the robotic system can provide reliable prior for instrument location, which is consistent regardless of different surgery types. To make use of such multi-modal information, we propose a novel visual-kinematics graph learning framework to accurately segment the instrument tip given various surgical procedures. Specifically, a graph learning framework is proposed to encode relational features of instrument parts from both image and kinematics. Next, a cross-modal contrastive loss is designed to incorporate robust geometric prior from kinematics to image for tip segmentation. We have conducted experiments on a private paired visual-kinematics dataset including multiple procedures, i.e., prostatectomy, total mesorectal excision, fundoplication and distal gastrectomy on cadaver, and distal gastrectomy on porcine. The leave-one-procedure-out cross validation demon-strated that our proposed multi-modal segmentation method significantly outperformed current image-based state-of-the-art approaches, exceeding averagely 11.2% on Dice. Yonghao Long 0001, Kai Chen 0028, Cheuk Hei Leung, Zerui Wang, Qi Dou 0001 |
IROS | 6 |
| 2023 | FedSoup: Improving Generalization and Personalization in Federated Learning via Selective Model Interpolation
Meirui Jiang, Qi Dou 0001 |
MICCAI (2) | 3 |
| 2023 | ArSDM: Colonoscopy Images Synthesis with Adaptive Refinement Semantic Diffusion Models
Yuncheng Jiang 0002, Shuangyi Tan, Xusheng Wu, Qi Dou 0001, Zhen Li 0026, Guanbin Li |
MICCAI (2) | 5 |
| 2023 | Client-Level Differential Privacy via Adaptive Intermediary in Federated Medical Imaging
Meirui Jiang, Yuan Zhong 0003, Anjie Le, Qi Dou 0001 |
MICCAI (2) | 5 |
| 2023 | Fast Non-Markovian Diffusion Model for Weakly Supervised Anomaly Detection in Brain MR Images
Jinpeng Li 0004, Hanqun Cao, Furui Liu, Qi Dou 0001, Guangyong Chen, Pheng-Ann Heng |
MICCAI (5) | 5 |
| 2023 | Learning Robust Classifier for Imbalanced Medical Image Dataset with Noisy Labels by Minimizing Invariant Risk
Jinpeng Li 0004, Hanqun Cao, Furui Liu, Qi Dou 0001, Guangyong Chen, Pheng-Ann Heng |
MICCAI (6) | 5 |
| 2023 | Imitation Learning from Expert Video Data for Dissection Trajectory Prediction in Endoscopic Surgical Procedure
Jianan Li 0006, Yueming Jin, Yueyao Chen, Hon-Chi Yip, Markus Scheppach, Philip W. Y. Chiu, Yeung Yam, Helen M. Meng, Qi Dou 0001 |
MICCAI (9) | 9 |
| 2023 | Treatment Outcome Prediction for Intracerebral Hemorrhage via Generative Prognostic Model with Imaging and Tabular Data
Wenao Ma, Cheng Chen 0013, Jill M. Abrigo, Calvin Hoi-Kwan Mak, Yuqi Gong, Nga Yan Chan, Chu Han, Zaiyi Liu, Qi Dou 0001 |
MICCAI (5) | 9 |
| 2023 | Foundation Model for Endoscopy Video Analysis via Large-Scale Self-supervised Pre-train
Zhao Wang 0006, Shaoting Zhang 0001, Qi Dou 0001 |
MICCAI (9) | 4 |
| 2023 | RecolorNeRF: Layer Decomposed Radiance Fields for Efficient Color Editing of 3D ScenesabstractRadiance fields have gradually become a main representation of media. Although its appearance editing has been studied, how to achieve view-consistent recoloring in an efficient manner is still under explored. We present RecolorNeRF, a novel user-friendly color editing approach for the neural radiance fields. Our key idea is to decompose the scene into a set of pure-colored layers, forming a palette. By this means, color manipulation can be conducted by altering the color components of the palette directly. To support efficient palette-based editing, the color of each layer needs to be as representative as possible. In the end, the problem is formulated as an optimization problem, where the layers and their blending weights are jointly optimized with the NeRF itself. Extensive experiments show that our jointly-optimized layer decomposition can be used against multiple backbones and produce photo-realistic recolored novel-view renderings. We demonstrate that RecolorNeRF outperforms baseline methods both quantitatively and qualitatively for color editing even in complex real-world scenes. Bingchen Gong, Yuehao Wang, Xiaoguang Han 0001, Qi Dou 0001 |
ACM Multimedia | 4 |
| 2023 | Uncertainty Estimation for Safety-critical Scene Segmentation via Fine-grained Reward MaximizationabstractUncertainty estimation plays an important role for future reliable deployment of deep segmentation models in safety-critical scenarios such as medical applications. However, existing methods for uncertainty estimation have been limited by the lack of explicit guidance for calibrating the prediction risk and model confidence. In this work, we propose a novel fine-grained reward maximization (FGRM) framework, to address uncertainty estimation by directly utilizing an uncertainty metric related reward function with a reinforcement learning based model tuning algorithm. This would benefit the model uncertainty estimation with direct optimization guidance for model calibration. Specifically, our method designs a new uncertainty estimation reward function using the calibration metric, which is maximized to fine-tune an evidential learning pre-trained segmentation model for calibrating prediction risk. Importantly, we innovate an effective fine-grained parameter update scheme, which imposes fine-grained reward-weighting of each network parameter according to the parameter importance quantified by the fisher information matrix. To the best of our knowledge, this is the first work exploring reward optimization for model uncertainty estimation in safety-critical vision tasks. The effectiveness of our method is demonstrated on two large safety-critical surgical scene segmentation datasets under two different uncertainty estimation settings. With real-time one forward pass at inference, our method outperforms state-of-the-art methods by a clear margin on all the calibration metrics of uncertainty estimation, while maintaining a high task accuracy for the segmentation results. Code is available at https://github.com/med-air/FGRM. Hongzheng Yang, Cheng Chen 0013, Yueyao Chen, Markus Scheppach, Hon-Chi Yip, Qi Dou 0001 |
NeurIPS | 6 |
| 2023 | SeamlessNeRF: Stitching Part NeRFs with Gradient PropagationabstractNeural Radiance Fields (NeRFs) have emerged as promising digital mediums of 3D objects and scenes, sparking a surge in research to extend the editing capabilities in this domain. The task of seamless editing and merging of multiple NeRFs, resembling the “Poisson blending” in 2D image editing, remains a critical operation that is under-explored by existing work. To fill this gap, we propose SeamlessNeRF, a novel approach for seamless appearance blending of multiple NeRFs. In specific, we aim to optimize the appearance of a target radiance field in order to harmonize its merge with a source field. We propose a well-tailored optimization procedure for blending, which is constrained by 1) pinning the radiance color in the intersecting boundary area between the source and target fields and 2) maintaining the original gradient of the target. Extensive experiments validate that our approach can effectively propagate the source appearance from the boundary area to the entire target field through the gradients. To the best of our knowledge, SeamlessNeRF is the first work that introduces gradient-guided appearance editing to radiance fields, offering solutions for seamless stitching of 3D objects represented in NeRFs. Our code and more results are available at https://sites.google.com/view/seamlessnerf. Bingchen Gong, Yuehao Wang, Xiaoguang Han 0001, Qi Dou 0001 |
SIGGRAPH Asia | 4 |
| 2023 | Federated Domain Generalization for Image Recognition via Cross-Client Style TransferabstractDomain generalization (DG) has been a hot topic in image recognition, with a goal to train a general model that can perform well on unseen domains. Recently, federated learning (FL), an emerging machine learning paradigm to train a global model from multiple decentralized clients without compromising data privacy, has brought new challenges and possibilities to DG. In the FL scenario, many existing state-of-the-art (SOTA) DG methods become ineffective because they require the centralization of data from different domains during training. In this paper, we propose a novel domain generalization method for image recognition under federated learning through cross-client style transfer (CCST) without exchanging data samples. Our CCST method can lead to more uniform distributions of source clients, and make each local model learn to fit the image styles of all the clients to avoid the different model biases. Two types of style (single image style and overall domain style) with corresponding mechanisms are proposed to be chosen according to different scenarios. Our style representation is exceptionally lightweight and can hardly be used to reconstruct the dataset. The level of diversity is also flexible to be controlled with a hyper-parameter. Our method outperforms recent SOTA DG methods on two DG benchmarks (PACS, OfficeHome) and a large-scale medical image dataset (Camelyon17) in the FL setting. Last but not least, our method is orthogonal to many classic DG methods, achieving additive performance by combined utilization. Our code is available at: https://chenjunming.ml/proj/CCST. Meirui Jiang, Qi Dou 0001, Qifeng Chen 0001 |
WACV | 3 |
| 2023 | Adaptive feature aggregation based multi-task learning for uncertainty-guided semi-supervised medical image segmentation
Bin Sui, Chengyan Wang, Qi Dou 0001, Harry Qin |
Expert Syst. Appl. | 4 |
| 2023 | The Liver Tumor Segmentation Benchmark (LiTS)abstractIn this work, we report the set-up and results of the Liver Tumor Segmentation Benchmark (LiTS), which was organized in conjunction with the IEEE International Symposium on Biomedical Imaging (ISBI) 2017 and the International Conferences on Medical Image Computing and Computer-Assisted Intervention (MICCAI) 2017 and 2018. The image dataset is diverse and contains primary and secondary tumors with varied sizes and appearances with various lesion-to-background levels (hyper-/hypo-dense), created in collaboration with seven hospitals and research institutions. Seventy-five submitted liver and liver tumor segmentation algorithms were trained on a set of 131 computed tomography (CT) volumes and were tested on 70 unseen test images acquired from different patients. We found that not a single algorithm performed best for both liver and liver tumors in the three events. The best liver segmentation algorithm achieved a Dice score of 0.963, whereas, for tumor segmentation, the best algorithms achieved Dices scores of 0.674 (ISBI 2017), 0.702 (MICCAI 2017), and 0.739 (MICCAI 2018). Retrospectively, we performed additional analysis on liver tumor detection and revealed that not all top-performing segmentation algorithms worked well for tumor detection. The best liver tumor detection method achieved a lesion-wise recall of 0.458 (ISBI 2017), 0.515 (MICCAI 2017), and 0.554 (MICCAI 2018), indicating the need for further research. LiTS remains an active benchmark and resource for research, e.g., contributing the liver-related segmentation tasks in http://medicaldecathlon.com/. In addition, both data and online evaluation are accessible via https://competitions.codalab.org/competitions/17094. Patrick Bilic, Patrick Ferdinand Christ, Hongwei Li 0004, Eugene Vorontsov, Avi Ben-Cohen, Georgios Kaissis, Adi Szeskin, Colin Jacobs, Gabriel Efrain Humpire Mamani, Gabriel Chartrand, Fabian Lohöfer, Julian Walter Holch, Wieland H. Sommer, Felix Hofmann, Alexandre Hostettler, Naama Lev-Cohain, Michal Drozdzal, Michal Amitai, Refael Vivanti, Jacob Sosna, Ivan Ezhov, Anjany Sekuboyina, Fernando Navarro, Florian Kofler, Johannes C. Paetzold, Suprosanna Shit, Xiaobin Hu, Jana Lipková, Markus Rempfler, Marie Piraud, Jan Kirschke, Benedikt Wiestler, Christian Hülsemeyer, Marcel Beetz, Florian Ettlinger, Michela Antonelli, Woong Bae, Miriam Bellver, Lei Bi 0001, Hao Chen 0011, Grzegorz Chlebus, Erik Dam, Qi Dou 0001, Chi-Wing Fu, Bogdan Georgescu, Xavier Giró-i-Nieto, Felix Grün, Xu Han 0009, Pheng-Ann Heng, Jürgen Hesser, Jan Hendrik Moltz, Christian Igel, Fabian Isensee, Paul F. Jaeger, Fucang Jia, Krishna Chaitanya Kaluva, Mahendra Khened, Ildoo Kim, Jae-Hun Kim, Sungwoong Kim, Simon Kohl, Tomasz K. Konopczynski, Avinash Kori, Ganapathy Krishnamurthi, Xiaomeng Li 0001, John S. Lowengrub, Jun Ma 0016, Klaus H. Maier-Hein, Kevis-Kokitsi Maninis, Hans Meine, Dorit Merhof, Akshay Pai, Mathias Perslev, Jens Petersen, Jordi Pont-Tuset, Xiaojuan Qi 0001, Oliver Rippel, Karsten Roth, Ignacio Sarasua, Andrea Schenk, Zengming Shen, Jordi Torres, Christian Wachinger, Chunliang Wang, Leon Weninger, Daguang Xu, Xiaoping Yang 0001, Simon C. H. Yu, Yading Yuan, Miao Yue, Liping Zhang 0009, Manuel Jorge Cardoso, Spyridon Bakas, Rickmer Braren, Volker Heinemann, Christopher Joseph Pal, An Tang, Samuel Kadoury, Luc Soler, Bram van Ginneken, Hayit Greenspan, Leo Joskowicz, Bjoern Menze |
Medical Image Anal. | 44 |
| 2023 | Region-focused multi-view transformer-based generative adversarial network for cardiac cine MRI reconstruction
Chengyan Wang, Chen Qin, Shuo Wang 0011, Qi Dou 0001, Harry Qin |
Medical Image Anal. | 6 |
| 2023 | Comparative validation of machine learning algorithms for surgical workflow and skill analysis with the HeiChole benchmarkabstractPURPOSE: Surgical workflow and skill analysis are key technologies for the next generation of cognitive surgical assistance systems. These systems could increase the safety of the operation through context-sensitive warnings and semi-autonomous robotic assistance or improve training of surgeons via data-driven feedback. In surgical workflow analysis up to 91% average precision has been reported for phase recognition on an open data single-center video dataset. In this work we investigated the generalizability of phase recognition algorithms in a multicenter setting including more difficult recognition tasks such as surgical action and surgical skill. METHODS: To achieve this goal, a dataset with 33 laparoscopic cholecystectomy videos from three surgical centers with a total operation time of 22 h was created. Labels included framewise annotation of seven surgical phases with 250 phase transitions, 5514 occurences of four surgical actions, 6980 occurences of 21 surgical instruments from seven instrument categories and 495 skill classifications in five skill dimensions. The dataset was used in the 2019 international Endoscopic Vision challenge, sub-challenge for surgical workflow and skill analysis. Here, 12 research teams trained and submitted their machine learning algorithms for recognition of phase, action, instrument and/or skill assessment. RESULTS: F1-scores were achieved for phase recognition between 23.9% and 67.7% (n = 9 teams), for instrument presence detection between 38.5% and 63.8% (n = 8 teams), but for action recognition only between 21.8% and 23.3% (n = 5 teams). The average absolute error for skill assessment was 0.78 (n = 1 team). CONCLUSION: Surgical workflow and skill analysis are promising technologies to support the surgical team, but there is still room for improvement, as shown by our comparison of machine learning algorithms. This novel HeiChole benchmark can be used for comparable evaluation and validation of future work. In future studies, it is of utmost importance to create more open, high-quality datasets in order to allow the development of artificial intelligence and cognitive robotics in surgery. Martin Wagner 0001, Beat P. Müller-Stich, Anna Kisilenko, Patrick Heger, Lars Mündermann, David M. Lubotsky, Tornike Davitashvili, Manuela Capek, Annika Reinke, Carissa Reid, Tong Yu 0009, Armine Vardazaryan, Chinedu Innocent Nwoye, Nicolas Padoy, Eungjoo Lee 0001, Constantin Disch, Hans Meine, Tong Xia, Fucang Jia, Satoshi Kondo, Wolfgang Reiter, Yueming Jin, Yonghao Long 0001, Meirui Jiang, Qi Dou 0001, Pheng-Ann Heng, Isabell Twick, Kadir Kirtaç, Enes Hosgor, Jon Lindström Bolmgren, Michael Stenzel, Björn von Siemens, Zhenxiao Ge, Haiming Sun, Di Xie, Mengqi Guo, Daochang Liu, Hannes Kenngott, Felix Nickel, Moritz von Frankenberg, Franziska Mathis-Ullrich, Annette Kopp-Schneider, Lena Maier-Hein, Stefanie Speidel, Sebastian Bodenstedt |
Medical Image Anal. | 28 |
| 2023 | Triplet attention and dual-pool contrastive learning for clinic-driven multi-label medical image classification
Yuhan Zhang 0001, Luyang Luo, Qi Dou 0001, Pheng-Ann Heng |
Medical Image Anal. | 3 |
| 2023 | IOP-FL: Inside-Outside Personalization for Federated Medical Image SegmentationabstractFederated learning (FL) allows multiple medical institutions to collaboratively learn a global model without centralizing client data. It is difficult, if possible at all, for such a global model to commonly achieve optimal performance for each individual client, due to the heterogeneity of medical images from various scanners and patient demographics. This problem becomes even more significant when deploying the global model to unseen clients outside the FL with unseen distributions not presented during federated training. To optimize the prediction accuracy of each individual client for medical imaging tasks, we propose a novel unified framework for both Inside and Outside model Personalization in FL (IOP-FL). Our inside personalization uses a lightweight gradient-based approach that exploits the local adapted model for each client, by accumulating both the global gradients for common knowledge and the local gradients for client-specific optimization. Moreover, and importantly, the obtained local personalized models and the global model can form a diverse and informative routing space to personalize an adapted model for outside FL clients. Hence, we design a new test-time routing scheme using the consistency loss with a shape constraint to dynamically incorporate the models, given the distribution information conveyed by the test data. Our extensive experimental results on two medical image segmentation tasks present significant improvements over SOTA methods on both inside and outside personalization, demonstrating the potential of our IOP-FL scheme for clinical practice. Code is available at https://github.com/med-air/IOP-FL. Meirui Jiang, Hongzheng Yang, Cheng Chen 0013, Qi Dou 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2023 | Joint Optimization of Class-Specific Training- and Test-Time Data Augmentation in SegmentationabstractThis paper presents an effective and general data augmentation framework for medical image segmentation. We adopt a computationally efficient and data-efficient gradient-based meta-learning scheme to explicitly align the distribution of training and validation data which is used as a proxy for unseen test data. We improve the current data augmentation strategies with two core designs. First, we learn class-specific training-time data augmentation (TRA) effectively increasing the heterogeneity within the training subsets and tackling the class imbalance common in segmentation. Second, we jointly optimize TRA and test-time data augmentation (TEA), which are closely connected as both aim to align the training and test data distribution but were so far considered separately in previous works. We demonstrate the effectiveness of our method on four medical image segmentation tasks across different scenarios with two state-of-the-art segmentation models, DeepMedic and nnU-Net. Extensive experimentation shows that the proposed data augmentation framework can significantly and consistently improve the segmentation performance when compared to existing solutions. Code is publicly available at https://github.com/ZerojumpLine/JCSAugment. Zeju Li, Konstantinos Kamnitsas, Qi Dou 0001, Chen Qin, Ben Glocker |
IEEE Trans. Medical Imaging | 3 |
| 2022 | HarmoFL: Harmonizing Local and Global Drifts in Federated Learning on Heterogeneous Medical ImagesabstractMultiple medical institutions collaboratively training a model using federated learning (FL) has become a promising solution for maximizing the potential of data-driven models, yet the non-independent and identically distributed (non-iid) data in medical images is still an outstanding challenge in real-world practice. The feature heterogeneity caused by diverse scanners or protocols introduces a drift in the learning process, in both local (client) and global (server) optimizations, which harms the convergence as well as model performance. Many previous works have attempted to address the non-iid issue by tackling the drift locally or globally, but how to jointly solve the two essentially coupled drifts is still unclear. In this work, we concentrate on handling both local and global drifts and introduce a new harmonizing framework called HarmoFL. First, we propose to mitigate the local update drift by normalizing amplitudes of images transformed into the frequency domain to mimic a unified imaging setting, in order to generate a harmonized feature space across local clients. Second, based on harmonized features, we design a client weight perturbation guiding each local model to reach a flat optimum, where a neighborhood area of the local optimal solution has a uniformly low loss. Without any extra communication cost, the perturbation assists the global model to optimize towards a converged optimal solution by aggregating several local flat optima. We have theoretically analyzed the proposed method and empirically conducted extensive experiments on three medical image classification and segmentation tasks, showing that HarmoFL outperforms a set of recent state-of-the-art methods with promising convergence behavior. Code is available at: https://github.com/med-air/HarmoFL Meirui Jiang, Qi Dou 0001 |
AAAI | 3 |
| 2022 | Single-Domain Generalization in Medical Image Segmentation via Test-Time Adaptation from Shape DictionaryabstractDomain generalization typically requires data from multiple source domains for model learning. However, such strong assumption may not always hold in practice, especially in medical field where the data sharing is highly concerned and sometimes prohibitive due to privacy issue. This paper studies the important yet challenging single domain generalization problem, in which a model is learned under the worst-case scenario with only one source domain to directly generalize to different unseen target domains. We present a novel approach to address this problem in medical image segmentation, which extracts and integrates the semantic shape prior information of segmentation that are invariant across domains and can be well-captured even from single domain data to facilitate segmentation under distribution shifts. Besides, a test-time adaptation strategy with dual-consistency regularization is further devised to promote dynamic incorporation of these shape priors under each unseen domain to improve model generalizability. Extensive experiments on two medical image segmentation tasks demonstrate the consistent improvements of our method across various unseen domains, as well as its superiority over state-of-the-art approaches in addressing domain generalization under the worst-case scenario. Quande Liu, Cheng Chen 0013, Qi Dou 0001, Pheng-Ann Heng |
AAAI | 3 |
| 2022 | Transformer-empowered Multi-scale Contextual Matching and Aggregation for Multi-contrast MRI Super-resolutionabstractMagnetic resonance imaging (MRI) can present multicontrast images of the same anatomical structures, enabling multi-contrast super-resolution (SR) techniques. Compared with SR reconstruction using a single-contrast, multicontrast SR reconstruction is promising to yield SR images with higher quality by leveraging diverse yet complementary information embedded in different imaging modalities. However, existing methods still have two shortcomings: (1) they neglect that the multi-contrast features at different scales contain different anatomical details and hence lack effective mechanisms to match and fuse these features for better reconstruction; and (2) they are still deficient in capturing long-range dependencies, which are essential for the regions with complicated anatomical structures. We propose a novel network to comprehensively address these problems by developing a set of innovative Transformer-empowered multi-scale contextual matching and aggregation techniques; we call it McMRSR. Firstly, we tame transformers to model long-range dependencies in both reference and target images. Then, a new multi-scale contextual matching method is proposed to capture corresponding contexts from reference features at different scales. Furthermore, we introduce a multi-scale aggregation mechanism to gradually and interactively aggregate multi-scale matched features for reconstructing the target SR MR image. Extensive experiments demonstrate that our network outperforms state-of-the-art approaches and has great potential to be applied in clinical practice. Codes are available at https://github.com/XAIMI-Lab/McMRSR. Yapeng Tian, Qi Dou 0001, Chengyan Wang, Chenliang Xu, Harry Qin |
CVPR | 4 |
| 2022 | Tackling Long-Tailed Category Distribution Under Domain Shifts
Xiao Gu 0003, Yao Guo 0002, Zeju Li, Jianing Qiu, Qi Dou 0001, Yuxuan Liu 0013, Benny P. L. Lo, Guang-Zhong Yang |
ECCV (23) | 5 |
| 2022 | Sim-to-Real 6D Object Pose Estimation via Iterative Self-training for Robotic Bin Picking
Kai Chen 0028, Stephen James, Yichuan Li 0002, Yun-Hui Liu 0001, Pieter Abbeel, Qi Dou 0001 |
ECCV (39) | 7 |
| 2022 | Federated Learning from Only Unlabeled Data with Class-conditional-sharing Clients
Nan Lu 0001, Zhao Wang 0006, Gang Niu 0001, Qi Dou 0001, Masashi Sugiyama |
ICLR | 5 |
| 2022 | Towards Robust Part-aware Instance Segmentation for Industrial Bin PickingabstractIndustrial bin picking is a challenging task that requires accurate and robust segmentation of individual object instances. Particularly, industrial objects can have irregular shapes, that is, thin and concave, whereas in bin-picking scenarios, objects are often closely packed with strong occlusion. To address these challenges, we formulate a novel part-aware instance segmentation pipeline. The key idea is to decompose industrial objects into correlated approximate convex parts and enhance the object-level segmentation with part-level segmentation. We design a part-aware network to predict part masks and part-to-part offsets, followed by a part aggregation module to assemble the recognized parts into instances. To guide the network learning, we also propose an automatic label decoupling scheme to generate ground-truth part-level labels from instance-level labels. Finally, we contribute the first instance segmentation dataset, which contains a variety of industrial objects that are thin and have non-trivial shapes. Extensive experimental results on various industrial objects demonstrate that our method can achieve the best segmentation results compared with the state-of-the-art approaches. Yidan Feng, Biqi Yang, Xianzhi Li 0001, Chi-Wing Fu, Kai Chen 0028, Qi Dou 0001, Mingqiang Wei, Yun-Hui Liu 0001, Pheng-Ann Heng |
ICRA | 7 |
| 2022 | 3D Perception based Imitation Learning under Limited Demonstration for Laparoscope Control in Robotic SurgeryabstractAutomatic laparoscope motion control is fundamentally important for surgeons to efficiently perform operations. However, its traditional control methods based on tool tracking without considering information hidden in surgical scenes are not intelligent enough, while the latest supervised imitation learning (IL)-based methods require expensive sensor data and suffer from distribution mismatch issues caused by limited demonstrations. In this paper, we propose a novel Imitation Learning framework for Laparoscope Control (ILLC) with reinforcement learning (RL), which can efficiently learn the control policy from limited surgical video clips. Specially, we first extract surgical laparoscope trajectories from unlabeled videos as the demonstrations and reconstruct the corresponding surgical scenes. To fully learn from limited motion trajectory demonstrations, we propose Shape Preserving Trajectory Augmentation (SPTA) to augment these data, and build a simulation environment that supports parallel RGB-D rendering to reinforce the RL policy for interacting with the environment efficiently. With adversarial training for IL, we obtain the laparoscope control policy based on the generated rollouts and surgical demonstrations. Extensive experiments are conducted in unseen reconstructed surgical scenes, and our method outperforms the previous IL methods, which proves the feasibility of our unified learning-based framework for laparoscope control. Bin Li 0082, Ruofeng Wei, Bo Lu 0001, Chi Hang Yee, Chi-Fai Ng, Pheng-Ann Heng, Qi Dou 0001, Yun-Hui Liu 0001 |
ICRA | 8 |
| 2022 | Distilled Visual and Robot Kinematics Embeddings for Metric Depth Estimation in Monocular Scene ReconstructionabstractEstimating precise metric depth and scene reconstruction from monocular endoscopy is a fundamental task for surgical navigation in robotic surgery. However, traditional stereo matching adopts binocular images to perceive the depth information, which is difficult to transfer to the soft robotics-based surgical systems due to the use of monocular endoscopy. In this paper, we present a novel framework that combines robot kinematics and monocular endoscope images with deep unsupervised learning into a single network for metric depth estimation and then achieve 3D reconstruction of complex anatomy. Specifically, we first obtain the relative depth maps of surgical scenes by leveraging a brightness-aware monocular depth estimation method. Then, the corresponding endoscope poses are computed based on non-linear optimization of geo-metric and photometric reprojection residuals. Afterwards, we develop a Depth-driven Sliding Optimization (DDSO) algorithm to extract the scaling coefficient from kinematics and calculated poses offline. By coupling the metric scale and relative depth data, we form a robust ensemble that represents the metric and consistent depth. Next, we treat the ensemble as supervisory labels to train a metric depth estimation network for surgeries (i.e., MetricDepthS-Net) that distills the embeddings from the robot kinematics, endoscopic videos, and poses. With accurate metric depth estimation, we utilize a dense visual reconstruction method to recover the 3D structure of the whole surgical site. We have extensively evaluated the proposed framework on public SCARED and achieved comparable performance with stereo-based depth estimation methods. Our results demon-strate the feasibility of the proposed approach to recover the metric depth and 3D structure with monocular inputs. Ruofeng Wei, Bin Li 0082, Hangjie Mo, Fangxun Zhong, Yonghao Long 0001, Qi Dou 0001, Yun-Hui Liu 0001, Dong Sun 0001 |
IROS | 6 |
| 2022 | SESR: Self-Ensembling Sim-to-Real Instance Segmentation for Auto-Store Bin PickingabstractInstance segmentation is an important task for supporting robotic grasping in auto-store scenarios. Accurate segmentation usually relies on the quantity and quality of available annotated training data. However, it requires tremendous cost to obtain these labels. In this work, without requiring any human annotations on real data, our proposed self-ensembling sim-to-real network, namely SESR, is able to generate precise instance masks for a wide variety of supermarket goods. We design our SESR with a teacher model and a student model trained with a self-ensembling strategy. We adopt different levels of consistency to bridge the sim-to-real gap and boost the model generalization ability. Also, we compile an auto-store bin-picking dataset covering various goods. Extensive experiments on both unseen scenarios and unseen objects validate the effectiveness and superiority of our method over others, and the robot arm demonstrations further show that our segmentation results can support real-time auto-store bin picking. Biqi Yang, Kai Chen 0028, Yidan Feng, Xianzhi Li 0001, Qi Dou 0001, Chi-Wing Fu, Yun-Hui Liu 0001, Pheng-Ann Heng |
IROS | 7 |
| 2022 | Pseudo-label Guided Cross-video Pixel Contrast for Robotic Surgical Scene Segmentation with Limited AnnotationsabstractSurgical scene segmentation is fundamentally crucial for prompting cognitive assistance in robotic surgery. However, pixel-wise annotating surgical video in a frame-by-frame manner is expensive and time consuming. To greatly reduce the labeling burden, in this work, we study semi-supervised scene segmentation from robotic surgical video, which is practically essential yet rarely explored before. We consider a clinically suitable annotation situation under the equidistant sampling. We then propose PGV-CL, a novel pseudo-label guided cross-video contrast learning method to boost scene segmentation. It effectively leverages unlabeled data for a trusty and global model regularization that produces more discriminative feature representation. Concretely, for trusty representation learning, we propose to incorporate pseudo labels to instruct the pair selection, obtaining more reliable representation pairs for pixel contrast. Moreover, we expand the representation learning space from previous image-level to cross-video, which can capture the global semantics to benefit the learning process. We extensively evaluate our method on a public robotic surgery dataset EndoVis18 and a public cataract dataset CaDIS. Experimental results demonstrate the effectiveness of our method, consistently outperforming the state-of-the-art semi-supervised methods under different labeling ratios, and even surpassing fully supervised training on EndoVis18 with 10.1% labeling. Our code is available at https://github.com/yangyu-cuhk/PGV-CL. Yang Yu 0070, Yueming Jin, Guangyong Chen, Qi Dou 0001, Pheng-Ann Heng |
IROS | 5 |
| 2022 | Dynamic Bank Learning for Semi-supervised Federated Image Diagnosis with Class Imbalance
Meirui Jiang, Hongzheng Yang, Quande Liu, Pheng-Ann Heng, Qi Dou 0001 |
MICCAI (3) | 6 |
| 2022 | Flat-Aware Cross-Stage Distilled Framework for Imbalanced Medical Image Classification
Jinpeng Li 0004, Guangyong Chen, Hangyu Mao, Danruo Deng, Dong Li 0016, Jianye Hao, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (3) | 7 |
| 2022 | WavTrans: Synergizing Wavelet and Cross-Attention Transformer for Multi-contrast MRI Super-Resolution
Chengyan Wang, Qi Dou 0001, Harry Qin |
MICCAI (6) | 4 |
| 2022 | DuDoCAF: Dual-Domain Cross-Attention Fusion with Recurrent Transformer for Fast Multi-contrast MR Imaging
Bin Sui, Chengyan Wang, Yapeng Tian, Qi Dou 0001, Harry Qin |
MICCAI (6) | 5 |
| 2022 | Test-Time Adaptation with Calibration of Medical Image Classification Nets for Label Distribution Shift
Wenao Ma, Cheng Chen 0013, Harry Qin, Huimao Zhang, Qi Dou 0001 |
MICCAI (3) | 6 |
| 2022 | Neural Rendering for Stereo 3D Reconstruction of Deformable Tissues in Robotic Surgery
Yuehao Wang, Yonghao Long 0001, Siu Hin Fan, Qi Dou 0001 |
MICCAI (8) | 4 |
| 2022 | AutoLaparo: A New Dataset of Integrated Multi-tasks for Image-guided Surgical Automation in Laparoscopic Hysterectomy
Ziyi Wang 0006, Bo Lu 0001, Yonghao Long 0001, Fangxun Zhong, Tak Hong Cheung, Qi Dou 0001, Yun-Hui Liu 0001 |
MICCAI (8) | 6 |
| 2022 | Unsupervised feature disentanglement for video retrieval in minimally invasive surgery
Ziyi Wang 0006, Bo Lu 0001, Yueming Jin, Zerui Wang, Tak Hong Cheung, Pheng-Ann Heng, Qi Dou 0001, Yun-Hui Liu 0001 |
Medical Image Anal. | 8 |
| 2022 | Toward Image-Guided Automated Suture Grasping Under Complex Environments: A Learning-Enabled and Optimization-Based Holistic FrameworkabstractTo realize a higher-level autonomy of surgical knot tying in minimally invasive surgery (MIS), automated suture grasping, which bridges the suture stitching and looping procedures, is an important yet challenging task needs to be achieved. This paper presents a holistic framework with image-guided and automation techniques to robotize this operation even under complex environments. The whole task is initialized by suture segmentation, in which we propose a novel semi-supervised learning architecture featured with a suture-aware loss to pertinently learn its slender information using both annotated and unannotated data. With successful segmentation in stereo-camera, we develop a Sampling-based Sliding Pairing (SSP) algorithm to online optimize the suture’s 3D shape. By jointly studying the robotic configuration and the suture’s spatial characteristics, a target function is introduced to find the optimal grasping pose of the surgical tool with Remote Center of Motion (RCM) constraints. To compensate for inherent errors and practical uncertainties, a unified grasping strategy with a novel vision-based mechanism is introduced to autonomously accomplish this grasping task. Our framework is extensively evaluated from learning-based segmentation, 3D reconstruction, and image-guided grasping on the da Vinci Research Kit (dVRK) platform, where we achieve high performances and successful rates in perceptions and robotic manipulations. These results prove the feasibility of our approach in automating the suture grasping task, and this work fills the gap between automated surgical stitching and looping, stepping towards a higher-level of task autonomy in surgical knot tying. Note to Practitioners—This paper aims to automate the suture grasping task in surgical knot tying by leveraging stereo visual guidance. To effectively robotize this procedure, it requires multidisciplinary knowledge to achieve suture segmentation, 3D shape reconstruction, and reliable automated grasping, while there are no existing works tackling this procedure especially using robots with RCM kinematics constraints and under complex environments. In this article, we propose a learning-driven method along with a 3D shape optimizer, which can conduct the suture segmentation and output its accurate spatial coordinates, serving as guidance for automated grasping operation. Apart from this, we introduce a unified function to optimize the grasping pose, and a vision-based grasping strategy is also proposed to intelligently complete this task. The experiments extensively validate the feasibility of our framework for automated suture grasp, and its successful completion can serve as a basis for the following looping manipulation, hence filling a step gap in robot-assisted knot tying. This framework can be also encapsulated into the medical robotic system, and by simply indicating (e.g. mouse click) the rough position of the suture’s tip in one camera frame, the overall framework can be initialized and further accomplish the suture grasping task, which further prompts a full autonomy of surgical knot tying in the near future. Bo Lu 0001, Bin Li 0082, Wei Chen 0068, Yueming Jin, Qi Dou 0001, Pheng-Ann Heng, Yun-Hui Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2022 | Learning With Privileged Multimodal Knowledge for Unimodal SegmentationabstractMultimodal learning usually requires a complete set of modalities during inference to maintain performance. Although training data can be well-prepared with high-quality multiple modalities, in many cases of clinical practice, only one modality can be acquired and important clinical evaluations have to be made based on the limited single modality information. In this work, we propose a privileged knowledge learning framework with the 'Teacher-Student' architecture, in which the complete multimodal knowledge that is only available in the training data (called privileged information) is transferred from a multimodal teacher network to a unimodal student network, via both a pixel-level and an image-level distillation scheme. Specifically, for the pixel-level distillation, we introduce a regularized knowledge distillation loss which encourages the student to mimic the teacher's softened outputs in a pixel-wise manner and incorporates a regularization factor to reduce the effect of incorrect predictions from the teacher. For the image-level distillation, we propose a contrastive knowledge distillation loss which encodes image-level structured information to enrich the knowledge encoding in combination with the pixel-level distillation. We extensively evaluate our method on two different multi-class segmentation tasks, i.e., cardiac substructure segmentation and brain tumor segmentation. Experimental results on both tasks demonstrate that our privileged knowledge learning is effective in improving unimodal segmentation and outperforms previous methods. Cheng Chen 0013, Qi Dou 0001, Yueming Jin, Quande Liu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Robust Medical Image Classification From Noisy Labeled Data With Global and Local Representation Guided Co-TrainingabstractDeep neural networks have achieved remarkable success in a wide variety of natural image and medical image computing tasks. However, these achievements indispensably rely on accurately annotated training data. If encountering some noisy-labeled images, the network training procedure would suffer from difficulties, leading to a sub-optimal classifier. This problem is even more severe in the medical image analysis field, as the annotation quality of medical images heavily relies on the expertise and experience of annotators. In this paper, we propose a novel collaborative training paradigm with global and local representation learning for robust medical image classification from noisy-labeled data to combat the lack of high quality annotated medical data. Specifically, we employ the self-ensemble model with a noisy label filter to efficiently select the clean and noisy samples. Then, the clean samples are trained by a collaborative training strategy to eliminate the disturbance from imperfect labeled samples. Notably, we further design a novel global and local representation learning scheme to implicitly regularize the networks to utilize noisy samples in a self-supervised manner. We evaluated our proposed robust learning strategy on four public medical image classification datasets with three types of label noise, i.e., random noise, computer-generated label noise, and inter-observer variability noise. Our method outperforms other learning from noisy label methods and we also conducted extensive experiments to analyze each component of our method. Cheng Xue 0003, Lequan Yu, Pengfei Chen 0003, Qi Dou 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 4 |
| 2022 | DLTTA: Dynamic Learning Rate for Test-Time Adaptation on Cross-Domain Medical ImagesabstractTest-time adaptation (TTA) has increasingly been an important topic to efficiently tackle the cross-domain distribution shift at test time for medical images from different institutions. Previous TTA methods have a common limitation of using a fixed learning rate for all the test samples. Such a practice would be sub-optimal for TTA, because test data may arrive sequentially therefore the scale of distribution shift would change frequently. To address this problem, we propose a novel dynamic learning rate adjustment method for test-time adaptation, called DLTTA, which dynamically modulates the amount of weights update for each test image to account for the differences in their distribution shift. Specifically, our DLTTA is equipped with a memory bank based estimation scheme to effectively measure the discrepancy of a given test sample. Based on this estimated discrepancy, a dynamic learning rate adjustment strategy is then developed to achieve a suitable degree of adaptation for each test sample. The effectiveness and general applicability of our DLTTA is extensively demonstrated on three tasks including retinal optical coherence tomography (OCT) segmentation, histopathological image classification, and prostate 3D MRI segmentation. Our method achieves effective and fast test-time adaptation with consistent performance improvement over current state-of-the-art test-time adaptation methods. Code is available at https://github.com/med-air/DLTTA. Hongzheng Yang, Cheng Chen 0013, Meirui Jiang, Quande Liu, Jianfeng Cao, Pheng-Ann Heng, Qi Dou 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2021 | FedDG: Federated Domain Generalization on Medical Image Segmentation via Episodic Learning in Continuous Frequency SpaceabstractFederated learning allows distributed medical institutions to collaboratively learn a shared prediction model with privacy protection. While at clinical deployment, the models trained in federated learning can still suffer from performance drop when applied to completely unseen hospitals outside the federation. In this paper, we point out and solve a novel problem setting of federated domain generalization (FedDG), which aims to learn a federated model from multiple distributed source domains such that it can directly generalize to unseen target domains. We present a novel approach, named as Episodic Learning in Continuous Frequency Space (ELCFS), for this problem by enabling each client to exploit multi-source data distributions under the challenging constraint of data decentralization. Our approach transmits the distribution information across clients in a privacy-protecting way through an effective continuous frequency space interpolation mechanism. With the transferred multi-source distributions, we further carefully design a boundary-oriented episodic learning paradigm to expose the local learning to domain distribution shifts and particularly meet the challenges of model generalization in medical image segmentation scenario. The effectiveness of our method is demonstrated with superior performance over state-of-the-arts and in-depth ablation experiments on two medical image segmentation tasks. The code is available at https://github.com/liuquande/FedDG-ELCFS. Quande Liu, Cheng Chen 0013, Harry Qin, Qi Dou 0001, Pheng-Ann Heng |
CVPR | 4 |
| 2021 | SGPA: Structure-Guided Prior Adaptation for Category-Level 6D Object Pose EstimationabstractCategory-level 6D object pose estimation aims to predict the position and orientation for unseen objects, which plays a pillar role in many scenarios such as robotics and augmented reality. The significant intra-class variation is the bottleneck challenge in this task yet remains unsolved so far. In this paper, we take advantage of category prior to overcome this problem by innovating a structure-guided prior adaptation scheme to accurately estimate 6D pose for individual objects. Different from existing prior based methods, given one object and its corresponding category prior, we propose to leverage their structure similarity to dynamically adapt the prior to the observed object. The prior adaptation intrinsically associates the adopted prior with different objects, from which we can accurately reconstruct the 3D canonical model of the specific object for pose estimation. To further enhance the structure characteristic of objects, we extract low-rank structure points from the dense object point cloud, therefore more efficiently incorporating sparse structural information during prior adaptation. Extensive experiments on CAMERA25 and REAL275 benchmarks demonstrate significant performance improvement. Project homepage: https://www.cse.cuhk.edu.hk/˜kaichen/projects/sgpa/sgpa.html. Kai Chen 0028, Qi Dou 0001 |
ICCV | 2 |
| 2021 | FedBN: Federated Learning on Non-IID Features via Local Batch Normalization
Meirui Jiang, Michael Kamp, Qi Dou 0001 |
ICLR | 5 |
| 2021 | Data-driven Holistic Framework for Automated Laparoscope Optimal View Control with Learning-based Depth PerceptionabstractLaparoscopic Field of View (FOV) control is one of the most fundamental and important components in Minimally Invasive Surgery (MIS), nevertheless the traditional manual holding paradigm may easily bring fatigue to surgical assistants, and misunderstanding between surgeons also hinders assistants to provide a high-quality FOV. Targeting this problem, we here present a data-driven framework to realize an automated laparoscopic optimal FOV control. To achieve this goal, we offline learn a motion strategy of laparoscope relative to the surgeon’s hand-held surgical tool from our in-house surgical videos, developing our control domain knowledge and an optimal view generator. To adjust the laparoscope online, we first adopt a learning-based method to segment the two-dimensional (2D) position of the surgical tool, and further leverage this outcome to obtain its scale-aware depth from dense depth estimation results calculated by our novel unsupervised RoboDepth model only with the monocular camera feedback, hence in return fusing the above real-time 3D position into our control loop. To eliminate the misorientation of FOV caused by Remote Center of Motion (RCM) constraints when moving the laparoscope, we propose a novel rotation constraint using an affine map to minimize the visual warping problem, and a null-space controller is also embedded into the framework to optimize all types of errors in a unified and decoupled manner. Experiments are conducted using Universal Robot (UR) and Karl Storz Laparoscope/Instruments, which prove the feasibility of our domain knowledge and learning enabled framework for automated camera control. Bin Li 0082, Bo Lu 0001, Yiang Lu, Qi Dou 0001, Yun-Hui Liu 0001 |
ICRA | 4 |
| 2021 | Relational Graph Learning on Visual and Kinematics Embeddings for Accurate Gesture Recognition in Robotic SurgeryabstractAutomatic surgical gesture recognition is fundamentally important to enable intelligent cognitive assistance in robotic surgery. With recent advancement in robot-assisted minimally invasive surgery, rich information including surgical videos and robotic kinematics can be recorded, which provide complementary knowledge for understanding surgical gestures. However, existing methods either solely adopt uni-modal data or directly concatenate multi-modal representations, which can not sufficiently exploit the informative correlations inherent in visual and kinematics data to boost gesture recognition accuracies. In this regard, we propose a novel online approach of multi-modal relational graph network (i.e., MRG-Net) to dynamically integrate visual and kinematics information through interactive message propagation in the latent feature space. In specific, we first extract embeddings from video and kinematics sequences with temporal convolutional networks and LSTM units. Next, we identify multi-relations in these multi-modal embeddings and leverage them through a hierarchical relational graph learning module. The effectiveness of our method is demonstrated with state-of-the-art results on the public JIGSAWS dataset, outperforming current uni-modal and multi-modal methods on both suturing and knot typing tasks. Furthermore, we validated our method on in-house visual-kinematics datasets collected with da Vinci Research Kit (dVRK) platforms in two centers, with consistent promising performance achieved. Our code and data are released at: https://www.cse.cuhk.edu.hk/~yhlong/mrgnet.html. Yonghao Long 0001, Jie Ying Wu, Bo Lu 0001, Yueming Jin, Mathias Unberath, Yun-Hui Liu 0001, Pheng-Ann Heng, Qi Dou 0001 |
ICRA | 8 |
| 2021 | One to Many: Adaptive Instrument Segmentation via Meta Learning and Dynamic Online Adaptation in Robotic Surgical VideoabstractSurgical instrument segmentation in robot-assisted surgery (RAS) - especially that using learning-based models - relies on the assumption that training and testing videos are sampled from the same domain. However, it is impractical and expensive to collect and annotate sufficient data from every new domain. To greatly increase the label efficiency, we explore a new problem, i.e., adaptive instrument segmentation, which is to effectively adapt one source model to new robotic surgical videos from multiple target domains, only given the annotated instruments in the first frame. We propose MDAL, a meta-learning based dynamic online adaptive learning scheme with a two-stage framework to fast adapt the model parameters on the first frame and partial subsequent frames while predicting the results. MDAL learns the general knowledge of instruments and the fast adaptation ability through the video-specific meta-learning paradigm. The added gradient gate excludes the noisy supervision from pseudo masks for dynamic online adaptation on target videos. We demonstrate empirically that MDAL outperforms other state-of-the-art methods on two datasets (including a real-world RAS dataset). The promising performance on ex-vivo scenes also benefits the downstream tasks such as robot-assisted suturing and camera control. Yueming Jin, Bo Lu 0001, Chi-Fai Ng, Qi Dou 0001, Yun-Hui Liu 0001, Pheng-Ann Heng |
ICRA | 5 |
| 2021 | Accurate Grid Keypoint Learning for Efficient Video PredictionabstractVideo prediction methods generally consume substantial computing resources in training and deployment, among which keypoint-based approaches show promising improvement in efficiency by simplifying dense image prediction to light keypoint prediction. However, keypoint locations are often modeled only as continuous coordinates, so noise from semantically insignificant deviations in videos easily disrupt learning stability, leading to inaccurate keypoint modeling. In this paper, we design a new grid keypoint learning framework, aiming at a robust and explainable intermediate keypoint representation for long-term efficient video prediction. We have two major technical contributions. First, we detect keypoints by jumping among candidate locations in our raised grid space and formulate a condensation loss to encourage meaningful keypoints with strong representative capability. Second, we introduce a 2D binary map to represent the detected grid keypoints and then suggest propagating keypoint locations with stochasticity by selecting entries in the discrete grid space, thus preserving the spatial structure of keypoints in the long-term horizon for better future frame generation. Extensive experiments verify that our method outperforms the state-of-the-art stochastic video prediction methods while saves more than 98% of computing resources. We also demonstrate our method on a robotic-assisted surgery dataset with promising results. Our code is available at https://github.com/xjgaocs/Grid-Keypoint-Learning. Yueming Jin, Qi Dou 0001, Chi-Wing Fu, Pheng-Ann Heng |
IROS | 3 |
| 2021 | Domain Adaptive Robotic Gesture Recognition with Unsupervised Kinematic-Visual Data AlignmentabstractAutomated surgical gesture recognition is of great importance in robot-assisted minimally invasive surgery. However, existing methods assume that training and testing data are from the same domain, which suffers from severe performance degradation when a domain gap exists, such as the simulator and real robot. In this paper, we propose a novel unsupervised domain adaptation framework which can simultaneously transfer multi-modality knowledge, i.e., both kinematic and visual data, from simulator to real robot. It remedies the domain gap with enhanced transferable features by using temporal cues in videos, and inherent correlations in multi-modal towards recognizing gesture. Specifically, we first propose a Motion Direction Oriented Kinematics feature alignment (MDO-K) to align kinematics, which exploits temporal continuity to transfer motion directions with smaller gap rather than position values, relieving the adaptation burden. Moreover, we propose a Kinematic and Visual Relation Attention (KV-Relation-ATT) to transfer the co-occurrence signals of kinematics and vision. Such features attended by correlation similarity are more informative for enhancing domain-irreverent of the model. Two feature alignment strategies benefit the model mutually during the end-to-end learning process. We extensively evaluate our method for gesture recognition using DESK dataset with peg transfer procedure. Results show that our approach recovers the performance with great improvement gains, up to 12.91% in Accuracy and 20.16% in F1score without using any annotations in real robot. Xueying Shi, Yueming Jin, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
IROS | 3 |
| 2021 | Category-Level 6D Object Pose Estimation via Cascaded Relation and Recurrent Reconstruction NetworksabstractCategory-level 6D pose estimation, aiming to predict the location and orientation of unseen object instances, is fundamental to many scenarios such as robotic manipulation and augmented reality, yet still remains unsolved. Precisely recovering instance 3D model in the canonical space and accurately matching it with the observation is an essential point when estimating 6D pose for unseen objects. In this paper, we achieve accurate category-level 6D pose estimation via cascaded relation and recurrent reconstruction networks. Specifically, a novel cascaded relation network is dedicated for advanced representation learning to explore the complex and informative relations among instance RGB image, instance point cloud and category shape prior. Furthermore, we design a recurrent reconstruction network for iterative residual refinement to progressively improve the reconstruction and correspondence estimations from coarse to fine. Finally, the instance 6D pose is obtained leveraging the estimated dense correspondences between the instance point cloud and the reconstructed 3D model in the canonical space. We have conducted extensive experiments on two well-acknowledged benchmarks of category-level 6D pose estimation, with significant performance improvement over existing approaches. On the representatively strict evaluation metrics of 3D75and 5°2cm, our method exceeds the latest state-of-the-art SPD [1] by 4.9% and 17.7% on the CAMERA25 dataset, and by 2.7% and 8.5% on the REAL275 dataset. Codes are avaliable at https://wangjiaze.cn/projects/6DPoseEstimation.html. Kai Chen 0028, Qi Dou 0001 |
IROS | 3 |
| 2021 | SurRoL: An Open-source Reinforcement Learning Centered and dVRK Compatible Platform for Surgical Robot LearningabstractAutonomous surgical execution relieves tedious routines and surgeon’s fatigue. Recent learning-based methods, especially reinforcement learning (RL) based methods, achieve promising performance for dexterous manipulation, which usually requires the simulation to collect data efficiently and reduce the hardware cost. The existing learning-based simulation platforms for medical robots suffer from limited scenarios and simplified physical interactions, which degrades the real-world performance of learned policies. In this work, we designed SurRoL, an RL-centered simulation platform for surgical robot learning compatible with the da Vinci Research Kit (dVRK). The designed SurRoL integrates a user-friendly RL library for algorithm development and a real-time physics engine, which is able to support more PSM/ECM scenarios and more realistic physical interactions. Ten learning-based surgical tasks are built in the platform, which are common in the real autonomous surgical execution. We evaluate SurRoL using RL algorithms in simulation, provide in-depth analysis, deploy the trained policies on the real dVRK, and show that our SurRoL achieves better transferability in the real world. Bin Li 0082, Bo Lu 0001, Yun-Hui Liu 0001, Qi Dou 0001, Pheng-Ann Heng |
IROS | 5 |
| 2021 | Source-Free Domain Adaptive Fundus Image Segmentation with Denoised Pseudo-Labeling
Cheng Chen 0013, Quande Liu, Yueming Jin, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (5) | 4 |
| 2021 | Trans-SVNet: Accurate Phase Recognition from Surgical Videos via Hybrid Embedding Aggregation Transformer
Yueming Jin, Yonghao Long 0001, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (4) | 4 |
| 2021 | Federated Semi-supervised Medical Image Classification via Inter-client Relation Matching
Quande Liu, Hongzheng Yang, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (3) | 3 |
| 2021 | E-DSSR: Efficient Dynamic Surgical Scene Reconstruction with Transformer-Based Stereoscopic Depth Perception
Yonghao Long 0001, Zhaoshuo Li, Chi Hang Yee, Chi-Fai Ng, Russell H. Taylor, Mathias Unberath, Qi Dou 0001 |
MICCAI (4) | 7 |
| 2021 | Semi-supervised learning with progressive unlabeled data excavation for label-efficient surgical workflow recognition
Xueying Shi, Yueming Jin, Qi Dou 0001, Pheng-Ann Heng |
Medical Image Anal. | 3 |
| 2021 | Anchor-guided online meta adaptation for fast one-Shot instrument segmentation from robotic surgical videos
Yueming Jin, Bo Lu 0001, Chi-Fai Ng, Yun-Hui Liu 0001, Qi Dou 0001, Pheng-Ann Heng |
Medical Image Anal. | 7 |
| 2021 | 3-D RoI-Aware U-Net for Accurate and Efficient Colorectal Tumor SegmentationabstractSegmentation of colorectal cancerous regions from 3-D magnetic resonance (MR) images is a crucial procedure for radiotherapy. Automatic delineation from 3-D whole volumes is in urgent demand yet very challenging. Drawbacks of existing deep-learning-based methods for this task are two-fold: 1) extensive graphics processing unit (GPU) memory footprint of 3-D tensor limits the trainable volume size, shrinks effective receptive field, and therefore, degrades speed and segmentation performance and 2) in-region segmentation methods supported by region-of-interest (RoI) detection are either blind to global contexts, detail richness compromising, or too expensive for 3-D tasks. To tackle these drawbacks, we propose a novel encoder-decoder-based framework for 3-D whole volume segmentation, referred to as 3-D RoI-aware U-Net (3-D RU-Net). 3-D RU-Net fully utilizes the global contexts covering large effective receptive fields. Specifically, the proposed model consists of a global image encoder for global understanding-based RoI localization, and a local region decoder that operates on pyramid-shaped in-region global features, which is GPU memory efficient and thereby enables training and prediction with large 3-D whole volumes. To facilitate the global-to-local learning procedure and enhance contour detail richness, we designed a dice-based multitask hybrid loss function. The efficiency of the proposed framework enables an extensive model ensemble for further performance gain at acceptable extra computational costs. Over a dataset of 64 T2-weighted MR images, the experimental results of four-fold cross-validation show that our method achieved 75.5% dice similarity coefficient (DSC) in 0.61 s per volume on a GPU, which significantly outperforms competing methods in terms of accuracy and efficiency. The code is publicly available. Yi-Jie Huang, Qi Dou 0001, Zi-Xian Wang, Li-Zhi Liu, Chao-Feng Li, Lisheng Wang, Hao Chen 0011, Rui-Hua Xu |
IEEE Trans. Cybern. | 2 |
| 2021 | Self-Ensembling Co-Training Framework for Semi-Supervised COVID-19 CT SegmentationabstractThe coronavirus disease 2019 (COVID-19) has become a severe worldwide health emergency and is spreading at a rapid rate. Segmentation of COVID lesions from computed tomography (CT) scans is of great importance for supervising disease progression and further clinical treatment. As labeling COVID-19 CT scans is labor-intensive and time-consuming, it is essential to develop a segmentation method based on limited labeled data to conduct this task. In this paper, we propose a self-ensembled co-training framework, which is trained by limited labeled data and large-scale unlabeled data, to automatically extract COVID lesions from CT scans. Specifically, to enrich the diversity of unsupervised information, we build a co-training framework consisting of two collaborative models, in which the two models teach each other during training by using their respective predicted pseudo-labels of unlabeled data. Moreover, to alleviate the adverse impacts of noisy pseudo-labels for each model, we propose a self-ensembling strategy to perform consistency regularization for the up-to-date predictions of unlabeled data, in which the predictions of unlabeled data are gradually ensembled via moving average at the end of every training epoch. We evaluate our framework on a COVID-19 dataset containing 103 CT scans. Experimental results show that our proposed method achieves better performance in the case of only 4 labeled CT scans compared to the state-of-the-art semi-supervised segmentation networks. Caizi Li, Qi Dou 0001, Fan Lin, Kebao Zhang, Zuxin Feng, Weixin Si, Xuesong Deng, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | Temporal Memory Relation Network for Workflow Recognition From Surgical VideoabstractAutomatic surgical workflow recognition is a key component for developing context-aware computer-assisted systems in the operating theatre. Previous works either jointly modeled the spatial features with short fixed-range temporal information, or separately learned visual and long temporal cues. In this paper, we propose a novel end-to-end temporal memory relation network (TMRNet) for relating long-range and multi-scale temporal patterns to augment the present features. We establish a long-range memory bank to serve as a memory cell storing the rich supportive information. Through our designed temporal variation layer, the supportive cues are further enhanced by multi-scale temporal-only convolutions. To effectively incorporate the two types of cues without disturbing the joint learning of spatio-temporal features, we introduce a non-local bank operator to attentively relate the past to the present. In this regard, our TMRNet enables the current feature to view the long-range temporal dependency, as well as tolerate complex temporal extents. We have extensively validated our approach on two benchmark surgical video datasets, M2CAI challenge dataset and Cholec80 dataset. Experimental results demonstrate the outstanding performance of our method, consistently exceeding the state-of-the-art methods by a large margin (e.g., 67.0% v.s. 78.9% Jaccard on Cholec80 dataset). Yueming Jin, Yonghao Long 0001, Cheng Chen 0013, Qi Dou 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Multi-Site Infant Brain Segmentation Algorithms: The iSeg-2019 ChallengeabstractTo better understand early brain development in health and disorder, it is critical to accurately segment infant brain magnetic resonance (MR) images into white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF). Deep learning-based methods have achieved state-of-the-art performance; h owever, one of the major limitations is that the learning-based methods may suffer from the multi-site issue, that is, the models trained on a dataset from one site may not be applicable to the datasets acquired from other sites with different imaging protocols/scanners. To promote methodological development in the community, the iSeg-2019 challenge (http://iseg2019.web.unc.edu) provides a set of 6-month infant subjects from multiple sites with different protocols/scanners for the participating methods. T raining/validation subjects are from UNC (MAP) and testing subjects are from UNC/UMN (BCP), Stanford University, and Emory University. By the time of writing, there are 30 automatic segmentation methods participated in the iSeg-2019. In this article, 8 top-ranked methods were reviewed by detailing their pipelines/implementations, presenting experimental results, and evaluating performance across different sites in terms of whole brain, regions of interest, and gyral landmark curves. We further pointed out their limitations and possible directions for addressing the multi-site issue. We find that multi-site consistency is still an open issue. We hope that the multi-site dataset in the iSeg-2019 and this review article will attract more researchers to address the challenging and critical multi-site issue in practice. Yue Sun 0001, Kun Gao 0002, Zhengwang Wu, Xiaopeng Zong, Zhihao Lei, Ying Wei 0007, Jun Ma 0016, Xiaoping Yang 0001, Xue Feng 0001, Li Zhao 0001, Trung Le Phan, Jitae Shin, Tao Zhong 0002, Yu Zhang 0064, Lequan Yu, Caizi Li, Ramesh Basnet, M. Omair Ahmad, M. N. S. Swamy 0001, Wenao Ma, Qi Dou 0001, Toan Duc Bui, Camilo Bermudez, Bennett A. Landman, Ian H. Gotlib, Kathryn L. Humphreys, Sarah Shultz, Longchuan Li, Sijie Niu, Weili Lin, Valerie Jewells, Dinggang Shen, Gang Li 0001, Li Wang 0026 |
IEEE Trans. Medical Imaging | 22 |
| 2020 | Harmonizing Transferability and Discriminability for Adapting Object DetectorsabstractRecent advances in adaptive object detection have achieved compelling results in virtue of adversarial feature adaptation to mitigate the distributional shifts along the detection pipeline. Whilst adversarial adaptation significantly enhances the transferability of feature representations, the feature discriminability of object detectors remains less investigated. Moreover, transferability and discriminability may come at a contradiction in adversarial adaptation given the complex combinations of objects and the differentiated scene layouts between domains. In this paper, we propose a Hierarchical Transferability Calibration Network (HTCN) that hierarchically (local-region/image/instance) calibrates the transferability of feature representations for harmonizing transferability and discriminability. The proposed model consists of three components: (1) Importance Weighted Adversarial Training with input Interpolation (IWAT-I), which strengthens the global discriminability by re-weighting the interpolated image-level features; (2) Context-aware Instance-Level Alignment (CILA) module, which enhances the local discriminability by capturing the underlying complementary effect between the instance-level feature and the global context information for the instance-level feature alignment; (3) local feature masks that calibrate the local transferability to provide semantic guidance for the following discriminative pattern alignment. Experimental results show that HTCN significantly outperforms the state-of-the-art methods on benchmark datasets. Chaoqi Chen, Zebiao Zheng, Xinghao Ding, Yue Huang 0001, Qi Dou 0001 |
CVPR | 5 |
| 2020 | Automatic Gesture Recognition in Robot-assisted Surgery with Reinforcement Learning and Tree SearchabstractAutomatic surgical gesture recognition is fundamental for improving intelligence in robot-assisted surgery, such as conducting complicated tasks of surgery surveillance and skill evaluation. However, current methods treat each frame individually and produce the outcomes without effective consideration on future information. In this paper, we propose a framework based on reinforcement learning and tree search for joint surgical gesture segmentation and classification. An agent is trained to segment and classify the surgical video in a human-like manner whose direct decisions are re-considered by tree search appropriately. Our proposed tree search algorithm unites the outputs from two designed neural networks, i.e., policy and value network. With the integration of complementary information from distinct models, our framework is able to achieve the better performance than baseline methods using either of the neural networks. For an overall evaluation, our developed approach consistently outperforms the existing methods on the suturing task of JIGSAWS dataset in terms of accuracy, edit score and F1 score. Our study highlights the utilization of tree search to refine actions in reinforcement learning framework for surgical robotic applications. Yueming Jin, Qi Dou 0001, Pheng-Ann Heng |
ICRA | 3 |
| 2020 | A Learning-Driven Framework with Spatial Optimization For Surgical Suture Thread Reconstruction and Autonomous Grasping Under Multiple Topologies and Environmental NoisesabstractSurgical knot tying is one of the most fundamental and important procedures in surgery, and a high-quality knot can significantly benefit the postoperative recovery of the patient. However, a longtime operation may easily cause fatigue to surgeons, especially during the tedious wound closure task. In this paper, we present a vision-based method to automate the suture thread grasping, which is a sub-task in surgical knot tying and an intermediate step between the stitching and looping manipulations. To achieve this goal, the acquisition of a suture's three-dimensional (3D) information is critical. Towards this objective, we adopt a transfer-learning strategy first to fine-tune a pre-trained model by learning the information from large legacy surgical data and images obtained by the onsite equipment. Thus, a robust suture segmentation can be achieved regardless of inherent environment noises. We further leverage a searching strategy with termination policies for a suture's sequence inference based on the analysis of multiple topologies. Exact results of the pixel-level sequence along a suture can be obtained, and they can be further applied for a 3D shape reconstruction using our optimized shortest path approach. The grasping point considering the suturing criterion can be ultimately acquired. Experiments regarding the suture 2D segmentation and ordering sequence inference under environmental noises were extensively evaluated. Results related to the automated grasping operation were demonstrated by simulations in V-REP and by robot experiments using Universal Robot (UR) together with the da Vinci Research Kit (dVRK) adopting our learning-driven framework. Bo Lu 0001, Wei Chen 0068, Yueming Jin, Qi Dou 0001, Henry K. Chu, Pheng-Ann Heng, Yun-Hui Liu 0001 |
IROS | 5 |
| 2020 | Shape-Aware Meta-learning for Generalizing Prostate MRI Segmentation to Unseen Domains
Quande Liu, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (2) | 2 |
| 2020 | Image-Level Harmonization of Multi-site Data Using Image-and-Spatial Transformer Networks
Robert Robinson, Qi Dou 0001, Daniel C. Castro, Konstantinos Kamnitsas, Marius de Groot, Ronald M. Summers, Daniel Rueckert, Ben Glocker |
MICCAI (7) | 2 |
| 2020 | Shape Mask Generator: Learning to Refine Shape Priors for Segmenting Overlapping Cervical Cytoplasms
Youyi Song, Lei Zhu 0003, Bai Ying Lei, Bin Sheng 0001, Qi Dou 0001, Harry Qin, Kup-Sze Choi |
MICCAI (4) | 5 |
| 2020 | Cascaded Robust Learning at Imperfect Labels for Chest X-ray Segmentation
Cheng Xue 0003, Xiaomeng Li 0001, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (6) | 4 |
| 2020 | Learning Motion Flows for Semi-supervised Instrument Segmentation from Robotic Surgical Video
Yueming Jin, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (3) | 4 |
| 2020 | Multi-task recurrent convolutional network with correlation loss for surgical video analysis
Yueming Jin, Huaxia Li, Qi Dou 0001, Hao Chen 0011, Harry Qin, Chi-Wing Fu, Pheng-Ann Heng |
Medical Image Anal. | 3 |
| 2020 | Weakly Supervised Deep Learning for Whole Slide Lung Cancer Image AnalysisabstractHistopathology image analysis serves as the gold standard for cancer diagnosis. Efficient and precise diagnosis is quite critical for the subsequent therapeutic treatment of patients. So far, computer-aided diagnosis has not been widely applied in pathological field yet as currently well-addressed tasks are only the tip of the iceberg. Whole slide image (WSI) classification is a quite challenging problem. First, the scarcity of annotations heavily impedes the pace of developing effective approaches. Pixelwise delineated annotations on WSIs are time consuming and tedious, which poses difficulties in building a large-scale training dataset. In addition, a variety of heterogeneous patterns of tumor existing in high magnification field are actually the major obstacle. Furthermore, a gigapixel scale WSI cannot be directly analyzed due to the immeasurable computational cost. How to design the weakly supervised learning methods to maximize the use of available WSI-level labels that can be readily obtained in clinical practice is quite appealing. To overcome these challenges, we present a weakly supervised approach in this article for fast and effective classification on the whole slide lung cancer images. Our method first takes advantage of a patch-based fully convolutional network (FCN) to retrieve discriminative blocks and provides representative deep features with high efficiency. Then, different context-aware block selection and feature aggregation strategies are explored to generate globally holistic WSI descriptor which is ultimately fed into a random forest (RF) classifier for the image-level prediction. To the best of our knowledge, this is the first study to exploit the potential of image-level labels along with some coarse annotations for weakly supervised learning. A large-scale lung cancer WSI dataset is constructed in this article for evaluation, which validates the effectiveness and feasibility of the proposed method. Extensive experiments demonstrate the superior performance of our method that surpasses the state-of-the-art approaches by a significant margin with an accuracy of 97.3%. In addition, our method also achieves the best performance on the public lung cancer WSIs dataset from The Cancer Genome Atlas (TCGA). We highlight that a small number of coarse annotations can contribute to further accuracy improvement. We believe that weakly supervised learning methods have great potential to assist pathologists in histology image diagnosis in the near future. Xi Wang 0013, Hao Chen 0011, Caixia Gan, Huangjing Lin, Qi Dou 0001, Efstratios Tsougenis, Qitao Huang, Muyan Cai, Pheng-Ann Heng |
IEEE Trans. Cybern. | 5 |
| 2020 | Contrastive Cross-Site Learning With Redesigned Net for COVID-19 CT ClassificationabstractThe pandemic of coronavirus disease 2019 (COVID-19) has lead to a global public health crisis spreading hundreds of countries. With the continuous growth of new infections, developing automated tools for COVID-19 identification with CT image is highly desired to assist the clinical diagnosis and reduce the tedious workload of image interpretation. To enlarge the datasets for developing machine learning methods, it is essentially helpful to aggregate the cases from different medical systems for learning robust and generalizable models. This paper proposes a novel joint learning framework to perform accurate COVID-19 identification by effectively learning with heterogeneous datasets with distribution discrepancy. We build a powerful backbone by redesigning the recently proposed COVID-Net in aspects of network architecture and learning strategy to improve the prediction accuracy and learning efficiency. On top of our improved backbone, we further explicitly tackle the cross-site domain shift by conducting separate feature normalization in latent space. Moreover, we propose to use a contrastive training objective to enhance the domain invariance of semantic embeddings for boosting the classification performance on each dataset. We develop and evaluate our method with two public large-scale COVID-19 diagnosis datasets made up of CT images. Extensive experiments show that our approach consistently improves the performanceson both datasets, outperforming the original COVID-Net trained on each dataset by 12.16% and 14.23% in AUC respectively, also exceeding existing state-of-the-art multi-site learning methods. Zhao Wang 0006, Quande Liu, Qi Dou 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Unsupervised Bidirectional Cross-Modality Adaptation via Deeply Synergistic Image and Feature Alignment for Medical Image SegmentationabstractUnsupervised domain adaptation has increasingly gained interest in medical image computing, aiming to tackle the performance degradation of deep neural networks when being deployed to unseen data with heterogeneous characteristics. In this work, we present a novel unsupervised domain adaptation framework, named as Synergistic Image and Feature Alignment (SIFA), to effectively adapt a segmentation network to an unlabeled target domain. Our proposed SIFA conducts synergistic alignment of domains from both image and feature perspectives. In particular, we simultaneously transform the appearance of images across domains and enhance domain-invariance of the extracted features by leveraging adversarial learning in multiple aspects and with a deeply supervised mechanism. The feature encoder is shared between both adaptive perspectives to leverage their mutual benefits via end-to-end learning. We have extensively evaluated our method with cardiac substructure segmentation and abdominal multi-organ segmentation for bidirectional cross-modality adaptation between MRI and CT images. Experimental results on two different tasks demonstrate that our SIFA method is effective in improving segmentation performance on unlabeled target images, and outperforms the state-of-the-art domain adaptation approaches by a large margin. Cheng Chen 0013, Qi Dou 0001, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Unpaired Multi-Modal Segmentation via Knowledge DistillationabstractMulti-modal learning is typically performed with network architectures containing modality-specific layers and shared layers, utilizing co-registered images of different modalities. We propose a novel learning scheme for unpaired cross-modality image segmentation, with a highly compact architecture achieving superior segmentation accuracy. In our method, we heavily reuse network parameters, by sharing all convolutional kernels across CT and MRI, and only employ modality-specific internal normalization layers which compute respective statistics. To effectively train such a highly compact model, we introduce a novel loss term inspired by knowledge distillation, by explicitly constraining the KL-divergence of our derived prediction distributions between modalities. We have extensively validated our approach on two multi-class segmentation problems: i) cardiac structure segmentation, and ii) abdominal organ segmentation. Different network settings, i.e., 2D dilated network and 3D U-net, are utilized to investigate our method's general efficacy. Experimental results on both tasks demonstrate that our novel multi-modal learning scheme consistently outperforms single-modal training and previous multi-modal approaches. Qi Dou 0001, Quande Liu, Pheng-Ann Heng, Ben Glocker |
IEEE Trans. Medical Imaging | 1 |
| 2020 | Multi-Task Deep Model With Margin Ranking Loss for Lung Nodule AnalysisabstractLung cancer is the leading cause of cancer deaths worldwide and early diagnosis of lung nodule is of great importance for therapeutic treatment and saving lives. Automated lung nodule analysis requires both accurate lung nodule benign-malignant classification and attribute score regression. However, this is quite challenging due to the considerable difficulty of lung nodule heterogeneity modeling and the limited discrimination capability on ambiguous cases. To solve these challenges, we propose a Multi-Task deep model with Margin Ranking loss (referred as MTMR-Net) for automated lung nodule analysis. Compared to existing methods which consider these two tasks separately, the relatedness between lung nodule classification and attribute score regression is explicitly explored in a cause-and-effect manner within our multi-task deep model, which can contribute to the performance gains of both tasks. The results of different tasks can be yielded simultaneously for assisting the radiologists in diagnosis interpretation. Furthermore, a Siamese network with a margin ranking loss is elaborately designed to enhance the discrimination capability on ambiguous nodule cases. To further explore the internal relationship between two tasks and validate the effectiveness of the proposed model, we use the recursive feature elimination method to iteratively rank the most malignancy-related features. We validate the efficacy of our method MTMR-Net on the public benchmark LIDC-IDRI dataset. Extensive experiments show that the diagnosis results with internal relationship explicitly explored in our model has met some similar patterns in clinical usage and also demonstrate that our approach can achieve competitive classification performance and more accurate scoring on attributes over the state-of-the-arts. Codes are publicly available at: https://github.com/CaptainWilliam/MTMR-NET. Qi Dou 0001, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2020 | MS-Net: Multi-Site Network for Improving Prostate Segmentation With Heterogeneous MRI DataabstractAutomated prostate segmentation in MRI is highly demanded for computer-assisted diagnosis. Recently, a variety of deep learning methods have achieved remarkable progress in this task, usually relying on large amounts of training data. Due to the nature of scarcity for medical images, it is important to effectively aggregate data from multiple sites for robust model training, to alleviate the insufficiency of single-site samples. However, the prostate MRIs from different sites present heterogeneity due to the differences in scanners and imaging protocols, raising challenges for effective ways of aggregating multi-site data for network training. In this paper, we propose a novel multi-site network (MS-Net) for improving prostate segmentation by learning robust representations, leveraging multiple sources of data. To compensate for the inter-site heterogeneity of different MRI datasets, we develop Domain-Specific Batch Normalization layers in the network backbone, enabling the network to estimate statistics and perform feature normalization for each site separately. Considering the difficulty of capturing the shared knowledge from multiple datasets, a novel learning paradigm, i.e., Multi-site-guided Knowledge Transfer, is proposed to enhance the kernels to extract more generic representations from multi-site data. Extensive experiments on three heterogeneous prostate MRI datasets demonstrate that our MS-Net improves the performance across all datasets consistently, and outperforms state-of-the-art methods for multi-site learning. Quande Liu, Qi Dou 0001, Lequan Yu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Semi-Supervised Medical Image Classification With Relation-Driven Self-Ensembling ModelabstractTraining deep neural networks usually requires a large amount of labeled data to obtain good performance. However, in medical image analysis, obtaining high-quality labels for the data is laborious and expensive, as accurately annotating medical images demands expertise knowledge of the clinicians. In this paper, we present a novel relation-driven semi-supervised framework for medical image classification. It is a consistency-based method which exploits the unlabeled data by encouraging the prediction consistency of given input under perturbations, and leverages a self-ensembling model to produce high-quality consistency targets for the unlabeled data. Considering that human diagnosis often refers to previous analogous cases to make reliable decisions, we introduce a novel sample relation consistency (SRC) paradigm to effectively exploit unlabeled data by modeling the relationship information among different samples. Superior to existing consistency-based methods which simply enforce consistency of individual predictions, our framework explicitly enforces the consistency of semantic relation among different samples under perturbations, encouraging the model to explore extra semantic information from unlabeled data. We have conducted extensive experiments to evaluate our method on two public benchmark medical image classification datasets, i.e., skin lesion diagnosis with ISIC 2018 challenge and thorax disease classification with ChestX-ray14. Our method outperforms many state-of-the-art semi-supervised learning methods on both single-label and multi-label image classification scenarios. Quande Liu, Lequan Yu, Luyang Luo, Qi Dou 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 4 |
| 2019 | Synergistic Image and Feature Adaptation: Towards Cross-Modality Domain Adaptation for Medical Image SegmentationabstractThis paper presents a novel unsupervised domain adaptation framework, called Synergistic Image and Feature Adaptation (SIFA), to effectively tackle the problem of domain shift. Domain adaptation has become an important and hot topic in recent studies on deep learning, aiming to recover performance degradation when applying the neural networks to new testing domains. Our proposed SIFA is an elegant learning diagram which presents synergistic fusion of adaptations from both image and feature perspectives. In particular, we simultaneously transform the appearance of images across domains and enhance domain-invariance of the extracted features towards the segmentation task. The feature encoder layers are shared by both perspectives to grasp their mutual benefits during the end-to-end learning procedure. Without using any annotation from the target domain, the learning of our unified model is guided by adversarial losses, with multiple discriminators employed from various aspects. We have extensively validated our method with a challenging application of crossmodality medical image segmentation of cardiac structures. Experimental results demonstrate that our SIFA model recovers the degraded performance from 17.2% to 73.0%, and outperforms the state-of-the-art methods by a significant margin. Cheng Chen 0013, Qi Dou 0001, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
AAAI | 2 |
| 2019 | Robust Multimodal Brain Tumor Segmentation via Feature Disentanglement and Gated Fusion
Cheng Chen 0013, Qi Dou 0001, Yueming Jin, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
MICCAI (3) | 2 |
| 2019 | Incorporating Temporal Prior from Motion Flow for Instrument Segmentation in Minimally Invasive Surgery Video
Yueming Jin, Keyun Cheng, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (5) | 3 |
| 2019 | Deep Angular Embedding and Feature Correlation Attention for Breast MRI Cancer Analysis
Luyang Luo, Hao Chen 0011, Xi Wang 0013, Qi Dou 0001, Huangjing Lin, Gongjie Li, Pheng-Ann Heng |
MICCAI (4) | 4 |
| 2019 | IRNet: Instance Relation Network for Overlapping Cervical Cell Segmentation
Yanning Zhou 0001, Hao Chen 0011, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (1) | 4 |
| 2019 | Improving RetinaNet for CT Lesion Detection with Dense Masks from Weak RECIST Labels
Martin Zlocha, Qi Dou 0001, Ben Glocker |
MICCAI (6) | 2 |
| 2019 | Domain Generalization via Model-Agnostic Learning of Semantic FeaturesabstractGeneralization capability to unseen domains is crucial for machine learning models when deploying to real-world conditions. We investigate the challenging problem of domain generalization, i.e., training a model on multi-domain source data such that it can directly generalize to target domains with unknown statistics. We adopt a model-agnostic learning paradigm with gradient-based meta-train and meta-test procedures to expose the optimization to domain shift. Further, we introduce two complementary losses which explicitly regularize the semantic structure of the feature space. Globally, we align a derived soft confusion matrix to preserve general knowledge of inter-class relationships. Locally, we promote domain-independent class-specific cohesion and separation of sample features with a metric-learning component. The effectiveness of our method is demonstrated with new state-of-the-art results on two common object recognition benchmarks. Our method also shows consistent improvement on a medical image segmentation task. Qi Dou 0001, Daniel C. Castro, Konstantinos Kamnitsas, Ben Glocker |
NeurIPS | 1 |
| 2019 | Webthetics: Quantifying webpage aesthetics with deep learning
Qi Dou 0001, Xianjun Sam Zheng, Tongfang Sun, Pheng-Ann Heng |
Int. J. Hum. Comput. Stud. | 1 |
| 2019 | MILD-Net: Minimal information loss dilated network for gland instance segmentation in colon histology images
Simon Graham, Hao Chen 0011, Jevgenij Gamper, Qi Dou 0001, Pheng-Ann Heng, David R. J. Snead, Yee-Wah Tsang, Nasir M. Rajpoot |
Medical Image Anal. | 4 |
| 2019 | Fast ScanNet: Fast and Dense Analysis of Multi-Gigapixel Whole-Slide Images for Cancer Metastasis DetectionabstractLymph node metastasis is one of the most important indicators in breast cancer diagnosis, that is traditionally observed under the microscope by pathologists. In recent years, with the dramatic advance of high-throughput scanning and deep learning technology, automatic analysis of histology from whole-slide images has received a wealth of interest in the field of medical image computing, which aims to alleviate pathologists' workload and simultaneously reduce misdiagnosis rate. However, the automatic detection of lymph node metastases from whole-slide images remains a key challenge because such images are typically very large, where they can often be multiple gigabytes in size. Also, the presence of hard mimics may result in a large number of false positives. In this paper, we propose a novel method with anchor layers for model conversion, which not only leverages the efficiency of fully convolutional architectures to meet the speed requirement in clinical practice but also densely scans the whole-slide image to achieve accurate predictions on both micro- and macro-metastases. Incorporating the strategies of asynchronous sample prefetching and hard negative mining, the network can be effectively trained. The efficacy of our method is corroborated on the benchmark dataset of 2016 Camelyon Grand Challenge. Our method achieved significant improvements in comparison with the state-of-the-art methods on tumor localization accuracy with a much faster speed and even surpassed human performance on both challenge tasks. Huangjing Lin, Hao Chen 0011, Simon Graham, Qi Dou 0001, Nasir M. Rajpoot, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 4 |
| 2018 | SFCN-OPI: Detection and Fine-Grained Classification of Nuclei Using Sibling FCN With Objectness Prior InteractionabstractCell nuclei detection and fine-grained classification have been fundamental yet challenging problems in histopathology image analysis. Due to the nuclei tiny size, significant inter-/intra-class variances, as well as the inferior image quality, previous automated methods would easily suffer from limited accuracy and robustness. In the meanwhile, existing approaches usually deal with these two tasks independently, which would neglect the close relatedness of them. In this paper, we present a novel method of sibling fully convolutional network with prior objectness interaction (called SFCN-OPI) to tackle the two tasks simultaneously and interactively using a unified end-to-end framework. Specifically, the sibling FCN branches share features in earlier layers while holding respective higher layers for specific tasks. More importantly, the detection branch outputs the objectness prior which dynamically interacts with the fine-grained classification sibling branch during the training and testing processes. With this mechanism, the fine-grained classification successfully focuses on regions with high confidence of nuclei existence and outputs the conditional probability, which in turn benefits the detection through back propagation. Extensive experiments on colon cancer histology images have validated the effectiveness of our proposed SFCN-OPI and our method has outperformed the state-of-the-art methods by a large margin. Yanning Zhou 0001, Qi Dou 0001, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
AAAI | 2 |
| 2018 | Unsupervised Cross-Modality Domain Adaptation of ConvNets for Biomedical Image Segmentations with Adversarial LossabstractConvolutional networks (ConvNets) have achieved great successes in various challenging vision tasks. However, the performance of ConvNets would degrade when encountering the domain shift. The domain adaptation is more significant while challenging in the field of biomedical image analysis, where cross-modality data have largely different distributions. Given that annotating the medical data is especially expensive, the supervised transfer learning approaches are not quite optimal. In this paper, we propose an unsupervised domain adaptation framework with adversarial learning for cross-modality biomedical image segmentations. Specifically, our model is based on a dilated fully convolutional network for pixel-wise prediction. Moreover, we build a plug-and-play domain adaptation module (DAM) to map the target input to features which are aligned with source domain feature space. A domain critic module (DCM) is set up for discriminating the feature space of both domains. We optimize the DAM and DCM via an adversarial loss without using any target domain label. Our proposed method is validated by adapting a ConvNet trained with MRI images to unpaired CT data for cardiac structures segmentations, and achieved very promising results. Qi Dou 0001, Cheng Ouyang, Cheng Chen 0013, Hao Chen 0011, Pheng-Ann Heng |
IJCAI | 1 |
| 2018 | ScanNet: A Fast and Dense Scanning Framework for Metastastic Breast Cancer Detection from Whole-Slide ImageabstractLymph node metastasis is one of the most significant diagnostic indicators in breast cancer, which is traditionally observed under the microscope by pathologists. In recent years, computerized histology diagnosis has become one of the most rapidly expanding directions in the field of medical image computing, which aims to alleviate pathologists' workload and simultaneously reduce misdiagnosis rate. However, automatic detection of lymph node metastases from whole slide images remains a challenging problem, due to the large-scale data with enormous resolutions and existence of hard mimics resulting in a large number of false positives. In this paper, we propose a novel framework by leveraging fully convolutional networks for efficient inference to meet the speed requirement for clinical practice, while reconstructing dense predictions under different offsets for ensuring accurate detection on both microand macro-metastases. Incorporating with the strategies of asynchronous sample prefetching and hard negative mining, the network can be effectively trained. Extensive experiments on the benchmark dataset of 2016 Camelyon Grand Challenge corroborated the efficacy of our method. Compared with the state-of-the-art methods, our method achieved superior performance with a faster speed on the tumor localization task and even surpassed human performance on the WSI classification task. Huangjing Lin, Hao Chen 0011, Qi Dou 0001, Liansheng Wang 0002, Harry Qin, Pheng-Ann Heng |
WACV | 3 |
| 2018 | 3D multi-scale FCN with random modality voxel dropout learning for Intervertebral Disc Localization and Segmentation from Multi-modality MR Images
Xiaomeng Li 0001, Qi Dou 0001, Hao Chen 0011, Chi-Wing Fu, Xiaojuan Qi 0001, Daniel L. Belavy, Gabriele Armbrecht, Dieter Felsenberg, Guoyan Zheng, Pheng-Ann Heng |
Medical Image Anal. | 2 |
| 2018 | SV-RCNet: Workflow Recognition From Surgical Videos Using Recurrent Convolutional NetworkabstractWe propose an analysis of surgical videos that is based on a novel recurrent convolutional network (SV-RCNet), specifically for automatic workflow recognition from surgical videos online, which is a key component for developing the context-aware computer-assisted intervention systems. Different from previous methods which harness visual and temporal information separately, the proposed SV-RCNet seamlessly integrates a convolutional neural network (CNN) and a recurrent neural network (RNN) to form a novel recurrent convolutional architecture in order to take full advantages of the complementary information of visual and temporal features learned from surgical videos. We effectively train the SV-RCNet in an end-to-end manner so that the visual representations and sequential dynamics can be jointly optimized in the learning process. In order to produce more discriminative spatio-temporal features, we exploit a deep residual network (ResNet) and a long short term memory (LSTM) network, to extract visual features and temporal dependencies, respectively, and integrate them into the SV-RCNet. Moreover, based on the phase transition-sensitive predictions from the SV-RCNet, we propose a simple yet effective inference scheme, namely the prior knowledge inference (PKI), by leveraging the natural characteristic of surgical video. Such a strategy further improves the consistency of results and largely boosts the recognition performance. Extensive experiments have been conducted with the MICCAI 2016 Modeling and Monitoring of Computer Assisted Interventions Workflow Challenge dataset and Cholec80 dataset to validate SV-RCNet. Our approach not only achieves superior performance on these two datasets but also outperforms the state-of-the-art methods by a significant margin. Yueming Jin, Qi Dou 0001, Hao Chen 0011, Lequan Yu, Harry Qin, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2018 | H-DenseUNet: Hybrid Densely Connected UNet for Liver and Tumor Segmentation From CT VolumesabstractLiver cancer is one of the leading causes of cancer death. To assist doctors in hepatocellular carcinoma diagnosis and treatment planning, an accurate and automatic liver and tumor segmentation method is highly demanded in the clinical practice. Recently, fully convolutional neural networks (FCNs), including 2-D and 3-D FCNs, serve as the backbone in many volumetric image segmentation. However, 2-D convolutions cannot fully leverage the spatial information along the third dimension while 3-D convolutions suffer from high computational cost and GPU memory consumption. To address these issues, we propose a novel hybrid densely connected UNet (H-DenseUNet), which consists of a 2-D DenseUNet for efficiently extracting intra-slice features and a 3-D counterpart for hierarchically aggregating volumetric contexts under the spirit of the auto-context algorithm for liver and tumor segmentation. We formulate the learning process of the H-DenseUNet in an end-to-end manner, where the intra-slice representations and inter-slice features can be jointly optimized through a hybrid feature fusion layer. We extensively evaluated our method on the data set of the MICCAI 2017 Liver Tumor Segmentation Challenge and 3DIRCADb data set. Our method outperformed other state-of-the-arts on the segmentation results of tumors and achieved very competitive performance for liver segmentation even with a single model. Xiaomeng Li 0001, Hao Chen 0011, Xiaojuan Qi 0001, Qi Dou 0001, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 4 |
| 2017 | Automated Pulmonary Nodule Detection via 3D ConvNets with Online Sample Filtering and Hybrid-Loss Residual Learning
Qi Dou 0001, Hao Chen 0011, Yueming Jin, Huangjing Lin, Harry Qin, Pheng-Ann Heng |
MICCAI (3) | 1 |
| 2017 | Automatic 3D Cardiovascular MR Segmentation with Densely-Connected Volumetric ConvNets
Lequan Yu, Jie-Zhi Cheng, Qi Dou 0001, Xin Yang 0009, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
MICCAI (2) | 3 |
| 2017 | DCAN: Deep contour-aware networks for object instance segmentation from histology images
Hao Chen 0011, Xiaojuan Qi 0001, Lequan Yu, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
Medical Image Anal. | 4 |
| 2017 | 3D deeply supervised network for automated segmentation of volumetric medical images
Qi Dou 0001, Lequan Yu, Hao Chen 0011, Yueming Jin, Xin Yang 0009, Harry Qin, Pheng-Ann Heng |
Medical Image Anal. | 1 |
| 2017 | Validation, comparison, and combination of algorithms for automatic detection of pulmonary nodules in computed tomography images: The LUNA16 challenge
Arnaud A. A. Setio, Alberto Traverso, Thomas de Bel, Moira S. N. Berens, Cas van den Bogaard, Piergiorgio Cerello, Hao Chen 0011, Qi Dou 0001, Maria Evelina Fantacci, Bram Geurts, Robbert van der Gugten, Pheng-Ann Heng, Bart Jansen 0001, Michael M. J. de Kaste, Valentin Kotov, Jack Yu-Hung Lin, Jeroen T. M. C. Manders, Alexander Sóñora-Mengana, Juan Carlos García-Naranjo, Evgenia Papavasileiou, Mathias Prokop |
Medical Image Anal. | 8 |
| 2017 | Evaluation and comparison of 3D intervertebral disc localization and segmentation methods for 3D T2 MR data: A grand challenge
Guoyan Zheng, Chengwen Chu, Daniel L. Belavy, Bulat Ibragimov, Robert Korez, Tomaz Vrtovec, Hugo Hutt, Richard M. Everson, Judith Meakin, Isabel Lopez Andrade, Ben Glocker, Hao Chen 0011, Qi Dou 0001, Pheng-Ann Heng, Chunliang Wang, Daniel Forsberg, Ales Neubert, Jurgen Fripp, Martin Urschler, Darko Stern, Maria Wimmer 0002 |
Medical Image Anal. | 13 |
| 2017 | Ultrasound Standard Plane Detection Using a Composite Neural Network FrameworkabstractUltrasound (US) imaging is a widely used screening tool for obstetric examination and diagnosis. Accurate acquisition of fetal standard planes with key anatomical structures is very crucial for substantial biometric measurement and diagnosis. However, the standard plane acquisition is a labor-intensive task and requires operator equipped with a thorough knowledge of fetal anatomy. Therefore, automatic approaches are highly demanded in clinical practice to alleviate the workload and boost the examination efficiency. The automatic detection of standard planes from US videos remains a challenging problem due to the high intraclass and low interclass variations of standard planes, and the relatively low image quality. Unlike previous studies which were specifically designed for individual anatomical standard planes, respectively, we present a general framework for the automatic identification of different standard planes from US videos. Distinct from conventional way that devises hand-crafted visual features for detection, our framework explores in- and between-plane feature learning with a novel composite framework of the convolutional and recurrent neural networks. To further address the issue of limited training data, a multitask learning framework is implemented to exploit common knowledge across detection tasks of distinctive standard planes for the augmentation of feature learning. Extensive experiments have been conducted on hundreds of US fetus videos to corroborate the better efficacy of the proposed framework on the difficult standard plane detection problem. Hao Chen 0011, Lingyun Wu, Qi Dou 0001, Harry Qin, Shengli Li 0001, Jie-Zhi Cheng, Dong Ni 0001, Pheng-Ann Heng |
IEEE Trans. Cybern. | 3 |
| 2017 | Integrating Online and Offline Three-Dimensional Deep Learning for Automated Polyp Detection in Colonoscopy VideosabstractAutomated polyp detection in colonoscopy videos has been demonstrated to be a promising way for colorectal cancer prevention and diagnosis. Traditional manual screening is time consuming, operator dependent, and error prone; hence, automated detection approach is highly demanded in clinical practice. However, automated polyp detection is very challenging due to high intraclass variations in polyp size, color, shape, and texture, and low interclass variations between polyps and hard mimics. In this paper, we propose a novel offline and online three-dimensional (3-D) deep learning integration framework by leveraging the 3-D fully convolutional network (3D-FCN) to tackle this challenging problem. Compared with the previous methods employing hand-crafted features or 2-D convolutional neural network, the 3D-FCN is capable of learning more representative spatio-temporal features from colonoscopy videos, and hence has more powerful discrimination capability. More importantly, we propose a novel online learning scheme to deal with the problem of limited training data by harnessing the specific information of an input video in the learning process. We integrate offline and online learning to effectively reduce the number of false positives generated by the offline network and further improve the detection performance. Extensive experiments on the dataset of MICCAI 2015 Challenge on Polyp Detection demonstrated the better performance of our method when compared with other competitors. Lequan Yu, Hao Chen 0011, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 3 |
| 2017 | Automated Melanoma Recognition in Dermoscopy Images via Very Deep Residual NetworksabstractAutomated melanoma recognition in dermoscopy images is a very challenging task due to the low contrast of skin lesions, the huge intraclass variation of melanomas, the high degree of visual similarity between melanoma and non-melanoma lesions, and the existence of many artifacts in the image. In order to meet these challenges, we propose a novel method for melanoma recognition by leveraging very deep convolutional neural networks (CNNs). Compared with existing methods employing either low-level hand-crafted features or CNNs with shallower architectures, our substantially deeper networks (more than 50 layers) can acquire richer and more discriminative features for more accurate recognition. To take full advantage of very deep networks, we propose a set of schemes to ensure effective training and learning under limited training data. First, we apply the residual learning to cope with the degradation and overfitting problems when a network goes deeper. This technique can ensure that our networks benefit from the performance gains achieved by increasing network depth. Then, we construct a fully convolutional residual network (FCRN) for accurate skin lesion segmentation, and further enhance its capability by incorporating a multi-scale contextual information integration scheme. Finally, we seamlessly integrate the proposed FCRN (for segmentation) and other very deep residual networks (for classification) to form a two-stage framework. This framework enables the classification network to extract more representative and specific features based on segmented results instead of the whole dermoscopy images, further alleviating the insufficiency of training data. The proposed framework is extensively evaluated on ISBI 2016 Skin Lesion Analysis Towards Melanoma Detection Challenge dataset. Experimental results demonstrate the significant performance gains of the proposed framework, ranking the first in classification and the second in segmentation among 25 teams and 28 teams, respectively. This study corroborates that very deep CNNs with effective training mechanisms can be employed to solve complicated medical image analysis tasks, even with limited training data. Lequan Yu, Hao Chen 0011, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2016 | Mitosis Detection in Breast Cancer Histology Images via Deep Cascaded NetworksabstractThe number of mitoses per tissue area gives an important aggressiveness indication of the invasive breast carcinoma.However, automatic mitosis detection in histology images remains a challenging problem. Traditional methods either employ hand-crafted features to discriminate mitoses from other cells or construct a pixel-wise classifier to label every pixel in a sliding window way. While the former suffers from the large shape variation of mitoses and the existence of many mimics with similar appearance, the slow speed of the later prohibits its use in clinical practice.In order to overcome these shortcomings, we propose a fast and accurate method to detect mitosis by designing a novel deep cascaded convolutional neural network, which is composed of two components. First, by leveraging the fully convolutional neural network, we propose a coarse retrieval model to identify and locate the candidates of mitosis while preserving a high sensitivity.Based on these candidates, a fine discrimination model utilizing knowledge transferred from cross-domain is developed to further single out mitoses from hard mimics.Our approach outperformed other methods by a large margin in 2014 ICPR MITOS-ATYPIA challenge in terms of detection accuracy. When compared with the state-of-the-art methods on the 2012 ICPR MITOSIS data (a smaller and less challenging dataset), our method achieved comparable or better results with a roughly 60 times faster speed. Hao Chen 0011, Qi Dou 0001, Xi Wang 0013, Harry Qin, Pheng-Ann Heng |
AAAI | 2 |
| 2016 | 3D Deeply Supervised Network for Automatic Liver Segmentation from CT Volumes
Qi Dou 0001, Hao Chen 0011, Yueming Jin, Lequan Yu, Harry Qin, Pheng-Ann Heng |
MICCAI (2) | 1 |
| 2016 | Automatic Detection of Cerebral Microbleeds From MR Images via 3D Convolutional Neural NetworksabstractCerebral microbleeds (CMBs) are small haemorrhages nearby blood vessels. They have been recognized as important diagnostic biomarkers for many cerebrovascular diseases and cognitive dysfunctions. In current clinical routine, CMBs are manually labelled by radiologists but this procedure is laborious, time-consuming, and error prone. In this paper, we propose a novel automatic method to detect CMBs from magnetic resonance (MR) images by exploiting the 3D convolutional neural network (CNN). Compared with previous methods that employed either low-level hand-crafted descriptors or 2D CNNs, our method can take full advantage of spatial contextual information in MR volumes to extract more representative high-level features for CMBs, and hence achieve a much better detection accuracy. To further improve the detection performance while reducing the computational cost, we propose a cascaded framework under 3D CNNs for the task of CMB detection. We first exploit a 3D fully convolutional network (FCN) strategy to retrieve the candidates with high probabilities of being CMBs, and then apply a well-trained 3D CNN discrimination model to distinguish CMBs from hard mimics. Compared with traditional sliding window strategy, the proposed 3D FCN strategy can remove massive redundant computations and dramatically speed up the detection process. We constructed a large dataset with 320 volumetric MR scans and performed extensive experiments to validate the proposed method, which achieved a high sensitivity of 93.16% with an average number of 2.74 false positives per subject, outperforming previous methods using low-level descriptors or 2D CNNs by a significant margin. The proposed method, in principle, can be adapted to other biomarker detection tasks from volumetric medical data. Qi Dou 0001, Hao Chen 0011, Lequan Yu, Lei Zhao 0003, Harry Qin, Defeng Wang, Vincent C. T. Mok, Lin Shi 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 1 |
| 2015 | Automatic Fetal Ultrasound Standard Plane Detection Using Knowledge Transferred Recurrent Neural Networks
Hao Chen 0011, Qi Dou 0001, Dong Ni 0001, Jie-Zhi Cheng, Harry Qin, Shengli Li 0001, Pheng-Ann Heng |
MICCAI (1) | 2 |