VLDB 2026 Research / reviewers in the wild / expert
Yueming Jin
dblp:183/6320
· DBLP profile ↗
77ranked-venue papers
5as first author
64since 2021 · last 2026
0000-0003-3775-3877ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 48 · 5 first-author · 39 since 2021Graphics, computer vision, multimedia, augmented reality and games · 34 · 1 first-author · 26 since 2021Artificial intelligence and machine learning · 24 · 22 since 2021Systems, architecture and hardware · 10 · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Unleashing the Power of Image-Tabular Self-Supervised Learning via Breaking Cross-Tabular BarriersabstractMulti-modal learning integrating medical images and tabular data has significantly advanced clinical decision-making in recent years. Self-Supervised Learning (SSL) has emerged as a powerful paradigm for pretraining these models on large-scale unlabeled image-tabular data, aiming to learn discriminative representations. However, existing SSL methods for image-tabular representation learning are often confined to specific data cohorts, mainly due to their rigid tabular modeling mechanisms when modeling heterogeneous tabular data. This inter-tabular barrier hinders the multi-modal SSL methods from effectively learning transferrable medical knowledge shared across diverse cohorts. In this paper, we propose a novel SSL framework, namely CITab, designed to learn powerful multi-modal feature representations in a cross-tabular manner. We design the tabular modeling mechanism from a semantic-awareness perspective by integrating column headers as semantic cues, which facilitates transferrable knowledge learning and the scalability in utilizing multiple data sources for pretraining. Additionally, we propose a prototype-guided mixture-of-linear layer (P-MoLin) module for tabular feature specialization, empowering the model to effectively handle the heterogeneity of tabular data and explore the underlying medical concepts. We conduct comprehensive evaluations on Alzheimer's disease diagnosis task across three publicly available data cohorts containing 4,461 subjects. Experimental results demonstrate that CITab outperforms state-of-the-art approaches, paving the way for effective and scalable cross-tabular multi-modal learning. Yibing Fu, Zhitao Zeng, Cheng Chen 0013, Yueming Jin |
AAAI | 5 |
| 2026 | Tackling Dual-stage Missing Modalities in Brain Tumor Segmentation via Robust Modality Reconstruction and Prompt-guided Modality AdaptationabstractAddressing missing modalities is a critical challenge in multimodal brain tumor segmentation. Most existing approaches merely handle modality-incomplete inputs during inference, assuming a full set of modalities for all training samples. However, this unrealistic assumption limits the usage of abundant modality-incomplete data commonly observed in clinical practice. In this paper, we explore a more practical task of tackling missing modalities during both training and inference. We propose a universal model featuring robust modality reconstruction and prompt-guided modality adaptation. Our mask-reconstruction pre-training enables robust modality-invariant representation learning, during which we design a novel distribution approximation method that supervises the reconstruction of absent modalities without requiring full-modal training data. Afterwards, when adapting our model to the segmentation task, we introduce the complete-then-distill (CTD) paradigm, which first estimates missing modalities in training samples from the available ones, and then distills the knowledge from the reconstructed full-modal representations to enhance learning from modality-incomplete data. Moreover, we propose prompt-guided modality adaptation to personalize a subset of model parameters during CTD, enabling the model to adapt to each distinct modality input scenario by using prompts with rich visual-textual information. Extensive experiments on two brain tumor segmentation benchmarks show our method consistently surpasses previous state-of-the-art approaches under dual-stage missing modality settings across various missing ratios. Cheng Chen 0013, Qing You Pang, Yibing Fu, Quanzheng Li, Carol Tang, Beng-Ti Ang, Yueming Jin |
AAAI | 8 |
| 2026 | Cellflow: Advancing pathological image augmentation from spatial views to temporal trajectories
Zeyu Liu 0013, Haoran Guo, Peng Zhang 0078, Chenbin Ma, Shangqing Lyu, Yunlu Feng, Yueming Jin, Dachun Zhao, Guanglei Zhang |
Medical Image Anal. | 11 |
| 2026 | Spatio-Temporal Representation Decoupling and Enhancement for Federated Instrument Segmentation in Surgical VideosabstractSurgical instrument segmentation under Federated Learning (FL) is a promising direction, which enables multiple surgical sites to collaboratively train the model without centralizing datasets. However, there exist very limited FL works in surgical data science, and FL methods for other modalities do not consider inherent characteristics in surgical domain: i) different scenarios show diverse anatomical backgrounds while highly similar instrument representation; ii) there exist surgical simulators which promote large-scale synthetic data generation with minimal efforts. In this paper, we propose a novel Personalized FL scheme, Spatio-Temporal Representation Decoupling and Enhancement (FedST), which wisely leverages surgical domain knowledge during both local-site and global-server training to boost segmentation. Concretely, our model embraces a Representation Separation and Cooperation (RSC) mechanism in local-site training, which decouples the query embedding layer to be trained privately, to encode respective backgrounds. Meanwhile, other parameters are optimized globally to capture the consistent representations of instruments, including the temporal layer to capture similar motion patterns. A textual-guided channel selection is further designed to highlight site-specific features, facilitating model adaptation to each site. Moreover, in global-server training, we propose Synthesis-based Explicit Representation Quantification (SERQ), which defines an explicit representation target based on synthetic data to synchronize the model convergence during fusion for improving model generalization. We construct a new PFL benchmark comprising five surgical sites from public datasets covering four types, with one out-of-federation site. FedST outperforms other state-of-the-art methods on federated sites (1.84% on IoU) and achieves a remarkable improvement on the out-of-federation site (45.29% on IoU). Our source code can be made available at: https://github.com/Meaw0415/FedST. Xiaoming Qi, Chun-Mei Feng 0001, Jialun Pei, Weixin Si, Yueming Jin |
IEEE Trans. Medical Imaging | 6 |
| 2026 | PathRWKV: Enhancing Whole Slide Image Inference With Asymmetric Recurrent ModelingabstractWhole Slide Imaging (WSI) has become a gold standard in cancer diagnosis, inspecting multi-scale information from cellular to tissue levels. Processing an entire WSI directly is infeasible due to GPU memory constraints; thus, Multiple Instance Learning (MIL) has emerged as the standard solution by partitioning WSIs into tiles. While recent two-stage MIL frameworks partially achieve memory efficiency by decoupling tile-level extraction from slide-level modeling, they still face four limitations: 1) the conflict between training throughput and inference memory efficiency, 2) the high susceptibility to overfitting on small-scale WSI datasets with sparse supervision, 3) the disruption of spatial structural integrity during sampling-based training, and 4) the inadequate modeling of multi-scale feature interactions within long sequences. We therefore introduce PathRWKV, a novel State Space Model designed for efficient and robust WSI analysis. To resolve the computational trade-off, we propose an asymmetric structure utilizing max pooling aggregation, enabling parallelized training for high throughput and recurrent inference with constant ( $\mathcal {O}\text {(}{1}\text {)}$ ) memory complexity. To mitigate overfitting, we employ random sampling to enhance data diversity, with a multi-task learning module to regularize feature learning on limited data. To restore spatial context, we introduce 2D sinusoidal position encoding to perceive the relative locations of tissue tiles. To capture comprehensive representations, we integrate TimeMix and ChannelMix modules, enabling dynamic multi-scale feature modeling across temporal and spatial dimensions. Experiments on 29,073 WSIs across 11 datasets demonstrate that PathRWKV outperforms 11 state-of-the-art methods on 10 datasets, establishing it as a scalable and solution with application potential. Sicheng Chen, Borui Kang, Dankai Liao, Qiaochu Xue, Bochong Zhang, Zeyu Liu 0013, Yueming Jin |
IEEE Trans. Medical Imaging | 9 |
| 2026 | DiffBulk: Enhancing Spatial Transcriptomic Prediction With Diffusion-Based TrainingabstractSpatial Transcriptomics (ST) technology detects gene expression from tissue biopsies, playing an emerging role in cancer diagnosis and precision medicine. However, the high cost of ST technology limits its broader application. Recently, deep learning approaches have provided insight into predicting gene expression based on H&E-stained histopathology images. Nevertheless, the relationship between morphological features and gene expression is highly complex. To address these challenges, we propose DiffBulk, a novel two-stage framework that leverages conditional diffusion models to learn expressive image representations enriched with gene expression information. In the first stage, we introduce a gene-to-image conditional diffusion model equipped with a permutation-invariant open-embedding gene encoder, which enables unified training across diverse gene panels. In the second stage, diffusion-derived features are fused with representations from a pathology foundation model, effectively bridging the domain gap and improving downstream gene expression prediction. We evaluate DiffBulk on high-quality Xenium ST data curated from the HEST dataset and the CrunchDAO challenge, constructing tile-level pseudo-bulk datasets for training and evaluation. Extensive experiments demonstrate that DiffBulk consistently outperforms state-of-the-art baselines across all metrics for gene expression prediction. These findings highlight the potential of diffusion-based gene-image representation learning and suggest promising directions for future research. Bochong Zhang, Qiaochu Xue, Zeyu Liu 0013, Dankai Liao, Timothy Antoni, Yeo Hui Ting Grace, Sicheng Chen, Hwee Kuan Lee, Shangqing Lyu, Yueming Jin |
IEEE Trans. Medical Imaging | 11 |
| 2025 | Agentic Reasoning: A Streamlined Framework for Enhancing LLM Reasoning with Agentic ToolsabstractWe introduce Agentic Reasoning, a framework that enhances large language model (LLM) reasoning by integrating external tool-using agents.Agentic Reasoning dynamically leverages web search, code execution, and structured memory to address complex problems requiring deep research.A key innovation in our framework is the Mind-Map agent, which constructs a structured knowledge graph to store reasoning context and track logical relationships, ensuring coherence in long reasoning chains with extensive tool usage.Additionally, we conduct a comprehensive exploration of the Web-Search agent, leading to a highly effective search mechanism that surpasses all prior approaches.When deployed on DeepSeek-R1, our method achieves a new state-of-the-art (SOTA) among public models and delivers performance comparable to OpenAI Deep Research, the leading proprietary model in this domain.Extensive ablation studies validate the optimal selection of agentic tools and confirm the effectiveness of our Mind-Map and Web-Search agents in enhancing LLM reasoning.Our code and data are publicly available. Jiayuan Zhu, Yuyuan Liu, Min Xu 0009, Yueming Jin |
ACL (1) | 5 |
| 2025 | Medical Graph RAG: Evidence-based Medical Large Language Model via Graph Retrieval-Augmented GenerationabstractJunde Wu, Jiayuan Zhu, Yunli Qi, Jingkun Chen, Min Xu, Filippo Menolascina, Yueming Jin, Vicente Grau. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jiayuan Zhu, Yunli Qi, Jingkun Chen, Min Xu 0009, Filippo Menolascina, Yueming Jin, Vicente Grau |
ACL (1) | 7 |
| 2025 | Structure Matters: Revisiting Boundary Refinement in Video Object SegmentationabstractGiven an object mask, Semi-supervised Video Object Segmentation (SVOS) technique aims to track and segment the object across video frames, serving as a fundamental task in computer vision. Although recent memory-based methods demonstrate potential, they often struggle with scenes involving occlusion, particularly in handling object interactions and high feature similarity. To address these issues and meet the real-time processing requirements of downstream applications, in this paper, we propose a novel bOundary Amendment video object Segmentation method with Inherent Structure refinement, hereby named OASIS. Specifically, a lightweight structure refinement module is proposed to enhance segmentation accuracy. With the fusion of rough edge priors captured by the Canny filter and stored object features, the module can generate an object-level structure map and refine the representations by highlighting boundary features. Evidential learning for uncertainty estimation is introduced to further address challenges in occluded regions. The proposed method, OASIS, maintains an efficient design, yet extensive experiments on challenging benchmarks demonstrate its superior performance and competitive inference speed compared to other state-of-the-art methods, i.e., achieving the F values of 91.6 (vs. 89.7 on DAVIS-17 validation set) and G values of 86.6 (vs. 86.2 on YouTubeVOS 2019 validation set) while maintaining a competitive speed of 48 FPS on DAVIS. Guanyi Qin, Ziyue Wang 0005, Daiyun Shen, Haofeng Liu, Hantao Zhou, Runze Hu, Yueming Jin |
ICCV | 8 |
| 2025 | Generalized Deep Multi-View Clustering Via Causal Learning With Partially Aligned Cross-View CorrespondenceabstractMulti-view clustering (MVC) aims to explore the common clustering structure across multiple views. Many existing MVC methods heavily rely on the assumption of view consistency, where alignments for corresponding samples across different views are ordered in advance. However, real-world scenarios often present a challenge as only partial data is consistently aligned across different views, restricting the overall clustering performance. In this work, we consider the model performance decreasing phenomenon caused by data order shift (i.e., from fully to partially aligned) as a generalized multi-view clustering problem. To tackle this problem, we design a causal multi-view clustering network, termed CauMVC. We adopt a causal modeling approach to understand multi-view clustering procedure. To be specific, we formulate the partially aligned data as an intervention and multi-view clustering with partially aligned data as an post-intervention inference. However, obtaining invariant features directly can be challenging. Thus, we design a Variational Auto-Encoder for causal learning by incorporating an encoder from existing information to estimate the invariant features. Moreover, a decoder is designed to perform the post-intervention inference. Lastly, we design a contrastive regularizer to capture sample correlations. To the best of our knowledge, this paper is the first work to deal generalized multi-view clustering via causal learning. Empirical experiments on both fully and partially aligned data illustrate the strong generalization and effectiveness of CauMVC. Xihong Yang, Siwei Wang 0001, Jiaqi Jin, Fangdi Wang, Tianrui Liu 0001, Yueming Jin, Xinwang Liu 0002, En Zhu, Kunlun He |
ICCV | 6 |
| 2025 | Automatically Identify and Rectify: Robust Deep Contrastive Multi-view Clustering in Noisy ScenariosabstractLeveraging the powerful representation learning capabilities, deep multi-view clustering methods have demonstrated reliable performance by effectively integrating multi-source information from diverse views in recent years. Most existing methods rely on the assumption of clean views. However, noise is pervasive in real-world scenarios, leading to a significant degradation in performance. To tackle this problem, we propose a novel multi-view clustering framework for the automatic identification and rectification of noisy data, termed AIRMVC. Specifically, we reformulate noisy identification as an anomaly identification problem using GMM. We then design a hybrid rectification strategy to mitigate the adverse effects of noisy data based on the identification results. Furthermore, we introduce a noise-robust contrastive mechanism to generate reliable representations. Additionally, we provide a theoretical proof demonstrating that these representations can discard noisy information, thereby improving the performance of downstream tasks. Extensive experiments on six benchmark datasets demonstrate that AIRMVC outperforms state-of-the-art algorithms in terms of robustness in noisy scenarios. The code of AIRMVC are available at https://github.com/xihongyang1999/AIRMVC on Github. Xihong Yang, Siwei Wang 0001, Fangdi Wang, Jiaqi Jin, Suyuan Liu, Yue Liu 0008, En Zhu, Xinwang Liu 0002, Yueming Jin |
ICML | 9 |
| 2025 | BCRNet: Enhancing Landmark Detection in Laparoscopic Liver Surgery via Bezier Curve Refinement
Qian Li 0036, Feng Liu 0062, Shuojue Yang, Daiyun Shen, Yueming Jin |
MICCAI (10) | 5 |
| 2025 | ReSurgSAM2: Referring Segment Anything in Surgical Video via Credible Long-Term Tracking
Haofeng Liu, Mingqi Gao 0003, Xuxiao Luo, Ziyue Wang 0005, Guanyi Qin, Yueming Jin |
MICCAI (10) | 7 |
| 2025 | Instrument-Splatting: Controllable Photorealistic Reconstruction of Surgical Instruments Using Gaussian Splatting
Shuojue Yang, Zijian Wu 0001, Mingxuan Hong, Qian Li 0036, Daiyun Shen, Tim Salcudean, Yueming Jin |
MICCAI (3) | 7 |
| 2025 | EIR-SDG: Explore Invariant Representation for Single-source Domain Generalization in Medical Image Segmentation
Ziwei Niu, Shiao Xie, Ziyue Wang 0005, Yen-Wei Chen 0001, Yueming Jin, Lanfen Lin |
ACM Multimedia | 5 |
| 2025 | Multi-scale Temporal Prediction via Incremental Generation and Multi-agent CollaborationabstractAccurate temporal prediction is the bridge between comprehensive scene understanding and embodied artificial intelligence. However, predicting multiple fine-grained states of scene at multiple temporal scales is difficult for vision-language models.
We formalize the Multi‐Scale Temporal Prediction (MSTP) task in general and surgical scene by decomposing multi‐scale into two orthogonal dimensions: the temporal scale, forecasting states of human and surgery at varying look‐ahead intervals, and the state scale, modeling a hierarchy of states in general and surgical scene. For instance in general scene, states of contacting relationship are finer-grained than states of spatial relationship. For instance in surgical scene, medium‐level steps are finer‐grained than high‐level phases yet remain constrained by their encompassing phase.
To support this unified task, we introduce the first MSTP Benchmark, featuring synchronized annotations across multiple state scales and temporal scales. We further propose a novel method, Incremental Generation and Multi‐agent Collaboration (IG-MC), which integrates two key innovations. Firstly, we propose an plug-and-play incremental generation to keep high-quality temporal prediction that continuously synthesizes up-to-date visual previews at expanding temporal scales to inform multiple decision-making agents, ensuring decision content and generated visuals remain synchronized and preventing performance degradation as look‐ahead intervals lengthen.
Secondly, we propose a decision‐driven multi‐agent collaboration framework for multiple states prediction, comprising generation, initiation, and multi‐state assessment agents that dynamically triggers and evaluates prediction cycles to balance global coherence and local fidelity. Extensive experiments on the MSTP Benchmark in general and surgical scene show that IG‐MC is a generalizable plug-and-play method for MSTP, demonstrating the effectiveness of incremental generation and the stability of decision‐driven multi‐agent collaboration. Zhitao Zeng, Guojian Yuan, Junyuan Mao, Yuxuan Wang 0004, Xiaoshuang Jia, Yueming Jin |
NeurIPS | 6 |
| 2025 | An objective comparison of methods for augmented reality in laparoscopic liver resection by preoperative-to-intraoperative image fusion from the MICCAI2022 challengeabstractAugmented reality for laparoscopic liver resection is a visualisation mode that allows a surgeon to localise tumours and vessels embedded within the liver by projecting them on top of a laparoscopic image. Preoperative 3D models extracted from Computed Tomography (CT) or Magnetic Resonance (MR) imaging data are registered to the intraoperative laparoscopic images during this process. Regarding 3D-2D fusion, most algorithms use anatomical landmarks to guide registration, such as the liver's inferior ridge, the falciform ligament, and the occluding contours. These are usually marked by hand in both the laparoscopic image and the 3D model, which is time-consuming and prone to error. Therefore, there is a need to automate this process so that augmented reality can be used effectively in the operating room. We present the Preoperative-to-Intraoperative Laparoscopic Fusion challenge (P2ILF), held during the Medical Image Computing and Computer Assisted Intervention (MICCAI 2022) conference, which investigates the possibilities of detecting these landmarks automatically and using them in registration. The challenge was divided into two tasks: (1) A 2D and 3D landmark segmentation task and (2) a 3D-2D registration task. The teams were provided with training data consisting of 167 laparoscopic images and 9 preoperative 3D models from 9 patients, with the corresponding 2D and 3D landmark annotations. A total of 6 teams from 4 countries participated in the challenge, whose results were assessed for each task independently. All the teams proposed deep learning-based methods for the 2D and 3D landmark segmentation tasks and differentiable rendering-based methods for the registration task. The proposed methods were evaluated on 16 test images and 2 preoperative 3D models from 2 patients. In Task 1, the teams were able to segment most of the 2D landmarks, while the 3D landmarks showed to be more challenging to segment. In Task 2, only one team obtained acceptable qualitative and quantitative registration results. Based on the experimental outcomes, we propose three key hypotheses that determine current limitations and future directions for research in this domain. Sharib Ali, Yamid Espinel, Yueming Jin, Peng Liu 0074, Bianca Güttner, Xukun Zhang, Lihua Zhang 0002, Thomas Dowrick, Matthew J. Clarkson, Shiting Xiao, Yifan Wu 0021, Lei Zhu 0003, Dai Sun, Micha Pfeiffer, Shahid Farid, Lena Maier-Hein, Emmanuel Buc, Adrien Bartoli |
Medical Image Anal. | 3 |
| 2025 | Pro-NeXt: An All-in-One Unified Model for General Fine-Grained Visual RecognitionabstractUnlike general visual classification (CLS) tasks, certain CLS problems are significantly more challenging as they involve recognizing professionally categorized or highly specialized images. Fine-Grained Visual Classification (FGVC) has emerged as a broad solution to address this complexity. However, most existing methods have been predominantly evaluated on a limited set of homogeneous benchmarks, such as bird species or vehicle brands. Moreover, these approaches often train separate models for each specific task, which restricts their generalizability. This paper proposes a scalable and explainable foundational model designed to tackle a wide range of FGVC tasks from a unified and generalizable perspective. We introduce a novel architecture named Pro-NeXt and reveal that Pro-NeXt exhibits substantial generalizability across diverse professional fields such as fashion, medicine, and art areas, previously considered disparate. Our basic-sized Pro-NeXt-B surpasses all preceding task-specific models across 12 distinct datasets within 5 diverse domains. Furthermore, we find its good scaling property that scaling up Pro-NeXt in depth and width with increasing GFlops can consistently enhance its accuracy. Beyond scalability and adaptability, the intermediate features of Pro-NeXt achieve reliable object detection and segmentation performance without extra training, highlighting its solid explainability. We will release the code to promote further research in this area. Jiayuan Zhu, Min Xu 0009, Yueming Jin |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | LMT++: Adaptively Collaborating LLMs With Multi-Specialized Teachers for Continual VQA in Robotic Surgical VideosabstractVisual question answering (VQA) plays a vital role in advancing surgical education. However, due to the privacy concern of patient data, training VQA model with previously used data becomes restricted, making it necessary to use the exemplar-free continual learning (CL) approach. Previous CL studies in the surgical field neglected two critical issues: i) significant domain shifts caused by the wide range of surgical procedures collected from various sources, and ii) the data imbalance problem caused by the unequal occurrence of medical instruments or surgical procedures. This paper addresses these challenges with a multimodal large language model (LLM) and an adaptive weight assignment strategy. First, we developed a novel LLM-assisted multi-teacher CL framework (named LMT++), which could harness the strength of a multimodal LLM as a supplementary teacher. The LLM's strong generalization ability, as well as its good understanding of the surgical domain, help to address the knowledge gap arising from domain shifts and data imbalances. To incorporate the LLM in our CL framework, we further proposed an innovative approach to process the training data, which involves the conversion of complex LLM embeddings into logits value used within our CL training framework. Moreover, we design an adaptive weight assignment approach that balances the generalization ability of the LLM and the domain expertise of conventional VQA models obtained in previous model training processes within the CL framework. Finally, we created a new surgical VQA dataset for model evaluation. Comprehensive experimental findings on these datasets show that our approach surpasses state-of-the-art CL methods. Yuyang Du 0001, Kexin Chen 0003, Yue Zhan, Chang Han Low, Mobarakol Islam, Yueming Jin, Guangyong Chen, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 7 |
| 2025 | S²Former-OR: Single-Stage Bi-Modal Transformer for Scene Graph Generation in ORabstractScene graph generation (SGG) of surgical procedures is crucial in enhancing holistically cognitive intelligence in the operating room (OR). However, previous works have primarily relied on multi-stage learning, where the generated semantic scene graphs depend on intermediate processes with pose estimation and object detection. This pipeline may potentially compromise the flexibility of learning multimodal representations, consequently constraining the overall effectiveness. In this study, we introduce a novel single-stage bi-modal transformer framework for SGG in the OR, termed S2Former-OR, aimed to complementally leverage multi-view 2D scenes and 3D point clouds for SGG in an end-to-end manner. Concretely, our model embraces a View-Sync Transfusion scheme to encourage multi-view visual information interaction. Concurrently, a Geometry-Visual Cohesion operation is designed to integrate the synergic 2D semantic features into 3D point cloud features. Moreover, based on the augmented feature, we propose a novel relation-sensitive transformer decoder that embeds dynamic entity-pair queries and relational trait priors, which enables the direct prediction of entity-pair relations for graph generation without intermediate steps. Extensive experiments have validated the superior SGG performance and lower computational cost of S2Former-OR on 4D-OR benchmark, compared with current OR-SGG methods, e.g., 3 percentage points Precision increase and 24.2M reduction in model parameters. We further compared our method with generic single-stage SGG methods with broader metrics for a comprehensive evaluation, with consistently better performance achieved. Our source code can be made available at: https://github.com/PJLallen/S2Former-OR. Jialun Pei, Diandian Guo, Jingyang Zhang, Manxi Lin, Yueming Jin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Instrument-Tissue-Guided Surgical Action Triplet Detection via Textual-Temporal Trail ExplorationabstractSurgical action triplet detection offers intuitive intraoperative scene analysis for dynamically perceiving laparoscopic surgical workflows and analyzing the interaction between instruments and tissues. The current challenge of this task lies in simultaneously localizing surgical instruments while performing more accurate surgical triplet recognition to enhance a comprehensive understanding of intraoperative surgical scenes. To fully leverage the spatial localization of surgical instruments for associating with triplet detection, we propose an Instrument-Tissue-Guided Triplet detector, termed ITG-Trip, which navigates the confluence of surgical action cues through instrument and tissue pseudo-localization labeling to optimize action triplet detection. For exploiting textual and temporal trails, our framework embraces a Visual-Linguistic Association (VLA) module that exploits a pre-trained text encoder to distill textual prior knowledge, enhancing semantic information in global visual features and compensating rare interaction class perception. Besides, we introduce a Mamba-enhanced Spatial-temporal Perception (MSP) decoder, which weaves Mamba and Transformer blocks to explore subject- and object-aware spatial and temporal information to improve the accuracy of action triplet detection in long-time sequence surgical videos. Experimental results on the CholecT50 benchmark indicate that our method significantly outperforms existing state-of-the-art methods in both instrument localization and action triplet detection. The code is available at: github.com/PJLallen/ITG-Trip. Jialun Pei, Jiaan Zhang, Guanyi Qin, Kai Wang 0092, Yueming Jin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 5 |
| 2025 | DC²T: Disentanglement-Guided Consolidation and Consistency Training for Semi-Supervised Cross-Site Continual SegmentationabstractContinual Learning (CL) is recognized to be a storage-efficient and privacy-protecting approach for learning from sequentially-arriving medical sites. However, most existing CL methods assume that each site is fully labeled, which is impractical due to budget and expertise constraint. This paper studies the Semi-Supervised Continual Learning (SSCL) that adopts partially-labeled sites arriving over time, with each site delivering only limited labeled data while the majority remains unlabeled. In this regard, it is challenging to effectively utilize unlabeled data under dynamic cross-site domain gaps, leading to intractable model forgetting on such unlabeled data. To address this problem, we introduce a novel Disentanglement-guided Consolidation and Consistency Training (DC2T) framework, which roots in an Online Semi-Supervised representation Disentanglement (OSSD) perspective to excavate content representations of partially labeled data from sites arriving over time. Moreover, these content representations are required to be consolidated for site-invariance and calibrated for style-robustness, in order to alleviate forgetting even in the absence of ground truth. Specifically, for the invariance on previous sites, we retain historical content representations when learning on a new site, via a Content-inspired Parameter Consolidation (CPC) method that prevents altering the model parameters crucial for content preservation. For the robustness against style variation, we develop a Style-induced Consistency Training (SCT) scheme that enforces segmentation consistency over style-related perturbations to recalibrate content encoding. We extensively evaluate our method on fundus and cardiac image segmentation, indicating the advantage over existing SSCL methods for alleviating forgetting on unlabeled data. Jingyang Zhang, Jialun Pei, Dunyuan Xu, Yueming Jin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 4 |
| 2024 | MedSegDiff-V2: Diffusion-Based Medical Image Segmentation with TransformerabstractThe Diffusion Probabilistic Model (DPM) has recently gained popularity in the field of computer vision, thanks to its image generation applications, such as Imagen, Latent Diffusion Models, and Stable Diffusion, which have demonstrated impressive capabilities and sparked much discussion within the community. Recent investigations have further unveiled the utility of DPM in the domain of medical image analysis, as underscored by the commendable performance exhibited by the medical image segmentation model across various tasks. Although these models were originally underpinned by a UNet architecture, there exists a potential avenue for enhancing their performance through the integration of vision transformer mechanisms. However, we discovered that simply combining these two models resulted in subpar performance. To effectively integrate these two cutting-edge techniques for the Medical image segmentation, we propose a novel Transformer-based Diffusion framework, called MedSegDiff-V2. We verify its effectiveness on 20 medical image segmentation tasks with different image modalities. Through comprehensive evaluation, our approach demonstrates superiority over prior state-of-the-art (SOTA) methodologies. Code is released at https://github.com/KidsWithTokens/MedSegDiff. Wei Ji 0011, Huazhu Fu, Min Xu 0009, Yueming Jin, Yanwu Xu 0001 |
AAAI | 5 |
| 2024 | ANEDL: Adaptive Negative Evidential Deep Learning for Open-Set Semi-supervised LearningabstractSemi-supervised learning (SSL) methods assume that labeled data, unlabeled data and test data are from the same distribution. Open-set semi-supervised learning (Open-set SSL) con- siders a more practical scenario, where unlabeled data and test data contain new categories (outliers) not observed in labeled data (inliers). Most previous works focused on out- lier detection via binary classifiers, which suffer from insufficient scalability and inability to distinguish different types of uncertainty. In this paper, we propose a novel framework, Adaptive Negative Evidential Deep Learning (ANEDL) to tackle these limitations. Concretely, we first introduce evidential deep learning (EDL) as an outlier detector to quantify different types of uncertainty, and design different uncertainty metrics for self-training and inference. Furthermore, we propose a novel adaptive negative optimization strategy, making EDL more tailored to the unlabeled dataset containing both inliers and outliers. As demonstrated empirically, our proposed method outperforms existing state-of-the-art methods across four datasets. Yang Yu 0070, Danruo Deng, Furui Liu, Qi Dou 0001, Yueming Jin, Guangyong Chen, Pheng-Ann Heng |
AAAI | 5 |
| 2024 | Energy-Induced Explicit Quantification for Multi-modality MRI Fusion
Xiaoming Qi, Yuan Zhang 0019, Tong Wang 0022, Guanyu Yang 0001, Yueming Jin, Shuo Li 0001 |
ECCV (7) | 5 |
| 2024 | LLM-Assisted Multi-Teacher Continual Learning for Visual Question Answering in Robotic SurgeryabstractVisual question answering (VQA) can be fundamentally crucial for promoting robotic-assisted surgical education. In practice, the needs of trainees are constantly evolving, such as learning more surgical types and adapting to new surgical instruments/techniques. Therefore, continually updating the VQA system by a sequential data stream from multiple resources is demanded in robotic surgery to address new tasks. In surgical scenarios, the privacy issue of patient data often restricts the availability of old data when updating the model, necessitating an exemplar-free continual learning (CL) setup. However, prior studies overlooked two vital problems of the surgical domain: i) large domain shifts from diverse surgical operations collected from multiple departments or clinical centers, and ii) severe data imbalance arising from the uneven presence of surgical instruments or activities during surgical procedures. This paper proposes to address these two problems with a multimodal large language model (LLM) and an adaptive weight assignment methodology. We first develop a new multi-teacher CL framework that leverages a multimodal LLM as the additional teacher. The strong generalization ability of the LLM can bridge the knowledge gap when domain shifts and data imbalances occur. We then put forth a novel data processing method that transforms complex LLM embeddings into logits compatible with our CL framework. We also design an adaptive weight assignment approach that balances the generalization ability of the LLM and the domain expertise of the old CL model. Finally, we construct a new dataset for surgical VQA tasks. Extensive experimental results demonstrate the superiority of our method to other advanced CL models. Kexin Chen 0003, Yuyang Du 0001, Tao You, Mobarakol Islam, Yueming Jin, Guangyong Chen, Pheng-Ann Heng |
ICRA | 6 |
| 2024 | Prompt Your Brain: Scaffold Prompt Tuning for Efficient Adaptation of fMRI Pre-trained Model
Zijian Dong 0001, Yilei Wu, Zijiao Chen, Yueming Jin, Juan Helen Zhou |
MICCAI (11) | 5 |
| 2024 | Tri-Modal Confluence with Temporal Dynamics for Scene Graph Generation in Operating Rooms
Diandian Guo, Manxi Lin, Jialun Pei, He Tang 0002, Yueming Jin, Pheng-Ann Heng |
MICCAI (6) | 5 |
| 2024 | Epicardium Prompt-Guided Real-Time Cardiac Ultrasound Frame-to-Volume Registration
Long Lei, Jun Zhou 0007, Jialun Pei, Baoliang Zhao, Yueming Jin, Jeremy Yuen-Chun Teoh, Harry Qin, Pheng-Ann Heng |
MICCAI (2) | 5 |
| 2024 | Deform3DGS: Flexible Deformation for Fast Surgical Scene Reconstruction with Gaussian Splatting
Shuojue Yang, Qian Li 0036, Daiyun Shen, Bingchen Gong, Qi Dou 0001, Yueming Jin |
MICCAI (6) | 6 |
| 2024 | SurgT challenge: Benchmark of soft-tissue trackers for robotic surgery
João Cartucho, Alistair Weld, Samyakh Tukra, Haozheng Xu, Hiroki Matsuzaki, Taiyo Ishikawa, Minjun Kwon, Yongeun Jang, Kwang-Ju Kim, Gwang Lee, Bizhe Bai, Lüder A. Kahrs, Lars Boecking, Simeon Allmendinger, Leopold Müller, Yueming Jin, Sophia Bano, Francisco Vasconcelos 0001, Wolfgang Reiter, Jonas Hajek, Estevão Lima, João L. Vilaça, Sandro F. Queiros, Stamatia Giannarou |
Medical Image Anal. | 17 |
| 2024 | SimCol3D - 3D reconstruction during colonoscopy challengeabstractColorectal cancer is one of the most common cancers in the world. While colonoscopy is an effective screening technique, navigating an endoscope through the colon to detect polyps is challenging. A 3D map of the observed surfaces could enhance the identification of unscreened colon tissue and serve as a training platform. However, reconstructing the colon from video footage remains difficult. Learning-based approaches hold promise as robust alternatives, but necessitate extensive datasets. Establishing a benchmark dataset, the 2022 EndoVis sub-challenge SimCol3D aimed to facilitate data-driven depth and pose prediction during colonoscopy. The challenge was hosted as part of MICCAI 2022 in Singapore. Six teams from around the world and representatives from academia and industry participated in the three sub-challenges: synthetic depth prediction, synthetic pose prediction, and real pose prediction. This paper describes the challenge, the submitted methods, and their results. We show that depth prediction from synthetic colonoscopy images is robustly solvable, while pose estimation remains an open research question. Anita Rau, Sophia Bano, Yueming Jin, Pablo Azagra, Javier Morlana, Rawen Kader, Edward Sanderson, Bogdan J. Matuszewski, Erez Posner, Netanel Frank, Varshini Elangovan, Sista Raviteja, Zhengwen Li, Jiquan Liu, Seenivasan Lalithkumar, Mobarakol Islam, Hongliang Ren 0001, Laurence B. Lovat, J. M. M. Montiel, Danail Stoyanov |
Medical Image Anal. | 3 |
| 2024 | CalibNet: Dual-Branch Cross-Modal Calibration for RGB-D Salient Instance SegmentationabstractIn this study, we propose a novel approach for RGB-D salient instance segmentation using a dual-branch cross-modal feature calibration architecture called CalibNet. Our method simultaneously calibrates depth and RGB features in the kernel and mask branches to generate instance-aware kernels and mask features. CalibNet consists of three simple modules, a dynamic interactive kernel (DIK) and a weight-sharing fusion (WSF), which work together to generate effective instance-aware kernels and integrate cross-modal features. To improve the quality of depth features, we incorporate a depth similarity assessment (DSA) module prior to DIK and WSF. In addition, we further contribute a new DSIS dataset, which contains 1,940 images with elaborate instance-level annotations. Extensive experiments on three challenging benchmarks show that CalibNet yields a promising result, i.e., 58.0% AP with 320×480 input size on the COME15K-E test set, which significantly surpasses the alternative frameworks. Our code and dataset will be publicly available at: https://github.com/PJLallen/CalibNet. Jialun Pei, Tao Jiang 0002, He Tang 0002, Nian Liu 0002, Yueming Jin, Deng-Ping Fan, Pheng-Ann Heng |
IEEE Trans. Image Process. | 5 |
| 2024 | Guest Editorial: Trustworthy Machine Learning for Health InformaticsabstractMachine learning (ML), the stem of today's artificial intelligence, has shown significant growth in the field of biomedical and health informatics. On the one hand, ML techniques are becoming more complex in order to deal with real-world data. On the other hand, ML is also more and more accessible to broader users. For example, automated machine learning products are enabling users to build their own ML models without writing code [1]. Luyang Luo, Daguang Xu, Harry Qin, Yueming Jin, Hao Chen 0011 |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | FedDP: Dual Personalization in Federated Medical Image SegmentationabstractPersonalized federated learning (PFL) addresses the data heterogeneity challenge faced by general federated learning (GFL). Rather than learning a single global model, with PFL a collection of models are adapted to the unique feature distribution of each site. However, current PFL methods rarely consider self-attention networks which can handle data heterogeneity by long-range dependency modeling and they do not utilize prediction inconsistencies in local models as an indicator of site uniqueness. In this paper, we propose FedDP, a novel fed erated learning scheme with d ual p ersonalization, which improves model personalization from both feature and prediction aspects to boost image segmentation results. We leverage long-range dependencies by designing a local query (LQ) that decouples the query embedding layer out of each local model, whose parameters are trained privately to better adapt to the respective feature distribution of the site. We then propose inconsistency-guided calibration (IGC), which exploits the inter-site prediction inconsistencies to accommodate the model learning concentration. By encouraging a model to penalize pixels with larger inconsistencies, we better tailor prediction-level patterns to each local site. Experimentally, we compare FedDP with the state-of-the-art PFL methods on two popular medical image segmentation tasks with different modalities, where our results consistently outperform others on both tasks. Our code and models are available at https://github.com/jcwang123/PFL-Seg-Trans. Jiacheng Wang 0002, Yueming Jin, Danail Stoyanov, Liansheng Wang 0002 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Video-Instrument Synergistic Network for Referring Video Instrument Segmentation in Robotic SurgeryabstractSurgical instrument segmentation is fundamentally important for facilitating cognitive intelligence in robot-assisted surgery. Although existing methods have achieved accurate instrument segmentation results, they simultaneously generate segmentation masks of all instruments, which lack the capability to specify a target object and allow an interactive experience. This paper focuses on a novel and essential task in robotic surgery, i.e., Referring Surgical Video Instrument Segmentation (RSVIS), which aims to automatically identify and segment the target surgical instruments from each video frame, referred by a given language expression. This interactive feature offers enhanced user engagement and customized experiences, greatly benefiting the development of the next generation of surgical education systems. To achieve this, this paper constructs two surgery video datasets to promote the RSVIS research. Then, we devise a novel Video-Instrument Synergistic Network (VIS-Net) to learn both video-level and instrument-level knowledge to boost performance, while previous work only utilized video-level information. Meanwhile, we design a Graph-based Relation-aware Module (GRM) to model the correlation between multi-modal information (i.e., textual description and video frame) to facilitate the extraction of instrument-level information. Extensive experimental results on two RSVIS datasets exhibit that the VIS-Net can significantly outperform existing state-of-the-art referring segmentation methods. We will release our code and dataset for future research (https://github.com/whq-xxh/RSVIS). Hongqiu Wang, Guang Yang 0006, Harry Qin, Yike Guo, Yueming Jin, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Calibrate the Inter-Observer Segmentation Uncertainty via Diagnosis-First PrincipleabstractMany of the tissues/lesions in the medical images may be ambiguous. Therefore, medical segmentation is typically annotated by a group of clinical experts to mitigate personal bias. A common solution to fuse different annotations is the majority vote, e.g., taking the average of multiple labels. However, such a strategy ignores the difference between the grader expertness. Inspired by the observation that medical image segmentation is usually used to assist the disease diagnosis in clinical practice, we propose the diagnosis-first principle, which is to take disease diagnosis as the criterion to calibrate the inter-observer segmentation uncertainty. Following this idea, a framework named Diagnosis-First segmentation Framework (DiFF) is proposed. Specifically, DiFF will first learn to fuse the multi-rater segmentation labels to a single ground-truth which could maximize the disease diagnosis performance. We dubbed the fused ground-truth as Diagnosis-First Ground-truth (DF-GT). Then, the Take and Give Model (T&G Model) to segment DF-GT from the raw image is proposed. With the T&G Model, DiFF can learn the segmentation with the calibrated uncertainty that facilitate the disease diagnosis. We verify the effectiveness of DiFF on three different medical segmentation tasks: optic-disc/optic-cup (OD/OC) segmentation on fundus images, thyroid nodule segmentation on ultrasound images, and skin lesion segmentation on dermoscopic images. Experimental results show that the proposed DiFF can effectively calibrate the segmentation uncertainty, and thus significantly facilitate the corresponding disease diagnosis, which outperforms previous state-of-the-art multi-rater learning methods. Yu Zhang 0091, Huihui Fang, Lixin Duan, Mingkui Tan, Weihua Yang, Yueming Jin, Yanwu Xu 0001 |
IEEE Trans. Medical Imaging | 9 |
| 2023 | Dynamic Interactive Relation Capturing via Scene Graph Learning for Robotic Surgical Report GenerationabstractFor robot-assisted surgery, an accurate surgical report reflects clinical operations during surgery and helps document entry tasks, post-operative analysis and follow-up treatment. It is a challenging task due to many complex and diverse interactions between instruments and tissues in the surgical scene. Although existing surgical report generation methods based on deep learning have achieved large success, they often ignore the interactive relation between tissues and instrumental tools, thereby degrading the report generation performance. This paper presents a neural network to boost surgical report generation by explicitly exploring the interactive relation between tissues and surgical instruments. To do so, we first devise a relational exploration (RE) module to model the interactive relation via graph learning, and an interaction perception (IP) module to assist the graph learning in RE module. In our IP module, we first devise a node tracking system to identify and append missing graph nodes of the current video frame for constructing graphs at RE module. Moreover, the IP module generates a global attention model to indicate the existence of the interactive relation on the whole scene of the current video frame to eliminate the graph learning at the current video frame. Furthermore, our IP module predicts a local attention model to more accurately identify the interaction relation of each graph node for assisting the graph updating at the RE module. After that, we concatenate features of all graph nodes of RE module and pass concatenated features into a transformer for generating the output surgical report. We validate the effectiveness of our method on a widely-used robotic surgery benchmark dataset, and experimental results show that our network can significantly outperform existing state-of-the-art surgical report generation methods (e.g., 7.48% and 5.43% higher for BLEU-1 and ROUGE). Hongqiu Wang, Yueming Jin, Lei Zhu 0003 |
ICRA | 2 |
| 2023 | Surgical Activity Triplet Recognition via Triplet Disentanglement
Yiliang Chen, Shengfeng He, Yueming Jin, Harry Qin |
MICCAI (9) | 3 |
| 2023 | Beyond the Snapshot: Brain Tokenized Graph Transformer for Longitudinal Brain Functional Connectome Embedding
Zijian Dong 0001, Yilei Wu, Joanna Su Xian Chong, Yueming Jin, Juan Helen Zhou |
MICCAI (5) | 5 |
| 2023 | Imitation Learning from Expert Video Data for Dissection Trajectory Prediction in Endoscopic Surgical Procedure
Jianan Li 0006, Yueming Jin, Yueyao Chen, Hon-Chi Yip, Markus Scheppach, Philip W. Y. Chiu, Yeung Yam, Helen M. Meng, Qi Dou 0001 |
MICCAI (9) | 2 |
| 2023 | Regressing Simulation to Real: Unsupervised Domain Adaptation for Automated Quality Assessment in Transoesophageal Echocardiography
Jialang Xu, Yueming Jin, Bruce Martin, Andrew P. T. Smith, Susan Wright, Danail Stoyanov, Evangelos B. Mazomenos |
MICCAI (9) | 2 |
| 2023 | Unite-Divide-Unite: Joint Boosting Trunk and Structure for High-accuracy Dichotomous Image SegmentationabstractHigh-accuracy Dichotomous Image Segmentation (DIS) aims to pinpoint category-agnostic foreground objects from natural scenes. The main challenge for DIS involves identifying the highly accurate dominant area while rendering detailed object structure. However, directly using a general encoder-decoder architecture may result in an oversupply of high-level features and neglect the shallow spatial information necessary for partitioning meticulous structures. To fill this gap, we introduce a novel Unite-Divide-Unite Network (UDUN) that restructures and bipartitely arranges complementary features to simultaneously boost the effectiveness of trunk and structure identification. The proposed UDUN proceeds from several strengths. First, a dual-size input feeds into the shared backbone to produce more holistic and detailed features while keeping the model lightweight. Second, a simple Divide-and-Conquer Module (DCM) is proposed to decouple multiscale low- and high-level features into our structure decoder and trunk decoder to obtain structure and trunk information respectively. Moreover, we design a Trunk-Structure Aggregation module (TSA) in our union decoder that performs cascade integration for uniform high-accuracy segmentation. As a result, UDUN performs favorably against state-of-the-art competitors in all six evaluation metrics on overall DIS-TE, i.e., achieving 0.772 weighted F-measure and 977 HCE. Using 1024X1024 input, our model enables real-time inference at 65.3 fps with ResNet-18. The source code is available at https://github.com/PJLallen/UDUN. Jialun Pei, Zhangjun Zhou, Yueming Jin, He Tang 0002, Pheng-Ann Heng |
ACM Multimedia | 3 |
| 2023 | Comparative validation of machine learning algorithms for surgical workflow and skill analysis with the HeiChole benchmarkabstractPURPOSE: Surgical workflow and skill analysis are key technologies for the next generation of cognitive surgical assistance systems. These systems could increase the safety of the operation through context-sensitive warnings and semi-autonomous robotic assistance or improve training of surgeons via data-driven feedback. In surgical workflow analysis up to 91% average precision has been reported for phase recognition on an open data single-center video dataset. In this work we investigated the generalizability of phase recognition algorithms in a multicenter setting including more difficult recognition tasks such as surgical action and surgical skill. METHODS: To achieve this goal, a dataset with 33 laparoscopic cholecystectomy videos from three surgical centers with a total operation time of 22 h was created. Labels included framewise annotation of seven surgical phases with 250 phase transitions, 5514 occurences of four surgical actions, 6980 occurences of 21 surgical instruments from seven instrument categories and 495 skill classifications in five skill dimensions. The dataset was used in the 2019 international Endoscopic Vision challenge, sub-challenge for surgical workflow and skill analysis. Here, 12 research teams trained and submitted their machine learning algorithms for recognition of phase, action, instrument and/or skill assessment. RESULTS: F1-scores were achieved for phase recognition between 23.9% and 67.7% (n = 9 teams), for instrument presence detection between 38.5% and 63.8% (n = 8 teams), but for action recognition only between 21.8% and 23.3% (n = 5 teams). The average absolute error for skill assessment was 0.78 (n = 1 team). CONCLUSION: Surgical workflow and skill analysis are promising technologies to support the surgical team, but there is still room for improvement, as shown by our comparison of machine learning algorithms. This novel HeiChole benchmark can be used for comparable evaluation and validation of future work. In future studies, it is of utmost importance to create more open, high-quality datasets in order to allow the development of artificial intelligence and cognitive robotics in surgery. Martin Wagner 0001, Beat P. Müller-Stich, Anna Kisilenko, Patrick Heger, Lars Mündermann, David M. Lubotsky, Tornike Davitashvili, Manuela Capek, Annika Reinke, Carissa Reid, Tong Yu 0009, Armine Vardazaryan, Chinedu Innocent Nwoye, Nicolas Padoy, Eungjoo Lee 0001, Constantin Disch, Hans Meine, Tong Xia, Fucang Jia, Satoshi Kondo, Wolfgang Reiter, Yueming Jin, Yonghao Long 0001, Meirui Jiang, Qi Dou 0001, Pheng-Ann Heng, Isabell Twick, Kadir Kirtaç, Enes Hosgor, Jon Lindström Bolmgren, Michael Stenzel, Björn von Siemens, Zhenxiao Ge, Haiming Sun, Di Xie, Mengqi Guo, Daochang Liu, Hannes Kenngott, Felix Nickel, Moritz von Frankenberg, Franziska Mathis-Ullrich, Annette Kopp-Schneider, Lena Maier-Hein, Stefanie Speidel, Sebastian Bodenstedt |
Medical Image Anal. | 25 |
| 2022 | Personalizing Federated Medical Image Segmentation via Local Calibration
Jiacheng Wang 0002, Yueming Jin, Liansheng Wang 0002 |
ECCV (21) | 2 |
| 2022 | TraSeTR: Track-to-Segment Transformer with Contrastive Query for Instance-level Instrument Segmentation in Robotic SurgeryabstractSurgical instrument segmentation - in general a pixel classification task - is fundamentally crucial for promoting cognitive intelligence in robot-assisted surgery (RAS). However, previous methods are struggling with discriminating instrument types and instances. To address above issues, we explore a mask classification paradigm that produces per-segment predictions. We propose TraSeTR, a novel Track-to-Segment Transformer that wisely exploits tracking cues to assist surgical instrument segmentation. TraSeTR jointly reasons about the instrument type, location, and identity with instance-level predictions i.e., a set of class-bbox-mask pairs, by decoding query embeddings. Specifically, we introduce the prior query that encoded with previous temporal knowledge, to transfer tracking signals to current instances via identity matching. A contrastive query learning strategy is further applied to reshape the query feature space, which greatly alleviates the tracking difficulty caused by large temporal variations. The effectiveness of our method is demonstrated with state-of-the-art instrument type segmentation results on three public datasets, including two RAS benchmarks from EndoVis Challenges and one cataract surgery dataset CaDIs. Yueming Jin, Pheng-Ann Heng |
ICRA | 2 |
| 2022 | Pseudo-label Guided Cross-video Pixel Contrast for Robotic Surgical Scene Segmentation with Limited AnnotationsabstractSurgical scene segmentation is fundamentally crucial for prompting cognitive assistance in robotic surgery. However, pixel-wise annotating surgical video in a frame-by-frame manner is expensive and time consuming. To greatly reduce the labeling burden, in this work, we study semi-supervised scene segmentation from robotic surgical video, which is practically essential yet rarely explored before. We consider a clinically suitable annotation situation under the equidistant sampling. We then propose PGV-CL, a novel pseudo-label guided cross-video contrast learning method to boost scene segmentation. It effectively leverages unlabeled data for a trusty and global model regularization that produces more discriminative feature representation. Concretely, for trusty representation learning, we propose to incorporate pseudo labels to instruct the pair selection, obtaining more reliable representation pairs for pixel contrast. Moreover, we expand the representation learning space from previous image-level to cross-video, which can capture the global semantics to benefit the learning process. We extensively evaluate our method on a public robotic surgery dataset EndoVis18 and a public cataract dataset CaDIS. Experimental results demonstrate the effectiveness of our method, consistently outperforming the state-of-the-art semi-supervised methods under different labeling ratios, and even surpassing fully supervised training on EndoVis18 with 10.1% labeling. Our code is available at https://github.com/yangyu-cuhk/PGV-CL. Yang Yu 0070, Yueming Jin, Guangyong Chen, Qi Dou 0001, Pheng-Ann Heng |
IROS | 3 |
| 2022 | Real-time landmark detection for precise endoscopic submucosal dissection via shape-aware relation network
Jiacheng Wang 0002, Yueming Jin, Shuntian Cai, Hongzhi Xu, Pheng-Ann Heng, Harry Qin, Liansheng Wang 0002 |
Medical Image Anal. | 2 |
| 2022 | Unsupervised feature disentanglement for video retrieval in minimally invasive surgery
Ziyi Wang 0006, Bo Lu 0001, Yueming Jin, Zerui Wang, Tak Hong Cheung, Pheng-Ann Heng, Qi Dou 0001, Yun-Hui Liu 0001 |
Medical Image Anal. | 4 |
| 2022 | Toward Image-Guided Automated Suture Grasping Under Complex Environments: A Learning-Enabled and Optimization-Based Holistic FrameworkabstractTo realize a higher-level autonomy of surgical knot tying in minimally invasive surgery (MIS), automated suture grasping, which bridges the suture stitching and looping procedures, is an important yet challenging task needs to be achieved. This paper presents a holistic framework with image-guided and automation techniques to robotize this operation even under complex environments. The whole task is initialized by suture segmentation, in which we propose a novel semi-supervised learning architecture featured with a suture-aware loss to pertinently learn its slender information using both annotated and unannotated data. With successful segmentation in stereo-camera, we develop a Sampling-based Sliding Pairing (SSP) algorithm to online optimize the suture’s 3D shape. By jointly studying the robotic configuration and the suture’s spatial characteristics, a target function is introduced to find the optimal grasping pose of the surgical tool with Remote Center of Motion (RCM) constraints. To compensate for inherent errors and practical uncertainties, a unified grasping strategy with a novel vision-based mechanism is introduced to autonomously accomplish this grasping task. Our framework is extensively evaluated from learning-based segmentation, 3D reconstruction, and image-guided grasping on the da Vinci Research Kit (dVRK) platform, where we achieve high performances and successful rates in perceptions and robotic manipulations. These results prove the feasibility of our approach in automating the suture grasping task, and this work fills the gap between automated surgical stitching and looping, stepping towards a higher-level of task autonomy in surgical knot tying. Note to Practitioners—This paper aims to automate the suture grasping task in surgical knot tying by leveraging stereo visual guidance. To effectively robotize this procedure, it requires multidisciplinary knowledge to achieve suture segmentation, 3D shape reconstruction, and reliable automated grasping, while there are no existing works tackling this procedure especially using robots with RCM kinematics constraints and under complex environments. In this article, we propose a learning-driven method along with a 3D shape optimizer, which can conduct the suture segmentation and output its accurate spatial coordinates, serving as guidance for automated grasping operation. Apart from this, we introduce a unified function to optimize the grasping pose, and a vision-based grasping strategy is also proposed to intelligently complete this task. The experiments extensively validate the feasibility of our framework for automated suture grasp, and its successful completion can serve as a basis for the following looping manipulation, hence filling a step gap in robot-assisted knot tying. This framework can be also encapsulated into the medical robotic system, and by simply indicating (e.g. mouse click) the rough position of the suture’s tip in one camera frame, the overall framework can be initialized and further accomplish the suture grasping task, which further prompts a full autonomy of surgical knot tying in the near future. Bo Lu 0001, Bin Li 0082, Wei Chen 0068, Yueming Jin, Qi Dou 0001, Pheng-Ann Heng, Yun-Hui Liu 0001 |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2022 | Learning With Privileged Multimodal Knowledge for Unimodal SegmentationabstractMultimodal learning usually requires a complete set of modalities during inference to maintain performance. Although training data can be well-prepared with high-quality multiple modalities, in many cases of clinical practice, only one modality can be acquired and important clinical evaluations have to be made based on the limited single modality information. In this work, we propose a privileged knowledge learning framework with the 'Teacher-Student' architecture, in which the complete multimodal knowledge that is only available in the training data (called privileged information) is transferred from a multimodal teacher network to a unimodal student network, via both a pixel-level and an image-level distillation scheme. Specifically, for the pixel-level distillation, we introduce a regularized knowledge distillation loss which encourages the student to mimic the teacher's softened outputs in a pixel-wise manner and incorporates a regularization factor to reduce the effect of incorrect predictions from the teacher. For the image-level distillation, we propose a contrastive knowledge distillation loss which encodes image-level structured information to enrich the knowledge encoding in combination with the pixel-level distillation. We extensively evaluate our method on two different multi-class segmentation tasks, i.e., cardiac substructure segmentation and brain tumor segmentation. Experimental results on both tasks demonstrate that our privileged knowledge learning is effective in improving unimodal segmentation and outperforms previous methods. Cheng Chen 0013, Qi Dou 0001, Yueming Jin, Quande Liu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2022 | Exploring Intra- and Inter-Video Relation for Surgical Semantic Scene SegmentationabstractAutomatic surgical scene segmentation is fundamental for facilitating cognitive intelligence in the modern operating theatre. Previous works rely on conventional aggregation modules (e.g., dilated convolution, convolutional LSTM), which only make use of the local context. In this paper, we propose a novel framework STswinCL that explores the complementary intra- and inter-video relations to boost segmentation performance, by progressively capturing the global context. We firstly develop a hierarchy Transformer to capture intra-video relation that includes richer spatial and temporal cues from neighbor pixels and previous frames. A joint space-time window shift scheme is proposed to efficiently aggregate these two cues into each pixel embedding. Then, we explore inter-video relation via pixel-to-pixel contrastive learning, which well structures the global embedding space. A multi-source contrast training objective is developed to group the pixel embeddings across videos with the ground-truth guidance, which is crucial for learning the global property of the whole data. We extensively validate our approach on two public surgical video benchmarks, including EndoVis18 Challenge and CaDIS dataset. Experimental results demonstrate the promising performance of our method, which consistently exceeds previous state-of-the-art approaches. Code is available at https://github.com/YuemingJin/STswinCL. Yueming Jin, Yang Yu 0070, Cheng Chen 0013, Pheng-Ann Heng, Danail Stoyanov |
IEEE Trans. Medical Imaging | 1 |
| 2021 | Modelling Neighbor Relation in Joint Space-Time Graph for Video Correspondence LearningabstractThis paper presents a self-supervised method for learning reliable visual correspondence from unlabeled videos. We formulate the correspondence as finding paths in a joint space-time graph, where nodes are grid patches sampled from frames, and are linked by two type of edges: (i) neighbor relations that determine the aggregation strength from intra-frame neighbors in space, and (ii) similarity relations that indicate the transition probability of inter-frame paths across time. Leveraging the cycle-consistency in videos, our contrastive learning objective discriminates dynamic objects from both their neighboring views and temporal views. Compared with prior works, our approach actively explores the neighbor relations of central instances to learn a latent association between center-neighbor pairs (e.g., "hand – arm") across time, thus improving the instance discrimination. Without fine-tuning, our learned representation outperforms the state-of-the-art self-supervised methods on a variety of visual tasks including video object propagation, part propagation, and pose keypoint tracking. Our self-supervised method also surpasses some fully supervised algorithms designed for the specific tasks. Yueming Jin, Pheng-Ann Heng |
ICCV | 2 |
| 2021 | Relational Graph Learning on Visual and Kinematics Embeddings for Accurate Gesture Recognition in Robotic SurgeryabstractAutomatic surgical gesture recognition is fundamentally important to enable intelligent cognitive assistance in robotic surgery. With recent advancement in robot-assisted minimally invasive surgery, rich information including surgical videos and robotic kinematics can be recorded, which provide complementary knowledge for understanding surgical gestures. However, existing methods either solely adopt uni-modal data or directly concatenate multi-modal representations, which can not sufficiently exploit the informative correlations inherent in visual and kinematics data to boost gesture recognition accuracies. In this regard, we propose a novel online approach of multi-modal relational graph network (i.e., MRG-Net) to dynamically integrate visual and kinematics information through interactive message propagation in the latent feature space. In specific, we first extract embeddings from video and kinematics sequences with temporal convolutional networks and LSTM units. Next, we identify multi-relations in these multi-modal embeddings and leverage them through a hierarchical relational graph learning module. The effectiveness of our method is demonstrated with state-of-the-art results on the public JIGSAWS dataset, outperforming current uni-modal and multi-modal methods on both suturing and knot typing tasks. Furthermore, we validated our method on in-house visual-kinematics datasets collected with da Vinci Research Kit (dVRK) platforms in two centers, with consistent promising performance achieved. Our code and data are released at: https://www.cse.cuhk.edu.hk/~yhlong/mrgnet.html. Yonghao Long 0001, Jie Ying Wu, Bo Lu 0001, Yueming Jin, Mathias Unberath, Yun-Hui Liu 0001, Pheng-Ann Heng, Qi Dou 0001 |
ICRA | 4 |
| 2021 | One to Many: Adaptive Instrument Segmentation via Meta Learning and Dynamic Online Adaptation in Robotic Surgical VideoabstractSurgical instrument segmentation in robot-assisted surgery (RAS) - especially that using learning-based models - relies on the assumption that training and testing videos are sampled from the same domain. However, it is impractical and expensive to collect and annotate sufficient data from every new domain. To greatly increase the label efficiency, we explore a new problem, i.e., adaptive instrument segmentation, which is to effectively adapt one source model to new robotic surgical videos from multiple target domains, only given the annotated instruments in the first frame. We propose MDAL, a meta-learning based dynamic online adaptive learning scheme with a two-stage framework to fast adapt the model parameters on the first frame and partial subsequent frames while predicting the results. MDAL learns the general knowledge of instruments and the fast adaptation ability through the video-specific meta-learning paradigm. The added gradient gate excludes the noisy supervision from pseudo masks for dynamic online adaptation on target videos. We demonstrate empirically that MDAL outperforms other state-of-the-art methods on two datasets (including a real-world RAS dataset). The promising performance on ex-vivo scenes also benefits the downstream tasks such as robot-assisted suturing and camera control. Yueming Jin, Bo Lu 0001, Chi-Fai Ng, Qi Dou 0001, Yun-Hui Liu 0001, Pheng-Ann Heng |
ICRA | 2 |
| 2021 | Accurate Grid Keypoint Learning for Efficient Video PredictionabstractVideo prediction methods generally consume substantial computing resources in training and deployment, among which keypoint-based approaches show promising improvement in efficiency by simplifying dense image prediction to light keypoint prediction. However, keypoint locations are often modeled only as continuous coordinates, so noise from semantically insignificant deviations in videos easily disrupt learning stability, leading to inaccurate keypoint modeling. In this paper, we design a new grid keypoint learning framework, aiming at a robust and explainable intermediate keypoint representation for long-term efficient video prediction. We have two major technical contributions. First, we detect keypoints by jumping among candidate locations in our raised grid space and formulate a condensation loss to encourage meaningful keypoints with strong representative capability. Second, we introduce a 2D binary map to represent the detected grid keypoints and then suggest propagating keypoint locations with stochasticity by selecting entries in the discrete grid space, thus preserving the spatial structure of keypoints in the long-term horizon for better future frame generation. Extensive experiments verify that our method outperforms the state-of-the-art stochastic video prediction methods while saves more than 98% of computing resources. We also demonstrate our method on a robotic-assisted surgery dataset with promising results. Our code is available at https://github.com/xjgaocs/Grid-Keypoint-Learning. Yueming Jin, Qi Dou 0001, Chi-Wing Fu, Pheng-Ann Heng |
IROS | 2 |
| 2021 | Domain Adaptive Robotic Gesture Recognition with Unsupervised Kinematic-Visual Data AlignmentabstractAutomated surgical gesture recognition is of great importance in robot-assisted minimally invasive surgery. However, existing methods assume that training and testing data are from the same domain, which suffers from severe performance degradation when a domain gap exists, such as the simulator and real robot. In this paper, we propose a novel unsupervised domain adaptation framework which can simultaneously transfer multi-modality knowledge, i.e., both kinematic and visual data, from simulator to real robot. It remedies the domain gap with enhanced transferable features by using temporal cues in videos, and inherent correlations in multi-modal towards recognizing gesture. Specifically, we first propose a Motion Direction Oriented Kinematics feature alignment (MDO-K) to align kinematics, which exploits temporal continuity to transfer motion directions with smaller gap rather than position values, relieving the adaptation burden. Moreover, we propose a Kinematic and Visual Relation Attention (KV-Relation-ATT) to transfer the co-occurrence signals of kinematics and vision. Such features attended by correlation similarity are more informative for enhancing domain-irreverent of the model. Two feature alignment strategies benefit the model mutually during the end-to-end learning process. We extensively evaluate our method for gesture recognition using DESK dataset with peg transfer procedure. Results show that our approach recovers the performance with great improvement gains, up to 12.91% in Accuracy and 20.16% in F1score without using any annotations in real robot. Xueying Shi, Yueming Jin, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
IROS | 2 |
| 2021 | Source-Free Domain Adaptive Fundus Image Segmentation with Denoised Pseudo-Labeling
Cheng Chen 0013, Quande Liu, Yueming Jin, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (5) | 3 |
| 2021 | Trans-SVNet: Accurate Phase Recognition from Surgical Videos via Hybrid Embedding Aggregation Transformer
Yueming Jin, Yonghao Long 0001, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (4) | 2 |
| 2021 | Efficient Global-Local Memory for Real-Time Instrument Segmentation of Robotic Surgical Video
Jiacheng Wang 0002, Yueming Jin, Liansheng Wang 0002, Shuntian Cai, Pheng-Ann Heng, Harry Qin |
MICCAI (4) | 2 |
| 2021 | Comparative validation of multi-instance instrument segmentation in endoscopy: Results of the ROBUST-MIS 2019 challengeabstractIntraoperative tracking of laparoscopic instruments is often a prerequisite for computer and robotic-assisted interventions. While numerous methods for detecting, segmenting and tracking of medical instruments based on endoscopic video images have been proposed in the literature, key limitations remain to be addressed: Firstly, robustness, that is, the reliable performance of state-of-the-art methods when run on challenging images (e.g. in the presence of blood, smoke or motion artifacts). Secondly, generalization; algorithms trained for a specific intervention in a specific hospital should generalize to other interventions or institutions. In an effort to promote solutions for these limitations, we organized the Robust Medical Instrument Segmentation (ROBUST-MIS) challenge as an international benchmarking competition with a specific focus on the robustness and generalization capabilities of algorithms. For the first time in the field of endoscopic image processing, our challenge included a task on binary segmentation and also addressed multi-instance detection and segmentation. The challenge was based on a surgical data set comprising 10,040 annotated images acquired from a total of 30 surgical procedures from three different types of surgery. The validation of the competing methods for the three tasks (binary segmentation, multi-instance detection and multi-instance segmentation) was performed in three different stages with an increasing domain gap between the training and the test data. The results confirm the initial hypothesis, namely that algorithm performance degrades with an increasing domain gap. While the average detection and segmentation quality of the best-performing algorithms is high, future research should concentrate on detection and segmentation of small, crossing, moving and transparent instrument(s) (parts). Tobias Roß, Annika Reinke, Peter M. Full, Martin Wagner 0001, Hannes Kenngott, Martin Apitz, Hellena Hempe, Diana Mîndroc-Filimon, Patrick Godau, Thuy Nuong Tran, Pierangela Bruno, Pablo Andrés Arbeláez, Guibin Bian, Sebastian Bodenstedt, Jon Lindström Bolmgren, Laura Bravo-Sánchez, Hua-Bin Chen, Cristina González, Pål Halvorsen, Pheng-Ann Heng, Enes Hosgor, Zeng-Guang Hou, Fabian Isensee, Debesh Jha, Tingting Jiang 0001, Yueming Jin, Kadir Kirtaç, Sabrina Kletz, Stefan Leger, Klaus H. Maier-Hein, Zhen-Liang Ni, Michael Riegler 0001, Klaus Schöffmann, Ruohua Shi, Stefanie Speidel, Michael Stenzel, Isabell Twick, Guotai Wang, Jiacheng Wang 0002, Liansheng Wang 0002, Lu Wang 0002, Yan-Jie Zhou, Lei Zhu 0003, Manuel Wiesenfarth, Annette Kopp-Schneider, Beat P. Müller-Stich, Lena Maier-Hein |
Medical Image Anal. | 27 |
| 2021 | Semi-supervised learning with progressive unlabeled data excavation for label-efficient surgical workflow recognition
Xueying Shi, Yueming Jin, Qi Dou 0001, Pheng-Ann Heng |
Medical Image Anal. | 2 |
| 2021 | Anchor-guided online meta adaptation for fast one-Shot instrument segmentation from robotic surgical videos
Yueming Jin, Bo Lu 0001, Chi-Fai Ng, Yun-Hui Liu 0001, Qi Dou 0001, Pheng-Ann Heng |
Medical Image Anal. | 2 |
| 2021 | Temporal Memory Relation Network for Workflow Recognition From Surgical VideoabstractAutomatic surgical workflow recognition is a key component for developing context-aware computer-assisted systems in the operating theatre. Previous works either jointly modeled the spatial features with short fixed-range temporal information, or separately learned visual and long temporal cues. In this paper, we propose a novel end-to-end temporal memory relation network (TMRNet) for relating long-range and multi-scale temporal patterns to augment the present features. We establish a long-range memory bank to serve as a memory cell storing the rich supportive information. Through our designed temporal variation layer, the supportive cues are further enhanced by multi-scale temporal-only convolutions. To effectively incorporate the two types of cues without disturbing the joint learning of spatio-temporal features, we introduce a non-local bank operator to attentively relate the past to the present. In this regard, our TMRNet enables the current feature to view the long-range temporal dependency, as well as tolerate complex temporal extents. We have extensively validated our approach on two benchmark surgical video datasets, M2CAI challenge dataset and Cholec80 dataset. Experimental results demonstrate the outstanding performance of our method, consistently exceeding the state-of-the-art methods by a large margin (e.g., 67.0% v.s. 78.9% Jaccard on Cholec80 dataset). Yueming Jin, Yonghao Long 0001, Cheng Chen 0013, Qi Dou 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 1 |
| 2020 | Automatic Gesture Recognition in Robot-assisted Surgery with Reinforcement Learning and Tree SearchabstractAutomatic surgical gesture recognition is fundamental for improving intelligence in robot-assisted surgery, such as conducting complicated tasks of surgery surveillance and skill evaluation. However, current methods treat each frame individually and produce the outcomes without effective consideration on future information. In this paper, we propose a framework based on reinforcement learning and tree search for joint surgical gesture segmentation and classification. An agent is trained to segment and classify the surgical video in a human-like manner whose direct decisions are re-considered by tree search appropriately. Our proposed tree search algorithm unites the outputs from two designed neural networks, i.e., policy and value network. With the integration of complementary information from distinct models, our framework is able to achieve the better performance than baseline methods using either of the neural networks. For an overall evaluation, our developed approach consistently outperforms the existing methods on the suturing task of JIGSAWS dataset in terms of accuracy, edit score and F1 score. Our study highlights the utilization of tree search to refine actions in reinforcement learning framework for surgical robotic applications. Yueming Jin, Qi Dou 0001, Pheng-Ann Heng |
ICRA | 2 |
| 2020 | A Learning-Driven Framework with Spatial Optimization For Surgical Suture Thread Reconstruction and Autonomous Grasping Under Multiple Topologies and Environmental NoisesabstractSurgical knot tying is one of the most fundamental and important procedures in surgery, and a high-quality knot can significantly benefit the postoperative recovery of the patient. However, a longtime operation may easily cause fatigue to surgeons, especially during the tedious wound closure task. In this paper, we present a vision-based method to automate the suture thread grasping, which is a sub-task in surgical knot tying and an intermediate step between the stitching and looping manipulations. To achieve this goal, the acquisition of a suture's three-dimensional (3D) information is critical. Towards this objective, we adopt a transfer-learning strategy first to fine-tune a pre-trained model by learning the information from large legacy surgical data and images obtained by the onsite equipment. Thus, a robust suture segmentation can be achieved regardless of inherent environment noises. We further leverage a searching strategy with termination policies for a suture's sequence inference based on the analysis of multiple topologies. Exact results of the pixel-level sequence along a suture can be obtained, and they can be further applied for a 3D shape reconstruction using our optimized shortest path approach. The grasping point considering the suturing criterion can be ultimately acquired. Experiments regarding the suture 2D segmentation and ordering sequence inference under environmental noises were extensively evaluated. Results related to the automated grasping operation were demonstrated by simulations in V-REP and by robot experiments using Universal Robot (UR) together with the da Vinci Research Kit (dVRK) adopting our learning-driven framework. Bo Lu 0001, Wei Chen 0068, Yueming Jin, Qi Dou 0001, Henry K. Chu, Pheng-Ann Heng, Yun-Hui Liu 0001 |
IROS | 3 |
| 2020 | Difficulty-Aware Meta-learning for Rare Disease Diagnosis
Xiaomeng Li 0001, Lequan Yu, Yueming Jin, Chi-Wing Fu, Lei Xing 0001, Pheng-Ann Heng |
MICCAI (1) | 3 |
| 2020 | Learning Motion Flows for Semi-supervised Instrument Segmentation from Robotic Surgical Video
Yueming Jin, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (3) | 2 |
| 2020 | Multi-task recurrent convolutional network with correlation loss for surgical video analysis
Yueming Jin, Huaxia Li, Qi Dou 0001, Hao Chen 0011, Harry Qin, Chi-Wing Fu, Pheng-Ann Heng |
Medical Image Anal. | 1 |
| 2019 | Robust Multimodal Brain Tumor Segmentation via Feature Disentanglement and Gated Fusion
Cheng Chen 0013, Qi Dou 0001, Yueming Jin, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
MICCAI (3) | 3 |
| 2019 | Incorporating Temporal Prior from Motion Flow for Instrument Segmentation in Minimally Invasive Surgery Video
Yueming Jin, Keyun Cheng, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (5) | 1 |
| 2018 | SV-RCNet: Workflow Recognition From Surgical Videos Using Recurrent Convolutional NetworkabstractWe propose an analysis of surgical videos that is based on a novel recurrent convolutional network (SV-RCNet), specifically for automatic workflow recognition from surgical videos online, which is a key component for developing the context-aware computer-assisted intervention systems. Different from previous methods which harness visual and temporal information separately, the proposed SV-RCNet seamlessly integrates a convolutional neural network (CNN) and a recurrent neural network (RNN) to form a novel recurrent convolutional architecture in order to take full advantages of the complementary information of visual and temporal features learned from surgical videos. We effectively train the SV-RCNet in an end-to-end manner so that the visual representations and sequential dynamics can be jointly optimized in the learning process. In order to produce more discriminative spatio-temporal features, we exploit a deep residual network (ResNet) and a long short term memory (LSTM) network, to extract visual features and temporal dependencies, respectively, and integrate them into the SV-RCNet. Moreover, based on the phase transition-sensitive predictions from the SV-RCNet, we propose a simple yet effective inference scheme, namely the prior knowledge inference (PKI), by leveraging the natural characteristic of surgical video. Such a strategy further improves the consistency of results and largely boosts the recognition performance. Extensive experiments have been conducted with the MICCAI 2016 Modeling and Monitoring of Computer Assisted Interventions Workflow Challenge dataset and Cholec80 dataset to validate SV-RCNet. Our approach not only achieves superior performance on these two datasets but also outperforms the state-of-the-art methods by a significant margin. Yueming Jin, Qi Dou 0001, Hao Chen 0011, Lequan Yu, Harry Qin, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 1 |
| 2017 | Automated Pulmonary Nodule Detection via 3D ConvNets with Online Sample Filtering and Hybrid-Loss Residual Learning
Qi Dou 0001, Hao Chen 0011, Yueming Jin, Huangjing Lin, Harry Qin, Pheng-Ann Heng |
MICCAI (3) | 3 |
| 2017 | 3D deeply supervised network for automated segmentation of volumetric medical images
Qi Dou 0001, Lequan Yu, Hao Chen 0011, Yueming Jin, Xin Yang 0009, Harry Qin, Pheng-Ann Heng |
Medical Image Anal. | 4 |
| 2017 | Reconfigurable interlocking furnitureabstractReconfigurable assemblies consist of a common set of parts that can be assembled into different forms for use in different situations. Designing these assemblies is a complex problem, since it requires a compatible decomposition of shapes with correspondence across forms, and a planning of well-matched joints to connect parts in each form. This paper presents computational methods as tools to assist the design and construction of reconfigurable assemblies, typically for furniture. There are three key contributions in this work. First, we present the compatible decomposition as a weakly-constrained dissection problem, and derive its solution based on a dynamic bipartite graph to construct parts across multiple forms; particularly, we optimize the parts reuse and preserve the geometric semantics. Second, we develop a joint connection graph to model the solution space of reconfigurable assemblies with part and joint compatibility across different forms. Third, we formulate the backward interlocking and multi-key interlocking models, with which we iteratively plan the joints consistently over multiple forms. We show the applicability of our approach by constructing reconfigurable furniture of various complexities, extend it with recursive connections to generate extensible and hierarchical structures, and fabricate a number of results using 3D printing, 2D laser cutting, and woodworking. Peng Song 0001, Chi-Wing Fu, Yueming Jin, Hongfei Xu, Ligang Liu 0001, Pheng-Ann Heng, Daniel Cohen-Or |
ACM Trans. Graph. | 3 |
| 2016 | 3D Deeply Supervised Network for Automatic Liver Segmentation from CT Volumes
Qi Dou 0001, Hao Chen 0011, Yueming Jin, Lequan Yu, Harry Qin, Pheng-Ann Heng |
MICCAI (2) | 3 |
| 2016 | Non-Local Sparse and Low-Rank Regularization for Structure-Preserving Image SmoothingabstractAbstract This paper presents a new image smoothing method that better preserves prominent structures. Our method is inspired by the recent non‐local image processing techniques on the patch grouping and filtering. Overall, it has three major contributions over previous works. First, we employ the diffusion map as the guidance image to improve the accuracy of patch similarity estimation using the region covariance descriptor. Second, we model structure‐preserving image smoothing as a low‐rank matrix recovery problem, aiming at effectively filtering the texture information in similar patches. Lastly, we devise an objective function, namely the weighted robust principle component analysis (WRPCA), by regularizing the low rank with the weighted nuclear norm and sparsity pursuit with L1norm, and solve this non‐convex WRPCA optimization problem by adopting the alternative direction method of multipliers (ADMM) technique. We experiment our method with a wide variety of images and compare it against several state‐of‐the‐art methods. The results show that our method achieves better structure preservation and texture suppression as compared to other methods. We also show the applicability of our method on several image processing tasks such as edge detection, texture enhancement and seam carving. Lei Zhu 0003, Chi-Wing Fu, Yueming Jin, Mingqiang Wei, Harry Qin, Pheng-Ann Heng |
Comput. Graph. Forum | 3 |