EDBT 2026 Demo / reviewers in the wild / expert
Zhen Chen 0018
dblp:11/1266-18
· DBLP profile ↗
16ranked-venue papers
5as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EndoChat: Grounded multimodal large language model for endoscopic surgeryabstractRecently, Multimodal Large Language Models (MLLMs) have demonstrated their immense potential in computer-aided diagnosis and decision-making. In the context of robotic-assisted surgery, MLLMs can serve as effective tools for surgical training and guidance. However, there is still a deficiency of MLLMs specialized for surgical scene understanding in endoscopic procedures. To this end, we present EndoChat, an MLLM tailored to address various dialogue paradigms and subtasks in understanding endoscopic procedures. To train our EndoChat, we construct the Surg-396K dataset through a novel pipeline that systematically extracts surgical information and generates structured annotations based on large-scale endoscopic surgery datasets. Furthermore, we introduce a multi-scale visual token interaction mechanism and a visual contrast-based reasoning mechanism to enhance the model's representation learning and reasoning capabilities. Our model achieves state-of-the-art performance across five dialogue paradigms and seven surgical scene understanding tasks. Additionally, we conduct evaluations with professional surgeons, who provide positive feedback on the majority of conversation cases generated by EndoChat. Overall, these results demonstrate that EndoChat has the potential to advance training and automation in robotic-assisted surgery. Our dataset and model are publicly available at https://github.com/gkw0010/EndoChat. Guankun Wang, Long Bai 0008, Kun Yuan 0004, Zhen Li 0026, Tianxu Jiang, Xiting He, Jinlin Wu, Zhen Chen 0018, Zhen Lei 0001, Hongbin Liu 0001, Fan Zhang 0016, Nicolas Padoy, Nassir Navab, Hongliang Ren 0001 |
Medical Image Anal. | 9 |
| 2026 | SA-Person: Text-Based Person Retrieval With Scene-Aware Re-Ranking
Yingjia Xu, Jinlin Wu, Daming Gao, Zhen Chen 0018, Yang Yang 0062, Min Cao 0005, Mang Ye, Zhen Lei 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Universal Domain Adaptive Object Detection via Dual Probabilistic AlignmentabstractDomain Adaptive Object Detection (DAOD) transfers knowledge from a labeled source domain to an unannotated target domain under closed-set assumption. Universal DAOD (UniDAOD) extends DAOD to handle open-set, partial-set, and closed-set domain adaptation. In this paper, we first unveil two issues: domain-private category alignment is crucial for global-level features, and the domain probability heterogeneity of features across different levels. To address these issues, we propose a novel Dual Probabilistic Alignment (DPA) framework to model domain probability as Gaussian distribution, enabling the heterogeneity domain distribution sampling and measurement. The DPA consists of three tailored modules: the Global-level Domain Private Alignment (GDPA), the Instance-level Domain Shared Alignment (IDSA), and the Private Class Constraint (PCC). GDPA utilizes the global-level sampling to mine domain-private category samples and calculate alignment weight through a cumulative distribution function to address the global-level private category alignment. IDSA utilizes instance-level sampling to mine domain-shared category samples and calculates alignment weight through Gaussian distribution to conduct the domain-shared category domain alignment to address the feature heterogeneity. The PCC aggregates domain-private category centroids between feature and probability spaces to mitigate negative transfer. Extensive experiments demonstrate that our DPA outperforms state-of-the-art UniDAOD and DAOD methods across various datasets and scenarios, including open, partial, and closed sets. Yuanfan Zheng, Jinlin Wu, Wuyang Li, Zhen Chen 0018 |
AAAI | 4 |
| 2025 | PRO-VPT: Distribution-Adaptive Visual Prompt Tuning via Prompt RelocationabstractVisual prompt tuning (VPT), i.e., fine-tuning some lightweight prompt tokens, provides an efficient and effective approach for adapting pre-trained models to various downstream tasks. However, most prior art indiscriminately uses a fixed prompt distribution across different tasks, neglecting the importance of each block varying depending on the task. In this paper, we introduce adaptive distribution optimization (ADO) by tackling two key questions: (1) How to appropriately and formally define ADO, and (2) How to design an adaptive distribution strategy guided by this definition? Through empirical analysis, we first confirm that properly adjusting the distribution significantly improves VPT performance, and further uncover a key insight that a nested relationship exists between ADO and VPT. Based on these findings, we propose a new VPT framework, termed PRO-VPT (iterative Prompt RelOcation-based VPT), which adaptively adjusts the distribution built upon a nested optimization formulation. Specifically, we develop a prompt relocation strategy derived from this formulation, comprising two steps: pruning idle prompts from prompt-saturated blocks, followed by allocating these prompts to the most prompt-needed blocks. By iteratively performing prompt relocation and VPT, our proposal can adaptively learn the optimal prompt distribution in a nested optimization-based manner, thereby unlocking the full potential of VPT. Extensive experiments demonstrate that our proposal significantly outperforms advanced VPT methods, e.g., PRO-VPT surpasses VPT by 1.6 pp and 2.0 pp average accuracy, leading prompt-based methods to state-of-the-art performance on VTAB-1k and FGVC benchmarks. The code is available at https://github.com/ckshang/PRO-VPT. Chikai Shang, Mengke Li 0001, Yiqun Zhang 0006, Zhen Chen 0018, Jinlin Wu, Fangqing Gu, Yang Lu 0009, Yiu-Ming Cheung |
ICCV | 4 |
| 2025 | SurgPLAN++: Universal Surgical Phase Localization Network for Online and Offline InferenceabstractSurgical phase recognition is critical for assisting surgeons in understanding surgical videos. Existing studies focused more on online surgical phase recognition, by leveraging preceding frames to predict the current frame. Despite great progress, they formulated the task as a series of frame-wise classification, which resulted in a lack of global context of the entire procedure and incoherent predictions. Moreover, besides online analysis, accurate offline surgical phase recognition is also in significant clinical need for retrospective analysis, and existing online algorithms do not fully analyze the entire video, thereby limiting accuracy in offline analysis. To over-come these challenges and enhance both online and offline inference capabilities, we propose a universal Surgical Phase LocalizAtion Network, named SurgPLAN++, with the principle of temporal detection. To ensure a global understanding of the surgical procedure, we devise a phase localization strategy for SurgPLAN ++ to predict phase segments across the entire video through phase proposals. For online analysis, to generate high-quality phase proposals, SurgPLAN++ incorporates a data augmentation strategy to extend the streaming video into a pseudo-complete video through mirroring, center-duplication, and down-sampling. For offline analysis, SurgPLAN++ capi-talizes on its global phase prediction framework to continu-ously refine preceding predictions during each online inference step, thereby significantly improving the accuracy of phase recognition. We perform extensive experiments to validate the effectiveness, and our SurgPLAN++ achieves remarkable performance in both online and offline modes, which outper-forms state-of-the-art methods. The source code is available at https://github.com/franciszchenlSurgPLAN-Plus. Zhen Chen 0018, Xingjian Luo, Jinlin Wu, Long Bai 0008, Zhen Lei 0001, Hongliang Ren 0001, Sébastien Ourselin, Hongbin Liu 0001 |
ICRA | 1 |
| 2025 | Advancing Dense Endoscopic Reconstruction with Gaussian Splatting-Driven Surface Normal-Aware Tracking and MappingabstractSimultaneous Localization and Mapping (SLAM) is essential for precise surgical interventions and robotic tasks in minimally invasive procedures. While recent advancements in 3D Gaussian Splatting (3DGS) have improved SLAM with high-quality novel view synthesis and fast rendering, these systems struggle with accurate depth and surface reconstruction due to multi-view inconsistencies. Simply incorporating SLAM and 3DGS leads to mismatches between the reconstructed frames. In this work, we present Endo-2DTAM, a real-time endoscopic SLAM system with 2D Gaussian Splatting (2DGS) to address these challenges. Endo-2DTAM incorporates a surface normal-aware pipeline, which consists of tracking, mapping, and bundle adjustment modules for geometrically accurate reconstruction. Our robust tracking module combines point-topoint and point-to-plane distance metrics, while the mapping module utilizes normal consistency and depth distortion to enhance surface reconstruction quality. We also introduce a pose-consistent strategy for efficient and geometrically coherent keyframe sampling. Extensive experiments on public endoscopic datasets demonstrate that Endo-2DTAM achieves an RMSE of$1.87 \pm 0.63 \mathbf{m m}$for depth reconstruction of surgical scenes while maintaining computationally efficient tracking, high-quality visual appearance, and real-time rendering. Our code will be released at github.com/lastbasket/Endo-2DTAM. Yiming Huang 0007, Beilei Cui, Long Bai 0008, Zhen Chen 0018, Jinlin Wu, Zhen Li 0026, Hongbin Liu 0001, Hongliang Ren 0001 |
ICRA | 4 |
| 2025 | PitVis-2023 challenge: Workflow recognition in videos of endoscopic pituitary surgeryabstractThe field of computer vision applied to videos of minimally invasive surgery is ever-growing. Workflow recognition pertains to the automated recognition of various aspects of a surgery, including: which surgical steps are performed; and which surgical instruments are used. This information can later be used to assist clinicians when learning the surgery or during live surgery. The Pituitary Vision (PitVis) 2023 Challenge tasks the community to step and instrument recognition in videos of endoscopic pituitary surgery. This is a particularly challenging task when compared to other minimally invasive surgeries due to: the smaller working space, which limits and distorts vision; and higher frequency of instrument and step switching, which requires more precise model predictions. Participants were provided with 25-videos, with results presented at the MICCAI-2023 conference as part of the Endoscopic Vision 2023 Challenge in Vancouver, Canada, on 08-Oct-2023. There were 18-submissions from 9-teams across 6-countries, using a variety of deep learning models. The top performing model for step recognition utilised a transformer based architecture, uniquely using an autoregressive decoder with a positional encoding input. The top performing model for instrument recognition utilised a spatial encoder followed by a temporal encoder, which uniquely used a 2-layer temporal architecture. In both cases, these models outperformed purely spatial based models, illustrating the importance of sequential and temporal information. This PitVis-2023 therefore demonstrates state-of-the-art computer vision models in minimally invasive surgery are transferable to a new dataset. Benchmark results are provided in the paper, and the dataset is publicly available at: https://doi.org/10.5522/04/26531686. Adrito Das, Danyal Z. Khan, Dimitris Psychogyios, John G. Hanrahan, Francisco Vasconcelos 0001, You Pang, Zhen Chen 0018, Jinlin Wu, Xiaoyang Zou, Guoyan Zheng, Abdul Qayyum 0002, Moona Mazher, Muhammad Imran Razzak, Tianbin Li, Jin Ye 0002, Junjun He, Szymon Plotka, Joanna Kaleta, Amine Yamlahi, Antoine Jund, Patrick Godau, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Dominik Rivoir, Stefanie Speidel, Alejandra Pérez, Santiago Rodríguez, Pablo Andrés Arbeláez, Danail Stoyanov, Hani J. Marcus, Sophia Bano |
Medical Image Anal. | 8 |
| 2024 | SurgFC: Multimodal Surgical Function Calling Framework on the Demand of SurgeonsabstractThe surgical intervention is crucial to patient healthcare, and many studies have developed advanced algorithms to provide understanding and decision-making assistance for surgeons. Despite great progress, these algorithms are developed for a single specific task and scenario, and in practice require the manual combination of different functions, thus limiting the applicability. Thus, an intelligent surgical assistant is expected to accurately understand the surgeon’s intentions and accordingly conduct the specific tasks to support the surgical process. In this work, by improving advanced multimodal large language models (MLLMs), we propose a multimodal Surgical Function Calling (SurgFC) framework that can accurately understand the surgeon’s intention and complete a series of surgical understanding tasks, e.g., surgical scene analysis, surgical instrument detection, and segmentation on demand. Specifically, to achieve superior surgical multimodal understanding, we devise a mixture-of-projectors (MOP) module to align the surgical MLLM in SurgFC to balance the natural and surgical knowledge. Moreover, we devise a surgical Function-Calling Tuning strategy to enable the SurgFC to understand surgical intentions, and thus make a series of surgical function calls on demand to meet the needs of the surgeons. Extensive experiments on neurosurgery data confirm that our SurgFC can understand the surgeon’s intention more accurately than the existing MLLM, resulting in overwhelming performance in textual analysis and visual tasks. The source code is available at https://github.com/franciszchen/SurgFC. Zhen Chen 0018, Xingjian Luo, Jinlin Wu, Danny T. M. Chan, Zhen Lei 0001, Sébastien Ourselin, Hongbin Liu 0001 |
BIBM | 1 |
| 2024 | SurgBox: Agent-Driven Operating Room Sandbox with Surgery CopilotabstractSurgical interventions, particularly in neurology, represent complex and high-stakes scenarios that impose substantial cognitive burdens on surgical teams. Although deliberate education and practice can enhance cognitive capabilities, surgical training opportunities remain limited due to patient safety concerns. To address these cognitive challenges in surgical training and operation, we propose SurgBox, an agent-driven sandbox framework to systematically enhance the cognitive capabilities of surgeons in immersive surgical simulations. Specifically, our SurgBox leverages large language models (LLMs) with tailored Retrieval-Augmented Generation (RAG) to authentically replicate various surgical roles, enabling realistic training environments for deliberate practice. In particular, we devise Surgery Copilot, an AI-driven assistant to actively coordinate the surgical information stream and support clinical decision-making, thereby diminishing the cognitive workload of surgical teams during surgery. By incorporating a novel Long-Short Memory mechanism, our Surgery Copilot can effectively balance immediate procedural assistance with comprehensive surgical knowledge. Extensive experiments using real neurosurgical procedure records validate our SurgBox framework in both enhancing surgical cognitive capabilities and supporting clinical decision-making. By providing an integrated solution for training and operational support to address cognitive challenges, our SurgBox framework advances surgical education and practice, potentially transforming surgical outcomes and healthcare quality. The code is available at https://github.com/franciszchen/SurgBox. Jinlin Wu, Xusheng Liang, Xuexue Bai, Zhen Chen 0018 |
IEEE Big Data | 4 |
| 2024 | PWISeg: Weakly-Supervised Surgical Instrument Instance SegmentationabstractAI-assisted operating room scene understanding is essential for the next generation of surgical interventions. Surgical instrument localization plays an important role in this context. However, existing instrument localization methods primarily focus on surgical instrument localization in endoscopy images and struggle with occlusions in broader operating room scenarios. In this work, we propose a weakly supervised instance segmentation framework, Pixel-driven Weakly-supervised Instance Segmentation (PWISeg), to solve the occluded instrument localization with low-cost annotations. Specifically, We utilize the projection relationship between the bounding box and the surgical instrument mask as a supervision signal to train PWISeg to predict coarse masks of surgical instruments. Then, we use the annotation of a few pixel points to train PWISeg to predict accurate masks of surgical instruments. To extensively validate the effectiveness, we collect and release a high-quality dataset, Surg-Inst that covers real-world hard cases of overlapping, dense placement, and various levels of instrument occlusion. Experiments demonstrate that our PWISeg achieves a remarkable performance advantage over state-of-the-art methods on both Surg-Inst and public HOSPI-Tools datasets. Zhen Sun 0001, Huan Xu 0003, Jinlin Wu, Zhen Chen 0018, Hongbin Liu 0001, Zhen Lei 0001 |
ICIP | 4 |
| 2024 | ASI-Seg: Audio-Driven Surgical Instrument Segmentation with Surgeon Intention UnderstandingabstractSurgical instrument segmentation is crucial in surgical scene understanding, thereby facilitating surgical safety. Existing algorithms directly detected all instruments of predefined categories in the input image, lacking the capability to segment specific instruments according to the surgeon’s intention. During different stages of surgery, surgeons exhibit varying preferences and focus toward different surgical instruments. Therefore, an instrument segmentation algorithm that adheres to the surgeon’s intention can minimize distractions from irrelevant instruments and assist surgeons to a great extent. The recent Segment Anything Model (SAM) reveals the capability to segment objects following prompts, but the manual annotations for prompts are impractical during the surgery. To address these limitations in operating rooms, we propose an audio-driven surgical instrument segmentation framework, named ASI-Seg, to accurately segment the required surgical instruments by parsing the audio commands of surgeons. Specifically, we propose an intention-oriented multimodal fusion to interpret the segmentation intention from audio commands and retrieve relevant instrument details to facilitate segmentation. Moreover, to guide our ASI-Seg segment of the required surgical instruments, we devise a contrastive learning prompt encoder to effectively distinguish the required instruments from the irrelevant ones. Therefore, our ASI-Seg promotes the workflow in the operating rooms, thereby providing targeted support and reducing the cognitive load on surgeons. Extensive experiments are performed to validate the ASI-Seg framework, which reveals remarkable advantages over classical state-of-the-art and medical SAMs in both semantic segmentation and intention-oriented segmentation. The source code is available at https://github.com/Zonmgin-Zhang/ASI-Seg. Zhen Chen 0018, Zongming Zhang, Wenwu Guo, Xingjian Luo, Long Bai 0008, Jinlin Wu, Hongliang Ren 0001, Hongbin Liu 0001 |
IROS | 1 |
| 2024 | EndoUIC: Promptable Diffusion Transformer for Unified Illumination Correction in Capsule Endoscopy
Long Bai 0008, Tong Chen 0011, Qiaozhi Tan, Wan Jun Nah, Yanheng Li 0002, Zhicheng He 0010, Sishen Yuan, Zhen Chen 0018, Jinlin Wu, Mobarakol Islam, Zhen Li 0026, Hongbin Liu 0001, Hongliang Ren 0001 |
MICCAI (7) | 8 |
| 2024 | PANS: Probabilistic Airway Navigation System for Real-Time Robust Bronchoscope Localization
Qingyao Tian, Zhen Chen 0018, Huai Liao, Bingyu Yang, Lujie Li, Hongbin Liu 0001 |
MICCAI (6) | 2 |
| 2024 | Transforming Surgical Interventions with Embodied Intelligence for Ultrasound Robotics
Huan Xu 0003, Jinlin Wu, Guanglin Cao, Zhen Chen 0018, Zhen Lei 0001, Hongbin Liu 0001 |
MICCAI (6) | 4 |
| 2023 | Temporal Action-aware Network with Sequence Regularization for Phase RecognitionabstractTo assist surgeons in the operating theatre, surgical phase recognition is critical for developing computer-assisted surgical systems, which requires comprehensive understanding of surgical videos. Although existing studies made great progress, there are still two significant limitations worthy of improvement. First, due to the compromise of resource consumption, framewise visual features are extracted by 2D networks and disregard spatial and temporal knowledge of surgical actions, which hinders subsequent inter-frame modeling for phase prediction. Second, these works simply utilize ordinary classification loss with one-hot phase labels to optimize the phase predictions, and cannot fully explore surgical videos under inadequate supervision. To overcome these two limitations, we propose a Surgical Temporal Action-aware Network with sequence Regularization, named STAR-Net, to recognize surgical phases more accurately from input videos. Specifically, we propose an efficient multi-scale surgical temporal action (MS-STA) module, which integrates visual features with spatial and temporal knowledge of surgical actions at the cost of 2D networks. Moreover, we devise the dual-classifier sequence regularization (DSR) to facilitate the training of STAR-Net by the sequence guidance of an auxiliary classifier with a smaller capacity. Our STAR-Net with MS-STA and DSR can exploit visual features of surgical actions with effective regularization, thereby leading to the superior performance of surgical phase recognition. Extensive experiments on a large-scale gastrectomy surgery dataset and the public Cholec80 benchmark prove that our STAR-Net significantly outperforms state-of-the-arts of surgical phase recognition. Zhen Chen 0018, Yuhao Zhai, Jun Zhang 0093, Jinqiao Wang |
BIBM | 1 |
| 2023 | Surgical Video Captioning with Mutual-Modal Concept Alignment
Zhen Chen 0018, Qingyu Guo, Leo K. T. Yeung, Danny T. M. Chan, Zhen Lei 0001, Hongbin Liu 0001, Jinqiao Wang |
MICCAI (9) | 1 |