EDBT 2026 Demo / reviewers in the wild / expert
An Wang 0007
dblp:06/4924-7
· DBLP profile ↗
13ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0001-5515-0653ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging Vision and Language for Robust Context-Aware Surgical Point Tracking: The VL-SurgPT Dataset and BenchmarkabstractAccurate point tracking in surgical environments remains challenging due to complex visual conditions, including smoke occlusion, specular reflections, and tissue deformation. While existing surgical tracking datasets provide coordinate information, they lack the semantic context necessary to understand tracking failure mechanisms. We introduce VL-SurgPT, the first large-scale multimodal dataset that bridges visual tracking with textual descriptions of point status in surgical scenes. The dataset comprises 908 in vivo video clips, including 754 for tissue tracking (17,171 annotated points across five challenging scenarios) and 154 for instrument tracking (covering seven instrument types with detailed keypoint annotations). We establish comprehensive benchmarks using eight state-of-the-art tracking methods and propose TG-SurgPT, a text-guided tracking approach that leverages semantic descriptions to improve robustness in visually challenging conditions. Experimental results demonstrate that incorporating point status information significantly improves tracking accuracy and reliability, particularly in adverse visual scenarios where conventional vision-only methods struggle. By bridging visual and linguistic modalities, VL-SurgPT enables the development of context-aware tracking systems crucial for advancing computer-assisted surgery applications that can maintain performance even under challenging intraoperative conditions. Rulin Zhou, Wenlong He, An Wang 0007, Jianhang Zhang, Xuanhui Zeng, Chaowei Zhu, Haijun Hu, Hongliang Ren 0001 |
AAAI | 3 |
| 2026 | EndoControlMag: Robust endoscopic vascular motion magnification with periodic reference resetting and hierarchical tissue-aware dual-mask controlabstractAccurate visualization of subtle vascular dynamics is a knowledge-intensive challenge in minimally invasive surgery. Conventional imaging systems struggle to reveal these imperceptible motions amidst the dynamic complexity of surgical scenes, limiting decision-making reliability. We introduce EndoControlMag , a Lagrangian framework that employs mask-conditioned magnification to selectively enhance vascular motion while preserving the structural integrity of surrounding tissues in endoscopic videos. Our approach integrates two key designs: Periodic Reference Resetting (PRR) , which divides videos into short overlapping clips with dynamically updated reference frames to alleviate error accumulation while maintaining temporal coherence, and Hierarchical Tissue-aware Magnification (HTM) , which combines pretrained visual tracking for accurate vessel localization with dual-mode adaptive softening strategies. HTM employs either motion-based softening that modulates magnification strength proportional to observed tissue displacement, or distance-based exponential decay that simulates biomechanical force attenuation. This strategy enables robust performance across diverse surgical scenarios where motion-based softening excels with complex tissue deformations and distance-based softening provides stability under unreliable optical flow conditions. To validate generality and scalability, we construct EndoVMM24, a benchmark dataset spanning four surgical specialties and diverse intraoperative scenarios. Extensive quantitative metrics, qualitative assessments, and expert surgeon evaluations demonstrate that EndoControlMag significantly outperforms existing methods in magnification accuracy, image quality, and robustness. This work advances engineering informatics for surgical vision by providing a reproducible, context-aware framework that supports reliable decision-making in minimally invasive procedures. The code, dataset, and video results are available at https://cho-haz.github.io/EndoControlMag/ . An Wang 0007, Rulin Zhou, Mengya Xu, Yiru Ye, Longfei Gou, Yiting Chang, Hao Chen 0011, Chwee Ming Lim, Jiankun Wang 0001, Hongliang Ren 0001 |
Adv. Eng. Informatics | 1 |
| 2026 | Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challengeabstractReliable recognition and localization of surgical instruments in endoscopic video recordings are foundational for a wide range of applications in computer- and robot-assisted minimally invasive surgery (RAMIS), including surgical training, skill assessment, and autonomous assistance. However, robust performance under real-world conditions remains a significant challenge. Incorporating surgical context - such as the current procedural phase - has emerged as a promising strategy to improve robustness and interpretability. To address these challenges, we organized the Surgical Procedure Phase, Keypoint, and Instrument Recognition (PhaKIR) sub-challenge as part of the Endoscopic Vision (EndoVis) challenge at MICCAI 2024. We introduced a novel, multi-center dataset comprising thirteen full-length laparoscopic cholecystectomy videos collected from three distinct medical institutions, with unified annotations for three interrelated tasks: surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation. Unlike existing datasets, ours enables joint investigation of instrument localization and procedural context within the same data while supporting the integration of temporal information across entire procedures. We report results and findings in accordance with the BIAS guidelines for biomedical image analysis challenges. The PhaKIR sub-challenge advances the field by providing a unique benchmark for developing temporally aware, context-driven methods in RAMIS and offers a high-quality resource to support future research in surgical scene understanding. Tobias Rueckert, David Rauber, Raphaela Maerkl, Leonard Klausmann, Suemeyye R. Yildiran, Max Gutbrod, Danilo Weber Nunes, Alvaro Fernandez Moreno, Imanol Luengo, Danail Stoyanov, Nicolas Toussaint, Enki Cho, Hyeon Bae Kim, Oh Sung Choo, Ka Young Kim, Seong Tae Kim 0001, Gonçalo Arantes, Kehan Song, Junchen Xiong, Tingyi Lin, Shunsuke Kikuchi, Hiroki Matsuzaki, Atsushi Kouno, João Renato Ribeiro Manesco, João Paulo Papa, Tae-Min Choi, Tae Kyeong Jeong, Oluwatosin Alabi, Tom Vercauteren, Runzhi Wu, Mengya Xu, An Wang 0007, Long Bai 0008, Hongliang Ren 0001, Amine Yamlahi, Jakob Hennighausen, Lena Maier-Hein, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Shu Yang 0004, Yihui Wang 0002, Hao Chen 0011, Santiago Rodríguez, Nicolás Aparicio, Leonardo Manrique, Juan Camilo Lyons, Olivia Hosie, Nicolás Ayobi, Pablo Andrés Arbeláez, Yiping Li 0002, Yasmina Alkhalil, Sahar Nasirihaghighi, Stefanie Speidel, Daniel Rueckert, Hubertus Feußner, Dirk Wilhelm, Christoph Palm |
Medical Image Anal. | 35 |
| 2025 | ETSM: Automating Dissection Trajectory Suggestion and Confidence Map-Based Safety Margin Prediction for Robot-Assisted Endoscopic Submucosal DissectionabstractRobot-assisted Endoscopic Submucosal Dissection (ESD) improves the surgical procedure by providing a more comprehensive view through advanced robotic instruments and bimanual operation, thereby enhancing dissection efficiency and accuracy. Accurate prediction of dissection trajectories is crucial for better decision-making, reducing intraoperative errors, and improving surgical training. Nevertheless, predicting these trajectories is challenging due to variable tumor margins and dynamic visual conditions. To address this issue, we create the ESD Trajectory and Confidence Map-based Safety Margin (ETSM) dataset with 1849 short clips, focusing on submucosal dissection with a dual-arm robotic system. We also introduce a framework that combines optimal dissection trajectory prediction with a confidence map-based safety margin, providing a more secure and intelligent decision-making tool to minimize surgical risks for ESD procedures. Additionally, we propose the Regression-based Confidence Map Prediction Network (RCMNet), which utilizes a regression approach to predict confidence maps for dissection areas, thereby delineating various levels of safety margins. We evaluate our RCMNet using three distinct experimental setups: in-domain evaluation, robustness assessment, and out-of-domain evaluation. Experimental results show that our approach excels in the confidence map-based safety margin prediction task, achieving a mean absolute error (MAE) of only 3.18. To the best of our knowledge, this is the first study to apply a regression approach for visual guidance concerning delineating varying safety levels of dissection areas. Our approach bridges gaps in current research by improving prediction accuracy and enhancing the safety of the dissection process, showing great clinical significance in practice. The dataset and code are available at https://github.com/FrankMOWJ/RCMNet. Mengya Xu, Wenjin Mo, Guankun Wang, Huxin Gao, An Wang 0007, Long Bai 0008, Chaoyang Lyu, Xiaoxiao Yang, Zhen Li 0026, Hongliang Ren 0001 |
ICRA | 5 |
| 2024 | OSSAR: Towards Open-Set Surgical Activity Recognition in Robot-assisted SurgeryabstractIn the realm of automated robotic surgery and computer-assisted interventions, understanding robotic surgical activities stands paramount. Existing algorithms dedicated to surgical activity recognition predominantly cater to pre-defined closed-set paradigms, ignoring the challenges of real-world open-set scenarios. Such algorithms often falter in the presence of test samples originating from classes unseen during training phases. To tackle this problem, we introduce an innovative Open-Set Surgical Activity Recognition (OSSAR) framework. Our solution leverages the hyperspherical reciprocal point strategy to enhance the distinction between known and unknown classes in the feature space. Additionally, we address the issue of over-confidence in the closed set by refining model calibration, avoiding misclassification of unknown classes as known ones. To support our assertions, we establish an open-set surgical activity benchmark utilizing the public JIGSAWS dataset. Besides, we also collect a novel dataset on endoscopic submucosal dissection for surgical activity tasks. Extensive comparisons and ablation experiments on these datasets demonstrate the significant outperformance of our method over existing state-of-the-art approaches. Our proposed solution can effectively address the challenges of real-world surgical scenarios. Our code is publicly accessible at github.com/longbai1006/OSSAR. Long Bai 0008, Guankun Wang, Jie Wang 0097, Xiaoxiao Yang, Huxin Gao, An Wang 0007, Mobarakol Islam, Hongliang Ren 0001 |
ICRA | 7 |
| 2024 | EndoDAC: Efficient Adapting Foundation Model for Self-Supervised Depth Estimation from Any Endoscopic Camera
Beilei Cui, Mobarakol Islam, Long Bai 0008, An Wang 0007, Hongliang Ren 0001 |
MICCAI (6) | 4 |
| 2024 | Curriculum-Based Augmented Fourier Domain Adaptation for Robust Medical Image SegmentationabstractAccurate and robust medical image segmentation is fundamental and crucial for enhancing the autonomy of computer-aided diagnosis and intervention systems. Medical data collection normally involves different scanners, protocols, and populations, making domain adaptation (DA) a highly demanding research field to alleviate model degradation in the deployment site. To preserve the model performance across multiple testing domains, this work proposes the Curriculum-based Augmented Fourier Domain Adaptation (Curri-AFDA) for robust medical image segmentation. In particular, our curriculum learning strategy is based on the causal relationship of a model under different levels of data shift in the deployment phase, where the higher the shift is, the harder to recognize the variance. Considering this, we progressively introduce more amplitude information from the target domain to the source domain in the frequency space during the curriculum-style training to smoothly schedule the semantic knowledge transfer in an easier-to-harder manner. Besides, we incorporate the training-time chained augmentation mixing to help expand the data distributions while preserving the domain-invariant semantics, which is beneficial for the acquired model to be more robust and generalize better to unseen domains. Extensive experiments on two segmentation tasks of Retina and Nuclei collected from multiple sites and scanners suggest that our proposed method yields superior adaptation and generalization performance. Meanwhile, our approach proves to be more robust under various corruption types and increasing severity levels. In addition, we show our method is also beneficial in the domain-adaptive classification task with skin lesion datasets. The code is available at https://github.com/lofrienger/Curri-AFDA.Note to Practitioners—Medical image segmentation is key to improving computer-assisted diagnosis and intervention autonomy. However, due to domain gaps between different medical sites, deep learning-based segmentation models frequently encounter performance degradation when deployed in a novel domain. Moreover, model robustness is also highly expected to mitigate the effects of data corruption. Considering all these demanding yet practical needs to automate medical applications and benefit healthcare, we propose the Curriculum-based Fourier Domain Adaptation (Curri-AFDA) for medical image segmentation. Extensive experiments on two segmentation tasks with cross-domain datasets show the consistent superiority of our method regarding adaptation and generalization on multiple testing domains and robustness against synthetic corrupted data. Besides, our approach is independent of image modalities because its efficacy does not rely on modality-specific characteristics. In addition, we demonstrate the benefit of our method for image classification besides segmentation in the ablation study. Therefore, our method can potentially be applied in many medical applications and yield improved performance. Future works may be extended by exploring the integration of curriculum learning regime with Fourier domain amplitude fusion in the testing time rather than in the training time like this work and most other existing domain adaptation works. An Wang 0007, Mobarakol Islam, Mengya Xu, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | An Efficient MLP-Based Point-Guided Segmentation Network for Ore Images With Ambiguous BoundaryabstractThe precise segmentation of ore images is critical to the successful execution of the beneficiation process. Due to the homogeneous appearance of the ores, which leads to low contrast and unclear boundaries, accurate segmentation becomes challenging, and recognition becomes problematic. This article proposes a lightweight framework based on multilayer perceptron (MLP), which focuses on solving the problem of edge blurring. Specifically, we introduce a lightweight backbone better suited for efficiently extracting low-level features. Besides, we design a feature pyramid network consisting of two MLP structures that balance local and global information, thus, enhancing detection accuracy. Furthermore, we propose a novel loss function that guides the prediction points to match the instance edge points to achieve clear object boundaries. We have conducted extensive experiments to validate the efficacy of our proposed method. Our approach achieves a remarkable processing speed of over 27 frames per second with a model size of only 73 MB. Moreover, our method delivers a consistently high level of accuracy, with impressive performance scores of 60.4 and 48.9 in$AP_{50}^{\text{box}}$and$AP_{50}^{\text{mask}}$, respectively, as compared with the currently available state-of-the-art techniques, when tested on the ORE image dataset. Guodong Sun 0002, Yuting Peng 0001, Mengya Xu, An Wang 0007, Hongliang Ren 0001, Yang Zhang 0053 |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | Generalizing Surgical Instruments Segmentation to Unseen Domains with One-to-Many SynthesisabstractDespite their impressive performance in various surgical scene understanding tasks, deep learning-based methods are frequently hindered from deploying to real-world surgical applications for various causes. Particularly, data collection, annotation, and domain shift in-between sites and patients are the most common obstacles. In this work, we mitigate data-related issues by efficiently leveraging minimal source images to generate synthetic surgical instrument segmentation datasets and achieve outstanding generalization performance on unseen real domains. Specifically, in our framework, only one background tissue image and at most three images of each foreground instrument are taken as the seed images. These source images are extensively transformed and employed to build up the foreground and background image pools, from which randomly sampled tissue and instrument images are composed with multiple blending techniques to generate new surgical scene images. Besides, we introduce hybrid training-time augmentations to diversify the training data further. Extensive evaluation on three real-world datasets, i.e., Endo2017, Endo2018, and RoboTool, demonstrates that our one-to-many synthetic surgical instruments datasets generation and segmentation framework can achieve encouraging performance compared with training with real data. Notably, on the RoboTool dataset, where a more significant domain gap exists, our framework shows its superiority of generalization by a considerable margin. We expect that our inspiring results will attract research attention to improving model generalization with data synthesizing. An Wang 0007, Mobarakol Islam, Mengya Xu, Hongliang Ren 0001 |
IROS | 1 |
| 2023 | LLCaps: Learning to Illuminate Low-Light Capsule Endoscopy with Curved Wavelet Attention and Reverse Diffusion
Long Bai 0008, Tong Chen 0011, Yanan Wu 0003, An Wang 0007, Mobarakol Islam, Hongliang Ren 0001 |
MICCAI (10) | 4 |
| 2023 | Rectifying Noisy Labels with Sequential Prior: Multi-scale Temporal Feature Affinity Learning for Robust Video Segmentation
Beilei Cui, Minqing Zhang, Mengya Xu, An Wang 0007, Wu Yuan 0001, Hongliang Ren 0001 |
MICCAI (9) | 4 |
| 2023 | S2ME: Spatial-Spectral Mutual Teaching and Ensemble Learning for Scribble-Supervised Polyp Segmentation
An Wang 0007, Mengya Xu, Yang Zhang 0053, Mobarakol Islam, Hongliang Ren 0001 |
MICCAI (1) | 1 |
| 2022 | Rethinking Surgical Instrument Segmentation: A Background Image Can Be All You Need
An Wang 0007, Mobarakol Islam, Mengya Xu, Hongliang Ren 0001 |
MICCAI (8) | 1 |