EDBT 2026 Demo / reviewers in the wild / expert
Mengya Xu
dblp:289/1917
· DBLP profile ↗
24ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0002-4338-7079ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 16 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 7 since 2021Systems, architecture and hardware · 6 · 4 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EndoControlMag: Robust endoscopic vascular motion magnification with periodic reference resetting and hierarchical tissue-aware dual-mask controlabstractAccurate visualization of subtle vascular dynamics is a knowledge-intensive challenge in minimally invasive surgery. Conventional imaging systems struggle to reveal these imperceptible motions amidst the dynamic complexity of surgical scenes, limiting decision-making reliability. We introduce EndoControlMag , a Lagrangian framework that employs mask-conditioned magnification to selectively enhance vascular motion while preserving the structural integrity of surrounding tissues in endoscopic videos. Our approach integrates two key designs: Periodic Reference Resetting (PRR) , which divides videos into short overlapping clips with dynamically updated reference frames to alleviate error accumulation while maintaining temporal coherence, and Hierarchical Tissue-aware Magnification (HTM) , which combines pretrained visual tracking for accurate vessel localization with dual-mode adaptive softening strategies. HTM employs either motion-based softening that modulates magnification strength proportional to observed tissue displacement, or distance-based exponential decay that simulates biomechanical force attenuation. This strategy enables robust performance across diverse surgical scenarios where motion-based softening excels with complex tissue deformations and distance-based softening provides stability under unreliable optical flow conditions. To validate generality and scalability, we construct EndoVMM24, a benchmark dataset spanning four surgical specialties and diverse intraoperative scenarios. Extensive quantitative metrics, qualitative assessments, and expert surgeon evaluations demonstrate that EndoControlMag significantly outperforms existing methods in magnification accuracy, image quality, and robustness. This work advances engineering informatics for surgical vision by providing a reproducible, context-aware framework that supports reliable decision-making in minimally invasive procedures. The code, dataset, and video results are available at https://cho-haz.github.io/EndoControlMag/ . An Wang 0007, Rulin Zhou, Mengya Xu, Yiru Ye, Longfei Gou, Yiting Chang, Hao Chen 0011, Chwee Ming Lim, Jiankun Wang 0001, Hongliang Ren 0001 |
Adv. Eng. Informatics | 3 |
| 2026 | Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challengeabstractReliable recognition and localization of surgical instruments in endoscopic video recordings are foundational for a wide range of applications in computer- and robot-assisted minimally invasive surgery (RAMIS), including surgical training, skill assessment, and autonomous assistance. However, robust performance under real-world conditions remains a significant challenge. Incorporating surgical context - such as the current procedural phase - has emerged as a promising strategy to improve robustness and interpretability. To address these challenges, we organized the Surgical Procedure Phase, Keypoint, and Instrument Recognition (PhaKIR) sub-challenge as part of the Endoscopic Vision (EndoVis) challenge at MICCAI 2024. We introduced a novel, multi-center dataset comprising thirteen full-length laparoscopic cholecystectomy videos collected from three distinct medical institutions, with unified annotations for three interrelated tasks: surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation. Unlike existing datasets, ours enables joint investigation of instrument localization and procedural context within the same data while supporting the integration of temporal information across entire procedures. We report results and findings in accordance with the BIAS guidelines for biomedical image analysis challenges. The PhaKIR sub-challenge advances the field by providing a unique benchmark for developing temporally aware, context-driven methods in RAMIS and offers a high-quality resource to support future research in surgical scene understanding. Tobias Rueckert, David Rauber, Raphaela Maerkl, Leonard Klausmann, Suemeyye R. Yildiran, Max Gutbrod, Danilo Weber Nunes, Alvaro Fernandez Moreno, Imanol Luengo, Danail Stoyanov, Nicolas Toussaint, Enki Cho, Hyeon Bae Kim, Oh Sung Choo, Ka Young Kim, Seong Tae Kim 0001, Gonçalo Arantes, Kehan Song, Junchen Xiong, Tingyi Lin, Shunsuke Kikuchi, Hiroki Matsuzaki, Atsushi Kouno, João Renato Ribeiro Manesco, João Paulo Papa, Tae-Min Choi, Tae Kyeong Jeong, Oluwatosin Alabi, Tom Vercauteren, Runzhi Wu, Mengya Xu, An Wang 0007, Long Bai 0008, Hongliang Ren 0001, Amine Yamlahi, Jakob Hennighausen, Lena Maier-Hein, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Shu Yang 0004, Yihui Wang 0002, Hao Chen 0011, Santiago Rodríguez, Nicolás Aparicio, Leonardo Manrique, Juan Camilo Lyons, Olivia Hosie, Nicolás Ayobi, Pablo Andrés Arbeláez, Yiping Li 0002, Yasmina Alkhalil, Sahar Nasirihaghighi, Stefanie Speidel, Daniel Rueckert, Hubertus Feußner, Dirk Wilhelm, Christoph Palm |
Medical Image Anal. | 34 |
| 2025 | ETSM: Automating Dissection Trajectory Suggestion and Confidence Map-Based Safety Margin Prediction for Robot-Assisted Endoscopic Submucosal DissectionabstractRobot-assisted Endoscopic Submucosal Dissection (ESD) improves the surgical procedure by providing a more comprehensive view through advanced robotic instruments and bimanual operation, thereby enhancing dissection efficiency and accuracy. Accurate prediction of dissection trajectories is crucial for better decision-making, reducing intraoperative errors, and improving surgical training. Nevertheless, predicting these trajectories is challenging due to variable tumor margins and dynamic visual conditions. To address this issue, we create the ESD Trajectory and Confidence Map-based Safety Margin (ETSM) dataset with 1849 short clips, focusing on submucosal dissection with a dual-arm robotic system. We also introduce a framework that combines optimal dissection trajectory prediction with a confidence map-based safety margin, providing a more secure and intelligent decision-making tool to minimize surgical risks for ESD procedures. Additionally, we propose the Regression-based Confidence Map Prediction Network (RCMNet), which utilizes a regression approach to predict confidence maps for dissection areas, thereby delineating various levels of safety margins. We evaluate our RCMNet using three distinct experimental setups: in-domain evaluation, robustness assessment, and out-of-domain evaluation. Experimental results show that our approach excels in the confidence map-based safety margin prediction task, achieving a mean absolute error (MAE) of only 3.18. To the best of our knowledge, this is the first study to apply a regression approach for visual guidance concerning delineating varying safety levels of dissection areas. Our approach bridges gaps in current research by improving prediction accuracy and enhancing the safety of the dissection process, showing great clinical significance in practice. The dataset and code are available at https://github.com/FrankMOWJ/RCMNet. Mengya Xu, Wenjin Mo, Guankun Wang, Huxin Gao, An Wang 0007, Long Bai 0008, Chaoyang Lyu, Xiaoxiao Yang, Zhen Li 0026, Hongliang Ren 0001 |
ICRA | 1 |
| 2025 | Surgical Action Planning with Large Language Models
Mengya Xu, Zhongzhen Huang, Xiaofan Zhang 0002, Qi Dou 0001 |
MICCAI (9) | 1 |
| 2025 | CSAP-Assist: Instrument-Agent Dialogue Empowered Vision-Language Models for Collaborative Surgical Action Planning
Mengya Xu, Qi Dou 0001 |
MICCAI (9) | 2 |
| 2024 | PitVQA: Image-Grounded Text Embedding LLM for Visual Question Answering in Pituitary Surgery
Runlong He, Mengya Xu, Adrito Das, Danyal Z. Khan, Sophia Bano, Hani J. Marcus, Danail Stoyanov, Matthew J. Clarkson, Mobarakol Islam |
MICCAI (6) | 2 |
| 2024 | Endo-4DGS: Endoscopic Monocular Scene Reconstruction with 4D Gaussian Splatting
Yiming Huang 0007, Beilei Cui, Long Bai 0008, Mengya Xu, Mobarakol Islam, Hongliang Ren 0001 |
MICCAI (6) | 5 |
| 2024 | AMIFN: Aspect-guided multi-view interactions and fusion network for multimodal aspect-based sentiment analysis
Juan Yang 0003, Mengya Xu, Yali Xiao |
Neurocomputing | 2 |
| 2024 | Curriculum-Based Augmented Fourier Domain Adaptation for Robust Medical Image SegmentationabstractAccurate and robust medical image segmentation is fundamental and crucial for enhancing the autonomy of computer-aided diagnosis and intervention systems. Medical data collection normally involves different scanners, protocols, and populations, making domain adaptation (DA) a highly demanding research field to alleviate model degradation in the deployment site. To preserve the model performance across multiple testing domains, this work proposes the Curriculum-based Augmented Fourier Domain Adaptation (Curri-AFDA) for robust medical image segmentation. In particular, our curriculum learning strategy is based on the causal relationship of a model under different levels of data shift in the deployment phase, where the higher the shift is, the harder to recognize the variance. Considering this, we progressively introduce more amplitude information from the target domain to the source domain in the frequency space during the curriculum-style training to smoothly schedule the semantic knowledge transfer in an easier-to-harder manner. Besides, we incorporate the training-time chained augmentation mixing to help expand the data distributions while preserving the domain-invariant semantics, which is beneficial for the acquired model to be more robust and generalize better to unseen domains. Extensive experiments on two segmentation tasks of Retina and Nuclei collected from multiple sites and scanners suggest that our proposed method yields superior adaptation and generalization performance. Meanwhile, our approach proves to be more robust under various corruption types and increasing severity levels. In addition, we show our method is also beneficial in the domain-adaptive classification task with skin lesion datasets. The code is available at https://github.com/lofrienger/Curri-AFDA.Note to Practitioners—Medical image segmentation is key to improving computer-assisted diagnosis and intervention autonomy. However, due to domain gaps between different medical sites, deep learning-based segmentation models frequently encounter performance degradation when deployed in a novel domain. Moreover, model robustness is also highly expected to mitigate the effects of data corruption. Considering all these demanding yet practical needs to automate medical applications and benefit healthcare, we propose the Curriculum-based Fourier Domain Adaptation (Curri-AFDA) for medical image segmentation. Extensive experiments on two segmentation tasks with cross-domain datasets show the consistent superiority of our method regarding adaptation and generalization on multiple testing domains and robustness against synthetic corrupted data. Besides, our approach is independent of image modalities because its efficacy does not rely on modality-specific characteristics. In addition, we demonstrate the benefit of our method for image classification besides segmentation in the ablation study. Therefore, our method can potentially be applied in many medical applications and yield improved performance. Future works may be extended by exploring the integration of curriculum learning regime with Fourier domain amplitude fusion in the testing time rather than in the training time like this work and most other existing domain adaptation works. An Wang 0007, Mobarakol Islam, Mengya Xu, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | Confidence-Aware Paced-Curriculum Learning by Label Smoothing for Surgical Scene UnderstandingabstractCurriculum learning and self-paced learning are the training strategies that gradually feed the samples from easy to more complex. They have captivated increasing attention due to their excellent performance in robotic vision. Most recent works focus on designing curricula based on difficulty levels in input samples or smoothing the feature maps. However, smoothing labels to control the learning utility in a curriculum manner is still unexplored. In this work, we design a paced curriculum by label smoothing (P-CBLS) using paced learning with uniform label smoothing (ULS) for classification tasks and fuse uniform and spatially varying label smoothing (SVLS) for semantic segmentation tasks in a curriculum manner. In ULS and SVLS, a bigger smoothing factor value enforces a heavy smoothing penalty in the true label and limits learning less information. Therefore, we design the curriculum by label smoothing (CBLS). We set a bigger smoothing value at the beginning of training and gradually decreased it to zero to control the model learning utility from lower to higher. We also designed a confidence-aware pacing function and combined it with our CBLS to investigate the benefits of various curricula. The proposed techniques are validated on four robotic surgery datasets of multi-class, multi-label classification, captioning, and segmentation tasks. We also investigate the robustness of our method by corrupting validation data into different severity levels. Our extensive analysis shows that the proposed method improves prediction accuracy and robustness. The code is publicly available at https://github.com/XuMengyaAmy/P-CBLS.Note to Practitioners—The motivation of this article is to improve the performance and robustness of deep neural networks in safety-critical applications such as robotic surgery by controlling the learning ability of the model in a curriculum learning manner and allowing the model to imitate the cognitive process of humans and animals. The designed approaches do not add parameters that require additional computational resources. Mengya Xu, Mobarakol Islam, Ben Glocker, Hongliang Ren 0001 |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2024 | An Efficient MLP-Based Point-Guided Segmentation Network for Ore Images With Ambiguous BoundaryabstractThe precise segmentation of ore images is critical to the successful execution of the beneficiation process. Due to the homogeneous appearance of the ores, which leads to low contrast and unclear boundaries, accurate segmentation becomes challenging, and recognition becomes problematic. This article proposes a lightweight framework based on multilayer perceptron (MLP), which focuses on solving the problem of edge blurring. Specifically, we introduce a lightweight backbone better suited for efficiently extracting low-level features. Besides, we design a feature pyramid network consisting of two MLP structures that balance local and global information, thus, enhancing detection accuracy. Furthermore, we propose a novel loss function that guides the prediction points to match the instance edge points to achieve clear object boundaries. We have conducted extensive experiments to validate the efficacy of our proposed method. Our approach achieves a remarkable processing speed of over 27 frames per second with a model size of only 73 MB. Moreover, our method delivers a consistently high level of accuracy, with impressive performance scores of 60.4 and 48.9 in$AP_{50}^{\text{box}}$and$AP_{50}^{\text{mask}}$, respectively, as compared with the currently available state-of-the-art techniques, when tested on the ORE image dataset. Guodong Sun 0002, Yuting Peng 0001, Mengya Xu, An Wang 0007, Hongliang Ren 0001, Yang Zhang 0053 |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Privacy-Preserving Synthetic Continual Semantic Segmentation for Robotic SurgeryabstractDeep Neural Networks (DNNs) based semantic segmentation of the robotic instruments and tissues can enhance the precision of surgical activities in robot-assisted surgery. However, in biological learning, DNNs cannot learn incremental tasks over time and exhibit catastrophic forgetting, which refers to the sharp decline in performance on previously learned tasks after learning a new one. Specifically, when data scarcity is the issue, the model shows a rapid drop in performance on previously learned instruments after learning new data with new instruments. The problem becomes worse when it limits releasing the dataset of the old instruments for the old model due to privacy concerns and the unavailability of the data for the new or updated version of the instruments for the continual learning model. For this purpose, we develop a privacy-preserving synthetic continual semantic segmentation framework by blending and harmonizing (i) open-source old instruments foreground to the synthesized background without revealing real patient data in public and (ii) new instruments foreground to extensively augmented real background. To boost the balanced logit distillation from the old model to the continual learning model, we design overlapping class-aware temperature normalization (CAT) by controlling model learning utility. We also introduce multi-scale shifted-feature distillation (SD) to maintain long and short-range spatial relationships among the semantic objects where conventional short-range spatial features with limited information reduce the power of feature distillation. We demonstrate the effectiveness of our framework on the EndoVis 2017 and 2018 instrument segmentation dataset with a generalized continual learning setting. Code is available at https://github.com/XuMengyaAmy/Synthetic_CAT_SD. Mengya Xu, Mobarakol Islam, Long Bai 0008, Hongliang Ren 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2023 | Generalizing Surgical Instruments Segmentation to Unseen Domains with One-to-Many SynthesisabstractDespite their impressive performance in various surgical scene understanding tasks, deep learning-based methods are frequently hindered from deploying to real-world surgical applications for various causes. Particularly, data collection, annotation, and domain shift in-between sites and patients are the most common obstacles. In this work, we mitigate data-related issues by efficiently leveraging minimal source images to generate synthetic surgical instrument segmentation datasets and achieve outstanding generalization performance on unseen real domains. Specifically, in our framework, only one background tissue image and at most three images of each foreground instrument are taken as the seed images. These source images are extensively transformed and employed to build up the foreground and background image pools, from which randomly sampled tissue and instrument images are composed with multiple blending techniques to generate new surgical scene images. Besides, we introduce hybrid training-time augmentations to diversify the training data further. Extensive evaluation on three real-world datasets, i.e., Endo2017, Endo2018, and RoboTool, demonstrates that our one-to-many synthetic surgical instruments datasets generation and segmentation framework can achieve encouraging performance compared with training with real data. Notably, on the RoboTool dataset, where a more significant domain gap exists, our framework shows its superiority of generalization by a considerable margin. We expect that our inspiring results will attract research attention to improving model generalization with data synthesizing. An Wang 0007, Mobarakol Islam, Mengya Xu, Hongliang Ren 0001 |
IROS | 3 |
| 2023 | Rectifying Noisy Labels with Sequential Prior: Multi-scale Temporal Feature Affinity Learning for Robust Video Segmentation
Beilei Cui, Minqing Zhang, Mengya Xu, An Wang 0007, Wu Yuan 0001, Hongliang Ren 0001 |
MICCAI (9) | 3 |
| 2023 | S2ME: Spatial-Spectral Mutual Teaching and Ensemble Learning for Scribble-Supervised Polyp Segmentation
An Wang 0007, Mengya Xu, Yang Zhang 0053, Mobarakol Islam, Hongliang Ren 0001 |
MICCAI (1) | 2 |
| 2023 | CholecTriplet2021: A benchmark challenge for surgical action triplet recognition
Chinedu Innocent Nwoye, Deepak Alapatt, Tong Yu 0009, Armine Vardazaryan, Fangfang Xia, Tong Xia, Fucang Jia, Yuxuan Yang 0007, Hao Wang 0081, Derong Yu, Guoyan Zheng, Xiaotian Duan, Neil Getty, Ricardo Sanchez-Matilla, Maria Robu, Li Zhang 0040, Huabin Chen, Jiacheng Wang 0002, Liansheng Wang 0002, Beerend G. A. Gerats, Sista Raviteja, Rachana Sathish, Rong Tao, Satoshi Kondo, Winnie Pang, Hongliang Ren 0001, Julian Ronald Abbing, Mohammad Hasan Sarhan, Sebastian Bodenstedt, Nithya Bhasker, Bruno Oliveira 0002, Helena R. Torres, Finn Gaida, Tobias Czempiel, João L. Vilaça, Pedro Morais, Jaime C. Fonseca 0001, Ruby Mae Egging, Inge Nicole Wijma, Chen Qian 0006, Guibin Bian, Zhen Li 0026, Velmurugan Balasubramanian, Debdoot Sheet, Imanol Luengo, Yuanbo Zhu, Shuai Ding 0001, Jakob-Anton Aschenbrenner, Nicolas Elini van der Kar, Mengya Xu, Mobarakol Islam, Seenivasan Lalithkumar, Alexander Jenke, Danail Stoyanov, Didier Mutter, Pietro Mascagni, Barbara Seeliger, Cristians Gonzalez, Nicolas Padoy |
Medical Image Anal. | 53 |
| 2023 | AMagPoseNet: Real-Time Six-DoF Magnet Pose Estimation by Dual-Domain Few-Shot Learning From Prior ModelabstractTraditional magnetic tracking approaches based on mathematical models and optimization algorithms are computationally intensive, depend on initial guesses, and do not guarantee convergence to a global optimum. Although fully supervised data-driven deep learning can solve the above issues, the demand for a comprehensive dataset hampers its applicability in magnetic tracking. Thus, we propose an annular magnet pose estimation network (called AMagPoseNet) based on dual-domain few-shot learning from a prior mathematical model, which consists of two subnetworks: PoseNet and CaliNet. PoseNet learns to estimate the magnet pose from the prior mathematical model, and CaliNet is designed to narrow the gap between the mathematical model domain and the real-world domain. Experimental results reveal that the AMagPoseNet outperforms the optimization-based method regarding localization accuracy (1.87$\pm$1.14 mm, 1.89$\pm \text{0.81}^{\circ }$), robustness (nondependence on initial guesses), and computational latency (2.08$\pm$0.02 ms). In addition, the six-degree-of-freedom pose of the magnet could be estimated when discriminative magnetic field features are provided. With the assistance of the mathematical model, the AMagPoseNet requires only a few real-world samples and has excellent performance, showing great potential for practical biomedical and industrial applications. Shijian Su, Sishen Yuan, Mengya Xu, Huxin Gao, Xiaoxiao Yang, Hongliang Ren 0001 |
IEEE Trans. Ind. Informatics | 3 |
| 2022 | Active SLAM in 3D deformable environmentsabstractThis paper considers active SLAM problem for 3D deformable environments where the trajectory of the robot is planned to optimize the SLAM results. A planning strategy combining an efficient global planner with an accurate local planner is proposed to solve the problem. Simulation results under different scenarios have shown that the proposed active SLAM algorithm provides a good balance between accuracy and efficiency as compared to the local planner and the global planner. The MATLAB code of this first active SLAM algorithm for 3D deformable environments is made publicly available4. Mengya Xu, Liang Zhao 0003, Shoudong Huang, Qi Hao 0003 |
IROS | 1 |
| 2022 | Rethinking Surgical Instrument Segmentation: A Background Image Can Be All You Need
An Wang 0007, Mobarakol Islam, Mengya Xu, Hongliang Ren 0001 |
MICCAI (8) | 3 |
| 2022 | Rethinking Surgical Captioning: End-to-End Window-Based MLP Transformer Using Patches
Mengya Xu, Mobarakol Islam, Hongliang Ren 0001 |
MICCAI (8) | 1 |
| 2021 | Learning Domain Adaptation with Model Calibration for Surgical Report Generation in Robotic SurgeryabstractGenerating a surgical report in robot-assisted surgery, in the form of natural language expression of surgical scene understanding, can play a significant role in document entry tasks, surgical training, and post-operative analysis. Despite the state-of-the-art accuracy of the deep learning algorithm, the deployment performance often drops when applied to the Target Domain (TD) data. For this purpose, we develop a multi-layer transformer-based model with the gradient reversal adversarial learning to generate a caption for the multi-domain surgical images that can describe the semantic relationship between instruments and surgical Region of Interest (ROI). In the gradient reversal adversarial learning scheme, the gradient multiplies with a negative constant and updates adversarially in backward propagation, discriminating between the source and target domains and emerging domain-invariant features. We also investigate model calibration with label smoothing technique and the effect of a well-calibrated model for the penultimate layer’s feature representation and Domain Adaptation (DA). We annotate two robotic surgery datasets of MICCAI robotic scene segmentation and Transoral Robotic Surgery (TORS) with the captions of procedures and empirically show that our proposed method improves the performance in both source and target domain surgical reports generation in the manners of unsupervised, zero-shot, one-shot, and few-shot learning. Mengya Xu, Mobarakol Islam, Chwee Ming Lim, Hongliang Ren 0001 |
ICRA | 1 |
| 2021 | Invariant EKF based 2D Active SLAM with Exploration TaskabstractRight invariant extended Kalman filter (RIEKF) based simultaneous localization and mapping (SLAM) proposed recently has shown to be able to produce more consistent SLAM estimates as compared with traditional EKF based SLAM methods, including some improved EKF SLAM methods such as observability constrained-EKF (OC-EKF) SLAM. Latest results have demonstrated that its performance is very close to optimization based SLAM algorithms such as iSAM. In this paper, we propose to use RIEKF SLAM algorithm in active SLAM where both the predicted SLAM results for choosing control actions and the actual estimated SLAM results applying the selected control actions are computed using RIEKF algorithms. The advantages over traditional EKF based active SLAM are the more accurate and consistent predicted uncertainty estimates which result in robustness of the active SLAM algorithm. The advantages over optimization based active SLAM is the reduced computational cost. Simulation results are presented to validate the advantages of the proposed algorithm3. Mengya Xu, Yang Song 0028, Yongbo Chen 0001, Shoudong Huang, Qi Hao 0003 |
ICRA | 1 |
| 2021 | Some Research Questions for SLAM in Deformable EnvironmentsabstractSLAM in deformable environments is a very challenging research topic. Some research works have been presented by different research groups in the past few years. However, there are still some challenging research questions remaining unanswered. This paper discusses some of these research questions focusing on the case when point features are used to describe the deformable environments. The SLAM problems are formulated as extensions of point feature based SLAM in static environments, including both optimisation based offline SLAM and filter based online SLAM. To illustrate the problems and questions more clearly, some concepts and results using simple 2D examples are presented. The MATLAB source codes of the results are made publicly available (https://github.com/cyb1212/DeformableSLAM2D.git) to help the readers understand the problems more clearly. Shoudong Huang, Yongbo Chen 0001, Liang Zhao 0003, Yanhao Zhang 0003, Mengya Xu |
IROS | 5 |
| 2021 | Class-Incremental Domain Adaptation with Smoothing and Calibration for Surgical Report Generation
Mengya Xu, Mobarakol Islam, Chwee Ming Lim, Hongliang Ren 0001 |
MICCAI (4) | 1 |