EDBT 2026 Demo / reviewers in the wild / expert
Junjie Yang 0001
dblp:41/3461-1
· DBLP profile ↗
14ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0001-7856-8120ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Systems, architecture and hardware · 6 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Software engineering, systems software and programming languages · 2Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CLAPS: A CLIP-Unified Auto-Prompt Segmentation for Multi-Modal Retinal ImagingabstractRecent advancements in foundation models, such as the Segment Anything Model (SAM), have significantly impacted medical image segmentation, especially in retinal imaging, where precise segmentation is vital for diagnosis. Despite this progress, current methods face critical challenges: 1) modality ambiguity in textual disease descriptions, 2) a continued reliance on manual prompting for SAM-based workflows, and 3) a lack of a unified framework, with most methods being modalityand task-specific. To overcome these hurdles, we propose CLIP-unified Auto-Prompt Segmentation (CLAPS), a novel method for unified segmentation across diverse tasks and modalities in retinal imaging. Our approach begins by pre-training a CLIP-based image encoder on a large, multi-modal retinal dataset to handle data scarcity and distribution imbalance. We then leverage GroundingDINO to automatically generate spatial bounding box prompts by detecting local lesions. To unify tasks and resolve ambiguity, we use text prompts enhanced with a unique “modality signature” for each imaging modality. Ultimately, these automated textual and spatial prompts guide SAM to execute precise segmentation, creating a fully automated and unified pipeline. Extensive experiments on 12 diverse datasets across 11 critical segmentation categories show that CLAPS achieves performance on par with specialized expert models while surpassing existing benchmarks across most metrics, demonstrating its broad generalizability as a foundation model. Yinzheng Zhao, Junjie Yang 0001, Xiangtong Yao, Quanmin Liang, Shahrooz Faghih Roohi, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
BIBM | 3 |
| 2025 | UOPSL: Unpaired OCT Predilection Sites Learning for Fundus Image Diagnosis AugmentationabstractSignificant advancements in AI-driven multimodal medical image diagnosis have led to substantial improvements in ophthalmic disease identification in recent years. However, acquiring paired multimodal ophthalmic images remains prohibitively expensive. While fundus photography is simple and cost-effective, the limited availability of OCT data and inherent modality imbalance hinder further progress. Conventional approaches that rely solely on fundus or textual features often fail to capture fine-grained spatial information, as each imaging modality provides distinct cues about lesion predilection sites. In this study, we propose a novel unpaired multimodal framework UOPSL that utilizes extensive OCT-derived spatial priors to dynamically identify predilection sites, enhancing fundus imagebased disease recognition. Our approach bridges unpaired fundus and OCTs via extended disease text descriptions. Initially, we employ contrastive learning on a large corpus of unpaired OCT and fundus images while simultaneously learning the predilection sites matrix in the OCT latent space. Through extensive optimization, this matrix captures lesion localization patterns within the OCT feature space. During the fine-tuning or inference phase of the downstream classification task based solely on fundus images, where paired OCT data is unavailable, we eliminate OCT input and utilize the predilection sites matrix to assist in fundus image classification learning. Extensive experiments conducted on 9 diverse datasets across 28 critical categories demonstrate that our framework outperforms existing benchmarks. Yinzheng Zhao, Junjie Yang 0001, Xiangtong Yao, Quanmin Liang, Daniel Zapp, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
BIBM | 3 |
| 2025 | Intraoperative Trocar-Based Eyeball Rotation Estimation Using Only 2D Microscope ImagesabstractIn ophthalmic surgery, surgeons or robots manipulate a light probe and an instrument around two separated trocars following sclerotomy to achieve orbital control for eyeball pose adjustment and subsequent surgical tasks referring to microscope frames. However, current methods face significant challenges in directly extracting the eyeball pose from real-time microscope frames due to the limited microscope perspective and the darkened operating room (OR). This paper decomposes eyeball rotations only along the x and y axes. Then, a method of calculating eyeball poses using eyeball geometry and microscopic trocar positions is presented. This method is tested by simulation and a phantom system with current [2.0, 2.8] degree error, providing assistant intraoperative eyeball status in the dark OR with extended method discussions. Junjie Yang 0001, Satoshi Inagaki, Daniel Zapp, Mathias Maier, Peter C. Issa, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
ICRA | 1 |
| 2024 | Extrapolating Prospective Glaucoma Fundus Images through Diffusion in Irregular Longitudinal SequencesabstractThe utilization of longitudinal datasets for glaucoma progression prediction offers a compelling approach to support early therapeutic interventions. Predominant methodologies in this domain have primarily focused on the direct prediction of glaucoma stage labels from longitudinal datasets. However, such methods may not adequately encapsulate the nuanced developmental trajectory of the disease. To enhance the diagnostic acumen of medical practitioners, we propose a novel diffusion-based model to predict prospective images by extrapolating from existing longitudinal fundus images of patients. The methodology delineated in this study distinctively leverages sequences of images as inputs. Subsequently, a time-aligned mask is employed to select a specific year for image generation. During the training phase, the time-aligned mask resolves the issue of irregular temporal intervals in longitudinal image sequence sampling. Additionally, we utilize a strategy of randomly masking a frame in the sequence to establish the ground truth. This methodology aids the network in continuously acquiring knowledge regarding the internal relationships among the sequences throughout the learning phase. Moreover, the introduction of textual labels is instrumental in categorizing images generated within the sequence. The empirical findings from the conducted experiments indicate that our proposed model not only effectively generates longitudinal data but also significantly improves the precision of downstream classification tasks. Junjie Yang 0001, Shahrooz Faghih Roohi, Yinzheng Zhao, Daniel Zapp, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
BIBM | 2 |
| 2024 | KLDD: Kalman Filter based Linear Deformable Diffusion Model in Retinal Image SegmentationabstractAI-based vascular segmentation is becoming increasingly common in enhancing the screening and treatment of ophthalmic diseases. Deep learning structures based on U-Net have achieved relatively good performance in vascular segmentation. However, small blood vessels and capillaries tend to be lost during segmentation when passed through the traditional U-Net downsampling module. To address this gap, this paper proposes a novel Kalman filter based Linear Deformable Diffusion (KLDD) model for retinal vessel segmentation. Our model employs a diffusion process that iteratively refines the segmentation, leveraging the flexible receptive fields of deformable convolutions in feature extraction modules to adapt to the detailed tubular vascular structures. More specifically, we first employ a feature extractor with linear deformable convolution to capture vascular structure information form the input images. To better optimize the coordinate positions of deformable convolution, we employ the Kalman filter to enhance the perception of vascular structures in linear deformable convolution. Subsequently, the features of the vascular structures extracted are utilized as a conditioning element within a diffusion model by the Cross-Attention Aggregation module (CAAM) and the Channel-wise Soft Attention module (CSAM). These aggregations are designed to enhance the diffusion model’s capability to generate vascular structures. Experiments are evaluated on retinal fundus image datasets (DRIVE, CHASE DB1) as well as the 3mm and 6mm of the OCTA-500 dataset, and the results show that the diffusion model proposed in this paper outperforms other methods. Yinzheng Zhao, Junjie Yang 0001, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
BIBM | 3 |
| 2024 | Shadow-Based 3D Pose Estimation of Intraocular Instrument Using Only 2D ImagesabstractIn ophthalmic surgeries, such as vitreoretinal operations, surgeons rely on imaging systems, primarily microscopes, for real-time instrument monitoring and motion planning. However, novice surgeons struggle to extract 3D instrument positions from 2D microscope frames, necessitating extensive trial-and-error experience with the background that additional imaging modalities such as iOCT remain inaccessible in most operating rooms. Targeting intraocular assessment within the current surgical setup, this paper presents an imagebased pose estimation method to obtain real-time instrument tip positions in a standard 12mm-radius spherical eyeball model, which links floating instruments with on-the-retinal objects based on the intraocular shadowing principle. We validate this estimation method in a Unity simulator and verify its depth estimation capability using a specially designed eyeball phantom. Both simulator and phantom experiments demonstrate an average needle-tip estimation error within [1.0, 2.0] mm using only 2D microscope frames. Junjie Yang 0001, Mathias Maier, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
ICRA | 1 |
| 2024 | Shadow Maintenance for Automatic Light-Probe Control in Ophthalmic Surgeries Using Only 2D informationabstractIn ophthalmic surgeries, the light probe is responsible for providing safe intraocular illumination and ensuring the visibility of the instrument and its shadow as the only available reference for qualitative depth estimation and landing point prediction in fundus microscopic images. To achieve sustainable shadow-based estimation during surgeries, we propose controlling the light probe automatically to limit the shadow position around the instrument tip using only 2D information from the microscope. We also integrate an intensity balancing sub-module to guarantee the normal intensity distribution and the safe depth of light-tip placement. Without motor-based pose coordination between the light probe and the instrument, experiments analyze the performance of our image-based shadow maintenance with only image information under the constraints of RCM and discuss the working volume and segmentation limitations during simulation and real-robot tests. Junjie Yang 0001, Satoshi Inagaki, Daniel Zapp, Mathias Maier, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
IROS | 1 |
| 2024 | Intraocular Reflection Modeling and Avoidance Planning in Image-Guided Ophthalmic SurgeriesabstractIntuitive enhancement of surgical precision in robotic retinal surgery highly depends on the stable acquisition of intraocular imaging data. Such acquisition requires segmenting intraocular components, especially instrument-tip positions, to achieve state estimation and subsequent navigation and motion control. However, intraocular light reflections and glares significantly impact instrument segmentation, state estimation, and subsequent visual servoing in retinal surgery. At the same time, light reflections are among the sources of information for intraoperative navigation. In this work, we propose a method for modeling and optimizing light reflections using microscopy as the standard surgical imaging modality. Beyond optimization, our approach seamlessly integrates the optimized reflection with path planning, strategically circumventing reflection areas and ensuring uninterrupted visibility of instrument tips throughout the surgical procedure. Experiments demonstrate the methodology’s efficacy in avoiding glare affections during eye surgeries. Junjie Yang 0001, Yinzheng Zhao, Daniel Zapp, Mathias Maier, Kai Huang 0001, Nassir Navab, M. Ali Nasseri |
IROS | 1 |
| 2023 | Label-Preserving Data Augmentation in Latent Space for Diabetic Retinopathy Recognition
Junjie Yang 0001, Shahrooz Faghih Roohi, Kai Huang 0001, Mathias Maier, Nassir Navab, M. Ali Nasseri |
MICCAI (3) | 2 |
| 2022 | ColibriDoc: an Eye-in-Hand Autonomous Trocar Docking SystemabstractRetinal surgery is a complex medical procedure that requires exceptional expertise and dexterity. For this purpose, several robotic platforms are currently under development to enable or improve the outcome of microsurgical tasks. Since the control of such robots is often designed for navigation inside the eye in proximity to the retina, successful trocar docking and insertion of the instrument into the eye represents an additional cognitive effort, and is therefore one of the open challenges in robotic retinal surgery. For this purpose, we present a platform for autonomous trocar docking that combines computer vision and a robotic setup. Inspired by the Cuban Colibri (hummingbird) aligning its beak to a flower using only vision, we mount a camera onto the endeffector of a robotic system. By estimating the position and pose of the trocar, the robot is able to autonomously align and navigate the instrument towards the Trocar Entry Point (TEP) and finally perform the insertion. Our experiments show that the proposed method is able to accurately estimate the position and pose of the trocar and achieve repeatable autonomous docking. The aim of this work is to reduce the complexity of the robotic setup prior to the surgical task and therefore, increase the intuitiveness of the system integration into clinical workflow. Shervin Dehghani, Michael Sommersperger, Junjie Yang 0001, Mehrdad Salehi, Benjamin Busam, Kai Huang 0001, Peter Gehlbach, Iulian Iordachita, Nassir Navab, M. Ali Nasseri |
ICRA | 3 |
| 2021 | Efficient runtime slack management for EDF-VD-based mixed-criticality scheduling
Junjie Yang 0001, Guangyi Xu 0004, Gang Chen 0023, Nan Guan, Kai Huang 0001 |
J. Syst. Archit. | 1 |
| 2018 | Resource-Aware Design for Reliable Autonomous Applications with Multiple Periods
Rongjie Yan, Yiqi Lv, Junjie Yang 0001, Kai Huang 0001 |
FM | 5 |
| 2018 | Design Verification and Validation for Reliable Safety-Critical Autonomous Control SystemsabstractProviding guarantees on the system behavior is mandatory for safety-critical autonomous vehicles. Among these guarantees, proving the fulfillment of real-time constraints and reliability requirements on the system is a key issue, as their violation could result in unexpected and unsafe behaviors. The violation may come from the complicated interaction between software and hardware modules, or transient hardware faults. AUTOSAR, the most popular industrial standard in the automotive domain, provides an open standardized architecture for software development, where an application can be deployed on multiple electronic control units (ECUs). We present a verification and validation method for the design of such safety-critical autonomous control systems that could tolerate transient faults. The embedded implementation of an AUTOSAR model is transformed into a three-layer system model in timed automata, so that system behavior can be evaluated and checked with hard real-time constraints and the implementing architecture. We demonstrate the feasibility of the method with a simplified controller developed for the autonomous vehicles. Rongjie Yan, Junjie Yang 0001, Kai Huang 0001 |
ICECCS | 2 |
| 2018 | LiSeg: Lightweight Road-object Semantic Segmentation In 3D LiDAR Scans For Autonomous DrivingabstractLiDAR based perception module plays an important role in autonomous driving. However, the present CNN models are designed for image processing but not LiDAR point clouds. The performances of such models are limited by the great memory consumption and heavy computation cost. In this work, a lightweight CNN model, Liseg, is proposed to perform real-time road-object semantic segmentation on LiDAR point cloud scans for autonomous driving. The model size of Liseg is several times smaller than others, while achieving high accuracy. Wenquan Zhang, Chancheng Zhou, Junjie Yang 0001, Kai Huang 0001 |
Intelligent Vehicles Symposium | 3 |