EDBT 2026 Demo / reviewers in the wild / expert
Adam Schmidt
dblp:29/1080
· DBLP profile ↗
13ranked-venue papers
9as first author
8since 2021 · last 2025
0000-0003-4769-4313ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SurgPose: a Dataset for Articulated Robotic Surgical Tool Pose Estimation and TrackingabstractAccurate and efficient surgical robotic tool pose estimation is of fundamental significance to downstream applications such as augmented reality (AR) in surgical training and learning-based autonomous manipulation. While significant advancements have been made in pose estimation for humans and animals, it is still a challenge in surgical robotics due to the scarcity of published data. The relatively large absolute error of the da Vinci end effector kinematics and arduous calibration procedure make calibrated kinematics data collection expensive. Driven by this limitation, we collected a dataset, dubbed SurgPose, providing instance-aware semantic keypoints for visual surgical tool pose estimation and tracking. By marking keypoints using ultraviolet (UV) reactive paint, which is invisible under white light and fluorescent under UV light, we execute the same trajectory under different lighting conditions to collect raw videos and keypoint annotations, respectively. The SurgPose dataset consists of approximately 120 K surgical instrument instances of 6 categories as shown in Fig. 1. Since the videos are collected in stereo pairs, the 2D pose can be lifted to 3D based on stereo-matching depth. In addition to releasing the dataset, we tested a few baseline approaches to surgical instrument tracking to demonstrate the utility of SurgPose. More details can be found at surgpose.github.io. Zijian Wu 0001, Adam Schmidt, Randy Moore, Haoying Zhou, Alexandre Banks, Peter Kazanzides, Tim Salcudean |
ICRA | 2 |
| 2024 | Tracking and mapping in medical computer vision: A reviewabstractAs computer vision algorithms increase in capability, their applications in clinical systems will become more pervasive. These applications include: diagnostics, such as colonoscopy and bronchoscopy; guiding biopsies, minimally invasive interventions, and surgery; automating instrument motion; and providing image guidance using pre-operative scans. Many of these applications depend on the specific visual nature of medical scenes and require designing algorithms to perform in this environment. In this review, we provide an update to the field of camera-based tracking and scene mapping in surgery and diagnostics in medical computer vision. We begin with describing our review process, which results in a final list of 515 papers that we cover. We then give a high-level summary of the state of the art and provide relevant background for those who need tracking and mapping for their clinical applications. After which, we review datasets provided in the field and the clinical needs that motivate their design. Then, we delve into the algorithmic side, and summarize recent developments. This summary should be especially useful for algorithm designers and to those looking to understand the capability of off-the-shelf methods. We maintain focus on algorithms for deformable environments while also reviewing the essential building blocks in rigid tracking and mapping since there is a large amount of crossover in methods. With the field summarized, we discuss the current state of the tracking and mapping methods along with needs for future algorithms, needs for quantification, and the viability of clinical applications. We then provide some research directions and questions. We conclude that new methods need to be designed or combined to support clinical applications in deformable environments, and more focus needs to be put into collecting datasets for training and evaluation. Adam Schmidt, Omid Mohareri, Simon P. DiMaio, Michael C. Yip, Tim Salcudean |
Medical Image Anal. | 1 |
| 2024 | Surgical Tattoos in Infrared: A Dataset for Quantifying Tissue Tracking and MappingabstractQuantifying performance of methods for tracking and mapping tissue in endoscopic environments is essential for enabling image guidance and automation of medical interventions and surgery. Datasets developed so far either use rigid environments, visible markers, or require annotators to label salient points in videos after collection. These are respectively: not general, visible to algorithms, or costly and error-prone. We introduce a novel labeling methodology along with a dataset that uses said methodology, Surgical Tattoos in Infrared (STIR). STIR has labels that are persistent but invisible to visible spectrum algorithms. This is done by labelling tissue points with IR-fluorescent dye, indocyanine green (ICG), and then collecting visible light video clips. STIR comprises hundreds of stereo video clips in both in vivo and ex vivo scenes with start and end points labelled in the IR spectrum. With over 3,000 labelled points, STIR will help to quantify and enable better analysis of tracking and mapping methods. After introducing STIR, we analyze multiple different frame-based tracking methods on STIR using both 3D and 2D endpoint error and accuracy metrics. STIR is available at https://dx.doi.org/10.21227/w8g4-g548. Adam Schmidt, Omid Mohareri, Simon P. DiMaio, Tim Salcudean |
IEEE Trans. Medical Imaging | 1 |
| 2023 | SENDD: Sparse Efficient Neural Depth and Deformation for Tissue Tracking
Adam Schmidt, Omid Mohareri, Simon P. DiMaio, Tim Salcudean |
MICCAI (9) | 1 |
| 2022 | Fast Graph Refinement and Implicit Neural Representation for Tissue TrackingabstractTracking of tissue in the surgical environment is often done via locating frame-to-frame keypoint correspondences, and then using these correspondences to warp a prior underlying model such as a spline, mesh, or embedded deformation. We introduce a novel learned model which takes keypoint correspondences as input and enables a prior-free estimation of deformation at any location. For fast point tracking, our model allows for sparse queries, unlike dense grid based CNNs, which run on full images. Our model begins with a novel graph-based point refinement scheme which refines matched keypoints, updating their features and movement instead of discarding possible outliers. Then, we use these refined matches to learn a novel neural implicit representation for estimating movement of any location given its k-nearest neighbor (k-NN) keypoints. We name our implicit deformation model KINFlow (k-NN implicit neural flow). We demonstrate the performance of KINFlow photometrically on three different datasets. KINFlow is the first model to use a graph network to estimate flow of arbitrary query points, and can estimate movement of 1024 points in under 3 ms. Adam Schmidt, Omid Mohareri, Simon P. DiMaio, Tim Salcudean |
ICRA | 1 |
| 2022 | Recurrent Implicit Neural Graph for Deformable Tracking in Endoscopic Videos
Adam Schmidt, Omid Mohareri, Simon P. DiMaio, Tim Salcudean |
MICCAI (4) | 1 |
| 2021 | Real-Time Rotated Convolutional Descriptor for Surgical Environments
Adam Schmidt, Tim Salcudean |
MICCAI (4) | 1 |
| 2021 | Multi-view Surgical Video Action Detection via Mixed Global View Attention
Adam Schmidt, Aidean Sharghi, Helene Haugerud, Daniel Oh, Omid Mohareri |
MICCAI (4) | 1 |
| 2017 | Toward evaluation of visual navigation algorithms on RGB-D data from the first- and second-generation KinectabstractAlthough the introduction of commercial RGB-D sensors has enabled significant progress in the visual navigation methods for mobile robots, the structured-light-based sensors, like Microsoft Kinect and Asus Xtion Pro Live, have some important limitations with respect to their range, field of view, and depth measurements accuracy. The recent introduction of the second- generation Kinect, which is based on the time-of-flight measurement principle, brought to the robotics and computer vision researchers a sensor that overcomes some of these limitations. However, as the new Kinect is, just like the older one, intended for computer games and human motion capture rather than for navigation, it is unclear how much the navigation methods, such as visual odometry and SLAM, can benefit from the improved parameters. While there are many publicly available RGB-D data sets, only few of them provide ground truth information necessary for evaluating navigation methods, and to the best of our knowledge, none of them contains sequences registered with the new version of Kinect. Therefore, this paper describes a new RGB-D data set, which is a first attempt to systematically evaluate the indoor navigation algorithms on data from two different sensors in the same environment and along the same trajectories. This data set contains synchronized RGB-D frames from both sensors and the appropriate ground truth from an external motion capture system based on distributed cameras. We describe in details the data registration procedure and then evaluate our RGB-D visual odometry algorithm on the obtained sequences, investigating how the specific properties and limitations of both sensors influence the performance of this navigation method. Marek Kraft, Michal R. Nowicki, Adam Schmidt, Michal Fularz, Piotr Skrzypczynski |
Mach. Vis. Appl. | 3 |
| 2016 | Collection of Visual Data in Climbing Experiments for Addressing the Role of Multi-modal Exploration in Motor Learning Efficiency
Adam Schmidt, Dominic Orth, Ludovic Seifert |
ACIVS | 1 |
| 2015 | Collaborative, Context Based Activity Control Method for Camera Networks
Marek Kraft, Michal Fularz, Adam Schmidt |
ACIVS | 3 |
| 2013 | An Indoor RGB-D Dataset for the Evaluation of Robot Navigation Algorithms
Adam Schmidt, Michal Fularz, Marek Kraft, Andrzej J. Kasinski, Michal R. Nowicki |
ACIVS | 1 |
| 2010 | The architecture and performance of the face and eyes detection system based on the Haar cascade classifiers
Andrzej J. Kasinski, Adam Schmidt |
Pattern Anal. Appl. | 2 |