EDBT 2026 Demo / reviewers in the wild / expert
Noah Stier
dblp:174/0292
· DBLP profile ↗
12ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0001-9602-4637ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prism: Semi-Supervised Multi-View Stereo with Monocular Structure PriorsabstractThe promise of unsupervised multi-view stereo (MVS) is to leverage large unlabeled datasets, yet current methods underperform when training on difficult data, such as handheld smartphone videos of indoor scenes. Meanwhile, high-quality synthetic datasets are available but MVS networks trained on these datasets fail to generalize to realworld examples. To bridge this gap, we propose a semisupervised learning framework that allows us to train on real and rendered images jointly, capturing structural priors from synthetic data while ensuring parity with the realworld domain. Central to our framework is a novel set of losses that leverages powerful existing monocular relativedepth estimators trained on the synthetic dataset, transferring the rich structure of this relative depth to the MVS predictions on unlabeled data. Inspired by perceptual image metrics, we compare the MVS and monocular predictions via a deep feature loss and a multi-scale statistical loss. Our full framework, which we call Prism, achieves large quantitative and qualitative improvements over current unsupervised and synthetic-supervised MVS networks. This is quite a useful result, opening the door to using both unlabeled smartphone videos and photorealistic synthetic datasets for training MVS networks. Alexander Rich 0001, Noah Stier, Pradeep Sen, Tobias Höllerer |
3DV | 2 |
| 2025 | AniGrad: Anisotropic Gradient-Adaptive Sampling for 3D Reconstruction From Monocular VideoabstractRecent image-based 3D reconstruction methods have achieved excellent quality for indoor scenes using 3D convolutional neural networks. However, they rely on a high-resolution grid in order to achieve detailed output surfaces, which is quite costly in terms of compute time, and it results in large mesh sizes that are more expensive to store, transmit, and render. In this paper we propose a new solution to this problem, using adaptive sampling. By re-formulating the final layers of the network, we are able to analytically bound the local surface complexity, and set the local sample rate accordingly. Our method, AniGrad1, achieves an order of magnitude reduction in both surface extraction latency and mesh size, while preserving mesh accuracy and detail. Noah Stier, Alexander Rich 0001, Pradeep Sen, Tobias Höllerer |
CVPR | 1 |
| 2024 | Smoothness, Synthesis, and Sampling: Re-thinking Unsupervised Multi-view Stereo with DIV Loss
Alexander Rich 0001, Noah Stier, Pradeep Sen, Tobias Höllerer |
ECCV (63) | 2 |
| 2024 | Multimodal 3D Fusion and In-Situ Learning for Spatially Aware AIabstractSeamless integration of virtual and physical worlds in augmented reality benefits from the system semantically “understanding” the physical environment. AR research has long focused on the potential of context awareness, demonstrating novel capabilities that leverage the semantics in the 3D environment for various object-level interactions. Meanwhile, the computer vision community has made leaps in neural vision-language understanding to enhance environment perception for autonomous tasks. In this work, we introduce a multimodal 3D object representation that unifies both semantic and linguistic knowledge with the geometric representation, enabling user-guided machine learning involving physical objects. We first present a fast multimodal 3D reconstruction pipeline that brings linguistic understanding to AR by fusing CLIP vision-language features into the environment and object models. We then propose “in-situ” machine learning, which, in conjunction with the multimodal representation, enables new tools and interfaces for users to interact with physical spaces and objects in a spatially and linguistically meaningful manner. We demonstrate the usefulness of the proposed system through two real-world AR applications on Magic Leap 2: a) spatial search in physical environments with natural language and b) an intelligent inventory system that tracks object changes over time. We also make our full implementation and demo data available at (https://github.com/cy-xu/spatially_aware_AI) to encourage further exploration and research in spatially aware AI. Radha Kumaran, Noah Stier, Kangyou Yu, Tobias Höllerer |
ISMAR | 3 |
| 2023 | LivePose: Online 3D Reconstruction from Monocular Video with Dynamic Camera PosesabstractDense 3D reconstruction from RGB images traditionally assumes static camera pose estimates. This assumption has endured, even as recent works have increasingly focused on real-time methods for mobile devices. However, the assumption of a fixed pose for each image does not hold for online execution: poses from real-time SLAM are dynamic and may be updated following events such as bundle adjustment and loop closure. This has been addressed in the RGB-D setting, by de-integrating past views and re-integrating them with updated poses, but it remains largely untreated in the RGB-only setting. We formalize this problem to define the new task of dense online reconstruction from dynamically-posed images. To support further research, we introduce a dataset called LivePose1containing the dynamic poses from a SLAM system running on Scan-Net [6]. We select three recent reconstruction systems and apply a framework based on de-integration to adapt each one to the dynamic-pose setting. In addition, we propose a novel, non-linear de-integration module that learns to remove stale scene content. We show that responding to pose updates is critical for high-quality reconstruction, and that our de-integration framework is an effective solution. Noah Stier, Baptiste Angles, Liang Yang 0005, Yajie Yan, Alex Colburn, Ming Chuang |
ICCV | 1 |
| 2023 | FineRecon: Depth-aware Feed-forward Network for Detailed 3D ReconstructionabstractRecent works on 3D reconstruction from posed images [17], [23], [24] have demonstrated that direct inference of scene-level 3D geometry without test-time optimization is feasible using deep neural networks, showing remarkable promise and high efficiency. However, the reconstructed geometry, typically represented as a 3D truncated signed distance function (TSDF), is often coarse without fine geometric details. To address this problem, we propose three effective solutions for improving the fidelity of inference-based 3D reconstructions. We first present a resolution-agnostic TSDF supervision strategy to provide the network with a more accurate learning signal during training, avoiding the pitfalls of TSDF interpolation seen in previous work. We then introduce a depth guidance strategy using multi-view depth estimates to enhance the scene representation and recover more accurate surfaces. Finally, we develop a novel architecture for the final layers of the network, conditioning the output TSDF prediction on high-resolution image features in addition to coarse voxel features, enabling sharper reconstruction of fine details. Our method, FineRecon1, produces smooth and highly accurate reconstructions, showing significant improvements across multiple depth and 3D reconstruction metrics. Noah Stier, Anurag Ranjan, Alex Colburn, Yajie Yan, Liang Yang 0005, Fangchang Ma, Baptiste Angles |
ICCV | 1 |
| 2022 | Interactive Segmentation and Visualization for Tiny Objects in Multi-megapixel ImagesabstractWe introduce an interactive image segmentation and visualization framework for identifying, inspecting, and editing tiny objects (just a few pixels wide) in large multi-megapixel high-dynamic-range (HDR) images. Detecting cosmic rays (CRs) in astronomical observations is a cum-bersome workflow that requires multiple tools, so we developed an interactive toolkit that unifies model inference, HDR image visualization, segmentation mask inspection and editing into a single graphical user interface. The feature set, initially designed for astronomical data, makes this work a useful research-supporting tool for human-in-the-loop tiny-object segmentation in scientific areas like biomedicine, materials science, remote sensing, etc., as well as computer vision. Our interface features mouse-controlled, synchronized, dual-window visualization of the image and the segmentation mask, a critical feature for locating tiny objects in multi-megapixel images. The browser-based tool can be readily hosted on the web to provide multi-user access and GPU acceleration for any device. The toolkit can also be used as a high-precision annotation tool, or adapted as the frontend for an interactive machine learning framework. Our open-source dataset, CR detection model, and visualization toolkit are available at https://github.com/cy-xu/cosmic-com. Boning Dong, Noah Stier, Curtis McCully, D. Andrew Howell, Pradeep Sen, Tobias Höllerer |
CVPR | 3 |
| 2022 | Weakly-Supervised Convolutional Neural Networks for Vessel Segmentation in Cerebral AngiographyabstractAutomated vessel segmentation in cerebral digital subtraction angiography (DSA) has significant clinical utility in the management of cerebrovascular diseases. Although deep learning has become the foundation for state-of-the-art image segmentation, a significant amount of labeled data is needed for training. Furthermore, due to domain differences, pre-trained networks cannot be applied to DSA data out-of-the-box. To address this, we propose a novel learning framework, which utilizes an active contour model for weak supervision and low-cost human-in-the-loop strategies to improve weak label quality. Our study produces several significant results, including state-of-the-art results for cerebral DSA vessel segmentation, which exceed human annotator quality, and an analysis of annotation cost and model performance trade-offs when utilizing weak supervision strategies. For comparison purposes, we also demonstrate our approach on the Digital Retinal Images for Vessel Extraction (DRIVE) dataset. Additionally, we will be publicly releasing code to reproduce our methodology and our dataset, the largest known high-quality annotated cerebral DSA vessel segmentation dataset. Arvind Vepa, Andrew Choi, Noor Nakhaei, Wonjun Lee 0004, Noah Stier, Andrew Vu, Greyson Jenkins, Manjot Shergill, Moira Desphy, Kevin Delao, Mia Levy, Cristopher Garduno, Lacy Nelson, Wandi Liu, Fan Hung, Fabien Scalzo |
WACV | 5 |
| 2021 | 3DVNet: Multi-View Depth Prediction and Volumetric RefinementabstractWe present 3DVNet, a novel multi-view stereo (MVS) depth-prediction method that combines the advantages of previous depth-based and volumetric MVS approaches. Our key idea is the use of a 3D scene-modeling network that iteratively updates a set of coarse depth predictions, resulting in highly accurate predictions which agree on the underlying scene geometry. Unlike existing depth-prediction techniques, our method uses a volumetric 3D convolutional neural network (CNN) that operates in world space on all depth maps jointly. The network can therefore learn meaningful scene-level priors. Furthermore, unlike existing volumetric MVS techniques, our 3D CNN operates on a feature-augmented point cloud, allowing for effective aggregation of multi-view information and flexible iterative refinement of depth maps. Experimental results show our method exceeds state-of-the-art accuracy in both depth prediction and 3D reconstruction metrics on the ScanNet dataset, as well as a selection of scenes from the TUM-RGBD and ICL-NUIM datasets. This shows that our method is both effective and generalizes to new settings. Alexander Rich 0001, Noah Stier, Pradeep Sen, Tobias Höllerer |
3DV | 2 |
| 2021 | VoRTX: Volumetric 3D Reconstruction With Transformers for Voxelwise View Selection and FusionabstractRecent volumetric 3D reconstruction methods can produce very accurate results, with plausible geometry even for unobserved surfaces. However, they face an undesirable trade-off when it comes to multi-view fusion. They can fuse all available view information by global averaging, thus losing fine detail, or they can heuristically cluster views for local fusion, thus restricting their ability to consider all views jointly. Our key insight is that greater detail can be retained without restricting view diversity by learning a view-fusion function conditioned on camera pose and image content. We propose to learn this multi-view fusion using a transformer. To this end, we introduce VoRTX,1an end-to-end volumetric 3D reconstruction network using transformers for wide-baseline, multi-view feature fusion. Our model is occlusion-aware, leveraging the transformer architecture to predict an initial, projective scene geometry estimate. This estimate is used to avoid back-projecting image features through surfaces into occluded regions. We train our model on ScanNet and show that it produces better reconstructions than state-of-the-art methods. We also demonstrate generalization without any fine-tuning, outperforming the same state-of-the-art methods on two other datasets, TUM-RGBD and ICL-NUIM. Noah Stier, Alexander Rich 0001, Pradeep Sen, Tobias Höllerer |
3DV | 1 |
| 2015 | Deep learning of tissue fate features in acute ischemic strokeabstractIn acute ischemic stroke treatment, prediction of tissue survival outcome plays a fundamental role in the clinical decision-making process, as it can be used to assess the balance of risk vs. possible benefit when considering endovascular clot-retrieval intervention. For the first time, we construct a deep learning model of tissue fate based on randomly sampled local patches from the hypoperfusion (Tmax) feature observed in MRI immediately after symptom onset. We evaluate the model with respect to the ground truth established by an expert neurologist four days after intervention. Experiments on 19 acute stroke patients evaluated the accuracy of the model in predicting tissue fate. Results show the superiority of the proposed regional learning framework versus a single-voxel-based regression model. Noah Stier, Nicholas Vincent, David S. Liebeskind, Fabien Scalzo |
BIBM | 1 |
| 2015 | Detection of hyperperfusion on arterial spin labeling using deep learningabstractHyperperfusion detected on arterial spin labeling (ASL) images acquired after acute stroke onset has been shown to correlate with development of subsequent intracerebral hemorrhage. We present in this study a quantitative hyperperfusion detection model that can provide an objective decision support for the interpretation of ASL cerebral blood flow (CBF) maps and rapidly delineate hyperperfusion regions. The detection problem is solved using Deep Learning such that the model relates ASL image patches to the corresponding label (normal or hyperperfused). Our method takes into account the regional intensity values of contralateral hemisphere during the labeling of a pixel. Each input vector is associated to a label corresponding to the presence of hyperperfusion that was manually established by a clinical researcher in Neurology. When compared to the manually established hyperperfusion, the predicted maps reached an accuracy of 97.45 ± 2.49% after crossvalidation. Pattern recognition based on deep learning can provide an accurate and objective measure of hyperperfusion on ASL CBF images and could therefore improve the detection of hemorrhagic transformation in acute stroke patients. Nicholas Vincent, Noah Stier, Songlin Yu, David S. Liebeskind, Danny J. J. Wang, Fabien Scalzo |
BIBM | 2 |