Benjamin Simon

dblp:210/9653 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MAISI-v2: Accelerated 3D High-Resolution Medical Image Synthesis with Rectified Flow and Region-specific Contrastive Loss
abstract
Medical image synthesis is an important topic for both clinical and research applications. Recently, diffusion models have become a leading approach in this area. Despite their strengths, many existing methods struggle with (1) limited generalizability, only working for specific body regions or voxel spacings, (2) slow inference, which is a common issue for diffusion models, and (3) weak alignment with input conditions, which is a critical issue for medical imaging. MAISI, a previously proposed framework, addresses generalizability issues but still suffers from slow inference and limited condition consistency. In this work, we present MAISI-v2, the first accelerated 3D medical image synthesis framework that integrates rectified flow to enable fast and high-quality generation. To further enhance condition fidelity, we introduce a novel region-specific contrastive loss to improve sensitivity to the region of interest. Our experiments show that MAISI-v2 can achieve state-of-the-art image quality with 33× acceleration for latent diffusion models. We also conducted a downstream segmentation experiment to show that the synthetic images can be used for data augmentation. We release our code, training details, model weights, and a GUI demo to facilitate reproducibility and promote further development within the community.
Can Zhao 0001, Dong Yang 0005, Yufan He, Yucheng Tang, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey, Daguang Xu
AAAI6
2025 VISTA3D: A Unified Segmentation Foundation Model For 3D Medical Imaging
abstract
Foundation models for interactive segmentation in 2D natural images and videos have sparked significant interest in building 3D foundation models for medical imaging. However, the domain gaps and clinical use cases for 3D medical imaging require a dedicated model that diverges from existing 2D solutions. Specifically, such foundation models should support a full workflow that can actually reduce human effort. Treating 3D medical images as sequences of 2D slices and reusing interactive 2D foundation models seems straightforward, but 2D annotation is too time-consuming for 3D tasks. Moreover, for large cohort analysis, it’s the highly accurate automatic segmentation models that reduce the most human effort. However, these models lack support for interactive corrections and lack zero-shot ability for novel structures, which is a key feature of "foundation". While reusing pre-trained 2D backbones in 3D enhances zero-shot potential, their performance on complex 3D structures still lags behind leading 3D models. To address these issues, we present VISTA3D, Versatile Imaging SegmenTation and Annotation model, that targets to solve all these challenges and requirements with one unified foundation model. VISTA3D is built on top of the well-established 3D segmentation pipeline, and it is the first model to achieve state-of-the-art performance in both 3D automatic (supporting 127 classes) and 3D interactive segmentation, even when compared with top 3D expert models on large and diverse benchmarks. Additionally, VISTA3D’s 3D interactive design allows efficient human correction, and a novel 3D supervoxel method that distills 2D pre-trained backbones grants VISTA3D top 3D zero-shot performance. We believe the model, recipe, and insights represent a promising step towards a clinically useful 3D foundation model. Code and weights are publicly available at https://github.com/Project-MONAI/VISTA.
Yufan He, Yucheng Tang, Andriy Myronenko, Vishwesh Nath, Ziyue Xu 0001, Dong Yang 0005, Can Zhao 0001, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey, Daguang Xu, Wenqi Li 0001
CVPR9
2025 VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge
abstract
Generalist vision language models (VLMs) have made significant strides in computer vision, but they fall short in specialized fields like healthcare, where expert knowledge is essential. Current large multimodal models like Gemini and GPT-4o are insufficient for medical tasks due to their reliance on memorized internet knowledge rather than the nuanced expertise required in healthcare. Meanwhile, existing medical VLMs (e.g. Med-Gemini) often lack expert consultation as part of their design, and many rely on outdated, static datasets that were not created with modern, large deep learning models in mind. VLMs are usually trained in three stages: vision pre-training, vision-language pre-training, and instruction fine-tuning (IFT). IFT has been typically applied using a mixture of generic and healthcare data. In contrast, we propose that for medical VLMs, a fourth stage of specialized IFT is necessary, which focuses on medical data and includes information from domain expert models. Domain expert models developed for medical use are crucial because they are specifically trained for certain clinical tasks, e.g. to detect tumors and classify abnormalities through segmentation and classification, which learn fine-grained features of medical data−features that are often too intricate for a VLM to capture effectively. This paper introduces a new framework, VILA-M3, for medical VLMs that utilizes domain knowledge via expert models. We argue that generic VLM architectures alone are not viable for real-world clinical applications and on-demand usage of domain-specialized expert model knowledge is critical for advancing AI in healthcare. Through our experiments, we show an improved state-of-the-art (SOTA) performance with an average improvement of ~9% over the prior SOTA model Med-Gemini and ~6% over models trained on the specific tasks. Our approach emphasizes the importance of domain expertise in creating precise, reliable VLMs for medical applications.
Vishwesh Nath, Wenqi Li 0001, Dong Yang 0005, Andriy Myronenko, Mingxin Zheng, Yao Lu 0006, Hongxu Yin, Yee Man Law, Yucheng Tang, Can Zhao 0001, Ziyue Xu 0001, Yufan He, Stephanie A. Harmon, Benjamin Simon, Greg Heinrich, Stephen R. Aylward, Marc Edgar, Michael Zephyr, Pavlo Molchanov 0001, Baris Turkbey, Holger Roth, Daguang Xu
CVPR16
2025 Funnelectro: Electrotactile Funneling Illusion and Localization Performance on the Forearm
abstract
On the skin, a small distance between two actuators can result in the perception of a single, centralized stimulus rather than two distinct stimuli – this phenomenon is known as the funneling illusion. In this work, we explore the electrotactile funneling illusion on the forearm in a user study with 16 participants. We placed an electrode strip with 9 electrode pairs along their forearm – from wrist to elbow. The calibration of the same perceived intensity of each electrode pair for each participant shows that the calibrated intensities near the wrist are significantly higher compared to the intensities calibrated near the elbow. A linear regression corresponds to this behavior as well as the qualitative feedback of our participants. Based on this, we created an equation that helps to reduce the calibration time considerably. The results of our study show that the funneling illusion can be reliably evoked at distances of up to 7.2 cm. Further, we provide detailed information on occurrence frequency and precision and explored whether an approach adapted for electrotactile feedback can create an apparent tactile motion.
Benjamin Simon, Dennis Stanke, Michael Rohs
MUM1
2025 MAISI: Medical AI for Synthetic Imaging
abstract
Medical imaging analysis faces challenges such as data scarcity, high annotation costs, and privacy concerns. This paper introduces the Medical AI for Synthetic Imaging (MAISI), an innovative approach using the diffusion model to generate synthetic 3D computed tomography (CT) images to address those challenges. MAISI leverages the foundation volume compression network and the latent diffusion model to produce high-resolution CT images (up to a landmark volume dimension of 512 × 512 × 768) with flexible volume dimensions and voxel spacing. By incorporating ControlNet, MAISI can process organ segmentation, including 127 anatomical structures, as additional conditions and enables the generation of accurately annotated synthetic images that can be used for various downstream tasks. Our experiment results show that MAISI's capabilities in generating realistic, anatomically accurate images for diverse regions and conditions reveal its promising potential to mitigate challenges using synthetic data.
Can Zhao 0001, Dong Yang 0005, Ziyue Xu 0001, Vishwesh Nath, Yucheng Tang, Benjamin Simon, Mason Belue, Stephanie A. Harmon, Baris Turkbey, Daguang Xu
WACV7
2024 Automated Detection and Characterization of Small Cell Lung Cancer Liver Metastases on CT
Sophia Ty, Fahmida Haque, Parth Desai, Nobuyuki Takahashi, Usamah Chaudhary, Benjamin Simon, Peter L. Choyke, Anish Thomas, Baris Turkbey, Stephanie A. Harmon
AIME (2)6
2024 RANsacked: A Domain-Informed Approach for Fuzzing LTE and 5G RAN-Core Interfaces
abstract
Cellular network infrastructure serves as the backbone of modern mobile wireless communication. As such, cellular cores must be proactively secured against external threats to ensure reliable service. Compromised base station attacks against the core are a rising threat to cellular networks, while user device inputs have long been considered as an attack vector; despite this, few techniques exist to comprehensively test RAN-Core interfaces against malicious input. In this work, we devise a fuzzing framework that performantly fuzzes cellular interfaces accessible from a base station or user device, overcoming several challenges in fuzzing specific to LTE/5G network components. We also introduce ASNFuzzGen, a tool that compiles ASN.1 specifications into structure-aware fuzzing modules, thereby facilitating effective fuzzing exploration of complex cellular protocols. We run fuzzing campaigns against seven open-source and commercial cores and discover 119 vulnerabilities, with 93 CVEs assigned. Our results reveal common implementation mistakes across several cores that lead to vulnerabilities, and the successful coordination of patches for these vulnerabilities across several vendors demonstrates the practical impact ASNFuzzGen has on hardening user-exposed cellular systems.
Nathaniel Bennett, Weidong Zhu 0002, Benjamin Simon, Ryon Kennedy, William Enck, Patrick Traynor, Kevin R. B. Butler
CCS3
2024 CaseTouch: Occlusion-Free Touch Input by adding a Thin Sensor Stripe to the Smartwatch Case
abstract
Operating small touchscreens with the finger occludes a large part of the screen. We propose using the watch case as the input space, without enlarging the smartwatch. Therefore, we created two prototypes, one with a touch surface on the watch case (CASE) and one with touch surfaces on the watch case and the wristband (CASE+BAND). In a comparative study, we analyze their suitability in a 1D list scrolling task and 2D map navigation task. The results show that occlusion is less of a problem for the list scrolling task, as visibility is sufficient. In the map navigation task, participants reached task completion times with CASE+BAND that are comparable to touch input. CASE was significantly slower, but only requires minimal additional hardware. However, the results of a subsequent longitudinal study demonstrates the learnability of CASE, which led to task completion times comparable to touch input, and provides insights in the gradual development of expert performance.
Dennis Stanke, Benjamin Simon, Sergej Löwen, Michael Rohs
MUM2
2024 Shock Me The Way: Directional Electrotactile Feedback under the Smartwatch as a Navigation Aid for Cyclists
abstract
Cycling navigation is a complex and stressful task as the cyclist needs to focus simultaneously on the navigation, the road, and other road users. We propose directional electrotactile feedback at the wrist to reduce the auditory and visual load during navigation-aided cycling. We designed a custom electrotactile grid with 9 electrodes that is clipped under a smartwatch. In a preliminary study we identified suitable calibration settings and gained first insights about a suitable electrode layout. In a subsequent laboratory study we showed that a direction can be encoded with a mean error of 19.28\,° (σ = 42.77°) by combining 2 adjacent electrodes. Additionally, by interpolating with 3 electrodes a direction can be conveyed with a similar mean error of 22.54° (σ = 43.57°). We evaluated our concept of directional electrotactile feedback for cyclists in an outdoor study, in which 98.8% of all junctions were taken correctly by eight study participants. Only one participant deviated substantially from the optimal path, but was successfully navigated back to the original route by our system.
Tim Dünte, Dennis Stanke, Moritz Klose, Benjamin Simon, Ibraheem Al-Azzawi, Michael Rohs
Proc. ACM Hum. Comput. Interact.4
2021 VRTactileDraw: A Virtual Reality Tactile Pattern Designer for Complex Spatial Arrangements of Actuators
Oliver Beren Kaul, Andreas Domin, Michael Rohs, Benjamin Simon, Maximilian Schrapel
INTERACT (5)4
2021 Around-the-Head Tactile System for Supporting Micro Navigation of People with Visual Impairments
abstract
Tactile patterns are a means to convey navigation instructions to pedestrians and are especially helpful for people with visual impairments. This article presents a concept to provide precise micro-navigation instructions through a tactile around-the-head display. Our system presents four tactile patterns for fundamental navigation instructions in conjunction with continuous directional guidance. We followed an iterative, user-centric approach to design the patterns for the fundamental navigation instructions, combined them with a continuous directional guidance stimulus, and tested our system with 13 sighted (blindfolded) and 2 blind participants in an obstacle course, including stairs. We optimized the patterns and validated the final prototype with another five blind participants in a follow-up study. The system steered our participants successfully with a 5.7 cm average absolute deviation from the optimal path. Our guidance is only a little less precise than the usual shoulder wobbling during normal walking and an order of magnitude more precise than previous tactile navigation systems. Our system allows various new use cases of micro-navigation for people with visual impairments, e.g., preventing collisions on a sidewalk or as an anti-veering tool. It also has applications in other areas, such as personnel working in low-vision environments (e.g., firefighters).
Oliver Beren Kaul, Michael Rohs, Marc Mogalle, Benjamin Simon
ACM Trans. Comput. Hum. Interact.4
2020 Vibrotactile Funneling Illusion and Localization Performance on the Head
abstract
The vibrotactile funneling illusion is the sensation of a single (non-existing) stimulus somewhere in-between the actual stimulus locations. Its occurrence depends upon body location, distance between the actuators, signal synchronization, and intensity. Related work has shown that the funneling illusion may occur on the forehead. We were able to reproduce these findings and explored five further regions to get a more complete picture of the occurrence of the funneling illusion on the head. The results of our study (24 participants) show that the actuator distance, for which the funneling illusion occurs, strongly depends upon the head region. Moreover, we evaluated the centralizing bias (smaller perceived than actual actuator distances) for different head regions, which also showed widely varying characteristics. We computed a detailed heat map of vibrotactile localization accuracies on the head. The results inform the design of future tactile head-mounted displays that aim to support the funneling illusion.
Oliver Beren Kaul, Michael Rohs, Benjamin Simon, Kerem Can Demir, Kamillo Ferry
CHI3
2017 Mission planning for a multi-robot team with a solar-powered charging station
abstract
This paper presents a mission planning problem for a cooperative team of unmanned ground vehicles (UGVs), which includes multiple rovers and a solar-powered mobile charging station. The team is required to start at an initial point and visit a series of objective points before arriving at the final point selected from the set of objective points, where the UGVs will be charged from the solar-powered mobile charging station. This mission is represented as a multi-Hamiltonian Path Problem (mHPP). In order to effectively coordinate the team, an understanding of the mission environment is first obtained by generating a scalar field representation of the solar insolation of the environment from a visual-spectrum image. Then, a cascaded heuristic optimization algorithm, using modified genetic algorithm and particle swarm optimization, is used to generate a time-optimized mission plan for the team of UGVs, which guides each UGV to its assigned objective points and then rendezvous at the final charging location while guaranteeing compliance with the net energy gain constraint. The feasibility and efficiency of the proposed algorithm are verified using an experimental testbed and constructed indoor simulation environments.
Nathaniel Kingry, Yen-Chen Liu, Matthew Martinez, Benjamin Simon, YunQi Bang, Ran Dai
IROS4