Jeroen van Baar

dblp:45/2996 · DBLP profile ↗
← Back
23ranked-venue papers
1as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 3 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 6Systems, architecture and hardware · 5 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2025 GeoDiffuser: Geometry-Based Image Editing with Diffusion Models
abstract
The success of image generative models has enabled us to build methods that can edit images based on text or other user input. However, these methods are imprecise, require additional information, or are limited to only 2D image edits. We present GeoDiffuser, a zero-shot optimization-based method that unifies common 2D and 3D image-based object editing capabilities into a single method. Our key insight is to view image editing operations as geometric transformations. We show that these transformations can be directly incorporated into the attention layers in diffusion models to implicitly perform editing operations. Our training-free optimization method uses an objective function that seeks to preserve object style but generate plausible images, for instance with accurate lighting and shadows. It also inpaints disoccluded parts of the image where the object was originally located. Given a natural image and user input, we segment the foreground object [27] and estimate a corresponding transform which is used by our optimization approach for editing. Figure 1 shows that GeoDiffuser can perform common 2D and 3D edits like object translation, 3D rotation, and removal. We present quantitative results, including a perceptual study, that shows how our approach is better than existing methods.
Rahul Sajnani, Jeroen van Baar, Jie Min, Kapil Katyal, Srinath Sridhar 0002
WACV2
2024 Scaling Object-centric Robotic Manipulation with Multimodal Object Identification
abstract
Robotic manipulation is a key enabler for automation in the fulfillment logistics sector. Such robotic systems require perception and manipulation capabilities to handle a wide variety of objects. Existing systems either operate on a closed set of objects or perform object-agnostic manipulation which lacks the capability for deliberate and reliable manipulation at scale. Object identification (ID) unlocks the ability for large-scale, object-centric manipulation by mapping object segments to one of the previously seen objects from a database. Nevertheless, it is often limited by the availability of reference data or coverage for objects in a database. In this work, we propose to perform object identification with multiple reference databases, including images and text references, each with a different coverage and matching challenge. We propose a training strategy that tackles the challenges of learning domain-invariant image embeddings, image-text matching and fusing predictions from different sources. We perform experiments over a recent benchmark with over 190K+ unique objects, extend the dataset with the additional reference sources and propose an evaluation strategy that simulates coverage for different reference sources. Model trained with the proposed learning pipeline shows robust performance over a range of simulation experiments.
Chaitanya Mitash, Mostafa Hussein, Jeroen van Baar, Vikedo Terhuja, Kapil Katyal
ICRA3
2022 Learning Occlusion-Aware Dense Correspondences for Multi-Modal Images
abstract
We introduce a scalable multi-modal approach to learn dense, i.e., pixel-level, correspondences and occlusion maps, between images in a video sequence. The problems of finding dense correspondences and occlusion maps are fundamental in computer vision. In this work we jointly train a deep network to tackle both, with a shared feature extraction stage. We use depth and color images with ground truth optical flow and occlusion maps to train the network end-to-end. From the multi-modal input, the network learns to estimate occlusion maps, optical flows, and a correspondence embedding providing a meaningful latent feature space. We evaluate the performance on a dataset of images derived from synthetic characters, and perform a thorough ablation study to demonstrate that the proposed components of our architecture combine to achieve the lowest correspondence error. The scalability of our proposed method comes from the ability to incorporate additional modalities, e.g., infrared images.
Ryosuke Shimoya, Takashi Morimoto, Jeroen van Baar, Petros Boufounos, Yanting Ma, Hassan Mansour
AVSS3
2022 Learning to Synthesize Volumetric Meshes from Vision-based Tactile Imprints
abstract
Vision-based tactile sensors typically utilize a deformable elastomer and a camera mounted above to provide high-resolution image observations of contacts. Obtaining accurate volumetric meshes for the deformed elastomer can provide direct contact information and benefit robotic grasping and manipulation. This paper focuses on learning to synthesize the volumetric mesh of the elastomer based on the image imprints acquired from vision-based tactile sensors. Synthetic image-mesh pairs and real-world images are gathered from 3D finite element methods (FEM) and physical sensors, respectively. A graph neural network (GNN) is introduced to learn the image-to-mesh mappings with supervised learning. A self-supervised adaptation method and image augmentation techniques are proposed to transfer networks from simulation to reality, from primitive contacts to unseen contacts, and from one sensor to another. Using these learned and adapted networks, our proposed method can accurately reconstruct the deformation of the real-world tactile sensor elastomer in various domains, as indicated by the quantitative and qualitative results.
Xinghao Zhu, Siddarth Jain, Masayoshi Tomizuka, Jeroen van Baar
ICRA4
2021 Joint 3D Human Shape Recovery and Pose Estimation from a Single Image with Bilayer Graph
abstract
The ability to estimate the 3D human shape and pose from images can be useful in many contexts. Recent approaches have explored using graph convolutional networks and achieved promising results. The fact that the 3D shape is represented by a mesh, an undirected graph, makes graph convolutional networks a natural fit for this problem. However, graph convolutional networks have limited representation power Information from nodes in the graph is passed to connected neighbors, and propagation of information requires successive graph convolutions. To overcome this limitation, we propose a dual-scale graph approach. We use a coarse graph, derived from a dense graph, to estimate the human’s 3D pose, and the dense graph to estimate the 3D shape. Information in coarse graphs can be propagated over longer distances compared to dense graphs. In addition, information about pose can guide to recover local shape detail and vice versa. We recognize that the connection between coarse and dense is itself a graph, and introduce graph fusion blocks to exchange information between graphs with different scales. We train our model end-to-end and show that we can achieve state-of-the-art results for several evaluation datasets. The code is available at the following link, https://github.com/yuxwind/BiGraphBody.
Xin Yu 0003, Jeroen van Baar, Siheng Chen
3DV2
2021 Cross-domain Imitation from Observations
abstract
Imitation learning seeks to circumvent the difficulty in designing proper reward functions for training agents by utilizing expert behavior. With environments modeled as Markov Decision Processes (MDP), most of the existing imitation algorithms are contingent on the availability of expert demonstrations in the same MDP as the one in which a new imitation policy is to be learned. In this paper, we study the problem of how to imitate tasks when discrepancies exist between the expert and agent MDP. These discrepancies across domains could include differing dynamics, viewpoint, or morphology; we present a novel framework to learn correspondences across such domains. Importantly, in contrast to prior works, we use unpaired and unaligned trajectories containing only states in the expert domain, to learn this correspondence. We utilize a cycle-consistency constraint on both the state space and a domain agnostic latent space to do this. In addition, we enforce consistency on the temporal position of states via a normalized position estimator function, to align the trajectories across the two domains. Once this correspondence is found, we can directly transfer the demonstrations on one domain to the other and use it for imitation. Experiments across a wide variety of challenging domains demonstrate the efficacy of our approach.
Dripta S. Raychaudhuri, Sujoy Paul, Jeroen van Baar, Amit K. Roy-Chowdhury
ICML3
2021 A Visual Inertial Odometry Framework for 3D Points, Lines and Planes
abstract
Recovering rigid registration between successive camera poses lies at the heart of 3D reconstruction, SLAM and visual odometry. Registration relies on the ability to compute discriminative 2D features in successive camera images for determining feature correspondences, which is very challenging in feature-poor environments, i.e. low-texture and/or low-light environments. In this paper, we aim to address the challenge of recovering rigid registration between successive camera poses in feature-poor environments in a Visual Inertial Odometry (VIO) setting. In addition to inertial sensing, we instrument a small aerial robot with an RGBD camera and propose a framework that unifies the incorporation of 3D geometric entities: points, lines, and planes. The tracked 3D geometric entities provide constraints in an Extended Kalman Filtering framework. We show that by directly exploiting 3D geometric entities, we can achieve improved registration. We demonstrate our approach on different texture-poor environments, with some containing only flat texture-less surfaces providing essentially no 2D features for tracking. In addition, we evaluate how the addition of different 3D geometric entities contributes to improved pose estimation by comparing an estimated pose trajectory to a ground truth pose trajectory obtained from a motion capture system. We consider computationally efficient methods for detecting 3D points, lines and planes, since our goal is to implement our approach on small mobile robots, such as drones.
Shenbagaraj Kannapiran, Jeroen van Baar, Spring Berman
IROS2
2020 DynamicsExplorer: Visual Analytics for Robot Control Tasks involving Dynamics and LSTM-based Control Policies
abstract
Deep reinforcement learning (RL), where a policy represented by a deep neural network is trained, has shown some success in playing video games and chess. However, applying RL to real-world tasks like robot control is still challenging. Because generating a massive number of samples to train control policies using RL on real robots is very expensive, hence impractical, it is common to train in simulations, and then transfer to real environments. The trained policy, however, may fail in the real world because of the difference between the training and the real environments, especially the difference in dynamics. To diagnose the problems, it is crucial for experts to understand (1) how the trained policy behaves under different dynamics settings, (2) which part of the policy affects the behaviors the most when the dynamics setting changes, and (3) how to adjust the training procedure to make the policy robust.This paper presents DynamicsExplorer, a visual analytics tool to diagnose the trained policy on robot control tasks under different dynamics settings. DynamicsExplorer allows experts to overview the results of multiple tests with different dynamics-related parameter settings so experts can visually detect failures and analyze the sensitivity of different parameters. Experts can further examine the internal activations of the policy for selected tests and compare the activations between success and failure tests. Such comparisons help experts form hypotheses about the policy and allows them to verify the hypotheses via DynamicsExplorer. Multiple use cases are presented to demonstrate the utility of DynamicsExplorer.
Teng-Yok Lee, Jeroen van Baar, Kent Wittenburg, Han-Wei Shen
PacificVis3
2020 Interactive Tactile Perception for Classification of Novel Object Instances
abstract
In this paper, we present a novel approach for classification of unseen object instances from interactive tactile feedback. Furthermore, we demonstrate the utility of a low resolution tactile sensor array for tactile perception that can potentially close the gap between vision and physical contact for manipulation. We contrast our sensor to high-resolution camera-based tactile sensors. Our proposed approach interactively learns a one-class classification model using 3D tactile descriptors, and thus demonstrates an advantage over the existing approaches, which require pre-training on objects. We describe how we derive 3D features from the tactile sensor inputs, and exploit them for learning one-class classifiers. In addition, since our proposed method uses unsupervised learning, we do not require ground truth labels. This makes our proposed method flexible and more practical for deployment on robotic systems. We validate our proposed method on a set of household objects and results indicate good classification performance in real-world experiments.
Radu Corcodel, Siddarth Jain, Jeroen van Baar
IROS3
2019 Sim-to-Real Transfer Learning using Robustified Controllers in Robotic Tasks involving Complex Dynamics
abstract
Learning robot tasks or controllers using deep reinforcement learning has been proven effective in simulations. Learning in simulation has several advantages. For example, one can fully control the simulated environment, including halting motions while performing computations. Another advantage when robots are involved, is that the amount of time a robot is occupied learning a task-rather than being productive-can be reduced by transferring the learned task to the real robot. Transfer learning requires some amount of fine-tuning on the real robot. For tasks which involve complex (non-linear) dynamics, the fine-tuning itself may take a substantial amount of time. In order to reduce the amount of fine-tuning we propose to learn robustified controllers in simulation. Robustified controllers are learned by exploiting the ability to change simulation parameters (both appearance and dynamics) for successive training episodes. An additional benefit for this approach is that it alleviates the precise determination of physics parameters for the simulator, which is a non-trivial task. We demonstrate our proposed approach on a real setup in which a robot aims to solve a maze game, which involves complex dynamics due to static friction and potentially large accelerations. We show that the amount of fine-tuning in transfer learning for a robustified controller is substantially reduced compared to a non-robustified controller.
Jeroen van Baar, Alan Sullivan, Radu Cordorel, Devesh K. Jha, Diego Romeres, Daniel Nikovski
ICRA1
2019 Learning from Trajectories via Subgoal Discovery
abstract
Learning to solve complex goal-oriented tasks with sparse terminal-only rewards often requires an enormous number of samples. In such cases, using a set of expert trajectories could help to learn faster. However, Imitation Learning (IL) via supervised pre-training with these trajectories may not perform as well and generally requires additional finetuning with expert-in-the-loop. In this paper, we propose an approach which uses the expert trajectories and learns to decompose the complex main task into smaller sub-goals. We learn a function which partitions the state-space into sub-goals, which can then be used to design an extrinsic reward function. We follow a strategy where the agent first learns from the trajectories using IL and then switches to Reinforcement Learning (RL) using the identified sub-goals, to alleviate the errors in the IL step. To deal with states which are under-represented by the trajectory set, we also learn a function to modulate the sub-goal predictions. We show that our method is able to solve complex goal-oriented tasks, which other RL, IL or their combinations in literature are not able to solve.
Sujoy Paul, Jeroen van Baar, Amit K. Roy-Chowdhury
NeurIPS2
2010 Stereoscopic 3D copy & paste
abstract
With the increase in popularity of stereoscopic 3D imagery for film, TV, and interactive entertainment, an urgent need for editing tools to support stereo content creation has become apparent. In this paper we present an end-to-end system for object copy & paste in a stereoscopic setting to address this need. There is no straightforward extension of 2D copy & paste to support the addition of the third dimension as we show in this paper. For stereoscopic copy & paste we need to handle depth, and our core objective is to obtain a convincing 3D viewing experience. As one of the main contributions of our system, we introduce a stereo billboard method for stereoscopic rendering of the copied selection. Our approach preserves the stereo volume and is robust to the inevitable inaccuracies of the depth maps computed from a stereo pair of images. Our system also includes an interactive stereoscopic segmentation tool to achieve high quality object selection. Hence, we focus on intuitive and minimal user interaction, and our editing operations perform within interactive rates to provide immediate feedback.
Wan-Yen Lo, Jeroen van Baar, Claude Knaus, Matthias Zwicker, Markus Gross 0001
ACM Trans. Graph.2
2009 A Unified Calibration Method with a Parametric Approach for Wide-Field-of-View Multiprojector Displays
abstract
In this paper, we describe techniques for supporting a wide-field-of-view multiprojector curved screen display system. Our main contribution is in achieving automatic geometric calibration and efficient rendering for seamless displays, which is effective even in the presence of panoramic surround screens with the multiview calibration method without polygonal representation of the display surface. We show several prototype systems that use a stereo camera for capturing and a new rendering method for quadric curved screens. Previous approaches have required a calibration camera at the sweet spot. Due to parameterized representation, however, our unified calibration method is independent of the orientation and field of view of the calibration camera. This method can simplify the tedious and complicated installation process as well as the maintenance of large multiprojector displays in planetariums, virtual reality systems, and other visualization venues.
Masato Ogata, Hiroyuki Wada, Jeroen van Baar, Ramesh Raskar
VR3
2006 A Handheld Projector Supported by Computer Vision
Akash Kushal, Jeroen van Baar, Ramesh Raskar, Paul A. Beardsley
ACCV (2)2
2005 Zoom-and-pick: facilitating visual zooming and precision pointing with interactive handheld projectors
abstract
Designing interfaces for interactive handheld projectors is an exiting new area of research that is currently limited by two problems: hand jitter resulting in poor input control, and possible reduction of image resolution due to the needs of image stabilization and warping algorithms. We present the design and evaluation of a new interaction technique, called zoom-and-pick, that addresses both problems by allowing the user to fluidly zoom in on areas of interest and make accurate target selections. Subtle design features of zoom-and-pick enable pixel-accurate pointing, which is not possible in most freehand interaction techniques. Our evaluation results indicate that zoom-and-pick is significantly more accurate than the standard pointing technique described in our previous work.
Clifton Forlines, Ravin Balakrishnan, Paul A. Beardsley, Jeroen van Baar, Ramesh Raskar
UIST4
2004 Multi-projectors and implicit interaction in persuasive public displays
abstract
Recent advances in computer video projection open up new possibilities for real-time interactive, persuasive displays. Now a display can continuously adapt to a viewer so as to maximize its effectiveness. However, by the very nature of persuasion, these displays must be both immersive and subtle. We have been working on technologies that support this application including multi-projector and implicit interaction techniques. These technologies have been used to create a series of interactive persuasive displays that are described.
Paul H. Dietz, Ramesh Raskar, Shane Booth, Jeroen van Baar, Kent Wittenburg, Brian Knep
AVI4
2004 Quadric Transfer for Immersive Curved Screen Displays
abstract
Abstract Curved screens are increasingly being used for high‐resolution immersive visualization environments. We describe a new technique to display seamless images using overlapping projectors on curved quadric surfaces such as spherical or cylindrical shape. We exploit a quadric image transfer function and show how it can be used to achieve sub‐pixel registration while interactively displaying two or three‐dimensional datasets for a head‐tracked user. Current techniques for automatically registered seamless displays have focused mainly on planar displays. On the other hand, techniques for curved screens currently involve cumbersome manual alignment to make the installation conform to the intended design. We show a seamless real‐time display system and discuss our methods for smooth intensity blending and efficient rendering. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Three‐Dimensional Graphics and Realism‐ Virtual reality
Ramesh Raskar, Jeroen van Baar, Thomas Willwacher, Srinivas Rao
Comput. Graph. Forum2
2004 RFIG lamps: interacting with a self-describing world via photosensing wireless tags and projectors
abstract
This paper describes how to instrument the physical world so that objects become self-describing, communicating their identity, geometry, and other information such as history or user annotation. The enabling technology is a wireless tag which acts as a radio frequency identity and geometry (RFIG) transponder. We show how addition of a photo-sensor to a wireless tag significantly extends its functionality to allow geometric operations - such as finding the 3D position of a tag, or detecting change in the shape of a tagged object. Tag data is presented to the user by direct projection using a handheld locale-aware mobile projector. We introduce a novel technique that we call interactive projection to allow a user to interact with projected information e.g. to navigate or update the projected information.The ideas are demonstrated using objects with active radio frequency (RF) tags. But the work was motivated by the advent of unpowered passive-RFID, a technology that promises to have significant impact in real-world applications. We discuss how our current prototypes could evolve to passive-RFID in the future.
Ramesh Raskar, Paul A. Beardsley, Jeroen van Baar, Paul H. Dietz, Johnny C. Lee, Darren Leigh, Thomas Willwacher
ACM Trans. Graph.3
2003 iLamps: geometrically aware and self-configuring projectors
abstract
Projectors are currently undergoing a transformation as they evolve from static output devices to portable, environment-aware, communicating systems. An enhanced projector can determine and respond to the geometry of the display surface, and can be used in an ad-hoc cluster to create a self-configuring display. Information display is such a prevailing part of everyday life that new and more flexible ways to present data are likely to have significant impact. This paper examines geometrical issues for enhanced projectors, relating to customized projection for different shapes of display surface, object augmentation, and co-operation between multiple units.We introduce a new technique for adaptive projection on nonplanar surfaces using conformal texture mapping. We describe object augmentation with a hand-held projector, including interaction techniques. We describe the concept of a display created by an ad-hoc cluster of heterogeneous enhanced projectors, with a new global alignment scheme, and new parametric image transfer methods for quadric surfaces, to make a seamless projection. The work is illustrated by several prototypes and applications.
Ramesh Raskar, Jeroen van Baar, Paul A. Beardsley, Thomas Willwacher, Srinivas Rao, Clifton Forlines
ACM Trans. Graph.2
2002 EWA Splatting
abstract
We present a framework for high quality splatting based on elliptical Gaussian kernels. To avoid aliasing artifacts, we introduce the concept of a resampling filter, combining a reconstruction kernel with a low-pass filter. Because of the similarity to Heckbert's (1989) EWA (elliptical weighted average) filter for texture mapping, we call our technique EWA splatting. Our framework allows us to derive EWA splat primitives for volume data and for point-sampled surface data. It provides high image quality without aliasing artifacts or excessive blurring for volume data and, additionally, features anisotropic texture filtering for point-sampled surfaces. It also handles nonspherical volume kernels efficiently; hence, it is suitable for regular, rectilinear, and irregular volume datasets. Moreover, our framework introduces a novel approach to compute the footprint function, facilitating efficient perspective projection of arbitrary elliptical kernels at very little additional cost. Finally, we show that EWA volume reconstruction kernels can be reduced to surface reconstruction kernels. This makes our splat primitive universal in rendering surface and volume data.
Matthias Zwicker, Hanspeter Pfister, Jeroen van Baar, Markus Gross 0001
IEEE Trans. Vis. Comput. Graph.3
2001 Surface splatting
abstract
Modern laser range and optical scanners need rendering techniques that can handle millions of points with high resolution textures. This paper describes a point rendering and texture filtering technique called surface splatting which directly renders opaque and transparent surfaces from point clouds without connectivity. It is based on a novel screen space formulation of the Elliptical Weighted Average (EWA) filter. Our rigorous mathematical analysis extends the texture resampling framework of Heckbert to irregularly spaced point samples. To render the points, we develop a surface splat primitive that implements the screen space EWA filter. Moreover, we show how to optimally sample image and procedural textures to irregular point data during pre-processing. We also compare the optimal algorithm with a more efficient view-independent EWA pre-filter. Surface splatting makes the benefits of EWA texture filtering available to point-based rendering. It provides high quality anisotropic texture filtering, hidden surface removal, edge anti-aliasing, and order-independent transparency.
Matthias Zwicker, Hanspeter Pfister, Jeroen van Baar, Markus Gross 0001
SIGGRAPH3
2001 EWA Volume Splatting
abstract
In this paper we present a novel framework for direct volume rendering using a splatting approach based on elliptical Gaussian kernels. To avoid aliasing artifacts, we introduce the concept of a resampling filter combining a reconstruction with a low-pass kernel. Because of the similarity to Heckbert's EWA (elliptical weighted average) filter for texture mapping we call our technique EWA volume splatting. It provides high image quality without aliasing artifacts or excessive blurring even with non-spherical kernels. Hence it is suitable for regular, rectilinear, and irregular volume data sets. Moreover, our framework introduces a novel approach to compute the footprint function. It facilitates efficient perspective projection of arbitrary elliptical kernels at very little additional cost. Finally, we show that EWA volume reconstruction kernels can be reduced to surface reconstruction kernels. This makes our splat primitive universal in reconstructing surface and volume data.
Matthias Zwicker, Hanspeter Pfister, Jeroen van Baar, Markus Gross 0001
IEEE Visualization3
2000 Surfels: surface elements as rendering primitives
abstract
Surface elements (surfels) are a powerful paradigm to efficiently render complex geometric objects at interactive frame rates. Unlike classical surface discretizations, i.e., triangles or quadrilateral meshes, surfels are point primitives without explicit connectivity. Surfel attributes comprise depth, texture color, normal, and others. As a pre-process, an octree-based surfel representation of a geometric object is computed. During sampling, surfel positions and normals are optionally perturbed, and different levels of texture colors are prefiltered and stored per surfel. During rendering, a hierarchical forward warping algorithm projects surfels to a z-buffer. A novel method called visibility splatting determines visible surfels and holes in the z-buffer. Visible surfels are shaded using texture filtering, Phong illumination, and environment mapping using per-surfel normals. Several methods of image reconstruction, including supersampling, offer flexible speed-quality tradeoffs. Due to the simplicity of the operations, the surfel rendering pipeline is amenable for hardware implementation. Surfel objects offer complex shape, low rendering cost and high image quality, which makes them specifically suited for low-cost, real-time graphics, such as games.
Hanspeter Pfister, Matthias Zwicker, Jeroen van Baar, Markus Gross 0001
SIGGRAPH3