Yizhak Ben-Shabat

dblp:195/6247 · DBLP profile ↗
← Back
21ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0001-7547-7493ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 16 · 6 first-author · 12 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 WiSAR3D - Aerial LiDAR dataset for 3D object detection
abstract
Wilderness Search and Rescue (WiSAR) operations aim to locate and rescue individuals in remote, rugged environments. We introduce WiSAR3D, the first aerial LiDAR 3D dataset specifically tailored for 3D object detection in rural scenes. WiSAR3D comprises 2633 strips—contiguous point clouds collected along the flight path, all fully annotated with 67.5K 3D bounding boxes for 22 categories. As the data is captured with multiple returns (echoes), it enables detection under foliage, which is impossible with RGB cameras. Additionally, as this is the first-of-its-kind dataset, we provide benchmark evaluations for state-of-the-art methods developed for autonomous cars on our aerial dataset. The results suggest that specialized models are needed in this domain, as finding individuals remains challenging for current 3D detectors. The code and dataset are publicly available at https://github.com/oshrout/WiSAR3D.
Oren Shrout, Ori Nizan, Yizhak Ben-Shabat, Ayellet Tal
WACV3
2025 VI^3NR: Variance Informed Initialization for Implicit Neural Representations
abstract
Implicit Neural Representations (INRs) are a versatile and powerful tool for encoding various forms of data, including images, videos, sound, and 3D shapes. A critical factor in the success of INRs is the initialization of the network, which can significantly impact the convergence and accuracy of the learned model. Unfortunately, commonly used neural network initializations are not widely applicable for many activation functions, especially those used by INRs. In this paper, we improve upon previous initialization methods by deriving an initialization that has stable variance across layers, and applies to any activation function. We show that this generalizes many previous initialization methods, and has even better stability for well studied activations. We also show that our initialization leads to improved results with INR activation functions in multiple signal modalities. Our approach is particularly effective for Gaussian INRs, where we demonstrate that the theory of our initialization matches with task performance in multiple experiments, allowing us to achieve improvements in image, audio, and 3D surface reconstruction.
Chamin Hewa Koneputugodage, Yizhak Ben-Shabat, Sameera Ramasinghe, Stephen Gould
CVPR2
2025 Less is More: Improving Motion Diffusion Models with Sparse Keyframes
abstract
Recent advances in motion diffusion models have led to remarkable progress in diverse motion generation tasks, including text-to-motion synthesis. However, existing approaches represent motions as dense frame sequences, requiring the model to process redundant or less informative frames. The processing of dense animation frames imposes significant training complexity, especially when learning intricate distributions of large motion datasets even with modern neural architectures. This severely limits the performance of generative motion models for downstream tasks. Inspired by professional animators who mainly focus on sparse keyframes, we propose a novel diffusion framework explicitly designed around sparse and geometrically meaningful keyframes. Our method reduces computation by masking non-keyframes and efficiently interpolating missing frames. We dynamically refine the keyframe mask during inference to prioritize informative frames in later diffusion steps. Extensive experiments show that our approach consistently outperforms state-of-the-art methods in text alignment and motion realism, while also effectively maintaining high performance at significantly fewer diffusion steps. We further validate the robustness of our framework by using it as a generative prior and adapting it to different downstream tasks.
Jinseok Bae, Inwoo Hwang, Young Yoon Lee, Yizhak Ben-Shabat, Young Min Kim 0001, Mubbasir Kapadia
ICCV6
2025 GEOPARD: Geometric Pretraining for Articulation Prediction in 3D Shapes
abstract
We present GEOPARD, a transformer-based architecture for predicting articulation from a single static snapshot of a 3D shape. The key idea of our method is a pretraining strategy that allows our transformer to learn plausible candidate articulations for 3D shapes based on a geometric-driven search without manual articulation annotation. The search automatically discovers physically valid part motions that do not cause detachments or collisions with other shape parts. Our experiments indicate that this geometric pretraining strategy, along with carefully designed choices in our transformer architecture, yields state-of-the-art results in articulation inference in the PartNet-Mobility dataset.
Pradyumn Goyal, Dmitry Petrov, Sheldon Andrews, Yizhak Ben-Shabat, Hsueh-Ti Derek Liu, Evangelos Kalogerakis
ICCV4
2025 StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross Fusion
abstract
We present StyleMotif, a novel Stylized Motion Latent Diffusion model, generating motion conditioned on both content and style from multiple modalities. Unlike existing approaches that either focus on generating diverse motion content or transferring style from sequences, StyleMotif seamlessly synthesizes motion across a wide range of content while incorporating stylistic cues from multi-modal inputs, including motion, text, image, video, and audio. To achieve this, we introduce a style-content cross fusion mechanism and align a style encoder with a pre-trained multi-modal model, ensuring that the generated motion accurately captures the reference style while preserving realism. Extensive experiments demonstrate that our framework surpasses existing methods in stylized motion generation and exhibits emergent capabilities for multi-modal motion stylization, enabling more nuanced motion synthesis. Source code and pre-trained models will be released upon acceptance. Project Page: https://stylemotif.github.io
Yizhak Ben-Shabat, Young Yoon Lee, Victor Zordan, Mubbasir Kapadia
ICCV2
2025 Temporally Grounding Instructional Diagrams in Unconstrained Videos
Frederic Z. Zhang, Cristian Rodriguez, Yizhak Ben-Shabat, Anoop Cherian, Stephen Gould
WACV4
2024 3DInAction: Understanding Human Actions in 3D Point Clouds
abstract
We propose a novel method for 3D point cloud action recognition. Understanding human actions in RGB videos has been widely studied in recent years, however, its 3D point cloud counterpart remains under-explored despite the clear value that 3D information may bring. This is mostly due to the inherent limitation of the point cloud data modality-lack of structure, permutation invariance, and varying number of points-which makes it difficult to learn a spatiotemporal representation. To address this limitation, we propose the 3DinAction pipeline that first estimates patches moving in time (t-patches) as a key building block, alongside a hierarchical architecture that learns an informative spatiotemporal representation. We show that our method achieves improved performance on existing datasets, including DFAUST and IKEA ASM. Code is publicly available at https://github.com/sitzikbs/3dincaction.
Yizhak Ben-Shabat, Oren Shrout, Stephen Gould
CVPR1
2024 Small Steps and Level Sets: Fitting Neural Surface Models with Point Guidance
abstract
A neural signed distance function (SDF) is a convenient shape representation for many tasks, such as surface recon-struction, editing and generation. However, neural SDFs are difficult to fit to raw point clouds, such as those sam-pled from the surface of a shape by a scanner. A major is-sue occurs when the shape's geometry is very differentfrom the structural biases implicit in the network's initialization. In this case, we observe that the standard loss formulation does not guide the network towards the correct SDF val-ues. We circumvent this problem by introducing guiding points, and use them to steer the optimization towards the true shape via small incremental changes for which the loss formulation has a good descent direction. We show that this point-guided homotopy-based optimization scheme fa-cilitates a deformation from an easy problem to the diffi-cult reconstruction problem. We also propose a metric to quantify the difference in surface geometry between a target shape and an initial surface, which helps indicate whether the standard loss formulation is guiding towards the target shape. Our method outperforms previous state-of-the-art approaches, with large improvements on shapes identified by this metric as particularly challenging.
Chamin Hewa Koneputugodage, Yizhak Ben-Shabat, Dylan Campbell, Stephen Gould
CVPR2
2024 Neural Experts: Mixture of Experts for Implicit Neural Representations
abstract
Implicit neural representations (INRs) have proven effective in various tasks including image, shape, audio, and video reconstruction. These INRs typically learn the implicit field from sampled input points. This is often done using a single network for the entire domain, imposing many global constraints on a single function. In this paper, we propose a mixture of experts (MoE) implicit neural representation approach that enables learning local piece-wise continuous functions that simultaneously learns to subdivide the domain and fit it locally. We show that incorporating a mixture of experts architecture into existing INR formulations provides a boost in speed, accuracy, and memory requirements. Additionally, we introduce novel conditioning and pretraining methods for the gating network that improves convergence to the desired solution. We evaluate the effectiveness of our approach on multiple reconstruction tasks, including surface reconstruction, image reconstruction, and audio signal reconstruction and show improved performance compared to non-MoE methods. Code is available at our project page https://sitzikbs.github.io/neural-experts-projectpage/ .
Yizhak Ben-Shabat, Chamin Hewa Koneputugodage, Sameera Ramasinghe, Stephen Gould
NeurIPS1
2024 IKEA Ego 3D Dataset: Understanding furniture assembly actions from ego-view 3D Point Clouds
abstract
We propose a novel dataset for ego-view 3D point cloud action recognition. While there has been extensive research on understanding human actions in RGB videos in recent years, the exploration of its 3D point cloud counterpart has been relatively limited. Furthermore, RGB ego-view datasets are rapidly growing, however, 3D point cloud ego-view datasets are scarce at best. Existing 3D datasets are limited in several ways, some include actions that are distinguishable by full-body motion while others use a distant static sensor that hinders the recognition of small objects. We introduce a new point cloud action recognition dataset—the IKEA Ego 3D dataset. It includes sequences of point clouds captured from an ego-view using a HoloLens 2 device. The dataset consists of approximately 493k frames and 56 classes of intricate furniture assembly actions of four different furniture types. We evaluate the performance of various state-of-the-art 3D action recognition methods on the proposed dataset and show that it is very challenging.
Yizhak Ben-Shabat, Jonathan Paul, Eviatar Segev, Oren Shrout, Stephen Gould
WACV1
2023 Octree Guided Unoriented Surface Reconstruction
abstract
We address the problem of surface reconstruction from unoriented point clouds. Implicit neural representations (INRs) have become popular for this task, but when information relating to the inside versus outside of a shape is not available (such as shape occupancy, signed distances or surface normal orientation) optimization relies on heuristics and regularizers to recover the surface. These methods can be slow to converge and easily get stuck in local minima. We propose a two-step approach, OG-INR, where we (1) construct a discrete octree and label what is inside and outside (2) optimize for a continuous and high-fidelity shape using an INR that is initially guided by the octree's labelling. To solve for our labelling, we propose an energy function over the discrete structure and provide an efficient move-making algorithm that explores many possible labellings. Furthermore we show that we can easily inject knowledge into the discrete octree, providing a simple way to influence the result from the continuous INR. We evaluate the effectiveness of our approach on two unoriented surface reconstruction datasets and show competitive performance compared to other unoriented, and some oriented, methods. Our results show that the exploration by the move-making algorithm avoids many of the bad local minima reached by purely gradient descent optimized methods (see Figure 1).
Chamin Hewa Koneputugodage, Yizhak Ben-Shabat, Stephen Gould
CVPR2
2023 GraVoS: Voxel Selection for 3D Point-Cloud Detection
abstract
3D object detection within large 3D scenes is challenging not only due to the sparsity and irregularity of 3D point clouds, but also due to both the extreme foreground-background scene imbalance and class imbalance. A common approach is to add ground-truth objects from other scenes. Differently, we propose to modify the scenes by removing elements (voxels), rather than adding ones. Our approach selects the “meaningful” voxels, in a manner that addresses both types of dataset imbalance. The approach is general and can be applied to any voxel-based detector, yet the meaningfulness of a voxel is network-dependent. Our voxel selection is shown to improve the performance of several prominent 3D detection methods.
Oren Shrout, Yizhak Ben-Shabat, Ayellet Tal
CVPR2
2023 Aligning Step-by-Step Instructional Diagrams to Video Demonstrations
abstract
Multimodal alignment facilitates the retrieval of instances from one modality when queried using another. In this paper, we consider a novel setting where such an alignment is between (i) instruction steps that are depicted as assembly diagrams (commonly seen in Ikea assembly manuals) and (ii) segments from in-the-wild videos; these videos comprising an enactment of the assembly actions in the real world. We introduce a supervised contrastive learning approach that learns to align videos with the subtle details of assembly diagrams, guided by a set of novel losses. To study this problem and evaluate the effectiveness of our method, we introduce a new dataset: IAW-for Ikea assembly in the wild-consisting of 183 hours of videos from diverse furniture assembly collections and nearly 8,300 illustrations from their associated instruction manuals and annotated for their ground truth alignments. We define two tasks on this dataset: First, nearest neighbor retrieval between video segments and illustrations, and, second, alignment of instruction steps and the segments for each video. Extensive experiments on IAW demonstrate superior performance of our approach against alternatives.
Anoop Cherian, Yizhak Ben-Shabat, Cristian Rodriguez Opazo, Stephen Gould
CVPR4
2022 DiGS : Divergence guided shape implicit neural representation for unoriented point clouds
abstract
Shape implicit neural representations (INRs) have recently shown to be effective in shape analysis and reconstruction tasks. Existing INRs require point coordinates to learn the implicit level sets of the shape. When a normal vector is available for each point, a higher fidelity representation can be learned, however normal vectors are often not provided as raw data. Furthermore, the method's initialization has been shown to play a crucial role for surface reconstruction. In this paper, we propose a divergence guided shape representation learning approach that does not require normal vectors as input. We show that incorporating a soft constraint on the divergence of the distance function favours smooth solutions that reliably orients gradients to match the unknown normal at each point, in some cases even better than approaches that use ground truth normal vectors directly. Additionally, we introduce a novel geometric initialization method for sinusoidal INRs that further improves convergence to the desired solution. We evaluate the effectiveness of our approach on the task of surface reconstruction and shape space learning and show SOTA performance compared to other unoriented methods. Code and model parameters available at our project page https://chumbyte.github.io/DiGS-Site/
Yizhak Ben-Shabat, Chamin Hewa Koneputugodage, Stephen Gould
CVPR1
2022 GoferBot: A Visual Guided Human-Robot Collaborative Assembly System
abstract
The current transformation towards smart manufacturing has led to a growing demand for human-robot collaboration (HRC) in the manufacturing process. Perceiving and understanding the human co-worker's behaviour introduces challenges for collaborative robots to efficiently and effectively perform tasks in unstructured and dynamic environments. Integrating recent data-driven machine vision capabilities into HRC systems is a logical next step in addressing these challenges. However, in these cases, off-the-shelf components struggle due to generalisation limitations. Real-world evaluation is required in order to fully appreciate the maturity and robustness of these approaches. Furthermore, understanding the pure-vision aspects is a crucial first step before combining multiple modalities in order to understand the limitations. In this paper, we propose GoferBot, a novel vision-based semantic HRC system for a real-world assembly task. It is composed of a visual servoing module that reaches and grasps assembly parts in an unstructured multi-instance and dynamic environment, an action recognition module that performs human action prediction for implicit communication, and a visual handover module that uses the perceptual understanding of human behaviour to produce an intuitive and efficient collaborative assembly experience. GoferBot is a novel assembly system that seamlessly integrates all sub-modules by utilising implicit semantic information purely from visual perception.
Zheyu Zhuang, Yizhak Ben-Shabat, Stephen Gould, Robert E. Mahony
IROS2
2022 CloudWalker: Random walks for 3D point cloud shape analysis
Adi Mesika, Yizhak Ben-Shabat, Ayellet Tal
Comput. Graph.2
2021 The IKEA ASM Dataset: Understanding People Assembling Furniture through Actions, Objects and Pose
abstract
The availability of a large labeled dataset is a key requirement for applying deep learning methods to solve various computer vision tasks. In the context of understanding human activities, existing public datasets, while large in size, are often limited to a single RGB camera and provide only per-frame or per-clip action annotations. To enable richer analysis and understanding of human activities, we introduce IKEA ASM-a three million frame, multi-view, furniture assembly video dataset that includes depth, atomic actions, object segmentation, and human poses. Additionally, we benchmark prominent methods for video action recognition, object segmentation and human pose estimation tasks on this challenging dataset. The dataset enables the development of holistic methods, which integrate multi-modal and multi-view data to better perform on these tasks.
Yizhak Ben-Shabat, Xin Yu 0002, Fatemehsadat Saleh, Dylan Campbell, Cristian Rodriguez Opazo, Hongdong Li, Stephen Gould
WACV1
2020 DeepFit: 3D Surface Fitting via Neural Network Weighted Least Squares
Yizhak Ben-Shabat, Stephen Gould
ECCV (1)1
2020 DPDist: Comparing Point Clouds Using Deep Point Cloud Distance
Dahlia Urbach, Yizhak Ben-Shabat, Michael Lindenbaum
ECCV (11)2
2019 Nesti-Net: Normal Estimation for Unstructured 3D Point Clouds Using Convolutional Neural Networks
abstract
In this paper, we propose a normal estimation method for unstructured 3D point clouds. This method, called Nesti-Net, builds on a new local point cloud representation which consists of multi-scale point statistics (MuPS), estimated on a local coarse Gaussian grid. This representation is a suitable input to a CNN architecture. The normals are estimated using a mixture-of-experts (MoE) architecture, which relies on a data-driven approach for selecting the optimal scale around each point and encourages sub-network specialization. Interesting insights into the network's resource distribution are provided. The scale prediction significantly improves robustness to different noise levels, point density variations and different levels of detail. We achieve state-of-the-art results on a benchmark synthetic dataset and present qualitative results on real scanned scenes.
Yizhak Ben-Shabat, Michael Lindenbaum, Anath Fischer
CVPR1
2018 Graph based over-segmentation methods for 3D point clouds
Yizhak Ben-Shabat, Tamar Avraham, Michael Lindenbaum, Anath Fischer
Comput. Vis. Image Underst.1