Mark Sheinin

dblp:173/5286 · DBLP profile ↗
← Back
14ranked-venue papers
10as first author
9since 2021 · last 2025
0000-0002-5462-5730ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 11 · 7 first-author · 8 since 2021
YearPublicationVenuePosition
2025 Learning to See Inside Opaque Liquid Containers Using Speckle Vibrometry
abstract
Computer vision seeks to infer a wide range of information about objects and events. However, vision systems based on conventional imaging are limited to extracting information only from the visible surfaces of scene objects. For instance, a vision system can detect and identify a Coke can in the scene, but it cannot determine whether the can is full or empty. In this paper, we aim to expand the scope of computer vision to include the novel task of inferring the hidden liquid levels of opaque containers by sensing the tiny vibrations on their surfaces. Our method provides a first-of-a-kind way to inspect the fill level of multiple sealed containers remotely, at once, without needing physical manipulation and manual weighing. First, we propose a novel speckle-based vibration sensing system for simultaneously capturing scene vibrations on a 2D grid of points. We use our system to efficiently and remotely capture a dataset of vibration responses for a variety of everyday liquid containers. Then, we develop a transformer-based approach for analyzing the captured vibrations and classifying the container type and its hidden liquid level at the time of measurement. Our architecture is invariant to the vibration source, yielding correct liquid level estimates for controlled and ambient scene sound sources. Moreover, our model generalizes to unseen container instances within known classes (e.g., training on five Coke cans of a six-pack, testing on a sixth) and fluid levels. We demonstrate our method by recovering liquid levels from various everyday containers.
Matan Kichler, Shai Bagon, Mark Sheinin
ICCV3
2025 Dual-Shutter Optical Vibration Sensing
abstract
Visual vibrometry is a highly useful tool for remote capture of audio, as well as the physical properties of materials, human heart rate, and more. While visually-observable vibrations can be captured directly with a high-speed camera, minute imperceptible object vibrations can be optically amplified by imaging the displacement of a speckle pattern created by shining a laser beam on the vibrating surface. In this paper, we propose a novel method for sensing vibrations at high speeds (up to 63 kHz), for multiple scene sources at once, using sensors rated for only 130 Hz operation. Our method relies on simultaneously capturing the scene with two cameras equipped with rolling and global shutter sensors, respectively. The rolling shutter camera captures distorted speckle images that encode the high-speed object vibrations. The global shutter camera captures undistorted reference images of the speckle pattern, helping to decode the source vibrations. We demonstrate our method by capturing vibration caused by audio sources (e.g., speakers, human voice, and musical instruments) and analyzing the vibration modes of a tuning fork.
Mark Sheinin, Dorian Chan, Matthew O'Toole, Srinivasa G. Narasimhan
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Projecting Trackable Thermal Patterns for Dynamic Computer Vision
abstract
Adding artificial patterns to objects, like QR codes, can ease tasks such as object tracking, robot navigation, and conveying information (e.g., a label or a website link). However, these patterns require a physical application and they alter the object's appearance. Conversely, projected patterns can temporarily change the object's appearance, aiding tasks like 3D scanning and retrieving object textures and shading. However, projected patterns impede dynamic tasks like object tracking because they do not ‘stick’ to the object's surface. Or do they? This paper introduces a novel approach combining the advantages of projected and persistent physical patterns. Our system projects heat patterns using a laser beam (similar in spirit to a LIDAR), which a thermal camera observes and tracks. Such thermal patterns enable tracking poorly-textured objects whose tracking is highly challenging with standard cameras while not affecting the object's appearance or physical properties. To avail these thermal patterns in existing vision frameworks, we train a network to reverse heat diffusion's effects and remove inconsistent pattern points between different thermal frames. We prototyped and tested this approach on dynamic vision tasks like structure from motion, optical flow, and object tracking of everyday textureless objects.
Mark Sheinin, Aswin C. Sankaranarayanan, Srinivasa G. Narasimhan
CVPR1
2024 Shape from Heat Conduction
Sriram Narayanan, Manikandasriram Srinivasan Ramanagopal, Mark Sheinin, Aswin C. Sankaranarayanan, Srinivasa G. Narasimhan
ECCV (38)3
2023 Analyzing Physical Impacts Using Transient Surface Wave Imaging
abstract
The subtle vibrations on an object's surface contain information about the object's physical properties and its interaction with the environment. Prior works imaged surface vibration to recover the object's material properties via modal analysis, which discards the transient vibrations propagating immediately after the object is disturbed. Conversely, prior works that captured transient vibrations focused on recovering localized signals (e.g., recording nearby sound sources), neglecting the spatiotemporal relationship between vibrations at different object points. In this paper, we extract information from the transient surface vibrations simultaneously measured at a sparse set of object points using the dual-shutter camera described by Sheinin et al. [37]. We model the geometry of an elastic wave generated at the moment an object's surface is disturbed (e.g., a knock or a footstep) and use the model to localize the disturbance source for various materials (e.g., wood, plastic, tile). We also show that transient object vibrations contain additional cues about the impact force and the impacting object's material properties. We demonstrate our approach in applications like localizing the strikes of a ping-pong ball on a table mid-play and recovering the footsteps' locations by imaging the floor vibrations they create.
Mark Sheinin, Dorian Chan, Mark Rau, Matthew O'Toole, Srinivasa G. Narasimhan
CVPR2
2023 SpinCam: High-Speed Imaging via a Rotating Point-Spread Function
abstract
High-speed cameras are an indispensable tool used for the slow-motion analysis of scenes. However, the fixed bandwidth of any imaging system quickly becomes a bottleneck, resulting in a fundamental trade-off between the camera’s spatial and temporal resolutions. In recent years, compressive high-speed imaging systems have been proposed to circumvent these issues by optically encoding the signal and using a reconstruction procedure to recover a video. Our work proposes a novel approach for compressive high-speed imaging based on temporally coding the camera’s point-spread function (PSF). By mechanically spinning a diffraction grating in front of a camera, the sensor integrates an image blurred by a PSF that continuously rotates over time. We also propose a deconvolution-based reconstruction algorithm to reconstruct videos from these measurements. Our method achieves superior light efficiency and handles a wider scene class than prior methods. Also, our mechanical design yields flexible temporal resolution that can be easily increased, potentially allowing capture at 192 kHz—far higher than prior works. We demonstrate a prototype for various applications, including motion capture and particle image velocimetry (PIV).
Dorian Chan, Mark Sheinin, Matthew O'Toole
ICCV2
2022 Dual-Shutter Optical Vibration Sensing
abstract
Visual vibrometry is a highly useful tool for remote capture of audio, as well as the physical properties of materials, human heart rate, and more. While visually-observable vibrations can be captured directly with a high-speed camera, minute imperceptible object vibrations can be optically amplified by imaging the displacement of a speckle pattern, created by shining a laser beam on the vibrating surface. In this paper, we propose a novel method for sensing vibrations at high speeds (up to 63kHz), for multiple scene sources at once, using sensors rated for only 130Hz operation. Our method relies on simultaneously capturing the scene with two cameras equipped with rolling and global shutter sensors, respectively. The rolling shutter camera captures distorted speckle images that encode the high-speed object vibrations. The global shutter camera captures undistorted reference images of the speckle pattern, helping to decode the source vibrations. We demonstrate our method by capturing vibration caused by audio sources (e.g. speakers, human voice, and musical instruments) and analyzing the vibration modes of a tuning fork.
Mark Sheinin, Dorian Chan, Matthew O'Toole, Srinivasa G. Narasimhan
CVPR1
2022 Computational Imaging on the Electric Grid
abstract
Night beats with alternating current (AC) illumination. By passively sensing this beat, we reveal new scene information which includes: the type of bulbs in the scene, the phases of the electric grid up to city scale, and the light transport matrix. This information yields unmixing of reflections and semi-reflections, nocturnal high dynamic range, and scene rendering with bulbs not observed during acquisition. The latter is facilitated by a dataset of bulb response functions for a range of sources, which we collected and provide. To do all this, we built a novel coded-exposure high-dynamic-range imaging technique, specifically designed to operate on the grid's AC lighting.
Mark Sheinin, Yoav Y. Schechner, Kiriakos N. Kutulakos
IEEE Trans. Pattern Anal. Mach. Intell.1
2021 Deconvolving Diffraction for Fast Imaging of Sparse Scenes
abstract
Most computer vision techniques rely on cameras which uniformly sample the 2D image plane. However, there exists a class of applications for which the standard uniform 2D sampling of the image plane is sub-optimal. This class consists of applications where the scene points of interest occupy the image plane sparsely (e.g., marker-based motion capture), and thus most pixels of the 2D camera sensor would be wasted. Recently, diffractive optics were used in conjunction with sparse (e.g., line) sensors to achieve high-speed capture of such sparse scenes. One such approach, called “Diffraction Line Imaging”, relies on the use of diffraction gratings to spread the point-spread-function (PSF) of scene points from a point to a color-coded shape (e.g., a horizontal line) whose intersection with a line sensor enables point positioning. In this paper, we extend this approach for arbitrary diffractive optical elements and arbitrary sampling of the sensor plane using a convolution-based image formation model. Sparse scenes are then recovered by formulating a convolutional coding inverse problem that can resolve mixtures of diffraction PSFs without the use of multiple sensors, extending the application of diffraction-based imaging to a new class of significantly denser scenes. For the case of a single-axis diffraction grating, we provide an approach to determine the minimal required sensor sub-sampling for accurate scene recovery. Compared to methods that use a speckle PSF from a narrow-band source or a diffuser-based PSF with a rolling shutter sensor, our approach uses spectrally-coded PSFs from broad-band sources and allows arbitrary sensor sampling, respectively. We demonstrate that the presented combination of the imaging approach and scene recovery method is well suited for high-speed marker based motion capture and particle image velocimetry (PIV) over long periods.
Mark Sheinin, Matthew O'Toole, Srinivasa G. Narasimhan
ICCP1
2020 Diffraction Line Imaging
Mark Sheinin, N. Dinesh Reddy, Matthew O'Toole, Srinivasa G. Narasimhan
ECCV (2)1
2019 Depth From Texture Integration
abstract
We present a new approach for active ranging, which can be compounded with traditional methods such as active depth from defocus or off-axis structured illumination. The object is illuminated by an active textured pattern having high spatial-frequency content. The illumination texture varies in time while the object undergoes a focal sweep. Consequently, in a single exposure, the illumination textures are encoded as a function of the object depth. Per-object depth, a particular illumination texture, with its high spatial frequency content, is focused; the other textures, projected when the system is defocused, are blurred. Analysis of the time-integrated image decodes the depth map. The plurality of projected and sensed color channels enhances the performance of the process, as we demonstrate experimentally. Using a wide aperture and only one or two readout frames, the method is particularly useful for imaging that requires high sensitivity to weak signals and high spatial resolution. Using a focal sweep during an exposure, the imaging has a wide dynamic depth range while being fast.
Mark Sheinin, Yoav Y. Schechner
ICCP1
2018 Rolling shutter imaging on the electric grid
abstract
Flicker of AC-powered lights is useful for probing the electric grid and unmixing reflected contributions of different sources. Flicker has been sensed in great detail with a specially-designed camera tethered to an AC outlet. We argue that even an untethered smartphone can achieve the same task. We exploit the inter-row exposure delay of the ubiquitous rolling-shutter sensor. When pixel exposure time is kept short, this delay creates a spatiotemporal wave pattern that encodes (1) the precise capture time relative to the AC, (2) the response function of individual bulbs, and (3) the AC phase that powers them. To sense point sources, we induce the spatiotemporal wave pattern by placing a star filter or a paper diffuser in front of the camera's lens. We demonstrate several new capabilities, including: high-rate acquisition of bulb response functions from one smartphone photo; recognition of bulb type and phase from one or two images; and rendering of live flicker video, as if it came from a high speed global-shutter camera.
Mark Sheinin, Yoav Y. Schechner, Kiriakos N. Kutulakos
ICCP1
2017 Computational Imaging on the Electric Grid
abstract
Night beats with alternating current (AC) illumination. By passively sensing this beat, we reveal new scene information which includes: the type of bulbs in the scene, the phases of the electric grid up to city scale, and the light transport matrix. This information yields unmixing of reflections and semi-reflections, nocturnal high dynamic range, and scene rendering with bulbs not observed during acquisition. The latter is facilitated by a database of bulb response functions for a range of sources, which we collected and provide. To do all this, we built a novel coded-exposure high-dynamic-range imaging technique, specifically designed to operate on the grids AC lighting.
Mark Sheinin, Yoav Y. Schechner, Kiriakos N. Kutulakos
CVPR1
2016 The Next Best Underwater View
abstract
To image in high resolution large and occlusion-prone scenes, a camera must move above and around. Degradation of visibility due to geometric occlusions and distances is exacerbated by scattering, when the scene is in a participating medium. Moreover, underwater and in other media, artificial lighting is needed. Overall, data quality depends on the observed surface, medium and the timevarying poses of the camera and light source (C&L). This work proposes to optimize C&L poses as they move, so that the surface is scanned efficiently and the descattered recovery has the highest quality. The work generalizes the next best view concept of robot vision to scattering media and cooperative movable lighting. It also extends descattering to platforms that move optimally. The optimization criterion is information gain, taken from information theory. We exploit the existence of a prior rough 3D model, since underwater such a model is routinely obtained using sonar. We demonstrate this principle in a scaled-down setup.
Mark Sheinin, Yoav Y. Schechner
CVPR1