Edward H. Adelson

dblp:73/143 · DBLP profile ↗
← Back
76ranked-venue papers
3as first author
12since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 54 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 2 first-authorSystems, architecture and hardware · 25 · 11 since 2021Human-computer interaction and ubiquitous computing · 4Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1Theory of computation · 1
YearPublicationVenuePosition
2025 PolyTouch: A Robust Multi-Modal Tactile Sensor for Contact-Rich Manipulation Using Tactile-Diffusion Policies
abstract
Achieving robust dexterous manipulation in un-structured domestic environments remains a significant challenge in robotics. Even with state-of-the-art robot learning methods, haptic-oblivious control strategies (i.e. those relying only on external vision and/or proprioception) often fall short due to occlusions, visual complexities, and the need for precise contact interaction control. To address these limitations, we introduce PolyTouch, a novel robot finger that integrates camera-based tactile sensing, acoustic sensing, and peripheral visual sensing into a single design that is compact and durable. PolyTouch provides high-resolution tactile feedback across multiple temporal scales, which is essential for efficiently learning complex manipulation tasks. Experiments demonstrate an at least 20-fold increase in lifespan over commercial tactile sensors, with a design that is both easy to manufacture and scalable. We then use this multimodal tactile feedback along with visuo-proprioceptive observations to synthesize a tactile-diffusion policy from human demonstrations; the resulting contact-aware control policy significantly outperforms haptic-oblivious policies in multiple contact-aware manipulation policies. This paper highlights how effectively integrating multimodal contact sensing can hasten the development of effective contact-aware manipulation policies, paving the way for more reliable and versatile domestic robots. More information can be found at https://polytouch.alanz.info/.
Jialiang Zhao, Naveen Kuppuswamy, Siyuan Feng 0003, Benjamin Burchfiel, Edward H. Adelson
ICRA5
2025 Grasp EveryThing (GET): 1-DoF, 3-Fingered Gripper with Tactile Sensing for Robust Grasping
abstract
We introduce the Grasp EveryThing (GET) gripper, a novel 1-DoF, 3-finger design for securely grasping objects of many shapes and sizes. Mounted on a standard parallel jaw actuator, the design features three narrow, tapered fingers arranged in a two-against-one configuration, where the two fingers converge into a V-shape. The GET gripper is more capable of conforming to object geometries and forming secure grasps than traditional designs with two flat fingers. Inspired by the principle of self-similarity, these V-shaped fingers enable secure grasping across a wide range of object sizes. Further to this end, fingers are parametrically designed for convenient resizing and interchangeability across robotic embodiments with a parallel jaw gripper. Additionally, we incorporate a rigid fingernail for ease in manipulating small objects. Tactile sensing can be integrated into the standalone finger via an externally-mounted camera. A neural network was trained to estimate normal force from tactile images with an average validation error of 1.3 N across a diverse set of geometries. In grasping 15 objects and performing 3 tasks via teleoperation, the GET fingers consistently outperformed standard flat fingers. All finger designs, compatible with multiple robotic embodiments, both incorporating and lacking tactile sensing, are available on GitHub.
Michael Burgess, Edward H. Adelson
IROS2
2025 Tactile-Reactive Roller Grasper
abstract
Manipulation of objects within a robot's hand is one of the most important challenges in achieving robot dexterity. To address this challenge, Roller Graspers use steerable rolling fingertips. The fingertips impart motions and exert forces to achieve six degree of freedom mobility and closed-loop grasp force control. The design reported here uses image processing from cameras placed inside steerable compliant rollers to track contact conditions and locations. Integration of this data into a controller enables a variety of robust in-hand manipulation capabilities. We demonstrate that the same information can be used to reconstruct object shape. In addition, we show that by converting in-hand manipulation from a discontinuous process, with fingers frequently attaching and detaching from the object surface, to a continuous process, we can implement a convergent control loop that minimizes errors that otherwise accumulate during large object motions. The difference is apparent when comparing the results of an object rotation using a discontinuous finger-gaiting approach, as would be required without rolling fingertips, to the results obtained with continuous rolling. The results suggest that hybrid rolling fingertip and finger-gaiting approaches to manipulation may be a promising future research direction.
Shenli Yuan, Shaoxiong Wang, Radhen Patel, Megha Tippur, Connor L. Yako, Mark R. Cutkosky, Edward H. Adelson, John Kenneth Salisbury Jr.
IEEE Trans. Robotics7
2024 A Passively Bendable, Compliant Tactile Palm with RObotic Modular Endoskeleton Optical (ROMEO) Fingers
abstract
Many robotic hands currently rely on extremely dexterous robotic fingers and a thumb joint to envelop themselves around an object. Few hands focus on the palm even though human hands greatly benefit from their central fold and soft surface. As such, we develop a novel structurally compliant soft palm, which enables more surface area contact for the objects that are pressed into it. Moreover, this design, along with the development of a new low-cost, flexible illumination system, is able to incorporate a high-resolution tactile sensing system inspired by the GelSight sensors. Concurrently, we design RObotic Modular Endoskeleton Optical (ROMEO) fingers, which are underactuated two-segment soft fingers that are able to house the new illumination system, and we integrate them into these various palm configurations. The resulting robotic hand is slightly bigger than a baseball and represents one of the first soft robotic hands with actuated fingers and a passively compliant palm, all of which have high-resolution tactile sensing. This design also potentially helps researchers discover and explore more soft-rigid tactile robotic hand designs with greater capabilities in the future.
Sandra Q. Liu, Edward H. Adelson
ICRA2
2024 GelLink: A Compact Multi-phalanx Finger with Vision-based Tactile Sensing and Proprioception
abstract
Compared to fully-actuated robotic end-effectors, underactuated ones are generally more adaptive, robust, and cost-effective. However, state estimation for underactuated hands is usually more challenging. Vision-based tactile sensors, like Gelsight, can mitigate this issue by providing high-resolution tactile sensing and accurate proprioceptive sensing. As such, we present GelLink, a compact, underactuated, linkage-driven robotic finger with low-cost, high-resolution vision-based tactile sensing and proprioceptive sensing capabilities. In order to reduce the amount of embedded hardware, i.e. the cameras and motors, we optimize the linkage transmission with a planar linkage mechanism simulator and develop a planar reflection simulator to simplify the tactile sensing hardware. As a result, GelLink only requires one motor to actuate the three phalanges, and one camera to capture tactile signals along the entire finger. Overall, GelLink is a compact robotic finger that shows adaptability and robustness when performing grasping tasks. The integration of vision- based tactile sensors can significantly enhance the capabilities of underactuated fingers and potentially broaden their future usage.
Jialiang Zhao, Edward H. Adelson
ICRA3
2024 RainbowSight: A Family of Generalizable, Curved, Camera-Based Tactile Sensors For Shape Reconstruction
abstract
Camera-based tactile sensors can provide high resolution positional and local geometry information for robotic manipulation. Curved and rounded fingers are often advantageous, but it can be difficult to derive illumination systems that work well within curved geometries. To address this issue, we introduce RainbowSight, a family of curved, compact, camera-based tactile sensors which use addressable RGB LEDs illuminated in a novel rainbow spectrum pattern. In addition to being able to scale the illumination scheme to different sensor sizes and shapes to fit on a variety of end effector configurations, the sensors can be easily manufactured and require minimal optical tuning to obtain high resolution depth reconstructions of an object deforming the sensor’s soft elastomer surface. Additionally, we show the advantages of our new hardware design and improvements in calibration methods for accurate depth map generation when compared to alternative lighting methods commonly implemented in previous camera-based tactile sensors. With these advancements, we make the integration of tactile sensors more accessible to roboticists by allowing them the flexibility to easily customize, fabricate, and calibrate camera-based tactile sensors to best fit the needs of their robotic systems.
Megha Tippur, Edward H. Adelson
ICRA2
2024 Learning incipient slip with GelSight sensors: Attention Classification with Video Vision Transformers
abstract
An important aspect of robotic grasping is the ability to detect incipient slip based on real-time information through tactile sensors. In this paper, we propose to use Video Vision Transformers to detect the onset of slip in grasping scenarios. The dynamic nature of slip makes Video Vision Transformers well-suited for capturing temporal correlations with relatively small datasets. The training data is acquired through two GelSight tactile sensors attached to the generic finger grippers of a Panda Franka Emika robot arm that grasps, lifts and shakes 30 everyday objects in order to induce slip. We further conducted an ablation study by considering 5, 4, 3, and 2 frames prior to slip onset, revealing consistent prediction accuracy. Our approach demonstrates the capability to predict slips well in advance, even up to the 5thframe before the onset. This underscores the predictive capability of our approach, indicating its effectiveness in slip detection well before of its occurrence. This advance prediction capability may be a valuable tool for undertaking preemptive corrective actions, such as implementing a more secure gripper closure. We evaluate the efficiency of our approach to predict onset of slip on 10 previously-unseen objects and achieve a zero-shot mean prediction accuracy of 99%.
Amit Parag, Edward H. Adelson, Ekrem Misimi
IROS2
2024 EyeSight Hand: Design of a Fully-Actuated Dexterous Robot Hand with Integrated Vision-Based Tactile Sensors and Compliant Actuation
abstract
In this work, we introduce the EyeSight Hand, a 7 degrees of freedom (DoF) humanoid hand featuring integrated vision-based tactile sensors tailored for enhanced whole-hand manipulation. Additionally, we introduce an actuation scheme centered around quasi-direct drive actuation to achieve human-like strength and speed while ensuring robustness for large-scale data collection. We evaluate the EyeSight Hand on three challenging tasks: bottle opening, plasticine cutting, and plate pick and place, which require a blend of complex manipulation, tool use, and precise force application. Imitation learning models trained on these tasks, with a vision dropout strategy, showcase the benefits of tactile feedback in enhancing task success rates. Our results reveal that the integration of tactile sensing dramatically improves task performance, underscoring the critical role of tactile information in dexterous manipulation.
Branden Romero, Haoshu Fang, Pulkit Agrawal 0001, Edward H. Adelson
IROS4
2023 TactoFind: A Tactile Only System for Object Retrieval
abstract
We study the problem of object retrieval in scenarios where visual sensing is absent, object shapes are unknown beforehand and objects can move freely, like grabbing objects out of a drawer. Successful solutions require localizing free objects, identifying specific object instances, and then grasping the identified objects, only using touch feedback. Unlike vision, where cameras can observe the entire scene, touch sensors are local and only observe parts of the scene that are in contact with the manipulator. Moreover, information gathering via touch sensors necessitates applying forces on the touched surface which may disturb the scene itself. Reasoning with touch, therefore, requires careful exploration and integration of information over time - a challenge we tackle. We present a system capable of using sparse tactile feedback from fingertip touch sensors on a dexterous hand to localize, identify and grasp novel objects without any visual feedback. Videos are available at https://sites.google.com/view/tactofind.
Sameer Pai, Tao Chen 0046, Megha Tippur, Edward H. Adelson, Abhishek Gupta 0004, Pulkit Agrawal 0001
ICRA4
2023 FingerSLAM: Closed-loop Unknown Object Localization and Reconstruction from Visuo-tactile Feedback
abstract
In this paper, we address the problem of using visuo-tactile feedback for 6-DoF localization and 3D reconstruction of unknown in-hand objects. We propose FingerSLAM, a closed-loop factor graph-based pose estimator that combines local tactile sensing at finger-tip and global vision sensing from a wrist-mount camera. FingerSLAM is constructed with two constituent pose estimators: a multi-pass refined tactile-based pose estimator that captures movements from detailed local textures, and a single-pass vision-based pose estimator that predicts from a global view of the object. We also design a loop closure mechanism that actively matches current vision and tactile images to previously stored key-frames to reduce accumulated error. FingerSLAM incorporates the two sensing modalities of tactile and vision, as well as the loop closure mechanism with a factor graph-based optimization framework. Such a framework produces an optimized pose estimation solution that is more accurate than the standalone estimators. The estimated poses are then used to reconstruct the shape of the unknown object incrementally by stitching the local point clouds recovered from tactile images. We train our system on real-world data collected with 20 objects. We demonstrate reliable visuo-tactile pose estimation and shape reconstruction through quantitative and qualitative real-world evaluations on 6 objects that are unseen during training.
Jialiang Zhao, Maria Bauzá 0001, Edward H. Adelson
ICRA3
2023 GelSight Svelte: A Human Finger-Shaped Single-Camera Tactile Robot Finger with Large Sensing Coverage and Proprioceptive Sensing
abstract
Camera-based tactile sensing is a low-cost, popular approach to obtain highly detailed contact geometry information. However, most existing camera-based tactile sensors are fingertip sensors, and longer fingers often require extraneous elements to obtain an extended sensing area similar to the full length of a human finger. Moreover, existing methods to estimate proprioceptive information such as total forces and torques applied on the finger from camera-based tactile sensors are not effective when the contact geometry is complex. We introduce GelSight Svelte, a curved, human finger-sized, single-camera tactile sensor that is capable of both tactile and proprioceptive sensing over a large area. GelSight Svelte uses curved mirrors to achieve the desired shape and sensing coverage. Proprioceptive information, such as the total bending and twisting torques applied on the finger, is reflected as deformations on the flexible backbone of GelSight Svelte, which are also captured by the camera. We train a convolutional neural network to estimate the bending and twisting torques from the captured images. We conduct gel deformation experiments at various locations of the finger to evaluate the tactile sensing capability and proprioceptive sensing accuracy. To demonstrate the capability and potential uses of GelSight Svelte, we conduct an object holding task with three different grasping modes that utilize different areas of the finger.
Jialiang Zhao, Edward H. Adelson
IROS2
2021 GelSight Wedge: Measuring High-Resolution 3D Contact Geometry with a Compact Robot Finger
abstract
Vision-based tactile sensors have the potential to provide important contact geometry to localize the objective with visual occlusion. However, it is challenging to measure high-resolution 3D contact geometry for a compact robot finger, to simultaneously meet optical and mechanical constraints. In this work, we present the GelSight Wedge sensor, which is optimized to have a compact shape for robot fingers, while achieving high-resolution 3D reconstruction. We evaluate the 3D reconstruction under different lighting configurations, and extend the method from 3 lights to 1 or 2 lights. We demonstrate the flexibility of the design by shrinking the sensor to the size of a human finger for fine manipulation tasks. We also show the effectiveness and potential of the reconstructed 3D geometry for pose tracking in the 3D space.
Shaoxiong Wang, Yu She, Branden Romero, Edward H. Adelson
ICRA4
2020 Soft, Round, High Resolution Tactile Fingertip Sensors for Dexterous Robotic Manipulation
abstract
High resolution tactile sensors are often bulky and have shape profiles that make them awkward for use in manipulation. This becomes important when using such sensors as fingertips for dexterous multi-fingered hands, where boxy or planar fingertips limit the available set of smooth manipulation strategies. High resolution optical based sensors such as GelSight have until now been constrained to relatively flat geometries due to constraints on illumination geometry. Here, we show how to construct a rounded fingertip that utilizes a form of light piping for directional illumination. Our sensors can replace the standard rounded fingertips of the Allegro hand. They can capture high resolution maps of the contact surfaces, and can be used to support various dexterous manipulation tasks.
Branden Romero, Filipe Veiga, Edward H. Adelson
ICRA3
2020 Exoskeleton-covered soft finger with vision-based proprioception and tactile sensing
abstract
Soft robots offer significant advantages in adaptability, safety, and dexterity compared to conventional rigid-body robots. However, it is challenging to equip soft robots with accurate proprioception and tactile sensing due to their high flexibility and elasticity. In this work, we describe the development of a vision-based proprioceptive and tactile sensor for soft robots called GelFlex, which is inspired by previous GelSight sensing techniques. More specifically, we develop a novel exoskeleton-covered soft finger with embedded cameras and deep learning methods that enable high-resolution proprioceptive sensing and rich tactile sensing. To do so, we design features along the axial direction of the finger, which enable high-resolution proprioceptive sensing, and incorporate a reflective ink coating on the surface of the finger to enable rich tactile sensing. We design a highly underactuated exoskeleton with a tendon-driven mechanism to actuate the finger. Finally, we assemble 2 of the fingers together to form a robotic gripper and successfully perform a bar stock classification task, which requires both shape and tactile information. We train neural networks for proprioception and shape (box versus cylinder) classification using data from the embedded sensors. The proprioception CNN had over 99% accuracy on our testing set (all six joint angles were within 1° of error) and had an average accumulative distance error of 0.77 mm during live testing, which is better than human finger proprioception. These proposed techniques offer soft robots the high-level ability to simultaneously perceive their proprioceptive state and peripheral environment, providing potential solutions for soft robots to solve everyday manipulation tasks. We believe the methods developed in this work can be widely applied to different designs and applications.
Yu She, Sandra Q. Liu, Peiyu Yu, Edward H. Adelson
ICRA4
2020 SwingBot: Learning Physical Features from In-hand Tactile Exploration for Dynamic Swing-up Manipulation
abstract
Several robot manipulation tasks are extremely sensitive to variations of the physical properties of the manipulated objects. One such task is manipulating objects by using gravity or arm accelerations, increasing the importance of mass, center of mass, and friction information. We present SwingBot, a robot that is able to learn the physical features of an held object through tactile exploration. Two exploration actions (tilting and shaking) provide the tactile information used to create a physical feature embedding space. With this embedding, SwingBot is able to predict the swing angle achieved by a robot performing dynamic swing-up manipulations on a previously unseen object. Using these predictions, it is able to search for the optimal control parameters for a desired swing-up angle. We show that with the learned physical features our end-to-end self-supervised learning pipeline is able to substantially improve the accuracy of swinging up unseen objects. We also show that objects with similar dynamics are closer to each other on the embedding space and that the embedding can be disentangled into values of specific physical properties.
Shaoxiong Wang, Branden Romero, Filipe Veiga, Edward H. Adelson
IROS5
2018 Slip Detection with Combined Tactile and Visual Information
abstract
Slip detection plays a vital role in robotic manipulation and it has long been a challenging problem in the robotic community. In this paper, we propose a new method based on deep neural network (DNN) to detect slip. The training data is acquired by a GelSight tactile sensor and a camera mounted on a gripper when we use a robot arm to grasp and lift 94 daily objects with different grasping forces and grasping positions. The DNN is trained to classify whether a slip occurred or not. To evaluate the performance of the DNN, we test 10 unseen objects in 152 grasps. A detection accuracy as high as 88.03 % is achieved. It is anticipated that the accuracy can be further improved with a larger dataset. This method is beneficial for robots to make stable grasps, which can be widely applied to automatic force control, grasping strategy selection and fine manipulation.
Siyuan Dong, Edward H. Adelson
ICRA3
2018 ViTac: Feature Sharing Between Vision and Tactile Sensing for Cloth Texture Recognition
abstract
Vision and touch are two of the important sensing modalities for humans and they offer complementary information for sensing the environment. Robots could also benefit from such multi-modal sensing ability. In this paper, addressing for the first time (to the best of our knowledge) texture recognition from tactile images and vision, we propose a new fusion method named Deep Maximum Covariance Analysis (DMCA) to learn a joint latent space for sharing features through vision and tactile sensing. The features of camera images and tactile data acquired from a GelSight sensor are learned by deep neural networks. But the learned features are of a high dimensionality and are redundant due to the differences between the two sensing modalities, which deteriorates the perception performance. To address this, the learned features are paired using maximum covariance analysis. Results of the algorithm on a newly collected dataset of paired visual and tactile data relating to cloth textures show that a good recognition performance of greater than 90% can be achieved by using the proposed DMCA framework. In addition, we find that the perception performance of either vision or tactile sensing can be improved by employing the shared representation space, compared to learning from unimodal data.
Shan Luo 0001, Wenzhen Yuan 0001, Edward H. Adelson, Anthony G. Cohn 0001, Raul A. Fuentes-Samaniego
ICRA3
2018 Active Clothing Material Perception Using Tactile Sensing and Deep Learning
abstract
Humans represent and discriminate the objects in the same category using their properties, and an intelligent robot should be able to do the same. In this paper, we build a robot system that can autonomously perceive the object properties through touch. We work on the common object category of clothing. The robot moves under the guidance of an external Kinect sensor, and squeezes the clothes with a GelSight tactile sensor, then it recognizes the 11 properties of the clothing according to the tactile data. Those properties include the physical properties, like thickness, fuzziness, softness and durability, and semantic properties, like wearing season and preferred washing methods. We collect a dataset of 153 varied pieces of clothes, and conduct 6616 robot exploring iterations on them. To extract the useful information from the high-dimensional sensory output, we applied Convolutional Neural Networks (CNN) on the tactile data for recognizing the clothing properties, and on the Kinect depth images for selecting exploration locations. Experiments show that using the trained neural networks, the robot can autonomously explore the unknown clothes and learn their properties. This work proposes a new framework for active tactile perception system with vision-touch system, and has potential to enable robots to help humans with varied clothing related housework.
Wenzhen Yuan 0001, Yuchen Mo, Shaoxiong Wang, Edward H. Adelson
ICRA4
2018 GelSlim: A High-Resolution, Compact, Robust, and Calibrated Tactile-sensing Finger
abstract
This work describes the development of a high-resolution tactile-sensing finger for robot grasping. This finger, inspired by previous GelSight sensing techniques (Johnson and Adelson 2009), features an integration that is slimmer, more robust, and with more homogeneous output than previous vision-based tactile sensors. To achieve a compact integration, we redesign the optical path from illumination source to camera by combining light guides and an arrangement of mirror reflections. We parameterize the optical path with geometric design variables and describe the tradeoffs between the finger thickness, camera depth of field, and size of the tactile sensing area. The sensor sustains the wear from continuous use - and abuse - in grasping tasks by combining tougher materials for the compliant gel, a textured fabric skin, a structurally rigid body, and a calibration process that maintains homogeneous illumination and contrast of the tactile images during use. Finally, we evaluate the sensor's durability along four metrics that track the signal quality during more than 3000 grasping experiments.
Elliott Donlon, Siyuan Dong, Melody Liu, Edward H. Adelson, Alberto Rodriguez 0003
IROS5
2018 3D Shape Perception from Monocular Vision, Touch, and Shape Priors
abstract
Perceiving accurate 3D object shape is important for robots to interact with the physical world. Current research along this direction has been primarily relying on visual observations. Vision, however useful, has inherent limitations due to occlusions and the 2D-3D ambiguities, especially for perception with a monocular camera. In contrast, touch gets precise local shape information, though its efficiency for reconstructing the entire shape could be low. In this paper, we propose a novel paradigm that efficiently perceives accurate 3D object shape by incorporating visual and tactile observations, as well as prior knowledge of common object shapes learned from large-scale shape repositories. We use vision first, applying neural networks with learned shape priors to predict an object's 3D shape from a single-view color image. We then use tactile sensing to refine the shape; the robot actively touches the object regions where the visual prediction has high uncertainty. Our method efficiently builds the 3D shape of common objects from a color image and a small number of tactile explorations (around 10). Our setup is easy to apply and has potentials to help robots better perform grasping or manipulation tasks on real-world objects.
Shaoxiong Wang, Jiajun Wu 0001, Xingyuan Sun, Wenzhen Yuan 0001, William T. Freeman, Josh Tenenbaum, Edward H. Adelson
IROS7
2017 Connecting Look and Feel: Associating the Visual and Tactile Properties of Physical Materials
abstract
For machines to interact with the physical world, they must understand the physical properties of objects and materials they encounter. We use fabrics as an example of a deformable material with a rich set of mechanical properties. A thin flexible fabric, when draped, tends to look different from a heavy stiff fabric. It also feels different when touched. Using a collection of 118 fabric samples, we captured color and depth images of draped fabrics along with tactile data from a high-resolution touch sensor. We then sought to associate the information from vision and touch by jointly training CNNs across the three modalities. Through the CNN, each input, regardless of the modality, generates an embedding vector that records the fabrics physical property. By comparing the embedding vectors, our system is able to look at a fabric image and predict how it will feel, and vice versa. We also show that a system jointly trained on vision and touch data can outperform a similar system trained only on visual data when tested purely with visual inputs.
Wenzhen Yuan 0001, Shaoxiong Wang, Siyuan Dong, Edward H. Adelson
CVPR4
2017 Tracking objects with point clouds from vision and touch
abstract
We present an object-tracking framework that fuses point cloud information from an RGB-D camera with tactile information from a GelSight contact sensor. GelSight can be treated as a source of dense local geometric information, which we incorporate directly into a conventional point-cloud-based articulated object tracker based on signed-distance functions. Our implementation runs at 12 Hz using an online depth reconstruction algorithm for GelSight and a modified second-order update for the tracking algorithm. We present data from hardware experiments demonstrating that the addition of contact-based geometric information significantly improves the pose accuracy during contact, and provides robustness to occlusions of small objects by the robot's end effector.
Gregory Izatt, Geronimo Mirano, Edward H. Adelson, Russ Tedrake
ICRA3
2017 Shape-independent hardness estimation using deep learning and a GelSight tactile sensor
abstract
Hardness is among the most important attributes of an object that humans learn about through touch. However, approaches for robots to estimate hardness are limited, due to the lack of information provided by current tactile sensors. In this work, we address these limitations by introducing a novel method for hardness estimation, based on the GelSight tactile sensor, and the method does not require accurate control of contact conditions or the shape of objects. A GelSight has a soft contact interface, and provides high resolution tactile images of contact geometry, as well as contact force and slip conditions. In this paper, we try to use the sensor to measure hardness of objects with multiple shapes, under a loosely controlled contact condition. The contact is made manually or by a robot hand, while the force and trajectory are unknown and uneven. We analyze the data using a deep constitutional (and recurrent) neural network. Experiments show that the neural net model can estimate the hardness of objects with different shapes and hardness ranging from 8 to 87 in Shore 00 scale.
Wenzhen Yuan 0001, Chenzhuo Zhu, Andrew Owens, Mandayam A. Srinivasan, Edward H. Adelson
ICRA5
2017 Improved GelSight tactile sensor for measuring geometry and slip
abstract
A GelSight sensor uses an elastomeric slab covered with a reflective membrane to measure tactile signals. It measures the 3D geometry and contact force information with high spacial resolution, and successfully helped many challenging robot tasks. A previous sensor [1], based on a semi-specular membrane, produces high resolution but with limited geometry accuracy. In this paper, we describe a new design of GelSight for robot gripper, using a Lambertian membrane and new illumination system, which gives greatly improved geometric accuracy while retaining the compact size. We demonstrate its use in measuring surface normals and reconstructing height maps using photometric stereo. We also use it for the task of slip detection, using a combination of information about relative motions on the membrane surface and the shear distortions. Using a robotic arm and a set of 37 everyday objects with varied properties, we find that the sensor can detect translational and rotational slip in general cases, and can be used to improve the stability of the grasp.
Siyuan Dong, Wenzhen Yuan 0001, Edward H. Adelson
IROS3
2016 Visually Indicated Sounds
abstract
Objects make distinctive sounds when they are hit or scratched. These sounds reveal aspects of an object's material properties, as well as the actions that produced them. In this paper, we propose the task of predicting what sound an object makes when struck as a way of studying physical interactions within a visual scene. We present an algorithm that synthesizes sound from silent videos of people hitting and scratching objects with a drumstick. This algorithm uses a recurrent neural network to predict sound features from videos and then produces a waveform from these features with an example-based synthesis procedure. We show that the sounds predicted by our model are realistic enough to fool participants in a "real or fake" psychophysical experiment, and that they convey significant information about material properties and physical interactions.
Andrew Owens, Phillip Isola, Josh H. McDermott, Antonio Torralba 0001, Edward H. Adelson, William T. Freeman
CVPR5
2016 Estimating object hardness with a GelSight touch sensor
abstract
Hardness sensing is a valuable capability for a robot touch sensor. We describe a novel method of hardness sensing that does not require accurate control of contact conditions. A GelSight sensor is a tactile sensor that provides high resolution tactile images, which enables a robot to infer object properties such as geometry and fine texture, as well as contact force and slip conditions. The sensor is pressed on silicone samples by a human or a robot and we measure the sample hardness only with data from the sensor, without a separate force sensor and without precise knowledge of the contact trajectory. We describe the features that show object hardness. For hemispherical objects, we develop a model to measure the sample hardness, and the estimation error is about 4% in the range of 8 Shore 00 to 45 Shore A. With this technology, a robot is able to more easily infer the hardness of the touched objects, thereby improving its object recognition as well as manipulation strategy.
Wenzhen Yuan 0001, Mandayam A. Srinivasan, Edward H. Adelson
IROS3
2015 On the appearance of translucent edges
abstract
Edges in images of translucent objects are very different from edges in images of opaque objects. The physical causes for these differences are hard to characterize analytically and are not well understood. This paper considers one class of translucency edges-those caused by a discontinuity in surface orientation-and describes the physical causes of their appearance. We simulate thousands of translucency edge profiles using many different scattering material parameters, and we explain the resulting variety of edge patterns by qualitatively analyzing light transport. We also discuss the existence of shape and material metamers, or combinations of distinct shape or material parameters that generate the same edge profile. This knowledge is relevant to visual inference tasks that involve translucent objects, such as shape or material estimation.
Ioannis Gkioulekas, Bruce Walter, Edward H. Adelson, Kavita Bala, Todd E. Zickler
CVPR3
2015 Discovering states and transformations in image collections
abstract
Objects in visual scenes come in a rich variety of transformed states. A few classes of transformation have been heavily studied in computer vision: mostly simple, parametric changes in color and geometry. However, transformations in the physical world occur in many more flavors, and they come with semantic meaning: e.g., bending, folding, aging, etc. The transformations an object can undergo tell us about its physical and functional properties. In this paper, we introduce a dataset of objects, scenes, and materials, each of which is found in a variety of transformed states. Given a novel collection of images, we show how to explain the collection in terms of the states and transformations it depicts. Our system works by generalizing across object classes: states and transformations learned on one set of objects are used to interpret the image collection for an entirely new object class.
Phillip Isola, Joseph J. Lim, Edward H. Adelson
CVPR3
2015 Talk abstract: Computational lighting design and band-sifting operators
abstract
In this talk, I present two projects that are inspired by how photographers work. These projects are in collaboration with Ivo Boyadzhiev and Kavita Bala at Cornell University and Ted Adelson at MIT.
Sylvain Paris, Ivaylo Boyadzhiev, Kavita Bala, Edward H. Adelson
ICIP4
2015 Measurement of shear and slip with a GelSight tactile sensor
abstract
Artificial tactile sensing is still underdeveloped, especially in sensing shear and slip on a contact surface. For a robot hand to manually explore the environment or perform a manipulation task such as grasping, sensing of shear forces and detecting incipient slip is important. In this paper, we introduce a method of sensing the normal, shear and torsional load on the contact surface with a GelSight tactile sensor [1]. In addition, we demonstrate the detection of incipient slip. The method consists of inferring the state of the contact interface based on analysis of the sequence of images of GelSights elastomer medium, whose deformation under the external load indicates the conditions of contact. Results with a robot gripper like experimental setup show that the method is effective in detecting interactions with an object during stable grasp as well as at incipient slip. The method is also applicable to other optical based tactile sensors.
Wenzhen Yuan 0001, Rui Li 0017, Mandayam A. Srinivasan, Edward H. Adelson
ICRA4
2015 Band-Sifting Decomposition for Image-Based Material Editing
abstract
Photographers often “prep” their subjects to achieve various effects; for example, toning down overly shiny skin, covering blotches, etc. Making such adjustments digitally after a shoot is possible, but difficult without good tools and good skills. Making such adjustments to video footage is harder still. We describe and study a set of 2D image operations, based on multiscale image analysis, that are easy and straightforward and that can consistently modify perceived material properties. These operators first build a subband decomposition of the image and then selectively modify the coefficients within the subbands. We call this selection process band sifting . We show that different siftings of the coefficients can be used to modify the appearance of properties such as gloss, smoothness, pigmentation, or weathering. The band-sifting operators have particularly striking effects when applied to faces; they can provide “knobs” to make a face look wetter or drier, younger or older, and with heavy or light variation in pigmentation. Through user studies, we identify a set of operators that yield consistent subjective effects for a variety of materials and scenes. We demonstrate that these operators are also useful for processing video sequences.
Ivaylo Boyadzhiev, Kavita Bala, Sylvain Paris, Edward H. Adelson
ACM Trans. Graph.4
2014 Crisp Boundary Detection Using Pointwise Mutual Information
Phillip Isola, Daniel Zoran, Dilip Krishnan, Edward H. Adelson
ECCV (3)4
2014 Localization and manipulation of small parts using GelSight tactile sensing
abstract
Robust manipulation and insertion of small parts can be challenging because of the small tolerances typically involved. The key to robust control of these kinds of manipulation interactions is accurate tracking and control of the parts involved. Typically, this is accomplished using visual servoing or force-based control. However, these approaches have drawbacks. Instead, we propose a new approach that uses tactile sensing to accurately localize the pose of a part grasped in the robot hand. Using a feature-based matching technique in conjunction with a newly developed tactile sensing technology known as GelSight that has much higher resolution than competing methods, we synthesize high-resolution height maps of object surfaces. As a result of these high-resolution tactile maps, we are able to localize small parts held in a robot hand very accurately. We quantify localization accuracy in benchtop experiments and experimentally demonstrate the practicality of the approach in the context of a small parts insertion problem.
Rui Li 0017, Robert Platt 0001, Wenzhen Yuan 0001, Andreas ten Pas, Nathan Roscup, Mandayam A. Srinivasan, Edward H. Adelson
IROS7
2013 Sensing and Recognizing Surface Textures Using a GelSight Sensor
abstract
Sensing surface textures by touch is a valuable capability for robots. Until recently it was difficult to build a compliant sensor with high sensitivity and high resolution. The GelSight sensor is compliant and offers sensitivity and resolution exceeding that of the human fingertips. This opens the possibility of measuring and recognizing highly detailed surface textures. The GelSight sensor, when pressed against a surface, delivers a height map. This can be treated as an image, and processed using the tools of visual texture analysis. We have devised a simple yet effective texture recognition system based on local binary patterns, and enhanced it by the use of a multi-scale pyramid and a Hellinger distance metric. We built a database with 40 classes of tactile textures using materials such as fabric, wood, and sandpaper. Our system can correctly categorize materials from this database with high accuracy. This suggests that the GelSight sensor can be useful for material recognition by robots.
Rui Li 0017, Edward H. Adelson
CVPR2
2013 Lump detection with a gelsight sensor
abstract
A GelSight sensor is a tactile sensing device comprising a clear elastomeric pad covered with a reflective membrane, coupled with optics to measure the membrane's deformations. When the pad is pressed against an object's surface, the membrane changes shape in accord with mechanical and geometrical properties of the object. Since soft tissue is more compliant than hard tissue, one can detect an embedded lump by pressing the GelSight pad against the tissue surface and observing the hump that forms over the lump. We tested this system's sensitivity by constructing phantoms of soft rubber with hard embedded lumps. The system is quite sensitive; for example it could detect a 2mm lump at a depth of 5mm. The sensor was more sensitive than previous tactile lump detectors. It was also better than human observers using their fingertips. Such a capability could help in tumor screening, and could augment the sensory information available in telemedicine or minimally invasive surgery.
Xiaodan Jia, Rui Li 0017, Mandayam A. Srinivasan, Edward H. Adelson
World Haptics4
2013 Recognizing Materials Using Perceptually Inspired Features
Lavanya Sharan, Ce Liu 0001, Ruth Rosenholtz, Edward H. Adelson
Int. J. Comput. Vis.4
2013 Understanding the role of phase function in translucent appearance
abstract
Multiple scattering contributes critically to the characteristic translucent appearance of food, liquids, skin, and crystals; but little is known about how it is perceived by human observers. This article explores the perception of translucency by studying the image effects of variations in one factor of multiple scattering: the phase function. We consider an expanded space of phase functions created by linear combinations of Henyey-Greenstein and von Mises-Fisher lobes, and we study this physical parameter space using computational data analysis and psychophysics. Our study identifies a two-dimensional embedding of the physical scattering parameters in a perceptually meaningful appearance space. Through our analysis of this space, we find uniform parameterizations of its two axes by analytical expressions of moments of the phase function, and provide an intuitive characterization of the visual effects that can be achieved at different parts of it. We show that our expansion of the space of phase functions enlarges the range of achievable translucent appearance compared to traditional single-parameter phase function models. Our findings highlight the important role phase function can have in controlling translucent appearance, and provide tools for manipulating its effect in material design applications.
Ioannis Gkioulekas, Bei Xiao, Edward H. Adelson, Todd E. Zickler, Kavita Bala
ACM Trans. Graph.4
2012 Playing with Puffball: simple scale-invariant inflation for use in vision and graphics
abstract
We describe how inflation, the act of mapping a 2D silhouette to a 3D region, can be applied in two disparate problems to offer insight and improvement: silhouette part segmentation and image-based material transfer. To demonstrate this, we introduce Puffball, a novel inflation technique, which achieves similar results to existing inflation approaches -- including smoothness, robustness, and scale and shift-invariance -- through an exceedingly simple and accessible formulation. The part segmentation algorithm avoids many of the pitfalls of previous approaches by finding part boundaries on a canonical 3-D shape rather than in the contour of the 2-D shape; the algorithm gives reliable and intuitive boundaries, even in cases where traditional approaches based on the 2D Minima Rule are misled. To demonstrate its effectiveness, we present data in which subjects prefer Puffball's segmentations to more traditional Minima Rule-based segmentations across several categories of silhouettes. The texture transfer algorithm utilizes Puffball's estimated shape information to produce visually pleasing and realistically synthesized surface textures with no explicit knowledge of either underlying shape.
Nathaniel R. Twarog, Marshall F. Tappen, Edward H. Adelson
SAP3
2012 Shapecollage: Occlusion-Aware, Example-Based Shape Interpretation
Forrester Cole, Phillip Isola, William T. Freeman, Frédo Durand, Edward H. Adelson
ECCV (3)5
2011 Shape estimation in natural illumination
abstract
The traditional shape-from-shading problem, with a single light source and Lambertian reflectance, is challenging since the constraints implied by the illumination are not sufficient to specify local orientation. Photometric stereo algorithms, a variant of shape-from-shading, simplify the problem by controlling the illumination to obtain additional constraints. In this paper, we demonstrate that many natural lighting environments already have sufficient variability to constrain local shape. We describe a novel optimization scheme that exploits this variability to estimate surface normals from a single image of a diffuse object in natural illumination. We demonstrate the effectiveness of our method on both simulated and real images.
Micah K. Johnson, Edward H. Adelson
CVPR2
2011 deForm: an interactive malleable surface for capturing 2.5D arbitrary objects, tools and touch
abstract
We introduce a novel input device, deForm, that supports 2.5D touch gestures, tangible tools, and arbitrary objects through real-time structured light scanning of a malleable surface of interaction. DeForm captures high-resolution surface deformations and 2D grey-scale textures of a gel surface through a three-phase structured light 3D scanner. This technique can be combined with IR projection to allow for invisible capture, providing the opportunity for co-located visual feedback on the deformable surface. We describe methods for tracking fingers, whole hand gestures, and arbitrary tangible tools. We outline a method for physically encoding fiducial marker information in the height map of tangible tools. In addition, we describe a novel method for distinguishing between human touch and tangible tools, through capacitive sensing on top of the input surface. Finally we motivate our device through a number of sample applications.
Sean Follmer, Micah K. Johnson, Edward H. Adelson, Hiroshi Ishii 0001
UIST3
2011 Microgeometry capture using an elastomeric sensor
abstract
We describe a system for capturing microscopic surface geometry. The system extends the retrographic sensor [Johnson and Adelson 2009] to the microscopic domain, demonstrating spatial resolution as small as 2 microns. In contrast to existing microgeometry capture techniques, the system is not affected by the optical characteristics of the surface being measured---it captures the same geometry whether the object is matte, glossy, or transparent. In addition, the hardware design allows for a variety of form factors, including a hand-held device that can be used to capture high-resolution surface geometry in the field. We achieve these results with a combination of improved sensor materials, illumination design, and reconstruction algorithm, as compared to the original sensor of Johnson and Adelson [2009].
Micah K. Johnson, Forrester Cole, Alvin Raj, Edward H. Adelson
ACM Trans. Graph.4
2010 Exploring features in a Bayesian framework for material recognition
abstract
We are interested in identifying the material category, e.g. glass, metal, fabric, plastic or wood, from a single image of a surface. Unlike other visual recognition tasks in computer vision, it is difficult to find good, reliable features that can tell material categories apart. Our strategy is to use a rich set of low and mid-level features that capture various aspects of material appearance. We propose an augmented Latent Dirichlet Allocation (aLDA) model to combine these features under a Bayesian generative framework and learn an optimal combination of features. Experimental results show that our system performs material recognition reasonably well on a challenging material database, outperforming state-of-the-art material/texture recognition systems.
Ce Liu 0001, Lavanya Sharan, Edward H. Adelson, Ruth Rosenholtz
CVPR3
2010 Personal photo enhancement using example images
abstract
We describe a framework for improving the quality of personal photos by using a person's favorite photographs as examples. We observe that the majority of a person's photographs include the faces of a photographer's family and friends and often the errors in these photographs are the most disconcerting. We focus on correcting these types of images and use common faces across images to automatically perform both global and face-specific corrections. Our system achieves this by using face detection to align faces between “good” and “bad” photos such that properties of the good examples can be used to correct a bad photo. These “personal” photos provide strong guidance for a number of operations and, as a result, enable a number of high-quality image processing operations. We illustrate the power and generality of our approach by presenting a novel deblurring algorithm, and we show corrections that perform sharpening, superresolution, in-painting of over- and underexposured regions, and white-balancing.
Neel Joshi, Wojciech Matusik, Edward H. Adelson, David J. Kriegman
ACM Trans. Graph.3
2009 Retrographic sensing for the measurement of surface texture and shape
abstract
We describe a novel device that can be used as a 2.5D “scanner” for acquiring surface texture and shape. The device consists of a slab of clear elastomer covered with a reflective skin. When an object presses on the skin, the skin distorts to take on the shape of the object's surface. When viewed from behind (through the elastomer slab), the skin appears as a relief replica of the surface. A camera records an image of this relief, using illumination from red, green, and blue light sources at three different positions. A photometric stereo algorithm that is tailored to the device is then used to reconstruct the surface. There is no problem dealing with transparent or specular materials because the skin supplies its own BRDF. Complete information is recorded in a single frame; therefore we can record video of the changing deformation of the skin, and then generate an animation of the changing surface. Our sensor has no moving parts (other than the elastomer slab), uses inexpensive materials, and can be made into a portable device that can be used “in the field” to record surface shape and texture.
Micah K. Johnson, Edward H. Adelson
CVPR2
2009 Ground truth dataset and baseline evaluations for intrinsic image algorithms
abstract
The intrinsic image decomposition aims to retrieve “intrinsic” properties of an image, such as shading and reflectance. To make it possible to quantitatively compare different approaches to this problem in realistic settings, we present a ground-truth dataset of intrinsic image decompositions for a variety of real-world objects. For each object, we separate an image of it into three components: Lambertian shading, reflectance, and specularities. We use our dataset to quantitatively compare several existing algorithms; we hope that this dataset will serve as a means for evaluating future work on intrinsic images.
Roger B. Grosse, Micah K. Johnson, Edward H. Adelson, William T. Freeman
ICCV3
2008 Human-assisted motion annotation
abstract
Obtaining ground-truth motion for arbitrary, real-world video sequences is a challenging but important task for both algorithm evaluation and model design. Existing ground-truth databases are either synthetic, such as the Yosemite sequence, or limited to indoor, experimental setups, such as the database developed by Baker et al (2007). We propose a human-in-loop methodology to create a ground-truth motion database for the videos taken with ordinary cameras in both indoor and outdoor scenes, using the fact that human beings are experts at segmenting objects and inspecting the match between two frames. We designed an interactive computer vision system to allow a user to efficiently annotate motion. Our methodology is cross-validated by showing that human annotated motion is repeatable, consistent across annotators, and close to the ground truth obtained by Baker et al (2007). Using our system, we collected and annotated 10 indoor and outdoor real-world videos to form a ground-truth motion database. The source code, annotation tool and database is online for public evaluation and benchmarking.
Ce Liu 0001, William T. Freeman, Edward H. Adelson, Yair Weiss
CVPR3
2008 ScribbleBoost: Adding Classification to Edge-Aware Interpolation of Local Image and Video Adjustments
abstract
Abstract One of the most common tasks in image and video editing is the local adjustment of various properties (e.g., saturation or brightness) of regions within an image or video. Edge‐aware interpolation of user‐drawn scribbles offers a less effort‐intensive approach to this problem than traditional region selection and matting. However, the technique suffers a number of limitations, such as reduced performance in the presence of texture contrast, and the inability to handle fragmented appearances. We significantly improve the performance of edge‐aware interpolation for this problem by adding a boosting‐based classification step that learns to discriminate between the appearance of scribbled pixels. We show that this novel data term in combination with an existing edge‐aware optimization technique achieves substantially better results for the local image and video adjustment problem than edge‐aware interpolation techniques without classification, or related methods such as matting techniques or graph cut segmentation.
Yuanzhen Li, Edward H. Adelson, Aseem Agarwala
Comput. Graph. Forum2
2007 Learning Gaussian Conditional Random Fields for Low-Level Vision
abstract
Markov random field (MRF) models are a popular tool for vision and image processing. Gaussian MRF models are particularly convenient to work with because they can be implemented using matrix and linear algebra routines. However, recent research has focused on on discrete-valued and non-convex MRF models because Gaussian models tend to over-smooth images and blur edges. In this paper, we show how to train a Gaussian conditional random field (GCRF) model that overcomes this weakness and can outperform the non-convex field of experts model on the task of denoising images. A key advantage of the GCRF model is that the parameters of the model can be optimized efficiently on relatively large images. The competitive performance of the GCRF model and the ease of optimizing its parameters make the GCRF model an attractive option for vision and image processing applications.
Marshall F. Tappen, Ce Liu 0001, Edward H. Adelson, William T. Freeman
CVPR3
2007 Apparent ridges for line drawing
abstract
Three-dimensional shape can be drawn using a variety of feature lines, but none of the current definitions alone seem to capture all visually-relevant lines. We introduce a new definition of feature lines based on two perceptual observations. First, human perception is sensitive to the variation of shading, and since shape perception is little affected by lighting and reflectance modification, we should focus on normal variation. Second, view-dependent lines better convey smooth surfaces. From this we define view-dependent curvature as the variation of the surface normal with respect to a viewing screen plane, and apparent ridges as the loci of points that maximize a view-dependent curvature. We present a formal definition of apparent ridges and an algorithm to render line drawings of 3D meshes. We show that our apparent ridges encompass or enhance aspects of several other feature lines.
Tilke Judd, Frédo Durand, Edward H. Adelson
ACM Trans. Graph.3
2006 Estimating Intrinsic Component Images using Non-Linear Regression
abstract
Images can be represented as the composition of multiple intrinsic component images, such as shading, albedo, and noise images. In this paper, we present a method for estimating intrinsic component images from a single image, which we apply to the problems of estimating shading and albedo images and image denoising. Our method is based on learning estimators that predict filtered versions of the desired image. Unlike previous approaches, our method does not require unnatural discretizations of the problem. We also demonstrate how to learn a weighting function that properly weights the local estimates when constructing the estimated image. For shading estimation, we introduce a new training set of real-world images. The accuracy of our method is measured both qualitatively and quantitatively, showing better performance on the shading/albedo separation problem than previous approaches. The performance on denoising is competitive with the current state of the art.
Marshall F. Tappen, Edward H. Adelson, William T. Freeman
CVPR (2)2
2006 Analysis of Contour Motions
abstract
A reliable motion estimation algorithm must function under a wide range of con- ditions. One regime, which we consider here, is the case of moving objects with contours but no visible texture. Tracking distinctive features such as corners can disambiguate the motion of contours, but spurious features such as T-junctions can be badly misleading. It is difficult to determine the reliability of motion from local measurements, since a full rank covariance matrix can result from both real and spurious features. We propose a novel approach that avoids these points al- together, and derives global motion estimates by utilizing information from three levels of contour analysis: edgelets, boundary fragments and contours. Boundary fragment are chains of orientated edgelets, for which we derive motion estimates from local evidence. The uncertainties of the local estimates are disambiguated after the boundary fragments are properly grouped into contours. The grouping is done by constructing a graphical model and marginalizing it using importance sampling. We propose two equivalent representations in this graphical model, re- versible switch variables attached to the ends of fragments and fragment chains, to capture both local and global statistics of boundaries. Our system is success- fully applied to both synthetic and real video sequences containing high-contrast boundaries and textureless regions. The system produces good motion estimates along with properly grouped and completed contours.
Ce Liu 0001, William T. Freeman, Edward H. Adelson
NIPS3
2005 Recovering Intrinsic Images from a Single Image
abstract
Interpreting real-world images requires the ability distinguish the different characteristics of the scene that lead to its final appearance. Two of the most important of these characteristics are the shading and reflectance of each point in the scene. We present an algorithm that uses multiple cues to recover shading and reflectance intrinsic images from a single image. Using both color information and a classifier trained to recognize gray-scale patterns, given the lighting direction, each image derivative is classified as being caused by shading or a change in the surface's reflectance. The classifiers gather local evidence about the surface's form and color, which is then propagated using the generalized belief propagation algorithm. The propagation step disambiguates areas of the image where the correct classification is not clear from local evidence. We use real-world images to demonstrate results and show how each component of the system affects the results.
Marshall F. Tappen, William T. Freeman, Edward H. Adelson
IEEE Trans. Pattern Anal. Mach. Intell.3
2005 Compressing and companding high dynamic range images with subband architectures
abstract
High dynamic range (HDR) imaging is an area of increasing importance, but most display devices still have limited dynamic range (LDR). Various techniques have been proposed for compressing the dynamic range while retaining important visual information. Multi-scale image processing techniques, which are widely used for many image processing tasks, have a reputation of causing halo artifacts when used for range compression. However, we demonstrate that they can work when properly implemented. We use a symmetrical analysis-synthesis filter bank, and apply local gain control to the subbands. We also show that the technique can be adapted for the related problem of "companding", in which an HDR image is converted to an LDR image, and later expanded back to high dynamic range.
Yuanzhen Li, Lavanya Sharan, Edward H. Adelson
ACM Trans. Graph.3
2005 Motion magnification
abstract
We present motion magnification, a technique that acts like a microscope for visual motion. It can amplify subtle motions in a video sequence, allowing for visualization of deformations that would otherwise be invisible. To achieve motion magnification, we need to accurately measure visual motions, and group the pixels to be modified. After an initial image registration step, we measure motion by a robust analysis of feature point trajectories, and segment pixels based on similarity of position, color, and motion. A novel measure of motion similarity groups even very small motions according to correlation over time, which often relates to physical cause. An outlier mask marks observations not explained by our layered motion model, and those pixels are simply reproduced on the output from the original registered observations.The motion of any selected layer may be magnified by a user-specified amount; texture synthesis fills-in unseen "holes" revealed by the amplified motions. The resulting motion-magnified images can reveal or emphasize small motions in the original sequence, as we demonstrate with deformations in load-bearing structures, subtle motions or balancing corrections of people, and "rigid" structures bending under hand pressure.
Ce Liu 0001, Antonio Torralba 0001, William T. Freeman, Frédo Durand, Edward H. Adelson
ACM Trans. Graph.5
2002 Recovering Intrinsic Images from a Single Image
abstract
We present an algorithm that uses multiple cues to recover shading and reflectance intrinsic images from a single image. Using both color in- formation and a classifier trained to recognize gray-scale patterns, each image derivative is classified as being caused by shading or a change in the surface’s reflectance. Generalized Belief Propagation is then used to propagate information from areas where the correct classification is clear to areas where it is ambiguous. We also show results on real images.
Marshall F. Tappen, William T. Freeman, Edward H. Adelson
NIPS3
2001 Statistics of Real-World Illumination
abstract
While computer vision systems often assume simple illumination models, real-world illumination is highly complex, consisting of reflected light from every direction as well as distributed and localized primary light sources. One can capture the illumination incident at a point in the real world from every direction photographically using a spherical illumination map. This paper illustrates, through analysis of photographically-acquired, high dynamic range illumination maps, that real-world illumination shares many of the statistical properties of natural images. In particular, the marginal and joint wavelet coefficient distributions, directional derivative distributions, and harmonic spectra of illumination maps resemble those documented in the natural image statistics literature. However, illumination maps differ from standard photographs in that illumination maps are statistically non-stationary and may contain localized light sources that dominate their power spectra. Our work provides a foundation for statistical models of real-world illumination that may facilitate robust estimation of shape, reflectance, and illumination from images.
Ron O. Dror, Thomas K. Leung, Edward H. Adelson, Alan S. Willsky
CVPR (2)3
1999 Separating Reflections and Lighting Using Independent Components Analysis
abstract
The image of an object can vary dramatically depending on lighting, specularities/reflections and shadows. It is often advantageous to separate these incidental variations from the intrinsic aspects of an image. This paper describes how the statistical tool of independent components analysis can be used to separate some of these incidental components. We describe the details of this method and show its efficacy with examples of separating reflections off glass, and separating the relative contributions of individual light sources.
Hany Farid, Edward H. Adelson
CVPR2
1996 A Unified Mixture Framework for Motion Segmentation: Incorporating Spatial Coherence and Estimating the Number of Models
abstract
Describing a video sequence in terms of a small number of coherently moving segments is useful for tasks ranging from video compression to event perception. A promising approach is to view the motion segmentation problem in a mixture estimation framework. However, existing formulations generally use only the motion, data and thus fail to make use of static cues when segmenting the sequence. Furthermore, the number of models is either specified in advance or estimated outside the mixture model framework. In this work we address both of these issues. We show how to add spatial constraints to the mixture formulations and present a variant of the EM algorithm that males use of both the form and the motion constraints. Moreover this algorithm estimates the number of segments given knowledge about the level of model failure expected in the sequence. The algorithm's performance is illustrated on synthetic and real image sequences.
Yair Weiss, Edward H. Adelson
CVPR2
1996 Noise removal via Bayesian wavelet coring
abstract
The classical solution to the noise removal problem is the Wiener filter, which utilizes the second-order statistics of the Fourier decomposition. Subband decompositions of natural images have significantly non-Gaussian higher-order point statistics; these statistics capture image properties that elude Fourier-based techniques. We develop a Bayesian estimator that is a natural extension of the Wiener solution, and that exploits these higher-order statistics. The resulting nonlinear estimator performs a "coring" operation. We provide a simple model for the subband statistics, and use it to develop a semi-blind noise removal algorithm based on a steerable wavelet pyramid.
Eero P. Simoncelli, Edward H. Adelson
ICIP (1)2
1994 Analyzing and recognizing walking figures in XYT
abstract
We describe a novel algorithm for gait analysis. A person walking frontoparallel to the image plane generates a characteristic "braided" pattern in a spatiotemporal (XYT) volume. Our algorithm detects this pattern, and fits it with a set of spatiotemporal snakes. The snakes can be used to find the bounding contours of the walker. The contours vary over time in a manner characteristic of each walker. Individual gaits can be recognized by applying standard pattern recognition techniques to the contour signals.>
Sourabh A. Niyogi, Edward H. Adelson
CVPR2
1994 MID-Level Vision: New Directions in Vision and Video
abstract
Human vision, machine vision, and image coding, all have to deal with the problem of finding representations that are useful and efficient. The best-known techniques are based on low-level processing, using signal processing concepts such as filters, transforms, and simple non-linearities. Low-level concepts are at the heart of standard vision systems for computing optic flow, texture, etc. Low-level image coding techniques include DCTs, pyramids, wavelets, etc. To advance to a new generation of image coding architectures one must work with new image representations that involve such concepts as surfaces, lighting, transparency, etc. These representations fall in the domain of "mid-level" vision. By representing images with these more sophisticated vocabularies one can increase the flexibility and efficiency of the vision and image coding systems. The authors decompose an image sequence into a set of overlapping layers, rather like the "cels" used by a traditional animator. These layers are ordered in depth, sliding over one another and being combined according to the rules of transparency and occlusion. For some test sequences the authors achieve data compression far better than is possible with standard techniques such as MPEG.>
Edward H. Adelson, John Y. A. Wang, Sourabh A. Niyogi
ICIP (2)1
1994 Representing moving images with layers
abstract
We describe a system for representing moving images with sets of overlapping layers. Each layer contains an intensity map that defines the additive values of each pixel, along with an alpha map that serves as a mask indicating the transparency. The layers are ordered in depth and they occlude each other in accord with the rules of compositing. Velocity maps define how the layers are to be warped over time. The layered representation is more flexible than standard image transforms and can capture many important properties of natural image sequences. We describe some methods for decomposing image sequences into layers using motion analysis, and we discuss how the representation may be used for image coding and other applications.
John Y. A. Wang, Edward H. Adelson
IEEE Trans. Image Process.2
1993 Layered representation for motion analysis
abstract
Standard approaches to motion analysis assume that the optic flow is smooth; such techniques have trouble dealing with occlusion boundaries. The image sequence can be decomposed into a set of overlapping layers, where each layer's motion is described by a smooth flow field. The discontinuities in the description are then attributed to object opacities rather than to the flow itself, mirroring the structure of the scene. A set of techniques is devised for segmenting images into coherently moving regions using affine motion analysis and clustering techniques. It is possible to decompose an image into a set of layers along with information about occlusion and depth ordering. The techniques are applied to a flower garden sequence. The scene can be analyzed into four layers, and, the entire 30-frame sequence can be represented with a single image of each layer, along with associated motion parameters.>
John Y. A. Wang, Edward H. Adelson
CVPR2
1993 Layered representation for image sequence coding
John Y. A. Wang, Edward H. Adelson
ICASSP (5)2
1993 Recovering reflectance and illumination in a world of painted polyhedra
abstract
To be immune to variations in illumination, a vision system needs to be able to decompose images into their illumination and surface reflectance components. Most computational studies thus far have been concerned with strategies for solving the problem in the restricted domain of 2-D Mondrians. This domain has the simplifying characteristic of permitting discontinuities only in the reflectance distribution while the illumination distribution is constrained to vary smoothly. Such approaches prove inadequate in a 3-D world of painted polyhedra which allows for the existence of discontinuities in both the reflectance and illumination distributions. The authors propose a two-stage computational strategy for interpreting images acquired in such a domain. The first stage attempts to use simple local gray-level junction analysis to classify the observed image edges into the illumination or reflectance categories. Subsequent processing verifies the global consistency of these local inferences while also reasoning about the 3-D structure of the object and the illumination source direction.>
Pawan Sinha, Edward H. Adelson
ICCV2
1992 Single Lens Stereo with a Plenoptic Camera
abstract
Ordinary cameras gather light across the area of their lens aperture, and the light striking a given subregion of the aperture is structured somewhat differently than the light striking an adjacent subregion. By analyzing this optical structure, one can infer the depths of the objects in the scene, i.e. one can achieve single lens stereo. The authors describe a camera for performing this analysis. It incorporates a single main lens along with a lenticular array placed at the sensor plane. The resulting plenoptic camera provides information about how the scene would look when viewed from a continuum of possible viewpoints bounded by the main lens aperture. Deriving depth information is simpler than in a binocular stereo system because the correspondence problem is minimized. The camera extracts information about both horizontal and vertical parallax, which improves the reliability of the depth estimates.>
Edward H. Adelson, John Y. A. Wang
IEEE Trans. Pattern Anal. Mach. Intell.1
1992 Shiftable multiscale transforms
abstract
One of the major drawbacks of orthogonal wavelet transforms is their lack of translation invariance: the content of wavelet subbands is unstable under translations of the input signal. Wavelet transforms are also unstable with respect to dilations of the input signal and, in two dimensions, rotations of the input signal. The authors formalize these problems by defining a type of translation invariance called shiftability. In the spatial domain, shiftability corresponds to a lack of aliasing; thus, the conditions under which the property holds are specified by the sampling theorem. Shiftability may also be applied in the context of other domains, particularly orientation and scale. Jointly shiftable transforms that are simultaneously shiftable in more than one domain are explored. Two examples of jointly shiftable transforms are designed and implemented: a 1-D transform that is jointly shiftable in position and scale, and a 2-D transform that is jointly shiftable in position and orientation. The usefulness of these image representations for scale-space analysis, stereo disparity measurement, and image enhancement is demonstrated.>
Eero P. Simoncelli, William T. Freeman, Edward H. Adelson, David J. Heeger
IEEE Trans. Inf. Theory3
1991 A stereoscopic camera employing a single main lens
abstract
A camera for extracting depth information from a scene is described. It incorporates a single main lens along with a lenticular array placed at the sensor plane. The resulting plenoptic camera provides information about how the scene would look when viewed from a continuum of possible viewpoints bounded by the main lens aperture. Deriving depth information is simpler than in a binocular stereo system because the correspondence problem is minimized. The camera extracts information about both horizontal and vertical parallax, which improves the reliability of the depth estimates.>
Edward H. Adelson, John Y. A. Wang
CVPR1
1991 Probability distributions of optical flow
abstract
Gradient methods are widely used in the computation of optical flow. The authors discuss extensions of these methods which compute probability distributions of optical flow. The use of distributions allows representation of the uncertainties inherent in the optical flow computation, facilitating the combination with information from other sources. Distributed optical flow for a synthetic image sequence is computed, and it is demonstrated that the probabilistic model accounts for the errors in the flow estimates. The distributed optical flow for a real image sequence is computed.>
Eero P. Simoncelli, Edward H. Adelson, David J. Heeger
CVPR2
1991 Motion without movement
abstract
We describe a technique for displaying patterns that appear to move continuously without changing their positions. The method uses a quadrature pair of oriented filters to vary the local phase, giving the sensation of motion. We have used this technique in various computer graphic and scientific visualization applications.
William T. Freeman, Edward H. Adelson, David J. Heeger
SIGGRAPH2
1991 The Design and Use of Steerable Filters
abstract
The authors present an efficient architecture to synthesize filters of arbitrary orientations from linear combinations of basis filters, allowing one to adaptively steer a filter to any orientation, and to determine analytically the filter output as a function of orientation. Steerable filters may be designed in quadrature pairs to allow adaptive control over phase as well as orientation. The authors show how to design and steer the filters and present examples of their use in the analysis of orientation and phase, angularly adaptive filtering, edge detection, and shape from shading. One can also build a self-similar steerable pyramid representation. The same concepts can be generalized to the design of 3-D steerable filters.>
William T. Freeman, Edward H. Adelson
IEEE Trans. Pattern Anal. Mach. Intell.2
1990 Steerable filters for early vision, image analysis, and wavelet decomposition
abstract
An efficient architecture is presented to synthesize filters of arbitrary orientations from linear combinations of basis filters, allowing one to adaptively 'steer' a filter to any orientation, and to determine analytically the filter output as a function of orientation. The authors show how to design and steer filters, and present examples of their use in several tasks: the analysis of orientation and phase, angularly adaptive filtering, edge detection, and shape-from-shading. It is also possible to build a self-similar steerable pyramid representation which may be considered to be a steerable wavelet transform. The same concepts can be generalized to the design of 3-D steerable filters, which should be useful in the analysis of image sequences and volumetric data.>
William T. Freeman, Edward H. Adelson
ICCV2
1990 Non-separable extensions of quadrature mirror filters to multiple dimensions
abstract
Generalized non-separable extensions of quadrature mirror filter (QMF) banks to two and three dimensions, in which the orientation specificity of the high-pass filters is greatly improved, are described. In particular, extensions to two dimensions with hexagonal symmetry, and 3-D spatiotemporal extensions with rhombic-dodecahedral symmetry, are discussed. Although these filters are conceived and designed on nonstandard sampling lattices, they can be applied to rectangularly sampled images. As in one dimension, these transformations can be hierarchically cascaded to form a multiscale pyramid representation. A set of example filters is designed and applied to the problems of image compression, progressive transmission, orientation analysis, and motion analysis.
Eero P. Simoncelli, Edward H. Adelson
Proc. IEEE2
1983 The Laplacian Pyramid as a Compact Image Code
abstract
We describe a technique for image encoding in which local operators of many scales but identical shape serve as the basis functions. The representation differs from established techniques in that the code elements are localized in spatial frequency as well as in space. Pixel-to-pixel correlations are first removed by subtracting a lowpass filtered copy of the image from the image itself. The result is a net data compression since the difference, or error, image has low variance and entropy, and the low-pass filtered image may represented at reduced sample density. Further data compression is achieved by quantizing the difference image. These steps are then repeated to compress the low-pass image. Iteration of the process at appropriately expanded scales generates a pyramid data structure. The encoding process is equivalent to sampling the image with Laplacian operators of many scales. Thus, the code tends to enhance salient image features. A further advantage of the present code is that it is well suited for many image analysis tasks as well as for image compression. Fast algorithms are described for coding and decoding.
Peter J. Burt, Edward H. Adelson
IEEE Trans. Commun.2
1983 A Multiresolution Spline with Application to Image Mosaics
abstract
We define a multiresolution spline technique for combining two or more images into a larger image mosaic.In this procedure, the images to be splined are first decomposed into a set of band-pass filtered component images.Next, the component images in each spatial frequency band are assembled into a corresponding band-pass mosaic.In this step, component images are joined using a weighted average within a transition zone which is proportional in size to the wave lengths represented in the band.Finally, these band-pass mosaic images are summed to obtain the desired image mosaic.In this way, the spline is matched to the scale of features within the images themselves.When coarse features occur near borders, these are blended gradually over a relatively large distance without blurring or otherwise degrading finer image details in the neighborhood of the border.
Peter J. Burt, Edward H. Adelson
ACM Trans. Graph.2